Skip to content

A 4.21 TFLOPS/W Memory-Efficient LLM Inference Accelerator with Bit-Layered Non-Uniform Quantization.

Byeongcheol Kim, Sangjin Kim, Sangwoo Ha, Soyeon Um, Kyomin Sohn, Hoi-Jun Yoo

VenueCISCAS
Year2025
ProceedingsISCAS

Browse the full ISCAS paper archive.