A 4.21 TFLOPS/W Memory-Efficient LLM Inference Accelerator with Bit-Layered Non-Uniform Quantization.
Byeongcheol Kim, Sangjin Kim, Sangwoo Ha, Soyeon Um, Kyomin Sohn, Hoi-Jun Yoo
Browse the full ISCAS paper archive.
Byeongcheol Kim, Sangjin Kim, Sangwoo Ha, Soyeon Um, Kyomin Sohn, Hoi-Jun Yoo
Browse the full ISCAS paper archive.