Progressive Mixed-Precision Decoding for Efficient LLM Inference.
Hao Mark Chen, Fuwen Tan, Alexandros Kouris, Royson Lee, Hongxiang Fan, Stylianos I. Venieris
Browse the full ICLR paper archive.
Hao Mark Chen, Fuwen Tan, Alexandros Kouris, Royson Lee, Hongxiang Fan, Stylianos I. Venieris
Browse the full ICLR paper archive.