From Reasoning to Answer: Empirical, Attention-Based and Mechanistic Insights into Distilled DeepSeek R1 Models.
Jue Zhang, Qingwei Lin, Saravan Rajmohan, Dongmei Zhang
Browse the full EMNLP paper archive.
Jue Zhang, Qingwei Lin, Saravan Rajmohan, Dongmei Zhang
Browse the full EMNLP paper archive.