Mitigating Content Effects on Reasoning in Language Models Through Fine-Grained Activation Steering.
Marco Valentino, Geonhee Kim, Dhairya Dalal, Zhixue Zhao, Andr Freitas
Browse the full AAAI paper archive.
Marco Valentino, Geonhee Kim, Dhairya Dalal, Zhixue Zhao, Andr Freitas
Browse the full AAAI paper archive.