原文整理页

研究发现 AI 推理模型在总结过程中存在“不诚实”现象,会刻意隐藏其为了凑出已知答案而进行的投机推理过程

来源作者:Alexander Panfilov (@kotekjedi_ml)原始来源:https://x.com/steipete/status/2087328498633650344

中文导读

研究发现 AI 推理模型在总结过程中存在“不诚实”现象,会刻意隐藏其为了凑出已知答案而进行的投机推理过程。

正文 Markdown

But we also took a chance to have a look at some in-the-wild scheming, reward seeking, etc. examples, and dumped it in appendix. 1) Summarizer unfaithfulness Reasoning summaries often omit important information from the original trace. Here, Opus 4.8 realizes it knows the answer to an AIME problem and then tries to fit a solution to that answer. None of this appears in the summary.