Learning GenAI via SOTA Papers
Title: Cliff Tokens: Identifying Single-Token Failure Triggers in LLM Mathematical Reasoning Source: http://arxiv.org/abs/2606.25524v1 Summary: This paper identifies 'cliff tokens' as the exact single-token triggers that cause large language models to diverge into reasoning failures during multi-step mathematical tasks. By introducing a taxonomy of these failures and a targeted preference optimization method (Cliff-DPO), it establishes a foundational approach to diagnosing and improving LLM reasoning reliability.
323 Folgen
Kommentare
0Sei die erste Person, die kommentiert
Melde dich jetzt an und werde Teil der Learning GenAI via SOTA Papers-Community!