Learning GenAI via SOTA Papers
Title: Cliff Tokens: Identifying Single-Token Failure Triggers in LLM Mathematical Reasoning Source: http://arxiv.org/abs/2606.25524v1 Summary: This paper identifies 'cliff tokens' as the exact single-token triggers that cause large language models to diverge into reasoning failures during multi-step mathematical tasks. By introducing a taxonomy of these failures and a targeted preference optimization method (Cliff-DPO), it establishes a foundational approach to diagnosing and improving LLM reasoning reliability.
323 episoder
Kommentarer
0Vær den første til at kommentere
Tilmeld dig nu og bliv en del af Learning GenAI via SOTA Papers-fællesskabet!