Women in AI Research (WiAIR)
In this episode of #WiAIRpodcast, we dive into a subtle but critical question: Does adding reasoning actually make LLMs safer and more reliable? Paper: https://arxiv.org/abs/2510.21049 [ https://arxiv.org/abs/2510.21049] Atoosa Chegini (University of Maryland, Apple) presents Reasoning's Razor (EACL 2026), where she and her collaborators examine how reasoning impacts high-stakes binary classification tasks, including safety filtering and hallucination detection. Their findings highlight an important nuance: * While reasoning can improve overall accuracy, it may degrade performance at low false positive rates -- exactly where real-world systems need to operate. This conversation covers: * Why accuracy is a misleading metric for safety-critical LLM applications * The importance of evaluating models at fixed false positive rates (FPR) * How two models with identical accuracy can behave completely differently in deployment * The impact of "think-on" (with reasoning) vs "think-off" (no reasoning) settings * Practical implications for RLHF, SFT, and post-training pipelines If you're working on: * LLM evaluation & reliability * AI safety or hallucination detection * Production deployment of language models — this discussion offers a perspective that is both technically grounded and immediately actionable. Atoosa: * https://www.linkedin.com/in/atoosa-chegini-6713741a3/ [https://www.linkedin.com/in/atoosa-chegini-6713741a3/] * https://scholar.google.com/citations?user=5nY9tagAAAAJ&hl=en&oi=ao [https://scholar.google.com/citations?user=5nY9tagAAAAJ&hl=en&oi=ao] 👍 Like & subscribe for more deep dives into cutting-edge AI research 🔔 New episodes from EACL 2026 coming soon
32 episodios
Comentarios
0Sé la primera persona en comentar
¡Regístrate ahora y únete a la comunidad de Women in AI Research (WiAIR)!