When AI reasoning goes wrong: Microsoft Research shows more tokens can mean more problems
Summary
A recent Microsoft Research study found that simply increasing computational resources during AI reasoning ("inference-time scaling") does not always yield better or more efficient results. The study, involving nine advanced foundation models like GPT-4o and Claude 3.7 Sonnet across tasks such as math, STEM reasoning, calendar planning, and navigation, used three inference-time scaling approaches: Chain-of-Thought (CoT), Parallel Scaling, and Sequential Scaling. The study discovered significant variations in performance benefits across models and tasks, high variability in token consumption, and no consistent correlation between longer reasoning chains and higher accuracy. Cost predictability is a challenge due to fluctuating token usage for repeated queries. The research suggests the potential of verification mechanisms to improve scaling performance consistently, which is crucial for enterprises to select cost-efficient models with low nondeterminism and develop robust verification methods for AI solutions.