Amplified Does Not Mean Predictive: Reasoning Behaviors in Thinking Models
Abstract
Reasoning training amplifies deliberative behaviors like self-correction more than high-correctness behaviors such as confidence calibration, revealing a gap between amplified and correctness-linked reasoning patterns.
Which reasoning behaviors are associated with correct answers in reasoning models, and does reasoning-oriented training amplify those behaviors? This distinction is important because reasoning-oriented training can make traces look more deliberative without amplifying the behaviors most tied to model correctness. We quantify this mismatch with Behavioral Lift, a metric that measures how much correctness changes when a behavior is present versus absent in a model's reasoning trace. Across 15 models and 6 benchmarks spanning text-only and vision-language reasoning, we annotate 15,282 traces with a taxonomy whose core behaviors are defined for both LLM and VLM traces. We find evidence for an Amplification-Lift Gap, in which thinking models strongly amplify self-correction, hypothesis testing, and uncertainty acknowledgment, while the highest-lift behaviors are confidence calibration, knowledge alignment, and self-awareness. Confidence calibration is among the strongest positive signals of correctness in both modalities, yet is barely amplified; uncertainty acknowledgment is amplified by 3--7times, yet is weakly or negatively associated with correctness. We find that reasoning-oriented training does not preferentially amplify the highest-Lift behaviors, motivating process-level objectives that reward calibrated and grounded reasoning rather than surface form alone.
Community
Amplified Does Not Mean Predictive: Reasoning Behaviors in Thinking Models
Thinking models produce more deliberative traces, but this added deliberation is not
concentrated in the behaviors most associated with reasoning correctness. The largest lifts are
associated with confidence calibration, knowledge alignment, and self-awareness rather than
visible search or hesitation.
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- CALIBER: Calibrating Confidence Before and After Reasoning in Language Models (2026)
- How Much Does a Reasoning Summary Reveal? An Observability Ladder for Large Language Models (2026)
- Reported Confidence in LLMs Tracks Commitment More Than Correctness (2026)
- Beyond Correctness: Benchmarking and Aligning Response Behaviors in Hybrid-Thinking MLLMs (2026)
- REIN: Bridging the Gap between Reasoning and Reliability via Reflection and Abstention Alignment (2026)
- Rewarding Better Thinking for LLM Preference Alignment (2026)
- Reasoning Jury: Multi-Model Consensus for Evaluating Reasoning Traces (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2608.13760 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 1
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper