Beyond the Commitment Boundary: Probing Epiphenomenal Chain-of-Thought in Large Reasoning Models Paper • 2606.13603 • Published Jun 11
Distilling Formal Logic into Neural Spaces: A Kernel Alignment Approach for Signal Temporal Logic Paper • 2603.05198 • Published Mar 5
Bridging Logic and Learning: Decoding Temporal Logic Embeddings via Transformers Paper • 2507.07808 • Published Jul 10, 2025
Predicting Future Behaviors in Reasoning Models Enables Better Steering Paper • 2606.11172 • Published Jun 9 • 1
Interpreto: An Explainability Library for Transformers Paper • 2512.09730 • Published Dec 10, 2025 • 1
Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models Paper • 2607.08317 • Published 21 days ago • 35
Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models Paper • 2607.08317 • Published 21 days ago • 35
Contextual Earnings-22: A Speech Recognition Benchmark with Custom Vocabulary in the Wild Paper • 2604.07354 • Published Mar 28 • 1
JacNet: Learning Functions with Structured Jacobians Paper • 2408.13237 • Published Aug 23, 2024 • 1
Variance Reduction for Expectations with Diffusion Teachers Paper • 2605.21489 • Published May 21 • 1
NVIDIA OmniDreams: Real-Time Generative World Model for Closed-Loop Autonomous Vehicle Simulation Paper • 2606.03159 • Published Jun 2 • 23
Who Annotates in NLP? A Large-scale Assessment of Human Annotation Reporting between 2018 and 2025 Paper • 2606.02255 • Published Jun 1
Who Annotates in NLP? A Large-scale Assessment of Human Annotation Reporting between 2018 and 2025 Paper • 2606.02255 • Published Jun 1
WhisperKit: On-device Real-time ASR with Billion-Scale Transformers Paper • 2507.10860 • Published Jul 14, 2025 • 2
SDBench: A Comprehensive Benchmark Suite for Speaker Diarization Paper • 2507.16136 • Published Aug 6, 2025 • 1
Models That Know How Evaluations Are Designed Score Safer Paper • 2605.28591 • Published May 27 • 10
Models That Know How Evaluations Are Designed Score Safer Paper • 2605.28591 • Published May 27 • 10