OS Agents: A Survey on MLLM-based Agents for General Computing Devices Use Paper • 2508.04482 • Published Aug 6, 2025 • 10
Exploring the Design Space of Reward Backpropagation for Flow Matching Paper • 2606.11075 • Published Jun 9 • 12
TempAct: Advancing Temporal Plausibility in Autoregressive Video Generation via Planner-Executor RL Paper • 2606.28016 • Published Jun 26 • 1
MeanFlowNFT: Bringing Forward-Process RL to Average-Velocity Generators Paper • 2607.15273 • Published Jul 16 • 17
Stale but Stable: Staleness-Adaptive Trust Regions for Stabilizing Asynchronous Reinforcement Learning Paper • 2607.18722 • Published Jul 21 • 36
Reinforcing Few-step Generators via Reward-Tilted Distribution Matching Paper • 2605.26108 • Published May 25 • 7
Beyond Uniform Token-Level Trust Region in LLM Reinforcement Learning Paper • 2606.10968 • Published Jun 9 • 42
Flow-DPPO: Divergence Proximal Policy Optimization for Flow Matching Models Paper • 2606.11025 • Published Jun 9 • 41
RTDMD Collection Reinforcing Few-step Generators via Reward-Tilted Distribution Matching • 5 items • Updated Jun 2 • 3
DeepSearch: Overcome the Bottleneck of Reinforcement Learning with Verifiable Rewards via Monte Carlo Tree Search Paper • 2509.25454 • Published Sep 29, 2025 • 147
Diagnose, Localize, Align: A Full-Stack Framework for Reliable LLM Multi-Agent Systems under Instruction Conflicts Paper • 2509.23188 • Published Sep 27, 2025 • 3
Beyond Pass@1: Self-Play with Variational Problem Synthesis Sustains RLVR Paper • 2508.14029 • Published Aug 19, 2025 • 119