Base Models Can Reason By Taking a Cue From Training Data Paper • 2610.06851 • Published 7 days ago • 25
TRACE: Rollout-Guided Quantization-Aware Training for FP4 Reinforcement Learning of MoE Language Models Paper • 2610.07767 • Published 6 days ago • 92
Making LLMs Say What They Think: Measuring and Improving CoT-Interpretability Alignment Paper • 2609.38972 • Published 12 days ago • 28
MetaRubric: Learning to Reward for Rubric-Based Reinforcement Learning Paper • 2610.02824 • Published 10 days ago • 33
Architect-Ant: Editable Automatic Furnishing of Architectural Floor Plans Paper • 2606.10953 • Published 12 days ago • 54
Think Before You Score: Thinking Reward Model for Visual Generation Paper • 2609.37372 • Published 13 days ago • 104
VLA-Precision: Asymmetric Co-Bootstrapping for Efficient Real-World Online RL of Vision-Language-Action Models Paper • 2609.04355 • Published 24 days ago • 9
Morphometric Imitation: From Morphology and Contact Aware Hand Retargeting to Sim-to-Real Visuomotor Policy Paper • 2609.28660 • Published 19 days ago • 16
Towards Full Pipeline FP8 Reinforcement Learning for LLMs Paper • 2609.22870 • Published 23 days ago • 18
One to More, More to One: Category-Aware Iterative Expert Training for Software Engineering Agents Paper • 2609.23377 • Published 22 days ago • 50
BI-Agent and BI-Bench: Towards Automating End-to-End Business Intelligence Paper • 2609.20886 • Published 26 days ago • 30
From Pretraining to Proficiency: Real-World Subtask RL for Long-Horizon Manipulation with Minimal Human Intervention Paper • 2609.21788 • Published 24 days ago • 13
SteerDuplex: Steerable Duplex Speech Dialogue Models Paper • 2609.12623 • Published about 1 month ago • 6
CodeMidas: Scaling Agentic Coding RL Environments from Code Itself Paper • 2609.22068 • Published 24 days ago • 138
MintAct: A Unified Visual Agent for Digital Environments Paper • 2609.22083 • Published 24 days ago • 35
EvoOntology: A Self-Evolving Ontology Layer for Data Agents Paper • 2609.15779 • Published 28 days ago • 145
RetireOPD: Self-Retiring On-Policy Distillation for Agentic Reinforcement Learning Paper • 2609.20784 • Published 25 days ago • 57
Region-Level Policy Optimization for Fine-grained MLLM Perception Paper • 2609.19745 • Published 25 days ago • 41
Don't Mask the Environment: Observation Supervision Changes How Agents Explore Under RL Paper • 2609.20715 • Published 25 days ago • 44