Making LLMs Say What They Think: Measuring and Improving CoT-Interpretability Alignment Paper • 2609.38972 • Published 11 days ago • 27
MetaRubric: Learning to Reward for Rubric-Based Reinforcement Learning Paper • 2610.02824 • Published 9 days ago • 33
James4Ever0/computer_agent_reinforcement_learning_trajectory_seagent_ai_assistant_tools_agent_mcp Updated Aug 10, 2025 • 153 • 6
Architect-Ant: Editable Automatic Furnishing of Architectural Floor Plans Paper • 2606.10953 • Published 11 days ago • 54
Think Before You Score: Thinking Reward Model for Visual Generation Paper • 2609.37372 • Published 12 days ago • 104
VLA-Precision: Asymmetric Co-Bootstrapping for Efficient Real-World Online RL of Vision-Language-Action Models Paper • 2609.04355 • Published 23 days ago • 9
Morphometric Imitation: From Morphology and Contact Aware Hand Retargeting to Sim-to-Real Visuomotor Policy Paper • 2609.28660 • Published 18 days ago • 16
Towards Full Pipeline FP8 Reinforcement Learning for LLMs Paper • 2609.22870 • Published 22 days ago • 18
One to More, More to One: Category-Aware Iterative Expert Training for Software Engineering Agents Paper • 2609.23377 • Published 21 days ago • 50