X-AuT: Progressive Audio-Encoder Compression for Speech LLMs with Cross-Scale Distillation Paper • 2609.11412 • Published 4 days ago • 41
X-AuT: Progressive Audio-Encoder Compression for Speech LLMs with Cross-Scale Distillation Paper • 2609.11412 • Published 4 days ago • 41
LD4WAM: Learning Latent Dynamics from Human Videos for World Action Models Paper • 2608.22403 • Published 22 days ago • 1
OpenWAM: An Open, Modular Exploration Towards Systematic World-Action Model Pretraining Paper • 2609.07398 • Published 7 days ago • 73
LD4WAM: Learning Latent Dynamics from Human Videos for World Action Models Paper • 2608.22403 • Published 22 days ago • 1
OpenWAM: An Open, Modular Exploration Towards Systematic World-Action Model Pretraining Paper • 2609.07398 • Published 7 days ago • 73
RynnBrain 1.1: Towards More Capable and Generalizable Embodied Foundation Model Paper • 2607.17977 • Published Jul 20 • 199
Benchmarking Visual State Tracking in Multimodal Video Understanding Paper • 2606.03920 • Published Jun 2 • 54
UniT: Toward a Unified Physical Language for Human-to-Humanoid Policy Learning and World Modeling Paper • 2604.19734 • Published Apr 21 • 34
AffordanceVLA: A Vision-Language-Action Model Empowering Action Generation through Affordance-Aware Understanding Paper • 2606.06155 • Published Jun 4 • 10
AffordanceVLA: A Vision-Language-Action Model Empowering Action Generation through Affordance-Aware Understanding Paper • 2606.06155 • Published Jun 4 • 10
LongAct: Harnessing Intrinsic Activation Patterns for Long-Context Reinforcement Learning Paper • 2604.14922 • Published Apr 16 • 7
SUGAR: A Scalable Human-Video-Driven Generalizable Humanoid Loco-Manipulation Learning Framework Paper • 2605.20373 • Published May 19
AffordanceVLA: A Vision-Language-Action Model Empowering Action Generation through Affordance-Aware Understanding Paper • 2606.06155 • Published Jun 4 • 10