Beyond Language Modeling: An Exploration of Multimodal Pretraining Paper ⢠2603.03276 ⢠Published Mar 3 ⢠109
Cosmos-Embed1 Collection Joint video-text embedding for physical AI ⢠5 items ⢠Updated Aug 11 ⢠9
DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression Paper ⢠2609.19969 ⢠Published 20 days ago ⢠224
On the Design of Qwen3.8-Next Architecture: Evaluation, Efficiency, and Training Stability Paper ⢠2608.30320 ⢠Published Aug 31 ⢠63
Qwen-Drive-1.0: An Initial Step towards a Vision-Language Foundation Model for Autonomous Driving Paper ⢠2609.00111 ⢠Published Aug 31 ⢠316
Cosmos-Reason2 Collection ā ļø This collection is archived. š https://huggingface.co/collections/nvidia/cosmos3 ⢠8 items ⢠Updated Aug 11 ⢠29
view article Article Welcome NVIDIA Cosmos 3: The First Open Omni-model for Physical AI Reasoning and Action nvidia ⢠Jun 1 ⢠91
NVIDIA OmniDreams Collection NVIDIA OmniDreams model checkpoints and sample datasets. ⢠3 items ⢠Updated Aug 11 ⢠10
view article Article Ulysses Sequence Parallelism: Training with Million-Token Contexts kashif, stas ⢠Mar 9 ⢠33
view article Article NVIDIA Cosmos-H-Dreams: Bringing Real-Time Generative Simulation to Surgical Robotics nvidia ⢠Jul 27 ⢠77
Knowledge Insulating Vision-Language-Action Models: Train Fast, Run Fast, Generalize Better Paper ⢠2505.23705 ⢠Published May 29, 2025 ⢠1
LingoQA: Video Question Answering for Autonomous Driving Paper ⢠2312.14115 ⢠Published Dec 21, 2023 ⢠4
Video Generation Models are General-Purpose Vision Learners Paper ⢠2607.09024 ⢠Published Jul 10 ⢠83