Papers
arxiv:2610.03252

COSMI: COmpositional Synthesis of Multi-object Interactions

Published on Oct 2
· Submitted by
Ilya Petrov
on Oct 6
Authors:

Abstract

Generative models of human-object interaction are bounded by the data that exists: everyday activities involve several objects, but most captured datasets record one at a time, as multi-object capture is combinatorially expensive. Our observation is that interactions are local, so single-object captures already contain the parts of multi-object activities. We compose them: contact-consistent clips of single interactions, mirrored to balance the hands, transfer between bodies, and a language model and geometric checks admit only the pairings that are plausible, semantically and physically. Therefore, the dataset grows combinatorially with the clips rather than recording time. The COSMI dataset holds 222k sequences and 275 hours with up to five objects, nearly thirty times the largest multi-object capture, and can be extended by adding datasets or even hand-object recordings. On this data we train the COSMI method, a text-to-interaction diffusion transformer that follows how the data is built: weight-shared object slots generate a variable number of objects, predicted relative to the body parts that move them. On a benchmark with an unseen object and unseen interaction combinations, models trained on the dataset generalize to the unseen combinations. COSMI outperforms baselines in text alignment and contact accuracy, where its margin is largest on the unseen object. Code, models, and the dataset pipeline will be released on the project page: https://ptrvilya.github.io/cosmi.

Community

Paper author Paper submitter

🧩 COSMI — multi-object human-object interaction in two parts: a dataset composed from single-object captures, and a text-to-interaction model trained on it.

📦 Data: interactions are local (a hand holds a cup, a chair supports the pelvis), so a limb and its object transfer between bodies. An LLM and geometric checks keep only plausible pairings, and mirroring balances left and right hands (49/51). The dataset has 222k sequences / 275 hours with up to 5 objects, nearly 30× the largest captured multi-object dataset.

🧠 Model: one 30M-parameter diffusion transformer generates interactions with 1 to 5 objects from text. Object slots share weights, and each object is predicted relative to the body part that moves it.

📊 Benchmark: holds out an unseen object and 5 unseen interaction combinations, such as walking with an apple. All models trained on the dataset generalize to the unseen combinations. COSMI leads in text alignment and contact accuracy, most clearly on the unseen object.

📈 The dataset grows by adding new datasets, even hand-object recordings (shown with HOT3D). Code, models and the dataset pipeline will be released.

🔗 Paper · Project page · GitHub · Hugging Face

This is an automated message from the Librarian Bot. I found the following papers similar to this paper.

The following papers were recommended by the Semantic Scholar API

Please give a thumbs up to this comment if you found it helpful!

If you want recommendations for any Paper on Hugging Face checkout this Space

You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend

Sign up or log in to comment

Models citing this paper 0

No model linking this paper

Cite arxiv.org/abs/2610.03252 in a model README.md to link it from this page.

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2610.03252 in a dataset README.md to link it from this page.

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2610.03252 in a Space README.md to link it from this page.

Collections including this paper 0

No Collection including this paper

Add this paper to a collection to link it from this page.