AI & ML interests

None defined yet.

Recent Activity

Enxin  updated a dataset about 21 hours ago
GMLRVigil/BenchCheck-FramesPixels
Enxin  published a dataset about 22 hours ago
GMLRVigil/BenchCheck-FramesPixels
Enxin  updated a dataset about 23 hours ago
GMLRVigil/BenchCheck-Pool-Frames
View all activity

Enxin 
posted an update over 1 year ago
view post
Post
1616
🎉 Introducing Video-MMLU, a new benchmark for evaluating large multimodal models on classroom-style lectures in math, physics, and chemistry!

🧑‍🏫📚Video-MMLU requires strong reasoning capabilities and world knowledge compared to the previous benchmarks for video LMMs.

Each video comes with two tasks:
📝 Take Notes — detailed captioning of multi-discipline lectures
🧠 Do Quiz — open-ended QA to test reasoning over visuals & proofs

We evaluated 90+ models, including vision-blind baselines, open-source models and proprietary ones.
📉 We find that existing models generally perform poorly, with accuracy ranging from only 10% to 50%.
📉We also explore how the number of visual tokens and the base LLMs influence performance, offering insights into the interplay between multimodal perception and reasoning in lecture comprehension.

For more details, please check below:
📄 Paper: https://arxiv.org/abs/2504.14693
💻 Code: https://github.com/Espere-1119-Song/Video-MMLU
🧠 Data: Enxin/Video-MMLU
🌐 Website: https://enxinsong.com/Video-MMLU-web/