I
InternVideo
OpenGVLab/InternVideo
[ECCV2024] Video Foundation Models & Data for Multimodal Understanding
★2.4kstars
Python
Apache-2.0
Updated: Today
📋 Project at a Glance
Tap to expand
What's this?Models project leveraging action-recognition and benchmark, built with Python, open-source
Who made it?Maintained by OpenGVLab team, 2.4K⭐ on GitHub, #596 out of 3201 in Models
Why does it exist?The OpenGVLab team recognized that existing action-recognition tools in Models were hard to use. InternVideo was designed to make benchmark more accessible.
What can it do?Key use cases: contrastive-learning, foundation-models, instruction-tuning
How to install with AI?Use an AI coding assistant (Claude Code, Cursor, Copilot) to automatically set up pip dependencies and virtual env. Follow the README — the AI handles the rest.
🔗 github.com/OpenGVLab/InternVideo
🔗 github.com/OpenGVLab/InternVideo
Topics
action-recognitionbenchmarkcontrastive-learningfoundation-modelsinstruction-tuningmasked-autoencodermultimodalopen-set-recognitionself-supervisedspatio-temporal-action-localizationtemporal-action-localizationvideo-clipvideo-datavideo-datasetvideo-question-answeringvideo-retrievalvideo-understandingvision-transformerzero-shot-classificationzero-shot-retrieval