X
xmodaler
YehLi/xmodaler
X-modaler is a versatile and high-performance codebase for cross-modal analytics(e.g., image captioning, video captioning, vision-language pre-training, visual question answering, visual commonsense reasoning, and cross-modal retrieval).
★972stars
Python
NOASSERTION
Updated: 3w ago
📋 Project at a Glance
Tap to expand
What's this?A , built with Python open-source project in the Models category, core strengths: cross-modal-retrieval/image-captioning
Who made it?Maintained by YehLi team, 972⭐ on GitHub, #1087 out of 3201 in Models
Why does it exist?In the Models space, cross-modal-retrieval workflows faced efficiency bottlenecks. xmodaler was built by YehLi to address these image-captioning challenges.
What can it do?Key use cases: pretraining, tden, video-captioning
How to install with AI?Use an AI coding assistant (Claude Code, Cursor, Copilot) to automatically set up pip dependencies and virtual env. Follow the README — the AI handles the rest.
🔗 github.com/YehLi/xmodaler
🔗 github.com/YehLi/xmodaler
Topics
cross-modal-retrievalimage-captioningpretrainingtdenvideo-captioningvision-and-languagevisual-question-answering