L
LLaVA-Mini
ictnlp/LLaVA-Mini
LLaVA-Mini is a unified large multimodal model (LMM) that can support the understanding of images, high-resolution images, and videos in an efficient manner.
★575stars
Python
Apache-2.0
Updated: 1w ago
📋 Project at a Glance
Tap to expand
What's this?A efficient/gpt4o tool in the Models category, built with Python, open-source
Who made it?Maintained by ictnlp team, 575⭐ on GitHub, #1675 out of 3201 in Models
Why does it exist?The ictnlp team recognized that existing efficient tools in Models were hard to use. LLaVA-Mini was designed to make gpt4o more accessible.
What can it do?Key use cases: gpt4v, large-language-models, large-multimodal-models
How to install with AI?Use an AI coding assistant (Claude Code, Cursor, Copilot) to automatically set up pip dependencies and virtual env. Follow the README — the AI handles the rest.
🔗 github.com/ictnlp/LLaVA-Mini
🔗 github.com/ictnlp/LLaVA-Mini
Topics
efficientgpt4ogpt4vlarge-language-modelslarge-multimodal-modelsllamallavamultimodalmultimodal-large-language-modelsvideovisionvision-language-modelvisual-instruction-tuning