F
FlexLLMGen
FMInference/FlexLLMGen
Running large language models on a single GPU for throughput-oriented scenarios.
★9.4kstars
Python
Apache-2.0
Updated: Today
📋 Project at a Glance
Tap to expand
What's this?A open-source Data & Infrastructure project, built with Python, focusing on deep-learning and gpt-3
Who made it?Maintained by FMInference team, 9.4K⭐ on GitHub, #216 out of 3133 in Data & Infrastructure
Why does it exist?With the growing demand for deep-learning and gpt-3 in Data & Infrastructure, FMInference created FlexLLMGen as a streamlined solution.
What can it do?Key use cases: high-throughput, large-language-models, machine-learning
How to install with AI?Use an AI coding assistant (Claude Code, Cursor, Copilot) to automatically set up pip dependencies and virtual env. Follow the README — the AI handles the rest.
🔗 github.com/FMInference/FlexLLMGen
🔗 github.com/FMInference/FlexLLMGen
Topics
deep-learninggpt-3high-throughputlarge-language-modelsmachine-learningoffloadingopt