F
flashinfer
flashinfer-ai/flashinfer
FlashInfer: Kernel Library for LLM Serving
★6.1kstars
Python
Apache-2.0
Updated: Today
📋 Project at a Glance
Tap to expand
What's this?A , built with Python open-source project in the Data & Infrastructure category, core strengths: attention/cuda
Who made it?Maintained by flashinfer-ai team, 6.1K⭐ on GitHub, #304 out of 3133 in Data & Infrastructure
Why does it exist?In the Data & Infrastructure space, attention workflows faced efficiency bottlenecks. flashinfer was built by flashinfer-ai to address these cuda challenges.
What can it do?Key use cases: distributed-inference, gpu, jit
How to install with AI?Use an AI coding assistant (Claude Code, Cursor, Copilot) to automatically set up pip dependencies and virtual env. Follow the README — the AI handles the rest.
🔗 github.com/flashinfer-ai/flashinfer | 官网 https://flashinfer.ai
🔗 github.com/flashinfer-ai/flashinfer | 官网 https://flashinfer.ai
Topics
attentioncudadistributed-inferencegpujitlarge-large-modelsllm-inferencemoenvidiapytorch