vllm-project/vllm
A high-throughput and memory-efficient inference and serving engine for LLMs
GitHub repository with 92,221 stars and 22,441 forks.
Language: Python
Topics: amd, blackwell, cuda, deepseek, deepseek-v3, gpt, gpt-oss, inference, kimi, llama