vllm-project
vllm
A high-throughput and memory-efficient inference and serving engine for LLMs
Stars
90,690
Forks
21,550
Language
Python
Licence
Apache-2.0
Created 9 Feb 2023Last push 1 Sept 2026
Stars and forks are a snapshot taken when this directory was last refreshed, not a live count — the figure here matches the one in our articles and videos. Last refreshed 1 Sept 2026.