vllm-project

vllm

A high-throughput and memory-efficient inference and serving engine for LLMs

Stars
90,690
Forks
21,550
Language
Python
Licence
Apache-2.0
Created 9 Feb 2023Last push 1 Sept 2026
Share:

Stars and forks are a snapshot taken when this directory was last refreshed, not a live count — the figure here matches the one in our articles and videos. Last refreshed 1 Sept 2026.