vLLM
High-throughput LLM inference/serving engine; the current standard for self-hosting open-weight LLMs at scale.
https://vllm.aiUpdate history
No updates recorded for vLLM yet. Check back after its next release.
Get the badge
Show that vLLM is tracked on StackFollow in your project's README.
[](https://stackfollow.xyz/tools/vllm)