vLLM
A high-throughput and memory-efficient inference and serving engine for LLMs.
Open-source software for high-throughput, memory-efficient inference and serving of large language models.

Recent stories
0 linked stories
No linked stories yet.
A high-throughput and memory-efficient inference and serving engine for LLMs.
Open-source software for high-throughput, memory-efficient inference and serving of large language models.
