Ray Serve LLM
Deploy LLMs with Ray Serve LLM.
An open-source Ray Serve framework for deploying large language models in production through an OpenAI-compatible API, with distributed multi-node serving, autoscaling, routing, parallelism, and prefill-decode disaggregation.

Recent stories
0 linked stories
No linked stories yet.