- Offline mode. vLLM loads the model on your machine, inside your Python process.
- Online mode. vLLM runs as a separate server, and Curator sends requests to it.
Prerequisites
- Python 3.10 or later.
- Curator, installed with
pip install bespokelabs-curator. - vLLM, installed with
pip install vllm, and a GPU.
Define the model and the generator
Both modes use this code.Offline mode
Setbackend="vllm" and pass the vLLM settings in backend_params.
Offline settings
Online mode
1
Start the vLLM server
2
Connect Curator to the server
Use the LiteLLM backend with the
hosted_vllm/ prefix on the model name.