> ## Documentation Index
> Fetch the complete documentation index at: https://docs.bespokelabs.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# vLLM

> Run open models with vLLM, either inside your Python process or as a server.

You can use vLLM with Curator in two ways.

* **Offline mode.** vLLM loads the model on your machine, inside your Python process.
* **Online mode.** vLLM runs as a separate server, and Curator sends requests to it.

This guide generates recipes with structured output in both modes.

## Prerequisites

* Python 3.10 or later.
* Curator, installed with `pip install bespokelabs-curator`.
* vLLM, installed with `pip install vllm`, and a GPU.

## Define the model and the generator

Both modes use this code.

```python theme={null}
from typing import List

from pydantic import BaseModel, Field

from bespokelabs import curator


class Recipe(BaseModel):
    title: str = Field(description="Title of the recipe")
    ingredients: List[str] = Field(description="List of ingredients needed")
    instructions: List[str] = Field(description="Step by step cooking instructions")
    prep_time: int = Field(description="Preparation time in minutes")
    cook_time: int = Field(description="Cooking time in minutes")
    servings: int = Field(description="Number of servings")


class RecipeGenerator(curator.LLM):
    response_format = Recipe

    def prompt(self, input: dict) -> str:
        return f"Generate a random {input['cuisine']} recipe. Be creative but keep it realistic."

    def parse(self, input: dict, response: Recipe) -> dict:
        return {
            "title": response.title,
            "ingredients": response.ingredients,
            "instructions": response.instructions,
            "prep_time": response.prep_time,
            "cook_time": response.cook_time,
            "servings": response.servings,
        }


cuisines = [{"cuisine": c} for c in ["Italian", "Chinese", "Mexican"]]
```

## Offline mode

Set `backend="vllm"` and pass the vLLM settings in `backend_params`.

```python theme={null}
generator = RecipeGenerator(
    model_name="Qwen/Qwen2.5-3B-Instruct",
    backend="vllm",
    backend_params={
        "tensor_parallel_size": 1,  # Set this to the number of GPUs.
        "gpu_memory_utilization": 0.7,
    },
)

recipes = generator(cuisines)
print(recipes.dataset.to_pandas())
```

Curator uses guided decoding in vLLM for structured output. If the model does not support it, Curator logs a warning.

### Offline settings

| Parameter | Default | Description |
| - | - | - |
| `tensor_parallel_size` | `1` | The number of GPUs for tensor parallelism. |
| `gpu_memory_utilization` | `0.95` | The share of GPU memory vLLM can use, from 0 to 1. |
| `max_model_length` | `4096` | The longest sequence the model can handle. |
| `max_tokens` | `4096` | The most tokens to generate. |
| `min_tokens` | `1` | The fewest tokens to generate. |
| `enforce_eager` | `False` | Whether to force eager execution. |
| `batch_size` | `256` | The number of prompts vLLM processes together. |
| `dtype` | `"auto"` | The data type of the model weights. |

## Online mode

<Steps>
  <Step title="Start the vLLM server">
    ```bash theme={null}
    vllm serve Qwen/Qwen2.5-3B-Instruct \
        --host localhost \
        --port 8787 \
        --api-key token-abc123
    ```
  </Step>

  <Step title="Connect Curator to the server">
    Use the LiteLLM backend with the `hosted_vllm/` prefix on the model name.

    ```python theme={null}
    import os

    os.environ["HOSTED_VLLM_API_KEY"] = "token-abc123"

    generator = RecipeGenerator(
        model_name="hosted_vllm/Qwen/Qwen2.5-3B-Instruct",
        backend="litellm",
        backend_params={
            "base_url": "http://localhost:8787/v1",
            "request_timeout": 30,
        },
    )

    recipes = generator(cuisines)
    print(recipes.dataset.to_pandas())
    ```
  </Step>
</Steps>

## Example output

Each row holds one recipe, like this one.

```json theme={null}
{
    "title": "Spicy Szechuan Noodles",
    "ingredients": [
        "400g wheat noodles",
        "2 tbsp Szechuan peppercorns",
        "3 cloves garlic, minced",
        "2 tbsp soy sauce"
    ],
    "instructions": [
        "Boil noodles according to package instructions",
        "Heat oil in a wok over medium-high heat",
        "Add peppercorns and garlic, stir-fry until fragrant",
        "Add noodles and soy sauce, toss to combine"
    ],
    "prep_time": 15,
    "cook_time": 20,
    "servings": 4
}
```
