> ## Documentation Index
> Fetch the complete documentation index at: https://docs.bespokelabs.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# API reference

> Reference for curator.LLM, its backends and parameters, CuratorResponse, and environment variables.

## curator.LLM

`curator.LLM` is the main class for prompting LLMs in Curator. Calling an `LLM` object returns a [`CuratorResponse`](#curatorresponse).

```python theme={null}
class LLM:
    response_format: Type[BaseModel] | None = None
    return_completions_object: bool = False

    def __init__(
        self,
        model_name: str,
        response_format: Type[BaseModel] | None = None,
        batch: bool = False,
        backend: Optional[str] = None,
        generation_params: dict | None = None,
        backend_params: BackendParamsType | None = None,
        system_prompt: str | None = None,
    ): ...
```

### Constructor parameters

| Parameter | Type | Default | Description |
| - | - | - | - |
| `model_name` | `str` | Required | The name of the model. |
| `response_format` | `Type[BaseModel] \| None` | `None` | A Pydantic model for [structured output](/curator/structured-output). You can also set it as a class attribute. |
| `batch` | `bool` | `False` | Use the provider's [batch API](/curator/batch). |
| `backend` | `str \| None` | `None` | The backend to use. If `None`, Curator picks one from the model name. See [Backends](#backends). |
| `generation_params` | `dict \| None` | `None` | Parameters that Curator passes to the model API, e.g., `temperature` or `max_tokens`. |
| `backend_params` | `dict \| None` | `None` | Settings for how Curator sends requests. See [Backend parameters](#backend-parameters). |
| `system_prompt` | `str \| None` | `None` | A system message that Curator adds to every request. |

### Class attributes

| Attribute | Default | Description |
| - | - | - |
| `response_format` | `None` | A Pydantic model for structured output. |
| `return_completions_object` | `False` | If `True`, `parse` gets the full API response as a dictionary instead of only the text. |

### Call parameters

| Parameter | Default | Description |
| - | - | - |
| `dataset` | `None` | The input. It can be a string, a list of messages, a list of strings or dictionaries, a Hugging Face `Dataset`, or a `CuratorResponse`. With no input, Curator sends one request. |
| `working_dir` | `None` | The cache directory for this run. The default is `CURATOR_CACHE_DIR`, or `~/.cache/curator`. |
| `batch_cancel` | `False` | In batch mode, cancel the running batches of this run instead of running it. |

### prompt()

```python theme={null}
def prompt(self, input: dict | BaseModel) -> str | list | tuple: ...
```

Builds the prompt for one input row. It returns one of these.

* A string, which Curator sends as a single user message.
* A list of messages, each a dictionary with `role` and `content`.
* A tuple of text and a `curator.types.Image`, for [multimodal prompts](/curator/guides/multimodal).

```python theme={null}
def prompt(self, input: dict) -> str:
    return f"Generate a {input['type']} about {input['topic']}"
```

### parse()

```python theme={null}
def parse(self, input: dict | BaseModel, response: str | BaseModel | dict) -> dict | list: ...
```

Turns the LLM response into output rows. It gets the input row and the response, and returns one row as a dictionary or several rows as a list of dictionaries. The response is a string by default, an instance of `response_format` if you set one, or a dictionary if `return_completions_object` is `True`.

```python theme={null}
def parse(self, input: dict, response: str) -> dict:
    return {"topic": input["topic"], "generated_text": response}
```

## Backends

| `backend` | Online | Batch | Chosen automatically when | API key |
| - | - | - | - | - |
| `openai` | Yes | Yes | LiteLLM identifies the model as an OpenAI model. | `OPENAI_API_KEY` |
| `anthropic` | Yes | Yes | The model name contains `claude`. | `ANTHROPIC_API_KEY` |
| `litellm` | Yes | No | The model does not match any other rule. LiteLLM must know the model, so use its provider prefix, e.g., `gemini/`. | Depends on the provider. See [LiteLLM](/curator/guides/litellm). |
| `mistral` | No | Yes | Never. Set `backend="mistral"`. | `MISTRAL_API_KEY` |
| `gemini` | No | Yes | Never. | Google Cloud credentials. See [Batch inference](/curator/batch). |
| `azure` | No | Yes | Never. | `AZURE_OPENAI_API_KEY` |
| `inference.net` | Yes | Yes | Never. | `INFERENCE_API_KEY` |
| `vllm` | Offline | No | Never. | None. See [vLLM](/curator/guides/vllm). |

To use any other API that is compatible with the OpenAI API, set `backend="openai"` and pass `base_url` and `api_key` in `backend_params`. The [Hugging Face Inference Providers](/curator/guides/hf-inference-providers) guide shows an example.

## Backend parameters

`backend_params` is a dictionary. The keys you can use depend on the mode.

### Common parameters

These work with every backend.

| Parameter | Type | Default | Description |
| - | - | - | - |
| `max_retries` | `int` | `10` | The number of times Curator retries a failed request. |
| `require_all_responses` | `bool` | `True` | If `True`, Curator raises an error when any request still fails after all retries. Set it to `False` to keep the successful responses. |
| `base_url` | `str` | `None` | The base URL of the API. |
| `api_key` | `str` | `None` | The API key. By default, Curator reads it from the provider's environment variable. |
| `request_timeout` | `int` | `600` | The timeout for one request, in seconds. |
| `in_mtok_cost` | `int` | `None` | The price per million input tokens, for models whose price Curator does not know. |
| `out_mtok_cost` | `int` | `None` | The price per million output tokens, for models whose price Curator does not know. |
| `invalid_finish_reasons` | `list` | `["content_filter", "length"]` | Finish reasons that Curator treats as failures. |

```python theme={null}
backend_params = {
    "max_retries": 3,
    "require_all_responses": True,
    "base_url": "https://custom-endpoint.com/v1",
    "request_timeout": 300,
}
```

### Online parameters

These work in online mode, which is the default.

| Parameter | Type | Default | Description |
| - | - | - | - |
| `max_requests_per_minute` | `int` | From provider headers, or 200 | The most requests in one minute. |
| `max_tokens_per_minute` | `int` | From provider headers, or 100,000 | The most input and output tokens in one minute. |
| `max_input_tokens_per_minute` | `int` | `None` | The most input tokens in one minute, for providers that limit input and output separately, e.g., Anthropic. |
| `max_output_tokens_per_minute` | `int` | `None` | The most output tokens in one minute, for the same providers. |
| `max_concurrent_requests` | `int` | `None` | The most requests waiting for a response at the same time. |
| `seconds_to_pause_on_rate_limit` | `int` | `10` | Seconds to wait after a rate limit error. |

```python theme={null}
backend_params = {
    "max_requests_per_minute": 2_000,
    "max_tokens_per_minute": 4_000_000,
    "seconds_to_pause_on_rate_limit": 15,
}
```

### Batch parameters

These work when `batch=True`.

| Parameter | Type | Default | Description |
| - | - | - | - |
| `batch_size` | `int` or `"auto"` | `10000` | The most requests in one batch. `"auto"` uses the provider's limit. |
| `batch_check_interval` | `int` | `60` | Seconds between status checks. |
| `delete_successful_batch_files` | `bool` | `False` | Delete the batch files at the provider after a batch succeeds. |
| `delete_failed_batch_files` | `bool` | `False` | Delete the batch files at the provider after a batch fails. |
| `completion_window` | `str` | `"24h"` | The time the provider has to finish a batch. Only some providers use it. |
| `azure_deployment` | `str` | `None` | The Azure OpenAI deployment name. The `azure` backend needs it. |

```python theme={null}
backend_params = {
    "batch_size": 100,
    "batch_check_interval": 30,
    "delete_successful_batch_files": True,
    "delete_failed_batch_files": False,
}
```

### Offline parameters

These work with the `vllm` backend.

| Parameter | Type | Default | Description |
| - | - | - | - |
| `tensor_parallel_size` | `int` | `1` | The number of GPUs for tensor parallelism. |
| `enforce_eager` | `bool` | `False` | Whether to force eager execution. |
| `max_model_length` | `int` | `4096` | The longest sequence the model can handle. |
| `max_tokens` | `int` | `4096` | The most tokens to generate. |
| `min_tokens` | `int` | `1` | The fewest tokens to generate. |
| `gpu_memory_utilization` | `float` | `0.95` | The share of GPU memory to use, from 0 to 1. |
| `batch_size` | `int` | `256` | The number of prompts vLLM processes together. |
| `dtype` | `str` | `"auto"` | The data type of the model weights. |

```python theme={null}
backend_params = {
    "tensor_parallel_size": 2,
    "max_model_length": 4096,
    "max_tokens": 2048,
    "gpu_memory_utilization": 0.85,
    "batch_size": 32,
}
```

## CuratorResponse

Every `LLM` call returns a `CuratorResponse` with these attributes.

### Data

| Attribute | Type | Description |
| - | - | - |
| `dataset` | `Dataset` | The output rows. |
| `cache_dir` | `str \| None` | The cache directory of the run. |
| `failed_requests_path` | `Path \| None` | The file with the failed requests, if any failed. |
| `viewer_url` | `str \| None` | The link to the run in the [Curator Viewer](/curator/viewer). |
| `batch_mode` | `bool` | Whether the run used batch mode. |

### Model

| Attribute | Type | Description |
| - | - | - |
| `model_name` | `str` | The model that ran. |
| `max_requests_per_minute` | `int \| None` | The request limit used in online mode. |
| `max_tokens_per_minute` | `int \| None` | The token limit used in online mode. |

### Statistics

| Attribute | Type | Description |
| - | - | - |
| `token_usage` | `TokenUsage` | The fields `input`, `output`, and `total`. |
| `cost_info` | `CostInfo` | The fields `total_cost`, `input_cost_per_million`, `output_cost_per_million`, and `projected_remaining_cost`. |
| `request_stats` | `RequestStats` | The fields `total`, `succeeded`, `failed`, `in_progress`, and `cached`. |
| `performance_stats` | `PerformanceStats` | The fields `total_time`, `requests_per_minute`, `input_tokens_per_minute`, `output_tokens_per_minute`, and `max_concurrent_requests`. |
| `metadata` | `dict` | More details about the run. |

## Other classes and functions

| Name | Description |
| - | - |
| `curator.CodeExecutor` | Runs code for each row of a dataset. See [Code execution](/curator/guides/code-execution). |
| `curator.types.Image` | An image in a multimodal prompt. See [Multimodal data](/curator/guides/multimodal). |
| `curator.push_to_viewer(dataset)` | Uploads a `Dataset` to the Curator Viewer and returns the link. |
| `curator.load_dataset(dataset_id)` | Downloads a dataset from the Curator Viewer. |
| `curator.TinkerTrainer`, `curator.FireworksTrainer` | Fine-tune a model on your data with Tinker or Fireworks AI. See the [Curator README](https://github.com/bespokelabsai/curator#-fine-tuning-with-tinker). |

## Environment variables

| Variable | Description | Default |
| - | - | - |
| `CURATOR_VIEWER` | Set to `1` to stream data to the [Curator Viewer](/curator/viewer). | Off |
| `BESPOKE_API_KEY` | Your Bespoke Labs API key. It links viewer datasets to your account. | None |
| `CURATOR_DISABLE_CACHE` | Set to `1` to turn off [caching](/curator/caching). | Off |
| `CURATOR_CACHE_DIR` | The cache directory. | `~/.cache/curator` |
| `CURATOR_DISABLE_RICH_DISPLAY` | Set to `1` to replace the Rich progress display with tqdm logging. This helps when you debug with `pdb` or breakpoints. | Off |
| `TELEMETRY_ENABLED` | Set to `False` to turn off anonymous usage telemetry. | `True` |
