> ## Documentation Index
> Fetch the complete documentation index at: https://docs.bespokelabs.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Batch inference

> Send large jobs to provider batch APIs by setting batch=True.

Several providers offer a batch API. You upload many prompts at once, the provider processes them within a time window, and the price is usually about half the normal price. Batch APIs take work to use directly.

* You have to write a batch file, upload it, and check for results until they are ready.
* Each provider limits the size of a batch, so you have to split a large dataset into several batches and track each one.

Curator does all of this for you. You set `batch=True` when you create an `LLM` object.

## Supported providers

| Provider | `backend` value | Chosen automatically when |
| - | - | - |
| OpenAI | `openai` | The model is an OpenAI model, e.g., `gpt-4o-mini`. |
| Anthropic | `anthropic` | The model name contains `claude`. |
| Gemini on Vertex AI | `gemini` | Never. Set `backend="gemini"`. |
| Mistral | `mistral` | Never. Set `backend="mistral"`. |
| Azure OpenAI | `azure` | Never. Set `backend="azure"`. |
| inference.net | `inference.net` | Never. Set `backend="inference.net"`. |

The `litellm` backend does not support batch mode. The `mistral` and `azure` backends support only batch mode.

## Example

This example writes new responses for the first messages in the [WildChat](https://huggingface.co/datasets/allenai/WildChat) dataset. The class is the same for every provider.

```python theme={null}
import logging

from datasets import load_dataset

from bespokelabs import curator

# Show details about how Curator processes the batches.
logger = logging.getLogger("bespokelabs.curator")
logger.setLevel(logging.INFO)


class WildChatReannotator(curator.LLM):
    """Write a new response to the first message of each conversation."""

    def prompt(self, input: dict) -> str:
        return input["conversation"][0]["content"]

    def parse(self, input: dict, response: str) -> dict:
        return {"instruction": input["conversation"][0]["content"], "new_response": response}


dataset = load_dataset("allenai/WildChat", split="train")
dataset = dataset.select(range(100))
```

Next, set up your provider and create the object with `batch=True`.

<Tabs>
  <Tab title="OpenAI">
    Set your API key.

    ```bash theme={null}
    export OPENAI_API_KEY=<your-api-key>
    ```

    ```python theme={null}
    reannotator = WildChatReannotator(model_name="gpt-4o-mini", batch=True)
    ```

    To use another API that is compatible with the OpenAI batch API, set `backend="openai"` and pass `base_url` and `api_key` in `backend_params`.
  </Tab>

  <Tab title="Anthropic">
    Set your API key.

    ```bash theme={null}
    export ANTHROPIC_API_KEY=<your-api-key>
    ```

    ```python theme={null}
    reannotator = WildChatReannotator(model_name="claude-haiku-4-5", batch=True)
    ```
  </Tab>

  <Tab title="Gemini">
    The `gemini` backend uses Vertex AI batch prediction. You need a Google Cloud project with Vertex AI turned on, and a Cloud Storage bucket that Curator can write to.

    ```bash theme={null}
    export GOOGLE_CLOUD_PROJECT=<project-id>
    export GEMINI_BUCKET_NAME=<bucket-name>
    export GOOGLE_CLOUD_REGION=us-central1  # optional, us-central1 is the default
    ```

    Sign in with Application Default Credentials.

    ```bash theme={null}
    gcloud auth application-default login
    ```

    ```python theme={null}
    reannotator = WildChatReannotator(model_name="gemini-2.5-flash", backend="gemini", batch=True)
    ```

    You can pass Gemini [generation parameters](https://cloud.google.com/vertex-ai/generative-ai/docs/reference/python/latest/vertexai.generative_models.GenerationConfig) in `generation_params`.
  </Tab>

  <Tab title="Mistral">
    Get a key from the [Mistral console](https://console.mistral.ai/api-keys). Batch mode needs a paid key.

    ```bash theme={null}
    export MISTRAL_API_KEY=<your-api-key>
    ```

    ```python theme={null}
    reannotator = WildChatReannotator(model_name="mistral-small-latest", backend="mistral", batch=True)
    ```
  </Tab>

  <Tab title="Azure OpenAI">
    Set your API key.

    ```bash theme={null}
    export AZURE_OPENAI_API_KEY=<your-api-key>
    ```

    Set `model_name` to the model behind your deployment. Curator uses it to look up prices. Set `azure_deployment` to the name of your batch deployment, and `base_url` to the v1 endpoint of your resource.

    ```python theme={null}
    reannotator = WildChatReannotator(
        model_name="gpt-4o-mini",
        backend="azure",
        batch=True,
        backend_params={
            "base_url": "https://<resource>.openai.azure.com/openai/v1/",
            "azure_deployment": "<deployment-name>",
        },
    )
    ```
  </Tab>
</Tabs>

Then run it.

```python theme={null}
result = reannotator(dataset)
print(result.dataset)
print(result.dataset[0])
```

The output looks like this.

| instruction | new\_response |
| - | - |
| Write a very long, elaborate, descriptive and ... | Scene: Omelette Apocalypse\n\n\*\*INT. DINER... |
| what are you? | I am a large language model, trained by ... |

## Batch settings

Pass these keys in `backend_params` when `batch=True`.

| Parameter | Default | Description |
| - | - | - |
| `batch_size` | `10000` | The largest number of requests in one batch. Set it to `"auto"` to use the provider's limit. |
| `batch_check_interval` | `60` | Seconds between checks on the status of running batches. |
| `delete_successful_batch_files` | `False` | Delete the batch files at the provider after a batch succeeds. |
| `delete_failed_batch_files` | `False` | Delete the batch files at the provider after a batch fails. Keep them if you want to debug the failure. |
| `completion_window` | `"24h"` | The time the provider has to finish the batch. Only some providers use this. |
| `azure_deployment` | None | The Azure OpenAI deployment name. Only the `azure` backend uses this. |

```python theme={null}
backend_params = {
    "batch_size": 1_000,
    "batch_check_interval": 10,
    "delete_successful_batch_files": True,
    "delete_failed_batch_files": False,
}
```

The [API reference](/curator/api-reference#backend-parameters) lists the settings that apply to every backend, e.g., `max_retries`.

## Cancel running batches

To cancel the batches of a run, call the object again with the same input and `batch_cancel=True`. Curator asks you to confirm before it cancels them.

```python theme={null}
reannotator(dataset, batch_cancel=True)
```
