> ## Documentation Index
> Fetch the complete documentation index at: https://docs.bespokelabs.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Caching and recovery

> How Curator caches responses so you can resume failed runs and reuse finished ones.

Curator caches the output of every `LLM` call. This helps in two cases.

* **Recovery.** A long run can fail or stop part way. When you start it again, Curator keeps the responses it already has and only sends the requests that are missing.
* **Reuse.** In a pipeline with several steps, you often change a later step and run the whole pipeline again. Curator reads the earlier steps from the cache, so they cost no time or money.

To see the cache work, run this code twice. The second run reads the response from the cache and does not call the model.

```python theme={null}
from bespokelabs import curator

llm = curator.LLM(model_name="gpt-4o-mini")
poem = llm("Write a poem about the importance of data in AI.")
print(poem.dataset.to_pandas())
```

## Turn off caching

Set `CURATOR_DISABLE_CACHE` before you run your code.

```bash theme={null}
export CURATOR_DISABLE_CACHE=1
```

## Change the cache directory

Curator saves the cache in `~/.cache/curator` by default. You can change this in two ways.

* Set the `CURATOR_CACHE_DIR` environment variable to the directory you want.
* Pass `working_dir` when you call the `LLM` object, e.g., `llm("Write a poem.", working_dir="/path/to/my/poems")`.

## What decides a cache hit

Curator gives each run a fingerprint. Two runs share a cache entry when all of these are the same.

* The input dataset.
* The `prompt` method.
* The model name.
* The response format.
* Whether the run uses batch mode.
* The generation parameters, e.g., `temperature`.

The `parse` method is not part of the fingerprint. If you change only `parse`, Curator reuses the cached responses and runs the new `parse` on them.

## What is in the cache directory

This part may change in future versions. The cache directory holds these items.

* `metadata.db`, a SQLite database with details about each run.
* One directory per run, named after the run's fingerprint.

```bash theme={null}
ls ~/.cache/curator
# 032bc5ead2892f8f        6d1f31229726231d
# 137851647e75f9a7        91f4ad23d5821c9f
# 24b1d8917f7ef6f1        a2a3c8e5a58e3fc3
# metadata.db
```

## Troubleshooting

If the cache directory gets too large or is corrupt, delete it.

```bash theme={null}
rm -rf ~/.cache/curator
```

<Warning>
  This deletes all cached responses, and Curator will call the models again on the next run.
</Warning>
