> ## Documentation Index
> Fetch the complete documentation index at: https://docs.bespokelabs.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Key concepts

> How curator.LLM uses the prompt and parse methods to turn one dataset into another.

You build a Curator pipeline from subclasses of `curator.LLM`. Each subclass has two methods.

* `prompt` takes one input row and returns the prompt for the LLM.
* `parse` takes the same input row and the LLM response, and returns one or more output rows.

Here is a small example.

```python theme={null}
from bespokelabs import curator


class Translator(curator.LLM):
    def prompt(self, input: dict) -> str:
        return f"Translate this sentence to French: {input['sentence']}"

    def parse(self, input: dict, response: str) -> dict:
        return {"sentence": input["sentence"], "french": response}


translator = Translator(model_name="gpt-4o-mini")
result = translator([
    {"sentence": "The cat is on the table."},
    {"sentence": "I like to read in the morning."},
])
print(result.dataset.to_pandas())
```

## prompt

Curator calls `prompt` once for each input row, and sends the requests in parallel. The method can return one of these.

* A string. Curator sends it as a single user message.
* A list of messages, e.g., `[{"role": "system", "content": "..."}, {"role": "user", "content": "..."}]`.
* A tuple of a string and an image or file, for [multimodal prompts](/curator/guides/multimodal).

If you do not override `prompt`, Curator sends the input row as the prompt. This is why `curator.LLM(model_name=...)` works with a plain list of strings.

To add the same system message to every request, pass `system_prompt` when you create the object. Do not also return a system message from `prompt`, because Curator raises an error when both are set.

## parse

Curator calls `parse` with two arguments.

* The input row that went into `prompt`.
* The LLM response. This is a string by default. If you set a [response format](/curator/structured-output), it is an instance of your Pydantic model.

The method returns a dictionary for one output row, or a list of dictionaries for several rows. If you do not override `parse`, Curator returns `{"response": response}`.

Curator does not include the parse function when it decides whether a run is cached. If you change only `parse`, Curator reuses the cached responses and runs your new `parse` on them without calling the model again.

## Inputs

You can call an `LLM` object with any of these inputs.

* A single string or a single list of messages.
* A list of strings or a list of dictionaries.
* A Hugging Face `Dataset`.
* The `CuratorResponse` from an earlier call. Curator uses its `dataset`.
* No input. Curator then sends one request, which is useful for a first step that generates seed data.

## Data flow

This is how two input rows become four output rows when `parse` returns two rows for each response.

```text theme={null}
Input dataset:
  Row A
  Row B

Processing:
  Row A -> prompt(A) -> Response R1 -> parse(A, R1) -> [C, D]
  Row B -> prompt(B) -> Response R2 -> parse(B, R2) -> [E, F]

Output dataset:
  Row C
  Row D
  Row E
  Row F
```

Curator prompts the LLM for rows A and B in parallel. Each call returns one response, and `parse` turns each response into two new rows. The output dataset holds all four rows.

Because the output of one `LLM` call can be the input of the next, you can chain several `LLM` objects to build a dataset step by step. [Structured output](/curator/structured-output#chain-llm-calls) shows an example.

## CuratorResponse

Every call returns a `CuratorResponse`. These are the attributes you will use most.

| Attribute | Description |
| - | - |
| `dataset` | The output rows as a Hugging Face `Dataset`. |
| `viewer_url` | The link to the run in the [Curator Viewer](/curator/viewer), if it is turned on. |
| `token_usage` | Input, output, and total token counts. |
| `cost_info` | The total cost and the price per million tokens, when Curator knows the price of the model. |
| `request_stats` | Counts of total, succeeded, failed, and cached requests. |
| `failed_requests_path` | The path to a file with the failed requests, if any failed. |

See the [API reference](/curator/api-reference#curatorresponse) for the full list.
