Skip to main content
You build a Curator pipeline from subclasses of curator.LLM. Each subclass has two methods.
  • prompt takes one input row and returns the prompt for the LLM.
  • parse takes the same input row and the LLM response, and returns one or more output rows.
Here is a small example.

prompt

Curator calls prompt once for each input row, and sends the requests in parallel. The method can return one of these.
  • A string. Curator sends it as a single user message.
  • A list of messages, e.g., [{"role": "system", "content": "..."}, {"role": "user", "content": "..."}].
  • A tuple of a string and an image or file, for multimodal prompts.
If you do not override prompt, Curator sends the input row as the prompt. This is why curator.LLM(model_name=...) works with a plain list of strings. To add the same system message to every request, pass system_prompt when you create the object. Do not also return a system message from prompt, because Curator raises an error when both are set.

parse

Curator calls parse with two arguments.
  • The input row that went into prompt.
  • The LLM response. This is a string by default. If you set a response format, it is an instance of your Pydantic model.
The method returns a dictionary for one output row, or a list of dictionaries for several rows. If you do not override parse, Curator returns {"response": response}. Curator does not include the parse function when it decides whether a run is cached. If you change only parse, Curator reuses the cached responses and runs your new parse on them without calling the model again.

Inputs

You can call an LLM object with any of these inputs.
  • A single string or a single list of messages.
  • A list of strings or a list of dictionaries.
  • A Hugging Face Dataset.
  • The CuratorResponse from an earlier call. Curator uses its dataset.
  • No input. Curator then sends one request, which is useful for a first step that generates seed data.

Data flow

This is how two input rows become four output rows when parse returns two rows for each response.
Curator prompts the LLM for rows A and B in parallel. Each call returns one response, and parse turns each response into two new rows. The output dataset holds all four rows. Because the output of one LLM call can be the input of the next, you can chain several LLM objects to build a dataset step by step. Structured output shows an example.

CuratorResponse

Every call returns a CuratorResponse. These are the attributes you will use most. See the API reference for the full list.