Skip to main content
This recipe asks Claude questions with thinking turned on, and saves both the thinking and the final answer for each question.

Prerequisites

  • Python 3.10 or later.
  • Curator, installed with pip install bespokelabs-curator.
  • An Anthropic API key.
1

Set the API key

2

Create an LLM subclass

Set return_completions_object = True so that parse gets the full API response instead of only the text. The response has a list of content blocks. A thinking block holds the thinking and a text block holds the answer.
3

Configure the model

Curator passes generation_params to the Anthropic Messages API as they are.
With adaptive thinking, Claude decides how much to think for each question. Set display to "summarized" to get a summary of Claude’s reasoning in each response.
4

Generate the data

The output looks like this.
Thinking settings differ between Claude models. Older models such as Claude Haiku 4.5 use {"type": "enabled", "budget_tokens": 8000} instead of adaptive thinking, and newer models reject budget_tokens. Check the Anthropic extended thinking docs for the model you use.

Use batch mode

For large datasets, set batch=True to use the Anthropic batch API at a lower price. The rest of the code stays the same. See Batch inference.
For more settings, see the API reference.