Skip to main content
This recipe builds a set of question and answer pairs that covers many subjects. It follows the approach of the CAMEL dataset. The pipeline first generates subjects, then subsubjects for each subject, and then questions and answers for each subsubject. The pairs are “ungrounded”, which means the LLM writes them instead of taking them from existing text. The full code is in the ungrounded QA example on GitHub.
1

Set up the environment

2

Define the data models

These Pydantic models describe the structured output of each step.
3

Generate subjects

This step generates broad subjects, e.g., “Math”.
4

Generate subsubjects

This step takes each subject and generates narrower topics. For “Math”, it might return “Calculus”, “Algebra”, and “Statistics”.
5

Generate questions and answers

This step writes question and answer pairs for each subsubject. Each row keeps its subject and subsubject.
6

Run the pipeline

Example output

With 3 subjects, 3 subsubjects each, and 3 pairs each, the pipeline returns 27 rows. The first rows look like this.

Customize the pipeline

To change a step, subclass it and override prompt.
You can use this pipeline to make training data for question answering models or to test how much a model knows across subjects. You can extend it with question types, difficulty levels, or new subject areas.