> ## Documentation Index
> Fetch the complete documentation index at: https://docs.bespokelabs.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Diverse QA dataset

> Build a dataset of question and answer pairs across many subjects with three chained LLM steps.

This recipe builds a set of question and answer pairs that covers many subjects. It follows the approach of the CAMEL dataset. The pipeline first generates subjects, then subsubjects for each subject, and then questions and answers for each subsubject.

The pairs are "ungrounded", which means the LLM writes them instead of taking them from existing text. The full code is in the [ungrounded QA example](https://github.com/bespokelabsai/curator/tree/main/examples/ungrounded-qa) on GitHub.

<Steps>
  <Step title="Set up the environment">
    ```python theme={null}
    # pip install bespokelabs-curator
    import os
    from typing import List

    from pydantic import BaseModel, Field

    from bespokelabs import curator

    # Remove this line if you do not want to use the Curator Viewer.
    os.environ["CURATOR_VIEWER"] = "1"
    ```
  </Step>

  <Step title="Define the data models">
    These Pydantic models describe the structured output of each step.

    ```python theme={null}
    class Subject(BaseModel):
        """A single subject."""

        subject: str = Field(description="A subject")


    class Subjects(BaseModel):
        """A list of subjects."""

        subjects: List[Subject] = Field(description="A list of subjects")


    class QA(BaseModel):
        """A question and answer pair."""

        question: str = Field(description="A question")
        answer: str = Field(description="An answer")


    class QAs(BaseModel):
        """A list of question and answer pairs."""

        qas: List[QA] = Field(description="A list of QAs")
    ```
  </Step>

  <Step title="Generate subjects">
    This step generates broad subjects, e.g., "Math".

    ```python theme={null}
    class SubjectGenerator(curator.LLM):
        """Generate diverse subjects."""

        response_format = Subjects

        def prompt(self, input: dict) -> str:
            return "Generate a diverse list of 3 subjects. Keep it high-level (e.g. Math, Science)."

        def parse(self, input: dict, response: Subjects) -> list:
            return [{"subject": s.subject} for s in response.subjects]
    ```
  </Step>

  <Step title="Generate subsubjects">
    This step takes each subject and generates narrower topics. For "Math", it might return "Calculus", "Algebra", and "Statistics".

    ```python theme={null}
    class SubsubjectGenerator(curator.LLM):
        """Generate diverse subsubjects for a given subject."""

        response_format = Subjects

        def prompt(self, input: dict) -> str:
            return f"For the given subject {input['subject']}. Generate 3 diverse subsubjects. No explanation."

        def parse(self, input: dict, response: Subjects) -> list:
            return [{"subject": input["subject"], "subsubject": s.subject} for s in response.subjects]
    ```
  </Step>

  <Step title="Generate questions and answers">
    This step writes question and answer pairs for each subsubject. Each row keeps its subject and subsubject.

    ```python theme={null}
    class QAGenerator(curator.LLM):
        """Generate diverse questions and answers for a given subsubject."""

        response_format = QAs

        def prompt(self, input: dict) -> str:
            return f"For the given subsubject {input['subsubject']}. Generate 3 diverse questions and answers. No explanation."

        def parse(self, input: dict, response: QAs) -> list:
            return [
                {
                    "subject": input["subject"],
                    "subsubject": input["subsubject"],
                    "question": qa.question,
                    "answer": qa.answer,
                }
                for qa in response.qas
            ]
    ```
  </Step>

  <Step title="Run the pipeline">
    ```python theme={null}
    subject_generator = SubjectGenerator(model_name="gpt-4o-mini")
    subjects = subject_generator()

    subsubject_generator = SubsubjectGenerator(model_name="gpt-4o-mini")
    subsubjects = subsubject_generator(subjects)

    qa_generator = QAGenerator(model_name="gpt-4o-mini")
    qas = qa_generator(subsubjects)

    # Strip whitespace from the answers.
    qa_dataset = qas.dataset.map(lambda row: {"answer": row["answer"].strip()}, num_proc=2)

    print(qa_dataset.to_pandas())
    ```
  </Step>
</Steps>

## Example output

With 3 subjects, 3 subsubjects each, and 3 pairs each, the pipeline returns 27 rows. The first rows look like this.

| subject | subsubject | question | answer |
| - | - | - | - |
| Mathematics | Calculus | What is the derivative of the function f(x) = e^x? | The derivative of f(x) = e^x is also e^x. |
| Mathematics | Calculus | What is the fundamental theorem of calculus? | It states that differentiation and integration are inverse processes. |
| Mathematics | Linear Algebra | What are eigenvalues and eigenvectors? | For a matrix A, if Av = λv, then λ is an eigenvalue and v is an eigenvector. |

## Customize the pipeline

To change a step, subclass it and override `prompt`.

```python theme={null}
class CustomSubjectGenerator(SubjectGenerator):
    def prompt(self, input: dict) -> str:
        return "Generate a diverse list of 5 subjects spanning sciences, arts, and humanities."


class ComplexQAGenerator(QAGenerator):
    def prompt(self, input: dict) -> str:
        return f"""For the given subsubject {input['subsubject']}, generate 3 diverse questions and answers.
Include at least one factual question, one conceptual question, and one application question.
Make the questions challenging but clear."""
```

You can use this pipeline to make training data for question answering models or to test how much a model knows across subjects. You can extend it with question types, difficulty levels, or new subject areas.
