Skip to main content
If you ask an LLM the same open question many times, e.g., “Name a periodic element”, it tends to give the same few answers. The StratifiedGenerator block makes the answers more varied. It is based on the SimpleStrat method from the paper Stratified Generation for Artificial Data in Question Answering.

Example

The input dataset needs a question column. The output dataset has a question column and an answer column.

How it works

StratifiedGenerator runs four curator.LLM steps in a row.
  1. For each question, the LLM lists true or false properties that an answer may have, e.g., “The element is a metal.”
  2. For each property, the LLM estimates how likely it is that a random answer has it.
  3. The LLM removes properties that overlap and keeps at most three, preferring ones close to a 50% chance.
  4. For each question, Curator picks one property at random, weighted by its probability, and adds it to the question. The LLM then answers the new question.
Because each question gets a different property, the answers cover more of the possible range.

Settings

StratifiedGenerator passes its arguments to each of the four steps, so it takes the same arguments as curator.LLM.

Save the results

The result is a Hugging Face Dataset, so you can push it to the Hub.

Troubleshooting

If the answers are still too similar, raise temperature in generation_params. If you get API errors, try another model or change the rate limits in backend_params. See the API reference.