StratifiedGenerator block makes the answers more varied. It is based on the SimpleStrat method from the paper Stratified Generation for Artificial Data in Question Answering.
Example
question column. The output dataset has a question column and an answer column.
How it works
StratifiedGenerator runs four curator.LLM steps in a row.
- For each question, the LLM lists true or false properties that an answer may have, e.g., “The element is a metal.”
- For each property, the LLM estimates how likely it is that a random answer has it.
- The LLM removes properties that overlap and keeps at most three, preferring ones close to a 50% chance.
- For each question, Curator picks one property at random, weighted by its probability, and adds it to the question. The LLM then answers the new question.
Settings
StratifiedGenerator passes its arguments to each of the four steps, so it takes the same arguments as curator.LLM.
Save the results
The result is a Hugging FaceDataset, so you can push it to the Hub.
Troubleshooting
If the answers are still too similar, raisetemperature in generation_params.
If you get API errors, try another model or change the rate limits in backend_params. See the API reference.