curator.CodeExecutor runs code for each row of a dataset and collects the output. This helps in cases like these.
- You want only code that runs without errors in your training data. Open Thoughts used this method.
- An LLM wrote code that draws a chart or renders a video, and you want the result.
- You are building agents that use tools.
Example
CodeExecutor has three methods.
codereturns the Python code to run for a row. The code usually comes from the row, e.g., from a column that acurator.LLMstep wrote.code_inputis optional. It returns the text that the code reads withinput().code_outputgets the row and the result of the run, and returns the output row. The result has the attributesstdout,stderr,message, anderror. The value ofmessageis"success","timeout", or"error".
CodeExecutor returns a Hugging Face Dataset.
Like curator.LLM, the code executor caches its results. If a run stops, you can start it again without losing the work that finished. A progress display shows how the run is going.
Execution settings
Passexecution_params when you call the executor. timeout is the most seconds one run can take. The default is 10.
Backends
The backend decides where the code runs. Curator uses thebespokelabs-sandbox package for this, and you choose the backend with the backend argument.
The old name
multiprocessing still works and means local.
Local
Docker
Install Docker, e.g., Docker Desktop, and make sure it is running. Curator uses thepython:3.12-slim image by default. To use a different image, set image in backend_params.
Ray
Use Ray when your dataset is too large to run on one machine. Set theRAY_ADDRESS environment variable to the address of your Ray cluster. If it is not set, Ray starts a local cluster.