> ## Documentation Index
> Fetch the complete documentation index at: https://docs.bespokelabs.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Online processing

> Send requests in real time and control rate limits with backend parameters.

In online mode, which is the default, Curator sends each request to the provider as soon as it can and gets the response right away. This guide generates poem topics and then poems, and shows the settings that control how fast Curator sends requests.

## Prerequisites

* Python 3.10 or later.
* Curator, installed with `pip install bespokelabs-curator`.
* An API key for an LLM provider. This guide uses OpenAI.

<Steps>
  <Step title="Define the response formats">
    ```python theme={null}
    from typing import List

    from pydantic import BaseModel, Field


    class Topics(BaseModel):
        """A list of topics."""

        topics_list: List[str] = Field(description="A list of topics.")


    class Poems(BaseModel):
        """A list of poems."""

        poems_list: List[str] = Field(description="A list of poems.")
    ```
  </Step>

  <Step title="Generate topics">
    This `LLM` subclass generates a list of topics.

    ```python theme={null}
    from bespokelabs import curator


    class Muse(curator.LLM):
        response_format = Topics

        def prompt(self, input: dict) -> str:
            return "Generate 10 diverse topics that are suitable for writing poems about."

        def parse(self, input: dict, response: Topics) -> list:
            return [{"topic": t} for t in response.topics_list]


    muse = Muse(model_name="gpt-4o-mini", backend_params={"max_requests_per_minute": 100})
    topics = muse()
    print(topics.dataset["topic"])
    ```
  </Step>

  <Step title="Write poems about the topics">
    This subclass writes two poems for each topic.

    ```python theme={null}
    class Poet(curator.LLM):
        """A poet that writes poems about given topics."""

        response_format = Poems

        def prompt(self, input: dict) -> str:
            return f"Write two poems about {input['topic']}."

        def parse(self, input: dict, response: Poems) -> list:
            return [{"topic": input["topic"], "poem": p} for p in response.poems_list]


    poet = Poet(model_name="gpt-4o-mini", backend_params={"max_requests_per_minute": 100})
    poems = poet(topics)
    print(poems.dataset.to_pandas())
    ```
  </Step>
</Steps>

The output looks like this.

```text theme={null}
                                 topic                                               poem
0                   Dreams vs. reality  In the realm where dreams take flight,\nWhere ...
1                   Dreams vs. reality  Reality stands with open eyes,\nA weighty thro...
2  Urban loneliness in a bustling city  In the city's heart where shadows blend,\nAmon...
3  Urban loneliness in a bustling city  Among the crowds, I walk alone,\nA sea of face...
```

## Rate limit settings

Pass these keys in `backend_params` to control how fast Curator sends requests.

| Parameter | Description |
| - | - |
| `max_requests_per_minute` | The most requests Curator sends in one minute. Retries count toward this limit. |
| `max_tokens_per_minute` | The most tokens Curator uses in one minute. Input and output tokens both count. |
| `max_input_tokens_per_minute` | The most input tokens in one minute. Use this with providers that limit input and output tokens separately, e.g., Anthropic. |
| `max_output_tokens_per_minute` | The most output tokens in one minute. Use this with the same providers. |
| `max_concurrent_requests` | The most requests that can wait for a response at the same time. |
| `seconds_to_pause_on_rate_limit` | Seconds to wait after the provider returns a rate limit error. The default is 10. |

```python theme={null}
backend_params = {
    "max_requests_per_minute": 60,
    "max_tokens_per_minute": 1_000_000,
    "seconds_to_pause_on_rate_limit": 30,
}
```

If you do not set the limits, Curator tries to read them from the provider's response headers. If it cannot, it uses 200 requests per minute and 100,000 tokens per minute.

The [API reference](/curator/api-reference#backend-parameters) lists the settings that apply to every backend, e.g., `max_retries` and `request_timeout`.
