> ## Documentation Index
> Fetch the complete documentation index at: https://docs.bespokelabs.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Pricing

> What each Nimble model costs, and how Nimble counts tokens.

Nimble charges for input tokens only. Output tokens are free. You pay from prepaid credit, which you buy in the [console](https://console.bespokelabs.ai).

| Model | Price per million input tokens |
| - | - |
| `nimble-latest`, `nimble-v3`, `bespokelabs/Bespoke-Nimble-9B` | \$0.04 |
| `nimble-factcheck-lite`, `nimble-codegrep-lite` | \$0.02 |
| `nimble-factcheck`, `nimble-codegrep` | \$0.04 |
| `nimble-factcheck-max`, `nimble-codegrep-max` | \$0.12 |

The `usage.input_tokens` field of each response is the number of input tokens that you pay for.

## How Nimble counts input tokens

* For the general models, Nimble builds one prompt for each question. Every prompt starts with the same shared text, which is the state and the questions. You pay for the shared text once per request, not once per question.
* For fact check, Nimble counts the tokens of the document and the claims.
* For code search, Nimble counts each item's prompt once. An item whose prompt is longer than 8,192 tokens is not scored, and you do not pay for it. With `nimble-codegrep`, Nimble also counts the prompts that it sends again to the larger model.

## Credit holds

When a request starts, Nimble holds back from your credit the most that the request could cost. When the request ends, you pay for the tokens that the request used, and Nimble releases the rest of the hold. A request that gets an error costs nothing.

If your available credit is less than the hold, the request gets `402`. For code search, the hold is 8,192 tokens for each question, and twice that for `nimble-codegrep`.
