The
usage.input_tokens field of each response is the number of input tokens that you pay for.
How Nimble counts input tokens
- For the general models, Nimble builds one prompt for each question. Every prompt starts with the same shared text, which is the state and the questions. You pay for the shared text once per request, not once per question.
- For fact check, Nimble counts the tokens of the document and the claims.
- For code search, Nimble counts each item’s prompt once. An item whose prompt is longer than 8,192 tokens is not scored, and you do not pay for it. With
nimble-codegrep, Nimble also counts the prompts that it sends again to the larger model.
Credit holds
When a request starts, Nimble holds back from your credit the most that the request could cost. When the request ends, you pay for the tokens that the request used, and Nimble releases the rest of the hold. A request that gets an error costs nothing. If your available credit is less than the hold, the request gets402. For code search, the hold is 8,192 tokens for each question, and twice that for nimble-codegrep.