Skip to main content

Limits

These limits apply to every model. nimble-latest and nimble-v3 also have these limits. The previous model, bespokelabs/Bespoke-Nimble-9B, allows 8,192 prompt tokens for one question. Nimble never cuts a prompt. A request above a limit gets an error. The fact check and code search guides list the limits of those models.

Errors

You are not charged for a request that gets an error. A 422 with the detail “Nimble request failed” comes from a general model itself. One cause is a prompt that is longer than the model allows. Another cause is a model name that Nimble does not know. These are the model names:
  • nimble-latest, which is nimble-v3 now
  • nimble-v3 (or bespokelabs/Bespoke-Nimble-9B-v3)
  • bespokelabs/Bespoke-Nimble-9B, the previous model
  • nimble-factcheck-lite, nimble-factcheck, nimble-factcheck-max
  • nimble-codegrep-lite, nimble-codegrep, nimble-codegrep-max

Request IDs

Every response has an x-request-id header, including an error response. The x-typesafe-request-id header has the same ID. Error responses also have it in the request_id field, and so do fact check and code search responses. Send this ID to Bespoke Labs when you report a problem with a request.