Limits
These limits apply to every model.nimble-latest and nimble-v3 also have these limits.
The previous model,
bespokelabs/Bespoke-Nimble-9B, allows 8,192 prompt tokens for one question.
Nimble never cuts a prompt. A request above a limit gets an error.
The fact check and code search guides list the limits of those models.
Errors
You are not charged for a request that gets an error.
A
422 with the detail “Nimble request failed” comes from a general model itself. One cause is a prompt that is longer than the model allows. Another cause is a model name that Nimble does not know. These are the model names:
nimble-latest, which isnimble-v3nownimble-v3(orbespokelabs/Bespoke-Nimble-9B-v3)bespokelabs/Bespoke-Nimble-9B, the previous modelnimble-factcheck-lite,nimble-factcheck,nimble-factcheck-maxnimble-codegrep-lite,nimble-codegrep,nimble-codegrep-max
Request IDs
Every response has anx-request-id header, including an error response. The x-typesafe-request-id header has the same ID. Error responses also have it in the request_id field, and so do fact check and code search responses. Send this ID to Bespoke Labs when you report a problem with a request.