Skip to main content
The model is on Hugging Face as bespokelabs/Bespoke-MiniCheck-7B. You can run it in two ways:
  • With the MiniCheck Python library. The library splits long documents into chunks and returns a support probability for each claim. Use this if you want scores in Python.
  • With a vLLM server. vLLM serves the model behind an OpenAI-compatible API. Use this if you want a server that other programs can call.
Both ways need an NVIDIA GPU. To try the model first, use this Colab notebook.
The model is under the CC BY-NC 4.0 license, which does not allow commercial use. For a commercial license, email company@bespokelabs.ai.

Use the MiniCheck library

1

Install the library

2

Score claims

The first run downloads the model to ./ckpts. By default, score splits each document into chunks of about 32K tokens. To use a different size, pass chunk_size. For more options, see the MiniCheck repository.

Run a vLLM server with Docker

This command starts the official vLLM image and serves the model on port 8000:
Notes on the command:
  • The model repository has custom code, so --trust-remote-code is required.
  • The model is public, so you do not need a Hugging Face token.
  • --api-key makes clients send that key. Set VLLM_API_KEY before you run the command.
  • bfloat16 needs a GPU with compute capability 8.0 or higher. On an older GPU, use --dtype float16.
The server returns the model’s text answer, which is “Yes” or “No”. To get a support probability instead, send the same prompt that the MiniCheck library sends and read the probability of the “Yes” token. The prompt and scoring code are in minicheck/inference.py.