bespokelabs/Bespoke-MiniCheck-7B. You can run it in two ways:
- With the MiniCheck Python library. The library splits long documents into chunks and returns a support probability for each claim. Use this if you want scores in Python.
- With a vLLM server. vLLM serves the model behind an OpenAI-compatible API. Use this if you want a server that other programs can call.
The model is under the CC BY-NC 4.0 license, which does not allow commercial use. For a commercial license, email company@bespokelabs.ai.
Use the MiniCheck library
1
Install the library
2
Score claims
./ckpts. By default, score splits each document into chunks of about 32K tokens. To use a different size, pass chunk_size. For more options, see the MiniCheck repository.
Run a vLLM server with Docker
This command starts the official vLLM image and serves the model on port 8000:- The model repository has custom code, so
--trust-remote-codeis required. - The model is public, so you do not need a Hugging Face token.
--api-keymakes clients send that key. SetVLLM_API_KEYbefore you run the command.bfloat16needs a GPU with compute capability 8.0 or higher. On an older GPU, use--dtype float16.
minicheck/inference.py.