> ## Documentation Index
> Fetch the complete documentation index at: https://docs.bespokelabs.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Self-hosting

> Run Bespoke-MiniCheck-7B on your own GPU.

The model is on Hugging Face as [`bespokelabs/Bespoke-MiniCheck-7B`](https://huggingface.co/bespokelabs/Bespoke-MiniCheck-7B). You can run it in two ways:

* With the MiniCheck Python library. The library splits long documents into chunks and returns a support probability for each claim. Use this if you want scores in Python.
* With a vLLM server. vLLM serves the model behind an OpenAI-compatible API. Use this if you want a server that other programs can call.

Both ways need an NVIDIA GPU. To try the model first, use this [Colab notebook](https://colab.research.google.com/drive/1s-5TYnGV3kGFMLp798r5N-FXPD8lt2dm?usp=sharing).

<Note>
  The model is under the CC BY-NC 4.0 license, which does not allow commercial use. For a commercial license, email [company@bespokelabs.ai](mailto:company@bespokelabs.ai).
</Note>

## Use the MiniCheck library

<Steps>
  <Step title="Install the library">
    ```bash theme={null}
    pip install "minicheck[llm] @ git+https://github.com/Liyan06/MiniCheck.git@main"
    ```
  </Step>

  <Step title="Score claims">
    ```python theme={null}
    from minicheck.minicheck import MiniCheck

    doc = "A group of students gather in the school library to study for their upcoming final exams."
    claim_1 = "The students are preparing for an examination."
    claim_2 = "The students are on vacation."

    scorer = MiniCheck(model_name="Bespoke-MiniCheck-7B", enable_prefix_caching=False, cache_dir="./ckpts")
    pred_label, raw_prob, _, _ = scorer.score(docs=[doc, doc], claims=[claim_1, claim_2])

    print(pred_label)  # [1, 0]
    print(raw_prob)    # support probability for each claim
    ```
  </Step>
</Steps>

The first run downloads the model to `./ckpts`. By default, `score` splits each document into chunks of about 32K tokens. To use a different size, pass `chunk_size`. For more options, see the [MiniCheck repository](https://github.com/Liyan06/MiniCheck).

## Run a vLLM server with Docker

This command starts the official vLLM image and serves the model on port 8000:

```bash theme={null}
docker run --runtime nvidia --gpus all \
  -v ~/.cache/huggingface:/root/.cache/huggingface \
  -p 8000:8000 \
  --ipc=host \
  vllm/vllm-openai:latest \
  --model bespokelabs/Bespoke-MiniCheck-7B \
  --trust-remote-code \
  --dtype bfloat16 \
  --max-model-len 32768 \
  --api-key "$VLLM_API_KEY"
```

Notes on the command:

* The model repository has custom code, so `--trust-remote-code` is required.
* The model is public, so you do not need a Hugging Face token.
* `--api-key` makes clients send that key. Set `VLLM_API_KEY` before you run the command.
* `bfloat16` needs a GPU with compute capability 8.0 or higher. On an older GPU, use `--dtype float16`.

The server returns the model's text answer, which is "Yes" or "No". To get a support probability instead, send the same prompt that the MiniCheck library sends and read the probability of the "Yes" token. The prompt and scoring code are in [`minicheck/inference.py`](https://github.com/Liyan06/MiniCheck/blob/main/minicheck/inference.py).
