Skip to main content
Nimble fact check takes a document and a list of claims. For each claim, it returns the probability that the document supports the claim. Send the request to POST /v1/systemone with one of the fact check models:
  • Put the document in state, as a string.
  • Add one noul question for each claim, with the claim in instructions.
  • The question IDs are your own. The answers use the same IDs.

Make a request

Read the answer

  • noul is the probability that the document supports every part of the claim.
  • A claim counts as supported when noul is above 0.5. You can use a higher threshold, e.g., 0.8, when a wrong “supported” costs you more than a wrong “not supported”.
  • A claim that the document contradicts and a claim that the document does not mention are both not supported.

Models

Accuracy

We measured the fact check models on the test split of LLM-AggreFact, a public benchmark of 11 datasets with 29,320 claims. No model was trained on LLM-AggreFact. Each number is the balanced accuracy at a threshold of 0.5. Balanced accuracy is the average of two shares. One is the share of supported claims that a model calls supported. The other is the share of unsupported claims that it calls unsupported. The table averages it over the 11 datasets, so each dataset counts the same. We also compared the models with Jev 1.13.0 from TypeSafe, on the same claims. Jev scores each claim whole, so the comparison uses whole claims. The last column gives each model’s difference from Jev in points, with a 95% interval from 1,000 resamples within each dataset. On whole claims:
  • nimble-factcheck-lite is 1.3 points behind Jev.
  • nimble-factcheck is 0.5 points behind Jev. The interval includes zero, so the two are about even.
  • nimble-factcheck-max is 0.8 points ahead of Jev, and the whole interval is above zero.
Two notes on these numbers:
  • nimble-factcheck and nimble-factcheck-max send a claim to larger models when the small model scores it between 0.2 and 0.8. We chose that range on this test set, so their numbers are slightly optimistic.
  • We measured Jev through TypeSafe’s System One API on 2026-09-25. Jev returns probabilities with two decimals, so some claims score exactly 0.50. Those claims count as supported.
CNN and XSum are from AggreFact, and MediaSum and MeetingBank are from TofuEval. RAGTruth has the most claims, 16,371, and Wice has the fewest, 358.

Options

split_claims is true by default. Nimble then scores each sentence of a claim on its own and keeps the lowest score. Set it to false to score each claim whole. Put it at the top level of the request, next to model. With TypeSafe’s SDK, pass extra_body={"split_claims": False} to system_one.

Limits

The questions must be noul questions with no criteria. Any other question gets 422. Nimble splits a long document into parts. A sentence counts as supported if any part supports it.