POST /v1/systemone with one of the fact check models:
- Put the document in
state, as a string. - Add one
noulquestion for each claim, with the claim ininstructions. - The question IDs are your own. The answers use the same IDs.
Make a request
Read the answer
noulis the probability that the document supports every part of the claim.- A claim counts as supported when
noulis above 0.5. You can use a higher threshold, e.g., 0.8, when a wrong “supported” costs you more than a wrong “not supported”. - A claim that the document contradicts and a claim that the document does not mention are both not supported.
Models
Accuracy
We measured the fact check models on the test split of LLM-AggreFact, a public benchmark of 11 datasets with 29,320 claims. No model was trained on LLM-AggreFact. Each number is the balanced accuracy at a threshold of 0.5. Balanced accuracy is the average of two shares. One is the share of supported claims that a model calls supported. The other is the share of unsupported claims that it calls unsupported. The table averages it over the 11 datasets, so each dataset counts the same. We also compared the models with Jev 1.13.0 from TypeSafe, on the same claims. Jev scores each claim whole, so the comparison uses whole claims. The last column gives each model’s difference from Jev in points, with a 95% interval from 1,000 resamples within each dataset.
On whole claims:
nimble-factcheck-liteis 1.3 points behind Jev.nimble-factcheckis 0.5 points behind Jev. The interval includes zero, so the two are about even.nimble-factcheck-maxis 0.8 points ahead of Jev, and the whole interval is above zero.
nimble-factcheckandnimble-factcheck-maxsend a claim to larger models when the small model scores it between 0.2 and 0.8. We chose that range on this test set, so their numbers are slightly optimistic.- We measured Jev through TypeSafe’s System One API on 2026-09-25. Jev returns probabilities with two decimals, so some claims score exactly 0.50. Those claims count as supported.
Results for each dataset, whole claims
Results for each dataset, whole claims
CNN and XSum are from AggreFact, and MediaSum and MeetingBank are from TofuEval. RAGTruth has the most claims, 16,371, and Wice has the fewest, 358.
Options
split_claims is true by default. Nimble then scores each sentence of a claim on its own and keeps the lowest score. Set it to false to score each claim whole. Put it at the top level of the request, next to model. With TypeSafe’s SDK, pass extra_body={"split_claims": False} to system_one.
Limits
The questions must be
noul questions with no criteria. Any other question gets 422.
Nimble splits a long document into parts. A sentence counts as supported if any part supports it.