A model that is not on this list cannot be used for anything. Each one
says how it is paid for: subscription and flat-rate run
on plans we already pay for, metered is billed per token and has
to clear a spend gate first, and local runs on our own machine
with no provider involved.
opencode-go/kimi-k3
opencode-go
·
262k
context
· flat-rate
·
no per-token billUsed for
- •
Extract claims and nodes from a reviewed ingest. (only if the first choice is unavailable)
- •
Run the same extraction across models to compare quality. (only if the first choice is unavailable)
Moonshot Kimi K3 through the opencode Go plan, a monthly allowance rather than a per-token bill. Added 2026-09-04 so extraction has a second provider and keeps moving when the Claude allowance is paced. Measured the same day over 8 digests it had the highest grounding of any model here (98.9%) and one unresolved quote, but it has not been tried on a book or a long transcript, so it sits above Haiku and below Sonnet and Opus until it has.
opencode-go/glm-5.2
opencode-go
·
200k
context
· flat-rate
·
no per-token billUsed for
- •
Run the same extraction across models to compare quality. (only if the first choice is unavailable)
Z.ai GLM 5.2 through the opencode Go plan, a monthly allowance rather than a per-token bill. Variant only for now - measured 2026-09-04 over 22 digests it had the lowest quote fidelity of any model here (0.835) and a quarter more claims per digest than Sonnet, and on one verbatim transcript a quarter of its quotes did not resolve.
claude-opus-5
anthropic
·
1M
context
· subscription
·
list $5 / $25 per million tokensnot for articles —
anthropic watermarks its output
Used for
- •
Extract claims and nodes from a reviewed ingest.
- •
Check that assertions in an assembled article trace to graph sources. (only if the first choice is unavailable)
The strongest model we have access to, at $5/$25 per million tokens - two and a half times Sonnet on input, not the outlier it was once priced at. Not used for anything at present. It cannot write articles because we cannot establish whether Anthropic watermarks its output.
claude-sonnet-5
anthropic
·
1M
context
· subscription
·
list $2 / $10 per million tokensnot for articles —
anthropic watermarks its output
wrote 253 of the pages here
Used for
- •
Extract claims and nodes from a reviewed ingest. (only if the first choice is unavailable)
- •
Match, merge and score entities across digests.
- •
Merge entities that are the same thing. (only if the first choice is unavailable)
- •
Judge whether two similar claims corroborate each other.
- •
Judge whether two records refer to the same specific subject (an incident, operation, programme or document), producing possibly-related record pairs for human review. Experimental. (only if the first choice is unavailable)
- •
Check that assertions in an assembled article trace to graph sources.
- •
Run the same extraction across models to compare quality.
Pulls claims out of source documents, and wrote most of the articles here before the watermarking rule barred it from writing. Its tokeniser changed with the 5 series, so older cost measurements do not carry over.
claude-haiku-4-5
anthropic
·
200k
context
· subscription
·
list $1 / $5 per million tokensnot for articles —
anthropic watermarks its output
Used for
- •
Extract claims and nodes from a reviewed ingest. (only if the first choice is unavailable)
- •
Match, merge and score entities across digests. (only if the first choice is unavailable)
- •
Merge entities that are the same thing.
- •
Judge whether two similar claims corroborate each other. (only if the first choice is unavailable)
- •
Judge whether two records refer to the same specific subject (an incident, operation, programme or document), producing possibly-related record pairs for human review. Experimental.
- •
Run the same extraction across models to compare quality. (only if the first choice is unavailable)
- •
Post-ingest metadata research for person names, known terms and source metadata.
Used where it is accurate enough, because it is cheaper. Its 200,000-token window is the smallest here, so a job that fits on Sonnet may not fit on it.
openai/gpt-5.6-luna
openai
·
1M
context
· metered
·
$0.2 / $1.2 per million tokensmay write articles
wrote 34 of the pages here
Used for
- •
Extract text faithfully from a source document.
- •
Run the same extraction across models to compare quality. (only if the first choice is unavailable)
- •
Write a reader-facing article or record page. (only if the first choice is unavailable)
- •
Render an article into another language. (only if the first choice is unavailable)
The fast, cheap end of its family. OpenAI points it at repetitive work with clear rules - extracting fields, tagging and classifying records - which is what our extraction stage does. It writes less than larger models when nothing tells it how long to be.
openai/gpt-5.6-sol
openai
·
1M
context
· metered
·
$2 / $10 per million tokensmay write articles
wrote 1 of the pages here
Used for
- •
Write a reader-facing article or record page.
- •
Render an article into another language.
The flagship of its family - the one OpenAI points at the hardest work. Its listed $2/$10 is promotional through at least 21 November 2026; standard is $4/$20, so this price will roughly double.
openai/gpt-5.6-terra
openai
·
1M
context
· metered
·
$2 / $12 per million tokensmay write articles
Used for
- •
Write a reader-facing article or record page. (only if the first choice is unavailable)
- •
Render an article into another language. (only if the first choice is unavailable)
The balanced middle of its family, the closest match to Sonnet in role. Untested here.
openai-subscription/gpt-5.6-sol
openai
·
272k
context
· openai-subscription
·
no per-token billNot used for anything yet.
Preferred execution route for a logical Sol assembly request when its allowance, input fit, transport qualification and production activation gates pass. Oversized payloads are refused on this route; the equivalent metered route may run only when separately enabled and spend-authorised.
openai-subscription/gpt-5.6-terra
openai
·
272k
context
· openai-subscription
·
no per-token billNot used for anything yet.
Candidate subscription route for manual assembly evaluation. It is outside automatic priority, so adding it does not change the production default. Oversized payloads are refused; a larger-context route requires separate selection and authorisation rather than automatic fallback.
openai-subscription/gpt-5.6-luna
openai
·
272k
context
· openai-subscription
·
no per-token billNot used for anything yet.
Candidate subscription route for manual assembly evaluation. It is outside automatic priority, so adding it does not change the production default. Oversized payloads are refused; a larger-context route requires separate selection and authorisation rather than automatic fallback.
moonshotai/kimi-k3
moonshot
·
1M
context
· metered
·
$2.648138063 / $13.28272425 per million tokensNot used for anything yet.
Second in independent rankings for English prose, ahead of every OpenAI model. Its weights are published, so it can be run without going through a provider at all.
deepseek/deepseek-v4-flash
deepseek
·
1M
context
· metered
·
$0.06566 / $0.13132 per million tokensNot used for anything yet.
A cheap fast model with published weights. Not tested here yet.
z-ai/glm-5.3-flash
z-ai
·
1M
context
· metered
·
$0.075 / $0.25 per million tokensNot used for anything yet.
Not yet documented.
qwen/qwen3.8-flash
qwen
·
1M
context
· metered
·
$0.15 / $0.47 per million tokensNot used for anything yet.
Not yet documented.
MoritzLaurer/DeBERTa-v3-base-mnli-fever-anli
moritzlaurer
·
512
context
· local
·
no per-token billUsed for
- •
Test whether each claim is warranted by its quote, or by the record around it.
A 184-million-parameter natural-language-inference classifier run on our own machine. It writes nothing; it labels a premise and a hypothesis as entails, neutral or contradicts. Stage one of the per-claim entailment check, the quote against the claim. Validated on 50 hand-labelled claims and 12 planted contradictions (2026-09-02).
MoritzLaurer/DeBERTa-v3-large-mnli-fever-anli-ling-wanli
moritzlaurer
·
512
context
· local
·
no per-token billUsed for
- •
Test whether each claim is warranted by its quote, or by the record around it. (only if the first choice is unavailable)
The larger sibling, used only for stage two of the entailment check, where a claim the quote alone leaves neutral is re-checked against the record text around the quote.
Qwen/Qwen3-Reranker-0.6B
qwen
·
32k
context
· local
·
no per-token billUsed for
- •
Rank candidate node merges for human review.
A small model run on our own machine that, given two names from the knowledge graph, judges whether they refer to the same thing - "Cmdr. Fravor" and "David Fravor", or two spellings of one country. It only orders the list a reviewer works through; nothing merges without a person. Tried on 2026-09-02 against 90 pairs a reviewer had already merged, it ranked them well ahead of our older name-matching rules and found most of the true pairs those rules had never noticed. Its known weakness is relatives and people from the same case, whom it tends to treat as one person, so a second check stays in front of any such merge.