Skip to content

Every model we use, and what for

Articles here are written by an AI model from claims taken out of source documents. This page lists every model involved, what each one does, which ones are not allowed near the writing, and why.

What wrote the 1104 pages on this site

Counted from the pages themselves. Each one records the model that produced it, so this is what happened rather than what we intend.

not recorded 816 pages · 74%
claude-sonnet-5 253 pages · 23%
openai/gpt-5.6-luna 34 pages · 3%
openai/gpt-5.6-sol 1 page · 0%

“Not recorded” is a page written before we started stamping the model into it, or one written by hand.

Who checks it

Currently set to write

openai/gpt-5.6-sol

openai

Currently set to check

claude-sonnet-5

anthropic

Two different companies.

Checking that an article traces back to its sources. Must not be the provider that wrote the article, so it refuses OpenAI while writing is OpenAI. Nothing it writes is published, so a watermarking provider is fine here - and Claude is the strongest option available for the job.

The models

A model that is not on this list cannot be used for anything. Each one says how it is paid for: subscription and flat-rate run on plans we already pay for, metered is billed per token and has to clear a spend gate first, and local runs on our own machine with no provider involved.

opencode-go/kimi-k3

opencode-go · 262k context · flat-rate · no per-token bill

Used for

  • • Extract claims and nodes from a reviewed ingest. (only if the first choice is unavailable)
  • • Run the same extraction across models to compare quality. (only if the first choice is unavailable)

Moonshot Kimi K3 through the opencode Go plan, a monthly allowance rather than a per-token bill. Added 2026-09-04 so extraction has a second provider and keeps moving when the Claude allowance is paced. Measured the same day over 8 digests it had the highest grounding of any model here (98.9%) and one unresolved quote, but it has not been tried on a book or a long transcript, so it sits above Haiku and below Sonnet and Opus until it has.

opencode-go/glm-5.2

opencode-go · 200k context · flat-rate · no per-token bill

Used for

  • • Run the same extraction across models to compare quality. (only if the first choice is unavailable)

Z.ai GLM 5.2 through the opencode Go plan, a monthly allowance rather than a per-token bill. Variant only for now - measured 2026-09-04 over 22 digests it had the lowest quote fidelity of any model here (0.835) and a quarter more claims per digest than Sonnet, and on one verbatim transcript a quarter of its quotes did not resolve.

claude-opus-5

anthropic · 1M context · subscription · list $5 / $25 per million tokens
not for articles — anthropic watermarks its output

Used for

  • • Extract claims and nodes from a reviewed ingest.
  • • Check that assertions in an assembled article trace to graph sources. (only if the first choice is unavailable)

The strongest model we have access to, at $5/$25 per million tokens - two and a half times Sonnet on input, not the outlier it was once priced at. Not used for anything at present. It cannot write articles because we cannot establish whether Anthropic watermarks its output.

claude-sonnet-5

anthropic · 1M context · subscription · list $2 / $10 per million tokens
not for articles — anthropic watermarks its output wrote 253 of the pages here

Used for

  • • Extract claims and nodes from a reviewed ingest. (only if the first choice is unavailable)
  • • Match, merge and score entities across digests.
  • • Merge entities that are the same thing. (only if the first choice is unavailable)
  • • Judge whether two similar claims corroborate each other.
  • • Judge whether two records refer to the same specific subject (an incident, operation, programme or document), producing possibly-related record pairs for human review. Experimental. (only if the first choice is unavailable)
  • • Check that assertions in an assembled article trace to graph sources.
  • • Run the same extraction across models to compare quality.

Pulls claims out of source documents, and wrote most of the articles here before the watermarking rule barred it from writing. Its tokeniser changed with the 5 series, so older cost measurements do not carry over.

claude-haiku-4-5

anthropic · 200k context · subscription · list $1 / $5 per million tokens
not for articles — anthropic watermarks its output

Used for

  • • Extract claims and nodes from a reviewed ingest. (only if the first choice is unavailable)
  • • Match, merge and score entities across digests. (only if the first choice is unavailable)
  • • Merge entities that are the same thing.
  • • Judge whether two similar claims corroborate each other. (only if the first choice is unavailable)
  • • Judge whether two records refer to the same specific subject (an incident, operation, programme or document), producing possibly-related record pairs for human review. Experimental.
  • • Run the same extraction across models to compare quality. (only if the first choice is unavailable)
  • • Post-ingest metadata research for person names, known terms and source metadata.

Used where it is accurate enough, because it is cheaper. Its 200,000-token window is the smallest here, so a job that fits on Sonnet may not fit on it.

openai/gpt-5.6-luna

openai · 1M context · metered · $0.2 / $1.2 per million tokens
may write articles wrote 34 of the pages here

Used for

  • • Extract text faithfully from a source document.
  • • Run the same extraction across models to compare quality. (only if the first choice is unavailable)
  • • Write a reader-facing article or record page. (only if the first choice is unavailable)
  • • Render an article into another language. (only if the first choice is unavailable)

The fast, cheap end of its family. OpenAI points it at repetitive work with clear rules - extracting fields, tagging and classifying records - which is what our extraction stage does. It writes less than larger models when nothing tells it how long to be.

openai/gpt-5.6-sol

openai · 1M context · metered · $2 / $10 per million tokens
may write articles wrote 1 of the pages here

Used for

  • • Write a reader-facing article or record page.
  • • Render an article into another language.

The flagship of its family - the one OpenAI points at the hardest work. Its listed $2/$10 is promotional through at least 21 November 2026; standard is $4/$20, so this price will roughly double.

openai/gpt-5.6-terra

openai · 1M context · metered · $2 / $12 per million tokens
may write articles

Used for

  • • Write a reader-facing article or record page. (only if the first choice is unavailable)
  • • Render an article into another language. (only if the first choice is unavailable)

The balanced middle of its family, the closest match to Sonnet in role. Untested here.

openai-subscription/gpt-5.6-sol

openai · 272k context · openai-subscription · no per-token bill

Not used for anything yet.

Preferred execution route for a logical Sol assembly request when its allowance, input fit, transport qualification and production activation gates pass. Oversized payloads are refused on this route; the equivalent metered route may run only when separately enabled and spend-authorised.

openai-subscription/gpt-5.6-terra

openai · 272k context · openai-subscription · no per-token bill

Not used for anything yet.

Candidate subscription route for manual assembly evaluation. It is outside automatic priority, so adding it does not change the production default. Oversized payloads are refused; a larger-context route requires separate selection and authorisation rather than automatic fallback.

openai-subscription/gpt-5.6-luna

openai · 272k context · openai-subscription · no per-token bill

Not used for anything yet.

Candidate subscription route for manual assembly evaluation. It is outside automatic priority, so adding it does not change the production default. Oversized payloads are refused; a larger-context route requires separate selection and authorisation rather than automatic fallback.

moonshotai/kimi-k3

moonshot · 1M context · metered · $2.648138063 / $13.28272425 per million tokens

Not used for anything yet.

Second in independent rankings for English prose, ahead of every OpenAI model. Its weights are published, so it can be run without going through a provider at all.

deepseek/deepseek-v4-flash

deepseek · 1M context · metered · $0.06566 / $0.13132 per million tokens

Not used for anything yet.

A cheap fast model with published weights. Not tested here yet.

z-ai/glm-5.3-flash

z-ai · 1M context · metered · $0.075 / $0.25 per million tokens

Not used for anything yet.

Not yet documented.

qwen/qwen3.8-flash

qwen · 1M context · metered · $0.15 / $0.47 per million tokens

Not used for anything yet.

Not yet documented.

MoritzLaurer/DeBERTa-v3-base-mnli-fever-anli

moritzlaurer · 512 context · local · no per-token bill

Used for

  • • Test whether each claim is warranted by its quote, or by the record around it.

A 184-million-parameter natural-language-inference classifier run on our own machine. It writes nothing; it labels a premise and a hypothesis as entails, neutral or contradicts. Stage one of the per-claim entailment check, the quote against the claim. Validated on 50 hand-labelled claims and 12 planted contradictions (2026-09-02).

MoritzLaurer/DeBERTa-v3-large-mnli-fever-anli-ling-wanli

moritzlaurer · 512 context · local · no per-token bill

Used for

  • • Test whether each claim is warranted by its quote, or by the record around it. (only if the first choice is unavailable)

The larger sibling, used only for stage two of the entailment check, where a claim the quote alone leaves neutral is re-checked against the record text around the quote.

Qwen/Qwen3-Reranker-0.6B

qwen · 32k context · local · no per-token bill

Used for

  • • Rank candidate node merges for human review.

A small model run on our own machine that, given two names from the knowledge graph, judges whether they refer to the same thing - "Cmdr. Fravor" and "David Fravor", or two spellings of one country. It only orders the list a reviewer works through; nothing merges without a person. Tried on 2026-09-02 against 90 pairs a reviewer had already merged, it ranked them well ahead of our older name-matching rules and found most of the true pairs those rules had never noticed. Its known weakness is relatives and people from the same case, whom it tends to treat as one person, so a second check stays in front of any such merge.

Why some models are kept away from the writing

A watermark is a hidden signal a provider can add to the text its models generate. This site reproduces what sources say, so its writing is kept clear of one. What each provider does was researched on 28 August 2026, with the sources listed so you can check it rather than take our word.

clean
We looked and found no text watermarking on the date shown. What we checked is recorded against each provider.
unknown
We could not establish it either way. Treated the same as watermarking.
watermarks
The provider puts a hidden signal in the text it generates, so it can later be identified as machine-written.

anthropic

high confidence · checked 2 September 2026
watermarks since 2 August 2026
What we found

Anthropic announced text watermarking on 2026-08-14: "Future Claude models will generate text that contains a watermark", and "Claude's text watermark is a version of the SynthID-Text approach published by Google DeepMind" - a sampling-time scheme that "changes the source of randomness used when making some of those choices during text generation", so "Nothing is added to the text and there are no hidden characters". The support article states "Claude models launched on or after August 2, 2026 support marking at launch" and that "Marking will apply to output from supported models wherever Claude is offered, worldwide", covering Claude Platform (API), Claude, Claude Code, Claude Cowork and Claude Tag, plus cloud partners (AWS, Google Cloud). The driver is EU AI Act Article 50 transparency; Anthropic says "We will soon be offering a watermark detection API. We're in the process of working out the details of its implementation." Separately, C2PA signed provenance metadata is applied to generated image files (.svg/.png/.jpg) - that is metadata, distinct from the in-text statistical watermark. Anthropic has never published downloadable weights for any Claude model; all releases are hosted-API only (its 2026 open-weights position paper argues against a ban but does not release Claude weights). Re-checked 2026-09-02: the support article now states "Models currently supported include Fable 5.1 and Mythos 5.1", that marking is in progress for models launched before 2026-08-02, and that "Embedded watermarks will apply to all generated text" - so the state is no longer unknown. Anthropic watermarks; the models this project uses (Sonnet 5, Opus 5, Haiku 4.5) are scheduled to follow.

What we could not establish

The three models named in the question all launched BEFORE the 2026-08-02 cutoff - Haiku 4.5 (Oct 2025), Sonnet 5 (2026-06-30), Opus 5 (2026-07-24) - so none of them was marked at launch. They sit in the EU AI Act transition period: "we're working to add watermarking for those models as well... This will be rolled out over the coming months." Anthropic has published no model-by-model coverage table and no retrofit date, so as of 2026-08-28 it is NOT established that output from Opus 5, Sonnet 5 or Haiku 4.5 carries a watermark today; only that it is intended to. No public detector exists yet (detection API still unreleased), so this cannot be verified independently - third-party claims of "tested, not watermarked" are unverifiable without the key. Two further scope limits Anthropic states: code carries little or no watermark (functional code has too little word-choice freedom), and short passages may not contain enough signal to detect. Direction of travel is clear: treat any post-2026-08-02 Claude model as watermarked, and assume the named three become watermarked without notice.

Sources: 1. www.anthropic.com 2. support.claude.com 3. techcrunch.com 4. techcrunch.com

google

high confidence · checked 28 August 2026
watermarks since 14 May 2024
What we found

Google DeepMind's own SynthID page states: "We've expanded SynthID to watermarking and identifying text generated by the Gemini app and web experience," and describes the mechanism as adjusting per-token probability scores during generation ("not noticeable to the human eye"), i.e. a statistical sampling-time watermark, not metadata. The peer-reviewed Nature paper (23 Oct 2024, "Scalable watermarking for identifying large language model outputs") says outright: "non-distortionary SynthID-Text has been productionized and is currently watermarking responses in Gemini and Gemini Advanced," validated over roughly 20 million live responses (thumbs-up rate differed by 0.01%). Text watermarking was announced on 14 May 2024. For the hosted API specifically, a Google staff member on the official Google AI Developers Forum first said on 2026-08-05 "Generated text from the API is NOT SynthID-watermarked," then posted a correction on 2026-08-19: "After checking further with the team, it turns out that text generated via the Gemini API IS actually SynthID-watermarked... This also applies to text generated through Google AI Studio and/or antigravity!" Google also open-sourced SynthID-Text as a logits processor (Hugging Face Transformers v4.46.0+) for developers to apply to their own models.

What we could not establish

Three real gaps. (1) The API is the weakest link evidentially: the only statement that Gemini API / AI Studio output is watermarked is a forum post by a Google employee that reversed his own answer two weeks earlier; no Google documentation, model card, or Vertex AI page states it, and the poster's follow-up question about when API watermarking started was left unanswered. Treat the app/web experience as documented fact and the API as staff-asserted but undocumented. (2) There is no way to verify independently: the SynthID Detector portal is waitlisted (journalists/researchers) and its published modality list is images, video and audio; the Gemini Apps verification help page covers "images, videos, and audio" only. Detection needs Google's secret key. (3) The watermark is fragile by design - it degrades under paraphrase, translation and heavy editing, and short outputs carry too little signal. On weights: the Gemini family is closed, API-only, so hosted generation is the only path and it carries the watermark. Gemma is Google's separate open-weights family - a self-hosted Gemma emits no watermark unless the operator deliberately enables the open-sourced SynthID-Text logits processor. The "since" date is the text-watermarking announcement for the Gemini app; the API start date is unknown.

Sources: 1. deepmind.google 2. deepmind.google 3. www.nature.com 4. pmc.ncbi.nlm.nih.gov

deepseek

high confidence · checked 28 August 2026
clean open weights
What we found

DeepSeek publishes V4-Pro and V4-Flash weights openly on Hugging Face under the MIT licence ("This repository and the model weights are licensed under the MIT License"), and neither model card mentions watermarking, content labelling, or any AI-generated-content identifier. DeepSeek's Open Platform Terms of Service place the disclosure duty on the developer, not on the model: "You shall clearly disclose to your end users that the Output content is generated by AI, and may contain errors or omissions" - DeepSeek makes no claim to embed anything in the output itself. DeepSeek's own 1 September 2025 compliance announcement under China's Measures for Labeling of AI-Generated Content and the mandatory standard GB 45438-2025 describes exactly two mechanisms: an explicit visible notice ("Content generated by AI, for reference only" shown in the interface and appended to generated text) and an implicit identifier written into FILE METADATA (content type, service-provider name/code, content number) - both explicitly outside the scope of statistical text watermarking, and neither survives copy-paste. No DeepSeek research publication describes a text watermarking scheme and DeepSeek publishes no detector; third-party detection of DeepSeek text (Originality.ai, Sapling) is stylometric classification, which is what you would expect in the absence of a watermark.

What we could not establish

Two limits. First, DeepSeek's hosted API sampling stack is closed and unaudited, so absence of documentation cannot strictly disprove a hidden sampling-time watermark on the API path - the finding rests on DeepSeek discussing AI-content marking openly and describing only labels plus metadata. Second, China's GB 45438-2025 says implicit identifiers should include "watermarks where feasible", so a future move is regulatorily plausible; watch for it. None of this affects self-hosted use: a statistical text watermark is applied at sampling time by whatever code runs the model, so MIT-licensed weights you run yourself cannot carry a DeepSeek watermark under any circumstances.

Sources: 1. huggingface.co 2. huggingface.co 3. cdn.deepseek.com 4. cdn.deepseek.com

moonshot

medium confidence · checked 28 August 2026
clean open weights
What we found

No Moonshot documentation, model card, research paper, or announcement describes a statistical text watermark, and they publish no detection tool. The Kimi API docs return nothing on watermark/水印/标识; the Kimi-K2-Thinking and Kimi-K3 model cards make "no mention of watermarking in generated outputs". The only adjacent clause is regulatory-compliance language in the terms of service: platform ToS s3.3(3) forbids "Deleting, altering, or concealing the identification marks of artificially generated content that we have labeled, including both explicit and implicit", and the consumer agreement says "您不得以任何方式删除、篡改、隐匿 Kimi 智能助手在输出内容中生成的深度合成服务标识" - this tracks China's 标识办法 / GB 45438-2025, where for text the mandated implicit label is file metadata (service-provider name, content number, timestamp, model info) and a digital watermark is "recommended... but this is not mandated". Weights are freely downloadable: K2/K2-Instruct/K2-Thinking under a Modified MIT Licence on Hugging Face, K3 (published 27 July 2026) under a bespoke Kimi K3 Licence with revenue/attribution thresholds but unrestricted download. Separately, the only prominent "watermark + Chinese AI" story is US Treasury's claim that American LLMs' watermarks were found in Chinese models via distillation - other labs' marks surviving into Chinese models, not Moonshot applying its own.

What we could not establish

Two gaps. (1) Moonshot has never made a public negative statement ("we do not watermark"), so this is an absence-of-evidence conclusion across their docs, ToS, privacy policy, model cards and press coverage, not a denial. (2) The ToS phrase "implicit" identification marks is the one genuine ambiguity: read narrowly it is the GB 45438-2025 file-metadata implicit label (out of scope here), but the standard also permits content-embedded digital watermarks as an implicit label, and Moonshot publishes no technical detail on which they use for text delivered as an API JSON string, where there is no file to carry metadata. No third-party analysis of Kimi output for token-distribution bias or zero-width characters was found either way. Regardless of the hosted API, open weights mean any self-hosted Kimi deployment carries no Moonshot watermark.

Sources: 1. platform.kimi.ai 2. www.kimi.com 3. platform.kimi.ai 4. huggingface.co

moritzlaurer

high confidence · checked 2 September 2026
clean open weights
What we found

Open-weight classifier checkpoints (MIT licence) published on Hugging Face by an individual researcher and run on our own hardware. A classifier emits a label and a probability, not text, so there is no generated text to watermark; and open weights run locally cannot carry a provider-side sampling watermark in any case.

Sources: 1. huggingface.co

openai

high confidence · checked 11 September 2026
clean
What we found

OpenAI's own API guide "Content provenance" enumerates exactly which signals it applies and to what: "C2PA Content Credentials | Images | Signed metadata with issuer and AI-use details" and "SynthID | Images and audio | A watermark embedded directly in supported media". Text is absent from that list; no GPT-5.6 model (Sol, Terra, Luna) is mentioned as carrying any provenance signal, and the verification API only checks images and audio. OpenAI's help-centre provenance page, updated 2026-08-02 (the day EU AI Act Article 50 transparency obligations began), states forward-looking intent rather than deployment: "Consistent with our commitments under the European Commission's Code of Practice on Transparency of AI-generated content, our goal is to expand provenance signals to all modalities including text" - and OpenAI told City A.M. it is "currently working through the details", giving no timeline. Reporting through late August 2026 agrees: InfoQ's roundup of Article 50 compliance credits statistical token-sampling text watermarking to Anthropic (Claude models released on or after 2 August 2026) and Google (SynthID-Text in Gemini), and describes OpenAI's measures as C2PA metadata plus SynthID on images/audio only. OpenAI built a cryptographic text watermark years ago (WSJ, August 2024) and deliberately shelved it over paraphrase/translation fragility and disproportionate flagging of non-native English writers; that decision has not been reversed.

What we could not establish

No OpenAI page says categorically "we do not watermark text" - the conclusion rests on their provenance docs enumerating covered modalities and excluding text, plus a goal statement that only makes sense if text is not yet covered. OpenAI has signed the EU Code of Practice on Transparency of AI-Generated Content and has publicly committed to extending provenance signals to text, so this is a status that could change with little notice; re-check before relying on it long-term. Separately, invisible Unicode characters (narrow no-break spaces and similar) are sometimes observed in ChatGPT web output - these are chat-interface rendering artefacts, not a watermark, and do not appear in API output. Note also that if a text watermark ships it would apply at hosted sampling time only. The authenticated OpenAI subscription route inherits this provider finding: authentication and allowance accounting change the route, not the model maker or its text-provenance behaviour. OpenCode's OpenAI OAuth route completed a no-tools generation on 2026-09-11; that establishes the route, not production quality or production readiness.

Sources: 1. developers.openai.com 2. help.openai.com 3. openai.com 4. www.infoq.com

opencode-go

high confidence · checked 4 September 2026
clean open weights
What we found

Not a model maker but a flat-rate route (the opencode Go plan) to two open-weight models already in this table, Moonshot Kimi K3 and Z.ai GLM 5.2. The route passes the model output through unchanged, so the state is the state of those providers, both clean. The plan is a monthly allowance, not a per-token bill, which is why these ids exist beside the metered OpenRouter ids for the same models.

Sources: 1. opencode.ai

qwen

medium confidence · checked 28 August 2026
clean open weights
What we found

No Alibaba/Qwen documentation, model card, technical report or announcement describes a statistical text watermark. Stanford's Foundation Model Transparency Index company report for Alibaba/Qwen3 (Dec 2025) records no disclosure at all on watermarking or provenance for text outputs. Positive counter-evidence: Alibaba's own way of spotting AI text is a CLASSIFIER, not a key-based watermark detector - the AI Guardrails / content-moderation docs describe "AI生成文本鉴别" via the `text_aigc_detector` service, which flags text that "疑似由AI生成合成" (appears to be AI-generated); a provider that watermarked its own text would verify the key instead of guessing. Alibaba does watermark other modalities, which sharpens the contrast: the Qwen-Image API exposes a `watermark` parameter, and the SASE `CreateWmEmbedTask` API embeds and extracts digital watermarks for images, audio, video and document files - none of it token-level watermarking of chat/completion text. The Chinese labelling regime Qwen complies with (Measures + mandatory standard GB 45438-2025, in force 1 Sept 2025) requires a visible label plus an implicit label in FILE METADATA; a digital watermark inside the content itself is explicitly encouraged but not mandated, so the widely-reported "Doubao, DeepSeek, Qwen and Ernie labelled 150 billion pieces of content" figure is metadata and visible labels, which the question excludes. Open weights: Qwen3 models are Apache 2.0 and freely downloadable on Hugging Face (e.g. Qwen/Qwen3-235B-A22B, ~406k downloads in the last month), so any self-hosted Qwen carries no provider-side watermark regardless of what the hosted API does.

What we could not establish

No published third-party statistical probe of Qwen's hosted API for a watermark. The ETH SRI black-box watermark-detection work (arXiv 2405.20777, ICLR 2025) tested GPT-4, Claude 3 and Gemini 1.0 Pro - finding no strong evidence of a watermark - but did not cover Qwen, and I found no later study that did. So "clean" rests on (a) absence across Alibaba's own docs, research output and transparency reporting, (b) their AI-text detection being classifier-based, and (c) the Chinese standard not requiring in-content text watermarks - not on a direct measurement of the API. Two things not established either way: whether Alibaba adds any hidden marker to Tongyi/Qwen Chat consumer output (a metadata/ID scheme would satisfy the law and is out of scope here anyway), and whether some future Qwen serving path adds one. Also note the flagship top-end model (Qwen3.x-Max) is served API-only under a bespoke licence rather than Apache 2.0, so "open weights" holds for the downloadable Qwen3 line, not literally every Qwen model.

Sources: 1. crfm.stanford.edu 2. help.aliyun.com 3. help.aliyun.com 4. www.alibabacloud.com

z-ai

medium confidence · checked 28 August 2026
clean open weights
What we found

No Z.ai surface describes a text watermark. The GLM-5 technical report (arXiv 2602.15763), the zai-org/GLM-5 README, the GLM-5.1 API guide, the GLM-5.2 and GLM-5.3-Flash model cards, and the BigModel platform docs contain no watermarking scheme, no announcement, and no detection tool; the only near-match is Z.ai's Terms of Use, which says "you may not remove, modify, or obscure any AI identifiers added to Outputs by Z.ai, regardless of the form in which such identifiers are presented" - a legal clause naming no technical method, consistent with China's labelling regime rather than a sampling-time watermark. That regime pushes away from text watermarking: GB 45438-2025 and the CAC Labelling Measures (in force 2025-09-01) require for TEXT an explicit visible label ("AI"/"生成") plus an implicit identifier in FILE METADATA, while content-embedded digital watermarks are encouraged for audio/image/video and are not mandated for text. Third-party surveys of who watermarks text as of August 2026 name only Anthropic (SynthID-Text-derived, every Claude model from 2026-08-02) and Google (SynthID); Z.ai/Zhipu/GLM appear in none of them, and Z.ai is not among the EU Transparency Code of Practice signatories. Structurally it could not bind anyway: GLM-5.1, GLM-5.2 and GLM-5.3-Flash (320B-A18B, released 2026-08-26) are MIT-licensed open weights on Hugging Face under zai-org, so self-hosted inference carries no provider-side sampling watermark; only the flagship GLM-5.3 weights were still held back pending safety evaluation as of late August 2026.

What we could not establish

Z.ai has never made an explicit public statement either way - there is no "we do not watermark" denial to point at, so this is an absence-of-evidence finding across their docs, papers, repo and model cards rather than a positive disclaimer. The Terms of Use "AI identifiers ... regardless of the form in which such identifiers are presented" clause is deliberately broad and would legally cover a hidden signal if they ever added one; it is the single piece of text that cannot be fully ruled out, though it names no method and matches the Chinese visible-label-plus-metadata scheme. No independent statistical test of GLM API output for a green-list/Gumbel-style watermark has been published, so the hosted API path is unverified empirically. SaferAI's external GLM-5.2 evaluation notes Z.ai published no safety framework, pre-deployment testing commitments or risk assessment at all, so silence on watermarking sits inside a broader silence on published safety posture - it is weaker evidence than silence from a lab that documents provenance work in detail. Behaviour could also differ between the international z.ai endpoint and the mainland open.bigmodel.cn endpoint, which faces the CAC rules directly; nothing found distinguishes the two. Open weights make the question moot for any locally run GLM regardless of what the API does.

Sources: 1. docs.z.ai 2. docs.z.ai 3. docs.bigmodel.cn 4. arxiv.org

Refused outright

google/*

Refused everywhere, not only for writing. The watermarking table records the evidence and the date it was checked.

The permissions here are generated from the file the pipeline reads, last revised 14 September 2026; the page counts are taken from the pages themselves. If something looks wrong, tell us.

Reading a source without paying for it twice

Reading one source takes many passes: the model is sent the same text over and over, each time told what it has already found and asked for more. Every provider will hold that text between passes at a fraction of the price, but only if the start of each request is identical, so the source comes first and everything that changes comes after it. This is what each provider offers, and when we last checked.

Claude, billed per token

marker required

claude-opus-5, claude-sonnet-5, claude-haiku-4-5

Nothing caches without a marker. There is no implicit fallback.

Held for
1h
Re-reading costs
10% of normal
Storing for 1h costs
200% of normal
Storing for 5m costs
125% of normal

Storing for an hour costs 2x against 1.25x for five minutes, so the hour pays for itself the first time a gap would let the short one lapse. Book extractions run for hours.

Claude, on the subscription plan

automatic

claude-opus-5, claude-sonnet-5, claude-haiku-4-5

The CLI places its own breakpoints and we cannot choose where. Measured on a subscription it uses the 1-hour cache (ephemeral_1h_input_tokens non-zero, 5m zero). Ordering is the only lever.

Held for
1h

GPT-5.6 - Luna, Sol and Terra

automatic, marker optional

openai/gpt-5.6-luna, openai/gpt-5.6-sol, openai/gpt-5.6-terra

Implicit by default on a stable prefix, no markers needed. GPT-5.6 also accepts prompt_cache_options.mode = explicit with prompt_cache_breakpoint on a content block, which we do not use yet.

Held for
30m
Storing for 30m costs
125% of normal

OpenAI discounts cached input by roughly 25-50% depending on model, NOT the 90% Anthropic and DeepSeek give. Confirm per model against a bill before relying on a figure.

GPT-5.6 through an authenticated OpenAI subscription

not checked

openai-subscription/gpt-5.6-luna, openai-subscription/gpt-5.6-sol, openai-subscription/gpt-5.6-terra

Not checked. Treat as uncached until usage from repeated stable-prefix calls establishes otherwise.

DeepSeek V4

automatic

deepseek/deepseek-v4-flash

DeepSeek watches for a repeated prefix and serves it from cache with no configuration.

Re-reading costs
2% of normal

Third-party trackers put cache-hit input at $0.0028/M against $0.14/M cache-miss for V4-Flash, about 98% off. Not confirmed against our own bill.

Kimi, GLM and Qwen

not checked

moonshotai/kimi-k3, z-ai/glm-5.3-flash, qwen/qwen3.8-flash

Not checked. Treat as uncached until measured - OpenRouter pins the provider between calls so a hit would show up in the ledger if there is one.

Kimi and GLM through the opencode plan

not checked

opencode-go/kimi-k3, opencode-go/glm-5.2

Routed through a local agent to providers whose caching we do not control or observe.

Last checked 11 September 2026. A provider marked “not checked” may well hold text between passes; nobody here has confirmed it, and we would rather say so than assume either way.

Language

30 languages covering 80% of the world's literate population

English English English (US) English (US) Spanish Español Portuguese Português Indonesian Bahasa Indonesia French Français Swahili Kiswahili Vietnamese Tiếng Việt Turkish Türkçe German Deutsch Italian Italiano Uzbek Oʻzbekcha Polish Polski Tagalog Tagalog
Mandarin 中文 Traditional Chinese 繁體中文 Japanese 日本語 Korean 한국어
Arabic العربية Urdu اردو Persian فارسی
Russian Русский Ukrainian Українська
Hindi हिन्दी Bengali বাংলা Thai ไทย Burmese မြန်မာ Telugu తెలుగు Marathi मराठी Tamil தமிழ்