Skip to content
Assay Layer
Menu

Accuracy

How accurate is Assay Layer, and how do we know

Accuracy is the first question anyone asks about a detector and the one most vendors answer with a single marketing percentage. Here is the whole method instead: what each layer does, where every threshold sits, what the independent literature says, and what we do about the places where detection is known to fail.

Updated

Methodology

What actually runs, in what order

An assay is one pipeline over one piece of text. The text is normalised — trimmed, with runs of whitespace collapsed — then hashed with SHA-256. Word count and a language guess come from the normalised text. Every layer then runs in parallel with a twelve-second budget and a single retry on provider error.

The layers do not vote. They are ranked, because they are not the same kind of evidence:

  1. Watermark and metadata come first. If a first-party watermark detector or a valid C2PA manifest returns found, the report’s signal is provenance regardless of what the classifiers say. Origin established beats origin inferred, every time.
  2. Classifiers run next, and only if there are at least fifty words. Below that they return too short, and the report’s signal is insufficient with no number attached.
  3. The summary is computed last, from the classifier scores, using the thresholds in the table beside this text.

A deep assay adds a second classifier from a different vendor and an agreement score, where 1.00 means the two models landed in the same place. Disagreement is reported, not averaged away.

Signal thresholds, as implemented in the pipeline.
Condition Signal
Any watermark or metadata layer returns found Provenance
Fewer than 50 words, or no classifier could run Insufficient
AI signal 0.80 and above AI
AI signal 0.20 and below None
Anything between 0.20 and 0.80 Mixed
Confidence, reported next to every signal.
Condition Confidence
Two classifiers with an agreement score of 0.80 or above High
One classifier with a score of 0.90 or above High
A classifier ran and none of the above applies Medium
Under 150 words, or a layer returned an error Low

Coverage

Coverage by generator

Which layer can actually answer for output from which system. The watermark column is the honest one: today it is almost entirely empty, and that is a fact about the industry rather than about us.

Coverage by generator, as of the date at the top of this page.
Generator Watermark check Classifier coverage
Anthropic (Claude) Pending — detector in private preview Yes, model-agnostic
Google (Gemini) No public detector for SynthID-Text Yes, model-agnostic
ChatGPT OpenAI (ChatGPT, GPT models) No public detector Yes, model-agnostic
Meta (Llama) No public detector Yes, model-agnostic
Mistral AI No public detector Yes, model-agnostic
DeepSeek No public detector Yes, model-agnostic
xAI Grok xAI (Grok) No public detector Yes, model-agnostic
Microsoft Copilot Microsoft Copilot No public detector Yes, model-agnostic
Perplexity No public detector Yes, model-agnostic
Cohere Cohere No public detector Yes, model-agnostic
Open weights via Hugging Face No public detector Yes, model-agnostic
Local models via Ollama No public detector Yes, model-agnostic

“Model-agnostic” in the third column means what it says: the classifier layer looks at the text, not at a list of systems, so it responds to output from models that did not exist when it was trained — and it is not tuned per generator, so none of these rows is more accurate than any other. Treat the column as “this layer runs”, not as “this generator is detected”.

The watermark column is the one worth watching. A first-party detector is the only thing that establishes origin rather than inferring it, and it counts here only when the provider exposes one we can call. A published paper is not a detector. When Anthropic’s preview opens, or a public SynthID-Text detector appears, this table changes and the pipeline version on every report changes with it.

Independent evidence

What the published research says

These are other people's results, not ours. They are the reason this product is built the way it is.

OpenAI withdrew its own detector, July 2023

OpenAI released a classifier for AI-written text on 31 January 2023 and reported that it correctly identified 26 per cent of AI-written English text while incorrectly labelling 9 per cent of human-written text as AI. It was withdrawn on 20 July 2023 for low accuracy. The company that trained the models could not make single-classifier detection work well enough to keep shipping it.

Stanford, 2023: bias against non-native writers

Liang and colleagues ran seven GPT detectors over TOEFL essays written by non-native English speakers and over essays by native US eighth-graders. The detectors flagged 61 per cent of the non-native essays as AI-generated, against under 5 per cent of the native ones — a bias that falls on exactly the people least able to contest it. Read the paper.

Chicago Booth / NBER, August 2025

NBER working paper 34223 evaluated commercial detectors and found the better ones held false positive rates under 1 per cent on long text — and found all of them degrading under fifty words. Both halves of that result matter: modern detectors are much better than the 2023 generation, and they still fall apart on short passages. Read the paper.

Read together, those three results say something fairly precise. Detection is not hopeless — it improved considerably between 2023 and 2025. But its errors are not evenly spread: they concentrate on short text and on non-native writing, which is to say on the cases where a false accusation does the most damage. A product built on one number, with no length floor and no standing caveats, will hurt somebody.

Policy

What we do about it

No verdict language, anywhere

No report, export, API response or MCP tool result ever says “this is AI” or “this is human”. Headlines take forms like “Strong AI signals across most sentences” or “Signals are weak or disagree”. The grammar is deliberate: we describe the evidence, you draw the conclusion.

A hard floor on short text

Under fifty words the classifier layer refuses to return a number. Not a low-confidence number — no number. This costs us conversions from people who want to check a tweet, and we are keeping it.

Caveats are part of the report

Under 150 words, a short-text caveat is attached. On English text, the non-native writing caveat is attached to every none and ai result. On mixed, the hybrid-editing caveat is attached. They travel with the JSON and the PDF, so they cannot be cropped out of a screenshot without it being obvious.

A human review path, always

A report is a prompt to look closer: ask for the draft history, compare against earlier work, talk to the writer. We say this on the page, in the caveats and in the PDF, because the alternative — an automated decision about a person based on a style score — is both unfair and, in several jurisdictions, legally fraught.

Changelog

Detector changelog

Every change to providers, thresholds or prompts bumps the pipeline version. The version is printed on every report, so an old export can always be read against the rules that produced it.

Version Date Change
2026.09.1 2026-09-17 Initial pipeline. Watermark layer with the Anthropic text watermark adapter and a SynthID-Text adapter awaiting a public detector. Metadata layer reading C2PA and XMP on uploads. Base classifier on every assay, deep classifier as well on deep assays, with an agreement score. Fifty-word floor, signal and confidence thresholds as documented above. Fixed: low-scoring classifier results were dropped from the summary.

Questions

Accuracy, answered plainly

Why do you not publish a single accuracy percentage?

Because a single percentage is meaningless without the test set, the text length and the base rate it was measured against. A detector that is 99 per cent accurate on thousand-word essays can be near useless on sixty-word support replies, and quoting the first number while selling the second is the standard trick in this market.

What is the false positive rate?

Any single figure would mislead you, because false positives depend far more on the genre and the length of the text than on the detector alone. Conversational and business prose is where classifiers are strongest. The weak spots are the same for every detector on the market, including the classifiers we build on: literary and narrative prose, very short text, heavily edited text, and English written by people who learned it later in life. That is why a report shows you every layer and its confidence instead of one verdict, and why a narrative text deserves a deep assay or a human read rather than a number. For published ranges, the independent studies on this page are the place to look.

Does a high score mean the text was definitely AI-generated?

No. A high classifier score means the text carries statistical patterns common in model output. Human writing that is formulaic, heavily edited, translated, or written to a template can produce the same patterns. Only the watermark and metadata layers can establish origin, and they either find something or they do not.