Method

What Nota is, how it measures, and what a result means.

No invented metrics on this page. The only numbers Nota publishes are the ones the API returns with a result — measured for the checkpoint that scored your text.

Updated 22 August 2026

The model — and what a reading means

Nota is a small language model with a classification head, trained on pre-2022-style human writing and AI mirrors of the same kinds of documents, and examined on frozen test sets. It reports how machine-like a passage looks under that model, with the false-positive rate we measured on held-out humans. It cannot name the generator, and it will not score anything under 250 words.

Nota is a document-level binary classifier: a 4-billion-parameter language model with a LoRA adapter and a linear classification head. It reads a document and returns a score of machine-likeness under this model and the band that score falls in — low, moderate, elevated or high — along with whether it abstained and the false-positive rate measured for the shipped checkpoint.

The verdict follows the band, and only high (score ≥ 0.5) reads “consistent with machine generation”. Moderate and elevated are the overlap between careful human prose and machine text, and both come back Inconclusive — an elevated reading is something to look at, never a verdict.

It does not name ChatGPT, Claude or Gemini. It does not highlight sentences. It does not split “human / mixed / AI” spans.

How the false-positive rate is measured

Human writing in the pool is pre-2022-style, so the human label is trustworthy. Machine text is generated to mirror those documents — the same kinds of writing, from many model families — and the test sets are frozen. The published false-positive rate is always on held-out humans of 250 words or more, and it is the number the API returns beside every score. We do not print a separate marketing figure that could drift from the model actually running.

At the verdict line, the frozen exam of 22 August 2026 measured 0.93% of 104,692 held-out human documents (250+ words) called machine-written, catching 98.8% of machine text. That pooled figure hides a spread by length, so the result card shows the measured rate for the scored document’s own length band: 1.81% at 250–400 words (n=11,658), 0.94% at 400–800 words (n=80,098), and 0.06% at 800+ words (n=12,936, longer documents folded in). None of this implies longer essays are safer — it is simply the measured rate for each band.

Abstention

Below 250 words there is not enough text to score responsibly, so Nota does not pretend to. It returns abstained: true and a reason instead of a score, and the interface shows the reason only. A tweet, a headline, a short paragraph under the floor: no number.

Published limits

Every detector has edges. We print ours, with the numbers attached — and we would rather you read them here than discover them in a meeting.

  • Paraphrase and “humanisers”. Text run through a rewriting tool drifts toward the machine ranges. On such text the score describes the tool, not the writer.
  • Short text. Abstained, by design.
  • Unseen generators. A model that writes unlike anything in the training mirrors can score low.
  • Formal, careful human prose. Can sit in the overlap and come back Inconclusive — moderate or elevated. This is why the false-positive rate is printed on the result, and why a result is a reason to look closer, not an answer.

The evidence panel

Optionally, a set of writing statistics — sentence-length variation, vocabulary richness, phrase variety and others — each placed against an envelope of human writing, each with the situation in which it misleads. Signals whose measured separation is below 20% render as context only. There is no aggregate evidence score, and the panel is never averaged into the Nota number. “Unremarkable” is the usual answer.

What every result shows

  • Scores are probabilistic and specific to this model version.
  • Below 250 words Nota abstains rather than guess.
  • The false-positive rate shown is the one measured for the shipped checkpoint on held-out human writing of 250 words or more — it comes from the API, never from a marketing page.
  • A low score is not proof a person wrote this. It means this model did not flag it.
  • Nota does not name a generator, does not highlight sentences, and scores documents whole.

Claims we do not make

  • “This was written by AI.” We say “consistent with machine generation under this model.”
  • “Detects ChatGPT / Claude / Gemini.”
  • “EU AI Act compliant / certified / approved.”
  • Any competitor’s accuracy or false-positive number as ours.
  • A giant “X% AI” traffic light as the only output.
  • “Fight AI cheating.” Nota is not sold as an accusation tool, and there is no published ESL false-positive rate to justify classroom use.
Not in v1: image detection, LMS plug-ins, browser extensions, sentence heatmaps, “which model wrote this”, watermarks. If a page here seems to promise one of those, it is a bug — tell us.

Try Nota free →