What Nota is, how it measures, and what a result means.
No invented metrics on this page. The only numbers Nota publishes are the ones the API returns with a result — measured for the checkpoint that scored your text.
The model — and what a reading means
Nota is a small language model with a classification head, trained on pre-2022-style human writing and AI mirrors of the same kinds of documents, and examined on frozen test sets. It reports how machine-like a passage looks under that model, with the false-positive rate we measured on held-out humans. It cannot name the generator, and it will not score anything under 250 words.
Nota is a document-level binary classifier: a 4-billion-parameter language model with a LoRA adapter and a linear classification head. It reads a document and returns a score of machine-likeness under this model and the band that score falls in — low, moderate, elevated or high — along with whether it abstained and the false-positive rate measured for the shipped checkpoint.
The verdict follows the band, and only high (score ≥ 0.5) reads “consistent with machine generation”. Moderate and elevated are the overlap between careful human prose and machine text, and both come back Inconclusive — an elevated reading is something to look at, never a verdict.
It does not name ChatGPT, Claude or Gemini. It does not highlight sentences. It does not split “human / mixed / AI” spans.
How the false-positive rate is measured
Human writing in the pool is pre-2022-style, so the human label is trustworthy. Machine text is generated to mirror those documents — the same kinds of writing, from many model families — and the test sets are frozen. The published false-positive rate is always on held-out humans of 250 words or more, and it is the number the API returns beside every score. We do not print a separate marketing figure that could drift from the model actually running.
At the verdict line, the frozen exam of 22 August 2026 measured 0.93% of 104,692 held-out human documents (250+ words) called machine-written, catching 98.8% of machine text. That pooled figure hides a spread by length, so the result card shows the measured rate for the scored document’s own length band: 1.81% at 250–400 words (n=11,658), 0.94% at 400–800 words (n=80,098), and 0.06% at 800+ words (n=12,936, longer documents folded in). None of this implies longer essays are safer — it is simply the measured rate for each band.
Abstention
Below 250 words there is not enough text to score responsibly, so Nota does not pretend to. It returns abstained: true and a reason instead of a score, and the interface shows the reason only. A tweet, a headline, a short paragraph under the floor: no number.
Published limits
Every detector has edges. We print ours, with the numbers attached — and we would rather you read them here than discover them in a meeting.
- Paraphrase and “humanisers”. Text run through a rewriting tool drifts toward the machine ranges. On such text the score describes the tool, not the writer.
- Short text. Abstained, by design.
- Unseen generators. A model that writes unlike anything in the training mirrors can score low.
- Formal, careful human prose. Can sit in the overlap and come back Inconclusive — moderate or elevated. This is why the false-positive rate is printed on the result, and why a result is a reason to look closer, not an answer.
The evidence panel
Optionally, a set of writing statistics — sentence-length variation, vocabulary richness, phrase variety and others — each placed against an envelope of human writing, each with the situation in which it misleads. Signals whose measured separation is below 20% render as context only. There is no aggregate evidence score, and the panel is never averaged into the Nota number. “Unremarkable” is the usual answer.
What every result shows
- Scores are probabilistic and specific to this model version.
- Below 250 words Nota abstains rather than guess.
- The false-positive rate shown is the one measured for the shipped checkpoint on held-out human writing of 250 words or more — it comes from the API, never from a marketing page.
- A low score is not proof a person wrote this. It means this model did not flag it.
- Nota does not name a generator, does not highlight sentences, and scores documents whole.
Claims we do not make
- “This was written by AI.” We say “consistent with machine generation under this model.”
- “Detects ChatGPT / Claude / Gemini.”
- “EU AI Act compliant / certified / approved.”
- Any competitor’s accuracy or false-positive number as ours.
- A giant “X% AI” traffic light as the only output.
- “Fight AI cheating.” Nota is not sold as an accusation tool, and there is no published ESL false-positive rate to justify classroom use.