How it works

An AI check, explained.

Paste your writing. Nota compares its patterns with human and AI writing and gives you an estimate. Here is what the result means and where its limits are.

Updated 27 September 2026

The model — and what a reading means

Nota compares patterns in your text with examples of human and AI writing. Its result is an estimate, and AI detectors can make mistakes. It will not score text under 10 words, and results under 50 words are often wrong. Long texts may be checked only in part. Measured error rates are shown where the text length is covered by the model’s evaluation.

Nota is a document-level binary classifier: a 4-billion-parameter language model with a LoRA adapter and a linear classification head. It returns a score of machine-likeness under this model and the band that score falls in — low, moderate, elevated or high — along with whether it abstained and available test information for the checkpoint used.

The result keeps the model’s four bands: low, moderate, elevated and high AI-likeness. Low is typical of the human writing in its pool; moderate is the overlap between careful human essays and light polish; elevated is more consistent with polished or generated text; high is typical of the AI examples it learned from. Only high produces a machine-consistent verdict in the API. None of these bands proves authorship.

It does not name ChatGPT, Claude or Gemini. It does not highlight sentences. It does not split “human / mixed / AI” spans.

What does the 0–100 number mean?

AI-likeness is a logarithmic display of the classifier’s raw score. It is not a percentage of AI-written words, a probability of authorship, or a linear confidence scale. For example, 40/100 corresponds to a raw score of about 0.000251 and falls in the model’s moderate band.

Model bandRaw scoreAI-likeness before rounding
LowBelow 0.0001Below 33⅓
Moderate0.0001 to below 0.00133⅓ to below 50
Elevated0.001 to below 0.550 to below about 94.983
High0.5 and aboveAbout 94.983 to 100

The interface rounds the displayed number. The band comes from the raw score, so two checks displaying 50 or 95 can fall on opposite sides of a band boundary. Nota keeps the band returned by the model.

How much of the text is checked?

The current Nota model reads the beginning of a document, up to 1,024 model tokens, including its special tokens. Tokens are pieces of words; there is no fixed conversion to a word count. A long upload can be accepted even when only its beginning is scored.

The result shows whether the model reported partial coverage, with token counts when available. If the scoring service does not provide coverage counts, Nota shows the known input limit and says the exact coverage is unknown. It does not silently turn a partial check into a score for every page.

How the false-positive rate is measured

A false-positive rate counts how often the model flags human writing in a test sample at a particular cutoff. The model’s published evaluation covers human writing of 250 words or more. This benchmark does not tell you the probability that your own result is wrong, and it does not establish equal performance for every language, subject or writer.

At the verdict line, the frozen exam of 27 Sep 2026 measured 0.033% of 104,692 held-out human documents (250+ words) called machine-written, catching 99.6% of machine text. That pooled figure hides a spread by length, so the result card shows the measured rate for the scored document’s own length band: 0.043% at 250–400 words (n=11,658), 0.037% at 400–800 words (n=80,098), and none of the 12,936 documents at 800+ words (longer documents folded in). All the numbers are on the accuracy page. A measured zero is the exam’s resolution at that length, not a promise of zero risk. None of this implies longer essays are safer — it is simply the measured rate for each band.

These figures come from a re-run of the exam on 27 September 2026. The first run read the model’s answer from the wrong position for most texts shorter than about 800 words and reported 1.28%. The live checker was never affected. What changed.

Abstention

Below 10 words there is not enough text to score responsibly, so Nota does not pretend to. It returns abstained: true and a reason instead of a score, and the interface shows the reason only. A headline or a short phrase under the floor: no number.

Short text

From 10 to 49 words Nota gives a score, but every result carries a warning: this result is not reliable. A few sentences give the model little to go on, and it leans towards reading them as AI. In an informal test on human-written movie reviews, about 1 in 3 texts cut to 15 or 35 words read high, and none read low. The same reviews at full length read low almost every time. Treat a short-text result as a rough hint, never as grounds for a conclusion, and check a longer piece of the same writing when you can. The full test is on the accuracy page. The API marks these results with shortTextNote.

Between 10 and 250 words Nota scores as usual — same bands, same verdict line — but publishes no false-positive rate. The frozen exam starts at 250 words, so there is no measured rate for shorter text, and we would rather say that than borrow the longer documents’ number. Those results carry measuredFpr: null and a note saying why.

Published limits

Every detector has edges. We print ours, with the numbers attached — and we would rather you read them here than discover them in a meeting.

  • Rewriting and AI assistance. Editing can change a score. The detector cannot reconstruct which parts a person wrote or which tools they used.
  • Short text. Under 10 words, no score. From 10 to 49 words, a score with a warning, because it is often wrong: in an informal test, about 1 in 3 human texts that short read high.
  • Unseen generators. A model that writes unlike anything in the training mirrors can score low.
  • Formal, careful human prose. Can sit in the overlap and come back Inconclusive — moderate or elevated. This is why the false-positive rate is printed on the result, and why a result is a reason to look closer, not an answer.
  • Liturgical, archaic, and highly formulaic prose (scripture, liturgy, some legal text). Can sit in the overlap and read high. The score describes how this model sees that register, not that a person did not write it.

The evidence panel

Optionally, a set of writing statistics — sentence-length variation, vocabulary richness, phrase variety and others — each placed against an envelope of human writing, each with the situation in which it misleads. Signals whose measured separation is below 20% render as context only. There is no aggregate evidence score, and the panel is never averaged into the Nota number. “Unremarkable” is the usual answer.

What every result shows

  • AI detectors can make mistakes. A result is an estimate, not proof of who wrote the text.
  • Use at least 10 words to get an AI result. Under 50 words, results are often wrong: in our test, about 1 in 3 human texts that short was read as AI.
  • Published error rates come from tests on human writing of at least 250 words. They are not the chance that your result is wrong.
  • Long texts may be checked only in part because the model has an input limit.
  • Nota does not identify a specific AI tool or mark individual sentences as AI-written.

Claims we do not make

  • “This was written by AI.” The result describes how the text looks to the detector, not who wrote it.
  • “Detects ChatGPT / Claude / Gemini.”
  • “EU AI Act compliant / certified / approved.”
  • Any competitor’s accuracy or false-positive number as ours.
  • A giant “X% AI” traffic light as the only output.
  • “Fight AI cheating.” A score alone does not establish misconduct. No separate false-positive rate is published for non-native English writers.
  • “Slop.” Nota reports resemblance, overlap and marks. It has no word for a text it dislikes — and no word for the person who wrote it.
Nota does not offer image detection, LMS plug-ins, sentence-level AI heatmaps, generator identification, or proprietary vendor watermark decoding.

Check for AI →