Accuracy

How often is Nota wrong?

Every AI detector makes mistakes. Here are ours, with the numbers: how often Nota calls human writing AI, how that changes with length, and what we have not tested yet.

Updated 27 September 2026 · Nota v0.1

1 in 78human texts of 250+ words wrongly read as AI in our main test
99.0%of AI texts of 250+ words caught in the same test
about 1 in 3human texts under 50 words read as AI in an informal test

How often does Nota call human writing AI?

We test each version of Nota on writing it never saw while it learned: 104,692 human documents and a set of AI documents, all 250 words or more. At the line where Nota says a text is consistent with machine generation (the high band), it wrongly flagged 1.28% of the human documents, about 1 in 78. It caught 99.0% of the AI documents.

The rate changes with length. Each result shows the rate for its own length group:

LengthHuman texts testedWrongly read as AI
250–400 words11,6582.28% (about 1 in 44)
400–800 words80,0981.34% (about 1 in 75)
800+ words12,936none

A zero means none in this test. It is not a promise that Nota will never be wrong on long text.

What about short text?

Our main test starts at 250 words, but Nota checks text from 10 words. So we ran an extra, informal test to see what happens below that. We took 200 movie reviews from a public research set (IMDB). People wrote them before AI writing tools existed. We cut each review to a set length and checked it with Nota v0.1 on 27 Sep 2026.

LengthHuman textsRead as AIRead as human-like
15 words20069 (35%)none
35 words19967 (34%)none
50 words19634 (17%)none
75 words1915 (3%)6 (3%)
100 words1835 (3%)31 (17%)
150 words133none50 (38%)
200 words86none52 (60%)
Whole review (250+)60none59 (98%)

What it shows: under 50 words, Nota could not recognise human writing. None of these texts read as human-like, and about 1 in 3 read as AI. At 50 words it was about 1 in 6, and at 75 to 100 words about 1 in 33. From 150 words, none read as AI. The same reviews at full length read as human-like 98% of the time, so the problem is length, not the kind of writing.

That is why every result under 50 words carries a warning, and why we ask for more writing when you can give it.

How far to trust this test. It is one kind of writing (casual reviews) in one language, cut mid-sentence, with 60 to 200 texts per length. We checked no AI texts in it, so it says nothing about how often Nota catches AI. Read it as the rough size of the problem, not a measured rate like the main test.

What do these numbers mean for my result?

  • A rate describes a test set. It is not the chance that your own result is wrong.
  • Longer is better. When you can, check 250 words or more of the same writing.
  • A result that says AI is a reason to look closer, never an accusation. Nota reads the writing, not the writer.

What have we not tested yet?

  • Writing by people whose first language is not English.
  • Languages other than English.
  • Human writing that was polished with grammar or rewriting tools.
  • Short text in a full test with both human and AI writing.
  • The end of long documents: Nota reads up to the first 1,024 tokens (small pieces of words).

Which versions of Nota have been tested?

VersionHuman texts read as AI (250+ words)AI texts caughtStatus
Nota v0.1
tested 29 Aug 2026
1.28%99.0%Live
Nota v0
tested 22 Aug 2026
0.93%98.8%Retired 29 Aug 2026

Each new version gets a new test before it goes live, and its row is added here. How the test works.

Questions people ask

Can AI detectors be wrong?

Yes. Every AI detector sometimes reads human writing as AI and AI writing as human. In our main test, Nota wrongly flagged about 1 in 78 human texts of 250 words or more. On short text it is wrong far more often.

Why do short texts get flagged as AI?

A few sentences give a detector very little to go on. In our informal test, Nota read about 1 in 3 human texts under 50 words as AI, and none of them as human-like. That is why every result under 50 words carries a warning. Check a longer piece of the same writing when you can.

Is 1.28% the chance that my result is wrong?

No. It is how often Nota wrongly flagged human writing in a test set. Your own writing may differ from that set in length, topic or style, so the chance for a single result can be higher or lower.

What should I do if a result says my writing looks like AI?

Look at the length first: short texts are often misread. Then look at other evidence, such as drafts, notes and version history. A detector result is a reason to look closer, never proof of who wrote something. Make a writing history report.

See how your writing reads.

Paste at least 250 words for the most reliable result.

Check for AI →