How we test the detector
Most tools publish an accuracy number and stop there. Below is the whole test behind ours — the texts, the threshold and what the result does not prove. Last run: August 2026.
What we ran
- 30 human-written excerpts — passages from public-domain books (Austen, Twain, Doyle, Dickens, Melville, Shelley, Thoreau, Stoker, Wilde, the Federalist Papers), taken from different points in each book.
- 22 AI-written texts — freshly generated for this test across five genres: essay, blog post, marketing copy, academic abstract and a casual forum post.
- A verdict counts as “AI” at 50% or above — the same threshold the product uses.
What came back
| Human texts wrongly flagged as AI | 0 of 30 |
| AI texts correctly caught | 22 of 22 |
| Highest score any human text received | 18.5% |
| Lowest score any AI text received | 86.7% |
The gap between those last two numbers is the point: on this corpus human and machine writing did not overlap, so the verdict did not depend on where exactly the threshold sits.
What this does not prove
Our own test on public-domain literary English against one AI model — not an independent audit.
- Literary classics are not student essays. A well-edited assignment sits closer to the middle of the range than Austen does.
- The AI half came from one model family. Other models — and text that has been rewritten to avoid detection — behave differently.
- No detector can be certain. A score is a probability, and it can be wrong in either direction, which is why we show it per sentence instead of as a single verdict.
We re-run this test as the detector changes and update the numbers here.