ParaTrace v7 — real API results Strict setting
seminar-essay.txtAI-written

In today's rapidly evolving digital landscape, remote work has fundamentally reshaped how organizations operate, offering employees unprecedented flexibility. However, it is important to note that this shift also presents unique challenges, from maintaining team cohesion to ensuring clear communication. Ultimately, by fostering a culture of trust, organizations can unlock the full potential of a hybrid workforce.

diary-1892.txtHuman · 1892

In the evening we went round to the Cummings’, to have a few fireworks. It began to rain, and I thought it rather dull. One of my squibs would not go off, and Gowing said: “Hit it on your boot, boy; it will go off then.” I gave it a few knocks on the end of my boot, and it went off with one loud explosion, and burnt my fingers rather badly.

study-tips-edited.txtAI-written + attack

Effective time managment is esential for аcademic success. By prioritizing tasks, setting realistc gоals, and minimizing distractions, students can significantly improve thier productivity. Moreover, regular breaks and adequate slеep play a crucial role in maintaining fоcus and overall well-being. Ultimatly, developing these habits early not only enhances academic performance but also lays a strong foundation for lifelong lеarning and personal growth.

AI likelihood >99%
Likely AI-generated High confidence 100% of the text
Likely human-written High confidence 100% of the text
Likely AI-generated High confidence 100% of the text, despite 5 misspellings and 5 look-alike letters
Reading the text…
99.6%
AI text caught
0.84%
human texts flagged
96.0%
caught under attack
90
AI models tested, from 26 developers
20ms
to score a 300-word text

Replays of real results from the production API (detector v7, Strict, 6 October 2026). The AI samples were written by an AI model for this demo; the human sample is from The Diary of a Nobody (1892).

Caught across 90 models

26,082 answers to real user prompts, collected from public chat logs up to October 2025. No developer falls below 98.8%, and 98% of AI answers score a probability of 99% or higher, so the verdicts are clear-cut rather than borderline.

Detection by developer
DeveloperDetectedAnswers
OpenAI99.4%4,595
Google99.4%3,782
Anthropic99.6%3,485
Meta99.8%2,138
Mistral99.7%1,958
Alibaba99.8%1,834
DeepSeek99.7%1,393
Reka99.9%1,200
Microsoft99.8%1,072
01.AI99.8%900
Detection by developer, continued
DeveloperDetectedAnswers
Cohere99.4%822
xAI99.8%579
Zhipu100%373
Amazon100%312
NVIDIA99.7%308
Nexusflow100%300
MiniMax100%183
Databricks98.8%168
Moonshot100%132
Other developers99.2%494

† fewer than 200 answers, so the rate is indicative. Tencent (54 answers, 100%) is left out of the table for space. Late-2025 flagships: Claude Sonnet 4.5 98.7% of 318, Gemini 2.5 Pro 99.3% of 667, GPT-5 99.2% of 357, Grok 4 99.0% of 102.

Models it never saw
98.6% of 1,172
Late-2025 model versions with no version in v7's training data
Developers it never saw
99.8% of 3,948
01.AI, Databricks, Microsoft, NVIDIA, Nexusflow and Reka had no model in training at all
Models held back on purpose
99.3% of 8,118
GPT-4 and Mistral-chat were kept out of training to test exactly this

Few false alarms on human writing

A detector is only as useful as its false-alarm rate. Across 22,691 held-out human texts from 43 sources, 0.84% were wrongly flagged. Longer texts are safer: above 400 words the rate drops to 0.2%, and no scientific paper, legal text or patent was flagged.

By lengthHuman texts wrongly flagged as AI
WordsFlaggedTexts95% CI
Under 502.8%24 of 8461.9–4.2%
50–991.6%30 of 1,9111.1–2.2%
100–1991.1%75 of 6,8250.9–1.4%
200–3990.6%54 of 9,0210.5–0.8%
400+0.2%7 of 4,0880.1–0.4%
By kind of writingFewest false alarms first
WritingFlaggedTextsWritingFlaggedTexts
Scientific papers0.0%432Encyclopedia1.3%1,718
Legal texts0.0%148Sci. abstracts1.4%1,816
Patents0.0%115News articles1.4%2,733
Essays0.2%5,562Other web1.5%2,037
Q&A0.4%3,369Blogs1.6%189
Reviews0.8%1,602Recipes2.0%293
Forum posts0.9%1,869Educational2.6%116
Book excerpts1.1%443Emails2.7%75

95% Wilson intervals. † fewer than 200 texts. Government reports (61), medical notes (17) and résumés (15) also had no false alarms; speeches and debates 1.2% of 81.

Hard to evade

We took 1,500 AI texts the detector never saw and applied 11 common tricks to each of them. Before any attack, 97.3% were caught; across all attacks, 96.0%. For chat-assistant text under attack the rate is 98.0%. AI paraphrasing is the strongest attack and still leaves 91.5% caught.

AttackDetected95% CI
No attack (baseline)97.3%1,500 texts
AI paraphrasing91.5%89.9–92.8%
Deleted articles (a, an, the)93.6%92.2–94.7%
Random upper/lower case95.2%94.0–96.2%
Synonym swaps95.9%94.8–96.8%
Deliberate misspellings96.1%95.0–96.9%
British/American spellings97.1%96.1–97.8%
Extra paragraph breaks97.2%96.2–97.9%
Look-alike letters97.3%96.4–98.0%
Invisible characters97.3%96.4–98.0%
Extra spaces97.3%96.4–98.0%
Changed numbers97.5%96.6–98.2%

RAID documents held out of training, 1,500 texts per attack, 95% Wilson intervals. Detection also holds across formats: 99.6% for answers containing code, 98.8% with math notation, 99.9% with tables, and 98.9% for answers as short as 50–99 words.

Scientific papers: the NeurIPS check

Before screening its 2026 position papers, NeurIPS tested Pangram on papers accepted at an AI ethics conference, some from 2022, before ChatGPT, and some from 2025. We ran the same check with ParaTrace v7. On papers from 2022 it flags none, exactly like Pangram. On papers from 2025 it flags 7.3% as at least half AI-written, against 1.0% for Pangram.

No false alarms
0 of 106 papers from 2022
flagged at any level, the same as Pangram. Even passage by passage, only 18 of 4,512 (0.4%) lean AI.
More AI flagged in 2025
7.3% vs 1.0% for Pangram
of papers from 2025 (9 of 123) reach an AI score of 50% or more
Close at the top
1.6% vs 1.0% for Pangram
of papers from 2025 reach 90% or more. That is 2 papers in each sample.
Papers from 2022Written before ChatGPT, so any paper flagged is a false alarm
DetectorPapers≥ 50%≥ 90%100%
Pangram 3.3.21590.0%0.0%0.0%
ParaTrace v71060.0%0.0%0.0%
Papers from 2025AI writing tools in wide use; no record of which papers used them
DetectorPapers≥ 50%≥ 90%100%
Pangram 3.3.22041.0%1.0%0.0%
ParaTrace v71237.3%1.6%0.8%

The AI score is the share of a paper's passages scored as AI; 100% means AI use across many parts of a paper, not that every word came from AI. Pangram's rows are from the NeurIPS post AI-generated papers in the NeurIPS 2026 position paper track (2 June 2026); the samples differ in size, so compare shares, not counts. Which 2025 papers used AI is not recorded: because ParaTrace flags no paper from before ChatGPT, false alarms are an unlikely cause of its extra flags, but that is an inference, not proof. NeurIPS and Pangram Labs, Inc. have not reviewed this comparison. Full write-up on paratrace.net →

Where it falls short

We publish the weak spots alongside the strengths. Head to head with Pangram on 21 public benchmark numbers, v7 is within 5 points on 15 of them; it trails clearly on text run through a commercial humanizer and on standard texts under 50 words. See the comparison →

Adult persuasive essays
18.4% 32 of 174 flagged
A kind of human writing kept out of training entirely, and the highest false-alarm rate we measured
Humanizer tools
21.4% TPR at 1% FPR
On StealthGPT-rewritten essays (UChicago benchmark), against 98.9% for Pangram 4
Very short texts
2.8% false alarms
Human texts under 50 words are flagged most often; treat short-text verdicts with care
Human text through a synonym swapper
16.4% of 1,185 flagged
Automatic rewording makes human writing look machine-made
AI-paraphrased human writing
42.8% of 1,185 flagged
Flagged by design: the words on the page were produced by an AI
A genre it never studied
90.6% of 5,000 AI poems
97.3% for poems from chat assistants, 83.2% from base models. Human poetry: 2.2% false alarms

How it works, and how we measured

Scoring. The text is cleaned (quotes, dashes and Markdown normalised, invisible characters removed), split into overlapping windows of 512 tokens, and each window is scored by our transformer classifier. Each sentence takes the label of the windows it lies in.
Verdict. The result is likely AI-generated, likely partly AI-generated, or likely human-written, depending on how much of the text is labelled AI. Long documents such as papers get a verdict per passage as well.
One setting throughout. Every figure uses the production service's Strict verdict: a single window is called AI at probability 0.4036, the 1% false-alarm point on human validation texts. A Lenient setting catches more AI text at about 1 in 20 human texts flagged.
AI-written answers. 26,082 chat answers from public chat logs, 2024 to October 2025, plus a small in-house set. Their prompts were excluded from training. Models released after October 2025 are not benchmarked yet.
Human writing. 22,691 held-out documents from v7's test split, 43 sources. Every essay corpus and every source with at most 1,123 texts is included in full; larger sources are sampled to 1,123.
Held out. No evaluation text is in v7's training split. Poetry, adult persuasive essays, GPT-4 and Mistral-chat are kept out of training entirely to measure generalisation honestly.
Production numbers. Every text went through the live API (detector v7) in October 2026, not an offline notebook. Ranges are 95% Wilson intervals; groups under 200 texts are marked †.
More detail. The benchmark report, the scientific-papers field test and the Pangram comparison are published on paratrace.net, along with a developer API.

Need AI-text checks inside your own workflow?

Try ParaTrace on paratrace.net, call it from your software through the API, or talk to us about a private deployment for your submissions, applications or reviews.

Try ParaTrace Book a technical call