Need AI-text checks inside your own workflow?
Try ParaTrace on paratrace.net, call it from your software through the API, or talk to us about a private deployment for your submissions, applications or reviews.
Try ParaTrace Book a technical callIn today's rapidly evolving digital landscape, remote work has fundamentally reshaped how organizations operate, offering employees unprecedented flexibility. However, it is important to note that this shift also presents unique challenges, from maintaining team cohesion to ensuring clear communication. Ultimately, by fostering a culture of trust, organizations can unlock the full potential of a hybrid workforce.
In the evening we went round to the Cummings’, to have a few fireworks. It began to rain, and I thought it rather dull. One of my squibs would not go off, and Gowing said: “Hit it on your boot, boy; it will go off then.” I gave it a few knocks on the end of my boot, and it went off with one loud explosion, and burnt my fingers rather badly.
Effective time managment is esential for аcademic success. By prioritizing tasks, setting realistc gоals, and minimizing distractions, students can significantly improve thier productivity. Moreover, regular breaks and adequate slеep play a crucial role in maintaining fоcus and overall well-being. Ultimatly, developing these habits early not only enhances academic performance but also lays a strong foundation for lifelong lеarning and personal growth.
Replays of real results from the production API (detector v7, Strict, 6 October 2026). The AI samples were written by an AI model for this demo; the human sample is from The Diary of a Nobody (1892).
26,082 answers to real user prompts, collected from public chat logs up to October 2025. No developer falls below 98.8%, and 98% of AI answers score a probability of 99% or higher, so the verdicts are clear-cut rather than borderline.
| Developer | Detected | Answers |
|---|---|---|
| OpenAI | 99.4% | 4,595 |
| 99.4% | 3,782 | |
| Anthropic | 99.6% | 3,485 |
| Meta | 99.8% | 2,138 |
| Mistral | 99.7% | 1,958 |
| Alibaba | 99.8% | 1,834 |
| DeepSeek | 99.7% | 1,393 |
| Reka | 99.9% | 1,200 |
| Microsoft | 99.8% | 1,072 |
| 01.AI | 99.8% | 900 |
| Developer | Detected | Answers |
|---|---|---|
| Cohere | 99.4% | 822 |
| xAI | 99.8% | 579 |
| Zhipu | 100% | 373 |
| Amazon | 100% | 312 |
| NVIDIA | 99.7% | 308 |
| Nexusflow | 100% | 300 |
| MiniMax | 100% | 183 |
| Databricks | 98.8% | 168 |
| Moonshot | 100% | 132 |
| Other developers | 99.2% | 494 |
† fewer than 200 answers, so the rate is indicative. Tencent (54 answers, 100%) is left out of the table for space. Late-2025 flagships: Claude Sonnet 4.5 98.7% of 318, Gemini 2.5 Pro 99.3% of 667, GPT-5 99.2% of 357, Grok 4 99.0% of 102.
A detector is only as useful as its false-alarm rate. Across 22,691 held-out human texts from 43 sources, 0.84% were wrongly flagged. Longer texts are safer: above 400 words the rate drops to 0.2%, and no scientific paper, legal text or patent was flagged.
| Words | Flagged | Texts | 95% CI |
|---|---|---|---|
| Under 50 | 2.8% | 24 of 846 | 1.9–4.2% |
| 50–99 | 1.6% | 30 of 1,911 | 1.1–2.2% |
| 100–199 | 1.1% | 75 of 6,825 | 0.9–1.4% |
| 200–399 | 0.6% | 54 of 9,021 | 0.5–0.8% |
| 400+ | 0.2% | 7 of 4,088 | 0.1–0.4% |
| Writing | Flagged | Texts | Writing | Flagged | Texts |
|---|---|---|---|---|---|
| Scientific papers | 0.0% | 432 | Encyclopedia | 1.3% | 1,718 |
| Legal texts | 0.0% | 148 | Sci. abstracts | 1.4% | 1,816 |
| Patents | 0.0% | 115 | News articles | 1.4% | 2,733 |
| Essays | 0.2% | 5,562 | Other web | 1.5% | 2,037 |
| Q&A | 0.4% | 3,369 | Blogs | 1.6% | 189 |
| Reviews | 0.8% | 1,602 | Recipes | 2.0% | 293 |
| Forum posts | 0.9% | 1,869 | Educational | 2.6% | 116 |
| Book excerpts | 1.1% | 443 | Emails | 2.7% | 75 |
95% Wilson intervals. † fewer than 200 texts. Government reports (61), medical notes (17) and résumés (15) also had no false alarms; speeches and debates 1.2% of 81.
We took 1,500 AI texts the detector never saw and applied 11 common tricks to each of them. Before any attack, 97.3% were caught; across all attacks, 96.0%. For chat-assistant text under attack the rate is 98.0%. AI paraphrasing is the strongest attack and still leaves 91.5% caught.
| Attack | Detected | 95% CI |
|---|---|---|
| No attack (baseline) | 97.3% | 1,500 texts |
| AI paraphrasing | 91.5% | 89.9–92.8% |
| Deleted articles (a, an, the) | 93.6% | 92.2–94.7% |
| Random upper/lower case | 95.2% | 94.0–96.2% |
| Synonym swaps | 95.9% | 94.8–96.8% |
| Deliberate misspellings | 96.1% | 95.0–96.9% |
| British/American spellings | 97.1% | 96.1–97.8% |
| Extra paragraph breaks | 97.2% | 96.2–97.9% |
| Look-alike letters | 97.3% | 96.4–98.0% |
| Invisible characters | 97.3% | 96.4–98.0% |
| Extra spaces | 97.3% | 96.4–98.0% |
| Changed numbers | 97.5% | 96.6–98.2% |
RAID documents held out of training, 1,500 texts per attack, 95% Wilson intervals. Detection also holds across formats: 99.6% for answers containing code, 98.8% with math notation, 99.9% with tables, and 98.9% for answers as short as 50–99 words.
Before screening its 2026 position papers, NeurIPS tested Pangram on papers accepted at an AI ethics conference, some from 2022, before ChatGPT, and some from 2025. We ran the same check with ParaTrace v7. On papers from 2022 it flags none, exactly like Pangram. On papers from 2025 it flags 7.3% as at least half AI-written, against 1.0% for Pangram.
| Detector | Papers | ≥ 50% | ≥ 90% | 100% |
|---|---|---|---|---|
| Pangram 3.3.2 | 159 | 0.0% | 0.0% | 0.0% |
| ParaTrace v7 | 106 | 0.0% | 0.0% | 0.0% |
| Detector | Papers | ≥ 50% | ≥ 90% | 100% |
|---|---|---|---|---|
| Pangram 3.3.2 | 204 | 1.0% | 1.0% | 0.0% |
| ParaTrace v7 | 123 | 7.3% | 1.6% | 0.8% |
The AI score is the share of a paper's passages scored as AI; 100% means AI use across many parts of a paper, not that every word came from AI. Pangram's rows are from the NeurIPS post AI-generated papers in the NeurIPS 2026 position paper track (2 June 2026); the samples differ in size, so compare shares, not counts. Which 2025 papers used AI is not recorded: because ParaTrace flags no paper from before ChatGPT, false alarms are an unlikely cause of its extra flags, but that is an inference, not proof. NeurIPS and Pangram Labs, Inc. have not reviewed this comparison. Full write-up on paratrace.net →
We publish the weak spots alongside the strengths. Head to head with Pangram on 21 public benchmark numbers, v7 is within 5 points on 15 of them; it trails clearly on text run through a commercial humanizer and on standard texts under 50 words. See the comparison →
Try ParaTrace on paratrace.net, call it from your software through the API, or talk to us about a private deployment for your submissions, applications or reviews.
Try ParaTrace Book a technical call
We use cookies
Some cookies are needed for this site to work. We would also like to set analytics cookies to understand how the site is used — but only if you agree. Privacy policy
Strictly necessary
Session and security cookies that keep forms and logins working. These cannot be switched off.
Analytics
Google Analytics, to count visits and see which pages are read. Off unless you turn it on.