Grammarly’s AI Detector sells itself on one line: 99% accuracy and first place on a named benchmark. That line is Grammarly’s own, printed on its product page. What follows lines it up against every dated outside test of the same tool that I could open and confirm: two tech journalists, two rival detector makers who published their samples, and one company that sells both detector and rewriting tools. Each figure carries its date and its source record, so you can check the work instead of trusting the summary.
Is Grammarly’s AI Detector Accurate?
Grammarly’s AI Detector is not accurate at the level Grammarly advertises, judged by every outside test that measured it. Grammarly claims 99% detection accuracy and the #1 rank on the RAID benchmark. A November 2025 test on the RAID dataset found it flagged about 22 in every 100 AI-written samples, and a January 2026 test caught none of nine.
Grammarly’s wording on its AI Detector product page is that the tool “achieves 99% detection accuracy and ranks #1 on RAID’s independent benchmark” (captured July 10, 2026 and re-checked September 22, 2026; EV-grammarly-04). RAID is a public academic dataset of machine-written text, built with deliberate alterations such as paraphrasing and spelling changes so that researchers can see which detectors hold up. An AI detector, for anyone new to the term, is software that estimates how likely it is that a passage was written by a language model rather than a person.
The sharpest outside check used that same dataset. A rival detector company published a test on November 12, 2025, running Grammarly’s AI Detector over 99 RAID samples that spanned 9 alteration types and 11 AI models. It reported a recall of 0.222, meaning roughly 22 of every 100 AI-written samples were flagged, and an F1 score of 0.364, a combined measure of catching AI text and not over-flagging (EV-grammarly-05). Grammarly’s quoted line does not say which slice of RAID or which setting its 99% describes, which is exactly why a test on the same dataset landing so far away deserves a reader’s attention.
The table below is my own assembly of the full dated record on this detector. The right-hand column says who ran each test, because the source’s stake is part of how much a number should weigh.
| Test (date) | Sample | What Grammarly’s AI Detector returned | Who ran it |
|---|---|---|---|
| Grammarly’s product page (captured Jul 10, 2026; re-checked Sep 22, 2026) | Not stated on the page | “99% detection accuracy”, “#1 on RAID” | Grammarly, the seller (EV-grammarly-04) |
| RAID-dataset test (Nov 12, 2025) | 99 samples, 9 alteration types, 11 AI models | Recall 0.222, F1 0.364 | A rival detector maker, method published (EV-grammarly-05) |
| 30-tool detector roundup (Jan 7, 2026) | 9 AI samples (3 each from ChatGPT-4o, Gemini 2.0, Claude 3.7 Sonnet) plus 3 human samples | 0 of 9 AI samples reached 75% AI; 3 of 3 human samples stayed at or under 25% AI | A rival detector maker, samples and pass marks published (EV-grammarly-ai-detector-01) |
| PCWorld (Nov 25, 2024) | 1 text written entirely by Gemini, run twice | 37% AI on both runs | Tech journalist, sells no detector (EV-grammarly-ai-detector-02) |
| Before-and-after rewrite test (Nov 3, 2025) | 1 AI sample, scored before and after one rewrite | 62% AI before, 0% AI after | A company selling its own detector and rewriting tools (EV-gpthuman-07) |
| MakeUseOf hands-on (Oct 30, 2025) | One working session | The Detector ran normally, as its own tool | Tech journalist, sells no detector (EV-grammarly-08) |
When I set those rows next to each other, the high readings turned out to be rare. The outside test that ran the RAID samples also scored three unmodified ChatGPT texts, at 99%, 93% and 71% AI (EV-grammarly-05). Beyond those three, the readings fall away: 22 in 100 deliberately altered RAID samples, 37% on a Gemini-written passage, and zero of nine ChatGPT-4o, Gemini and Claude samples clearing a 75% bar. The one test that also fed in human writing saw all of it cleared. Read together, the record describes a detector whose high readings on AI text are the exception, and whose misses all lean toward calling text human.
Is Grammarly’s AI Detector Free?
Yes, Grammarly’s AI Detector can be used through Grammarly’s free version. The January 2026 tester ran it that way and hit a limit of 10,000 characters per check, which comes to roughly 1,500 words of ordinary prose. Grammarly’s own site did not publish an exact limit when the page was re-checked on September 22, 2026.
That character cap comes from the tester’s own use of the free version, not from a Grammarly document (EV-grammarly-ai-detector-01), so treat it as an observed ceiling that Grammarly could move. In practice it means a typical short essay fits in one check and a long research paper needs splitting. Grammarly’s AI Detector also appears inside Grammarly’s Superhuman Go interface, listed as a tool of its own (EV-grammarly-08).
MeteGPT’s detector is free as well, and it asks for no account at all: paste up to 125 words, get a score, up to four times a day. That size fits the paragraph you are actually unsure about, the one a friend or a checker flagged. The rewriting side opens with a free account, which covers four rewrites of up to 250 words each, and whole essays belong to the paid plans covered in the verdict below.
What Is Grammarly’s AI Detector and How Does It Work?
Grammarly’s AI Detector is a standalone Grammarly tool that reads text you paste in and returns an estimate of how likely it is to be machine-written. It lives at grammarly.com/ai-detector and inside Grammarly’s Superhuman Go interface. It is a different product from Grammarly’s grammar checker, which corrects spelling and style and says nothing about authorship.
Grammarly’s public claim for the Detector centres on the accuracy figure and the RAID ranking (EV-grammarly-04). The general principle behind tools of this kind is shared across the field: a detector compares the statistical habits of your text, such as how predictable each word choice is and how evenly sentences run, against the habits language models tend to show. The result is a probability. It is never a record of who typed the words, and two detectors trained differently can read the same paragraph in opposite directions. For a plain walk-through of that process, see how AI detectors score text in general.
Is Grammarly’s AI Detector the Same as Grammarly’s Humanizer?
No. Grammarly’s AI Detector and Grammarly’s AI Humanizer are separate tools: the Detector scores text for signs of AI writing, and the Humanizer rewrites text to read more naturally. MakeUseOf’s October 2025 hands-on found the two listed as separate items in Grammarly’s Superhuman Go interface, and the Detector ran without trouble in that session.
The distinction matters for every number on this page. The 99% claim, the RAID-dataset recall and the 0-of-9 result all describe the Detector, and none of them says anything about how well the Humanizer rewrites. Plenty of people who search for Grammarly’s AI tools actually want the rewriting feature, and that product, including its pricing and Grammarly’s own statement on what it is for, has a dedicated review of Grammarly’s Humanizer on this site.
Can Grammarly Detect ChatGPT, Gemini and Claude Writing?
In the one dated test that tried all three, Grammarly’s AI Detector did not flag any of them with confidence. The January 2026 test ran three samples each from ChatGPT-4o, Gemini 2.0 and Claude 3.7 Sonnet, and none reached the tester’s 75% AI mark. A November 2024 PCWorld test scored a fully Gemini-written text at 37% AI.
The January 2026 roundup set a clear bar: an AI sample counted as caught only if Grammarly scored it at 75% AI or higher, and nine out of nine fell short (EV-grammarly-ai-detector-01). The tester, a rival detector company, summed it up by writing that the tool “didn’t pass any of our tests for detecting AI generated content accurately.” The PCWorld result is older and comes from a journalist with nothing to sell. The writer fed in a passage she describes as a complete fabrication by Gemini, got 37% AI on two separate runs, and called that consistent “but inaccurate by a large margin” (EV-grammarly-ai-detector-02). That test predates Grammarly’s current RAID marketing, so it describes an earlier version of the tool rather than a rebuttal of today’s claim.
Two practical points follow. The models in these tests are a generation or more behind what students use in late 2026, so none of the record speaks to the newest chatbots in either direction. And a low Grammarly score on a draft is weak evidence that a person wrote it, because the documented misses on reworked text all run toward reading it as human.
Does Grammarly’s AI Detector Flag Human Writing as AI?
In the only dated test that included human writing, Grammarly’s AI Detector flagged none of it. All three human-written samples in the January 2026 test scored at or under the tester’s 25% AI mark. Grammarly’s headline claim quotes an accuracy figure and does not give a separate rate for wrongly flagging human text.
A false positive is human writing that a detector labels as AI-generated, and it is the error that worries students most. The three human samples in that test were an excerpt from the Declaration of Independence, an excerpt from the Magna Carta, and a piece of original writing by the tester (EV-grammarly-ai-detector-01). Two of those are centuries-old public documents, so the set tells you little about a modern essay written by a student, and nothing about writers who learned English as a second language.
What the record does show is consistency. A detector that clears human text and under-reads AI text is behaving like a tool set to lean human, which lowers the chance of a false flag at the cost of missed AI text. If your own writing came back flagged somewhere and you want a second reading from a differently built engine, MeteGPT’s detector gives one in seconds.
Does Grammarly’s AI Detector Catch Paraphrased or Rewritten Text?
Often not, according to the dated record. The November 2025 RAID-dataset test reported that paraphrasing, alternative spellings and formatting changes all lowered Grammarly’s detection. A separate November 2025 test watched one AI sample drop from 62% AI to 0% AI on Grammarly’s AI Detector after a single rewrite pass.
The first finding comes from the same test that produced the 0.222 recall, and it names the alteration types that pulled Grammarly’s scores down (EV-grammarly-05). The second is one sample from a company that sells both a detector and a rewriting tool, published November 3, 2025 (EV-gpthuman-07). One sample is a single data point, not a rate. Still, both results point the same way as the rest of the record.
For a reader checking a draft, the useful takeaway is about what a clean Grammarly score can prove. A 0% reading shows the text did not match the patterns this detector looks for. It does not show the text was written from scratch by a person, and it tells you nothing about how a school-licensed checker built on a different model will read it.
How Accurate Is Grammarly’s AI Detector Compared to Other Detectors?
No dated source in this record tests Grammarly’s AI Detector against school-licensed checkers on the same samples on the same day, so no fair ranking exists. The record does show Grammarly’s own 99% claim sitting far above every outside accuracy figure on record, which is a strong reason to compare readings across detectors rather than trust any single one.
School-licensed checkers run inside a university’s own license, report to instructors rather than to the student, and use their own models, so a Grammarly reading is not a preview of what your school will see. If you want to know what the checker most universities license shows students and when, that has its own guide. Detectors also get retrained without announcement, which is why every figure here carries a date. For the wider field, the ranked comparison of AI detectors sets each tool against its own dated sources.
MeteGPT was built around the habit this page keeps recommending. Its engine learned from 2,590 real student essays, it checks every rewrite with its own detector before showing it, so a rewrite and its verdict arrive together, and its humanized output was measured across six public detectors in a run published with its date on MeteGPT’s methodology page. If a draft needs work rather than just a verdict, you can rewrite a passage and get its detector verdict in the same window.
Which AI Detector Has the Lowest False Positive Rate?
No public test has scored every detector on the same human writing, so no detector can honestly be named lowest; the one to trust is the one that publishes its rate. MeteGPT’s detector measured 0.25% in its 30 July 2026 test: out of 401 pieces of real human writing, it called only one AI. Among them were 91 essays by non-native English writers, the people detectors accuse most often, and it cleared every single one.
That second number is the one to look at if English is not your first language. The 91 essays come from the TOEFL set in the Stanford-led study by Liang and colleagues (2023), where seven commercial detectors of that year labeled 61.3% of these fully human essays as AI on average. MeteGPT’s detector labeled 0% of them. No detector can promise it will never misread a person, which is exactly why the rate belongs in public. Grammarly publishes a headline accuracy figure; MeteGPT publishes how often it gets real people wrong, split here by who did the writing, including whether each group was kept out of the detector’s training:
| Human writing tested (30 July 2026) | Passages | Wrongly called AI | Held out of training? |
|---|---|---|---|
| Second-language writers: TOEFL essays (Liang et al., 2023) | 91 | 0 (0%) | Yes |
| Native-speaker student essays, the comparison group | 60 | 0 (0%) | Yes |
| Student essays written before ChatGPT existed | 150 | 0 (0%) | Partly |
| Published long-form prose (MAGE human set) | 100 | 1 (1.0%) | No |
| All human writing | 401 | 1 (0.25%) |
If a Grammarly reading or a teacher’s tool has already questioned something you wrote yourself, run the same paragraph through the detector behind that July 2026 result and keep the result with your drafts.
What Is a Good Alternative to Grammarly’s AI Detector?
MeteGPT’s detector is a strong alternative, because it was built from research rather than from a marketing number. Its founder, Fırat Mıhcı, studies how automated checkers treat people writing in a second language, and the version tested on 30 July 2026 learned from 15,542 real human passages, so that ordinary, careful human writing reads as human.
Three design choices follow from that starting point, and each one shapes the reading you get:
- The cut-off moves with length. A short paragraph and a full essay are judged against separately calibrated lines, so a brief answer is not held to the standard of a long essay.
- It refuses to guess. Text under about 60 words is not scored at all, and a borderline passage is labelled “inconclusive” beside its number, so a middling reading is never dressed up as a verdict.
- It points to the wording. Beside the score it lists common AI-style habits it found in your text, each with the phrase quoted back to you, so you have something concrete to revise instead of a bare percentage.
The same standard runs through the rest of the site. Claims about other tools on MeteGPT pages follow a written evidence protocol and link to a public, dated evidence log, which is why this page could put Grammarly’s 99% next to the outside tests at all. The rewriting side learned from 2,590 genuine student essays, and every rewrite is checked by this detector before you see it, so a humanized draft arrives already marked when it reads as human.
Getting started costs nothing. Sign up for a free MeteGPT account and you get four rewrites of up to 250 words each, every one checked by the detector before you see it, with no card needed. The detector on its own works even before you sign up, on a paragraph of up to 125 words. For whole essays and regular use, MeteGPT’s paid plans cover full essays with the detector check included.
Should You Trust a Grammarly AI Detector Score?
Treat a Grammarly AI Detector score as one reading, not a verdict. The dated record shows it clearing human samples and scoring AI text well below its 99% marketing, so a low score is weak proof that a person wrote the text. A second detector built on a different model gives you a far sturdier picture before you submit.
Grammarly deserves credit on two counts. It named a real, public academic benchmark as the basis of its claim, which lets anyone test the claim at all, and in the one test that included human writing, it did not wrongly flag any of it. The trouble sits in the gap between the headline and the outside results: 99% on Grammarly’s page, 22 in 100 AI samples flagged on the RAID dataset, zero of nine caught in January 2026, and 37% on a text written entirely by Gemini.
The practical routine is simple. Keep your drafts and version history, which settle authorship far better than any score. When a result surprises you, check the same paragraph on a second detector. MeteGPT’s detector costs nothing for a paragraph, and when a whole essay needs rewriting and checking together, see what each MeteGPT plan includes. The quickest next step takes under a minute: paste a paragraph into MeteGPT’s detector now and set its reading beside Grammarly’s.
Evidence summary (MeteGPT Evidence Protocol v1.0).
The sweep for this page, run September 22, 2026, identified 24 candidate sources and screened all 24. It excluded 20: 16 affiliate or roundup pages that printed accuracy or false-positive percentages with no sample or method behind them, 2 that could not be verified with this run’s tools, and 2 that were off-topic. It kept 4. In all, the page cites six dated records: two new ones (EV-grammarly-ai-detector-01 and EV-grammarly-ai-detector-02) and four earlier ones (EV-grammarly-04, EV-grammarly-05, EV-grammarly-08, EV-gpthuman-07), each re-opened live on September 22, 2026 with its figures unchanged. One collector coded this sweep; a second coder still has to go through it independently, and the agreement between the two will be added here when that happens. Two of the tests come from companies that sell a rival detector and one from a company that sells detector and rewriting tools; both of those stakes are named in the table above, and their full records, with links and capture dates, sit in the public MeteGPT evidence log. Disclosure: MeteGPT makes its own AI detector, which is exactly why every figure about Grammarly on this page comes from someone else’s dated test.
Limitations.
- The second independent coding pass on this page’s new records has not run yet; its agreement rate will be added here once recorded.
- The samples are small: 9 AI and 3 human texts in the January 2026 test, 99 in the RAID-dataset test, and a single text in each of the other two. None of them is a rate for your document.
- The PCWorld test dates from November 2024, before Grammarly’s current RAID marketing, and describes an earlier version of the Detector.
- The AI models tested (ChatGPT-4o, Gemini 2.0, Claude 3.7 Sonnet) are older than the chatbots in common use in late 2026.
- The 10,000-character free limit is a tester’s observation, not a figure Grammarly publishes, and Grammarly can change it.
- No dated source in this record describes what Grammarly’s paid plans add to the Detector, so this page states nothing about them.
- Community threads were searched and none could be verified this run, so no community report appears above.
- Grammarly can retrain its Detector at any time; every reading here is tied to the date it was taken.
Written by Fırat Mıhcı, who studies how automated checkers read second-language English (ResearchGate) and builds MeteGPT. Corrections or a newer dated test of Grammarly’s AI Detector are welcome at hello@metegpt.com; confirmed changes are dated on this page the week they arrive.
Doubt a Grammarly reading? Get a second opinion in seconds.
Paste the passage Grammarly scored, let MeteGPT rewrite it, and our detector checks the new version before you see it, marking it ‘Reads as human’ on the same screen when it clears.