HomeQuillBot AI Detector

QuillBot AI Detector: What Its Cited RAID Score Leaves Out

By Fırat Mıhcı. I build MeteGPT, and the part of this field I keep returning to is the space between a vendor’s headline percentage and the data sitting underneath it. QuillBot’s detector turned out to be the most interesting case I have worked on for that reason: the number it advertises is real, and the paperwork behind it is public, so in this case the whole claim can be opened up and read line by line. That is what this page does, with a date on every figure. My ResearchGate profile is here. Published August 8, 2026.

TL;DR: QuillBot’s AI detector advertises a 99% detection rate credited to RAID, an academic benchmark, and RAID’s own published data mostly backs that up: 99.62% at a 5% false-positive setting, still 94.07% against paraphrased AI text at a 1% setting. What QuillBot’s page never states is that setting, or which half of RAID’s split leaderboard its entry sits in.

How this page was built (MeteGPT Evidence Protocol v1.0).

The way sources get chosen, weighted, and revisited here is written out on our methodology page. This one weighs a single product, QuillBot’s AI detector, against QuillBot’s own live pages, against the academic benchmark QuillBot names as its accuracy source, and against every outside test of that detector I could open and confirm on August 8, 2026. Logged retrieval produced 23 candidate sources; all 23 were screened and 4 were set aside for reasons kept in the record: two competing-vendor reviews that print percentages with no test disclosed behind them, one community branch I could not verify at all, and one duplicate of an article already counted. The kept sources yielded 21 evidence records, 16 written for this page plus 5 earlier records re-opened and checked against their live sources. Records and sources do not map one to one, since five of those records come out of files in the same public benchmark repository. Anything that survives screening is meant to be labelled twice, by two coders working apart from each other, and exactly where that second pass stands for this page is written into the limitations list near the end. Behind every source named above sits a record with its own capture date, and those records are published in full in the public evidence log.

My stake, stated before any number. I am not a bystander in this market. MeteGPT sells a detector and a humanizer both, so a reader who walks away reaching for either one has done my business a favor, and you should read every paragraph below with that in mind. It is also the reason this page labels each figure by who produced it: a vendor’s marketing line, a rival detector’s in-house trial, a journalist’s repeat testing, and a public academic benchmark are worth different amounts, and I say which is which as I go. Two things this page is not. It is not an argument that QuillBot’s detector fails, because the evidence I gathered does not support that argument. And it is not a claim about anyone’s motives; where I found QuillBot’s own pages saying different things, I report the difference and stop there.

What the record shows, strongest evidence first.

  1. The 99% holds up at the setting it comes from. QuillBot’s page reads, verbatim, “Detect AI text reliably with a 99% detection rate, according to independent evaluations from RAID.” QuillBot’s own submission file in RAID’s public repository puts the figure at 99.62% when the benchmark is tuned so that 5 in 100 human documents get wrongly flagged, and 98.10% at the stricter 1-in-100 setting. Against AI text that RAID had deliberately paraphrased, at that stricter setting, it still reads 94.07%. That is a drop of about five and a half points at the hardest condition I checked, which is a smaller penalty than a skeptical reading expects.
  2. One distinction that would change what the number means cannot be checked at all. RAID’s paper states that its leaderboard is divided between detectors whose owners report having trained on the RAID dataset and detectors whose owners do not. The metadata file every submitter fills in, QuillBot’s included, has no field that records which of the two applies, and the live leaderboard page returned nothing my tools could read. So which side QuillBot’s entry belongs to is not something I can establish, and I do not assert either answer anywhere below.
  3. QuillBot’s own two pages about this detector do not agree on where the number comes from. The product page leads with the RAID citation. QuillBot’s separate blog post about the same detector, which ranks near the top for this search, contains zero mentions of RAID and zero mentions of the 99% figure, checked against the raw page on August 8, 2026.
  4. The independent tests are real evidence and much smaller evidence. Every outside test of this detector I could verify ran between 3 and 50 texts in a single pass, landing anywhere from roughly 78% to 100%. Those are not bigger numbers beating a smaller one; they are a different and thinner kind of measurement than a benchmark that scores detectors across 11 generative models, 8 content domains, and 11 adversarial attacks.
  5. QuillBot and Scribbr are sister brands, and the page that grades one against the other does not say so. Learneo’s own announcement names QuillBot and Scribbr among its operating businesses. Scribbr’s detector ranking scores QuillBot’s free detector at 78% and ties it with Scribbr’s own free tier, without disclosing the shared parent on that page.
  6. Almost nobody writing about this detector has no stake in the answer. Of the pages that rank for this question, most are published either by rival detector vendors or by companies selling their own writing tools inside the review. That includes me, which is why the disclosure above sits above the findings rather than under them.

People usually type “quillbot ai detector” or some version of “is it accurate” in the few minutes after a scan handed back a number they did not expect, or just before pasting a document they care about. This page answers that in the order the question actually gets asked: what the tool is, what the outside record says, what QuillBot’s cited benchmark score really measures, what the free tier will and will not do, and what a single score from it settles.

What Is QuillBot’s AI Detector?

QuillBot’s AI detector is a Learneo product that answers one question about a passage you paste in: what share of it looks machine-written, expressed as a figure from 0 to 100 with the individual sentences driving that figure marked up. It lives at quillbot.com/ai-content-detector, and hardly anyone arrives there on purpose. It shares an account with QuillBot’s paraphraser, a tool millions of students and writers already keep open, so for most people the detector shows up as one more tab beside something they were already using rather than as a decision to evaluate a detector at all.

Three things need separating before any accuracy talk, because this brand carries several products with similar names. First, QuillBot’s AI detector is not QuillBot’s paraphraser or its Humanizer. One scores text, the others rewrite it, and they are different products with different jobs; I keep QuillBot’s paraphraser and Humanizer on their own page and this page holds no rewriting content at all. Second, QuillBot’s detector is not Scribbr’s detector, even though both belong to the same parent company, Learneo, a relationship that gets its own section later. Third, RAID here means the academic benchmark by that name, not any other use of the acronym.

On output, QuillBot’s product page describes what you get in some detail: a score from 0 to 100% for how likely the text was AI-generated, highlighting at the sentence level, explainer cards next to the result, and a four-way authorship label covering written by AI, generated then refined by AI, written by a human and refined by AI, and written by a human, on the product page read August 8, 2026. That page also tells readers not to lean on AI detection alone when a decision could affect somebody’s career or academic standing, which is a fair thing for a vendor to print and not every vendor prints it.

One piece of context about the writing you will find on this subject, mine included. Of the pages ranking for this question, most come from rival AI detectors reviewing a competitor, or from companies selling their own writing products and recommending them inside the review; the genuinely disinterested coverage exists but is a minority. That is a fact about who publishes, not an accusation about anyone’s honesty, and the practical consequence for a reader is simple: check who owns the tool being praised at the bottom of any review, including this one.

Is QuillBot’s AI Detector Accurate? What RAID Says and What Hands-On Testers Found

QuillBot states its detector achieves a 99% detection rate “according to independent evaluations from RAID.” QuillBot’s own help center is more cautious about detection generally, telling users in writing that AI detection tools are not 100% accurate and that paraphrased or even human-written content can be flagged as AI-generated. Both statements come from the same company, and holding them together is the honest starting point: a strong benchmark figure, plus a vendor admission that the category makes mistakes.

The outside record is smaller than the benchmark and less consistent. Here is every independent test of this detector I could open and verify on August 8, 2026, with who ran it and what they disclosed.

Test (date)Who ran itSampleQuillBot result
ZDNet ongoing re-test (published Oct 29, 2025)David Gewirtz, ZDNet; sells no detectorfixed text set, re-run on each pass100% correct on the latest pass, after earlier passes the writer called “wildly inconsistent” across repeats of the same text
Scribbr’s own comparison (snapshot Aug 11, 2025)Scribbr, QuillBot’s sister brand under Learneo30 texts, six categories, 12 detectorsfree detector 78%, tied with Scribbr’s free tier; Scribbr’s paid tool highest at 84%
BeLikeNative comparison (updated Aug 3, 2026)BeLikeNative, sells its own writing extension50 texts, composition not disclosedabout 80%, behind GPTZero at 99.5% and Scribbr at 84% in the same run
Textero review (updated Jun 19, 2026)Textero, sells its own writing tools3 textsclassic literature read 100% human; raw GPT-4 text 97% AI; that same text after QuillBot’s own paraphraser 100% AI
Originality.ai review (Oct 28, 2025)Originality.ai, a rival detector3 ChatGPT samplesall three read 100% AI, matching the reviewer’s own tool; author flags the sample as too small to generalize
Independent writer’s informal run (Apr 29, 2025)Himanshu Kumar, Medium; no tool of his own3 scenarios, no percentages givena genuinely human piece came back partly AI; an AI draft with roughly 40% manual rewriting read as human; raw AI text was flagged correctly
GPTZero’s review (Mar 2, 2026)GPTZero, a rival detectorno test of its own; cites a journalist’s“around 80% accuracy”

Read that band for what it is. The largest of those tests ran 50 texts in one pass with its composition undisclosed; three of them ran three texts. Set against RAID’s public benchmark, these are ad hoc single-run trials, and the honest framing is not that a smaller number beats a bigger one. It is that two different kinds of evidence exist here, and neither has been reconciled with the other by anyone, including me. Where a test comes from a company selling a competing product, the table says so; that is a documented ownership fact and not a claim that any of those results were cooked.

One thread inside that table is worth pulling, because it shows how quickly an accuracy figure decays. GPTZero’s review of QuillBot cites “around 80% accuracy” and hyperlinks that phrase to a specific ZDNet article. When I re-fetched that exact article, the ZDNet writer’s current pass scores QuillBot at 100%, and he describes his earlier runs of the same text as wildly inconsistent. Nobody did anything wrong there: an ongoing test moved, and a page citing it did not. It is a good argument for asking of any detector percentage you meet, including every one on this page, when it was measured and against what.

Has QuillBot’s Detector Ever Flagged Human Writing as AI?

Yes, in documented individual cases, and no published rate exists to tell you how often. A false positive means genuine human writing scored as machine-made. The clearest recorded instance in my set comes from an independent writer with no tool of his own to sell: a piece he had actually written came back labelled partly AI-generated, while an AI draft he had rewritten by roughly 40% passed as human. BeLikeNative’s comparison separately reports the detector struggling with paraphrased AI text and with inputs under 80 words. QuillBot itself states plainly that human-written content can be flagged.

What I could not find, on any QuillBot page I read this run, is a published false-positive rate for this detector. The closest figure that exists is the operating point inside QuillBot’s RAID submission, which is a false-positive setting of 5% or 1% depending on the column, and that number lives in a benchmark data file rather than on the marketing page. I am reporting an absence, not inventing a substitute for it. If you came to English later in life and keep getting flagged, one piece of peer-reviewed context belongs here with its scope intact: a Stanford-led study measured heavy false-positive rates on non-native TOEFL essays across the detectors it tested, and QuillBot was not among the tools in that study. That finding describes the category this detector belongs to. It is not a QuillBot figure and I will not repackage it as one.

MeteGPT’s Own Result on QuillBot’s Detector (FC-METEGPT-001)

I hold exactly one first-party number that touches this tool, and it needs its boundaries drawn before the figure itself. In a small controlled in-house run from mid-May 2026, across about 30 academic-style passages, the rewritten output of our own humanizer came back clean on QuillBot’s detector on all 30 of them. Now the boundaries. That result measures how our engine’s output happened to read on their tool on one date, on one narrow set of passages. It says nothing about how accurately QuillBot’s detector judges anyone else’s writing, and treating it as an accuracy score for QuillBot would be a category error, so I am not making that move and neither should a reader. Five limits ride with that figure and none of them are optional. Thirty passages is a small run. The data is our own, open to the owner on request but never audited by anyone outside this company. Inputs that sit far from ordinary academic prose, source code most obviously, read very differently on strict detectors. Detectors get retrained without announcing it. And whatever a tool reported on one afternoon in May says nothing binding about the document you are about to submit.

What Does QuillBot’s Cited RAID Score Actually Measure?

RAID is a real academic benchmark, and this is the part of QuillBot’s claim that almost no page covering this detector opens up. The paper behind it, Dugan et al., arXiv:2405.07940, was published at ACL 2024 and, as described at publication, evaluated detectors across 11 generative models, 8 content domains, 11 adversarial attacks, and 4 decoding strategies. An adversarial attack in this sense is a deliberate alteration the benchmark applies to AI text before testing, such as rewriting it or inserting invisible characters, to see whether a detector still recognizes it. QuillBot’s participation is why any of this is checkable: the benchmark publishes each submission’s raw results in a public repository, so QuillBot’s figures can be read directly rather than taken on faith.

Read directly, they say something more specific than “99%”. A detection percentage is only meaningful next to its operating point, meaning the trade-off the tester chose between catching AI text and wrongly flagging human text. QuillBot’s submission reports both settings:

Slice of QuillBot’s own RAID submissionSettingResult
All conditions aggregated5 human documents wrongly flagged per 10099.62%
All conditions aggregated1 human document wrongly flagged per 10098.10%
AI text rewritten by the benchmark’s paraphrase attack1 per 10094.07%
AI text with zero-width characters inserted1 per 10091.39%

Two honest readings come out of that table, and the first is the one I did not expect to be writing. The figure survives its own scoping better than a skeptical reading assumes: moving from the friendliest condition to the hardest one I checked costs it about five and a half points, not thirty. For context I pulled one rival’s submission file out of the same repository and ran the identical arithmetic on it. GPTZero’s entry, filed October 7, 2025, comes back at 71.60% on that same paraphrase slice at the same 1-in-100 setting, against QuillBot’s 94.07%. Read as a like-for-like comparison of two self-filed result files, and nothing broader than that, QuillBot’s weakest measured condition is still ahead of a rival’s on the identical test.

The second reading is the disclosure gap. The 99% that appears in QuillBot’s marketing is the aggregate at the 5% setting, and that setting is the price of the number: at that operating point, roughly 5 in every 100 human documents get flagged by design. QuillBot’s page states neither the setting nor the condition. Nor does it mention a detail sitting in its own submission file that shapes how the benchmark result generalizes: the threshold is calibrated separately for each of RAID’s content domains, and those thresholds span roughly four orders of magnitude, from 0.000126 for academic abstracts to 0.936 for product reviews at the 5% setting. That is standard benchmark practice and I am reporting it as methodology, not as a fault. It does mean a benchmark score is tuned per domain in a way your single document never is.

One last piece of context, quotable and easy to verify at the link above. RAID’s own abstract opens by naming commercial detectors that claim to detect machine-generated text with very high accuracy, 99% or higher, as precisely the kind of claim the benchmark was built to interrogate. QuillBot then cites RAID in support of a 99% claim. I draw no conclusion from that. It is simply worth knowing that the source and the marketing line are aimed in slightly different directions.

The Trained-vs-Not-Trained Split QuillBot’s Page Doesn’t Show

Here is the question I most wanted to answer on this page and could not. RAID’s paper is explicit that its leaderboard has two halves: “The leaderboard is split up into two sections—one for those who self-report having trained on the RAID dataset and one for those who do not. This is important to ensure that a clear distinction is made between detectors that are generalizing to out-of-domain data and those that are not”. The distinction matters because a detector tested on data resembling what it trained on is doing an easier job than one meeting the material cold, and a reader deciding how much to trust a 99% figure has a real interest in knowing which situation produced it.

That answer is not available to me. The metadata file every submitter to RAID fills in has no field recording the distinction at all, in QuillBot’s submission or in the repository’s own template, and RAID’s live leaderboard page returned nothing readable to the tools I use, so I could not confirm how, or whether, the split is displayed there. I am stating that as an open question rather than guessing at it, because guessing would put an implication in your head that no source of mine supports. What follows from it is narrow and worth carrying: a reader who clicks the link on QuillBot’s own page cannot determine which half of that leaderboard the 99% comes from, and neither can I.

One checkable fact does attach to the submission, and it is a date. QuillBot’s entry is dated January 30, 2025, which makes it the oldest of the major commercial submissions in that repository; GPTZero’s is dated October 7, 2025, It’s AI’s December 25, 2025, and Grammarly’s February 9, 2026. A detector marketed as continually improving is a moving object, while a benchmark submission is frozen at the day it was filed. The 99% therefore describes the detector as it stood in January 2025, which is not a criticism of the figure so much as a boundary on it.

Does QuillBot’s AI Detector Give a Score or a Category? Its Own Pages Disagree

The answer is both, and knowing that in advance keeps a result from reading as more definite than it is. QuillBot’s product page describes a percentage from 0 to 100 for AI likelihood, sentence-level highlighting so you can see which passages drove the score, explainer cards beside the result, and a four-way authorship label that separates AI-written text, AI-written and then refined text, human-written and AI-refined text, and human-written text. None of those four labels is a finding of fact about who wrote your document. All of them are a model’s estimate.

Where QuillBot’s own pages come apart is on the sourcing of the accuracy claim, not the output. The product page leads with the RAID citation and links the leaderboard. QuillBot’s separate blog post about the same detector, a page that ranks near the top for this exact search, mentions RAID zero times and the 99% figure zero times; I checked the raw page for both strings on August 8, 2026. So whether you ever learn that this detector’s headline number rests on an academic benchmark depends entirely on which of QuillBot’s two pages you happen to land on. I am not claiming that is deliberate, and I have no way to know. I am recording that two live pages from one company, about one product, describe the evidence behind it differently, and that a reader trying to trace the number should make sure they are reading the page that actually cites it.

Is QuillBot’s AI Detector Free? Word Limits, Daily Scans, and What Premium Actually Changes

The detector is free to use, with three limits that QuillBot publishes on the product page and that most reviews state only partially. Your text needs to be at least 80 words before the tool will scan it. A single free scan takes up to 1,200 words. And free accounts get up to 6 scans per day, verbatim from the product page read August 8, 2026.

Those numbers have practical consequences worth planning around. A standard 500-word assignment fits in one scan. A 4,000-word paper does not, so you would be splitting it into four chunks and burning four of six daily scans, and four chunk-level readings are not the same thing as one whole-document reading, since a score computed over a short passage carries more noise than a score over a long one. The 80-word floor also means the abstract, the cover note, or the two-paragraph discussion post you were most worried about may be too short to scan at all.

Free vs Premium: What Actually Changes (the Help-Center Contradiction)

If you are considering paying QuillBot specifically to lift those limits, check the current pages first, because the two QuillBot sources I read on August 8, 2026 do not describe Premium the same way. The product page sells Premium on no word limit for detection scans, unlimited scans, and detailed explainers across the full text. QuillBot’s own help-center article on whether the detector is free or premium, updated about a month before I read it, says something narrower: that the detector is free to all users, that “the only difference for Premium users is in file uploads,” free users uploading one file at a time against 20 at once for Premium, and that “the detection itself works the same for both Free and Premium accounts.” Same company, same week, two accounts of what the money buys. I am not picking a winner between them and I make no claim about why they differ. One disclosure belongs on that quote: I captured it on August 8, 2026, and QuillBot’s help domain now refuses automated requests from my tools, so I cannot re-pull it on demand. Treat it as one dated reading of a page that can be edited at any time, and check the live article yourself before you pay. The reader-level takeaway is that the word cap and the daily scan count are exactly the things to verify on the live page before paying to remove them.

Since I am about to point at my own tool, let me name its ceiling in the same breath. Our detector at MeteGPT costs nothing and asks nobody to sign in, and an anonymous run stops at 125 words, roughly a tenth of what QuillBot’s free scan allows, with 4 runs a day. That is deliberately a small box. It is built for a second opinion on a paragraph you are unsure about, not for a thesis chapter, and if you paste 2,000 words into it you will be disappointed rather than served. The same ceiling applies to the humanizer side of our free tier, which is where most people run into it: a 125-word rewrite is enough to see what the engine does to your voice and not enough to process a real assignment, and whole-document work is what a paid plan is actually for. That is the honest trade, laid out on our pricing page rather than hidden behind a signup. For a short passage, though, free is genuinely free: you can check a passage in our own detector and hold the two readings side by side, which is a better habit than trusting either tool alone.

QuillBot’s AI Detector vs Turnitin, GPTZero, and Scribbr

No same-day, same-samples head-to-head test of these tools exists anywhere in the record I gathered, so I will not manufacture a ranking. What is documented is how each of these comparisons actually differs, which matters more than a league table for someone deciding which reading to worry about.

QuillBot vs Turnitin

These two answer to different people, and that difference outweighs any percentage. QuillBot’s detector is a self-serve scan you run on yourself. Turnitin’s AI indicator runs inside an institution’s license and produces its reading for an instructor, not for you. Turnitin also reports in an unusual way that is worth knowing before you compare readings: below a 20% threshold its AI writing report displays an asterisk with no numeric score and no highlighted text, which is why I treat Turnitin as an under-20% bound rather than a precise figure anywhere on this site. The consequence is blunt. A clean QuillBot result is a self-run signal from a different model, not a preview of the reading your school will see. I keep Turnitin’s AI writing report, and why it prints a bound instead of a number, on its own dated page.

QuillBot vs GPTZero

Both detectors have entries in RAID’s public repository, QuillBot’s filed January 2025 and GPTZero’s October 2025, which is the closest thing to a common yardstick either of them has. Beyond that, the comparison most often quoted is GPTZero’s own review of QuillBot, which puts it at “around 80% accuracy” and links a journalist’s test that now reads 100% at the same URL. Weigh that as a rival’s page citing a source that has since moved, and weigh GPTZero’s own accuracy claims with the same suspicion you would apply to QuillBot’s; I put GPTZero’s own accuracy claims through the same check separately.

QuillBot vs Scribbr (Same Parent Company, Different Detector)

This is the comparison most worth understanding, because it is the one people use as a cross-check without realizing what they are doing. Learneo’s own announcement lists QuillBot and Scribbr among its operating businesses, so the two are sister brands under one parent. Scribbr publishes a detector ranking, authored by a Scribbr co-founder, that discloses a real method, 30 texts across six categories run through 12 detectors, and reports QuillBot’s free detector at 78%, tied with Scribbr’s own free tier, with Scribbr’s paid tool on top at 84%. Two observations, both about composition rather than motive. The shared parent company is not disclosed on that ranking page. And a vendor’s own comparison naming its own paid product the winner is a self-test, which is a ceiling on how much weight the 78% can carry, not proof it is wrong. The practical upshot for a reader: running a QuillBot result past Scribbr’s detector for a second opinion is not the independent check it looks like, so reach for a third tool with no relation to either. I cover Scribbr’s detector, including what it does and does not disclose, on its own page, and the wider field sits in my ranked record of every detector I have checked.

Can You Rely on a QuillBot AI Detector Score? The Verdict

Not on its own, and the reason is narrower and more interesting than the usual verdict in this niche. QuillBot’s headline number is not inflated. Checked against the benchmark it credits, it reads 99.62% at the 5% false-positive setting and holds at 94.07% under paraphrased AI text at the 1% setting, which is a solid record and a better one than a suspicious reader would predict. What is missing is not accuracy. It is the labelling: no operating point beside the number, no condition, and no way for a reader following QuillBot’s own link to learn which half of RAID’s split leaderboard the figure comes from. A percentage without its operating point cannot be compared to any other percentage, which is exactly what a person choosing between detectors is trying to do.

Credit where the evidence supports it, and it supports more than I expected. QuillBot named a genuine peer-reviewed benchmark rather than an in-house test, linked it, and submitted results that anyone can now read line by line in a public repository. That is why this page could be written at all. QuillBot is not alone in that: Grammarly credits RAID for its own detector and filed a submission in the same repository, dated February 2026. What makes the pair unusual is the rest of the field, where a headline percentage usually traces back to nothing a reader can open. QuillBot also publishes its free-tier limits openly and states in its own help center that detection tools are not 100% accurate and can flag human writing. Set against that: two of its own pages describe the evidence differently, two of its own pages describe Premium differently, and the outside tests of it are all small single-run trials that scatter between roughly 78% and 100%.

So treat a QuillBot reading as one dated, unaudited signal from one model. If it comes back clean, that is genuinely reassuring and not a clearance from whatever system will actually judge your work. If it comes back flagged on writing you produced yourself, that is a documented failure mode of the category and of this tool specifically, and the response that holds up is a record rather than a rewrite: keep your drafts and version history, ask which system produced the score and at what length, and get a second reading from a differently built engine before anything becomes formal.

The conflict, and what I do about it.

MeteGPT sells both a detector and a humanizer, so no direction a reader turns here is bad for me, and I would rather you weigh that against everything above than discover it at the bottom of a footer. That stake is why my own single measurement on this page carries more caveats than the number deserves, why I refuse to convert it into a verdict on QuillBot’s accuracy, and why I have not claimed anywhere that our detector reads QuillBot’s cases better, having never run the two side by side. It is also the sharpest difference between this page and much of what ranks around it: most of the detailed writing on this question is published by companies with a competing tool to sell, and that stake usually goes unstated. Mine is stated on every page, and the protocol behind every record here is set out in full on our methodology page. If what you actually want is to reshape an AI draft into something that sounds like you, rather than to audit a score, that is a different job and our own tool puts the rewrite and the score on one screen.

The limits of this page, stated plainly.

  • Which half of RAID’s split leaderboard QuillBot’s entry sits in is unresolved. The submission metadata carries no field for it and the live leaderboard page returned nothing readable to my tools this run, so no answer is asserted in either direction.
  • The rival comparison figure, GPTZero at 71.60% on the paraphrase slice, was computed by me from GPTZero’s own submission file using the same aggregation applied to QuillBot’s. Running that method against QuillBot’s file reproduced the 94.07% already on record, which is the check that makes the pair comparable. It compares two self-filed result files on one slice and is not a general ranking of either detector.
  • Every independent test in the accuracy table ran between 3 and 50 texts in a single pass, and several disclose neither their sample composition nor a repeat run. Treat the 78% to 100% spread as a set of unrelated snapshots taken under unrelated conditions, never as a range one detector was measured across.
  • Four of the tests come from companies selling a competing detector or writing product, which is an ownership fact recorded next to each result rather than a judgment about those results.
  • QuillBot publishes no false-positive rate for this detector on any page I read this run, so none is stated here; the operating points quoted come from its RAID submission, not its marketing.
  • The community side of this question could not be sourced this cycle. Reddit is blocked to my fetch tools, and five documented query variants surfaced no verifiable thread URLs, so no community-reported claim about this detector appears above rather than being filled in from memory.
  • The Stanford non-native-English study measured a set of detectors that did not include QuillBot, so it stands as background for the category and never as a score for this tool.
  • Our own 30-of-30 result is a small controlled in-house reading of what our humanizer produces, verifiable by me but not audited externally, describing one date and one narrow sample.
  • The independent second-coder pass for this run is scheduled under the protocol and is not complete; sources were verified and coded in a single curator pass, and the reconciled figure gets added when that pass runs.
  • Product limits, prices, and page wording are capture-date facts from August 8, 2026 and can change without notice, including the two QuillBot pages that currently disagree with each other.

Common Questions

Is QuillBot’s AI detector accurate? By the benchmark it cites, yes at a stated setting: 99.62% where 5 in 100 human documents are wrongly flagged, 98.10% at 1 in 100, and 94.07% against paraphrased AI text at that stricter setting. Small independent trials of 3 to 50 texts land between roughly 78% and 100%. Read any single score as one signal, not a verdict.

Is QuillBot’s AI detector free? Yes, with published limits: an 80-word minimum, up to 1,200 words per scan, and up to 6 scans per day on the free tier. What Premium changes is genuinely unclear from QuillBot’s own material, since the product page advertises unlimited scans and no word cap while its help center says the only difference is bulk file uploads.

Is QuillBot’s AI detector as accurate as Turnitin? No comparison in my record tests them on the same samples on the same day. They also serve different people: Turnitin runs inside your institution’s license and, below a 20% reading, shows an asterisk instead of a number. A clean QuillBot score is a self-run signal, not the reading your school will see.

Is QuillBot’s AI detector the same as Scribbr’s? No. They are separate tools with separate models as far as any source I verified shows, but both brands belong to Learneo, so a Scribbr cross-check on a QuillBot result is not an independent second opinion. Scribbr’s own testing puts the two free detectors level at 78%.

Can QuillBot’s detector flag my own writing as AI? It can, and QuillBot says so itself. One independent writer documented a human-written piece coming back partly AI-generated, alongside an AI draft with about 40% rewriting passing as human. Short inputs and paraphrased text are the reported weak spots.

Does QuillBot’s detector check for plagiarism too? Its AI detector and its plagiarism checker are separate products with separate outputs, so an AI likelihood score is not a similarity report. If someone tells you QuillBot flagged your work, ask which of the two produced the flag before you plan a response.

Last updated August 8, 2026. Every figure above was pulled on that date from sources that can move underneath it: a product page, a help-center article that currently contradicts that product page, a blog post from the same company, and a benchmark leaderboard whose data is public but whose rendering is not stable to my tools. A monthly promise would be the wrong commitment for a page built on live artifacts like those, so here is the one I will keep instead: when any of those sources changes what it says, this page gets re-pulled, re-dated, and rewritten to match, and if a figure here stops being true, it gets corrected in place rather than quietly surviving a refresh. I am Fırat Mıhcı, I build AI writing tools, and I study how detection systems handle second-language writing (ResearchGate). MeteGPT, which I run, sells a detector and a humanizer, which is precisely why every claim above is tied to a source and a date you can open yourself.

Two free tools on MeteGPT

Humanize a draft, then check the score yourself.

MeteGPT keeps a humanizer and an independent AI detector on one screen, so you can rewrite an AI-flagged passage and read a detector score on the result before anyone else does. Free daily runs, no signup.