Two things people expect on a page like this are missing on purpose, so I will flag them once here. First, a pair of paid-tier prices ($14.99 and roughly $23.99 per month) circulates across review posts, but this run could not trace either figure to a single fetchable source, and they conflicted with the numbers I did pull, so I left them out (the pricing section says what is and is not vendor-verified). Second, no Reddit or Quora thread about GPTZero could be fetched and verified live this run, so this file makes no “people online report” claim at all. What follows is only what I could open, read, and date.
What Is GPTZero?
GPTZero is an AI-writing detector at gptzero.me: you paste in a block of text and it returns an estimate of how likely a machine wrote it, expressed as a probability rather than a plain yes or no. It was launched in early 2023 by Edward Tian and grew into one of the most-recognized names in the category, which is exactly why a GPTZero result carries weight it may not have earned. GPTZero also ships a Chrome extension and a “Writing Replay” feature that reconstructs a document’s editing history, and it markets itself heavily to schools. Everything on this page is about that one product, GPTZero’s detector, not about writing tools in general.
One distinction is worth fixing before anyone acts on a reading. A GPTZero result of “80% AI” is not a claim that 80% of your words are machine-written; it is GPTZero’s stated confidence that the passage as a whole is AI-generated. That gap matters the moment a score gets used against a student, because a confidence level and a proportion are not the same measurement, and they invite very different reactions.
GPTZero’s Superhuman Acquisition (June 2026)
If you are reading an older review of GPTZero, note that the company changed hands recently. GPTZero was acquired by Superhuman, announced June 23, 2026; deal terms were not disclosed, though GPTZero’s founder said the company had passed 19 million registered users and $30 million in annual recurring revenue before the sale, and GPTZero continues to run as a standalone product ( TechCrunch on the acquisition; EV-gptzero-ai-detector-09). Superhuman already operated its own detection feature and framed the deal, in TechCrunch’s words, as “two AI detectors are better than one.” That is an ownership fact and a currency marker for this page, not a signal about accuracy one way or the other.
How Does GPTZero Detect AI Writing?
GPTZero reads two statistical properties of your text. The first is perplexity, roughly a measure of how surprising each next word is to a language model: writing that a model finds very predictable scores as more machine-like. The second is burstiness, a measure of how much sentence length and rhythm vary across a passage: humans tend to mix long and short sentences unevenly, while machine drafts often run smoother and flatter. GPTZero’s newer models fold these into what it describes as a lexical-predictability approach, but the underlying idea is unchanged: the flatter and more predictable your writing looks to the model, the more AI-like GPTZero rates it.
The problem hidden in that mechanism is the one this whole page circles back to. Predictability is not the same thing as authorship. A careful, plain, heavily-revised human writer can produce low-perplexity, low-burstiness text, and a second-language writer using a smaller, more cautious vocabulary can too. GPTZero is scoring a texture, and some genuine human writing has exactly the texture it penalizes.
Can GPTZero Detect Humanized or Paraphrased Text?
This is where the published record gets specific, and I want to be careful about how I present it, because the point is reliability, not a workaround. In the NYU Abu Dhabi study covered below, researchers took AI-generated answers that GPTZero had already scored, ran them through a paraphrasing tool, and re-tested: GPTZero’s false-negative rate rose to 95%, meaning it missed almost all of the reworded machine text ( Ibrahim et al., Scientific Reports; EV-gptzero-ai-detector-04). Read that as a measured limitation of the tool, documented by academics, not as a method. What it establishes for a reader is the opposite of reassurance for anyone leaning on GPTZero as a verdict: a detector that can be thrown off this easily by ordinary rewording is not a system whose single reading should decide anything on its own.
Is GPTZero Accurate? Self-Reported vs the Outside Tests
Start with GPTZero’s own number, clearly labeled for what it is. In a February 2026 benchmarking post authored by GPTZero’s own ML lead and co-founder, GPTZero reports a 0.08% false-positive rate, 99.60% recall and 99.76% accuracy for its current 4.3b model, with a perfect 0.00% false-positive rate and 100% recall in the essays category specifically ( GPTZero’s benchmarking post; EV-gptzero-ai-detector-01). Here is the part the headline number leaves out: by the post’s own description, GPTZero selects the human texts from its own database, generates the matching AI versions, and scores itself, with no external or peer audit mentioned anywhere. That is a self-reported, self-graded figure. It is not independent, and it should be read the way you would read any company grading its own exam.
There is a reason self-reported detector numbers draw scrutiny by default. In July 2023, OpenAI quietly discontinued its own AI-text classifier, citing a “low rate of accuracy,” about six months after launching it ( TechCrunch; EV-gptzero-ai-detector-08). The company that builds the models could not make reliable detection work and said so. That is the industry backdrop against which GPTZero’s near-perfect self-score sits.
Now the independent record, which is smaller and older than most people assume but not empty. There are two peer-reviewed studies that tested GPTZero by name, and for a long stretch of these reviews only the first got cited:
- Habibzadeh, 2023 (Journal of Korean Medical Science). On a 50-text medical-writing sample, GPTZero scored 65% sensitivity, 90% specificity, 80% overall accuracy, a 10% false-positive rate and a 35% false-negative rate ( DOI 10.3346/jkms.2023.38.e319; EV-gptzero-ai-detector-02). The author is a past president of the World Association of Medical Editors, which makes this a high-credibility source, and it comes with an important caveat: it was published in September 2023, on medical text only, and GPTZero has updated its model several times since, so treat it as dated. The study also adds a Bayesian point worth carrying: at an assumed 20% rate of AI text in the pile, a GPTZero “positive” is only about 62% likely to be correct, which is why the authors frame GPTZero as better at raising a flag than at settling one.
- Ibrahim et al., 2023 (Scientific Reports, NYU Abu Dhabi). This is the second peer-reviewed test, and it matters because it means Habibzadeh is no longer the only one. Researchers ran GPTZero over 280 real exam questions across 32 university courses and measured an 18% false-positive rate and a 32% false-negative rate ( DOI 10.1038/s41598-023-38964-3; EV-gptzero-ai-detector-04). It is the same study whose paraphrase re-test pushed the false-negative rate to 95%.
Two more measured tests sit outside the peer-reviewed tier and both cut against the 99% story. A third-party review site, TwainGPT (January 2026, no byline), ran a disclosed 300-sample test and found a 29% false-positive rate on human writing, 17 of 100 human samples wrongly flagged ( TwainGPT’s test; EV-gptzero-ai-detector-06). And AI Busted (June 2026) reported GPTZero catching AI text 83% of the time and misfiring on human text 11% of the time, with only 68% of mixed human-and-AI samples resolved correctly ( AI Busted’s write-up; EV-gptzero-ai-detector-05). One caution on that last one: AI Busted sells its own competing detector and does not state that conflict of interest in the article, so I count it as a self-interested source, not a neutral test. Here is the whole record in one place.
| Test (date) | Who ran it | What it measured | False-positive rate | Source |
|---|---|---|---|---|
| GPTZero self-benchmark (Feb 2026) | GPTZero (4.3b), self-selected and self-scored | Its own essay / academic / review set | 0.08% (0.00% on essays) | EV-01 |
| Habibzadeh, J Korean Med Sci (Sep 2023) | Peer-reviewed; 50 medical texts | 20 AI + 30 human medical texts | 10% (35% false-negative) | EV-02 |
| Ibrahim et al., Sci Reports (Aug 2023) | Peer-reviewed; NYU Abu Dhabi | 280 real exam questions, 32 courses | 18% (32% false-negative) | EV-04 |
| TwainGPT (Jan 2026) | Third-party blog; 300 samples, no byline | 100 human writing samples | 29% (17 of 100 flagged) | EV-06 |
| AI Busted (Jun 2026) | Competing detector (self-interested); 200 samples | Human and mixed text | 11% (68% of mixed resolved) | EV-05 |
Each row measured different text under a different method, so the rows are not head-to-head, and only the top row is GPTZero grading itself. Read across the outside tests, GPTZero’s real-world numbers land between roughly 62% and 90% depending on what each study measured (from Habibzadeh’s 62% chance that a positive is truly AI, up to his 90% specificity on human text), never the 99%-plus GPTZero prints for itself. That gap, self-reported near-perfection against a documented 10-29% false-positive range, is the single thing to carry away from this page.
What Is a Good GPTZero Score, and What If GPTZero Flagged Your Essay?
There is no universal “safe” GPTZero score, and any page that hands you one is inventing it. GPTZero returns a probability, and what a given number means depends entirely on who is reading it and what threshold they apply. So the more useful question is what a single flag actually proves, and the honest answer is: on its own, not much. A detector that misfires on 10 to 29 percent of human writing in outside tests will, by simple arithmetic, flag real students, and one documented case shows exactly how far that can go. In a federal lawsuit filed in February 2025, an EMBA student alleged that Yale School of Management flagged his final exam using GPTZero and suspended him for a year; his complaint states that when decades-old, human-published academic writing (including work by a former Yale president) was run through GPTZero, it returned a “100% probability” AI verdict on text no machine could have written ( Poets&Quants coverage of the filed complaint; EV-gptzero-ai-detector-10). Read that carefully for what it is: one documented, dated, court-filed incident, with the allegations as filed and not adjudicated. A federal judge denied the student’s request for a preliminary injunction in May 2025, so the suspension stands and the false-positive allegation has not been confirmed in court. It is an illustration of the failure mode, not a proven rate.
If GPTZero flagged your essay and you wrote it, the productive move is documentation, not panic: keep your drafts, outlines, notes, and version history, because the record of how a piece took shape is the thing a probability score cannot produce and cannot argue with. A single flag is a prompt to show your work, not a confession. The same defense travels beyond the classroom. If you are a freelancer and a client runs your delivered copy through GPTZero and it comes back flagged, your saved drafts and revision history protect your reputation the same way they protect a student’s grade, and the false-positive spread documented above is exactly what to point the client toward. If you want a second reading from a different engine before any of that becomes a conversation, you can paste a passage into the detector we run, free and without an account; it is candid about this same weakness on its own page.
Does GPTZero Have a Bias Against Non-Native English Writers?
If you learned English as a second language and your honest work got flagged, this section is the one written for you. There is a well-documented pattern here, and it is measured at the level of the detector class rather than GPTZero alone. A Stanford-led team tested seven commercial GPT detectors against genuine TOEFL essays from writers whose first language was not English, and on average the tools misread 61.3% of that real human writing as AI, even as they cleared native-speaker essays almost every time ( Liang et al., Patterns / Cell Press, DOI 10.1016/j.patter.2023.100779; EV-gptzero-ai-detector-03). One honesty note keeps this accurate: that study anonymized the seven detectors it tested and did not name GPTZero, so I am not presenting 61.3% as a GPTZero score, because it is not one. What it establishes is a measured baseline for the entire family of predictability-scoring detectors that GPTZero belongs to, and the mechanism section above explains why second-language writing, often built from a more cautious and predictable vocabulary, lands squarely in the texture these tools misread.
The Yale complaint reaches for the same pattern from the real-world side: alongside the specific incident, it alleges that GPTZero’s high false-positive rate disproportionately affects non-native English speakers (EV-gptzero-ai-detector-10). That is an allegation in a filing, not a finding, so I weight it as such, but it points in the same direction as the Stanford measurement. The practical takeaway is not despair; it is evidence. If English is a language you came to later and your own work was flagged, that result sits inside a bias researchers have already published, and your saved drafts and revision history are the strongest counter to a number that was never designed to account for how you write.
Is GPTZero Free? Pricing and Plans
GPTZero does run a genuine free tier, and it is the only pricing figure I can confirm straight from GPTZero. The free education plan lets you scan up to 10,000 words per month and use the Chrome extension ( GPTZero’s students page; EV-gptzero-ai-detector-07). GPTZero’s help center adds that across its plans, usage over your allowance is billed at $0.00015 per word, and that an account can run up to roughly one million words past its plan before being asked to upgrade (EV-gptzero-ai-detector-07).
On the paid tiers I have to be straight about a gap: GPTZero’s own pricing page is rendered client-side and returned no legible dollar figures on two direct attempts this run, so I cannot state paid prices as GPTZero’s confirmed numbers. A third-party SaaS-pricing aggregator, spotsaas.com, lists GPTZero’s paid plans as Essential at $15/month (150,000 words), Premium at $24/month (300,000 words) and Professional at $46/month (500,000 words), with a discount for annual billing ( spotsaas listing; EV-gptzero-ai-detector-11). Treat those as an aggregator’s figures that may have changed; confirm the live numbers at gptzero.me/pricing before you rely on them. I would rather tell you the source is second-hand than dress an unverified price up as fact.
For contrast on the free side, and in the interest of disclosing my own stake: the detector we run at MeteGPT is free with no account, and I will name its ceiling rather than sell around it. The anonymous free run caps each check at 125 words, which makes it a spot-read on a short passage, not a whole-essay scan. GPTZero’s 10,000-words-a-month free tier and our 125-word free run solve different-sized problems, and I make no claim that our detector reads more accurately than GPTZero’s: no head-to-head test has been run on my side, and I will not publish a figure I did not measure. Where our tool goes further is written up plainly on our pricing and free-tier page.
How Does GPTZero Compare to Turnitin and Originality?
People search “GPTZero vs Turnitin” and “GPTZero vs Originality” expecting a winner, but the more useful split is who each tool is built for, because they are aimed at different buyers and no honest single accuracy ranking exists across them in the record I gathered. GPTZero is a self-serve consumer product: anyone can open it, paste text, and read a score themselves. Turnitin is an institutional product that lives inside a school’s license, generating its AI writing score for instructors and administrators rather than for the student, who usually cannot run it at all. Originality.ai is a paid commercial detector marketed mainly at web publishers and content agencies checking their own or freelancers’ work. That difference in buyer is the real fork: a clean GPTZero result is a self-run signal, not the reading your institution will actually see.
I am deliberately not printing a GPTZero-beats-Turnitin accuracy number, because nothing in the sources I verified backs a specific head-to-head figure, and inventing one would defeat the purpose of this page. What the broader record does support is a single posture across all three: every one of these detectors misfires often enough on human writing, and second-language writing in particular, that a lone score from any of them is one dated signal rather than a verdict. If your real worry is what a university’s system will report, GPTZero is not that system, and I keep a separate dated record of Turnitin’s AI checker built the same way as this one for exactly that question.
Is GPTZero the Same as ZeroGPT?
No, and the mix-up is common enough to be worth one plain paragraph here. GPTZero (gptzero.me), the tool this page is about, is a different company and a different product from ZeroGPT (zerogpt.com), a separate detector whose name is GPTZero’s two syllables in reverse. The two are routinely confused: on the bare search “zerogpt,” GPTZero itself surfaces near the top of the first page even though it is the other product, and GPTZero has publicly written that ZeroGPT “popped up on January 18th, 2023, capitalizing on brand confusion” (EV-zerogpt-ai-detector-05), which I pass along as GPTZero’s own characterization of a rival rather than a settled fact. Their accuracy records, their pricing, and their false-positive histories are not shared, so a review or a result about one tells you nothing reliable about the other.
If the tool that actually flagged you was the one at zerogpt.com, this GPTZero page is not the right record for it. I keep ZeroGPT, the copycat-named detector people confuse this with, on its own dated page built under the same evidence rules, so you can read the tool you actually used instead of the one that merely shares its syllables.
Limitations
Every evidence page has soft spots; here are this one’s, named up front instead of buried.
- The 99.76% accuracy and 0.08% false-positive rate are GPTZero’s own, from a benchmark GPTZero built and scored itself; no outside audit of that particular figure has been published anywhere I could find, so this page traces where the number came from rather than confirming or disproving it.
- The two peer-reviewed studies (Habibzadeh and NYU Abu Dhabi) are both from 2023 and predate GPTZero’s current model by several generations; they are the best direct evidence available, but they are dated, and one is medical-text-only, so neither is the last word on the 4.3b model shipping today.
- The Stanford 61.3% figure is detector-class, not a GPTZero measurement: that study anonymized the seven tools it tested and named none of them, so it is used here strictly as context for the family GPTZero belongs to.
- The TwainGPT and AI Busted tests are non-peer-reviewed, and AI Busted sells a competing detector, so both are presented as single dated tests with their stakes labeled, never as settled rates.
- No Reddit or Quora thread about GPTZero could be fetched and verified live this run, so this page makes no “users online report” claim; community anecdote, where it exists, tends to over-represent unhappy outcomes anyway, and a sample of complaints is not a population rate.
- Every vendor figure here, the free-tier word cap and the overage rate, is a capture-date fact stamped July 12, 2026, and GPTZero can change any of it without notice; the paid prices are an aggregator’s, not GPTZero’s confirmed numbers.
Does MeteGPT’s Own Humanized Text Pass GPTZero?
You are entitled to know how the tool I build performs against the detector this whole page is about, so here is the figure, wrapped in the caveats that give it meaning. On a small, controlled in-house run of roughly 30 academic-style passages in mid-May 2026, GPTZero scored our humanized output at 4% AI. Read that precisely: 4% is how much of our rewritten text GPTZero flagged on that sample, on that day, on the model versions then live. It is not a grade on GPTZero’s accuracy, and it is emphatically not a promise that your own submission will land at 4%.
| What was measured | MeteGPT’s humanized output, read by GPTZero’s detector |
| Result | 4% AI (how GPTZero rated our output, not a score of its accuracy) |
| Sample and date | About 30 academic-style passages, one controlled run, mid-May 2026 |
| Known spread | Unusual inputs (code, dense technical prose) can climb into a 30-60% AI band on strict detectors; one headline average conceals that range |
| Status | Our own measurement (FC-METEGPT-001), verifiable by the owner on request, not an independent audit |
The caveats are the substance here, not the disclaimer:
- The sample is small and controlled. Around 30 academic passages, identical settings across the run, on the detector build that was live in mid-May 2026. That makes it our own data, checkable by request, not an external audit and not an industry benchmark.
- Out-of-distribution text behaves differently. On atypical subjects, source code, or heavily technical writing, the same humanized text can score far higher, up into a 30-60% AI range on the strictest detectors. A single flattering average will not tell you that.
- A pass is not a guarantee. Detectors update, and one favorable run in May does not promise the same reading on your text next month. Treat 4% as one dated data point, confirm it with a second independent reading, and keep your drafting history either way.
Is GPTZero Worth Using? The Verdict
Would I trust a GPTZero result? It is a widely-used tool with a real free tier, and it did at least publish a named benchmark, which not every rival does. What the evidence will not support is the 99.76% headline. The two peer-reviewed tests that measured GPTZero directly found 10% and 18% false-positive rates; a disclosed third-party test found 29%; a competitor’s test found 11%; and the family of detectors GPTZero belongs to has a published, measured bias against second-language writers that GPTZero’s own marketing does not mention. So use GPTZero the way its own numbers actually justify: as one dated signal that can raise a question, never as a system whose single score should decide a grade, a suspension, or an accusation. If you are a student or a second-language writer whose honest work got flagged, that spread of false-positive numbers is your evidence, and your saved drafts are your defense.
One bias to weigh against everything above.
I am not neutral here: I build MeteGPT, which ships both a detector and a humanizer, so I profit from this exact market from two directions. That is precisely why this verdict refuses to claim our detector out-measures GPTZero’s, and why I decline to print any comparison number I have not measured myself under a disclosed method. The narrow thing this page will defend about itself is that it set GPTZero’s marketing against the outside studies, dated and linked every figure, and flagged the self-graded number for what it is, which is work most write-ups on this tool skip. The rule it follows is laid out in the protocol. If you came here actually wanting to rewrite an AI draft into your own register, that is a different job than checking one, and our dated review of the humanizer field is candid that no tool guarantees anything; MeteGPT itself pairs both sides, so you can humanize a passage and then read a score on the result.
Common Questions
Is GPTZero accurate? GPTZero self-reports 99.76% accuracy and a 0.08% false-positive rate on a benchmark it built and graded itself (EV-gptzero-ai-detector-01). The two peer-reviewed studies that tested GPTZero directly measured 10% and 18% false-positive rates instead (EV-gptzero-ai-detector-02, EV-gptzero-ai-detector-04). Read the self-reported figure as a claim, not a measurement.
Is GPTZero free? Yes, GPTZero runs a free education plan that scans up to 10,000 words per month and includes the Chrome extension (EV-gptzero-ai-detector-07). Paid tiers exist above that; the exact prices could not be verified from GPTZero directly this run, so confirm them at gptzero.me/pricing.
GPTZero flagged my essay, is that proof I used AI? No. GPTZero returns a probability, and outside tests put its false-positive rate on human writing between 10% and 29% (EV-gptzero-ai-detector-02, EV-gptzero-ai-detector-06). A single flag is a reason to show your drafts and revision history, not a confession.
Is GPTZero biased against non-native English speakers? The best evidence is detector-class: a Stanford-led study found commercial detectors flagged 61.3% of authentic non-native-English essays as AI, though it did not name GPTZero specifically (EV-gptzero-ai-detector-03). A filed lawsuit alleges the same pattern for GPTZero, as an allegation not a finding (EV-gptzero-ai-detector-10).
Who owns GPTZero now? GPTZero was acquired by Superhuman in June 2026 and continues to operate as a standalone detector (EV-gptzero-ai-detector-09).
Last revised July 12, 2026. I treat this as a live file: a newer dated study replaces an older one, and any figure that stops surviving a re-check gets corrected in place rather than left standing. I am Fırat Mıhcı, an applied-linguistics researcher who also builds AI-writing tools and studies where detectors misread people writing in a second language ( ResearchGate). The disclosure that governs all of it: MeteGPT, which I run, sells both a detector and a humanizer, so every number here is sourced and dated on purpose, for you to check rather than take on trust.
Humanize a draft, then check the score yourself.
MeteGPT keeps a humanizer and an independent AI detector on one screen, so you can rewrite an AI-flagged passage and read a detector score on the result before anyone else does. Free daily runs, no signup.