HomePangram AI Detector

Pangram AI Detector: Two Questions the Marketing Doesn’t Answer

Fırat Mıhcı, writing. I build MeteGPT, and I also study how AI-detection systems read the work of people who learned English later in life. Pangram’s detector is the name that keeps getting pushed as the one you cannot beat, so I wanted a single record that puts its own numbers next to the outside studies, marks which figures are the company’s marketing and which come from independent researchers, and says plainly where the evidence stops and the marketing takes over. My ResearchGate profile is here. Published July 14, 2026.

TL;DR: A Pangram flag is not proof, and this page is about the two things its marketing skips. On clean, medium-to-long text, one independent University of Chicago Booth study did measure its false-positive rate near zero, and Pangram’s own claim is 1-in-10,000, a vendor figure, not an audited one. But hybrid, lightly-AI-edited writing is a documented soft spot, one adversarial finding could not be verified, and MeteGPT has never run its own humanizer against Pangram’s detector, so this page makes no claim either way.

How this page was built (MeteGPT Evidence Protocol v1.0; the full method is written up here). Everything below is drawn from Pangram Labs’ own live pages, the University of Chicago Booth research review and its underlying NBER working paper, one peer-reviewed Stanford study, a 2026 preprint from researchers at Imperial College London and Stanford, an independent nonprofit’s audit blog, one independent hands-on test, and a tech-news aggregation, with a university teaching-center statement for context. The dated record runs from a July 2023 peer-reviewed study to Pangram’s live pricing page pulled on July 14, 2026. This sweep started from 93 possible sources found through the logged queries and screened every one; 74 were set aside (27 that would not verify live this run, 27 duplicates, 11 off-topic, and 9 affiliate or no-method review posts), which left 19 that qualified. Two coder agents labelled each kept source independently across five fields; their 65 labels matched on 63, a 96.9% agreement rate, with the two splits, both on how to classify a single source’s claim type, resolved on a third pass. The full public evidence record sits at github.com/metegpt/metegpt-evidence.

One disclosure before any number.

MeteGPT ships a humanizer as well as a detector, so I have a direct commercial stake in this exact market from two sides. That is the reason this page never claims our tool beats Pangram’s detector, and the reason every figure below carries its source and a tier: a company’s own marketing claim, an independently corroborated study, or a figure a founder gave to a single journalist. Those three are not the same weight, and I mark which is which every time. That kind of conflict-of-interest disclosure is the one thing the closest write-up to this one, published on a rival humanizer’s blog, left out entirely.

How Does Pangram AI Detector Work?

Pangram’s detector, made by Pangram Labs at pangram.com, is a trained AI-text classifier. You give it a block of writing and it returns a judgment about whether a machine produced it, and increasingly about how much of a machine’s hand is in it. It is a different design from the older, perplexity-based detectors that score how statistically predictable your word choices look: those tools measure a texture, while a trained classifier learns the difference between human and AI writing from large sets of labeled examples. I am not going to re-derive Pangram Labs’ internal method here, because the useful facts for a reader are what it outputs and how it performs, and both of those are documented.

On performance, the placement is not just the company’s own word. In the University of Chicago Booth and NBER research covered below, Pangram’s detector sat in the top tier of the commercial tools tested, well clear of an open-source baseline that misclassified human text at rates between 30% and 78% ( the study, via the Becker Friedman Institute; EV-detect-03). A separate 2026 preprint from researchers at Imperial College London, the Internet Archive and Stanford, who needed a detector for an unrelated study of AI text online, compared four options and picked Pangram’s version 3 as the most consistent across text length and multiple languages ( arXiv preprint; EV-pangram-ai-detector-09). Neither of those is Pangram scoring its own homework, which is the only reason its accuracy claim gets weighed here at all rather than taken on the company’s word, and note what both measured: long, clean text, not the hybrid, short, and non-native cases where the record gets shakier below.

What Does a Pangram AI Detector Score Mean?

Here is the detail most reviews skip, and it is the one that decides whether a flag feels fair. Pangram’s detector does not just return a flat pass or fail. Its EditLens output places a passage on a four-tier scale: Fully Human-Written, Lightly AI-Assisted, Moderately AI-Assisted, and Fully AI-Generated ( Pangram’s own documentation, December 2025; EV-pangram-ai-detector-01). In Pangram Labs’ own words, a “light” reading means surface-level changes like grammar fixes or rephrasing that do not touch the underlying ideas, while a “moderate” reading means the AI rewrote significant portions or added content of its own (EV-pangram-ai-detector-01).

That distinction between AI-generated and AI-assisted is not a technicality; it is where fairness lives. A student who ran a grammar checker over their own paragraph and a paper drafted wholesale by a chatbot are different situations, and a four-tier score is built to tell them apart. The failure mode is human, not technical: a “Lightly AI-Assisted” reading gets reported up the chain, or screenshotted, as if it were “Fully AI-Generated,” and a low false-positive tool still ends up attached to a damning-feeling accusation. If a Pangram result ever lands on you, the tier it actually returned is the first thing to read, and the first thing to ask about, because those two tiers should not carry the same consequences.

Is Pangram AI Detector Accurate?

Start with the company’s own headline, labeled for what it is. Pangram Labs’ comparison blog states a false-positive rate of 0.01%, which it frames as 1-in-10,000 overall and 1-in-25,000 on academic essays, alongside 99.9% recall on GPT-4-generated text ( Pangram’s comparison post; EV-pangram-ai-detector-06). A tech-news aggregation of a May 2026 Atlantic investigation attributes the same 1-in-10,000 false-positive figure to Pangram’s founder ( Techmeme’s pull-quote of the report; EV-pangram-ai-detector-08). Read both as the same claim reaching you through two routes: a vendor’s marketing and a figure a founder gave a journalist. Neither is an independent measurement, and I hold Pangram’s numbers to the standard I would hold my own to.

The one place the independent record does back Pangram is narrow, and it happens to be the case that matters most to a wrongly-accused writer: false positives on clean, medium-to-long human text. The strongest single source is a University of Chicago Booth research review of a study by Brian Jabarian and Alex Imas. I read the primary source directly rather than repeat the paraphrase that circulates: it tested roughly 2,000 human-written passages across four length bands, from under 50 words to about 1,000, and found Pangram’s accuracy never dropped below 99.8%, with false positives “essentially 0 across most decision thresholds” ( Chicago Booth Review; EV-pangram-ai-detector-02). That is a materially different, and more precise, description than the widely-quoted secondary version of “some 3,000 texts of 500 to 1,000 words,” which is why I am citing what the primary source actually says.

Two honest limits sit inside even that strong result. First, the same research is explicit that accuracy for every detector, Pangram included, degrades on short passages, with texts under roughly 50 words the hardest case for all of them ( the length finding; EV-detect-04). A short quote or a one-line answer is a noisy signal for any tool of this kind. Second, near-zero false positives on clean, machine-generated-versus-human text is not the same as near-zero false positives on hybrid writing, where a human wrote most of it and a machine touched part of it. That second case is the one where the record gets genuinely unsettled, and it has its own section below.

Does Pangram’s Detector Catch Humanized or Paraphrased Text?

This is the question that made me want to write the page, and it is the one where I have to be most careful, because I sell a humanizer and the honest answer is not the one that would sell it.

Pangram Labs’ own position is confident. An August 2025 post claims its detector scored 97% accurate on humanized text against a cited external study, and reports it performing above 90% across a self-run benchmark of 19 named humanizer tools ( Pangram’s humanizer benchmark; EV-pangram-ai-detector-07). That is a vendor’s own counter-claim to the worry that rewording gets past it, and I count it as exactly that: self-interested, not independently reproduced.

Against the company’s confidence sits a more mixed independent record on the harder, more common case, which is not fully machine-written text but hybrid text. In one documented test, a writer took a human-written published article, ran it through an AI “refine” step, and put the result into Pangram’s detector: it returned a 99.3% AI score and flagged phrases and lines that had originated in the author’s own human draft ( CompleteAITraining’s test; EV-pangram-ai-detector-04). That is one case, not a rate, but it points the same way as Wiki Education’s audit, which credited Pangram highly and still disclosed that it “struggled with false positives” on bibliography and citation formatting and on outline-style bullet lists ( Wiki Education’s report; EV-pangram-ai-detector-11). Read those two together and the picture is honest rather than tidy: on clean long-form text Pangram’s detector is very good, and on lightly-AI-touched or unusually-formatted human text it is not a solved problem.

What The Atlantic Reported, and What This Page Cannot Confirm

There is a widely-cited part of this story I have to hold at arm’s length. The Atlantic published a May 2026 investigation into Pangram’s detector, and a much-repeated passage from it describes a named humanizer whose output was consistently scored as human-written, alongside a false-negative figure of roughly 1-in-70 attributed to Pangram’s founder. I could not open and independently verify The Atlantic’s article through any route available to me this run, so I am not going to state either the humanizer result or the 1-in-70 figure as established fact on this page. I flag them as reported but not independently verified here. The only piece of that reporting I could corroborate through a secondary aggregation is the 1-in-10,000 false-positive attribution already cited above (EV-pangram-ai-detector-08). If you see that Atlantic finding quoted elsewhere as settled, treat it the way I do: a serious report worth reading, not a number I have confirmed.

What MeteGPT Has, and Has Not, Tested Against Pangram

Now the part where a humanizer company is supposed to claim a win, and will not. MeteGPT runs its own detector benchmark, and that benchmark did not include Pangram’s detector. I have not run our humanizer’s output through Pangram, under any method, on any sample. So I make zero claim, in either direction, about whether text from our tool would be scored as human-written or flagged by Pangram’s detector. Not a hint, not an implied “it holds up.” I would rather leave the most conversion-friendly sentence on this whole page unwritten than print a result I never measured. If you want to see how a rewrite tool actually works before you judge any of this, our review of the humanizer field is candid that no tool guarantees anything against any detector, and it makes the same refusal to invent numbers.

Is Pangram AI Detector Biased Against Non-Native English Writers?

For a writer who learned English as a second language and has already watched honest work get flagged, the warning that AI detectors are biased against prose like theirs is familiar, and as a broad pattern it holds up. A peer-reviewed Stanford-led study ran seven commercial detectors over real TOEFL essays from non-native English speakers and found they wrongly tagged, on average, 61.3% of that genuine human writing as AI, while scoring native-English essays correctly almost every time ( Liang et al., Patterns / Cell Press; EV-best-ai-humanizer-01). That study anonymized the seven tools it tested and did not name Pangram, so 61.3% is a measure of the detector class, not a Pangram score.

On Pangram’s detector specifically, the one number available cuts the other way, and it is the company’s own. Pangram Labs published an ESL audit reporting a 0.078% false-positive rate across 25,021 pooled non-native-English samples drawn from four datasets, ELLIPSE, ICNALE, PELIC and the Stanford TOEFL set, with each dataset scoring between 0% and 0.019% ( Pangram’s ESL audit; EV-pangram-ai-detector-03). If that figure holds, it runs sharply counter to the general pattern the Stanford study documented. But the caveat is the whole point: it is a vendor-reported result on a benchmark the vendor chose, not an independently reproduced study, and I present it as a claim pointing in a hopeful direction, not as proof that Pangram’s detector is free of the bias the rest of the category has shown. If you write in English as a later language and a detector flags you, your saved drafts and revision history remain the strongest answer to any single score, whichever tool produced it.

How Does Pangram AI Detector Compare to Turnitin and GPTZero?

People searching “Pangram vs Turnitin” or “Pangram vs GPTZero” want a leaderboard. I can’t honestly hand them one: nothing in the evidence I collected runs all three detectors over a single shared set of passages on one day, so a genuine head-to-head simply does not exist yet. What the table below does instead is stack each tool’s dated results next to each other, tagged with whoever generated the number. Every row meets the same standard, and our own detector would carry a self-reported tag here too, since I have never scored it against these tools directly.

DetectorBest independent evidence in this sweepThe vendor’s own claimSource(s)
Pangram’s detectorEssentially zero false positives on medium and long text; ~2,000-passage academic test, accuracy never below 99.8% (Chicago Booth / NBER)0.01% false-positive rate (1-in-10,000); 99.9% recall on GPT-4 textEV-02, EV-detect-03 / EV-06
TurnitinVanderbilt University disabled Turnitin’s AI detector in 2023 over false-positive risk; no independent false-positive rate surfaced this runCited by Pangram (not by Turnitin here) at 0.51%, roughly 1-in-200, with 76.8% GPT-4 recallEV-turnitin-01 / EV-06
GPTZero“Secondary tier” among commercial detectors, tuned to minimize false positives (Chicago Booth / NBER)See our separate GPTZero record for its self-reported figureEV-detect-05 / EV-detect-06
Originality.ai“Secondary tier,” tuned instead to maximize AI detection (Chicago Booth / NBER)Not covered on this pageEV-detect-05 / EV-detect-06
Open-source (RoBERTa baseline)Misclassified human text at 30% to 78% depending on scenarioOpen-source, no vendor claimEV-detect-03

Read the table as evidence at different dates under different methods, not a live shoot-out. The Turnitin figures in the third column are Pangram’s characterization of a competitor, which is why they carry Pangram’s name, not Turnitin’s; the Chicago Booth study’s own framing is that GPTZero and Originality.ai occupy a secondary tier below Pangram, with a real trade-off between them, GPTZero favoring fewer false positives and Originality.ai favoring more detection (EV-detect-05). If your actual worry is the system your school runs, that is usually Turnitin, for which I maintain its own dated Turnitin AI-checker record assembled on the same method as this page; for the tool most flagged students meet first, the same kind of evidence file on GPTZero’s detector is next door.

Pangram’s detector also ships beyond a paste box: it runs inside Canvas for instructors who use it there, which is how at least one named faculty user describes working with it (EV-pangram-ai-detector-10), and the same 2026 preprint that praised its consistency did so partly on the strength of its multilingual handling (EV-pangram-ai-detector-09). Those are integration and coverage facts, not accuracy signals, and I keep them here rather than dress them up as performance.

Which Colleges Use Pangram AI Detector?

This is where a punchy claim is available and I am going to decline it. You will see “12+ universities switched from Turnitin to Pangram” repeated around this topic. I went looking for a source and could not find one, so I will not print it as fact.

Verified: the institutions actually documented

The named, dated record is narrower and worth stating precisely. Wellesley College is piloting Pangram’s detector, which is a Provost-office-approved trial run in consultation with the college’s advisory and AI working committees, explicitly a pilot and not a full institutional switch ( Pangram’s press page, citing The Wellesley News; EV-pangram-ai-detector-12). Delaware County Community College is a named, active user through a faculty member, Susan Ray, an Associate Professor of English who runs it inside her Canvas courses ( Pangram’s homepage testimonial; EV-pangram-ai-detector-10). Wiki Education, the nonprofit, is a documented user that ran thousands of its articles through the tool and wrote up the results ( Wiki Education’s blog; EV-pangram-ai-detector-11). Separately, researchers affiliated with Imperial College London and Stanford chose Pangram’s detector for a study, which is a research selection, not an academic-integrity deployment at those schools (EV-pangram-ai-detector-09). Each of those is a specific, checkable fact, and none of them adds up to an aggregate count.

Unverified: the “12+ universities switched” claim

The aggregate stays out because no source supports it. An admissions-consulting blog that actively advocates for Pangram over Turnitin, and would have every reason to cite a switch count, names none, and its own strongest example is only that Vanderbilt University disabled Turnitin’s AI detector, with no replacement tool named ( GradPilot’s article; EV-pangram-ai-detector-13). That Vanderbilt decision is real and better-sourced than any “switch” story: the university turned off Turnitin’s AI detector in August 2023, citing the false-positive risk against its own submission volume plus non-native-speaker bias and a lack of transparency ( Vanderbilt’s own statement; EV-turnitin-01). Note what it is and is not: a school stepping away from one detector is not the same event as a school adopting Pangram, and the honest version keeps those two stories separate.

Pangram AI Detector Pricing: Is There a Free Plan?

Yes, there is a free plan, and the paid prices are worth pinning down because older reviews disagree on them. On Pangram’s live pricing page, pulled July 14, 2026, the free tier gives 4 credits per day with no payment method required; the Individual paid plan is $20 per month for 600 credits (annual billing saves $60); and a Professional tier is $65 per month for 3,000 credits ( Pangram’s pricing page; EV-pangram-ai-detector-05). If you have seen $15 per month quoted for the 600-credit tier, that figure is stale; the current listed price is $20. One honest caveat: a vendor’s pricing page shows no visible last-updated stamp and may change at any time without notice, so confirm the live number at pangram.com before you plan around it rather than trusting any review, this one included.

For contrast, and to keep my own stake in view: MeteGPT’s detector runs free of charge and account-free, and rather than selling around its ceiling I will state it outright. An anonymous free check caps at 125 words, which makes it a quick read on a short passage rather than a whole-essay scan. That solves a different-sized problem than Pangram’s daily-credit free tier, and I am not claiming ours measures better, because I have run no comparison between them. What our free run is for, and where the paid limits sit, is spelled out on our own pricing and free-tier page.

Limitations of this page, stated plainly rather than buried.

  • Pangram’s false-positive rate (1-in-10,000), its false-negative rate, and its 0.078% ESL figure all trace back to Pangram Labs’ own pages as vendor claims; no outside party has audited any of them, so this page interrogates where each number came from instead of vouching for it.
  • The most load-bearing adversarial finding people associate with this topic, The Atlantic’s report of a humanizer’s output going unflagged and a roughly 1-in-70 false-negative figure, could not be independently verified through any route this run, so it is flagged as reported-but-unconfirmed and is not treated as fact anywhere above.
  • The University of Chicago Booth result is strong and independent, but re-verify it from chicagobooth.edu directly if you are relying on it for a decision; secondary paraphrases of it already circulate with the sample described incorrectly.
  • No Reddit or Quora thread about Pangram’s detector could be found and verified live this run, an unusual gap for this category, so this page makes no “users online report” claim at all.
  • Pangram’s pricing is a capture-date fact from July 14, 2026, and can change without notice; confirm the live figures at pangram.com.

Should You Trust a Pangram AI Detector Flag? The Verdict

Should a Pangram flag be trusted? Never on its own, whatever its record shows, and the two questions the marketing does not answer are the reason. The one thing the independent research does settle is narrow: on false positives on clean, medium-to-long human text, the Chicago Booth and NBER work measured that rate near zero (EV-pangram-ai-detector-02, EV-detect-03). That is the number a wrongly-accused writer most needs, but it is a measurement on one kind of text on one day, not a verdict on your paper. First, hybrid and lightly-AI-edited writing is a documented soft spot, not a solved case, and a single high AI score on part-human text should start a conversation, not end one (EV-pangram-ai-detector-04, EV-pangram-ai-detector-11). Second, the strongest reported evidence that a humanizer’s output can go unflagged sits behind a source I could not verify, so it stays an open question rather than a settled fact.

So use a Pangram result the way its own evidence justifies: as a strong signal that still needs a human to read the tier it actually returned, weigh the length and type of text, and ask before acting. Before any of that hardens into something official, it costs nothing to get a second opinion from a different engine: you can paste the passage into the free detector we run, a page that is equally frank about the limits I have laid out here.

A conflict to weigh against everything above.

Treat me as an interested party, because I am one. MeteGPT, the product I operate, earns on both ends: it sells a detector and a humanizer, so a reader who leans either way still sends business my direction. That is the reason this write-up won’t assert that our detector outperforms Pangram’s, and the reason I left the one sentence that would flatter me most off the page entirely. Here is the part I will stand behind, because you can check it: the page lines Pangram’s own marketing up against the independent studies, stamps each number with its author and its date, fixes a paraphrase that had misquoted the primary source, and points at the spots where the evidence simply stops. The procedure behind all of that is written out in the protocol. And if your actual goal was to reshape an AI draft into your own prose instead of auditing one, that is a separate task; MeteGPT runs both halves, letting you humanize a passage and then check a score on what comes out.

Common Questions

Is Pangram’s detector accurate? Narrowed to how often it wrongly flags clean, medium-to-long human writing, one independent University of Chicago Booth study measured its rate near zero (EV-pangram-ai-detector-02). Its own 1-in-10,000 false-positive figure is a vendor claim, not an audited one (EV-pangram-ai-detector-06), and neither number settles how it reads hybrid, short, or non-native-English writing, where the record is weaker or unverified.

Does Pangram catch humanized or paraphrased text? Pangram claims high accuracy against humanizers (EV-pangram-ai-detector-07), but the independent record on hybrid, lightly-AI-edited text is unsettled, with documented false-positive and near-total-flag cases (EV-pangram-ai-detector-04, EV-pangram-ai-detector-11). MeteGPT has run no test of its own tool against Pangram, so this page makes no claim either way.

Is Pangram free? Yes, the free tier is 4 credits per day with no payment method required; paid plans start at $20 per month for 600 credits, as listed July 14, 2026 (EV-pangram-ai-detector-05). Verify the live price at pangram.com.

Is Pangram biased against non-native English writers? In the peer-reviewed study most often cited here, AI detectors as a category flagged 61.3% of genuine non-native TOEFL essays as machine-written, a figure measured across the whole class of tools rather than on Pangram by name (EV-best-ai-humanizer-01). Pangram’s in-house ESL audit puts its own rate far lower at 0.078%, though that is a vendor number on a benchmark it selected rather than an independently reproduced result (EV-pangram-ai-detector-03).

Do universities use Pangram? Some do in documented, named ways: Wellesley is piloting it, and Delaware County Community College and Wiki Education are named users (EV-pangram-ai-detector-12, EV-pangram-ai-detector-10, EV-pangram-ai-detector-11). No source names any school that switched from Turnitin to Pangram (EV-pangram-ai-detector-13).

Last revised July 14, 2026. I treat this page as an open ledger: a newer dated source gets folded in as it surfaces, and any figure that stops surviving scrutiny gets rewritten instead of quietly propped up. Author: Fırat Mıhcı, who builds writing tools and studies how detection models read the work of second-language authors ( ResearchGate). To say it once more without hedging: MeteGPT, which I run, ships a detector and a humanizer alike, and because I profit either way, every figure above comes tagged with a source and a date so you can audit the work rather than trust it on faith.

Two free tools on MeteGPT

Humanize a draft, then check the score yourself.

MeteGPT keeps a humanizer and an independent AI detector on one screen, so you can rewrite an AI-flagged passage and read a detector score on the result before anyone else does. Free daily runs, no signup.