How AI detectors work is a question with a strange property: the standard answer you will read almost everywhere is one the biggest detector vendor stopped relying on in 2023. Most guides still present perplexity and burstiness as the whole mechanism; the vendor that made those two words famous states, on its own site, that it migrated to a different architecture years ago (EV-how-ai-detectors-work-05). This page covers what text detectors actually do now, and why the mechanism predicts its own failures, including flags on work people wrote themselves. To watch these signals fire on your own writing as you read, MeteGPT’s free scanner gives you a pattern-based reading of any paragraph, no signup. It reports a signal, not a verdict on your integrity, and that distinction is half of what this article is about.
How this page was built (MeteGPT Evidence Protocol v1.0; the full protocol is written up on our methodology page). Ten logged search queries ran across academic archives and the vendors’ own documentation, covering the ACL Anthology, Springer, Nature, arXiv, PubMed Central, and the help centers and blogs of the detector companies named below, plus two direct re-verification fetches. Counting the census: 84 candidate sources identified, all 84 screened, 74 excluded (33 duplicates, 18 off-topic, 22 explainer or review pages with no original evidence behind them, and 1 that would not verify live this run), 10 included. Each of the 10 backs one new dated record, cited in the text by its (EV-…) tag. Two older registry entries, the Stanford false-positive study and Turnitin’s below-20% display rule, were re-opened live the same day and extended to this page; the remaining tags reuse records already built for this site’s per-detector pages. No Reddit or Quora sweep ran this time, on purpose: this is a mechanism question, and the community reports on individual tools already sit in the records those pages cite. The dated record runs from January 2023 to vendor pages captured live on August 20, 2026. Every tag resolves to a dated entry in the public evidence log.
How Do AI Detectors Work?
An AI detector is a statistical model that reads text and outputs a probability of machine authorship. It cannot know who typed a document. What it can do follows from how language models write: an LLM produces text by repeatedly choosing a likely next word. A detector inverts that, asking how closely your wording tracks what a model would have chosen, and converts the answer into a score.
In its own explainer, one major detector vendor lists the technique families the industry draws on: pattern recognition across word use, sentence structure, and meaning; perplexity; burstiness; trained classifiers; and vector embeddings, which represent words and phrases as numbers a model can compare (GPTZero’s explainer, October 14, 2024; EV-how-ai-detectors-work-06). The same page concedes the limits on the record: false results in both directions, English-centric training bias, weak short-text performance, and outputs that never amount to certainty.
A detector score is therefore a thresholded estimate, not an eyewitness account: one vendor’s technical write-up describes the final step as a probability that gets called fake once it crosses a chosen cutoff (EV-how-ai-detectors-work-09). The estimate also needs enough text to stand on: the same vendor puts its reliability floor at roughly 50 tokens of input, naming short text as its own weak point (EV-how-ai-detectors-work-09). Everything interesting about detectors lives in what that probability responds to and what it cannot tell apart. This page covers text detection only; image and video detection are different systems with different failure modes.
What Is Perplexity in AI Detection?
Perplexity measures how surprising each word in a text is to a language model; the vendor that popularized the term defines it as how likely an AI model would have chosen the exact same set of words found in the document (GPTZero’s perplexity and burstiness explainer, first published March 2023; EV-how-ai-detectors-work-05). Low perplexity means the model saw each next word coming; high means the wording kept surprising it.
A concrete pair makes the idea visible. “The results of the study indicate that the intervention was effective” is a sentence a language model completes almost word for word. “The intervention worked, which mostly irritated the committee that had funded its rival” contains turns no model would rank as the likely continuation. The first scores low perplexity, the second high, and a perplexity-based detector reads low as machine-like. To see what a per-word predictability score becomes inside a real product, take Turnitin’s launch-era account of its own pipeline: overlapping segments of about five to ten sentences, each sentence rated 0 to 1, rolled up into the document percentage (EV-how-ai-detectors-work-07).
The limit is built into the measure. Plenty of genuinely human writing is supposed to be predictable: methods sections, legal boilerplate, exam prose, the careful sentences of someone writing in a second language. Perplexity cannot distinguish “a machine chose these words” from “a person chose the safest words,” a gap that resurfaces in the false-positive section below because it is the mechanism working as built, not an edge case.
What Is Burstiness in AI Detection?
Burstiness measures variation: how much sentence length, rhythm, and per-sentence predictability swing across a passage (EV-how-ai-detectors-work-05). People drift between a long winding sentence and a blunt short one; model output, left unedited, tends to hold one cadence from first paragraph to last. A burstiness signal treats steadiness as evidence of machine authorship and variation as evidence of a person.
Read your own messages and you will see the fingerprint: fragments, an interruption, a sentence that runs on because you were thinking while typing. A corporate FAQ, by contrast, is human-written and nearly flat, and that is the problem: lab reports, technical documentation, and the structured five-paragraph essay taught in composition classes are low-burstiness on purpose. A detector leaning on rhythm variance reads disciplined formatting as a machine signature.
There is also a dated caveat that most explainers skip: perplexity and burstiness are no longer the core of the tool that made them famous. As of autumn 2023, that vendor states it moved to a deep-learning architecture in which the two classic signals survive only as one indicator among seven (EV-how-ai-detectors-work-05). Which raises the obvious question: what replaced them?
How Do Trained Classifier Detectors Work?
A trained classifier is a model taught the difference between human and AI text by example rather than by formula: engineers collect writing labeled human or machine, train a neural network to separate the piles, and let it find whatever regularities distinguish them, from word-choice distributions to discourse habits no one has named. This is the mechanism behind most detectors you will meet in 2026, and the least explained part of every popular guide.
The most architecture-specific public description from any commercial vendor is Originality.ai’s own technical explainer (updated October 18, 2025): a modified BERT-lineage model, pre-trained with an ELECTRA-style scheme over 160GB of text, fine-tuned on roughly one million samples labeled human-generated or AI-generated, producing a probability that a threshold turns into a verdict, reliable by the vendor’s own account only from about 50 tokens of input (Originality.ai’s how-detection-works post; EV-how-ai-detectors-work-09). That is the vendor describing itself, but the structural lesson holds: the verdict is only as good as the labeled pile the model learned from.
Turnitin’s launch-era FAQ (March 2023) is similarly mechanical: submissions get split into overlapping segments of roughly five to ten sentences, each sentence scored between 0 and 1, segment results aggregated into the document percentage. The same document states the model was trained at the time on GPT-3 and GPT-3.5 output, flags only at 98% confidence, accepts likely missing up to 15% of AI text to hold its stated false-positive rate under 1%, and warns that in short documents “the prediction will be mostly all or nothing” (Turnitin’s institutional FAQ PDF, March 2023; EV-how-ai-detectors-work-07). Treat that as a dated snapshot, not current architecture: Turnitin’s live documentation no longer restates these internals. Its current FAQ specifies a 300-word prose minimum, a 30,000-word maximum, exclusion of code and bullets, coverage of long-form English, Spanish, and Japanese, and a percentage that means the share of qualifying text likely AI-generated or AI-paraphrased, not the share of the whole file (Turnitin’s current capabilities FAQ, captured August 20, 2026; EV-how-ai-detectors-work-08).
What Detectors Changed Recently
Classifier detectors are moving targets, which is why mechanism explainers rot. In January 2026, GPTZero added a third detection signal modeling how predictable each word is given its surrounding context, and refreshed its training data with newer model outputs, according to a dated review by a rival vendor (EV-best-ai-humanizer-10). The running record on that particular tool is here; the general lesson is that any un-dated claim about how a specific detector works should be treated as expired.
What Words Trigger AI Detection?
No published detector documentation names a list of words that trigger a flag, and that fact is worth more than the lists circulating online. What is documented is the technique class: detectors analyze word-use patterns and stylistic regularities as part of their feature set (EV-how-ai-detectors-work-06), and a classifier fine-tuned on a million labeled samples (EV-how-ai-detectors-work-09) absorbs whatever vocabulary differences separate its training piles. If chatbot output overuses certain connectives and ornamental verbs, those habits become weak evidence in the model’s weighting. That is stylometry: reading authorship from measurable style.
The popular version, that a specific word gives you away, does not survive contact with the architecture: a classifier scores a probability over the whole passage, so no single token flips a document-level verdict on its own (EV-how-ai-detectors-work-09). The honest summary is that word choice is a real input, published trigger-word lists are folklore, and the strongest lexical signal is not any one word but the sustained pattern of always choosing the expected one.
Do AI Models Watermark Their Text?
Watermarking is the one detection approach that does not guess after the fact. Google DeepMind’s SynthID-Text, published in Nature on October 23, 2024, embeds a statistical signature while the text is being generated, by subtly modifying how the model samples each next token; a matching detector later verifies the signature without access to the underlying model, and the system was evaluated live on roughly 20 million Gemini production responses with no meaningful quality difference reported (the SynthID-Text paper, Nature 2024; EV-how-ai-detectors-work-04). Those deployment figures come from vendor-affiliated authors; read them as the builder’s own account.
Mechanically this is a different species: a classifier scores arbitrary text and can be wrong about anyone, while a watermark check identifies text produced with the watermark switched on. That precision is also why it has not replaced classifiers at scale: the signature exists only in text from a cooperating generator, and no scheme covers the open ecosystem of models a writer can actually reach. It is not indestructible either; a University of Maryland study found recursive paraphrasing degrades watermark-based detection along with every other family it stress-tested (Sadasivan et al., TMLR; EV-how-ai-detectors-work-10). Watermarking answers “did this cooperating model produce this exact text,” a narrower question than the one your instructor is asking.
Can AI Detectors Detect Paraphrased AI Text?
Sometimes, and the mechanism explains both halves of that answer. Synonym-swapping paraphrase keeps each sentence’s skeleton, so most of what a classifier reads survives it: structure, rhythm, paragraph logic, the even distribution of safe choices. That is why detectors keep catching lightly reworded chatbot output, and why Turnitin explicitly claims to score text “likely generated and modified by an AI paraphraser” as part of its qualifying-text percentage (EV-how-ai-detectors-work-08).
The independent record shows the other half. The largest academic multi-tool test in education, Weber-Wulff and colleagues (2023, DOI 10.1007/s40979-023-00146-z), found roughly 50% of AI-generated texts that undergo obfuscation get misattributed to humans, in a study where all 14 tested tools scored below 80% accuracy and the authors’ verdict was “neither accurate nor reliable” (the Weber-Wulff study; EV-how-ai-detectors-work-02). At the theoretical end, the Maryland group argued that as human and model text distributions converge, even an optimal detector trends toward chance (EV-how-ai-detectors-work-10). That is a limit argument, not a promise that any rewriting tool wins; treat anyone quoting it as a guarantee with suspicion.
Where does an ordinary grammar checker fall in all this? Sentence-level fixes do not add the signals trained classifiers read. Correcting commas, articles, and agreement leaves your sentence skeletons, rhythm, and ordering exactly where you put them, so text you wrote stays human-shaped after a grammar pass, and an AI draft stays machine-shaped after one. The line sits where correction becomes regeneration: a rewrite feature that regenerates whole sentences is an AI paraphraser, and that, not hand-corrected grammar, is what Turnitin’s qualifying-text percentage claims to catch (EV-how-ai-detectors-work-08).
Here is what this means if your workflow is an AI draft you then revise: swapping synonyms attacks the layer detectors stopped relying on years ago. What moves the deeper signals is rewriting at the level of structure and rhythm, the way a person actually drafts: uneven sentences, reordered logic, word choices a model would not rank first. That is the design idea behind MeteGPT’s rewriter, trained on 2,590 real student essays rather than on model output; judge the difference on your own text rather than taking my word for it. If what you really want to know is what happens when people try to get around detectors entirely, that question has its own evidence page; this one stays with the mechanics.
Why Do Two AI Detectors Score the Same Text Differently?
Disagreement between detectors is structural, not a glitch, and three design choices guarantee it. First, training data: a classifier learns the boundary between its own labeled piles, so a model fine-tuned on one million mixed samples (EV-how-ai-detectors-work-09) and a model trained at launch on GPT-3-era output (EV-how-ai-detectors-work-07) have learned different boundaries and disagree precisely on borderline text. Second, thresholds: each vendor picks its own cutoff for turning probability into a verdict (EV-how-ai-detectors-work-09), and Turnitin’s launch documentation describes flagging only at 98% confidence while accepting up to 15% missed AI text (EV-how-ai-detectors-work-07). Same essay, different cutoffs, different verdicts.
Third, and least understood: the percentages are not even the same kind of number. Winston AI’s own scoring page states its figure is a confidence level about authorship, so 80% human means 80% confidence the text is human-written, not that 20% of it is AI (Winston’s score-interpretation page; EV-winston-ai-detector-03). Turnitin’s percentage means the share of qualifying prose flagged (EV-how-ai-detectors-work-08), and since July 8, 2024 it shows an asterisk instead of any number for scores between 0% and 20%, by its own statement “to avoid potential incidence of false positives” (EV-turnitin-10). Comparing a confidence score to a proportion score is comparing units, not accuracy.
How Should You Read a Vendor Accuracy Claim?
With the question: who measured this, when, on what text? Turnitin’s under-1% false-positive figure is the vendor’s own stated number (Turnitin’s false-positives post, March 16, 2023; EV-turnitin-07), and its Chief Product Officer acknowledged in June 2023 that real-world use “is yielding different results from our lab” (K-12 Dive, June 7, 2023; EV-turnitin-12). GPTZero’s homepage advertises 99% accuracy (EV-detect-12) while its own blog ranks itself first using a benchmark it designed and ran (EV-best-ai-detector-04). The largest independent benchmark, RAID at ACL 2024, built from over six million generations across 11 models and 11 adversarial attacks, concluded that current detectors are “easily fooled” by adversarial edits, sampling changes, and unseen generators (the RAID benchmark paper; EV-how-ai-detectors-work-01). Vendor numbers are marketing until an independent test repeats them; what a score is actually worth, and how to contest one that is wrong about you, lives on its own dated page.
Why Do AI Detectors Flag Human Writing?
Because for certain writers, the flag is the mechanism operating exactly as designed. Every signal above rewards unpredictability, and whole populations of honest writers produce predictable text for structural reasons: formulaic academic prose is trained into students deliberately, technical registers compress variation on purpose, and a person writing in a second language reaches, sentence after sentence, for the safest available word. A model cannot tell machine probability from human caution; it only sees low surprise.
My own research sits in this section: I started logging detector behavior toward non-native English writing because the pattern in the appeal cases I kept reading was too consistent to ignore, careful and correct but unidiomatic prose drawing high AI percentages. The strongest published evidence agrees. A Stanford team led by Weixin Liang tested commercial GPT detectors on TOEFL essays by non-native speakers and found more than half incorrectly labeled AI-generated, an average false-positive rate of 61.3%, peer-reviewed in Patterns (Cell Press, DOI 10.1016/j.patter.2023.100779) (the study on PubMed Central; EV-best-ai-humanizer-01).
What the 61.3% Stanford Figure Actually Measured
The number is specific and gets misquoted in both directions. It is the average false-positive rate across the commercial detectors the study tested, on essays real people wrote for the TOEFL exam; the same paper reports the detectors handled native-speaker US student essays far more accurately, without publishing a comparable numeric baseline in the journal version. It is evidence of a mechanism-level bias against predictable, second-language prose, not a claim about any single tool’s current version. Weber-Wulff’s team points the same direction from another angle: machine-translated text dropped detection accuracy by about 20 percentage points (EV-how-ai-detectors-work-02). Community reports echo it as lived experience: a widely shared r/ChatGPT thread from May 2026 argued detectors punish students for writing “too correctly,” one thread and a sentiment rather than a measured rate, but the same design property this section derives from the mechanism (EV-best-ai-humanizer-12).
If you wrote something yourself and a detector flagged it, the mechanism is your defense, not your problem. Do not rewrite honest work to satisfy a probability model; document your drafting process and contest the number instead. The accuracy page walks through what a flag does and does not prove, and how to appeal one.
Why Did OpenAI Shut Down Its Own AI Detector?
In its own words: “due to its low rate of accuracy.” OpenAI launched an AI-text classifier on January 31, 2023, and retired it on July 20, 2023, noting on the same page that it correctly identified only 26% of AI-written text while mislabeling human writing as AI 9% of the time (OpenAI’s classifier announcement and discontinuation note; EV-how-ai-detectors-work-03; TechCrunch’s report; EV-gptzero-ai-detector-08).
Sit with what that means: the organization with the deepest access to how its own model generates text could not build a post-hoc classifier reliable enough to keep public. This is the event that quietly reorganized the field. The first-generation statistical meters gave way on two fronts in the same year: commercial vendors moved to trained deep-learning classifiers, the best-known dating its migration to autumn 2023 (EV-how-ai-detectors-work-05), and researchers pushed toward generation-time watermarking to stop guessing entirely (EV-how-ai-detectors-work-04). Every verdict you see today comes from a field that watched its most-resourced participant fold and rebuilt around harder methods, a piece of context the explainers currently answering this question leave out entirely.
Do Universities Still Use AI Detectors?
Most do, and a documented minority have switched theirs off; both halves of that sentence matter. The restriction record with primary sources runs through four named institutions. Vanderbilt disabled Turnitin’s AI detector in August 2023, doing arithmetic in public: against roughly 75,000 annual submissions, even the vendor’s claimed 1% false-positive rate implies hundreds of wrongly flagged papers (Vanderbilt’s announcement, August 16, 2023; EV-turnitin-01). Yale’s Canvas help center confirms the feature is disabled there (Yale’s Turnitin overview page, updated January 4, 2024; EV-turnitin-22). Waterloo discontinued the functionality campus-wide effective September 2025 (Waterloo’s notice, posted May 5, 2025; EV-turnitin-20). Curtin announced in September 2025 that it would disable the feature across all campuses effective January 1, 2026, framing the move around assessment fairness (Curtin’s update; EV-turnitin-02).
Around that chain sits an advisory layer. Pittsburgh’s teaching center disabled the tool, stating current AI detection software “is not yet reliable enough to be deployed without a substantial risk of false positives” (EV-turnitin-03). UT Austin’s provost office prohibits third-party detection tools from evaluating student work absent a university contract (EV-turnitin-21). Cornell’s Center for Teaching Innovation recommends against using detection algorithms for academic-integrity decisions at all (EV-ai-detector-accuracy-02). None of this means detectors are gone from education; it means the institutions that examined the mechanism most closely split from the ones that did not. The Turnitin-specific record, score bands and update history included, lives on its own dated evidence page.
AI Detector vs Plagiarism Checker: What Is the Difference?
An AI detector and a plagiarism checker answer different questions with different machinery, and confusing them derails real appeals. A plagiarism checker compares your text against a database of sources and points at matches, so its output is checkable. An AI detector estimates authorship statistically, matches against nothing, and can point at no source, because there is none to point at.
| AI detector | Plagiarism checker | |
|---|---|---|
| Question asked | Was this text machine-generated? | Does this text match existing sources? |
| Method | Statistical estimate from a trained model | Database and web comparison |
| Output | Probability or percentage, no source | Matched passages with the source shown |
| Verifiable? | No; the estimate cannot show its evidence | Yes; you can inspect every match |
Vendors that sell both keep them separate internally: Copyleaks’ own FAQ confirms its AI detector and its similarity checker are two independent classifiers producing two unrelated percentages (Copyleaks’ detector FAQ document; EV-copyleaks-ai-detector-11). When a report shows both scores, you are looking at two systems, and an answer that clears one says nothing about the other. In practice that means one report can show 0% similarity next to a high AI percentage (EV-copyleaks-ai-detector-11), and it tells you where an appeal has traction: the similarity match shows its sources and can be contested with evidence, while the AI estimate cannot.
The limits of this page, spelled out.
- The scope is text detection alone. Detectors for images, video, and audio are separate systems with separate failure modes, and nothing on this page transfers to them.
- No accuracy percentage for MeteGPT’s own detector appears anywhere above, and none should: nobody independent has measured it, so every accuracy figure on this page belongs to a named outside source with a date attached.
- Two coder agents label each included source independently under the protocol; dual-independent coding: pending for this page. No agreement rate exists yet, and the figure joins this record once that pass completes.
How Do You Check Your Own Text for AI Signals?
Reading about the mechanism and watching it score your own paragraph are different kinds of understanding, and the second one is free. Paste a passage into MeteGPT’s scanner and get a pattern-based signal reading in seconds, no account needed. Read the result the way this page has taught you to read every detector output: one signal, from one method, on one day. I publish no accuracy percentage for the MeteGPT scanner, because no measurement I would stand behind exists, and quoting an unmeasured accuracy number is exactly the vendor habit this article told you to distrust. The dated-sourcing rules every claim here follows are public too; the evidence protocol is documented on the methodology page.
One more habit before you paste unpublished work into any detector, this one included: check what the tool does with your text afterward. A thesis chapter or an unsubmitted manuscript is valuable precisely because nobody else has it yet, and detectors differ in what their policies permit; the retention section of a privacy policy, not the marketing page, is where the answer lives. MeteGPT’s answer is on the record: text you paste into the scanner is not retained after the request completes and is never used to train our models (the privacy policy states both). Hold any other tool to the same written standard.
The free tier’s limits are real and worth stating plainly. Anonymous scans cap at 125 words and four runs a day: enough to triage the paragraph you are most worried about, nowhere near enough for a full essay, let alone a thesis chapter. Checking longer work section by section, or re-checking after every revision pass, hits that ceiling immediately; the paid plans exist precisely to lift those caps, and that is the honest shape of the trade.
If you are staring at a percentage and wondering whether 40% is bad or whether some 30% rule exists, the score-threshold guide covers what the bands actually mean. Whatever number any tool hands you, the habit this page argues for stays the same: never let a single detector’s verdict stand alone, because the disagreement between tools is designed in. Run your paragraph through the free scanner and see which signals it trips; that firsthand reading, plus the dated record above, is a better foundation than any vendor’s marketing page.
Humanize a draft, then check the score yourself.
MeteGPT keeps a humanizer and an independent AI detector on one screen, so you can rewrite an AI-flagged passage and read a detector score on the result before anyone else does. Free daily runs, no signup.