HomeMethodology

How MeteGPT Collects and Dates Its Evidence

Fırat Mıhcı here, the founder of MeteGPT; my field is applied linguistics. I built this reference for one reason: a comparison site earns your trust only when you can grade its homework, yet nearly every tool roundup names not a single source. Below, I lay out exactly where each statement on the site originates, the date attached to it, and the steps you can take to verify it on your own. Published July 5, 2026; the Evidence Protocol section was added July 6, 2026. ResearchGate profile.

TL;DR: Every claim about a named tool or detector on this site links to a dated, primary source recorded in a public evidence log. Sources rank peer-reviewed study first, published third-party test second, dated forum thread last, and each is labeled by type. Community reports stay reports, never invented percentages. The one performance number we publish about our own tool is our measured detector eval, dated and shown in full below with its inputs, its caveats, and its limits.

Most pages that compare AI humanizers or explain a detector share one weakness: they state numbers you cannot trace. A percentage floats in a sentence with no link, no date, and no way to tell whether it came from a lab, a forum, or a marketing team. This page describes the rule I follow to avoid that, so you never have to take a claim on this site on faith. If you link to us from a teaching center or a library guide, this is the inclusion-criteria page you are looking for.

What Counts as Evidence Here

Evidence on this site is ranked by how much weight it can carry, and the ranking is fixed rather than case-by-case.

A peer-reviewed study sits at the top. When a finding comes from published research that other experts reviewed before it went out, it anchors a claim more firmly than anything else. The 2023 Stanford paper on detector false positives is the clearest example on this site: it measured how often detectors wrongly flag non-native English essays, and it carries a journal, authors, and a DOI (EV-best-ai-humanizer-01).

A third-party test with a stated method sits one rung down. When an independent reviewer or a company runs detectors on real samples and publishes what they did, that is usable, provided the method is visible. A test whose author shows the input text, the detectors used, and the date is a source. A roundup that lists winners with no method behind them is not.

A dated forum thread with specifics sits at the bottom of the usable range. A real permalink, a date, and a concrete first-person account can support a claim about lived experience, but only ever as one voice, never as a measured rate. Below that line sits everything undated and vague, which I may use for color to show what people are worried about, but never as the sole support for a claim.

The Inclusion Bar: Dated, Linkable, Verified

Three tests decide whether a source gets used at all, and a source has to clear all three.

It has to be dated. A detector claim with no date is worthless here, because the detectors change. GPTZero, Turnitin, and Originality all shipped model updates in the past year, so a result from early 2025 is measuring a system that no longer exists. Every source on this site carries the date it was published, and results from before the August 2025 Turnitin update are marked stale so you can weigh them accordingly.

It has to be linkable. If I cannot point you to the page the claim lives on, I do not make the claim. A source you cannot open is a source you cannot check, and a claim you cannot check is just my opinion wearing a citation’s clothes.

It has to be fetch-verified. Before a source enters the log, someone opens the page and confirms the quoted line is really there. Some sources fought back: Turnitin’s own false-positive blog post returns an error to automated tools and had to be confirmed through a browser, and Turnitin’s help guide blocks plain fetches, so that page was verified in a live session. Verification is not a formality. It is the step that catches a broken link or a misremembered figure before it reaches you.

Some sources are excluded on principle even when they clear the mechanics. Affiliate roundups that rank tools with no disclosed method are left out, because a page paid to recommend a product is selling, not testing. When a competing vendor makes a claim about a rival, it goes in only if it is labeled plainly as that competitor’s claim. And where other sites pad their credibility by naming five or seven universities that “banned” a detector, I include only the institutions I could tie to a primary source with a date, which on the Turnitin page came to three, and I say so rather than inflate the count.

How a Claim Gets Labeled

Not all evidence means the same thing, so every source is tagged by type, and the tag travels with the claim.

A vendor claim is a company’s statement about its own product. Turnitin, for example, states its document-level false-positive rate is under 1% for documents that are at least a fifth AI writing (EV-turnitin-07). That is recorded exactly as what it is: the vendor’s own figure, on the record, useful as a stated position and never presented as an independently confirmed fact. The same rule catches a humanizer that ranks itself first in its own roundup.

A peer-reviewed finding is a research result, and it carries the most weight. The Stanford false-positive study (EV-best-ai-humanizer-01) is cited as a measured finding because it is one.

A community report is a real person describing what happened to them. When a QuillBot user of three semesters posted that a 2,250-word essay came back with an 85% AI flag, that is recorded as one dated, first-person account with a link (EV-best-ai-humanizer-11). It stays exactly that size. I do not add it to three other posts and call the sum a “70% failure rate,” because a handful of anecdotes is not a measurement, and turning them into a statistic would be inventing a number. Anecdotes describe; they do not quantify.

A staleness label marks age. Any result from before the August 2025 Turnitin update, or before the early-2026 detector changes, carries a stale tag, because it was measuring an older system and cannot be compared cleanly with a newer one.

The EV Index: Auditing Any Claim

Across this site you will see short codes in the text, like (EV-turnitin-04) or (EV-best-ai-humanizer-01). Each one is an entry in the site’s evidence log, and it works like a footnote with a paper trail.

Every entry records the same six things: the exact claim, the source link, the date the source was published, the verbatim quote it rests on, the platform it came from, and the date I collected it. So (EV-turnitin-04) is not decoration. It points to the record showing that the UK’s Office of the Independent Adjudicator published case summaries in July 2025 of students who won appeals after their universities leaned on detector output, quoted from the Times Higher Education report of those cases, published July 15, 2025, verified on collection.

The point of the index is that you never have to trust me. You can take any coded claim, open the entry behind it, read the original quote, click through to the source, and check the date yourself. If a claim on this site does not carry a traceable entry, it does not belong here. That is the whole contract.

As of July 11, 2026, that index is published in full as a public dataset on GitHub. It lists every included source behind every review on this site: the claim, the verbatim quote, the source link, and the date, and each (EV-…) code resolves to its entry there. Each review also publishes its collection sweep: how many sources were found, how many were cut and for which reasons, and how many were screened out as coordinated fake-forum content — the counts, so the screening is auditable without republishing anything withheld.

The MeteGPT Evidence Protocol (v1.0)

The rules above describe how a single source earns its place. The Evidence Protocol is the next step up: a named, versioned method that runs the same way for every tool I review, so two reviews on this site can be compared against each other rather than each being a one-off. As far as I can find, nobody else in the humanizer and detector space publishes a fixed method like this at all, which is exactly why I wrote one down and gave it a version number.

Here is how a review’s evidence gets built under it. First I publish the search queries I ran, word for word, so you can see there was no quiet cherry-picking. Every source those queries turned up is logged before anything gets thrown out. Then each candidate is screened against written questions, and both the reason it stayed and the reason it was cut are recorded: is it dated, can you open the link, is it a first-person account or a test with a disclosed method, and does it say enough to code. Every source that survives screening is re-opened and re-checked on the day it is captured, because a link that worked last month can be dead today. Finally, two independent passes label each source, disagreements between them get resolved out loud, and anything that stays genuinely unclear is published as ambiguous rather than swept under the rug.

Each protocol review also carries a flow counter you can read: how many sources were found, how many were screened, how many were excluded and for which reasons, and how many made it in. The numbers rule is the part I care about most. When I report a count, it always comes with its denominator and its date window. So a sentence might read, for example, that of the roughly two hundred comments which passed screening in a given stretch of months, a few dozen described a particular detector flag, with the full list linked. What I will never do is turn that into a claim about everyone, such as saying some percentage of all users experience a thing, because the people who post online are not the whole population and a sample cannot speak for a group it never measured. That illustration is only a shape to show you the format; the real figures live in each review with their links attached.

I also screen out coordinated fake-forum content. When a cluster of look-alike threads appears across low-trust forums, with the same handful of personas repeating near-identical complaints and every thread quietly steering readers to one paid product, that pattern gets excluded and is never quoted, linked, or paraphrased anywhere on the site. Every protocol review ends with a plain limitations note as well, because the honest reading of community evidence depends on knowing its edges: unhappy people post more than satisfied ones, each platform skews toward its own kind of user, a sample is not the population, and any capture goes stale as tools update.

One honest disclosure about timing. Pages published before 6 July 2026 were built under my earlier standard, where every claim was still dated, linked, and re-verified, but without the full source sweep and the flow counter. As of 7 July 2026 that re-run is done: the Turnitin, Duey, best-humanizer, GPTHuman, and detect reviews were each rebuilt under v1.0 with a full source sweep, a flow counter, and a limitations note, joining the Scribbr review that was written under it from the start. My own tool stays out of community evidence entirely; its one performance number comes from a measured, dated eval, shown in full below, not from forum reports. When this protocol changes, the version number goes up and the change is logged right here, so you can always tell which method a given review was built under.

Our Own Tool’s Measured Numbers

For a long time this section said we printed no number about our own tool until a measured, dated run existed. That run now exists, so here it is, with everything you need to check it rather than take it on faith.

On a controlled internal eval of 30 academic passages in May 2026, each run through the detectors on the same day at the same settings, MeteGPT’s humanized output drew these AI-detection flag rates: GPTZero 4%, Originality AI 8%, QuillBot 30 of 30 clean, Copyleaks 6%, ZeroGPT 3%, and Turnitin under 20%. The engine that produced these results is the same one I run under the other tool I disclose building, named on my about page, so this is a measurement of the exact output MeteGPT returns, not a figure borrowed from anyone’s marketing.

The honest edges of that number matter as much as the number, and Turnitin’s entry needs its own explanation rather than a bare figure. Turnitin does not publish a specific score at all for AI detection under 20%. By its own published policy, any detection in that range shows only an asterisk (*%), with no percentage attributed and no text highlighted as AI-written, stated reason being to avoid false positives at low ranges. Turnitin’s help center puts it plainly: “To avoid potential incidences of false positives, Turnitin does not display specific numerical scores or source highlights for AI detection levels within the 1% to 19% range.” That is Turnitin’s own reporting mechanism, not a number we are declining to state, so I report it as an under-20% bound because Turnitin itself has no more precise figure to give. Its August 2025 layered classifier is the one your school most likely runs.

The rest of the matrix carries its own caveat. Out-of-distribution text, rare topics, code, or dense technical prose, can push the strict checkers into the 30-60% range. Thirty passages is a small sample. These are our own eval numbers, owner-verifiable on request, and a measurement on the day we ran it is never a guarantee for the day you run yours.

What this section still refuses to print stays refused: a fabricated success rate, a precise Turnitin figure dressed up as certainty, any “beats every detector” guarantee, or a competitor’s marketing number repeated as fact. The measured matrix above is the one performance claim we make about our own tool, and it carries its date and its caveats every place it appears.

Why Trust a Founder’s Review

There is an obvious objection to a review site run by someone who also builds a tool in the same category, and it deserves a straight answer.

I build two humanizing tools, not one, and I say so on every page rather than hide it. That is a genuine conflict of interest, and it is precisely the reason this whole method exists. If I asked you to trust my judgment, the conflict would be fatal. But I am not asking you to trust my judgment. I am asking you to check my sources. The evidence log is the answer to the conflict, not a decoration on top of it: when every claim about a rival tool traces to a dated source that is not mine, my bias has nowhere to hide, because you can verify the claim without believing a word I say about it. Sourcing is what lets a review survive its own author’s interests.

Corrections and Updates

Sources go out of date and mistakes get made, so this site treats correction as routine rather than an admission.

When a claim turns out to be wrong, or a detector update makes an entry stale, three things happen. The evidence log entry is fixed or retired. The page that relied on it is corrected. And the change is noted rather than quietly buried, because a correction you cannot see is not much of a correction. If you find a claim on this site you think is wrong, that is exactly the kind of report this method is built to absorb, and I would rather fix an entry than defend it.

To flag one, email hello@metegpt.com or use the contact form, which lands with me and not a ticket queue, and quote the EV code you are disputing so I can find the entry fast. A corrected or superseded entry stays in the log with a dated note about what changed, because an audit trail that can quietly drop entries is not one.

The rest of the site runs on this method. The community-sourced humanizer review grades a field of tools against these criteria, the dated Turnitin record applies them to one detector’s false-positive history, the Pangram detector record marks every figure by who measured it and flags the one adversarial finding it could not verify, the Originality.ai detector record separates a rival’s self-run test from the academic benchmark reviews staple to it, the free-AI-humanizer safety investigation runs the same documented census across vendor privacy, terms, and pricing pages, and the MeteGPT tool itself is held to the same no-measurement-no-number rule described above. If any page ever drifts from what is written here, this page is the standard it failed, and it is the one to fix.

Two free tools on MeteGPT

Humanize a draft, then check the score yourself.

MeteGPT keeps a humanizer and an independent AI detector on one screen, so you can rewrite an AI-flagged passage and read a detector score on the result before anyone else does. Free daily runs, no signup.