How Do Teachers Actually Check for AI? (And What Still Works in 2026)
A calm, honest map of what the tools can and can’t do, where false positives hurt the most, and what good colleagues are quietly switching to instead.
The question gets typed into Google by teachers and students simultaneously, which tells you everything about where we are. Teachers are exhausted by the arms race. They want a practical map of what actually works — not another opinion piece. Here it is.
Quick answer
The four tools teachers actually use
These are the tools that come up in staffroom conversations, not the ones that come up in vendor press releases.
| Tool | Cost model | How it scores | Known weaknesses |
|---|---|---|---|
| Turnitin AI Detection | Institutional license only — Turnitin does not publish public per-seat pricing. Schools and universities license it through Turnitin sales, and AI Detection is typically bundled into an existing Turnitin subscription rather than priced separately. | Returns a percentage of text “likely AI-generated”; highlights flagged passages | Higher false-positive rate on non-native English writers; paraphrased AI often passes; does not identify which model was used |
| GPTZero | Free tier (limited scans). Paid plans (annual billing): Essential $8.33/month (150,000 words), Premium $12.99/month (300,000 words), Professional $24.99/month (500,000 words). Monthly billing is roughly 45% higher across tiers. | Perplexity + burstiness scoring; sentence-level highlighting | Free tier scan limits are low for classroom use; known false positives on highly technical or academic prose; no guarantee of accuracy on newest models |
| Copyleaks | Personal: $13.99/month (annual) or $16.99/month (monthly) — 1,200 credits, ~300,000 words. Pro: $74.99/month (annual) or $99.99/month (monthly), includes 25 user seats. Education plans: custom-priced per institution size, integrates with Canvas, Moodle, D2L Brightspace, Blackboard and others. | AI probability score + plagiarism overlap in one report | Combined report can conflate AI use with plagiarism, which are different issues; accuracy claims vary by model version |
| Your LMS’s native detector (Canvas, Blackboard, Google Classroom) |
Often bundled — check your LMS contract | Varies; most route through a third-party engine (frequently Turnitin or similar) | Transparency about the underlying engine varies; some institutions do not know which vendor powers their LMS detector |
The honest accuracy picture
Vendors publish headline accuracy numbers that deserve careful reading. A claim of “92% accurate” typically means 92% of known AI text was flagged — not that 92% of flags are correct. The false-positive side of that equation matters far more in a classroom, where a wrongly accused student faces real consequences.
The groups most affected by false positives are ESL and EAL writers, whose controlled, repetitive sentence structures can pattern-match to AI text; students with certain writing disorders; and anyone writing in a highly formal register (legal, academic, technical) that also happens to resemble model output.
False negatives — AI-generated text that passes undetected — are the flip side. A student who paraphrases model output through a second tool, or writes in a conversational style the model rarely mimics, will often clear any detector. A widely-cited 2023 Stanford study by Liang et al. — GPT Detectors Are Biased Against Non-Native English Writers, published in Patterns (Cell Press) — found that seven commercial AI detectors flagged 61.3% of non-native English TOEFL essays as machine-written, versus only 5.1% of native English samples. It remains the clearest empirical evidence to date that current detection skews against ESL writers.
The upshot: detection scores are evidence, not verdicts. Every academic-integrity framework that leans on a single tool’s percentage is one false positive away from a serious injustice.
What works better than detection
The educators who report the least friction in 2026 are not the ones who found a better detector. They’re the ones who redesigned the assignment.
- Process documentation — require students to submit drafts, revision notes, or Google Docs version history alongside the final piece. AI use tends to show up as a jump from nothing to polished; authentic writing leaves a trail.
- In-class writing — timed, supervised, handwritten or on a locked browser. The simplest method remains the most reliable.
- Oral defence — a five-minute conversation about the submitted work will surface whether a student understands it. This also happens to be good pedagogy.
- AI-permitted assignments — tasks that explicitly use AI as a tool, where the assessed work is the student’s editing, critique, or extension of an AI draft. Removes the detection problem by design.
- Intrinsically AI-resistant prompts — personal narrative, hyperlocal research, primary interviews, reflections on classroom discussion. Difficult to fake convincingly because the source material isn’t in any training set.
The arms race and where it goes
Every improvement in detection has been followed, within months, by model and prompt improvements that reduce its effectiveness. This has been the pattern since 2023 and mid-2026 shows no sign of breaking it. Newer models produce text with more natural perplexity variance; steerable outputs can be prompted to avoid detector fingerprints; and the same students who ask ChatGPT for an essay will ask it how to make the essay undetectable.
The honest projection: pure statistical detection becomes less viable each year as a primary enforcement mechanism. That does not mean detectors are useless — they remain useful as one signal among several, and they catch the unsophisticated cases. But institutions that have built their entire academic-integrity policy on detection are building on sand.
AI humanizers: what they are and whether they work
No honest article about AI detection can skip this. Search volume for terms like best AI humanizer runs to thousands of queries a month, and the tools behind those searches — Undetectable AI, StealthGPT, WriteHuman, Humbot and a long tail of imitators — exist for one purpose: to rewrite AI-generated text so it scores as human on detectors.
This section is written from the detection side, because that is what this article is about and because the marketing claims on humanizer sites are worth checking against reality. We are not linking to these tools or ranking them.
What a humanizer actually does
Nothing mysterious. A humanizer is itself a language model, prompted to rewrite input text so that its statistical fingerprint — the perplexity and burstiness patterns detectors measure — looks less machine-like. In practice that means: varying sentence length more aggressively, swapping common words for less predictable synonyms, introducing minor grammatical looseness, and breaking up the even rhythm that instruction-tuned models default to.
Do they work?
Partially, temporarily, and at a cost. Three things are consistently true as of mid-2026:
- They do reduce detector scores — often substantially, on the specific detectors the vendor optimised against. That is the honest part of the marketing.
- Detectors adapt. Humanizer output has its own fingerprint: the synonym substitutions skew toward unusual register, and the injected variance is itself statistically regular. Major detectors have added humanizer-specific classifiers, and vendors’ own “undetectable” claims tend to lag their testing by several months.
- The writing gets worse. This is the part nobody advertises. Substituting rarer synonyms for common words produces prose that reads oddly — technically fluent, semantically slightly off. Teachers who have read a student’s earlier work notice it, and it is the kind of thing that triggers a closer look rather than deflecting one.
The third point matters more than the first two. A humanizer optimises for one metric — a detector score — at the expense of the thing being assessed. Text that passes GPTZero but reads like it was translated twice has not solved the student’s problem; it has changed which signal gives them away.
What this means for teachers
If you suspect humanizer use, the tells are stylistic rather than statistical: vocabulary that sits above the student’s demonstrated range, synonym choices that are technically correct but contextually strange, and sentence-level variety that does not match paragraph-level structure. A low AI-detection score does not rule out AI use — and this is precisely why detection scores should never be the sole basis for an academic-integrity decision.
The design response is the same one that works for undisguised AI use: process documentation, drafts, revision history, oral defence, and assignments anchored in specific lived or local material. None of those are affected by how the text was post-processed.
What this means for students
Two things worth being straight about. First, the tools are less reliable than their marketing suggests, and the failure mode is not a neutral one — it produces writing that draws attention. Second, and more importantly: most institutional academic-integrity policies treat deliberate evasion of detection as an aggravating factor, not a neutral act. Being caught having used AI is one conversation; being caught having actively tried to disguise it is a materially worse one.
If your institution permits AI use with disclosure — and a growing number do — disclosure is both easier and safer than evasion. If it does not, the honest calculation is that humanizers shift the risk rather than removing it.
What schools are actually doing in 2026
There is no consensus, and that is itself the story. A rough landscape as of mid-2026:
- Allow with disclosure: A growing number of universities — particularly in technology and business faculties — permit AI use provided it is documented in a methods or tools section, similar to any other research tool.
- Blanket ban: Common in secondary schools and in disciplines where the writing process is itself the learning objective (writing courses, law, medicine). Enforcement relies on a mix of detection and design.
- Undecided or inconsistent: The most common situation. Individual teachers make their own calls; policy documents exist but predate current model capabilities; enforcement is uneven.
geminy.ai does not advocate for any of these positions. The landscape is genuinely contested and the right answer depends on the educational context, subject, level, and institutional values.
One consistent data point: schools that invested in teacher professional development around AI — not just detection, but use, ethics, and assignment design — report fewer integrity incidents than those that invested only in tools. See also: AI for accessibility in education, where AI use is unambiguously beneficial and where blanket bans cause real harm.
A note for students reading this
This article will be read from both sides of the desk. If you’re a student: detection is only one part of how academic integrity is assessed, and it’s the part that is getting less reliable, not more. The bigger shift is in assignment design — and the teachers who are best at catching AI use are the ones who have redesigned their assignments so the question rarely arises. Understanding AI tools well — including how ChatGPT and Claude actually work — is more valuable than knowing how to evade a detector.
If you’re uploading your work to a detection service: read the privacy implications first. Student work submitted to third-party tools may be retained, processed, or used in ways your institution’s policy has not fully addressed.
Frequently asked questions
How do teachers check for AI writing?
Most use one of four tools: Turnitin AI Detection, GPTZero, Copyleaks, or their LMS’s built-in detector. They typically look for a high AI-probability score combined with other signals — inconsistent voice, sudden quality jumps, or a lack of process documentation. A score alone is not considered sufficient by most academic-integrity policies.
Is Turnitin AI detection accurate?
Turnitin reports high accuracy on known AI text, but false-positive rates — flagging genuine student writing as AI — are a documented concern, particularly for ESL writers and highly formal prose. Most institutions treat a Turnitin AI score as one signal, not a verdict.
What is the best free AI checker for teachers?
GPTZero has the most widely used free tier among educators. Copyleaks also has a limited free option. Both have scan limits on free plans that may be restrictive for high-volume classroom use. Accuracy varies; no free tool has demonstrated consistent reliability across all model outputs in 2026.
Can AI-generated text be detected reliably?
Not reliably enough to be the sole basis for an academic-integrity decision. All current detectors have meaningful false-positive and false-negative rates. Detection works best as one signal among several, combined with process documentation and assignment design.
What assignment designs are AI-resistant?
Personal narrative drawing on specific lived experience, hyperlocal research using primary sources, oral defences, in-class writing, and assignments that require students to document their process (drafts, revision history). AI-permitted assignments — where AI use is the explicit task and the student’s editing is assessed — sidestep the detection problem entirely.
What is the best AI humanizer?
We don’t rank them, and the premise is weaker than the marketing suggests. AI humanizers are language models that rewrite AI-generated text to lower its detector score, typically by varying sentence length and substituting less predictable synonyms. They do reduce scores on the detectors they were optimised against, but detectors have added humanizer-specific classifiers, and the rewriting reliably degrades writing quality — producing prose that is technically fluent but contextually odd, which is itself a signal teachers notice. Most academic-integrity policies also treat deliberate evasion as an aggravating factor rather than a neutral act.
Can teachers detect AI humanizer use?
Often, but through stylistic rather than statistical signals. Humanizer output tends to show vocabulary above the student’s demonstrated range, synonym choices that are technically correct but contextually strange, and sentence variety that doesn’t match paragraph structure. A low AI-detection score does not rule out AI use — which is one of the main reasons detection scores should never be the sole basis for an academic-integrity decision.
Is it a privacy risk to upload student work to an AI detector?
Potentially. Third-party detection services may retain submitted documents for model training or product improvement. Check your institution’s data-processing agreements and the vendor’s privacy policy before uploading student work, especially for minors. See geminy.ai’s privacy guide for the broader context on AI tool data handling.
Sources
- GPTZero — Pricing
- Copyleaks — Pricing
- Liang et al. (2023). GPT detectors are biased against non-native English writers. Patterns, Cell Press.
Leave a Reply