ChatGPT Detector
Paste a draft and every phrase that reads as ChatGPT gets marked, with the reason named — the stock vocabulary, the recycled transitions, the em dashes, the conclusion that only restates — plus the hidden characters a paste carries that never appear on screen. Not a percentage. The specific words, and what to do about each. Which matters because of the fact below: OpenAI built a classifier for its own model and withdrew it in 2023 for low accuracy.
OpenAI built a ChatGPT detector and shut it down.
In early 2023 OpenAI released an AI text classifier. In July 2023 it withdrew it, stating plainly that the reason was a low rate of accuracy. It is worth sitting with what that means before reading any other claim in this category.
They had every advantage. Access to the model, to its training data, to enormous volumes of its output, and to the people who built it. If reliable detection of ChatGPT text were achievable, they were the best-placed organisation on earth to achieve it.
The limits they published are the limits of the whole field. The classifier needed at least a thousand characters to say anything, and performed noticeably worse on text that was not in English. Both constraints apply to every tool that has since claimed to solve this, and most of them do not mention either.
Independent research agrees. Studies evaluating the commercial detectors — ZeroGPT, GPTZero and others — report false positives and false negatives at rates that make them unsuitable as evidence of authorship, which is precisely the use they are most often put to.
So this page does not publish an accuracy figure, and it is candid that the check below is a reading of how predictable your prose is rather than a determination of who wrote it. The next two sections are more useful than any score, because they are things you can verify yourself.
Five markers of ChatGPT prose, without a tool.
Three of these come from comparative research on ChatGPT answers against human ones. All five are checkable by eye.
- The three-part answer, whatever was askedThe most reliable marker and the one visible without any tool. An opening restating the question, a body of evenly weighted points, a closing that summarises rather than concludes. Every section arrives at a similar length, which flattens the sentence-length variation detectors measure and which a reader registers in seconds.
- More nouns and longer sentences than people writeComparative studies of ChatGPT answers against human ones find measurably higher noun density, longer average sentences, and more determiners and conjunctions holding those sentences together. That is what nominalised, carefully hedged prose looks like when you count it rather than read it.
- Sentiment that stays neutralThe same research finds ChatGPT output sits closer to neutral sentiment than human writing on the same questions. People writing about something they care about drift positive or negative; a model trained to be balanced does not. Across a long document that evenness is audible.
- A small set of favourite wordsDelve, leverage, robust, seamless, landscape, tapestry, testament, crucial. Ordinary words used at a rate writers do not use them. One means nothing. Four on a page is a pattern, and it is the tell most people notice first without being able to name it.
- Transitions and hedges at a fixed rateIt is important to note, moreover, furthermore, in conclusion. Not wrong individually, and used at a frequency no person writes at. Three furthermores in four paragraphs reads as machine-made to an editor long before it reaches any detector.
Detecting GPT-4, GPT-4.5, GPT-5, GPT-5 Pro and GPT-5.1 text.
Scores fall with each release, and the structure that gives the output away does not change with them.
GPT-4 and GPT-4.5 detector
The output most represented in every detector's training, and the most confidently scored.
Detectors were built during the GPT-4 era and largely on GPT-4 text, so this is where they perform best and where their published accuracy figures come from. If you are reading a study about detection accuracy, there is a good chance the corpus was GPT-4 output.
The structural habit is at its most pronounced here: the restated introduction, the three evenly weighted body sections, the summarising conclusion. GPT-4.5 softened the vocabulary slightly without changing the shape.
GPT-5 detector
Writes with more range, which moves scores down without changing what a reader sees.
GPT-5 varies sentence length more than GPT-4 did and reaches for the stock vocabulary less often, so it tends to score lower on tools calibrated against earlier output. That is a genuine change and it is a change in the statistics rather than in the structure.
The three-part shape persists. A GPT-5 draft often reads as less mechanical sentence by sentence and lands in the same overall form, which is why a document-level percentage tells you less about it than a sentence-level read does — and why a reader can still recognise it when a detector does not.
GPT-5 Pro detector
The hardest of the family to score, for the same reason it is the most useful.
Reasoning-heavy output is longer, more varied and more specific, which pushes it further from the flat statistical profile detectors were built to catch. Expect lower scores on text that is unmistakably generated to anyone who reads it.
It also sometimes carries planning language into the answer — a paragraph discussing how the question should be approached before answering it. That is not statistical and it is the clearest evidence available in this family.
GPT-5.1 and GPT-5.2 detector
The moving-target problem, stated plainly.
Each version writes slightly differently, and detectors are retrained after the fact rather than in advance. So there is always a window where current output scores lower than it should, simply because the tool has not seen it yet. Anyone quoting an accuracy figure for a model released this quarter is extrapolating.
The practical consequence: a low score on very recent output is weak evidence of human authorship, and it is exactly the case where reading the document yourself matters most.
The human writing that gets flagged as ChatGPT.
These are not edge cases. They are large, identifiable groups whose ordinary writing scores high for reasons that have nothing to do with how it was produced.
Non-native English speakers. The best-documented failure in the category, and one OpenAI named in its own classifier’s limitations. Competent second-language English tends to be more textbook-regular than a native speaker’s, with a narrower vocabulary and closer adherence to taught structures. All three raise the score.
Anyone writing to a template. Taught essay structures, standard report formats, the five-paragraph shape. Following the structure you were taught is the thing most likely to get you flagged for not having written it, which is a genuinely perverse outcome.
Technical and legal writing. Uniform by professional convention, because ambiguity in a specification is a defect. The measurement cannot separate disciplined prose from generated prose, since at the level it operates they look the same.
Short passages, from anyone. The measurement is statistical and needs volume. OpenAI’s own classifier required a thousand characters before it would return anything, and tools that will score a paragraph are answering a question they do not have the data for.
ChatGPT detector questions.
Is there an official ChatGPT detector from OpenAI?
Not any more, and the history is the most useful fact on this page. OpenAI released an AI text classifier and withdrew it in July 2023, citing a low rate of accuracy. It had required at least 1,000 characters and performed poorly on non-English text. The company with the most access to how ChatGPT writes could not build a reliable detector for it, which is worth remembering when a third party claims to have done so.
How accurate are ChatGPT detectors?
Less accurate than their marketing, and independent research is blunter about it than the vendors are. Studies evaluating ZeroGPT, GPTZero and similar tools report both false positives and false negatives at rates that make them unsuitable as evidence of authorship. They are useful as a signal about how predictable prose is and not as proof of who wrote it.
How can I tell if text was written by ChatGPT without a tool?
Read for the shape rather than the words. An opening that restates the question, three body sections of near-identical length, a conclusion that summarises instead of concluding. Then look for the vocabulary — delve, leverage, robust, seamless, tapestry — and count the transitions. Four stock connectives in a page is a stronger signal than most scores.
Does GPT-5 text score differently from GPT-4?
Yes, generally lower. GPT-5 varies sentence length more and uses the stock vocabulary less, so it sits further from the flat profile detectors were calibrated on. The structure is unchanged, which means a reader can often still recognise it when a tool does not.
Why do detectors get worse with each new model?
Because they are trained after the fact. A model ships, then detectors are retrained on its output, so there is always a window where current text scores lower than it should. Any accuracy claim about a model released this quarter is an extrapolation rather than a measurement.
Can it detect ChatGPT text that has been edited?
Less reliably, and the mixed document is the normal case rather than the exception. A generated draft a person has revised can score low, and a hand-written draft a person has tightened can score high, because tightening prose is the same operation as making it regular. Sentence-level reading is the only version of this that helps.
Does removing em dashes or hidden characters change a score?
No. Detectors read the words and their predictability. They do not count punctuation, read metadata, or look for zero-width characters. All of those are worth cleaning for other reasons and none of it moves a score.
My own writing was flagged as ChatGPT. What now?
It happens constantly, and to a predictable group: people writing in English as a second language, people writing to a taught structure, and anyone in a field whose conventions demand a formal register. A high score means the prose is predictable. Show the process instead — version history, dated drafts, notes — because that is what a percentage cannot produce.
What can I check here?
A sentence-level read of which exact phrases read as AI-written, and why each was flagged — the stock vocabulary, the recycled transitions, the em dashes — rather than a statistical score of the kind the tools above return. We publish no accuracy figure, for the reason at the top of this page: the organisation best placed to build one gave up on it.
Is my text stored?
The document is attached to your account so you can return to it, and you can delete it whenever you like. It is not used to train anything.