AI Assignment Checker

Most people searching this are the ones marking, so this page is written for them first. A detector reads how predictable prose is — useful as a prompt to look closer, and not capable of showing who wrote something. It is also meaningless on half of what arrives called an assignment: problem sets, lab reports, short reflections and code give it almost nothing to measure and it returns a confident number anyway. What to look at instead is further down. If you are the student, the last two sections are for you.

The submission0 words
What scores as AI
Your findings will appear here
01 · If you are marking

What a score can support, and what it cannot.

Everything else on this page follows from one distinction, so it is worth putting first and plainly.

A detector measures predictability, not authorship. It reports how closely each word follows from the ones before it and how much sentence length varies. Generated text scores high because that is what a language model produces. So does a great deal of honest work, for entirely different reasons.

Which means a percentage cannot carry a decision on its own. It can reasonably prompt you to read something more carefully. It cannot substantiate an allegation, and a process that treats it as though it can is one that will eventually be appealed by a student who did nothing wrong. GPTZero, which sells one of these tools, states in its own materials that results should not be used to punish or as a final verdict.

And the errors are not randomly distributed. This is the part that matters most. Students writing in English as a second language are penalised systematically, because competent second-language English is often more textbook-regular than a native speaker’s. Students who follow the taught structure closely score higher, which frequently means the conscientious ones. Routine use of a detector applies pressure unevenly, and it applies it to the students least able to push back.

02 · By format

Assignments are not all essays, and detection knows it.

A tool will return a number for every one of these. It is only measuring anything meaningful on the first.

  • Essays and written answersThe only format these tools were really built for, and the only one where a score means much. Continuous prose of a few hundred words or more gives a statistical measurement enough to work with. Everything below gives it less, and the tool will still return a confident-looking number.
  • Short answers and reflectionsToo short to score reliably, and reflections are the worst case of all. A student writing about their own learning in a taught format produces exactly the regular, conventional prose that reads as generated. Personal reflection is also the format where a false accusation does the most damage to a working relationship.
  • Problem sets and calculationsDetectors have essentially nothing to read here. Working, notation and short justifications are not the continuous prose the measurement needs, and a percentage on a maths assignment is close to noise. If the concern is that a model produced the solution, the method shown is far more informative than any score — models tend to produce correct answers by routes nobody was taught.
  • Lab and technical reportsWritten to a fixed structure in a deliberately flat register, which is to say written to be predictable. This is the format where honest work scores highest, because following the template correctly is the assessed skill. Treat a high score on a lab report as evidence the student followed the template.
  • Code and coursework with codeText detectors do not read code meaningfully, and the comments are too sparse to score. Version history, commit patterns and the ability to explain a function are the things that actually answer the question here, and all three are available to you.
03 · What to look at instead

Four things you have that no detector does.

These are all stronger than a percentage, and all of them are available to somebody marking a set of submissions rather than analysing one file in isolation.

Fit to the brief. The most reliable signal there is. Generated work answers the general topic competently and misses the specific instruction: your module’s terminology, the set text you required, the framework from week six you asked them to apply. A model does not know what happened in your seminar. An answer that could have been written for any course on the subject usually was.

The student’s own previous work. A sudden change in voice, vocabulary or structure across submissions is far more informative than any absolute figure, because it compares a student against themselves rather than against a general model of writing. This is also the comparison that catches the opposite case: a flagged student whose writing has been consistent all term.

Whether it can be explained. A short conversation resolves most of these. Someone who wrote an argument can say why they took that line, what they considered and dropped, where a source came from. Someone who did not runs out of depth quickly, and does so in a way neither of you needs to formalise.

The references, opened. Models invent plausible citations, and search-based tools cite real sources for claims they do not make. A fabricated reference is concrete in a way a percentage is not, and it is the finding that actually holds up.

04 · By assistant

What ChatGPT, Claude, Gemini, LLaMA and Perplexity leave in submitted work.

Read as things a marker can see, since most of them are not what a detector is measuring.

ChatGPT assignment checker

The assistant behind most generated coursework, and the one whose habits a marker learns to recognise fastest.

The recognisable thing is shape rather than vocabulary: an opening that restates the brief, evenly weighted middle sections, a closing that summarises. Across a set of thirty submissions it stands out immediately, because real student writing varies enormously in structure and this does not.

The more useful signal for a marker is fit to the brief. Generated answers address the general topic competently and skip the specific instruction — the module's own terminology, the required reference to a set text, the thing you asked them to apply from week six. That gap is visible to you and invisible to any detector.

Claude assignment checker

The hardest to call, because what it produces is genuinely good writing.

Claude hedges carefully and structures arguments properly, so a submission reads as a strong student. The tell is repetition of shape across paragraphs: a claim, a clause narrowing it, a clause admitting an exception, again and again. Individually each sentence is defensible, which is exactly why a score on it is a weak basis for a decision.

The practical marker is not statistical. Claude output arrives carrying Markdown, so submitted work sometimes contains bolded lead-ins or bulleted breakdowns in a section meant to be continuous. That says something about how the file was assembled and nothing about who thought about the question.

Gemini assignment checker

Gemini output interacts with the format problem above more than any other assistant.

It returns headings and bullets by default, so submissions drafted with it arrive fragmented. Fragmented text gives a detector very little to measure, which means Gemini-drafted work can return a lower score than a genuinely human essay written in continuous prose. That inversion is worth knowing before a number is used to compare two students.

What is visible instead is that the answer is a list where an argument was asked for. That is markable on its own terms, and it is a judgement about the work rather than about its origin.

LLaMA assignment checker

The case where scores are least predictable and least meaningful.

LLaMA is weights rather than a product, so what a student used depends on which app served it and how heavily the model was quantised. Two students using tools with the same model name submit differently shaped prose, and a detector calibrated on one generalises poorly to the other.

Smaller deployments also produce genuine errors the hosted assistants do not: tense that wanders, arguments that lose their thread mid-page. Those are markable as writing problems, which is a firmer footing than a percentage.

Perplexity assignment checker

The one where the reference list deserves more attention than the prose.

Because it searches, its citations usually point at real documents, which sounds reassuring and introduces the harder problem: a real source cited for a claim it does not make. Checking that a reference exists returns a pass, so the error survives the verification most markers have time for.

Its output is also built around citation markers. Students strip them before submitting, which leaves sentences hedged toward an authority named nowhere in the work — the shape to read for, and more informative than any score.

05 · If you are the student

Checking your own work, and what to do if you are flagged.

You can paste an assignment above and see which passages read as statistically predictable. Useful for knowing whether your writing sounds like a template, and it will not tell you what your institution’s tool will report — detectors disagree constantly and you probably cannot access the one being used on you.

If you have been flagged and you wrote it, do not argue the number. Treating it as a quantity of evidence concedes that it is evidence. It reports that your prose is predictable, and there are several reasons for that which have nothing to do with how it was produced.

Show the process. Version history, dated drafts, your outline, reading notes, the browser history where you found sources, the paragraph you cut and why. Work that was written leaves a trail behind it. This is more persuasive than any counter-check, because it is the thing a score cannot produce.

And keep that trail before you need it. Writing in a tool with version history costs nothing and means the evidence accumulates on its own. It is the single most useful habit for anyone submitting into a system that runs these checks, and it takes no effort once it is a default.

06 · Related

For other kinds of work.

Assignments cover a lot of ground, and several of these are handled better elsewhere.

For essays specifically, the AI essay checker covers what a marker reads for and the invented-citation problem in more detail. For postgraduate work, the AI thesis checker covers why a single score across a long document is meaningless, and the AI research paper checker covers what journal policies permit and require.

For how the measurement itself works and how far to trust it, the AI detector page sets it out, and the Turnitin AI checker page explains what institutions are usually running and why students never see the report.

If the writing is yours and reads as flat, that is editing rather than detection: the AI grammar checker fixes errors without touching your voice, and the AI academic humanizer varies rhythm while keeping the register and every citation exactly as written.

07 · FAQ

AI assignment checker questions.

Can an AI checker prove a student used AI?

No, and this is the most important thing on the page for anyone marking. A detector measures how predictable prose is. Predictable prose has several causes, of which generation is one and following a taught structure is another. A percentage is a prompt to look more closely, and it is not evidence of authorship in any process that would survive an appeal.

Which students get falsely flagged most?

Students writing in English as a second language, by a wide and well-documented margin, because competent second-language English is often more textbook-regular than a native speaker's. Then students who follow the taught structure closely, which tends to mean the conscientious ones. Then anyone writing in a format that requires a flat register, such as a lab report. If a detector is used routinely, it will apply pressure unevenly to exactly those groups.

Does it work on problem sets, lab reports and code?

Not usefully. Detection needs continuous prose of a few hundred words to have anything to measure. Working and notation, short technical sections and code give it almost nothing, and it will still return a confident number. For those formats the method shown, the version history, and whether a student can explain their own work are all more informative.

What should I look at instead of the score?

Fit to the brief. Generated work answers the general topic and misses the specific instruction — your module's terminology, the set text you required, the thing from week six you asked them to apply. Compare against the student's own earlier work too, since a sudden change in voice is more telling than any absolute figure. Both of those are things you have and no tool does.

Can it check assignments written with ChatGPT, Claude, Gemini or LLaMA?

It reads all of them, and the sections above set out what each leaves behind. Worth knowing that Gemini-drafted work often scores lower than genuinely human prose because it arrives fragmented into bullets, which means a score can rank two submissions the wrong way round.

I am a student — can I check my own assignment?

Yes, and the useful expectation is modest. You will see the exact phrases that read as AI-written, which is worth knowing if you want your writing to sound less like a template. It will not tell you what your institution's tool will report, since detectors disagree with each other constantly and you may not have access to the one being used on you.

My assignment was flagged and I wrote it. What do I do?

Do not argue about the percentage; that concedes it is evidence. Show the process instead: version history in Google Docs or Word, dated drafts, your notes, the browser history where you found sources, the paragraph you cut and why. Work that was written leaves a trail of being written. Keeping that trail routinely, before you need it, costs nothing.

Should I tell students I use an AI checker?

It is worth thinking about, and the case for saying so is strong. A check students know about is a policy; one they discover after an accusation is a surprise that damages trust. Being explicit about what the tool does, what it cannot show, and what happens after a flag turns a confrontation into a conversation, which is where these are actually resolved.

Do hidden characters or em dashes indicate AI?

They indicate a copy and paste, which is not the same claim. Text pasted from a chat window can carry characters that were never typed, and that says something about a clipboard rather than about authorship — a student may have drafted in one app and pasted into another. Detectors do not read them either way.

Is submitted work stored?

The document is attached to your account so you can return to it, and you can delete it whenever you like. It is not used to train anything and it is not submitted anywhere.