Chinese AI Detector
In Chinese academia the check that matters is 知网 — CNKI, the database every thesis already passes through — which now reports an AIGC rate alongside its similarity score. An entire industry has grown around lowering that number, and the most widely circulated technique is worth understanding before you use it: adding charts, tables and formulas is reported to cut detection weight substantially, because those are not prose and prose is what gets scored. The percentage falls and not one sentence has changed. Paste your text and we will mark the phrases, with a reason for each.
CNKI is the check that decides things, and an industry exists to game it.
Understanding the Chinese situation means starting with who is doing the checking, because it is not the tools that dominate English-language coverage.
知网 already had the institutional position. CNKI is the database Chinese theses pass through for similarity checking, so when it added AIGC detection the verdict arrived inside a process universities were running anyway. That distribution is what gives a number weight — the same reason Turnitin matters in the anglophone world regardless of how accurate it is.
It analyses language patterns rather than matching keywords. By its own account it applies deep-learning analysis across several dimensions, which means the old advice about swapping individual words was never going to work against it for long.
And around it sits a whole commercial layer. PaperPass, 零感AI, BunnyScholar, 68爱写 and others sell 降AIGC率 — lowering your AIGC rate — with optimisation aimed at CNKI specifically. That is a narrower product than it appears: tools tuned to one system’s current behaviour work until that system changes.
The advice circulating alongside them splits in two. Some of it is ordinary editing with a target attached — vary your sentence length, add concrete detail, replace characteristic vocabulary — and that genuinely improves writing whether or not it moves a number. The rest is the denominator trick in the next section, which improves nothing.
Why adding charts lowers the percentage.
The most widely shared 降AIGC率 technique is to add tables, figures, formulas and code, with reported reductions in detection weight of a fifth to two fifths. It works. It is also not what people think it is, and knowing the mechanism changes whether you should rely on it.
Detection scores prose, and only prose. A formula is not prose. A data table is not prose. A flowchart is not prose. None of it can be assessed for linguistic predictability, so none of it enters the calculation.
So adding them shrinks what is being measured. The same generated paragraphs become a smaller fraction of a larger document, and the reported percentage falls accordingly. Nothing about the writing improved. The measurement simply covers less of the thing.
Which means the number stops describing the document. A thesis at eight per cent because it is mostly tables tells a reader far less than one at eight per cent that is entirely prose, and no report distinguishes them. Anyone making a decision on that figure is comparing quantities that are not comparable.
The honest version: if the charts and data belong in the work, add them because they belong, and the score falling is a side effect. If they are being added to move a number, the underlying prose is unchanged and a supervisor reading it will notice what a percentage did not.
What makes Chinese genuinely hard to measure.
Underneath the institutional question sits a technical one, and it applies to every tool including ours.
- No spaces, so the words have to be guessed firstChinese is written as a continuous run of characters. Any tool measuring word-level predictability must segment the text before it can measure anything, and segmentation is itself a statistical guess. Its errors propagate into every number downstream. This uncertainty sits underneath a Chinese score before the text is read at all.
- Tokenisation closer to the character than the wordA model breaks Chinese into units nearer to single characters than to words, so the probability distribution it works from has a different shape from English. A threshold carried over from English is not stricter or looser here; it is measuring a different quantity.
- Chengyu are fixed by definition and correct by conventionFour-character idioms mark educated written Chinese. They are also, unavoidably, predictable multi-character sequences. A measure that penalises fixed patterns penalises the exact feature signalling a competent writer, which is a perverse outcome nobody designed and everybody inherits.
- Simplified and traditional are different character setsMainland and Singapore use simplified; Taiwan, Hong Kong and Macau use traditional, with vocabulary diverging alongside — 软件 against 軟體, 网络 against 網路. A document mixing them is visible to any reader, and conversion is not purely mechanical since some simplified forms map to more than one traditional character.
What marks generated Chinese, without any tool.
Chinese sources call it AI味 — the taste or flavour of AI. It is recognisable before any score is involved, and these are the components.
首先、其次、最后 imposed on something that was not three-part. The strongest structural marker. Generated Chinese reaches for enumeration as a substitute for having a structure, producing an essay-shaped answer to a question that was not an essay. If the three items would not survive being written as continuous argument, they were one point split into three.
值得注意的是, 需要指出的是. Announcing that something is worth noting rather than noting it. Real writing uses these occasionally; generated writing uses them as paragraph hinges.
此外 and the other connectives, on a schedule. Chinese sources cataloguing AI writing patterns name 此外 specifically as overused. The tell is frequency and position — heading consecutive paragraphs — rather than the word itself.
总之, 综上所述 closing by restatement. A conclusion that repeats the piece rather than concluding it. Chinese writing can legitimately close this way, and generated Chinese does it invariably.
And uniform sentence length. Chinese handles very short sentences well and generated Chinese rarely produces one. A page with no brief sentence anywhere in it is flat in a way readers feel without naming.
What our check does on Chinese.
No percentage, and on Chinese that matters more than usual — a percentage here is of qualifying prose rather than of your document, and adding a table changes it. What we mark is the specific phrases that read as generated, each with a reason.
The named signal list is English — delve, tapestry, in the realm of. On Chinese you get the reading rather than the dictionary match, which is why the markers above are written out: those are the checks to run yourself.
We cannot tell you what CNKI will report. Nothing outside CNKI can. It is a different model with different training and it is not available to third parties, exactly as Turnitin is not.
Hidden characters work identically, and character-set mixing is worth a separate look: a document containing both simplified and traditional forms has been assembled from more than one source.
One thing that is not a caveat. The language is a setting rather than a guess: arriving from this page puts the tool in Chinese, and everything it gives back — the reasons, the replacements, the rewrite — comes back in Chinese. It does not answer you in English about your Chinese.
Chinese AI detector questions.
What does 知网 AIGC detection actually do?
CNKI runs a deep-learning analysis of language patterns rather than matching keywords, reporting across several dimensions. Because it is the database Chinese theses already pass through for similarity checking, its AIGC verdict arrives inside a process that institutions were using anyway — which is what gives it weight, in the same way Turnitin has weight in the anglophone world.
Does adding charts and formulas really lower an AIGC rate?
It lowers the reported percentage, and it is worth understanding why before treating that as a result. Detection scores prose; tables, formulas, code and figures are not prose. Adding them shrinks the amount of text being measured, so the same generated writing becomes a smaller share of a larger document. The figure moves and the writing has not changed.
Is the 降AIGC率 industry selling anything real?
It is selling optimisation against one specific system, which is a narrower thing than it sounds. Tools tuned to CNKI's current behaviour work until CNKI changes, and the advice circulating alongside them — replace characteristic words, add subjective expressions, vary sentence length — is ordinary editing advice with a target attached. The editing helps. The targeting expires.
Why is Chinese harder to measure than English?
Because it has no spaces between words. Every tool must segment the text before measuring it, segmentation is a statistical guess, and its errors carry into everything computed afterwards. That uncertainty is present in any Chinese score and no interface reports it.
Do chengyu raise the score?
They can, and it is the clearest case of a measure working against the language. Four-character idioms are a marker of educated written Chinese and are by definition fixed sequences, so a system penalising predictable patterns penalises the feature that signals a good writer.
What can I check by eye?
首先/其次/最后 imposed on an argument that was not three-part — generated Chinese uses enumeration as a substitute for structure. 值得注意的是 announcing significance instead of demonstrating it. 在当今快速发展的时代 as an opener. And 总之 or 综上所述 closing by restating what was already said.
My thesis was flagged and I wrote it. What now?
Ask what the percentage is of, since the figure covers qualifying prose rather than your whole document. Then show the process: drafts, version history, notes, data. Chinese academic supervision is close enough that your supervisor has usually seen the work develop, and that account is stronger than any counter-check.
Is my text stored?
The document is attached to your account so you can return to it, and you can delete it whenever you like. It is not used to train anything and it is not submitted anywhere.