How to Tell if a Text Was Written by AI: Signs and Checks
How to Tell if a Text Was Written by AI: Signs and Checks
With full certainty, you cannot. No single sign and no detector gives you proof, only a probability. But a set of signs plus a factual check does work: after five minutes of reading you can usually tell whether someone wrote with material in hand or a model wrote around a topic.
Below, in order: signs at the level of words, structure and meaning, why detectors get it wrong, how to check by hand, and how to put a language model to work on the review.
Can AI-written text be identified reliably?
A generative model picks the next word by probability, so generated text is on average more predictable than human text: fewer rare words, fewer awkward transitions, fewer personal decisions like a sudden one-sentence paragraph.
But predictability is not a signature. A person filling in a routine report or a product card produces the same even surface. And an AI draft that an editor has gone through and filled with facts is close to impossible to separate out - which is fine, because human work is already in it. So the practical question is not "human or machine" but "is there a human behind this who is answerable for the facts". The signs tell you where to look; the factual check and a conversation with the author decide.
Where this approach fails: short texts. On a 400-character paragraph any conclusion is guesswork. Signs become visible from roughly one or two pages up.
If you want to see the raw material for yourself, the same text models used for drafting are available in one place on Moleculs.ai - useful when you want to compare how a generated paragraph reads next to your own.
Word-level signs: filler phrases, clichés, repeated connectors
A single phrase means nothing. Density is what matters. Five or six items from this list on one page is already a signal.
- Empty connectors: "it is important to note", "it is worth understanding", "thus".
- Evaluative clichés with no content: "plays a key role", "a wide range of possibilities".
- Paired constructions: "not only, but also", "both in terms of A and in terms of B".
- Neat three-item runs: "convenient, fast and effective".
- Soft hedged generalities instead of a claim: "can significantly improve".
The second marker is monotone rhythm: generated sentences cluster around 12 to 18 words, while human spread is wider. The third is identical paragraph openings - three paragraphs in a row starting with noun plus verb. A live writer rarely does that; one paragraph opens with a question, the next with "this is where the trouble starts".
Structure-level signs: symmetric sections, lists of three, circular endings
Structure gives generation away more reliably than wording, because it is harder to reproduce by accident.
Symmetry. Every section is roughly the same length, three paragraphs each, three sentences per paragraph. People write unevenly: where they know more, the section is twice as long. Same goes for lists of exactly three items of equal length - if a text has four lists and all four have three bullets, that is not a coincidence.
Circular paragraphs. A paragraph opens with a claim, explains it, then restates the same thought in different words. Cut the last sentence of a paragraph: if nothing is lost, that was a machine loop.
A conclusion that retells the article. "So, we have looked at" followed by a recap of what you just read is a standard generated ending. A person adds something new at the end: a caveat, a piece of advice, a limit of applicability.
[[СХЕМА: три уровня проверки текста - words (clichés, rhythm), structure (symmetry, threes), meaning (generalities, no facts), and a separate block below them for "fact check and questions for the author" | Review order: from a quick read to digging into the facts]]
Meaning-level signs: empty generalities, no details, no sources
An honest test: retell a paragraph in one sentence so that it reads like news. If what comes out is "the tool is useful but there are nuances to consider", the paragraph is empty. What to look for specifically:
- Claims that are true of anything. "Saves time and cuts costs" fits a CRM and a vegetable garden equally well.
- Numbers with no source and no baseline: "30% more efficient" compared to what?
- Names and terms that get introduced once and never used again.
- Promises instead of description: "will take you to the next level".
- No negatives anywhere. A real author almost always says where the method fails.
A separate sign is uniform confidence. In human text you can see where the author knows firmly and where they are guessing; in generated text every claim, including the debatable ones, sounds equally level.
The main test: numbers, names, dates, first-hand experience
Checking the facts beats any stylistic sign, because editing the style does not get around it. A model with no access to your material does not know what your contractor charged or which day the launch slipped.
- Write out everything concrete: numbers, dates, names, titles, links, quotes. One page should yield at least three or four.
- Search for the two or three prettiest numbers. A round number with no source is the first candidate for invention.
- Check quotes and links. Nonexistent studies and dead links to "the company report" show up in generated text regularly.
- Find the places that require experience: how long it took, what broke, what was done wrong the first time. Zero such places means the author has not done the work.
The cost of a miss is easy to estimate: one invented number that slips into commercial copy means a correction after publication and an awkward conversation with the client. Half an hour of fact checking is cheaper.
Signs that point away from AI: typos, uneven rhythm, opinions
Things you rarely see in raw generated text:
- Typos, missing commas, accidentally doubled words.
- A thought that breaks off and gets picked up again two paragraphs later.
- A personal opinion you can argue with: "I dislike this approach, and here is why".
- Local specifics: a colleague's name, a neighbourhood, an equipment brand.
- Uneven rhythm and humour tied to the context, not a generic joke about robots.
Be careful with the reverse inference: typos can be added by hand, and a model will write a personal opinion if you ask it to. One such sign only lowers the default probability of generation.
AI text detectors: how they work and why they misfire
Detectors measure statistics: how predictable each next word is and how evenly that predictability is spread. Both standard errors follow from this. False positive: formulaic human text - an instruction, a legal paragraph, a translation - is statistically even, so the detector returns a high score, especially for a non-native writer. Miss: generation that a person rewrote and filled with facts lands in the "human" zone. Detectors are tuned mainly for English, and results for other languages are weaker still; the methodology is not shown to you, and two detectors will score the same passage differently.
The practical takeaway: a detector score is a hint about where to look, not a verdict. Failing, fining or refusing payment on the basis of a percentage is not defensible, because there is no proof in it.
| Method | What it gives | What it does not give |
|---|---|---|
| Detector | A fast signal about where to look closer | Proof; it errs in both directions |
| Manual signs | An understanding of what exactly is wrong | Certainty on short texts |
| Fact checking | A verifiable result: the fact holds or it is invented | An answer about who wrote the style |
| Questions for the author | The most reliable signal | Only works if you can reach the author |
A seven-step manual check
This order has been worked out on a stream of texts from contractors. It takes 10 to 15 minutes per page.
- Read the whole text without editing and mark the spots where you got bored. Generalities usually sit there.
- Count the filler connectors. More than five per page is a signal.
- Look at the skeleton: section lengths, number of bullets per list, number of paragraphs. Hunt for symmetry.
- Cut the last sentence of three random paragraphs. If nothing suffers, those were loops.
- Write out the facts and check two or three numbers and every link.
- Find at least one negative: a limit, a mistake, a case where the method does not work.
- Ask the author two questions about specific paragraphs: where this number came from and why this wording was chosen.
Step seven decides more than the first six. Whoever wrote it answers immediately and in detail; whoever handed in a generation retells the text in the same words. Reviewing a student submission works the same way: ask for two debatable paragraphs to be explained in their own words, with sources and edit history.
How to review a text with a language model and how to read its answer
A model answers "is this AI?" badly and "where is this empty?" well. Ask for an analysis by sign, not a score:
Role: you are an editor doing a final pass before publication.
Task: in the text below, find filler connectors, empty
generalities, repeated constructions, circular paragraphs and
claims with no supporting facts. Separately mark passages that
look generated and explain which sign led you there.
Output: a three-column table - quote from the text, what is
wrong, how to rewrite. No more than 15 rows.
Constraints: do not rewrite the whole text, do not judge the
topic, do not give an "AI probability" percentage, no praise.
Text:
<paste text>How to read the answer. Look at the quotes: each one must be a real line from your text. If the model quotes something that is not there, the analysis is unreliable and the request is worth repeating. Discard any remark without a quote. A useful second pass: ask which facts in the text cannot be verified from the text itself. You get a list for manual checking.
The prompt works on any decent text model - for example ChatGPT or Claude, both of which handle long passages and quoting reasonably well. The run cost is shown next to the model, and the quotas per plan are listed on the pricing page. For long material, split the review into parts: one page at a time is more accurate than a whole article at once.
[[СКРИНШОТ: text model with the editor prompt pasted in | Input field with the review prompt and the name of the selected text model]]
What to do if the text really was generated: edit, do not ban
A generated draft is not defective in itself. What is defective is a text with no facts and no human answerable for it. Fixing it is cheaper than sending it back with "redo this, the detector said 87%".
- Conclusion: remove the recap, put in a thought the text has not stated yet.
- Generalities: replace every "significantly improves efficiency" with a number, a timeframe or an example.
- Facts: add what only the author knows - project numbers, names, dates, what broke.
- Filler connectors: delete them. The text gets shorter and clearer.
- Rhythm: break up long sentences, leave a short one standing alone somewhere.
- Limits: write a paragraph about where the approach does not work.
If you are the one writing with a model, the mistake is usually at the input: the prompt contained a topic, not material. Give it facts, your own examples and a ready structure, and you will get fewer generalities.
Images follow similar logic: generated pictures give themselves away through hands, text on signs and light that is too clean. Working through those signs is a separate subject - you can see what current image models produce on the Nano Banana page - but the principle holds: check the specifics, not the general impression.

Corporate access to AI models
Invoice for legal entities, centralised payment, priority support
- Access to ChatGPT, Gemini, Grok, Claude and DeepSeek
- Prompt library and shared access inside the team