Few emails frighten authors more than one saying their manuscript "returned a high similarity score". The panic is usually unnecessary, because the number by itself means very little. Similarity software does not detect plagiarism; it detects matching text. Understanding the difference — and how an editor actually reads the report — turns a scary percentage into a manageable checklist.
What the software actually does
Screening tools compare your manuscript against a large corpus of published articles, web pages, and often previously submitted student and author work. They return two things: an overall similarity index, and a colour-coded report showing which passages matched which sources. That is all. The software makes no judgement about intent, attribution, or whether a match is legitimate.
Why the headline percentage misleads
A high score is often entirely innocent, because these routinely match:
- The reference list — dozens of citations that are identical everywhere by design.
- Standard methods — "samples were centrifuged at 3,000 g for 10 minutes" has only so many phrasings.
- Common terminology and institutional names, ethics-statement boilerplate, and instrument descriptions.
- Your own preprint or thesis, which may match at 90% and is perfectly legitimate.
Conversely, a low overall score can hide a serious problem: a single copied paragraph in your discussion, taken verbatim from another author, may register a few per cent — and it is still plagiarism.
Editors do not read the number. They open the report and look at where the matches are, how long the matched strings are, and whether they are attributed.
How to read your own report
- Exclude the bibliography and quotations first — most tools can do this, and it immediately clears the noise.
- Look at the longest single match. One 200-word run against a single source matters more than fifty scattered five-word phrases.
- Check where the matches sit. Methods matches are usually forgivable; matches in the introduction, discussion or conclusions — the parts that should be your own thinking — are not.
- Identify the source. Matching your own earlier paper raises text-recycling questions; matching someone else's raises plagiarism ones.
Fixing it properly
Where the text is genuinely someone else's, the fix is not to shuffle synonyms until the software stops flagging it — that is still plagiarism, just harder to detect. Either quote it explicitly and attribute it, or close the source, understand the idea, and write it in your own words. For unavoidable standard phrasing in methods, cite the original protocol paper.
Prevent it while you write
Most similarity problems begin as sloppy note-taking: pasting sentences into a draft "to rewrite later" and forgetting which were yours. Keep quoted material clearly marked from the start, record the source with every note, and see our guide on types of plagiarism for the boundaries that catch people out.