
[Image by Circe Denyer from source]
Not text, ideas. As peer reviewer, I have been flooded in the last two years with requests to review papers in philosophy which I suspect bear a LLM-heavy influence in their design and interpretation. But how to show this to the journal editors who send these papers out for review in good faith? Many such journals have an explicit policy of no LLM generated texts, but LLMs can be used to “polish” and “improve” and “clarify structure”, effectively introducing the LLM permission on the back door and undermining their own criteria with vague conditions. This permissive policy makes it very hard for a reader to discern between the LLM that merely “polished” the text and the one that fiddled with the very design and interpretation. We currently have AI-text generators, imperfect, I know, but we have zero ways of showing when an LLM introduced its mediocre reasoning into a text, ruining it.
Below is my modest attempt to showcase the marks of LLM thinking in a humanities text, but also works for any academic text that argues for a claim. The list is open, and I think I will keep adding to it.
So, here is my list of the markers of LLM reasoning that will ruin a paper irreversibly:
- A. vague and general claims about gaps in the scholarship out there. These are made in such a way that one cannot prove them nor disprove them. Example: There is research on the effects of X on Y, but no research has comprehensively explained this effect. Or: not in depth, insufficient study, etc. These claims contain a general diagnosis verdict (not enough, not comprehensive enough, not in depth) that is oddly subjective and has no metrics attached to it. what would a comprehensive treatment look like exactly? The paper has no answer, but it does claim that it is the one doing it while gesturing vaguely at the papers out there who fail to do that.
- B. Gesturing at unnamed lit review and sources. The papers out there that fall under the previous diagnosis (of lacking or failing to engage deeply enough with the topic) are not realy named. We should believe, in all honesty, that the author has read these papers, found them lacking in comprehensiveness and depth, and yet does not name them. We are supposed to trust that they’ve read and diagnosed the lack, then tossed the papers.
- Explanation about A and B: why do these occur? because LLMs are next-word predictors and they predict which phrases need to apper in a paper. LLMs noticed that their training dataset – published papers out there – always signalled its novelty and contribution, so the LLM knows it has to have these words in there. Contribution is easiest to claim by saying that nobody else out there has looked at X. But LLMs have no idea what is a significant or meaningful contribution, so they just throw these words on a page and hope it sticks.
- C. What about the actual proposal of the paper? the main idea, the contribution. Suppose it is novel and A and B really are the case. How does the paper frame its own contribution? If written by an LLM, it will make a tiny, almost surgical distinction between two views and frame that as big. For example: author X says supercalifragilistic is the case, and author Y says no, supercalifragilistic is totally not the case. The LLM-generated paper will give both views their due, disagree with both and combine them into a hybrid: actually, it’s supercalifragilistic+. Reality+, extended+, with some conditions added by the paper. it’s very tiny, and small contribution, and then it’s supposed to add unto both X and Y, disagree with them, but also continue their work. It’s a tiny renaming or agreeing with all that makes this paper show signs of the LLM writing. Why is this LLM? after all, analytic philosophers have been doing this for centuries, it is a mark of analytic phil to be precise and surgical to the point of irrelevance. But… the issue is why would anyone want to write that paper? who would set out to write a paper arguing that supercalifragilistic+ is actually the case? Either someone really into the supercalifragilistic debates, who is kept awake at night about the debates, rehearsing them in their mind, or someone who set out a paper, didn’t know what the main idea was and discovered it while writing. That can happen. But, if the paper was about the narrow debate on the supercalifragilistic issue, the author should know the debate, map it like a pro, and speak accordingly. If, however, one sees signs of points A and B above, no knowledge of the literature out there, and yet a minuscule contribution, then it’s probably an LLM. One cannot perform surgery before mastering the anatomy and using the right language of the field, knowing the names, mastering the debates. So this C always works in conjunction with A and B. If A and B are not present, it is possible that the paper started with an open question, a debate that interested the author, and then they could not solve it, so they settled for this tiny distinction. In this case, it’s not LLM.
- D. inventing terms. taxonomies. categories. words. If the whole paper hangs on a taxonomy that the author invented, and it is super-important that the audience picks that up, and this is the only major contribution of the paper, without showing how this taxonomy/ category clarifies previous discussions, it’s probably an LLM. Words are cheap, and for LLMs these are cheaper than for humans. Inventing words as the main contribution of a paper is really not enough, but LLMs don’t know that and, again, having been trained on papers that make conceptual distinctions and keep proposing terms, LLMs start to assume that this is what philosophy is about. But LLMs cannot draw the boundary between meaningful and meaningless contributions. They cannot see it or hear it. Because, as I have argued elsewhere, the main distinction between a human and an LLM is the lack of taste. Taste comes with mastery of your field. LLMs are helpless rookies that keep speaking and gesturing like professionals, mimicking the language of expertise in whatever field one unleashes them in.
All the signs I mentioned can also be attributed to inexperience and beginner mistakes. But all four combined are a good indication that an LLM messed up the study’s design and contribution, not just the polishing of the words.

Leave a Reply