The copy is not the work
A subject the papers are about. The loosest grouping, and the one to reach for last.
Records whose held file is a different document from the work described: an arXiv preprint, a technical report, a manuscript, a reprint typesetting, a fragment standing in for the published version -- or, the other way about, a whole bound issue or proceedings standing in for one paper inside it. The record is right in each case and the file is not, so nothing here should be paginated or quoted as the published text.
At 141 records this is the corpus's largest set and it is a defect register rather than a subject. What it is for is to stop the same file being trusted twice.
The kinds, which want different remedies:
*Wrong version.* The bulk of it. A preprint, extended version, technical report or manuscript standing in for the version of record. Harmless to read, wrong to cite, and the page numbers will not match.
*The volume, not the item.* A bound issue or proceedings where one paper was wanted. These are the ones the queue surfaces on its own, under 'the copy is the volume, not the work' -- currently fifteen, not the six this description used to name; the number moves as copies arrive and the old figure was not revised with it. hartley1928transmission is the clearest case, 29 declared pages against a 267-page file.
*Truncation.* The file stops before the work does.
*A text layer that is dead, or worse, lying.* Some scans extract nothing. A smaller and more dangerous group extracts fluent prose that is not what the page says -- characters dropped from line ends, or a substitution cipher in the embedded font. Those produce quotable, plausible, wrong text, and they cannot be caught by reading the extraction alone.
How the categories were found, since the methods transfer. Page count against the record's declared range catches the volume cases. Tokens per page separates a healthy text layer from a dead one; a dictionary hit rate against a word list catches the lying ones, which score far below the 0.87 or so that honest English extraction gives. Page aspect ratio near 16:9 identifies a slide deck filed as a paper. None of this requires reading the document, which is why it scales, and all of it requires opening the file, which is why no catalogue supplies it.