Posthoc Verification and the Fallibility of the Ground Truth
The work
| Authors | Yifan Ding; Nicholas Botzer; Tim Weninger |
|---|---|
| Editors | |
| Type | inproceedings |
| Year | 2022 |
| Citekey | ding2022posthoc |
Where it appeared
| Published in | Proceedings of the First Workshop on Dynamic Adversarial Data Collection |
|---|---|
| Publisher | Association for Computational Linguistics |
| Pages | 23--29 |
Identifiers
| DOI | 10.18653/v1/2022.dadc-1.3 |
|---|---|
| OpenAlex | W3170137791 |
Access
| Landing page | https://doi.org/10.18653/v1/2022.dadc-1.3 |
|---|---|
| Free full text | https://aclanthology.org/2022.dadc-1.3.pdf |
Related
| Distinct from | ding2021posthoc The published version and its preprint. ding2021posthoc is arXiv:2106.07353v1 of 2 June 2021; this is the version of record, in the Proceedings of the First Workshop on Dynamic Adversarial Data Collection, 2022, same three authors and same title. Created from the doi recorded on the preprint's note by an earlier reading, not from a reconstruction. The corpus has no link kind for preprint-and-published -- issue 122 -- so distinct-from carries it, as it does for the two Piedeleu records. |
|---|
Abstract
Classifiers commonly make use of preannotated datasets, wherein a model is evaluated by pre-defined metrics on a held-out test set typically made of human-annotated labels. Metrics used in these evaluations are tied to the availability of well-defined ground truth labels, and these metrics typically do not allow for inexact matches. These noisy ground truth labels and strict evaluation metrics may compromise the validity and realism of evaluation results. In the present work, we conduct a systematic label verification experiment on the entity linking (EL) task. Specifically, we ask annotators to verify the correctness of annotations after the fact (i.e., posthoc). Compared to pre-annotation evaluation, state-of-the-art EL models performed extremely well according to the posthoc evaluation methodology. Surprisingly, we find predictions from EL models had a similar or higher verification rate than the ground truth. We conclude with a discussion on these findings and recommendations for future evaluations. The source code, raw results, and evaluation scripts
A copy is held
pdf, 284.1 kB. Not published — it may be under copyright. The facts and links here are.
How it got here
| How it got here | agent via crossref |
|---|---|
| Added | 2026-09-02 07:14 UTC |
| Approved by | a person 2026-09-02 08:55 UTC |
Cite it as
@inproceedings{ding2022posthoc,
title = {Posthoc Verification and the Fallibility of the Ground Truth},
author = {Yifan Ding and Nicholas Botzer and Tim Weninger},
year = {2022},
booktitle = {Proceedings of the First Workshop on Dynamic Adversarial Data Collection},
publisher = {Association for Computational Linguistics},
pages = {23--29},
doi = {10.18653/v1/2022.dadc-1.3},
}
This record lives at https://refs.drheap.org/ding2022posthoc/ and will keep doing so.