The500Feed.Live

Everything going on in AI - updated daily from 500+ sources

← Back to The 500 Feed
📄 ResearchJuly 23, 2026

Impact of Inaccurately Labeled Data on the Performance of Multi-label Classification for Disease Recognition

The process of medical diagnostics is challenging, especially since patients can simultaneously suffer from several diseases with similar, contradictory, or even opposing diagnoses. Statistical prediction can support physicians in this task; however, the quality of data used for predicition as well as the chosen statistical model can affect the reliability of data-driven decision support. Data quality can, in particular, be reduced by incomplete medical diagnoses, that is, the termination of the diagnostic process once a patient has tested positive for one disease that explains the symptoms. When interpreting missing diagnoses as negative, this leads to potentially false negative health data. Another source of low data quality lies in diagnoses being made through a principle of elimination, i.e., after several negative results, one opts for the seemingly last remaining possibility. This may lead to false positive health data. In our work, we investigate how such inaccurately labeled data affects the predictive ability of multi-label classification (MLC) for disease recognition. Unlike single-label classification (SLC), MLC allows the simultaneous assignment of multiple diseases to a patient and can therefore describe clinical conditions more holistically. To that end, we conduct a synthetic-data simulation study as well as a real-data case study on the example of chronic pain patients. In this regard, we compare MLC performance on accurately and inaccurately labeled data. We manipulate the data such that it corresponds to different diagnostic test sensitivities as well as to different examination sequences, thus paying special attention to resulting uncertainty within the process of medical diagnostics. Our results show that inaccurate labeling substantially decreases MLC prediction ability. Furthermore, low diagnostic test-sensitivity, the order of disease examination and covariate effects have a strong impact on MLC performance. These findings contribute to a better understanding of the interplay and impact of diagnostic procedures, data documentation and interpretation, and statistical modeling. This underlines the need for careful data collection as a basis for model development; special consideration should be given to the extensive examination of patients as well as the targeted collection of covariates. This is particularly crucial when models are transferred into everyday clinical practice.

Read Original Article →

Source

https://www.medrxiv.org/content/10.64898/2026.07.22.26358665v1?rss=1