| Published Version Download ( PDF | 2MB) | License: Creative Commons Attribution 4.0 |
Handling imperfection: A taxonomy for machine learning on data with data quality defects
Hagn, Michael, Heinrich, Bernd
, Krapf, Thomas and Schiller, Alexander
(2025)
Handling imperfection: A taxonomy for machine learning on data with data quality defects.
Decision Support Systems 196, p. 114493.
Date of publication of this fulltext: 24 Jun 2025 05:57
Article
DOI to cite this document: 10.5283/epub.76908
Abstract
In recent years, machine learning (ML) has become ubiquitous in sectors including transportation, security, health, and finance to analyze large amounts of data and support decision-making. However, real-world datasets used in ML often exhibit various data quality (DQ) defects that can significantly impair the performance and validity of ML models and thus also the decisions derived from them. ...
In recent years, machine learning (ML) has become ubiquitous in sectors including transportation, security, health, and finance to analyze large amounts of data and support decision-making. However, real-world datasets used in ML often exhibit various data quality (DQ) defects that can significantly impair the performance and validity of ML models and thus also the decisions derived from them. Therefore, a plethora of methods across various research strands have been proposed to address DQ defects and mitigate their negative impact on ML-based data analysis and decision support. This has resulted in a fragmented research landscape, where comparisons and classifications of methods dealing with ML on data with DQ defects are very challenging for both researchers and practitioners. Thus, based on a structured design process, we develop and present a taxonomy for this research field. The taxonomy serves as a systematic framework to classify and organize existing research and methods according to relevant dimensions and facilitates future work in this area. Its reliability, understandability, completeness, and usefulness are supported by an evaluation with external researchers and practitioners. Finally, we identify current trends and research gaps and derive challenges and directions for future research.
Alternative links to fulltext
Involved Institutions
Details
| Item type | Article | ||||
| Journal or Publication Title | Decision Support Systems | ||||
| Publisher: | Elsevier | ||||
|---|---|---|---|---|---|
| Open Access Type: | DEAL (Elsevier) | ||||
| Volume: | 196 | ||||
| Page Range: | p. 114493 | ||||
| Date | 16 June 2025 | ||||
| Institutions | Business, Economics and Information Systems > Institut für Wirtschaftsinformatik > Lehrstuhl für Wirtschaftsinformatik II (Prof. Dr. Bernd Heinrich) Informatics and Data Science > Department Information Systems > Lehrstuhl für Wirtschaftsinformatik II (Prof. Dr. Bernd Heinrich) | ||||
| Projects |
Funded by:
Deutsche Forschungsgemeinschaft (DFG)
(494840328)
| ||||
| Identification Number |
| ||||
| Keywords | Taxonomy, Machine learning, Data quality, Data uncertainty | ||||
| Dewey Decimal Classification | 000 Computer science, information & general works > 004 Computer science | ||||
| Status | Published | ||||
| Refereed | Yes, this version has been refereed | ||||
| Created at the University of Regensburg | Yes | ||||
| URN of the UB Regensburg | urn:nbn:de:bvb:355-epub-769086 | ||||
| Item ID | 76908 |
Download Statistics
Download Statistics