| Published Version Download ( PDF | 444kB) | License: Creative Commons Attribution 4.0 |
Kurz erklärt: Measuring Data Changes in Data Engineering and their Impact on Explainability and Algorithm Fairness
Klettke, Meike
, Lutsch, Adrian and Störl, Uta
(2021)
Kurz erklärt: Measuring Data Changes in Data Engineering and their Impact on Explainability and Algorithm Fairness.
Datenbank-Spektrum 21 (3), pp. 245-249.
Date of publication of this fulltext: 13 Aug 2025 06:57
Article
DOI to cite this document: 10.5283/epub.77290
Abstract
Data engineering is an integral part of any data science and ML process. It consists of several subtasks that are performed to improve data quality and to transform data into a target format suitable for analysis. The quality and correctness of the data engineering steps is therefore important to ensure the quality of the overall process. In machine learning processes requirements such as ...
Data engineering is an integral part of any data science and ML process. It consists of several subtasks that are performed to improve data quality and to transform data into a target format suitable for analysis. The quality and correctness of the data engineering steps is therefore important to ensure the quality of the overall process.
In machine learning processes requirements such as fairness and explainability are essential. The answers to these must also be provided by the data engineering subtasks. In this article, we will show how these can be achieved by logging, monitoring and controlling the data changes in order to evaluate their correctness. However, since data preprocessing algorithms are part of any machine learning pipeline, they must obviously also guarantee that they do not produce data biases.
In this article we will briefly introduce three classes of methods for measuring data changes in data engineering and present which research questions still remain unanswered in this area.
Alternative links to fulltext
Involved Institutions
Details
| Item type | Article | ||||
| Journal or Publication Title | Datenbank-Spektrum | ||||
| Publisher: | Springer Nature | ||||
|---|---|---|---|---|---|
| Open Access Type: | CC-License | ||||
| Volume: | 21 | ||||
| Number of Issue or Book Chapter: | 3 | ||||
| Page Range: | pp. 245-249 | ||||
| Date | October 2021 | ||||
| Institutions | Informatics and Data Science > General computer science > Data Engineering (Prof. Dr.-Ing. Meike Klettke) | ||||
| Identification Number |
| ||||
| Keywords | Data engineering pipelines · Reliability · Explainability · Data bias · Degree of data changes | ||||
| Dewey Decimal Classification | 000 Computer science, information & general works > 004 Computer science | ||||
| Status | Published | ||||
| Refereed | Yes, this version has been refereed | ||||
| Created at the University of Regensburg | No | ||||
| URN of the UB Regensburg | urn:nbn:de:bvb:355-epub-772900 | ||||
| Item ID | 77290 |
Download Statistics
Download Statistics