Direkt zum Inhalt

Owner only: item control page
Restat, Valerie ; Diestelkämper, Indra ; Klettke, Meike ; Störl, Uta

FONDUE - Fine-Tuned Optimization: Nurturing Data Usability & Efficiency

Restat, Valerie, Diestelkämper, Indra, Klettke, Meike and Störl, Uta (2025) FONDUE - Fine-Tuned Optimization: Nurturing Data Usability & Efficiency. Journal of Big Data 12 (1).

Date of publication of this fulltext: 10 Jul 2025 09:04
Article
DOI to cite this document: 10.5283/epub.77126


Abstract

To provide good results and decisions in data-driven systems, data quality must be ensured as a primary consideration. An important aspect of this is data cleaning. Although many different algorithms and tools already exist for data cleaning, an end-to-end data quality solution is still needed. In this paper, we present FONDUE, our vision of a well-founded end-to-end data quality optimizer. In ...

To provide good results and decisions in data-driven systems, data quality must be ensured as a primary consideration. An important aspect of this is data cleaning. Although many different algorithms and tools already exist for data cleaning, an end-to-end data quality solution is still needed. In this paper, we present FONDUE, our vision of a well-founded end-to-end data quality optimizer. In contrast to many studies that consider data cleaning in the context of machine learning, our approach focuses on various scenarios, such as when preprocessing and downstream analysis are separated. As an adaptive and easily extendable framework, FONDUE operates similarly to proven methods of database query optimization. Analogously, it consists of the following parts: Rule-based optimization, where the appropriate data cleaning algorithms are selected based on use case constraints, optimizer hints in the form of best practices, and cost-based optimization, where the costs are measured in terms of data quality. Accordingly, the result is an optimized data cleaning pipeline. The choice of different optimization goals enables further flexibility, e.g. for environments with limited resources. As a first building block of FONDUE, we present CheDDaR, which is used to detect errors and measure data quality. Both are important tasks for improving data quality with FONDUE.



Involved Institutions


Details

Item typeArticle
Journal or Publication TitleJournal of Big Data
Publisher:Springer
Open Access Type:CC-License
Volume:12
Number of Issue or Book Chapter:1
Date23 May 2025
InstitutionsInformatics and Data Science > General computer science > Data Engineering (Prof. Dr.-Ing. Meike Klettke)
Identification Number
ValueType
10.1186/s40537-025-01158-xDOI
KeywordsData quality, Data cleaning, Optimization
Dewey Decimal Classification000 Computer science, information & general works > 004 Computer science
StatusPublished
RefereedYes, this version has been refereed
Created at the University of RegensburgPartially
URN of the UB Regensburgurn:nbn:de:bvb:355-epub-771261
Item ID77126

Export bibliographical data

Owner only: item control page

nach oben