Skip to Main content Skip to Navigation
Conference papers

ADESIT: Visualize the Limits of your Data in a Machine Learning Process

Pierre Faure-Giovagnoli 1, 2 Marie Le Guilly 2 Jean-Marc Petit 2 Vasile-Marian Scuturici 2 
2 BD - Base de Données
LIRIS - Laboratoire d'InfoRmatique en Image et Systèmes d'information
Abstract : Thanks to the numerous machine learning tools available to us nowadays, it is easier than ever to derive a model from a dataset in the frame of a supervised learning problem. However, when this model behaves poorly compared with an expected performance, the underlying question of the existence of such a model is often underlooked and one might just be tempted to try different parameters or just choose another model architecture. This is why the quality of the learning examples should be considered as early as possible as it acts as a go/no go signal for the following potentially costly learning process. With ADESIT, we provide a way to evaluate the ability of a dataset to perform well for a given supervised learning problem through statistics and visual exploration. Notably, we base our work on recent studies proposing the use of functional dependencies and specifically counterexample analysis to provide dataset cleanliness statistics but also a theoretical upper bound on the prediction accuracy directly linked to the problem settings (measurement uncertainty, expected generalization...). In brief, ADESIT is intended as a go/no go step right after data selection and right before the machine learning process itself. With further analysis for a given problem, the user can characterize, clean and export dynamically selected subsets, allowing to better understand what regions of the data could be refined and where the data precision must be improved by using, for example, new or more precise sensors.
Complete list of metadata

https://hal.archives-ouvertes.fr/hal-03242380
Contributor : Pierre Faure--Giovagnoli Connect in order to contact the contributor
Submitted on : Thursday, September 2, 2021 - 10:06:06 AM
Last modification on : Saturday, September 4, 2021 - 3:31:09 AM

File

ADESIT_camReady_round2.pdf
Publisher files allowed on an open archive

Identifiers

  • HAL Id : hal-03242380, version 2

Citation

Pierre Faure-Giovagnoli, Marie Le Guilly, Jean-Marc Petit, Vasile-Marian Scuturici. ADESIT: Visualize the Limits of your Data in a Machine Learning Process. International Conference on Very Large Data Bases, Aug 2021, Copenhague, Denmark. ⟨hal-03242380v2⟩

Share

Metrics

Record views

160

Files downloads

100