- AutorIn
- Prof. Dr.-Ing. Wolfgang Lehner Technische Universität Dresden, Fakultät Informatik, Institut für Systemarchitektur, Professur Datenbanken
- Ahmad AhmadovTechnische Universität Dresden, Fakultät Informatik, Institut für Systemarchitektur, Dresden Database Research Group
- Dr.-Ing. Maik ThieleTechnische Universität Dresden, Fakultät Informatik, Institut für Systemarchitektur, Dresden Database Research Group
- Dr.-Ing. Julian Eberius
- Robert Wrembel
- Titel
- Towards a Hybrid Imputation Approach Using Web Tables
- Zitierfähige Url:
- https://nbn-resolving.org/urn:nbn:de:bsz:14-qucosa2-820994
- Konferenz
- 2015 IEEE/ACM 2nd International Symposium on Big Data Computing (BDC). Limassol, 07.-10.12.2015
- Quellenangabe
- 2015 IEEE/ACM 2nd International Symposium on Big Data Computing (BDC). Proceedings
Erscheinungsort: New York
Verlag: IEEE
Erscheinungsjahr: 2015
Seiten: 21-30 - Erstveröffentlichung
- 2015
- Abstract (EN)
- Data completeness is one of the most important data quality dimensions and an essential premise in data analytics. With new emerging Big Data trends such as the data lake concept, which provides a low cost data preparation repository instead of moving curated data into a data warehouse, the problem of data completeness is additionally reinforced. While traditionally the process of filling in missing values is addressed by the data imputation community using statistical techniques, we complement these approaches by using external data sources from the data lake or even the Web to lookup missing values. In this paper we propose a novel hybrid data imputation strategy that, takes into account the characteristics of an incomplete dataset and based on that chooses the best imputation approach, i.e. either a statistical approach such as regression analysis or a Web-based lookup or a combination of both. We formalize and implement both imputation approaches, including a Web table retrieval and matching system and evaluate them extensively using a corpus with 125M Web tables. We show that applying statistical techniques in conjunction with external data sources will lead to a imputation system which is robust, accurate, and has high coverage at the same time.
- Andere Ausgabe
- Link zum Artikel, der zuerst in der IEEE Xplore Digital Library erschienen ist.
DOI: 10.1109/BDC.2015.38 - Freie Schlagwörter (DE)
- Internetanalyse, Datenvorverarbeitung, maschinelles Lernen
- Freie Schlagwörter (EN)
- Web mining, Data preprocessing, Machine learning
- Klassifikation (DDC)
- 004
- Verlag
- IEEE, New York
- Förder- / Projektangaben
- Erasmus Mundus Association Erasmus Mundus Joint Doctorate
Information Technologies for Business Intelligence - Doctoral College
(IT4BI-DC) - Version / Begutachtungsstatus
- angenommene Version / Postprint / Autorenversion
- URN Qucosa
- urn:nbn:de:bsz:14-qucosa2-820994
- Veröffentlichungsdatum Qucosa
- 12.01.2023
- Dokumenttyp
- Konferenzbeitrag
- Sprache des Dokumentes
- Englisch
- Lizenz / Rechtehinweis