- AutorIn
- Dr.-Ing. Julian Eberius Technische Universität Dresden, Fakultät Informatik, Institut für Systemarchitektur, Professur für Datenbanken
- Dr.-Ing. Maik ThieleTechnische Universität Dresden, Fakultät Informatik, Institut für Systemarchitektur, Professur für Datenbanken
- Dr.-Ing. Katrin BraunschweigTechnische Universität Dresden, Fakultät Informatik, Institut für Systemarchitektur, Professur für Datenbanken
- Prof. Dr.-Ing. Wolfgang Lehner
- Titel
- Top-k Entity Augmentation using Consistent Set Covering
- Zitierfähige Url:
- https://nbn-resolving.org/urn:nbn:de:bsz:14-qucosa2-806674
- Konferenz
- SSDBM 2015: International Conference on Scientific and Statistical Database Management. La Jolla, 29. Juni - 01. Juli 2015
- Quellenangabe
- SSDBM '15: Proceedings of the 27th International Conference on Scientific and Statistical Database Management
Herausgeber: Amarnath Gupta
Herausgeber: Susan Rathbun
Erscheinungsort: New York
Verlag: ACM
Erscheinungsjahr: 2015
ISBN: 978-1-4503-3709-0
Artikelnummer: 8 - Erstveröffentlichung
- 2015
- Abstract (EN)
- Entity augmentation is a query type in which, given a set of entities and a large corpus of possible data sources, the values of a missing attribute are to be retrieved. State of the art methods return a single result that, to cover all queried entities, is fused from a potentially large set of data sources. We argue that queries on large corpora of heterogeneous sources using information retrieval and automatic schema matching methods can not easily return a single result that the user can trust, especially if the result is composed from a large number of sources that user has to verify manually. We therefore propose to process these queries in a Top-k fashion, in which the system produces multiple minimal consistent solutions from which the user can choose to resolve the uncertainty of the data sources and methods used. In this paper, we introduce and formalize the problem of consistent, multi-solution set covering, and present algorithms based on a greedy and a genetic optimization approach. We then apply these algorithms to Web table-based entity augmentation. The publication further includes a Web table corpus with 100M tables, and a Web table retrieval and matching system in which these algorithms are implemented. Our experiments show that the consistency and minimality of the augmentation results can be improved using our set covering approach, without loss of precision or coverage and while producing multiple alternative query results.
- Andere Ausgabe
- Link zum Artikel, der zuerst in der ACM Digital Library erschienen ist.
DOI: 10.1145/2791347.2791353 - Freie Schlagwörter (DE)
- heterogene Quellen abfragen, Information Retrieval, automatische Schema-Matching-Methoden, Top-k-Methode, Algorithmen
- Freie Schlagwörter (EN)
- Query heterogeneous sources, information retrieval, automatic schema matching methods, top-k method, algorithms
- Klassifikation (DDC)
- 004
- Verlag
- ACM, New York
- Version / Begutachtungsstatus
- angenommene Version / Postprint / Autorenversion
- URN Qucosa
- urn:nbn:de:bsz:14-qucosa2-806674
- Veröffentlichungsdatum Qucosa
- 19.09.2022
- Dokumenttyp
- Konferenzbeitrag
- Sprache des Dokumentes
- Englisch
- Lizenz / Rechtehinweis