Extraction and integration of partially overlapping web sources

Mirko Bronzi; Valter Crescenzi; Paolo Merialdo; Paolo Papotti

doi:10.14778/2536206.2536209

Extraction and integration of partially overlapping web sources

Mirko Bronzi, Valter Crescenzi, Paolo Merialdo, Paolo Papotti

Research output: Contribution to journal › Conference article › peer-review

49 Scopus citations

Abstract

We present an unsupervised approach for harvesting the data exposed by a set of structured and partially overlapping data-intensive web sources. Our proposal comes within a formal framework tackling two problems: the data extraction problem, to generate extraction rules based on the input websites, and the data integration problem, to integrate the extracted data in a unified schema. We introduce an original algorithm, WEIR, to solve the stated problems and formally prove its correctness. WEIR leverages the overlapping data among sources to make better decisions both in the data extraction (by pruning rules that do not lead to redundant information) and in the data integration (by reflecting local properties of a source over the mediated schema). Along the way, we characterize the amount of redundancy needed by our algorithm to produce a solution, and present experimental results to show the benefits of our approach with respect to existing solutions.

Original language	English (US)
Pages (from-to)	805-816
Number of pages	12
Journal	Proceedings of the VLDB Endowment
Volume	6
Issue number	10
DOIs	https://doi.org/10.14778/2536206.2536209
State	Published - Aug 2013
Externally published	Yes
Event	39th International Conference on Very Large Data Bases, VLDB 2012 - Trento, Italy Duration: Aug 26 2013 → Aug 30 2013

ASJC Scopus subject areas

Computer Science (miscellaneous)
General Computer Science

Access to Document

10.14778/2536206.2536209

Cite this

@article{da09c9020aa543b2a2954644aa7ef778,

title = "Extraction and integration of partially overlapping web sources",

abstract = "We present an unsupervised approach for harvesting the data exposed by a set of structured and partially overlapping data-intensive web sources. Our proposal comes within a formal framework tackling two problems: the data extraction problem, to generate extraction rules based on the input websites, and the data integration problem, to integrate the extracted data in a unified schema. We introduce an original algorithm, WEIR, to solve the stated problems and formally prove its correctness. WEIR leverages the overlapping data among sources to make better decisions both in the data extraction (by pruning rules that do not lead to redundant information) and in the data integration (by reflecting local properties of a source over the mediated schema). Along the way, we characterize the amount of redundancy needed by our algorithm to produce a solution, and present experimental results to show the benefits of our approach with respect to existing solutions.",

author = "Mirko Bronzi and Valter Crescenzi and Paolo Merialdo and Paolo Papotti",

year = "2013",

month = aug,

doi = "10.14778/2536206.2536209",

language = "English (US)",

volume = "6",

pages = "805--816",

journal = "Proceedings of the VLDB Endowment",

issn = "2150-8097",

publisher = "Very Large Data Base Endowment Inc.",

number = "10",

note = "39th International Conference on Very Large Data Bases, VLDB 2012 ; Conference date: 26-08-2013 Through 30-08-2013",

}

TY - JOUR

T1 - Extraction and integration of partially overlapping web sources

AU - Bronzi, Mirko

AU - Crescenzi, Valter

AU - Merialdo, Paolo

AU - Papotti, Paolo

PY - 2013/8

Y1 - 2013/8

N2 - We present an unsupervised approach for harvesting the data exposed by a set of structured and partially overlapping data-intensive web sources. Our proposal comes within a formal framework tackling two problems: the data extraction problem, to generate extraction rules based on the input websites, and the data integration problem, to integrate the extracted data in a unified schema. We introduce an original algorithm, WEIR, to solve the stated problems and formally prove its correctness. WEIR leverages the overlapping data among sources to make better decisions both in the data extraction (by pruning rules that do not lead to redundant information) and in the data integration (by reflecting local properties of a source over the mediated schema). Along the way, we characterize the amount of redundancy needed by our algorithm to produce a solution, and present experimental results to show the benefits of our approach with respect to existing solutions.

AB - We present an unsupervised approach for harvesting the data exposed by a set of structured and partially overlapping data-intensive web sources. Our proposal comes within a formal framework tackling two problems: the data extraction problem, to generate extraction rules based on the input websites, and the data integration problem, to integrate the extracted data in a unified schema. We introduce an original algorithm, WEIR, to solve the stated problems and formally prove its correctness. WEIR leverages the overlapping data among sources to make better decisions both in the data extraction (by pruning rules that do not lead to redundant information) and in the data integration (by reflecting local properties of a source over the mediated schema). Along the way, we characterize the amount of redundancy needed by our algorithm to produce a solution, and present experimental results to show the benefits of our approach with respect to existing solutions.

UR - http://www.scopus.com/inward/record.url?scp=84891126419&partnerID=8YFLogxK

UR - http://www.scopus.com/inward/citedby.url?scp=84891126419&partnerID=8YFLogxK

U2 - 10.14778/2536206.2536209

DO - 10.14778/2536206.2536209

M3 - Conference article

AN - SCOPUS:84891126419

SN - 2150-8097

VL - 6

SP - 805

EP - 816

JO - Proceedings of the VLDB Endowment

JF - Proceedings of the VLDB Endowment

IS - 10

T2 - 39th International Conference on Very Large Data Bases, VLDB 2012

Y2 - 26 August 2013 through 30 August 2013

ER -

Extraction and integration of partially overlapping web sources

Abstract

ASJC Scopus subject areas

Access to Document

Other files and links

Fingerprint

Cite this