Query processing over incomplete autonomous databases: Query rewriting using learned data dependencies

Garrett Wolf; Aravind Kalavagattu; Hemal Khatri; Raju Balakrishnan; Bhaumik Chokshi; Jianchun Fan; Yi Chen; Subbarao Kambhampati

doi:10.1007/s00778-009-0155-0

Query processing over incomplete autonomous databases: Query rewriting using learned data dependencies

Garrett Wolf, Aravind Kalavagattu, Hemal Khatri, Raju Balakrishnan, Bhaumik Chokshi, Jianchun Fan, Yi Chen, Subbarao Kambhampati

Research output: Contribution to journal › Article › peer-review

25 Scopus citations

Abstract

Incompleteness due to missing attribute values (aka "null values") is very common in autonomous web databases, on which user accesses are usually supported through mediators. Traditional query processing techniques that focus on the strict soundness of answer tuples often ignore tuples with critical missing attributes, even if they wind up being relevant to a user query. Ideally we would like the mediator to retrieve such possibleanswers and gauge their relevance by accessing their likelihood of being pertinent answers to the query. The autonomous nature of web databases poses several challenges in realizing this objective. Such challenges include the restricted access privileges imposed on the data, the limited support for query patterns, and the bounded pool of database and network resources in the web environment. We introduce a novel query rewriting and optimization framework QPIAD that tackles these challenges. Our technique involves reformulating the user query based on mined correlations among the database attributes. The reformulated queries are aimed at retrieving the relevant possibleanswers in addition to the certain answers. QPIAD is able to gauge the relevance of such queries allowing tradeoffs in reducing the costs of database query processing and answer transmission. To support this framework, we develop methods for mining attribute correlations (in terms of Approximate Functional Dependencies), value distributions (in the form of Naïve Bayes Classifiers), and selectivity estimates. We present empirical studies to demonstrate that our approach is able to effectively retrieve relevant possibleanswers with high precision, high recall, and manageable cost.

Original language	English (US)
Pages (from-to)	1167-1190
Number of pages	24
Journal	VLDB Journal
Volume	18
Issue number	5
DOIs	https://doi.org/10.1007/s00778-009-0155-0
State	Published - Oct 2009

Keywords

Incomplete databases
Query rewriting
Uncertainty

ASJC Scopus subject areas

Information Systems
Hardware and Architecture

Access to Document

10.1007/s00778-009-0155-0

Cite this

@article{37c6a8439e1b43f08fd27ee86e41c869,

title = "Query processing over incomplete autonomous databases: Query rewriting using learned data dependencies",

abstract = "Incompleteness due to missing attribute values (aka {"}null values{"}) is very common in autonomous web databases, on which user accesses are usually supported through mediators. Traditional query processing techniques that focus on the strict soundness of answer tuples often ignore tuples with critical missing attributes, even if they wind up being relevant to a user query. Ideally we would like the mediator to retrieve such possibleanswers and gauge their relevance by accessing their likelihood of being pertinent answers to the query. The autonomous nature of web databases poses several challenges in realizing this objective. Such challenges include the restricted access privileges imposed on the data, the limited support for query patterns, and the bounded pool of database and network resources in the web environment. We introduce a novel query rewriting and optimization framework QPIAD that tackles these challenges. Our technique involves reformulating the user query based on mined correlations among the database attributes. The reformulated queries are aimed at retrieving the relevant possibleanswers in addition to the certain answers. QPIAD is able to gauge the relevance of such queries allowing tradeoffs in reducing the costs of database query processing and answer transmission. To support this framework, we develop methods for mining attribute correlations (in terms of Approximate Functional Dependencies), value distributions (in the form of Na{\"i}ve Bayes Classifiers), and selectivity estimates. We present empirical studies to demonstrate that our approach is able to effectively retrieve relevant possibleanswers with high precision, high recall, and manageable cost.",

keywords = "Incomplete databases, Query rewriting, Uncertainty",

author = "Garrett Wolf and Aravind Kalavagattu and Hemal Khatri and Raju Balakrishnan and Bhaumik Chokshi and Jianchun Fan and Yi Chen and Subbarao Kambhampati",

note = "Funding Information: This research was supported in part by the NSF grants IIS 308139, IIS 0624341, IIS 0740129 and IIS 0845647 (CAREER); the ONR grants N000140610058 and N000140910032, a Google research award, as well as support from ASU (via ECR A601), the ASU Prop 301 grant to ET-I3 initiative.",

year = "2009",

month = oct,

doi = "10.1007/s00778-009-0155-0",

language = "English (US)",

volume = "18",

pages = "1167--1190",

journal = "VLDB Journal",

issn = "1066-8888",

publisher = "Springer New York",

number = "5",

}

TY - JOUR

T1 - Query processing over incomplete autonomous databases

T2 - Query rewriting using learned data dependencies

AU - Wolf, Garrett

AU - Kalavagattu, Aravind

AU - Khatri, Hemal

AU - Balakrishnan, Raju

AU - Chokshi, Bhaumik

AU - Fan, Jianchun

AU - Chen, Yi

AU - Kambhampati, Subbarao

N1 - Funding Information: This research was supported in part by the NSF grants IIS 308139, IIS 0624341, IIS 0740129 and IIS 0845647 (CAREER); the ONR grants N000140610058 and N000140910032, a Google research award, as well as support from ASU (via ECR A601), the ASU Prop 301 grant to ET-I3 initiative.

PY - 2009/10

Y1 - 2009/10

N2 - Incompleteness due to missing attribute values (aka "null values") is very common in autonomous web databases, on which user accesses are usually supported through mediators. Traditional query processing techniques that focus on the strict soundness of answer tuples often ignore tuples with critical missing attributes, even if they wind up being relevant to a user query. Ideally we would like the mediator to retrieve such possibleanswers and gauge their relevance by accessing their likelihood of being pertinent answers to the query. The autonomous nature of web databases poses several challenges in realizing this objective. Such challenges include the restricted access privileges imposed on the data, the limited support for query patterns, and the bounded pool of database and network resources in the web environment. We introduce a novel query rewriting and optimization framework QPIAD that tackles these challenges. Our technique involves reformulating the user query based on mined correlations among the database attributes. The reformulated queries are aimed at retrieving the relevant possibleanswers in addition to the certain answers. QPIAD is able to gauge the relevance of such queries allowing tradeoffs in reducing the costs of database query processing and answer transmission. To support this framework, we develop methods for mining attribute correlations (in terms of Approximate Functional Dependencies), value distributions (in the form of Naïve Bayes Classifiers), and selectivity estimates. We present empirical studies to demonstrate that our approach is able to effectively retrieve relevant possibleanswers with high precision, high recall, and manageable cost.

AB - Incompleteness due to missing attribute values (aka "null values") is very common in autonomous web databases, on which user accesses are usually supported through mediators. Traditional query processing techniques that focus on the strict soundness of answer tuples often ignore tuples with critical missing attributes, even if they wind up being relevant to a user query. Ideally we would like the mediator to retrieve such possibleanswers and gauge their relevance by accessing their likelihood of being pertinent answers to the query. The autonomous nature of web databases poses several challenges in realizing this objective. Such challenges include the restricted access privileges imposed on the data, the limited support for query patterns, and the bounded pool of database and network resources in the web environment. We introduce a novel query rewriting and optimization framework QPIAD that tackles these challenges. Our technique involves reformulating the user query based on mined correlations among the database attributes. The reformulated queries are aimed at retrieving the relevant possibleanswers in addition to the certain answers. QPIAD is able to gauge the relevance of such queries allowing tradeoffs in reducing the costs of database query processing and answer transmission. To support this framework, we develop methods for mining attribute correlations (in terms of Approximate Functional Dependencies), value distributions (in the form of Naïve Bayes Classifiers), and selectivity estimates. We present empirical studies to demonstrate that our approach is able to effectively retrieve relevant possibleanswers with high precision, high recall, and manageable cost.

KW - Incomplete databases

KW - Query rewriting

KW - Uncertainty

UR - http://www.scopus.com/inward/record.url?scp=70349835611&partnerID=8YFLogxK

UR - http://www.scopus.com/inward/citedby.url?scp=70349835611&partnerID=8YFLogxK

U2 - 10.1007/s00778-009-0155-0

DO - 10.1007/s00778-009-0155-0

M3 - Article

AN - SCOPUS:70349835611

SN - 1066-8888

VL - 18

SP - 1167

EP - 1190

JO - VLDB Journal

JF - VLDB Journal

IS - 5

ER -

Query processing over incomplete autonomous databases: Query rewriting using learned data dependencies

Abstract

Keywords

ASJC Scopus subject areas

Access to Document

Other files and links

Fingerprint

Cite this