Abstract
RDF/XML has been widely recognized as the standard for annotating online Web documents and for transforming the HTML Web to the so called Semantic Web. In order to enable widespread usability for the Semantic Web there is a need to bootstrap large, rich and up-to-date domain ontologies that organize most relevant concepts, their relationships and instances. In this paper, we present automated tech-niques for bootstrapping and populating specialized domain ontologies by organizing and mining a set of relevant Web sites provided by the user. We develop algorithms that detect and utilize HTML regularities in the Web documents to turn them into hierarchical semantic structures encoded as XML. Next, we present tree-mining algorithms that identify key domain concepts and their taxonomical relationships.We also extract semi-structured concept instances annotated with their labels whenever they are available. Experimental evaluation for the News and Hotels do-main indicates that our algorithms can bootstrap and populate domain specific ontologies with high precision and recall.
Original language | English (US) |
---|---|
Title of host publication | Proceedings of the 1st International Conference on Semantic Web and Databases, SWDB 2003 |
Publisher | Association for Computing Machinery, Inc |
Pages | 245-262 |
Number of pages | 18 |
State | Published - 2003 |
Event | 1st International Conference on Semantic Web and Databases, SWDB 2003 - Berlin, Germany Duration: Sep 7 2003 → Sep 8 2003 |
Other
Other | 1st International Conference on Semantic Web and Databases, SWDB 2003 |
---|---|
Country/Territory | Germany |
City | Berlin |
Period | 9/7/03 → 9/8/03 |
ASJC Scopus subject areas
- Information Systems
- Computer Networks and Communications