Issue No.09 - September (2005 vol.17)
James Caverlee , IEEE
Ling Liu , IEEE Computer Society
DOI Bookmark: http://doi.ieeecomputersociety.org/10.1109/TKDE.2005.151
This paper presents the QA-Pagelet as a fundamental data preparation technique for large-scale data analysis of the Deep Web. To support QA-Pagelet extraction, we present the Thor framework for sampling, locating, and partioning the QA-Pagelets from the Deep Web. Two unique features of the Thor framework are 1) the novel page clustering for grouping pages from a Deep Web source into distinct clusters of control-flow dependent pages and 2) the novel subtree filtering algorithm that exploits the structural and content similarity at subtree level to identify the QA-Pagelets within highly ranked page clusters. We evaluate the effectiveness of the Thor framework through experiments using both simulation and real data sets. We show that Thor performs well over millions of Deep Web pages and over a wide range of sources, including e-Commerce sites, general and specialized search engines, corporate Web sites, medical and legal resources, and several others. Our experiments also show that the proposed page clustering algorithm achieves low-entropy clusters, and the subtree filtering algorithm identifies QA-Pagelets with excellent precision and recall.
Index Terms- Deep Web, data preparation, data extraction, pagelets, clustering.
James Caverlee, Ling Liu, "QA-Pagelet: Data Preparation Techniques for Large-Scale Data Analysis of the Deep Web", IEEE Transactions on Knowledge & Data Engineering, vol.17, no. 9, pp. 1247-1262, September 2005, doi:10.1109/TKDE.2005.151