The Community for Technology Leaders
2014 IEEE 30th International Conference on Data Engineering Workshops (ICDEW) (2014)
Chicago, IL, USA
March 31, 2014 to April 4, 2014
ISBN: 978-1-4799-3481-2
pp: 68-73
Kyle Williams , Information Sciences and Technology, Pennsylvania State University, University Park, PA 16802, USA
Jian Wu , Information Sciences and Technology, Pennsylvania State University, University Park, PA 16802, USA
Sagnik Ray Choudhury , Information Sciences and Technology, Pennsylvania State University, University Park, PA 16802, USA
Madian Khabsa , Computer Science and Engineering, Pennsylvania State University, University Park, PA 16802, USA
C. Lee Giles , Computer Science and Engineering, Pennsylvania State University, University Park, PA 16802, USA
ABSTRACT
CiteSeerχ is a digital library that contains approximately 3.5 million scholarly documents and receives between 2 and 4 million requests per day. In addition to making documents available via a public Website, the data is also used to facilitate research in areas like citation analysis, co-author network analysis, scalability evaluation and information extraction. The papers in CiteSeerχ are gathered from the Web by means of continuous automatic focused crawling and go through a series of automatic processing steps as part of the ingestion process. Given the size of the collection, the fact that it is constantly expanding, and the multiple ways in which it is used both by the public to access scholarly documents and for research, there are several big data challenges. In this paper, we provide a case study description of how we address these challenges when it comes to information extraction, data integration and entity linking in CiteSeerχ. We describe how we: aggregate data from multiple sources on the Web; store and manage data; process data as part of an automatic ingestion pipeline that includes automatic metadata and information extraction; perform document and citation clustering; perform entity linking and name disambiguation; and make our data and source code available to enable research and collaboration.
INDEX TERMS
CITATION
Kyle Williams, Jian Wu, Sagnik Ray Choudhury, Madian Khabsa, C. Lee Giles, "Scholarly big data information extraction and integration in the CiteSeerχ digital library", 2014 IEEE 30th International Conference on Data Engineering Workshops (ICDEW), vol. 00, no. , pp. 68-73, 2014, doi:10.1109/ICDEW.2014.6818305
93 ms
(Ver 3.3 (11022016))