Issue No. 02 - February (2012 vol. 24)
DOI Bookmark: http://doi.ieeecomputersociety.org/10.1109/TKDE.2010.231
Chien-Chih Chen , National Taiwan University, Taipei
Kai-Hsiang Yang , National Taipei University of Education, Taipei
Chuen-Liang Chen , National Taiwan University, Taipei
Jan-Ming Ho , Academia Sinica, Taipei
Dramatic increase in the number of academic publications has led to growing demand for efficient organization of the resources to meet researchers' needs. As a result, a number of network services have compiled databases from the public resources scattered over the Internet. However, publications by different conferences and journals adopt different citation styles. It is an interesting problem to accurately extract metadata from a citation string which is formatted in one of thousands of different styles. It has attracted a great deal of attention in research in recent years. In this paper, based on the notion of sequence alignment, we present a citation parser called BibPro that extracts components of a citation string. To demonstrate the efficacy of BibPro, we conducted experiments on three benchmark data sets. The results show that BibPro achieved over 90 percent accuracy on each benchmark. Even with citations and associated metadata retrieved from the web as training data, our experiments show that BibPro still achieves a reasonable performance.
Data integration, digital libraries, information extraction, sequence alignment.
C. Chen, K. Yang, C. Chen and J. Ho, "BibPro: A Citation Parser Based on Sequence Alignment," in IEEE Transactions on Knowledge & Data Engineering, vol. 24, no. , pp. 236-250, 2010.