Issue No. 04 - July/August (2010 vol. 14)
DOI Bookmark: http://doi.ieeecomputersociety.org/10.1109/MIC.2010.58
Hanna Köpcke , University of Leipzig
Andreas Thor , University of Leipzig
Erhard Rahm , University of Leipzig
Entity matching is a key task for data integration and especially challenging for Web data. Effective entity matching typically requires combining several match techniques and finding suitable configuration parameters, such as similarity thresholds. The authors investigate to what degree machine learning helps semi-automatically determine suitable match strategies with a limited amount of manual training effort. They use a new framework, Fever, to evaluate several learning-based approaches for matching different sets of Web data entities. In particular, they study different approaches for training-data selection and how much training is needed to find effective combined match strategies and configurations.
Web data integration, entity matching, machine learning
H. Köpcke, A. Thor and E. Rahm, "Learning-Based Approaches for Matching Web Data Entities," in IEEE Internet Computing, vol. 14, no. , pp. 23-31, 2010.