The Community for Technology Leaders
Green Image
Issue No. 03 - March (2014 vol. 26)
ISSN: 1041-4347
pp: 682-697
Ma'ayan Dror , Dept. of Inf. Syst. Eng., Ben-Gurion Univ. of the Negev, Beer-Sheva, Israel
Asaf Shabtai , Dept. of Inf. Syst. Eng., Ben-Gurion Univ. of the Negev, Beer-Sheva, Israel
Lior Rokach , Dept. of Inf. Syst. Eng., Ben-Gurion Univ. of the Negev, Beer-Sheva, Israel
Yuval Elovici , Dept. of Inf. Syst. Eng., Ben-Gurion Univ. of the Negev, Beer-Sheva, Israel
ABSTRACT
One-to-many data linkage is an essential task in many domains, yet only a handful of prior publications have addressed this issue. Furthermore, while traditionally data linkage is performed among entities of the same type, it is extremely necessary to develop linkage techniques that link between matching entities of different types as well. In this paper, we propose a new one-to-many data linkage method that links between entities of different natures. The proposed method is based on a one-class clustering tree (OCCT) that characterizes the entities that should be linked together. The tree is built such that it is easy to understand and transform into association rules, i.e., the inner nodes consist only of features describing the first set of entities, while the leaves of the tree represent features of their matching entities from the second data set. We propose four splitting criteria and two different pruning methods which can be used for inducing the OCCT. The method was evaluated using data sets from three different domains. The results affirm the effectiveness of the proposed method and show that the OCCT yields better performance in terms of precision and recall (in most cases it is statistically significant) when compared to a C4.5 decision tree-based linkage method.
INDEX TERMS
Couplings, Decision trees, Vegetation, Training, Classification algorithms, Numerical models, Buildings,decision tree induction, Clustering, classification, data matching
CITATION
Ma'ayan Dror, Asaf Shabtai, Lior Rokach, Yuval Elovici, "OCCT: A One-Class Clustering Tree for Implementing One-to-Many Data Linkage", IEEE Transactions on Knowledge & Data Engineering, vol. 26, no. , pp. 682-697, March 2014, doi:10.1109/TKDE.2013.23
86 ms
(Ver )