2006 Eighth International Symposium on Symbolic and Numeric Algorithms for Scientific Computing (2006)
Sept. 26, 2006 to Sept. 29, 2006
Mirel Cosulschi , University of Craiova, Romania
Adrian Giurca , Brandenburg Technical University Cottbus, Germany
Bogdan Udrescu , University of Craiova, Romania
Nicolae Constantinescu , University of Craiova, Romania
Mihai Gabroveanu , University of Craiova, Romania
Existing methods of information extraction from HTML documents include manual approach, supervised learning and automatic techniques. The manual method has high precision and recall values but it is difficult to apply it for large number of pages. Supervised learning involves human interaction to create positive and negative samples. Automatic techniques benefit from less human effort but they are not highly reliable regarding the information retrieved.
M. Gabroveanu, M. Cosulschi, N. Constantinescu, B. Udrescu and A. Giurca, "HTML Pattern Generator--Automatic Data Extraction from Web Pages," 2006 Eighth International Symposium on Symbolic and Numeric Algorithms for Scientific Computing(SYNASC), Timisoara, Romania, 2006, pp. 75-78.