Eighth International Symposium on Symbolic and Numeric Algorithms for Scientific Computing (SYNASC'06)
HTML Pattern Generator--Automatic Data Extraction from Web Pages
Timisoara, Romania
September 26-September 29
ISBN: 0-7695-2740-X
Existing methods of information extraction from HTML documents include manual approach, supervised learning and automatic techniques. The manual method has high precision and recall values but it is difficult to apply it for large number of pages. Supervised learning involves human interaction to create positive and negative samples. Automatic techniques benefit from less human effort but they are not highly reliable regarding the information retrieved.
Citation:
Mirel Cosulschi, Adrian Giurca, Bogdan Udrescu, Nicolae Constantinescu, Mihai Gabroveanu, "HTML Pattern Generator--Automatic Data Extraction from Web Pages," synasc, pp.75-78, Eighth International Symposium on Symbolic and Numeric Algorithms for Scientific Computing (SYNASC'06), 2006