loading...
 This Article 
   
 Share 
   
 Bibliographic References 
   
 Add to: 
 
Digg
Furl
Spurl
Blink
Simpy
Google
Del.icio.us
Y!MyWeb
 
 Search 
   
Eighth International Symposium on Symbolic and Numeric Algorithms for Scientific Computing (SYNASC'06)
HTML Pattern Generator--Automatic Data Extraction from Web Pages
Timisoara, Romania
September 26-September 29
ISBN: 0-7695-2740-X
Mirel Cosulschi, University of Craiova, Romania
Adrian Giurca, Brandenburg Technical University Cottbus, Germany
Bogdan Udrescu, University of Craiova, Romania
Nicolae Constantinescu, University of Craiova, Romania
Mihai Gabroveanu, University of Craiova, Romania
Existing methods of information extraction from HTML documents include manual approach, supervised learning and automatic techniques. The manual method has high precision and recall values but it is difficult to apply it for large number of pages. Supervised learning involves human interaction to create positive and negative samples. Automatic techniques benefit from less human effort but they are not highly reliable regarding the information retrieved.
Citation:
Mirel Cosulschi, Adrian Giurca, Bogdan Udrescu, Nicolae Constantinescu, Mihai Gabroveanu, "HTML Pattern Generator--Automatic Data Extraction from Web Pages," synasc, pp.75-78, Eighth International Symposium on Symbolic and Numeric Algorithms for Scientific Computing (SYNASC'06), 2006
Usage of this product signifies your acceptance of the Terms of Use.