The Community for Technology Leaders
RSS Icon
Subscribe
Issue No.07 - July (1993 vol.15)
pp: 737-747
ABSTRACT
<p>A method for extracting alternating horizontal and vertical projection profiles are from nested sub-blocks of scanned page images of technical documents is discussed. The thresholded profile strings are parsed using the compiler utilities Lex and Yacc. The significant document components are demarcated and identified by the recursive application of block grammars. Backtracking for error recovery and branch and bound for maximum-area labeling are implemented with Unix Shell programs. Results of the segmentation and labeling process are stored in a labeled x-y tree. It is shown that families of technical documents that share the same layout conventions can be readily analyzed. Results from experiments in which more than 20 types of document entities were identified in sample pages from two journals are presented.</p>
INDEX TERMS
syntactic segmentation; image recognition; document image processing; horizontal projection profiles; digitized pages; vertical projection profiles; scanned page images; technical documents; thresholded profile strings; compiler utilities; Lex; Yacc; block grammars; error recovery; branch and bound; Unix Shell; labeling; labeled x-y tree; document image processing; feature extraction; grammars; image recognition; image segmentation
CITATION
M. Krishnamoorthy, G. Nagy, S. Seth, M. Viswanathan, "Syntactic Segmentation and Labeling of Digitized Pages from Technical Journals", IEEE Transactions on Pattern Analysis & Machine Intelligence, vol.15, no. 7, pp. 737-747, July 1993, doi:10.1109/34.221173
19 ms
(Ver 2.0)

Marketing Automation Platform Marketing Automation Tool