Seventh IEEE International Conference on Data Mining Workshops (ICDMW 2007)
High-Speed Identification of Language and Script
Omaha, Nebraska, USA
October 28-October 31
ISBN: 0-7695-3033-8
Humans communicate with text in thousands of languages, in dozens of scripts, and a wide variety of binary codes. There is a need to identify the language, script and code of this text to enable follow-on processing such as transcoding, translation, transliteration, routing and prioritization. This paper deals with the implementation of real-time language and script identification on high-speed hardware (principally a ternary content addressable memory) capable of processing network data streams at several gigabits per second.