This Article 
 Bibliographic References 
 Add to: 
An Insight into the Entropy and Redundancy of the English Dictionary
November 1988 (vol. 10 no. 6)
pp. 960-970

The inherent statistical characteristics, including the economy, entropy, and redundancy, of a very large set containing 93681 words from the Shorter Oxford English Dictionary is investigated. Analytical n-gram statistics are also presented for applications in natural language understanding, text processing, test compression, error detection and correction, and speech synthesis and recognition. Experimental results show how the distribution of n-grams in the dictionary varies from the ideal as n increases from 2 to 5, that is, from bigrams to pentagrams; it is shown that the corresponding redundancy increases from 0.1067 to 0.3409. The results are of interest because, (1) the dictionary provides a finite list for deterministic analyses, (2) each entry (word) appears once, compared to free-running text where words are repeated, and (3) all entries, even rarely occurring ones, have equal weight.

[1] F. M. Lynch, J. H. Petrie, and J. M. Snell, "Analysis of the micro-structure of titles in the INSPEC database,"Inform. Storage Retrieval, vol. 9, pp. 331-337, 1973.
[2] C. E. Shannon, "A mathematical theory of communication,"Bell Syst. Tech. J., vol. 27, pp. 379-423, July 1948, and pp. 623-656, Oct. 1948.
[3] C. E. Shannon, "Prediction and entropy of printed English,"Bell Syst. Tech. J., vol. 30, pp. 50-64, 1951.
[4] C. Y. Suen, "n-gram statistics for natural language understanding and text processing,"IEEE Trans. Pattern Anal. Machine Intell., vol. PAMI-1, no. 2, pp. 164-172, Apr. 1979.
[5] E. J. Yannakoudakis, "The generation and use of text fragments for data compression,"Inform. Processing Management, vol. 18, no. 1, pp. 15-21, 1982.
[6] E. J. Yannakoudakis, "Towards a universal record identification and retrieval scheme,"J. Inform., vol. 3, no. 1, pp. 7-11, 1979.
[7] E. J. Yannakoudakis, "The effectiveness of compression techniques on the English dictionary," in preparation.
[8] E. J. Yannakoudakis and A. K. P. Wu, "Quasi-equifrequent group generation and evaluation,"Comput. J., vol. 25, no. 2, pp. 183-187, 1982.
[9] G. K. Zipf,Human Behavior and the Principle of Least Effort. Reading, MA: Addison-Wesley, 1949.

Index Terms:
entropy; redundancy; inherent statistical characteristics; Shorter Oxford English Dictionary; natural language understanding; text processing; test compression; error detection; glossaries; information analysis; natural languages; redundancy; statistical analysis
E.J. Yannakoudakis, G. Angelidakis, "An Insight into the Entropy and Redundancy of the English Dictionary," IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 10, no. 6, pp. 960-970, Nov. 1988, doi:10.1109/34.9119
Usage of this product signifies your acceptance of the Terms of Use.