The Community for Technology Leaders
RSS Icon
Subscribe
Issue No.11 - November (2005 vol.17)
pp: 1505-1517
ABSTRACT
This paper establishes a formal connection between two common, but previously unconnected methods for analyzing data streams: discovering frequent episodes in a computer science framework and learning generative models in a statistics framework. We introduce a special class of discrete Hidden Markov Models (HMMs), called Episode Generating HMMs (EGHs), and associate each episode with a unique EGH. We prove that, given any two episodes, the EGH that is more likely to generate a given data sequence is the one associated with the more frequent episode. To be able to establish such a relationship, we define a new measure of frequency of an episode, based on what we call nonoverlapping occurrences of the episode in the data. An efficient algorithm is proposed for counting the frequencies for a set of episodes. Through extensive simulations, we show that our algorithm is both effective and more efficient than current methods for frequent episode discovery. We also show how the association between frequent episodes and EGHs can be exploited to assess the significance of frequent episodes discovered and illustrate empirically how this idea may be used to improve the efficiency of the frequent episode discovery.
INDEX TERMS
Index Terms- Temporal data mining, sequential data, frequent episodes, Hidden Markov Models, statistical significance.
CITATION
Srivatsan Laxman, P.S. Sastry, K.P. Unnikrishnan, "Discovering Frequent Episodes and Learning Hidden Markov Models: A Formal Connection", IEEE Transactions on Knowledge & Data Engineering, vol.17, no. 11, pp. 1505-1517, November 2005, doi:10.1109/TKDE.2005.181
680 ms
(Ver 2.0)

Marketing Automation Platform Marketing Automation Tool