2001/11/30 by Anand Venkataraman
Computer Science · #cs.CL
published as Computational Linguistics, 27(3), pp.352--372, 2001 · Expanded version of ICML-01 paper (pp.569--576)
arxiv created 2001/11/30 · arxiv updated 2009/11/30
A statistical model for segmentation and word discovery in continuous speech is presented. An incremental unsupervised learning algorithm to infer word boundaries based on this model is described. Results of empirical tests showing that the algorithm is competitive with other models that have been used for similar tasks are also presented.