An implicit segmentation approach for Telugu text recognition based on Hidden Markov models

Rao, D. Koteswara; Negi, Atul

An implicit segmentation approach for Telugu text recognition based on Hidden Markov models

Date

2016-01-01

Authors

Rao, D. Koteswara

Negi, Atul

Abstract

Telugu text is composed of aksharas (characters). The presence of split and connected aksharas in Telugu document images causes segmentation difficulties and the performance of the Telugu OCR systems is affected. Our novel approach to solve this problem is using an implicit segmentation for recognizing words. The implicit segmentation approach does not need prior segmentation of the words into aksharas before they are recognized. Since the Hidden Markov models (HMM) are successfully applied for phoneme recognition with no prior segmentation of the speech into phonemes in the automatic speech recognition applications. In this paper, we report on the use of continuous density Hidden Markov Models for representing the shape of aksharas to build Telugu text recognition system. The sliding window method is used for computing simple statistical features and 450 akshara HMMs are trained. We use word bigram language model as contextual information. The word recognition relies on akshara models and contextual information of words. The word recognition involves finding the maximum likelihood sequence of akshara models that matches against the feature vector sequence. Our system recognizes words with split and connected aksharas. The performance of the system is encouraging.

Citation

Advances in Intelligent Systems and Computing. v.425

URI

10.1007/978-3-319-28658-7_54
http://link.springer.com/10.1007/978-3-319-28658-7_54
https://dspace.uohyd.ac.in/handle/1/8581

Collections

Computer and Information Sciences - Publications

Full item page