An implicit segmentation approach for Telugu text recognition based on Hidden Markov models

dc.contributor.author Rao, D. Koteswara
dc.contributor.author Negi, Atul
dc.date.accessioned 2022-03-27T05:52:57Z
dc.date.available 2022-03-27T05:52:57Z
dc.date.issued 2016-01-01
dc.description.abstract Telugu text is composed of aksharas (characters). The presence of split and connected aksharas in Telugu document images causes segmentation difficulties and the performance of the Telugu OCR systems is affected. Our novel approach to solve this problem is using an implicit segmentation for recognizing words. The implicit segmentation approach does not need prior segmentation of the words into aksharas before they are recognized. Since the Hidden Markov models (HMM) are successfully applied for phoneme recognition with no prior segmentation of the speech into phonemes in the automatic speech recognition applications. In this paper, we report on the use of continuous density Hidden Markov Models for representing the shape of aksharas to build Telugu text recognition system. The sliding window method is used for computing simple statistical features and 450 akshara HMMs are trained. We use word bigram language model as contextual information. The word recognition relies on akshara models and contextual information of words. The word recognition involves finding the maximum likelihood sequence of akshara models that matches against the feature vector sequence. Our system recognizes words with split and connected aksharas. The performance of the system is encouraging.
dc.identifier.citation Advances in Intelligent Systems and Computing. v.425
dc.identifier.issn 21945357
dc.identifier.uri 10.1007/978-3-319-28658-7_54
dc.identifier.uri http://link.springer.com/10.1007/978-3-319-28658-7_54
dc.identifier.uri https://dspace.uohyd.ac.in/handle/1/8581
dc.title An implicit segmentation approach for Telugu text recognition based on Hidden Markov models
dc.type Book Series. Conference Paper
dspace.entity.type
Files
License bundle
Now showing 1 - 1 of 1
No Thumbnail Available
Name:
license.txt
Size:
1.71 KB
Format:
Plain Text
Description: