SciELO - Scientific Electronic Library Online

 
vol.23 issue3Semi-Automatic Knowledge Graph Construction by Relation Pattern ExtractionExtracting Context of Math Formulae Contained inside Scientific Documents author indexsubject indexsearch form
Home Pagealphabetic serial listing  

Services on Demand

Journal

Article

Indicators

Related links

  • Have no similar articlesSimilars in SciELO

Share


Computación y Sistemas

On-line version ISSN 2007-9737Print version ISSN 1405-5546

Abstract

WARJRI, Sunita; PAKRAY, Partha; LYNGDOH, Saralin  and  KUMAR MAJI, Arnab. Identification of POS Tag for Khasi Language based on Hidden Markov Model POS Tagger. Comp. y Sist. [online]. 2019, vol.23, n.3, pp.795-802.  Epub Aug 09, 2021. ISSN 2007-9737.  https://doi.org/10.13053/cys23-3-3248.

Computational Linguistic (CL) becomes an essential and important amenity in the present scenarios, as many different technologies are involved in making machines to understand human languages. Khasi is the language which is spoken in Meghalaya, India. Many Indian languages have been researched in different fields of Natural Language Processing (NLP), whereas Khasi lacks substantial research from the NLP perspectives. Therefore, in this paper, taking POS tagging as one of the key aspects of NLP, we present POS tagger based on Hidden Markov Model (HMM) for Khasi language. In this present preliminary stage of building NLP system for Khasi, with the analyses of the categories and structures of the words is started. Therefore, we have designed specific POS tagsets to categories Khasi words and vocabularies. Then, the POS system based on HMM is trained by using Khasi words which have been tagged manually using the designed tagsets. As ambiguity is one of the main challenges in POS tagging in Khasi, we anticipated difficulties in tagging. However, by running with the first few sets of data in the experimental data by using the HMM tagger we found out that the result yielded by this model is 76.70% of accurate.

Keywords : Natural language processing (NLP); computational linguistic; part of speech (POS); POS tagger; hidden Markov model (HMM).

        · text in English     · English ( pdf )