We apologize for a recent technical issue with our email system, which temporarily affected account activations. Accounts have now been activated. Authors may proceed with paper submissions. PhDFocusTM
CFP last date
20 November 2024
Reseach Article

Part of Speech Tagging in Manipuri: A Rule based Approach

by Kh Raju Singha, Bipul Syam Purkayastha, Kh Dhiren Singha
International Journal of Computer Applications
Foundation of Computer Science (FCS), NY, USA
Volume 51 - Number 14
Year of Publication: 2012
Authors: Kh Raju Singha, Bipul Syam Purkayastha, Kh Dhiren Singha
10.5120/8111-1727

Kh Raju Singha, Bipul Syam Purkayastha, Kh Dhiren Singha . Part of Speech Tagging in Manipuri: A Rule based Approach. International Journal of Computer Applications. 51, 14 ( August 2012), 31-36. DOI=10.5120/8111-1727

@article{ 10.5120/8111-1727,
author = { Kh Raju Singha, Bipul Syam Purkayastha, Kh Dhiren Singha },
title = { Part of Speech Tagging in Manipuri: A Rule based Approach },
journal = { International Journal of Computer Applications },
issue_date = { August 2012 },
volume = { 51 },
number = { 14 },
month = { August },
year = { 2012 },
issn = { 0975-8887 },
pages = { 31-36 },
numpages = {9},
url = { https://ijcaonline.org/archives/volume51/number14/8111-1727/ },
doi = { 10.5120/8111-1727 },
publisher = {Foundation of Computer Science (FCS), NY, USA},
address = {New York, USA}
}
%0 Journal Article
%1 2024-02-06T20:50:24.151432+05:30
%A Kh Raju Singha
%A Bipul Syam Purkayastha
%A Kh Dhiren Singha
%T Part of Speech Tagging in Manipuri: A Rule based Approach
%J International Journal of Computer Applications
%@ 0975-8887
%V 51
%N 14
%P 31-36
%D 2012
%I Foundation of Computer Science (FCS), NY, USA
Abstract

The process of assigning morpho-syntactic categories of each morpheme including punctuation marks in a given text document according to the context is called Part of Speech (POS) tagging. In this paper we represent the rule-based Part of Speech Tagger of Manipuri by applying a set of hand written linguistic rules of Manipuri language. Nevertheless, it is very difficult to classify the lexical categories of Manipuri, an agglutinating Tibeto-Burman language of Northeast India. So, in this tagger we are using the affix stripping technique to segment the affixes from the root. As Manipuri has limited POS tagged corpus, the tagged output of this tagger will be very helpful to analyze Manipuri Part of speech by using many statistical models.

References
  1. G. A. Grierson's Linguistic Survey of India. Vol. III, Pt. III, 1976.
  2. Eric Brill. A simple rule-based part of speech tagger. In Proceedings Third Conference on Applied Natural Language Processing, ACL, Trento, Italy, 1992.
  3. K. Oflazer, I Kuruoz, "Tagging and morphological disambiguation of Turkish text". In Proceedings of 4th ACL conference on Applied Natural Language Processing Conference, 1994.
  4. Ch. Yashawanta Singh "Manipuri Grammar. " Rajesh Publications, New Delhi, 2000.
  5. J. Hajic, P. Krbec, P. Kveton, K. Oliva, V. Petkevic, "A Case Study in Czech Tagging". In proceedings of the 39th Annual Meeting of the ACL, 2001.
  6. Sandipan Dandapat, Sudeshna Sarkar and Anupam Basu "A Hybrid Model for Part-of-Speech Tagging and its Application to Bengali", Transactions on Engineering, Computing and Technology V1 December, 2004.
  7. Sirajul Islam Choudhury, Leihaorambam Sarbajit Singh, Samir Borgohain, P. K. Das, "Morphological Analyzer for Manipuri: Design and Implementation". In Proceedings of AACC, Kathmandu, Nepal, pp 123-129, 2004.
  8. S. Imoba. "Manipuri to English Dictionary". S. Ibetombi Devi, Imphal, 2004.
  9. P. C. Thoudam. "Problems in the Analysis of Manipuri Language. "www. ciil-ebooks. net, CIIL, Mysore, 2006.
  10. Sachin Burange, Sushant Devlakar, Pushpak Bhattacharyya, "Rule Governed Marathi POS Tagging". In Proceeding of MSPIL, IIT Bombay, pp 69- 78, 2006.
  11. Kh. Dhiren Singha, "Loan Words in Manipuri", Bilingualism and North-East India, an Assam University Publication, 2008.
  12. Manish Shrivastava and Pushpak Bhattacharyya, Hindi POS Tagger Using Naïve Stemming: Harnessing Morphological Information without Extensive Linguistic Knowledge, International Conference on NLP (ICON08), Pune, India, 2008.
  13. S. Baskaran et al. " Designing a Common POS-Tagset Framework for Indian Languages" The 6th Workshop on Asian Languae Resources, 2008.
  14. Thoudam Doren Singh & Sivaji Bandyopadhyay "Morphology Driven Manipuri POS Tagger", Proceedings of the IJCNLP-08 Workshop on NLP for Less Privileged Languages, pages 91–98, Hyderabad, India, January 2008.
  15. D. Jurafsky, and J. H. Martin, "Speech and Language Processing", Second edition, Published by Pearson Education, 2009.
  16. Dinesh Kumar and Gurpreet Singh Josan, "Part of Speech Taggers for Morphologically Rich Indian Languages: A Survey", International Journal of Computer Applications (0975 – 8887) Volume6–No. 5, September, 2010.
  17. Kh Raju Singha, Bipul Syam Purkayastha, Kh Dhiren Singha, Arindam Roy "Developing a Tagset for Manipuri Part of Speech Tagging" Journal of Computer Science and Engineering, Volume-5-issue-1-january 2011. http://sites. google. com/site. /jcseuk/volume-5-issue 1-january-2011.
  18. http://language. worldofcomputing. net/pos-tagging/rule based-pos-tagging. html. 2012.
Index Terms

Computer Science
Information Sciences

Keywords

Tagset tokenizer lexicon corpus affix stemmer information retrieval