International Journal of Computer Applications |
Foundation of Computer Science (FCS), NY, USA |
Volume 120 - Number 9 |
Year of Publication: 2015 |
Authors: Nitesh Pradhan, Manasi Gyanchandani, Rajesh Wadhvani |
10.5120/21257-4109 |
Nitesh Pradhan, Manasi Gyanchandani, Rajesh Wadhvani . A Review on Text Similarity Technique used in IR and its Application. International Journal of Computer Applications. 120, 9 ( June 2015), 29-34. DOI=10.5120/21257-4109
With large number of documents on the web, there is a increasing need to be able to retrieve the best relevant document. There are different techniques through which we can retrieve most relevant document from the large corpus. Similarity between words, sentences, paragraphs and documents is an important component in various tasks such as information retrieval, document clustering, word-sense disambiguation, automatic essay scoring, short answer grading, machine translation and text summarization. Text similarity means user's query text is matched with the document text and on the basis on this matching user retrieves the most relevant documents. Text similarity also plays an important role in the categorization of text as well as document. We can measure the similarity between sentences, words, paragraphs and documents to categorize them in an efficient way. On the basis of this categorization, we can retrieve the best relevant document corresponding to user's query. This paper describes different types of similarity like lexical similarity, semantic similarity etc.