A Computational Approach to Plagiarism Detection in Kannada Texts Using Sentence, Bigram, and Trigram Similarity Models

U B Pavanaja * and Prajna Devadiga

Vishvakannada Foundation, Bengaluru, Karnataka, India.
* Corresponding Author
ORCID Details                                                                                                                        

 

Research Article
Open Access Research Journal of Engineering and Technology, 2026, 11(01), 077–084.
Article DOI: 10.53022/oarjet.2026.11.1.0060
Publication history: 
Received on 10 July 2026; revised on 16 August 2026; accepted on 18 August 2026
 
Abstract: 
This paper documents a Python implementation for comparing Kannada Unicode texts using sentence-level sequence similarity and character- and word-level bigram and trigram similarities. The application normalizes text with NFKC (Normalization Form Compatibility Composition), collapses whitespace, segments sentences using Kannada danda and common punctuation, computes Jaccard and cosine n-gram scores, and produces an interpretable report. In the supplied combined-mode run, Kadyanata contained 1,388 segmented source sentences and Karavaliya Saviradondu Daivagalu contained 219 target sentences. At a sentence threshold of 0.60, 138 target sentences matched their best source sentence, giving a sentence-match rate of 63.01%. Character bigram and trigram cosine similarities were 96.05% and 88.89%, while word bigram and trigram cosine similarities were 15.43% and 6.25%. The mean of the eight implemented n-gram scores was 36.08%. The code-defined combined score, 40% sentence-match rate plus 60% mean n-gram similarity, was 46.85%.
 
Keywords: 
Kannada Text Processing; Plagiarism Screening; Sentence Similarity; Bigram; Trigram; Jaccard Similarity; Cosine Similarity.
 
Full text article in PDF: