A Computational Approach to Plagiarism Detection in Kannada Texts Using Sentence, Bigram, and Trigram Similarity Models
Vishvakannada Foundation, Bengaluru, Karnataka, India.
* Corresponding Author
ORCID Details
U B Pavanaja: https://orcid.org/my-orcid?orcid=0009-0000-2148-9536
Prajna Devadiga: https://orcid.org/0009-0005-7627-2234
Research Article
Open Access Research Journal of Engineering and Technology, 2026, 11(01), 077–084.
Article DOI: 10.53022/oarjet.2026.11.1.0060
Publication history:
Received on 10 July 2026; revised on 16 August 2026; accepted on 18 August 2026
Abstract:
This paper documents a Python implementation for comparing Kannada Unicode texts using sentence-level sequence similarity and character- and word-level bigram and trigram similarities. The application normalizes text with NFKC (Normalization Form Compatibility Composition), collapses whitespace, segments sentences using Kannada danda and common punctuation, computes Jaccard and cosine n-gram scores, and produces an interpretable report. In the supplied combined-mode run, Kadyanata contained 1,388 segmented source sentences and Karavaliya Saviradondu Daivagalu contained 219 target sentences. At a sentence threshold of 0.60, 138 target sentences matched their best source sentence, giving a sentence-match rate of 63.01%. Character bigram and trigram cosine similarities were 96.05% and 88.89%, while word bigram and trigram cosine similarities were 15.43% and 6.25%. The mean of the eight implemented n-gram scores was 36.08%. The code-defined combined score, 40% sentence-match rate plus 60% mean n-gram similarity, was 46.85%.
Keywords:
Kannada Text Processing; Plagiarism Screening; Sentence Similarity; Bigram; Trigram; Jaccard Similarity; Cosine Similarity.
Full text article in PDF:
Copyright information:
Copyright © 2026 Author(s) retain the copyright of this article. This article is published under the terms of the Creative Commons Attribution Liscense 4.0
