Why random forest fails in small-sample academic grade prediction
Department of Artificial Intelligence, Shri Vishnu Engineering College for Women, Bhimavaram, Andhra Pradesh, India 534202.
Research Article
Open Access Research Journal of Multidisciplinary Studies, 2025, 09(02), 061-069.
Article DOI: 10.53022/oarjms.2025.9.2.0033
Publication history:
Received on 29 May 2025; revised on 14 June 2025; accepted on 16 June 2025
Abstract:
Academic performance prediction plays a crucial role in education planning, student intervention strategies, and institutional decision-making. As universities and schools seek data-driven methods to assess student progress, understanding grade trends and performance patterns has become essential. Traditional evaluation methods rely on historical records, but they often fail to account for complex relationships between subjects, student learning behaviors, and external influences. This creates a need for predictive analysis, which can provide early insights into student outcomes, enabling institutions to offer targeted academic support. Machine learning techniques have increasingly been explored for academic forecasting, allowing educators to identify high-risk students, optimize curriculum design, and improve assessment models. However, the challenge lies in selecting an appropriate model that effectively handles structured academic datasets with limited records. This study evaluates predictive modeling approaches in the context of student grade forecasting, examining how variations in test size, dataset structure, and subject dependencies influence prediction accuracy. By analyzing scatter plots, R² values, and feature correlations, this research highlights the strengths and weaknesses of machine learning-driven grade prediction. Through this evaluation, we aim to refine methodologies that enhance academic forecasting accuracy, ensuring that predictive models align with real-world institutional needs. Our findings underscore the importance of selecting robust approaches that account for dataset constraints and generalization challenges rather than relying solely on automated algorithms.
Keywords:
Random forest; Academic performance; Engineering education; Machine learning
Full text article in PDF:
Copyright information:
Copyright © 2025 Author(s) retain the copyright of this article. This article is published under the terms of the Creative Commons Attribution Liscense 4.0
