Predictive modeling of solar energy generation using supervised machine learning algorithms
1School of Technology Management and Engineering, SVKM NMIMS Global University, Dhule, Maharashtra 424001, India
2School of Technology Management and Engineering, SVKM NMIMS Global University, Dhule, Maharashtra 424001, India
3Dr. Vishwanath Karad MIT World Peace University, Pune, 411038, Maharashtra, India
4Government College of Engineering, Nagpur, Maharashtra 441108, India
5School of Technology Management and Engineering, SVKM NMIMS Global University, Dhule, Maharashtra 424001, India
6School of Technology Management and Engineering, SVKM NMIMS Global University, Dhule, Maharashtra 424001, India
7Nutan Maharashtra Institute of Engineering and Technology, Talegaon, Maharashtra 410507, India
J Ther Eng 2026; 12(5): 1748-1759 DOI: 10.47481/jten.0063
Full Text PDF

Abstract

The nonlinear and dynamic nature of solar energy generation makes effective energy management, operational planning, and grid stability challenging. Purely numerical and statistical approaches are often unable to pick up the nonlinearity in the photovoltaic power generation. The present work tests the predictive capability of supervised regression algorithms for photovoltaic energy forecasting. This research is a comparative analysis of five algorithms, which are Linear Regression, k-Nearest Neighbor, Support Vector Regression, Decision Tree, and Random Forest. Unlike earlier studies that just targeted irradiation parameters and focused on short-term datasets, this study incorporates a long-term dataset, consisting of 95,950 hourly observations collected from a rooftop photovoltaic plant found in the Western Australia region from 1990 to 2014. The dataset includes multiple variables that are categorized into solar-irradiance and meteorological variables. Solar-irradiance variables are Direct-normal, Global-horizontal, and Diffused-horizontal. Meteorological variables are Wet-bulb temperature and Dew-point temperature. The photovoltaic energy produced is the output variable. The model is trained strategically with hyperparameter tuning performed using a hold-out k-fold cross-validation. From the first correlation heatmap, it is found that Direct-normal irradiance affects the photovoltaic energy the most, followed by Global-horizontal, and Diffused-horizontal irradiance. Wet-bulb temperature and the Dew-point temperature are the least influential parameters. The performance metrics evaluated to assess the developed models are Mean-Squared Error, Root Mean-Squared Error, Mean Absolute Error, and the Coefficient of Determination. Overall, the ensemble models showed superior performance to the regression-based models. Ensemble models pick up the non-linearity among the variables well. Compared to other models, Random Forest achieves the highest value of the Coefficient of Determination. Also, the lowest values of Mean Absolute Error, Mean-Squared Error, and Root Mean-Squared Error. Compared to the linear regression model, the random forest model has reduced the Mean-Squared Error by 45.51%, Mean Absolute Error by 41.95%, and increased the Coefficient of Determination by 1.34%. The average reduction in error metrics in all the models stays below 38%. The proposed framework can support utility operators, photovoltaic system engineers, long-term energy planners, and policymakers in efficient photovoltaic energy monitoring and forecasting to enhance societal energy security.