Computationally efficient CNN–vision transformer fusion framework for efficient fault diagnosis in solar photovoltaic system
1Department of Electronics and Telecommunication, Pimpri Chinchwad College of Engineering, Savitribai Phule Pune University, Pune 411044, Maharashtra, India
2Department of Electronics and Telecommunication, Pimpri Chinchwad College of Engineering, Savitribai Phule Pune University, Pune 411044, Maharashtra, India
3Department of Electronics and Telecommunication, Pimpri Chinchwad College of Engineering, Savitribai Phule Pune University, Pune 411044, Maharashtra, India
J Ther Eng 2026; 12(5): 1760-1777 DOI: 10.47481/jten.0064
Full Text PDF

Abstract

Solar photovoltaic (PV) power plants are a key pillar of the global transition to renewable energy, and rapid technological development has led to widespread deployment. As deployment increases, reliable fault detection is essential to ensure long-term performance, safety, and financial sustainability of PV systems. Common defects such as cracking, delamination, discoloration, and shadowing can significantly reduce power output and cause irreversible failure of the module unless these defects are detected in the early stages. Current fault-detection systems are mostly
based on classical image-processing methods or single deep-learning models, which cannot adapt to changing environmental and operational conditions and cannot cope with complex visual variations. Moreover, the limited variety and quality of datasets hinder the robustness and
generalizability of the model. To examine four key PV defects under different lighting and weather conditions, a high-quality RGB dataset was created. The proposed work introduces a new hybrid CNN-ViT architecture that incorporates depth wise separable convolutions to dramatically
decrease the computational complexity while keeping discriminative spatial features. Squeeze-and-excitation blocks are also added to improve defect-aware feature learning and provide adaptive channel-wise recalibration, which can be used to effectively highlight photovoltaic defects
such as cracks, delamination, and discoloration. Moreover, a multi-scale attention mechanism with different kernel sizes (1×1, 3×3, and 5×5) is proposed to encode defect patterns at different spatial resolutions. To overcome these shortcomings, this paper proposes a hybrid deep learning
architecture that combines a customized convolutional neural network (CNN, ResNet-50V2) and a Vision Transformer (ViT). The CNN part captures fine-grained local spatial features, while the ViT uses global self-attention to be long-range contextual dependencies. The experimental
performance is high, with an accuracy of 95.85, a precision of 95.62, a recall of 95.90, and an F1-score of 95.76. The model proposed in this research paper is compared with ultramodern hybrid and transformer-based architectures, including DL-LGM, Cat Boost-GB, DeepF-SVM, DeiT, CCT, and Swin Transformer. Overall, the framework offers a powerful, scalable, intelligent PV monitoring solution that enables prompt fault detection and enhances system reliability and efficiency.