| Journal of Smart Sensors and Computing
Received: 28 May 2025; Revised: 12 August 2025; Accepted: 27 August 2025; Published Online: 01 September 2025.
J. Smart Sens. Comput., 2025, 1(2), 25207 | Volume 1 Issue 2 (September 2025) | DOI: https://doi.org/10.64189/ssc.25207
© The Author(s) 2025
This article is licensed under Creative Commons Attribution NonCommercial 4.0 International (CC-BY-NC 4.0)
Smart Farming and Crop Protection by Evaluating the
Performance of Convolutional Neural Networks and
YOLOv4 for Plant Leaf Disease Detection
Sushilkumar S. Salve,* Sujal Ghone, Sanath Lokhande and Rajesh Pandit
Department of Electronics and Telecommunications Engineering, Sinhgad Institute of Technology, Lonavala, Maharashtra, 410401,
India
*Email: sushil.472@gmail.com (Sushilkumar S. Salve)
Abstract
Agriculture plays a significant role in India due to population growth and increased food demands. Hence, there
is a need to enhance the yield of crops. Vegetation is frequently susceptible to a wide range of diseases that arise
due to various seasonal and environmental conditions. These plant diseases not only jeopardize the quality and

Traditionally, the identity and treatment of plant diseases have relied closely on guide inspection and
professional understanding, which may be time-consuming and susceptible to human mistakes. With recent
advancements in technology, there is a growing interest in automated disease detection systems that leverage
artificial intelligence and machine learning techniques. These contemporary solutions provide faster, more
accurate and cost-effective techniques for identifying plant diseases, permitting farmers to take properly timed
preventive and corrective measures. This study presents a novel approach to plant leaf disease detection and
severity classification by leveraging the capabilities of YOLOv4 and Convolutional Neural Networks (CNNs).
These machine learning algorithms have proven great potential in image processing and pattern recognition
tasks, making them appropriate for diagnosing plant situations from visual information. We have used a dataset
containing images of four various plant species, each suffering from different kinds of infections. By training
these models by available datasets, the proposed system can recognize and classify diverse plant diseases with
high accuracy. The performance parameters are evaluated extensively and results are derived. The accuracy of
the CNN and YOLOv4 obtained around 95.5% & 91.0%, respectively.
Keywords: Plant disease; Severity classification; Convolutional neural networks; Support vector machines.
1. Introduction


            





     

 


















2. Related work








et al.












Models which provided higher accuracy were widely used by researchers in plant disease detection. Junction
extraction is used to get better results with neutral community models while using low computer resources than
conventional models. Its average accuracy of 94.8% shows that it is effective even in unfavorable circumstances.
Majeed et al.
[8]
put out a model that is predicated on the residual connection and inception layer. An image
processing structure comprising three stages like image segmentation, Feature extraction, and classification
was used to identify and classify plant diseases. The multi-threshold, and other techniques were used in the
trials, which were conducted on four distinct tomato leaf sections. This approach had a 10-fold cross-
convenience and an overall accuracy of 98.3%.
The primary objective is to inform users of the diagnosed disease name and direct them to an online marketplace
where they can purchase pesticides for the ailments and use them exactly as prescribed. In this study, Support
Vector Machine (SVM) and Artificial Neural Network (ANN) are used to choose two plants such as corn and tomato
for diseasedentification and alerts the customers of the ailment. SVM attains 6273% accuracy, while ANN attains
85%. Table 1 summarises the related wok used for detecting diseases in plant leaves.
Table 1: Summary of related works.
Source
Dataset/Crops
Methods/Results
KC et al.
[1]
Garden Village
Exception with Adam; 99.7%
Saleem et al.
[6]
Garden Village
VGG16;94.8%
M. Bhagat et al.
[3]
Leaf of the Plant
Densenet121(Removed Background); 93%
Dhaka et al.
[4]
Tomato
SVM;98.3%
Ananthi et al.
[5]
Garden Village
DL models; 95%
Rinu et al.
[9]
Tomato and corn
SVM; 60-80%
Hassan et al.
[10]
Fruits e.g., Apple
CNN;70-80%
Kaur P. et al.
[11]
Garden Village
GAN
Zou K. et al.
[7]
Plant Village,
Cassava and Rice
Around 99%for garden Village and rice,75%
for cassava
3. Dataset
Images of plant leaves from Plant Village were used to see the overall performance. A total of 65,345 plant leaves,
including both healthy and diseased samples, were collected and are known as the Plant Village Dataset (Plant
Village).
Apple, blue berry, etc. are among the fourteen amazing crop varieties that are included in the databases. A
selection of example photos is shown in Fig. 1, which displays the number of image files collected for lesion
diagnosis and detection. The availability of water, vitamins, microorganisms, viruses, and fungus are examples of
common stressors that lead to sickness.
[12]
Fig. 1 shows sample examples of plant leaf images. In this study, we have used the Garden Village database.
[13,14]
This dataset contains 58,432 images of 13 different plants, divided into 40 categories of healthy leaves and
various types of diseases. This study used 32,878 photos of 8 kinds of vegetation, apples (9,123 pics), corn
(8,987pics), potato (4,898pics), tomato (11,125 images), and rice (123).
Fig. 1: Sample examples of plant leaf images.
To compare individual sample of plant leaf images taken from dataset are feed to deep learning models. The
process includes classification, feature selection, feature extraction, and preprocessing. Because inaccurate data
in a dataset might alter the appearance of a test, the information series technique is essential in real-time
operations. As a result, it is essential to follow the unusual norm and standard while gathering statistics. Subsets
of the datasets are created using an 80:20 training-to-testing ratio. While the last twenty-seven demonstrate
unique plant leaf diseases, thirteen of the forty trainings that comprise our information are the healthful classes.
The data have 256×256-pixel RGB images showing results of leaves. Based on their image classifications, care
was taken during the photo shoot to ensure that each image captured a single centered leaf. Additionally, the
environment for shooting photos and the lighting are consistent. It is significantly more beneficial to ask
questions about how to use the knowledge effectively after analyzing a variety of data.
[14]
Fig. 2 shows block
diagram of plant leaves disease detection mechanism.
Fig. 2: 
4. Pre-processing and augmentation

        









This is used encoder-decoder architecture and highly defined neural networks to apply semantic leaf disease
division to a collection of plant pictures. Three distinct semantic segmentation models like Lonate-34, Pyramid
Scene Parsing Network (PSPNet), and Seagate
[15]
were employed to detect wounds to provide a high density. After
the lesions are detected, they are classified using several classifiers.
Fig. 3: Grayscale conversion of input dataset.
The plant blades are the input for both semantic segmentation techniques. PSPNet,
[16]
Seagate,
[17]
and Longett-
34,
[18]
are two semantic segmentation models, are used to recreate model. To utilize global reference
information, this modular semantic partition paradigm uses reference aggregation based on many domains.
Local and global cues work together to strengthen the final restriction. Moreover, the U-Net architecture is
extensively recognized for its effectiveness in the responsibilities of semantic division, creating the foundation
of this technique. It can divide a wide range of gadgets, which include clinical imaging for PC pictures and
prescribed in satellite TV. By combining decoder and encoder approaches, the Llenge-34 design aims to enhance
partition model training and increase productivity and efficiency. It consists of a down-sampling server which
shrinks input photos to exclude top level information generate predictions at the pixel level. Linknet-34 connects
the coder and decoder via a jump connection. In addition, low-level features may be communicated without
delay into decoders and integrated with high-degree statistics via the use of jump connections, akin to U-Net.
Division is contemplated inside the object. When it comes to segmental segmentation problems, the PSP Net
plays properly. By assigning a semantic label to a given image, each pixel attempts to share the semantic
segmentation of the image into regions corresponding to different types of objects. Pyramid basin modules are
used by the PSPN to acquire multi-paan reference facts from wonderful components of the doorway picture.
[1 9]
This
makes it easier for the model to assume more pixels than are necessary, particularly for objects of different sizes.
To collect contextual records, PSP Net repeatedly down samples the input characteristic map using a pyramid
shape and uses global pooling at various scales. As an alternative, the conventional layers are used for
characteristic combining and up sampling. Examples of jobs where PSPNET performs well and where pixel-level
segmentation is required are Sean Parsing, Image Segmentation, and Pleasant-Green Object recognition. With a
total of 26 convolutional layers, Seg-Net is an encoder-decoder version. The VGG16 community's development
and contraction routes have thirteen Convo layers. The encoder and decoder networks are separated by two
fully connected (FC) layers. The Rectified Linear Unit (ReLU) is the system that is used to easily and quickly
construct function mappings.
A max-pooling operation with a stride of two comes after each layer for the down sampling of the feature map.
Down sampling increases the number of channels and filter banks, typically doubling them at each step. Each
encoder layer has a corresponding decoder layer, where the decoder samples the data by a factor of two before
passing it to the next feature map. The primary encoder handiest has a multichannel characteristic map, but the
decoder has just three channels. After map output, a multi-dimensional feature is employed to solve a 2-class
problem by using the Sign ID Ed capability to separate plant pixels from the background.
[20]
4.1 Image acquisition
The input information first is captured the use of a Xiaomi USB 2.0 HD webcam that supports taking pictures
video datasets up to 720p and a body fee of 30 frames per 2d (fps). This enters records then undergo photograph
pre-processing where the statistics are normalized into a scale of [0,1] because it consisted of pixels starting from
0-255. Upon normalization, the performance of the CNN model improves ensuring better numerical balance and
quicker convergence. The input records also undergo grayscale conversion as weed detection is predicated
greater upon shapes and textures than color.
[11]
The machine is made greater efficient by way of resizing the facts
to 64×64 pixels for that reason reducing the photograph size and reducing the computational fee.
[21]
4.2 Feature extraction
Now functions are being extracted from the pre- processed photo using 2D Convolution that extracts out all
critical features and styles like edges and so forth.
[16]
It applies 32 filters of the dimensions (3×3 i.e., 64×64 in
grayscale) at the input image. The 2d convolutional layer once more applies to 64 filters of the equal size.
Rectified Linear Unit (ReLU) right here acts because the activation function which converts the terrible values
to zero therefore introducing non- linearity. Fig. 4 shows demonstration of the Rectified Linear Unit (ReLU).
Fig. 4: A demonstration of the Rectified Linear Unit (ReLU).
The non-linearity added through the ReLU activation feature lets in the CNN network to learn extra complex
patterns and functions which are beyond the linear relationships. This makes the network computationally
extra green as fewer neurons prompt immediately, enhancing generalization, appearing as simple threshold
capability. Compared to other activation functions such as sigmoid and tanh, it avoids costly exponential
calculations, thereby enabling faster convergence during network training and helping gradients remain large
during backpropagation. Fig. 5 illustrates how negative inputs are converted to zeros, introducing sparsity in the
activations.
[22]
5. Proposed methodology
We begin by outline the dataset, emphasizing the splitting processes, and discussing the methods of training
models. Also, the suggested model's records flow diagram is displayed in Fig. 5.
[23]
Fig. 5: Flow diagram of plant leaf disease classification mechanism.
5.1 Dataset description
The proposed experiment applied the Garden Village Dataset present in Kaggle which includes 20,639 photos
of high decision of 38 exceptional healthful and diseased leaves bearing on 18 different species of plants, as
shown in Table 2. The model implementation considers segmented images of four plants along with their
diseases.
Table 2: 



Corn
  spot
443
 rust
2193
Fit (Healthy)
1162

 
























5.2 Image segmentation
One crucial aspect of image processing is image division. There are several techniques for dividing pictures,
including the Otsu method, K-peins clustering, borders and spot detection algorithms, etc. One of the best edge
detectors is the Edge Detection as it offers the best, most dependable, and least error-prone real age point
detection.
The following procedures are used to identify edges with the clever edge detector:
1. Smoothing: order to smooth the photo and minimize noise, Gaussian clear out is used.
2. Finding depth gradients: Wherever the picture's gradients have significant magnitudes, the edges are
indicated. Large- magnitude photos' gradients are emphasized as edges.
3. Non-maximum suppression: This technique eliminates erroneous reactions to component detection.
4. Double Threshold: This is a criterion used to determine true edges and abilities.
5. Edge tracking: The weak edges attached to the strong edge are the original or the actual edge, while the weak
edges that are not attached to the strong edge are pressed.
Fig. 5: Flow diagram of plant leaf disease classification mechanism.
6. Classification
6.1 Support Vector Machine (SVM) algorithm










6.2 Convolutional Neural Networks (CNN)
Convolutional Neural Networks (CNNs) are a class of Artificial Neural Network (ANN) designed to automatically
and adaptively learn the features at different levels of detail from the input data (generally images). Image
processing and recognition are two main applications of CNN. They are carried out by means of an optimizer
and activation capabilities.
[25]
DenseNet: One shape of CNN that employs dense connections between layers is referred to as a Dense Net. These
layers are linked to each other through Dense Blocks. If there are similar plant picture variants, this technique
is appropriate. Apple and tomato leaves, as an instance.
[26]
Efficient Net: A specific predetermined set of scaling coefficients is used by the CNN structure Efficient Net to
evenly scale each dimension using the compound coefficient approach. It attains greater precision and efficiency.
Activation functions: The activation function in neural networks overseas transforming the weighted sum of
inputs from nodes in a community layer into an output. ReLU (Rectified Linear Unit) is used in this model
because the hidden layer's activation characteristic outputs the input instantly if it is of good quality; else, it
would output 0
f(x)=max (0, x)
  

 







    



    





Fig. 6:
6.3 Convolution operation




    

󰇛󰇟

󰇠󰇜
 

󰇛󰇟

󰇠󰇜






   






. 2D convolution level function of a CNN.



             
Fig. 7: 2D Convolution level function of a CNN.
Fig. 8:











      








1. LeNet: Designed for the purpose of handwritten digit recognition, this network has two layers of convolution
which are pooled using max pooling to get features. Finally, to the outlet category, we apply a final convolutional
layer using dense layers.
[17]
2. AlexNet: This Network consists of five preliminary convolutional layers. screen 3 layers which only two
layers Z-layers is not fully attached within the quit to provide the classification. It aims to use convolutional,
neural network architecture with right overall performance mentioned in the corresponding studies.
3. MobileNet: This convolutional neural network is designed on deep separable convolution operations, which
reduces the burden of workload to execute the internal operations the initial layers of this mobile- targeted
devices and embedded devices.
[28]
4. ShuffieNet: It is primarily built upon two operations, which the authors defined: the so-called the group
convolutions that could be foreroi4ing, and may be multiple convolutions on part of the input channels, and the
channel shuffie, which uses a random blend the output channels of the convolutions inside the organization.
This structure, advocates say supports a reasonable accuracy with a low computing cost.
[32]
5. EffNet: Along the lines of utilizing in-depth separable convolution, which is akin to MobileNet design.
[32]
and
ShuffieNet networks; however, it presents a new convolution block that reduces the computational cost and
outperforms state-of-the-art for certain known databases.
[33]
6. Sheaf Attention Network: Efficient segmentation and classification using Convolutional neural networks with
Sheaf Attention Networks (CSAN).
[34]
7. Results and discussion















7.1 Using deep learning (Image-based approach):







Workflow:

o 

o 
o 
o 

o 

o 
o 
o 
7.2 Evaluation metrics







 











































7.3 Experimental analysis
7.3.1 Metric value






Table 3:
True cases
False cases
% Error
% Success
TN
TP
FN
FP
6
11
2
1
15
85
6
12
2
0
10
90
6
12
0
2
10
90
3
16
0
1
5
95
7
12
1
0
5
95
11
9
0
0
0
100
8
12
0
0
0
100
4
16
0
0
0
100
14
6
0
0
0
100
10
10
0
0
0
100
75
116
5
4
-
-
-
-
-
-
4.5
95.5
Fig. 9:


 
  








 






Table 4 summarizes the result metrics of CNN technique that was implemented on the very same setup for a
through comparison. Fig. 10 shows results using CNN technique.

 
 



















Table 4: Results during field testing using CNN.
Field
Trial
True Cases
False Cases
% Error
% Success
TN
TP
FN
FP
1
6
11
2
1
15
85
2
6
12
2
0
10
90
3
6
12
0
2
10
90
4
3
16
0
1
5
95
5
7
12
1
0
5
95
6
11
9
0
0
0
100
7
8
12
0
0
0
100
8
4
16
0
0
0
100
9
14
6
0
0
0
100
10
10
10
0
0
0
100
Total
75
116
5
4
-
-
Average
-
-
-
-
4.5
95.5
Fig. 10: Results of field trail using CNN.
CNN classification and YOLOv4 supervised algorithms for a comparison-based study and detailed analysis in
the search for the best algorithm to be implemented.
Fig. 11 shows the Confusion Matrix for YOLOv4 technique YOLOv4 in field testing lacks true positive cases (TP =
103), whereas its greater true negative (TN = 79) and false negative (FN = 12)
  




Fig. 11: Confusion matrix for YOLOv4 technique.
Fig. 12:


Extract features like color histograms, shape descriptors, texture.
Web App Integration:
Once we have a model, we have:
Deploy via Flask backend.
The homepage layout of this graphical user interface is displayed in Fig. 13, while the functional user interface
of the proposed prediction model is shown in Fig. 14.
We looked at different plant disease datasets from Kaggle, including apples, corn, tomato, and grapes. The test
scores for the disease and severity classifiers are shown in Table 5.
This study was to determine the optimal performance of various deep learning (DL) algorithms in classification
of plant leaf diseases. CNN classification and YOLOv4 supervised algorithms for a comparison-based study and
detailed analysis in the search for the best algorithm to be implemented.
Table 5: 
Plant
Disease name
Accuracy (%)
Apple
Healthy
92.5
Apple scab
100
Black rot
100
Corn
Healthy
100
Common rust
100
Gray leaf spot
57.5


90
Black rot
High
100
Low
85
Esca
High
90
Low
100
Siriasis
High
95
Low
100
Tomato
Healthy
100
Leaf mold
100
Septoria
100
Early blight
95
Fig. 13: Homepage of graphical user interface.
Fig. 14: 
7.4 Discussion
Deep learning (DL) is changing the scenario when it comes to diagnosing plant diseases using digital images. To
get the right response quickly, the models need to be fast and accurate. The design of the network chosen will
depend on to improve or reduce accuracy of prediction model. If we need to change things up often, Dense Net
(CNN) is the quickest option, though it can be a bit unstable. On the other hand, if we want top-notch accuracy,
we might want to look at GoogLeNet or AlexNet. They found that while InceptionV3 had the lowest accuracy in
their tests, AlexNet, though not performing as well as expected, still outperformed it overall.
The prototype in this study is being developed and implemented using CNN classification and YOLOv4
supervised learning algorithms for a comparison-based study and detailed analysis in the search for the best
algorithm to be implemented. This step is particularly necessary for accurate classification of plant leaf diseases
detection on different geographical locations and regions. Fig. 15 shows the accuracy comparison by various
detection techniques. Fig. 16 shows the results of comparision between YOLOv4 vs CNN.
Now as we observe in Table 6, a comparison has been stated amongst accuracies of four other methods being
generally used in effective object detection and classifications with the two root methods mentioned in this
study. Jun Zhang et al.
[33]
mentioned in his study about the higher accuracy of the original ViT model due to its
stronger sequence modelling abilities and unique capabilities to capture long-range dependencies. But when
we carefully consider both CNN, ViT and YOLOv4 in a comprehensive way, then the CNN model due to its better
balance for local and global features, results in an overall better performance and improved classification.
Additionally, despite having several depth layers, the deepest networks are ResNet50, ResNet101, and
InceptionV3 are showed comparatively low accuracy. Lastly, a time and performance analysis of several CNNs
was suggested.
Table 6: Comparison with other detection techniques.
Model Name
Accuracy (%)
VGG16
86.21
GoogleNet
79.23
AlexNet
80.09
ViT
89.09
YOLOv4
91.00%
CNN
95.50%
Fig. 15: Accuracy comparison by various detection techniques (deep learning algorithms).
Fig. 16: Results of YOLOv4 vs CNN.
8. Conclusion and future scope
The main goal of the plant disease detection and severity classification model is to uses images of leaves of infected
plant species to precisely identify plant diseases and their severity levels. This model focuses on using advanced
image processing techniques based on CNN to determine multi-class detection of plant leaf disease detection,
which helps to extract important characteristics required for efficient categorization. Utilizing these extracted
metrics, the model facilitates early and accurate disease identification in a variety of plants, allowing producers
to take prompt and suitable corrective action. Apple, corn, grapes, and tomatoes are the four plant species on
which the system has been extensively tested; each of these plant species has two to three different diseases. This
makes the model approachable and useful for actual agricultural applications by enabling effective and
straightforward forecasts. Both YOLOv4 and CNN-based models exhibit notable gains in plant disease detection
accuracy, according to the experimental investigation. The addition of severity categorization improves the
model's usefulness in practice by revealing information about the disease's course in addition to its
identification.
Funding Declaration
This research did not receive any specific grant from funding agencies in the public, commercial, or not-for-
profit sectors.
Data Availability Statement
The datasets generated and/or analyzed during the current study that support the findings are available from
the corresponding author upon reasonable request.
Conflict of Interest
There is no conflict of interest.
Artificial Intelligence (AI) Use Disclosure
The authors declare that artificial intelligence (AI)-assisted tools were used only for language refinement,
grammar improvement, and manuscript structuring purposes during the preparation of this work. All technical
content, experimental implementation, results, and interpretations were independently developed and verified
by the authors.
Supporting Information
Not applicable.
References


11

Y. Gai, H. Wang, Plant disease: a growing threat to global food security, Agronomy, 2024, 14, 1615, doi:
10.3390/agronomy14081615.


 



          21  


V. P. Ananthi, Fused segmentation algorithm for the detection of nutrient deficiency in crops using SAR
images,  


          
         9  


      
24


            
 
170





10

 
     
14



3


   
13






27


       

14


 



       2020  



 




25

           
       119  



       75  



  
105

S. S. Salve, S. P. Narote, Iris recognition using SVM and ANN, In 2016 International Conference on
Wireless Communications, Signal Processing and Networking (WiSPNET), IEEE, 2016, 474-478, doi:
10.1109/WiSPNET.2016.7566179.


82



214


82



         2070  






70



        
218


         12  



11

S. S. Salve, S. P. Narote, Performance evaluation of efficient segmentation and classification-based iris
recognition using sheaf attention network, Journal of Visual Communication and Image Representation,
2024, 103, 104262, doi: 10.1016/j.jvcir.2024.104262.

S. S. Salve, S. S. Chakraborty, S. Gandhewar, S. S. Girhe, A deep learning framework for smart agriculture:
real time weed classification using convolutional neural network, Journal of Smart Sensors and
Computing, 2025, 1, 25205, doi: 10.64189/ssc.25205.
Publisher Note: The views, statements, and data in all publications solely belong to the authors and
contributors. GR Scholastic is not responsible for any injury resulting from the ideas, methods, or products
mentioned. GR Scholastic remains neutral regarding jurisdictional claims in published maps and institutional
affiliations.
Open Access
This article is licensed under a Creative Commons Attribution-NonCommercial 4.0 International License, which
permits the non-commercial use, sharing, adaptation, distribution and reproduction in any medium or format,
as long as appropriate credit to the original author(s) and the source is given by providing a link to the Creative
Commons License and changes need to be indicated if there are any. The images or other third-party material
in this article are included in the article's Creative Commons License, unless indicated otherwise in a credit line
to the material. If material is not included in the article's Creative Commons License and your intended use is
not permitted by statutory regulation or exceeds the permitted use, you will need to obtain permission directly
from the copyright holder. To view a copy of this License, visit: https://creativecommons.org/licenses/by-
nc/4.0/
© The Author(s) 2025