A Transfer Learning Approach For Efficient
Classification of Waste Materials
Md Humaion Kabir Mehedi , Irfana Arafin , Murad Hasan , Farhin Rahman , Rufaida Tasin , and
Annajiat Alim Rasel
Department of Computer Science and Engineering
Brac University
66 Mohakhali, Dhaka - 1212, Bangladesh
{humaion.kabir.mehedi, irfana.arafin, murad.hasan, farhin.rahman, rufaida.tasin}@g.bracu.ac.bd
annajiat@gmail.com
Abstract—The authors of this study have used the Waste
Classification Dataset to build a highly accurate model that
classified rubbish into two distinct groups in an effort to
address the problem of waste classification for various classes
of discarded material. VGG16, MobileNetV2, and a baseline 6
layer CNN model are used in the experiments. The VGG16 model
have achieved 96.00% accuracy, while the MobileNetV2 model
achieved 95.51%, and the baseline CNN model achieved 90.61%
accuracy. The garbage in the input picture can be correctly
classified by the neural network model. The experimental findings
are compared to other studies in the same area. In addition,
LIME is also implemented to make our models’s prediction more
explainable. This investigation’s experimental applications are
geared on facilitating more precise trash classification.
Index Terms—VGG16, CNN, MobileNetV2, Transfer Learning,
waste classification, deep learning
I. I NTRODUCTION
With the growing population, proper actions are required to
prevent the accumulation of non-recyclable waste in landfills
around the world and the length of time it takes for the
majority of its materials to biodegrade will affect our lifestyle
in the near future. Accumulation of waste can facilitate
the transmission of disease by flies, mosquitoes, and other
insects. Landfill construction and deforestation destroy natural
habitats and contaminate soil and water with toxic chemicals.
This environmental pollution from waste materials can wreak
havoc on the food chain, leading to increased human and
animal health problems. There are three primary reasons why
waste has accumulated over the past fifty years. First, there
are not enough recyclable items on the market, despite the
fact that businesses are developing more environmentally
friendly products. Overpopulation is the second contributing
factor. A large population requiring a variety of resources
poses a complex logistical challenge while dealing with waste
generation; consequently, more recyclable products end up in
landfills or the ocean, affecting tens of thousands of marine
species. Third, the lack of involvement of our society in issues
such as climate change. The global population generates
approximately 7 to 9 billion tons of waste annually, of which
979-8-3503-3286-5/23/$31.00 ©2023 IEEE
70 percent is mishandled and ends up in landfills, polluting
natural environments and posing new health risks such as
microplastics in the ocean. This is the total amount of used,
unwanted, and discarded items that are man-made. Municipal
Solid Waste (MSW) comprises trash from urban areas and
their influence zones. MSW accounts [1] for approximately
33% of the 2 billion tons of urban waste produced annually.
Daily garbage production per person ranges from 0.1 to
4.5 kilos, with an average of 0.7 kilograms. Population
expansion and the necessity for extensive use of natural
resources for industrial development and the survival of our
civilization are predicted to increase MSW to 3,4 billion
tons by 2050. A model of a circular economy that is fully
implemented could solve the accumulation problem, climate
change, and supply crises in certain regions of the world.
Its three principles, reducing waste and pollution, reusing
materials, and regenerating nature, result in a more efficient
method of managing natural resources, which are only valued
sometimes. The implementation of a complex plan is hindered
by technological, engineering, and logistical limitations.
In this work, we have implemented both transfer learning
and convolutional neural network (CNN) [2] approaches to
compare their performance on waste classification datasets.
VGG16 and MobileNetV2 have been fine-tuned to achieve
the best possible performance. After that, we have also implemented Explainable AI to validate our proposed model with
the use of LIME. Therefore, in Section 2, we have provided a
brief overview of the existing literature on waste classification;
in Section 3, the methodology of the study, the construction
of the deep learning model of trash classification, the transfer
learning strategy, and the model evaluation procedure have
been described in detail. The experimental findings and model
comparison are presented in Section 4. Lastly, section 5
concludes with a brief summary and future projections.
II. R ELATED W ORKS
Waste management has emerged as a major concern in
recent years as a global challenge for climate change as
it is known for directly impacting earth’s climate change.
Recycling, reuse, and waste reduction are essential practices
for effective waste management. Waste classification is
required in the modern era in order to manage waste and
carry out these activities with the aid of technology. Many
researchers have suggested various methods to accumulate
this process, with machine learning and deep learning being
the most popular. Olugboja Adedeji and Zenghui Wang
[3] have discussed a system in which a 50-layer residual
Convolution Neural Network (ResNet-50) has been proposed
to classify trash images into a variety of different groups.
The proposed model makes use of Gary Thung and Mindy
Yang’s trash image dataset [4], which they developed in
order to facilitate the model. Without involving any humans
in the process, this method achieved an accuracy of 87%
when categorizing images of trash. ConvNext was the main
architecture used in a different study [5] to categorize
various waste images. The proposed Mask R-CNN with
Convnext by the authors outperformed the currently used
waste classification techniques. The study used a customized
dataset that was collected and consisted of 1660 images, with
about 400 images per class. Four categories of trash exist:
dry, moist, hazardous, and recyclable.The authors claim to
achieve 79.88% accuracy in waste classification using their
method. An approach to classifying solid waste using five
layer DCNN architectures was suggested in [6], which was
written by its author. When classifying the waste dataset, both
a four-layer and a five-layer deep convolution neural network
architecture were utilized; however, the five-layer architecture
performed significantly better than the four-layer architecture.
After being cleaned, the images in the waste dataset, which
included pictures of paper, glass, plastic, and organic waste,
were converted to a size of 224 by 224 pixels. According to
the findings of this research, the DCNN architecture, which
consists of various layers, was able to achieve 70% accuracy
when distinguishing between various types of waste.
With regard to performing convolutional neural networks
on a waste image dataset, Qiang Zhang and his co - authors
[7] primarily focused on a DensNet169 pre-trained model
and transfer learning approach. NWNU-TRASH has used the
waste image dataset to examine the model. This dataset, which
the authors claim is well balanced and diverse, consists of
18911 images of various classes. The image classification
models AlexNet, VGG, and GoogLeNet were chosen for this
study to use on this dataset. Accordingly, AlexNet, VGG 16,
and GoogleNet V2 had accuracy scores of 68.78%, 70.12%,
and 75.42%. The DenseNet169 deep learning model was also
tested by the authors, both with and without a transfer learning
strategy. The outcomes demonstrate that accuracy of 82.80%
was significantly higher with the use of DenseNet169 with
transfer learning than the other approaches. This analysis
demonstrates that DenseNet169, following transfer learning, is
a better image classification algorithm. Deep learning has been
used to the problem of detecting coastal rubbish, according
to the authors of [8] but it has experienced difficulties due
to the difficulty of detecting small things and the poor performance of the object detection model. Using feature fusion,
dense blocks, and focal loss, the Multi-Strategy Deconvolution
Single Shot Multibox Detector (MS-DSSD) is an end-to-end
trained feed-forward network that overcomes these obstacles.
MS-DSSD513 gets a higher mAP on PASCAL VOC2007
and our coastal rubbish dataset than state-of-the-art object
identification methods, with 82.2% and 84.4%, respectively.
Object detection is a computer vision problem that asks what
object is located at a given set of coordinates. To meet
this need, CNN-based object identification systems have been
developed, such as the You Only Look Once (YOLO) series,
Single Shot MultiBox Detector (SSD), and RetinaNet.
Data augmentation approaches are used to enhance the
data with previous information, enabling the model to learn
a broader range of visual characteristics and transformations.
Ma et al proposal of the.’s L-SSD framework for automatic
trash sorting recognition, Panwar et alcreation of AquVision
for identifying hazardous litter, and Shi et al network topology’s deep network development, and usage of short-circuit
connections. Kraft et al. presented a low-cost method for
detecting trash in low-level images collected by unmanned
aerial vehicles (UAVs) during automated patrol missions. The
convolutional YOLO network, the SSD network structure of
24 convolutional layers and 2 full connection layers, and the
DSSD multitasking network each collect features in a distinct
manner. Both the probability that the bounding box contains
the target and the precision with which this box is drawn affect
the confidence score.
SSD is a multiscale detection model with a variety of
semantic qualities and positional sensitivity that detects the
same image through the use of many scaled feature maps.
It employs a data augmentation approach to improve generalization ability and accuracy, and it filters positive and
negative samples based on the IoU between ground-truth
boxes and preceding boxes. Both MS-DSSD and Residual101 consist of convolution and deconvolution layers. Several
sliding window detection algorithms, including YOLO, SSD,
and DSSD, have difficulty drawing conclusions from a subset
of an image’s characteristics. SSD is one of the most effective
single-stage object detection algorithms, however, it fails badly
at recognizing small targets.
MS-DSSD is a set of convolution and deconvolution layers
that utilizes Residual-101 as its foundation and a deconvolution module to integrate spatial contextual information. By
employing dense blocks after the Conv2 x and Conv5 x layers,
more global and local feature information is supplied to the
model, and the capacity of dense blocks to migrate features
is enhanced. Improving detection performance requires the
addition of two essential image feature representation characteristics: image invariance and equivalence transformation.
The purpose of this study is to develop a deep learning
system that can automatically collect photos or videos of waste
from a camera equipped with object recognition, detection, and
prediction, and then classify the waste into multiple categories,
including cardboard, glass, metal, paper, plastic, and trash. The
authors in [9] suggests using a Convolutional Neural Network
(CNN) with a grey-level co-occurrence matrix to identify
and detect rubbish (GLCM). After testing multiple iterations
of the well-known deep convolutional neural network architecture, Recycle Net’s developers determined that Inceptionv4 had the highest test accuracy at 90%. Auto Trash is an
entirely automated trash compactor that separates recyclable
and biodegradable waste into separate bins. The input is
received by one group of neurons in the network, processed
by their axons, and then transmitted to another group.
Convolutional neural networks (CNN) are artificial neural
networks used to detect objects in images, classify images
into groups, and cluster images according to similarities. CNN
extension that utilizes Region Proposal Networks (RPN). It
consists of a deep convolutional network and a layer for
combining features from a region of interest (ROI). With the
Object Detection API of TensorFlow, you can rapidly build,
train, and deploy object detection models. Multiplying the
accuracy and recall numbers while identifying bounding boxes
yields the mAP.
The created approach consists of four stages: data gathering,
model creation, model training, and model validation. NumPy,
Matplotlib, OS, TensorFlow, Utils, and OpenCV are used
as libraries. The newly trained waste detection classifier is
evaluated with the Python program OpenCV. The Tarfile file
format is utilized for storing huge datasets. Cardboard, glass,
metal, paper, plastic, and trash are the six parent categories for
which the Improved R-CNN Model for Garbage Classification
is complete. In this study, the Faster R-CNN algorithm is used
to classify several objects within a single image into one of
six waste categories (Cardboard, Metal, Glass, Paper, Plastic,
Garbage).
The results demonstrate that the model accurately classifies
six types of trash and that it takes around 8.05 seconds to
predict a single object from an image. It is advised that future
studies incorporate photographs of locally collected waste
goods, as the dataset contains images that differ slightly from
those of local garbage.
III. M ETHODOLOGY
A. Dataset Description
There are 24,705 photos in the dataset, split evenly between organic (13,880) and recyclable (10,825) solid trash.
This dataset is known as “Waste Classification Dataset”
[10]. These data have been reorganized and represented from
the original dataset created by Sashaank Sekar and hosted
at https://www.kaggle.com/techsash/waste-classification-data.
Out of the first 25,077 photos in the Kaggle dataset, 13,966
are organic and 11,111 are recyclable. Figure 1 shows some
the sample data from our dataset.
B. Dataset Preprocessing
To facilitate training, testing, and validating our deep learning models, we have shrunk our photos to 224 × 224. We have
also separated our data into a training set, a validation set, and
a test set. There were a total of 19764 pictures in our training
Fig. 1. Sample images from datset. Organic (upper), and Recyclable (bottom)
set, 2470 in our validation set, and 2471 in our test set. Our
models are complete after training and validation on the data
we provided. The models never saw the test set before, but it
was used to compare their results using a variety of metrics.
As a solution to the problem of our dataset’s lack of data, we
have normalized our models by dividing the pixel values by
255 in order to get better accuracy from the models, and we
have enhanced the photos.
C. Model Architectures
In this study, we used a 6 layer baseline convolutional neural
network (CNN) in addition to two pretrained convolutional
neural network (CNN) models: VGG16 [11] and MobileNetV2
[12] [13]. For the models, we have used the pretrained weights
obtained from ImageNet. In order to eliminate any kind of bias
during our quantitative comparison of the different models,
we have used the exact identical process to customize each
of these three models so that they can categorize garbage.
An Adam optimizer with a learning rate of 0.00001, binary
cross entropy as the loss function, and a sigmoid activation
function were used in the output layer, which was comprised
of 2 neurons.
We have used an Adam optimiser with a learning rate of
0.00001, binary cross entropy as loss function, and sigmoid
activation function in the output layer consisting of 2 neurons.
IV. R ESULT A NALYSIS
By making predictions on our test set, we have evaluated
the efficacy of our models using quantitative performance
evaluation metrics such as accuracy (1), precision (3), recall
(3), and f1 score (4) and confusion matrix.
Accuracy =
TN + TP
TP + FP + TN + FN
(1)
Here, TN = True negative, TP = True positive, FN = False
negative, FP = False positive.
P recision =
TP
TP + FP
(2)
Here, TP = True positive, FP = False positive.
Recall =
TP
TP + FN
(3)
Here, TP = True positive, FN = False negative.
F 1Score =
2 ∗ P recision ∗ Recall
P recision + Recall
(4)
Fig. 2. Confusion matrix of VGG16
A comparative analysis between VGG16, MobileNetV2,
and baseline CNN in the basis of these evaluation parameters
is depicted in Table I.
TABLE I
A COMPARISON TABLE BETWEEN VGG16, M OBEILE N ET V2, AND CNN
ON PERFORMANCE EVALUATION METRICS
Model
Accuracy
Precision
Recall
F1 Score
VGG16
96.00%
96.00%
96.00%
96.00%
MobileNetV2
95.511%
95.00%
96.00%
95.00%
CNN
90.61%
90.38%
90.75%
90.51%
After analyzing our data on various models, we discovered
that some performed admirably, achieving high accuracy with
impressive precision, recall, and F1-score rate. CNN and
MobileNetV2 have the lowest accuracy among all models at
95.51% and 90.61% respectively, as shown in table I. The
highest accuracy is 96.00% for model VGG16. Precision,
recall, and F1 all have very high scores of almost 96% in
both the VGG16 and MobileNetV2 models. CNN has the
lowest Precision, Recall, and F1 scores, with 90.38, 90.75,
and 90.51 percent, respectively. As we can see from the
table, even though the CNN model has the lowest accuracy,
the overall model has over 90% accuracy, which is quite
impressive.
Fig. 3. Confusion matrix of MobileNEtV2
In fig.2 our proposed VGG16 correctly classified 1347
organic and 1023 recyclable waste. On the other hand fig.
3 MobileNetV2 misclassified 69 and 42 for both organic
and recyclable respectively. In addition, baseline CNN also
performs well to classify organic and recyclable waste which
is shown in fig. 4.
Fig. 5. Explainability of VGG16 model
Fig. 4. Confusion matrix of CNN
Here, the explainability of VGG16 model has been checked
by using LIME. As can be seen from the fig.5 above, Our
model can identify the proper regions of the image to identify
it as an organic or recyclable waste. The heatmap shows the
regions of the image that are more prioritized as blue and less
significant areas as red.
V. C ONCLUSION
Finally, we have suggested a deep learning-based approach
to waste categorization that can effectively identify and isolate
individual waste components. One can use this technique to
automatically sort trash, which will cut down on human labor
and help avoid contamination and pollution. Using the results,
VGG16 was able to achieve an accuracy of 96% on the waste
classification dataset. Assembling Two Entities
without or with little human intervention, adopting our
technology will make garbage disposal quicker and more
intelligent. Increases in system accuracy are possible when
more images are added to the collection. More waste items
will be sorted into categories in the future as we work to refine
our system by adjusting some of the current parameters. In
addition, we will implement more advance model to increase
the model’s performance.
R EFERENCES
[1] D. G. Solla, “Advanced waste classification with Machine Learning,”
https://towardsdatascience.com/advanced-waste-classification-withmachine-learning-6445bff1304f, jan 14 2022, [Online; accessed
2022-12-1].
[2] M. Al-Amin, D. Z. Karim, and T. A. Bushra, “Prediction of rice
disease from leaves using deep convolution neural network towards a
digital agricultural system,” in 2019 22nd International Conference on
Computer and Information Technology (ICCIT), 2019, pp. 1–5.
[3] A. Olugboja and Z. Wang, “Intelligent waste classification system using
deep learning convolutional neural network,” Procedia Manufacturing,
vol. 35, pp. 607–612, 01 2019.
[4] G. Thung and M. Yang, “Classification of trash for recyclability status,”
2016.
[5] J. Qi, M. Nguyen, and W. Yan, “Waste classification from digital
images using convnext,” in Pacific-Rim Symposium on Image and Video
Technology, 2022.
[6] A. Altikat, A. Gulbe, and S. Altikat, “Intelligent solid waste classification using deep convolutional neural networks,” International Journal of
Environmental Science and Technology, vol. 19, no. 3, pp. 1285–1292,
2022.
[7] Q. Zhang, Q. Yang, X. Zhang, Q. Bao, J. Su, and X. Liu, “Waste
image classification based on transfer learning and convolutional neural
network,” Waste Management, vol. 135, pp. 150–157, 2021.
[8] C. Ren, S. Lee, D.-K. Kim, G. Zhang, and D. Jeong, “A multi-strategy
framework for coastal waste detection,” Journal of Marine Science and
Engineering, vol. 10, no. 9, p. 1330, 2022.
[9] ——, “A multi-strategy framework for coastal waste detection,” Journal
of Marine Science and Engineering, vol. 10, no. 9, p. 1330, 2022.
[10] N. Nnamoko, J. Barrowclough, and J. Procter, “Waste Classification
Dataset,” /datasets/n3gtgm9jxj/2, [Online; accessed 2022-12-01].
[11] M. H. K. Mehedi, E. Haque, S. Radin, M. Rahman, M. T. Reza, and
M. G. R. Alam, “Kidney tumor segmentation and classification using
deep neural network on ct images,” 12 2022.
[12] M. H. K. Mehedi, A. S. Hosain, S. Ahmed, S. T. Promita, R. K. Muna,
M. Hasan, and M. T. Reza, “Plant leaf disease detection using transfer
learning and explainable ai,” in 2022 IEEE 13th Annual Information
Technology, Electronics and Mobile Communication Conference (IEMCON), 2022, pp. 0166–0170.
[13] K. M. Hasib, S. Sakib, J. A. Mahmud, K. Mithu, M. S. Rahman, and
M. S. Alam, “Covid-19 prediction based on infected cases and deaths
of bangladesh using deep transfer learning,” in 2022 IEEE World AI IoT
Congress (AIIoT), 2022, pp. 296–302.
0
You can add this document to your study collection(s)
Sign in Available only to authorized usersYou can add this document to your saved list
Sign in Available only to authorized users(For complaints, use another form )