Available online at www.sciencedirect.com Available online at www.sciencedirect.com Available online at www.sciencedirect.com ScienceDirect Procedia Computer Science 00 (2025) 000–000 Procedia Computer Science 00 (2025) 000–000 Procedia Computer Science 258 (2025) 3490–3499 www.elsevier.com/locate/procedia www.elsevier.com/locate/procedia International Conference on Machine Learning and Data Engineering International Conference on Machine Learning and Data Engineering A Deep Neural Framework for Self-Injurious Behavior Detection in A Deep Neural Framework for Self-Injurious Autistic Children Behavior Detection in Autistic Children Uday Singha , Shailendra Shukla1 , Manoj Madhava Gore1,∗ a Computer Science and Engineering Department, a 1 1,∗ Utter Pradesh,India Motilal Nehru National Institute of Technology Allahabad,Prayagraj Uday Singh , Shailendra Shukla , Manoj Madhava Gore a Computer Science and Engineering Department, Motilal Nehru National Institute of Technology Allahabad,Prayagraj Utter Pradesh,India Abstract Video Action Detection (VAD) is becoming increasingly common, with distributed methods shifting towards edge computing for Abstract real-time processing. The limited diversity and size of the existing Self-Stimulatory Behaviours Dataset (SSBD) hinder the genVideo ActionofDetection (VAD)Behaviors is becoming increasingly common, withdetection distributed shifting towards edge computingand for eralizability Self-Injurious (SIB) detection models. Early of methods SIB is imperative for timely intervention real-time processing. The limited diversity and size of the existing Self-Stimulatory Behaviours Dataset (SSBD) hinder the gensupport for individuals with Autism Spectrum Disorder (ASD), emphasizing the need for accurate recognition systems. Addressing eralizability of Self-Injurious Behaviors in (SIB) detection models. detection SIB is imperative for timely intervention and these challenges requires advancements dataset diversity, modelEarly accuracy, and of generalizability to diverse populations and envisupport forThis individuals with Autism Spectrumfor Disorder (ASD),of emphasizing the need systems. Addressing ronments. paper proposes a framework the detection SIB in children. Wefor useaccurate a hybridrecognition approach combining CNN and these requires advancements in dataset diversity, model accuracy, to diverse populations and enviLSTMchallenges to detect SIB effectively. To enhance the recognition capabilities, we and add generalizability new actions to the SSBD dataset to increase its ronments. This paper proposes a framework for the detection of SIB in children. We use a hybrid approach combining CNN and size. This addition enriches the dataset, enabling more comprehensive training of our recognition models. We then extract frames LSTM to detect effectively. To enhance the recognition capabilities, we add new actions to and the SSBD datasetTerm to increase its from videos andSIB apply augmentation techniques to the frames. The ConvLSTM, EfficientNet, Long-Short Recurrent size. This addition enriches the dataset, enabling more comprehensive training of our recognition models. We then extract frames Convolutional Networks (LRCN) models are used for SIB action detection. Among these, the LRCN model demonstrates superior from videos and apply augmentation to the frames. The ConvLSTM, and (77.17%). Long-Short Term Recurrent performance, achieving an accuracy oftechniques 92.62%, surpassing ConvLSTM (80.33%) EfficientNet, and EfficientNet The LRCN model Convolutional Networks modelsofare used highlighting for SIB action these, theprediction LRCN model superior achieved a Mean Squared(LRCN) Error (MSE) 0.045, itsdetection. reliabilityAmong in minimizing errorsdemonstrates for action detection. performance, achieving an accuracy 92.62%, surpassing (80.33%) and EfficientNet (77.17%).ofThe LRCN model This underscores the effectiveness of of hybrid models for videoConvLSTM action recognition, emphasizing the importance early detection in achieved a Mean Squared Error (MSE) of 0.045, highlighting its reliability in minimizing prediction errors for action detection. supporting individuals with ASD. This underscores the effectiveness of hybrid models for video action recognition, emphasizing the importance of early detection in supporting individuals with ASD. © 2025 The Authors. Published by Elsevier B.V. This is an open accessPublished article under the CC BY-NC-ND license (http://creativecommons.org/licenses/by-nc-nd/4.0/) © 2025 The Authors. by Elsevier B.V. © 2025 The under Authors. Published by B.V. committee of the International Conference on Machine Learning and Data EngiPeer-review responsibility of Elsevier the This is an open access article under the scientific CC BY-NC-ND license (https://creativecommons.org/licenses/by-nc-nd/4.0) This is an open access article under CC BY-NC-ND (http://creativecommons.org/licenses/by-nc-nd/4.0/) neering. Peer-review under responsibility of thethe scientific committeelicense of the International Conference on Machine Learning and Data Engineering Peer-review under responsibility of the scientific committee of the International Conference on Machine Learning and Data Engineering. Keywords: Machine Learning, Autism, Self Injurious Behaviour, video classification, Deep Learning repetitive behavior; Keywords: Machine Learning, Autism, Self Injurious Behaviour, video classification, Deep Learning repetitive behavior; 1. Introduction 1. Introduction Autism Spectrum Disorder (ASD) has become an increasingly significant health issue for infants and toddlers [29]. A key behavioral characteristic in children with ASD is self-stimulatory behavior, known as stimming, which Autism Spectrum Disorder (ASD) has become an increasingly significant health issue for infants and toddlers can sometimes escalate to self-injurious behavior (SIB). SIB involves intentional actions that cause harm to oneself, [29]. A key behavioral characteristic in children with ASD is self-stimulatory behavior, known as stimming, which such as head-banging, biting, and scratching [30]. These behaviors are not typically suicidal but can be mistaken for can sometimes escalate to self-injurious behavior (SIB). SIB involves intentional actions that cause harm to oneself, such as head-banging, biting, and scratching [30]. These behaviors are not typically suicidal but can be mistaken for ∗ Corresponding author. Tel.: +0-000-000-0000 ; fax: +0-000-000-0000. E-mail address: uday.2020rcs07@mnnit.ac.in ∗ Corresponding author. Tel.: +0-000-000-0000 ; fax: +0-000-000-0000. 1877-0509 © 2025 The Authors. Published by Elsevier B.V. address: uday.2020rcs07@mnnit.ac.in ThisE-mail is an open access article under the CC BY-NC-ND license (http://creativecommons.org/licenses/by-nc-nd/4.0/) 1877-0509 © 2025 2025 The Published by B.V. 1877-0509 © The Authors. Authors.of Published by Elsevier Elsevier B.V. Peer-review under responsibility the scientific committee oflicense the International Conference on Machine Learning and Data Engineering. This is an open access article under the CC BY-NC-ND (https://creativecommons.org/licenses/by-nc-nd/4.0) This is an open access article under the CC BY-NC-ND license (http://creativecommons.org/licenses/by-nc-nd/4.0/) Peer-review under responsibility of the scientific committee of the International Conference on Machine Learning and Peer-review under responsibility of the scientific committee of the International Conference on Machine Learning and Data Engineering. Data Engineering 10.1016/j.procs.2025.04.605 2 Uday Singh et al. / Procedia Computer Science 258 (2025) 3490–3499 Author name / Procedia Computer Science 00 (2025) 000–000 3491 suicidal attempts. SIB is a significant concern for individuals with severe ASD or intellectual disabilities, with around 30-50% of people with ASD engaging in some form of SIB during their lives [37]. Research indicates that 66% of children engage in self-injurious behavior, with 50% displaying self-head hitting and head-banging against a wall, pillow, or floor [32]. Studies have noted that SIB can vary based on a child’s clinical condition and repetitive behavioral patterns [33]. Early signs indicating the emergence of SIB in autistic children include repetitive spinning and arm flapping [34]. Detecting SIB at an early stage is crucial for caregivers and healthcare professionals working with individuals with ASD. Recognizing these behaviors is essential for timely intervention and support [7, 45]. Traditional methods of diagnosis focus on social and communicative difficulties in ASD [7]. However, recent advancements in neural network algorithms and hardware have empowered the application of artificial intelligence (AI) to automatically collect data on self-stimulatory actions. This integration of AI in the field of computer vision has led to remarkable advancements in recognition models aimed at both images [35] and videos [5]. Researchers have utilized computer vision models for the detection of SIB from video data [24, 39, 8]. Techniques such as I3D [13], 3DCNN [21], and ConvLSTM [36] have been employed, achieving varying degrees of accuracy. To address these issues, a hybrid model of CNN-LSTM is used for detecting self-injurious behaviors from video datasets. The Self-Stimulatory Behavior Dataset (SSBD) [24], enhanced with additional categories such as biting and slapping, provides a more comprehensive dataset for training. Using models like LRCN and EfficientNet, this approach aims to improve the prediction accuracy of SIB detection from video data. For children diagnosed with ASD, self-injurious behaviors are a major source of concern. Understanding and detecting these behaviors early on is crucial for ensuring the safety and well-being of individuals with ASD [37]. The study has several notable drawbacks that may impact the validity and applicability of its findings. One limitation is that participants were monitored within a specific and controlled setting, which may hinder the generalizability of the results to other environments [17]. The research also suffers from a lack of large and diverse training datasets for pediatric developmental delays, constraining the robustness of the model [13]. Additionally, the study does not address the challenge of recognizing stimming actions among numerous normal actions, an issue more akin to anomaly detection. The small size of the dataset used further restricts the generalizability of the findings [20]. Another significant issue is the unreliable extraction of robust pose information from RGB images, which can compromise the accuracy of the recognition system [21]. The dataset’s lack of ground-truth information about the condition of subjects in the videos, such as healthy versus pathological annotations, limits the ability to make accurate classifications and comparisons in behavior analysis [23]. The limited dataset size may fail to capture the full diversity of ASD cases, impacting the generalizability of the results and the detection of variations within the ASD population. Additionally, neuroimaging techniques like fMRI [44, 43]used in the study have inherent limitations, such as sensitivity to motion artifacts and low temporal resolution [23]. The objective of this paper is to use video analysis to create a predictive model that can recognize self-injurious behaviors (SIB) in people with autism spectrum disorder (ASD). Our research aims to contribute to the early detection and understanding of SIB in ASD, leveraging innovative approaches in video-based analysis. The primary objective is to create a proficient system for identifying initial indicators of SIB in autism, focusing on detecting stimming behavior, facilitating accurate diagnoses, and enabling timely intervention for effective treatment. To sum up, our contributions are as follows: • We enhanced the Self-Stimulatory Behavior Dataset (SSBD) by adding two more categories of self-injurious actions: biting and slapping. This expansion provides a more comprehensive dataset that captures a wider range of SIB behaviors, thus offering a robust foundation for training more accurate and reliable detection models. • We used augmentation techniques like Temporal Jittering and spatiotemporal transformation and added noise to the video to address limitations and expand its diversity. This augmentation enriches the dataset’s content, improving the accuracy of SIB detection in ASD and facilitating the development of more robust deep-learning models for SIB intervention. • Our Proposed framework employs various preprocessing techniques, including resizing, normalization, and data augmentation, to make an even dataset for the deep neural model. These preprocessed datasets enhance the effectiveness of deep learning models in meticulously scrutinizing video datasets, covering a spectrum of behaviors, including arm flapping, head banging, and spinning. 3492 Uday Singh et al. / Procedia Computer Science 258 (2025) 3490–3499 Author name / Procedia Computer Science 00 (2025) 000–000 3 • We conducted a thorough comparative analysis of our proposed hybrid framework against existing state-of-theart methods, such as I3D, 3DCNN, and ConvLSTM. Our analysis demonstrates the superior performance of the proposed framework in terms of accuracy and reliability, underscoring its potential for real-world application in SIB detection and intervention. The paper comprises multiple sections. Section 2 provides an overview of related work, popular SIB dataset, SIB treatment, the difference between normal and stimming behaviors, and challenges in SIB. Section 3- Method and materials explicates the dataset employed in our research and outlines the methodological details underpinning our model. Section 4 showcases the results derived from our study, emphasizing significant findings. Finally, Section 5 outlines the future research directions and conclusion. 2. Related Work In recent years, machine learning (ML) and deep learning (DL) approaches have been extensively employed to identify self-injurious behaviors (SIB) through video analysis. Various studies have made significant contributions in this field, leveraging different techniques and datasets to improve detection accuracy and reliability. Rajagopalan et al. [24] introduced the SSBD dataset, achieving a preliminary classification accuracy of 47.1% for 5-fold validation. This foundational work highlighted the challenges of video analysis in uncontrolled environments. In contrast, Cook et al. [20] developed a clinical decision support system using YouTube video clips, with deep neural networks analyzing body movements to achieve a 71% accuracy. Despite its practical implications, the moderate accuracy suggested the need for further enhancements to ensure clinical reliability. Advancements in model architecture have been notable. Zunino et al. [16] employed an LSTM network with an attentional mechanism to analyze video sequences of grasping gestures, achieving a detection accuracy of 0.82. This approach underscored the benefits of attention mechanisms in focusing on relevant video frame areas. Similarly, Washington et al. [25] utilized CNN and LSTM networks to detect headbanging behaviors in home videos, achieving a 90.77% F1-score, but faced challenges with limited training data and camera motion impacts. Incorporating traditional image descriptors with advanced neural networks, Negin et al. [21] combined optical flow descriptors with ConvLSTM and 3DCNN, achieving 56.7% accuracy in early ASD diagnosis. This study demonstrated the potential of hybrid approaches in enhancing action recognition performance. Furthermore, Liang et al. [17] introduced Temporal Coherency Deep Networks (TCDN) for feature learning from unlabeled videos, achieving 98.3% accuracy, although it lacked information on the generalizability of its findings. For aggressive action detection, Shanmughapriya et al. [26] used a 3DCNN model with Skeleton Joint features, achieving 83.56% accuracy. This study emphasized the importance of continuous monitoring to prevent self-injuries caused by aggressive behavior. On the other hand, Ali et al. [13] proposed a multi-modal fusion network, achieving over 85% accuracy on the Activis dataset, yet faced challenges with real-world applicability and generalization. Deng et al. [27] demonstrated the effectiveness of the Video Swin Transformer (VST) model, achieving accuracy between 88.50% and 97.09% on autism behavior datasets by incorporating language descriptions to enhance visual feature representation. Ribeiro et al. [23] used an I3D + MS-TCN model to classify autism-related behaviors in the ASBD dataset, achieving a 0.83 Weighted F1-score, though further exploration of potential limitations is needed. Steenfeldt-Kristensen et al. [12] utilized on-body accelerometers and machine learning for behavior segmentation and classification, achieving up to 99.6% accuracy for ADL, highlighting the effectiveness of wearable sensors combined with ML. However, challenges in detecting specific behaviors accurately remain. Cantin-Garside et al. [22] employed ML algorithms for high-accuracy classification of SIB in children with ASD, noting limitations in controlled data settings. The paper [40] proposes a novel approach combining U-Net and a Convolution Vision Transformer (CViT) for video anomaly detection and localization, enhancing the extraction of both local and global features from RGB frames. Limitations and Research Gaps Despite these advancements, several limitations persist. Many models, such as those by Rajagopalan et al. [24] and Liang et al. [17], face challenges with real-world applicability and generalization due to controlled data settings or limited training datasets. The varying accuracy levels reported, such as the 47.1% by Rajagopalan et al. [24], and 56.7% by Negin et al. [21], indicate the need for further refinement in model precision. Additionally, the lack of diverse dataset utilization and sensitivity to environmental factors, as noted in studies like Washington et al. [25], and Karuppasamy et al. [19], highlight the necessity for more robust and adaptable Uday Singh et al. / Procedia Computer Science 258 (2025) 3490–3499 Author name / Procedia Computer Science 00 (2025) 000–000 4 3493 detection systems. Hasana et al. [41],xplores convolutional spatiotemporal LSTM networks for skeleton-based human action recognition, enhancing action detection in video datasets through spatiotemporal integration. Additionally, a new approach combines ConvLSTM and LRCN models for human activity recognition, yielding competitive results, particularly with high-resolution video data. Sharma et al [42] introduces a genetic algorithm to optimize update blocks in EfficientNet, improving transfer learning efficiency by reducing training time while maintaining accuracy. Despite advances in SIB detection, limitations persist, including small and less diverse datasets hindering generalizability, moderate accuracy insufficient for clinical use, and challenges in model validation across broader populations. Environmental factors like uncontrolled settings and camera motion further complicate accurate behavior classification. Addressing these gaps is crucial for developing reliable ASD detection systems. 3. Method and Materials This section describes the proposed framework for SIB detection using the ESSBD dataset. In our study, we implemented a comprehensive approach to video classification by leveraging different deep-learning models. Figure 1 shows the proposed pipeline for SIB action detection.. Fig. 1. Proposed pipeline for SIB action detection using deep learning models 3.1. Data Collection The Self-Stimulatory Behavior Dataset (SSBD), a publicly accessible dataset [24], was primarily used to identify self-injurious actions. The information was gathered from a variety of internet resources, including Vimeo, Dailymotion, and YouTube. Out of the 75 videos reported in the original paper, only 58 could be downloaded because of YouTube’s privacy policies. The videos’ minimum resolution is 320x240 pixels, and their average runtime is ninety seconds. The dataset consists of three categories of actions: Armflapping, HeadBanging, and Spinning. 3.2. Enhanced SSBD (ESSBD) To enhance the SSBD, two additional categories—Slapping and Biting—were incorporated, expanding the dataset beyond its original categories of Arm Flapping, Head Banging, and Spinning. Data augmentation techniques were applied to each category to further increase the dataset’s diversity and robustness. These techniques included horizontal flipping, which mirrors the video along the vertical axis, vertical flipping, which mirrors the video along the horizontal axis, and the addition of Gaussian noise to simulate different recording conditions. This augmentation process increased the total number of videos to 400, significantly enriching the dataset. These enhancements improve the dataset’s utility for identifying and analyzing self-injurious behaviors, making it more comprehensive and versatile for research and practical applications. Table 1 shows the details of the SSBD and ESSBD datasets. 3.3. Frame Extraction: The action part is cut from the video to make even dataset. 3494 Uday Singh et al. / Procedia Computer Science 258 (2025) 3490–3499 Author name / Procedia Computer Science 00 (2025) 000–000 Table 1. Details of SSBD and ESSBD Dataset Action Arm Flapping Head Banging Spinning Slapping Biting 5 ESSBD 130 Video 90 Video 142 Video 12 Video 26 Video SSBD 25 Video 25 Video 25 Video – – 3.4. Action Selection and Frame Extraction: We first isolate the specific action segments from the video to focus on Self-Injurious Behaviors (SIB). Once the action portions are identified, we proceed to extract individual frames from these segments, ensuring that we capture key moments that represent the behavior accurately. 3.5. Data Augmentation To enhance the dataset and improve model robustness, we applied three augmentation techniques. These techniques help in increasing the variability of the training data, which in turn aids in making the model more generalizable and resilient to different types of input data. The Figure 2 dipict the snapshot of augmentation picture. Specific augmentation methods employed are: In the Data Augmentation phase, we employ various techniques, including horizontal flipping, vertical flipping, upsampling, down-sampling, and adding noise to the videos, to expand the dataset. First, we apply horizontal flipping to capture the mirror effect, ensuring that our model is robust to different orientations of the actions. Next, we consider temporal augmentations to further diversify the data. To increase the number of frames within each video, we upsample every video by a factor of 1.5, enhancing temporal resolution. Finally, we apply down-sampling by a factor of 0.5 to remove some information from the video clips, creating a more varied dataset while maintaining essential features for action detection. Fig. 2. Snapshot of action 6 Uday Singh et al. / Procedia Computer Science 258 (2025) 3490–3499 Author name / Procedia Computer Science 00 (2025) 000–000 3495 3.6. Frame Sequence Length and Frame Resizing: To ensure consistent input size for the deep learning model, all video frames in the dataset are resized to a specified height and width. Preprocessing steps include resizing frames to a uniform dimension of 224x224 pixels and normalizing pixel values to a range of [0, 1], which is essential for training deep learning models. 3.7. Used Models for training and testing In this section, we utilized a combination of ConvLSTM, EfficientNet, and Long-Short Term Recurrent Convolutional Networks (LRCN) models. These models were selected for their ability to effectively capture spatial and temporal features, enhancing the accuracy of Self-Injurious Behavior (SIB) detection in videos. 3.7.1. ConvLSTM Approach The ConvLSTM model for video action recognition is constructed sequentially. It commences with the initialization of a Sequential model, followed by the addition of the first ConvLSTM layer. This initial layer, configured with four filters and a (3, 3) kernel, captures spatiotemporal features from video sequences. Subsequently, 3D MaxPooling is applied to reduce spatial dimensions, and TimeDistributed Dropout is introduced to prevent overfitting. The second ConvLSTM layer, consisting of eight filters, refines feature extraction, while additional MaxPooling and Dropout layers continue the process of downsampling and regularization. The architecture further deepens by including the third and fourth ConvLSTM layers, each employing 14 and 16 filters, respectively. MaxPooling following the fourth layer further reduces dimensions before a flattened layer prepares the output for fully connected layers. The subsequent Dense layer, utilizing softmax activation, classifies the video sequence into distinct actions. The model’s architecture is summarized, including layer types, output shapes, and parameter details. The function concludes by returning the fully constructed ConvLSTM model. This systematic sequence of layers enables the model to learn intricate spatiotemporal representations effectively for accurate video action recognition. 3.7.2. Long-term recurrent convolutional network (LRCN) The Long-term Recurrent Convolutional Network (LRCN) is an advanced neural network architecture designed to effectively handle spatiotemporal data, particularly in tasks such as video analysis and action recognition. This model combines the strengths of convolutional neural networks (CNNs) for spatial feature extraction and recurrent neural networks (RNNs) for capturing temporal dependencies within sequential data. In the context of SIB detection in children with Autism Spectrum Disorder (ASD), LRCN offers several advantages. By leveraging both spatial and temporal information, LRCN can effectively analyze video data and discern subtle behavioral patterns indicative of SIB occurrences. Its ability to capture spatial and temporal dependencies simultaneously makes it a powerful tool for identifying relevant behaviors in video sequences. 3.7.3. EfficientNet EfficientNet models scale three aspects of neural networks—depth, width, and resolution—in a balanced way to maximize performance while minimizing the computational cost. EfficientNet, with its compound scaling approach, is a highly efficient and accurate model for video action detection tasks, making it ideal for detecting self-injurious behaviors in individuals with ASD using the ESSBD dataset. Its ability to handle high-resolution data, combined with its lightweight nature, makes it suitable for real-time action detection in resource-constrained environments. Additionally, its superior feature extraction capabilities enable it to detect subtle and complex behaviors indicative of self-harm. Self-injurious behaviors can vary in intensity and duration. EfficientNet’s ability to process high-resolution videos efficiently ensures that the model can scale to different kinds of video data (from low-resolution home surveillance videos to high-resolution clinical recordings) found in ESSBD. 3496 Uday Singh et al. / Procedia Computer Science 258 (2025) 3490–3499 Author name / Procedia Computer Science 00 (2025) 000–000 7 4. Result 4.1. Experimental Setup This section deals with experimental results on learning models using the ESSBD datasets. We also conducted a comparative analysis of our method and other approaches for SIB behavior. In our experiments with the ESSBD dataset, we extracted features from the video frames and applied normalization to balance the dataset. For training, we adhered to standard data augmentation techniques, including resizing frames to ensure the shortest side is 256 pixels, followed by random cropping of a sub-region resized to 224 × 224 pixels. A frame-flipping technique was applied with a probability of 0.5 during training, and frames were normalized to the [0, 1] range for consistency. After preparing the dataset, we split it into training and testing sets with an 80:20 ratio, adhering to a common data split strategy. The model was optimized using the Adam optimizer, with an initial learning rate of 0.001 and a weight decay of 0.05 to prevent overfitting. The total number of parameters in the model was 73,027, all of which were trainable. The model’s performance was compared against a baseline model using stratified cross-validation, ensuring that each fold maintained the proportion of action classes. This approach ensured a reliable evaluation of the model’s performance across all action types. The training methodology followed in this work is based on the standard procedures employed in previous works such as [38] and [39], forming a strong foundation for the comparative analysis. We employed stratified cross-validation, ensuring that each fold maintained the proportion of action classes, thereby providing a more reliable evaluation of the model’s performance across all action types. Table 2 shows the parameter of model. Table 2. Experimental Results and Pseudo Code for ESSBD Dataset Parameter Training/Testing Data Split Model Comparison Optimizer Learning Rate Weight Decay Total Parameters Trainable Parameters Non-trainable Parameters Details 80:20 Compared the performance of the proposed model with a baseline model Adam Optimizer Initial learning rate of 0.001 0.05 73,027 73,027 0 Pseudo Code for Action Detection • Extract features and normalize each frame in ESSBD dataset. • Apply data augmentation techniques to each normalized frame. • Split the dataset into training and testing sets (80:20). • Initialize the model for training. • For each epoch, iterate through batches of training data. • Perform a forward pass to get predictions for each batch. • Compute the loss based on the predictions and actual labels. • Update the model parameters using the optimizer. • For each fold in cross-validation, train the model on the training data. • Evaluate the model performance on the test data. 4.2. Result Discussion The models were evaluated using two key metrics: Accuracy, which reflects the percentage of correct predictions, and F1 Score, which balances precision and recall—particularly valuable for handling imbalanced datasets. EfficientNet demonstrates a moderate performance with an accuracy of 77.17% and an F1 score of 71.90. ConvoLSTM Uday Singh et al. / Procedia Computer Science 258 (2025) 3490–3499 Author name / Procedia Computer Science 00 (2025) 000–000 8 3497 Table 3. Accuracy and F1 Score of models with ESSBD dataset Model EfficientNet ConvoLSTM LRCN Accuracy 77.17 80.33 92.6 F1 Score 71.90 81.67 83.59 performs better, achieving 80.33% accuracy and an improved F1 score of 81.67. However, LRCN delivers the best results, with the highest accuracy of 92.6% and an F1 score of 83.59. Among the models, LRCN stands out as the most effective for the ESSBD dataset, excelling in both accuracy and F1 score. ConvoLSTM also performs well, particularly in balancing precision and recall. While EfficientNet shows reasonable capability, it ranks the lowest in both metrics. Overall, the table highlights LRCN as the most suitable model for this task. Table 3 shows the accuracy and F1 score of models. Fig. 3. Accuracy of ConvLSTM models Fig. 5. Loss of ConvoLSTM models Fig. 4. Accuracy of LRCN models Fig. 6. Loss of LRCN models Fig. 7. EfficientNet Model Accuracy In Fig-5, the ConvLSTM model shows a low training loss of 0.00737 (blue line), indicating a good fit, but a higher validation loss of 0.44 (red line), suggesting possible overfitting. Similarly, in Fig-6, the LRCN model has a training loss of 0.007 (blue line), but the validation loss of 0.38 (red line) indicates weaker generalization to unseen data. Both models show a gap between training and validation performance, pointing to potential overfitting. This suggests the models may be memorizing the training data rather than generalizing well. Figure 7 shows the accuracy of EfficientNet. 4.3. Comparison with the State of the Art The proposed framework result is compared with different state of art method . The LRCN model achieves the best performance with the highest accuracy (92.6%) and a strong F1 score (83.59), followed by the Two Stream CNN with Uday Singh et al. / Procedia Computer Science 258 (2025) 3490–3499 Author name / Procedia Computer Science 00 (2025) 000–000 3498 9 Table 4. Result comparison with state of the art method Model 3DCNN I3D(RGB+FLOW) [13] I3D(RGB) [13] Two Stream CNN EfficientNet ConvoLSTM LRCN Accuracy 42 75.32 76.92 87.76 77.17 80.33 92.6 F1 Score 69 60 85.34 71.90 81.67 83.59 an accuracy of 87.76% and the highest F1 score (85.34). In contrast, the 3DCNN shows the lowest performance with an accuracy of 42%, without an F1 score reported. Table 4 compares outcomes comprehensively against state-of-theart methodologies, focusing on accuracy and F1 score metrics. 5. Conclusion and Future Works In this paper, we enhanced the SSBD dataset to ESSBD dataset to add two more action vedio and successfully used the ConvLSTM, LRCN, and EfficientNet for detecting stereotypic behaviors (SIB) in autistic children. The LRCN model outperforms the others, achieving an accuracy of 92.62 Future research directions should prioritize augmenting datasets with diverse and contemporary video samples that encompass interactions involving children with ASD and healthcare providers. This expanded dataset could enhance model generalization and robustness, thereby improving real-world applicability. Furthermore, refining loss functions and optimization techniques tailored to the nuances of video data could further elevate model performance. By systematically exploring these avenues, researchers aim to advance the accuracy, reliability, and practical utility of video classification models in clinical and educational settings. References [1] Li, Jing, Yihao Zhong, and Gaoxiang Ouyang. ”Identification of ASD children based on video data.” 2018 24th International conference on pattern recognition (ICPR). IEEE, 2018. [2] Zunino, Andrea, et al. ”Video gesture analysis for autism spectrum disorder detection.” 2018 24th international conference on pattern recognition (ICPR). IEEE, 2018. [3] O’Roak, Brian J., and Matthew W. State. ”Autism genetics: strategies, challenges, and opportunities.” Autism Research 1.1 (2008): 4-17. [4] Elsabbagh, Mayada, et al. ”Global prevalence of autism and other pervasive developmental disorders.” Autism research 5.3 (2012): 160-179. [5] Li, Jing, et al. ”Classifying ASD children with LSTM based on raw videos.” Neurocomputing 390 (2020): 226-238. [6] Li, Beibin, et al. ”A facial affect analysis system for autism spectrum disorder.” 2019 IEEE International Conference on Image Processing (ICIP). IEEE, 2019. [7] Bodfish, James W., et al. ”Varieties of repetitive behavior in autism: Comparisons to mental retardation.” Journal of autism and developmental disorders 30 (2000): 237-243. [8] Chen, Shi, and Qi Zhao. ”Attention-based autism spectrum disorder screening with privileged modality.” Proceedings of the IEEE/CVF International Conference on Computer Vision. 2019. [9] Bryson, Susan E., et al. ”The Autism Observation Scale for Infants: scale development and reliability data.” Journal of autism and developmental disorders 38 (2008): 731-738. [10] Graves, Alex, and Alex Graves. ”Long short-term memory.” Supervised sequence labelling with recurrent neural networks (2012): 37-45. [11] Wan, Guobin, et al. ”Applying eye tracking to identify autism spectrum disorder in children.” Journal of autism and developmental disorders 49 (2019): 209-215. [12] Plötz, Thomas, et al. ”Automatic assessment of problem behavior in individuals with developmental disabilities.” Proceedings of the 2012 ACM conference on ubiquitous computing. 2012. [13] Ali, Abid, et al. ”Video-based behavior understanding of children for objective diagnosis of autism.” VISAPP 2022-17th International Conference on Computer Vision Theory and Applications. 2022. [14] Stewart, Gavin R., et al. ”Self-harm and suicidality experiences of middle-age and older adults with vs. without high autistic traits.” Journal of autism and developmental disorders 53.8 (2023): 3034-3046. [15] Cantin-Garside, Kristine D., et al. ”Understanding the experiences of self-injurious behavior in autism spectrum disorder: Implications for monitoring technology design.” Journal of the American Medical Informatics Association 28.2 (2021): 303-310. 10 Uday Singh et al. / Procedia Computer Science 258 (2025) 3490–3499 Author name / Procedia Computer Science 00 (2025) 000–000 3499 [16] Zunino, Andrea, et al. ”Video gesture analysis for autism spectrum disorder detection.” 2018 24th international conference on pattern recognition (ICPR). IEEE, 2018. [17] Liang, Shuaibing, et al. ”Autism spectrum self-stimulatory behaviors classification using explainable temporal coherency deep features and svm classifier.” IEEE Access 9 (2021): 34264-34275. [18] Steenfeldt-Kristensen, Catherine, Chris A. Jones, and Caroline Richards. ”The prevalence of self-injurious behaviour in autism: a meta-analytic study.” Journal of autism and developmental disorders 50 (2020): 3857-3873. [19] Karuppasamy, Sankar Ganesh, et al. ”Prediction of autism spectrum disorder using convolution neural network.” 2022 6th International Conference on Trends in Electronics and Informatics (ICOEI). IEEE, 2022. [20] Cook, Andrew, et al. ”Towards automatic screening of typical and atypical behaviors in children with autism.” 2019 IEEE International Conference on Data Science and Advanced Analytics (DSAA). IEEE, 2019. [21] Negin, Farhood, et al. ”Vision-assisted recognition of stereotype behaviors for early diagnosis of autism spectrum disorders.” Neurocomputing 446 (2021): 145-155. [22] Cantin-Garside, Kristine D., et al. ”Detecting and classifying self-injurious behavior in autism spectrum disorder using machine learning techniques.” Journal of autism and developmental disorders 50 (2020): 4039-4052. [23] Ribeiro, Guilherme Ocker, Mateus Grellert, and Jonata Tyska Carvalho. ”Stimming behavior dataset-unifying stereotype behavior dataset in the wild.” 2023 IEEE 36th International Symposium on Computer-Based Medical Systems (CBMS). IEEE, 2023. [24] Rajagopalan, Shyam, Abhinav Dhall, and Roland Goecke. ”Self-stimulatory behaviours in the wild for autism diagnosis.” Proceedings of the IEEE International Conference on Computer Vision Workshops. 2013. [25] Washington, Peter, et al. ”Activity recognition with moving cameras and few training examples: applications for detection of autism-related headbanging.” Extended abstracts of the 2021 CHI conference on human factors in computing systems. 2021. [26] Shanmughapriya, M., et al. ”AGGRESSIVE ACTION IDENTIFICATION IN AUTISM SPECTRUM DISORDER USING VIDEO ANALYSIS.” EPRA Int. J. Res. Dev. 7 (2022): 19-27. [27] Deng, Andong, et al. ”Problem behaviors recognition in videos using language-assisted deep learning model for children with autism.” arXiv preprint arXiv:2211.09310 (2022). [28] Kumar, N. Sunil, et al. ”Restricted and Repetitive Behaviors and Interests in Young Children with Autism: A Comparative Study.” Indian Journal of Pediatrics 89.12 (2022): 1216-1221. [29] Rutter, Michael. ”Concepts of autism: a review of research.” Child Psychology Psychiatry Allied Disciplines (1968) [30] Huisman, Sylvia, et al. ”Self-injurious behavior.” Neuroscience Biobehavioral Reviews 84 (2018): 483-491. [31] Minshawi, Noha F., et al. ”The association between self-injurious behaviors and autism spectrum disorders.” Psychology research and behavior management (2014): 125-136. [32] Berkson, Gershon, and Megan Tupa. ”Early development of stereotyped and self-injurious behaviors.” Journal of Early Intervention 23.1 (2000): 1-19. [33] Richman, David M., and Steven E. Lindauer. ”Longitudinal assessment of stereotypic, proto-injurious, and self-injurious behavior exhibited by young children with developmental delays.” American Journal on Mental Retardation 110.6 (2005): 439-450. [34] Fodstad, Jill C., Johannes Rojahn, and Johnny L. Matson. ”The emergence of challenging behaviors in at-risk toddlers with and without autism spectrum disorder: A cross-sectional study.” Journal of Developmental and Physical Disabilities 24 (2012): 217-234. [35] Krizhevsky, Alex, Ilya Sutskever, and Geoffrey E. Hinton. ”Imagenet classification with deep convolutional neural networks.” Advances in neural information processing systems 25 (2012). [36] Pandey, Prashant, et al. ”Guided weak supervision for action recognition with scarce data to assess skills of children with autism.” Proceedings of the AAAI Conference on Artificial Intelligence. Vol. 34. No. 01. 2020. [37] Minshawi, Noha F., et al. ”The association between self-injurious behaviors and autism spectrum disorders.” Psychology research and behavior management (2014): 125-136. [38] Duta, Ionut Cosmin, et al. ”Improved residual networks for image and video recognition.” 2020 25th International Conference on Pattern Recognition (ICPR). IEEE, 2021. [39] Rajagopalan, Shyam Sundar, and Roland Goecke. ”Detecting self-stimulatory behaviours for autism diagnosis.” 2014 IEEE International Conference on Image Processing (ICIP). IEEE, 2014. [40] Roka, Sanjay, and Manoj Diwakar. ”Cvit: a convolution vision transformer for video abnormal behavior detection and localization.” SN Computer Science 4.6 (2023): 829. [41] Hasana, M. M., Ibrahim, M., Ali, M. S. (2023). Speeding Up EfficientNet: Selecting Update Blocks of Convolutional Neural Networks using Genetic Algorithm in Transfer Learning. arXiv preprint arXiv:2303.00261. [42] ConvST-LSTM-Net: convolutional spatiotemporal LSTM networks for skeleton-based human action recognition [43] Singh, Uday, Shailendra Shukla, and Manoj Madhava Gore. ”Functional Connectivity and Graph Embedding-Based Domain Adaptation for Autism Classification from Multi-site Data.” Arabian Journal for Science and Engineering (2024): 1-20. [44] Singh, Uday, Shailendra Shukla, and Manoj Madhava Gore. ”Detection of autism spectrum disorder using multi-scale enhanced graph convolutional network.” Cognitive Computation and Systems (2024). [45] Singh, Uday, Shailendra Shukla, and Manoj Madhava Gore. ”An Improved Feature Selection Algorithm for Autism Detection.” 2022 IEEE 9th Uttar Pradesh Section International Conference on Electrical, Electronics and Computer Engineering (UPCON). IEEE, 2022.
0
You can add this document to your study collection(s)
Sign in Available only to authorized usersYou can add this document to your saved list
Sign in Available only to authorized users(For complaints, use another form )