Volume 8 | Issue - 8
Volume 8 | Issue - 8
Volume 8 | Issue - 8
Volume 8 | Issue - 7
Volume 8 | Issue - 7
Several studies have been led on video footage using AI deep learning methods. Human activity recognition (HAR), object localization, and event recognition are some of the most common applications. Among these, HAR stands out as a particularly difficult task and a major focus topic for future study in video data processing. Many AI-based models have been created for activity identification. Still, their poor presentation of real-world long-term HAR can be attributed to their inability to extract spatial and temporal data. To improve the system's prediction skills, the authors of this study suggest a hybrid deep learning perfect called the Long-Range Context Network (LRCN), which combines a convolutional neural network with a long-short-term memory network (CNN LSTM) and is driven by a self-attention algorithm. The transformer's built-in self attention mechanism can compete with high-end convolutional neural networks with long short-term memory by expressing unique dependencies among signal levels within a time series. The human activity recognition process is then finalized by feeding the collected features into the SoftMax layer. Experiments determine the best values for the model's parameters. To find the best option for HAR, a thorough ablation study is achieved over many classic machine learning and deep learning models. A few of the measures examined in this experimental study of the publicly accessible modified UCF50 dataset (UCF50mini). Using the LRCN method, a 99% accuracy is attained, proving the proposed model's viability for HAR use.