Dual-Mode Feature Fusion for Dynamic Gesture Recognition Using GRU-LSTM in Complex Backgrounds
Abstract
To address the problems of poor gesture segmentation, incomplete feature extraction, and low recognition accuracy caused by cluttered backgrounds, light intensity fluctuations, and skin-like background interference in dynamic gesture recognition, a dual-mode feature fusion dynamic gesture recognition method based on GRU-LSTM is proposed for complex backgrounds. First, adaptive threshold segmentation and motion gesture segmentation are combined to realize fine gesture segmentation, with the Intersection over Union (IoU) and Dice coefficient reaching 92.1% and 93.5%, respectively. For feature extraction, a dual-mode fusion strategy is adopted: the MediaPipe model extracts 21 hand joint local features (processed by relative coordinate transformation and mean-variance normalization), and an improved CNN embedded with Inception and CBAM attention modules extracts global image features. A GRU-LSTM combined network is designed for sequence learning, which reduces the computational complexity by 28% compared with a single LSTM and improves the fitting accuracy by 2.7% compared with a single GRU. The Softmax-SVM combined algorithm and a three-layer MLP are used for dual-mode feature combined classification. A dedicated Seven-Step Handwashing Technique (SWT) dataset is constructed in collaboration with Hubei Provincial People's Hospital, and the proposed method achieves an average recognition accuracy of 97.73% on this dataset (with Precision 97.21%, Recall 96.89%, and F1-score 97.05%). Comparative experiments show that the method outperforms state-of-the-art (SOTA) methods (e.g., 5.2% higher accuracy than LSTM and 9.93% higher than HOG+SVM), and the inference speed reaches 28 frames per second (fps) meeting real-time recognition requirements. Experimental results confirm that the method can accurately segment and recognize dynamic gestures in complex backgrounds, and has strong robustness to light and background changes.References
[1]A. A. Barbhuiya, R. K. Karsh, and R. Jain. A convolutional neural network and classical moments-based feature fusion model for gesture recognition. Multimed. Syst., 28(5):1779–1792, 2022. https://doi.org/10.1007/s00530-022-00951-5
[2]A. A. Ilham, and I. Nurtanio. Applying LSTM and GRU methods to recognize and interpret hand gestures, poses, and face-based sign language in real time. J. Adv. Comput. Intell. Intell. Inform., 28(2):265–272, 2024. https://doi.org/10.20965/jaciii.2024.p0265
[3]A. Nimbekar, Y. S. Dinesh, A. Gautam, V. Hunsigida, A. R. Nali, and A. Acharyya. Reconfigurable VLSI design architecture for deep learning established forelimb and hindlimb gesture recognition for rehabilitation application. IEEE Access, 11:70061–70070, 2023. https://doi.org/10.1109/access.2023.3293422
[4]F. Zhang. Human–Computer Interactive Gesture Feature Capture and Recognition in Virtual Reality. Ergonomics Des., 29(2):19–25, 2021. https://doi.org/10.1177/1064804620924133
[5]H. Haroon, S. Altaf, Z. U. Rehman, M. W. Soomro, and S. Iqbal. Human hand gesture identification framework using SIFT and knowledge‐level technique. ETRI J., 45(6):1022–1034, 2023. https://doi.org/10.4218/etrij.2022-0281
[6]H. Mahmud, M. M. Morshed, and M. K. Hasan. Quantized depth image and skeleton-based multimodal dynamic hand gesture recognition. Vis. Comput., 40(1):11–25, 2024. https://doi.org/10.1007/s00371-022-02762-1
[7]H. Zhang, H. Qu, L. Teng, and C. Y. Tang. LSTM-MSA: A novel deep learning model with dual-stage attention mechanisms forearm EMG-based hand gesture recognition. IEEE Trans. Neural Syst. Rehabil. Eng., 31:4749–4759, 2023. https://doi.org/10.1109/tnsre.2023.3336865
[8]J. Ni, Y. Wang, G. Tang, W. Cao, and S. X. Yang. A lightweight GRU-based gesture recognition model for skeleton dynamic graphs. Multimed. Tools Appl., 83(27):70545–70570, 2024. https://doi.org/10.1007/s11042-024-18313-w
[9]J. Zhang, and X. Zeng. Multi-touch gesture recognition of Braille input based on Petri Net and RBF Net. Multimed. Tools Appl., 81(14):19395–19413, 2022. https://doi.org/10.1007/s11042-021-11156-9
[10]M. U. Rehman, F. Ahmed, M. A. Khan, U. Tariq, F. A. Alfouzan, N. M. Alzahrani, and J. Ahmad. Dynamic hand gesture recognition using 3D-CNN and LSTM networks. Comput. Mater. Contin., 70(3):4675–4690, 2022. https://doi.org/10.32604/cmc.2022.019586
[11]S. M. S. Shah, J. I. Khan, S. H. Abbas, and A. Ghani. Symmetric mean binary pattern-based Pakistan sign language recognition using multiclass support vector machines. Neural Comput. Appl., 35(1):949–972, 2023. https://doi.org/10.1007/s00521-022-07804-2
[12]W. Qi, S. E. Ovur, Z. Li, A. Marzullo, and R. Song. Multi-sensor guided hand gesture recognition for a teleoperated robot using a recurrent neural network. IEEE Robot. Autom. Lett., 6(3):6039–6045, 2021. https://doi.org/10.1109/lra.2021.3089999
[13]Y. Yang, and W. Bao. Application of human-computer interaction technology in remote language learning platform. Int. J. Hum.-Comput. Interact., 41(3):1751–1761, 2025. https://doi.org/10.1080/10447318.2023.2194709
[14]B. Jin, X. Ma, Z. Zhang, Z. Lian, and B. Wang. Interference-robust millimeter-wave radar-based dynamic hand gesture recognition using 2-d cnn-transformer networks. IEEE Internet Things J., 11(2):2741–2752, 2023. https://doi.org/10.1109/jiot.2023.3293092
[15]S. H. Yuanyuan, L. I. Yunan, F. U. Xiaolong, K. Miao, and Q. Miao. Review of dynamic gesture recognition. Virtual Real. Intell. Hardw., 3(3):183–206, 2021.
[16]U. Sayed, S. Bakheet, M. A. Mofaddel, and Z. El-Zohry. Robust Hand Gesture Recognition Using HOG Features and machine learning. Sohag J. Sci., 9(3):226–233, 2024. https://doi.org/10.21608/sjsci.2024.248702.1151
DOI:
https://doi.org/10.31449/inf.v50i14.13729Keywords:
Dynamic gesture recognition, deep learning, dual-mode feature fusion, GRU-LSTM, complex background, seven-step handwashing methodDownloads
Published
Issue
Section
License
Authors retain copyright in their work. By submitting to and publishing with Informatica, authors grant the publisher (Slovene Society Informatika) the non-exclusive right to publish, reproduce, and distribute the article and to identify itself as the original publisher.
All articles are published under the Creative Commons Attribution license CC BY 3.0. Under this license, others may share and adapt the work for any purpose, provided appropriate credit is given and changes (if any) are indicated.
Authors may deposit and share the submitted version, accepted manuscript, and published version, provided the original publication in Informatica is properly cited.







