ARTICULATORY AND VISUAL REPRESENTATIONS OF SPEECH PHONEMES: SYNCHRONIZATION ISSUES IN ANIMATION
Published 2026-10-04
Keywords
- phoneme,
- articulation,
- consonant sound,
- vowel sound,
- plosive sound
- fricative sound,
- affricate,
- lateral sound,
- nasal sound,
- speech animation,
- visual speech,
- synchronization ...More
How to Cite
Abstract
This article analyzes the articulatory, acoustic, and visual characteristics of speech phonemes, as well as issues related to their representation in animation. Consonant and vowel phonemes, plosive, fricative, affricate, lateral, trill, and nasal sounds are classified. The process of synchronizing a character’s lip and facial muscle movements with speech, preparing visual representations of phonemes, and placing them in animation frames are discussed, along with their pedagogical and technological significance. The findings of the study can be applied to speech animation, interactive educational tools, and artificial intelligence-based systems
References
- Avriel, M. (2003). Nonlinear Programming: Analysis and Methods. M. Avriel. Courier Corporation. 512 p. ISBN: 978-0-486-43227-4. Retrieved September 06, 2026 from https://books.google.co.uz/books?id=byF4Xb1QbvMC&printsec=frontcover&hl=ru#v=onepage&q&f=false
- Bressem, Jana & Ladewig, Silva. (2011). Rethinking gesture phases: Articulatory features of gestural movement?. Semiotica. 184. 53–91. DOI: 10.1515/semi.2011.022. Retrieved September 06, 2026 from https://philpapers.org/rec/BRERGP?utm_source
- Cao, C., et al. (2013). Facewarehouse: A 3D facial expression database for visual computing. IEEE Transactions on Visualization and Computer Graphics, 20(3), 413–425. DOI:10.1109/TVCG.2013.249. https://pubmed.ncbi.nlm.nih.gov/24434222/
- Beknazarova S., Sadullaeva S., Bazhenov R., Qayumova G., Jaumitbayeva M. Application of nonlinear splitting algorithm to the method of reference equations. AIP Conf. Proc. 16 June 2022; 2432 (1): 060003. https://doi.org/10.1063/5.0089494
- Beknazarova S., Yunusova D., Qayumova G. et al. Adaptive video compression and transmission algorithms for smart surveillance in IOT networks, Proc. SPIE 14014, Advanced Materials for Optics and Photonics: Chemistry and Engineering Perspectives (AMOP 2025), 1401403 (18 Dec 2025); https://doi.org/10.1117/12.3091984
- Cudeiro, D. et al. (2019). Capture, Learning, and Synthesis of 3D Speaking Styles. Proceedings IEEE Conf. on Computer Vision and Pattern Recognition (CVPR). - С. 10101-10111. Retrieved September 06, 2026 from http://voca.is.tue.mpg.de/.
- Graves, A., et al. (2006). Connectionist temporal classification: Labelling unsegmented sequence data with recurrent neural networks. In Proceedings of the 23rd International Conference on Machine Learning (pp. 369–376). DOI: 10.1145/1143844.1143891. Retrieved September 06, 2026 from https://mlanthology.org/icml/2006/graves2006icml-connectionist/?utm_source
- Kendon, A. (2004). Gesture: Visible Action as Utterance. Cambridge University Press. Retrieved September 06, 2026 from 10.1017/CBO9780511807572.
- Nguyen, N. (2000). Perceiving Talking Faces: From Speech Perception to a Behavioral Principle by Massaro, D. W. Journal of Phonetics, 28(1), 103–109. DOI: 10.1006/jpho.2000.0108.
- Pelachaud, C. (2009). Studies on gesture expressivity for a virtual agent. Speech Communication, 51(7), 630–639. https://doi.org/10.1016/j.specom.2008.04.009
- Pham, H. X., Pavlovic, V., Cai, J., & Cham, T. (2016). Robust real-time performance-driven 3D face tracking. In L. Davis, A. Del Bimbo, & B. C. Lovell (Eds.), 2016 23rd International Conference on Pattern Recognition (ICPR 2016) (pp. 1851-1856). Article 7899906 IEEE, Institute of Electrical and Electronics Engineers. https://doi.org/10.1109/ICPR.2016.7899906
- Rublee, E., Rabaud, V., Konolige, K., Bradski, G.. (2011) ORB: An efficient alternative to SIFT or SUFT. ICCV ‘11 Proceedings of the 2011 International Conference on Computer Vision. – P. 2564–2571. DOI: 10.1109/ICCV.2011.6126544. https://ieeexplore.ieee.org/document/6126544?utm_source
- Taylor, S., Kim, T., Yue, Y., Mahler, M., Krahe, J., Rodriguez, A., ... & Matthews, I. (2020). A deep learning approach for generalized speech animation. ACM Transactions on Graphics, 39(4), 93:1–93:15. https://doi.org/10.1145/3386569.3392450
- Tucker, L. R. (1966). Some mathematical notes on three-mode factor analysis. Psychometrika, 31(3), 279–311. DOI: https://doi.org/10.1007/BF02289464
- Wagner, Petra & Malisz, Zofia & Kopp, Stefan. (2014). Gesture and speech in interaction: An overview. Speech Communication. 57. 209-232. Retrieved September 06, 2026 from 10.1016/j.specom.2013.09.008.
- Kayumova, G. & Boymurodov, B. (2025). Uch o‘lchamli personajlarning yuz holatini modellashtirish va animatsiyalashning zamonaviy yondashuvlari. Al-Farg’oniy avlodlari, 1 (3), 151-158. doi: 10.5281/zenodo.17295640
- Корзун В.А. (2022). Генерация мимики для виртуальных ассистентов. Труды Московского физико-технического института, 14(3/55), 57–62. URL: https://sciup.org/generacija-mimiki-dlja-virtualnyh-assistentov-142236478?utm_source
