No. 14 (2026): International journal of media and communications in Central Asia
Статьи

ARTICULATORY AND VISUAL REPRESENTATIONS OF SPEECH PHONEMES: SYNCHRONIZATION ISSUES IN ANIMATION

Gulshan Kayumova
University of Journalism and Mass Communications of Uzbekistan
Munojat Sultonova
Tashkent University of Information Technologies

Published 2026-10-04

Keywords

  • phoneme,
  • articulation,
  • consonant sound,
  • vowel sound,
  • plosive sound,
  • fricative sound,
  • affricate,
  • lateral sound,
  • nasal sound,
  • speech animation,
  • visual speech,
  • synchronization
  • ...More
    Less

How to Cite

Kayumova, G., & Sultonova , M. (2026). ARTICULATORY AND VISUAL REPRESENTATIONS OF SPEECH PHONEMES: SYNCHRONIZATION ISSUES IN ANIMATION. INTERNATIONAL SCIENTIFIC JOURNAL OF MEDIA AND COMMUNICATIONS IN CENTRAL ASIA, (14). https://doi.org/10.62499/ijmcc.vi14.316

Abstract

This article analyzes the articulatory, acoustic, and visual characteristics of speech phonemes, as well as issues related to their representation in animation. Consonant and vowel phonemes, plosive, fricative, affricate, lateral, trill, and nasal sounds are classified. The process of synchronizing a character’s lip and facial muscle movements with speech, preparing visual representations of phonemes, and placing them in animation frames are discussed, along with their pedagogical and technological significance. The findings of the study can be applied to speech animation, interactive educational tools, and artificial intelligence-based systems

 

 

 

 

References

  1. Avriel, M. (2003). Nonlinear Programming: Analysis and Methods. M. Avriel. Courier Corporation. 512 p. ISBN: 978-0-486-43227-4. Retrieved September 06, 2026 from https://books.google.co.uz/books?id=byF4Xb1QbvMC&printsec=frontcover&hl=ru#v=onepage&q&f=false
  2. Bressem, Jana & Ladewig, Silva. (2011). Rethinking gesture phases: Articulatory features of gestural movement?. Semiotica. 184. 53–91. DOI: 10.1515/semi.2011.022. Retrieved September 06, 2026 from https://philpapers.org/rec/BRERGP?utm_source
  3. Cao, C., et al. (2013). Facewarehouse: A 3D facial expression database for visual computing. IEEE Transactions on Visualization and Computer Graphics, 20(3), 413–425. DOI:10.1109/TVCG.2013.249. https://pubmed.ncbi.nlm.nih.gov/24434222/
  4. Beknazarova S., Sadullaeva S., Bazhenov R., Qayumova G., Jaumitbayeva M. Application of nonlinear splitting algorithm to the method of reference equations. AIP Conf. Proc. 16 June 2022; 2432 (1): 060003. https://doi.org/10.1063/5.0089494
  5. Beknazarova S., Yunusova D., Qayumova G. et al. Adaptive video compression and transmission algorithms for smart surveillance in IOT networks, Proc. SPIE 14014, Advanced Materials for Optics and Photonics: Chemistry and Engineering Perspectives (AMOP 2025), 1401403 (18 Dec 2025); https://doi.org/10.1117/12.3091984
  6. Cudeiro, D. et al. (2019). Capture, Learning, and Synthesis of 3D Speaking Styles. Proceedings IEEE Conf. on Computer Vision and Pattern Recognition (CVPR). - С. 10101-10111. Retrieved September 06, 2026 from http://voca.is.tue.mpg.de/.
  7. Graves, A., et al. (2006). Connectionist temporal classification: Labelling unsegmented sequence data with recurrent neural networks. In Proceedings of the 23rd International Conference on Machine Learning (pp. 369–376). DOI: 10.1145/1143844.1143891. Retrieved September 06, 2026 from https://mlanthology.org/icml/2006/graves2006icml-connectionist/?utm_source
  8. Kendon, A. (2004). Gesture: Visible Action as Utterance. Cambridge University Press. Retrieved September 06, 2026 from 10.1017/CBO9780511807572.
  9. Nguyen, N. (2000). Perceiving Talking Faces: From Speech Perception to a Behavioral Principle by Massaro, D. W. Journal of Phonetics, 28(1), 103–109. DOI: 10.1006/jpho.2000.0108.
  10. Pelachaud, C. (2009). Studies on gesture expressivity for a virtual agent. Speech Communication, 51(7), 630–639. https://doi.org/10.1016/j.specom.2008.04.009
  11. Pham, H. X., Pavlovic, V., Cai, J., & Cham, T. (2016). Robust real-time performance-driven 3D face tracking. In L. Davis, A. Del Bimbo, & B. C. Lovell (Eds.), 2016 23rd International Conference on Pattern Recognition (ICPR 2016) (pp. 1851-1856). Article 7899906 IEEE, Institute of Electrical and Electronics Engineers. https://doi.org/10.1109/ICPR.2016.7899906
  12. Rublee, E., Rabaud, V., Konolige, K., Bradski, G.. (2011) ORB: An efficient alternative to SIFT or SUFT. ICCV ‘11 Proceedings of the 2011 International Conference on Computer Vision. – P. 2564–2571. DOI: 10.1109/ICCV.2011.6126544. https://ieeexplore.ieee.org/document/6126544?utm_source
  13. Taylor, S., Kim, T., Yue, Y., Mahler, M., Krahe, J., Rodriguez, A., ... & Matthews, I. (2020). A deep learning approach for generalized speech animation. ACM Transactions on Graphics, 39(4), 93:1–93:15. https://doi.org/10.1145/3386569.3392450
  14. Tucker, L. R. (1966). Some mathematical notes on three-mode factor analysis. Psychometrika, 31(3), 279–311. DOI: https://doi.org/10.1007/BF02289464
  15. Wagner, Petra & Malisz, Zofia & Kopp, Stefan. (2014). Gesture and speech in interaction: An overview. Speech Communication. 57. 209-232. Retrieved September 06, 2026 from 10.1016/j.specom.2013.09.008.
  16. Kayumova, G. & Boymurodov, B. (2025). Uch o‘lchamli personajlarning yuz holatini modellashtirish va animatsiyalashning zamonaviy yondashuvlari. Al-Farg’oniy avlodlari, 1 (3), 151-158. doi: 10.5281/zenodo.17295640
  17. Корзун В.А. (2022). Генерация мимики для виртуальных ассистентов. Труды Московского физико-технического института, 14(3/55), 57–62. URL: https://sciup.org/generacija-mimiki-dlja-virtualnyh-assistentov-142236478?utm_source