Deep learning architectures in physics-based character animation: a bibliometric mapping of research trajectories, 2016-2025

Abstract

Perkembangan kecerdasan buatan (artificial intelligence/AI) telah mengubah lanskap produksi animasi secara fundamental, mulai dari sintesis gerak karakter, animasi wajah, hingga simulasi visual berbasis pembelajaran mendalam. Meskipun literatur pada domain ini berkembang pesat, pemetaan struktur intelektual dan arah evolusinya secara menyeluruh masih terbatas. Penelitian ini bertujuan memetakan struktur bibliometrik penelitian AI-assisted animation melalui analisis terhadap 983 dokumen yang diperoleh dari basis data Scopus periode 2016–2025, menggunakan kombinasi analisis deskriptif kuantitatif dan visualisasi jaringan kata kunci berbasis VOSviewer. Hasil analisis menunjukkan tren publikasi yang meningkat tajam sejak 2021, dengan dominasi kontribusi dari Tiongkok, Amerika Serikat, dan Korea Selatan, serta konsentrasi topik pada deep learning, reinforcement learning, dan character animation sebagai inti kluster keilmuan. Peta densitas dan jaringan kata kunci mengindikasikan pergeseran fokus riset dari pendekatan machine learning konvensional menuju arsitektur pembelajaran mendalam yang terintegrasi dengan reinforcement learning untuk pengendalian gerak fisik karakter (physics-based character control). Studi ini memberikan kontribusi berupa peta jalan tematik yang dapat menjadi rujukan bagi peneliti, praktisi industri animasi, maupun pengembang kebijakan riset teknologi kreatif dalam merumuskan agenda penelitian AI-assisted animation ke depan.

Keywords
  • AI-assisted animation, Bibliometrik, Deep learning, Character animation,
References
  1. Abdi, L., & Meddeb, A. (2018). Driver information system: A combination of augmented reality, deep learning and vehicular ad-hoc networks. Multimedia Tools and Applications, 77(12), 14673–14703. https://doi.org/10.1007/s11042-017-5054-6
  2. Alexanderson, S., Henter, G. E., Kucherenko, T., & Beskow, J. (2020). Style-controllable speech-driven gesture synthesis using normalising flows. Computer Graphics Forum, 39(2), 487–496. https://doi.org/10.1111/cgf.13946
  3. Bailey, S. W., Omens, D., Dilorenzo, P., & O'Brien, J. F. (2020). Fast and deep facial deformations. ACM Transactions on Graphics, 39(4). https://doi.org/10.1145/3386569.3392397
  4. Bergamin, K., Clavet, S., Holden, D., & Richard Forbes, J. (2019). DReCon: Data-driven responsive control of physics-based characters. ACM Transactions on Graphics, 38(6). https://doi.org/10.1145/3355089.3356536
  5. Bhavana, S., & Vijayalakshmi, V. (2022). AI-based metaverse technologies advancement impact on higher education learners. WSEAS Transactions on Systems, 21, 178–184. https://doi.org/10.37394/23202.2022.21.19
  6. Cao, Q., Zhang, W., & Zhu, Y. (2021). Deep learning-based classification of the polar emotions of 'moe'-style cartoon pictures. Tsinghua Science and Technology, 26(3), 275–286. https://doi.org/10.26599/TST.2019.9010035
  7. Choi, J.-Y., Ha, E., Son, M., Jeon, J.-H., & Kim, J.-W. (2024). Human joint angle estimation using deep learning-based three-dimensional human pose estimation for application in a real environment. Sensors, 24(12), Article 3823. https://doi.org/10.3390/s24123823
  8. Dong, Y., Tan, W., Tao, D., Zheng, L., & Li, X. (2022). CartoonLossGAN: Learning surface and coloring of images for cartoonization. IEEE Transactions on Image Processing, 31, 485–498. https://doi.org/10.1109/TIP.2021.3130539
  9. Fujino, S., Hatanaka, T., Mori, N., & Matsumoto, K. (2019). Evolutionary deep learning based on deep convolutional neural network for anime storyboard recognition. Neurocomputing, 338, 393–398. https://doi.org/10.1016/j.neucom.2018.05.124
  10. Gao, P. (2022). Key technologies of human–computer interaction for immersive somatosensory interactive games using VR technology. Soft Computing, 26(20), 10947–10956. https://doi.org/10.1007/s00500-022-07240-3
  11. He, F., Ong, S. K., & Nee, A. Y. C. (2021). An integrated mobile augmented reality digital twin monitoring system. Computers, 10(8), Article 99. https://doi.org/10.3390/computers10080099
  12. Holden, D., Saito, J., & Komura, T. (2016). A deep learning framework for character motion synthesis and editing. ACM Transactions on Graphics, 35(4). https://doi.org/10.1145/2897824.2925975
  13. Holden, D., Komura, T., & Saito, J. (2017). Phase-functioned neural networks for character control. ACM Transactions on Graphics, 36(4). https://doi.org/10.1145/3072959.3073663
  14. Jang, D.-K., Park, S., & Lee, S.-H. (2022). Motion puzzle: Arbitrary motion style transfer by body part. ACM Transactions on Graphics, 41(3), Article 33. https://doi.org/10.1145/3516429
  15. Karras, T., Aila, T., Laine, S., Herva, A., & Lehtinen, J. (2017). Audio-driven facial animation by joint end-to-end learning of pose and emotion. ACM Transactions on Graphics, 36(4). https://doi.org/10.1145/3072959.3073658
  16. Kim, H., Garrido, P., Tewari, A., Xu, W., Thies, J., Nießner, M., Pérez, P., Richardt, C., Zollhöfer, M., & Theobalt, C. (2018). Deep video portraits. ACM Transactions on Graphics, 37(4). https://doi.org/10.1145/3197517.3201283
  17. Kumar, A., Saudagar, A. K. J., Alkhathami, M., Alsamani, B., Khan, M. B., Hasanat, M. H. A., Ahmed, Z. H., Kumar, A., & Srinivasan, B. (2023). Gamified learning and assessment using ARCS with next-generation AIoMT integrated 3D animation and virtual reality simulation. Electronics, 12(4), Article 835. https://doi.org/10.3390/electronics12040835
  18. Leo-Liu, J., & Wu-Ouyang, B. (2024). A "soul" emerges when AI, AR, and Anime converge: A case study on users of the new anime-stylized hologram social robot "Hupo." New Media & Society, 26(7), 3810–3832. https://doi.org/10.1177/14614448221106030
  19. Ling, H. Y., Zinno, F., Cheng, G., & Van de Panne, M. (2020). Character controllers using motion VAEs. ACM Transactions on Graphics, 39(4). https://doi.org/10.1145/3386569.3392422
  20. Lopes, A. T., de Aguiar, E., De Souza, A. F., & Oliveira-Santos, T. (2017). Facial expression recognition with Convolutional Neural Networks: Coping with few data and the training sample order. Pattern Recognition, 61, 610–628. https://doi.org/10.1016/j.patcog.2016.07.026
  21. Luo, Y.-S., Soeseno, J. H., Chen, T. P.-C., & Chen, W.-C. (2020). CARL: Controllable agent with reinforcement learning for quadruped locomotion. ACM Transactions on Graphics, 39(4), Article 38. https://doi.org/10.1145/3386569.3392433
  22. Mason, I., Starke, S., & Komura, T. (2022). Real-time style modelling of human locomotion via feature-wise transformations and local motion phases. Proceedings of the ACM on Computer Graphics and Interactive Techniques, 5(1). https://doi.org/10.1145/3522618
  23. Metzger, S. L., Littlejohn, K. T., Silva, A. B., Moses, D. A., Seaton, M. P., Wang, R., Dougherty, M. E., Liu, J. R., Wu, P., Berger, M. A., Zhuravleva, I., Tu-Chan, A., Ganguly, K., Anumanchipalli, G. K., & Chang, E. F. (2023). A high-performance neuroprosthesis for speech decoding and avatar control. Nature, 620, 1037–1046. https://doi.org/10.1038/s41586-023-06443-4
  24. Niu, Z., Lu, K., Xue, J., Qin, X., Wang, J., & Shao, L. (2024). From methods to applications: A review of deep 3D human motion capture. IEEE Transactions on Circuits and Systems for Video Technology, 34(11), 11340–11359. https://doi.org/10.1109/TCSVT.2024.3423411
  25. Oshiba, J., Iwata, M., & Kise, K. (2023). Face image generation of anime characters using an advanced first order motion model with facial landmarks. IEICE Transactions on Information and Systems, E106-D(1), 22–30. https://doi.org/10.1587/transinf.2022MUP0004
  26. Peng, X. B., Berseth, G., Yin, K., & Van de Panne, M. (2017). DeepLoco: Dynamic locomotion skills using hierarchical deep reinforcement learning. ACM Transactions on Graphics, 36(4). https://doi.org/10.1145/3072959.3073602
  27. Peng, X. B., Abbeel, P., Levine, S., & Van de Panne, M. (2018a). DeepMimic: Example-guided deep reinforcement learning of physics-based character skills. ACM Transactions on Graphics, 37(4). https://doi.org/10.1145/3197517.3201311
  28. Peng, X. B., Kanazawa, A., Malik, J., Abbeel, P., & Levine, S. (2018b). SFV: Reinforcement learning of physical skills from videos. ACM Transactions on Graphics, 37(6), Article 178. https://doi.org/10.1145/3272127.3275014
  29. Qin, J., Zheng, Y., & Zhou, K. (2022). Motion in-betweening via two-stage transformers. ACM Transactions on Graphics, 41(6), Article 184. https://doi.org/10.1145/3550454.3555454
  30. Smith, H. J., Cao, C., Neff, M., & Wang, Y. (2019). Efficient neural networks for real-time motion style transfer. Proceedings of the ACM on Computer Graphics and Interactive Techniques, 2(2), Article 13. https://doi.org/10.1145/3340254
  31. Sun, L., Chen, P., Xiang, W., Chen, P., Gao, W.-Y., & Zhang, K.-J. (2019). SmartPaint: A co-creative drawing system based on generative adversarial networks. Frontiers of Information Technology & Electronic Engineering, 20(12), 1644–1656. https://doi.org/10.1631/FITEE.1900386
  32. Taylor, S., Kim, T., Yue, Y., Mahler, M., Krahe, J., Rodriguez, A. G., Hodgins, J., & Matthews, I. (2017). A deep learning approach for generalized speech animation. ACM Transactions on Graphics, 36(4). https://doi.org/10.1145/3072959.3073699
  33. Wang, B., & Shi, Y. (2023). Expression dynamic capture and 3D animation generation method based on deep learning. Neural Computing and Applications, 35(12), 8797–8808. https://doi.org/10.1007/s00521-022-07644-0
  34. Wang, C. (2024). Research on prompt engineering for large model art image generation. Journal of Graphics, 45(6), 1243–1255. https://doi.org/10.11996/JG.j.2095-302X.2024061243
  35. Wang, Y., Wang, R., Shi, H., & Liu, D. (2024). MS-HRNet: Multi-scale high-resolution network for human pose estimation. The Journal of Supercomputing, 80(12), 17269–17291. https://doi.org/10.1007/s11227-024-06125-6
  36. Wiatowski, T., & Bölcskei, H. (2018). A mathematical theory of deep convolutional neural networks for feature extraction. IEEE Transactions on Information Theory, 64(3), 1845–1866. https://doi.org/10.1109/TIT.2017.2776228
  37. Xie, Y., Franz, E., Chu, M., & Thuerey, N. (2018). TempoGAN: A temporally coherent, volumetric GAN for super-resolution fluid flow. ACM Transactions on Graphics, 37(4). https://doi.org/10.1145/3197517.3201304
  38. Xie, Y., Takikawa, T., Saito, S., Litany, O., Yan, S., Khan, N., Tombari, F., Tompkin, J., Sitzmann, V., & Sridhar, S. (2022). Neural fields in visual computing and beyond. Computer Graphics Forum, 41(2), 641–676. https://doi.org/10.1111/cgf.14505
  39. Yang, D., Zhang, J., Sun, Y., & Huang, Z. (2024a). Showing usage behavior or not? The effect of virtual influencers' product usage behavior on consumers. Journal of Retailing and Consumer Services, 79, Article 103859. https://doi.org/10.1016/j.jretconser.2024.103859
  40. Yang, Z., Wen, Y.-H., Chen, S.-Y., Liu, X., Gao, Y., Liu, Y.-J., Gao, L., & Fu, H. (2024b). Keyframe control of music-driven 3D dance generation. IEEE Transactions on Visualization and Computer Graphics, 30(7), 3474–3486. https://doi.org/10.1109/TVCG.2023.3235538
  41. Zhang, H., Ye, Y., Shiratori, T., & Komura, T. (2021). ManipNet: Neural manipulation synthesis with a hand-object spatial representation. ACM Transactions on Graphics, 40(4), Article 121. https://doi.org/10.1145/3450626.3459830
  42. Zhao, Y., Ren, D., Chen, Y., Jia, W., Wang, R., & Liu, X. (2022). Cartoon image processing: A survey. International Journal of Computer Vision, 130(11), 2733–2769. https://doi.org/10.1007/s11263-022-01645-1
  43. Zhou, Y., Xu, Z., Landreth, C., Kalogerakis, E., Maji, S., & Singh, K. (2018). VisemeNet: Audio-driven animator-centric speech animation. ACM Transactions on Graphics, 37(4), Article 161. https://doi.org/10.1145/3197517.3201292.