Quality gated multimodal fusion improves scene construction and interaction experience in augmented reality digital media
Augmented reality now carries a great deal of digital media production, yet its scene construction still suffers from registration drift, asynchronous sensor streams, unnatural interaction and excessive cognitive load. We address these coupled problems with an integrated framework. Hardware-triggered acquisition and timestamping disciplined by the Precision Time Protocol feed a quality-gated cross-modal attention model whose fusion weights follow per-modality reliability rather than fixed priors.
A dynamic scene construction algorithm then couples semantic matching, layout optimization by covariance matrix adaptation and harmonic relighting, so that virtual content stays geometrically and photometrically consistent with the room around it, while an adaptive feedback policy driven by the decoder’s uncertainty estimates tunes interaction cues to the demand of the current task. On a corpus of 184 controlled sequences together with public benchmark subsets, running on a HoloLens 2 platform, the system reaches 91.4% fusion accuracy at a median end-to-end latency of 16.8 ms with a 95th percentile of 18.9 ms, and its robustness margin widens when two modalities degrade at once. A within-subjects user study (n = 50) recorded reliable gains in immersion, presence, usability, satisfaction and cognitive load, with an exploratory analysis suggesting that first-time users benefit more than experienced ones.
The approach offers a coherent path toward AR digital media that remains usable beyond enthusiast audiences, although outdoor scalability, hardware dependence and long-session fatigue all remain open. Covariance matrix adaptation evolution strategy Bidirectional reflectance distribution function Learned perceptual image patch similarity Questionnaire for user interaction satisfaction Special Euclidean group in three dimensions The authors used a generative language model to check grammar and improve the readability of the English text. No content, analysis, result or reference was generated by that tool, every citation has been verified against the primary source, and the authors take full responsibility for the entire manuscript.
The authors declare that no funding, grants, or other financial support was received during the preparation of this manuscript. Department of Formative Fusion Art, General Graduate School, Hoseo University, Cheonan, Chungcheongnam-Do, 31518, South Korea School of Fine Arts and Design, Huizhou University, Huizhou, 516000, Guangdong, China The authors declare no competing interests. This study was approved by the Research Ethics Committee of Hoseo University (Reference Number: HSU-IRB-2024-0427).
All participants provided written informed consent prior to enrollment. The study was conducted in accordance with the Declaration of Helsinki and relevant national regulations. All authors have reviewed the manuscript and consent to its publication.
No identifiable information regarding participants has been included. Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations. Open Access This article is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License, which permits any non-commercial use, sharing, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons licence, and indicate if you modified the licensed material.
Extract — continue reading at the source.