A multimodal framework of continuous music emotion recognition for adaptive cockpit lighting
Continuous music emotion recognition remains sensitive to dataset shift and imperfect lyric alignment. We propose Cockpit-EmoNet, which combines a frozen MERT acoustic encoder and a partially fine-tuned GTE text encoder through bidirectional cross-attention, feature-wise gating, and joint regression–alignment learning. On the pooled PMEmo–DEAM benchmark, Cockpit-EmoNet achieves valence PCC/CCC of 0.69/0.67 and arousal PCC/CCC of 0.81/0.79.
Across five runs, it significantly reduces song-level errors relative to Cross-Attention Only (\(p_}
Extract — continue reading at the source.