Research / 03
DynaMER
Dynamic Representation Augmentation for Multimodal Emotion Recognition with MLLMs
Making time-varying visual and acoustic evidence explicit before language decoding.
- Question
- How can multimodal emotion recognition preserve short-term visual and acoustic dynamics instead of relying on the language model to reconstruct them from compressed semantic representations?
- Approach
- DynaMER constructs modality-specific Dynamic Representations and adaptively integrates them with Global Representations before multimodal language decoding.
- Contribution
- A dynamic representation augmentation framework for explicitly exposing temporally discriminative audiovisual evidence to multimodal fusion.
Research Problem
Multimodal emotion is not only about what is present, but how visual and acoustic signals change over time.
Existing emotion-oriented MLLMs can capture high-level semantic information from video and audio, but short-term dynamics may be compressed or left for the language-model backbone to infer implicitly.
For emotion recognition, these temporal changes can be important: facial motion, local appearance changes, prosody, loudness, pitch and voice quality all evolve over time.
DynaMER asks whether these dynamics can instead be represented explicitly before language decoding.
Research Approach
DynaMER separates each modality into two complementary views.
Global Representations capture higher-level semantic context.
Dynamic Representations encode short-term temporal evidence that may otherwise be difficult to recover after semantic compression.
For video, DynaMER uses codec-native motion vectors and prediction residuals to represent local motion and appearance changes.
For audio, it models trajectories of acoustic descriptors related to prosody and voice quality.
These representations are temporally aligned and combined through Dynamic–Global Routed Prefusion before entering the language-model interface.
Method
Dynamic Evidence Construction
Visual and acoustic streams are divided into aligned temporal windows. Global encoders capture semantic context, while dedicated dynamic encoders model motion, appearance change and short-term acoustic variation.
Dynamic–Global Routed Prefusion
Within each modality, DynaMER first estimates how much Dynamic evidence should augment the Global representation. It then performs sample-level allocation across visual and acoustic streams.
Controlled Evaluation
The evaluation separates overall benchmark performance from mechanism-focused controls, including dynamic-evidence, routing and temporal-structure ablations.
Key Findings
Explicit dynamic evidence provides information that is not captured by Global representations alone.
Interpretation
Short-term motion, appearance change and acoustic variation contribute emotionally relevant cues that may be weakened when only high-level semantic representations are preserved.
Adaptive routing matters: dynamic evidence should not contribute equally for every sample or modality.
Interpretation
The role of dynamic evidence depends on the sample and on the modality, so effective fusion benefits from adaptive allocation rather than fixed equal treatment.
Explicit domain-native evidence can reduce dependence on LLM backbone scale.
Interpretation
By exposing temporally discriminative audiovisual evidence earlier in the pipeline, the system can rely less on backbone size alone to recover emotion-relevant dynamics.
Contribution
Representation
Explicit Dynamic Representations for visual and acoustic temporal evidence.
Fusion
Dynamic–Global Routed Prefusion for adaptive integration of Global and Dynamic evidence.
Evaluation
A controlled evaluation framework for separating performance gains from the effects of dynamic evidence, routing and model capacity.