Research / 03

DynaMER

Dynamic Representation Augmentation for Multimodal Emotion Recognition with MLLMs

  • Multimodal Learning
  • Emotion Recognition
  • MLLMs

Making time-varying visual and acoustic evidence explicit before language decoding.

Question
How can multimodal emotion recognition preserve short-term visual and acoustic dynamics instead of relying on the language model to reconstruct them from compressed semantic representations?
Approach
DynaMER constructs modality-specific Dynamic Representations and adaptively integrates them with Global Representations before multimodal language decoding.
Contribution
A dynamic representation augmentation framework for explicitly exposing temporally discriminative audiovisual evidence to multimodal fusion.
DynaMER architecture overview diagram showing dynamic representation augmentation and adaptive routed prefusion before language decodingDynaMER architecture overview diagram showing dynamic representation augmentation and adaptive routed prefusion before language decoding
DynaMER explicitly augments high-level Global Representations with modality-specific Dynamic Representations before multimodal language decoding.

Research Problem

Multimodal emotion is not only about what is present, but how visual and acoustic signals change over time.

Existing emotion-oriented MLLMs can capture high-level semantic information from video and audio, but short-term dynamics may be compressed or left for the language-model backbone to infer implicitly.

For emotion recognition, these temporal changes can be important: facial motion, local appearance changes, prosody, loudness, pitch and voice quality all evolve over time.

DynaMER asks whether these dynamics can instead be represented explicitly before language decoding.

Research Approach

DynaMER dynamic evidence construction across visual and acoustic streams
Dynamic evidence construction: extracting motion vectors, prediction residuals, and acoustic trajectory descriptors to expose short-term dynamics.

DynaMER separates each modality into two complementary views.

Global Representations capture higher-level semantic context.

Dynamic Representations encode short-term temporal evidence that may otherwise be difficult to recover after semantic compression.

For video, DynaMER uses codec-native motion vectors and prediction residuals to represent local motion and appearance changes.

For audio, it models trajectories of acoustic descriptors related to prosody and voice quality.

These representations are temporally aligned and combined through Dynamic–Global Routed Prefusion before entering the language-model interface.

Method

  1. Dynamic Evidence Construction

    Visual and acoustic streams are divided into aligned temporal windows. Global encoders capture semantic context, while dedicated dynamic encoders model motion, appearance change and short-term acoustic variation.

  2. Dynamic–Global Routed Prefusion

    Within each modality, DynaMER first estimates how much Dynamic evidence should augment the Global representation. It then performs sample-level allocation across visual and acoustic streams.

  3. Controlled Evaluation

    The evaluation separates overall benchmark performance from mechanism-focused controls, including dynamic-evidence, routing and temporal-structure ablations.

Key Findings

  1. Explicit dynamic evidence provides information that is not captured by Global representations alone.

    Interpretation

    Short-term motion, appearance change and acoustic variation contribute emotionally relevant cues that may be weakened when only high-level semantic representations are preserved.

  2. Adaptive routing matters: dynamic evidence should not contribute equally for every sample or modality.

    Interpretation

    The role of dynamic evidence depends on the sample and on the modality, so effective fusion benefits from adaptive allocation rather than fixed equal treatment.

  3. Explicit domain-native evidence can reduce dependence on LLM backbone scale.

    Interpretation

    By exposing temporally discriminative audiovisual evidence earlier in the pipeline, the system can rely less on backbone size alone to recover emotion-relevant dynamics.

Contribution

  1. Representation

    Explicit Dynamic Representations for visual and acoustic temporal evidence.

  2. Fusion

    Dynamic–Global Routed Prefusion for adaptive integration of Global and Dynamic evidence.

  3. Evaluation

    A controlled evaluation framework for separating performance gains from the effects of dynamic evidence, routing and model capacity.