EGCAF: Expert-Guided Cross-Modal Attention Fusion (MELD)

This repository contains EGCAF, the multimodal fusion architecture presented in Chapter 7 of the PhD thesis A Multimodal Approach for Emotion Recognition from Multimedia Content (Jihed Jabnoun, Faculty of Sciences of Monastir, 2026) and in the accompanying ACIIDS 2026 paper. EGCAF recognises seven emotions (Anger, Disgust, Fear, Happiness, Neutral, Sadness, Surprise) from text, audio, and facial video in conversational utterances.

EGCAF keeps three pretrained emotion experts frozen and trains only a small fusion module on top of them: 8.07M trainable parameters out of 270.6M (97.0% frozen). The frozen experts cannot drift toward the rebalanced training distribution, which is what allows minority-class oversampling without the accuracy penalty observed in jointly trained early fusion.

The implementation class is named HybridLateFusionToEarlyFusion in modeling_hybrid.py, a working name from development. It is the EGCAF model described in the thesis.

Architecture

EGCAF architecture

  • Frozen experts:
  • Cross-modal attention: 8-head attention over the three projected modality vectors (one vector per modality per utterance), with a residual connection.
  • Emotion-specific gating: seven two-layer gating networks (1,557β†’256β†’3), one per emotion, each combining the attended representations with the stacked expert logits. The gate scores are multiplied by a learned confidence adjustment and by fixed affinity priors derived from each expert's recall on the MELD validation set, then renormalised over the available modalities.
  • Emotion-specific classification: seven independent heads (1,557β†’256β†’128β†’1). Each head receives the three attended representations, scaled by their fusion weights, concatenated with the unweighted logits of all three experts.
  • Loss: focal loss (Ξ± = 0.25, Ξ³ = 2.0) plus a KL-divergence term (Ξ» = 0.1) anchoring the output to the average of the expert predictions.

Results on MELD

Test set, 2,610 utterances. All figures come from a single run with a fixed seed (42).

Approach Acc Macro-F1 Weighted-F1 Disgust F1 Fear F1
Text expert only 58.3 48.2 59.4 – –
Audio expert only 17.7 13.4 13.1 – –
Vision expert only 13.2 10.7 16.4 – –
Late fusion, equal averaging (MELD-fine-tuned experts) 61.1 35.1 56.0 0.078 0.068
Early fusion, concatenation 63.9 38.1 61.2 0.000 0.000
Early fusion, cross-attention 65.1 40.0 62.9 0.000 0.000
Early fusion, concatenation + augmentation 58.5 42.3 59.2 0.241 0.122
Early fusion, emotion-specific projections + focal loss 63.5 41.9 61.9 0.222 0.000
EGCAF 66.7 51.1 66.2 0.292 0.272

Per-emotion results

Emotion Precision Recall F1 Support Prior weights (T/A/V) Learned weights (T/A/V)
Anger 0.51 0.60 0.552 345 0.72 / 0.18 / 0.10 0.56 / 0.21 / 0.23
Disgust 0.50 0.21 0.292 68 0.84 / 0.09 / 0.07 0.16 / 0.00 / 0.84
Fear 0.35 0.22 0.272 50 0.90 / 0.06 / 0.04 0.07 / 0.03 / 0.91
Happiness 0.61 0.58 0.595 402 0.20 / 0.07 / 0.73 0.99 / 0.01 / 0.00
Neutral 0.79 0.79 0.793 1256 0.65 / 0.25 / 0.10 0.88 / 0.06 / 0.06
Sadness 0.59 0.39 0.470 208 0.23 / 0.73 / 0.04 0.68 / 0.23 / 0.09
Surprise 0.53 0.70 0.606 281 0.47 / 0.12 / 0.41 0.29 / 0.06 / 0.64

Learned weights are the final fusion weights averaged over MELD test samples of each emotion. They scale the attended representations, which already mix information across modalities, and the expert logits of all three modalities reach every classifier head unweighted. A high "vision" weight therefore means the vision-anchored representation is scaled up, not that the decision is made from facial evidence alone.

Confusion matrix for EGCAF on MELD

The main error patterns are:

  • Absorption by Neutral: 30.8% of Sadness, 27.9% of Disgust and 24.0% of Fear are predicted as Neutral.
  • Disgust β†’ Anger: 30.9% of Disgust samples.
  • Fear: errors split between Anger and Neutral (24.0% each). Fear and Surprise confuse rarely in the fused predictions.

Context: published systems on MELD

Weighted-F1 is used here because every system reports it. Published figures are as compiled by Fu et al. (BeMERC, 2025).

System Weighted-F1
BeMERC (LLM-based, ~7B parameters) 69.78
SpeechCueLLM 67.60
DGODE 67.20
M3Net 67.05
MultiEMO 66.74
EGCAF (271M total, 8.07M trainable) 66.2
UniMSE 65.51

EGCAF is not state of the art: it is 3.6 weighted-F1 points below BeMERC. It is competitive with the strongest graph- and attention-based systems while training only 8M parameters. Its Disgust and Fear F1 are comparable to, not above, the strongest published systems.

Cross-dataset evaluation

These results use the MELD-trained model with no retraining on the new data.

Dataset Text expert only (Acc / M-F1) EGCAF (Acc / M-F1)
MELD (in-domain) 58.3 / 48.2 66.7 / 51.1
IEMOCAP Session 5 (vision replaced by a blank frame) 39.2 / 31.0 47.3 / 38.5
EmoVid (film and TV clips) 50.4 / 45.4 39.1 / 35.5
  • IEMOCAP: EGCAF transfers positively to IEMOCAP, even with the visual input withheld.
  • EmoVid: EGCAF falls 11.3 points below its own text expert, while simple averaging over the same kind of experts does well on this corpus. The thesis calls this failure mode reliability inversion: for three of seven emotions (Surprise, Happiness, Fear), the gates learned on MELD give the dominant weight to the modality that is least reliable on EmoVid.
  • Priors vs learned gates: the fixed priors match EmoVid's reliability ordering for six of seven emotions. What fails to transfer is the learned override of those priors. Validate on the target domain before any deployment.

Missing modalities (MELD)

Available modalities Accuracy Macro-F1
Text + Audio + Vision 66.7 51.1
Text + Audio 65.2 48.3
Text + Vision 64.8 47.9
Text only 60.1 45.2
Audio + Vision 38.2 28.1

An absent modality is replaced by a blank input of the expected shape (for example, a zero-valued frame); the model has no explicit availability mask. Degradation is small while text is available and severe without it.

Efficiency

  • Parameters: 270,569,012 total, 8,070,815 trainable (97.0% frozen).
  • Latency: about 158 ms mean per utterance on a Tesla T4. Expert forward passes account for about 73% of this; the fusion module adds about 11 ms.
  • Weight memory: about 1.03 GB in FP32 and 0.52 GB in FP16. Activation memory is additional and grows with batch size.

Training configuration

  • Data: MELD, with Disgust and Fear oversampled 4Γ— and Sadness 2Γ—.
  • Optimiser: AdamW, learning rate 2e-5, weight decay 0.01, 500 warmup steps.
  • Batch: batch size 16 with gradient accumulation over 4 steps (effective 64).
  • Schedule: up to 4 epochs, early stopping with patience 3 on validation weighted-F1.
  • Loss: focal loss (Ξ± = 0.25, Ξ³ = 2.0) + 0.1 Γ— KL divergence to the expert ensemble.
  • Hardware: FP16 mixed precision on a single Tesla T4 (16 GB); training time 101.3 minutes.

Usage

pip install torch transformers librosa pillow numpy huggingface_hub
import torch
from huggingface_hub import hf_hub_download

REPO_ID = "jihedjabnoun/egcaf-meld"

model_def_path = hf_hub_download(repo_id=REPO_ID, filename="modeling_hybrid.py")
model_path = hf_hub_download(repo_id=REPO_ID, filename="pytorch_model.bin")

# Defines HybridLateFusionToEarlyFusion (EGCAF), EMOTION_LABELS and ID_TO_LABEL
exec(open(model_def_path).read())

model = HybridLateFusionToEarlyFusion(
    num_labels=7,
    hidden_size=512,
    num_heads=8,
    freeze_experts=True,
)
model.load_state_dict(torch.load(model_path, map_location="cpu"))
model.eval()

from transformers import AutoTokenizer, Wav2Vec2FeatureExtractor, ViTImageProcessor

text_tokenizer = AutoTokenizer.from_pretrained("jihedjabnoun/text-emotion-distilroberta")
audio_extractor = Wav2Vec2FeatureExtractor.from_pretrained("jihedjabnoun/hubert-emotion-recognition-v2")
vision_processor = ViTImageProcessor.from_pretrained("jihedjabnoun/vit-face-emotion-recognition")
import librosa
import numpy as np
from PIL import Image

def predict_emotion(text, audio_path, face_images):
    # Text
    text_inputs = text_tokenizer(text, padding="max_length", truncation=True,
                                 max_length=128, return_tensors="pt")

    # Audio: 3 seconds at 16 kHz
    audio, _ = librosa.load(audio_path, sr=16000)
    target_length = 3 * 16000
    audio = audio[:target_length] if len(audio) > target_length else \
        np.pad(audio, (0, target_length - len(audio)), mode="constant")
    audio_inputs = audio_extractor(audio, sampling_rate=16000,
                                   return_tensors="pt", padding=True)

    # Vision: up to 5 face frames, averaged pixel-wise into one image (as in training)
    faces = [f.resize((224, 224), Image.Resampling.LANCZOS) for f in face_images[:5]]
    vision_inputs = vision_processor(images=faces, return_tensors="pt")
    pixel_values = vision_inputs["pixel_values"].mean(dim=0, keepdim=True)

    with torch.no_grad():
        outputs = model(
            input_ids=text_inputs["input_ids"],
            attention_mask=text_inputs["attention_mask"],
            audio_values=audio_inputs["input_values"],
            audio_attention_mask=torch.ones_like(audio_inputs["input_values"]),
            pixel_values=pixel_values,
        )

    logits = outputs["logits"]
    probs = torch.softmax(logits, dim=-1)
    predicted_id = torch.argmax(logits, dim=-1).item()

    result = {
        "predicted_emotion": ID_TO_LABEL[predicted_id],
        "confidence": probs[0][predicted_id].item(),
        "all_probabilities": {ID_TO_LABEL[i]: probs[0][i].item() for i in range(7)},
    }
    if hasattr(model, "last_expert_logits"):
        result["expert_predictions"] = {
            m: ID_TO_LABEL[torch.argmax(model.last_expert_logits[m]).item()]
            for m in ["text", "audio", "vision"]
        }
    if hasattr(model, "last_gate_values"):
        # Gate scores for the predicted emotion, before the confidence adjustment and priors
        gates = model.last_gate_values[predicted_id, 0, :].numpy()
        result["gate_values"] = {"text": float(gates[0]), "audio": float(gates[1]),
                                 "vision": float(gates[2])}
    return result

Intended use and limitations

EGCAF is a research model. It is not suitable for clinical diagnosis, hiring decisions, surveillance, or any high-stakes decision without human review.

  • Minority classes: recall on Disgust and Fear is 21–22%, so most instances of these emotions are missed.
  • Text dependence: performance collapses without the text channel (38.2% accuracy with audio and vision only).
  • Domain shift: the learned gates encode MELD's reliability structure and can invert on new domains (39.1% on EmoVid versus 50.4% for text alone). Check per-emotion performance on a labelled sample of the target domain before use.
  • Training data: MELD consists of scripted English dialogue from one television series, and its official splits reuse the same main characters, so MELD results are not speaker-independent.
  • Utterance-level only: predictions are made per utterance from pooled representations. Emotional context across turns and transient cues within an utterance are not modelled.
  • Single runs: all figures come from single runs, without confidence intervals.

Repository contents

  • pytorch_model.bin: model weights
  • config.json: model configuration
  • modeling_hybrid.py: model implementation (class HybridLateFusionToEarlyFusion)
  • detailed_results.json: evaluation metrics
  • test_predictions_detailed.csv: all MELD test predictions
  • egcaf_arch.png: architecture diagram
  • confusion_matrix.png: MELD test confusion matrix
  • augmentation_log.json: data augmentation details

Citation

@phdthesis{jabnoun2026multimodal,
  title  = {A Multimodal Approach for Emotion Recognition from Multimedia Content},
  author = {Jabnoun, Jihed},
  school = {Faculty of Sciences of Monastir, University of Monastir},
  year   = {2026}
}

@inproceedings{jabnoun2026adaptive,
  title     = {Adaptive Multimodal Fusion for Interpretable and Efficient Conversational Emotion Recognition},
  author    = {Jabnoun, Jihed and Maraoui, Mohsen and Zrigui, Mounir},
  booktitle = {18th Asian Conference on Intelligent Information and Database Systems (ACIIDS)},
  year      = {2026},
  note      = {Accepted}
}
Downloads last month
24
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Evaluation results