EGCAF: Expert-Guided Cross-Modal Attention Fusion (MELD)
This repository contains EGCAF, the multimodal fusion architecture presented in Chapter 7 of the PhD thesis A Multimodal Approach for Emotion Recognition from Multimedia Content (Jihed Jabnoun, Faculty of Sciences of Monastir, 2026) and in the accompanying ACIIDS 2026 paper. EGCAF recognises seven emotions (Anger, Disgust, Fear, Happiness, Neutral, Sadness, Surprise) from text, audio, and facial video in conversational utterances.
EGCAF keeps three pretrained emotion experts frozen and trains only a small fusion module on top of them: 8.07M trainable parameters out of 270.6M (97.0% frozen). The frozen experts cannot drift toward the rebalanced training distribution, which is what allows minority-class oversampling without the accuracy penalty observed in jointly trained early fusion.
The implementation class is named
HybridLateFusionToEarlyFusioninmodeling_hybrid.py, a working name from development. It is the EGCAF model described in the thesis.
Architecture
- Frozen experts:
- Cross-modal attention: 8-head attention over the three projected modality vectors (one vector per modality per utterance), with a residual connection.
- Emotion-specific gating: seven two-layer gating networks (1,557β256β3), one per emotion, each combining the attended representations with the stacked expert logits. The gate scores are multiplied by a learned confidence adjustment and by fixed affinity priors derived from each expert's recall on the MELD validation set, then renormalised over the available modalities.
- Emotion-specific classification: seven independent heads (1,557β256β128β1). Each head receives the three attended representations, scaled by their fusion weights, concatenated with the unweighted logits of all three experts.
- Loss: focal loss (Ξ± = 0.25, Ξ³ = 2.0) plus a KL-divergence term (Ξ» = 0.1) anchoring the output to the average of the expert predictions.
Results on MELD
Test set, 2,610 utterances. All figures come from a single run with a fixed seed (42).
| Approach | Acc | Macro-F1 | Weighted-F1 | Disgust F1 | Fear F1 |
|---|---|---|---|---|---|
| Text expert only | 58.3 | 48.2 | 59.4 | β | β |
| Audio expert only | 17.7 | 13.4 | 13.1 | β | β |
| Vision expert only | 13.2 | 10.7 | 16.4 | β | β |
| Late fusion, equal averaging (MELD-fine-tuned experts) | 61.1 | 35.1 | 56.0 | 0.078 | 0.068 |
| Early fusion, concatenation | 63.9 | 38.1 | 61.2 | 0.000 | 0.000 |
| Early fusion, cross-attention | 65.1 | 40.0 | 62.9 | 0.000 | 0.000 |
| Early fusion, concatenation + augmentation | 58.5 | 42.3 | 59.2 | 0.241 | 0.122 |
| Early fusion, emotion-specific projections + focal loss | 63.5 | 41.9 | 61.9 | 0.222 | 0.000 |
| EGCAF | 66.7 | 51.1 | 66.2 | 0.292 | 0.272 |
Per-emotion results
| Emotion | Precision | Recall | F1 | Support | Prior weights (T/A/V) | Learned weights (T/A/V) |
|---|---|---|---|---|---|---|
| Anger | 0.51 | 0.60 | 0.552 | 345 | 0.72 / 0.18 / 0.10 | 0.56 / 0.21 / 0.23 |
| Disgust | 0.50 | 0.21 | 0.292 | 68 | 0.84 / 0.09 / 0.07 | 0.16 / 0.00 / 0.84 |
| Fear | 0.35 | 0.22 | 0.272 | 50 | 0.90 / 0.06 / 0.04 | 0.07 / 0.03 / 0.91 |
| Happiness | 0.61 | 0.58 | 0.595 | 402 | 0.20 / 0.07 / 0.73 | 0.99 / 0.01 / 0.00 |
| Neutral | 0.79 | 0.79 | 0.793 | 1256 | 0.65 / 0.25 / 0.10 | 0.88 / 0.06 / 0.06 |
| Sadness | 0.59 | 0.39 | 0.470 | 208 | 0.23 / 0.73 / 0.04 | 0.68 / 0.23 / 0.09 |
| Surprise | 0.53 | 0.70 | 0.606 | 281 | 0.47 / 0.12 / 0.41 | 0.29 / 0.06 / 0.64 |
Learned weights are the final fusion weights averaged over MELD test samples of each emotion. They scale the attended representations, which already mix information across modalities, and the expert logits of all three modalities reach every classifier head unweighted. A high "vision" weight therefore means the vision-anchored representation is scaled up, not that the decision is made from facial evidence alone.
The main error patterns are:
- Absorption by Neutral: 30.8% of Sadness, 27.9% of Disgust and 24.0% of Fear are predicted as Neutral.
- Disgust β Anger: 30.9% of Disgust samples.
- Fear: errors split between Anger and Neutral (24.0% each). Fear and Surprise confuse rarely in the fused predictions.
Context: published systems on MELD
Weighted-F1 is used here because every system reports it. Published figures are as compiled by Fu et al. (BeMERC, 2025).
| System | Weighted-F1 |
|---|---|
| BeMERC (LLM-based, ~7B parameters) | 69.78 |
| SpeechCueLLM | 67.60 |
| DGODE | 67.20 |
| M3Net | 67.05 |
| MultiEMO | 66.74 |
| EGCAF (271M total, 8.07M trainable) | 66.2 |
| UniMSE | 65.51 |
EGCAF is not state of the art: it is 3.6 weighted-F1 points below BeMERC. It is competitive with the strongest graph- and attention-based systems while training only 8M parameters. Its Disgust and Fear F1 are comparable to, not above, the strongest published systems.
Cross-dataset evaluation
These results use the MELD-trained model with no retraining on the new data.
| Dataset | Text expert only (Acc / M-F1) | EGCAF (Acc / M-F1) |
|---|---|---|
| MELD (in-domain) | 58.3 / 48.2 | 66.7 / 51.1 |
| IEMOCAP Session 5 (vision replaced by a blank frame) | 39.2 / 31.0 | 47.3 / 38.5 |
| EmoVid (film and TV clips) | 50.4 / 45.4 | 39.1 / 35.5 |
- IEMOCAP: EGCAF transfers positively to IEMOCAP, even with the visual input withheld.
- EmoVid: EGCAF falls 11.3 points below its own text expert, while simple averaging over the same kind of experts does well on this corpus. The thesis calls this failure mode reliability inversion: for three of seven emotions (Surprise, Happiness, Fear), the gates learned on MELD give the dominant weight to the modality that is least reliable on EmoVid.
- Priors vs learned gates: the fixed priors match EmoVid's reliability ordering for six of seven emotions. What fails to transfer is the learned override of those priors. Validate on the target domain before any deployment.
Missing modalities (MELD)
| Available modalities | Accuracy | Macro-F1 |
|---|---|---|
| Text + Audio + Vision | 66.7 | 51.1 |
| Text + Audio | 65.2 | 48.3 |
| Text + Vision | 64.8 | 47.9 |
| Text only | 60.1 | 45.2 |
| Audio + Vision | 38.2 | 28.1 |
An absent modality is replaced by a blank input of the expected shape (for example, a zero-valued frame); the model has no explicit availability mask. Degradation is small while text is available and severe without it.
Efficiency
- Parameters: 270,569,012 total, 8,070,815 trainable (97.0% frozen).
- Latency: about 158 ms mean per utterance on a Tesla T4. Expert forward passes account for about 73% of this; the fusion module adds about 11 ms.
- Weight memory: about 1.03 GB in FP32 and 0.52 GB in FP16. Activation memory is additional and grows with batch size.
Training configuration
- Data: MELD, with Disgust and Fear oversampled 4Γ and Sadness 2Γ.
- Optimiser: AdamW, learning rate 2e-5, weight decay 0.01, 500 warmup steps.
- Batch: batch size 16 with gradient accumulation over 4 steps (effective 64).
- Schedule: up to 4 epochs, early stopping with patience 3 on validation weighted-F1.
- Loss: focal loss (Ξ± = 0.25, Ξ³ = 2.0) + 0.1 Γ KL divergence to the expert ensemble.
- Hardware: FP16 mixed precision on a single Tesla T4 (16 GB); training time 101.3 minutes.
Usage
pip install torch transformers librosa pillow numpy huggingface_hub
import torch
from huggingface_hub import hf_hub_download
REPO_ID = "jihedjabnoun/egcaf-meld"
model_def_path = hf_hub_download(repo_id=REPO_ID, filename="modeling_hybrid.py")
model_path = hf_hub_download(repo_id=REPO_ID, filename="pytorch_model.bin")
# Defines HybridLateFusionToEarlyFusion (EGCAF), EMOTION_LABELS and ID_TO_LABEL
exec(open(model_def_path).read())
model = HybridLateFusionToEarlyFusion(
num_labels=7,
hidden_size=512,
num_heads=8,
freeze_experts=True,
)
model.load_state_dict(torch.load(model_path, map_location="cpu"))
model.eval()
from transformers import AutoTokenizer, Wav2Vec2FeatureExtractor, ViTImageProcessor
text_tokenizer = AutoTokenizer.from_pretrained("jihedjabnoun/text-emotion-distilroberta")
audio_extractor = Wav2Vec2FeatureExtractor.from_pretrained("jihedjabnoun/hubert-emotion-recognition-v2")
vision_processor = ViTImageProcessor.from_pretrained("jihedjabnoun/vit-face-emotion-recognition")
import librosa
import numpy as np
from PIL import Image
def predict_emotion(text, audio_path, face_images):
# Text
text_inputs = text_tokenizer(text, padding="max_length", truncation=True,
max_length=128, return_tensors="pt")
# Audio: 3 seconds at 16 kHz
audio, _ = librosa.load(audio_path, sr=16000)
target_length = 3 * 16000
audio = audio[:target_length] if len(audio) > target_length else \
np.pad(audio, (0, target_length - len(audio)), mode="constant")
audio_inputs = audio_extractor(audio, sampling_rate=16000,
return_tensors="pt", padding=True)
# Vision: up to 5 face frames, averaged pixel-wise into one image (as in training)
faces = [f.resize((224, 224), Image.Resampling.LANCZOS) for f in face_images[:5]]
vision_inputs = vision_processor(images=faces, return_tensors="pt")
pixel_values = vision_inputs["pixel_values"].mean(dim=0, keepdim=True)
with torch.no_grad():
outputs = model(
input_ids=text_inputs["input_ids"],
attention_mask=text_inputs["attention_mask"],
audio_values=audio_inputs["input_values"],
audio_attention_mask=torch.ones_like(audio_inputs["input_values"]),
pixel_values=pixel_values,
)
logits = outputs["logits"]
probs = torch.softmax(logits, dim=-1)
predicted_id = torch.argmax(logits, dim=-1).item()
result = {
"predicted_emotion": ID_TO_LABEL[predicted_id],
"confidence": probs[0][predicted_id].item(),
"all_probabilities": {ID_TO_LABEL[i]: probs[0][i].item() for i in range(7)},
}
if hasattr(model, "last_expert_logits"):
result["expert_predictions"] = {
m: ID_TO_LABEL[torch.argmax(model.last_expert_logits[m]).item()]
for m in ["text", "audio", "vision"]
}
if hasattr(model, "last_gate_values"):
# Gate scores for the predicted emotion, before the confidence adjustment and priors
gates = model.last_gate_values[predicted_id, 0, :].numpy()
result["gate_values"] = {"text": float(gates[0]), "audio": float(gates[1]),
"vision": float(gates[2])}
return result
Intended use and limitations
EGCAF is a research model. It is not suitable for clinical diagnosis, hiring decisions, surveillance, or any high-stakes decision without human review.
- Minority classes: recall on Disgust and Fear is 21β22%, so most instances of these emotions are missed.
- Text dependence: performance collapses without the text channel (38.2% accuracy with audio and vision only).
- Domain shift: the learned gates encode MELD's reliability structure and can invert on new domains (39.1% on EmoVid versus 50.4% for text alone). Check per-emotion performance on a labelled sample of the target domain before use.
- Training data: MELD consists of scripted English dialogue from one television series, and its official splits reuse the same main characters, so MELD results are not speaker-independent.
- Utterance-level only: predictions are made per utterance from pooled representations. Emotional context across turns and transient cues within an utterance are not modelled.
- Single runs: all figures come from single runs, without confidence intervals.
Repository contents
pytorch_model.bin: model weightsconfig.json: model configurationmodeling_hybrid.py: model implementation (classHybridLateFusionToEarlyFusion)detailed_results.json: evaluation metricstest_predictions_detailed.csv: all MELD test predictionsegcaf_arch.png: architecture diagramconfusion_matrix.png: MELD test confusion matrixaugmentation_log.json: data augmentation details
Citation
@phdthesis{jabnoun2026multimodal,
title = {A Multimodal Approach for Emotion Recognition from Multimedia Content},
author = {Jabnoun, Jihed},
school = {Faculty of Sciences of Monastir, University of Monastir},
year = {2026}
}
@inproceedings{jabnoun2026adaptive,
title = {Adaptive Multimodal Fusion for Interpretable and Efficient Conversational Emotion Recognition},
author = {Jabnoun, Jihed and Maraoui, Mohsen and Zrigui, Mounir},
booktitle = {18th Asian Conference on Intelligent Information and Database Systems (ACIIDS)},
year = {2026},
note = {Accepted}
}
- Downloads last month
- 24
Evaluation results
- Accuracy on MELD (Multimodal EmotionLines Dataset)test set self-reported0.667
- Macro F1-Score on MELD (Multimodal EmotionLines Dataset)test set self-reported0.511
- Weighted F1-Score on MELD (Multimodal EmotionLines Dataset)test set self-reported0.662

