You need to agree to share your contact information to access this model

This repository is publicly accessible, but you have to accept the conditions to access its files and content.

Access is granted manually by the owner. State your name, affiliation and intended use.

Log in or Sign Up to review the conditions and access this model content.

VideoFace-X-LLM โ€” ViT-only fine-tune (step 3406 of 3406)

InternVideo2.5-Chat-8B fine-tuned on the ExplainFace corpus (Gemini descriptions of CelebV-HQ videos) for the ACCV 2026 paper "When, Where, and Why: Hierarchical Causal Explainability for Video Face Understanding with MLLMs". This is the "Ours (ViT-Only)" row of the results table.

Training

  • Trainable: vision encoder (lr 1e-5) + connector mlp1 (lr 4e-5); language model frozen.
  • 1 epoch = 3,406 steps over the 27,304-clip identity-independent training split (Aak579/CelebV-HQ-identity-split20), micro-batch 1, global batch 8, AdamW, weight decay 0.01, 3 % linear warm-up then cosine decay, gradient clip 1.0, 8โ€“32 frames per clip, seed 42, FSDP full sharding (xtuner unify_internvl2_train_r16.py). Full args: train_run_config.json.
  • This checkpoint is the end of the epoch.

Test results (6,923-clip identity-independent test split)

Appearance mA / F1mac Action mA / F1mac Emotion video Acc / F1mac Emotion temporal Acc / F1mac Deepfake (FF++) Acc / F1mac AU (DISFA) mA / frame F1 / clip F1
86.16 / 26.29 94.50 / 17.88 55.61 / 22.00 71.47 / 23.83 50.00 / 35.23 71.0 / 16.7 / 16.8

Loading

The folder is a full HF checkpoint (trust_remote_code=True; the modeling_*.py / conversation.py files are copied from the base snapshot). For the JSON tasks the model was decoded with the training-style prompt prefix and a { prefill, as in the paper's inference scripts.

Downloads last month
-
Safetensors
Model size
8B params
Tensor type
BF16
ยท
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for Aak579/VideoFace-X-LLM-ViT-only-3406

Finetuned
(3)
this model