You need to agree to share your contact information to access this model
This repository is publicly accessible, but you have to accept the conditions to access its files and content.
Access is granted manually by the owner. State your name, affiliation and intended use.
Log in or Sign Up to review the conditions and access this model content.
VideoFace-X-LLM โ ViT-only fine-tune (step 3406 of 3406)
InternVideo2.5-Chat-8B fine-tuned on the ExplainFace corpus (Gemini descriptions of CelebV-HQ videos) for the ACCV 2026 paper "When, Where, and Why: Hierarchical Causal Explainability for Video Face Understanding with MLLMs". This is the "Ours (ViT-Only)" row of the results table.
Training
- Trainable: vision encoder (lr 1e-5) + connector
mlp1(lr 4e-5); language model frozen. - 1 epoch = 3,406 steps over the 27,304-clip identity-independent training split
(
Aak579/CelebV-HQ-identity-split20), micro-batch 1, global batch 8, AdamW, weight decay 0.01, 3 % linear warm-up then cosine decay, gradient clip 1.0, 8โ32 frames per clip, seed 42, FSDP full sharding (xtunerunify_internvl2_train_r16.py). Full args:train_run_config.json. - This checkpoint is the end of the epoch.
Test results (6,923-clip identity-independent test split)
| Appearance mA / F1mac | Action mA / F1mac | Emotion video Acc / F1mac | Emotion temporal Acc / F1mac | Deepfake (FF++) Acc / F1mac | AU (DISFA) mA / frame F1 / clip F1 |
|---|---|---|---|---|---|
| 86.16 / 26.29 | 94.50 / 17.88 | 55.61 / 22.00 | 71.47 / 23.83 | 50.00 / 35.23 | 71.0 / 16.7 / 16.8 |
Loading
The folder is a full HF checkpoint (trust_remote_code=True; the modeling_*.py / conversation.py
files are copied from the base snapshot). For the JSON tasks the model was decoded with the
training-style prompt prefix and a { prefill, as in the paper's inference scripts.
- Downloads last month
- -
Model tree for Aak579/VideoFace-X-LLM-ViT-only-3406
Base model
OpenGVLab/InternVideo2_5_Chat_8B