CM-CAT: A Cross-Modal Contextual Attention Transformer for Robust Mental Health Prediction via Textual and Facial Temporal Dynamics
Keywords:
Affective computing; Cross-modal fusion; Transformer; Depression detection; Explainable AI (XAI); Facial action units.Abstract
The mental health crisis due to the COVID-19 pandemic has been getting worse globally, and there is an urgent need for new easy-to-use diagnostic techniques for objective depression screening. Deep learning has been used to automate the process of identifying emotions, but the sensitivity and specificity needed to capture the important micro-expressions are not often addressed. Most models fail to capture the nuanced micro-dynamics that can differentiate true depressed affect from social masking. This paper presents Cross-Modal Contextual Attention Transformer (CM-CAT) for enhanced mental health prediction that integrates sentiment in text and facial expression temporal dynamics. The architecture is based on a dual-stream long-range cross attention (visual and linguistic) capture model. The proposed transformer architecture model is used with DepVidMood for validation. CM-CAT model is further enhanced with an Explainable AI (XAI) which allows non-expert users to gain insight into the model and its prediction. The CM-CAT model framework achieved the highest accuracy of 92.4% and F1-score of 0.89 across 4 severity classes of mental health (Healthy, Mild, Moderate and Severe). The CM-CAT model was also able to surface and focus temporal saliency around the clinically significant periorbital and perioral areas salient consistent with established biomarkers. This study enhanced CM-CAT model has a clinically interpretable design that supports the earliest intervention for mental health concerns and enables objective psychiatric evaluations.
Objectives: The expanding global mental health burden has intensified the critical need for objective, scalable diagnostic screening tools for clinical depression. While conventional deep learning models can automate basic affect recognition, they frequently fail to capture the highly subtle facial micro-dynamics that differentiate genuinely depressed affect from superficial social masking. This study presents a novel, robust multimodal framework designed to solve cross-modal temporal dependencies and enhance prediction reliability.
Methods: We present the Cross-Modal Contextual Attention Transformer (CM-CAT), which integrates contextual text semantics with detailed facial temporal dynamics through a distinct dual-stream architecture. The pipeline merges a spatio-temporal visual stream built on facial Action Units and a Temporal Convolutional Network with a contextual textual stream utilizing RoBERTa embeddings. Evaluation was executed on the DepVidMood clinical video corpus across four severity tiers, utilizing a person-independent 70%/15%/15% train/validation/test split.
Results: On the held-out test partition, CM-CAT achieves a high 92.4% validation accuracy and a macro-averaged F1-score of 0.89, outperforming contemporary multimodal baseline models. Furthermore, an Explainable AI (XAI) module based on integrated gradients demonstrates that attention accurately concentrates on the periorbital and perioral zones associated with established physiological markers of depressive affect.
Conclusions: CM-CAT introduces a highly interpretable, reliable, and clinically viable digital biomarker tool that successfully advances objective mental health tracking methodologies inside resource-constrained psychiatric informatics networks.





