A train-once, evaluate-twice protocol for cross-lingual depression detection with large language models
Depression affects more than 280 million people worldwide, and social media text provides a large-scale resource for automated mental health assessment. Large language model (LLM)-based approaches have shown promising performance for social media-based depression detection; however, their cross-lingual robustness for underrepresented languages such as Arabic remains insufficiently explored. We proposed the “train once, evaluate twice” protocol to evaluate the cross-lingual robustness of eight LLMs for detecting depression in both English and Arabic.
For six open-source LLMs, we used QLoRA-based parameter-efficient fine-tuning (PEFT) to adapt them; for two closed-source GPT models, we fine-tuned them with the OpenAI supervised fine-tuning (SFT) API. To provide context for understanding transfer performance, we also evaluated a translation-based pipeline (MarianMT + XLM-R) and a bilingual fine-tuning ablation using data-mixing techniques. The models performed nearly at ceiling in their monolingual evaluation (English), with the best Macro-F1 score of 0.990 in GPT-4.1 Mini, and moderately well in Arabic (best Macro-F1 of 0.800 in Qwen2).
Cross-lingual evaluation revealed significant directional asymmetry. The strongest transfer performance was observed in AR → EN (Macro-F1 = 0.793 for GPT-4.1 mini), whereas EN → AR remained considerably more challenging (best Macro-F1 = 0.638 for GPT-3.5 Turbo). The MarianMT + XLM-R translate-then-classify baseline achieved consistently lower cross-lingual performance than the fine-tuned generative LLMs evaluated in this study.
The results of a bilingual fine-tuning ablation experiment conducted using Qwen2-7B-Instruct suggested that bilingual data mixing is sufficient to preserve English performance at near-ceiling levels. However, bilingual data mixing did not improve Arabic performance relative to monolingual Arabic fine-tuning. Thus, for this family of models, simple bilingual exposure appears insufficient to address the cross-lingual performance gap.
Error analysis showed that EN→AR transfer predominantly increased false negatives, whereas AR → EN transfer exhibited more balanced precision–recall degradation profiles. Notably, the two Arabic-specialized models evaluated did not consistently outperform the strongest general-purpose multilingual models under cross-lingual transfer settings. PEFT produced compact adapters with low local inference latency, whereas API-based GPT inference exhibited substantially higher end-to-end latency.
These findings support the notion that standard single-language fine-tuning does not afford much robustness when models are deployed across languages, especially in the EN → AR direction; thus, providing impetus for further research into new methods for adapting LLMs across multiple languages. Trial registration: Not applicable. This work was supported by the College of Science and Engineering, Hamad Bin Khalifa University, Doha, Qatar.
Extract — continue reading at the source.