FINE-TUNING SMALL LANGUAGE MODELS FOR MENTAL HEALTH TEXT CLASSIFICATION: A PARAMETER-EFFICIENT APPROACH OUTPERFORMING ZERO-SHOT LARGE LANGUAGE MODELS

Authors

DOI:

https://doi.org/10.37943/XSAJ1029

Keywords:

mental health classification , small language models , parameter-efficient fine-tuning , LoRA , text analysis , stress detection , depression detection , zero-shot learning

Abstract

Automated analysis of user-generated text is increasingly used to support scalable mental-health assessment, but deploying large language models for classification remains costly and latency-sensitive. This study evaluates whether parameter-efficient fine-tuning (PEFT) of a small model can provide a better accuracy–efficiency trade-off. We fine-tune SmolLM2-1.7B on two benchmarks: SWMH (5-class mental-health classification) and Dreaddit (binary stress detection). Evaluation follows two protocols: (1) primary full held-out test evaluation for the fine-tuned model, and (2) paired evaluation on identical 200-example subsets for fair cross-model comparison with zero-shot GPT-4o-mini and GPT-4o. On full held-out tests, SmolLM2-1.7B achieves 0.723 accuracy, 0.720 weighted F1, and 0.727 macro F1 on SWMH, and 0.808 accuracy, 0.808 weighted F1, and 0.808 macro F1 on Dreaddit. On paired subsets, SmolLM2-1.7B outperforms both zero-shot GPT baselines on both datasets (SWMH accuracy: 0.705 vs 0.600/0.645; Dreaddit accuracy: 0.755 vs 0.650/0.675). These results indicate that task-specific PEFT can make small models practically deployable and, under controlled paired evaluation, stronger than zero-shot larger models.

References

World Health Organization. (2022). World mental health report: Transforming mental health for all. WHO Press.

Gulliver, A., Griffiths, K. M., & Christensen, H. (2010). Perceived barriers and facilitators to mental health help-seeking in young people: A systematic review. BMC Psychiatry, 10(1), 113. https://doi.org/10.1186/1471-244X-10-113

Chancellor, S., & De Choudhury, M. (2020). Methods in predictive techniques for mental health status on social media: A critical review. NPJ Digital Medicine, 3(1), 43. https://doi.org/10.1038/s41746-020-0233-7

OpenAI. (2023). GPT-4 technical report. arXiv preprint arXiv:2303.08774. https://doi.org/10.48550/arXiv.2303.08774

Bender, E. M., Gebru, T., McMillan-Major, A., & Shmitchell, S. (2021). On the dangers of stochastic parrots: Can language models be too big? Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency, 610-623. https://doi.org/10.1145/3442188.3445922

Schick, T., & Schütze, H. (2021). It’s not just size that matters: Small language models are also few-shot learners. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics, 2339-2352. https://doi.org/10.18653/v1/2021.naacl-main.185

Allal, L. B., Lozhkov, A., Penedo, G., Wolf, T., & von Werra, L. (2024). SmolLM2: A family of compact language models. Hugging Face Blog. https://huggingface.co/blog/smollm2

Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., … & Chen, W. (2021). LoRA: Low-rank adaptation of large language models. arXiv preprint arXiv:2106.09685. https://doi.org/10.48550/arXiv.2106.09685

Yang, K., Ji, S., Zhang, T., Xie, Q., Ananiadou, S., & Bian, J. (2022). Towards objective-guided semantic understanding for early mental health intervention. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 3949-3962. https://doi.org/10.18653/v1/2022.emnlp-main.261

Losada, D. E., Crestani, F., & Parapar, J. (2020). Overview of eRisk 2020: Early risk prediction on the internet. Experimental IR Meets Multilinguality, Multimodality, and Interaction, 12260, 272-287. https://doi.org/10.1007/978-3-030-58219-7_20

De Choudhury, M., Gamon, M., Counts, S., & Horvitz, E. (2013). Predicting depression via social media. Proceedings of the 7th International AAAI Conference on Weblogs and Social Media, 128-137.

Lin, C., Hu, P., Su, H., Li, S., Mei, J., Zhou, J., & Leung, H. (2020). SenseMood: Depression detection on social media. Proceedings of the 2020 International Conference on Multimedia Retrieval, 407-411. https://doi.org/10.1145/3372278.3390670

Ji, S., Yu, C. P., Fung, S. F., Pan, S., & Long, G. (2020). Supervised learning for suicidal ideation detection in online user content. Complexity, 2020, 6157249. https://doi.org/10.1155/2020/6157249

Gaur, M., Alambo, A., Sain, J. P., Kursuncu, U., Thirunarayan, K., Kavuluru, R., … & Sheth, A. (2019). Knowledge-aware assessment of severity of suicide risk for early intervention. The World Wide Web Conference, 514-525. https://doi.org/10.1145/3308558.3313698

Turcan, E., & McKeown, K. (2019). Dreaddit: A Reddit dataset for stress analysis in social media. Proceedings of the Tenth International Workshop on Health Text Mining and Information Analysis, 97-107. https://doi.org/10.18653/v1/D19-6212

Xu, L., Zhou, X., Tao, C., Liu, H., & Xu, H. (2024). Can ChatGPT answer mental health questions? An evaluation study. JMIR Mental Health, 11, e52758. https://doi.org/10.2196/52758

Wei, J., Bosma, M., Zhao, V. Y., Guu, K., Yu, A. W., Lester, B., … & Le, Q. V. (2022). Finetuned language models are zero-shot learners. International Conference on Learning Representations. https://openreview.net/forum?id=gEZrGCozdqR

Sclar, M., Choi, Y., Tsvetkov, Y., & Suhr, A. (2023). Quantifying language models’ sensitivity to spurious features in prompt design or: How I learned to start worrying about prompt formatting. arXiv preprint arXiv:2310.11324. https://doi.org/10.48550/arXiv.2310.11324

Nori, H., King, N., McKinney, S. M., Carignan, D., & Horvitz, E. (2023). Capabilities of GPT-4 on medical challenge problems. arXiv preprint arXiv:2303.13375. https://doi.org/10.48550/arXiv.2303.13375

Ding, N., Qin, Y., Yang, G., Wei, F., Yang, Z., Su, Y., … & Liu, Z. (2023). Parameter-efficient fine-tuning of large-scale pre-trained language models. Nature Machine Intelligence, 5(3), 220-235. https://doi.org/10.1038/s42256-023-00626-4

Dettmers, T., Pagnoni, A., Holtzman, A., & Zettlemoyer, L. (2023). QLoRA: Efficient finetuning of quantized LLMs. Advances in Neural Information Processing Systems, 36, 10088-10115.

Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., … & Scialom, T. (2023). Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288. https://doi.org/10.48550/arXiv.2307.09288

Low, D. M., Rumker, L., Talkar, T., Torous, J., Cecchi, G., & Ghosh, S. S. (2020). Natural language processing reveals vulnerable mental health support groups and heightened health anxiety on Reddit during COVID-19: Observational study. Journal of Medical Internet Research, 22(10), e22635. https://doi.org/10.2196/22635

Liang, P., Bommasani, R., Lee, T., Tsipras, D., Soylu, D., Yasunaga, M., … & Leskovec, J. (2023). Holistic evaluation of language models. Transactions on Machine Learning Research. https://openreview.net/forum?id=iO4LZibEqW

Strubell, E., Ganesh, A., & McCallum, A. (2019). Energy and policy considerations for deep learning in NLP. Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, 3645-3650. https://doi.org/10.18653/v1/P19-1355

Gu, Y., Tinn, R., Cheng, H., Lucas, M., Usuyama, N., Liu, X., … & Poon, H. (2021). Domain-specific language model pretraining for biomedical natural language processing. ACM Transactions on Computing for Healthcare, 3(1), 1-23. https://doi.org/10.1145/3458754

Ji, Z., Lee, N., Frieske, R., Yu, T., Su, D., Xu, Y., … & Fung, P. (2023). Survey of hallucination in natural language generation. ACM Computing Surveys, 55(12), 1-38. https://doi.org/10.1145/3571730

Bitew, S. K., Deleu, J., Develder, C., & Demeester, T. (2022). Predicting fine-grained subreddit relevance for mental health queries. Proceedings of the Eighth Workshop on Computational Linguistics and Clinical Psychology, 133-143. https://doi.org/10.18653/v1/2022.clpsych-1.12

Matero, M., Idnani, A., Son, Y., Giorgi, S., Vu, H., Zamani, M., … & Schwartz, H. A. (2019). Suicide risk assessment with multi-level dual-context language and BERT. Proceedings of the Sixth Workshop on Computational Linguistics and Clinical Psychology, 39-44. https://doi.org/10.18653/v1/W19-3005

Patterson, D., Gonzalez, J., Le, Q., Liang, C., Munguia, L. M., Rothchild, D., … & Dean, J. (2021). Carbon emissions and large neural network training. arXiv preprint arXiv:2104.10350. https://doi.org/10.48550/arXiv.2104.10350

Chancellor, S., Birnbaum, M. L., Caine, E. D., Silenzio, V. M., & De Choudhury, M. (2019). A taxonomy of ethical tensions in inferring mental health states from social media. Proceedings of the Conference on Fairness, Accountability, and Transparency, 79-88. https://doi.org/10.1145/3287560.3287587

Zhou, Y., Muresanu, A. I., Han, Z., Paster, K., Pitis, S., Chan, H., & Ba, J. (2023). Large language models are human-level prompt engineers. International Conference on Learning Representations. https://openreview.net/forum?id=92gvk82DE-

Lin, T. Y., Goyal, P., Girshick, R., He, K., & Dollár, P. (2017). Focal loss for dense object detection. Proceedings of the IEEE International Conference on Computer Vision, 2980-2988. https://doi.org/10.1109/ICCV.2017.324

Manikonda, L., & De Choudhury, M. (2017). Modeling and understanding visual attributes of mental health disclosures in social media. Proceedings of the 2017 CHI Conference on Human Factors in Computing Systems, 170-181. https://doi.org/10.1145/3025453.3025932

Zhang, M. L., & Zhou, Z. H. (2014). A review on multi-label learning algorithms. IEEE Transactions on Knowledge and Data Engineering, 26(8), 1819-1837. https://doi.org/10.1109/TKDE.2013.39

Downloads

Published

2026-06-30

How to Cite

Kapizov, D., & Kumar, P. (2026). FINE-TUNING SMALL LANGUAGE MODELS FOR MENTAL HEALTH TEXT CLASSIFICATION: A PARAMETER-EFFICIENT APPROACH OUTPERFORMING ZERO-SHOT LARGE LANGUAGE MODELS. Scientific Journal of Astana IT University, 26(2), 121–136. https://doi.org/10.37943/XSAJ1029

Issue

Section

Information Technologies