2026
GPT outperforms BERT and LIWC for stance and anger detection in German news articles and user comments
Read abstract
We live in an era of abundant qualitative language data that allows insights into societal trends and public attitudes, such as social media posts and news articles. Natural language processing (NLP) has become essential for analyzing this data, with recent advances in Large Language Models (LLMs) driving rapid progress. These models show strong potential for psychological research, though their validity in detecting attitudes varies across domains. This study evaluates GPT-3.5 and GPT-4 Turbo for identifying stances and expressions of anger in texts about climate activism, comparing them with traditional methods. The dataset included news articles and user comments from eight German outlets. A subset of 320 articles and 330 comments was manually annotated by three human raters for (1) stance toward Fridays for Future and (2) anger expression. GPT models outperformed BERT (Bidirectional Encoder Representations from Transformers) in stance detection and LIWC (Linguistic Inquiry and Word Count) in anger detection. Moreover, GPT achieved higher agreement with the human majority vote than individual raters agreed with each other. GPT-4 Turbo performed slightly better than GPT-3.5, though differences were minor and domain-specific. Overall, GPT models proved reliable and efficient for annotating nuanced German texts where conventional NLP tools often struggle.