This is an outdated version published on 08/11/2026. Read the most recent version.
Preprint / Version 1

Do LLMs Understand Idioms? Evidence for a Theory of Simulated Artificial Phraseological Competence

##article.authors##

DOI:

https://doi.org/10.1590/SciELOPreprints.16937

Keywords:

phraseology, phraseological competence, large language models

Abstract

This study examines, by means of a quasi-controlled experiment, the performance of the Claude 4.5 Haiku large language model in processing Brazilian Portuguese idiomatic expressions. Grounded in the intersection of Phraseology and Cognitive Linguistics, the study proposes the concept of simulated artificial phraseological competence, defined as a functional performance based on statistical-distributional regularities, devoid of the embodied experience, sociocultural immersion, and pragmatic inference characteristic of human language users. The experimental corpus, consisting of 140 phraseological units (including opaque, somatic, and cultural idioms, alongside experimental control samples), was evaluated across three prompting conditions—zero-shot, contextualized few-shot, and Chain-of-Thought with role-playing. Four dimensions were assessed: idiomaticity detection, semantic precision, pragmatic appropriateness, and cultural sensitivity. The findings demonstrate high semantic accuracy for conventionalized units, yet reveal asymmetrical performance across the pragmatic and cultural dimensions, alongside a pronounced tendency toward figurative hallucination when encountering non-existent idioms. The Chain-of-Thought condition yielded consistent qualitative improvements, reducing false positives. These results reinforce the hypothesis that the analyzed LLM lacks mechanisms equivalent to human phraseological competence, although evidence suggests that prompt engineering strategies can partially mitigate these architectural limitations.

Downloads

Download data is not yet available.

Author Biography

Thyago Jose da Cruz, Federal University of Mato Grosso do Sul

Tem experiência na área de Letras e Educação do Campo, com ênfase em Língua Portuguesa e Espanhola , atuando principalmente nos seguintes temas: Língua Portuguesa e sua Prática de Ensino, Língua Espanhola e sua Prática de Ensino e Educação do Campo. É doutor em Letras, mestre em Estudos de Linguagens e desenvolve pesquisas no âmbito da Fraseologia, Fraseografia, da Semântica e da Linguística Aplicada ao Ensino. Atualmente, exerce a função de professor, lotado na Faculdade de Educação (FAED) da Universidade Federal de Mato Grosso do Sul (UFMS), nos cursos de Licenciatura em Educação do Campo ([área de Linguagens e Códigos) e de Letras (PRILEI). Além disso, é Professor Permanente do Programa de Pós-graduação stricto sensu em Letras (Unidade Universitária de Campo Grande/UEMS)

Submitted

07/16/2026

Posted

08/11/2026

Versions

How to Cite

Do LLMs Understand Idioms? Evidence for a Theory of Simulated Artificial Phraseological Competence. (2026). In SciELO Preprints. https://doi.org/10.1590/SciELOPreprints.16937

Section

Linguistic, literature and arts

Plaudit

Data statement

  • The research data is available on demand, condition justified in the manuscript