Preprint / Version 2

The Performance of Large Language Models in Synthesising Educational Research Reports: A Comparative Study of Articles in Portuguese

##article.authors##

DOI:

https://doi.org/10.1590/SciELOPreprints.17611

Keywords:

Generative Artificial Intelligence, Scientific Abstract Processing, Algorithmic Biases

Abstract

This study compares the performance of ChatGPT 5.0, Gemini 3.0 Flash, and DeepSeek V3 in producing abstracts of scientific articles published in Portuguese. A mixed-methods approach was adopted, involving the submission of five articles in the field of Education to the three language models using three prompt variations. The outputs were evaluated by 15 specialists, with an ordinal perception of quality (PQ) scale to assess the dimensions of content quality and freedom from bias. The results indicated that the models performed well, with PQ medians ranging from 84% to 94% of the total score. The Kruskal–Wallis test identified statistically significant differences when the results were compared by article, H(4)=24.040,p<0.001, but found no significant differences among the models, H(2)=1.604,p=0.448, or among the prompt types, H(8) = 2,506, p = 0,961. These findings suggest a technological convergence among the models in the task of scientific text synthesis. However, the qualitative analysis identified  generalizations’ tendency and the suppression of uncertainty markers, indicating shortcomings in the validity of the generated abstracts. The findings also suggest the possibility of biases arising from the languages represented in the models’ training data, which could hypothetically affect their ability to capture the narrative texture and specialized terminology of scientific knowledge produced in Portuguese. It was concluded that, although the generative AI models tested demonstrated good syntactic performance, human oversight remains indispensable for ensuring the accuracy and completeness of scientific communication in educational research.

Downloads

Download data is not yet available.

Submitted

08/23/2026

Posted

09/17/2026 — Updated on 09/28/2026

Versions

How to Cite

The Performance of Large Language Models in Synthesising Educational Research Reports: A Comparative Study of Articles in Portuguese. (2026). In SciELO Preprints. https://doi.org/10.1590/SciELOPreprints.17611 (Original work published 2026)

Section

Human Sciences

Plaudit

Version justification

Esta versão foi atualizada em 26/9/2026 para correção de erros na digitação das estatísticas dos testes de Kruskal-Wallis destinados à comparação de resultados por modelo de IAG e tipo de prompt. Também foi inserida, no Quadro 1, a íntegra do conteúdo dos prompts utilizados e foi ajustado o tópico dos resultados referentes à análise baseada nos tipos de prompt, pois os gráficos das figuras 2 e 3 continham imprecisões. Eles foram substituídos por uma tabela. Os ajustes não invalidam a versão anterior no que se refere às conclusões ou as discussões decorrentes da interpretação dos resultados.

Data statement