Preprint / Version 1

The Indeterminacy of Authorship: Sociodemographic Modeling and the Limits of Artificial Intelligence in Predicting ENEM Essay Scores in the State of Rio de Janeiro

##article.authors##

  • Francisco Alves de Freitas Neto IFF / UENF https://orcid.org/0000-0001-8359-004X
    • Software
    • Conceptualization
    • Data Curation
    • Formal Analysis
    • Funding Acquisition
    • Investigation
    • Methodology
    • Validation
    • Writing – Original Draft Preparation
    • Writing – Review & Editing
  • Carlos Henrique Medeiro de Souza State University of Norte Fluminense image/svg+xml https://orcid.org/0000-0002-3774-0323
    • Writing – Review & Editing
    • Writing – Original Draft Preparation
    • Visualization
    • Validation
    • Supervision
    • Project Administration
    • Methodology
    • Investigation
  • Fabio Machado de Oliveira State University of Norte Fluminense image/svg+xml https://orcid.org/0000-0003-1336-2994
    • Writing – Review & Editing
    • Writing – Original Draft Preparation
    • Visualization
    • Validation
    • Supervision
    • Software
    • Project Administration
  • Leonard Barreto Moreira Fluminense Federal University image/svg+xml https://orcid.org/0000-0002-5588-2923
    • Writing – Review & Editing
    • Writing – Original Draft Preparation
    • Visualization
    • Validation
    • Software
    • Resources
    • Project Administration
    • Methodology

DOI:

https://doi.org/10.1590/SciELOPreprints.17449

Keywords:

Artificial Intelligence, Discourse Analysis, Fairness AI

Abstract

This work investigates the theoretical and operational limits of Artificial Intelligence in predicting discursive performance in the assessment of ENEM essay scores, integrating Artificial Intelligence Systems Modeling and Discourse Analysis. The central hypothesis postulates that predictive models fed by sociodemographic proxies are ineffective in fully mapping structural literacy, particularly within spaces of median performance—a category where polyphonic agency and authorship override material determinism. The methodological approach consisted of developing Machine Learning models to classify subjects into three analytical strata ("Below", "Above", and "Median"). Building upon a Baseline model structured on Random Forest, an optimized Multilayer Perceptron (MLP) deep neural network was implemented. To ensure evaluation rigor, the training dataset was synthetically equalized across the three classes using the SMOTE-NC algorithm. Given the initial algorithmic failure to identify the median stratum in real-world data, Fairness AI techniques were applied, initially through asymmetric error penalization via Cost-Sensitive Learning, and subsequently through a post-processing adjustment using Threshold Moving. Empirical results demonstrated that the model maintained strong deterministic traction at the socioeconomic extremes, yet suffered from a lack of accuracy in the central class. Ethical calibrations merely resulted in the mathematical redistribution of the error and an increase in entropy, limiting the overall accuracy to a ceiling of approximately 54%. It is concluded that median discursive performance possesses non-linear properties that resist matricial reduction. The machine’s predictive stagnation provides empirical evidence that human authorship, when subjected to conditions of material equality, presents itself as computationally difficult to predict.

Downloads

Download data is not yet available.

Submitted

08/17/2026

Posted

08/18/2026

How to Cite

The Indeterminacy of Authorship: Sociodemographic Modeling and the Limits of Artificial Intelligence in Predicting ENEM Essay Scores in the State of Rio de Janeiro. (2026). In SciELO Preprints. https://doi.org/10.1590/SciELOPreprints.17449

Section

Applied Social Sciences

Plaudit

Data statement