Advanced Architectures for AI Self-Training:State-of-the-Art Methodologies and Performance Simulation
DOI:
https://doi.org/10.1590/SciELOPreprints.17305Keywords:
Artificial Intelligence,, Self-training,, RLAIF,, Self-Instruct,Resumen
We develop a parametric simulation framework that models the training dynamics of two self-training architectures: a classical pseudo-labeling baseline and a state-
of-the-art (SOTA) pipeline combining RLAIF (Reinforcement Learning from AI Feedback), Self-
Instruct generation, and multi-agent consensus filtering. Each observable of interest —
validation accuracy, loss decay, per-domain F1-Score, and curation latency — is expressed
as a closed-form function of the training epoch or batch size, and the free parameters
are calibrated to reflect the qualitative behaviour reported for each architecture. By
numerically integrating these models we quantify the convergence ceiling, optimisation
stability, domain-wise accuracy, and throughput scaling of the two pipelines. The simulation
shows that consensus-filtered curation raises the accuracy ceiling by roughly seven points,
suppresses the residual loss by an order of magnitude, and keeps curation latency sub-linear
in batch size. This work is a modeling and simulation study; all curves are generated from
the stated equations rather than measured on a physical training run.
Downloads
Enviado
Postado
Cómo citar
Serie
Derechos de autor 2026 Dheiver Francisco Santos

Esta obra está bajo una licencia internacional Creative Commons Atribución 4.0.
Plaudit
Declaración de datos
-
Los datos de investigación están incluidos en el propio manuscrito


