Preprint / Version 1

Brazilian Linguistic Diversity Platform: a repository of linguistic data for sovereign artificial intelligence

##article.authors##

  • Marcia dos Santos Machado Vieira Federal University of Rio de Janeiro image/svg+xml https://orcid.org/0000-0002-2320-5055
    • Conceptualization
    • Data Curation
    • Formal Analysis
    • Project Administration
    • Writing – Original Draft Preparation
    • Writing – Review & Editing
  • Raquel Meister Ko Freitag Federal University of Sergipe image/svg+xml https://orcid.org/0000-0002-4972-4320
    • Conceptualization
    • Data Curation
    • Formal Analysis
    • Project Administration
    • Writing – Original Draft Preparation
    • Writing – Review & Editing
  • Juliana Bertucci Barbosa Federal University of Triângulo Mineiro image/svg+xml https://orcid.org/0000-0002-1510-633X
    • Conceptualization
    • Data Curation
    • Formal Analysis
    • Project Administration
    • Writing – Original Draft Preparation
    • Writing – Review & Editing

DOI:

https://doi.org/10.1590/SciELOPreprints.18232

Keywords:

Digital repositories, Data curation, Linguistic diversity, Open science, Artificial intelligence

Abstract

Objective: To present the proposal of the Brazilian Linguistic Diversity Platform, an initiative of the Brazilian Association of Linguistics (ABRALIN) and the Sociolinguistics Working Group of the National Association of Postgraduate Studies and Research in Letters and Linguistics (ANPOLL), and to discuss its contribution to the development of Portuguese language models foreseen in the Brazilian Artificial Intelligence Plan (PBIA) 2024-2028.

Methods: Theoretical and propositional essay, based on documentary analysis of the PBIA and on studies of linguistic bias in language models and of the management of sociolinguistic data collections in Brazil.

Results: The language models currently available are trained on data dominated by English, in which Portuguese appears mostly in written, formal, and urban registers, which results in the reproduction of biases, stereotypes, and the invisibility of underrepresented varieties and speakers; meanwhile, linguistic data collections produced by specialists remain fragmented and regionally and institutionally concentrated. The Platform is proposed as a national repository of linguistic data collections, guided by the FAIR principles and supported by a collaborative network of laboratories.

Conclusions: The Platform represents a strategic bridge between linguistic science, data science, information science, and technological development, with impacts on education, language policies, and Brazilian digital sovereignty.

Downloads

Download data is not yet available.

Submitted

09/30/2026

Posted

10/02/2026

How to Cite

Brazilian Linguistic Diversity Platform: a repository of linguistic data for sovereign artificial intelligence. (2026). In SciELO Preprints. https://doi.org/10.1590/SciELOPreprints.18232

Section

Linguistic, literature and arts

Plaudit

Data statement

  • The research data is contained in the manuscript