Thematic modeling of digital literacy using language models: a qualitative-computational analysis of interviews on hearing health
DOI:
https://doi.org/10.1590/SciELOPreprints.17666Keywords:
Speech and Language Pathology and Audiology, Health Literacy, Natural Language Processing, Generative Artificial Intelligence, Qualitative ResearchAbstract
Objective: To describe thematic patterns of digital health literacy among adults with hearing loss and to evaluate the feasibility of using small-scale language models for the qualitative categorization of clinical interviews. Method: This is a cross-sectional, qualicomputational methodological proof-of-concept study. Transcripts of 15 semi-structured interviews with adults diagnosed with hearing loss were analyzed. A model based on MPNet (Masked and Permuted Pre-training for Language Understanding) was employed. Topic extraction was performed using the BERTopic method in zero-shot mode (automatic classification without prior training on the categories). The responses were organized into eight predefined thematic categories, derived directly from the domains addressed in the interview script. Results: The computational approach demonstrated feasibility in the extraction and structured clustering of the responses provided by the participants. The findings indicate a predominantly reactive trigger for digital literacy: information seeking begins mainly after diagnosis, with behaviors ranging from incidental to proactive and technologically proficient. Conclusion: The data indicate that family mediation contributes to these patients’ access to digital technologies and to the filtering of information. Audiological rehabilitation may benefit from training the family support network to also act as mediators of digital literacy. The findings suggest that small-scale language models constitute a feasible methodological tool to support qualitative research in speech-language pathology.
Downloads
Submitted
Posted
How to Cite
Section
Copyright (c) 2026 Hector Gabriel Corrale de Matos, Sammia Klann Vieira, Kátia de Freitas Alvarenga, Ana Paula Berberian, Lilian Cássia Bórnia Jacob

This work is licensed under a Creative Commons Attribution 4.0 International License.
Funding data
-
Fundação de Amparo à Pesquisa do Estado de São Paulo
Grant numbers 2024/05572-7
Plaudit
Data statement
-
The research data is available in one or more data repository(ies)
-
The research data is contained in the manuscript
-
The research data is available on demand, condition justified in the manuscript


