PPgSC/UFRN PROGRAMA DE PÓS-GRADUAÇÃO EM SISTEMAS E COMPUTAÇÃO ADMINISTRAÇÃO DO CCET Téléphone/Extension: (84) 99919-3640 https://posgraduacao.ufrn.br/ppgsc

Banca de QUALIFICAÇÃO: JOSÉ EDIVANDRO DE SOUSA JUNIOR

Uma banca de QUALIFICAÇÃO de MESTRADO foi cadastrada pelo programa.
STUDENT : JOSÉ EDIVANDRO DE SOUSA JUNIOR
DATE: 31/07/2026
TIME: 10:00
LOCAL: Remoto
TITLE:

I-AUTOML – An iterative and adaptive approach to automatic optimization in standard classification problems


KEY WORDS:

AutoML, Supervised Classification, Adaptive Hyperparameter Tuning, Statistical Validation, Large Language Models.


PAGES: 56
BIG AREA: Ciências Exatas e da Terra
AREA: Ciência da Computação
SUBÁREA: Metodologia e Técnicas da Computação
SPECIALTY: Sistemas de Informação
SUMMARY:

The advancement of Automated Machine Learning (AutoML) tools has expanded access
to supervised classification techniques by reducing the need for manual configuration
of models and hyperparameters. However, most existing approaches focus primarily on
optimizing individual performance metrics and rely on fixed stopping criteria or rigid
heuristics during hyperparameter search, offering limited support for systematic statistical
comparison among models and for structured interpretation of results, especially when
dealing with heterogeneous datasets.
In this context, this work proposes and evaluates I-AutoM, an automated system for
supervised classification that integrates preprocessing, model training, adaptive hyperpa-
rameter tuning with early stopping, comparative statistical evaluation, and an interpreta-
bility layer based on Large Language Models (LLMs) within a unified functional pipeline.
The experimental methodology involved applying a catalog of eleven classification algo-
rithms to 30 real-world datasets collected from Kaggle, organized into three thematic
domains and encompassing diverse structural characteristics, class imbalance scenarios,
and multiple target variables. Three train-test split proportions (10%, 20%, and 30%)
were analyzed, with repeated executions to assess performance stability, alongside statis-
tical validation via the Friedman test and a targeted pairwise post-hoc test comparing
the system’s recommended model against the remaining algorithms in the catalog.
The results indicate that a 20% test split achieved the best average balance between
performance and stability across datasets, while the 30% scenario showed the largest
incremental gains from the adaptive tuning mechanism, suggesting that optimization be-
comes more impactful under more constrained training conditions. Furthermore, metrics
such as F1-score and ROC-AUC proved more informative than accuracy in imbalanced
scenarios. The integration of a large language model to support statistical interpretation
also contributed to translating quantitative metrics and hypothesis tests into explanations
accessible to non-expert users. The findings demonstrate the feasibility of I-AutoM as an automated, exploratory, and educational tool capable of consistently processing heterogeneous datasets while combi-
ning quantitative evaluation, rigorous statistical validation, and natural-language-assisted
interpretation.


COMMITTEE MEMBERS:
Interna - 1350250 - ANNE MAGALY DE PAULA CANUTO
Externo ao Programa - 1669545 - DANIEL SABINO AMORIM DE ARAUJO - UFRNExterno à Instituição - GISELE LOBO PAPPA - UFMG
Notícia cadastrada em: 06/07/2026 16:56
SIGAA | Superintendência de Tecnologia da Informação - (84) 3342 2210 | Copyright © 2006-2026 - UFRN - sigaa04-producao.info.ufrn.br.sigaa04-producao