Prediction of Cervical Cytopathological Changes Using Artificial Intelligence in Primary Health Care
Nursing; Cervical neoplasms; Machine learning; Artificial intelligence; Screening; Primary health care.
Cervical cancer remains a significant public health challenge in Brazil,
primarily due to the low effectiveness of opportunistic cytopathological screening and
regional disparities in access. The transition to molecular HPV DNA testing as
a primary strategy, combined with the incorporation of Artificial Intelligence technologies,
represents a promising strategy for improving early detection and risk stratification.
Thus, this study aims to develop a machine learning model
for predicting cervical cytopathological abnormalities based on
sociodemographic, clinical, and reproductive data. This is a cross-sectional, diagnostic prediction
study with a quantitative approach, conducted in primary care units
in four municipalities in the state of Rio Grande do Norte, with data collection carried out between June and
December 2025. The sample comprised 483 women aged 25 to 64 years who underwent a
structured interview and standardized cytopathological examination. Descriptive,
bivariate (chi-square/Fisher), and multivariate logistic regression analyses were performed. For predictive modeling,
the data were partitioned into training (70%) and testing (30%) sets with stratification. Feature selection
employed the Boruta algorithm, and five classifiers were trained and compared:
Random Forest, XGBoost, LightGBM, CatBoost, and TabPFN, with Bayesian optimization of
hyperparameters via Optuna and stratified cross-validation (k=5). The interpretability of the
selected model was assessed using the SHAP method. The study was approved by the
Research Ethics Committee under opinion no. 7.296.33. Of the 483 samples evaluated, 99.17% were
satisfactory, with an overall prevalence of cellular atypia in 14.91% (n=72) of the samples and
a predominance of ASC-US in 9.52%. The following were identified as independent predictors of
cytopathological abnormalities: treatment for vaginal infection in the past six months (adjusted OR=2.39;
95% CI 1.31–4.33; p=0.005), parity (Adj. OR=3.20; p=0.032), absence of prior treatment
for HPV (Adj. OR=2.67; p=0.047), current smoking (Adj. OR=2.67; p=0.017), and age up to 42 years
(Adj. PR=1.92; p=0.019). Boruta confirmed five predictors (smoking, recent treatment for
vaginal infection, income up to one minimum wage, parity, and vaginal delivery). The five models
showed accuracy around 0.841, with TabPFN reaching 0.848. CatBoost achieved
the highest discriminative power (AUC-ROC=0.634), followed by Random Forest (0.618) and
XGBoost (0.617). The SHAP analysis of CatBoost highlighted a greater predictive contribution from
treatment for recent vaginal infection (≈0.180), median income (≈0.140), and current smoking
(≈0.105), partially corroborating Boruta’s selection. The developed model demonstrated
moderate discriminative performance. The convergence among the predictors identified by
logistic regression, Boruta, and SHAP reinforces the epidemiological consistency of the findings
and signals the potential of machine learning models as a complementary tool
for risk stratification in reflex cytology, requiring sample expansion and the incorporation
of additional variables for future refinement.