To cite this paper use one of the standards below:
Urinary tract infections (UTIs) are a common complication in patients with compromised immune systems and can significantly affect their quality of life. Previous epidemiological studies have identified several clinical risk factors associated with increased infection probability; however, the combined or synergistic effect among these variables remains poorly understood. In this work, we propose a machine learning framework based on mutual information to capture the interaction between clinical risk factors and the probability of developing UTIs in patients with systemic lupus erythematosus (SLE).
The outcome variable corresponds to the presence or absence of a urinary tract infection. Clinical and infectious variables are collected from a dataset specifically constructed for this study, which includes 269 baseline patients and 149 variables, of which 25 correspond to clinical predictors. To identify the most relevant predictors, a feature selection stage based on Symmetrical Uncertainty is applied. This information-theoretic metric quantifies the dependency between each candidate predictor and the outcome, allowing the ranking of variables according to their predictive relevance. The five variables with the highest scores are selected as inputs for the predictive models.
Following the variable selection stage, several machine learning classifiers are trained to estimate the probability of infection. The evaluated models include Logistic Regression, Random Forest, Support Vector Machines, and Extreme Gradient Boosting. To ensure a reliable evaluation, a grouped cross-validation strategy is implemented so that all observations belonging to the same patient remain within the same fold, preventing information leakage between training and validation sets.
The dataset presents a strong class imbalance, with approximately 6% of observations corresponding to positive infection cases. To address this issue, the Synthetic Minority Oversampling Technique (SMOTE) is applied during the training phase to balance the minority class. Model performance is assessed using multiple evaluation metrics, including the area under the ROC curve (AUC), sensitivity, specificity, and the geometric mean, while confidence intervals for AUC are estimated using bootstrap resampling.
Experimental results demonstrate high predictive performance across all evaluated models, with AUC values consistently above 0.95. Sensitivity and specificity show stable behavior, indicating a balanced capacity to correctly identify both infected and non-infected patients. Model interpretability analysis using SHAP reveals that the most influential predictors include previous UTI episodes, patient age, glucocorticoid therapy, antibiotic exposure, and low fluid intake. These findings are consistent with previously reported epidemiological evidence and highlight the clinical relevance of the proposed machine learning framework for identifying patients at risk of urinary tract infection. Future work will focus on extending the longitudinal analysis to improve model calibration and predictive robustness.
With nearly 200,000 papers published, Galoá empowers scholars to share and discover cutting-edge research through our streamlined and accessible academic publishing platform.
Learn more about our products:
This proceedings is identified by a DOI , for use in citations or bibliographic references. Attention: this is not a DOI for the paper and as such cannot be used in Lattes to identify a particular work.
Check the link "How to cite" in the paper's page, to see how to properly cite the paper