A machine learning-based early in-hospital risk stratification model for severe mycoplasma pneumoniae pneumonia in children
Highlight box
Key findings
• An exploratory machine learning-based model was developed for early in-hospital severity risk stratification of pediatric severe mycoplasma pneumoniae pneumonia (SMPP).
• The model retained five routine variables: basophil absolute count, platelet count, atelectasis, pleural effusion, and extrapulmonary complications.
What is known and what is new?
• Early identification of children at high risk of SMPP remains clinically important.
• This study provides a structured risk stratification framework based on routine laboratory, imaging, and complication-related information, but the findings require cautious interpretation.
What is the implication, and what should change now?
• The model may serve as an auxiliary tool to support early in-hospital assessment, rather than replace physician judgment.
• External validation and prospective clinical utility studies are needed before clinical application.
Introduction
Mycoplasma pneumoniae pneumonia (MPP) is one of the common causes of community-acquired pneumonia in children (1). Most children with MPP have a mild disease course and a favorable prognosis after standard treatment. However, some cases may progress to severe mycoplasma pneumoniae pneumonia (SMPP), which is characterized by aggravated pulmonary lesions, pleural effusion, and an increased incidence of extrapulmonary complications (EXCP), including encephalitis, myocarditis, and hemolytic anemia (2,3). In addition, with the increase in drug resistance and the inappropriate use of antibiotics, some patients may develop serious complications such as acute respiratory distress syndrome and may even require extracorporeal membrane oxygenation support (4).
In recent years, MPP infection in children has risen again markedly in the post-pandemic era. Relevant data showed that the number of MPP cases increased significantly after 2023, with the epidemic peak mainly occurring from October to December (5). Studies from some regions reported that the proportion of SMPP could reach 23.7% (6). These findings suggest that SMPP has become a disease burden that cannot be ignored in the clinical management of pediatric MPP. At present, the assessment of disease severity in pediatric MPP mainly relies on a comprehensive evaluation of clinical symptoms, imaging findings, and laboratory results. However, the progression from non-severe MPP to SMPP is influenced by multiple factors, and the definition and diagnostic criteria of SMPP have not yet been fully standardized (7). Therefore, establishing a comprehensive early prediction model based on multidimensional clinical features is of great importance for early risk stratification, timely intervention, and reduction of severe outcomes in children with MPP.
Previous risk assessment models for SMPP have included different variables, but their main focus has generally been consistent. Most of them concentrate on inflammatory burden, the extent of lung involvement, and the overall complexity of systemic illness. For example, an SMPP risk model based on admission laboratory indicators included erythrocyte sedimentation rate (ESR) and other markers reflecting inflammatory status (8). Another nomogram model for pediatric SMPP included respiratory rate, decreased breath sounds, duration of fever, co-infection with other pathogens, ferritin, and lactate dehydrogenase (LDH) (9). A dynamic nomogram further used C-reactive protein (CRP), neutrophil-to-lymphocyte ratio (NLR), D-dimer, alanine aminotransferase (ALT), LDH, and the extent of computed tomography (CT) involvement for risk identification (10). Therefore, although basophil absolute count (BASO), platelet count (PLT), atelectasis, pleural effusion, and EXCP differ from the variables commonly included in previous studies, the information reflected by these variables is still in line with the main direction of existing research. Specifically, they correspond to immune-inflammatory status, the extent of pulmonary involvement, and the complexity of systemic disease.
In recent years, machine learning has shown good potential in clinical risk assessment studies. It can integrate multivariable information and identify potential nonlinear relationships, thereby improving predictive performance. However, existing studies have mainly focused on general severity assessment or the construction of a single model, and studies specifically aiming at early prediction of SMPP in children remain limited. In the present study, we retrospectively collected the clinical data of hospitalized children with MPP and developed an early prediction model for SMPP using the disease status during hospitalization as the study outcome. Through feature selection, we finally identified five key variables, namely BASO, PLT, atelectasis, pleural effusion, and EXCP. We then compared the predictive performance of different models and performed interpretable analysis using SHapley Additive exPlanations (SHAP). Our aim was to provide supportive evidence for early risk prediction and stratified management of pediatric SMPP. We present this article in accordance with the TRIPOD reporting checklist (available at https://tp.amegroups.com/article/view/10.21037/tp-2026-0333/rc).
Methods
Study design and participants
This was a single-center retrospective study. We retrospectively collected the clinical data of children with MPP who were admitted to Chuzhou Hospital Affiliated to Anhui Medical University between February 2022 and July 2023. No additional post-discharge follow-up was included because the predicted outcome was SMPP status assessed during the index hospitalization. All information was obtained from the hospital electronic medical record system, including demographic characteristics, clinical symptoms and signs, imaging findings, and laboratory test results. The inclusion criteria were as follows: (I) age <18 years; (II) clinical diagnosis of MPP during hospitalization; (III) presence of lower respiratory tract infection-related clinical manifestations and chest imaging findings suggestive of pneumonia; (IV) pathogen-related evidence supporting mycoplasma pneumoniae infection; (V) clear grouping information as severe or non-severe disease; and (VI) availability of core clinical data required for model construction.
The exclusion criteria were as follows: (I) missing key data that made severity assessment or model construction impossible; (II) concomitant major respiratory diseases, hematologic diseases, immune-related diseases, or other severe underlying diseases that might markedly affect laboratory indicators or disease assessment; (III) presence of another clearly identified pneumonia pathogen as the main cause, with insufficient evidence for MPP; (IV) repeated hospitalization of the same child, in which case only the first eligible admission record was included; and (V) insufficient medical record information for reliable retrospective analysis. Finally, 98 children with MPP were included, including 49 cases in the SMPP group and 49 cases in the non-severe MPP group. SMPP status during hospitalization was used as the study outcome, and the model was developed using routinely available clinical variables obtained during the early stage of in-hospital assessment. Both model development and internal evaluation were based on the same retrospective single-center electronic medical record dataset, with internal random partitioning used to form the training and validation subsets.
This study was conducted in accordance with the Declaration of Helsinki and its subsequent amendments. The study protocol was reviewed and approved by the Ethics Committee of Chuzhou Hospital Affiliated to Anhui Medical University [Approval No. (2026); Bioethics Review No. 29]. Owing to the retrospective design of the study and the use of de-identified routinely collected clinical data, the requirement for individual written informed consent was waived by the Ethics Committee.
Diagnostic criteria and grouping
The diagnosis of MPP was based on the Evidence-based Guideline for the Diagnosis and Treatment of Mycoplasma pneumoniae Pneumonia in Children (2023). Children were required to have lower respiratory tract infection-related clinical manifestations, chest imaging findings suggestive of pneumonia, and microbiological evidence supporting mycoplasma pneumoniae infection. In this retrospective study, microbiological evidence was defined as a positive mycoplasma pneumoniae nucleic acid test from a respiratory specimen and/or a positive serum mycoplasma pneumoniae-specific IgM antibody result recorded during hospitalization. All included cases were required to have a clear clinical diagnosis of MPP in the medical records together with at least one of these laboratory findings. Cases without documented microbiological evidence of mycoplasma pneumoniae infection were not included (11,12).
The classification of SMPP was also based on the Evidence-based Guideline for the Diagnosis and Treatment of Mycoplasma pneumoniae Pneumonia in Children (2023) and was further interpreted in combination with the severity assessment framework described in the Guidelines for the Management of Community-Acquired Pneumonia in Children (2024 revision). Patients who met the diagnostic criteria for MPP and also fulfilled the criteria for severe pneumonia were assigned to the SMPP group, whereas the remaining patients were assigned to the non-severe MPP group. Severe grouping was based on a comprehensive assessment of overall clinical severity, including respiratory status, oxygenation status, extent of imaging involvement, and complication burden, rather than on a single indicator. In the subsequent analyses, SMPP status was treated as a binary outcome variable (11-13). Because this was a retrospective study based on routinely recorded electronic medical records, no formal blinding procedure was applied during outcome ascertainment.
Data collection
The candidate variables included routine clinical information recorded during the early in-hospital assessment period, including baseline characteristics, symptoms and signs, imaging features, and laboratory indicators. In this retrospective study, the early in-hospital assessment period referred to the initial clinical evaluation after admission, during which routine blood tests, chest imaging review, and complication assessment were completed as part of standard care. Because the exact time point of clinical recognition of SMPP was not always recorded in a structured format, the temporal relationship between some imaging or complication-related predictors and final SMPP classification could not be strictly separated. After feature selection, five key variables were finally retained for model construction: BASO, PLT, atelectasis, pleural effusion, and EXCP. Because some of these variables may be clinically related to the severity assessment framework of SMPP, the present model should be interpreted as an early in-hospital prediction model developed for real-world risk stratification based on routinely available admission-stage information, rather than as a purely pre-event progression model independent of clinical severity evaluation. These five variables were consistently included in all subsequent models.
Predictor definition and assessment
All candidate predictors were collected from routinely available clinical information during the early stage of hospitalization. BASO and PLT were obtained from the first routine blood test results available after admission. Atelectasis and pleural effusion were determined according to chest imaging findings documented during the early in-hospital assessment period. EXCP was identified from the hospitalization records based on clinically documented organ-system involvement outside the respiratory tract during the same early assessment period. For these imaging and complication-related variables, the recorded findings reflected information available during initial in-hospital clinical evaluation. However, because this was a retrospective electronic medical record study, the exact temporal order between the documentation of atelectasis, pleural effusion, EXCP, and the clinical recognition of SMPP could not be fully reconstructed for every patient. Therefore, these variables were treated as early in-hospital severity-related features rather than strictly pre-outcome prognostic predictors.
Sample size
This was a single-center retrospective study, and the study size was determined by the number of eligible hospitalized children with MPP during the predefined study period who met the inclusion criteria and had sufficient data for analysis. No formal prospective sample size calculation was performed. Therefore, the development and comparison of multiple machine learning models in this study should be regarded as exploratory. In particular, the validation subset contained only 30 patients, which limited the precision and stability of the estimated performance metrics.
Data preprocessing
Missing data were imputed using the mice package in R. The classification and regression trees (CART) method was used for imputation, and the number of imputations was set to five. After multiple imputation, a complete dataset was obtained for subsequent statistical analysis and model construction. In the present exploratory analysis, missingness was handled in the full dataset before subsequent feature selection and model development. We acknowledge that this preprocessing strategy may introduce information from the validation subset into the earlier analytical steps and may therefore contribute to optimistic internal performance estimates.
Feature selection
To reduce variable dimensionality and identify candidate predictors, a two-stage feature selection strategy was used in the imputed full dataset before the random train-validation split. In the first stage, an RF algorithm was applied to assess the importance of all candidate variables, and variables with an importance score greater than 0 were retained for the next step. In the second stage, the preselected variables were further analyzed using LASSO regression, and the optimal penalty parameter was determined by five-fold cross-validation. Variables with non-zero coefficients at the optimal lambda were retained as key predictive features. After this two-step selection process, five variables were identified: BASO, PLT, atelectasis, pleural effusion, and EXCP. These variables were then used in subsequent model development and internal validation. Because feature selection was performed before the random split, the validation subset was not fully independent from the predictor selection process. Therefore, the model performance should be interpreted as exploratory internal performance and may be affected by feature-selection-related data leakage.
Model development
The complete dataset was randomly divided into a training set and a validation set at a ratio of 7:3, including 68 cases in the training set and 30 cases in the validation set. Based on the selected key predictive variables, seven supervised learning prediction models were constructed, including Gaussian naive Bayes (GNB), logistic regression (LR), random forest (RF), multilayer perceptron (MLP), LightGBM, support vector machine (SVM), and extreme gradient boosting (XGBoost). The initial hyperparameters were determined by grid search combined with manual tuning.
Model evaluation
Model performance was evaluated in both the training set and the validation set. The primary evaluation metric was the area under the receiver operating characteristic curve (ROC). Sensitivity, specificity, PPV (positive predictive value), and accuracy were calculated. Because this was an exploratory analysis based on a small retrospective dataset, the performance metrics were reported as point estimates. Confidence intervals for AUC, sensitivity, specificity, PPV, and accuracy were not calculated, which limited the precision and interpretability of the model evaluation. Precision-recall (PR) curves were plotted to evaluate predictive performance under different thresholds. Calibration curves were used to assess the consistency between predicted probabilities and actual outcomes. Decision curve analysis (DCA) was performed to evaluate the clinical net benefit of each model at different threshold probabilities. Based on the overall comparison across multiple evaluation metrics and visualization analyses, the MLP model was selected for subsequent optimization and interpretability analysis.
Model optimization and interpretation
Among the seven candidate models, the MLP model was selected for subsequent optimization and interpretability analysis based on the overall comparison across multiple evaluation metrics and visualization analyses in the current dataset. Bayesian optimization was then applied to fine-tune its hyperparameters and establish the final prediction model. The optimization objective was the AUC in the validation process, and the main hyperparameters of the MLP were tuned within predefined search ranges. To improve model interpretability, SHapley Additive exPlanations (SHAP) was used to quantify the relative contribution of each feature to the model output. A SHAP summary plot was further used to show variable importance and the direction of their effects on prediction.
Statistical analysis
Baseline analysis was performed using the complete dataset after imputation, with severe MPP and non-severe MPP as the grouping variables. Continuous variables were expressed as mean ± standard deviation or median (interquartile range), according to their distribution. Categorical variables were expressed as number (percentage). For between-group comparisons, continuous variables were analyzed using the independent-samples t test or Mann-Whitney U test, and categorical variables were analyzed using the Chi-squared test or Fisher’s exact test. Statistical analysis and model construction were performed using R and Python. A two-sided P<0.05 was considered statistically significant.
Results
Study flowchart and feature selection
A total of 98 children with MPP were included in this study, with 49 cases in the SMPP group and 49 cases in the non-severe MPP group (the enrollment flowchart is shown in Figure 1A). To identify key clinical features associated with SMPP risk in this exploratory analysis, we sequentially used the RF algorithm (Figure 1B) and LASSO regression with cross-validation results (Figure 1C,1D) in the imputed full dataset to reduce dimensionality and screen candidate variables. We further compared the overlap between the variables selected by RF and LASSO regression. The Venn diagram showed that the two methods shared five overlapping variables (Figure 1E), namely BASO, PLT, atelectasis, pleural effusion, and EXCP. These five variables were included in all subsequent machine learning prediction model development and performance comparison analyses.
Baseline characteristics of severe and non-severe MPP
The baseline clinical characteristics of the two groups are shown in Table 1. Compared with the non-severe MPP group, the SMPP group showed differences in several hematological and inflammation-related indicators, as well as higher rates of pleural effusion and EXCP. Some variables retained by the machine learning-based feature selection procedure were not necessarily the most statistically significant in univariable comparisons, suggesting that their predictive value may depend on joint effects and nonlinear relationships rather than isolated between-group differences.
Table 1
| Variable | Total (N=98) | SMPP (n=49) | Non-severe MPP (n=49) | P value |
|---|---|---|---|---|
| WBC, ×109/L | 6.620 [5.183–8.678] | 7.510 [5.910–9.500] | 5.960 [4.600–7.920] | 0.004 |
| NEUT#, ×109/L | 3.760 [2.465–5.348] | 4.430 [3.110–5.970] | 3.320 [2.070–4.250] | 0.001 |
| LYM#, ×109/L | 1.995 [1.620–2.953] | 1.970 [1.610–3.010] | 2.040 [1.660–2.930] | >0.99 |
| MONO#, ×109/L | 0.540 [0.412–0.718] | 0.560 [0.450–0.910] | 0.520 [0.370–0.640] | 0.03 |
| EO#, ×109/L | 0.050 [0.010–0.175] | 0.060 [0.010–0.180] | 0.040 [0.010–0.160] | 0.65 |
| BASO#, ×109/L | 0.010 [0.010–0.020] | 0.010 [0.010–0.020] | 0.010 [0.010–0.020] | 0.94 |
| PLT, ×109/L | 251.000 [207.000–314.000] | 255.000 [208.000–331.000] | 240.000 [201.000–306.000] | 0.44 |
| CRP, mg/L | 10.800 [3.725–23.925] | 12.000 [5.000–26.100] | 9.300 [3.000–18.200] | 0.07 |
| NLR | 1.808 [1.252–2.541] | 2.181 [1.418–2.837] | 1.561 [1.064–2.108] | 0.006 |
| MLR | 0.254 [0.208–0.342] | 0.267 [0.231–0.397] | 0.222 [0.192–0.321] | 0.01 |
| PLR | 121.611 [91.873–169.153] | 122.727 [88.372–177.612] | 111.246 [95.618–168.072] | 0.60 |
| SII | 495.801 [286.295–661.199] | 540.819 [368.596–777.359] | 369.130 [263.041–571.530] | 0.009 |
| AGR | 1.647±0.227 | 1.604±0.255 | 1.690±0.187 | 0.06 |
| LDH, U/L | 255.000 [232.000–278.750] | 247.000 [217.000–271.000] | 259.000 [236.000–285.000] | 0.11 |
| Age (months) | 111.510±32.954 | 114.245±34.532 | 108.776±31.415 | 0.41 |
| D-dimer, mg/L | 0.359 [0.283–0.457] | 0.363 [0.251–0.530] | 0.356 [0.293–0.403] | 0.55 |
| Admission body temperature, ℃ | 36.700 [36.400–37.925] | 36.700 [36.300–37.950] | 36.800 [36.500–37.875] | 0.45 |
| Duration of fever (days) | 4.000 [3.000–6.000] | 5.000 [3.250–6.000] | 4.000 [3.000–5.000] | 0.09 |
| Length of hospital stay (days) | 6.987 [6.501–8.670] | 7.022 [6.655–8.082] | 6.962 [5.978–8.741] | 0.47 |
| Male sex | 53 (54.1) | 26 (53.1) | 27 (55.1) | >0.99 |
| Wheezing | 2 (2.0) | 1 (2.0) | 1 (2.0) | >0.99 |
| Gastrointestinal symptoms | 1 (1.0) | 1 (2.0) | 0 (0.0) | >0.99 |
| Moist rales | 25 (25.5) | 18 (36.7) | 7 (14.3) | 0.02 |
| Pulmonary consolidation | 35 (35.7) | 35 (71.4) | 0 (0.0) | <0.001 |
| Atelectasis | 4 (4.1) | 4 (8.2) | 0 (0.0) | 0.12 |
| Pleural effusion | 22 (22.4) | 22 (44.9) | 0 (0.0) | <0.001 |
| Extrapulmonary complications | 6 (6.1) | 6 (12.2) | 0 (0.0) | 0.03 |
| Coinfection | 2 (2.0) | 2 (4.1) | 0 (0.0) | 0.50 |
Data are presented as mean ± SD, median [Q1–Q3] or n (%). AGR, albumin-to-globulin ratio; BASO, basophil absolute count; CRP, C-reactive protein; EO, eosinophil absolute count; LDH, lactate dehydrogenase; LYM, lymphocyte absolute count; MLR, monocyte-to-lymphocyte ratio; MONO, monocyte absolute count; NEUT, neutrophil absolute count; NLR, neutrophil-to-lymphocyte ratio; PLR, platelet-to-lymphocyte ratio; SD, standard deviation; SII, systemic immune-inflammation index; WBC, white blood cell count.
Comparison of machine learning models
Based on the five selected key predictive variables, we constructed seven machine learning prediction models, including GNB, LR, RF, MLP, LightGBM, SVM, and XGBoost. Their performance was evaluated using quantitative metrics and visualization analyses (Table 2). In the training set, ROC curve analysis showed that XGBoost, LR, GNB, SVM, and RF all achieved an AUC of 1.000, whereas the MLP and LightGBM models showed AUC values of 0.976 and 0.947, respectively (Figure 2A). In the validation set, ROC curves showed that XGBoost, LR, GNB, SVM, and RF all achieved an AUC of 1.000, whereas the MLP and LightGBM models showed AUC values of 0.900 and 0.813, respectively (Figure 2B). The validation-set accuracy was 1.000 for XGBoost, LR, GNB, and RF, 0.900 for MLP, 0.850 for SVM, and 0.700 for LightGBM. The MLP model also showed a sensitivity of 1.000, specificity of 0.750, and PPV of 0.857 in the validation set. PR curve analysis showed that the candidate models generally maintained good PR performance, while LightGBM performed relatively less well (Figure 2C). DCA further demonstrated differences in potential clinical net benefit among the models across a range of threshold probabilities (Figure 2D). Overall, several models showed high discriminative performance in the current dataset. Based on the overall comparison across multiple evaluation metrics and visualization analyses, the MLP model was selected for subsequent optimization and interpretability analysis.
Table 2
| Model | Dataset | AUC | Cutoff | Accuracy | Sensitivity | Specificity | PPV |
|---|---|---|---|---|---|---|---|
| XGBoost | Training set | 1 | 0.689 | 1 | 1 | 1 | 1 |
| Validation set | 1 | 0.689 | 1 | 1 | 1 | 1 | |
| Logistic regression | Training set | 1 | 0.844 | 1 | 1 | 1 | 1 |
| Validation set | 1 | 0.844 | 1 | 1 | 1 | 1 | |
| GNB | Training set | 1 | 1 | 1 | 1 | 1 | 1 |
| Validation set | 1 | 1 | 1 | 1 | 1 | 1 | |
| MLP | Training set | 0.976 | 0.374 | 0.91 | 1 | 0.829 | 0.841 |
| Validation set | 0.900 | 0.374 | 0.900 | 1.000 | 0.750 | 0.857 | |
| SVM | Training set | 1 | 0.98 | 1 | 1 | 1 | 1 |
| Validation set | 1 | 0.98 | 0.85 | 0.75 | 1 | 1 | |
| LightGBM | Training set | 0.947 | 0.432 | 0.923 | 0.865 | 0.976 | 0.97 |
| Validation set | 0.813 | 0.432 | 0.7 | 0.75 | 0.625 | 0.75 | |
| Random forest | Training set | 1 | 0.45 | 1 | 1 | 1 | 1 |
| Validation set | 1 | 0.45 | 1 | 1 | 1 | 1 |
AUC, area under the curve; GNB, Gaussian naive Bayes; LightGBM, light gradient boosting machine; MLP, multilayer perceptron; PPV, positive predictive value; SVM, support vector machine; XGBoost, extreme gradient boosting.
Performance of the final optimized model
After the MLP was identified as the optimal baseline model, Bayesian optimization was further applied to optimize its internal settings and establish the final prediction model (detailed performance metrics are shown in Table 3). The optimized MLP model showed an AUC of 0.971 in the training set and 0.933 in the validation set (Figure 3A,3B). These values suggested possible discrimination within the current dataset but should not be interpreted as robust evidence of generalizable predictive performance. The calibration curve further showed acceptable agreement between the predicted probabilities and the actual clinical outcomes (Figure 3C). In addition, the learning curve suggested that model performance remained relatively stable as the number of included cases increased (Figure 3D). These findings suggested that the optimized MLP model had potential utility for early in-hospital risk stratification within the current dataset. However, because the validation subset was small and was generated by a single random split, the reported AUC should be interpreted as an exploratory internal estimate rather than a robust estimate of external performance. The lack of repeated cross-validation or bootstrap validation further limits the stability of the performance estimates. Taken together, these multidimensional evaluation results suggest possible model utility, but they require confirmation using more robust internal validation strategies and larger independent cohorts.
Table 3
| Model | Dataset | AUC |
|---|---|---|
| Optimized MLP | Training set | 0.971 |
| Optimized MLP | Validation set | 0.933 |
AUC, area under the curve; MLP, multilayer perceptron.
SHAP interpretation of the final model
To clarify the basis of the final prediction model, SHAP was used to analyze the optimized MLP model. The SHAP summary plot (Figure 4A) and feature importance bar plot (Figure 4B) showed the magnitude and direction of the contribution of each variable to severe risk prediction. The results showed that a higher PLT, as well as the presence of atelectasis, pleural effusion, and EXCP during early in-hospital assessment, contributed to a higher model-assigned severe-risk classification. In contrast, BASO was associated with a lower predicted risk of severe disease. Overall, the final model mainly integrated hematological indicators, pulmonary involvement features, and complication-related information for severe risk prediction.
Discussion
Over the past 30 years, although the overall number of pediatric pneumonia cases has decreased worldwide, this disease remains one of the leading causes of death in children under 5 years of age (14). MPP is one of the major types of community-acquired pneumonia in children, but early identification of high-risk patients remains difficult in clinical practice. On the one hand, the early clinical manifestations are often non-specific. On the other hand, some imaging findings or complications that suggest disease worsening may not be fully apparent at the initial stage of the disease. Therefore, early prediction based on routine clinical data available during the initial stage of hospitalization is of practical value for improving early stratified management of SMPP.
Against this background, the present study developed a machine learning prediction model based on retrospective clinical data from hospitalized children with MPP, with the aim of early identification of patients at high risk of SMPP. After feature selection and comparison of multiple models using clinical data from 98 children with MPP, we finally selected the MLP model for further optimization. We also identified five core predictive variables, namely BASO, PLT, atelectasis, pleural effusion, and EXCP. These findings suggest that the model has potential value for early risk prediction and may provide a reference for stratified management of pediatric SMPP in real-world clinical settings.
Previous studies have shown that the severity of MPP is closely related to the host immune-inflammatory response, and most existing risk assessment models identify high-risk children from three main aspects: inflammatory burden, extent of pulmonary involvement, and disease complexity (10,15). Although the final variables retained in our study, including BASO, PLT, atelectasis, pleural effusion, and EXCP, are not exactly the same as commonly reported indicators such as CRP, LDH, D-dimer, NLR, or CT involvement range, the clinical dimensions reflected by these variables are generally consistent with the main direction of previous studies. Earlier studies have suggested that pediatric SMPP is not determined only by the pathogen itself, but may be more closely related to dysregulation of the host immune-inflammatory response (15). Among the retained variables, BASO and PLT mainly reflect hematological and inflammation-related information. Although BASO accounts for only a small proportion of peripheral blood cells, mycoplasma pneumoniae infection may be accompanied by changes in basophil-related immune function. Previous studies have also shown that children with allergic constitution who develop MPP may have more severe disease. Therefore, the inclusion of BASO in the model suggests that basic blood cell classification data may reflect differences in immune status related to SMPP (16-18). PLT are involved not only in coagulation, but also in inflammatory amplification and immune regulation. Previous studies found that PLT in children with SMPP may be higher than those in children with non-severe MPP, which is consistent with the finding that the inclusion of PLT may reflect inflammatory burden (15). Therefore, the retention of these two variables suggests that basic hematological indicators may provide additional value in early severe risk prediction.
In addition, atelectasis, pleural effusion, and EXCP are more closely related to the clinical phenotype after disease progression. Previous studies have shown that children with SMPP are more likely to develop pulmonary complications such as atelectasis and pleural effusion, suggesting more severe airway obstruction, inflammatory exudation, and local ventilation impairment (19). Pleural effusion is considered one of the imaging features more clearly associated with SMPP. Children with MPP accompanied by pleural effusion usually have a more complex clinical course, longer duration of fever and hospitalization, and poorer treatment response. When SMPP is accompanied by moderate to large pleural effusion, it often indicates more severe pulmonary complications and a greater disease burden (20,21). EXCP have long been regarded as an important manifestation of systemic involvement in SMPP. Their presence suggests that the disease is no longer limited to the lungs, but has entered a more complex stage of systemic inflammation and multi-organ involvement (17,22). Therefore, the retention of these three variables suggests that imaging abnormalities and complication-related information have important value in clinical risk prediction under real-world admission conditions. It should be noted that atelectasis, pleural effusion, and EXCP have some clinical proximity to the stratification of SMPP itself. Therefore, the present model is more appropriately understood as an early in-hospital prediction tool for clinical risk stratification that integrates laboratory indicators, imaging information, and complication burden, rather than as a purely pre-symptomatic progression model independent of current clinical severity assessment. Overall, although the variables included in different studies may vary, the underlying logic of severe risk prediction is largely consistent, namely, it focuses on inflammatory status, extent of lung involvement, and overall systemic complexity.
From a methodological perspective, most existing studies on SMPP risk assessment have used LR or nomogram models and have mainly emphasized the linear integration of multiple clinical variables (9,23). Through comparison of multiple models and Bayesian optimization, our study found that the MLP model showed relatively better overall predictive performance in internal validation of the current dataset. This finding suggests that, in early risk prediction of pediatric SMPP, machine learning methods may help integrate multidimensional information and capture potential nonlinear relationships. At the same time, we further used SHAP to interpret the model, and the results showed that the model output was determined by the combined contribution of multiple features rather than any single variable. It should be emphasized that this interpretability analysis was used to explain the basis of the model output, rather than to provide direct evidence of the biological mechanisms of these variables. Notably, several candidate models achieved near-perfect performance in the current internal split, including AUC values of 1.000. These findings should be interpreted with caution and should not be presented as definitive evidence of excellent model performance. In a small retrospective dataset, extremely high performance metrics may indicate overfitting, model selection bias, incorporation bias, or optimistic internal validation. The absence of external validation further limits the interpretation of these results. In addition, because confidence intervals were not calculated for the performance metrics, the uncertainty around the reported AUC, sensitivity, specificity, PPV, and accuracy could not be quantified. The real-world clinical utility and incremental value of this model should also be interpreted cautiously. Several retained predictors, especially atelectasis, pleural effusion, and EXCP, are clinically evident markers of more severe disease and are commonly considered by experienced clinicians during routine severity assessment. Therefore, the present model may not provide substantial new information beyond expert clinical judgment when these findings are already apparent. Its potential value may instead lie in integrating routinely available laboratory, imaging, and complication-related information into a structured and reproducible risk stratification framework. Such a framework may help standardize early in-hospital assessment, support less experienced clinicians, and provide a quantitative summary of severity-related information in settings where clinical evaluation may vary across physicians or institutions. However, this potential incremental value was not directly tested in the present study. We did not compare model-assisted assessment with clinician-only assessment, nor did we perform decision-impact analysis. Therefore, the model should currently be regarded as an exploratory auxiliary tool rather than a validated replacement for routine clinical judgment. The total sample size was limited, and the validation subset included only 30 patients. Under this condition, complex algorithms such as RF, XGBoost, LightGBM, and MLP may capture sample-specific patterns rather than stable disease-related signals. Therefore, the near-perfect AUC values observed in several models may reflect overfitting and optimistic internal validation. The model comparison in this study should be regarded as exploratory, and the reported performance metrics should not be interpreted as definitive evidence of generalizable predictive accuracy.
Several limitations of this study should be acknowledged. First, this was a single-center retrospective study with a relatively small sample size. Although an internal random split was used, the validation subset included only 30 patients. Therefore, the reported AUC and other performance metrics should be interpreted as preliminary exploratory internal estimates. Larger multicenter and prospective cohorts are needed to evaluate the generalizability of the model. Second, the internal validation strategy was limited. A single 7:3 random split may be sensitive to sample allocation, especially in a small dataset. Repeated cross-validation, bootstrap validation, or nested cross-validation may provide more stable estimates of model performance. These approaches were not performed in the present analysis and should be considered in future studies. Third, the analytical workflow may have introduced optimistic bias. Missing data imputation, RF-based feature screening, and LASSO regression were performed in the full dataset before the random train-validation split. Therefore, the validation subset may have contributed information to predictor selection, and the subsequent validation process was not fully independent. Future studies should perform imputation, feature selection, hyperparameter tuning, and model evaluation within a strictly separated training-validation framework. Fourth, several candidate models showed extremely high performance metrics, including AUC values of 1.000. These findings should be interpreted with caution because multiple machine learning algorithms, including complex models such as RF, XGBoost, LightGBM, and MLP, were compared in a limited dataset. This may increase the risk of overfitting, model selection bias, and optimistic internal validation. In addition, confidence intervals were not calculated for the performance metrics, which limited the assessment of statistical uncertainty around the reported AUC, sensitivity, specificity, PPV, and accuracy. Fifth, incorporation bias, outcome leakage, and temporal ambiguity should be considered when interpreting the model. Three retained predictors, namely atelectasis, pleural effusion, and EXCP, are clinically close to the severity assessment framework used to define SMPP. Although these variables were extracted from the early in-hospital assessment period, their exact timing relative to the clinical recognition of SMPP could not be fully reconstructed from the retrospective electronic medical records. Therefore, the model should not be interpreted as a pure prognostic model predicting later progression before severe manifestations appear. It is more appropriately regarded as an early in-hospital severity risk stratification tool. Sixth, the incremental clinical value of the model beyond routine physician assessment was not directly evaluated. Several selected predictors are clinically evident markers of severe disease and may already be recognized by experienced clinicians. We did not compare the model with clinician-only assessment, existing severity criteria, or simpler rule-based approaches. We also did not assess whether model-assisted decision-making would improve treatment timing, resource allocation, or patient outcomes. Seventh, the microbiological confirmation of MPP was based on routinely available clinical test results in a retrospective setting. Mycoplasma pneumoniae nucleic acid testing and serum IgM antibody testing have different diagnostic windows and may vary in sensitivity and specificity. IgM positivity may persist after recent infection, whereas nucleic acid detection may be influenced by sampling quality, timing, and pathogen load. Therefore, differences in microbiological testing methods may have introduced diagnostic heterogeneity or misclassification into the cohort. Future prospective studies should use standardized microbiological criteria, ideally combining molecular testing with paired serology when feasible. Finally, this was a clinical prediction model study, and the SHAP results only reflect the contribution of variables to model output. They cannot be used to infer direct causal relationships between BASO, PLT, atelectasis, pleural effusion, EXCP, and SMPP. Future external validation and prospective clinical utility studies are needed to further optimize the model and clarify its appropriate role in clinical decision-making.
Conclusions
This study developed an exploratory machine learning model integrating routine hematological indicators, imaging findings, and EXCP for early in-hospital severity risk stratification of children with MPP. The optimized MLP model showed promising discrimination within the current dataset. However, because of the small single-center sample, limited internal validation, potential information leakage, and the clinical proximity of several predictors to the SMPP severity assessment framework, the model should be regarded as an auxiliary risk stratification tool rather than a validated prognostic model or a replacement for clinical judgment. Larger multicenter cohorts and prospective external validation are required before clinical application
Acknowledgments
None.
Footnote
Reporting Checklist: The authors have completed the TRIPOD reporting checklist. Available at https://tp.amegroups.com/article/view/10.21037/tp-2026-0333/rc
Data Sharing Statement: Available at https://tp.amegroups.com/article/view/10.21037/tp-2026-0333/dss
Peer Review File: Available at https://tp.amegroups.com/article/view/10.21037/tp-2026-0333/prf
Funding: This study was supported by
Conflicts of Interest: All authors have completed the ICMJE uniform disclosure form (available at https://tp.amegroups.com/article/view/10.21037/tp-2026-0333/coif). All authors report research funding from and a research collaboration with Adicon (Hefei) Clinical Laboratories Co., Ltd. The authors have no other conflicts of interest to declare.
Ethical Statement: The authors are accountable for all aspects of the work in ensuring that questions related to the accuracy or integrity of any part of the work are appropriately investigated and resolved. This study was conducted in accordance with the Declaration of Helsinki and its subsequent amendments. The study protocol was reviewed and approved by the Ethics Committee of Chuzhou Hospital Affiliated to Anhui Medical University [Approval No. (2026); Bioethics Review No. 29]. Owing to the retrospective design of the study and the use of de-identified routinely collected clinical data, the requirement for individual written informed consent was waived by the Ethics Committee.
Open Access Statement: This is an Open Access article distributed in accordance with the Creative Commons Attribution-NonCommercial-NoDerivs 4.0 International License (CC BY-NC-ND 4.0), which permits the non-commercial replication and distribution of the article with the strict proviso that no changes or edits are made and the original work is properly cited (including links to both the formal publication through the relevant DOI and the license). See: https://creativecommons.org/licenses/by-nc-nd/4.0/.
References
- Mendez R, Banerjee S, Bhattacharya SK, et al. Lung inflammation and disease: A perspective on microbial homeostasis and metabolism. IUBMB Life 2019;71:152-65. [Crossref] [PubMed]
- Lee KL, Lee CM, Yang TL, et al. Severe Mycoplasma pneumoniae pneumonia requiring intensive care in children, 2010-2019. J Formos Med Assoc 2021;120:281-91. [Crossref] [PubMed]
- Hon KL, Leung AS, Cheung KL, et al. Typical or atypical pneumonia and severe acute respiratory symptoms in PICU. Clin Respir J 2015;9:366-71. [Crossref] [PubMed]
- Zhang X, Yu Y. Severe pediatric Mycoplasma pneumonia as the cause of diffuse alveolar hemorrhage requiring veno-venous extracorporeal membrane oxygenation: A case report. Front Pediatr 2022;10:925655. [Crossref] [PubMed]
- Sun Y, Li P, Jin R, et al. Characterizing the epidemiology of Mycoplasma pneumoniae infections in China in 2022-2024: a nationwide cross-sectional study of over 1.6 million cases. Emerg Microbes Infect 2025;14:2482703. [Crossref] [PubMed]
- Li Y, Wu M, Liang Y, et al. Mycoplasma pneumoniae infection outbreak in Guangzhou, China after COVID-19 pandemic. Virol J 2024;21:183. [Crossref] [PubMed]
- Søndergaard MJ, Friis MB, Hansen DS, et al. Clinical manifestations in infants and children with Mycoplasma pneumoniae infection. PLoS One 2018;13:e0195288. [Crossref] [PubMed]
- Chang Q, Chen HL, Wu NS, et al. Prediction Model for Severe Mycoplasma pneumoniae Pneumonia in Pediatric Patients by Admission Laboratory Indicators. J Trop Pediatr 2022;68:fmac059. [Crossref] [PubMed]
- Li L, Guo R, Zou Y, et al. Construction and Validation of a Nomogram Model to Predict the Severity of Mycoplasma pneumoniae Pneumonia in Children. J Inflamm Res 2024;17:1183-91. [Crossref] [PubMed]
- Zhang X, Sun R, Jia W, et al. A new dynamic nomogram for predicting the risk of severe Mycoplasma pneumoniae pneumonia in children. Sci Rep 2024;14:8260. [Crossref] [PubMed]
- Subspecialty Group of Respiratory. the Society of Pediatrics, Chinese Medical Association; China National Clinical Research Center of Respiratory Diseases; Editorial Board, Chinese Journal of Pediatrics. Evidence-based guideline for the diagnosis and treatment of Mycoplasma pneumoniae pneumonia in children (2023). Pediatr Investig 2025;9:1-11.
- Subspecialty Group of Respiratory. the Society of Pediatrics, Chinese Medical Association; China National Clinical Research Center of Respiratory Diseases; Editorial Board, Chinese Journal of Pediatrics. Evidence-based guideline for the diagnosis and treatment of Mycoplasma pneumoniae pneumonia in children (2023). Zhonghua Er Ke Za Zhi 2024;62:1137-44.
- Subspecialty Group of Respiratory. the Society of Pediatrics, Chinese Medical Association; Editorial Board, Chinese Journal of Pediatrics; China Medicine Education Association Committee on Pediatrics. Guidelines for the management of community-acquired pneumonia in children (2024 revision). Zhonghua Er Ke Za Zhi 2024;62:920-30.
- Kok HC, Chang AB, Fong SM, et al. Antibiotics for Paediatric Community-Acquired Pneumonia: What is the Optimal Course Duration? Paediatr Drugs 2025;27:261-72. [Crossref] [PubMed]
- Jiang C, Bao S, Shen W, et al. Predictive value of immune-related parameters in severe Mycoplasma pneumoniae pneumonia in children. Transl Pediatr 2024;13:1521-8. [Crossref] [PubMed]
- Bian C, Li S, Huo S, et al. Association of atopy with disease severity in children with Mycoplasma pneumoniae pneumonia. Front Pediatr 2023;11:1281479. [Crossref] [PubMed]
- Wang Z, Sun J, Liu Y, et al. Impact of atopy on the severity and extrapulmonary manifestations of childhood Mycoplasma pneumoniae pneumonia. J Clin Lab Anal 2019;33:e22887. [Crossref] [PubMed]
- Jeong YC, Yeo MS, Kim JH, et al. Mycoplasma pneumoniae Infection Affects the Serum Levels of Vascular Endothelial Growth Factor and Interleukin-5 in Atopic Children. Allergy Asthma Immunol Res 2012;4:92-7. [Crossref] [PubMed]
- Huang X, Gu H, Wu R, et al. Chest imaging classification in Mycoplasma pneumoniae pneumonia is associated with its clinical features and outcomes. Respir Med 2024;221:107480. [Crossref] [PubMed]
- Kim SH, Lee E, Song ES, et al. Clinical Significance of Pleural Effusion in Mycoplasma pneumoniae Pneumonia in Children. Pathogens 2021;10:1075. [Crossref] [PubMed]
- Luo XQ, Luo J, Wang CJ, et al. Clinical features of severe Mycoplasma pneumoniae pneumonia with pulmonary complications in childhood: A retrospective study. Pediatr Pulmonol 2023;58:2815-22. [Crossref] [PubMed]
- Butpech T, Tovichien P. Mycoplasma pneumoniae pneumonia in children. World J Clin Cases 2025;13:99149. [Crossref] [PubMed]
- Wu X, Lu W, Wang T, et al. Optimization strategy for the early timing of bronchoalveolar lavage treatment for children with severe mycoplasma pneumoniae pneumonia. BMC Infect Dis 2023;23:661. [Crossref] [PubMed]

