Random Forest and Support Vector Machine Models for Predicting Drug Release Profiles from Polymeric Nanocarriers Targeting Multi-Drug Resistant Bacterial Infections

Authors

  • Ravula Arun Kumar
  • Tarana Singh
  • L. Chandra Sekhar Reddy
  • A. Jyothi
  • Nikhila Kathirisetty
  • Amita Mishra
  • Babita Verma

Keywords:

Polymer-based nanocarriers, drug release prediction, support vector machine, random forest, multidrug-resistant bacteria, machine learning, antimicrobial drug delivery, nanoparticle optimization.

Abstract

Polymeric nanocarriers are being investigated for the targeted delivery of antibiotics to combat multi-drug resistant (MDR) bacterial infections. Predicting in vitro drug release is a crucial step in early formulation screening, although in vivo validation remains necessary to establish clinical relevance. Because of the intricate interactions among formulation variables, forecasting the drug release behavior of these carriers is complex. This study employs machine learning models, namely Random Forest (RF) and Support Vector Machine (SVM), to estimate the 24-hour cumulative drug release (CDR₂₄) from polymeric nanocarriers designed for MDR infections. The dataset comprised 412 experimental records from 87 studies published between 2005 and 2023, with seven input features: polymer concentration, drug loading, particle size, zeta potential, pH, temperature, and polymer type. Both models were optimized through a 5-fold cross-validation grid search after preprocessing steps (one-hot encoding and min–max scaling) and a 70:15:15 split for training, validation, and testing. On the held-out test set (n = 62), RF outperformed SVM (R² = 0.89, RMSE = 4.87%, MAE = 3.62%) as well as the baseline methods, the Korsmeyer–Peppas kinetic model (R² = 0.71) and linear regression (R² = 0.58), achieving R² = 0.94, RMSE = 3.21%, and MAE = 2.47%. A leave-one-study-out cross-validation, used to assess generalization across the 87 source studies, gave a more conservative average R² of 0.881 ± 0.067. Feature importance analysis showed that polymer concentration, drug loading, and particle size were the most influential predictors, together accounting for 69.2% of the RF model's predictive power. With an RMSE of 3.21% and MAE of 2.47%, the RF model allows formulation scientists to computationally screen candidate nanocarriers and reserve experimental dissolution testing for the subset flagged as most promising. In a simulated screening exercise involving 100 candidate formulations, applying a decision threshold of predicted CDR₂₄ ≥ 70% with RMSE < 5% flagged 12 formulations for experimental testing, a reduction in required dissolution experiments of 88% under this specific threshold; a more conservative estimate across broader formulation spaces and less favorable thresholds would be in the range of 30–50%. These reductions depend on the model's precision (0.91) and recall (0.87) being sustained on future datasets, and prospective validation is necessary to confirm the effectiveness of this approach.

Downloads

Published

2026-09-22

How to Cite

Kumar, R. A., Singh, T., Reddy, L. C. S., Jyothi, A., Kathirisetty, N., Mishra, A., & Verma, B. (2026). Random Forest and Support Vector Machine Models for Predicting Drug Release Profiles from Polymeric Nanocarriers Targeting Multi-Drug Resistant Bacterial Infections. International Journal of Artificial Intelligence and Machine Learning, 6(11s), 995–1009. Retrieved from https://mail.svedbergopen.com/index.php/ijaiml/article/view/2220