CloudSafe-TP: Hybrid AI-Based SLO-Aware Predictive Thread-Pool Control for Distributed Cloud Backends
DOI:
https://doi.org/10.33411/IJIST/1989Keywords:
Cloud-based distributed services, Thread-pool control, P99 latency SLO, Gaussian process, conformal calibration, model predictive control, Cloud AI Integration, Cloud Computing, Distributed ComputingAbstract
loud-based distributed services use backend thread-pool controllers to process changing request workloads. Fixed and reactive controllers may allocate too few threads during bursts or maintain unnecessary threads during normal operation. Existing schemes often depend on current measurements, fixed rules, or predictions without calibrated uncertainty, which may cause queue growth and P99 latency SLO violations. This paper presents CloudSafe-TP, a hybrid AI-based and SLO-aware predictive thread-pool control framework for distributed cloud backends. The novelty of CloudSafe-TP lies in combining a sparse Gaussian-process state-space model, conformal calibration, and model predictive control within an independent local controller that evaluates candidate pool sizes before they are applied. At each control interval, CloudSafe-TP predicts P99 latency, queue length, CPU usage, and memory usage over a short horizon, removes candidates that may violate the P99 SLO or resource limits, and selects a lower-cost safe configuration. The framework was evaluated on Oracle Cloud Infrastructure against analytical, neural-network-based, and reinforcement-learning controllers through five independent runs using random seeds 0–4. Under the bursty workload, CloudSafe-TP reduced mean P99 latency by 7.17–23.01%, queue length by 39.74–55.01%, and SLO violation rate by 77.11–89.57% compared with the baselines. Under moderate load, it used 20.70–26.60% fewer threads, reduced CPU usage by 3.90–7.64%, and reduced heap usage by 2.68–23.01%, while maintaining a P99 latency comparable to that of the strongest baseline. Paired t-tests further confirmed statistically significant improvements in P99 latency, queue length, and SLO violation rate under the bursty workload. Overall, CloudSafe-TP provides predictive and uncertainty-aware pool-size adjustment that improves P99 SLO protection during bursty workloads while avoiding unnecessary concurrency and resource usage under moderate load.
References
Z. Abdiramanov, “Design and implementation of a scalable microservice-based cloud computing service for stochastic fluid flow modeling,” Array, vol. 29, p. 100679, Mar. 2026, doi: 10.1016/J.ARRAY.2026.100679.
L. Zanussi, D. Tessera, L. Massari, B. Bermejo, C. Juiz, and M. Calzarossa, “Load-aware predictive auto-scaling framework for cloud environments,” Clust. Comput. 2026 293, vol. 29, no. 3, pp. 197-, Mar. 2026, doi: 10.1007/S10586-026-05944-X.
D. Charlak, J. Brzeziński, and G. Kozieł, “Comparative analysis of reactive programming and Java virtual threads,” J. Comput. Sci. Inst., vol. 39, pp. 154–160, Jun. 2026, doi: 10.35784/JCSI.9409.
C. Costa and J. H. Brito, “Throughput impact of software multithreading for deep-learning inference on the AMD Kria KV260,” J. Real-Time Image Process. 2026 232, vol. 23, no. 2, pp. 89-, Apr. 2026, doi: 10.1007/S11554-026-01889-X.
S. Harutyunyan Gevorgyan, E. César, A. Sikora, J. Filipovič, and J. Alcaraz, “Automatic tuning based on hardware performance counters and machine learning,” Futur. Gener. Comput. Syst., vol. 179, p. 108358, Jun. 2026, doi: 10.1016/J.FUTURE.2025.108358.
W. Du, H. Yuan, Y. Ren, J. Yu, and X. Ren, “Tiered Scheduling and Cost-Benefit-Aware Speculation for Optimizing Spark in Heterogeneous Clusters,” Concurr. Comput. Pract. Exp., vol. 38, no. 13, p. e70844, Jul. 2026, doi: 10.1002/CPE.70844.
“Chapter 9. Distributed Executor Service.” Accessed: Aug. 14, 2026. [Online]. Available: https://docs.hazelcast.org/docs/2.3/manual/html/ch09.html
L. Akash, D. Fernando, M. Jayasinghe, C. Keppitiyagama, and K. Thangarajah, “Machine Learning Based Thread Pool Tuning via Program Analysis,” 2021 IEEE 23rd Int. Conf. High Perform. Comput. Commun. 7th Int. Conf. Data Sci. Syst. 19th Int. Conf. Smart City 7th Int. Conf. Dependability Sensor, C…, pp. 648–653, 2022, doi: 10.1109/HPCC-DSS-SMARTCITY-DEPENDSYS53884.2021.00108.
J. Alcaraz et al., “Predicting number of threads using balanced datasets for openMP regions,” Comput. 2022 1055, vol. 105, no. 5, pp. 999–1017, Apr. 2022, doi: 10.1007/S00607-022-01081-6.
F. H. S. da Silva, J. B. Fernandes, I. M. Sardina, T. Barros, S. Xavier-de-Souza, and I. A. S. Assis, “Auto-Tuning for OpenMP Dynamic Scheduling applied to Full Waveform Inversion,” May 2025, doi: 10.1016/j.cageo.2025.105932.
A. Dutta, J. Alcaraz, A. Tehranijamsaz, A. Sikora, E. Cesar, and A. Jannesari, “Pattern-based Autotuning of OpenMP Loops using Graph Neural Networks,” Proc. AI4S 2022 Artif. Intell. Mach. Learn. Sci. Appl. Held conjunction with SC 2022 Int. Conf. High Perform. Comput. Networking, Storage Anal., pp. 26–31, 2022, doi: 10.1109/AI4S56813.2022.00010.
A. Dutta, J. Choi, and A. Jannesari, “Power Constrained Autotuning using Graph Neural Networks,” Proc. - 2023 IEEE Int. Parallel Distrib. Process. Symp. IPDPS 2023, pp. 535–545, 2023, doi: 10.1109/IPDPS54959.2023.00060.
E. C. de Lima, F. D. Rossi, M. C. Luizelli, R. N. Calheiros, and A. F. Lorenzon, “A neural network framework for optimizing parallel computing in cloud servers,” J. Syst. Archit., vol. 150, p. 103131, May 2024, doi: 10.1016/J.SYSARC.2024.103131.
“CAPES: Unsupervised Storage Performance Tuning Using Neural Network-Based Deep Reinforcement Learning | IEEE Conference Publication | IEEE Xplore.” Accessed: Aug. 14, 2026. [Online]. Available: https://ieeexplore.ieee.org/document/9926120
H. T. Nguyen, T. V. Do, and C. Rotter, “Optimizing the resource usage of actor-based systems,” J. Netw. Comput. Appl., vol. 190, p. 103143, Sep. 2021, doi: 10.1016/J.JNCA.2021.103143.
J. Schwarzrock, C. C. De Oliveira, M. Ritt, A. F. Lorenzon, and A. C. S. Beck, “A Runtime and Non-Intrusive Approach to Optimize EDP by Tuning Threads and CPU Frequency for OpenMP Applications,” IEEE Trans. Parallel Distrib. Syst., vol. 32, no. 7, pp. 1713–1724, Jul. 2021, doi: 10.1109/TPDS.2020.3046537.
T. S. Medeiros, L. Pereira, F. D. Rossi, M. C. Luizelli, A. C. S. Beck, and A. F. Lorenzon, “Mitigating the processor aging through dynamic concurrency throttling,” J. Parallel Distrib. Comput., vol. 156, pp. 86–100, Oct. 2021, doi: 10.1016/J.JPDC.2021.05.006.
J. Liu, Q. Wang, S. Zhang, L. Hu, and D. Da Silva, “Sora: A Latency Sensitive Approach for Microservice Soft Resource Adaptation,” Middlew. 2023 - Proc. 24th ACM/IFIP Int. Middlew. Conf., pp. 43–56, Nov. 2023, doi: 10.1145/3590140.3592851.
“Cloud Infrastructure | Oracle.” Accessed: Aug. 14, 2026. [Online]. Available: https://www.oracle.com/cloud/
F. Bahadur, A. I. Umar, I. Ullah, F. Algarni, and M. A. Khan, “The Double Edge Sword Based Distributed Executor Service,” Comput. Syst. Sci. Eng., vol. 42, no. 2, pp. 589–604, Jan. 2022, doi: 10.32604/CSSE.2022.022319.
J. Park, B. Choi, C. Lee, and D. Han, “Graph Neural Network-Based SLO-Aware Proactive Resource Autoscaling Framework for Microservices,” IEEE/ACM Trans. Netw., vol. 32, no. 4, pp. 3331–3346, 2024, doi: 10.1109/TNET.2024.3393427.
J. P. K. S. Nunes, S. Nejati, M. Sabetzadeh, and E. Y. Nakagawa, “Self-adaptive, Requirements-driven Autoscaling of Microservices,” Proc. - 2024 IEEE/ACM 19th Symp. Softw. Eng. Adapt. Self-Managing Syst. SEAMS 2024, pp. 168–174, Apr. 2024, doi: 10.1145/3643915.3644094
K. Hu, M. Xu, K. Ye, and C. Xu, “LSRAM: A Lightweight Autoscaling and SLO Resource Allocation Framework for Microservices Based on Gradient Descent,” Softw. - Pract. Exp., vol. 55, no. 4, pp. 714–730, Nov. 2024, doi: 10.1002/spe.3395.
F. Bahadur, Z. Ahmad, and A. Algarni, “PoolRunner: An Extensible Performance Testing Simulation Tool for Thread-Pool Middleware,” IEEE Access, vol. 13, pp. 96660–96680, 2025, doi: 10.1109/ACCESS.2025.3575437.
“Compute Shapes.” Accessed: Aug. 14, 2026. [Online]. Available: https://docs.oracle.com/enus/iaas/Content/Compute/References/computeshapes.htm
“JDK 17 Readme.” Accessed: Aug. 14, 2026. [Online]. Available: https://www.oracle.com/java/technologies/javase/jdk17-readme-downloads.html
“Networking Overview.” Accessed: Aug. 14, 2026. [Online]. Available: https://docs.oracle.com/en-us/iaas/Content/Network/Concepts/overview.htm
Downloads
Published
How to Cite
Issue
Section
License
Copyright (c) 2026 50SEA

This work is licensed under a Creative Commons Attribution 4.0 International License.


















