HARP: An Adaptive Hybrid Autoscaling Framework with Online Policy Arbitration for SLO-Aware Cloud-Native ML Inference

Authors

  • Yuchen Luo South China University of Technology, China
  • Zihan Peng South China University of Technology, China
  • Rui Zhang South China University of Technology, China

DOI:

https://doi.org/10.54097/8vyfba92

Keywords:

Autoscaling, machine learning inference, service level objectives, online learning, conformal prediction, Kubernetes, cloud-native systems

Abstract

Cloud-native machine learning inference services must hold tail-latency service level objectives (SLOs) while paying for as few replicas as possible. Two families of autoscalers compete for this job: reactive threshold controllers such as the Kubernetes horizontal pod autoscaler, and predictive controllers that provision ahead of demand. Neither family dominates. Which one wins is decided by operating conditions—replica activation delay, the price an operator places on a replica relative to an SLO violation, and the statistical regime of the arriving workload—that are unknown at design time and change during a deployment. We present HARP, an adaptive hybrid framework that treats the choice of control philosophy as an online decision rather than a design-time commitment. HARP maintains a pool of six heterogeneous demand hypotheses, scores every hypothesis at every epoch with a counterfactual shadow loss computed from telemetry a production deployment already collects, and arbitrates among them with a fixed-share Hedge rule that provably tracks the best hypothesis under regime change. Two further mechanisms make the arbitration actionable: an adaptive conformal margin whose coverage target is driven by an SLO error budget, and a risk-averse aggregation that maps the arbiter's belief to a replica count. On three real production traces (AWS ELB request counts, NYC taxi demand and tweet volume) driven by a measured CPU inference profile, HARP cuts SLO violations by 26.8–92.4% against reactive and forecasting baselines. Across 27 operating conditions its worst-case cost regret against the best fixed policy of each condition is 4.4%, versus 42.5% for a strong confidence-aware predictive controller and over 300% for threshold controllers.

Downloads

Download data is not yet available.

References

[1] Zhang, H. (2024). Graph aware multi task learning for early prediction of timing violations and routing congestion in digital IC physical design. World Journal of Information Technology, 2(1), 13–20.

[2] Wang, J., Tan, Y., Jiang, B., Wu, B., & Liu, W. (2025). Dynamic marketing uplift modeling: A symmetry preserving framework integrating causal forests with deep reinforcement learning for personalized intervention strategies. Symmetry, 17(4), 610.

[3] Zhao, X., Liu, J., Wang, Y., & Wang, J. (2026). CryptoMamba SSM: Linear complexity state space models for cryptocurrency volatility prediction. IEEE Open Journal of the Computer Society, 7, 226–243.

[4] Wang, J., Liu, J., Zheng, W., & Ge, Y. (2025). Temporal heterogeneous graph contrastive learning for fraud detection in credit card transactions. IEEE Access.

[5] Sun, T., Yang, J., Li, J., Chen, J., Liu, M., Fan, L., & Wang, X. (2024). Enhancing auto insurance risk evaluation with transformer and SHAP. IEEE Access, 12, 116546–116557.

[6] Zhang, H. (2025). Reinforcement learning approaches for layout optimization in electronic design automation with electromagnetic compatibility constraints. Frontiers in Robotics and Automation, 2(2), 77–93.

[7] Han, X., Yang, Y., Chen, J., Wang, M., & Zhou, M. (2025). Symmetry aware credit risk modeling: A deep learning framework exploiting financial data balance and invariance. Symmetry, 17(3), 341.

[8] Chen, J., Cui, Y., Zhang, X., Yang, J., & Zhou, M. (2024). Temporal convolutional network for carbon tax projection: A data driven approach. Applied Sciences, 14(20), 9213.

[9] Sun, T., Wang, M., & Chen, J. (2025). Leveraging machine learning for tax fraud detection and risk scoring in corporate filings. Asian Business Research Journal, 10(11), 1–13.

[10] Chen, J., Wang, M., & Sun, T. (2025). Intelligent tax systems and the role of natural language processing in regulatory interpretation. American Journal of Machine Learning, 6(4), 74–94.

[11] Chen, Z., Liu, J., & Chen, J. (2025). Machine learning methods for financial forecasting in enterprise planning: Transitioning from rule based models to predictive analytics. Frontiers in Artificial Intelligence Research, 2(3), 541–564.

[12] Chen, J., Liu, J., Liang, Y., & Zhou, M. (2026). KE MLLM: A knowledge enhanced multi sensor learning framework for explainable fake review detection. Applied Sciences, 16(6), 2909.

[13] Yue, X., Yang, J., & Liu, W. (2026). FraudDebate Agent: A multi agent LLM framework with an evidence based debate mechanism for financial statement fraud detection. Mathematics, 14(15), 2695.

[14] Zhu, R., & Yue, X. (2024). An explainable machine learning framework for predicting dividend sustainability and financial risk of U.S. REITs under interest rate shocks. Journal of Trends in Finance and Economics, 1(2), 48–57.

[15] Yue, X., & Zhu, R. (2024). An explainable hybrid machine learning framework for cash flow conditioned multi horizon liquidity risk early warning: An SME oriented benchmark study. World Journal of Management Science, 2(1), 63–72.

[16] Zi, B. (2024). Cloud native distributed systems for real time payment intelligence. AI and Data Science Journal, 1(1), 51–56.

[17] Zhang, S., & Qiu, L. (2023). An interpretable, uncertainty aware framework for bridge deterioration prediction and risk based maintenance prioritization. Academic Journal of Architecture and Civil Engineering, 1(2), 21–28.

[18] Jin, J., Xing, S., Ji, E., & Liu, W. (2025). Xgate: Explainable reinforcement learning for transparent and trustworthy API traffic management in IoT sensor networks. Sensors, 25(7), 2183.

[19] Xing, S., Wang, Y., & Liu, W. (2025). Multi dimensional anomaly detection and fault localization in microservice architectures: A dual channel deep learning approach with causal inference for intelligent sensing. Sensors, 25(11), 3396.

[20] Xing, S., & Wang, Y. (2025). Proactive data placement in heterogeneous storage systems via predictive multi objective reinforcement learning. IEEE Access, 13, 117986–117998.

[21] Ji, E., Wang, Y., Xing, S., & Jin, J. (2025). Hierarchical reinforcement learning for energy efficient API traffic optimization in large scale advertising systems. IEEE Access.

[22] Xing, S., & Wang, Y. (2025). Cross modal attention networks for multi modal anomaly detection in system software. IEEE Open Journal of the Computer Society.

[23] Liu, Y., Ren, S., Wang, X., & Zhou, M. (2024). Temporal logical attention network for log based anomaly detection in distributed systems. Sensors, 24(24), 7949.

[24] Ren, S., Jin, J., Niu, G., & Liu, Y. (2025). ARCS: Adaptive reinforcement learning framework for automated cybersecurity incident response strategy optimization. Applied Sciences, 15(2), 951.

[25] Li, P., Ren, S., Zhang, Q., Wang, X., & Liu, Y. (2024). Think4SCND: Reinforcement learning with thinking model for dynamic supply chain network design. IEEE Access, 12, 195974–195985.

[26] Hu, X., Zhao, X., Wang, J., & Yang, Y. (2025). Information theoretic multi scale geometric pre training for enhanced molecular property prediction. PLOS ONE, 20(10), e0332640.

[27] Wang, M., Zhang, X., Yang, Y., & Wang, J. (2025). Explainable machine learning in risk management: Balancing accuracy and interpretability. Journal of Financial Risk Management, 14(3), 185–198.

[28] Yang, J., Li, P., Cui, Y., Han, X., & Zhou, M. (2025). Multi sensor temporal fusion transformer for stock performance prediction: An adaptive Sharpe ratio approach. Sensors, 25(3), 976.

[29] Zhao, X., Sun, T., Ren, S., Yang, J., & Liu, Y. (2025). RAG Based AI Agents for Enterprise Software Development: Implementation Patterns and Production Deployment. Frontiers in Artificial Intelligence Research, 2(3), 501–520.

[30] Sun, T., Wang, M., & Chen, J. (2025). Leveraging machine learning for tax fraud detection and risk scoring in corporate filings. Asian Business Research Journal, 10(11), 1–13.

[31] Shang, W., Teng, D., Chen, T., Zhang, F., & Ding, J. (2026). Neuro Elastic: A unified framework for hardware aware adaptive quantization and dynamic sparsity in real time edge intent prediction. IEEE Access.

[32] Teng, D. (2025). TEAS: Token and energy aware autoscaling for cost efficient LLM serving. AI and Data Science Journal, 6(3), 1–10.

[33] Qin, Y., Rhee, M., Teng, D., Zhang, C., & Zi, B. (2026). Neuro symbolic constraint verification for LLM driven internet finance transaction execution. Scientific Reports.

[34] Wang, B., Zhao, W., & Wang, Z. (2024). DAS GNN: A scalable disagreement aware graph neural framework for anomaly detection in large scale networks. AI and Data Science Journal, 1(1), 67–75.

[35] Wang, Z., & Shen, Z. (2024). Flex Talent: An explainable and fair machine learning framework for high stakes leadership pipeline evaluation. World Journal of Management Science, 2(1), 73–82.

[36] Guo, Z., Chen, T., & Shang, W. (2022). Class conditional intermediate domain adversarial adaptation for cross device acoustic scene classification. Journal of Computer Science and Electrical Engineering, 4(1), 26–35.

[37] Wang, Y., Bian, M., & Lin, Y. (2022). An explainable temporal graph neural network for disruption prediction and risk propagation in multi tier supply chains. Journal of Computer Science and Electrical Engineering, 4(1), 17–25.

[38] Zhao, W., & Wang, B. (2022). CAPA: A prediction driven autoscaling framework for SLO aware cloud native machine learning inference. Journal of Computer Science and Electrical Engineering, 4(1), 8–16.

Downloads

Published

2026-09-07

Issue

Section

Articles

How to Cite

Luo, Y., Peng, Z., & Zhang, R. (2026). HARP: An Adaptive Hybrid Autoscaling Framework with Online Policy Arbitration for SLO-Aware Cloud-Native ML Inference. International Journal of Advanced Engineering and Technology Research, 3(2), 16-24. https://doi.org/10.54097/8vyfba92