A Self-Evolving Agentic Reinforcement Meta-Learning Framework for Adaptive Cloud Scheduling under Dynamic Workloads
Keywords:
Argentic AI, Reinforcement, Learning; Meta-Learning, Goal Adaptation, Cloud Scheduling, Deep Q-Network, Multi-Objective Optimization, Dynamic WorkloadsAbstract
Cloud platforms operate under fluctuating workloads, heterogeneous resource demands, and changing service priorities. Effective scheduling must therefore balance latency, operational cost, resource utilization, throughput, and energy consumption. Conventional schedulers follow predetermined rules, while many reinforcement-learning methods use reward functions whose objective weights remain fixed during training and deployment. These restrictions can produce inefficient decisions when workload conditions or performance priorities change. This study introduces a self-evolving argentic reinforcement meta-learning framework for adaptive cloud scheduling. Its dual-level architecture integrates an action-learning layer with a goal-evolution layer. At the lower level, a goal-conditioned Deep Q-Network selects task-to-resource assignments from observed system states, including queue length, task requirements, and available CPU, GPU, and memory capacity. At the upper level, a meta-learning mechanism examines long-term performance feedback and adjusts the importance assigned to individual scheduling objectives. The interaction between both layers allows the scheduler to update its decision policy and optimization priorities without manually redesigning the reward function. The framework is evaluated in a simulated heterogeneous cloud environment using workload characteristics derived from Alibaba Cluster Trace data. Performance is compared with First-Come, First-Served, Round Robin, heuristic scheduling, and a conventional Deep Q-Network. The reported results show that the proposed scheduler achieves 69 ms average latency, an operational cost of $61, 94% resource utilization, throughput of 86 tasks per second, and energy consumption of 53 kWh. Compared with First-Come, First-Served, latency decreases by 50.7%, cost by 47.0%, and energy use by 44.2%, while throughput increases by 72.0%. Resource utilization rises from 60% to 94%, representing a 34-percentage-point improvement. Ablation results indicate that dynamic goal evolution performs better than static multi-objective weighting. Overall, jointly adapting actions and objective priorities offers a promising approach for responsive and efficient cloud scheduling, although independent validation on operational cloud infrastructure remains necessary across broader workloads, platforms, and diverse resource configurations
Downloads
References
[1] Sutton, R. S., & Barto, A. G. (2018). Reinforcement learning: An introduction (2nd ed.). MIT Press.
[2] Watkins, C. J. C. H., & Dayan, P. (1992). Q-learning. Machine Learning, 8(3–4), 279–292.
[3] Mnih, V., Kavukcuoglu, K., Silver, D., et al. (2015). Human-level control through deep reinforcement learning. Nature, 518(7540), 529–533.
[4] Mnih, V., Kavukcuoglu, K., Silver, D., et al. (2013). Playing Atari with deep reinforcement learning. arXiv preprint arXiv:1312.5602.
[5] Van Hasselt, H., Guez, A., & Silver, D. (2016). Deep reinforcement learning with Double Q-learning. Proceedings of AAAI Conference on Artificial Intelligence, 30(1), 2094–2100.
[6] Wang, Z., Schaul, T., Hessel, M., Van Hasselt, H., Lanctot, M., & De Freitas, N. (2016). Dueling network architectures for deep reinforcement learning. Proceedings of ICML, 1995–2003.
[7] Schaul, T., Quan, J., Antonoglou, I., & Silver, D. (2016). Prioritized experience replay. Proceedings of ICLR.
[8] Bellemare, M. G., Dabney, W., & Munos, R. (2017). A distributional perspective on reinforcement learning. Proceedings of ICML, 449–458.
[9] Hessel, M., Modayil, J., Van Hasselt, H., et al. (2018). Rainbow: Combining improvements in deep reinforcement learning. Proceedings of AAAI Conference on Artificial Intelligence, 32(1).
[10] Lillicrap, T. P., Hunt, J. J., Pritzel, A., et al. (2016). Continuous control with deep reinforcement learning. Proceedings of ICLR.
[11] Schulman, J., Levine, S., Abbeel, P., Jordan, M., & Moritz, P. (2015). Trust region policy optimization. Proceedings of ICML, 1889–1897.
[12] Schulman, J., Wolski, F., Dhariwal, P., Radford, A., & Klimov, O. (2017). Proximal policy optimization algorithms. arXiv preprint arXiv:1707.06347.
[13] Silver, D., Lever, G., Heess, N., et al. (2014). Deterministic policy gradient algorithms. Proceedings of ICML, 387–395.
[14] Konda, V. R., & Tsitsiklis, J. N. (2000). Actor-critic algorithms. Proceedings of NeurIPS, 1008–1014.
[15] Tsitsiklis, J. N., & Van Roy, B. (1997). An analysis of temporal-difference learning with function approximation. IEEE Transactions on Automatic Control, 42(5), 674–690.
[16] Bertsekas, D. P. (2019). Reinforcement learning and optimal control. Athena Scientific.
[17] Kaelbling, L. P., Littman, M. L., & Moore, A. W. (1996). Reinforcement learning: A survey. Journal of Artificial Intelligence Research, 4, 237–285.
[18] Busoniu, L., Babuska, R., De Schutter, B., & Ernst, D. (2010). Reinforcement learning and dynamic programming using function approximators. CRC Press.
[19] Szepesvári, C. (2010). Algorithms for reinforcement learning. Morgan & Claypool.
[20] Silver, D., Huang, A., Maddison, C. J., et al. (2016). Mastering the game of Go with deep neural networks and tree search. Nature, 529(7587), 484–489.
[21] Finn, C., Abbeel, P., & Levine, S. (2017). Model-agnostic meta-learning for fast adaptation of deep networks. Proceedings of ICML, 1126–1135.
[22] Ravi, S., & Larochelle, H. (2017). Optimization as a model for few-shot learning. Proceedings of ICLR.
[23] Andrychowicz, M., Denil, M., Gomez, S., et al. (2016). Learning to learn by gradient descent by gradient descent. Proceedings of NeurIPS, 3981–3989.
[24] Hospedales, T., Antoniou, A., Micaelli, P., & Storkey, A. (2021). Meta-learning in neural networks: A survey. IEEE Transactions on Pattern Analysis and Machine Intelligence, 44(9), 5149–5169.
[25] Wang, J. X., Kurth-Nelson, Z., Tirumala, D., et al. (2016). Learning to reinforcement learn. arXiv preprint arXiv:1611.05763.
[26] Sutton, R. S., Precup, D., & Singh, S. (1999). Between MDPs and semi-MDPs: A framework for temporal abstraction in reinforcement learning. Artificial Intelligence, 112(1–2), 181–211.
[27] Dietterich, T. G. (2000). Hierarchical reinforcement learning with the MAXQ value function decomposition. Journal of Artificial Intelligence Research, 13, 227–303.
[28] Vezhnevets, A. S., Osindero, S., Schaul, T., et al. (2017). FeUdal networks for hierarchical reinforcement learning. Proceedings of ICML, 3540–3549.
[29] Nachum, O., Gu, S., Lee, H., & Levine, S. (2018). Data-efficient hierarchical reinforcement learning. Proceedings of NeurIPS, 3303–3313.
[30] Levy, A., Konidaris, G., Platt, R., & Saenko, K. (2019). Hierarchical reinforcement learning with hindsight. Proceedings of ICLR.
[31] Jennings, B., & Stadler, R. (2015). Resource management in clouds: Survey and research challenges. Journal of Network and Systems Management, 23, 567–619.
[32] Singh, S., & Chana, I. (2016). Cloud resource provisioning: Survey, status and future research directions. Knowledge and Information Systems, 49, 1005–1069.
[33] Manvi, S. S., & Shyam, G. K. (2014). Resource management for IaaS in cloud computing: A survey. Journal of Network and Computer Applications, 41, 424–440.
[34] Madni, S. H. H., Abd Latiff, M. S., Coulibaly, Y., & Abdulhamid, S. M. (2017). Resource allocation techniques in cloud computing: A review. Cluster Computing, 20, 2489–2533.
[35] Calheiros, R. N., Ranjan, R., Beloglazov, A., et al. (2011). CloudSim: A toolkit for modeling and simulation. Software: Practice and Experience, 41(1), 23–50.
[36] Mao, H., Alizadeh, M., Menache, I., & Kandula, S. (2016). Resource management with deep reinforcement learning. Proceedings of HotNets, 50–56.
[37] Mao, H., Schwarzkopf, M., Venkatakrishnan, S. B., et al. (2019). Learning scheduling algorithms for data processing clusters. Proceedings of SIGCOMM, 270–288.
[38] Xu, C., Rao, J., & Bu, X. (2012). A unified reinforcement learning approach for cloud management. Journal of Parallel and Distributed Computing, 72(2), 95–105.
[39] Tesauro, G., Das, R., Chan, H., et al. (2007). Reinforcement learning for system optimization. Proceedings of NeurIPS, 1497–1504.
[40] Barrett, E., Howley, E., & Duggan, J. (2013). RL for cloud resource allocation. Concurrency and Computation: Practice and Experience, 25(12), 1656–1674.
[41] Schaul, T., Horgan, D., Gregor, K., & Silver, D. (2015). Universal value function approximators. Proceedings of the 32nd International Conference on Machine Learning, 1312–1320.
[42] Andrychowicz, M., Wolski, F., Ray, A., et al. (2017). Hindsight experience replay. Advances in Neural Information Processing Systems, 30, 5048–5058.
[43] Pateria, S., Subagdja, B., Tan, A.-h., & Quek, C. (2021). Hierarchical reinforcement learning: A comprehensive survey. ACM Computing Surveys, 54(5), 1–35.
[44] Barreto, A., Dabney, W., Munos, R., et al. (2017). Successor features for transfer in reinforcement learning. Advances in Neural Information Processing Systems, 30, 4055–4065.
[45] Kulkarni, T. D., Narasimhan, K., Saeedi, A., & Tenenbaum, J. (2016). Hierarchical deep reinforcement learning: Integrating temporal abstraction and intrinsic motivation. Advances in Neural Information Processing Systems, 29, 3675–3683.
[46] Florensa, C., Held, D., Wulfmeier, M., Zhang, M., & Abbeel, P. (2018). Reverse curriculum generation for reinforcement learning. Proceedings of the Conference on Robot Learning, 482–495.
[47] Achiam, J., Edwards, H., Amodei, D., & Abbeel, P. (2018). Variational option discovery algorithms. arXiv preprint arXiv:1807.10299.
[48] Haarnoja, T., Tang, H., Abbeel, P., & Levine, S. (2017). Reinforcement learning with deep energy-based policies. Proceedings of the 34th International Conference on Machine Learning, 1352–1361.
[49] Haarnoja, T., Zhou, A., Abbeel, P., & Levine, S. (2018). Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor. Proceedings of the 35th International Conference on Machine Learning, 1861–1870.
[50] Nachum, O., Norouzi, M., Xu, K., & Schuurmans, D. (2017). Bridging the gap between value and policy based reinforcement learning. Advances in Neural Information Processing Systems, 30, 2775–2785.
[51] Xiao, Z., Song, W., & Chen, Q. (2013). Dynamic resource allocation using virtual machines for cloud computing environment. IEEE Transactions on Parallel and Distributed Systems, 24(6), 1107–1117.
[52] Xu, J., & Fortes, J. A. B. (2010). Multi-objective virtual machine placement in virtualized data centers. Proceedings of IEEE Green Computing Conference, 179–188.
[53] Beloglazov, A., & Buyya, R. (2012). Energy-efficient management of data centers. Concurrency and Computation: Practice and Experience, 24(13), 1397–1420.
[54] Ghosh, R., Naik, V. K., & Trivedi, K. S. (2011). Power-performance trade-offs in cloud computing. Proceedings of IEEE DSN Workshops, 1–6.
[55] Wood, T., Shenoy, P., Venkataramani, A., & Yousif, M. (2009). Sandpiper: Black-box resource management for virtual machines. Computer Networks, 53(17), 2923–2938.
[56] Verma, A., Ahuja, P., & Neogi, A. (2008). pMapper: Power-aware application placement. Proceedings of Middleware, 243–264.
[57] Islam, S., Keung, J., Lee, K., & Liu, A. (2012). Adaptive resource provisioning in cloud. Future Generation Computer Systems, 28(1), 155–162.
[58] Lorido-Botran, T., Miguel-Alonso, J., & Lozano, J. A. (2014). Auto-scaling techniques in cloud environments. Journal of Grid Computing, 12, 559–592.
[59] Mao, M., Li, J., & Humphrey, M. (2010). Cloud auto-scaling with constraints. Proceedings of IEEE GRID, 41–48.
[60] Zaharia, M., Borthakur, D., Sen Sarma, J., et al. (2010). Delay scheduling for cluster computing. Proceedings of EuroSys, 265–278.
[61] Hindman, B., Konwinski, A., Zaharia, M., et al. (2011). Mesos: Resource sharing in data centers. Proceedings of NSDI, 295–308.
[62] Ousterhout, K., Wendell, P., Zaharia, M., & Stoica, I. (2013). Sparrow: Low latency scheduling. Proceedings of SOSP, 69–84.
[63] Reiss, C., Wilkes, J., & Hellerstein, J. L. (2011). Google cluster-usage traces. Google Technical Report.
[64] Tirmazi, Y., Lao, R., Harchol-Balter, M., et al. (2020). Borg: Cluster management at scale. Proceedings of EuroSys.
[65] Armbrust, M., Fox, A., Griffith, R., et al. (2010). A view of cloud computing. Communications of the ACM, 53(4), 50–58.
[66] Buyya, R., Yeo, C. S., Venugopal, S., et al. (2009). Cloud computing: Vision and architecture. Future Generation Computer Systems, 25(6), 599–616.
[67] Mao, H., Alizadeh, M., Menache, I., & Kandula, S. (2016). Deep reinforcement learning for resource management. Proceedings of HotNets, 50–56.
[68] Mao, H., Schwarzkopf, M., Venkatakrishnan, S. B., et al. (2019). Learning scheduling algorithms. Proceedings of SIGCOMM, 270–288.
[69] Rjoub, G., Wahab, O. A., Bataineh, A., et al. (2021). Deep reinforcement learning for task scheduling. Future Internet, 13(9), 230.
[70] Jiao, J., Zhou, Y., Li, H., & Luo, X. (2023). DRL-based workflow scheduling. Expert Systems with Applications, 213, 118997.
[71] Liu, C., Zou, C., Wu, P., et al. (2017). Hybrid RL scheduling in cloud. Proceedings of ICCCS.
[72] Choudhary, A., Rana, N., & Matahai, K. (2017). Energy-efficient scheduling in cloud. Procedia Computer Science, 78, 132–138.
[73] Deb, K. (2001). Multi-objective optimization using evolutionary algorithms. Wiley.
[74] Marler, R. T., & Arora, J. S. (2004). Survey of multi-objective optimization methods. Structural and Multidisciplinary Optimization, 26, 369–395.
[75] Coello Coello, C. A. (2006). Evolutionary multi-objective optimization. IEEE Computational Intelligence Magazine, 1(1), 28–36.
[76] Beloglazov, A., Abawajy, J., & Buyya, R. (2012). Energy-aware resource allocation. Future Generation Computer Systems, 28(5), 755–768.
[77] Mashayekhy, L., Nejad, M. M., Grosu, D., et al. (2015). Energy-aware scheduling of workflows. Future Generation Computer Systems, 43–44, 87–97.
[78] Arabnejad, H., & Barbosa, J. G. (2014). Scheduling in heterogeneous systems. IEEE Transactions on Parallel and Distributed Systems, 25(3), 682–694.



