The Counterfactual Digital Twin Portfolio Scheduling for Heterogeneous HPC and Multi-cloud Systems
Keywords:
counterfactual simulation, digital twin, high-performance computing, multi-cloud, Pareto optimizationAbstract
Selecting where to execute a computational job across institutional high-performance computing systems and public clouds is a decision under uncertainty. Queue delay, runtime, data-transfer time, monetary cost, accelerator availability, energy use, carbon intensity, failure probability, and policy constraints can change between submission and execution. Existing selection methods often assign a resource class or calculate a scalar fitness score from current descriptions. This paper proposes a different approach: counterfactual digital-twin portfolio scheduling. A compact digital twin represents each feasible execution path and predicts a joint outcome distribution rather than one resource label. The broker simulates alternative placements under the same workload fingerprint, removes solutions that violate hard constraints, constructs a robust Pareto set, and selects either one placement or a bounded hedge when prediction uncertainty is high. Observed queue, runtime, cost, transfer, energy, and failure outcomes update twin calibration without allowing an online model to override policy. The method supports complete jobs and workflow stages, data locality, reservations, spot interruption, and portability risk. A safe scheduling algorithm, architecture, equations, decision explanations, and a reproducible evaluation protocol are provided. The study does not claim unmeasured numerical improvement; final results must be generated through controlled execution on declared HPC and cloud resources. The proposed idea changes resource selection from fuzzy class prediction to auditable counterfactual decision-making under multiple objectives and uncertainty.
Downloads
References
[1] R. Tracey, M. O. Akinsolu, V. Elisseev, and Y. Vagapov, “Hybrid framework for resource allocation in heterogeneous HPC and cloud environments using fuzzy logic and machine learning: A comparison of classification models,” IEEE Access, vol. 14, 2026, doi: 10.1109/ACCESS.2026.3702420.
[2] A. B. Yoo, M. A. Jette, and M. Grondona, “SLURM: Simple Linux utility for resource management,” in Job Scheduling Strategies for Parallel Processing. Berlin, Germany: Springer, 2003, pp. 44–60.
[3] M. A. Jette and A. B. Yoo, “Slurm: Simple Linux utility for resource management,” in Active Middleware Services, 2002.
[4] A. W. Mu’alem and D. G. Feitelson, “Utilization, predictability, workloads, and user runtime estimates in scheduling the IBM SP2 with backfilling,” IEEE Trans. Parallel Distrib. Syst., vol. 12, no. 6, pp. 529–543, Jun. 2001.
[5] D. G. Feitelson, Workload Modeling for Computer Systems Performance Evaluation. Cambridge, U.K.: Cambridge Univ. Press, 2015.
[6] H. Topcuoglu, S. Hariri, and M.-Y. Wu, “Performance-effective and low-complexity task scheduling for heterogeneous computing,” IEEE Trans. Parallel Distrib. Syst., vol. 13, no. 3, pp. 260–274, Mar. 2002.
[7] J. J. Durillo and R. Prodan, “Multi-objective workflow scheduling in Amazon EC2,” Cluster Comput., vol. 17, pp. 169–189, 2014.
[8] E. Zitzler, M. Laumanns, and L. Thiele, “SPEA2: Improving the strength Pareto evolutionary algorithm,” ETH Zurich, Tech. Rep. 103, 2001.
[9] K. Deb, Multi-Objective Optimization Using Evolutionary Algorithms. Chichester, U.K.: Wiley, 2001.
[10] E. Negri, L. Fumagalli, and M. Macchi, “A review of the roles of digital twin in CPS-based production systems,” Procedia Manuf., vol. 11, pp. 939–948, 2017.
[11] F. Tao, H. Zhang, A. Liu, and A. Y. C. Nee, “Digital twin in industry: State-of-the-art,” IEEE Trans. Ind. Informat., vol. 15, no. 4, pp. 2405–2415, Apr. 2019.
[12] J. Snoek, H. Larochelle, and R. P. Adams, “Practical Bayesian optimization of machine learning algorithms,” in Adv. Neural Inf. Process. Syst., vol. 25, 2012.
[13] C. E. Rasmussen and C. K. I. Williams, Gaussian Processes for Machine Learning. Cambridge, MA, USA: MIT Press, 2006.
[14] A. N. Angelopoulos and S. Bates, “A gentle introduction to conformal prediction and distribution-free uncertainty quantification,” Found. Trends Mach. Learn., vol. 16, no. 4, pp. 494–591, 2023.
[15] J. Pearl, Causality, 2nd ed. Cambridge, U.K.: Cambridge Univ. Press, 2009.
[16] D. Bertsimas and M. Sim, “The price of robustness,” Oper. Res., vol. 52, no. 1, pp. 35–53, 2004.
[17] A. D. Ferguson et al., “Participatory networking: An API for application control of SDNs,” in Proc. ACM SIGCOMM, 2013, pp. 327–338.
[18] A. Verma et al., “Large-scale cluster management at Google with Borg,” in Proc. EuroSys, 2015, pp. 1–17.
[19] B. Burns, B. Grant, D. Oppenheimer, E. Brewer, and J. Wilkes, “Borg, Omega, and Kubernetes,” Commun. ACM, vol. 59, no. 5, pp. 50–57, May 2016.
[20] A. Ghodsi et al., “Dominant resource fairness,” in Proc. USENIX NSDI, 2011, pp. 323–336.
[21] L. A. Barroso, U. Hölzle, and P. Ranganathan, The Datacenter as a Computer, 3rd ed. Cham, Switzerland: Springer, 2018.
[22] M. Armbrust et al., “A view of cloud computing,” Commun. ACM, vol. 53, no. 4, pp. 50–58, Apr. 2010.
[23] P. Mell and T. Grance, The NIST Definition of Cloud Computing, NIST Special Publication 800-145, Sep. 2011.
[24] B. Beyer, C. Jones, J. Petoff, and N. R. Murphy, Eds., Site Reliability Engineering. Sebastopol, CA, USA: O’Reilly Media, 2016.
[25] D. Sculley et al., “Hidden technical debt in machine learning systems,” in Adv. Neural Inf. Process. Syst., vol. 28, 2015.
[26] A. Dean and D. Voss, Design and Analysis of Experiments, 2nd ed. New York, NY, USA: Springer, 2017.



