Engineering PapersSearch

SEARCH · Engineering Papers

Results for “Task scheduling”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Ensemble Simulations on Leadership Computing Systems

Scientific productivity can be enhanced through workflow management tools, relieving large High Performance Computing (HPC) system users from the tedious tasks of scheduling and designing the complex computational execution of scientific applications. This paper presents a study on the usage of ensemble workflow tools to accelerate science using the Summit and Frontier supercomputing systems. The research aims to connect science domain simulations using Oak Ridge Leadership Computing Facility (OLCF) supercomputing platforms with ensemble workflow methods in order to accelerate HPC-enabled discovery and boost scientific impact. We present the coupling, porting and optimization of Radical-Cybertools on three applications: Chroma, NAMD and LAMMPS. The tools augment traditional HPC monolithic runs with a pilot scheduler. Lessons-learned are discussed for physics, biology and materials science applications. We discuss intrinsic limitations of coupling and porting ensemble workflow tools to applications that run on large HPC systems. The origins of technical challenges and their solutions developed during the implementation process are discussed. Data management strategies, OLCF’s policies for ensembles, and natively supported workflow tools are also summarized.

Georgiadou, Antigoni [ORNL] (ORCID:000000020977631

Priority-BF: A Task Manager for Priority-Based Scheduling

The increasing demand for computational resources, particularly in High-Performance Computing environments, necessitates to rethink how we handle job scheduling strategies. This work addresses the challenge of managing concurrent jobs with differing priorities on overloaded parallel systems, where strict QoS constraints are often difficult for users to define. Our solution relies on a qualitative description of priorities and pulls from two key approaches: the Easy-BF algorithm and the Conservative Backfilling algorithms. This solution improves the response time for high-priority jobs by 50% without affecting the overall system utilization. We show its applicability in several critical scenarios such as High-Performance Computing (HPC) resource management and in-situ computing.

Gainaru, Ana [ORNL]

CHARM-SYCL & IRIS: A Tool Chain for Performance Portability on Extremely Heterogeneous Systems

Performance portability is becoming crucial as high-performance computing systems become increasingly heterogeneous. We have many options for CPUs and accelerators (e.g., GPUs) but also for non-Von Neumann architectures such as field-programmable gate arrays. This paper presents the CHARM-SYCL unified programming environment for multiple accelerator types as a performance-portable programming environment. It uses the IRIS library developed at Oak Ridge National Laboratory as the back end accelerator runtime. IRIS has a high-performance scheduler to distribute tasks across accelerators. This design allows us to run an application from the same source on multiple systems with multiple configurations. We provide three types of portability with CHARM-SYCL: Portable Workflow, Compiler and Runtime Portability, and Application and Performance Portability. We implement a Monte Carlo simulation benchmark code on the CHARM-SYCL execution environment and demonstrate that our programming environment can accommodate extremely heterogeneous systems.

Fujita, Norihisa

Reported Energy and Cost Savings from the DOE ESPC IDIQ Program: FY 2023

The objective of this work was to determine the realization rate of energy and cost savings from the U.S. Department of Energy’s (DOE’s) Energy Savings Performance Contract (ESPC) program based on information reported by the energy services companies (ESCOs) that are carrying out ESPC projects at federal sites. Information was extracted from 201 measurement and verification (M&V) reports covering 191 projects to determine reported, estimated, and guaranteed cost savings and the associated reported and estimated energy savings for the previous contract performance year. This report covers projects that had a performance year ending in fiscal year 2023, between October 1, 2022 and September 30, 2023, and had an M&V report issued. Additionally, the annual cost to perform M&V was extracted from the individual project Task Order (TO) Schedules.

29 ENERGY PLANNING, POLICY, AND ECONOMY

Quantum/AI Topology-Aware Latency-Adaptive HPC Workflow Scheduling Optimization

The growing demand for more powerful high-performance computing (HPC) systems has led to a steady rise in energy consumption by supercomputing worldwide. This study is focused on comparing our Application-Topology Mapper (ATMapper) to the popular Simple Linux Utility for Resource Management (SLURM) for the purpose of exploring methods that can further optimize job-scheduling within HPC systems. ATMapper is an Artificial-Intelligence based approach to job-scheduling that is currently being enhanced with quantum annealing (QA) to generate optimal schedules faster. We are applying QA to speedup our ATMapper process to achieve higher computing efficiency, thereby reducing HPC energy consumption. Here, we examine how four job-scheduling approaches perform in processor node assignment when using an example network architecture of 4 interconnected nodes. Using a specialized script, we are assessing the schedule of a computation flow with 11 interdependent tasks. The data movements among nodes were tracked to count for the number of interactions (network hops) between nodes needed to complete the tasks. The total number of hops and the job completion time were then used to quantify the efficiency of the different mapping approaches. In addition to SLURM, we also compare our ATMapper to the QA-enabled LBNL TIGER and the D-Wave Distributed Computing processor assignment approaches. The preliminary results showed that our topology-aware, latency-adaptive ATMapper is significantly more efficient when compared to the other scheduling approaches due to its load-imbalance network allocation. The scheduler displayed a computing efficiency of 53% by performing significantly fewer network hops than its alternatives. By reducing the number of hops, ATMapper was able to perform all 11 tasks by using only 3 nodes out of given 4. This research indicates the potential to use QA/AI for HPC job-scheduling. Later, we will test a SLURM simulator program to draw further comparisons on the effectiveness of ATMapper's scheduling approach. The results of this comparison will serve as a baseline for later improving SLURM's performance using a QA-enhanced ATMapper approach.

Caraveo, Braulio [University of Huston - Clear Lak

IRIS-GNN: Leveraging Graph Neural Networks for Scheduling on Truly Heterogeneous Runtime Systems

The diversity of accelerators in computer systems poses significant challenges for software developers, such as managing vendor-specific compiler toolchains, code fragmentation requiring different kernel implementations, and performance portability issues. To address these, the Intelligent Runtime System (IRIS) was developed. IRIS works across various systems, from smartphones to supercomputers, enabling automatic performance scaling based on available accelerators. It introduces abstract tasks for seamless execution transitions between accelerators while ensuring memory consistency and task dependencies. Although IRIS simplifies system details, optimal dynamic scheduling still requires user input to understand workload structures. To address this, we introduce a new scheduling policy for IRIS, termed IRIS-GNN, which is the first IRIS hybrid policy that operates in conjunction with the dynamic policies. This policy employs a Graph-Neural Network (GNN) to conduct Graph Classification of any task graphs submitted to IRIS. This GNN analyzes the structure and attributes of the task graph, categorizing it as either locality, concurrency, or mixed. This classification subsequently guides the selection of the dynamic policy used by IRIS. We provide a comparison of the performance of IRIS-GNN against the complete spectrum of IRIS’s dynamic policies, assess the overhead introduced by the GNN within this scheduling framework, and ultimately explore its practical application in real-world scenarios.

Johnston, Beau

E-Area Low-Level Waste Facility Inadvertent Human Intruder Limits and Doses in Support of the PA2022

This report documents the inadvertent human intruder (IHI) analysis for the E-Area Low-Level Waste Facility (ELLWF) at the Savannah River Site (SRS), near Aiken, South Carolina. This analysis supports the revised ELLWF Performance Assessment (PA), complying with the Department of Energy standard for operation of low-level waste disposal facilities (USDOE, 2017). The ELLWF is an operating waste disposal facility and is scheduled to continue accepting waste to 2065. One task of the revised PA is to establish waste inventory limits for the various disposal units at ELLWF. This is done by modeling future contaminant release and transport through applicable pathways to human receptors, comparing predicted doses per disposed curie with applicable performance measures, to obtain inventory limits which will assure that doses to receptors do not exceed performance measures. This report documents results of modeling future doses to one class of receptor, the inadvertent human intruder. It is assumed that after site closure, public knowledge of the site is lost, and IHIs will engage in activities on the ELLWF that will disrupt the closure cap, causing dose to the IHI. Following USDOE (2017), six different stylized exposure scenarios are considered, simulating activities by an IHI which could result in a radiological dose. The six scenarios are: • Acute – Basement Construction: IHI constructs a basement and encounters waste during excavation which is inadvertently mixed with clean soil and diluted. • Acute – Well Drilling: IHI drills a water well through waste and is exposed to drill cuttings mixed with clean soil that are brought to the surface. • Acute – Discovery: IHI begins constructing a basement but stops when encountering the riprap in the final closure cap and is exposed to photon radiation from unexcavated material residing in the undisturbed waste zone. • Chronic – Agriculture: Resident IHI is exposed to waste that was excavated for basement construction and mixed with native soil in the intruder’s vegetable garden. • Chronic – Post-Drilling: Resident IHI is exposed to waste from drill cuttings mixed with native soil and scattered in the garden area. • Chronic – Residential: Resident IHI is exposed to external radiation while in home located above waste with shielding provided by the concrete basement floor and any soil or engineered material remaining between the basement and waste. Dose calculations are performed using the SRNL Dose Toolkit (Aleman, 2023), following the approach of Smith et al (2019). Calculations are performed separately for 27 of the 33 disposal units (DUs) at ELLWF and are radionuclide specific. The results of the IHI analysis include: • Dose Factors: mrem per disposed curie (acute) and mrem/yr per disposed curie (chronic) for each parent radionuclide, for each DU. • Inventory Limits: in curies, for each parent radionuclide, for each DU. • Estimated Dose to IHI: mrem (acute) and mrem/yr (chronic), for each DU, given its projected closure inventory without inventory biases applied. Most DU-specific IHI inventory limits are in the range of 10 3 to 10 7 curies per nuclide. The lowest inventory limits are associated with gamma-emitters such as Sn-126, Ra-226, Th-232, and Cm-248. Radionuclides with short half-lives such as Pu-241, and nuclides which are pure beta emitters or which decay by electron capture, such as Ni-59 and Ni-63, have the highest limits. For the 27 evaluated DUs, predicted IHI doses are shown in Table ES-1. The maximum acute dose is 1.18 mrem, at ST23, much less than the DOE performance measure of 500 mrem (USDOE, 2017). The highest chronic dose is 37.2 mrem/yr at ST02, below the DOE performance measure of 100 mrem/yr. Also shown are estimated inventory sums of fractions (SOFs) at closure in 2065, for groundwater (GW) and IHI pathways. For each DU, the inventory is constrained by the GW pathway. For most DUs, the IHI SOFs are approximately 1000 times lower than the GW SOF values, and the IHI pathway does not drive risk for any disposal unit.

12 MANAGEMENT OF RADIOACTIVE AND NON-RADIOACTIVE W

Energy Improvements of Fire Station 71

Since 2018, the City of Shawnee, Kansas has completed two phases of the State of Kansas Facility Conservation Improvement Program (FCIP), an initiative that guarantees operational cost and energy savings through targeted construction improvements on City facilities and infrastructure. The City is currently in the third phase of this FCIP, where one of the projects included an investment in energy improvements for Fire Station 71 (FS 71). The City partnered with Navitas, an Energy Service Company (ESCO), to implement a Photovoltaic Solar Array on FS71. The purpose of this project was to invest in sustainable building improvements with Energy Conservation Measures (ECM) to bring cost savings to the City and to provide sustainable benefits to the residents of Shawnee. In the first task of the project, Navitas collaborated with the City of Shawnee and the Community Development Department to determine the optimal layout and schedule for the installation of the solar array on FS 71. In the second task of the project, Navitas installed the 99.8 kW DC Photovoltaic solar array system. This system installation comprised of racking, inverters, optimizers, load center, and disconnect, which were all installed at a total ECM price of $\$$247,948. The third task focused on start-up and commissioning of the array. Navitas installed a real-time data analytics information management system integrated with utility meters, which evaluates the operations of the utility system and verifies operation of equipment and ensures optimum operation for energy efficiency. In the final task of this project, this analytics system was used for monitoring and verification, which will continue to be used to evaluate the success of the project for the coming years. The primary goal of the project was to install the 99.8 kW DC PV solar array at FS 71 to demonstrate the viability of solar energy systems in essential municipal facilities. Fire stations are energy demanding structures, as they require a constant intake of power and have a high baseline energy usage. The success of solar arrays on a fire station exemplifies their energy efficiency and effectiveness and displays their potential for application on other city facilities. By installing a solar array at such a facility, the City sought not only to offset electricity usage but also to serve as a model for ECMs in other municipal facilities and infrastructure projects. From an economic standpoint, this project demonstrates the feasibility of renewable energy at the municipal level. The total project cost of $\$$247,948 was split evenly between city funds and award funding, minimizing financial risk while ensuring guaranteed long-term savings. Any excess savings that are beyond the guaranteed minimums remain with the city, which enables future investment in sustainable energy initiatives. This project provides many benefits to the public. In addition to reducing the environmental footprint of city operations, it lowers taxpayer-funded utility spending and improves the energy security of a critical facility. The knowledge gained from this implementation motivates the City to focus on similar efforts across other public facilities in future FCIP phases and other City projects.

14 SOLAR ENERGY

Q-IRIS: The Evolution of the IRIS Task-Based Runtime to Enable Classical-Quantum Workflows

Extreme heterogeneity in emerging HPC systems are starting to include quantum accelerators, motivating runtimes that can coordinate between classical and quantum workloads. We present a proof-of-concept hybrid execution framework integrating the IRIS asynchronous task-based runtime with the XACC quantum programming framework via the Quantum Intermediate Representation Execution Engine (QIR-EE). IRIS orchestrates multiple programs written in the quantum intermediate representation (QIR) across heterogeneous backends (including multiple quantum simulators), enabling concurrent execution of classical and quantum tasks. Although not a performance study, we report measurable outcomes through the successful asynchronous scheduling and execution of multiple quantum workloads. To illustrate practical runtime implications, we decompose a four-qubit circuit into smaller subcircuits through a process known as quantum circuit cutting, reducing per-task quantum simulation load and demonstrating how task granularity can improve simulator throughput and reduce queueing behavior -- effects directly relevant to early quantum hardware environments. We conclude by outlining key challenges for scaling hybrid runtimes, including coordinated scheduling, classical-quantum interaction management, and support for diverse backend resources in heterogeneous systems.

Miniskar, Narasinga Rao [ORNL] (ORCID:000000018259

Improving I/O-aware Workflow Scheduling via Data Flow Characterization and trade-off Analysis

The scientific computing paradigm has transitioned from compute-intensive to I/O-intensive and memory-intensive in the past decade, especially when data-driven science has become common practice. Numerous empirical I/O-aware scheduling optimizations have been developed by incorporating I/O capacity and bandwidth as constraints into scheduling. Unfortunately, there is a lack of data flow (I/O) characterization tool and an understanding of trade-offs between concurrency, locality, and I/O bandwidth. To bridge the gap, this work 1) presents a set of descriptors to characterize, organize, and visualize I/O profiles, including flow size, I/O bandwidth, and operation count, which group data flows by I/O types, tasks, and files; 2) proposes an I/O Roofline model-based trade-off analysis to find the optimal trade-off between flow operational intensity, concurrency, and flow performance. The I/O descriptors generate useful insights into complicated I/O behaviors, suggesting distinct concurrency, storage, and scheduling to be used by types, tasks, and files. The proposed trade-off analysis guides scheduling decisions that generate resource assignment with the best flow parallelism. We evaluate our I/O-aware scheduling methodology on a highly I/O-intensive workflow–1000 Genomes. The experimental results demonstrate speedups of up to 2.4× compared to the state-of-the- art methods.

Guo, Luanzheng [BATTELLE (PACIFIC NW LAB)]

Classic and Quantum Task-Based Intelligent Runtime for QIRs Running on Multiple QPUs

High-performance computing systems are rapidly evolving into heterogeneous platforms that fuse quantum accelerators with traditional classical processing units (CPUs) and graphical processing units (GPUs). This convergence calls for runtimes capable of managing both classical and quantum workloads in a unified manner. We introduce an intelligent, task-based runtime that marries the Intelligent RuntIme System (IRIS) asynchronous scheduler with a quantum programming stack through the Quantum Intermediate Representation Execution Engine (QIR-EE). Our design allows programs written in the quantum intermediate representation (QIR) to be dispatched concurrently to a variety of back-ends, including multiple quantum simulators and nascent quantum processors, enabling genuine hybrid execution on a single node. To illustrate its practicality, we partition a 4-qubit and 20-qubit circuit into three sub-circuits using quantum circuit cutting via the QCut library. Each sub-circuit is simulated independently by the QIR-EE driver within IRIS, after which a classical post-processing step merges the simulation results to recover the outcome of the original full-circuit computation. This case study demonstrates how finer task granularity can enable the parallel execution and lower the simulation burden per quantum task while preserving overall accuracy, highlighting the feasibility of our hybrid approach.

Miniskar, Narasinga Rao [ORNL] (ORCID:000000018259

ExaWorks software development kit: a robust and scalable collection of interoperable workflows technologies

Scientific discovery increasingly requires executing heterogeneous scientific workflows on high-performance computing (HPC) platforms. Heterogeneous workflows contain different types of tasks (e.g., simulation, analysis, and learning) that need to be mapped, scheduled, and launched on different computing. That requires a software stack that enables users to code their workflows and automate resource management and workflow execution. Currently, there are many workflow technologies with diverse levels of robustness and capabilities, and users face difficult choices of software that can effectively and efficiently support their use cases on HPC machines, especially when considering the latest exascale platforms. We contributed to addressing this issue by developing the ExaWorks Software Development Kit (SDK). The SDK is a curated collection of workflow technologies engineered following current best practices and specifically designed to work on HPC platforms. We present our experience with (1) curating those technologies, (2) integrating them to provide users with new capabilities, (3) developing a continuous integration platform to test the SDK on DOE HPC platforms, (4) designing a dashboard to publish the results of those tests, and (5) devising an innovative documentation platform to help users to use those technologies. Our experience details the requirements and the best practices needed to curate workflow technologies, and it also serves as a blueprint for the capabilities and services that DOE will have to offer to support a variety of scientific heterogeneous workflows on the newly available exascale HPC platforms.

97 MATHEMATICS AND COMPUTING

Validation Exercise of a Coarse Finite Element Model of Laser Welds

The objective of this project is to validate low-fidelity models of 304L to 304L stainless steel partial-penetration laser welds for thin sheets. Low-fidelity means that the weld is represented by coarsely meshed element blocks. Here, the hexahedral element size is approx imately half the weld penetration depth. The material behavior of the block is represented by a J2 plasticity model with a Voce hardening function. The source of the data used in this work is an extensive experimental study conducted by Sharlotte Kramer (1528) and published in 2015. Figure 1 shows a cross-section of the weld of interest. The nominal thickness of the sheets is 0.063 in. while the target penetration depth of the weld is in the range of 0.028 to 0.032 in., extending about half the sheet thickness. Uniaxial tension tests provided data for calibration of base material and weld models. Results of two validation geometries were also provided. The principal validation geometry is shown in Fig. 2. It consists of a plate specimen with in-plane dimensions 6 in × 2.875 in loaded in tension. A circular plug with a 1.5 in. diameter was cut from the center of the plate and then welded in place. The details of the welding schedule are given. An important assumption is that the welds in the calibration and validation specimens have similar geometric and material properties as those in the validation tests. The task was to first calibrate models for the base material and the welds and then simulate the validation tests until the point of weld first failure.

36 MATERIALS SCIENCE

The Myth of Fungible FTE: A Quantitative Assessment of Matrixed Resource Allocation

Matrix organizations allow scientific facilities to share specialized personnel across projects, operations, maintenance, and strategic initiatives. Nominal staffing allocations, however, may not capture the schedule consequences of fragmented individual commitments, limited access to specialist groups, and intermittent availability of key decision makers. We developed a stochas- tic, daily-time-step simulation of a hypothetical medium-sized accelerator-facility project com- prising sequential phases and parallel tasks. Each task requires role-specific work measured in FTE-days. Ordinary personnel may be unavailable because they contribute concurrently to other institutional activities, while designated key roles have independently specified daily un- availability probabilities. An organization-wide priority factor scales the number of people from each functional group who can effectively contribute to the project. It is interpreted as a composite proxy for project access and workforce fragmentation across competing commit- ments. We examined project completion time as a function of this factor and Project Lead unavailability using 100 Monte Carlo runs per condition. Increasing priority factor from 0.1 to 1.0 reduced median completion time from 1708.5 days (interquartile range 1681.5–1735.25) to 390 days (interquartile range 379–399). At priority factor = 0.1, increasing Project Lead unavailability from 0.5 to 0.9 increased median completion time from 1713.5 days (interquartile range 1691–1733.25) to 4,417 days (interquartile range 4271.75–4550.5). The model quantifies the commonly expected sensitivity of project schedules to fragmented resource commitments and limited coordination availability. Within this model, the results also indicate a possible threshold regime in which small increases in workforce availability yield only modest sched- ule improvements until sufficient capacity becomes accessible, after which project performance improves sharply. With further validation and calibration, this quantitative framework could support resource-allocation decisions during initial project planning and subsequent schedule rebaselining.

Bai, Mei [SLAC National Accelerator Laboratory (SL

Bridging Equipment Reliability Data and Risk Informed Decisions in a Plant Operation Context

Industry equipment reliability and asset management programs are essential elements that help ensure the safe and economical operation of nuclear power plants. The effectiveness of these programs is addressed in several industry-developed and regulatory programs. The Risk-Informed Asset Management (RIAM) project is tasked to develop tools in support of the equipment reliability and asset management programs at nuclear power plants. These tools are designed to create a direct bridge between component health/lifecycle data and decision making (e.g., maintenance scheduling and project prioritization). The goal of this article is to provide a guide for specific use cases that the RIAM project is targeting. We have grouped uses cases into three main areas. The first area focuses on the analysis of equipment reliability data with a particular emphasis on condition-based data, such as test/surveillance reports and component monitoring data. The second area focuses on the integration of equipment reliability into system/plant reliability models to determine system/plant health and identify the components that are critical to maintain an operational system. Lastly, the third area manages plant resources, such as maintenance activities and replacement scheduling using optimization methods. Here the primary focus is on supporting typical system engineer decisions regarding maintenance activity scheduling and component aging management. This is performed in a risk-informed context where the term “risk” is broadly constructed to include both plant reliability and economics. This framework combines data analytics tools to analyze equipment reliability data with risk-informed methods designed to support system engineer decisions (e.g., maintenance and replacement schedules, optimal maintenance posture) in a customizable workflow.

97 - MATHEMATICS AND COMPUTING

matsim-agents v1.0

matsim-agents is a multi-agent AI framework for atomistic materials simulation and discovery. It orchestrates large language models (LLMs), machine-learned interatomic potentials (MLIPs), and DFT codes into a single agentic loop running on laptops and DOE leadership-class supercomputers. MULTI-AGENT ORCHESTRATION A LangGraph state machine with three nodes: a Planner that converts a natural-language research objective into structured tasks; an Executor that dispatches atomistic tools and loops until the queue is empty; and an Analyst that summarizes results into a human-readable report. State is checkpointed after every step and human-in-the-loop gates can be inserted at any edge. HYPOTHESIS-DRIVEN DISCOVERY CHAT An interactive REPL (matsim-agents chat) that couples LLM dialogue with atomistic simulation. Chemical formulas are automatically detected in conversation turns and trigger a full crystal-phase exploration: structure generation → relaxation → stability scoring → result injection back into the conversation, creating a closed hypothesis-refinement loop. CRYSTAL PHASE ENUMERATION Given a composition, the phase explorer enumerates prototypes by stoichiometry: elemental (fcc/bcc/hcp/sc/diamond), binary 1:1 (rocksalt/CsCl/zincblende/ wurtzite/fluorite/rutile), ternary 1:1:3 (cubic perovskite), ternary 1:2:4 (perovskite + spinel), quaternary 1:1:2:6 (Fm-3m double perovskite). 2-D prototypes (graphene, h-BN, MoS2 2H/1T) and multilayer stacking are also supported via --include-2d and --num-layers. SUPERCELL GENERATION AND SITE DECORATION Auto-tiling to a minimum atom count (--min-atoms), explicit NxNxN tiling (--supercell), symmetry-distinct site decorations (--n-orderings), and isotropic lattice-scale sweeps (--lattice-scales) for volume bracketing. MLFF RELAXATION AND STABILITY SCORING HydraGNN (multi-headed GNN) drives structure relaxation via ASE with FIRE, BFGS, or BFGSLineSearch. Stability output: delta-E/atom ranking across phases and a max-residual-force dynamical-stability proxy. Other MLIPs (MACE, NequIP, Orb) can be plugged in through the same interface. DFT BACKENDS Quantum ESPRESSO pw.x and VASP 6.6 are first-class labellers. Both have validated GPU builds and SLURM/PBS launchers for three DOE platforms: Frontier (AMD MI250X, ROCm), Aurora (Intel PVC, oneAPI), Perlmutter (NVIDIA A100, CUDA). QE produces ~100 binaries (pw.x, ph.x, epw.x, ...). VASP supports scf, relax, vc-relax, and vc-relax-shape run types. ACTIVE-LEARNING LOOP matsim-agents al run CONFIG.yaml drives an iterative HydraGNN-DFT loop: MD generates candidates → ensemble/MC-dropout uncertainty selects the most informative → DFT labels them in parallel inside one allocation → dataset grows → HydraGNN retrains → repeat. DFT backend is a single YAML toggle (dft.backend: vasp | qe). LLM-generated seed structures are supported (no curated POSCAR library needed). Config uses ${VAR}, ${VAR:-default}, ${VAR:?msg} shell-style substitution for cross-user/cross-site portability. LLM BACKENDS Ollama (local, default), vLLM (HPC multi-GPU serving), OpenAI, Anthropic, HuggingFace Transformers+Accelerate. Selected at runtime via flag or env var with no code changes. HPC PORTABILITY Same Python entry points run on Frontier (ROCm 7.2), Aurora (oneAPI), and Perlmutter (CUDA 12). DFT and ML stacks are never co-loaded in the same shell; they couple through the scheduler and filesystem. Advanced multi-node launchers (serve, discovery-chat, single-relaxation, active-learning, QE warm-start) are provided for all three platforms. CODABENCH COMPETITION BUNDLE A self-contained benchmark: 159 atomistic test structures across 11 material classes, 5 tasks (formation energy, forces, ML relaxation, AI-DFT relaxation, phase stability ranking), public/private leaderboard split (30/70), and four ready-to-run baselines: MACE-MP-0, HydraGNN, UMA, AllScAIP.

Lupo Pasini, Massimiliano [Oak Ridge National Labo

iDDS: intelligent distributed dispatch and scheduling for workflow orchestration

The intelligent distributed dispatch and scheduling (iDDS) service is a versatile workflow orchestration system designed for large-scale, distributed scientific computing. iDDS extends traditional workload and data management by integrating data-aware execution, conditional logic, and programmable workflows, enabling automation of complex and dynamic processing pipelines. Originally developed for the ATLAS experiment at the large hadron collider, iDDS has evolved into an experiment-agnostic platform that supports both template-driven workflows and a Function-as-a-Task model for Python-based orchestration. This paper presents the architecture and core components of iDDS, highlighting its scalability, modular message-driven design, and integration with systems such as PanDA and Rucio. We demonstrate its versatility through real-world use cases: fine-grained tape resource optimization for ATLAS, orchestration of large Directed Acyclic Graph (DAG) workflows for the Rubin Observatory, distributed hyperparameter optimization for machine learning applications, active learning for physics analyses, and AI-assisted detector design at the electron–ion collider. By unifying workload scheduling, data movement, and adaptive decision-making, iDDS reduces operational overhead and enables reproducible, high-throughput workflows across heterogeneous infrastructures. We conclude with current challenges and future directions, including interactive, cloud-native, and serverless workflow support.

97 MATHEMATICS AND COMPUTING

Preparation of the Multi-Site Data Processing at the Vera C. Rubin Observatory

The Vera C. Rubin Observatory’s Legacy Survey of Space and Time (LSST) Camera is scheduled to start taking data in the summer of 2025. The Data Release Production will run the LSST Science Pipe software at data facilities in the US, France and the UK. The LSST Science Pipeline consists of complex directed acyclic graphs (DAGs) of tasks. Rubin will use the Production and Distributed Analysis (PanDA) workflow and workload management system to orchestrate this complex workflow and the distribution of workloads to the data facilities. When run end-to-end by a team of data production staff, this processing (the Science Pipelines, distributed by the workflow and workload management system) is referred to as a 'campaign'. This paper describes the central services and data facility specific services that support this multi-site data process model, including the service deployment infrastructure, the workload and workflow system, the Campaign Management tools, and connection to Rubin Data Management. This paper will also mention the experience of processing the Rubin Commissioning Camera data. All these are part of the effort to scale up the processing capabilities for the expected very large data volume from the LSST Camera.

Yang, Wei [SLAC]