Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “job execution”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

pyRMG: A framework for high-throughput, large-cell DFT calculations on supercomputers

Exascale computing delivers the raw power to simulate ever larger and more chemically realistic systems, but realizing this potential requires codes that can efficiently use thousands of processors. Our real-space multigrid (RMG) density functional theory (DFT) code’s grid-decomposition approach scales nearly linearly with the number of graphics processing units (GPUs), even for simulations exceeding thousands of atoms. This scalability makes RMG a compelling tool for high-throughput DFT studies of materials that would otherwise be bottlenecked in other codes (for example, by global fast Fourier transforms in plane-wave DFT). However, the limited workflow infrastructure for RMG has thus far constrained its adoption to a small user community. In this work, we present pyRMG, a Python package designed to streamline the setup and execution of RMG DFT calculations. Built on the pymatgen and ASE (Atomic Simulation Environment) computational materials science Python packages, pyRMG automates input generation and convergence checking, and it integrates with modern job schedulers (e.g., Flux) on leadership-class platforms such as Frontier and Perlmutter. Here, we demonstrate pyRMG for a high-throughput study of strain effects in 2D 2L-Bi 2 Se 3 /2L-NbSe 2 heterostructures, which offers chemical insights into this system and shows that RMG-based workflows can converge with limited user intervention.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

An end-to-end workflow for executing a classically bootstrapped variational quantum algorithm on an academic quantum computer

Academic quantum computing platforms often face unique challenges in executing quantum workloads due to fragmented software environments and limited engineering support. Unlike commercial ecosystems, academic devices typically evolve without full-stack integration in mind, making it difficult to run complex applications—such as variational quantum algorithms (VQA)—reliably and efficiently. Issues such as incompatible software layers and lack of automated job management significantly increase the overhead of theory-experiment collaboration. To address these challenges, we develop a modular, end-to-end workflow that decouples application-layer code from low-level hardware control, automates circuit submission and result collection, and supports fine-grained circuit-level job scheduling and recovery. The architecture employs a dual-end application programming interface (API) design, enabling robust operation across unstable or resource-constrained hardware backends. For practical use, the framework is lightweight and user-friendly, allowing rapid prototyping of full-stack workflows using basic Python tools. We validate this workflow on a high-fidelity trapped-ion quantum computer by demonstrating a variational quantum eigensolver (VQE) experiment with a classically bootstrapped ansatz initialization technique. The system successfully executed over 60,000 circuits across multiple molecular test cases with minimal human intervention, highlighting the framework’s effectiveness in enabling reproducible, resilient quantum experimentation in academic settings.

Clifford↗

Evaluating HPC Scheduling Strategies for Urgent Workloads

Scientific computing centers increasingly face workloads with diverse urgency requirements, driven by applications that demand rapid or even immediate execution. Appropriately configured scheduling policies can significantly improve both user satisfaction and overall cluster utilization. In this work, we present a systematic analysis of scheduler configurations under scenarios where a fraction of jobs have urgent computing needs. We evaluate multiple job scheduling simulators, develop a lightweight job-submission emulation framework, and create tools to analyze and visualize the resulting scheduling data. Our study identifies key trade-offs between responsiveness, fairness, and efficiency, and offers a set of practical scheduling configurations (particularly for Slurm) that can be tailored to HPC environments supporting mixed-urgency workloads.

Maheshwari, Ketan [ORNL] (ORCID:000000033800662X)↗

Performance Characterization and Provenance of Distributed Task-based Workflows on HPC Platforms

Understanding performance and provenance of task-based workflows poses significant challenges, particularly in distributed configurations where resources are shared by multiple applications. Task-based workflow management systems further complicate performance predictability because of their dynamicity that subtly alters task execution order from run to run. In this paper we propose a layered characterization framework for performance and task provenance for Dask.distributed workflows running on high-performance computing (HPC) platforms. It collects data from jobs, the workflow management system, and the operating system to aid in understanding the performance of these workflows. Our approach encompasses three main contributions: first, an extension of Dask.distributed to capture high-fidelity task provenance using Mochi data services; second, the adaptation of the established HPC I/O characterization tool Darshan to gather high-fidelity I/O data, thereby enhancing the granularity of our analysis; and third, a framework to combine and process the collected data and provide helpful insights into performance characterization and reproducibility, alongside our lessons learned.

Dask↗

Oak Ridge National Laboratory FY 2023 Site Sustainability Plan With FY 2022 Performance Data

At the close of each fiscal year, the US Department of Energy (DOE) Sustainability Performance Division (SPD) issues guidance documents and technical resource aids/tools necessary for DOE sites and national laboratories to complete sustainability reporting requirements. SPD is part of the DOE Office of Asset Management. As required by DOE Order 436.1, Departmental Sustainability, “each site will develop and commit to an annual Site Sustainability Plan (SSP) that identifies its respective contribution toward meeting the DOE’s sustainability goals.” SPD collects and compiles information reported by each site to develop an agency-wide Sustainability Report and Implementation Plan, which is used to report DOE sustainability progress to the federal government as required by all major federal agencies. DOE launched a formal Sustainability Office and annual SSP process in 2011. Each year, Oak Ridge National Laboratory (ORNL), in concert with the Office of Science (SC), provides the resources essential to fulfill its commitment to deliver a complete and accurate SSP report and quality performance data for entry into the DOE Sustainability Dashboard as managed by SPD. The performance data entered by each DOE site are then combined to disclose the progress of each DOE Program Office and are further combined to show comprehensive progress for the agency. The Office of Asset Management provides assistance to program offices in sustaining their missions, freeing up resources by reducing waste, avoiding excess expenditure on utilities, maximizing productivity, and improving the efficiency of facilities and processes. By focusing on mission needs, programs and associated DOE sites can help the agency meet its sustainability goals, as outlined in federal statutory and regulatory requirements. In FY 2022, the SSP guidance was updated to capture requirements from Executive Order (EO) 14008, Tackling the Climate Crisis at Home and Abroad, the Energy Act of 2020 (EAct 20), actions outlined in DOE’s Climate Adaptation & Resilience Plan and Sustainability Plan, and EO 14057, Catalyzing Clean Energy Industries and Jobs Through Federal Sustainability. Updates in SSP guidance help to minimize and streamline reporting while simultaneously addressing updated federal requirements. Per DOE, each SSP report should provide an overview of the site’s planned actions, as well as an overview of efforts and accomplishments during the reporting period. SPD collects and compiles information reported by each site to develop DOE’s Annual Sustainability Report, Climate Adaptation & Resilience Plan, and Annual Energy Management Report to Congress. The agency goal has been to lower the reporting burden for sites and increase and improve the consistency of information available to decision makers, allowing them to better identify projects and potential for increased efficiency, as well as to reduce waste, lower emissions, and enhance operational resilience. Sites may elect to produce a more polished publication for their leadership and stakeholders, but this step is no longer required. The ORNL SSP narrative report (this document) and the reporting of DOE SPD Sustainability Dashboard performance data is a collaborative effort of approximately 30 subject matter experts (SMEs) from ORNL facility management and research divisions. Annually, these associates come together to provide a report that can be used by DOE to demonstrate continued agency progress in energy efficiency and sustainable federal operations.

99 GENERAL AND MISCELLANEOUS↗

Producing High-fidelity Synthetic Population Ensembles at Scale

Used within social simulations, synthetic population ensembles enable uncertainty quantification (UQ) methods for obtaining more robust model inference and prediction. A synthetic population ensemble is a series of plausible virtual reconstructions of an area’s population at the granularity of people and residences, generated stochastically to preserve privacy of the source population survey’s respondents. In this paper, we demonstrate the production of large synthetic population ensembles for the US via Oak Ridge National Laboratory’s UrbanPop framework to support modeling of high spatial resolution energy affordability metrics from nationwide social surveys in collaboration with the fusionACS project. Our initial task involves creating ensembles for 17 US metropolitan areas, each consisting of 41 population instances (a base realization and 40 replicates). To accomplish this task at scale, we configured an integrated system comprised of a research cloud, virtual containerization, GPU-enhanced functionality, and a dual API/CLI to interact with UrbanPop’s maturing Likeness Python ecosystem. We observe a reduction in theoretical execution time while maintaining high-fidelity approximations of residential totals by metropolitan area and the demographic characteristics of neighborhoods. We discuss expansion of our approach to produce synthetic population ensembles for the entire US, particularly plans to establish automated workflows for job orchestration to increase computational efficiency, as well as provide outlook for broadening applications of the ensembles.

Gaboardi, James [ORNL] (ORCID:0000000247766826)↗

Generalizable coordination of large multiscale workflows: challenges and learnings at scale

The advancement of machine learning techniques and the heterogeneous architectures of most current supercomputers are propelling the demand for large multiscale simulations that can automatically and autonomously couple diverse components and map them to relevant resources to solve complex problems at multiple scales. Nevertheless, despite the recent progress in workflow technologies, current capabilities are limited to coupling two scales. In the first-ever demonstration of using three scales of resolution, we present a scalable and generalizable framework that couples pairs of models using machine learning and in situ feedback. We expand upon the massively parallel Multiscale Machine-Learned Modeling Infrastructure (MuMMI), a recent, award-winning workflow, and generalize the framework beyond its original design. We discuss the challenges and learnings in executing a massive multiscale simulation campaign that utilized over 600,000 node hours on Summit and achieved more than 98% GPU occupancy for more than 83% of the time. We present innovations to enable several orders of magnitude scaling, including simultaneously coordinating 24,000 jobs, and managing several TBs of new data per day and over a billion files in total. Finally, we describe the generalizability of our framework and, with an upcoming open-source release, discuss how the presented framework may be used for new applications.

Bhatia, Harsh↗

X-composer: enabling cross-environments in-situ workflows between HPC and cloud

As large-scale scientific simulations and big data analyses become more popular, it is increasingly more expensive to store huge amounts of raw simulation results to perform post-analysis. To minimize the expensive data I/O, "in-situ" analysis is a promising approach, where data analysis applications analyze the simulation generated data on the fly without storing it first. However, it is challenging to organize, transform, and transport data at scales between two semantically different ecosystems due to the distinct software and hardware difference. To tackle these challenges, we design and implement the X-Composer framework. X-Composer connects cross-ecosystem applications to form an "in-situ" scientific workflow, and provides a unified approach and recipe for supporting such hybrid in-situ workflows on distributed heterogeneous resources. X-Composer reorganizes simulation data as continuous data streams and feeds them seamlessly into the Cloud-based stream processing services to minimize I/O overheads. For evaluation, we use X-Composer to set up and execute a cross-ecosystem workflow, which consists of a parallel Computational Fluid Dynamics simulation running on HPC, and a distributed Dynamic Mode Decomposition analysis application running on Cloud. Our experimental results show that X-Composer can seamlessly couple HPC and Big Data jobs in their own native environments, achieve good scalability, and provide high-fidelity analytics for ongoing simulations in real-time.

Wang, Dali↗

Critical Materials Supply Chain White Paper (April 2020)

The DOE Office of Energy Efficiency and Renewable Energy (EERE) Advanced Manufacturing Office (AMO) partners with industry, small business, universities, and other stakeholders to identify and invest in emerging technologies with the potential to create high-quality domestic manufacturing jobs and enhance the global competitiveness of the United States

Advanced Manufacturing Office↗

Unveiling User Behavior on Summit Login Nodes as a User

We observe and analyze usage of the login nodes of the leadership class Summit supercomputer from the perspective of an ordinary user—not a system administrator—by periodically sampling user activities (job queues, running processes, etc.) for two full years (2020–2021). Our findings unveil key usage patterns that evidence misuse of the system, including gaming the policies, impairing I/O performance, and using login nodes as a sole computing resource. Our analysis highlights observed patterns for the execution of complex computations (workflows), which are key for processing large-scale applications.

Wilkinson, Sean↗

Enabling machine learning-ready HPC ensembles with Merlin

With the growing complexity of computational and experimental facilities, many scientific researchers are turning to machine learning (ML) techniques to analyze large scale ensemble data. With complexities such as multi-component workflows, heterogeneous machine architectures, parallel file systems, and batch scheduling, care must be taken to facilitate this analysis in a high performance computing (HPC) environment. Here, we present Merlin, a workflow framework to enable large ML-friendly ensembles of scientific HPC simulations. By augmenting traditional HPC with distributed compute technologies, Merlin aims to lower the barrier for scientific subject matter experts to incorporate ML into their analysis. As a producer–consumer workflow model, Merlin enables multi-machine, cross-batch job, dynamically allocated yet persistent workflows capable of utilizing surge-compute resources. Key features of Merlin are a flexible HPC-centric interface, low per-task overhead, multi-tiered fault recovery, and a hierarchical sampling algorithm that allows for $\mathscr{O}$(N) task execution and $\mathscr{O}$(N ln N) task queuing to ensembles of millions of tasks. In addition to Merlin’s design, we test the algorithm’s performance in an HPC center and demonstrate the ability to enqueue 40 million simulations in 100 s, with a 30 millisecond per-task overhead that is independent of ensemble size. Finally, we describe some example applications that Merlin has enabled on leadership-class HPC resources, such as the ML-augmented optimization of nuclear fusion experiments and the calibration of infectious disease models to study the progression of and possible mitigation strategies for COVID-19.

97 MATHEMATICS AND COMPUTING↗

Implementing interdisciplinary sustainability education with the food-energy-water (FEW) nexus

Growth in the green jobs sector has increased demand for college graduates who are prepared to enter the workforce with interdisciplinary sustainability skills. Simultaneously, scholarly calls for interdisciplinary collaboration in the service of addressing the societal challenges of enhancing resilience and sustainability have also increased in recent years. However, developing, executing, and assessing interdisciplinary content and skills at the post-secondary level has been challenging. The objective of this paper is to offer the Food-Energy-Water (FEW) Nexus as a powerful way to achieve sustainability competencies and matriculate graduates who will be equipped to facilitate the transformation of the global society by meeting the targets set by the United Nations Sustainable Development Goals. The paper presents 10 curricular design examples that span multiple levels, including modules, courses, and programs. These modules enable clear evaluation and assessment of key sustainability competencies, helping to prepare graduates with well-defined skillsets who are equipped to address current and future workforce needs.

54 ENVIRONMENTAL SCIENCES↗

Jay : A software framework for prototyping and evaluating offloading applications in hybrid edge clouds

Abstract We present Jay , a software framework for offloading applications in hybrid edge clouds. Jay provides an API, services, and tools that enable mobile application developers to implement, instrument, and evaluate offloading applications using configurable cloud topologies, offloading strategies, and job types. We start by presenting Jay 's job model and the concrete architecture of the framework. We then present the programming API with several examples of customization. Then, we turn to the description of the internal implementation of Jay instances and their components. Finally, we describe the Jay Workbench, a tool that allows the setup, execution, and reproduction of experiments with networks of hosts with different resource capabilities organized with specific topologies. The complete source code for the framework and workbench is provided in a GitHub repository.

Silva, Joaquim↗

Execute BEE workflows on private cloud infrastructure-2.3.6.01 - LANL ATDM ST / STNS01-22 Milestone Completion Documentation (BEE-FY21 P6-2) [Slides]

This work involves the creation of the Cloud Launcher, a new subcomponent of BEE, and the extension of the BEETaskManager to run on Cloud systems. BEE will be able to interact with the Google Compute Engine and OpenStack cloud APIs to set up simple Cloud clusters for launching HPC job scripts. BEE will use existing functionality to launch jobs that previously could only be launched on HPC systems. The BEETaskManager will handle launching tasks on the Cloud cluster.

97 MATHEMATICS AND COMPUTING↗

Bridging paradigms: Designing for HPC-Quantum convergence

Here, this paper presents a comprehensive software stack architecture for integrating quantum computing (QC) capabilities with High-Performance Computing (HPC) environments. While quantum computers show promise as specialized accelerators for scientific computing, their effective integration with classical HPC systems presents significant technical challenges. We propose a hardware-agnostic software framework that supports both current noisy intermediate-scale quantum devices and future fault-tolerant quantum computers, while maintaining compatibility with existing HPC workflows. The architecture includes a quantum gateway interface, standardized APIs for resource management, and robust scheduling mechanisms to handle both simultaneous and interleaved quantum–classical workloads. Key innovations include: (1) a unified resource management system that efficiently coordinates quantum and classical resources, (2) a flexible quantum programming interface that abstracts hardware-specific details, (3) A Quantum Platform Manager API that simplifies the integration of various quantum hardware systems, and (4) a comprehensive tool chain for quantum circuit optimization and execution. We demonstrate our architecture through implementation of quantum–classical algorithms, including the variational quantum linear solver, showcasing the framework’s ability to handle complex hybrid workflows while maximizing resource utilization. This work provides a foundational blueprint for integrating QC capabilities into existing HPC infrastructures, addressing critical challenges in resource management, job scheduling, and efficient data movement between classical and quantum resources.

97 MATHEMATICS AND COMPUTING↗

Carbon Capture from ArcelorMittal Hot Briquetted Iron Plant Using Air Liquide Cryocap™ FG Technology – FEED Study

The process of steel production is energy and carbon intensive with global average energy consumption of 5.5 MWh/tonne of steel and CO2 emission intensity of 1.83 tonne CO2/tonne of steel. The steel making process has inherent CO2 emissions from mineral conversion and is considered major contributors to the global carbon emissions. The steel industry is responsible for 8% of global carbon emissions. The main objective of this research project is to execute and complete a front-end engineering and design (FEED) study for a commercial-scale, carbon capture project that separates 95% of the total CO2 emissions at the ArcelorMittal’s Hot Briquetted Iron (HBI) plant in Portland, TX (Figure 1). The HBI is an ore-based metallic that is used as high-grade feedstock for high-quality steel via an Electric Arc Furnace (EAF) route. The HBI plant produces 2.0 million metric tonnes of high-quality HBI and emits approximately 1 million tonnes CO2/yr. The capture system is a Pressure Swing Adsorption (PSA) system assisted Cryocap™ FG technology (Figure 2). The captured CO2 will be pipeline grade and will be geologically stored in a facility within 10 miles of the CO2 source. The Host Site location in Corpus Christi, TX, is near hydrocarbon processing facilities and near Environmental Justice (EJ) and Qualified Opportunity Zone (QOZ) communities. Due to the location of the Host Site, the retrofit project offers the ability to demonstrate how a workforce focused on the fossil energy sector can be redirected to the clean- energy sector. The Air Liquide Cryocap™ capture technology is a proven technology and has been extensively examined for large industrial applications. It has been shown to be applicable to a variety of industrial applications including the steel industry. Cryocap™ FG (specific setup for Flue Gas application) consists of a Pressure Swing Adsorption (PSA) unit coupled with a Cryogenic System. The PSA pre-concentrates the CO2 from the flue gas, while the cryogenic unit enables the CO2 purity to be increased to the desired level. The scope of this study incorporates completing FEED study of the CO2 capture system which includes point-source CO2 capture and balance-of-plant; Business Case Analysis (BCA) outlining the current and projected volumes of the steel plant’s point sources of CO2 and the potential utilization of tax credits, including its projected revenue and duration; Life Cycle Analysis (LCA); Environmental Justice Analysis; Economic Revitalization and Job Creation Outcomes Analysis; and Workforce Readiness Plan. The plant design work was divided into two components: Inside Battery Limits (ISBL) and Outside Battery Limits (OSBL). The ISBL focuses on the capture system, while the OSBL focuses on the utility feeds and ducting from the plant to the capture system. Various design and engineering deliverables will be developed to define commodity quantities, equipment specifications, and labour effort required to execute the project. These FEED study deliverables will be prepared with the intent to develop an overall project capital cost estimate consistent with an AACE Class 3 estimate. The modular approach for the Cryocap™ FG that is being designed for this study integrates compression, PSA, and cryogenic “bricks” to achieve the desired CO2 capture rates. This carbon capture system integrates easily with the existing plant, thus reducing project costs and risks. It is also capable of managing impurities such as nitrogen oxides (NOx), sulfur oxides (SOx), mercury, hydrocarbons, and particulate matter. The capture system has a smaller footprint than amine-based systems. The two-step process uses PSA to preconcentrate the CO2 in the feedstream and then uses the cryogenic portion to purify and compress the resulting high purity CO2 product. This combination of purification and compression (i.e., process intensification) significantly reduces the CAPEX associated with use of a separate compressor commonly utilized for amine solvent-based systems. Successful completion of the FEED study will provide DOE with a detailed understanding of the costs for scaling up this proven capture technology for commercial applications at industrial facilities.

42 ENGINEERING↗

Heterogeneous Reconstruction of Tracks and Primary Vertices With the CMS Pixel Tracker

The High-Luminosity upgrade of the Large Hadron Collider (LHC) will see the accelerator reach an instantaneous luminosity of 7 × 10 34 cm −2 s −1 with an average pileup of 200 proton-proton collisions. These conditions will pose an unprecedented challenge to the online and offline reconstruction software developed by the experiments. The computational complexity will exceed by far the expected increase in processing power for conventional CPUs, demanding an alternative approach. Industry and High-Performance Computing (HPC) centers are successfully using heterogeneous computing platforms to achieve higher throughput and better energy efficiency by matching each job to the most appropriate architecture. In this paper we will describe the results of a heterogeneous implementation of pixel tracks and vertices reconstruction chain on Graphics Processing Units (GPUs). The framework has been designed and developed to be integrated in the CMS reconstruction software, CMSSW. The speed up achieved by leveraging GPUs allows for more complex algorithms to be executed, obtaining better physics output and a higher throughput.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Clippy

Clippy (CLI + PYthon) is a Python language interface to HPC resources. Precompiled binaries that execute on HPC systems are exposed as methods to a dynamically-created Clippy Python object, where they present a familiar interface to researchers, data scientists, and others. Clippy allows these users to interact with HPC resources in an easy, straightforward environment - at the REPL, for example, or within a notebook - without the need to learn complex HPC behavior and arcane job submission commands.

Bromberger, SethA.↗