Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “runtime”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 253 records · Page 14

The Case for Co-Designing Model Architectures with Hardware

While GPUs are responsible for training the vast majority of state-of-the-art deep learning models, the implications of their architecture are often overlooked when designing new deep learning (DL) models. As a consequence, modifying a DL model to be more amenable to the target hardware can significantly improve the runtime performance of DL training and inference. In this paper, we provide a set of guidelines for users to maximize the runtime performance of their transformer models. These guidelines have been created by carefully considering the impact of various model hyperparameters controlling model shape on the efficiency of the underlying computation kernels executed on the GPU. We find the throughput of models with “efficient” model shapes is up to 39% higher while preserving accuracy compared to models with a similar number of parameters but with unoptimized shapes.

Yin, Junqi↗

SpecSims: A Scalable Speculative Tree-based Simulation Cloning Framework for Finite Memory Machines

Simulation cloning is a technique in which cloned simulations whose state spaces differ partially from their parent simulation due to intervening events are spawned at runtime and concurrently advanced. It is a powerful method to carry out what-if analysis by speculatively exploring and evaluating the impact of various permutations of intervening cascade of events. Due to the exponential growth in the number of possible clones even for a small number of distinct intervening events, the practical efficacy of the approach is often severely limited by the maximum available memory of the computing host. In this paper, we introduce a novel speculative simulation cloning framework that executes a simulation cloning campaign capable of efficiently exploring an exponentially large space of clone simulations created by permutation of intervening events under a finite memory constraint. We provide a theoretical analysis of the runtime characteristics of our proposed approach and highlight its novel advantages such as memory-aware and as-long-as-needed execution. Furthermore, in support of our analytical findings and to demonstrate its practical feasibility, we implement a prototype of the cloning framework on a shared memory system and report its performance characteristics in the context of a heat diffusion simulation, and a power grid simulation subject to cascading disruptions from geomagnetic disturbances.

Simulation framework↗

Poster Abstract: Leveraging Large Language Models to Reveal Interpretable Cooling Behaviors from Smart Thermostat Data

Frequent heatwaves and hot summers increasingly challenge occupant comfort, health, and energy grid stability. Addressing these challenges requires a detailed understanding of household cooling behaviors, such as thermostat adjustments and adaptive responses to extreme conditions. Traditional analyses often rely on aggregated numerical metrics that overlook subtle but important household-specific variations. In this study, we introduce a generalizable methodology that integrates large language models (LLMs) with vision capabilities to enable scalable and detailed analysis of residential thermostat data. Using Ecobee's Donate Your Data (DYD) dataset—which provides five-minute records of indoor temperatures, thermostat setpoints, and HVAC runtimes—we focus on two U.S. cities with contrasting summer climates : Austin (TX) and Phoenix (AZ). Because raw time-series data are not well suited for direct LLM analysis, we transform them into visual representations, such as daily indoor temperature trajectories and weekly runtime histograms, to better capture behavioral variations. Leveraging LLMs' visual interpretation, we extract descriptive behavioral features, including temperature preferences, time-of-day cooling orientation, anticipatory versus reactive heatwave responses, and behavioral consistency. These semantic features support unsupervised clustering to identify distinct occupant archetypes at scale, revealing differences—such as morning-centric anticipatory coolers versus households that shift toward warmer setpoints during heatwaves—that can inform demand response, resilience planning, and health-aware interventions. By converting raw numerical data into interpretable behavioral patterns, this methodology enables scalable and practical analysis of occupant behavior, supporting actionable insights for comfort, resilience, and energy management.

Nihar, Kopal↗

IRIS-MASH: Efficient Multi-device Asynchronous Multi-Stream Heterogeneous Computing

In the rapidly evolving field of high-performance computing (HPC), effectively leveraging heterogeneous devices through asynchronous task programming is paramount. This paper presents a robust asynchronous task programming model tailored for a multi-device, multi-stream execution environment that incorporates a diverse array of heterogeneous computing units, including GPUs from various vendors and other accelerators. Current state-of-the-art task programming models provide methodologies to support asynchronous task executions, but they typically handle homogeneous devices using native programming languages, while support for heterogeneous devices is limited to frameworks like OpenCL. This gap presents significant challenges in abstracting heterogeneous devices to harness their true asynchronous capabilities effectively using their native programming languages. By implementing asynchronous task execution, our model significantly boosts the performance of tiled algorithm task graphs through overlapping data transfers with computation and enabling the simultaneous execution of multiple kernels. We integrate this approach into a heterogeneous Intelligent Runtime System (IRIS) and assess its performance using a suite of tiled algorithm benchmarks from the heterogeneous math kernels library (MatRIS) based on IRIS. Experimental results demonstrate a performance improvement ranging from 1.6 × to 2 × over IRIS without asynchronous support, and a notable 22% performance enhancement compared to established runtime systems such as StarPU and PaRSEC. This approach significantly improves computation efficiency of HPC workflows and provides a solid base for future exploration and development in the area of asynchronous task programming in heterogeneous systems.

Miniskar, Narasinga Rao [ORNL] (ORCID:000000018259↗

Scalable Federated Learning for Scientific Foundation Models on Leadership-Class Systems

Federated learning (FL) at leadership-class HPC systems remains largely unexplored, despite growing interest in deploying federated workflows on modern HPC systems. This paper provides the first system-level empirical characterization of federated fine-tuning of pretrained foundation models on an exascale supercomputer under a multi-node deployment. Using up to 96 concurrent FL clients deployed across Frontier nodes, we study the impact of client scale, model size, data heterogeneity, partial participation, and differential privacy on runtime, communication overhead, and convergence stability. Our results show that pretrained transformer models remain robust to heterogeneity, client dropout, and privacy noise, while system efficiency degrades rapidly with scale as synchronizat and orchestration dominate runtime. We further demonstrate that system-aware execution strategies, including intra-node aggregation and early aggregation, significantly reduce wall-clock time without degrading model quality. These findings establish a practical performance baseline and inform the design of communication-efficient FL systems on leadership-class HPC platforms.

Kotevska, Olivera [ORNL] (ORCID:0000000316772243)↗

Quantum Search Compiler (Qsearch) v2.0

Quantum Gate Synthesis is the process of taking a desired quantum gate operation, in this case specified as a unitary matrix, and decomposing it into a list of operations that can actually be performed on a quantum computer. Qsearch is an implementation of a quantum synthesis algorithm based on combining A* search with numerical optimization. It is designed to minimize CNOT count at the expense of having a long runtime. It produces particularly efficient quantum circuits in terms of CNOT count, up until about 4 qubits. It is generally impractical to run Qsearch on circuits with more than 4 qubits, but it can be used as a subroutine in algorithms designed to work with these larger circuits. Qsearch takes the form of a Python package, taking advantage of Numpy and Scipy, with part of the package written in Rust and taking advantage of OpenBLAS, NLOPT, and Google Ceres for faster runtime.

Lancu, Costin↗

High Performance Approximate Computing

This code repository contains the implementation of the "High-Performance Approximate Computing" (HPAC) toolkit. The toolkit allows you to approximate your own C/C++. The developer uses "pragma's" to annotate code regions as approximate. The compiler extensions lower these pragmas to either compiletime approximate techniques or runtime approximation techniques. At execution time, the implemented runtime system decides which annotated regions it should approximate. HPAC also provides a set of script utilities. The utilities perform a grid search within approximation parameters and performance. The user can analyze the raw data to identify optimal approximation techniques for the application.

Parasyris, Konstantinos↗

XPlacer/Tracer

XPlacer/Tracer is a dynamic analysis tool that finds bad memory access patterns in heterogeneous codes. XPlacer/Tracer consists of two components. (1) A configurable ROSE plugin for source code instrumentation of heterogeneous code written in the C/C++ and CUDA programming languages. (2) A sample configuration that instruments memory accesses to dynamic memory in CPU and GPU codes, and a runtime library that tracks these memory accesses at runtime. The collected information is reported in form of a textual summary or as memory map that can be converted to images.

Pirkelbauer, PeterM.↗

Contra

written in Contra can run without modification on a variety of parallel runtimes. Currently supported parallel runtimes include: MPI, Legion, CUDA, AMD ROCm/HIP, and PThreads.

Charest, Marc↗

Clang UPC2C Translator (Clang UPC2C) v9.0.1-1

Clang Unified Parallel C 2 C (Clang UPC2C) translator compiles programs written in the UPC (Unified Parallel C) language to ISO C99, with calls to the Berkeley UPC runtime system. The Clang UPC2C compiler extends the capabilities of the Clang LLVM C frontend to comply with the UPC Language Specification version 1.3. The compiler generates programs that run on a wide variety of systems ranging from workstations to leadership-class supercomputers, in conjunction with the Berkeley UPC runtime and GASNet communication system. In addition to the standard UPC libraries, Clang UPC2C also provides access to Berkeley UPC library extensions.

Hargrove, Paul↗

Resiliency in numerical algorithm design for extreme scale simulations

Here this work is based on the seminar titled ‘Resiliency in Numerical Algorithm Design for Extreme Scale Simulations’ held March 1–6, 2020, at Schloss Dagstuhl, that was attended by all the authors. Advanced supercomputing is characterized by very high computation speeds at the cost of involving an enormous amount of resources and costs. A typical large-scale computation running for 48 h on a system consuming 20 MW, as predicted for exascale systems, would consume a million kWh, corresponding to about 100k Euro in energy cost for executing 10 23 floating-point operations. It is clearly unacceptable to lose the whole computation if any of the several million parallel processes fails during the execution. Moreover, if a single operation suffers from a bit-flip error, should the whole computation be declared invalid? What about the notion of reproducibility itself: should this core paradigm of science be revised and refined for results that are obtained by large-scale simulation? Naive versions of conventional resilience techniques will not scale to the exascale regime: with a main memory footprint of tens of Petabytes, synchronously writing checkpoint data all the way to background storage at frequent intervals will create intolerable overheads in runtime and energy consumption. Forecasts show that the mean time between failures could be lower than the time to recover from such a checkpoint, so that large calculations at scale might not make any progress if robust alternatives are not investigated. More advanced resilience techniques must be devised. The key may lie in exploiting both advanced system features as well as specific application knowledge. Research will face two essential questions: (1) what are the reliability requirements for a particular computation and (2) how do we best design the algorithms and software to meet these requirements? While the analysis of use cases can help understand the particular reliability requirements, the construction of remedies is currently wide open. One avenue would be to refine and improve on system- or application-level checkpointing and rollback strategies in the case an error is detected. Developers might use fault notification interfaces and flexible runtime systems to respond to node failures in an application-dependent fashion. Novel numerical algorithms or more stochastic computational approaches may be required to meet accuracy requirements in the face of undetectable soft errors. These ideas constituted an essential topic of the seminar. The goal of this Dagstuhl Seminar was to bring together a diverse group of scientists with expertise in exascale computing to discuss novel ways to make applications resilient against detected and undetected faults. In particular, participants explored the role that algorithms and applications play in the holistic approach needed to tackle this challenge. This article gathers a broad range of perspectives on the role of algorithms, applications and systems in achieving resilience for extreme scale simulations. The ultimate goal is to spark novel ideas and encourage the development of concrete solutions for achieving such resilience holistically.

79 ASTRONOMY AND ASTROPHYSICS↗

Designing and prototyping extensions to the Message Passing Interface in MPICH

As HPC system architectures and the applications running on them continue to evolve, the MPI standard itself must evolve. The trend in current and future HPC systems toward powerful nodes with multiple CPU cores and multiple GPU accelerators makes efficient support for hybrid programming critical for applications to achieve high performance. However, the support for hybrid programming in the MPI standard has not kept up with recent trends. The MPICH implementation of MPI provides a platform for implementing and experimenting with new proposals and extensions to fill this gap and to gain valuable experience and feedback before the MPI Forum can consider them for standardization. Here, in this work, we detail six extensions implemented in MPICH to increase MPI interoperability with other runtimes, with a specific focus on heterogeneous architectures. First, the extension to MPI generalized requests lets applications integrate asynchronous tasks into MPI’s progress engine. Second, the iovec extension to datatypes lets applications use MPI datatypes as a general-purpose data layout API beyond just MPI communications. Third, a new MPI object, MPIX_Stream, can be used by applications to identify execution contexts beyond MPI processes, including threads and GPU streams. MPIX stream communicators can be created to make existing MPI functions thread-aware and GPU-aware, thus providing applications with explicit ways to achieve higher performance. Fourth, MPIX Streams are extended to support the enqueue semantics for offloading MPI communications onto a GPU stream context. Fifth, thread communicators allow MPI communicators to be constructed with individual threads, thus providing a new level of interoperability between MPI and on-node runtimes such as OpenMP. Lastly, we present an extension to invoke MPI progress, which lets users spawn progress threads with fine-grained control to adapt the communication performance to their application designs. We describe the design and implementation of these extensions, provide usage examples, and highlight their expected benefits with performance results.

97 MATHEMATICS AND COMPUTING↗

PaRSEC: Scalability, flexibility, and hybrid architecture support for task-based applications in ECP

This paper highlights the most significant enhancements made to PaRSEC, a scalable task-based runtime system designed for hybrid machines, during the Exascale Computing Project (ECP). The enhancements focus on expanding the capabilities of PaRSEC to address the evolving landscape of parallel computing. Notable achievements include the integration of support for three major types of accelerators (NVIDIA, AMD, and Intel GPUs), the refinement and increased flexibility of the communication subsystem, and the introduction of new programming interfaces tailored for irregular applications. Additionally, the project resulted in the development of powerful debugging and performance analysis tools aimed at assisting users in understanding and optimizing their applications. We present a comprehensive demonstration of these advancements through a series of benchmarks and applications within ECP and beyond, thereby showcasing the enhanced capabilities of PaRSEC across the diverse architectures within the ECP, providing valuable insights into the runtime system’s adaptability and performance across varied computing environments.

Bouteiller, Aurelien↗

Speeding genomic island discovery through systematic design of reference database composition

Background Genomic islands (GIs) are mobile genetic elements that integrate site-specifically into bacterial chromosomes, bearing genes that affect phenotypes such as pathogenicity and metabolism. GIs typically occur sporadically among related bacterial strains, enabling comparative genomic approaches to GI identification. For a candidate GI in a query genome, the number of reference genomes with a precise deletion of the GI serves as a support value for the GI. Our comparative software for GI identification was slowed by our original use of large reference genome databases (DBs). Here we explore smaller species-focused DBs. Results With increasing DB size, recovery of our reliable prophage GI calls reached a plateau, while recovery of less reliable GI calls (FPs) increased rapidly as DB sizes exceeded ~500 genomes; i.e., overlarge DBs can increase FP rates. Paradoxically, relative to prophages, FPs were both more frequently supported only by genomes outside the species and more frequently supported only by genomes inside the species; this may be due to their generally lower support values. Setting a DB size limit for our SMA ll R anked T ailored (SMART) DB design speeded runtime ~65-fold. Strictly intra-species DBs would tend to lower yields of prophages for small species (with few genomes available); simulations with large species showed that this could be partially overcome by reaching outside the species to closely related taxa, without an FP burden. Employing such taxonomic outreach in DB design generated redundancy in the DB set; as few as 2984 DBs were needed to cover all 47894 prokaryotic species. Conclusions Runtime decreased dramatically with SMART DB design, with only minor losses of prophages. We also describe potential utility in other comparative genomics projects.

59 BASIC BIOLOGICAL SCIENCES↗

Characterizing electrical demand and load diversity of low-power water and space heating appliances in US homes

Home renovation and remodeling projects can involve costly and time consuming electrical infrastructure upgrades at the household level. From the grid perspective they also lead to costly replacement of local infrastructure, such as transformers, and can add stress to the grid at peak times. The emergence of innovative, power-efficient household appliances offers a way to minimize these problems. These appliances are designed for lower power consumption, simplifying installation through standard plug-in connections, reducing the need for new electric circuits/panels/service, and minimizing the peak power demand for the home. Key examples include low-power heat pump water heaters (HPWHs) and cold climate window heat pumps that operate on standard 120V outlets. To assess the real-world impact of these solutions, we compiled and analyzed power metering data from several US field studies. This data provides insights into the effects on peak power demand of selecting lower-power appliances. Our analysis focuses on several key metrics, including the maximum power demand of individual appliances, their operational runtime, continuous operation and load diversity. While individual low-power 120V space and water heating appliances offer significant peak demand reductions compared to 240V heat pump or resistance alternatives, their longer runtimes might increase the likelihood of operation during whole-dwelling peak events, albeit at lower power levels.

Less, Brennan↗

Performance Analysis of Traditional and Data-Parallel Primitive Implementations of Visualization and Analysis Kernels

Measurements of absolute runtime are useful as a summary of performance when studying parallel visualization and analysis methods on computational platforms of increasing concurrency and complexity. We can obtain even more insights by measuring and examining more detailed measures from hardware performance counters, such as the number of instructions executed by an algorithm implemented in a particular way, the amount of data moved to/from memory, memory hierarchy utilization levels via cache hit/miss ratios, and so forth. This work focuses on performance analysis on modern multi-core platforms of three different visualization and analysis kernels that are implemented in different ways: one is "traditional", using combinations of C++ and VTK, and the other uses a data-parallel approach using VTK-m. Our performance study consists of measurement and reporting of several different hardware performance counters on two different multi-core CPU platforms. The results reveal interesting performance differences between these two different approaches for implementing these kernels, results that would not be apparent using runtime as the only metric.

97 MATHEMATICS AND COMPUTING↗

Detection of Defects in Additively Manufactured Metallic Materials with Machine Learning of Pulsed Thermography Images

Additive manufacturing (AM) is an emerging method for cost-efficient fabrication of nuclear reactor parts. AM of metallic structures for nuclear energy applications is currently based on laser powder bed fusion (LPBF) process, which can introduce internal material flaws, such as pores and anisotropy. Integrity of AM structures needs to be evaluated nondestructively because material flaws could lead to premature failures due to exposure to high temperature, radiation and corrosive environment in a nuclear reactor. Quality control (QC) requires nondestructive evaluation (NDE) of actual AM structures. Pulsed thermography is a potentially promising QC technique because it is scalable to arbitrary structure size. However, detection sensitivity of this method is limited by noises. We investigate separation of signal from noise in thermography images using several machine learning (ML) methods, including new spatio-temporal blind source separation (STBSS) and spatio-temporal sparse dictionary learning (STSDL) methods. Performance of the ML methods is benchmarked using thermography data obtained from imaging stainless steel 316L and Inconel 718 specimens produced LPBF method with imprinted calibrated porosity defects. The ML methods are ranked by F-score and execution runtime. The ML methods with higher accuracy require longer run time. However, this runtime is sufficiently short to perform QC within a realistic time frame.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Development and Validation of Algorithms That Analyze Communicating Thermostat Data to Identify Enclosure Retrofit Opportunities

Annual energy savings of up to $\$ 4$ to $\$ 5$ billion could be achieved nationwide through basic insulation and heating system retrofits of existing homes. However, current utility energy efficiency programs are costly and challenging to scale. Customer acquisition occurs primarily through energy bill mailers, mass media, and online advertising that lack specificity about home-specific retrofit opportunities, expected energy savings, and cost-effectiveness. Specific retrofit opportunities are identified via on-site home energy assessments (HEAs) that are inconvenient to homeowners, expensive, and of variable accuracy. We developed computational algorithms that automatically analyze communicating thermostat (CT) heating data that could be used to increase the customer uptake of insulation and air sealing energy conservation measures (ECMs) by identifying homes with the most significant retrofit opportunities, estimating post-retrofit energy savings, and formulating home-specific outreach. The algorithms are based on an extended second-order grey-box model that characterizes a building’s thermal response using lumped elements, coupled with an empirical model of infiltration that accounts for both wind and stack effects. The basic parameters of the model correspond to actual physical parameters of the home, i.e., the home’s overall R-value of and the building envelope ACH50. Unlike the conventional approach, which estimates model parameters based on the best fit to the observed time-dependent room temperature, our approach derives correlations between the daily heating system runtime and temperature difference (indoor-outdoor) that are more robust to data quality issues in real-world applications. We also used HEA data for algorithm development and validation. With the help of our utility partners, Eversource and National Grid, we obtained data sets for hundreds of Massachusetts homes. For each home, these data sets included three sets of information anonymized by the utility: (1) CT data (HVAC runtime, room temperature, and, for some vendors, outdoor temperature and wind speed) collected by the CT vendor (one of three) over a heating season, (2) HEA report performed by the HEA vendor (same vendor for all homes), (3) Monthly utility gas bills coincident with the CT data (3 to 24 per home, depending on availability). For some homes, we also obtained blower-door test results. Initially, we applied the algorithms developed to homes with a single CT and then extended them to homes with two CTs by using an equivalent home approach. Finally, we developed algorithms for prediction of energy savings and a methodology of comparing our predictions with those generated by HEAs. The main technical results indicate that we can reliably identify homes with insulation and/or air sealing retrofit opportunities and provide accurate savings predictions. Our hypothesis is that the algorithms could be applied to utility energy efficiency programs to identify homes that could realize significant energy savings from insulation and/or air sealing retrofits. This information could then be used to reach out to those homes with highly customized outreach, thereby delivering increased program energy savings and cost-effectiveness. This would: Significantly increase the uptake rate of on-site HEAs, and Significantly increase the fraction of HEAs resulting in ECM implementation. To test these hypotheses, we designed and conducted a randomized controlled trial (RCT). The RCT results suggest that personal messaging leads to a two- to five-fold increase in the HEA uptake rate.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗