Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “runtime”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10

Designing and prototyping extensions to the Message Passing Interface in MPICH

As HPC system architectures and the applications running on them continue to evolve, the MPI standard itself must evolve. The trend in current and future HPC systems toward powerful nodes with multiple CPU cores and multiple GPU accelerators makes efficient support for hybrid programming critical for applications to achieve high performance. However, the support for hybrid programming in the MPI standard has not kept up with recent trends. The MPICH implementation of MPI provides a platform for implementing and experimenting with new proposals and extensions to fill this gap and to gain valuable experience and feedback before the MPI Forum can consider them for standardization. Here, in this work, we detail six extensions implemented in MPICH to increase MPI interoperability with other runtimes, with a specific focus on heterogeneous architectures. First, the extension to MPI generalized requests lets applications integrate asynchronous tasks into MPI’s progress engine. Second, the iovec extension to datatypes lets applications use MPI datatypes as a general-purpose data layout API beyond just MPI communications. Third, a new MPI object, MPIX_Stream, can be used by applications to identify execution contexts beyond MPI processes, including threads and GPU streams. MPIX stream communicators can be created to make existing MPI functions thread-aware and GPU-aware, thus providing applications with explicit ways to achieve higher performance. Fourth, MPIX Streams are extended to support the enqueue semantics for offloading MPI communications onto a GPU stream context. Fifth, thread communicators allow MPI communicators to be constructed with individual threads, thus providing a new level of interoperability between MPI and on-node runtimes such as OpenMP. Lastly, we present an extension to invoke MPI progress, which lets users spawn progress threads with fine-grained control to adapt the communication performance to their application designs. We describe the design and implementation of these extensions, provide usage examples, and highlight their expected benefits with performance results.

97 MATHEMATICS AND COMPUTING↗

PaRSEC: Scalability, flexibility, and hybrid architecture support for task-based applications in ECP

This paper highlights the most significant enhancements made to PaRSEC, a scalable task-based runtime system designed for hybrid machines, during the Exascale Computing Project (ECP). The enhancements focus on expanding the capabilities of PaRSEC to address the evolving landscape of parallel computing. Notable achievements include the integration of support for three major types of accelerators (NVIDIA, AMD, and Intel GPUs), the refinement and increased flexibility of the communication subsystem, and the introduction of new programming interfaces tailored for irregular applications. Additionally, the project resulted in the development of powerful debugging and performance analysis tools aimed at assisting users in understanding and optimizing their applications. We present a comprehensive demonstration of these advancements through a series of benchmarks and applications within ECP and beyond, thereby showcasing the enhanced capabilities of PaRSEC across the diverse architectures within the ECP, providing valuable insights into the runtime system’s adaptability and performance across varied computing environments.

Bouteiller, Aurelien↗

Speeding genomic island discovery through systematic design of reference database composition

Background Genomic islands (GIs) are mobile genetic elements that integrate site-specifically into bacterial chromosomes, bearing genes that affect phenotypes such as pathogenicity and metabolism. GIs typically occur sporadically among related bacterial strains, enabling comparative genomic approaches to GI identification. For a candidate GI in a query genome, the number of reference genomes with a precise deletion of the GI serves as a support value for the GI. Our comparative software for GI identification was slowed by our original use of large reference genome databases (DBs). Here we explore smaller species-focused DBs. Results With increasing DB size, recovery of our reliable prophage GI calls reached a plateau, while recovery of less reliable GI calls (FPs) increased rapidly as DB sizes exceeded ~500 genomes; i.e., overlarge DBs can increase FP rates. Paradoxically, relative to prophages, FPs were both more frequently supported only by genomes outside the species and more frequently supported only by genomes inside the species; this may be due to their generally lower support values. Setting a DB size limit for our SMA ll R anked T ailored (SMART) DB design speeded runtime ~65-fold. Strictly intra-species DBs would tend to lower yields of prophages for small species (with few genomes available); simulations with large species showed that this could be partially overcome by reaching outside the species to closely related taxa, without an FP burden. Employing such taxonomic outreach in DB design generated redundancy in the DB set; as few as 2984 DBs were needed to cover all 47894 prokaryotic species. Conclusions Runtime decreased dramatically with SMART DB design, with only minor losses of prophages. We also describe potential utility in other comparative genomics projects.

59 BASIC BIOLOGICAL SCIENCES↗

Characterizing electrical demand and load diversity of low-power water and space heating appliances in US homes

Home renovation and remodeling projects can involve costly and time consuming electrical infrastructure upgrades at the household level. From the grid perspective they also lead to costly replacement of local infrastructure, such as transformers, and can add stress to the grid at peak times. The emergence of innovative, power-efficient household appliances offers a way to minimize these problems. These appliances are designed for lower power consumption, simplifying installation through standard plug-in connections, reducing the need for new electric circuits/panels/service, and minimizing the peak power demand for the home. Key examples include low-power heat pump water heaters (HPWHs) and cold climate window heat pumps that operate on standard 120V outlets. To assess the real-world impact of these solutions, we compiled and analyzed power metering data from several US field studies. This data provides insights into the effects on peak power demand of selecting lower-power appliances. Our analysis focuses on several key metrics, including the maximum power demand of individual appliances, their operational runtime, continuous operation and load diversity. While individual low-power 120V space and water heating appliances offer significant peak demand reductions compared to 240V heat pump or resistance alternatives, their longer runtimes might increase the likelihood of operation during whole-dwelling peak events, albeit at lower power levels.

Less, Brennan↗

Performance Analysis of Traditional and Data-Parallel Primitive Implementations of Visualization and Analysis Kernels

Measurements of absolute runtime are useful as a summary of performance when studying parallel visualization and analysis methods on computational platforms of increasing concurrency and complexity. We can obtain even more insights by measuring and examining more detailed measures from hardware performance counters, such as the number of instructions executed by an algorithm implemented in a particular way, the amount of data moved to/from memory, memory hierarchy utilization levels via cache hit/miss ratios, and so forth. This work focuses on performance analysis on modern multi-core platforms of three different visualization and analysis kernels that are implemented in different ways: one is "traditional", using combinations of C++ and VTK, and the other uses a data-parallel approach using VTK-m. Our performance study consists of measurement and reporting of several different hardware performance counters on two different multi-core CPU platforms. The results reveal interesting performance differences between these two different approaches for implementing these kernels, results that would not be apparent using runtime as the only metric.

97 MATHEMATICS AND COMPUTING↗

Detection of Defects in Additively Manufactured Metallic Materials with Machine Learning of Pulsed Thermography Images

Additive manufacturing (AM) is an emerging method for cost-efficient fabrication of nuclear reactor parts. AM of metallic structures for nuclear energy applications is currently based on laser powder bed fusion (LPBF) process, which can introduce internal material flaws, such as pores and anisotropy. Integrity of AM structures needs to be evaluated nondestructively because material flaws could lead to premature failures due to exposure to high temperature, radiation and corrosive environment in a nuclear reactor. Quality control (QC) requires nondestructive evaluation (NDE) of actual AM structures. Pulsed thermography is a potentially promising QC technique because it is scalable to arbitrary structure size. However, detection sensitivity of this method is limited by noises. We investigate separation of signal from noise in thermography images using several machine learning (ML) methods, including new spatio-temporal blind source separation (STBSS) and spatio-temporal sparse dictionary learning (STSDL) methods. Performance of the ML methods is benchmarked using thermography data obtained from imaging stainless steel 316L and Inconel 718 specimens produced LPBF method with imprinted calibrated porosity defects. The ML methods are ranked by F-score and execution runtime. The ML methods with higher accuracy require longer run time. However, this runtime is sufficiently short to perform QC within a realistic time frame.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Development and Validation of Algorithms That Analyze Communicating Thermostat Data to Identify Enclosure Retrofit Opportunities

Annual energy savings of up to $\$ 4$ to $\$ 5$ billion could be achieved nationwide through basic insulation and heating system retrofits of existing homes. However, current utility energy efficiency programs are costly and challenging to scale. Customer acquisition occurs primarily through energy bill mailers, mass media, and online advertising that lack specificity about home-specific retrofit opportunities, expected energy savings, and cost-effectiveness. Specific retrofit opportunities are identified via on-site home energy assessments (HEAs) that are inconvenient to homeowners, expensive, and of variable accuracy. We developed computational algorithms that automatically analyze communicating thermostat (CT) heating data that could be used to increase the customer uptake of insulation and air sealing energy conservation measures (ECMs) by identifying homes with the most significant retrofit opportunities, estimating post-retrofit energy savings, and formulating home-specific outreach. The algorithms are based on an extended second-order grey-box model that characterizes a building’s thermal response using lumped elements, coupled with an empirical model of infiltration that accounts for both wind and stack effects. The basic parameters of the model correspond to actual physical parameters of the home, i.e., the home’s overall R-value of and the building envelope ACH50. Unlike the conventional approach, which estimates model parameters based on the best fit to the observed time-dependent room temperature, our approach derives correlations between the daily heating system runtime and temperature difference (indoor-outdoor) that are more robust to data quality issues in real-world applications. We also used HEA data for algorithm development and validation. With the help of our utility partners, Eversource and National Grid, we obtained data sets for hundreds of Massachusetts homes. For each home, these data sets included three sets of information anonymized by the utility: (1) CT data (HVAC runtime, room temperature, and, for some vendors, outdoor temperature and wind speed) collected by the CT vendor (one of three) over a heating season, (2) HEA report performed by the HEA vendor (same vendor for all homes), (3) Monthly utility gas bills coincident with the CT data (3 to 24 per home, depending on availability). For some homes, we also obtained blower-door test results. Initially, we applied the algorithms developed to homes with a single CT and then extended them to homes with two CTs by using an equivalent home approach. Finally, we developed algorithms for prediction of energy savings and a methodology of comparing our predictions with those generated by HEAs. The main technical results indicate that we can reliably identify homes with insulation and/or air sealing retrofit opportunities and provide accurate savings predictions. Our hypothesis is that the algorithms could be applied to utility energy efficiency programs to identify homes that could realize significant energy savings from insulation and/or air sealing retrofits. This information could then be used to reach out to those homes with highly customized outreach, thereby delivering increased program energy savings and cost-effectiveness. This would: Significantly increase the uptake rate of on-site HEAs, and Significantly increase the fraction of HEAs resulting in ECM implementation. To test these hypotheses, we designed and conducted a randomized controlled trial (RCT). The RCT results suggest that personal messaging leads to a two- to five-fold increase in the HEA uptake rate.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Dynamic Species Reduction for Multi-Cycle CFD Simulations (Final Technical Report)

This project primarily sought to address some of the computational cost concerns of detailed simulations by developing improved methods of handling species transport, and chemical kinetics evaluation in a commercial 3D Computational Fluid Dynamics (CFD) environment. Secondary goals were to apply these techniques to fuels and conditions of interest to better understand the key species and reactions required to adequately capture cycle to cycle coupling. Improved modeling will lead to better understanding of these combustion modes and their dependence on fuel composition, which can then enable more clean and efficient engines, ultimately benefiting the consumer as well as the general public. Two approaches were used to address the computational cost. A “Dynamic Species Reduction” (DSR) procedure was developed to remove chemical species from the simulation during periods when chemical reactions were not expected to be important, particularly during gas exchange when temperatures are low and little fuel remains. This modelling procedure automatically detects relevant species to retain in the simulation domain based on their local concentration, removes species below a specified concentration threshold, and then adjusts the remaining species mass to conserve not only the number of H, C, and O atoms in each cell, but also the relative proportions between species and the heating value of the mixture in every cell. The second procedure was a “Product Directed Remapping” (PDR) updates the algorithm used to group individual computational cells for chemical kinetics evaluation to account for non-uniform temperature distributions and the CO to CO 2 ratio in the cells. The model techniques developed in this work successfully demonstrated computational performance improvements for a range of conditions relevant for Highly Dilute SI and HCCI engine operation. Runtime reductions of 10% were observed for small mechanisms, with further reductions of up to 36% observed for larger mechanisms. Improvements were primarily related to reducing the number of chemical species tracked during the gas exchange process using the Dynamic Species Reduction method. With smaller benefits observed from changes to the kinetics binning and evaluation strategy in the post combustion region using the PDR method. The results of this work show the potential for improved computational runtime for complicated simulations. They can and should be extended to additional conditions and new bio-derived and renewable fuels as they are developed and new kinetic mechanisms become available.

10 SYNTHETIC FUELS↗

GPU Acceleration of a Diagnostic Wind Solver

Reducing QUIC-Fire simulation runtimes is crucial in enabling simulation ensembles to guide science-driven prescribed fire planning in regions of complex terrain. Utilizing GPUs through OpenACC and Kokkos frameworks to accelerate generation of 3D windfields in QUIC-Fire would provide substantial speedup to the runtime of the code. Using GPUs for the Successive Over-Relaxation (SOR) algorithm utilized in QUIC-Fire will require deriving a ‘four colored’ version of the well explored ‘Red-Black’ SOR parallelization scheme. This effort could open the door for utilizing GPUs in other portions of the QUIC-Fire algorithm, increasing its viability in more complex and larger domains.

58 GEOSCIENCES↗

Faster Tensor Network Decoding for Topological Quantum Codes

We present a fast and Bayes-optimal-approximating tensor network decoder for planar quantum LDPC codes based on the tensor renormalization group algorithm, originally proposed by Levin, and Nave. By precomputing the renormalization group flow for the null syndrome, we need only recompute tensor contractions in the causal cone of the measured syndrome at the time of decoding. This allows us to achieve an overall runtime complexity of ($pnχ^6$) where p is the depolarizing noise rate, and χ is the cutoff value used to control singular value decomposition approximations used in the algorithm. We apply our decoder to the surface code in the code capacity noise model and compare its performance to the original matrix product state (MPS) tensor network decoder introduced by Bravyi, Suchara, and Vargo. The MPS decoder has a p-independent runtime complexity of $\mathcal{O}(nχ^3)$ resulting in significantly slower decoding times compared to our algorithm in the low-p regime.

97 MATHEMATICS AND COMPUTING↗

Improving the Performance of NEML2 with Modern Graph Compilation Backends

NEML2 vectorizes constitutive-model evaluation for large-scale multiphysics simulation, using PyTorch as its tensor backend so that a batch of material-point updates runs on CPU or GPU through a single implementation. In the two prior reports in this series it was a C++-native library, deployed through TorchScript tracing and just-in-time (JIT) compilation; it has since been rewritten from the ground up into a Python-native library deployed through Ahead-of-Time Inductor (AOTInductor), a modern PyTorch graph-compilation backend. The rewrite is driven by a persistent tension, not a language preference: NEML2 composes constitutive models at runtime from a registry of small, independently-authored pieces, and that flexibility is difficult to reconcile with the compile-time knowledge an efficient GPU kernel needs. This report documents the rewrite and the investment that accompanied it: the AOTInductor export pipeline that turns a Python-authored model into a portable, Python-free compiled artifact loadable from pure C++; the eager and compiled runtimes and the new implicit solver layer built on them; a head-to-head benchmark of legacy JIT against AOTInductor; the physics-model catalog and its worked examples; the developer tooling; and the corresponding overhaul of MOOSE’s NEML2 integration that lets MOOSE consume it. A central objective is to examine whether modern PyTorch graph-compilation backends are effective for MOOSE GPU integration. The benchmark answers directly: AOTInductor outperforms legacy JIT on every GPU scenario measured, by 1.0–4.5×. Modern graph-compilation backends are effective for MOOSE GPU integration, and AOTInductor specifically – not compilation in the abstract – is why.

Hu, Gary (Tianchen) [Argonne National Laboratory (↗

Caffeine: CoArray Fortran Framework of Efficient Interfaces to Network Environments

This paper provides an introduction to the CoArray Fortran Framework of Efficient Interfaces to Network Environments (Caffeine), a parallel runtime library built atop the GASNet-EX exascale networking library. Caffeine leverages several non-parallel Fortran features to write type- and rank-agnostic interfaces and corresponding procedure definitions that support parallel Fortran 2018 features, including communication, collective operations, and related services. One major goal is to develop a runtime library that can eventually be considered for adoption by LLVM Flang, enabling that compiler to support the parallel features of Fortran. The paper describes the motivations behind Caffeine's design and implementation decisions, details the current state of Caffeine's development, and previews future work. We explain how the design and implementation offer benefits related to software sustainability by lowering the barrier to user contributions, reducing complexity through the use of Fortran 2018 C-interoperability features, and high performance through the use of a lightweight communication substrate.

Rouson, Damian↗

Idiomatic Correctness-Checking via Julienne in Fortran 2023

This paper presents a unified approach to unit testing and runtime assertion checking using Fortran 2023. The paper describes the support for our approach in the Julienne framework. Julienne leverages recent Fortran standards to implement object-oriented design patterns, support testing parallel programs, and implement functional programming patterns in order to craft idioms inspired by natural-language expressions. The presented idioms employ novel operators to write expressions that evaluate to a test-diagnosis object encapsulating two components: (1) the test outcome or assertion outcome and (2) an automatically generated diagnostic string. Two other novel aspects of the approach include (1) the ability to enforce assertions inside pure procedures and (2) the ability to output rich diagnostic information inside pure procedures during error termination when assertions fail. The latter capability mitigates against a reason that Fortran programmers commonly cite for not writing pure procedures: difficulty obtaining useful program output inside pure procedures when debugging code. This paper demonstrates how the adoption of the proposed idioms leads naturally to a unifying theme across two otherwise disparate technologies: unit testing and runtime assertion checking. Finally, this paper describes the usage of the Julienne testing framework for writing unit tests and assertions in the Matcha high-performance computing application and the Fiats deep learning library.

Rouson, Damian↗

Agile Acceleration of LLVM Flang Support for Fortran 2018 Parallel Programming

The LLVM Flang compiler ("Flang") is currently Fortran 95 compliant, and the frontend can parse Fortran 2018. However, Flang does not have a comprehensive 2018 test suite and does not fully implement the static semantics of the 2018 standard. We are investigating whether agile software development techniques, such as pair programming and test-driven development (TDD), can help Flang to rapidly progress to Fortran 2018 compliance. Because of the paramount importance of parallelism in high-performance computing, we are focusing on Fortran’s parallel features, commonly denoted “Coarray Fortran.” We are developing what we believe are the first exhaustive, open-source tests for the static semantics of Fortran 2018 parallel features, and contributing them to the LLVM project. A related effort involves writing runtime tests for parallel 2018 features and supporting those tests by developing a new parallel runtime library: the CoArray Fortran Framework of Efficient Interfaces to Network Environments (Caffeine).

Rasmussen, Katherine↗

Opportunities for enhancing MLCommons efforts while leveraging insights from educational MLCommons earthquake benchmarks efforts

MLCommons is an effort to develop and improve the artificial intelligence (AI) ecosystem through benchmarks, public data sets, and research. It consists of members from start-ups, leading companies, academics, and non-profits from around the world. The goal is to make machine learning better for everyone. In order to increase participation by others, educational institutions provide valuable opportunities for engagement. In this article, we identify numerous insights obtained from different viewpoints as part of efforts to utilize high-performance computing (HPC) big data systems in existing education while developing and conducting science benchmarks for earthquake prediction. As this activity was conducted across multiple educational efforts, we project if and how it is possible to make such efforts available on a wider scale. This includes the integration of sophisticated benchmarks into courses and research activities at universities, exposing the students and researchers to topics that are otherwise typically not sufficiently covered in current course curricula as we witnessed from our practical experience across multiple organizations. As such, we have outlined the many lessons we learned throughout these efforts, culminating in the need for benchmark carpentry for scientists using advanced computational resources. The article also presents the analysis of an earthquake prediction code benchmark while focusing on the accuracy of the results and not only on the runtime; notedly, this benchmark was created as a result of our lessons learned. Energy traces were produced throughout these benchmarks, which are vital to analyzing the power expenditure within HPC environments. Additionally, one of the insights is that in the short time of the project with limited student availability, the activity was only possible by utilizing a benchmark runtime pipeline while developing and using software to generate jobs from the permutation of hyperparameters automatically. It integrates a templated job management framework for executing tasks and experiments based on hyperparameters while leveraging hybrid compute resources available at different institutions. The software is part of a collection called cloudmesh with its newly developed components, cloudmesh-ee (experiment executor) and cloudmesh-cc (compute coordinator).

58 GEOSCIENCES↗

A Scalable Approach to Minimize Charging Costs for Electric Bus Fleets

Incorporating battery electric buses into bus fleets faces three primary challenges: a BEB’s extended refuel time, the cost of charging, both by the consumer and the power provider, and large compute demands for planning methods. When BEBs charge, the additional demands on the grid may exceed hardware limitations, so power providers divide a consumer’s energy needs into separate meters even though doing so is expensive for both power providers and consumers. Prior work has developed a number of strategies for computing charge schedules for bus fleets; however, prior work has not worked to reduce costs by aggregating meters. Additionally, because many works use mixed integer linear programs, their compute needs make planning for commercial-sized bus fleets intractable. This work presents a multi-program approach to computing charge plans for electric bus fleets. The proposed method solves a series of subproblems where the solution to the charge problem becomes more refined with each problem, moving closer to the optimal schedule. The results demonstrate how runtimes are reduced by using intermediate subproblems to refine the bus charge solution so that the proposed method can be applied to large bus fleets of 100+ buses. Not only will we demonstrate that runtimes scale linearly with the number of buses but we will also show how the proposed method scales to large bus fleets of over 100 buses while managing the monthly cost of energy.

Mortensen, Daniel (ORCID:0000000276494452)↗

Productive Programming of Distributed Systems with the SHAD C++ Library

High-performance computing (HPC) is often perceived as a matter of making large-scale systems (e.g., clusters) run as fast as possible, regardless the required programming effort. However, the idea of "bringing HPC to the masses" has recently emerged. Inspired by this vision, we have designed SHAD, the Scalable High-performance Algorithms and Data-structures library. SHAD is open source software, written in C++, for C++ developers. Unlike other HPC libraries for distributed systems, which rely on SPMD models, SHAD adopts a shared-memory programming abstraction, to make C++ programmers feel at home. Underneath, SHAD manages tasking and data-movements, moving the computation where data resides and taking advantage of asynchrony to tolerate network latency. At the bottom of his stack, SHAD can interface with multiple runtime systems: this not only improves developer’s productivity, by hiding the complexity of such software and of the underlying hardware, but also greatly enhance code portability. Thanks to its abstraction layers, SHAD can indeed target different systems, ranging from laptops to HPC clusters, without any need for modifying the user-level code. We have prototyped and open-sourced the implementation of (a subset of) the C++ standard library (STL) targeting multi-node HPC clusters. Our work allows plain STL-based C++ code to scale on HPC systems, with no need for rewriting the code to exploit the complex hardware. SHAD is available under Apache v2 License at https://github.com/pnnl/SHAD. In this paper we overview the design of the SHAD library, depicting its main components: runtime systems abstractions for tasking; parallel and distributed data-structures; STL-compliant interfaces and algorithms.

Castellana, Vito G.↗

System and method for characterization of retrofit opportunities in building using data from communicating thermostats

Systems and methods for characterization of retrofit opportunities are described. The methods may comprise computing, using at least one computing device disposed remote from a building and based at least in part on heating, ventilation and air conditioning (HVAC) runtime data associated with the building, one or more thermal characteristics of the building. In some embodiments, a model-predicted indoor temperature may be fitted against thermal data measured by a thermostat at the building. The thermal characteristic of the building may comprise a thermal insulation, an air leakage rate and/or an HVAC efficiency. The method may be used to determine, using the at least one computing device, suitability of the building for a retrofit opportunity to improve energy efficiency of the building. Determining the suitability may comprise evaluating the one or more thermal characteristics. The HVAC runtime data may be computed based on data received from a thermostat or a meter, such as an electric or a gas meter.

Zeifman, Michael↗