Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “runtime”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

Summary of CTF Modeling and Numerical Improvements for Boiling Water Reactor Simulation

This report documents geometry and numerical improvements made to CTF for the modeling of boiling water reactor (BWR) geometry and operating conditions. These activities are part of a larger program to extend the Virtual Environment for Reactor Applications (VERA) to better support BWR modeling and simulation. The activities documented in this report added features to CTF, including support for mixed-fuel cores, modeling of the upper plenum, and modeling of the lower tie plate form losses. A review of the spacer grid modeling approach was also performed, and a plan was discussed for future improvement. The parallelization of the model was improved, leading to a roughly 2× improvement in CTF runtime and a 1.6× improvement in total VERA runtime. An in-depth review of the governing equations and their linearization was performed and is documented in this report. Once implemented, this new linearization will allow CTF to take much larger timesteps, leading to more significant reductions in CTF and VERA runtimes.

21 SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLANTS↗

Unified Memory: GPGPU-Sim/UVM Smart Integration

CPU/GPU heterogeneous compute platforms are an ubiquitous element in computing and a programming model specified for this heterogeneous computing model is important for both performance and programmability. A programming model that exposes the shared, unified, address space between the heterogeneous units is a necessary step in this direction as it removes the burden of explicit data movement from the programmer while maintaining performance. GPU vendors, such as AMD and NVIDIA, have released software-managed runtimes that can provide programmers the illusion of unified CPU and GPU memory by automatically migrating data in and out of the GPU memory. However, this runtime support is not included in GPGPU-Sim, a commonly used framework that models the features of a modern graphics processor that are relevant to non-graphics applications. UVM Smart was developed, which extended GPGPU-Sim 3.x to in- corporate the modeling of on-demand pageing and data migration through the runtime. This report discusses the integration of UVM Smart and GPGPU-Sim 4.0 and the modifications to improve simulation performance and accuracy.

97 MATHEMATICS AND COMPUTING↗

UPC++ as_eager Working Group Draft, Revision 2020.6.2

This draft proposes an extension for a new future-based completion variant that can be more effectively streamlined for RMA and atomic access operations that happen to be satisfied at runtime using purely node-local resources. Many such operations are most efficiently performed synchronously using load/store instructions on shared-memory mappings, where the actual access may only require a few CPU instructions. In such cases we believe it’s critical to minimize the overheads imposed by the UPC++ runtime and completion queues, in order to enable efficient operation on hierarchical node hardware using shared-memory bypass. The new upcxx::{source,operation}_cx::as_eager_future() completion variant accomplishes this goal by relaxing the current restriction that future-returning access operations must return a non-ready future whose completion is deferred until a subsequent explicit invocation of user-level progress. This relaxation allows access operations that are completed synchronously to instead return a ready future, thereby avoiding most or all of the runtime costs associated with deferment of future completion and subsequent mandatory entry into the progress engine. We additionally propose to make this new as_eager_future() completion variant the new default completion for communication operations that currently default to returning a future. This should encourage use of the streamlined variant, and may provide performance improvements to some codes without source changes. A mechanism is proposed to restore the legacy behavior on-demand for codes that might happen to rely on deferred completion for correctness. Finally, we propose a new as_eager_promise() completion variant that extends analogous improvements to promise-based completion, and corresponding changes to the default behavior of as_promise().

97 MATHEMATICS AND COMPUTING↗

System and method of storing and analyzing information

A system and method of storing and analyzing information is disclosed. The system includes a compiler layer to convert user queries to data parallel executable code. The system further includes a library of multithreaded algorithms, processes, and data structures. The system also includes a multithreaded runtime library for implementing compiled code at runtime. The executable code is dynamically loaded on computing elements and contains calls to the library of multithreaded algorithms, processes, and data structures and the multithreaded runtime library.

Feo, John T.↗

Reducing the cost of energy estimation in the variational quantum eigensolver algorithm with robust amplitude estimation

Quantum chemistry and materials is one of the most promising applications of quantum computing. Yet much work is still to be done in matching industry-relevant problems in these areas with quantum algorithms that can solve them. Most previous efforts have carried out resource estimations for quantum algorithms run on large-scale fault-tolerant architectures, which include the quantum phase estimation algorithm. In contrast, few have assessed the performance of near-term quantum algorithms, which include the variational quantum eigensolver (VQE) algorithm. Recently, a large-scale benchmark study [Gonthier et al. 2020] found evidence that the performance of the variational quantum eigensolver for a set of industry-relevant molecules may be too inefficient to be of practical use. This motivates the need for developing and assessing methods that improve the efficiency of VQE. In this work, we predict the runtime of the energy estimation subroutine of VQE when using robust amplitude estimation (RAE) to estimate Pauli expectation values. Under conservative assumptions, our resource estimation predicts that RAE can reduce the runtime over the standard estimation method in VQE by one to two orders of magnitude. Despite this improvement, we find that the runtimes are still too large to be practical. These findings motivate two complementary efforts towards quantum advantage: 1) the investigation of more efficient near-term methods for ground state energy estimation and 2) the development of problem instances that are of industrial value and classically challenging, but better suited to quantum computation.

Johnson, Peter D.↗

System and method for characterization of air leakage in building using data from communicating thermostats and/or interval meters

Systems and methods for characterization of retrofit opportunities are described. Some embodiments are directed to methods for determining the air leakage rate of a building, and accordingly, for determining suitability of sealing of air leaks to improve the energy efficiency of a building. The methods may comprise computing, using at least one computing device disposed remote from a building and based at least in part on heating, ventilation and air conditioning (HVAC) runtime data associated with the building, one or more thermal characteristics of the building. The HVAC runtime data may be computed based on data received from a thermostat or a meter, such as an electric or a gas meter. To isolate the impact of air leakage, subsets of the HVAC runtime data at time intervals selected to have substantially the same conditions, but different wind speeds, may be computed.

Zeifman, Michael↗

Application of linear prolongation to coarse mesh finite difference acceleration in CASMO5

The Coarse-Mesh Finite Difference (CMFD) method has been used for over a decade to accelerate the convergence of the Method of Characteristics (MOC) solution to the two- dimensional particle transport equation in CASMO5. Numerical testing, along with widespread use in production-level calculations, have shown that the current CMFD implementation provides stability and robustness for a wide range of realistic reactor physics problems. However, the recent development of linear prolongation has attracted attention from the community as a way to further improve the performance and stability of CMFD. Two interpolation methods for linear prolongation are presented in this work and implemented into a test version of CASMO5. The performance of the proposed interpolations, referred to as the bilinear and linear directional schemes, is evaluated in terms of runtime relative to the default constant or uniform scaling update. Numerical results indicate that the use of linear prolongation can reduce the transport solver runtime on average by approximately 10% when tested with two hundred randomly selected cases. The new directional linear interpolation, combined with default constant boundary updates, is found to provide the highest reduction in runtime for the cases analyzed. (authors)

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

LC-MEMENTO: A Memory Model for Accelerated Architectures

With the advent of heterogeneous architectures, in particular, with the ubiquity of multi-GPU systems, it is becoming increasingly important to manage device memory efficiently in order to reap the benefits of the additional core count. To date, such responsibility mainly falls on the programmer where device-to-host data communication (and vice versa), if not done properly, may incur costly memory transfer operations and synchronization. The problem may be compounded by additional requirement to maintain system-wide memory consistency that may involve expensive synchronization overhead. In this paper, we present Location Consistency Memory Model for Enhanced Transfer Operations (LC-MEMENTO). This framework considers incorporating runtime techniques for multi-GPU memory management to support relaxed synchronization semantics and memory transfer operations automatically. Specifically, we implement a relaxed form of a memory consistency model based on the Location Consistency (LC) in an Asynchronous Many-Task Runtime (ARTS) and demonstrate that, this memory model enables additional optimization opportunities for the three representative applications encompassing different computational patterns (scientific computation, graphs, data streaming, etc.).

Memory Models, Accelerators, Adaptive Optimization↗

Domain-Specific Type-Safe APIs for Hierarchical Scientific Data with Modern C++

General-purpose library application programming interfaces (APIs) for self-describing hierarchical scientific data storage, such as the HDF5 and NetCDF libraries, are traditionally of runtime nature. Runtime errors for entry existence and data types are typically caught later in the development process of higher-level application-specific APIs. In this paper, we propose exploiting modern C++ metaprogramming features to add compile-time type-safety to improve the interaction with a well-defined metadata-rich scientific schema in domain-specific hierarchical datasets. We tackle two aspects of common use: (i) direct data access, (ii) flexible “in-memory” index models for efficient search and data processing. The proposed APIs use C++17’s template type auto deduction features, C++11’s enum class for type-safety and C-style preprocessor macros for generative templated code. We showcase the pros and cons of our initial work on the standard NeXus schema used for annotating and storing experimental neutron scattering data at several facilities around the world on top of HDF5. Extendable compile-time type-safe APIs are a desirable feature that could be indexed by any modern integrated development environment (IDE). Hence, such APIs can help ease the learning curve for domain scientists using a less error-prone software interaction to enhance the findability of their data without resorting to a domain-specific language (DSL).

Godoy, William↗

$\mathrm{PPT}$-Multicore: performance prediction of Open$\mathrm{MP}$ applications using reuse profiles and analytical modeling

In this report we present PPT-Multicore, an analytical model embedded in the Performance Prediction Toolkit (PPT) to predict parallel applications’ performance running on a multicore processor. PPT-Multicore builds upon our previous work towards a multicore cache model. We extract LLVM basic block labeled memory trace using an architecture-independent LLVM-based instrumentation tool only once in an application’s lifetime. The model uses the memory trace and other parameters from an instrumented sequentially executed binary. We use probabilistic and computationally efficient reuse profiles to predict the cache hit rates and runtimes of OpenMP programs’ parallel sections. We model Intel’s Broadwell, Haswell, and AMD’s Zen2 architectures and validate our framework using different applications from PolyBench and PARSEC benchmark suites. The results show that PPT-Multicore can predict cache hit rates with an overall average error rate of 1.23% while predicting the runtime with an error rate of 9.08%.

97 MATHEMATICS AND COMPUTING↗

Optimal operation of multi-plant steam district heating systems for enhanced efficiency and sustainability

Despite their crucial role in supplying heat and power to universities, industries, and healthcare facilities, many steam-based district heating systems rely on outdated control methods. Among these, multi-central plant districts are particularly challenging due to the complexities of coordinating multiple plants, optimizing load distributions, and managing system downtime. In response, new operational strategies are developed to enhance the efficiency and sustainability of steam districts while utilizing existing resources. These strategies include reducing plant operational pressure without compromising the reliable supply to buildings and optimizing load allocation across multiple plants. The load allocation considers boiler part-load efficiency, runtime, network losses, and building pressure set points, and is compared with traditional multi-boiler controls. To support this exploration, new dynamic Modelica models are developed. In addition, methods to reduce modeling complexities are incorporated, enhancing their suitability for practical applications. A holistic district-wide analysis using a real university case study demonstrates a 4.7% fuel savings by lowering boiler operational pressure from 900 kPa to 600 kPa, along with a 13.3% reduction in condensation losses across the distribution network. Furthermore, the load allocation approach results in a 13.1% reduction in fuel consumption during peak winter periods and 15.3% during shoulder periods, with corresponding decreases in carbon emissions and fuel costs. This approach can also save maintenance costs by reducing the boiler runtime by 49.6%. In conclusion, this research underscores the benefits of retrofitting aging steam district heating systems, offering immediate operational improvements by enhancing efficiency, meeting regulatory compliance, and extending infrastructure lifespans while delaying costly overhauls.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Efficient distributed continual learning for steering experiments in real-time

Deep learning has emerged as a powerful method for extracting valuable information from large volumes of data. However, when new training data arrives continuously (i.e., is not fully available from the beginning), incremental training suffers from catastrophic forgetting (i.e., new patterns are reinforced at the expense of previously acquired knowledge). Training from scratch each time new training data becomes available would result in extremely long training times and massive data accumulation. Rehearsal-based continual learning has shown promise for addressing the catastrophic forgetting challenge, but research to date has not addressed performance and scalability. To fill this gap, we propose an approach based on a distributed rehearsal buffer that efficiently complements data-parallel training on multiple GPUs to achieve high accuracy, short runtime, and scalability. It leverages a set of buffers (local to each GPU) and uses several asynchronous techniques for updating these local buffers in an embarrassingly parallel fashion, all while handling the communication overheads necessary to augment input minibatches using unbiased, global sampling. We further propose a generalization of rehearsal buffers to support both classification and generative learning tasks, as well as more advanced rehearsal strategies (notably Dark Experience Replay, leveraging knowledge distillation). We illustrate this approach with a real-life HPC streaming application from the domain of ptychographic image reconstruction. Furthermore, we run extensive experiments on up to 128 GPUs of the ThetaGPU supercomputer to compare our approach with baselines representative of training-from-scratch (the upper bound in terms of accuracy) and incremental training (the lower bound). Results show that rehearsal-based continual learning achieves a top-5 validation accuracy close to the upper bound, while simultaneously exhibiting a runtime close to the lower bound.

Asynchronous data management↗

Secure boot, trusted boot and remote attestation for ARM TrustZone-based IoT Nodes

With the extensive application of IoT techniques, IoT devices have become ubiquitous in daily lives. Meanwhile, attacks against IoT devices have emerged to compromise IoT devices by tampering with system pre-installed programs or injecting new malware. To mitigate these attacks, integrity enforcement of IoT systems has been proposed. The integrity of an IoT device system includes load-time integrity and runtime integrity. In this paper, we design an IoT system based on ARM TrustZone to enforce the system integrity. First, we establish the root of trust and propose a hybrid booting approach consisting of both secure boot and trusted boot to enforce the system load-time integrity. Second, we investigate a paging-based process integrity measurement method to measure the NW processes and conduct remote attestation based on the measurement results ensuring the NW runtime process integrity. We implement an IoT prototype system on a NXP i.MX6Q SABRE SD development board to assess its feasibility. Finally, real-world experiment results demonstrate that our prototype introduces negligible performance overhead to the original system.

97 MATHEMATICS AND COMPUTING↗

Darshan for HEP applications

Modern HEP workflows must manage increasingly large and complex data collections. HPC facilities may be employed to help meet these workflows’ growing data processing needs. However, a better understanding of the I/O patterns and underlying bottlenecks of these workflows is necessary to meet the performance expectations of HPC systems.Darshan is a lightweight I/O characterization tool that captures concise views of HPC application I/O behavior. It intercepts application I/O calls at runtime, records file access statistics for each process, and generates log files detailing application I/O access patterns.Typical HEP workflows include event generation, detector simulation, event reconstruction, and subsequent analysis stages. A study of the I/O behavior of the ATLAS simulation and filtering stage, and the CMS simulation workflow using Darshan is presented, including insights into the I/O operations and data access size.

Wang, Rui↗

On-the-fly data set combinations with RNTuple

With the expected data volume increase for HL-LHC and the even more complex computing challenges set by future colliders, the need for efficient data storage and processing becomes more pressing. ROOT’s next-generation data format and I/O subsystem, RNTTuple, is designed to address these challenges. RNTTuple already demonstrates a clear improvement in storage and I/O efficiency, as well as overall stability and robustness with respect to its predecessor, TTTree. These improvements provide a solid baseline to introduce novel extensions to common high-energy and nuclear physics (HENP) workflows. Notably, many workflows could benefit from the ability to arbitrarily join and chain data set samples at runtime, which could reduce overall storage requirements and improve application runtime and ergonomics. In this paper, we present the RNTupleProcessor, which enables HENP data set combinations with RNTuple. We will discuss the main design considerations, present the interfaces to support data set combinations and show how they integrate in typical workflows.

de Geus, Florine Willemijn [CERN; Twente U., Ensch↗

Development of a metamodelling framework for building energy models with application to fifth-generation district heating and cooling networks

Fully defined physics-based building energy models can accurately represent building systems; however, generating models based on high-level parameters is time consuming and simulation time of complex models can be slow. This article discusses the development of a Metamodelling Framework to create metamodels from a building energy modelling dataset. The framework generates metamodels using either linear regression, random forests, or support vector regressions. A fifth-generation district heating and cooling system analysis use case was used to motivate the development of the framework. The use case required quick and accurate representations of annual building loads reported hourly. Typical annual building modelling approaches can result in a runtime of 10 min. The metamodels runtime was reduced to less than 10 s to load and run an annual simulation with user-defined covariates. The results of the metamodel performance and an abbreviated topology analysis based on the motivating use case will be presented.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

HPC-driven computational reproducibility in numerical relativity codes: a use case study with IllinoisGRMHD

Abstract Reproducibility of results is a cornerstone of the scientific method. Scientific computing encounters two challenges when aiming for this goal. Firstly, reproducibility should not depend on details of the runtime environment, such as the compiler version or computing environment, so results are verifiable by third-parties. Secondly, different versions of software code executed in the same runtime environment should produceconsistent numerical results for physical quantities. In this manuscript, we test the feasibility of reproducing scientific results obtained using theIllinoisGRMHDcode that is part of an open-source community software for simulation in relativistic astrophysics, theEinstein Toolkit. We verify that numerical results of simulating a single isolated neutron star withIllinoisGRMHDcan be reproduced, and compare them to results reported by the code authors in 2015. We use two different supercomputers: Expanse at SDSC, and Stampede2 at TACC. By compiling the source code archived along with the paper on both Expanse and Stampede2, we find thatIllinoisGRMHDreproduces results published in its announcement paper up to errors comparable to round-off level changes in initial data parameters. We also verify that a current version ofIllinoisGRMHDreproduces these results once we account for bug fixes which have occurred since the original publication.

Astronomy & Astrophysics↗

An open-source framework for balancing computational speed and fidelity in production cost models

Studies of bulk power system operations need to incorporate uncertainty and sensitivity analyses, especially around exposure to weather and climate variability and extremes, but this remains a computational modeling challenge. Commercial production cost models (PCMs) have shorter runtimes, but also important limitations (opacity, license restrictions) that do not fully support stochastic simulation. Open-source PCMs represent a potential solution. They allow for multiple, simultaneous runs in high-performance computing environments and offer flexibility in model parameterization. Yet, developers must balance computational speed (i.e. runtime) with model fidelity (i.e. accuracy). In this paper, we present Grid Operations (GO), a framework for instantiating open-source, scale-adaptive PCMs. GO allows users to search across parameter spaces to identify model versions that appropriately balance computational speed and fidelity based on experimental needs and resource limits. Results provide generalizable insights on how to navigate the fidelity and computational speed tradeoff through parameter selection. We show that models with coarser network topologies can accurately mimic market operations, sometimes better than higher-resolution models. It is thus possible to conduct large simulation experiments that characterize operational risks related to climate and weather extremes while maintaining sufficient model accuracy.

42 ENGINEERING↗