Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “runtime”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 577 records · Page 32

GANpiler

Focal Area(s): Rather than augment or replace physical models with machine learning, we instead preserve the existing models and augment the underlying code with faster surrogate models created by generative adversarial networks (GAN). To leverage the performance optimization, we also propose a runtime system and user interfaces that allow prediction and tracking of accumulated error as well as dynamic, per-process decision making as to which model (if any) to use for each iteration. Science Challenge: This proposal sits at the nexus of two hard problems. First, climate models based on machine learning will be, by their nature, difficult to trust once their predictions begin diverging from the consensus. Second, compilers and hardware have been making only incremental performance gains for decades. GPGPUs have provided a welcome performance boost for codes that can take advantage of them, but there is no similar technology on the horizon to provide the next performance leap.

54 ENVIRONMENTAL SCIENCES↗

Concurrent Optimization of Capital Cost and Expected O&M

Concentrating solar power (CSP) technologies can utilize heat from concentrated sunlight from a field of tracking mirrors to generate electricity, reform fuel, provide process heat, or augment fossil plant heat sources. Electricity-generating power tower systems focus light from thousands of independent heliostats onto a thermal receiver, which uses the focused light to warm a heat transfer fluid (HTF), typically, a molten nitrate salt. The HTF is then sent to a power generation cycle or diverted into thermal energy storage (TES) for later use. Thermal storage is – in principle – a straightforward proposition. However, optimal utilization of a TES resource is complex and multi-faceted: thermal energy may be dispatched to produce electricity immediately upon first availability, or thermal energy may be reserved for next-day peak periods at risk of filling storage and dumping energy, or a portion of the thermal energy can be reserved to maintain equipment temperatures, reducing power cycle startup time, etc. Many possible dispatch permutations variously emphasize producing peak power, operating through transients, expediting daily startup, etc. The best operation strategy can change day-to-day throughout the year, depending on the weather and market pricing forecasts. The project we describe in this report develops a software package that allows users to explore design optimization, operations decisions, and performance characterization of concentrating solar power tower plants. Users interface with the tool through a scripting language, and results are reported in time series tables, plots, runtime logs, and design outputs. Users choose from a list of variables such as tower height, solar multiple, design-point irradiance, thermal storage size, etc., and specify information about the system using a list of parameters. The software can then optimize the specified variables to reduce the cost of energy produced by the system while meeting certain production requirements, accounting for uncertain weather and electricity price forecasts, and correcting for equipment failures or repair time. The software we develop is the first comprehensive design tool of its kind to incorporate all of these aspects while being deployed as open source.

14 SOLAR ENERGY↗

Field Validation of a Smart Energy Recovery Ventilation System Using Low-Cost Indoor Air Quality Sensors

This project is a field validation, using low-cost indoor air quality (IAQ) sensors, of a smart ventilation system that can help low-load homes in humid environments maintain acceptable indoor humidity conditions while providing adequate ventilation according to ASHRAE 62.2. The objectives of this research were to (1) address builders’ concerns with mechanical ventilation in humid environments and (2) answer the question of whether smart control logic helps with occupant comfort and the creation of a more acceptable indoor environment. To address the objectives of the study, the Southface team collected field data for one year in four Charleston, South Carolina, new construction homes in order to determine the differences in occupant comfort; comfort metrics; IAQ; and heating, ventilating, and air-conditioning (HVAC) energy consumption when toggling biweekly between an energy recovery ventilator (ERV) operating continuously and an ERV operating with smart, time-varying humidity control logic. The smart ventilation algorithm under consideration in this field test did create a less humid indoor environment on an annual basis as quantitatively measured through temperature and relative humidity (T/RH) readings, expressed most discernably as “percentage of time above 60% RH” and “percentage of time above 55°F dewpoint.” However, the difference it made was inconsistent during the spring, summer, and fall months, and it was only directionally consistent during the winter months. We suspect that this is primarily due to the long runtimes and concomitant dehumidification activity of the air-conditioning (A/C) units in response to the high sensible loads in Charleston. The effect of the smart ventilation algorithm was not discernable to the occupants in this study, as recorded through seasonal surveys.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

STNS01-21 BEE - FY21 P6-2: Archive, clone, and re-run workflows [Slide]

BEE will give ECP a tool that great simplifies the deployment of containerized workflows on the next generation of pre-exascale and exascale systems, as well as public and private clouds. BEE allows scientists to describe their workflow using the Common Workflow Language and then deploy that workflow across the entire spectrum of systems without having to learn the specifics of each container runtime, HPC resource manager, or cloud API. BEE also streamlines the curation and sharing of common workflows among the scientific community.

97 MATHEMATICS AND COMPUTING↗

Evaluating Spatial Accelerator Architectures with Tiled Matrix-Matrix Multiplication

There is a growing interest in custom spatial accelerators for machine learning applications. These accelerators employ a spatial array of processing elements (PEs) interacting via custom buffer hierarchies and networks-on-chip. The efficiency of these accelerators comes from employing optimized dataflow (i.e., spatial/temporal partitioning of data across the PEs and fine-grained scheduling) strategies to optimize data reuse. The focus of this work is to evaluate these accelerator architectures using a tiled general matrix-matrix multiplication (GEMM) kernel. To do so, we develop a framework that finds optimized mappings (dataflow and tile sizes) for a tiled GEMM for a given spatial accelerator and workload combination, leveraging an analytical cost model for runtime and energy. Our evaluations over five spatial accelerators demonstrate that the tiled GEMM mappings systematically generated by our framework achieve high performance on various GEMM workloads and accelerators.

43 PARTICLE ACCELERATORS↗

LDMS-GPU: Lightweight Distributed Metric Service (LDMS) for NVIDIA GPGPUs

GPUs are now a fundamental accelerator for many high-performance computing applications. They are viewed by many as a technology facilitator for the surge in fields like machine learning and Convolutional Neural Networks. To deliver the best performance on a GPU, we need to create monitoring tools to ensure that we optimize the code to get the most performance and efficiency out of a GPU. Since NVIDIA GPUs are currently the most commonly implemented in HPC applications and systems, NVIDIA tools are the solution for performance monitoring. The Light-Weight Distributed Metric System (LDMS) at Sandia is an infrastructure widely adopted for large-scale systems and application monitoring. Sandia has developed CPU application monitoring capability within LDMS. Therefore, we chose to develop a GPU monitoring capability within the same framework. In this report, we discuss the current limitations in the NVIDIA monitoring tools, how we overcame such limitations, and present an overview of the tool we built to monitor GPU performance in LDMS and its capabilities. Also, we discuss our current validation results. Most of the performance counter results are the same in both vendor tools and our tool when using LDMS to collect these results. Furthermore, our tool provides these statistics during the entire runtime of the tool as a time series and not just aggregate statistics at the end of the application run. This allows the user to see the progress of the behavior of the applications during their lifetime.

97 MATHEMATICS AND COMPUTING↗

Task Parallelism to Optimize Performance of Environmental Modeling Software

Climate modeling is an integral part of environmental research, from studying rare phenomena to predicting future climate trends. The need for more accurate models is only growing, but as climate modeling capabilities advance, existing workflows require optimization to recoup performance. A solution comes in the form of task parallelism, a novel programming capability that provides an opportunity for optimization at execution time by allowing tasks to be executed in parallel, reducing runtime significantly. Using Parsl, an intuitive and scalable parallel scripting library for Python, we implement task parallelism within support software to aid in the continuous advancement of climate modeling technology.

54 ENVIRONMENTAL SCIENCES↗

FleCSI Intro 2.0 [Slides]

FleCSI is a programming system or framework for developing mulitphysics application codes. FleCSI provides two primary capabilities--runtime abstraction layer and topology data structures.

97 MATHEMATICS AND COMPUTING↗

A novel MC TRT method: vectorizable variance reduction for the energy spectra

Data produced up to now gives us confidence that we will continue to see variance reduction behavior at amenable runtimes as we move this test bed forward. This method has characteristics of event and history based Monte Carlo transport, without the bulky modifications required by event based Monte Carlo while avoiding some overhead.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Assessment of the Griffin Reactor Multiphysics Application Using the Empire Micro Reactor Design Concept

In late 2019, INL and ANL agreed to jointly develop the reactor physics code named Griffin based on the integration of the two code suites, MAMMOTH/Rattlesnake (INL) and MC2 - 3/PROTEUS (ANL). Griffin is being developed based on the MOOSE framework and MOOSE quality assurance procedures. This decision was made to be able to allow DOE-NE to efficiently invest funding to this area and to provide effective and timely support for existing and potential users; the latter includes industry and government organizations who are developing various types of advanced reactors in the near and long term. Since MAMMOTH/Rattlesnake has been developed based on the MOOSE framework, the INL/ANL Griffin development team agreed to build Griffin beginning with a merger of MAMMOTH and Rattlesnake into a single code and moving forward by implementing capabilities from the PROTEUS suite into Griffin. Moving forward, both ANL and INL efforts are equally invested in the Griffin project, with management support, to provide an advanced reactor multiphysics tool to assist in reactor design, optimization, and safety analysis. Much work remains in moving Griffin forward to migrate PROTEUS capabilities and to optimize performance to meet user needs. The main objective of this work is to assess the current status of Griffin capabilities in terms of performance and accuracy, to determine priorities for PROTEUS migration, and to identify capabilities and features to improve for supporting the code integration effort. For this assessment, the Empire micro reactor problem that was developed in the ARPA-E MEITNER program was selected as an advance reactor concept of interest to the technical community. The Empire reactor problem was expanded from its original incomplete specification to be a small heat-pipe-cooled micro reactor core with ~113 cm radius and 70 cm in height, composed of 18 fuel assemblies, 12 control drums, and beryllium radial and axial reflectors. In the current model, using 5 cm axial reflectors specified in the original Empire assembly model, more than 10% of neutrons leak axially and through the empty center safety hole, as well as through heat pipe channels in fuel assembly elements that extend through the top reflector region. Several calculation models of the core were defined for systematic assessment, including 2-D and 3-D fuel assemblies and whole cores with cylindrical boundaries. Cross sections were generated using Serpent 2, and meshes were produced using the Argonne mesh tool or the INL neutronics meshing tools combined with CUBIT. Cross sections and meshes were converted to the ISOXML and Exodus formats, respectively, so that Griffin and PROTEUS could use consistent data for solving the reactor problems. With the prepared cross sections and meshes, PROTEUS was run first to ensure that all input data were correctly generated and input options in terms of angle, mesh, and energy group were accurately determined. Comparisons against Serpent 2 solutions were made in terms of eigenvalue and pin power. The same calculations and comparisons were then conducted using Griffin. For the fuel assembly and whole core problems, the PROTEUS eigenvalues agreed well with reference Serpent 2 solutions within 100 and 30 pcm, respectively, and pin power differences relative to Serpent 2 were overall less than 2.2% and RMS 0.8% for the whole core models. This indicated that all input data were properly prepared. Using the same data, Griffin was run selecting the SAAF-CFEM SN solver with Legendre-Gaussian quadrature and NDA and DSA for acceleration. It was found that the SAAF-CFEM solver of Griffin required finer meshes to achieve eigenvalue and pin power solutions in good agreement with Serpent 2, consequently requiring more memory requirement and longer computation time. On the other hand, the SPH-Diffusion 2-D core calculations performed using Griffin were able to recover the exact eigenvalue from the reference Serpent 2 solutions, resulting in a pin-power distribution with an RMS of 0.6% and maximum absolute difference of less than 1.4%. The runtimes for SPH-Diffusion for the 2-D core were less than 3 minutes on 40 cores. During this evolution of this evaluation, many updates were made in Griffin by the Griffin development team of INL (focusing on software updates) and ANL (reviewing and supporting software updates) to complete this assessment. Observations from the code assessment are presented in the conclusion section of this report, followed by a discussion of recommendations for future work.

21 SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLANTS↗

Position Papers for the ASCR Workshop on Cybersecurity and Privacy for Scientific Computing Ecosystems

At the request of the Department of Energy's (DOE) Office of Advanced Scientific Computing Research (ASCR), this program committee has been tasked with organizing a workshop to identify basic research needs in cybersecurity and privacy to better support DOE's science and energy mission. As part of the process, the program committee is soliciting community input in the form of position papers to help identify significant use cases, facility issues, and other barriers to enabling verifiably trustworthy computational science while preserving data confidentiality as appropriate for scientific workflows of interest to DOE. The program committee will review these position papers and based on the fit of their area of expertise and interest, selected contributors will have the opportunity to participate in the workshop currently planned as a virtual event November 3-5th, 2021. The thrust areas that will be explored by this workshop are the following: (1) Algorithms for secure, scalable, privacy-enhancing technologies and frameworks, including: Federated AI/ML, Differential privacy, Randomized algorithms, Adversarial modeling & simulation, Graph algorithms, and Formal methods; (2) Platforms to support the entire scientific-computing ecosystem, including edge computing for large-scale experiments, focusing on heterogeneous systems and distributed systems, including: Heterogeneous computing systems, Distributed computing systems, and Secure data architectures; and (3) Data workflows to allow agile use of data while preserving integrity and privacy, making the important properties verifiable either at runtime or post-computation, including: Integrity and provenance and Data management infrastructure. Topics that are out-of-scope for the workshop include discussing specific proposed solutions or areas that are clearly out of DOE's fundamental and applied-sciences mission scope, e.g., cryptography, enterprise security, and general-operations technology.

97 MATHEMATICS AND COMPUTING↗

Barriers to Broader Utilization of Fault Detection Technologies for Improving Residential HVAC Equipment Efficiency

Faults in residential heating, ventilating, and air conditioning (HVAC) equipment may occur due to poor installation practices or develop over time, and these faults can negatively impact system efficiency, thermal comfort, and equipment lifespan. Automated fault detection and diagnostic (AFDD) technologies identify energy wasting HVAC faults, such as low indoor airflow and improper refrigerant charge, and guide technicians in improving system efficiency. For residential HVAC, AFDD consists of a range of fault detecting and diagnostic capabilities, sensor configurations, and target applications. AFDD technology can either be permanently installed by the original equipment manufacturer (OEM) using embedded sensors or as an add-on product either during or after installation. Additionally, several advanced installation tools and refrigerant gauge sets include AFDD features for temporary use during equipment installation and tune-ups. Some technologies can detect a fault but have limited diagnostic capabilities. For example, a single-point measurement from the home's thermostat or energy monitor can provide certain fault detection capability by analyzing the equipment runtime or energy consumption. These technologies, though limited at determining the cause of a given fault, may have significant energy savings potential due to their low cost and prevalence in the residential HVAC market. Despite the potential benefits, fault detection technologies face many technical and market barriers preventing broad adoption. Beyond the cost barriers due to the added sensor requirements and technology development, fault detection technologies face many implementation and adoption barriers such as installer training, customer awareness, standardized communication protocols, and methods of test for evaluating accuracy. The purpose of this whitepaper is to characterize market and technical barriers impeding broader utilization of fault detection technology for residential HVAC energy efficiency applications.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Hierarchical Division and Clustering of Group Structures

When using multigroup neutron transport models, it is important to choose a suitable group structure due to the impact on runtime and accuracy. Group structures at LANL have been largely pared down to a handful of commonly used group structures, such as the LANL-30 and LANL-70 group structures. By identifying the key boundaries for a given problem, more insight can be gained into why the commonly used group structures are accurate and when they are inaccurate. This work focuses on using hierarchical division and agglomeration methods to identify the most critical boundaries in these group structures, using the ICSBEP HEU-MET-FAST-001-001 benchmark as a case study.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Status of the NEAMS and ARC fast reactor tools integration to the NEAMS Workbench

The Workbench initiative was launched in FY-2017 within the Nuclear Energy Advanced Modeling and Simulation (NEAMS) program to facilitate the transition from conventional tools to high-fidelity tools. The NEAMS Workbench provides a common user interface for model creation, real-time validation, execution, output processing, and visualization for integrated codes. The integration of the Argonne Reactor Computation (ARC) suite of codes into the NEAMS Workbench through the PyARC module was initiated in FY-2017. The ARC codes, which focus on fast reactor multiphysics analysis, contain both legacy codes like DIF3D and REBUS-3 that were developed with over 30 years of experience, and newer NEAMS additions like MC 2 -3, PERSENT, and PROTEUS. Recent work expended this integration to other tools to support the U.S. fast reactor community such as DASSH for sub-channel thermal-hydraulics, Griffin for high-fidelity deterministic transport calculations, and OpenMC for Monte Carlo simulations (with Shift integration planned for FY-2023). The integration of the “extended ARC” suite of codes into the NEAMS Workbench interface relies on the PyARC and PyGriffin modules to handles the pre- and post-processing of these codes input, and the runtime environment. The PyARC module together with the NEAMS Workbench interface are both released under Open Source Software licenses.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Anomaly Detection in Gamma Spectra Using Hopfield Neural Network with B-SAT and Grover’s Algorithm on a Quantum Computing Simulator

Environmental screening of gamma radiation consists of detecting weak nuisance and anomaly signal in the presence of strong and highly varying background. In a typical scenario, a mobile detector-spectrometer continuously measures gamma radiation spectra in short, e.g., one-second, signal acquisition intervals. In principle, gamma radiation sources can be detected and identified by their unique spectral lines. However, detecting sources from data measured in a search scenario is difficult due to the highly varying background because of naturally occurring radioactive material (NORM), and low signal-to-noise ratio (S/N) of spectral signal measured during one-second acquisition intervals. In recent prior work, we have developed a Hopfield Neural Network (HNN) in conjunction with an image processing algorithm to detect a weak signal anomaly hidden among the highly fluctuating background spectra. The objective of this work is to explore quantum computing methods to increase the speed of HNN. The approach is based on the Grover’s search algorithm in conjunction with a 3-SAT problem formalism. The Grover’s algorithm is implemented on a quantum computing simulator using Qiskit software. Performance of HNN algorithm is benchmarked using search data from an environmental screening campaign, where the anomaly is a subset of measurements containing a 137 Cs source. Results indicate that using Grover’s algorithm on a quantum simulator reduces runtime of HNN by two orders of magnitude.

61 RADIATION PROTECTION AND DOSIMETRY↗

Performance Analysis and Optimization for Scientific Data Workloads

Scientific data generated at experimental and observational facilities are increasingly being processed on large-scale compute systems. Most of the experimental data analysis workflows are not designed or implemented to run on large scale environments and take full advantage of HPC compute and storage resources. These applications are unlike the traditional tightly-coupled scientific applications and hence face significant performance and scalability challenges as the volume of data increases exponentially. In this paper, we conduct a performance and scalability analysis for experimental analysis applications and workflows operating on data from light sources. Our analysis detects and quantifies I/O performance, scalability and runtime bottlenecks for three data analysis applications that run on NERSC resources. Based on our analysis we propose and implement a set of optimizations that lead to reducing the amount of time spent on I/O operations by almost 90%.

97 MATHEMATICS AND COMPUTING↗

Modeling, Performance Assessment, and Nodal Data Analysis of TRISO-Fueled Systems with Shift

This technical report documents several enhancements to the Shift Monte Carlo (MC) code under the US Department of Energy (DOE) Nuclear Energy Advanced Modeling and Simulation (NEAMS) program in fiscal year (FY) 2022. Performance enhancements were added to Shift specifically for tristructural isotropic (TRISO)–fueled reactor systems and guided based on performance analysis in FY 2021. For the pebble performance model developed in previous studies, the runtime improved by ~ 91× compared to the original model and ~ 2× compared to the user-optimized model. Compared to Serpent, Shift is ~ 3× slower if Serpent delta-tracking is enabled but ~ 2× faster when delta-tracking is disabled. The multigroup cross section generation was improved through simplifying tally input definitions, porting several post-processing tally operations from Python scripts into the Shift code base, and accounting for production reactions in the scattering multiplicity. Progress was also made on two emerging capabilities: (1) the development of Titan (a Shift reactor physics user interface) and (2) initial investigation into path-length tallies for computing multigroup scattering matrices.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

EM Project v9

This video is a sizzle reel showcasing the new capabilities MCMQ has acquired after the expansion of their machine shop and quality control testing. The total video runtime is one minute and fifty-three seconds.

42 ENGINEERING↗