Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Benchmarking Software”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10

A Benchmark Suite for Evaluating Scientific AI Workloads on GPUs

AI applications have been steadily increasing in the allocation portfolio among leadership computing facilities. These applications depend on deep learning frameworks with hardware acceleration and underlying software systems. With the rapid development of applications, software stacks, and hardware devices, it is essential to evaluate the performance of core operations in AI workloads for direction of optimizations and procurement of next-generation high-performance computing (HPC) infrastructures. Currently, most benchmarks lack scientific AI workloads. So, we present DeepKernelBench and the experimental results of evaluating the benchmark suite for early observations and performance comparisons on datacenter GPUs using representative workloads for scientific AI, including Attentions, General matrix multiplications, Geometrics and Fourier neural operations.

Jin, Zheming [Advanced Micro Devices (AMD)]↗

Automatic Crack Segmentation and Feature Extraction in Electroluminescence Images of Solar Modules

The effect of cracks in solar cells on the long-term degradation of photovoltaic (PV) modules remains to be determined. To investigate this effect in future studies, it is necessary to quantitatively describe the crack features (e.g., length) and correlate them with module power loss. Electroluminescence (EL) imaging is a common technique for identifying cracks. However, it is currently challenging and time-consuming to identify cracks in a large number of EL images and quantify complex crack features by human inspection. This article introduces a fast semantic segmentation method (~0.18 s/cell) to automatically segment cracks from EL images and algorithms to extract crack features. Here we fine-tuned a UNet neural network model using pretrained VGG16 as the encoder and obtained an average F1 score of 0.875 and an intersection over union score of 0.782 on the testing set. With cracks and busbars segmented, we developed algorithms for extracting crack features, including the crack-isolated area, the brightness inside the isolated area, and the crack length. We also developed an automatic preprocessing tool for cropping individual cell images from EL images of PV modules (~0.72 s/module). Our codes are published as open-source an software, and our annotated dataset composed of various types of cells is published as a benchmark for crack segmentation in EL images.

14 SOLAR ENERGY↗

Establishing model credibility for process-microstructure-property relationships in additive manufacturing using exascale computing

Additive Manufacturing (AM) of alloys holds significant promise as a disruptive technology in various industries, yet its adoption is often hindered by challenges in achieving consistent part quality. These issues are primarily due to the complex process-microstructure-property (PSP) relationships inherent to AM. Computational models can greatly aid in understanding these relationships, but their widespread impact and adoption has been limited by a lack of validated, open-source, and computationally efficient PSP modeling frameworks and hardware limitations. Here, this study leverages the ExaAM software suite and data from the AMBench-2018 series of laser powder bed fusion (LPBF) benchmark experiments to perform a comprehensive model assessment, including verification, validation, sensitivity analysis, and uncertainty quantification. The RADICAL-EnTK workflow manager was used to perform an ensemble of heat transport, solidification, and mechanical response simulations on the exascale computer Frontier, considering uncertainties in critical model inputs such as laser spot size and nucleation parameters, and consisting of 125 explicit grain structure simulations and 7875 crystal plasticity simulations. For a selected location within the Inconel 625 AMBench-2018 test artifact, sensitivity analysis and uncertainty quantification were performed using the predicted distributions of grain structure and mechanical properties. Qualitative agreement was found between the predicted grain size and texture and the observed AMBench-2018 microstructure, the mean predicted yield stress was within 5% of the experimental measurement mean, and the mean predicted engineering stress at 5% strain was within 10% of the experimental measurement mean. The insights gained from development and validation of the ExaAM PSP modeling framework will help guide future directions for enhancing the credibility and reliability of PSP models in AM, thereby accelerating the adoption of AM technologies in various industries.

Additive manufacturing↗

BISON TRISO Modeling Advancements and Validation to AGR-1 Data

BISON is a finite element-based nuclear fuel performance code. Among its unique characteristics are its ability to model 1D, 2D, and 3D geometries and its applicability to a wide variety of nuclear fuels. For the last eight years, BISON has included a beginning capability to model tri-structural isotropic (TRISO) fuel. Recently, interest in TRISO fuel has grown, and a significant effort has been made to improve BISON’s capabilities in this area. Capability development has occurred for each material present in TRISO fuel particles: the buffer, inner pyrolytic carbon, silicon carbide, and outer pyrolytic carbon layers, as well as the fuel kernel. New elastic, creep, swelling, thermal expansion, thermal conductivity, and fission gas release (FGR) models are available. New models for the graphite matrix are also now available. Another important addition is the ability to perform statistical failure analysis of large samples of fuel particles. This new capability, which continues to grow, enables evaluation of failure due to pressure or crack formation by analyzing many thousands of particles. This enables realistic calculations of fission product release from the many particles in a TRISO-fueled reactor. These capabilities were checked via regression and verification tests. A large number of code benchmarking problems were also run, showing that BISON’s results closely match those of other software tools. Finally, a significant validation effort was completed in which fission product release, measured as part of the AGR-1 capsule experiments, was compared to BISON outputs. BISON outputs compared very well to the experimental data and to PARFUME results. Interest in BISON’s TRISO capabilities is growing, with the U.S. Nuclear Regulatory Commission (NRC) and Westinghouse Electric Company receiving training during the past year. Multiple other entities have expressed interest in or are actively using BISON. Kairos Power, LLC, has a strong partnership with Idaho National Laboratory (INL) regarding the use of BISON for TRISO analysis. While its capabilities still continue to grow, BISON has already become a powerful tool for TRISO analysis.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Comparing the OpenMP, MPI, and Hybrid Programming Paradigm on an SMP Cluster

Clusters of SMP (Symmetric Multi-Processors) nodes provide support for a wide range of parallel programming paradigms. The shared address space within each node is suitable for OpenMP parallelization. Message passing can be employed within and across the nodes of a cluster. Multiple levels of parallelism can be achieved by combining message passing and OpenMP parallelization. Which programming paradigm is the best will depend on the nature of the given problem, the hardware components of the cluster, the network, and the available software. In this study we compare the performance of different implementations of the same CFD benchmark application, using the same numerical algorithm but employing different programming paradigms.

Jost, Gabriele↗

DISARM: Target Electronic Device Informed Mitigation of Software Runtime Side-Channel Vulnerabilities

Program runtime/timing attacks exploit variations in a program’s execution times to extract sensitive information from the program (e.g. encryption keys, sensitive variable data, intellectual property). State-of-the-art solutions to runtime side-channel attacks attempt to balance the execution time of the sensitive code for different control flow paths to eliminate the timing leakage. However, during the mitigation process, most techniques do not consider the underlying hardware/device on which the target program is supposed to run on. This can lead to over-fixing (unnecessary extra operations), under-fixing (not solving the imbalance properly), and even failures. Here, we propose DISARM, a joint hardware-software methodology (unlike any existing solution) for mitigating runtime side-channel vulnerabilities that utilizes timing values from real embedded devices to generate targeted software fixes. We implement DISARM to support C/C++/Java source codes and validate it across 22 standard benchmarks. DISARM outperforms state-of-the-art solutions such as PENDULUM and DifFuzzaR in terms of execution time overhead, code size overhead, and correctness on five different embedded/edge devices.

Timing/runtime side-channel↗

DLIO: A DATA-CENTRIC BENCHMARK FOR DEEP LEARNING APPLICATIONS

SF-22-136 Deep learning has been shown as a successful method for various tasks, and its popularity results in numerous open-source deep learning software tools. Deep learning has been applied to a broad spectrum of scientific domains such as cosmology, particle physics, computer vision, fusion, and astrophysics. Scientists have performed a great deal of work to optimize the computational performance of deep learning frameworks. However, the same cannot be said for I/O performance. As deep learning algorithms rely on big-data volume and variety to effectively train neural networks accurately, I/O is a significant bottleneck on large-scale distributed deep learning training. DLIO, is a novel representative benchmark suite built based on the I/O profiling of the selected workloads. DLIO can be utilized to accurately emulate the I/O behavior of modern deep learning applications. Using DLIO, application developers and system software solution architects can identify potential I/O bottlenecks in their applications and guide optimizations to boost the I/O performance leading to lower training times. The storage vendor can also use DLIO as a guide for designing and optimize the storage and filesystem targeting at deep learning application.

ZHENG, HUIHUO↗

The QICK (Quantum Instrumentation Control Kit): Readout and control for qubits and detectors

We introduce a Xilinx RF System-on-Chip (RFSoC)-based qubit controller (called the Quantum Instrumentation Control Kit, or QICK for short), which supports the direct synthesis of control pulses with carrier frequencies of up to 6 GHz. The QICK can control multiple qubits or other quantum devices. The QICK consists of a digital board hosting an RFSoC field-programmable gate array, custom firmware, and software and an optional companion custom-designed analog front-end board. We characterize the analog performance of the system as well as its digital latency, important for quantum error correction and feedback protocols. We benchmark the controller by performing standard characterizations of a transmon qubit. We achieve an average gate fidelity of ℱ avg =99.93%. All of the schematics, firmware, and software are open-source.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Support CEED-enabled ECP applications in their preparation for Aurora/Frontier

The goal of this milestone was to help CEED-enabled ECP applications (particularly ExaSMR, MARBL, ExaAM, ExaWind and E3SM) in their preparations for the Aurora and Frontier architectures. This work included collaboration with ECP vendors and porting and optimization of CEED’s benchmarks and miniapps to early access hardware. As part of this milestone, we also made the best bake-off problems and bake-off kernel implementation from Nek, MFEM, libParanumal and the external community available in the latest libCEED release, libCEED-0.7. During the milestone period we also organized, in virtual form, the fourth CEED Annual meeting (CEED4AM) which included representatives from ECP applications, vendors and software technology projects. The specific tasks addressed in this milestone were to: (1) Work with vendors to port and run CEED benchmarks on early access systems for Aurora and Frontier; (2) Make the best BP/BK implementations from Nek, MFEM, libParanumal and external community available in libCEED; (3) Organize the next CEED Annual meeting (CEED4AM); and (4) Optimize CEED applications and miniapps for Aurora and Frontier architectures. The artifacts delivered include the next libCEED release, libCEED-0.7, and a number of developments integrated within applications to improve their GPU performance and capabilities. See the CEED website, http://ceed.exascaleproject.org and the CEED GitHub organization, http://github.com/ceed for more details.

97 MATHEMATICS AND COMPUTING↗

Performance Measurement, Visualization and Modeling of Parallel and Distributed Programs

This paper presents a methodology for debugging the performance of message-passing programs on both tightly coupled and loosely coupled distributed-memory machines. The AIMS (Automated Instrumentation and Monitoring System) toolkit, a suite of software tools for measurement and analysis of performance, is introduced and its application illustrated using several benchmark programs drawn from the field of computational fluid dynamics. AIMS includes (i) Xinstrument, a powerful source-code instrumentor, which supports both Fortran77 and C as well as a number of different message-passing libraries including Intel's NX Thinking Machines' CMMD, and PVM; (ii) Monitor, a library of timestamping and trace -collection routines that run on supercomputers (such as Intel's iPSC/860, Delta, and Paragon and Thinking Machines' CM5) as well as on networks of workstations (including Convex Cluster and SparcStations connected by a LAN); (iii) Visualization Kernel, a trace-animation facility that supports source-code clickback, simultaneous visualization of computation and communication patterns, as well as analysis of data movements; (iv) Statistics Kernel, an advanced profiling facility, that associates a variety of performance data with various syntactic components of a parallel program; (v) Index Kernel, a diagnostic tool that helps pinpoint performance bottlenecks through the use of abstract indices; (vi) Modeling Kernel, a facility for automated modeling of message-passing programs that supports both simulation -based and analytical approaches to performance prediction and scalability analysis; (vii) Intrusion Compensator, a utility for recovering true performance from observed performance by removing the overheads of monitoring and their effects on the communication pattern of the program; and (viii) Compatibility Tools, that convert AIMS-generated traces into formats used by other performance-visualization tools, such as ParaGraph, Pablo, and certain AVS/Explorer modules.

Yan, Jerry C.↗

Diagnostic Algorithm Benchmarking

A poster for the NASA Aviation Safety Program Annual Technical Meeting. It describes empirical benchmarking on diagnostic algorithms using data from the ADAPT Electrical Power System testbed and a diagnostic software framework.

Poll, Scott↗

SoMoGym: A Toolkit for Developing and Evaluating Controllers and Reinforcement Learning Algorithms for Soft Robots

Soft robotsoffer a host of benefits over traditional rigid robots, including inherent compliance that lets them passively adapt to variable environments and operate safely around humans and fragile objects. However, that same compliance makes it hard to use model-based methods in planning tasks requiring high precision or complex actuation sequences. Reinforcement learning (RL) can potentially find effective control policies, but training RL using physical soft robots is often infeasible, and training using simulations has had a high barrier to adoption. To accelerate research in control and RL for soft robotic systems, we introduce SoMoGym ( So ft Mo tion Gym ), a software toolkit that facilitates training and evaluating controllers for continuum robots. SoMoGym provides a set of benchmark tasks in which soft robots interact with various objects and environments. It allows evaluation of performance on these tasks for controllers of interest, and enables the use of RL to generate new controllers. Custom environments and robots can likewise be added easily. We provide and evaluate baseline RL policies for each of the benchmark tasks. These results show that SoMoGym enables the use of RL for continuum robots, a class of robots not covered by existing benchmarks, giving them the capability to autonomously solve tasks that were previously unattainable.

Moritz A. Graule↗

Metrics and Benchmarks for Visualization

What is a "good" visualization? How can the quality of a visualization be measured? How can one tell whether one visualization is "better" than another? I claim that the true quality of a visualization can only be measured in the context of a particular purpose. The same image generated from the same data may be excellent for one purpose and abysmal for another. A good measure of visualization quality will correspond to the performance of users in accomplishing the intended purpose, so the "gold standard" is user testing. As a user of visualization software (or at least a consultant to such users) I don't expect visualization software to have been tested in this way for every possible use. In fact, scientific visualization (as distinct from more "production oriented" uses of visualization) will continually encounter new data, new questions and new purposes; user testing can never keep up. User need software they can trust, and advice on appropriate visualizations of particular purposes. Considering the following four processes, and their impact on visualization trustworthiness, reveals important work needed to create worthwhile metrics and benchmarks for visualization. These four processes are (1) complete system testing (user-in-loop), (2) software testing, (3) software design and (4) information dissemination. Additional information is contained in the original extended abstract.

Uselton, Samuel P.↗

Software Verification and Validation Guidelines for Non-Linear Soil-Structure Interaction Analysis

Seismic analysis of structures, systems, and components (SSCs), including the consideration of soil-structure interaction (SSI) effects, is an important and required step in the design and licensing of nuclear power plant SSCs important to safety. Historically, the SSI analysis of nuclear structures has been performed using equivalent linear methods. However, there has been considerable industry investment in alternative seismic design and analysis approaches to reduce the construction cost of new reactors. Toward that goal and in alignment with the Licensing Modernization Project (LMP) framework, the Nuclear Regulatory Commission (US NRC) has proposed a risk-informed, performance-based approach to seismic design that allows inelastic response in those nuclear plant structures not required for confinement. As a complementary effort, reactor designers are exploring nonlinear seismic analysis methods to optimize structural designs and reduce construction costs.

21 SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLANTS↗

JARVIS-Leaderboard: a large scale benchmark of materials design methods

Abstract Lack of rigorous reproducibility and validation are significant hurdles for scientific development across many fields. Materials science, in particular, encompasses a variety of experimental and theoretical approaches that require careful benchmarking. Leaderboard efforts have been developed previously to mitigate these issues. However, a comprehensive comparison and benchmarking on an integrated platform with multiple data modalities with perfect and defect materials data is still lacking. This work introduces JARVIS-Leaderboard, an open-source and community-driven platform that facilitates benchmarking and enhances reproducibility. The platform allows users to set up benchmarks with custom tasks and enables contributions in the form of dataset, code, and meta-data submissions. We cover the following materials design categories: Artificial Intelligence (AI), Electronic Structure (ES), Force-fields (FF), Quantum Computation (QC), and Experiments (EXP). For AI, we cover several types of input data, including atomic structures, atomistic images, spectra, and text. For ES, we consider multiple ES approaches, software packages, pseudopotentials, materials, and properties, comparing results to experiment. For FF, we compare multiple approaches for material property predictions. For QC, we benchmark Hamiltonian simulations using various quantum algorithms and circuits. Finally, for experiments, we use the inter-laboratory approach to establish benchmarks. There are 1281 contributions to 274 benchmarks using 152 methods with more than 8 million data points, and the leaderboard is continuously expanding. The JARVIS-Leaderboard is available at the website: https://pages.nist.gov/jarvis_leaderboard/

36 MATERIALS SCIENCE↗

Tough Errors Are no Match (TEAM): Optimizing the quantum compiler for noise resilience

This report summarizes Unitary Fund’s contributions to the Department of Energy’s TEAM project (DE-SC0020266) under Thrust 2: Quantum Programming and Compilation. The central outcomes of this work have been the development of Mitiq, an open-source Python toolkit for applying quantum error mitigation (QEM) techniques to noisy quantum programs, and the invention, benchmarking and theoretical investigation of novel QEM techniques. Additional outcomes include the development of other open source software packages for the usage, simulation and control of quantum computers.

97 MATHEMATICS AND COMPUTING↗

DSMC analysis in a heterogeneous parallel computing environment

A methodology for implementing parallel DSMC codes in a heterogeneous computing environment is described. The methodology involves the use of a common message-passing software library together with recently developed software that handles the actual interprocessor communications in a standard manner across a variety of computing platforms. Benchmark tests using a simple DSMC model problem were performed on an Intel iPSC/860, a Cray-YMP and a group of Sun workstations. The approach was found to give speedups that scaled linearly with problem size on all the computing platforms tested. This methodology was then incorporated into a production-type DSMC code to allow the simulation of problems that would not otherwise have been practical. The application of this production code to simulations of hypersonic shear flows and shock-lip interactions under near-continuum conditions is described. Synchronous and asynchronous models for implementing parallelism into DSMC simulations are also described and both models are shown to produce the same steady-state result.

Wilmoth, R. G.↗

Branch recovery with compiler-assisted multiple instruction retry

In processing systems where rapid recovery from transient faults is important, schemes for multiple instruction rollback recovery may be appropriate. Multiple instruction retry has been implemented in hardware by researchers and also in mainframe computers. This paper extends compiler-assisted instruction retry to a broad class of code execution failures. Five benchmarks were used to measure the performance penalty of hazard resolution. Results indicate that the enhanced pure software approach can produce performance penalties consistent with existing hardware techniques. A combined compiler/hardware resolution strategy is also described and evaluated. Experimental results indicate a lower performance penalty than with either a totally hardware or totally software approach.

Alewine, N. J.↗