Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Performance Analysis”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Metallic Fuel Performance Analysis for the European Sodium Fast Reactor (ESFR-SIMPLE): Analysis of metallic fuel performance using SAS4A/SASSYS-1 $-$ MFUEL

The European Sodium Fast Reactor - Safety by Innovative Monitoring, Power Level flexibility and Experimental research (ESFR-SIMPLE) project was initiated in 2022 and includes assessment of a metallic-fueled version of the ESFR concept. Argonne National Laboratory (ANL) has been partnering with the ESFR-SIMPLE project to share its expertise on metallic fueled SFR designs and support some of its analysis. This report focuses on metallic fuel behavior analysis for ESFR-SIMPLE design conditions under base irradiation and transients (ULOF and UTOP).

21 SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLANTS↗

Traveler: Navigating Task Parallel Traces for Performance Analysis

Understanding the behavior of software in execution is a key step in identifying and fixing performance issues. This is especially important in high performance computing contexts where even minor performance tweaks can translate into large savings in terms of computational resource use. To aid performance analysis, developers may collect an execution trace —a chronological log of program activity during execution. As traces represent the full history, developers can discover a wide array of possibly previously unknown performance issues, making them an important artifact for exploratory performance analysis. However, interactive trace visualization is difficult due to issues of data size and complexity of meaning. Traces represent nanosecond-level events across many parallel processes, meaning the collected data is often large and difficult to explore. The rise of asynchronous task parallel programming paradigms complicates the relation between events and their probable cause. Here, to address these challenges, we conduct a continuing design study in collaboration with high performance computing researchers. We develop diverse and hierarchical ways to navigate and represent execution trace data in support of their trace analysis tasks. Through an iterative design process, we developed Traveler , an integrated visualization platform for task parallel traces. Traveler provides multiple linked interfaces to help navigate trace data from multiple contexts. We evaluate the utility of Traveler through feedback from users and a case study, finding that integrating multiple modes of navigation in our design supported performance analysis tasks and led to the discovery of previously unknown behavior in a distributed array library.

97 MATHEMATICS AND COMPUTING↗

Empirical Performance Analysis of HPC Applications with Portable Hardware Counter Metrics [Thesis]

In this dissertation, we demonstrate that it is possible to develop methods of empirical hardware-counter-based performance analysis for scientific applications running on diverse types of CPUs. Although hardware counters have been used in performance analysis for at least 30 years, the methods used are still limited to particular CPU vendors or even particular generations of CPUs from the same vendor. Our motivating hypothesis is that hardware counter-based measurements could be developed to provide consistent performance information on diverse CPU types. This dissertation proves the hypothesis was correct by demonstrating one such set of metrics.

97 MATHEMATICS AND COMPUTING↗

Performance analysis and data reduction for exascale scientific workflows

Chimbuko is the first in situ, scalable, workflow-level performance analysis tool for trace-level analysis and visualization of application performance. This tool was developed by the Co-design Center for Online Data Analysis and Reduction and funded by the U.S. Department of Energy’s Exascale Computing Project. We provide a detailed description of Chimbuko’s architecture and illustrate our online and offline visualization with multiple use cases. We also present results for the deployment and scalability of the tool as applied to a high-energy physics workflow running at large scale on the Frontier supercomputer.

97 MATHEMATICS AND COMPUTING↗

Transient fuel performance analysis for the preliminary fuel concept of general atomics fast modular reactor

This study investigates the transient fuel performance of General Atomics Fast Modular Reactor (GA-FMR) during accident scenarios, focusing on the behavior of its innovative fuel system that combines high-assay low enriched uranium dioxide (HALEUO2) fuel with SiGA® ceramic matrix composite silicon carbide cladding. The preliminary fuel design’s response was analyzed during reactivity-initiated accidents (RIA) and loss of coolant accidents (LOCA) using BISON fuel performance analysis code, which included both the diffusion enhanced and BISON-FASTGRASS coupled UO 2 models. The RIA analysis demonstrated that effective reactivity control reduced fuel temperature, though with transient fission gas release resulting in additional tensile stress state on the cladding. LOCA simulations revealed differing predictions between the two models: the BISON UO 2 model showed more transient fission gas release but minimal pellet expansion, while the BISON-FASTGRASS UO 2 model predicted less pronounced fission gas release but more fuel swelling and thermal expansion, potentially leading to pellet-cladding mechanical interaction. Here, these findings highlight critical areas for fuel design optimization and identify knowledge gaps requiring further experimental and computational investigation to advance GA-FMR fuel development.

Lee, Soon K. [Argonne National Laboratory (ANL), A↗

Unveiling Temporal Performance Deviation: Leveraging Clustering in Microservices Performance Analysis

As the market for cloud computing continues to grow, an increasing number of users are deploying applications as microservices. The shift introduces unique challenges in identifying and addressing performance issues, particularly within large and complex infrastructures. To address this challenge, we propose a methodology that unveils temporal performance deviations in microservices by clustering containers based on their performance characteristics at different time intervals. Showcasing our methodology on the Alibaba dataset, we found both stable and dynamic performance patterns, providing a valuable tool for enhancing overall performance and reliability in modern application landscapes.

Clustering↗

New machine protection system at the Spallation Neutron Source – design process and performance analysis

A New Machine Protection System (MPS) at the Spallation Neutron Source (SNS) was developed and implemented on µTCA-based hardware platforms. The system monitors more than 2500 field inputs and shuts off the beam within 10 µs if adverse events occur. We will present system level design process of various firmware and software components as well as the system integration into EPICS environment. The performance analysis of the MPS after two SNS run cycles will also be presented.

Bobrek, Miljko [ORNL] (ORCID:0000000332763451)↗

Development of Accelerated Steady-state Test Capsule Experiments to Replicate EBR-II Fuel Behavior Using BISON Fuel Performance Analysis

Here in this work, BISON fuel performance calculations were performed to predict the fuel behavior of accelerated burnup U-Pu-Zr fuel, with temperature operation conditions of the fuel and the cladding mirroring conditions within EBR-II fuel pins. The temperature operating conditions within the FAST accelerated burnup rods were aimed at replicating EBR-II X447/X447A fuel surface and inner cladding surface temperatures. Due to the FAST capsule design, these temperatures can be replicated with fission rate densities being significantly increased. fuel performance modeling has not been assessed for novel experiments such as accelerated burnup utilizing the FAST capsule within ATR. This is an important step in understanding accelerated irradiation methods as many performance models are empirical models conforming to the results of PIE but do not always include physical models that would represent the changes in irradiation tests.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Refining HPCToolkit for application performance analysis at exascale

As part of the US Department of Energy’s Exascale Computing Project (ECP), Rice University has been refining its HPCToolkit performance tools to better support measurement and analysis of applications executing on exascale supercomputers. To efficiently collect performance measurements of GPU-accelerated applications, HPCToolkit employs novel non-blocking data structures to communicate performance measurements between tool threads and application threads. To attribute performance information in detail to source lines, loop nests, and inlined call chains, HPCToolkit performs parallel analysis of large CPU and GPU binaries involved in the execution of an exascale application to rapidly recover mappings between machine instructions and source code. To analyze terabytes of performance measurements gathered during executions at exascale, HPCToolkit employs distributed-memory parallelism, multithreading, sparse data structures, and out-of-core streaming analysis algorithms. To support interactive exploration of profiles up to terabytes in size, HPCToolkit’s hpcviewer graphical user interface uses out-of-core methods to visualize performance data. The result of these efforts is that HPCToolkit now supports collection, analysis, and presentation of profiles and traces of GPU-accelerated applications at exascale. These improvements have enabled HPCToolkit to efficiently measure, analyze and explore terabytes of performance data for executions using as many as 64K MPI ranks and 64K GPU tiles on ORNL’s Frontier supercomputer. HPCToolkit’s support for measurement and analysis of GPU-accelerated applications has been employed to study a collection of open-science applications developed as part of ECP. This paper reports on these experiences, which provided insight into opportunities for tuning applications, strengths and weaknesses of HPCToolkit itself, as well as unexpected behaviors in executions at exascale.

Adhianto, Laksono↗

Performance Analysis of an Optimization Algorithm for Metamaterial Design on the Integrated High-Performance Computing and Quantum Systems

Optimizing metamaterials with complex geometries is a big challenge. Although an active learning algorithm, combining machine learning (ML), quantum computing, and optical simulation, has emerged as an efficient optimization tool, it still faces difficulties in optimizing complex structures that have potentially high performance. In this work, we comprehensively analyze the performance of an optimization algorithm for metamaterial design on the integrated HPC and quantum systems. We demonstrate significant time advantages through message-passing interface (MPI) parallelization on the high-performance computing (HPC) system showing approximately 54% faster ML tasks and 67 times faster optical simulation against serial workloads. Furthermore, we analyze the performance of a quantum algorithm designed for optimization, which runs with various quantum simulators on a local computer or HPC-quantum system. Results showcase ~24 times speedup when executing the optimization algorithm on the HPC-quantum hybrid system. This study paves a way to optimize complex metamaterials using the integrated HPC-quantum system.

Kim, Seongmin↗

Performance Analysis and Optimization for Scientific Data Workloads

Scientific data generated at experimental and observational facilities are increasingly being processed on large-scale compute systems. Most of the experimental data analysis workflows are not designed or implemented to run on large scale environments and take full advantage of HPC compute and storage resources. These applications are unlike the traditional tightly-coupled scientific applications and hence face significant performance and scalability challenges as the volume of data increases exponentially. In this paper, we conduct a performance and scalability analysis for experimental analysis applications and workflows operating on data from light sources. Our analysis detects and quantifies I/O performance, scalability and runtime bottlenecks for three data analysis applications that run on NERSC resources. Based on our analysis we propose and implement a set of optimizations that lead to reducing the amount of time spent on I/O operations by almost 90%.

97 MATHEMATICS AND COMPUTING↗

MDPCT1 Quench Data and Performance Analysis

MDPCT1 is a four-layer cos-theta Nb3Sn dipole demonstrator developed and tested at FNAL in the framework of the U.S. Magnet Development Program. The magnet reached record fields for accelerator magnets of 14.1 T at 4.5 K in the first test and 14.5 T at 1.9 K in the second test and then showed large degradation. While its inner coils performed exceptionally well with only two quenches up to 14.5 T and no evidence of degradation, the outer coils degraded over the course of testing. By adopting new measurement and analysis techniques at FNAL we are discussing in detail what happened. Both success and failure in our diagnostics are discussed. The evolution of techniques over the course of two tests (and three thermal cycles) shows the path to address challenges brought by the first four-layer magnet tested at FNAL. This paper presents the analysis of quench data along with diagnostic features and complementary measurements taken in support of the magnet performance analysis.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

Performance Analysis of the Expected Track Length Estimator

This work investigates whether performance of the MCNP code is improved by the implementation of the expected track length estimator. The expected track length estimator is known to decrease variance when compared to the track length estimator in many classes of Monte Carlo particle transport simulations. However, the computational cost of computing an exponent for every particle collision within a tallied cell has previously been considered prohibitively expensive. Therefore, the goal of this work is to determine whether the variance improvements seen by the expected track length estimator in most cases makes up for the increase in transport time when implemented in the MCNP code.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗