Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Benchmarking Software”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

Agilent CRADA (Abstract)

The CRADA between Agilent Technologies Inc. and Battelle will focus on five software components as listed below: Prototype 4D Feature Finding functionality with a particular focus on recovering low level features and extending the bottom end dynamic range of IM-MS technology. Compare and contrast developments to current 4D Feature Finding capabilities. Highlight important algorithmic aspects employed. Implement the PNNL saturation correction algorithm. Agilent will give PNNL the needed data file access API and assistance in understanding it implementation and any needed instrumental aspects. Supported high resolution products to include Agilent’s TOF, QTOF and IM-QTOF mass spectrometers. PNNL will then work with Agilent to benchmark performance. Implementation of the PNNL Hadamard de-multiplexing algorithm. Agilent will give provide PNNL the needed date file access API access and as needed assistance in understanding the current Agilent multiplexed IM offering. PNNL will then work with Agilent on benchmark performance. Add ion mobility collision cross sections to existing and new metabolomic libraries for data analysis with Agilent’s informatics program MPP/ID Browser. PNNL will work with Agilent to create a software pipeline that takes data from chemical and metabolic standards and properly formats it for inclusion in MPP accessible libraries, using the collision cross section as a new separation dimension. Improvements of MPP multidimensional matching to identify metabolomic features using multiple characteristics beyond retention time and accurate mass. Most significantly matching will include analyte collision cross section with proposed support for sample fraction or RapidFire cartridge and fragmentation spectra. PNNL will work with Agilent to modify and improve the current MPP analysis pipeline to allow for creating, aligning, and identifying MS features defined by accurate mass, collision cross section and chromatographic retention time. As additional criteria such as fraction or RapidFire cartridge type are supported in the identification process, then they also will become part of the automation workflow. This includes the automation of said system to work with command line program (i.e. not a GUI) sufficient for programmatic execution in a pipeline.

97 MATHEMATICS AND COMPUTING↗

Benchmark for two-dimensional large scale coherent structures in partially magnetized E × B plasmas—community collaboration & lessons learned

Low-temperature plasmas (LTPs) are essential to both fundamental scientific research and critical industrial applications. As in many areas of science, numerical simulations have become a vital tool for uncovering new physical phenomena and guiding technological development. Code benchmarking remains crucial for verifying implementations and evaluating performance. This work continues the Landmark benchmark initiative, a series specifically designed to support the verification of LTP codes. In this study, seventeen simulation codes from a collaborative community of nineteen international institutions modeled a partially magnetized E × B Penning discharge. The emergence of large scale coherent structures, or rotating plasma spokes, endows this configuration with an enormous range of time scales, making it particularly challenging to simulate. The codes showed excellent agreement on the rotation frequency of the spoke as well as key plasma properties, including time-averaged ion density, plasma potential, and electron temperature profiles. Achieving this level of agreement came with challenges, and we share lessons learned on how to conduct future benchmarking campaigns. Comparing code implementations, computational hardware, and simulation runtimes also revealed interesting trends, which are summarized with the aim of guiding future plasma simulation software development.

benchmarking↗

Noise-aware circuit compilations for a continuously parameterized two-qubit gateset

State-of-the-art noisy-intermediate-scale quantum processors are currently implemented across a variety of hardware platforms, each with their own distinct gatesets. As such, circuit compilation should not only be aware of but also deeply connect to the native gateset and noise properties of each. Trapped-ion processors are one such platform that provides a gateset that can be continuously parameterized across both one- and two-qubit gates. Here we use the Quantum Scientific Computing Open User Testbed to study noise-aware compilations focused on continuously parameterized two-qubit 𝑍⁢𝑍 gates (based on the Mølmer-Sørensen interaction) using $\scriptsize{SUPERSTAQ}$, a quantum software platform for hardware-aware circuit compiler optimizations. We discuss the realization of 𝑍⁢𝑍 gates with arbitrary angle on the all-to-all connected trapped-ion system. Then we discuss a variety of different compiler optimizations that innately target these 𝑍⁢𝑍 gates and their noise properties. These optimizations include moving from a restricted maximally entangling gateset to a continuously parameterized one, swap mirroring to further reduce the total entangling angle of the operations, focusing the heaviest 𝑍⁢𝑍 angle participation on the best-performing gate pairs, and circuit approximation to remove the least impactful 𝑍⁢𝑍 gates. We demonstrate these compilation approaches on the hardware with randomized quantum volume circuits, observing the potential to realize a larger quantum volume as a result of these optimizations. Using differing yet complementary analysis techniques, we observe the distinct improvements in system performance provided by these noise-aware compilations and study the role of stochastic and coherent error channels for each compilation choice.

Noise↗

Foundation model framework for all tasks involving jet physics

Foundation models use large datasets to build an effective representation of data that can be deployed on diverse downstream tasks. Previous research developed the omnilearn foundation model for jet physics, using unique properties of particle physics, and showed that it could significantly advance discovery potential across collider experiments. This paper introduces a major upgrade, resulting in the omnilearned framework. This framework has three new elements: (1) updates to the model architecture and training, (2) using over 1 × 10 9 jets used for training, and (3) providing well-documented software for accessing all datasets and models. We demonstrate omnilearned with three representative tasks: top-quark jet tagging with the community delphes-based benchmark dataset, b tagging with ATLAS full simulation, and anomaly detection with CMS experimental data. In each case, omnilearned is the state of the art, further expanding the discovery potential of past, current, and future collider experiments.

Bhimji, Wahid [Lawrence Berkeley National Laborato↗

Unified differentiable digital twin for the IOTA/FAST facility

As the design complexity of modern accelerators grows, there is more interest in using advanced simulations that have fast execution time or produce insights about accelerator state. One notable example of additional information are gradients of physical observables with respect to design parameters produced by differentiable simulations. The IOTA/FAST facility has recently begun a program to implement and experimentally validate a unified start-to-end differentiable digital twin to serve as a virtual accelerator test stand, allowing for rapid prototyping of new software and experiments with minimal beam time costs. In this contribution we will discuss our plans and progress. Specifically, we will cover the selection and benchmarking of both physics and ML codes, the development of generic interfaces between device models and surrogate or physics-based sections, and the export of the parameters through either a deterministic event loop or a fully asynchronous EPICS soft input/output controller. We will also discuss challenges in model calibration and uncertainty quantification, as well as future plans to support larger proton accelerators like PIPII and Booster.

Kuklev, Nikita [Fermilab]↗

End-to-end differentiable digital twin for the IOTA/FAST facility

As the design complexity of modern accelerators grows, there is more interest in using controllable-fidelity simulations that have fast execution time or can yield additional insights about accelerator state. One notable example of additional information are gradients of physical observables with respect to design parameters produced by differentiable simulations. The IOTA/FAST facility has recently begun a program to implement and experimentally validate an end-to-end digital twin to serve as a virtual accelerator test stand, allowing for rapid prototyping of new software and experiments with minimal beam time costs. In this contribution we will discuss our plans and progress. Specifically, we will cover the selection and benchmarking of both physics and ML codes for linac and ring simulation, the development of generic interfaces between surrogate and physics-based sections, and presenting the control interface as either a deterministic event loop or a fully asynchronous EPICS soft input/output controller. We will also discuss challenges in model calibration and uncertainty quantification, as well as future plans to implement larger proton accelerators like PIPII and Booster.

Kuklev, N. [Fermilab]↗

Asymptotic hydrogen redistribution analysis in yttrium-hydride-moderated heat-pipe-cooled microreactors using DireWolf

Yttrium hydride (YH x ) is one of the materials being considered for moderating thermal and epithermal nuclear microreactors. One potential issue with YH x use is that the hydrogen redistributes in the hydride when thermal and concentration gradients are present. This hydrogen redistribution leads to spatial gradients in the hydrogen concentration, thus affecting neutron transport in the reactor. Here, by building upon observations in prior works, this paper aims to gain a better understanding of the reactivity feedback associated with such hydrogen redistributions. In particular, we wish to understand the sign (+/–) of the hydrogen redistribution neutronic feedback, its order of magnitude, and its underlying physical causes. To achieve this goal, the DireWolf multiphysics software driver was used to solve the coupled radiation transport, heat transfer, heat pipe two-phase flow, and hydrogen redistribution equations for the Simplified Microreactor Benchmark Assessment (SiMBA) problem, a full-core microreactor numerical benchmark developed at Idaho National Laboratory.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Machine learning framework for predicting uranium enrichments from M400 CZT gamma spectra

A machine learning framework was developed for predicting uranium enrichments from M400 CZT gamma spectra. This framework leverages the availability of a large amount of measured M400 gamma spectra and uses a recently updated version of Gamma Detector Response and Analysis Software (GADRAS) for gamma spectrum analysis and generation. It also leverages the existing machine learning modules in Python for gamma spectrum data processing, curation, model training, benchmarking, and optimization of the deep machine learning models. The framework is used to develop a deep learning model to analyze gamma spectra from a set of U 3 O 8 samples with enrichments ranging from 0.31 to 93.17% and UF 6 cylinders with enrichments ranging from 0.2 to 4.95%, and the model performance is tested using a set of measured spectra and the respective declared enrichment values. Results show that the model can correctly classify 99.35% of the U 3 O 8 sample enrichments, and can predict the samples’ enrichments within an average absolute error of 0.099% (in percentage points of enrichment). For the UF 6 cylinders, the average absolute error was approximately 0.03%, with an accuracy of 98% in classifying discrete enrichment values of UF 6 samples. Finally, the results also show that the model has performed significantly better in terms of predicting enrichments in UF 6 cylinders based on measured gamma spectra than the GEM code, with a standard deviation (of the relative errors) of 2.23% (compared with the 11.51% value for the GEM code) based on results from a set of test data.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

The ePIC Simulation Campaign Workflow on the Open Science Grid

The ePIC collaboration is realizing the first experiment of the future Electron-Ion Collider (EIC) at the Brookhaven National Laboratory that will allow for a precision study of the nucleons and the nucleus at the scale of sea quarks and gluons through the study of electron-proton/ion collisions. This paper will discuss the current workflow for running centralized simulation campaigns for ePIC on the Open Science Grid (OSG) infrastructure. This involves monthly releases of ePIC software and container deployments to CVMFS, generation of input datasets in HepMC format according to collaboration-defined policy, using Snakemake in CI/CD for validation and benchmarking, and submitting jobs to the OSG condor scheduler for opportunistic running on available resources. File transfers utilize XrootD, and Rucio is used for data management. The workflow is continuously refined to improve daily throughput (currently 50-100k core hours per day) and minimize job failures. Since May 2023, monthly simulation campaigns employing the workflow have cumulatively used over 20 million core hours on the OSG and produced over 350 TB of simulation data. The campaigns incorporate simulations for the broad science program of the EIC and are actively used for the detector and physics studies in preparation of the Technical Design Report (TDR).

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Direction-optimizing Label Propagation Framework for Structure Detection in Graphs: Design, Implementation, and Experimental Analysis

Label Propagation is not only a well-known machine learning algorithm for classification but also an effective method for discovering communities and connected components in networks. We propose a new Direction-optimizing Label Propagation Algorithm (DOLPA) framework that enhances the performance of the standard Label Propagation Algorithm (LPA), increases its scalability, and extends its versatility and application scope. As a central feature, the DOLPA framework relies on the use of frontiers and alternates between label push and label pull operations to attain high performance. It is formulated in such a way that the same basic algorithm can be used for finding communities or connected components in graphs by only changing the objective function used. Additionally, DOLPA has parameters for tuning the processing order of vertices in a graph to reduce the number of edges visited and improve the quality of solution obtained. We present the design and implementation of the enhanced algorithm as well as our shared-memory parallelization of it using OpenMP. We also present an extensive experimental evaluation of our implementations using the LFR benchmark and real-world networks drawn from various domains. Compared with an implementation of LPA for community detection available in a widely used network analysis software, we achieve at most five times the F-Score while maintaining similar runtime for graphs with overlapping communities. We also compare DOLPA against an implementation of the Louvain method for community detection using the same LFR-graphs and show that DOLPA achieves about three times the F-Score at just 10% of the runtime. For connected component decomposition, our algorithm achieves orders of magnitude speedups over the basic LP-based algorithm on large-diameter graphs, up to 13.2× speedup over the Shiloach-Vishkin algorithm, and up to 1.6× speedup over Afforest on an Intel Xeon processor using 40 threads.

97 MATHEMATICS AND COMPUTING↗

Implementation of real‐time TDDFT for periodic systems in the open‐source PySCF software package

Abstract We present a new implementation of real‐time time‐dependent density functional theory (RT‐TDDFT) for calculating excited‐state dynamics of periodic systems in the open‐source Python‐based PySCF software package. Our implementation uses Gaussian basis functions in a velocity gauge formalism and can be applied to periodic surfaces, condensed‐phase, and molecular systems. As representative benchmark applications, we present optical absorption calculations of various molecular and bulk systems and a real‐time simulation of field‐induced dynamics of a (ZnO) 4 molecular cluster on a periodic graphene sheet. We present representative calculations on optical response of solids to infinitesimal external fields as well as real‐time charge‐transfer dynamics induced by strong pulsed laser fields. Due to the widespread use of the Python language, our RT‐TDDFT implementation can be easily modified and provides a new capability in the PySCF code for real‐time excited‐state calculations of chemical and material systems.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

CommBench: Micro-Benchmarking Hierarchical Networks with Multi-GPU, Multi-NIC Nodes

Modern high-performance computing systems have multiple GPUs and network interface cards (NICs) per node. The resulting network architectures have multilevel hierarchies of subnetworks with different interconnect and software technologies. These systems offer multiple vendor-provided communication capabilities and library implementations (IPC, MPI, NCCL, RCCL, OneCCL) with APIs providing varying levels of performance across the different levels. Understanding this performance is currently difficult because of the wide range of architectures and programming models (CUDA, HIP, OneAPI). We present CommBench, a library with cross-system portability and a high-level API that enables developers to easily build microbenchmarks relevant to their use cases and gain insight into the performance (bandwidth & latency) of multiple implementation libraries on different networks. We demonstrate CommBench with three sets of microbenchmarks that profile the performance of six systems. Our experimental results reveal the effect of multiple NICs on optimizing the bandwidth across nodes and also present the performance characteristics of four available communication libraries within and across nodes of NVIDIA, AMD, and Intel GPU networks.

Hidayetoglu, Mert↗

Benchmark Modeling and Simulation of the FFTF LOFWOS Test #13 Using SAM

The Fast Flux Test Facility (FFTF) was a 400 MW thermal powered, oxide-fueled, liquid sodium cooled test reactor, built to assist development and testing of advanced fuels and materials for fast breeder reactors. In July 1986, a series of unprotected Loss of Flow Without Scram (LOFWOS) transients were performed in FFTF as part of the Passive Safety Testing (PST) program. The LOFWOS Test #13, which was initiated at 50% power and 100% flow with the pump pony motors left off, has been chosen as a benchmark case by IAEA to support collaborative efforts within international partnerships on the validation of simulation tools and models in the area of sodium fast reactor passive safety in an IAEA Coordinated Research Project (CRP), launched in October 2018. The System Analysis Module (SAM) is an advanced and modern system analysis tool under development at Argonne National Laboratory for advanced non-LWR safety analysis. It utilizes the object-oriented application framework MOOSE to leverage the modern software environment and advanced numerical methods. The capabilities of SAM are being extended to enable the transient modeling, analysis, and design of various advanced nuclear reactor systems. To participate the IAEA CRP and enhance the SAM validation base for advanced reactor transient safety analysis, benchmark simulations of the FFTF LOFWOS Test #13 are performed using the SAM code. In this first phase of the validation effort, the thermal-hydraulic behavior of the reactor system is the focus and the reactor kinetics is not considered in the SAM FFTF model. Instead, the results of Argonne’s neutronics calculations are directly used, including the power shape of the active core region and the power history during the transient. The simulation results of FFTF at steady state agreed well with the measured data from the test. During the transient, reasonably good agreement were also obtained. Future work to improve the model will focus on introducing the reactivity predictions into the model, as well as better understanding or resolving the current discrepancies with the measured data.

21 SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLANTS↗

Simulations and analysis tools for charge-exchange (d, 2 He) reactions in inverse kinematics with the AT-TPC

Charge-exchange (d, 2 He) reactions in inverse kinematics at intermediate energies are a very promising method to investigate the Gamow–Teller transition strength in unstable nuclei. A simulation and analysis software based on the attpcroot package was developed to study this type of reactions with the active-target time projection chamber (AT-TPC). The simulation routines provide a realistic detector response that can be used to understand and benchmark experimental data. Analysis tools and correction routines can be developed and tested from simulations in ATTPCROOT , because they are processed in the same way as the real data. In particular, we study the feasibility of using coincidences with beam-like particles to unambiguously identify the (d, 2 He) reaction channel, and to develop a kinematic fitting routine for future applications. More technically, the impact of space-charge effects in the track reconstruction, and a possible correction method are investigated in detail. Finally, this analysis and simulation package constitutes an essential part of the software development for the fast-beams program with the AT-TPC.

(d,2He)↗

IER 500: AWE-LLNL Measurement Campaign at DAF [Slides]

Presentation Hosted at Device Assembly Facility (DAF). Joint collaboration by Atomic Weapons Establishment (AWE) and Lawrence Livermore National Laboratory (LLNL). Multiple measurements were recorded through a two-week period with 12 unique objects measured. LLNL deployed Machine learning software program for diagnostic assessments. Current status notes the dedicated DAF team LLNL maintains. Additionally, LLNL is compiling a report on the AWE-LLNL measurements and designing security benchmark experiments. The presentation concludes with future work envisioned.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Systematic Crosstalk Mitigation for Superconducting Qubits via Frequency-Aware Compilation

One of the key challenges in current Noisy Intermediate-Scale Quantum (NISQ) computers is to control a quantum system with high-fidelity quantum gates. There are many reasons a quantum gate can go wrong - for superconducting transmon qubits in particular, one major source of gate error is the unwanted crosstalk between neighboring qubits due to a phenomenon called frequency crowding. We motivate a systematic approach for understanding and mitigating the crosstalk noise when executing near-term quantum programs on superconducting NISQ computers. Here, we present a general software solution to alleviate frequency crowding by systematically tuning qubit frequencies according to input programs, trading parallelism for higher gate fidelity when necessary. The net result is that our work dramatically improves the crosstalk resilience of tunable-qubit, fixed-coupler hardware, matching or surpassing other more complex architectural designs such as tunable-coupler systems. On NISQ benchmarks, we improve worst-case program success rate by 13.3x on average, compared to existing traditional serialization strategies.

Computer architecture↗

An Integrated Framework for Memory-Centric Analysis: From Trace Collection to Co-Design

The memory wall phenomenon—where advances in processor performance significantly outpace those in memory subsystems—poses a fundamental challenge for contemporary computing systems. In memory-bound applications, memory subsystem behavior dominates performance, yet existing analysis approaches present significant limitations: detailed microarchitectural simulators require days to weeks to simulate modest workloads; hardware performance counters provide only aggregate statistics that obscure temporal and spatial access patterns; and scaled simulation approaches face challenges in capturing certain behaviors that emerge at larger scales. These limitations reflect a processor-centric design philosophy increasingly misaligned with memory-bound workloads where detailed understanding of memory access patterns, cache hierarchy interactions, and contention is critical for effective optimization. This paper presents an integrated framework for memory-centric analysis that enables effective hardware-software co-design. We describe practical trace collection techniques, including hardware-assisted processor tracing with minimal overhead and portable software-based instrumentation with statistical sampling. We present multi-perspective analysis methods that examine memory behavior from temporal, sequential, spatial, and relational viewpoints, revealing distinct optimization opportunities invisible in aggregate metrics. We detail an architectural modeling framework that uses sampled traces with temporal interpolation and confidence-based filtering to evaluate cache and memory configurations. Evaluation on representative benchmarks demonstrates that this framework achieves practical accuracy (L2 cache errors of 2.64\%, confidence-filtered L3 errors of 9.92\%, bandwidth errors of 7.33\%) while providing substantial speedup (26.8×) over cycle-accurate simulation, enabling rapid design space exploration. We demonstrate how this integrated framework enables systematic identification of both hardware optimizations (memory controller tuning, bank partitioning, NUMA configuration) and software optimizations (data layout restructuring, prefetching strategies, memory-aware scheduling). Through this comprehensive treatment of the memory-centric analysis pipeline—from trace collection through architectural modeling to co-design application—we provide researchers and practitioners with practical techniques for addressing memory bottlenecks in contemporary computing systems.

Gajaria, Dhruv Mayur↗

Implementation of a Binary Neural Network on a Passive Array of Magnetic Tunnel Junctions

The increasing scale of neural networks and their growing application space have produced demand for more energy- and memory-efficient artificial-intelligence-specific hardware. Avenues to mitigate the main issue, the von Neumann bottleneck, include in-memory and near-memory architectures, as well as algorithmic approaches. In this report we leverage the low-power and the inherently binary operation of magnetic tunnel junctions (MTJs) to demonstrate neural network hardware inference based on passive arrays of MTJs. In general, transferring a trained network model to hardware for inference is confronted by degradation in performance due to device-to-device variations, write errors, parasitic resistance, and nonidealities in the substrate. To quantify the effect of these hardware realities, we benchmark 300 unique weight matrix solutions of a two-layer perceptron to classify the Wine dataset for both classification accuracy and write fidelity. Despite device imperfections, we achieve software-equivalent accuracy of up to 95.3% with proper tuning of network parameters in 15 x 15 MTJ arrays having a range of device sizes. The success of this tuning process shows that new metrics are needed to characterize the performance and quality of networks reproduced in mixed signal hardware.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗