Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Benchmarking Software”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11

Accelerating Radiation Computations for Dynamical Models With Targeted Machine Learning and Code Optimization

Abstract Atmospheric radiation is the main driver of weather and climate, yet due to a complicated absorption spectrum, the precise treatment of radiative transfer in numerical weather and climate models is computationally unfeasible. Radiation parameterizations need to maximize computational efficiency as well as accuracy, and for predicting the future climate many greenhouse gases need to be included. In this work, neural networks (NNs) were developed to replace the gas optics computations in a modern radiation scheme (RTE+RRTMGP) by using carefully constructed models and training data. The NNs, implemented in Fortran and utilizing BLAS for batched inference, are faster by a factor of 1–6, depending on the software and hardware platforms. We combined the accelerated gas optics with a refactored radiative transfer solver, resulting in clear‐sky longwave (shortwave) fluxes being 3.5 (1.8) faster to compute on an Intel platform. The accuracy, evaluated with benchmark line‐by‐line computations across a large range of atmospheric conditions, is very similar to the original scheme with errors in heating rates and top‐of‐atmosphere radiative forcings typically below 0.1 K day −1 and 0.5 W m −2 , respectively. These results show that targeted machine learning, code restructuring techniques, and the use of numerical libraries can yield material gains in efficiency while retaining accuracy.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Data from: "Towards CONUS-Wide ML-Augmented Conceptually-Interpretable Modeling of Catchment-Scale Precipitation-Storage-Runoff Dynamics"

This data package was generated to support the manuscript “Towards CONUS-Wide Machine Learning-Augmented Conceptually Interpretable Modeling of Catchment-Scale Precipitation-Storage-Runoff Dynamics.” It provides input files, model outputs, plotting data, scripts, notebooks, and documentation used to develop, evaluate, and reproduce Mass-Conserving Perceptron (MCP)-based hydrologic modeling experiments across 513 selected Catchment Attributes and Meteorology for Large-sample Studies in the United States (CAMELS-US) basins. The files are organized by modeling component and analysis purpose, including rainfall–runoff experiments, snow module experiments, coupled hydrologic-snow experiments, Long Short-Term Memory (LSTM) benchmark results, model skill metrics, initialization and epoch records, cell-state normalization files, Akaike Information Criterion (AIC)-based model comparison files, and data used to generate manuscript figures. Tabular files can be opened using standard spreadsheet software or Python/R data-analysis tools. Python scripts, Jupyter notebooks, and selected MATLAB scripts are included for model execution, postprocessing, plotting, and statistical analysis. Quality assurance and quality control were conducted through the source-data selection and modeling workflow. Meteorological forcing, streamflow, and static catchment attributes were derived from the CAMELS-US dataset, and snow water equivalent data were derived from the University of Arizona (UA) Snow Water Equivalent dataset. Selected basins and time periods were screened during the associated research workflow to avoid missing observations or poor-quality cases. Static geospatial features were processed primarily using Quantum Geographic Information System (QGIS) and Geospatial Data Abstraction Library (GDAL) workflows. Additional details are provided in the associated manuscript and documentation.

ESS-DIVE CSV File Formatting Guidelines Reporting ↗

Using Likwid and Byfl to Benchmark Hardware Performance [Poster]

Benchmark Study conducted focusing on CPU and program performance analysis. Performance data gathered using 2 different programs and comparisons made based on performance. After comparisons are made, conclusions can be drawn and improvements are made upon hardware and software.

97 MATHEMATICS AND COMPUTING↗

Quantum Electrodynamics Coupled-Cluster at Scale: High-Performance Implementation for Complex Systems

Coupled-cluster theory (CC) is a highly accurate and versatile method for simulating complex interactions within quantum systems. The extension of CC theory to model mixed electron-photon processes with quantum electrodynamics (QED) has improved our capability to predict cavity-modified chemistry, a field where photons are used as cost-effective and eco-friendly alternatives to catalyze/inhibit chemical reactions. However, calculations with CC methods, even without incorporating QED effects, are often prohibitively expensive. Simulations of larger systems require scalable infrastructures that exist for traditional CC methods but not for QED-CC methods. As such, we present a GPU-enabled, high-performance, open-source implementation of the quantum electrodynamics coupled-cluster method with single and double excitations (QED-CCSD) within the ExaChem quantum chemistry software package. ExaChem relies on the Tensor Algebra for Many-body Methods (TAMM) infrastructure: a parallel heterogeneous tensor library designed to achieve scalable performance on modern heterogeneous supercomputing platforms. Furthermore, we discuss theoretical foundations, algorithmic details, and numerical benchmarks to showcase the larger systems that ExaChem can simulate and how the integration of photonic degrees-of-freedom alters their ground-state properties.

Basis sets↗

Did You Win the GPU Cloud Lottery? Benchmarking from TFLOPS to Tokens/$

Cloud GPUs are commonly assumed to deliver consistent performance for a given GPU model. This assumption does not always hold: cloud providers employ diverse system configurations and virtualization mechanisms, and GPUs themselves exhibit non-negligible manufacturing variability (the silicon lottery). In this work, we present a large-scale measurement study of GPU performance variability across 11 cloud providers, covering over 3,500 physical GPUs and 6,800 benchmark runs. Our hierarchical analysis shows that while execution-level variation stays below 9%, performance varies by up to 38% across devices and providers for the same GPU model. Regression analysis indicates that driver- and OS-related software factors contribute less than 1% of the variance; instead, silicon lottery effects dominate observed performance variation, and cloud providers further amplify them through persistent, systematic second-order effects.

Slynko, Platon [Silicon Data, New York, USA] (ORCI↗

Particle Swarm Optimization Algorithm for Critical Experiment Design

Nuclear criticality experiments are used to validate nuclear cross section data used by simulation software. This is typically achieved by designing a critical system with a high sensitivity to a certain material’s cross section. Once the experiment has been carried out, a high fidelity model of the system is developed into a benchmark. When this benchmark model is simulated by a transport code, some of the difference between the experimental and computational effective neutron multiplication factor can be attributed to inaccurate nuclear data. Nuclear data evaluators then can make adjustments accordingly to improve cross section data.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

NCCS High Performance GMRES Mixed Precision

HPG-MxP is a software package that performs a fixed number of multigrid preconditioned (using a Gauss-Seidel smoother) Generalized minimal residual (PGMRES) iterations in order to solve a possibly nonsymmetric large sparse linear system of equations. It is designed to be a benchmark to measure a computer's performance for sparse linear algebra workloads typical in scientific computing while allowing the use of mixed precision methods. The solution is required to have convergence characteristics and accuracy similar to double precision GMRES. It is based on the High Performance Conjugate Gradient Benchmark (HPCG) which restricts all implementations to use only the IEEE double precision format (FP64). The original implementation (https://github.com/hpg-mxp/hpg-mxp) was written by Ichitaro Yamazaki, Jennifer Loe, Christian Glusa, Sivasankaran Rajamanickam, Piotr Luszczek, and Jack Dongarra. Please refer to that repository for documentation on the original implementation. This version is maintained by the National Center for Computational Sciences at Oak Ridge National Laboratory. It is highly scalable and optimized for Oak Ridge Leadership Computing Facility (OLCF) systems, particularly Frontier.

Kashi, Aditya [Oak Ridge National Laboratory (ORNL↗

Stereo-DIC Challenge 1.0 – Rigid Body Motion of a Complex Shape

Background Stereo-DIC is a widely used optical measurement technique that provides a dense full-field 3D measurement of the shape, displacement, and strain of a solid sample. When compared with 2D-DIC, Stereo-DIC provides greater flexibility and expands its use beyond flat, planar specimens. Furthermore, the widespread availability of commercial systems has led to the adoption of the technique throughout industry, academia, and government research labs. Objective Even though some research has been done to understand the effects of different experimental and stereo-DIC parameters, no reference is available to benchmark and compare the performance of current stereo-DIC algorithms to each other. Methods This paper provides the description and analysis of a carefully controlled 3D experiment and associated images used to compare the results from five subset based DIC software packages. Both the images and analysis codes used in this paper to compare the results are described here and are available for download and use for continued research. Results We show that over a very large range of motion, the 3D errors are very small, less than 80μm over a travel of ±20 mm out-of-plane and ±20 mm in-plane. While all codes performed similarly, there are important differences noted in the paper. Conclusion The image sets and results comparison software are hosted by the International DIC Society (www.iDICs.org) and are freely available for download and analysis for comparison with results in this paper. Furthermore, it is hoped that this set of images can be used for future research in improving stereo-DIC by future authors.

Algorithms comparison↗

AmeriFlux FLUXNET-1F US-ADR Amargosa Desert Research Site (ADRS)

This is the AmeriFlux Management Project (AMP) created FLUXNET-1F version of the carbon flux data for the site US-ADR Amargosa Desert Research Site (ADRS). This is the FLUXNET version of the carbon flux data for the site US-ADR Amargosa Desert Research Site (ADRS) produced by applying the standard ONEFlux (1F) software. Site Description - This tower is located at the Amargosa Desert Research Site (ADRS). The U.S. Geological Survey (USGS) began studies of unsaturated zone hydrology at ADRS in 1976. Over the years, USGS investigations at ADRS have provided long-term "benchmark" information about the hydraulic characteristics and soil-water movement for both natural-site conditions and simulated waste-site conditions in an arid environment. The ADRS is located in a creosote-bush community adjacent to disposal trenches for low-level radioactive waste.

Moreo, Michael [U.S. Geological Survey]↗

Off-design performance of molten salt-driven Rankine cycles and its impact on the optimal dispatch of concentrating solar power systems

This paper presents a model for improving off-design performance predictions for molten salt-driven Rankine power cycles, such as in concentrating solar power tower applications. The model predicts cycle off-design performance under various boundary conditions, including molten salt inlet temperature, mass flow rate, and ambient temperature. The model is validated using industry performance data and benchmarked with results from the literature. A complete concentrating solar power plant, inclusive of solar heliostat field and receiver, is then considered, by implementing the Rankine cycle off-design performance results into the National Renewable Energy Laboratory’s System Advisor Model software, which includes a tool that determines optimal power production schedules. The work improves upon the current System Advisor Model by updating off-design performance characteristics. A case study demonstrates the impact of cycle off-design behavior on annual performance for a stand-alone concentrating solar power system and a concentrating solar power-photovoltaic hybrid system. In addition, we demonstrate how cycle off-design performance influences optimal operator dispatch decisions and, thereby, overall system design and economics. We conclude that off-design cycle performance impacts “optimal” sub-system sizing, especially for a concentrating solar power-photovoltaic hybrid configuration in which concentrating solar power must dispatch in conjunction with photovoltaic generation.

14 SOLAR ENERGY↗

Multi-Core Microcontroller Hardware In the Loop System for Electric Machine Control

Hardware in the Loop (HIL) is a simulation technique used to reduce the software development cycle and test control systems in a non-destructive environment. This work describes a cost effective HIL simulator on a dual core microcontroller in which one core acts as a controller and the other emulates the system under control. The emulator runs one step per Pulse Width Modulation (PWM) period in real time. To handle the computational burden and prioritize execution of simulation and control tasks, an interrupt-based software architecture with task prioritization has been developed. As a demonstration, the HIL has been implemented on a Texas Instruments TMS320F28379D dual core microcontroller, which emulates a Permanent Magnet Synchronous Machine (PMSM) with resolver feedback. Hardware peripherals are developed and tested concurrently with the control system, providing higher confidence in the software. By using the peripherals in the HIL development, the controller exercises either the HIL emulation or a pin compatible PMSM testbench. To quantify performance and validate the processor based emulator, the HIL results are compared to the preexisting testbench for accuracy benchmarking at no-load and under load for a range of operating points.

33 ADVANCED PROPULSION SYSTEMS↗

A general Bayesian algorithm for the autonomous alignment of beamlines

Autonomous methods to align beamlines can decrease the amount of time spent on diagnostics, and also uncover better global optima leading to better beam quality. The alignment of these beamlines is a high-dimensional expensive-to-sample optimization problem involving the simultaneous treatment of many optical elements with correlated and nonlinear dynamics. Bayesian optimization is a strategy of efficient global optimization that has proved successful in similar regimes in a wide variety of beamline alignment applications, though it has typically been implemented for particular beamlines and optimization tasks. In this paper, we present a basic formulation of Bayesian inference and Gaussian process models as they relate to multi-objective Bayesian optimization, as well as the practical challenges presented by beamline alignment. We show that the same general implementation of Bayesian optimization with special consideration for beamline alignment can quickly learn the dynamics of particular beamlines in an online fashion through hyperparameter fitting with no prior information. We present the implementation of a concise software framework for beamline alignment and test it on four different optimization problems for experiments on X-ray beamlines at the National Synchrotron Light Source II and the Advanced Light Source, and an electron beam at the Accelerator Test Facility, along with benchmarking on a simulated digital twin. We discuss new applications of the framework, and the potential for a unified approach to beamline alignment at synchrotron facilities.

47 OTHER INSTRUMENTATION↗

Performance Potential of Mixed Data Management Modes for Heterogeneous Memory Systems

Many high-performance systems now include different types of memory devices within the same compute platform to meet strict performance and cost constraints. Such heterogeneous memory systems often include an upper-level tier with better performance, but limited capacity, and lower-level tiers with higher capacity, but less bandwidth and longer latencies for reads and writes. To utilize the different memory layers efficiently, current systems rely on hardware-directed, memory -side caching or they provide facilities in the operating system (OS) that allow applications to make their own data-tier assignments. Since these data management options each come with their own set of trade-offs, many systems also include mixed data management configurations that allow applications to employ hardware- and software-directed management simultaneously, but for different portions of their address space. Despite the opportunity to address limitations of stand-alone data management options, such mixed management modes are under-utilized in practice, and have not been evaluated in prior studies of complex memory hardware. In this work, we develop custom program profiling, configurations, and policies to study the potential of mixed data management modes to outperform hardware- or software-based management schemes alone. Our experiments, conducted on an Intel ® Knights Landing platform with high-bandwidth memory, demonstrate that the mixed data management mode achieves the same or better performance than the best stand-alone option for five memory intensive benchmark applications (run separately and in isolation), resulting in an average speedup compared to the best stand-alone policy of over 10 %, on average.

Effler, Chad↗

IsoForma: An R Package for Quantifying and Visualizing Positional Isomers in Top-Down LC-MS/MS Data

Proteoforms, the different forms of a protein with sequence variations including post-translational modifications (PTMs), execute vital functions in biological systems such as cell signaling and epigenetic regulation. Precisely defining the stoichiometry of PTMs has been challenging because, in the widely used bottom-up proteomics methods, the detection occurs at the peptide level and thus the link between peptides and their specific modification site is lost, resulting in proteoform ambiguity. Advances in top-down mass spectrometry (MS) technology have permitted the direct characterization of intact proteoforms and their exact number of modification sites, allowing for the relative quantification of positional isomers (PI). Proteins with positional isomers refers to proteoforms with identical total mass and set of modifications but varying PTM site combinations. The relative abundance of PI can be estimated by matching proteoform-specific fragment ions to top-down tandem MS (MS2) data to localize and quantify modifications. However, current approaches heavily rely on manual annotation. Here, we present IsoForma, an open-source R package for relative quantification of PI within a single tool. We benchmarked IsoForma’s performance against two existing workflows and highlight the similarity of the results and improvements in speed. Overall, IsoForma provides a streamlined process, reduces the time of conducting isoform-based analyses, and offers an essential framework for developing customized proteoform analysis workflows. Finally, the software is open source and available at https://github.com/EMSL-Computing/isoforma-lib.

59 BASIC BIOLOGICAL SCIENCES↗

Keeping LAMMPS cutting edge

Since its inception 30 years ago, LAMMPS has grown to be a world-class molecular dynamics code and a cornerstone of computational materials science research. This project aimed to keep LAMMPS at the forefront of molecular dynamics simulations by adapting LAMMPS to the latest developments in machine learning technology and hardware. Initially, the project set out to provide a unified implementation of active learning for efficient training data generation in LAMMPS, but the research trajectory pivoted to address more immediate and impactful opportunities. On the hardware side, recent record-breaking molecular dynamics simulations were developed on the Cerebras wafer-scale AI chip, and this project has developed an interface between LAMMPS and the hardware-specific molecular dynamics code to accelerate and simplify development and user adoption. On the software side, PyTorch’s Ahead-of-Time (AOT) compilation features promised increased performance for state-of-the-art equivariant neural network potentials, and this project laid the groundwork for their adoption in LAMMPS, resulting in a nearly 20x acceleration in extreme cases. Combined with a comprehensive benchmark study of LAMMPS across all current exascale systems, this project has reinforced LAMMPS’s role as a versatile, high-performance tool for current and future materials science applications.

36 MATERIALS SCIENCE↗

A Full-Induction Magnetohydrodynamics Solver for Liquid Metal Fusion Blankets in Vertex-CFD

Multiphysics modeling of liquid metal fusion blankets, which produce tritium and convert energy of neutrons created via fusion reactions into heat, is crucial for predicting performance, ensuring structural integrity, and optimizing energy production. While traditional blanket modeling of liquid metal flows during normal steady operating conditions commonly employs the inductionless approximation of the magnetohydrodynamics (MHD) equations, transient scenarios, when the plasma-confining magnetic field varies on millisecond time scales, require a full-induction MHD approach that dynamically evolves the magnetic field via the time-dependent induction equation. This paper presents the formulation, implementation, and initial verification of a full-induction MHD solver integrated within the open-source Vertex-CFD framework, which aims to achieve tight multiphysics coupling, a flexible software design enabling easy extension and addition of physics models, and performance portability across computing platforms. The solver utilizes finite element spatial discretization, implicit Runge–Kutta time integration, and an inexact Newton method to solve the resulting discrete nonlinear system, leveraging Trilinos packages for efficient computation. Verification against selected benchmark problems demonstrates accuracy and robustness of the solver. Furthermore, when the solver is applied to an idealized blanket model in 2.5D and full 3D, results obtained with Vertex-CFD are in good agreement with recently published quasi-2D simulations. These findings establish a computational foundation for future simulations of transient MHD phenomena in liquid metal blankets with Vertex-CFD, and open avenues for future extensions and performance optimizations.

Endeve, Eirik [ORNL] (ORCID:0000000312519507)↗

Cryo-EM model validation recommendations based on outcomes of the 2019 EMDataResource challenge

This paper describes outcomes of the 2019 Cryo-EM Model Challenge. The goals were to (1) assess the quality of models that can be produced from cryogenic electron microscopy (cryo-EM) maps using current modeling software, (2) evaluate reproducibility of modeling results from different software developers and users and (3) compare performance of current metrics used for model evaluation, particularly Fit-to-Map metrics, with focus on near-atomic resolution. Our findings demonstrate the relatively high accuracy and reproducibility of cryo-EM models derived by 13 participating teams from four benchmark maps, including three forming a resolution series (1.8 to 3.1 Å). The results permit specific recommendations to be made about validating near-atomic cryo-EM structures both in the context of individual experiments and structure data archives such as the Protein Data Bank. We recommend the adoption of multiple scoring parameters to provide full and objective annotation and assessment of the model, reflective of the observed cryo-EM map density.

59 BASIC BIOLOGICAL SCIENCES↗

RAG for FLAG: AI Assistance for a Physics Code

Artificial intelligence (AI) has quickly become an important tool in scientific research, where significant efforts are underway to develop tools that will expedite the research process. One area of particular impact is scientific software, which can be particularly complex, and therefore time consuming to learn and use effectively. AI assistants are increasingly helping to streamline the process by performing tasks such as interactively answering user questions or suggesting solutions. Los Alamos National Laboratory (LANL) develops several advanced scientific codes, such as FLAG, which can be used to run multiphysics simulations. With this study, our goal was to develop an AI assistant for FLAG that could help make the process of understanding the software and running physics simulations more efficient. To develop an AI assistant for FLAG, we used a method called retrieval-augmented generation (RAG), which is a technique that uses information from relevant data sources to enhance the accuracy of large language models (LLMs). We used the FLAG user manual and other FLAG documentation as the knowledge base for the RAG system. When a user provides a query, RAG retrieves relevant sections from the knowledge base in response, then uses those excerpts to generate grounded and contextually rich answers. We found that our AI assistant was able to provide context aware answers and source references to user queries. To evaluate performance, we developed a set of 40 benchmark questions and compared the accuracy of the responses to those of two standard LLMs without retrieval. Our AI assistant significantly outperformed the standard LLMs at answering FLAG-related questions, with an 82.5% accuracy rate, compared to 47.5% for both of the standard LLMs. This has the potential to make the process of learning and using FLAG much easier, especially for new users. Ultimately, it supports LANL’s broader mission by empowering scientists and engineers to focus more on discovery and analysis rather than on navigating complex software systems.

97 MATHEMATICS AND COMPUTING↗