Engineering PapersSearch

SEARCH · Engineering Papers

Results for “code generation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Kinetic Model of HoxEFU reduction by NADH [SWR-26-087]

This repository is used to release code generated for manuscripts on the Photosynthetic Energy Transduction core program. This code simulates the reduction of HoxEFU by NADH. The electro transfer rate constants for the simulation are specified in the .csv files. The two .csv files correspond tot he two models described in Dawson et al. Cell. Rep. Phys. Sci. 2026. The code utilizes a chemical master equation, a set of differential equations, defining the time evolution of the oxidation and reduction kinetics of NAD+, NADH, a FMN flavin, and a set of iron sulfur clusters. The kinetics of HoxEFU reduction by NADH are evaluated by numerical integration of the chemical master equation using a variable-time-step Runge-Kutta algorithm.

Dahl, Peter [National Laboratory of the Rockies (N

Kinetic modeling of hot tail runaway electron generation during plasma disruptions using the JOREK code

The generation of runaway electrons (REs) during disruptions poses a significant challenge for the operation of tokamaks. The production of these high-energy electrons can cause substantial damage, particularly when the plasma current is high, making it a critical concern for ITER. For the high-temperature plasmas anticipated in ITER, the primary generation of REs may be dominated by the hot tail mechanism, which consists of the acceleration of hot electrons from the pre-disruption population which have not yet thermalized with the bulk following the rapid cooling of the plasma. To account for the significant 3D effects on RE production, a hot tail modeling framework has been developed within the non-linear 3D extended MHD code JOREK. This paper presents the structure of this framework, which is based on test electrons evolving in MHD fields. The verification of the method shows good agreement with the reference DREAM code for 0D test cases, as well as for axisymmetric simulations of 15 MA ITER H-mode disruption scenarios. Furthermore, a proof-of-principle application to a DIII-D case demonstrates the framework’s capability to capture for the first time the hot tail generation in 3D MHD simulations in realistic geometry. Preliminary results suggest that the production of REs is significantly reduced by stochastic losses.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY

Verification of the ENDF/B-VII.1 Based MC 2 -3 Library Rev.1

The MC 2 -3 code, developed by Argonne National Laboratory under the DOE-NE NEAMS program, is a multigroup cross section generation code for fast reactor applications. Last year, the ENDF/B-VII.0 (E70) MC 2 -3 library, which has been extensively used, verified, and validated over a long period, was intensively reverified and updated to support the commercial grade dedication (CGD) requirement of the TerraPower Natrium project. This year, the ENDF/B-VII.1 (E71) MC 2 -3 library, the preliminary version of which was generated several years ago, was regenerated and rigorously verified to support the Natrium project as well as the completion of verification of the E71 library. The E71 library was verified using the process developed during the verification of the E70 library, including comparisons of cross sections with the NJOY-generated cross sections, comparisons of the resolved resonance cross sections with those using the PEDNF library, and comparison of total cross sections with the sum of partial cross sections. Additional verifications were conducted to ensure that the benchmark problem solutions with the E71 library are reasonable compared to the corresponding Monte Carlo solutions. Furthermore, the E71 gamma library was generated, which includes data for prompt gamma, delayed gamma, and delayed beta as well as neutron and gamma heating. The gamma library was verified at the level of individual isotopes. The EBR-II core solutions from MC 2 -3/ DIF3D and MCNP were compared, demonstrating that those solutions in terms of k-effective and assembly powers were in good agreement.

22 GENERAL STUDIES OF NUCLEAR REACTORS

Fiscal Year 2024 Software Quality Assurance Activities for the ARC Software

The continued goal of the ARC SQA project in the Advanced Reactor Technologies program of DOE is to resolve the QA gaps for the ARC software that limit, or prevent, commercialization of the software for industry users. This project started in earnest in fiscal year 2023 which saw the entire code system moved from a SVN repository to a GitLab repository and an associated software quality assurance plan (SQAP) developed and ratified. Most of the QA gaps in the ARC software were identified in collaboration with industry partners and work begin in fiscal year 2023 and continued in 2024. The primary documentation that is missing includes user manuals, user guides, software verification reports, and code coverage assessments. The SUMMAR manual was completed this fiscal year and work was started on creating manuals for SE2ANL, SE2RCT, and DASSH. Software verification work was carried out for DIF3D and REBUS in a previous program and the current fiscal year saw the completion of software verification reports for GAMSOR, GAMSRC, VARPOW, EvaluateFlux, and SUMMAR. The goal for the next fiscal year is to complete the PERSENT software verification work and begin planning the software verification work for DASSH, SE2ANL, and SE2RCT. The code coverage reports for DIF3D and MC2-3 were completed in the previous fiscal year and the goal is to generate code coverage reports for REBUS, GAMSOR, PERSENT, and DASSH in the coming fiscal year. A considerable amount of effort was spent in the current fiscal year working on the continuous integration capability for automated regression testing in GitLab. The first version of the testing was created in the previous fiscal year and applied to DIF3D and its utility programs. That testing was extended this year to cover GAMSOR, REBUS, and PERSENT. To accomplish this, the first version of the new testing methodology had to be updated to make a single output checking methodology viable for all of the ARC software. This will result in a single document to detail the automated regression testing methodology and minor documents to detail the tolerance settings that have been applied to the output for each ARC code. The previous methodology put into place with SVN would have required a separate document for each ARC code to detail the output checking methodology and the tolerance settings for the output from each code. Because some of our industry partners are providing funds to add new capabilities to the ARC software to meet their needs, all of which must be reviewed and approved by the SQA program funded by this project, a summary of that development work is detailed in this report. Overall progress on resolving the QA gaps has been good this year with the most impactful improvement for our industry partners in capability being the creation of a threaded version of DIF3D-VARIANT that allows the DIF3D, REBUS, and GAMSOR run times to be reduced by a factor of 4-6. The most impactful QA gap that was resolved was the software verification of GAMSRC and VARPOW.

97 MATHEMATICS AND COMPUTING

wa-hls4ml and lui-gnn: A benchmark and GNN-based surrogate model for hls4ml resource and latency estimation

As machine learning (ML) increasingly serves as a tool for addressing real-time challenges in scientific applications, the development of advanced tooling has significantly reduced the time required to iterate on various designs. These advancements have solved major obstacles, but also exposed new challenges. For example, processes that were not previously considered bottlenecks, such as model synthesis, are now becoming limiting factors in the rapid iteration of designs. To reduce these emerging constraints, multiple efforts are being launched toward designing an ML-based surrogate model that estimates resource usage of synthesized accelerator architectures. This model would reduce the design iteration time, especially when designing within a set of given hardware constraints. This approach shows considerable potential, but as it stands, the effort is early and would benefit from coordination and standardization to assist future work as it emerges. We introduce wa-hls4ml, a benchmark for ML accelerator resource and latency estimation, and its corresponding initial dataset of more than 100,000 fully connected neural networks, all synthesized using hls4ml and targeting Xilinx FPGAs. In addition to the resource utilization and latency data provided, the dataset includes generated artifacts and log files for many of the synthesized neural networks, in order to support future research in ML-based code generation. The benchmark evaluates the performance of resource and latency predictors against several common ML model architectures, primarily originating from scientific domains, as exemplar models, as well as the average performance across a subset of the dataset. We measure the performance of a given predictor model through multiple metrics, including $R^2$ score and SMAPE on regression tasks, as well as inference time to further characterize the estimator under test. Additionally, we introduce the latency/utilization inference graph neural network (lui-gnn), a surrogate model that uses a graph neural network to represent input architectures in the form of a directed graph. This graph representation allows for a diverse set of model architectures to all be effectively handled by a surrogate model. We present the architecture and performance of the model, as evaluated by the new proposed benchmark, including SMAPE, $R^2$ score, and inference times, and find that lui-gnn generally predicts latency and utilization for the 75\% quantile within several percent of the synthesized resources on the synthetic test dataset, indicating that this approach of estimating resource and latency via a surrogate models has promise and warrants further research.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS

LibraryX: A Framework for Cross-Library-Call Optimization

Scientific applications utilize performance libraries as a software engineering concept: these libraries encapsulate important and well-understood (mathematical) operations, allow for reuse, and are implemented and tuned by experts. Domain scientists then implement complex algorithms based on these domainspecific libraries. While individual library calls are optimized, larger performance gains across sequences of calls—sometimes spanning multiple libraries—are often unrealized, forcing a trade-off between performance and implementation complexity.To overcome this issue, we propose LibraryX, an approach and a system that allows for cross-library-call optimization even when library calls stem from multiple performance libraries. LibraryX annotates library calls with semantic information and optimizes entire directed acyclic graphs (DAGs) of calls dynamically using the SPIRAL code generation system. We demonstrate its effectiveness across a range of memory bound workloads, achieving significant speedups on Nvidia, AMD, and Intel accelerators compared to code using native libraries without cross-call optimization.

Rao, Sanil [Carnegie Mellon University,Department

Data and Code for Understanding Generative AI Content with Embedding Models

This repository contains code for the experiments in the paper "Understanding Generative AI Content with Embedding Models". Constructing high-quality features is critical to any quantitative data analysis. While feature engineering was historically addressed by carefully hand-crafting data representations based on domain expertise, deep neural networks (DNNs) now offer a radically different approach. DNNs implicitly engineer features by transforming their input data into hidden feature vectors called embeddings. For embedding vectors produced by foundation models -- which are trained to be useful across many contexts -- we demonstrate that simple and well-studied dimensionality-reduction techniques such as Principal Component Analysis uncover inherent heterogeneity in input data concordant with human-understandable explanations. Of the many applications for this framework, we find empirical evidence that there is intrinsic separability between real samples and those generated by artificial intelligence (AI).

Vargas, Max [Pacific Northwest National Laboratory

MFC 5.0: An exascale many-physics flow solver

Many problems of interest in engineering, medicine, and the fundamental sciences rely on high-fidelity flow simulation, making performant computational fluid dynamics solvers a mainstay of the open-source software community. Previous work MFC 3.0 was made a published, documented, and open-source solver via Bryngelson et al. Comp. Phys. Comm. (2021) with numerous physical features, numerical methods, and scalable infrastructure. MFC 5.0 is a significant update to MFC 3.0, featuring a broad set of well-established and novel physical models and numerical methods, as well as the introduction of GPU and APU (or superchip) acceleration. Here, we exhibit state-of-the-art performance and ideal scaling on the first two exascale supercomputers, OLCF Frontier and LLNL El Capitan. Combined with MFC’s single-accelerator performance, MFC achieves exascale computation in practice, and achieved the largest-to-date public CFD simulation at 200 trillion grid points as a 2025 ACM Gordon Bell Prize finalist. New physical features include the immersed boundary method, N-fluid phase change, Euler–Euler and Euler–Lagrange sub-grid bubble models, fluid-structure interaction, hypo- and hyper-elastic materials, chemically reacting flow, two-material surface tension, magnetohydrodynamics (MHD), and more. Numerical techniques now represent the current state-of-the-art, including general relaxation characteristic boundary conditions, WENO variants, Strang splitting for stiff sub-grid flow features, and low Mach number treatments. Weak scaling to tens of thousands of GPUs on OLCF Summit and Frontier and LLNL El Capitan achieves efficiencies within 5% of ideal to over 90% of their respective system sizes. Strong scaling results for a 16-times increase in device count show parallel efficiencies over 90% on OLCF Frontier. MFC’s software stack has undergone further improvements, including continuous integration, which ensures code resilience and correctness through over 300 regression tests; metaprogramming, which reduces code length while maintaining performance portability; and code generation for computing chemical reactions

Computational fluid dynamics

Evaluation of PBR Spent Fuel Criticality and Dose Rate Compliance for Storage and Transportation

Spent tri-structural isotropic (TRISO)–based fuels have a strong track record in storage and transportation without documented incidents. This work seeks to reduce uncertainty to aid in more informed spent fuel management of TRISO-based fuels by modeling both fresh and spent pebble bed reactor (PBR) fuel and comparing the results to the regulatory standards from 10 CFR 71. SCALE was used for all modeling due to it having fast and accurate methods for handling PBR fuel modeling, as well as having an efficient method for shielding calculations in monaco with automated variance reduction using importance calculations (MAVRIC), which utilizes the consistent adjoint-driven importance sampling (CADIS) and the forward-weighted consistent adjoint-driven importance sampling (FW-CADIS) methods. KENO-VI was used for all criticality calculations, TSUNAMI was used for uncertainty quantification on k-effective, TRITON and the Oak Ridge isotope generation code (ORIGEN) were both used for depletion of the fuel, and MAVRIC was used for shielding calculations. For criticality assessments, this study focused on the requirement that the value of the neutron multiplication factor, k-effective (k-eff), would not exceed a peak value of 0.95, including uncertainty, with 95% confidence. Criticality was initially examined by modeling fresh fuel from three different designs—HTR-10 fuel, PBMR-400 fuel, and demonstration fuel representative of a TRISO-fueled modern high-temperature gas reactor (HTGR) design, henceforth referred to as Demo HTGR—and placing them into various sized containers with conditions described in 10 CFR 71 to quantify the peak k-eff state. When the peak value of 0.95 k-eff was exceeded, mitigation methods were examined in those scenarios. Burnup credit, pebble displacement in areas of strong neutron multiplication, and random pebble replacement using pebbles of various compositions and replacement fractions were examined. In summary, the criticality of PBR fuels can be well accounted for by restricting container size, taking credit for burnup, or by displacing/replacing pebbles. Uncertainty of the k-eff due to nuclear data uncertainties was recorded at ~0.6644%Δk/k, or roughly 664% mil (pcm). The nuclear data–induced uncertainty was relatively small and should not require significant modification in the design to be accounted for. Revisions to the evaluated nuclear data file values have been shown to have a larger impact than nuclear data–induced uncertainty. For dose rate aspects, U.S. Nuclear Regulatory Commission regulations require a maximum dose rate of 10 millirem per hour (mrem/h) at 2 meters. In examining the dose rate behavior of spent PBR fuel, the representative Demo HTGR fuel was modeled exclusively due to it possessing the highest target burnup of the examined fuels. Equilibrium cycle modeling methods were used to produce a higher-fidelity discharge isotopic composition than simple assumptions, such as reflected pebbles. The discharge composition was used as a source term in the fixed-source transport shielding calculations, and dose rates were calculated at 2 m for the shortest possible cooling time. The low concentration of fuel material led to dose rates that were in line with regulatory limits, despite the high burnup when compared to traditional light water reactor fuels. In conclusion, the methods employed in this study would require more work to further verify and validate and are limited to the criticality and dose rate analyses performed.

11 - NUCLEAR FUEL CYCLE AND FUEL MATERIALS

tether

Tether is a python module for benchmarking and assessing large language model (LLMs) performance at generic scientific tasks. The code generates benchmarks, uses the benchmark to prompt LLMs through automatic programming interfaces (APIs), and then logs the number of prompts an LLM correctly answers and presents the results as a completed benchmark.

Kaiser, Bryan [Los Alamos National Laboratory]

TInCup

SAND2025-11464O TInCuP (Tag Invoked Customization Points) is a modern, header-only C++20 library that addresses the boilerplate problem in tag invoke based customization points. It provides comprehensive code generation and verification tools. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

von Winckel, Gregory [Sandia National Lab. (SNL-CA

ellora-spack-gen

The project contains software to analyze the behavior of large language models at code generation tasks. Publicly available information is used to generate Spack package recipes for HPC developers. The software contains orchestration tooling, analysis, and plotting functionality.

Melone, CaetanoN

Flash-X Recipe Tools

SF-24-102 A code generation tool for Flash-X to support their performance portability using a domain-specific runtime library.

Lee, Youngjun

GridSTIX

SF-25-112 Grid-STIX is a comprehensive extension of the STIX (Structured Threat Information Expression) 2.1 ontology specifically designed for electrical grid cybersecurity applications. This ontology provides a standardized, machine-readable framework for modeling grid assets, operational technology devices, threats, vulnerabilities, supply chain risks, and security relationships in electrical power systems. ## Key Features - **Comprehensive Grid Coverage**: Physical assets, OT devices, grid components, sensors, and energy storage systems - **Zero Trust Architecture**: Policy decision points, enforcement points, trust brokers, and continuous monitoring - **AMI Infrastructure**: Advanced metering networks, head-end systems, mesh gateways, and MDM systems - **Advanced Security Modeling**: Attack patterns, vulnerabilities, mitigations, and supply chain risks - **Critical Grid Relationships**: Power flow, protection, control, and synchronization relationships - **Supply Chain Security**: Supplier modeling, country of origin tracking, and risk assessment - **Protocol Support**: DNP3, Modbus, IEC 61850, IEC 60870-5-104, OPC-UA, and IEEE standards - **Python Code Generation**: Automated STIX-compliant Python class generation from ontologies - **Interactive Visualization**: Enhanced HTML network graphs with grid-specific categorization - **STIX 2.1 Compliance**: Full compatibility with STIX threat intelligence ecosystem

Blakely, Benjamin [Argonne National Laboratory (AN

Power Electronics Manufacturing Improvements for Heavy-Duty Fuel Cell Vehicles

The Marel Power Solutions project, funded by the U.S. Department of Energy under Award DE-SC0023801, focused on advancing manufacturing techniques for power electronics in heavy-duty fuel cell vehicles. The research aimed to enhance system efficiency, reduce costs, and support broader adoption of hydrogen fuel cell technology. Key areas of investigation included power topology, thermal modeling, system architecture, and accessibility through software tools. Key Accomplishments: 1. Power Topology: - Developed an interleaved boost converter with optimized phase count, leveraging Marel’s proprietary Power Stacks. - Achieved reduced parasitic inductance and resistance, enabling high efficiency in DC-DC converters. 2. Thermal Modeling: - Integrated innovative cooling systems into compact Silicon Carbide (SiC) modules. - Simulations demonstrated the ability to dissipate significant heat (up to 7.5 kW), ensuring device reliability under heavy loads. 3. System Architecture: - Utilized simulation tools to analyze the impact of various fuel cell and vehicle parameters on efficiency. - Highlighted the role of smaller, modular improvements, such as enhanced DC-DC converters, in achieving system-wide gains. 4. Accessibility: - Evaluated and implemented MATLAB/Simulink code generation tools for real-world hardware applications. - Demonstrated the potential for rapid prototyping of custom power systems with reduced development costs. Impact and Benefits: - Efficiency and Cost Reduction: Marel’s cooling technology enhances SiC die performance, reducing the number of dies required and overall system size. - Scalability and Flexibility: The innovations support tailored solutions for diverse applications, from mass transit to mining vehicles. - Sustainability: The research promotes the integration of electrification technologies, helping meet rising energy demands sustainably. Conclusion: The project’s outcomes advance the state of power electronics for hydrogen fuel cell vehicles, enabling more efficient, compact, and cost-effective solutions. These developments lay a foundation for future innovation, contributing to the broader adoption of clean energy technologies in transportation and other industries.

08 HYDROGEN

Burnup Monitoring for Pebble Bed Reactor Systems

A pebble burnup monitoring system is a required component for domestic reactor safety and safeguards applications associated with pebble bed reactors (PBRs). One of the main requirements of a PBR burnup monitoring system is that it needs to be capable of rapid measurements to assess the burnup of each individual pebble to determine whether to recirculate it in the reactor or discard it as spent fuel. This report considers three different approaches for a burnup monitoring system for pebbles discharged from the reactor core in a pebble bed modular reactor-400: • passive gamma spectrometry measurement, • passive neutron coincidence measurement, and • active neutron counter based on the differential die-away technique. Conceptual designs have been created for each of these detectors, and preliminary analysis has been performed using Monte Carlo N-Particle and Oak Ridge Isotope Generation code simulations. The advantages and practical limitations (e.g., high radiation background) of each system were identified. Simulations suggest that each of the three measurement techniques can be successfully employed to distinguish between pebbles based on their number of passes through the core and to quantify the burnup of pebbles.

22 GENERAL STUDIES OF NUCLEAR REACTORS

Experiences with SYCL on AMD GPUs with Kokkos

With the recent diversification of the hardware landscape in the high-performance computing (HPC) community, performance-portability solutions are becoming more and more important. One of the most popular choices is Kokkos, which recently became a Linux Foundation project. Most of its development is supported by the US Department of Energy and the French Alternative Energies and Atomic Energy Commission. Kokkos is implemented as a C++ library with multiple backends to support CPUs as well as various GPU architectures. These backends include OpenMP, CUDA, HIP, and also SCYL. This approach enables users to leverage the preferred vendor toolchain for the respective platform (e.g. CUDA, ROCm, OneAPI). The SYCL backend is used to target Intel GPUs, in particular to support the Aurora exascale supercomputer. However, SYCL itself also offers a large degree of portability, and in fact Kokkos’ CI for SYCL has been running on NVIDIA hardware due to a lack of access to Intel GPUs. In this report, we describe our experience with using Kokkos SYCL backend on AMD GPUs targeting the Frontier supercomputer at Oak Ridge National Laboratory. The two major SYCL implementations are DPC++ and AdaptiveCpp. While the Kokkos SYCL backend has been implemented using the former, the latter was the first implementation to target AMD GPUs. We will discuss the experience with both of these SYCL implementations in terms of functionality and performance. Using Kokkos to evaluate SYCL toolchains has a number of benefits. Kokkos’ use of SYCL is fairly complex, exercising features such as graphs, relocatable device functions, atomics – including for non-arithmetic types, as well as pinned and page migratable memory allocations. Kokkos also needs to implement capabilities such as Kokkos’ hierarchical parallelism that are not a straight-forward mapping to SYCL capabilities. Furthermore, a large number of libraries and applications that represent diverse use cases are implemented in Kokkos, providing readily available test cases for a toolchain evaluation. Preliminary results show that support for AMD GPUs in DPC++ is much less mature than for NVIDIA GPUs or Intel GPUs. While the situation has improved significantly over the last year, we still encounter many runtime failures, dispatching problems, and code generation issues. With AdaptiveCpp the challenges arise even earlier in the evaluation process. Since Kokkos’ SYCL implementation is largely focused on supporting Intel GPUs, we opted to leverage SYCL extensions which are available in DPC++ but not in AdaptiveCpp. Furthermore, AdaptiveCpp appears to be less conformant with the SYCL2020 standard which Kokkos relies on. In some cases, we are able to work around the lack of feature support, in other cases we have to disable certain Kokkos capabilities to evaluate the toolchain. Our evaluation will leverage Kokkos’ unit tests to establish basic functionality and feature completeness. We then use simple benchmarks for components of a CG implementation as a measure of usability and performance of the SYCL toolchains.

97 MATHEMATICS AND COMPUTING

Multiphysics Time-Integration for Turbulent Combustion at the Exascale

Turbulent reacting flow systems are often modeled with coupled time-dependent partial differential equations (PDEs). Solving such equations can easily tax the world's largest supercomputers. One pragmatic strategy for attacking such problems is to split the PDEs into components that can more easily be solved in isolation. This generic operator-splitting strategy leads to a set of ordinary differential equations (ODEs) that need to be solved as part of an "outer-loop" time-stepping approach. In many combustion applications, the ODEs to be solved can be very stiff, exhibiting timescales that span many orders of magnitude. The SUNDIALS library provides a plethora of robust time integration algorithms for solving these ODEs on exascale-capable computing hardware, yet for many complex applications (such multicomponent fuels or emissions predictions), the chemical models remain too complex to solve using reasonable resources. The Quasi-Steady State Approximation (QSSA) can be an effective tool for reducing the size and stiffness of the simulations. In this talk, I will discuss the use of the SUDIALS library of ODE solvers together with automatic code generation tools to solve complex turbulent reacting flow problems using QSSA models.

chemistry