Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “custom software”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13

Intelligent Experiments through Real-Time AI: Fast Data Processing and Autonomous Detector Control for High-Energy Nuclear Experiments

The aim of this project is to develop software and hardware for fast real-time data processing and autonomous detector control and calibration for the sPHENIX and the future EIC experiments. Below summarizes Georgia Tech team efforts in the past year: 1. We developed a real-time clustering algorithm and FPGA-based pipeline architecture for processing fired pixel data from ALPIDE sensors in sPHENIX experiments. Our Columnar Clustering Co-Design introduces a hardware-aware, stream-friendly approach that segments pixel data by column pairs using a Column Pair Clustering (CPC) strategy, followed by Cluster Stitching to merge adjacent subclusters. Implemented in Vitis HLS, the pipeline comprises five stages—read-in, subclustering, stitching, analysis, and write-out—connected by tagged HLS streams with custom end-of-event signaling for robust synchronization. We designed a pipelined dataflow model optimized for throughput, low latency, and minimal buffering, enabling scalable clustering across events of arbitrary size. Our system maintains spatial precision via center-of-mass and shape key extraction and efficiently handles edge cases such as fragmented or nested clusters. Compared against DBSCAN in both software and hardware, our approach demonstrates competitive performance under FPGA constraints. 2. We also conducted a comprehensive algorithm-to-hardware co-design of connected component analysis tailored for sPHENIX experiments, focusing on real-time, low-latency processing using FPGAs and High-Level Synthesis (HLS). Starting from a Python-based particle tracking pipeline, the team translated the core logic—graph traversal via DFS and Union-Find—into an HLS-compatible C++ model, replacing dynamic memory and recursion with static arrays and pipelined control flow. The final design includes a fully streamed and dataflow-compatible Union-Find kernel optimized across five iterations, incorporating loop pipelining, array partitioning, AXI/FIFO interface tuning, and function flattening. Experimental results show up to 14.8× speedup over the CPU baseline, reducing per-graph latency to 1.58 μs and demonstrating strong resource efficiency with only ~7k LUTs and zero BRAM usage. The design maintains functional correctness against the Python reference using a Python-based C-simulation framework and Mean Squared Error metrics. This work validates the potential of HLS-driven FPGA designs for edge-level HEP data acquisition, laying a scalable foundation for future integration with real-time detector pipelines and multi-graph processing systems.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Modeling Tool Development and Validation for Solar Industry Process Heat Using Particle Thermal Energy Storage

U.S. industry sectors used 26.2 quadrillion Btu and accounted for 33% of total energy consumption in 2021 according to the Energy Information Agency. Industrial process heat accounts for 70% of industrial energy use with application temperatures ranging from 60 degrees -1100 degrees C. Industry processes, heavily relying on fossil fuels of cheap coal or natural gas, differ widely in operating conditions and load requirements which makes them difficult to standardize and imposes great challenges in decarbonization. Industry processes require reliable energy supply and vary widely in temperature ranges. Storing energy from renewable sources is necessary to improve reliability and to mitigate renewable intermittency when replacing carbon fuel-based heat supplies to achieve energy savings and reduce emissions. To this end, we have developed a particle-based thermal energy storage (TES) technology using low-cost and highly stable silica sand as a storage medium. The economic and performance-based analysis is key for renewable energy sources to reliably supply industry process heat and ultimately displace fossil fuels for decarbonization. The diversified industrial processes need case-by-case analysis and design. Therefore, an adaptive modeling tool is key for renewable power with energy storage to meet industry demands. Thus, a modeling tool to simulate a solar industry process heat system using the particle TES has been developed using the object-oriented equation-based language Modelica and the commercial platform of Modelon Impact. The Modelica-based software tool provides a general simulation environment for the design of reliable solar energy sources integrated with TES for various industrial process applications at different temperatures for economic competence with fossil fuels such as coal and natural gases. It uses both customized and standard component modeling modules in Modelon libraries for the flexibility to be adapted to a specific energy demand application. The particle TES system establishes a uniform energy supply platform with an efficient heat exchanger and particle thermal energy reservoir integrated with renewable powers. The particle TES system can provide a wide temperature range and can have a large storage temperature difference that increases storage energy density; therefore, it can be an adaptable energy storage system integrated with renewable power to supply 24/7 heat for industry decarbonization.

concentrated solar thermal↗

Towards Superior Software Portability with SHAD and HPX C++ Libraries

As hardware architectures and software stacks complexity grows, development productivity, performance and software portability, quickly evolve from desirable features to actual needs. SHAD, the Scalable High-performance Algorithms and Data-structures C++ library is designed to mitigate these issues: it provides general purpose building blocks as well as high-level custom utilities, and offers a shared-memory programming abstraction which facilitates the programming of complex systems, scaling up to High Performance Computing clusters. SHAD’s portability is achieved through an abstract runtime interface, which decouples the upper layers of the library and hides the low level details of the underlying architecture. This layer enables SHAD to interface with different runtime/threading systems, e.g. Intel TBB and Global Memory and Threading (GMT). However, current backends targeting distributed systems, rely on a centralized controller which may possibly limit scalability up to hundreds of nodes and creates a network hot spot due to all to one communication for synchronization, and possibly resulting in degraded performance at high process counts. In this research, we explore HPX, the C++ standard library for parallelism and concurrency, as an additional backend in support of the SHAD library, and present the methodologies in support of local and remote task executions in SHAD with respect to HPX. Finally, we evaluate the proposed system by comparing against existing backends of SHAD and analyzing their performance on C++ Standard Template Library algorithms.

Wu, Nanmiao↗

High performance sparse multifrontal solvers on modern GPUs

Here, we have ported the numerical factorization and triangular solve phases of the sparse direct solver STRUMPACK to GPU. STRUMPACK implements sparse LU factorization using the multifrontal algorithm, which performs most of its operations in dense linear algebra operations on so-called frontal matrices of various sizes. Our GPU implementation off-loads these dense linear algebra operations, as well as the sparse scatter–gather operations between frontal matrices. For the larger frontal matrices, our GPU implementation relies on vendor libraries such as cuBLAS and cuSOLVER for NVIDIA GPUs and rocBLAS and rocSOLVER for AMD GPUs. For the smaller frontal matrices we developed custom CUDA and HIP kernels to reduce kernel launch overhead. Overall, high performance is achieved by identifying submatrix factorizations corresponding to sub-trees of the multifrontal assembly tree which fit entirely in GPU memory. The multi-GPU setting uses SLATE (Software for Linear Algebra Targeting Exascale) as a modern GPU-aware replacement for ScaLAPACK. On 4 nodes of SUMMIT the code runs ~10X faster when using all 24 V100 GPUs compared to when it only uses the 168 POWER9 cores. On 8 SUMMIT nodes, using 48 V100 GPUs, the sparse solver reaches over 50TFlop/s. Compared to SuperLU, on a single V100, for a set of 17 matrices our implementation is faster for all but one matrix, and is on average 5X (median 4X) faster

97 MATHEMATICS AND COMPUTING↗

SYMBIOSYS: A Methodology for Performance Analysis of Composable HPC Data Services

Microservices are a powerful new way of building, customizing, and deploying distributed services owing to their flexibility and maintainability. Several large-scale distributed platforms have emerged to serve the growing needs of data-centric workloads and services in commercial computing. Concurrently, high-performance computing (HPC) systems and software are rapidly evolving to meet the demands of diversified applications and heterogeneity. The interplay of hardware factors, software configuration parameters, and the flexibility offered with a microservice architecture makes it nontrivial to estimate the optimal service instantiation for a given application workload. Further, this problem is exacerbated when considering that these services operate in a dynamic and heterogeneous HPC environment. An optimally integrated service can be vastly more performant than a haphazardly integrated one. Existing performance tools for HPC either fail to understand the request-response model of communication inherent to microservices or they operate within a narrow scope, limiting the insight that can be gleaned from employing them in isolation. We propose a methodology for integrated performance analysis of HPC microservices frameworks and applications called SYMBIOSYS. We describe its design and implementation within the context of the Mochi framework. This integration is achieved by combining distributed callpath profiling and tracing with a performance data exchange strategy that collects fine-grained, low-level metrics from the RPC communication library and network layers. The result is a portable, low-overhead performance analysis setup that provides a holistic profile of the dependencies among microservices and how they interact with the Mochi RPC software stack. Using HEPnOS, a production-quality Mochi data service, we demonstrate the low-overhead operation of SYMBIOSYS at scale and use it to identify the root causes of poorly performing service configurations.

microservices↗

Prime VI

SAND2025-03757O Prime VI is a distribution-of-disease outbreak model calibration code based on variational inference. It accompanies a publication for submission to Statistics in Medicine journal, and the code will be maintained for open-source use on Sandia's GitLab. The software provides methods for calibrating an epidemiological model to measured case-count data for a multitude of correlated spatial regions. The code solves a Bayesian inverse problem for model calibration where the posterior over-model parameters are approximated through a custom implementation of variational inference. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Safta, Cosmin↗

Redis-Based Streaming Architecture for Accelerator Beam Instrumentation DAQ Systems

The Fermilab Acceleraor Division, Beam Instrumentation Department, is always adopting modern and current software methodologies for complex DAQ architectures. This paper highlights the Redis Adapter (RA) as the key software component enabling high performance, modular communication between digitizers and distributed control systems by leveraging Redis and containerization. The RA provides a unified, efficient interface between Redis based data streams and consumer systems. In the legacy architecture, digitized data flowed through the custom, UDP based Distributed Data Communication Protocol in the middle layer. In the current system, DDCP remains the ingestion path, while the RA serves as the decoupling layer. The proposed system replaces old VME digitizers with a SOM-based digitizer that communicates with Redis using the RA. The RA acts as both a performance-critical bridge and a protocol-agnostic adapter, ensuring compatibility with legacy control frameworks while enabling future scalability and modularity. This restructuring of the middle layer also helps the system achieve high throughput, reduce latency, and simplify the data path. Finally, we will demonstrate how RA is utilized in our two core products to deliver both legacy compatibility and future flexibility.

Joshi, S. [Fermilab]↗

DER Cybersecurity Detection and Response Suite

SAND2024-08475O The Distributed Energy Resource (DER) Cybersecurity Detection and Response Suite is a solution for distributed energy resource (DER) systems. The DER Security Orchestration, Automation, and Response (SOAR) solution that uses alerts from signature- and behavior-based Intrusion Detection Systems are intended to be deployed as bump-in-the-wire (BITW) devices in front of DER equipment. The fielded application would use multiple intrusion detection systems that report data to SOAR to respond to cyberattacks. The suite consists of two software components: • The proactive intrusion detection and mitigation system (PIDMS) secures grid-edge photovoltaic smart inverters and other equipment in distributed energy resource systems. It is a distributed BITW solution; cyber and physical data are automatically processed using network inspection tools and custom machine learning algorithms to detect abnormal events and correlate cyber-physical events. • The Security Orchestration, Automation, and Response for Distributed Energy Resources (SOAR4DER) application ingests data from several intrusion detection systems to quickly block attacks and revert DER systems to good states. Using a collection of intrusion detection system technologies on a BITW device, it incorporates physical and cyber data to detect abnormal and potential malicious behaviors. Multiple SOAR playbooks then use the intrusion detection system data streams to automatically defend the system. SOAR4DER system testing showed detection and response times under 30 seconds for all adversary reconnaissance, denial-of-service attacks, malicious Modbus commands, brute-force logins, and machine-in-the-middle attacks. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Johnson, Jay↗

CCSI Toolset 3.17 Release

CCSI Toolset 3.17 Release Highlights A workaround was developed to allow complex Aspen Custom Modeler (ACM) models to be used in FOQUS. This workaround uses Visual Basic for Applications to connect the ACM models to FOQUS. The ability for User plugins to be uploaded to FOQUS Cloud was added. The documentation was updated to include Optional Software Install and Tutorial Notes to clarify the usage of Turbine and SimSinter in installation instructions and adds a link to the relevant tutorial page. The Sequential Design of Experiments documentation was updated with current screenshots. The copyright was updated to include 2023.

AS↗

Graph Neural Network-based Tracking as a Service

Recent studies have shown promising results for track finding in dense environments using Graph Neural Network (GNN)-based algorithms. However, GNN-based track finding is computationally slow on CPUs, necessitating the use of coprocessors to accelerate the inference time. Additionally, the large input graph size demands a large device memory for efficient computation, a requirement not met by all computing facilities used for particle physics experiments, particularly those lacking advanced GPUs. Furthermore, deploying the GNN-based track-finding algorithm in a production environment requires the installation of all dependent software packages, exclusively utilized by this algorithm. These computing challenges must be addressed for the successful implementation of GNN-based track-finding algorithm into production settings. In response, we introduce a ``GNN-based tracking as a service'' approach, incorporating a custom backend within the NVIDIA Triton inference server to facilitate GNN-based tracking. This paper presents the performance of this approach using the Perlmutter supercomputer at NERSC.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Enabling New Flexibility in the SUNDIALS Suite of Nonlinear and Differential/Algebraic Equation Solvers

In recent years, the SUite of Nonlinear and DIfferential/ALgebraic equation Solvers (SUNDIALS) has been redesigned to better enable the use of application-specific and third-party algebraic solvers and data structures. Throughout this work, we have adhered to specific guiding principles that minimized the impact to current users while providing maximum flexibility for later evolution of solvers and data structures. The redesign was done through the addition of new linear and nonlinear solvers classes, enhancements to the vector class, and the creation of modern Fortran interfaces. The vast majority of this work has been performed “behind-the-scenes,” with minimal changes to the user interface and no reduction in solver capabilities or performance. These changes allow SUNDIALS users to more easily utilize external solver libraries and create highly customized solvers, enabling greater flexibility on extreme-scale, heterogeneous computational architectures.

97 MATHEMATICS AND COMPUTING↗

China’s Bid To Lead Artificial Intelligence Chip Development within the Decade

In 2017, the Chinese government announced an ambitious, broad strategy to become the global leader in artificial intelligence (AI) theories, technologies, and applications by 2030 and indicated that China’s ability to indigenously produce cutting-edge AI chips would be integral to its success. China’s goal to become a dominant producer of AI hardware within a decade is a bold undertaking because China faces domestic chip production challenges and fierce competition from US chip producers. Our review of Chinese AI chip product lines shows various levels where Chinese companies compete with global firms in the production chain. China lacks a robust indigenous AI chip production infrastructure, and the United States and its allies are tightening controls on the supply of advanced semiconductor manufacturing equipment (SME) and electronic design automation (EDA) software to China. In contrast to their Chinese counterparts, US companies operate in all areas of AI chip production and have consistently driven the development of new chip designs. AI chip production is a costly endeavor, and China’s chip producers currently lack the customer base to lead global AI chip development by the 2030 goal. An expanded customer base may help China attract the necessary expertise to bolster its AI chip production chain. This would also help Chinese producers make their investments in the requisite technology economically sustainable. Chinese researchers continue to pursue promising next-generation AI chip designs, such as neuromorphic computing-based architectures, which in the next five to ten years could position Chinese companies to make or contribute to key innovations in the field.

74 ATOMIC AND MOLECULAR PHYSICS↗

SD-WAN Evaluation Criteria for a Defense Information Systems Network Expeditionary Customer Edge

The Defense Information Systems Agency (DISA) has identified a service provision gap midway between the capacity and capability of a Defense Information Systems Network (DISN) transport Edge Points of Presence (POP) and U.S. Department of Defense (DoD) Enterprise Classified Travel Kit (DECTK). Some deployed force 100-user Tactical Operations Centers are in field locations far removed from the DISN transport core, in challenging environmental conditions with limitations on space, weight, power, and cooling for network equipment. As part of a multi-phase project, Pacific Northwest National Laboratory (PNNL) will gather and analyze requirements, design, develop, prototype, and test a miniaturized and ruggedized DISN Customer Edge POP scaled to support approximately 100 users in a field Tactical Operations Center. An Expeditionary Customer Edge (ECE) will be more suitable for field deployment than a DISN transport Edge POP, using ruggedized hardware and network function virtualization (NFV) to operate in challenging field environmental conditions, plus reduce weight, power, and cooling requirements. A key enabling technology for ECE is Software Defined Wide Area Networking (SD-WAN). This document provides: • A brief overview of SD-WAN use cases and how they apply to ECE • How SD-WAN technology compares to existing networks like DISN that use Optical Transport Networking (OTN) and Multiprotocol Label Switching (MPLS) • SD-WAN functional requirements relevant to ECE • A description of an ECE Virtual Prototype, including simulated wide area network paths and a DISN-representative implementation of MPLS, in Cisco Modeling Labs • Detailed network flow walkthroughs • Evaluation criteria based on ECE-relevant SD-WAN functional requirements Finally, an appendix provides an informal evaluation of Speedify—a commercial retail Virtual Private Network service—against the ECE SD-WAN evaluation criteria.

42 ENGINEERING↗

Dashboard for Marine Energy Site Assessment and Monitoring

The marine energy (ME) industry presently relies upon fragmented site assessment solutions that require high resource expenditure for deployment at each site and do not leverage the wealth of readily available tools and information. A wave energy resource assessment dashboard, currently in development, will substantially improve siting, permitting, operations, and maintenance of ME projects by providing an integrated solution that is a one-stop-shop for a developer’s needs. The Site Energy Assessment and MOnitoring Dashboard (SEAMOD) will be of commercial interest to anyone seeking to deploy an ME project and is easily expandable to include tidal and wind energy site assessments. The integrated dashboard is being developed using state-of-the-art database and cloud computing methods and data-assimilative modeling tools that can be coupled with low-cost, rapidly deployable wave buoys and environmental sensing hardware. The combined software and hardware dashboard will reduce wave energy site characterization and wave climate monitoring costs by more than 60 percent and provide assessments that meet international industry standards. To realize a thriving global ME industry, the physical environment at a potential deployment site must be understood, not only for resource characterization, but also for optimization of device and power conversion performance. SEAMOD directly addresses these needs with a commercially marketable product. SEAMOD is a low-cost solution that provides comprehensive ME resource assessments, baseline environmental monitoring, and offshore characterizations required for successful ME development. The key technical objectives for Phase I were a series of software development goals, which when implemented with monitoring solutions, produced an initial proof-of-concept low-cost wave energy resources dashboard. In Phase II, the development of the prototype SEAMOD continued. The basic framework employed was the development of a revised dashboard and monitoring tool customized for ME applications by focusing on IEC site assessment and method requirements. Development was focused on the integration of full hindcast metocean products to provide hindcast resource characterization and environmental information. The final integrated dashboard provides a low-cost solution that delivers comprehensive, scalable, industry-standard energy resource assessments and offshore characterizations required for successful ME development. The integrated dashboard offers visibility of the most recent site modeling, measurements, and historical data. The application and integration of consensus-based standards for wave energy resource assessment, as determined by the International Electrotechnical Commission (IEC), are crucial for the impact and value of SEAMOD. SEAMOD includes monthly, seasonal, and yearly statistics, as well as the total 30-year record, offering temporal resolution of the IEC parameters to aid potential developers in determining the available wave energy resources in their area of interest.

16 TIDAL AND WAVE POWER↗

Software-Defined Network for End-to-end Networked Science at the Exascale

Domain science applications and workflow processes are currently forced to view the network as an opaque infrastructure into which they inject data and hope that it emerges at the destination with an acceptable Quality of Experience. There is little ability for applications to interact with the network to exchange information, negotiate performance parameters, discover expected performance metrics, or receive status/troubleshooting information in real time. The work we presen here is motivated by a vision for a new smart network and smart application ecosystem that will provide a more deterministic and interactive environment for domain science workflows. The Software-Defined Network for End-to-end Networked Science at Exascale (SENSE) system includes a model-based architecture, implementation, and deployment which enables automated end-to-end network service instantiation across administrative domains. An intent based interface allows applications to express their high-level service requirements, an intelligent orchestrator and resource control systems allow for custom tailoring of scalability and real-time responsiveness based on individual application and infrastructure operator requirements. This allows the science applications to manage the network as a first-class schedulable resource as is the current practice for instruments, compute, and storage systems. Deployment and experiments on production networks and testbeds have validated SENSE functions and performance. Emulation based testing verified the scalability needed to support research and education infrastructures. Key contributions of this work include an architecture definition, reference implementation, and deployment. This provides the basis for further innovation of smart network services to accelerate scientific discovery in the era of big data, cloud computing, machine learning and artificial intelligence.

97 MATHEMATICS AND COMPUTING↗

Software-defined network for end-to-end networked science at the exascale

Domain science applications and workflow processes are currently forced to view the network as an opaque infrastructure into which they inject data and hope that it emerges at the destination with an acceptable Quality of Experience. There is little ability for applications to interact with the network to exchange information, negotiate performance parameters, discover expected performance metrics, or receive status/troubleshooting information in real time. The work presented here is motivated by a vision for a new smart network and smart application ecosystem that will provide a more deterministic and interactive environment for domain science workflows. The Software-Defined Network for End-to-end Networked Science at Exascale (SENSE) system includes a model-based architecture, implementation, and deployment which enables automated end- to-end network service instantiation across administrative domains. An intent based interface allows applications to express their high-level service requirements, an intelligent orchestrator and resource control systems allow for custom tailoring of scalability and real-time responsiveness based on individual application and infrastructure operator requirements. This allows the science applications to manage the network as a first-class schedulable resource as is the current practice for instruments, compute, and storage systems. Deployment and experiments on production networks and testbeds have validated SENSE functions and performance. Emulation based testing verified the scalability needed to support research and education infrastructures. Key contributions of this work include an architecture definition, reference implementation, and deployment. This provides the basis for further innovation of smart network services to accelerate scientific discovery in the era of big data, cloud computing, machine learning and artificial intelligence.

47 OTHER INSTRUMENTATION↗

Versatile TRISO fuel particle modeling in Bison

Tri-structural isotropic (TRISO) fuel particles are a key component in several previous and current reactors as well as in a variety of novel nuclear reactor designs. Interest in TRISO fuel is on the rise, necessitating considerable computer modeling of TRISO fuel behavior in order to support related design and licensing activities. The Bison nuclear fuel performance code, which offers a full set of capabilities for modeling TRISO fuels, makes it easier to explore the various important aspects of TRISO fuel behavior. One key advantage of Bison is its ability to create meshes in 1D, 2D, and 3D. Users can customize these meshes for specific geometries, mesh densities, and use cases. This enables a wide variety of analyses, including thermal, structural, mass diffusion, homogenization, and statistical failure analyses. Furthermore, the meshing capability simplifies analysts’ workflows. The inherent mesh generation capability eliminates the need for separate mesh-generating software and mesh file management. Also, the fact that the meshes are customizable makes it straightforward to automate an investigation over a range of geometric parameters or mesh densities. Here, the present paper highlights the ease with which Bison may be used to create meshes for both simple and relatively complex TRISO fuel particles, and it explores the types of analyses enabled by these meshes.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Zero-Order Reaction Kinetics (Zero-RK): Enabling the Use of Detailed Chemical Kinetics in Combustion Simulations (CRADA Final Report)

This was a collaborative effort between Lawrence Livermore National Security, LLC (LLNS), as manager and operator of Lawrence Livermore National Laboratory (LLNL) and Gamma Technologies, LLC (GT or Participante), to incorporate the ability to access LLNL chemical kinetics technologies while using GT-SUITE, GT’s market leading engine simulation software. At the end of the project, LLNL has released Zero-RK version 3.5 with zero- and one-dimensional (0-D and 1-D) solver functionality that interfaces with GT’s GT-SUITE v2023 and later releases. GT has tested its product to assure their customers that the interface can provide reduction in chemistry solution time for detailed chemistry simulations. The process has also positioned GT to easily benefit from future improvements of the Zero-RK suite of tools developed under the DOE Vehicle Technologies Office.

33 ADVANCED PROPULSION SYSTEMS↗