Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “scalable performance”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Performance of an Astrophysical Radiation Hydrodynamics Code under Scalable Vector Extension Optimization

We present results of a performance study of an astrophysical radiation hydrodynamics code, V2D, on the Arm-based A64FX processor developed by Fujitsu. The code solves sparse linear systems, a task for which the A64FX architecture should be well suited. Here, we performed the performance analysis study on Ookami, an Apollo 80 platform utilizing the A64FX processor. We explored several compilers and performance anal-ysis packages and found the code did not perform as expected under scalable vector extension optimization, suggesting that a “deeper dive” into analyzing the code is worthwhile. However, a simple driver program that exercised basic sparse linear algebra routines used by V2D did show significant speedup with the use of the scalable vector extension optimization. We present the initial results from the study which used V2D on a relatively simple test problem that emphasized the repeated solution of sparse linear systems.

79 ASTRONOMY AND ASTROPHYSICS↗

Energy-efficient scientific computing using chemical reservoirs

The rapid growth of computing demands driven by scientific computing, data analytics, and artificial intelligence (AI) advancements has exposed the limitations of traditional digital processing systems. These systems are nearing physical energy barriers, making significant gains in energy efficiency increasingly unattainable. As we advance toward post-exascale computing, disruptive approaches are critical to overcoming these limitations. Among emerging analog solutions, biochemical computing offers a transformative path for achieving orders-of-magnitude improvements in energy efficiency. By leveraging the natural optimization capabilities of chemical reaction networks (CRNs), biochemical systems have the potential to meet high-performance computing needs through natural scalability. However, numerous challenges remain, including theoretical limitations in mapping computational problems to CRNs and practical barriers in implementing biochemical computing devices. In this paper, we present a framework for chemical computation using biochemical systems and introduce key components of our approach for energy-efficient scientific computing. We showcase the feasibility of this framework by solving a system of ordinary differential equations by emulating a chemical reservoir device, demonstrating its potential for addressing modern computing challenges. This work lays a foundational step toward harnessing the computational power of chemistry to design energy-efficient, scalable, high-performance next-generation computing systems.

Johnson, Connah G. M. [Pacific Northwest National ↗

Multi-modal Energy-optimal Trip Scheduling in Real-time (METS-R) for Transportation Hubs (Final Report)

This report summarizes the work performed under the award number EE0008524. The project develops the Multi-modal Energy-optimal Trip Scheduling in Real-time (METS-R) platform as the next-generation transportation solution based on autonomous electric vehicles (AEV) serving passenger trips from and to urban transportation hubs, to substantially reduce transportation energy consumption. Extensive data collection and analyses were first conducted to understand the demand patterns and energy consumption of hub-based on-road trips. Then, a data-driven framework that consists of an analytical module and a simulation module was proposed. For the analytical module, five planning + operation tools were developed to support the planning and energy-efficient operations of urban AEV services: the charging station planning that robotically allocates charging supplies based on the stationary charging demand distribution; the transit planning and demand adaptive scheduling model that efficiently generates\ candidate transit routes from hubs to other places and dynamically adjusts the transit time table to fit the current demand; the online energy-efficient routing that learns the energy-optimal paths from observations of link-level energy consumption in real-time; the hub-based ridesharing that matches trip requests together with account for the uncertainty of future trip demand and vehicle supply; and finally, the integrated demand prediction and anomaly detection pipeline that leverages the flight/train time table and support other planning/operation tools. To demonstrate the performance of these tools, a scalable high-performance agent-based simulator was built. We divided the urban space into multiple service zones where each zone was considered as an agent for passenger generation and vehicle charging. Two types of AEV agents were coded to model two types of mobility services: AEV taxi and AEV transit. For the AEV taxi, the team implemented the functions of pickup/drop-off passengers, energy-efficient routing, ridesharing, fleet rebalancing, and recharging. For the AEV bus, the team implemented the functions of demand-adaptive route scheduling, passenger boarding, and recharging. A high-performance computing framework was introduced to receive various profiling information (such as link energy updates, vehicle speed) from the simulator instances and communicate the operational commands back to the instances. The numerical experiments show that each of the proposed operational algorithms can reduce energy consumption and improve system efficiency. Furthermore, there exists the need to collectively consider multiple planning + operational strategies as multiple strategies can influence each other in terms of performance impacts. Recommendations for future work related to AEV planning and simulation are discussed.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Scalable, low-cost ink-based processing of high-performance silver selenide thermoelectrics

The growing global energy demand and its accelerating contribution to climate change emphasize the urgent need for sustainable energy conversion/harvesting technologies. Thermoelectric (TE) devices offer a compelling route to directly convert waste heat into electricity and enable solid-state cooling without moving parts or harmful refrigerants. Achieving their full potential requires not only higher TE performance (zT) but also scalable, low-cost manufacturing processes. Here, we introduce a transformative ink-based processing approach for scalable manufacturing of high-performance silver selenide-based TE materials and devices. Using a simple, high-throughput ink-mixing and blade coating strategy, our Ag 2 Se-based materials under the optimized composition and processing conditions yield an ultrahigh room-temperature power factor of 2.8 mW m −1 K −2 , over 100% higher than baseline samples and a reproducible figure of merit zT of 1 at room temperature. A thermoelectric generator (TEG) achieves a very competitive power density of 112 mW cm −2 at a 90 °C temperature difference between the hot and cold sides of the device, which is among the highest reported for silver selenide-based TE devices to date. This facile, scalable ink-based processing establishes a practical pathway toward industrial-scale manufacturing and widespread adoption of thermoelectric devices, advancing sustainable energy technologies.

Bappy, Md. Omarsany [University of Notre Dame, IN ↗

Exploring Architectural-Aware Affinity Policies in Modern HPC Runtimes

Modern commodity and High-Performance Computing (HPC) systems are evolving with complex CPU architectures. These architectures now feature higher core and NUMA domain counts and implement features such as hyperthreading. When considering significant differences in hardware configurations, library availability, and hardware-tailored system/software stacks, which could substantially vary from one system to another, performance portability is hard to achieve. Throughout the years, this trend resulted in an increasingly high burden on application developers to fine-tune their workloads for each architecture. This work explores how hardware-dependent aspects such as locality/process/thread affinity affect performance in modern CPU architectures. We focus our study on the Global Memory and Threading (GMT) distributed runtime system as a representative of Partitioned Global Address Space (PGAS) software stacks commonly adopted for productivity. In particular, to appreciate performance implications, we evaluate GMT’s thread affinity policies, and, introduce two new ones which exploit architectural awareness. Finally, we explore alternative NUMA configurations via different process bindings and perform a scalability study on three HPC clusters with varying CPU architectures and NUMA layouts. Our analysis indicates that more complex architectures are more affected by affinity and binding policies and highlights the importance of setting proper runtime configurations to achieve superior performance.

Di Dio Lavore, Ian↗

Additive manufacturing for electrocaloric terpolymer thin films

Current heating, venting, and air conditioning (HVAC) systems have drawbacks of high energy consumption, large CO 2 emissions, and low efficiency. Electrocaloric (EC) cycles present an eco-friendly alternative by converting thermal energy to electrical energy. Defect-free thin films with uniform thickness are required to achieve optimal EC performance. A scalable thin film fabrication process is essential for integrating EC cycles into HVAC systems. This study introduces EC thin films prepared by electrospray (ES) processing, a manufacturing method that deposits EC polymer layer by layer using high voltage. The resulting films show superior thickness control, smoother surfaces, and improved thermal and electrical properties compared with solution casting. In addition, post-annealing at 120°C enhances the thermal and EC performance, with films achieving a temperature change (ΔT) of 3.6°C at 100 MV/m when tested near room temperature. With the potential for future scalability, the ES method offers a promising approach for fabricating EC thin films.

36 MATERIALS SCIENCE↗

Introduction to Special Section: Machine Learning for Image-based Geologic Interpretation

Image-based geological interpretation has been a labor-intensive and time-consuming process because it requires well-trained geoscientists to identify geological structures, features, and textures from various types of images. These images include scanning electron microscopic images, optical microscopic images, optical photos, resistivity images, seismic volumes, remote-sensing images, etc. With fast-evolving machine learning (ML) technology and computing power in recent decades, computers can achieve nearhuman-level to super-human-level performance with scalable high efficiency in the computer vision field. These technological revolutions facilitated image-based geological interpretation in petroleum exploration and production. For example, a fault picking method applied to 3-D seismic volume data using deep learning can achieve superior performance in comparison to conventional auto-picking methods. In addition, under the new normal of low oil prices, the petroleum industry seeks cost-effective strategies such as automating traditionally labor-intensive processes. Nevertheless, the potential of applying ML to geological image interpretation is still facing a few key challenges including data scarcity, data distribution, poor data and/or label quality, data leakage, learning algorithms, model architecture, training methodologies, testing and evaluation metrics, hyper-parameters optimization, model drift, production deployment, and the like.

58 GEOSCIENCES↗

Scalable mechanochemical synthesis of high-quality Prussian blue analogues for high-energy and durable potassium-ion batteries

Prussian blue analogues (PBAs) are recognized as promising cathode materials for potassium-ion batteries (PIBs), particularly the low-cost and high-energy K 2 Mn[Fe(CN) 6 ](KMnF). However, conventional solution-based synthesis inevitably introduces [Fe(CN) 6 ] 4− defects and lattice water while suffering low synthesis efficiency, unfavorable to the improvement of electrochemical performance and scalability. Here, in this work, we report a simple solvent-free mechanochemical strategy for the synthesis of a wide variety of K 2 M[Fe(CN) 6 ] (M = Mn, Mg, Ca, etc.) with negligible defects and water, and it is unprecedented to achieve kilogram-level products of high-quality KMnF within just 10 minutes. The as-prepared KMnF delivers a high energy density of 590 Wh kg −1 at 0.2 C and exhibits an astonishing stability over 10 000 cycles and rate ability up to 50 C in a potassium metal half-cell. Encouragingly, a high-areal-capacity pouch cell with 2.2 mAh cm −2 (16.5 mg cm −2 ) exhibits a capacity retention of 80.7% after 500 cycles. Furthermore, systematic in situ characterization reveals underlying mechanism insights into structure–performance relationships. Specifically, the fully coordinated Mn–N 6 octahedral configuration effectively suppresses Mn 3+ Jahn–Teller distortion, enabling reversible phase transitions under both high-voltage and long-term cycling conditions. In addition, minimal defects provide sufficient redox centers, while the continuous three-dimensional framework facilitates rapid K + diffusion kinetics. This work provides a new opportunity for the ultrafast, universal and scalable synthesis of high-quality PBAs, facilitating the practical application of PIBs while enabling precise structural and compositional design of novel PBAs.

Mechanochemical method↗

ExaTN: Scalable GPU-Accelerated High-Performance Processing of General Tensor Networks at Exascale

We present ExaTN (Exascale Tensor Networks), a scalable GPU-accelerated C++ library which can express and process tensor networks on shared- as well as distributed-memory high-performance computing platforms, including those equipped with GPU accelerators. Specifically, ExaTN provides the ability to build, transform, and numerically evaluate tensor networks with arbitrary graph structures and complexity. It also provides algorithmic primitives for the optimization of tensor factors inside a given tensor network in order to find an extremum of a chosen tensor network functional, which is one of the key numerical procedures in quantum many-body theory and quantum-inspired machine learning. Numerical primitives exposed by ExaTN provide the foundation for composing rather complex tensor network algorithms. We enumerate multiple application domains which can benefit from the capabilities of our library, including condensed matter physics, quantum chemistry, quantum circuit simulations, as well as quantum and classical machine learning, for some of which we provide preliminary demonstrations and performance benchmarks just to emphasize a broad utility of our library.

97 MATHEMATICS AND COMPUTING↗

Two 28-nm front-end ASICs for ultra-fine spatial resolution and precision timing to be 3D integrated with 12 LGADs

The 3DIntSenS Collaboration—a joint effort between SLAC, Fermilab, and LLNL—is developing enabling technologies for next-generation radiation imaging detectors that combine ultra-fine spatial resolution (about 10 µm) with precision timing (<20 ps), while maintaining low power <1 W/cm2 and high data throughput. The approach leverages 3D integration between advanced CMOS readout ASICs and finely pixelated LGAD sensors to achieve the performance and scalability required for large-area, high-rate applications. High-granularity, precision-timing detectors are essential for scientific advances in HEP, NP, BES, and FES, but widespread adoption is limited by the cost and complexity of 3D integration. To close this gap, the collaboration is developing LGAD sensors compatible with 12-inch commercial CMOS processes, enabling cost-effective integration with high-performance ASICs under development. We present two 28 nm CMOS ASIC prototypes, including a low-jitter front end, and in-pixel TDC demonstrating sub-10 ps timing resolution. These advances represent a critical step toward scalable, high-resolution radiation imaging systems for future scientific instrumentation.

England, Troy [Fermilab] (ORCID:0000000154405255)↗

Two 28-nm front-end ASICs for ultra-fine spatial resolution and precision timing to be 3D integrated with 12 LGADs

The 3DIntSenS Collaboration—a joint effort between SLAC, Fermilab, and LLNL—is developing enabling technologies for next-generation radiation imaging detectors that combine ultra-fine spatial resolution (about 10 µm) with precision timing (<20 ps), while maintaining low power <1 W/cm2 and high data throughput. The approach leverages 3D integration between advanced CMOS readout ASICs and finely pixelated LGAD sensors to achieve the performance and scalability required for large-area, high-rate applications. High-granularity, precision-timing detectors are essential for scientific advances in HEP, NP, BES, and FES, but widespread adoption is limited by the cost and complexity of 3D integration. To close this gap, the collaboration is developing LGAD sensors compatible with 12-inch commercial CMOS processes, enabling cost-effective integration with high-performance ASICs under development. We present two 28 nm CMOS ASIC prototypes, including a low-jitter front end, and in-pixel TDC demonstrating sub-10 ps timing resolution. These advances represent a critical step toward scalable, high-resolution radiation imaging systems for future scientific instrumentation.

England, Troy [Fermilab] (ORCID:0000000154405255)↗

2025 Advances in NekRS: Supporting improved performance for nuclear applications

This report presents several 2025 advancements in NekRS, a high-fidelity spectral element CFD code developed at Argonne National Laboratory to support the NEAMS thermal-hydraulics program. The forthcoming v25 release consolidates several of these advances, adding new features for portability across heterogeneous GPU architectures, real-time in situ visualization, improved turbulence modeling, and conjugate heat transfer coupling. Over the past year, NekRS has demonstrated strong scalability and performance on DOE’s leading exascale platforms, including Aurora and Frontier, confirming its readiness for some of the largest and most complex simulations attempted to date. These achievements provide a powerful new platform for high-fidelity data generation, which in turn supports the development and validation of advanced closure models critical for reactor safety and design. Significant algorithmic innovations have also been introduced. A new global runtime h-refinement capability simplifies workflows by reducing mesh preparation burdens and enabling coarse-to-fine restarts. Building on this, a novel multigrid strategy was implemented to accelerate pressure and transport solves at scale, addressing long-standing bottlenecks in exascale CFD. Together, these developments improve both the efficiency and accessibility of high-fidelity simulations for reactor-relevant problems. Collectively, these enhancements represent a major step forward in simulation technology, positioning NekRS as a cornerstone of NEAMS efforts to enable accurate, efficient, and scalable high-fidelity analysis of advanced nuclear systems.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

OpenSNAPI: Toward a Unified API for SmartNICs

The end of Moore’s Law and Dennard Scaling has produced a renaissance in the field of computer architecture. Unable to continue leveraging silicon-level processor improvements to further enhance performance and scalability, system architects have been forced to explore other options. In this new era of heterogeneous architectures and hardware/software codesign, a new class of devices known as “accelerators” has emerged. Independently designed for optimized execution of distinct workloads, these devices have proven critical to the continued advancement of application performance. SmartNICs, accelerator devices integrated with a network controller, have conventionally been utilized to offload low-level networking functionality. However, newer SmartNIC variants, which incorporate a system-on-chip (SoC) with traditional designs, are challenging this precedent. Leveraging significantly augmented resources, these new devices offer increased versatility and the potential to more effectively complement a given architecture’s CPU. In this talk, we introduce the motivation underlying acceleration, explore the fundamentals of SmartNICs, and discuss traditional use cases. We also detail our initial efforts to investigate the feasibility and benefits of SmartNICs as general-purpose accelerators. We present the OpenSNAPI project created to define a uniform application programming interface (API) for this emerging class of devices. Finally, we provide a brief tutorial regarding development of SmartNIC-accelerated applications on Los Alamos National Laboratory’s SmartNIC-enabled platforms.

97 MATHEMATICS AND COMPUTING↗

On the Efficient Evaluation of the Exchange Correlation Potential on Graphics Processing Unit Clusters

The predominance of Kohn–Sham density functional theory (KS-DFT) for the theoretical treatment of large experimentally relevant systems in molecular chemistry and materials science relies primarily on the existence of efficient software implementations which are capable of leveraging the latest advances in modern high-performance computing (HPC). With recent trends in HPC leading toward increasing reliance on heterogeneous accelerator-based architectures such as graphics processing units (GPU), existing code bases must embrace these architectural advances to maintain the high levels of performance that have come to be expected for these methods. In this work, we purpose a three-level parallelism scheme for the distributed numerical integration of the exchange-correlation (XC) potential in the Gaussian basis set discretization of the Kohn–Sham equations on large computing clusters consisting of multiple GPUs per compute node. In addition, we purpose and demonstrate the efficacy of the use of batched kernels, including batched level-3 BLAS operations, in achieving high levels of performance on the GPU. We demonstrate the performance and scalability of the implementation of the purposed method in the NWChemEx software package by comparing to the existing scalable CPU XC integration in NWChem.

97 MATHEMATICS AND COMPUTING↗

28nm front end ASIC and 12” LGADs for 3D integration

The 3DIntSenS Collaboration—a joint effort between SLAC, Fermilab, and LLNL—is developing enabling technologies for next-generation radiation imaging detectors that combine ultra-fine spatial resolution (≈10 μm) with precision timing (<20 ps), while maintaining low power <1 W/cm2 and high data throughput. The approach leverages 3D integration between advanced CMOS readout ASICs and finely pixelated LGAD sensors to achieve the performance and scalability required for large-area, high-rate applications. High-granularity, precision-timing detectors are essential for scientific advances in HEP, NP, BES, and FES, but widespread adoption is limited by the cost and complexity of 3D integration. To close this gap, the collaboration is developing LGAD sensors compatible with 12-inch commercial CMOS processes, enabling cost-effective integration with high-performance ASICs under development. We present the design and results from a 28 nm CMOS ASIC prototype, including a low-jitter front end, and in-pixel TDC demonstrating sub-10 ps timing resolution. We also report on the co-design and characterization of reticle-scale LGAD sensors with 50 μm and 100 μm pixels and introduce the next 10k-pixel ASIC designed for full 3D integration. These advances represent a critical step toward scalable, high-resolution radiation imaging systems for future scientific instrumentation.

England, Troy [Fermilab] (ORCID:0000000154405255)↗

MatRIS: Multi-level Math Library Abstraction for Heterogeneity and Performance Portability using IRIS Runtime

Vendor libraries are tuned for a specific architecture and are not portable to others. Moreover, they lack support for heterogeneity and multi-device orchestration, which is required for efficient use of contemporary HPC and cloud resources. To address these challenges, we introduce MatRIS—a multilevel math library abstraction for scalable and performance-portable sparse/dense BLAS/LAPACK operations using IRIS runtime. The MatRIS-IRIS co-design introduces three levels of abstraction to make the implementation completely architecture agnostic and provide highly productive programming. We demonstrate that MatRIS is portable without any change in source code and can fully utilize multi-device heterogeneous systems by achieving high performance and scalability on Summit, Frontier, and a CADES cloud node equipped with four NVIDIA A100 GPUs and four AMD MI100 GPUs. A detailed performance study is presented in which MatRIS demonstrates multi-device scalability. When compared, MatRIS provides competitive and even better performance than libraries from vendors and other third parties.

Monil, M. A. H.↗

Optimizing High Performance Markov Clustering for Pre-Exascale Architectures

HipMCL is a high-performance distributed memory implementation of the popular Markov Cluster Algorithm (MCL) and can cluster large-scale networks within hours using a few thousand CPU-equipped nodes. It relies on sparse matrix computations and heavily makes use of the sparse matrix-sparse matrix multiplication kernel (SpGEMM). The existing parallel algorithms in HipMCL are not scalable to Exascale architectures, both due to their communication costs dominating the runtime at large concurrencies and also due to their inability to take advantage of accelerators that are increasingly popular. In this work, we systematically remove scalability and performance bottlenecks of HipMCL. We enable GPUs by performing the expensive expansion phase of the MCL algorithm on GPU. Additionally, we propose a CPU-GPU joint distributed SpGEMM algorithm called pipelined Sparse SUMMA and integrate a probabilistic memory requirement estimator that is fast and accurate. Furthermore, we develop a new merging algorithm for the incremental processing of partial results produced by the GPUs, which improves the overlap efficiency and the peak memory usage. We also integrate a recent and faster algorithm for performing SpGEMM on CPUs. We validate our new algorithms and optimizations with extensive evaluations. With the enabling of the GPUs and integration of new algorithms, HipMCL is up to 12.4x faster, being able to cluster a network with 70 million proteins and 68 billion connections just under 15 minutes using 1024 nodes of ORNL's Summit supercomputer.

97 MATHEMATICS AND COMPUTING↗

Neuromorphic Photonics Based on Phase Change Materials

Neuromorphic photonics devices based on phase change materials (PCMs) and silicon photonics technology have emerged as promising solutions for addressing the limitations of traditional spiking neural networks in terms of scalability, response delay, and energy consumption. In this review, we provide a comprehensive analysis of various PCMs used in neuromorphic devices, comparing their optical properties and discussing their applications. We explore materials such as GST (Ge 2 Sb 2 Te 5 ), GeTe-Sb 2 Te 3 , GSST (Ge 2 Sb 2 Se 4 Te 1 ), Sb 2 S 3 /Sb 2 Se 3 , Sc 0.2 Sb 2 Te 3 (SST), and In 2 Se 3 , highlighting their advantages and challenges in terms of erasure power consumption, response rate, material lifetime, and on-chip insertion loss. By investigating the integration of different PCMs with silicon-based optoelectronics, this review aims to identify potential breakthroughs in computational performance and scalability of photonic spiking neural networks. Further research and development are essential to optimize these materials and overcome their limitations, paving the way for more efficient and high-performance photonic neuromorphic devices in artificial intelligence and high-performance computing applications.

36 MATERIALS SCIENCE↗