Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “scalable performance”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

Exploring Architectural-Aware Affinity Policies in Modern HPC Runtimes

Modern commodity and High-Performance Computing (HPC) systems are evolving with complex CPU architectures. These architectures now feature higher core and NUMA domain counts and implement features such as hyperthreading. When considering significant differences in hardware configurations, library availability, and hardware-tailored system/software stacks, which could substantially vary from one system to another, performance portability is hard to achieve. Throughout the years, this trend resulted in an increasingly high burden on application developers to fine-tune their workloads for each architecture. This work explores how hardware-dependent aspects such as locality/process/thread affinity affect performance in modern CPU architectures. We focus our study on the Global Memory and Threading (GMT) distributed runtime system as a representative of Partitioned Global Address Space (PGAS) software stacks commonly adopted for productivity. In particular, to appreciate performance implications, we evaluate GMT’s thread affinity policies, and, introduce two new ones which exploit architectural awareness. Finally, we explore alternative NUMA configurations via different process bindings and perform a scalability study on three HPC clusters with varying CPU architectures and NUMA layouts. Our analysis indicates that more complex architectures are more affected by affinity and binding policies and highlights the importance of setting proper runtime configurations to achieve superior performance.

Di Dio Lavore, Ian↗

Additive manufacturing for electrocaloric terpolymer thin films

Current heating, venting, and air conditioning (HVAC) systems have drawbacks of high energy consumption, large CO 2 emissions, and low efficiency. Electrocaloric (EC) cycles present an eco-friendly alternative by converting thermal energy to electrical energy. Defect-free thin films with uniform thickness are required to achieve optimal EC performance. A scalable thin film fabrication process is essential for integrating EC cycles into HVAC systems. This study introduces EC thin films prepared by electrospray (ES) processing, a manufacturing method that deposits EC polymer layer by layer using high voltage. The resulting films show superior thickness control, smoother surfaces, and improved thermal and electrical properties compared with solution casting. In addition, post-annealing at 120°C enhances the thermal and EC performance, with films achieving a temperature change (ΔT) of 3.6°C at 100 MV/m when tested near room temperature. With the potential for future scalability, the ES method offers a promising approach for fabricating EC thin films.

36 MATERIALS SCIENCE↗

Introduction to Special Section: Machine Learning for Image-based Geologic Interpretation

Image-based geological interpretation has been a labor-intensive and time-consuming process because it requires well-trained geoscientists to identify geological structures, features, and textures from various types of images. These images include scanning electron microscopic images, optical microscopic images, optical photos, resistivity images, seismic volumes, remote-sensing images, etc. With fast-evolving machine learning (ML) technology and computing power in recent decades, computers can achieve nearhuman-level to super-human-level performance with scalable high efficiency in the computer vision field. These technological revolutions facilitated image-based geological interpretation in petroleum exploration and production. For example, a fault picking method applied to 3-D seismic volume data using deep learning can achieve superior performance in comparison to conventional auto-picking methods. In addition, under the new normal of low oil prices, the petroleum industry seeks cost-effective strategies such as automating traditionally labor-intensive processes. Nevertheless, the potential of applying ML to geological image interpretation is still facing a few key challenges including data scarcity, data distribution, poor data and/or label quality, data leakage, learning algorithms, model architecture, training methodologies, testing and evaluation metrics, hyper-parameters optimization, model drift, production deployment, and the like.

58 GEOSCIENCES↗

Scalable mechanochemical synthesis of high-quality Prussian blue analogues for high-energy and durable potassium-ion batteries

Prussian blue analogues (PBAs) are recognized as promising cathode materials for potassium-ion batteries (PIBs), particularly the low-cost and high-energy K 2 Mn[Fe(CN) 6 ](KMnF). However, conventional solution-based synthesis inevitably introduces [Fe(CN) 6 ] 4− defects and lattice water while suffering low synthesis efficiency, unfavorable to the improvement of electrochemical performance and scalability. Here, in this work, we report a simple solvent-free mechanochemical strategy for the synthesis of a wide variety of K 2 M[Fe(CN) 6 ] (M = Mn, Mg, Ca, etc.) with negligible defects and water, and it is unprecedented to achieve kilogram-level products of high-quality KMnF within just 10 minutes. The as-prepared KMnF delivers a high energy density of 590 Wh kg −1 at 0.2 C and exhibits an astonishing stability over 10 000 cycles and rate ability up to 50 C in a potassium metal half-cell. Encouragingly, a high-areal-capacity pouch cell with 2.2 mAh cm −2 (16.5 mg cm −2 ) exhibits a capacity retention of 80.7% after 500 cycles. Furthermore, systematic in situ characterization reveals underlying mechanism insights into structure–performance relationships. Specifically, the fully coordinated Mn–N 6 octahedral configuration effectively suppresses Mn 3+ Jahn–Teller distortion, enabling reversible phase transitions under both high-voltage and long-term cycling conditions. In addition, minimal defects provide sufficient redox centers, while the continuous three-dimensional framework facilitates rapid K + diffusion kinetics. This work provides a new opportunity for the ultrafast, universal and scalable synthesis of high-quality PBAs, facilitating the practical application of PIBs while enabling precise structural and compositional design of novel PBAs.

Mechanochemical method↗

ExaTN: Scalable GPU-Accelerated High-Performance Processing of General Tensor Networks at Exascale

We present ExaTN (Exascale Tensor Networks), a scalable GPU-accelerated C++ library which can express and process tensor networks on shared- as well as distributed-memory high-performance computing platforms, including those equipped with GPU accelerators. Specifically, ExaTN provides the ability to build, transform, and numerically evaluate tensor networks with arbitrary graph structures and complexity. It also provides algorithmic primitives for the optimization of tensor factors inside a given tensor network in order to find an extremum of a chosen tensor network functional, which is one of the key numerical procedures in quantum many-body theory and quantum-inspired machine learning. Numerical primitives exposed by ExaTN provide the foundation for composing rather complex tensor network algorithms. We enumerate multiple application domains which can benefit from the capabilities of our library, including condensed matter physics, quantum chemistry, quantum circuit simulations, as well as quantum and classical machine learning, for some of which we provide preliminary demonstrations and performance benchmarks just to emphasize a broad utility of our library.

97 MATHEMATICS AND COMPUTING↗

Two 28-nm front-end ASICs for ultra-fine spatial resolution and precision timing to be 3D integrated with 12 LGADs

The 3DIntSenS Collaboration—a joint effort between SLAC, Fermilab, and LLNL—is developing enabling technologies for next-generation radiation imaging detectors that combine ultra-fine spatial resolution (about 10 µm) with precision timing (<20 ps), while maintaining low power <1 W/cm2 and high data throughput. The approach leverages 3D integration between advanced CMOS readout ASICs and finely pixelated LGAD sensors to achieve the performance and scalability required for large-area, high-rate applications. High-granularity, precision-timing detectors are essential for scientific advances in HEP, NP, BES, and FES, but widespread adoption is limited by the cost and complexity of 3D integration. To close this gap, the collaboration is developing LGAD sensors compatible with 12-inch commercial CMOS processes, enabling cost-effective integration with high-performance ASICs under development. We present two 28 nm CMOS ASIC prototypes, including a low-jitter front end, and in-pixel TDC demonstrating sub-10 ps timing resolution. These advances represent a critical step toward scalable, high-resolution radiation imaging systems for future scientific instrumentation.

England, Troy [Fermilab] (ORCID:0000000154405255)↗

Two 28-nm front-end ASICs for ultra-fine spatial resolution and precision timing to be 3D integrated with 12 LGADs

The 3DIntSenS Collaboration—a joint effort between SLAC, Fermilab, and LLNL—is developing enabling technologies for next-generation radiation imaging detectors that combine ultra-fine spatial resolution (about 10 µm) with precision timing (<20 ps), while maintaining low power <1 W/cm2 and high data throughput. The approach leverages 3D integration between advanced CMOS readout ASICs and finely pixelated LGAD sensors to achieve the performance and scalability required for large-area, high-rate applications. High-granularity, precision-timing detectors are essential for scientific advances in HEP, NP, BES, and FES, but widespread adoption is limited by the cost and complexity of 3D integration. To close this gap, the collaboration is developing LGAD sensors compatible with 12-inch commercial CMOS processes, enabling cost-effective integration with high-performance ASICs under development. We present two 28 nm CMOS ASIC prototypes, including a low-jitter front end, and in-pixel TDC demonstrating sub-10 ps timing resolution. These advances represent a critical step toward scalable, high-resolution radiation imaging systems for future scientific instrumentation.

England, Troy [Fermilab] (ORCID:0000000154405255)↗

2025 Advances in NekRS: Supporting improved performance for nuclear applications

This report presents several 2025 advancements in NekRS, a high-fidelity spectral element CFD code developed at Argonne National Laboratory to support the NEAMS thermal-hydraulics program. The forthcoming v25 release consolidates several of these advances, adding new features for portability across heterogeneous GPU architectures, real-time in situ visualization, improved turbulence modeling, and conjugate heat transfer coupling. Over the past year, NekRS has demonstrated strong scalability and performance on DOE’s leading exascale platforms, including Aurora and Frontier, confirming its readiness for some of the largest and most complex simulations attempted to date. These achievements provide a powerful new platform for high-fidelity data generation, which in turn supports the development and validation of advanced closure models critical for reactor safety and design. Significant algorithmic innovations have also been introduced. A new global runtime h-refinement capability simplifies workflows by reducing mesh preparation burdens and enabling coarse-to-fine restarts. Building on this, a novel multigrid strategy was implemented to accelerate pressure and transport solves at scale, addressing long-standing bottlenecks in exascale CFD. Together, these developments improve both the efficiency and accessibility of high-fidelity simulations for reactor-relevant problems. Collectively, these enhancements represent a major step forward in simulation technology, positioning NekRS as a cornerstone of NEAMS efforts to enable accurate, efficient, and scalable high-fidelity analysis of advanced nuclear systems.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

OpenSNAPI: Toward a Unified API for SmartNICs

The end of Moore’s Law and Dennard Scaling has produced a renaissance in the field of computer architecture. Unable to continue leveraging silicon-level processor improvements to further enhance performance and scalability, system architects have been forced to explore other options. In this new era of heterogeneous architectures and hardware/software codesign, a new class of devices known as “accelerators” has emerged. Independently designed for optimized execution of distinct workloads, these devices have proven critical to the continued advancement of application performance. SmartNICs, accelerator devices integrated with a network controller, have conventionally been utilized to offload low-level networking functionality. However, newer SmartNIC variants, which incorporate a system-on-chip (SoC) with traditional designs, are challenging this precedent. Leveraging significantly augmented resources, these new devices offer increased versatility and the potential to more effectively complement a given architecture’s CPU. In this talk, we introduce the motivation underlying acceleration, explore the fundamentals of SmartNICs, and discuss traditional use cases. We also detail our initial efforts to investigate the feasibility and benefits of SmartNICs as general-purpose accelerators. We present the OpenSNAPI project created to define a uniform application programming interface (API) for this emerging class of devices. Finally, we provide a brief tutorial regarding development of SmartNIC-accelerated applications on Los Alamos National Laboratory’s SmartNIC-enabled platforms.

97 MATHEMATICS AND COMPUTING↗

On the Efficient Evaluation of the Exchange Correlation Potential on Graphics Processing Unit Clusters

The predominance of Kohn–Sham density functional theory (KS-DFT) for the theoretical treatment of large experimentally relevant systems in molecular chemistry and materials science relies primarily on the existence of efficient software implementations which are capable of leveraging the latest advances in modern high-performance computing (HPC). With recent trends in HPC leading toward increasing reliance on heterogeneous accelerator-based architectures such as graphics processing units (GPU), existing code bases must embrace these architectural advances to maintain the high levels of performance that have come to be expected for these methods. In this work, we purpose a three-level parallelism scheme for the distributed numerical integration of the exchange-correlation (XC) potential in the Gaussian basis set discretization of the Kohn–Sham equations on large computing clusters consisting of multiple GPUs per compute node. In addition, we purpose and demonstrate the efficacy of the use of batched kernels, including batched level-3 BLAS operations, in achieving high levels of performance on the GPU. We demonstrate the performance and scalability of the implementation of the purposed method in the NWChemEx software package by comparing to the existing scalable CPU XC integration in NWChem.

97 MATHEMATICS AND COMPUTING↗

Highly Scalable Matching Pursuit Signal Decomposition Algorithm

Matching Pursuit Decomposition (MPD) is a powerful iterative algorithm for signal decomposition and feature extraction. MPD decomposes any signal into linear combinations of its dictionary elements or atoms . A best fit atom from an arbitrarily defined dictionary is determined through cross-correlation. The selected atom is subtracted from the signal and this procedure is repeated on the residual in the subsequent iterations until a stopping criterion is met. The reconstructed signal reveals the waveform structure of the original signal. However, a sufficiently large dictionary is required for an accurate reconstruction; this in return increases the computational burden of the algorithm, thus limiting its applicability and level of adoption. The purpose of this research is to improve the scalability and performance of the classical MPD algorithm. Correlation thresholds were defined to prune insignificant atoms from the dictionary. The Coarse-Fine Grids and Multiple Atom Extraction techniques were proposed to decrease the computational burden of the algorithm. The Coarse-Fine Grids method enabled the approximation and refinement of the parameters for the best fit atom. The ability to extract multiple atoms within a single iteration enhanced the effectiveness and efficiency of each iteration. These improvements were implemented to produce an improved Matching Pursuit Decomposition algorithm entitled MPD++. Disparate signal decomposition applications may require a particular emphasis of accuracy or computational efficiency. The prominence of the key signal features required for the proper signal classification dictates the level of accuracy necessary in the decomposition. The MPD++ algorithm may be easily adapted to accommodate the imposed requirements. Certain feature extraction applications may require rapid signal decomposition. The full potential of MPD++ may be utilized to produce incredible performance gains while extracting only slightly less energy than the standard algorithm. When the utmost accuracy must be achieved, the modified algorithm extracts atoms more conservatively but still exhibits computational gains over classical MPD. The MPD++ algorithm was demonstrated using an over-complete dictionary on real life data. Computational times were reduced by factors of 1.9 and 44 for the emphases of accuracy and performance, respectively. The modified algorithm extracted similar amounts of energy compared to classical MPD. The degree of the improvement in computational time depends on the complexity of the data, the initialization parameters, and the breadth of the dictionary. The results of the research confirm that the three modifications successfully improved the scalability and computational efficiency of the MPD algorithm. Correlation Thresholding decreased the time complexity by reducing the dictionary size. Multiple Atom Extraction also reduced the time complexity by decreasing the number of iterations required for a stopping criterion to be reached. The Course-Fine Grids technique enabled complicated atoms with numerous variable parameters to be effectively represented in the dictionary. Due to the nature of the three proposed modifications, they are capable of being stacked and have cumulative effects on the reduction of the time complexity.

Christensen, Daniel↗

28nm front end ASIC and 12” LGADs for 3D integration

The 3DIntSenS Collaboration—a joint effort between SLAC, Fermilab, and LLNL—is developing enabling technologies for next-generation radiation imaging detectors that combine ultra-fine spatial resolution (≈10 μm) with precision timing (<20 ps), while maintaining low power <1 W/cm2 and high data throughput. The approach leverages 3D integration between advanced CMOS readout ASICs and finely pixelated LGAD sensors to achieve the performance and scalability required for large-area, high-rate applications. High-granularity, precision-timing detectors are essential for scientific advances in HEP, NP, BES, and FES, but widespread adoption is limited by the cost and complexity of 3D integration. To close this gap, the collaboration is developing LGAD sensors compatible with 12-inch commercial CMOS processes, enabling cost-effective integration with high-performance ASICs under development. We present the design and results from a 28 nm CMOS ASIC prototype, including a low-jitter front end, and in-pixel TDC demonstrating sub-10 ps timing resolution. We also report on the co-design and characterization of reticle-scale LGAD sensors with 50 μm and 100 μm pixels and introduce the next 10k-pixel ASIC designed for full 3D integration. These advances represent a critical step toward scalable, high-resolution radiation imaging systems for future scientific instrumentation.

England, Troy [Fermilab] (ORCID:0000000154405255)↗

MatRIS: Multi-level Math Library Abstraction for Heterogeneity and Performance Portability using IRIS Runtime

Vendor libraries are tuned for a specific architecture and are not portable to others. Moreover, they lack support for heterogeneity and multi-device orchestration, which is required for efficient use of contemporary HPC and cloud resources. To address these challenges, we introduce MatRIS—a multilevel math library abstraction for scalable and performance-portable sparse/dense BLAS/LAPACK operations using IRIS runtime. The MatRIS-IRIS co-design introduces three levels of abstraction to make the implementation completely architecture agnostic and provide highly productive programming. We demonstrate that MatRIS is portable without any change in source code and can fully utilize multi-device heterogeneous systems by achieving high performance and scalability on Summit, Frontier, and a CADES cloud node equipped with four NVIDIA A100 GPUs and four AMD MI100 GPUs. A detailed performance study is presented in which MatRIS demonstrates multi-device scalability. When compared, MatRIS provides competitive and even better performance than libraries from vendors and other third parties.

Monil, M. A. H.↗

Optimizing High Performance Markov Clustering for Pre-Exascale Architectures

HipMCL is a high-performance distributed memory implementation of the popular Markov Cluster Algorithm (MCL) and can cluster large-scale networks within hours using a few thousand CPU-equipped nodes. It relies on sparse matrix computations and heavily makes use of the sparse matrix-sparse matrix multiplication kernel (SpGEMM). The existing parallel algorithms in HipMCL are not scalable to Exascale architectures, both due to their communication costs dominating the runtime at large concurrencies and also due to their inability to take advantage of accelerators that are increasingly popular. In this work, we systematically remove scalability and performance bottlenecks of HipMCL. We enable GPUs by performing the expensive expansion phase of the MCL algorithm on GPU. Additionally, we propose a CPU-GPU joint distributed SpGEMM algorithm called pipelined Sparse SUMMA and integrate a probabilistic memory requirement estimator that is fast and accurate. Furthermore, we develop a new merging algorithm for the incremental processing of partial results produced by the GPUs, which improves the overlap efficiency and the peak memory usage. We also integrate a recent and faster algorithm for performing SpGEMM on CPUs. We validate our new algorithms and optimizations with extensive evaluations. With the enabling of the GPUs and integration of new algorithms, HipMCL is up to 12.4x faster, being able to cluster a network with 70 million proteins and 68 billion connections just under 15 minutes using 1024 nodes of ORNL's Summit supercomputer.

97 MATHEMATICS AND COMPUTING↗

An Application-Based Performance Evaluation of NASAs Nebula Cloud Computing Platform

The high performance computing (HPC) community has shown tremendous interest in exploring cloud computing as it promises high potential. In this paper, we examine the feasibility, performance, and scalability of production quality scientific and engineering applications of interest to NASA on NASA's cloud computing platform, called Nebula, hosted at Ames Research Center. This work represents the comprehensive evaluation of Nebula using NUTTCP, HPCC, NPB, I/O, and MPI function benchmarks as well as four applications representative of the NASA HPC workload. Specifically, we compare Nebula performance on some of these benchmarks and applications to that of NASA s Pleiades supercomputer, a traditional HPC system. We also investigate the impact of virtIO and jumbo frames on interconnect performance. Overall results indicate that on Nebula (i) virtIO and jumbo frames improve network bandwidth by a factor of 5x, (ii) there is a significant virtualization layer overhead of about 10% to 25%, (iii) write performance is lower by a factor of 25x, (iv) latency for short MPI messages is very high, and (v) overall performance is 15% to 48% lower than that on Pleiades for NASA HPC applications. We also comment on the usability of the cloud platform.

Saini, Subhash↗

Neuromorphic Photonics Based on Phase Change Materials

Neuromorphic photonics devices based on phase change materials (PCMs) and silicon photonics technology have emerged as promising solutions for addressing the limitations of traditional spiking neural networks in terms of scalability, response delay, and energy consumption. In this review, we provide a comprehensive analysis of various PCMs used in neuromorphic devices, comparing their optical properties and discussing their applications. We explore materials such as GST (Ge 2 Sb 2 Te 5 ), GeTe-Sb 2 Te 3 , GSST (Ge 2 Sb 2 Se 4 Te 1 ), Sb 2 S 3 /Sb 2 Se 3 , Sc 0.2 Sb 2 Te 3 (SST), and In 2 Se 3 , highlighting their advantages and challenges in terms of erasure power consumption, response rate, material lifetime, and on-chip insertion loss. By investigating the integration of different PCMs with silicon-based optoelectronics, this review aims to identify potential breakthroughs in computational performance and scalability of photonic spiking neural networks. Further research and development are essential to optimize these materials and overcome their limitations, paving the way for more efficient and high-performance photonic neuromorphic devices in artificial intelligence and high-performance computing applications.

36 MATERIALS SCIENCE↗

Optimizing Irregular Communication with Neighborhood Collectives and Locality-Aware Parallelism

Irregular communication often limits both the performance and scalability of parallel applications. Typically, applications individually implement irregular communication as point-to-point, and any optimizations are integrated directly into the application. As a result, these optimizations lack portability. It is difficult to optimize point-to-point messages within MPI, as the interface for single messages provides no information on the collection of all communication to be performed. However, the persistent neighbor collective API, released in the MPI 4 standard, provides an interface for portable optimizations of irregular communication within MPI libraries. This paper presents methods for implementing existing optimizations for irregular communication within neighborhood collectives, analyzes the impact of replacing point-to-point communication in existing codebases such as Hypre BoomerAMG with neighborhood collectives, and finally shows up to a 1.38x speedup on sparse matrix-vector multiplication communication within a BoomerAMG solve through the use of our optimized neighbor collectives. Here, the authors analyze three implementations of persistent neighborhood collectives for Alltoallv: an unoptimized wrapper of standard point-to-point communication, and two locality-aware aggregating methods. The second locality-aware implementation exposes an non-standard interface to perform additional optimization, and the authors present the additional 0.07x speedup from the extended interface. All optimizations are available in an open-source codebase, MPI Advance, which sits on top of MPI, allowing for optimizations to be added into existing codebases regardless of the system MPI install.

AMG↗

Design and Performance of Kokkos Staging Space toward Scalable Resilient Application Couplings

With the growing number of applications designed for heterogeneous HPC devices, application programmers and users are finding it challenging to compose scalable workflows as ensembles of these applications, that are portable, performant and resilient. The Kokkos C++ library has been designed to simplify this cumbersome procedure by providing an intra-application uniform programming model and portable performance. However, assembling multiple Kokkos-enabled applications into a complex workflow is still a challenge. Although Kokkos enables a uniform programming model, the inter-application data exchange still remains a challenge from both performance and software development cost perspectives. In order to address this issue, we propose Kokkos data staging memory space, an extension of Kokkos' data abstraction (memory space) for heterogeneous computing systems. This new abstraction allows to express data on a virtual shared-space for multiple Kokkos applications, thus extending Kokkos to support inter-application data exchange to build an efficient application workflow. Additionally, we study the effectiveness of asynchronous data layout conversions for applications requiring different memory access patterns for the shared data. Our preliminary evaluation with a synthetic benchmark indicate the effectiveness of this conversion adapted to three different scenarios representing access frequency and use patterns of the shared data.

97 MATHEMATICS AND COMPUTING↗