Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “scalable performance”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Message Passing and Shared Address Space Parallelism on an SMP Cluster

Currently, message passing (MP) and shared address space (SAS) are the two leading parallel programming paradigms. MP has been standardized with MPI, and is the more common and mature approach; however, code development can be extremely difficult, especially for irregularly structured computations. SAS offers substantial ease of programming, but may suffer from performance limitations due to poor spatial locality and high protocol overhead. In this paper, we compare the performance of and the programming effort required for six applications under both programming models on a 32-processor PC-SMP cluster, a platform that is becoming increasingly attractive for high-end scientific computing. Our application suite consists of codes that typically do not exhibit scalable performance under shared-memory programming due to their high communication-to-computation ratios and/or complex communication patterns. Results indicate that SAS can achieve about half the parallel efficiency of MPI for most of our applications, while being competitive for the others. A hybrid MPI+SAS strategy shows only a small performance advantage over pure MPI in some cases. Finally, improved implementations of two MPI collective operations on PC-SMP clusters are presented.

Shan, Hongzhang↗

hPIC2: A hardware-accelerated, hybrid particle-in-cell code for dynamic plasma-material interactions

The exascale era of high performance computing promises to bring the field of computational plasma physics ever closer to the goal of accurate multiscale modeling. Such computers will rely on hardware acceleration to offload work to dedicated components, notably general-purpose graphics processing units (GPUs). However, devices from different manufacturers require software to be written with different parallel programming models, greatly increasing the code maintenance burden of applications designed to perform on more than one such device. hPIC2 is a hybrid plasma simulation code developed with the Kokkos performance portability framework to target the architectures that will drive exascale computing for the foreseeable future. As a hybrid simulation code, hPIC2 investigates the simultaneous use of various plasma models on the same domain, at the same time. hPIC2 also optionally couples to RustBCA, which accurately models ion-material interactions using the binary collision approximation (BCA) method. In conclusion, hPIC2 therefore achieves scalable performance on a variety of computing architectures when simulating complex and diverse plasmas, particularly near plasma-material interfaces.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Summer Internship Report: ARA2 Benchmarking

Over the past decade, the RISC-V Instruction Set Architecture (ISA) has emerged as a significant player in both academic and industrial processor design due to its open-source nature, modular extension system, and versatility across domains ranging from microcontrollers to high-performance computing (HPC). One of its most important recent advancements is the RISC-V Vector Extension (RVV), which enables explicit data-level parallelism through vector registers and vectorized instructions. Unlike traditional SIMD (Single Instruction, Multiple Data) architectures that fix vector lengths at design time, RVV uses the concept of VLEN (vector register length) as a hardware-independent parameter and allows software to adapt dynamically to the available vector width. This flexible approach ensures portability across implementations while enabling scalable performance. The ARA2 core is a parameterizable RISC-V vector processor developed at the Integrated Systems Lab at ETH Zürich and the University of Bologna. Designed as a tightly-coupled accelerator to a scalar RISC-V core, ARA2 implements the RVV 1.0 specification and offers tunable architectural parameters such as the number of vector lanes, VLEN, and cache sizes.

97 MATHEMATICS AND COMPUTING↗

Robotic Specialization in Autonomous Robotic Structural Assembly

Robotic in-space assembly of large space structures is a long-term NASA goal to reduce launch costs and enable larger scale missions. Recently, researchers have proposed using discrete lattice building blocks and co-designed robots to build high-performance, scalable primary structure for various on-orbit and surface applications. These robots would locomote on the lattice and work in teams to build and reconfigure building-blocks into functional structure. However, the most reliable and efficient robotic system architecture, characterized by the number of different robotic 'species' and the allocation of functionality between species, is an open question. To address this problem, we decompose the robotic building-block assembly task into functional primitives and, in simulation, study the performance of the the variety of possible resulting architectures. For a set consisting of five process types (move self, move block, move friend, align bock, fasten block), we describe a method of feature space exploration and ranking based on energy and reliability cost functions. The solution space is enumerated, filtered for unique solutions, and evaluated against energy and reliability cost functions for various simulated build sizes. We find that a 2 species system, dividing the five mentioned process types between one unit cell transport robot and one fastening robot, results in the lowest energy cost system, at some cost to reliability. This system enables fastening functionality to occupy the build front while reducing the need for that functional mass to travel back and forth from a feed station. Because the details of a robot design affect the weighting and final allocation of functionality, a sensitivity analysis was conducted to evaluate the effect of changing mass allocations on architecture performance. Future systems with additional functionalities such as repair, inspection, and others may use this process to analyze and determine alternative robot architectures.

Bernus, Borbala↗

High Performance, High Fidelity: A GPU‐Accelerated Doubly‐Periodic Configuration of the Simple Cloud‐Resolving E3SM Atmosphere Model Version 1 (DP‐SCREAMv1)

The development of the Simplified Cloud Resolving Energy Exascale Earth System Atmosphere Model (SCREAMv1) enables global storm-resolving simulations on modern GPU-based supercomputers. However, the high computational cost of SCREAMv1 limits its routine use for process-level studies, creating a need for efficient proxy configurations. This study addresses this gap by introducing DP-SCREAMv1, a doubly periodic cloud-resolving model designed to be fully consistent with SCREAMv1 while enabling high-resolution, long-duration simulations at significantly reduced computational expense by simulating a limited doubly periodic domain rather than the entire globe. Built on a C++/Kokkos architecture, DP-SCREAMv1 achieves exceptional performance scalability on GPU systems and includes a rich library of cases for validation and scientific exploration. In this work, we demonstrate short wall-clock times at SCREAMv1's default resolution and show that DP-SCREAMv1 supports routine execution of large-domain, high-resolution experiments that were previously challenging in practice. Furthermore, we show that DP-SCREAMv1 enables routine execution of “Giga-LES” style simulations and facilitates large-domain, high-resolution simulations that were recently considered burdensome to perform. These results document an efficient, fully consistent process-level configuration for SCREAMv1 (DP-SCREAMv1) and illustrate its use for long-duration and large-domain experiments at cloud-resolving to eddy-permitting resolution.

Environmental sciences↗

A new communication protocol family for a distributed spacecraft control system

In this paper we describe the concepts behind and architecture of a communication protocol family, which was designed to fulfill the communication requirements of ESOC's new distributed spacecraft control system SCOS 2. A distributed spacecraft control system needs a data delivery subsystem to be used for telemetry (TLM) distribution, telecommand (TLC) dispatch and inter-application communication, characterized by the following properties: reliability, so that any operational workstation is guaranteed to receive the data it needs to accomplish its role; efficiency, so that the telemetry distribution, even for missions with high telemetry rates, does not cause a degradation of the overall control system performance; scalability, so that the network is not the bottleneck both in terms of bandwidth and reconfiguration; flexibility, so that it can be efficiently used in many different situations. The new protocol family which satisfies the above requirements is built on top of widely used communication protocols (UDP and TCP), provides reliable point-to-point and broadcast communication (UDP+) and is implemented in C++. Reliability is achieved using a retransmission mechanism based on a sequence numbering scheme. Such a scheme allows to have cost-effective performances compared to the traditional protocols, because retransmission is only triggered by applications which explicitly need reliability. This flexibility enables applications with different profiles to take advantage of the available protocols, so that the best rate between sped and reliability can be achieved case by case.

Baldi, Andrea↗

COLLABORATIVE DEVELOPMENT PROJECTS - PHOTONIC MEMORY CONTROLLER MODULE (P-MCM)

As computational density for high-performance computing and big-data services continues to scale, performance scalability of next generation computing systems is becoming increasingly constrained by limitations in memory access, power dissipation and chip packaging. The processor-memory communication bottleneck, a major challenge in current multicore processors due to limited pin-out and power budget, presents a detrimental scaling barrier to data-intensive computing. A consortium team of small businesses and leading researchers that includes experts from photonics processor-memory architecture, III/V photonic laser design/fabrication, silicon photonics design/fabrication, photonics packaging and assembly, and FPGA-based high-performance memory controller IP development – to collaboratively develop a commercialization path for a Photonic Memory Controller Module (P-MCM).

97 MATHEMATICS AND COMPUTING↗

The PetscSF Scalable Communication Layer

PetscSF, the communication component of the Portable, Extensible Toolkit for Scientific Computation (PETSc), is designed to provide PETSc's communication infrastructure suitable for exascale computers that utilize GPUs and other accelerators. PetscSF provides a simple application programming interface (API) for managing common communication patterns in scientific computations by using a star-forest graph representation. PetscSF supports several implementations based on MPI and NVSHMEM, whose selection is based on the characteristics of the application or the target architecture. An efficient and portable model for network and intra-node communication is essential for implementing large-scale applications. The Message Passing Interface, which has been the de facto standard for distributed memory systems, has developed into a large complex API that does not yet provide high performance on the emerging heterogeneous CPU-GPU-based exascale systems. Here, we discuss the design of PetscSF, how it can overcome some difficulties of working directly with MPI on GPUs, and we demonstrate its performance, scalability, and novel features.

97 MATHEMATICS AND COMPUTING↗

Multiscale characterization of phase change materials for building thermal energy storage applications

Phase change materials (PCMs) store and release large amounts of thermal energy because of their high latent energy storage capacity. However, long-term cyclic stability, supercooling and performance-scalability are some of the major challenges for their use in building thermal energy storage (TES) applications. Here, in this study, we present a comprehensive multiscale characterization of two commercially available organic PCMs, Puretemp 18 and Puretemp 23. At the microscale, differential scanning calorimetry (DSC) was used to characterize phase change temperature, specific heat, and latent heat. At the mesoscale, a heat flow meter apparatus (HFMA), following the ASTM C1784 standard, was employed to measure the phase change temperature, specific heat, and latent heat properties. A comparative analysis of latent heat as a function of temperature was conducted by integrating the DSC and HFMA results. At the macroscale, the thermal performance and cyclic stability of the TES system was evaluated using Puretemp 23. The TES system consisted of a finned tube heat exchanger with a storage volume of 0.0189 m 3 (5 gal), which represents a compact, real-world TES solution suitable for building energy storage. The results showed consistent thermal stability of the PCM over 200 cycles, and the supercooling temperature remained within 0.2 °C, which was not detected in smaller-scale characterization methods. Additionally, the macroscale testing methodology of the PCM revealed that the TES is able to charge and discharge stored latent energy within 2 h under a temperature differential of 16.67 °C measured between the inlet water temperature and the phase transition temperature of the PCM. The proposed multiscale PCM characterization method provides a systematic basis for comparing important thermal storage properties while also investigating the scalability, reliability and integration challenges in large scale TES applications.

Latent heat↗

Pre-exascale accelerated application development: The ORNL Summit experience

High-performance computing (HPC) increasingly relies on heterogeneous architectures to achieve higher performance. In the Oak Ridge Leadership Facility (OLCF), Oak Ridge, TN, USA, this trend continues as its latest supercomputer, Summit, entered production in early 2019. The combination of IBM POWER9 CPU and NVIDIA V100 GPU, along with a fast NVLink2 interconnect and other latest technologies, pushes system performance to a new height and breaks the exascale barrier by certain measures. Due to Summit's powerful GPUs and much higher GPU–CPU ratio, offloading to accelerators becomes a requirement for any application, which intends to effectively use the system. To facilitate navigating a complex landscape of competing heterogeneous architectures, a collection of applications from a wide spectrum of scientific domains is selected for early adoption on Summit. In this article, the experience and lessons learned are summarized, in the hope of providing useful guidance to address new programming challenges, such as scalability, performance portability, and software maintainability, for future application development efforts on heterogeneous HPC systems.

97 MATHEMATICS AND COMPUTING↗

PICSAR-QED: a Monte Carlo module to simulate strong-field quantum electrodynamics in particle-in-cell codes for exascale architectures

Abstract Physical scenarios where the electromagnetic fields are so strong that quantum electrodynamics (QED) plays a substantial role are one of the frontiers of contemporary plasma physics research. Investigating those scenarios requires state-of-the-art particle-in-cell (PIC) codes able to run on top high-performance computing (HPC) machines and, at the same time, able to simulate strong-field QED processes. This work presents the PICSAR-QED library, an open-source, portable implementation of a Monte Carlo module designed to provide modern PIC codes with the capability to simulate such processes, and optimized for HPC. Detailed tests and benchmarks are carried out to validate the physical models in PICSAR-QED, to study how numerical parameters affect such models, and to demonstrate its capability to run on different architectures (CPUs and GPUs). Its integration with WarpX, a state-of-the-art PIC code designed to deliver scalable performance on upcoming exascale supercomputers, is also discussed and validated against results from the existing literature.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

A compute-bound formulation of Galerkin model reduction for linear time-invariant dynamical systems

This work aims to advance computational methods for projection-based reduced-order models (ROMs) of linear time-invariant (LTI) dynamical systems. For such systems, current practice relies on ROM formulations expressing the state as a rank-1 tensor (i.e., a vector), leading to computational kernels that are memory bandwidth bound and, therefore, ill-suited for scalable performance on modern architectures. This weakness can be particularly limiting when tackling many-query studies, where one needs to run a large number of simulations. This work introduces a reformulation, called rank-2 Galerkin, of the Galerkin ROM for LTI dynamical systems which converts the nature of the ROM problem from memory bandwidth to compute bound. We present the details of the formulation and its implementation, and demonstrate its utility through numerical experiments using, as a test case, the simulation of elastic seismic shear waves in an axisymmetric domain. We quantify and analyze performance and scaling results for varying numbers of threads and problem sizes. In conclusion, we present an end-to-end demonstration of using the rank-2 Galerkin ROM for a Monte Carlo sampling study. We show that the rank-2 Galerkin ROM is one order of magnitude more efficient than the rank-1 Galerkin ROM (the current practice) and about 970 times more efficient than the full-order model, while maintaining accuracy in both the mean and statistics of the field.

97 MATHEMATICS AND COMPUTING↗

Revolutionizing Energy Storage: AI, Automation, and Advanced Modeling as Catalysts for Next-Generation Breakthroughs

The Presidential Symposium (PRES) at the 2025 Fall Meeting, hosted by the President’s Office and Energy and Fuels Division, American Chemical Society (ACS) in Washington, DC, brought together a diverse group of chemists, engineers, and materials scientists working in battery materials & systems, automation and artificial intelligence from academia, industry, and national laboratories. The accelerating demand for high-performance, scalable, and sustainable energy storage has catalyzed a paradigm shift in how materials are dis-covered, devices are engineered, and systems are optimized. This Presidential Symposium, entitled “Revolutionizing Energy Storage: AI, Automation, and Advanced Modeling Driving Next-Gen Breakthroughs”, brings together global leaders to unveil transformative strategies anchored in the AAA framework: Artificial Intelligence, Automation, and Advanced Modeling. Artificial Intelligence is redefining the frontiers of energy storage by enabling predictive design, real-time optimization, and intelligent control across diverse chemistries and architectures. Automation is streamlining the synthesis, characterization, and testing of battery materials, dramatically accelerating innovation cycles and unlocking scalable solutions for grid and mobility applications. Advanced Modeling, spanning atomic to system-level scales, provides unprecedented insight into electrochemical dynamics, degradation pathways, and thermal behavior, particularly when coupled with physics-informed machine learning and digital twin technologies. Digital twins, in turn, leverage the AAA framework by integrating real-time data, physics-based models, and AI predictions into dynamic virtual replicas, enabling proactive diagnostics, optimization, and system resilience. Together, these synergistic pillars are not only re-shaping the scientific landscape but also forging a new era of reproducible, data-driven, and resilient energy storage innovation. In conclusion, this symposium marks a pivotal moment in the convergence of computational intelligence and experimental rigor, charting the course for next-generation breakthroughs in lithium-ion, solid-state, and flow battery technologies.

Artificial Intelligence (AI)↗

Synthetic residential load models for smart city energy management simulations

The ability to control tens of thousands of residential electricity customers in a coordinated manner has the potential to enact system-wide electric load changes, such as reduce congestion and peak demand, among other benefits. To quantify the potential benefits of demand-side management and other power system simulation studies (e.g. home energy management, large-scale residential demand response), synthetic load datasets that accurately characterize the system load are required. This study designs a combined top-down and bottom-up approach for modelling individual residential customers and their individual electric assets, each possessing their own characteristics, using time-varying queueing models. The aggregation of all customer loads created by the queueing models represents a known city-sized load curve to be used in simulation studies. The three presented residential queueing load models use only publicly available data. An open-source Python tool to allow researchers to generate residential load data for their studies is also provided. The simulation results presented consider the ComEd region (utility company from Chicago, IL) and demonstrate the characteristics of the three proposed residential queueing load models, the impact of the choice of model parameters, and scalability performance of the Python tool.

24 POWER TRANSMISSION AND DISTRIBUTION↗

NWChem: Past, present, and future

Specialized computational chemistry packages have permanently reshaped the landscape of chemical sciences by providing tools to support and guide the experimental effort and for prediction of chemical and materials properties. In this regard, a special role has been played by electronic structure packages where complex chemical and materials processes can be modeled using first-principle-driven methodologies. Over the last few decades, the rapid development of computing technologies and a tremendous increase in computational power has offered a unique chance to study complex chemical transformations using sophisticated and predictive many-body techniques to describe correlated behavior of electrons in molecular and condensed phase systems at different levels of theory. In enabling these simulations, a critical role has been played by novel parallel algorithms capable of taking advantage of computational resources to address polynomial scaling of electronic structure methods. NWChem was among the first electronic structure codes that focused on delivering scalable performance for electronic structure simulations. Herein, we briefly review the NWChem suite of computational codes including its history, design principles, parallel tools, current capabilities, outreach and outlook.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Self-radiography of imploded shells on OMEGA based on additive-free multi-monochromatic continuum spectral analysis

Radiographs of pure-DT cryogenic imploding shells provide critical validation of progress toward ignition-scalable performance of inertial confinement fusion implosions. Cryogenic implosions on the OMEGA Laser System can be self-radiographed by their own core spectral emission near ≈2 keV. Utilizing the distinct spectral dependences of continuum emissivity and opacity, the projected optical-thickness distribution of imploded shells, i.e., the shell radiograph, can be distinguished from the structure of the core emission distribution in images.Importantly, this can be done without relying on spectral additives (shell dopants), as in previous applications of implosion self-radiography. Furthermore, demonstrations with simulated data show that this technique is remarkably well-suited to cryogenic implosions and can also be applied to self-radiography of imploded room-temperature CH shells at higher spectral energy (hv ≈ 3–5 keV) based on the very similar continuum spectrum of carbon. Experimental demonstration of additive-free self-radiography with warm CH shell implosions on OMEGA will provide an important proof of principle for future applications to cryogenic DT implosions.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Advancements in NbTiN based circuits for Superconducting Digital Logic

Superconducting (SC) electronics have emerged as a promising platform for high-speed, energy-efficient computing and quantum information processing. This work, centered on NbTiN, presents recent advances in material science and fabrication methods leading to significant improvements in performance, scalability and vertical integration. We specifically report on fabrication and characterization of key components, including Josephson junctions (JJs), flux trapping structures and SC interconnects. Together, these efforts represent critical steps towards realizing practical, complex, dense and large-scale SC integrated circuits.

Pokhrel, A. [Imec,Heverlee,Belgium]↗

A Backend-agnostic, Quantum-classical Framework for Simulations of Chemistry in C ++

As quantum computing hardware systems continue to advance, the research and development of performant, scalable, and extensible software architectures, languages, models, and compilers is equally as important to bring this novel coprocessing capability to a diverse group of domain computational scientists. For the field of quantum chemistry, applications and frameworks exist for modeling and simulation tasks that scale on heterogeneous classical architectures, and we envision the need for similar frameworks on heterogeneous quantum-classical platforms. Furthermore, we present the XACC system-level quantum computing framework as a platform for prototyping, developing, and deploying quantum-classical software that specifically targets chemistry applications. We review the fundamental design features in XACC, with special attention to its extensibility and modularity for key quantum programming workflow interfaces and provide an overview of the interfaces most relevant to simulations of chemistry. A series of examples demonstrating some of the state-of-the-art chemistry algorithms currently implemented in XACC are presented, while also illustrating the various APIs that would enable the community to extend, modify, and devise new algorithms and applications in the realm of chemistry.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗