Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “scalable performance”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Scaling Ensembles of Data-Intensive Quantum Chemical Calculations for Millions of Molecules

Deep learning models are efficient computational tools that can accelerate the inverse design of molecules with desired functional properties by generating predictions at a fraction of the time required by traditional quantum chemical approaches. To ensure that a model maintains accuracy and transferability across broad regions of the chemical space explored during the inverse design, it must be trained on massively large volumes of simulation data. This requires running large-scale ensemble quantum chemical calculations on high-performance computing (HPC) systems for data collection. However, the efficient execution of such large ensemble calculations and the management of large volumes of output data require tools that can judiciously utilize computational resources and manage metadata overhead on the file system. Therefore, we present a high-performance, scalable, ensemble management framework for performing data-intensive quantum chemical electronic structure calculations for organic molecules. This framework provides abstractions to plug different ab initio, first principles, and first principles-based semi-empirical methods and executes them efficiently at large scale on HPC systems. It dynamically distributes tasks to resources and uses tiered storage for managing large collections of files. We employed this framework to process over ten million organic molecules and generate open-source datasets that provide UV-vis absorption spectra by running time-dependent density-functional tight-binding calculations. It is the largest database containing molecular optical spectra that were simulated with quantum chemical methods in a consistent manner.

Mehta, Kshitij↗

Fast HARDI Uncertainty Quantification and Visualization with Spherical Sampling

In this paper, we study uncertainty quantification and visualization of orientation distribution functions (ODF), which corresponds to the diffusion profile of high angular resolution diffusion imaging (HARDI) data. The shape inclusion probability (SIP) function is the state‐of‐the‐art method for capturing the uncertainty of ODF ensembles. The current method of computing the SIP function with a volumetric basis exhibits high computational and memory costs, which can be a bottleneck to integrating uncertainty into HARDI visualization techniques and tools. We propose a novel spherical sampling framework for faster computation of the SIP function with lower memory usage and increased accuracy. In particular, we propose direct extraction of SIP isosurfaces, which represent confidence intervals indicating spatial uncertainty of HARDI glyphs, by performing spherical sampling of ODFs. Our spherical sampling approach requires much less sampling than the state‐of‐the‐art volume sampling method, thus providing significantly enhanced performance, scalability, and the ability to perform implicit ray tracing. Our experiments demonstrate that the SIP isosurfaces extracted with our spherical sampling approach can achieve up to 8164× speedup, 37282× memory reduction, and 50.2% less SIP isosurface error compared to the classical volume sampling approach. We demonstrate the efficacy of our methods through experiments on synthetic and human‐brain HARDI datasets.

97 MATHEMATICS AND COMPUTING↗

Scalable membrane-less microbial electrolysis cell with multiple compact electrode assemblies for high performance hydrogen production

Bioelectrochemical hydrogen production via microbial electrolysis cells (MECs) is a promising method for sustainable energy production and decarbonization of energy systems. However, the application of MECs is limited by the electrochemical performance, scalability, and the cost associated with expensive materials. Here, in this study, a scalable MEC (500 mL) with novel compact electrode assemblies and high electrode surface area to volume ratio (160 m 2 /m 3 ) was designed and constructed. The use of membranes, precious metal catalyst, and current collectors with high costs was avoided. A high current density at the steady state of 49.5 ± 5.3 A/m 2 was achieved using acetate as the substrate with phosphate buffer under the applied voltage of 1.01 V. The corresponding volumetric current density was 3948 ± 422 A/m 3 . The compact electrode assembly design limited methane production rate to 3.9 ± 0.2 L/L/D, while achieving a hydrogen production rate of 33.7 ± 1.7 L/L/D. With the suppression of microbial hydrogen consumption, the hydrogen production rate was 39.8 ± 1.9 L/L/D, higher by almost one order of magnitude than those of MECs with scaling up attempts. The compact electrode configuration reduced internal resistance to 88.5 ± 4.4 Ω cm 2 . The energy efficiency based on input electricity was 146 ± 7 % to 189 ± 9 % within the applied voltage range of 0.71 to 1.05 V. The results in this study demonstrated successful scaling up of high performance small MECs and offered a new possible approach of scaling up MECs by stacking high-performance subunits, with no trade-offs on electrochemical performance.

08 HYDROGEN↗

Scalable filesystem enumeration and metadata operations

Systems, apparatus, and methods are disclosed for performing scalable operations in a file system, including POSIX-like file systems. Metadata entries in a namespace or directory tree are sharded across multiple file metadata servers. An enumeration operation, such as listing a directory, is parallelized across the multiple file metadata servers, while retaining standard functionality transparently to clients. Other enumeration operations include no-output operations such as changing file attributes or deleting a file, and cumulative operations such as counting total disk space usage. The parallelization is compatible with tree-level parallelization and storage-level parallelization. Disclosed technologies can be applied to other fields requiring scalable enumeration, such as database and network applications.

Grider, Gary A.↗

Scalable augmented enumeration and metadata operations for large filesystems

Systems, apparatus, and methods are disclosed for performing scalable operations in a file system. Metadata entries in a namespace or directory tree are sharded across multiple file metadata servers. An augmented enumeration operation, such as listing a directory, is parallelized across the multiple file metadata servers, transparently to clients. Exemplary augmentation features can include filtering and sorting. Augmentation features can be executed concurrently with enumeration, prior to enumeration, after enumeration, or as a combination of these, and can utilize pre-built index structures or holding structures for intermediate results. Augmented enumeration operations can also include no-output operations such as changing file attributes or deleting a file, and cumulative operations such as counting total disk space usage. The parallelization is compatible with tree-level parallelization and storage-level parallelization. Disclosed technologies can be applied to other fields requiring scalable enumeration, such as database and network applications.

Grider, Gary A.↗

Scalable and Regenerable Fibrous Amine-functionalized Matrix (FAM) sorbent for Efficient Enrichment of Critical Minerals from Coal Wastewaters

The poster presents the latest progress on utilizing a commercially scalable flat sheet sorbent for the effective enrichment of critical minerals from coal wastewater. It highlights the performance, scalability, and potential for industrial applications, addressing key challenges in critical recovery from complex wastewater streams.

critical metals↗

Scalable and Actionable Performance Measures for Traffic Signal Systems using Probe Vehicle Trajectory Data

Scalable and actionable performance measures for traffic signal systems provide opportunities for practitioners to measure and improve the transportation network. Historically, traffic signal improvements have relied on scheduled signal retiming based on limited data collection, or on the public to call and alert engineers of an issue. This inefficient method of improving signal timing led to the creation of automated traffic signal performance measures (ATSPMs). These metrics rely on expensive infrastructure, including detection and communications, which has produced barriers for numerous agencies to fully adopt. Recently, third-party data providers have begun to release vehicle trajectory data, which allows for enhanced signal metrics with no investment in physical equipment. The purpose of this study is to demonstrate the use of these data and summarize the scalability of the created metrics. This work builds on previous efforts to quantify signal performance on nine intersections in Michigan, U.S. Ten signalized corridors in Columbus, Ohio, were chosen to scale a performance assessment using crowdsourced trajectory data. A total of 136 intersections were assessed in 2-h intervals using data from all weekdays in 2017. High-level corridor summary metrics including average percent of vehicles stopping (18%–32%), average delay (9.4–20.5 s), and level of travel time reliability (1.23–2.73) were calculated for each corridor direction. Intersection-level metrics were also introduced, which can be used by practitioners to identify problems, improve signal timings, and prioritize future infrastructure investments.

99 GENERAL AND MISCELLANEOUS↗

Non-dimensional performance and safety parameters for heat pipes

The use of heat pipes in safety-critical systems such as nuclear microreactors dictates the development of generalized, practical, scalable performance and safety parameters. Traditional dimensional metrics, while informative, lack the universality required for comparative analysis across varying designs and operating regimes. Here, this work introduces a comprehensive set of non-dimensional parameters to characterize heat pipe performance and safety, including capillary performance, effective thermal conductivity, response time, exergetic efficiency, allowable temperature gradients, allowable rate of temperature change, priming coefficients, and factor of safety. A reference heat pipe design representative of microreactor applications was analyzed via the developed parameters using both traditional analytical models and Sockeye simulations under transient and steady-state conditions. Sodium, potassium, and water were evaluated as working fluids to demonstrate the applicability of the framework across a broad temperature range. The proposed non-dimensional parameters effectively captured key thermal-hydraulic behaviors and safety concerns, as was demonstrated via Sockeye simulations. This framework supports the development of design optimization strategies, operational protocols, and safety assurance practices for advanced reactor systems and other high-reliability applications.

42 - ENGINEERING↗

Nek5000/RS performance on advanced GPU architectures

The authors explore performance scalability of the open-source thermal-fluids code, NekRS, on the U.S. Department of Energy's leadership computers, Crusher, Frontier, Summit, Perlmutter, and Polaris. Particular attention is given to analyzing performance and time-to-solution at the strong-scale limit for a target efficiency of 80%, which is typical for production runs on the DOE's high-performance computing systems. Several examples of anomalous behavior are also discussed and analyzed.

97 MATHEMATICS AND COMPUTING↗

hPIC2: A hardware-accelerated, hybrid particle-in-cell code for dynamic plasma-material interactions

The exascale era of high performance computing promises to bring the field of computational plasma physics ever closer to the goal of accurate multiscale modeling. Such computers will rely on hardware acceleration to offload work to dedicated components, notably general-purpose graphics processing units (GPUs). However, devices from different manufacturers require software to be written with different parallel programming models, greatly increasing the code maintenance burden of applications designed to perform on more than one such device. hPIC2 is a hybrid plasma simulation code developed with the Kokkos performance portability framework to target the architectures that will drive exascale computing for the foreseeable future. As a hybrid simulation code, hPIC2 investigates the simultaneous use of various plasma models on the same domain, at the same time. hPIC2 also optionally couples to RustBCA, which accurately models ion-material interactions using the binary collision approximation (BCA) method. In conclusion, hPIC2 therefore achieves scalable performance on a variety of computing architectures when simulating complex and diverse plasmas, particularly near plasma-material interfaces.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Summer Internship Report: ARA2 Benchmarking

Over the past decade, the RISC-V Instruction Set Architecture (ISA) has emerged as a significant player in both academic and industrial processor design due to its open-source nature, modular extension system, and versatility across domains ranging from microcontrollers to high-performance computing (HPC). One of its most important recent advancements is the RISC-V Vector Extension (RVV), which enables explicit data-level parallelism through vector registers and vectorized instructions. Unlike traditional SIMD (Single Instruction, Multiple Data) architectures that fix vector lengths at design time, RVV uses the concept of VLEN (vector register length) as a hardware-independent parameter and allows software to adapt dynamically to the available vector width. This flexible approach ensures portability across implementations while enabling scalable performance. The ARA2 core is a parameterizable RISC-V vector processor developed at the Integrated Systems Lab at ETH Zürich and the University of Bologna. Designed as a tightly-coupled accelerator to a scalar RISC-V core, ARA2 implements the RVV 1.0 specification and offers tunable architectural parameters such as the number of vector lanes, VLEN, and cache sizes.

97 MATHEMATICS AND COMPUTING↗

High Performance, High Fidelity: A GPU‐Accelerated Doubly‐Periodic Configuration of the Simple Cloud‐Resolving E3SM Atmosphere Model Version 1 (DP‐SCREAMv1)

The development of the Simplified Cloud Resolving Energy Exascale Earth System Atmosphere Model (SCREAMv1) enables global storm-resolving simulations on modern GPU-based supercomputers. However, the high computational cost of SCREAMv1 limits its routine use for process-level studies, creating a need for efficient proxy configurations. This study addresses this gap by introducing DP-SCREAMv1, a doubly periodic cloud-resolving model designed to be fully consistent with SCREAMv1 while enabling high-resolution, long-duration simulations at significantly reduced computational expense by simulating a limited doubly periodic domain rather than the entire globe. Built on a C++/Kokkos architecture, DP-SCREAMv1 achieves exceptional performance scalability on GPU systems and includes a rich library of cases for validation and scientific exploration. In this work, we demonstrate short wall-clock times at SCREAMv1's default resolution and show that DP-SCREAMv1 supports routine execution of large-domain, high-resolution experiments that were previously challenging in practice. Furthermore, we show that DP-SCREAMv1 enables routine execution of “Giga-LES” style simulations and facilitates large-domain, high-resolution simulations that were recently considered burdensome to perform. These results document an efficient, fully consistent process-level configuration for SCREAMv1 (DP-SCREAMv1) and illustrate its use for long-duration and large-domain experiments at cloud-resolving to eddy-permitting resolution.

Environmental sciences↗

COLLABORATIVE DEVELOPMENT PROJECTS - PHOTONIC MEMORY CONTROLLER MODULE (P-MCM)

As computational density for high-performance computing and big-data services continues to scale, performance scalability of next generation computing systems is becoming increasingly constrained by limitations in memory access, power dissipation and chip packaging. The processor-memory communication bottleneck, a major challenge in current multicore processors due to limited pin-out and power budget, presents a detrimental scaling barrier to data-intensive computing. A consortium team of small businesses and leading researchers that includes experts from photonics processor-memory architecture, III/V photonic laser design/fabrication, silicon photonics design/fabrication, photonics packaging and assembly, and FPGA-based high-performance memory controller IP development – to collaboratively develop a commercialization path for a Photonic Memory Controller Module (P-MCM).

97 MATHEMATICS AND COMPUTING↗

The PetscSF Scalable Communication Layer

PetscSF, the communication component of the Portable, Extensible Toolkit for Scientific Computation (PETSc), is designed to provide PETSc's communication infrastructure suitable for exascale computers that utilize GPUs and other accelerators. PetscSF provides a simple application programming interface (API) for managing common communication patterns in scientific computations by using a star-forest graph representation. PetscSF supports several implementations based on MPI and NVSHMEM, whose selection is based on the characteristics of the application or the target architecture. An efficient and portable model for network and intra-node communication is essential for implementing large-scale applications. The Message Passing Interface, which has been the de facto standard for distributed memory systems, has developed into a large complex API that does not yet provide high performance on the emerging heterogeneous CPU-GPU-based exascale systems. Here, we discuss the design of PetscSF, how it can overcome some difficulties of working directly with MPI on GPUs, and we demonstrate its performance, scalability, and novel features.

97 MATHEMATICS AND COMPUTING↗

Multiscale characterization of phase change materials for building thermal energy storage applications

Phase change materials (PCMs) store and release large amounts of thermal energy because of their high latent energy storage capacity. However, long-term cyclic stability, supercooling and performance-scalability are some of the major challenges for their use in building thermal energy storage (TES) applications. Here, in this study, we present a comprehensive multiscale characterization of two commercially available organic PCMs, Puretemp 18 and Puretemp 23. At the microscale, differential scanning calorimetry (DSC) was used to characterize phase change temperature, specific heat, and latent heat. At the mesoscale, a heat flow meter apparatus (HFMA), following the ASTM C1784 standard, was employed to measure the phase change temperature, specific heat, and latent heat properties. A comparative analysis of latent heat as a function of temperature was conducted by integrating the DSC and HFMA results. At the macroscale, the thermal performance and cyclic stability of the TES system was evaluated using Puretemp 23. The TES system consisted of a finned tube heat exchanger with a storage volume of 0.0189 m 3 (5 gal), which represents a compact, real-world TES solution suitable for building energy storage. The results showed consistent thermal stability of the PCM over 200 cycles, and the supercooling temperature remained within 0.2 °C, which was not detected in smaller-scale characterization methods. Additionally, the macroscale testing methodology of the PCM revealed that the TES is able to charge and discharge stored latent energy within 2 h under a temperature differential of 16.67 °C measured between the inlet water temperature and the phase transition temperature of the PCM. The proposed multiscale PCM characterization method provides a systematic basis for comparing important thermal storage properties while also investigating the scalability, reliability and integration challenges in large scale TES applications.

Latent heat↗

Pre-exascale accelerated application development: The ORNL Summit experience

High-performance computing (HPC) increasingly relies on heterogeneous architectures to achieve higher performance. In the Oak Ridge Leadership Facility (OLCF), Oak Ridge, TN, USA, this trend continues as its latest supercomputer, Summit, entered production in early 2019. The combination of IBM POWER9 CPU and NVIDIA V100 GPU, along with a fast NVLink2 interconnect and other latest technologies, pushes system performance to a new height and breaks the exascale barrier by certain measures. Due to Summit's powerful GPUs and much higher GPU–CPU ratio, offloading to accelerators becomes a requirement for any application, which intends to effectively use the system. To facilitate navigating a complex landscape of competing heterogeneous architectures, a collection of applications from a wide spectrum of scientific domains is selected for early adoption on Summit. In this article, the experience and lessons learned are summarized, in the hope of providing useful guidance to address new programming challenges, such as scalability, performance portability, and software maintainability, for future application development efforts on heterogeneous HPC systems.

97 MATHEMATICS AND COMPUTING↗

PICSAR-QED: a Monte Carlo module to simulate strong-field quantum electrodynamics in particle-in-cell codes for exascale architectures

Abstract Physical scenarios where the electromagnetic fields are so strong that quantum electrodynamics (QED) plays a substantial role are one of the frontiers of contemporary plasma physics research. Investigating those scenarios requires state-of-the-art particle-in-cell (PIC) codes able to run on top high-performance computing (HPC) machines and, at the same time, able to simulate strong-field QED processes. This work presents the PICSAR-QED library, an open-source, portable implementation of a Monte Carlo module designed to provide modern PIC codes with the capability to simulate such processes, and optimized for HPC. Detailed tests and benchmarks are carried out to validate the physical models in PICSAR-QED, to study how numerical parameters affect such models, and to demonstrate its capability to run on different architectures (CPUs and GPUs). Its integration with WarpX, a state-of-the-art PIC code designed to deliver scalable performance on upcoming exascale supercomputers, is also discussed and validated against results from the existing literature.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

A compute-bound formulation of Galerkin model reduction for linear time-invariant dynamical systems

This work aims to advance computational methods for projection-based reduced-order models (ROMs) of linear time-invariant (LTI) dynamical systems. For such systems, current practice relies on ROM formulations expressing the state as a rank-1 tensor (i.e., a vector), leading to computational kernels that are memory bandwidth bound and, therefore, ill-suited for scalable performance on modern architectures. This weakness can be particularly limiting when tackling many-query studies, where one needs to run a large number of simulations. This work introduces a reformulation, called rank-2 Galerkin, of the Galerkin ROM for LTI dynamical systems which converts the nature of the ROM problem from memory bandwidth to compute bound. We present the details of the formulation and its implementation, and demonstrate its utility through numerical experiments using, as a test case, the simulation of elastic seismic shear waves in an axisymmetric domain. We quantify and analyze performance and scaling results for varying numbers of threads and problem sizes. In conclusion, we present an end-to-end demonstration of using the rank-2 Galerkin ROM for a Monte Carlo sampling study. We show that the rank-2 Galerkin ROM is one order of magnitude more efficient than the rank-1 Galerkin ROM (the current practice) and about 970 times more efficient than the full-order model, while maintaining accuracy in both the mean and statistics of the field.

97 MATHEMATICS AND COMPUTING↗