Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Parallel processing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 415 records · Page 23

Inference-Engine v0.1.0

Given a pre-trained neural network, Inference-Engine performs maps network inputs to outputs by executing the forward pass through the provided network. Although the predominant programming language for machine-learning is Python, most high-performance computing (HPC) applications are written in Fortran, C, or C++. Inference-Engine aims to support HPC programs and is written in Fortran, a language with a large feature set supporting interoperability with C. This software exposes concurrency in a portable way by using standard language features that some modern Fortran compilers can exploit with various optimizations, including offloading computation to a Graphics Processing Unit (GPU). In particular, this software makes extensive use of Fortran's "do concurrent" parallel loop construct, implicitly parallel array statements, and pure procedures that can be invoked inside "do concurrent" blocks. Inference-Engine also supports dynamic choice of inference methods at runtime. Two current options include one method that uses Fortran's "dot_product" intrinsic function inside "do concurrent" blocks and another method that instead uses Fortran' "matmul" array intrinsic function. We plan to investigate automatic compiler offloading of "do concurrent" calculations to GPUs and compile-time substitution of optimized libraries such as the Basic Linear Algebra Library (BLAS) for "matmul" invocations. We also envision the potential for the choice of which method to use could happen at program launch based on in situ performance measurements on any given platform.

Rouson, Damian↗

Practical procedures for sensor quality assessment

Sensors are increasingly deployed for process monitoring and control. These produce on-line measurements at a high frequency, in parallel with low-frequency laboratory measurements. Compared to laboratory practices, sensor data quality assessment and control practices are far less structured at most utilities. This leads to inaccurate sensor data with unknown uncertainty factors.This chapter shows how to establish standard operating procedures (SOPs) to support sensor data quality assessment and control and subsequent maintenance actions by producing relevant sensor metadata. Furthermore, SOPs are provided for the most commonly used wastewater quality sensors, inspired by utility and academic best practices. This chapter builds on definitions provided in Chapter 3 and provides additional definitions specifically related to sensors maintenance. Chapter 6 complements the methods in this chapter, which are based on reference measurements, with data-analytical techniques.

Alferes, Janelcy↗

Digital Assurance for Grid Reliability in the Era of Large Load Growth

The rapid expansion of large electric loads is reshaping the operational and regulatory landscape of the U.S. electric grid. These facilities are reaching new scales of expansion, now exceeding a gigawatt per site, and their highly sensitive, digitally driven behaviors introduce new reliability risks. Recent grid events, including large load losses following routine transmission disturbances, highlight the consequences of limited ride-through capability, inconsistent protection settings, inadequate modeling, and lack of behind-the-meter visibility. Parallels to earlier integration challenges of new grid technologies suggest that the grid’s existing processes, standards, and interconnection frameworks are no longer adequate for emerging large loads. This brief synthesizes lessons from the evolution of inverter-based resource regulation and applies them to large-load integration. It identifies critical gaps in modeling accuracy, interconnection processes, performance standards, and compliance mechanisms. Technical recommendations emphasize advanced monitoring, improved modeling, coordinated communication protocols, modernized substations, and structured behind-the-meter control schemes. Collectively, these measures provide a roadmap to maintain bulk power system reliability while enabling the continued growth of large, electrified digital infrastructure.

24 - POWER TRANSMISSION AND DISTRIBUTION↗

Understanding cold electron impact on parallel-propagating whistler chorus waves via moment-based quasilinear theory

Earth's magnetosphere hosts a wide range of collisionless particle populations that interact through various wave-particle processes. Among these, cold electrons, with energies below 100 eV, often dominate the plasma density but remain poorly characterized due to measurement challenges such as spacecraft charging and photoelectron contamination. Understanding the contribution of these cold populations to wave–particle interaction is of significant interest. Recent kinetic simulations identified a secondary drift-driven instability, in which parallel-propagating whistler-mode chorus waves excite oblique electrostatic whistler waves near the resonance cone and Bernstein-mode turbulence. These secondary modes enable a new channel of energy transfer from the parallel-propagating whistler wave to the cold electrons. In this work, we develop a moment-based quasilinear theory of the secondary instabilities to quantify such energy exchange. Our results show that these secondary instabilities persist for a wide range of parameters and, in many cases, lead to nearly complete damping of the primary wave. Such secondary instability might limit the amplitude of parallel-propagating whistler waves in Earth's magnetosphere and might explain why high-amplitude oblique whistler or electron Bernstein waves are rarely observed simultaneously with high-amplitude field-aligned whistler waves in the inner magnetosphere.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

rustpix

rustpix is a high-performance, open-source Rust library with first-class Python bindings (via PyO3) for processing pixel-detector data in neutron imaging. It targets time-stamping detectors such as Timepix3 (TPX3) at ORNL's Spallation Neutron Source (VENUS beamline), where each detected neutron deposits charge across a cluster of pixels within a very high-rate event stream (96M+ hits/sec). rustpix parses TPX3 event data in parallel using memory-mapped I/O, offers four interchangeable clustering algorithms (ABS adjacency-based search, DBSCAN, graph/union-find connected components, and a parallel grid method), and extracts weighted, super-resolved centroids to produce neutron-event lists. A streaming architecture lets it process files larger than available memory. rustpix is distributed as a pip-installable Python package (with NumPy integration), Rust crates, a command-line tool, and an interactive GUI; it writes HDF5, Apache Arrow, and CSV; and it is designed to extend to TPX4 and other detector types. Released as open-source under the MIT License.

Zhang, Chen [Oak Ridge National Laboratory (ORNL),↗

Process Variability Effects on Tensile Response in Injection Molded, Fluorinated Thermoplastics

The mechanical properties of fluorinated thermoplastics (i.e., tensile strength and elongation) can vary with changes in injection molding processing parameters. Four fluoropolymers are examined: poly(vinylidene fluoride) (PVDF) and random poly(vinylidene fluoride-co-chlorotrifluoroethylene) (PVDF-CTFE) with three CTFE concentrations. Dog bones were manufactured with various cylinder dwell times and mold cooling times to assess the manufacturing sensitivity to the tensile response. Dwell and cooling times increasingly impact mechanical performance as CTFE concentration increases. Specimens exhibit higher tensile strength as a function of injection order. The first injected specimen exhibits the lowest tensile strength and highest elongation in all copolymers. This trend becomes more pronounced among fluoropolymers with higher CTFE concentration and lower weight-averaged molecular weight. Parallel plate rheology was used to obtain the zero-shear viscosity as a function of material type, process, and injection order. We found that in the copolymers, the first injected sample exhibited a lower zero-shear viscosity than the next, which indicates a lower molecular weight in the first injected specimen. This phenomenon was not presented for the PVDF homopolymer. Copolymer mechanical uncertainties are hypothesized to result from the shorter molecular weight chains extruding out of the specimens' sides as a flash due to higher mobility with CTFE segments.

36 MATERIALS SCIENCE↗

Process Variability Effects on Tensile Response in Injection Molded, Fluorinated Thermoplastics

The mechanical properties of fluorinated thermoplastics (i.e., tensile strength and elongation) can vary with changes in injection molding processing parameters. Four fluoropolymers are examined: poly(vinylidene fluoride) (PVDF) and random poly(vinylidene fluoride-co-chlorotrifluoroethylene) (PVDF-CTFE) with three CTFE concentrations. Dog bones were manufactured with various cylinder dwell times and mold cooling times to assess the manufacturing sensitivity to the tensile response. Dwell and cooling times increasingly impact mechanical performance as CTFE concentration increases. Specimens exhibit higher tensile strength as a function of injection order. The first injected specimen exhibits the lowest tensile strength and highest elongation in all copolymers. This trend becomes more pronounced among fluoropolymers with higher CTFE concentration and lower weight-averaged molecular weight. Parallel plate rheology was used to obtain the zero-shear viscosity as a function of material type, process, and injection order. We found that in the copolymers, the first injected sample exhibited a lower zero-shear viscosity than the next, which indicates a lower molecular weight in the first injected specimen. This phenomenon was not present for the PVDF homopolymer. Copolymer mechanical uncertainties are hypothesized to result from the shorter molecular weight chains extruding out of the specimens’ sides as a flash due to higher mobility with CTFE segments.

36 MATERIALS SCIENCE↗

Heterogeneous graphics processing unit for scheduling thread groups for execution on variable width SIMD units

A compute unit configured to execute multiple threads in parallel is presented. The compute unit includes one or more single instruction multiple data (SIMD) units and a fetch and decode logic. The SIMD units have differing numbers of arithmetic logic units (ALUs), such that each SIMD unit can execute a different number of threads. The fetch and decode logic is in communication with each of the SIMD units, and is configured to assign the threads to the SIMD units for execution based on such differing numbers of ALUs.

97 MATHEMATICS AND COMPUTING↗

CORE-BFS: Communication-Optimized REctangular-partitioned BFS Achieving 160.845 TeraTEPS on Frontier Supercomputer

Distributed Breadth-First Search (BFS) is fundamental to many large-scale graph applications, but its performance on parallel systems is often limited by high communication overhead. This paper presents CORE-BFS, an extremely scalable GPU-based BFS implementation that introduces a unique rectangular 2D partitioning-based design for Frontier supercomputer. To further improve performance, we propose four key optimizations: (1) Rectangular 2D-partition specific data formats that use two compressed row and one compressed column status array bitmaps combined with a Double Compressed Sparse Row (DCSR) format per partition, reducing memory footprint and inter-rank traffic; (2) Adaptive frontier & communication strategy that unifies top-down and bottom-up traversal on the rectangular layout, uses lazy synchronization in top-down levels, and switches variants based on frontier size to minimize communication overhead; (3) Frontier-split degree-aware update that maps frontier vertices to thread-centric, wavefront-centric, and block-centric kernels based on their degree to improve GPU utilization and memory coalescing; (4) Row-reduction pipeline that overlaps bottom-up adjacency list processing with row-wise bitmap reduction to hide inter-rank latency. Together, these techniques increase parallelism while reducing memory and communication overhead. On the Graph500 benchmark, CORE - BFS scales up to 9,248 Frontier nodes with scale-42 graphs and reaches 160.845 TTEPS, delivering a 5.42 × speedup over our previous Frontier implementation.

Yang, Haoshen [Rutgers University]↗

NMDQi Nuclear Materials Discovery and Qualification Initiative Conference Overview

The Nuclear Materials Discovery and Qualification Initiative (NMDQi) is designed to accelerate nuclear materials qualification to fulfill the promises of early and advanced reactor technologies as a safe, clean, and low-cost baseload energy. Materials development and qualification in the nuclear industry is by definition challenging due to stringent safety requirements, limited availability of specialized facilities for materials irradiation and testing, and the challenging high-temperature, high-radiation environment. NMDQi will establish tools and capabilities that will greatly accelerate the nuclear fuels and materials development process. These tools will provide both computationally informed insights and high-throughput infrastructure with the goal of completing qualification in a single pass. In this approach, materials must be fabricated with a range of properties of interest so that materials performance can be examined in parallel, rather than with multiple discrete specimens. Advanced manufacturing (AM) techniques, which can produce complex component geometries, microstructures, and compositions, are ideal. The improved process monitoring and control that AM can provide are ideal for ensuring and reducing material variability, which will be particularly important as standards committees pursue regulations for AM components. However, gaining these benefits requires overcoming the substantial barrier of qualifying AM processes and the resulting materials and components for supporting research campaigns and eventually nuclear service.

36 MATERIALS SCIENCE↗

A library of calcium mineral reference spectra recorded by parallel imaging using NEXAFS spectromicroscopy

Calcium minerals are ubiquitous in geology and life chemistry. Understanding the phase and chemical state of calcium minerals is important for numerous processes including materials chemistry, hard tissue biogenesis and geological processes. Photoemission spectroscopies such as near edge X-ray absorption fine structure (NEXAFS) and scanning transmission X-ray microscopy have been instrumental in identifying and characterizing calcium minerals in all these areas. In this work, we have recorded reference spectra for a range of different calcium minerals including a series of calcium carbonates, calcium oxalates and calcium phosphates. While collections of reference spectra for several calcium minerals can be found in the literature, these spectra have been reported in different contexts using a variety of instruments. We, here, report a comprehensive list of references recorded in parallel in a single experiment by imaging an array of calcium minerals using a NEXAFS microscope. We present reference NEXAFS spectra at the calcium L-, carbon K- and oxygen K-edges.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Scale-dependent spatial variabilities of hydrological exchange flows and transit time in a large regulated river

Hydrological exchange flows (HEF) across the river-aquifer interface and the associated residence time of river water in the aquifer have important implications for contaminant plume migration and biogeochemical processes in the river corridor. HEFs and residence time are influenced by both subsurface physical features and hydrologic forcing related to the transport process, which can exhibit complex spatial and temporal variations. In this study, we used a massively parallel subsurface flow model and a particle-tracking model to study the influences of different control factors on spatial variability of HEFs and residence time distributions (RTD) in the Hanford Reach of the Columbia River in Washington State. A total number of 100M particles were randomly injected in time and space and then tracked in a model domain that covers a 51-km 2 area (15.1M model cells). We used hourly river stages and groundwater levels to drive the model to provide dynamic velocity fields for the particle tracking in the simulation period that was longer than 2 years. The groundwater flow simulation and particle-tracking results provide the first comprehensive assessment of the spatial distribution of HEFs and residence time in large complex river corridors. Overall, our results show that the aquifer hydrogeological structure has the strongest correlation with the extent and magnitude of exchange flux. The residence time exhibits complex patterns that are impacted by all the river geomorphologic, hydrodynamic, and hydrogeologic factors and are strongly correlated with the downwelling ratio of exchange flux. The new insights gained through this study can be used to support the development of reduced-order models of HEFs and RTDs for large complex river systems.

54 ENVIRONMENTAL SCIENCES↗

Enhanced quantum state transfer by circumventing quantum chaotic behavior

The ability to realize high-fidelity quantum communication is one of the many facets required to build generic quantum computing devices. In addition to quantum processing, sensing, and storage, transferring the resulting quantum states demands a careful design that finds no parallel in classical communication. Existing experimental demonstrations of quantum information transfer in solid-state quantum systems are largely confined to small chains with few qubits, often relying upon non-generic schemes. Here, by using a superconducting quantum circuit featuring thirty-six tunable qubits, accompanied by general optimization procedures deeply rooted in overcoming quantum chaotic behavior, we demonstrate a scalable protocol for transferring few-particle quantum states in a two-dimensional quantum network. These include single-qubit excitation, two-qubit entangled states, and two excitations for which many-body effects are present. Our approach, combined with the quantum circuit’s versatility, paves the way to short-distance quantum communication for connecting distributed quantum processors or registers, even if hampered by inherent imperfections in actual quantum devices.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

Magnetic island formation and rotation braking induced by low-Z impurity penetration in an EAST plasma

Abstract Recent observations of the successive formations of the 4 / 1 , 3 / 1 , and 2 / 1 magnetic islands as well as the subsequent braking of the 2 / 1 mode during a low- Z impurity penetration process in EAST experiments are well reproduced in our 3 D resistive MHD simulations. The enhanced parallel current perturbation induced by impurity radiation predominately contributes to the tearing mode growth, and the 2 / 1 island rotation is mainly damped by the impurity accumulation as results of the influence from high n modes.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Unprecedented cloud resolution in a GPU-enabled full-physics atmospheric climate simulation on OLCF’s summit supercomputer

Clouds represent a key uncertainty in future climate projection. While explicit cloud resolution remains beyond our computational grasp for global climate, we can incorporate important cloud effects through a computational middle ground called the Multi-scale Modeling Framework (MMF), also known as Super Parameterization. This algorithmic approach embeds high-resolution Cloud Resolving Models (CRMs) to represent moist convective processes within each grid column in a Global Climate Model (GCM). The MMF code requires no parallel data transfers and provides a self-contained target for acceleration. This study investigates the performance of the Energy Exascale Earth System Model-MMF (E3SM-MMF) code on the OLCF Summit supercomputer at an unprecedented scale of simulation. Hundreds of kernels in the roughly 10K lines of code in the E3SM-MMF CRM were ported to GPUs with OpenACC directives. A high-resolution benchmark using 4600 nodes on Summit demonstrates the computational capability of the GPU-enabled E3SM-MMF code in a full physics climate simulation.

58 GEOSCIENCES↗

UPC++ v1.0 Programmer’s Guide, Revision 2020.10.0

UPC++ is a C++11 library that provides Partitioned Global Address Space (PGAS) programming. It is designed for writing parallel programs that run efficiently and scale well on distributed-memory parallel computers. The PGAS model is single program, multiple-data (SPMD), with each separate constituent process having access to local memory as it would in C++. However, PGAS also provides access to a global address space, which is allocated in shared segments that are distributed over the processes. UPC++ provides numerous methods for accessing and using global memory. In UPC++, all operations that access remote memory are explicit, which encourages programmers to be aware of the cost of communication and data movement. Moreover, all remote-memory access operations are by default asynchronous, to enable programmers to write code that scales well even on hundreds of thousands of cores.

97 MATHEMATICS AND COMPUTING↗

UPC++ v1.0 Programmer’s Guide, Revision 2021.9.0

UPC++ is a C++ library that provides Partitioned Global Address Space (PGAS) programming. It is designed for writing parallel programs that run efficiently and scale well on distributed-memory parallel computers. The PGAS model is single program, multiple-data (SPMD), with each separate constituent process having access to local memory as it would in C++. PGAS additionally provides one-sided Remote Memory Access (RMA) to a global address space, which is allocated in shared segments that are distributed over the processes. UPC++ also features Remote Procedure Call (RPC) communication, making it easy to move computation to operate on data that resides on remote processes. In UPC++, all communication operations are explicit, which encourages programmers to be aware of the cost of communication and data movement. Moreover, all communication operations are asynchronous by default, to enable programmers to write code that scales well even on hundreds of thousands of cores.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗