Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Parallel processing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 361 records · Page 20

Pervasive shifts in forest dynamics in a changing world

Forest dynamics arise from the interplay of environmental drivers and disturbances with the demographic processes of recruitment, growth, and mortality, subsequently driving biomass and species composition. However, forest disturbances and subsequent recovery are shifting with global changes in climate and land use, altering these dynamics. Changes in environmental drivers, land use, and disturbance regimes are forcing forests toward younger, shorter stands. Rising carbon dioxide, acclimation, adaptation, and migration can influence these impacts. Recent developments in Earth system models support increasingly realistic simulations of vegetation dynamics. In parallel, emerging remote sensing datasets promise qualitatively new and more abundant data on the underlying processes and consequences for vegetation structure. In combination, these advances hold promise for improving the scientific understanding of changes in vegetation demographics and disturbances.

54 ENVIRONMENTAL SCIENCES↗

Markov Chain Monte Carlo Predictions of Neutron-rich Lanthanide Properties as a Probe of r -process Dynamics

Lanthanide element signatures are key to understanding many astrophysical observables, from merger kilonova light curves to stellar and solar abundances. To learn about the lanthanide element synthesis that enriched our solar system, we apply the statistical method of Markov Chain Monte Carlo to examine the nuclear masses capable of forming the r-process rare-earth abundance peak. We describe the physical constraints we implement with this statistical approach and demonstrate the use of the parallel chains method to explore the multidimensional parameter space. We apply our procedure to three moderately neutron-rich astrophysical outflows with distinct types of r-process dynamics. We show that the mass solutions found are dependent on outflow conditions and are related to the r-process path. We describe in detail the mechanism behind peak formation in each case. We then compare our mass predictions for neutron-rich neodymium and samarium isotopes to the latest experimental data from the CPT at CARIBU. Here, we find our mass predictions given outflows that undergo an extended (n,γ)⇌(γ,n) equilibrium to be those most compatible with both observational solar abundances and neutron-rich mass measurements.

79 ASTRONOMY AND ASTROPHYSICS↗

HBMax: Optimizing Memory Efficiency for Parallel Influence Maximization on Multicore Architectures

The goal of influence maximization is to select k most-influential vertices or seeds in a network, where influence is defined by a given diffusion process. The problem has a number of important applications such as viral marketing, information spread, and epidemic control. Although computing optimal seed set is NP-Hard, due to the submodular nature of the problem efficient approximation algorithms exist. However, even state-of-the-art parallel implementations are limited by a sampling step that incurs large memory footprints. This in turn limits the problem size reach and approximation quality. In this work, we study the memory footprint of the sampling process collecting reverse reachability information in the IMM algorithm over large real-world social networks. We present an adaptive and memory-efficient optimization approach for a state-of-the-art multi-threaded parallel influence maximization algorithm. Our approach,HuffMax, uses a portion of the reverse reachable (RR) sets collected by the algorithm to learn the characteristics of the graph. Then, it compresses the intermediate reverse reachability information with Huffman coding, and queries directly on the compressed data to preserve the memory savings obtained through compression. We also propose an efficient sampling strategy based on the distribution of RR sets, which can further reduce the computation time for typical social networks with long-tail distributions. Considering a NUMA architecture, we scale up our solution on 128-core CPUs and reduce the memory footprint by up to 45.7% with negligible time overhead (or even faster) and without perceivable loss of accuracy.

Chen, Xinyu↗

Electrochemical Reduction of Flue Gas Carbon Dioxide to Commercially Viable C2-C4 Products (Final Report)

This is the final scientific/technical report for a DOE project focused on the electrochemical conversion of CO 2 in non-aqueous solvents to novel products. Electrochemical reduction of CO 2 provides an attractive route to produce valuable fuels and chemicals that can simultaneously lower greenhouse gas emissions when powered by renewable electricity. While recent technological advances have shown the feasibility of industrial CO 2 electroreduction, many challenges remain to improve this technology and expand the list of economically viable products. The vast majority of electrochemical CO 2 reduction research has been conducted in aqueous media under neutral to alkaline conditions, leading to commonly reported products including carbon monoxide, formic acid, methane, methanol, ethylene, acetic acid, and ethanol. In comparison, non-aqueous media for CO 2 reduction has been underexplored but represents a possible avenue to yield new products and improved operating conditions. The aim of the project was to convert waste CO 2 in the form of flue gas to a multicarbon C2 - C4 chemical product in a reactor designed to achieve economically competitive values of current density and selectivity. The project strived to advance the technology readiness of an electrochemical CO 2 reduction process in alcohol solvents from the proof-of-concept stage to a device capable of meeting performance metrics for commercial viability. In the initial plan, the University of Louisville researchers were to focus on investigating the electrochemical process and improving the faradaic efficiency for novel C2 – C4 species, while also working on a parallel effort to build a practical electrolysis reactor to markedly increase the CO 2 reduction current density. The reactor development effort also aimed to engineer a dual-electrolyte feed strategy with non-aqueous catholyte and aqueous anolyte to promote water oxidation as the coupling anodic half-reaction to enable a sustainable and economical overall process. At the outset, the University of North Dakota was to investigate the feasibility of operating directly from coal-derived flue gas without separate capture and purification. The research team sought to determine impurity effects and test mitigation strategies, as well as engineer the gaseous feed system for high reactor tolerance to lower CO 2 concentration. In the last half year of the project, the focus was planned to shift to integrating the advances in the catalysis, electrochemical conditions, reactor design, and flue gas compatibility into a fully functional device and improve it for maximum current density and faradaic efficiency for C2 – C4 species. Knowledge of the full system components, constraints, and maximum performance was then to be used as the basis for a thorough technoeconomic analysis (TEA) and life cycle analysis (LCA) at the end of the project.

01 COAL, LIGNITE, AND PEAT↗

Dynamics of rapidly spinning blob-filaments: Fluid theory with a parallel kinetic extension

Blob-filaments (or simply “blobs”) are coherent structures formed by turbulence and sustained by nonlinear processes in the edge and scrape-off layer (SOL) of tokamaks and other magnetically confined plasmas. The dynamics of these blob-filaments, in particular, their radial motion, can influence the scrape-off layer width and plasma interactions with both the divertor target and with the main chamber walls. Motivated by recent results from the XGC1 gyrokinetic simulation code reported on elsewhere [J. Cheng et al., Nucl. Fusion 63, 086015 (2023)], a theory of rapidly spinning blob-filaments has been developed for this work. The theory treats blob-filaments in the closed flux surface region or the region that is disconnected from sheaths in the SOL. It extends previous work by treating blob spin, arising from partially or fully adiabatic electrons, as the leading-order effect and retaining inertial (ion charge polarization) physics in next order. Spin helps to maintain blob coherency and affects the blob's propagation speed. Dipole charge polarization, treated perturbatively, gives rise to blob-filaments with relatively slow radial velocity, comparable to that observed in the simulations. The theory also treats the interaction of rapidly spinning blob-filaments with a zonal flow layer. It is shown analytically that the flow layer can act like a transport barrier for these structures. Finally, parallel electron kinetic effects are incorporated into the theory. Various asymptotic parameter regimes are discussed, and asymptotic expressions for the radial and poloidal motion of the blob-filaments are obtained.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

High-count rate effects in event processing for XRISM/ Resolve X-ray microcalorimeter: I. Ground test

The spectroscopic performance of an X-ray microcalorimeter is compromised at high count rates. We utilize the Resolve X-ray microcalorimeter onboard the XRISM satellite to examine the effects observed during high-count rate measurements and propose modeling approaches to mitigate them. We specifically address the following instrumental effects that impact performance: CPU limit, pile-up, and untriggered electrical cross-talk. Experimental data at high count rates were acquired during ground testing using the flight model instrument and a calibration X-ray source. In the experiment, data processing not limited by the performance of the onboard CPU was run in parallel, which cannot be done in orbit. This makes it possible to access the data degradation caused by limited CPU performance. We use these data to develop models that allow for a more accurate estimation of the aforementioned effects. To illustrate the application of these models in observation planning, we present a simulated observation of GX 13+1. Understanding and addressing these issues is crucial to enhancing the reliability and precision of X-ray spectroscopy in situations characterized by elevated count rates.

47 OTHER INSTRUMENTATION↗

The LISE package: solvers for static and time-dependent superfluid local density approximation equations in three dimensions

Nuclear implementation of the density functional theory (DFT) is at present the only microscopic framework applicable to the whole nuclear landscape. The extension of DFT to superfluid systems in the spirit of the Kohn-Sham approach, the superfluid local density approximation (SLDA) and its extension to time-dependent situations, time-dependent superfluid local density ap- proximation (TDSLDA), have been extensively used to describe various static and dynamical problems in nuclear physics, neutron star crust, and cold atom systems. In this paper, we present the codes that solve the static and time-dependent SLDA equations in three-dimensional coordinate space without any symmetry restriction. These codes are fully parallelized with the message passing interface (MPI) library and take advantage of graphic processing units (GPU) for accelerating execution. The dynamic codes have checkpoint/restart capabilities and for initial conditions one can use any generalized Slater determinant type of wave function. The code can describe a large number of physical problems: nuclear fission, collisions of heavy ions, the interaction of quantized vor- tices with nuclei in the nuclear star crust, excitation of superfluid fermion systems by time dependent external fields, quantum shock waves, domain wall generation and propagation, the dynamics of the Anderson-Bogoliubov-Higgs mode, dynamics of fragmented condensates, vortex rings dynamics, generation and dynamics of quantized vortices, their crossing and recombinations and the incipient phases of quantum turbulence.

Jin, Shi↗

Multi-Source Machine Learning and Thermoplastics Enhanced Aerostructure Manufacturing (mTEAM)

RTX Technology Research Center (RTRC), together with Collins Aerospace (Collins) and Oak Ridge National Laboratory (ORNL) has developed an Artificial Intelligence (AI) / Machine Learning (ML) guided solution to advance the manufacturing and assembly of high performance and lightweight thermoplastic composite (TPC) aerospace products. The solution aims to lower risk, cost and lead time for induction heating based welding and consolidation processes for TPC structure. The cost and lead time of part and material specific process development for induction welding (IW) and induction consolidation will be reduced by replacing traditional empirical methods with optimization methods that merge AI/ML and physics-based process simulations and process experiments with sensing and controls. TPC-IW process development is empirical in nature, and uncertainties in material & process behavior exist near & far from the induction coil. Physics-based simulations can be leveraged directly for process optimization but can be too computationally expensive to run in high fidelity and real time to do robust process optimization. The key impact of successful TPC induction consolidation and welding is cost & lead time reduction for part & material specific consolidation and welding recipes. This is an enabler for more rapid deployment of TPC structures via joining assembly, which can reduce energy & cost intensive usage of autoclaves & ovens. The solution aimed to advance the U.S. Department of Energy’s interests in using thermoplastics and automation in composite manufacturing for improvement of products for existing markets via increased production speeds, reduced costs, and lowered use of energy. Welded TPC structures can offer significant weight & energy savings for high-value commercial aerospace & industrial applications compared to metal & thermoset composite structures assembled by mechanical fastening and/or adhesive bonding. The project was organized into two Budget Periods. Budget Period 1 (BP1) was 15 months and its goal was to perform ML process optimization framework development & deployment on lab-coupon aerostructure components. A Go/No-Go Review was performed at the end of BP1 to verify fulfilment of key tasks & milestones to justify a Go Decision to move into the next Budget Period. Budget Period 2 (BP2) was 12 months and its goal was the deployment of the ML framework for ML process optimization of pilot industrial scale aerostructure components. The overall project aim was to develop & demonstrate ML-enhanced modeling framework that learns process-property mapping from multiple data sources at different fidelities. During BP1, the team accomplished key tasks & milestones to demonstrate the concept of multi-source ML for TPC aerostructure consolidation and assembly. First, the team completed documentation of induction based TPC heating requirements including baseline metrics to compare measured results against. Next the team completed demonstration of data generation from physics-based simulations for ML surrogate model generation and demonstrated the integration of physics-based simulation data into multi-source AI/ML algorithms. In parallel, the team established the lab-coupon scale induction welding system and completed a process to label and reduce generated data from physics-based simulation and experiments for ML surrogate models to enable multi-source ML model training & testing. To complete BP1, the team integrated physics-based simulation data and experimental data into multi-source ML algorithms. This was based on the team completing ML deployment of the induction welding on a lab system at RTRC and AI/ML deployment on existing induction welding line at Collins. ORNL visited both Collins and RTRC sites to witness the TPC induction welding process. Then, ORNL designed and constructed a new version of their vision-based sensing system better adapted to acquire process signals of the TPC induction welding process for process anomaly and defect detection. In BP2, the team accomplished key tasks & milestones to scale up multi-source ML for TPC aerostructure consolidation and assembly from the lab-coupon scale to the pilot-industrial scale. In BP2, the team demonstrated real time anomaly & defect detection via experiments performed by ORNL & RTRC. The team completed ML-optimization heating trials for TPC induction consolidation at Collins, and the team confirmed pilot industrial scale experimental data from Collins was compatible with the developed ML pipeline from RTRC. The team completed sub-element scale ML process optimization demonstration at RTRC, where the team leveraged RTRC’s robotic TPC welding setup to de-risk the ML process optimization by performing ML analysis of recorded temperatures to account for complex part features. Then, the team applied its ML-derived control strategies and ML process optimization framework at Collins to the pilot-industrial scale on a demo skin-stiffener part representative of a nacelle aerostructure fan cowl section. The key innovation is the AI/ML framework enabling effective process development of high performance, lightweight, energy efficient TPCs for composite aircraft structures.

36 MATERIALS SCIENCE↗

Parallel Programming in MCNP6

Monte Carlo N-Particle (MCNP)1 is a general-purpose Monte Carlo particle transport code developed by Los Alamos National Laboratory (LANL). To efficiently handle long simulations, MCNP version 6 (MCNP6) supports parallel execution using two primary programming models: • Shared-memory task-based threading using OpenMP (Open Multi-Processing), and • Distributed-memory calculations using MPI (Message Passing Interface). The OpenMP and MPI programming models enable MCNP6 to scale from desktop systems to high-performance computing (HPC) clusters, allowing users to run MCNP in one of three parallel modes: • OpenMP-only, • MPI-only, and • Hybrid (MPI + OpenMP). The choice of parallelization mode depends on the underlying computer architecture and the characteristics of the simulation problem.

97 MATHEMATICS AND COMPUTING↗

Parallel quantum annealing

Quantum annealers of D-Wave Systems, Inc., offer an efficient way to compute high quality solutions of NP-hard problems. This is done by mapping a problem onto the physical qubits of the quantum chip, from which a solution is obtained after quantum annealing. However, since the connectivity of the physical qubits on the chip is limited, a minor embedding of the problem structure onto the chip is required. In this process, and especially for smaller problems, many qubits will stay unused. We propose a novel method, called parallel quantum annealing, to make better use of available qubits, wherein either the same or several independent problems are solved in the same annealing cycle of a quantum annealer, assuming enough physical qubits are available to embed more than one problem. Although the individual solution quality may be slightly decreased when solving several problems in parallel (as opposed to solving each problem separately), we demonstrate that our method may give dramatic speed-ups in terms of the Time-To-Solution (TTS) metric for solving instances of the Maximum Clique problem when compared to solving each problem sequentially on the quantum annealer. Additionally, we show that solving a single Maximum Clique problem using parallel quantum annealing reduces the TTS significantly.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Virtual Time III, Part 2: Combining Conservative and Optimistic Synchronization

This is Part 2 of a trio of works intended to provide a unifying framework in which conservative and optimistic synchronization for parallel discrete event simulations can be freely and transparently combined in the same logical process on an event-by-event basis. Here, in this article, we continue the outline of an approach called Unified Virtual Time (UVT) that was introduced in Part 1, showing in detail via two extended examples how conservative synchronization can be refactored and combined with optimistic synchronization in the UVT framework. We describe UVT versions of both a basic time windowing algorithm called Unified Simple Time Windows and a refactored version of the Chandy-Misra-Bryant Null Message algorithm called Unified CMB.

97 MATHEMATICS AND COMPUTING↗

Glass-bonded ceramic waste forms for immobilization of radioiodine from caustic scrubber wastes

Glass-bonded sodalite composite waste forms have been developed for the immobilization of liquid radioactive wastes resulting from off-gas treatment during aqueous reprocessing of used nuclear fuel, with a particular focus on 129I. The proposed composite waste form is comprised of aluminosilicate ceramic phases containing volatile radionuclides bonded with a glassy matrix. In this work, a suite of ten candidate low-temperature glass binders (ZnO-Bi2O3-based glasses and a Na2O-B2O3-SiO2 glass) were examined. Six glasses were mixed with caustic scrubber waste simulant previously converted into a sodalite-rich material (to provide glass fractions of 10 and 20 wt.%), uniaxially pressed into pellets, and sintered at 350 °C or 550 °C for 8 h in air. Iodine retention after heat treatment was assessed by neutron activation analysis, showing retention of 67-100 % of expected iodine. The aqueous durabilities of the resulting materials were then determined, following the ASTM C1308 standard test, showing iodine releases of 1 to 23 g m-2 after 4 d. The cumulative iodine release for the best performing system (a zinc-bismuth-borate glass binder) was <1 g m-2, and its iodine retention from processing was 67 %. The iodine releases compared favorably with other waste forms. In parallel, this best-performing composition was also consolidated via hot isostatic pressing (HIP) in a stainless-steel canister at 550 °C for 2 h under 100 MPa pressure. The HIPed sample was produced at the ~20 g scale and showed improved densification and minimal reaction with the canister.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Potential virus-mediated nitrogen cycling in oxygen-depleted oceanic waters

Viruses play an important role in the ecology and biogeochemistry of marine ecosystems. Beyond mortality and gene transfer, viruses can reprogram microbial metabolism during infection by expressing auxiliary metabolic genes (AMGs) involved in photosynthesis, central carbon metabolism, and nutrient cycling. While previous studies have focused on AMG diversity in the sunlit and dark ocean, less is known about the role of viruses in shaping metabolic networks along redox gradients associated with marine oxygen minimum zones (OMZs). Here, we analyzed relatively quantitative viral metagenomic datasets that profiled the oxygen gradient across Eastern Tropical South Pacific (ETSP) OMZ waters, assessing whether OMZ viruses might impact nitrogen (N) cycling via AMGs. Identified viral genomes encoded six N-cycle AMGs associated with denitrification, nitrification, assimilatory nitrate reduction, and nitrite transport. The majority of these AMGs (80%) were identified in T4-like Myoviridae phages, predicted to infect Cyanobacteria and Proteobacteria, or in unclassified archaeal viruses predicted to infect Thaumarchaeota. Four AMGs were exclusive to anoxic waters and had distributions that paralleled homologous microbial genes. Together, these findings suggest viruses modulate N-cycling processes within the ETSP OMZ and may contribute to nitrogen loss throughout the global oceans thus providing a baseline for their inclusion in the ecosystem and geochemical models.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Toward active disruption avoidance via real-time estimation of the safe operating region and disruption proximity in tokamaks

This paper describes a real-time capable algorithm for identifying the safe operating region around a tokamak operating point. The region is defined by a convex set of linear constraints, from which the distance of a point from a disruptive boundary can be calculated. The disruptivity of points is calculated from an empirical machine learning predictor that generates the likelihood of disruption. While the likelihood generated by such empirical models can be compared to a threshold to trigger a disruption mitigation system, the safe operating region calculation enables active optimization of the operating point to maintain a safe margin from disruptive boundaries. The proposed algorithm is tested using a random forest disruption predictor fit on data from DIII-D. The safe operating region identification algorithm is applied to historical data from DIII-D showing the evolution of disruptive boundaries and the potential impact of optimization of the operating point. Real-time relevant execution times are made possible by parallelizing many of the calculation steps and implementing the algorithm on a graphics processing unit. Lastly, a real-time capable algorithm for optimizing the target operating point within the identified constraints is also proposed and simulated.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Towards Superior Software Portability with SHAD and HPX C++ Libraries

As hardware architectures and software stacks complexity grows, development productivity, performance and software portability, quickly evolve from desirable features to actual needs. SHAD, the Scalable High-performance Algorithms and Data-structures C++ library is designed to mitigate these issues: it provides general purpose building blocks as well as high-level custom utilities, and offers a shared-memory programming abstraction which facilitates the programming of complex systems, scaling up to High Performance Computing clusters. SHAD’s portability is achieved through an abstract runtime interface, which decouples the upper layers of the library and hides the low level details of the underlying architecture. This layer enables SHAD to interface with different runtime/threading systems, e.g. Intel TBB and Global Memory and Threading (GMT). However, current backends targeting distributed systems, rely on a centralized controller which may possibly limit scalability up to hundreds of nodes and creates a network hot spot due to all to one communication for synchronization, and possibly resulting in degraded performance at high process counts. In this research, we explore HPX, the C++ standard library for parallelism and concurrency, as an additional backend in support of the SHAD library, and present the methodologies in support of local and remote task executions in SHAD with respect to HPX. Finally, we evaluate the proposed system by comparing against existing backends of SHAD and analyzing their performance on C++ Standard Template Library algorithms.

Wu, Nanmiao↗

Validation of Modern Nuclear Data Processing in SCALE

The nuclear data (ND) community is continuously developing more accurate, diversified, and comprehensive data for radiation transport modeling to support the nuclear science community. As these community efforts progress, it is crucial that ND processing tools like AMPX (used for SCALE [1] ND) also be developed in parallel to incorporate these new data into transport codes and actually deliver those data to end users. AMPX is a mature, well-tested code that was developed by many people at Oak Ridge National Laboratory (ORNL) over the course of the past few decades. A large portion of the AMPX codebase, however, was outdated, difficult to maintain, and incompatible with modern code development tools. Some of the most important parts of the AMPX code have now been replaced with modern C++ code that can be maintained more cost-effectively and can be tested more rigorously.

AMPX↗

Amplitude Analysis of ωπ0 Photoproduction at GlueX

spectrum of light mesons produced from a linearly polarized photon beam. The production and decays of a light meson resonance X such as γp → Xp′ →ωπ0p′ can be modeled with polarized vector-pseudoscalar ampli- tudes, which can describe the contribution of individual amplitudes to the total measured intensity. The status of mass-independent fits to the ωπ0 mass spectrum over a wide range of mass and momentum transfer−twill be presented, with an emphasis on interactions with the b1(1235) meson. We will also present in parallel an analysis of moments of the angular dis- tributions for the same process, as a means of verifying the stability of the amplitude-based results. These results will help to improve the broader knowledge of light meson states and how they are produced.

Scheuer, Kevin [College of William and Mary, Willi↗

FPGA-based computing system for processing data in size, weight, and power constrained environments

Technologies that are well-suited for use in size, weight, and power (SWAP)-constrained environments are described herein. A host controller dispatches data processing instructions to hardware acceleration engines (HAEs) of one or more field programmable gate arrays (FPGAs) and further dispatches data transfer instructions to a memory controller, such that the HAEs perform processing operations on data stored in local memory devices of the HAEs in parallel with other data being transferred from external memory devices coupled to the FPGA(s) to the local memory devices.

Napier, Matthew↗