Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “asynchronous”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

Formally Verified ZTA Requirements for OT/ICS Environments with Isabelle/HOL

The clean energy transformation includes the integration of distributed energy resources with the power grid, which has led to a substantial increase in the complexity of power grids infrastructure and the underlying operational technology environment. Power grids infrastructure represents an operational technology environment that has become a system of systems, integrating heterogeneous devices which are both software-and hardware-intensive; as a result, there are increasing demands to exploit advances in the commodity of software-hardware infrastructures to improve energy systems requirements such as cybersecurity and resilience. In such a setting, system requirements at different levels mix, which leads to vulnerabilities and undesirable outcomes. The use of formal methods to characterize and prove system requirements removes ambiguity, increases automation, and provides high levels of assurance and reliability. In this paper, we contribute a methodology and a framework for the system-level verification of zero trust architecture requirements in operational technology environments. We define a formal specification for the core functionalities of operational technology environments, the corresponding invariants, and security proofs. Of particular note is our modular approach for the formal verification of asynchronous interactions in operational technology environments. The formal specification and the proofs have been mechanized using the interactive theorem proving environment Isabelle/HOL.

formal methods↗

Lessons Learned and Scalability Achieved When Porting Uintah to DOE Exascale Systems

A key challenge faced when preparing codes for Department of Energy (DOE) exascale systems was designing scalable applications for systems featuring hardware and software not yet available at leadership-class scale. With such systems now available, it is important to evaluate scalability of the resulting software solutions on these target systems. One such code designed with the exascale DOE Aurora and DOE Frontier systems in mind is the Uintah Computational Framework, an open-source asynchronous many-task (AMT) runtime system. To prepare for exascale, Uintah adopted a portable MPI+X hybrid parallelism approach using the Kokkos performance portability library (i.e., MPI+Kokkos). This paper complements recent work with additional details and an evaluation of the resulting approach on Aurora and Frontier. Results are shown for a challenging benchmark demonstrating interoperability of 3 portable codes essential to Uintah-related combustion research. These results demonstrate single-source portability across Aurora and Frontier with scaling characteristics shown to 3,072 Aurora nodes and 9,216 Frontier nodes. In addition to showing results run to new scales on new systems, this paper also discusses lessons learned through efforts preparing Uintah for exascale systems.

Holmen, John [ORNL] (ORCID:0000000259342641)↗

Distributed out-of-memory NMF on CPU/GPU architectures

We propose an efficient distributed out-of-memory implementation of the non-negative matrix factorization (NMF) algorithm for heterogeneous high-performance-computing systems. The proposed implementation is based on prior work on NMFk, which can perform automatic model selection and extract latent variables and patterns from data. In this work, we extend NMFk by adding support for dense and sparse matrix operation on multi-node, multi-GPU systems. The resulting algorithm is optimized for out-of-memory problems where the memory required to factorize a given matrix is greater than the available GPU memory. Memory complexity is reduced by batching/tiling strategies, and sparse and dense matrix operations are significantly accelerated with GPU cores (or tensor cores when available). Input/output latency associated with batch copies between host and device is hidden using CUDA streams to overlap data transfers and compute asynchronously, and latency associated with collective communications (both intra-node and inter-node) is reduced using optimized NVIDIA Collective Communication Library (NCCL) based communicators. Benchmark results show significant improvement, from 32X to 76x speedup, with the new implementation using GPUs over the CPU-based NMFk. Good weak scaling was demonstrated on up to 4096 multi-GPU cluster nodes with approximately 25,000 GPUs when decomposing a dense 340 Terabyte-size matrix and an 11 Exabyte-size sparse matrix of density 10 -6 .

97 MATHEMATICS AND COMPUTING↗

Chapter 9: Impact of Variable Renewable Energy Sources on Bulk Power System Planning and Operations

Wind and solar photovoltaics (PV) have experienced remarkable growth in recent years, with many consequent benefits within and outside of power systems. At the same time, wind and solar PV have unique characteristics relative to the historically dominant dispatchable technologies like coal, gas, and nuclear power plants that have required and will continue to require changes in power system planning and operations. This chapter discusses planning and operational challenges of integrating wind and solar PV into bulk power systems. We first present the key characteristics of wind and solar PV that differentiate it from conventional technologies, such as variable and uncertain electricity generation, asynchronous interconnection to the power system, and near-zero marginal costs. We then link these characteristics to power system planning and operational challenges at low through high wind and solar penetrations. Finally, we discuss near- and long-term solutions to those challenges, such as diversifying the generation mix and wind and solar fleets, improving system flexibility, diversifying ancillary service products, and integrating generation and transmission planning.

bulk power system↗

An active learning high-throughput microstructure calibration framework for solving inverse structure–process problems in materials informatics

Determining a process–structure–property relationship is the holy grail of materials science, where both computational prediction in the forward direction and materials design in the inverse direction are essential. Problems in materials design are often considered in the context of process–property linkage by bypassing the materials structure, or in the context of structure–property linkage as in microstructure-sensitive design problems. However, there is a lack of research effort in studying materials design problems in the context of process–structure linkage, which has a great implication in reverse engineering. In this paper, given a target microstructure, we propose an active learning high-throughput microstructure calibration framework to derive a set of processing parameters, which can produce an optimal microstructure that is statistically equivalent to the target microstructure. The proposed framework is formulated as a noisy multi-objective optimization problem, where each objective function measures a deterministic or statistical difference of the same microstructure descriptor between a candidate microstructure and a target microstructure. Furthermore, to significantly reduce the physical waiting wall-time, we enable the high-throughput feature of the microstructure calibration framework by adopting an asynchronously parallel Bayesian optimization to exploit high-performance computing resources. Case studies in additive manufacturing and grain growth are used to demonstrate the applicability of the proposed framework, where kinetic Monte Carlo (kMC) simulation is used as a forward predictive model, such that for a given target microstructure, the target processing parameters that produced this microstructure are successfully recovered.

36 MATERIALS SCIENCE↗

The use of teledentistry in facilitating oral health for older adults: A scoping review

Teledentistry is used in many countries to provide oral health care services. However, using teledentistry to provide oral health care services for older adults is not well documented. This knowledge gap needs to be addressed, especially when accessing a dental clinic is not possible and teledentistry might be the only way for many older adults to receive oral health care services. Nine databases were searched and 3,396 studies were screened using established eligibility criteria. Included studies were original research or review articles in which the intervention of interest was delivered to an older adult population (≥ 60 years) via teledentistry. The authors followed the Preferred Reporting Items for Systematic Reviews and Meta-Analyses Extension for Scoping Review criteria. Nineteen studies were identified that met the criteria for inclusion. Only 1 study was from the United States. Seven studies had results focusing on older adult participants only, with most of those conducted in elder care facilities. The remainder consisted of studies with mixed-age populations reporting distinct results or information for older adults. The included studies used teledentistry, in both synchronous and asynchronous modes, to provide services such as diagnosis, oral hygiene promotion, assessment and referral of oral emergencies, and postintervention follow-up. Teledentistry comprises a variety of promising apps. Lastly, the authors identified and described uses, promising possibilities, and limitations of teledentistry to improve the oral health of older adults.

60 APPLIED LIFE SCIENCES↗

Causality guided machine learning model on wetland CH 4 emissions across global wetlands

Wetland CH 4 emissions are among the most uncertain components of the global CH 4 budget. The complex nature of wetland CH 4 processes makes it challenging to identify causal relationships for improving our understanding and predictability of CH 4 emissions. In this study, we used the flux measurements of CH 4 from eddy covariance towers (30 sites from 4 wetlands types: bog, fen, marsh, and wet tundra) to construct a causality-constrained machine learning (ML) framework to explain the regulative factors and to capture CH 4 emissions at sub-seasonal scale. We found that soil temperature is the dominant factor for CH 4 emissions in all studied wetland types. Ecosystem respiration (CO 2 ) and gross primary productivity exert controls at bog, fen, and marsh sites with lagged responses of days to weeks. Integrating these asynchronous environmental and biological causal relationships in predictive models significantly improved model performance. More importantly, modeled CH 4 emissions differed by up to a factor of 4 under a +1°C warming scenario when causality constraints were considered. These results highlight the significant role of causality in modeling wetland CH 4 emissions especially under future warming conditions, while traditional data-driven ML models may reproduce observations for the wrong reasons. Our proposed causality-guided model could benefit predictive modeling, large-scale upscaling, data gap-filling, and surrogate modeling of wetland CH 4 emissions within earth system land models.

54 ENVIRONMENTAL SCIENCES↗

Neutron transport methods for multiphysics heterogeneous reactor core simulation in Griffin

Griffin is a reactor physics application based on the Multiphysics Object-Oriented Simulation Environment (MOOSE). This work discloses the methods, algorithms, and implementation for simulating heterogeneous reactor dynamics models. Griffin utilizes a discontinuous finite-element method with discrete ordinates (DFEM-S ) to discretize the field variable of the multigroup neutron transport equation. Multiphysics feedback is handled using two-step tabulated cross-section methodology. Feedback quantities are evaluated using the MOOSE-MultiApp system to couple various engineering phenomena, such as heat conduction and thermal fluids. The multiphysics DFEM-S system is solved using fixed-point iteration with a fully asynchronous parallel sweeper, unstructured coarse-mesh finite difference acceleration, and a multi-timescale improved quasi-static method scheme. The implementation is applied to a multiphysics microreactor model, with two transients: one initiated by a single heat-pipe failure and another by control drum rotation. Importantly, these examples demonstrate the ability of Griffin to tractably solve the neutron transport equation considering seven independent variables and feedback.

97 MATHEMATICS AND COMPUTING↗

Myna: Connecting powder bed fusion build data to simulation tools for digital twin applications

Additive manufacturing (AM), as a digital process, can generate a detailed digital thread linking a part’s design and manufacturing to its operational performance. As AM systems advance, an increasing amount of process data is stored in manufacturing databases. In principle, this data can be utilized by simulation-based digital twin approaches, such as real-time process control and asynchronous post-processing guidance. However, few tools currently exist for systematically integrating digital thread data with computational tools. Here, in this study, we propose a software package, called Myna, for connecting data from powder bed fusion processes to simulation tools. The utility of such a platform is demonstrated using build data from the Oak Ridge National Laboratory Manufacturing Demonstration Facility “Peregrine v2023-10” public dataset to automatically configure and run 54 semi-analytical 3DThesis melt pool simulations, 78 numerical Additive FOAM melt pool simulations, and 3 ExaCA microstructure simulations. The simulated, spatially registered microstructures are then compared directly with electron backscatter diffraction characterization of the corresponding as-built part locations. The resulting simulated microstructure showed variation as a function of process parameters, particularly stripe width; however, the experimental data had little variation between the microstructure texture and grain size resulting from different processing conditions. Analysis of the discrepancies suggest that it is possible a two-phase ferritic-austenitic solidification model is needed to accurately predict grain size and texture for certain stainless steel 316L feedstock compositions under powder bed fusion conditions, providing direction for future research. As illustrated here, due to the number and complexity of the simulations involved in AM process-structure–property predictions, automated methods to connect process data and simulations will remain necessary tools for testing hypotheses and implementing digital twin applications.

Knapp, Gerald L. [Oak Ridge National Laboratory (O↗

Code modernization strategies for short-range non-bonded molecular dynamics simulations

Modern HPC systems are increasingly relying on greater core counts and wider vector registers. Thus, applications need to be adapted to fully utilize these hardware capabilities. One class of applications that can benefit from this increase in parallelism are molecular dynamics simulations. In this paper, we describe our efforts at modernizing the ESPResSo++ simulation package for molecular dynamics by restructuring its particle data layout for efficient memory accesses and applying vectorization techniques to benefit the calculation of short-range non-bonded forces, which results in an overall three times speedup and serves as a baseline for further optimizations. We also implement fine-grained parallelism for multi-core CPUs through HPX, a C++ runtime system which uses lightweight threads and an asynchronous many-task approach to maximize concurrency. Our goal is to evaluate the performance of an HPX-based approach compared to the bulk-synchronous MPI-based implementation. This requires the introduction of an additional layer to the domain decomposition scheme that defines the task granularity. On spatially inhomogeneous systems, which impose a corresponding load-imbalance in traditional MPI-based approaches, we demonstrate that by choosing an optimal task size, the efficient work-stealing mechanisms of HPX can overcome the overhead of communication resulting in an overall 1.4 times speedup compared to the baseline MPI version.

97 MATHEMATICS AND COMPUTING↗

OpenSn: A massively parallel, open-source simulation environment for discrete ordinates radiation transport

OpenSn is an open-source, massively parallel deterministic radiation transport code for solving the discrete-ordinates ( S N ) form of the Boltzmann transport equation on unstructured, arbitrary polyhedral meshes. It supports high-fidelity simulations involving steady-state, eigenvalue, and adjoint problems for neutral particles (e.g., neutrons, photons, multi-particles), using the multigroup approximation in energy. OpenSn combines angular discretization via discrete ordinates with a discontinuous Galerkin finite element method (DGFEM) in space, enabling accurate resolution of transport physics on arbitrary polyhedral cells, included locally refined spatial grids. It includes multiple angular quadrature types, including locally refined angular quadratures. Written in modern C++ with a Python API, OpenSn runs efficiently on platforms ranging from laptops to supercomputers. The transport sweep algorithm is implemented using a task-based, directed-acyclic-graph (DAG) approach for each angle and supports asynchronous parallelism across thousands of MPI ranks. Group-set aggregation improves compute intensity, and synthetic acceleration techniques (e.g., diffusion synthetic acceleration, second-moment method) enhance solver convergence. OpenSn has been verified on reactor physics problems and demonstrated excellent weak and strong scaling performance on more than 32,768 processes, making it a versatile and robust platform for large-scale transport simulations in complex geometries.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Towards high data-rate diffusive molecular communications: A review on performance enhancement strategies

Diffusive molecular communications (DiMC) have recently gained attention as a candidate for nano- to micro- and macro-scale communications due to its simplicity and energy efficiency. As signal propagation is solely enabled by Brownian motion mechanics, DiMC faces severe inter-symbol interference (ISI), which limits reliable and high data-rate communications. Herein, recent literature on DiMC performance enhancement strategies is surveyed; key research directions are identified. Signaling design and associated design constraints are presented. Here, studies on fundamental information theoretic limits of DiMC channel are reviewed. Classical and novel transceiver designs are discussed with an emphasis on methods for ISI mitigation and performance-complexity tradeoffs. Key parameter estimation strategies such as synchronization and channel estimation are considered in conjunction with asynchronous and timing error robust receiver methods. Finally, source and channel coding in the context of DiMC is presented.

42 ENGINEERING↗

Co-simulation of transactive energy markets: A framework for market testing and evaluation

The proliferation of distributed energy resources (DER)—and the ability to intelligently control these assets—is re-defining the electrical distribution system. As the number of controllable devices rapidly expands, grid operators must determine how to incorporate these assets while delivering reliable, equitable, and affordable electricity. One possible approach is to establish distribution-level electricity markets and allow devices/aggregations of devices to participate in price establishment. While this approach purports some of the same benefits as the highly successful wholesale electricity markets (i.e., open competition, efficient price discovery, reduced communication overhead), this needs to be researched and quantified via an analysis platform that models distribution-level markets at the appropriate fidelity. Specifically, the simultaneous evaluation of market performance, DER performance, DER bidding approaches, and distribution feeder power quality requires modeling that spans multiple technical areas. Co-simulation has emerged as a powerful tool in addressing this type of problem, where outputs depend on a range of underlying areas of expertise and associated models. In this paper we describe a solution, as implemented in the HELICS co-simulation platform, where we include (1) high fidelity house models, (2) intelligent bidding agents, (3) a modular market integration/design, and (4) a distribution feeder model. We then present a case study where we test two different market designs: (1) a pseudo-wholesale double-blind auction, and (2) an asynchronous matching market. In this work, the markets are run under two DER penetration levels and economic results are compared to full retail net energy metering and avoided cost net metering scenarios that bookend current approaches to remuneration of DER participation. We show the potential for transactive markets to provide increased value for most customers relative to net metering (and all customers relative to avoided cost scenarios) while decreasing costs for the utility.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Integrating a ponderomotive guiding center algorithm into a quasi-static particle-in-cell code based on azimuthal mode decomposition

High fidelity modeling of plasma based acceleration (PBA) requires the use of three dimensional, fully nonlinear, and kinetic descriptions based on the particle-in-cell (PIC) method. In PBA an intense particle beam or laser (driver) propagates through a tenuous plasma whereby it excites a plasma wave wake. Three-dimensional PIC algorithms based on the quasi-static approximation (QSA) have been successfully applied to efficiently model the interaction between relativistic charged particle beams and plasma. In a QSA PIC algorithm, the plasma response to a charged particle beam or laser driver is calculated based on forces from the driver and self-consistent forces from the QSA form of Maxwell's equations. These fields are then used to advance the charged particle beam or laser forward by a large time step. Since the time step is not limited by the regular Courant-Friedrichs-Lewy (CFL) condition that constrains a standard 3D fully electromagnetic PIC code, a 3D QSA PIC code can achieve orders of magnitude speedup in performance. Recently, a new hybrid QSA PIC algorithm that combines another speedup technique known as an azimuthal Fourier decomposition has been proposed and implemented. This hybrid algorithm decomposes the electromagnetic fields, charge and current density into azimuthal harmonics and only the Fourier coefficients need to be updated, which can reduce the algorithmic complexity of a 3D code to that of a 2D code. Modeling the laser-plasma interaction in a full 3D electromagnetic PIC algorithm is very computationally expensive due the enormous disparity of physical scales to be resolved. In the QSA the laser is modeled using the ponderomotive guiding center (PGC) approach. We describe how to implement a PGC algorithm compatible for the QSA PIC algorithms based on the azimuthal mode expansion. Here this algorithm permits time steps orders of magnitude larger than the cell size and it can be asynchronously parallelized. Details on how this is implemented into the QSA PIC code that utilizes an azimuthal mode expansion, QPAD, are also described. Benchmarks and comparisons between a fully 3D explicit PIC code (OSIRIS), as well as a few examples related to laser wakefield acceleration, are presented.

97 MATHEMATICS AND COMPUTING↗

Proton-detected solid-state NMR spectroscopy of spin-1/2 nuclei with large chemical shift anisotropy

Constant-time (CT) dipolar heteronuclear multiple quantum coherence (D-HMQC) has previously been demonstrated as a method for proton detection of high-resolution wideline NMR spectra of spin-1/2 nuclei with large chemical shift anisotropy (CSA). However, 1 H transverse relaxation and t 1 -noise often reduce the sensitivity of D-HMQC experiments, preventing the theoretical gains in sensitivity provided by 1 H detection from being realized. In this paper, we demonstrate a series of improved pulse sequences for 1 H detection of spin-1/2 nuclei under fast MAS, with 195 Pt SSNMR experiments on cisplatin as an example. First, a t 1 -incrementation protocol for D-HMQC dubbed Arbitrary Indirect Dwell (AID) is demonstrated. AID allows the use of arbitrary, rotor asynchronous t 1 -increments, but removes the constant time period from CT D-HMQC, resulting in improved sensitivity by reducing transverse relaxation losses. Next, we show that short high-power adiabatic pulses (SHAPs), which efficiently invert broad MAS sideband manifolds, can be effectively incorporated into 1 H detected symmetry-based resonance echo double resonance (S-REDOR) and t 1 -noise eliminated (TONE) D-HMQC experiments. The S-REDOR experiments with SHAPs provide approximately double the dipolar dephasing, as compared to experiments with rectangular inversion pulses. We lastly show that sensitivity and resolution can be further enhanced with the use of swept excitation pulses as well as adiabatic magic angle turning (aMAT).

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Distributed approximate minimal Steiner trees with millions of seed vertices on billion-edge graphs

In this report, we present a parallel 2-approximation Steiner minimal tree algorithm and its MPI-based distributed implementation. In place of expensive distance computations between all pairs of seed vertices, the solution we employ exploits a cheaper Voronoi cell computation. Our design leverages asynchronous processing and message prioritization to accelerate convergence of distance computations, and harnesses vertex and edge centric processing to offer fast time-to-solution. We demonstrate scalability and performance using real-world graphs with up to 128 billion edges and 512 compute nodes, and show the ability to find Steiner trees with up to one million seed vertices. Using 12 data instances, we present comparison with the state-of-the-art exact solver, SCIP-Jack, and two sequential 2-approximate algorithms. We empirically show that, on average, the total distance of the Steiner tree identified by our solution is 1.1290 times greater than the Steiner minimal tree – well within the theoretical approximation bound of 2.

97 MATHEMATICS AND COMPUTING↗

Bicontinuous nanoporous design induced homogenization of strain localization in metallic glasses

Bicontinuous nanoporous metallic glasses (MG) synergize the outstanding properties of MGs and open-cell nanoporous materials. The low-density and high-specific-surface-area of bicontinuous nanoporous structures have the potential to enhance the applicability of MGs in catalysis, sensors, and lightweight structural designs. In this work, we report molecular dynamics simulations of tensile loading deformation and failure of bicontinuous nanoporous Cu 64 Zr 36 MG with 55% porosity and 4.4 nm ligament size. Results indicate an anomalous mechanical behavior featuring delocalized plastic deformation preceding ductile failure. The deformation follows two mechanisms: i) Necking of ligaments aligned with the loading direction and ii) progressive alignment of randomly oriented ligaments. Failure occurs at 0.16 strain, following massive rupture of ligaments. This work indicates that a bicontinuous nanoporous design is able to effectively delocalize strain localization in a MG due to a combination of size effect on the ductility of MGs resulting in nano ligaments necking and progressive asynchronous alignment of ligaments.

36 MATERIALS SCIENCE↗

Flash-X: A multiphysics simulation software instrument

Flash-X is a highly composable multiphysics software system that can be used to simulate physical phenomena in several scientific domains. It derives some of its solvers from FLASH, which was first released in 2000. Flash-X has a new framework that relies on abstractions and asynchronous communications for performance portability across a range of increasingly heterogeneous hardware platforms. Flash-X is meant primarily for solving Eulerian formulations of applications with compressible and/or incompressible reactive flows. It also has a built-in, versatile Lagrangian framework that can be used in many different ways, including implementing tracers, particle-in-cell simulations, and immersed boundary methods.

97 MATHEMATICS AND COMPUTING↗