Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “parallel systems”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 523 records · Page 29

Method of making a distributed optical fiber sensor having enhanced Rayleigh scattering and enhanced temperature stability, and monitoring systems employing same

A method of making an optical fiber sensor device for distributed sensing includes generating a laser beam comprising a plurality of ultrafast pulses, and focusing the laser beam into a core of an optical fiber to form a nanograting structure within the core, wherein the nanograting structure includes a plurality of spaced nanograting elements each extending substantially parallel to a longitudinal axis of optical fiber. Also, an optical fiber sensor device for distributed sensing includes an optical fiber having a longitudinal axis, a core, and a nanograting structure within the core, wherein the nanograting structure includes a plurality of spaced nanograting elements each extending substantially parallel to the longitudinal axis of the optical fiber. Also, a distributed sensing method and system and an energy production system that employs such an optical fiber sensor device.

42 ENGINEERING↗

Method of making a distributed optical fiber sensor having enhanced Rayleigh scattering and enhanced temperature stability, and monitoring systems employing same

A method of making an optical fiber sensor device for distributed sensing includes generating a laser beam comprising a plurality of ultrafast pulses, and focusing the laser beam into a core of an optical fiber to form a nanograting structure within the core, wherein the nanograting structure includes a plurality of spaced nanograting elements each extending substantially parallel to a longitudinal axis of optical fiber. Also, an optical fiber sensor device for distributed sensing includes an optical fiber having a longitudinal axis, a core, and a nanograting structure within the core, wherein the nanograting structure includes a plurality of spaced nanograting elements each extending substantially parallel to the longitudinal axis of the optical fiber. Also, a distributed sensing method and system and an energy production system that employs such an optical fiber sensor device.

Chen, Peng Kevin↗

The Integrated Reference Region Analysis for Parallel DFIGs’ Interfacing Inductors

Although the traditional design of doubly-fed induction generators (DFIGs)’ interfacing inductors consider the peak ripple of the grid-side converter (GSC)’s output current, it does not consider the inductors’ impact on DFIGs’ smallsignal stability. Located in series between the GSC and the stator, the inappropriate selection of the interfacing inductor can easily result to system instability. Therefore, this paper first proposes the integrated reference region analysis for small wind farm's interfacing inductors to improve the traditional design method. Firstly, a new detailed parallel DFIGs’ smallsignal model that considers output currents’ coupling is built in d-q coordinate system with state-space approach. The model focuses on representing the operating states of parallel DFIGs in wind farm. Secondly, the linearized state-space matrix of parallel DFIGs is decomposed into nominal-value matrix and location matrix. Considering the traditional design requirements, the integrated reference region for interfacing inductor is proposed through spectral radius and bialternate matrix sum (BMS). Furthermore, it can provide better guidance for parameter selecting and stabilization method researches. Finally, the simulation and experimental results show that the proposed integrated reference region for parallel DFIGs’ interfacing inductors is accurate and instructive.

47 OTHER INSTRUMENTATION↗

Coupled cluster theory on modern heterogeneous supercomputers

This study examines the computational challenges in elucidating intricate chemical systems, particularly through ab-initio methodologies. This work highlights the Divide-Expand-Consolidate (DEC) approach for coupled cluster (CC) theory—a linear-scaling, massively parallel framework—as a viable solution. Detailed scrutiny of the DEC framework reveals its extensive applicability for large chemical systems, yet it also acknowledges inherent limitations. To mitigate these constraints, the cluster perturbation theory is presented as an effective remedy. Attention is then directed towards the CPS (D-3) model, explicitly derived from a CC singles parent and a doubles auxiliary excitation space, for computing excitation energies. The reviewed new algorithms for the CPS (D-3) method efficiently capitalize on multiple nodes and graphical processing units, expediting heavy tensor contractions. As a result, CPS (D-3) emerges as a scalable, rapid, and precise solution for computing molecular properties in large molecular systems, marking it an efficient contender to conventional CC models.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Asynchronous distributed-memory task-parallel algorithm for compressible flows on unstructured 3D Eulerian grids

Here, we discuss the implementation of a finite element method, used to numerically solve the Euler equations of compressible flows, using an asynchronous runtime system (RTS). The algorithm is implemented for distributed-memory machines, using stationary unstructured 3D meshes, combining data-, and task-parallelism on top of the Charm++ RTS. Charm++’s execution model is asynchronous by default, allowing arbitrary overlap of computation and communication. Task-parallelism allows scheduling parts of an algorithm independently of, or dependent on, each other. Built-in automatic load balancing enables continuous redistribution of computational load by migration of work units based on real-time CPU load measurement. The RTS also features automatic checkpointing, fault tolerance, resilience against hardware failure, and supports power-, and energy-aware computation. We demonstrate scalability up to 25 x 10 9 cells at $\mathscr{O}$10 4 compute cores and the benefits of automatic load balancing for irregular workloads. The full source code with documentation is available at https://quinoacomputing.org.

42 ENGINEERING↗

GIGA-Lens: Fast Bayesian Inference for Strong Gravitational Lens Modeling

We present GIGA-Lens: a gradient-informed, GPU-accelerated Bayesian framework for modeling strong gravitational lensing systems, implemented in TensorFlow and JAX. The three components, optimization using multistart gradient descent, posterior covariance estimation with variational inference, and sampling via Hamiltonian Monte Carlo, all take advantage of gradient information through automatic differentiation and massive parallelization on graphics processing units (GPUs). We test our pipeline on a large set of simulated systems and demonstrate in detail its high level of performance. The average time to model a single system on four Nvidia A100 GPUs is 105 s. The robustness, speed, and scalability offered by this framework make it possible to model the large number of strong lenses found in current surveys and present a very promising prospect for the modeling of ${ \mathcal O }({10}^{5})$ lensing systems expected to be discovered in the era of the Vera C. Rubin Observatory, Euclid, and the Nancy Grace Roman Space Telescope.

79 ASTRONOMY AND ASTROPHYSICS↗

2023 Southeast Decarbonization Workshop

Decarbonization refers to a large-scale shift away from fossil fuel sources for energy production and toward energy sources, energy end-use practices, and land management approaches that do not result in a net increase of carbon dioxide in the atmosphere. Decarbonization in the Southeastern United States is distinguished from that of other regions in the country by its potential impact on historically underserved populations and the ways this region’s human–environmental systems are predicted to fare in a climate-altered world. In parallel, ensuring a clean energy transition requires the ability to engage entire communities that are motivated to learn, build, and encourage the spread of so-called clean tech. In contrast to other technology trends from the past century, the foundation of clean tech is a shared sense of purpose—a collective strategy to mitigate the threats of climate change and support a better environment for everyone. As seen in the Office of Science and Technology Policy’s (OSTP’s) Net-Zero Technology Action Plan (2023), decarbonization is best accelerated by simultaneous investments in Innovation, Demonstration, and Deployment of technology in tandem with intentional policy and community-based solutions. The 2023 Southeast Decarbonization Workshop, hosted by the Georgia Institute of Technology (Georgia Tech) and Oak Ridge National Laboratory (ORNL), aimed to bring together members of our communities and a group of regional experts to strategize about opportunities in this important area, as well as to incorporate the important pillar of System Interactions to bridge the gap between clean tech’s intent and its impact.

54 ENVIRONMENTAL SCIENCES↗

Quadratic pseudospectrum for identifying localized states

Here we examine the utility of the quadratic pseudospectrum for understanding and detecting states that are somewhat localized in position and energy, in particular, in the context of condensed matter physics. Specifically, the quadratic pseudospectrum represents a method for approaching systems with incompatible observables {A j |1 ≤ j ≤ d} as it minimizes collectively the errors $\parallel$A j v - λ j v$\parallel$ while defining a joint approximate spectrum of incompatible observables. Moreover, we derive an important estimate relating the Clifford and quadratic pseudospectra. Finally, we prove that the quadratic pseudospectrum is local and derive the bounds on the errors that are incurred by truncating the system in the vicinity of where the pseudospectrum is being calculated.

97 MATHEMATICS AND COMPUTING↗

Layer-Parallel Training of Deep Residual Neural Networks

Residual neural networks (ResNets) are a promising class of deep neural networks that have shown excellent performance for a number of learning tasks, e.g., image classification and recognition. Mathematically, ResNet architectures can be interpreted as forward Euler discretizations of a nonlinear initial value problem whose time-dependent control variables represent the weights of the neural network. Hence, training a ResNet can be cast as an optimal control problem of the associated dynamical system. For similar time-dependent optimal control problems arising in engineering applications, parallel-in-time methods have shown notable improvements in scalability. This paper demonstrates the use of those techniques for efficient and effective training of ResNets. The proposed algorithms replace the classical (sequential) forward and backward propagation through the network layers with a parallel nonlinear multigrid iteration applied to the layer domain. This adds a new dimension of parallelism across layers that is attractive when training very deep networks. From this basic idea, we derive multiple layer-parallel methods. The most efficient version employs a simultaneous optimization approach where updates to the network parameters are based on inexact gradient information in order to speed up the training process. Finally, using numerical examples from supervised classification, we demonstrate that the new approach achieves a training performance similar to that of traditional methods, but enables layer-parallelism and thus provides speedup over layer-serial methods through greater concurrency.

97 MATHEMATICS AND COMPUTING↗

Simulating coupled surface–subsurface flows with ParFlow v3.5.0: capabilities, applications, and ongoing development of an open-source, massively parallel, integrated hydrologic model

Surface flow and subsurface flow constitute a naturally linked hydrologic continuum that has not traditionally been simulated in an integrated fashion. Recognizing the interactions between these systems has encouraged the development of integrated hydrologic models (IHMs) capable of treating surface and subsurface systems as a single integrated resource. IHMs are dynamically evolving with improvements in technology, and the extent of their current capabilities are often only known to the developers and not general users. This article provides an overview of the core functionality, capability, applications, and ongoing development of one open-source IHM, ParFlow. ParFlow is a parallel, integrated, hydrologic model that simulates surface and subsurface flows. ParFlow solves the Richards equation for three-dimensional variably saturated groundwater flow and the two-dimensional kinematic wave approximation of the shallow water equations for overland flow. The model employs a conservative centered finite-difference scheme and a conservative finite-volume method for subsurface flow and transport, respectively. ParFlow uses multigrid-preconditioned Krylov and Newton–Krylov methods to solve the linear and nonlinear systems within each time step of the flow simulations. The code has demonstrated very efficient parallel solution capabilities. ParFlow has been coupled to geochemical reaction, land surface (e.g., the Common Land Model), and atmospheric models to study the interactions among the subsurface, land surface, and atmosphere systems across different spatial scales. This overview focuses on the current capabilities of the code, the core simulation engine, and the primary couplings of the subsurface model to other codes, taking a high-level perspective.

58 GEOSCIENCES↗

High-order matrix-free incompressible flow solvers with GPU acceleration and low-order refined preconditioners

In this work, we present a matrix-free flow solver for high-order finite element discretizations of the incompressible Navier-Stokes and Stokes equations with GPU acceleration. For high polynomial degrees, assembling the matrix for the linear systems resulting from the finite element discretization can be prohibitively expensive, both in terms of computational complexity and memory. For this reason, it is necessary to develop matrix-free operators and preconditioners, which can be used to efficiently solve these linear systems without access to the matrix entries themselves. The matrix-free operator evaluations utilize GPU-accelerated sum-factorization techniques to minimize memory movement and maximize throughput. The preconditioners developed in this work are based on a low-order refined methodology with parallel subspace corrections, as described for diffusion problems in [1]. The saddle-point Stokes system is solved using block-preconditioning techniques, which are robust in mesh size, polynomial degree, time step, and viscosity. For the incompressible Navier-Stokes equations, we make use of projection (fractional step) methods, which require Helmholtz and Poisson solves at each time step. The performance of our flow solvers is assessed on several benchmark problems in two and three spatial dimensions.

97 MATHEMATICS AND COMPUTING↗

Accelerating the density-functional tight-binding method using graphical processing units

Acceleration of the density-functional tight-binding (DFTB) method on single and multiple graphical processing units (GPUs) was accomplished using the MAGMA linear algebra library. Herein two major computational bottlenecks of DFTB ground-state calculations were addressed in our implementation: the Hamiltonian matrix diagonalization and the density matrix construction. The code was implemented and benchmarked on two different computer systems: (1) the SUMMIT IBM Power9 supercomputer at the Oak Ridge National Laboratory Leadership Computing Facility with 1–6 NVIDIA Volta V100 GPUs per computer node and (2) an in-house Intel Xeon computer with 1–2 NVIDIA Tesla P100 GPUs. The performance and parallel scalability were measured for three molecular models of 1-, 2-, and 3-dimensional chemical systems, represented by carbon nanotubes, covalent organic frameworks, and water clusters.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Controls Status of Fermilab's PIP-II Project

The Fermilab Proton Improvement Project II (PIP-II) is building a new Super Conducting Linear Accelerator (SCL) accelerating protons to 800 MeV for injection into the rest of the FNAL beam complex. Key progress since the last status report given at ICALEPCS includes the adoption of modern DevOps practices with continuous integration and GitOps-based deployments, commissioning of EPICS-based systems at the Cryomodule Test Facility, and integration of a Virtual Accelerator framework for application development ahead of installation. In parallel, web-based applications using Dart and Flutter have matured, providing secure, unified access to both EPICS and legacy ACNET data. Data acquisition and timing systems have also evolved. This paper presents the current state of controls, emphasizing these recent developments and outlining upcoming milestones as PIP-II approaches commissioning of its cryoplant in 2026 and the Warm Front End in 2027.

Crisp, D. B. [Fermilab]↗

ornladios/ADIOS2

ADIOS 2: The Adaptable Input Output (I/O) System version 2 is an open-source framework that addresses scientific data management challenges, e.g. scalable parallel I/O, as we approach the exascale era in high-performance computing (HPC). ADIOS 2 bindings are available in C++, C, Fortran, Python and can be used on supercomputers, personal computers, and cloud systems running on Linux, macOS and Windows. ADIOS 2 has out-of-the-box support for MPI and serial environments.

ECP↗

Software Control Program For Transportable Microgrid State-of-charge Balancing And Frequency Stability Controls

A deterministic state-of-charge (SOC) balancing approach software control code is introduced as an integral secondary management to primary control layer of an islanded small microgrid or nanogrid system made up of multiple grid-forming inverter/battery/solar combination systems, where each set of batteries with each inverter are on independent DC buses (i.e. non-paralleled on the DC sides). A DERMS-level control approach, algorithm and automation controller program was developed to improve coordination and enable microgrid asset compliance and SOC balancing, enabling provision of a system-level power stability support architecture, load support, and asset scalability. The architecture is configured to treat each unit or micro/nano-grid as a node in a microgrid network, allowing for autonomous DERMS control regarding load and SOC balancing and power stability. As the network grows with the addition of units, greater coordination efforts may be required. The ideal small network microgrid ranges from 2-10 inverter/battery units before additional control parameters must be considered in the existing architecture. The control approach focuses on a deterministic state-of-charge analysis as the primary level control process followed by a secondary control loop using a forced frequency-watt droop strategy to conform off-the-shelf components into behaving under a leader-follower configuration. Adopting this control scheme has been shown to allow for a balanced, unit-coordinated microgrid network, enabling stable power flow. The deterministic state-of-charge approach is introduced as an integral primary control layer of an islanded small network microgrid. A standard strategy for SOC balancing is implementing a battery management system (BMS) to control SOC on the DC side. An alternative approach is to determine how to coordinate sending and receiving power on the AC side with multiple units. The latter approach assesses all the integrated units in the microgrid network. Once the individual units are identified, further system data is required to calculate each unit's total kWh, provided information about its capability to supply or consume kWh and availability. The secondary control layer in the multi-layered small network microgrid methodology uses the primary layer’s decision to initiate frequency setpoint changes, initializing the SOC balancing. The secondary control layer considers numerous system-dependent variables to enable a charging and discharging profile based on adjustable frequency setpoints. The combined architecture will result in stable, coordinated power flow enhancing an AC microgrid's functionalities.

Myers, KurtS [Idaho National Laboratory (INL), Ida↗

Multi-mode power split hybrid transmission with two planetary gear mechanisms

A multi-mode, power-split hybrid transmission system having two planetary gear (PG) sets connected to one engine, two electric motors, one output shaft, and each other by several clutches, brakes, and direct connection elements. Depending on the specific location and actuation of the various clutch and brake elements, the multi-mode, power-split hybrid transmission system can be run in one of several modes (e.g. electric drive, power-split, parallel hybrid, series hybrid, electronic continuously variable transmission (eCVT), generator, neutral, and the like).

33 ADVANCED PROPULSION SYSTEMS↗

Parallel simulation via SPPARKS of on-lattice kinetic and Metropolis Monte Carlo models for materials processing

Abstract SPPARKS is an open-source parallel simulation code for developing and running various kinds of on-lattice Monte Carlo models at the atomic or meso scales. It can be used to study the properties of solid-state materials as well as model their dynamic evolution during processing. The modular nature of the code allows new models and diagnostic computations to be added without modification to its core functionality, including its parallel algorithms. A variety of models for microstructural evolution (grain growth), solid-state diffusion, thin film deposition, and additive manufacturing (AM) processes are included in the code. SPPARKS can also be used to implement grid-based algorithms such as phase field or cellular automata models, to run either in tandem with a Monte Carlo method or independently. For very large systems such as AM applications, the Stitch I/O library is included, which enables only a small portion of a huge system to be resident in memory. In this paper we describe SPPARKS and its parallel algorithms and performance, explain how new Monte Carlo models can be added, and highlight a variety of applications which have been developed within the code.

36 MATERIALS SCIENCE↗

Multiphysics analysis system for heat pipe cooled micro-reactors employing PRAGMA-OpenFOAM-ANLHTP

A multiphysics analysis system for neutronics/thermo-mechanical/heat-pipe analysis of heat pipe cooled micro-reactors was developed using the PRAGMA code as the neutronics engine. PRAGMA, which used to be a GPU-based continuous-energy MC code for power reactor applications, now has an extended geometry package to handle geometries with unstructured meshes generated by the ANSYS Design-Modeler and Meshing. The NVIDIA ray tracing engine OptiX was exploited for efficient neutron transport on unstructured geometry. On the multiphysics side, the open-source CFD tool OpenFOAM and one-dimensional heat pipe analysis code ANLHTP were adopted. The manager-worker system based on the MPI dynamic process management (DPM) model enables efficient coupling of codes employing different parallelization schemes. With all the features, the multiphysics analysis of the one-sixth symmetrical MegaPower 2D core was performed and it demonstrated the benefit of the tight integration of the three-way coupling system and one-to-one geometry coupling strategy. (authors)

21 SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLANTS↗