Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “parallel systems”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 541 records · Page 30

Diagnostic development for parallel wave-number measurement of lower hybrid waves in EAST

In this study, an eight-channel magnetic probe diagnostic system has been designed and installed adjacent to the 4.6 GHz lower hybrid (LH) grill antenna in the low-field side of the Experimental Advanced Superconducting Tokamak (EAST) in order to study the n ∥ evolution of LH waves in the first pass from the launcher to the core plasma. The magnetic probes are separated by 6.6 mm, which allows measurement of the dominant parallel refractive index n ∥ up to n ∥ = 5 for 4.6 GHz LH waves. The magnetic probes are designed to be sensitive to the magnetic field component perpendicular to the background magnetic field with a slit on the casing that encloses the probe. The intermediate frequency stage, which consists of two mixing stages, down-coverts the frequency of the measured wave signals at 4.6 GHz to 20 MHz. A bench test demonstrates the phase stability of the magnetic probe diagnostic system. By evaluating the phase variation of the measured signals along the background magnetic field, the dominant n ∥ of the LH wave in the scrape-off layer has been deduced during the 2019 experimental campaign. In the low density plasma, the measured dominant n ∥ of the LH waves is about 2.1, corresponding to the main peak 2.04 of the launched n ∥ spectrum. n ∥ deduced by the least-squares linear fit method remains near this value in the low density plasma with a high spatial correlation magnitude of 0.9. With an eight-channel probe system, a wave-number spectrum has also been deduced, which has a peak near to the measured dominant n ∥ .

47 OTHER INSTRUMENTATION↗

A Compact 50kW High Power Density, Hybrid 3-Level Paralleled T-type Inverter for More Electric Aircraft Applications

The demand for high performing, lightweight, reliable inverters, increased the scope of wide bandgap and high-switching frequency based solutions. To achieve such high efficiency inverters, it is vital to focus on the system level design considerations to maximize the benefits of these advanced technologies. This paper presents an improved design based on considerations to further reap the benefits of choosing the right inverter topology; increased capabilities through paralleling devices, with reduced total number of switches; and designing a planarized inverter with PCB based busbar. Appropriate thermal analysis and heatsink design has aided in increased system power density along with the overall efficiency. Demonstration of a 50kW 3-phase 3-level paralleled T-type SiC inverter operating at 40kHz switching frequency for aircraft applications is shown to evaluate the benefits of proposed design methodology. Here, the prototype achieves a high power density of 11kW/L.

42 ENGINEERING↗

On the potentially transformative role of auxiliary-field quantum Monte Carlo in quantum chemistry: A highly accurate method for transition metals and beyond

Approximate solutions to the ab initio electronic structure problem have been a focus of theoretical and computational chemistry research for much of the past century, with the goal of predicting relevant energy differences to within “chemical accuracy” (1 kcal/mol). For small organic molecules, or in general, for weakly correlated main group chemistry, a hierarchy of single-reference wave function methods has been rigorously established, spanning perturbation theory and the coupled cluster (CC) formalism. For these systems, CC with singles, doubles, and perturbative triples is known to achieve chemical accuracy, albeit at O(N7) computational cost. In addition, a hierarchy of density functional approximations of increasing formal sophistication, known as Jacob’s ladder, has been shown to systematically reduce average errors over large datasets representing weakly correlated chemistry. However, the accuracy of such computational models is less clear in the increasingly important frontiers of chemical space including transition metals and f-block compounds, in which strong correlation can play an important role in reactivity. A stochastic method, phaseless auxiliary-field quantum Monte Carlo (ph-AFQMC), has been shown to be capable of producing chemically accurate predictions even for challenging molecular systems beyond the main group, with relatively low O(N3 − N4) cost and near-perfect parallel efficiency. Herein, we present our perspectives on the past, present, and future of the ph-AFQMC method. We focus on its potential in transition metal quantum chemistry to be a highly accurate, systematically improvable method that can reliably probe strongly correlated systems in biology and chemical catalysis and provide reference thermochemical values (for future development of density functionals or interatomic potentials) when experiments are either noisy or absent. Finally, we discuss the present limitations of the method and where we expect near-term development to be most fruitful.

Chemistry↗

Space-Time Block Preconditioning for Incompressible Flow

Parallel-in-time methods have become increasingly popular in the simulation of time-dependent numerical PDEs, allowing for the efficient use of additional message passing interface processes when spatial parallelism saturates. Most methods treat the solution and parallelism in space and time separately. In contrast, all-at-once methods solve the full space-time system directly, largely treating time as simply another spatial dimension. All-at-once methods offer a number of benefits over separate treatment of space and time, most notably significantly increased parallelism and faster time to solution (when applicable). However, the development of fast, scalable all-at-once methods has largely been limited to time-dependent (advection-)diffusion problems. This paper introduces the concept of space-time block preconditioning for the all-at-once solution of incompressible flow. By extending well-known concepts of spatial block preconditioning to the space-time setting, we develop a block preconditioner whose application requires the solution of a space-time (advection-)diffusion equation in the velocity block, coupled with a pressure Schur complement approximation consisting of independent spatial solves at each time-step, and a space-time matrix-vector multiplication. The new method is tested on four classical models in incompressible flow. Finally, the results indicate perfect scalability in refinement of spatial and temporal mesh spacing, perfect scalability in nonlinear Picard iteration count when applied to a nonlinear Navier--Stokes problem, and minimal overhead in terms of number of preconditioner applications compared with sequential time-stepping.

97 MATHEMATICS AND COMPUTING↗

Enabling Ultra-Compact, Lightweight, Efficient, and Reliable 6.6 kW On-Board Bi-Directional Electric Vehicle Charger with Advanced Topology and Control

The research explored new topologies, control methods, mechanical integration, and thermal management methods for electric vehicle (EV) on-board chargers. The team investigated capacitor-based power conversion, leveraging the high energy densities inherent to capacitive energy storage compared to inductive methods. The proposed topologies simultaneously enabled high power density and high efficiency of the design. The proposed architecture was demonstrated in a 6.6 kW bi-directional charger prototype. The research pursued several directions to improve system performance. Innovative topologies were studied for both the main power conversion stage as well as the single-phase twice-line-frequency energy buffer. To ensure robust and efficient operation, new control methods were developed to integrate these two subsystems. To achieve high power density in the full system solution, the mechanical structure of the charger is highly optimized to maximally fill the converter box volume. In parallel with the mechanical design effort, the converter was packaged with high-performance cooling methods which removed heat from key areas of power dissipation in the converter. The thermal management system was optimized to minimize its weight and volume, ultimately motivating the design of a custom additively manufactured cold-plate. The full system achieves a peak power of 7 kW with less than 0.3% total harmonic distortion (THD) and greater than 0.994 power factor in power factor correction (PFC) operation, corresponding to a total box-volume power density of 47.9 kW/L and gravimetric power density of 24.6 W/g. The system achieves a peak efficiency of 98.9%, with 97.9% efficiency at maximum power.

33 ADVANCED PROPULSION SYSTEMS↗

A Fine-grained Asynchronous Bulk Synchronous parallelism model for PGAS applications

The Partitioned Global Address Space (PGAS) model is well suited for executing irregular applications on cluster-based systems, due to its efficient support for short, one-sided messages. Separately, the actor model has been gaining popularity as a productive asynchronous message-passing approach for distributed objects in enterprise and cloud computing platforms, typically implemented in languages such as Erlang, Scala or Rust. To the best of our knowledge, there has been no past work on using the actor model to deliver both productivity and scalability to irregular PGAS applications with large number of small messages. In this paper, we introduce a new programming system for PGAS applications, in which point-to-point remote operations can be expressed as fine-grained asynchronous actor messages. In our approach, the programmer does not need to worry about programming complexities related to message aggregation and termination detection. Our approach can be viewed as extending the classical Bulk Synchronous Parallelism model with fine-grained asynchronous communications within a phase or superstep. Here, we believe that our approach offers a desirable point in the productivity-performance space for PGAS applications, with more scalable performance and higher productivity relative to past approaches. Specifically, for seven irregular mini-applications from the Bale Kernels and three graph kernels executed using 2048 cores in the NERSC Cori system, our approach shows geometric mean performance improvements of ≥ 20X relative to standard PGAS versions (UPC and OpenSHMEM) while maintaining comparable productivity to those versions.

97 MATHEMATICS AND COMPUTING↗

Parallel Algorithms for Efficient Computation of High-Order Line Graphs of Hypergraphs

This paper considers structures of systems beyond dyadic (pairwise) interactions and investigates mathematical modeling of multi-way interactions and connections as hypergraphs, where captured relationships among system entities are set-valued. To date, in most situations, entities in a hypergraph are considered connected as long as there is at least one common ``neighbor''. However, minimal commonality sometimes discards the ``strength'' of connections and interactions among groups. To this end, considering the ``width'' of a connection, referred to as the \emph{$s$-overlap} of neighbors, provides more meaningful insights into how closely the communities or entities interact with each other. In addition, $s$-overlap computation is the fundamental kernel to construct the line graph of a hypergraph, a low-order approximation of the hypergraph which can carry significant information about the original hypergraph. Subsequent stages of a data analytics pipeline then can apply highly-tuned graph algorithms on the line graph to reveal important features. Given a hypergraph, computing the $s$-overlaps by exhaustively considering all pairwise entities can be computationally prohibitive. To tackle this challenge, we develop efficient algorithms to compute $s$-overlaps and the corresponding line graph of a hypergraph. We propose several heuristics to avoid execution of redundant work and improve performance of the $s$-overlap computation. Our parallel algorithm, combined with these heuristics, is orders of magnitude (more than $10\times$) faster than the naive algorithm in all cases and the SpGEMM algorithm with filtration in most cases (especially with large $s$ value).

hypergraph algorithms, graph algorithms, parallel ↗

Extreme-scale stochastic optimization and simulation via learning-enhanced decomposition and parallelization (Final Technical Report)

Stochastic optimization and simulation models ubiquitously arise in designing and operating complex service/engineering systems. They can be extreme in scale due to high-dimensional data and decisions, and can also involve decisions made sequentially in response to newly revealed data, both causing significant computational challenge. The objective of this research is to explore a unified framework that integrates machine learning with discrete optimization and risk-averse modeling, to improve the efficiency of decomposition paradigms for stochastic optimization and simulations at extreme scale. The models we consider represent a broad class of complex decision-making problems, where 0-1 or continuous decisions are made before and/or after knowing multiple sources of uncertainties that could be correlated. We will employ machine learning methods to dynamically decide and prioritize computational procedures, including cut generation, branching, and bounding of the optimal objective. Furthermore, the research will shed new lights on the traditional decomposition algorithms for extreme-scale computing. Deliverables of the research include new modeling and computational methods for advancing the state-of-the-art research in optimization and simulation, bringing many relevant risk-averse, data-driven optimization problems in practice within the range of tractability. Examples include distributed computing server scheduling and sensor deployment for monitoring critical infrastructures. Success in this effort will enable progress in solving multiple extreme-scale problems in the complex system design and operations arising from DoE missions in energy, environment, and national security.

24 POWER TRANSMISSION AND DISTRIBUTION↗

PRO-X Parallelization Study

The proliferation resistance optimization (PRO-X) program is actively supporting the design of nuclear systems by developing a framework to both optimize the fuel cycle infrastructure for nuclear reactor (including both advanced reactors (ARs) and research reactors (RRs)) and minimize the potential for production of weapons-usable nuclear material (Figure 1). One area of interest is in the impact a modular approach to bulk handling fuel cycle facilities could have on meeting safeguards requirements to identify future areas of growth within the proliferation resistance space. This study evaluates how changing the number of streams within a fuel cycle facility could impact a facilities ability to meet both domestic and international safeguards requirements.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Systems and methods for tensor scheduling

A technique for efficient scheduling of operations in a program for parallelized execution thereof using a multi-processor runtime environment having two or more processors includes constraining the type or number of loop optimization transforms that may be explored such that memory and processing capacity available for the scheduling task are not exceeded, while facilitating a tradeoff between memory locality, parallelization, and/or data communication between memory modules of the multi-processor runtime environment.

Meister, Benoit J.↗

Transformational Challenge Reactor preconceptual core design studies

In the nuclear industry, a manufacturing-informed design approach has the potential to yield the most benefit from advanced manufacturing. By leveraging advanced materials, data science, and rapid testing and deployment, manufacturing-informed design can drive down costs and development times, ultimately improving future commercial viability. This approach is being demonstrated in the US Department of Energy Office of Nuclear Energy (DOE-NE) Transformational Challenge Reactor (TCR) program. Preconceptual design activities for TCR have been focused on analyzing and maturing four reactor core design concepts: two fast-spectrum and two thermal-spectrum systems. Furthermore, the designs were iteratively modified and analyzed, and subcomponents were manufactured in parallel over weeks instead of months or years. To meet key program initiatives (e.g., timeline and material use), several constraints—including fissile material availability, component availability, materials compatibility, and additive manufacturing capabilities—were factored into the design effort, yielding small cores less than one cubic meter in volume with near-term viability. Additionally, the TCR program has made significant progress on development of advanced moderator materials such as yttrium hydride, advancing the feasibility of gas-cooled thermal spectrum systems using less than 250 kg of high-assay low enriched uranium (HALEU) and occupying less than 1 m3. Each of the two resulting thermal designs uses a different fuel form: traditional UO2 ceramic fuel and tristructural isotropic (advanced TRISO) fuel particles embedded inside a SiC matrix. Core neutronics and thermal performance for these systems were assessed and summarized. Evaluation of the performance metrics for these two moderated designs has yielded the downselected TCR design: a TRISO-fueled and yttrium hydride moderated gas-cooled reactor.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Engine and fuel cell system including first and second turbochargers

An engine system includes an internal combustion engine, a fuel cell system, a first turbocharger and a second turbocharger. The internal combustion engine has an intake passage, and a first exhaust passage fluidly connected to the first set of combustion chambers. The first turbocharger has a first compressor and a first turbine. The second turbocharger has a second compressor and a second turbine, the second compressor connected in series with the first compressor, and the second turbine being in fluid communication with the second exhaust passage. The first and second turbines are connected in parallel such that the first turbine only receives exhaust flow from the fuel cell system, and the second turbine only receives exhaust flow from the internal combustion engine.

Lusardi, Christopher↗

Parallel performance of algebraic multigrid domain decomposition

Algebraic multigrid (AMG) is a widely used scalable solver and preconditioner for large-scale linear systems resulting from the discretization of a wide class of elliptic PDEs. While AMG has optimal computational complexity, the cost of communication has become a significant bottleneck that limits its scalability as processor counts continue to grow on modern machines. This article examines the design, implementation, and parallel performance of a novel algorithm, algebraic multigrid domain decomposition (AMG-DD), designed specifically to limit communication. The goal of AMG-DD is to provide a low-communication alternative to standard AMG V-cycles by trading some additional computational overhead for a significant reduction in communication cost. Numerical results show that AMG-DD achieves superior accuracy per communication cost compared with AMG, and speedup over AMG is demonstrated on a large GPU cluster.

97 MATHEMATICS AND COMPUTING↗

Neutron transport methods for multiphysics heterogeneous reactor core simulation in Griffin

Griffin is a reactor physics application based on the Multiphysics Object-Oriented Simulation Environment (MOOSE). This work discloses the methods, algorithms, and implementation for simulating heterogeneous reactor dynamics models. Griffin utilizes a discontinuous finite-element method with discrete ordinates (DFEM-S ) to discretize the field variable of the multigroup neutron transport equation. Multiphysics feedback is handled using two-step tabulated cross-section methodology. Feedback quantities are evaluated using the MOOSE-MultiApp system to couple various engineering phenomena, such as heat conduction and thermal fluids. The multiphysics DFEM-S system is solved using fixed-point iteration with a fully asynchronous parallel sweeper, unstructured coarse-mesh finite difference acceleration, and a multi-timescale improved quasi-static method scheme. The implementation is applied to a multiphysics microreactor model, with two transients: one initiated by a single heat-pipe failure and another by control drum rotation. Importantly, these examples demonstrate the ability of Griffin to tractably solve the neutron transport equation considering seven independent variables and feedback.

97 MATHEMATICS AND COMPUTING↗

Second Generation Readout For Large Format Photon Counting Microwave Kinetic Inductance Detectors

We present the development of a second generation digital readout system for photon counting microwave kinetic inductance detector (MKID) arrays operating in the optical and near-infrared wavelength bands. Our system retains much of the core signal processing architecture from the first generation system but with a significantly higher bandwidth, enabling the readout of kilopixel MKID arrays. Each set of readout boards is capable of reading out 1024 MKID pixels multiplexed over 2 GHz of bandwidth; two such units can be placed in parallel to read out a full 2048 pixel microwave feedline over a 4 GHz–8 GHz band. As in the first generation readout, our system is capable of identifying, analyzing, and recording photon detection events in real time with a time resolution of order a few microseconds. Here, we describe the hardware and firmware, and present an analysis of the noise properties of the system. We also present a novel algorithm for efficiently suppressing IQ mixer sidebands to below −30 dBc.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Equilipy: a python package for calculating phase equilibria

The CALPHAD (CALculation of PHAse Diagram) approach (Nigel Saunders & Miodownik, 1998) provides predictions for thermodynamically stable phases in multicomponent-multiphase materials across a wide range of temperatures. Consequently, the CALPHAD calculations became an essential tool in materials and process design (Luo, 2015). Such design tasks frequently require navigating a high-dimensional space due to multiple components involved in the system. This increasing complexity demands high-throughput CALPHAD calculations, especially in the rapidly evolving field of alloy design. In response to the need, we developed Equilipy an open-source Python package designed for calculating phase equilibria of multicomponent-multiphase systems. Equilipy is specifically tailored for high-throughput CALPHAD calculations, offering parallel computations across multiple processors and nodes with the given NPT input conditions namely elemental compositions (N), pressure (P), and temperature (T). Equilipy utilizes the program structure and Gibbs energy functions from the Fortran-based program, Thermochimica (Piro et al., 2013), with incorporating a new Gibbs energy minimization algorithm. This algorithm, originally developed by Capitani and Brown in 1987 (Capitani & Brown, 1987), has been revised and implemented to enhance the stability and performance of calculations. The Fortran codes are precompiled and interfaced with Python via F2PY, ensuring high computation speed. Benchmark tests shown in Figure 1 demonstrate that Equilipy’s computation speed is comparable to those of established commercial software, TC-Python and PanPython. This result highlights its efficiency and potential applications in various scientific and industrial fields.

97 MATHEMATICS AND COMPUTING↗

Quantum-inspired tempering for ground state approximation using artificial neural networks

A large body of work has demonstrated that parameterized artificial neural networks (ANNs) can efficiently describe ground states of numerous interesting quantum many-body Hamiltonians. However, the standard variational algorithms used to update or train the ANN parameters can get trapped in local minima, especially for frustrated systems and even if the representation is sufficiently expressive. We propose a parallel tempering method that facilitates escape from such local minima. This methods involves training multiple ANNs independently, with each simulation governed by a Hamiltonian with a different "driver" strength, in analogy to quantum parallel tempering, and it incorporates an update step into the training that allows for the exchange of neighboring ANN configurations. We study instances from two classes of Hamiltonians to demonstrate the utility of our approach using Restricted Boltzmann Machines as our parameterized ANN. The first instance is based on a permutation-invariant Hamiltonian whose landscape stymies the standard training algorithm by drawing it increasingly to a false local minimum. The second instance is four hydrogen atoms arranged in a rectangle, which is an instance of the second quantized electronic structure Hamiltonian discretized using Gaussian basis functions. We study this problem in a minimal basis set, which exhibits false minima that can trap the standard variational algorithm despite the problem’s small size. We show that augmenting the training with quantum parallel tempering becomes useful to finding good approximations to the ground states of these problem instances.

Albash, Tameem↗

UPC++ v1.0 Specification, Revision 2020.10.0

UPC++ is a C++11 library providing classes and functions that support Partitioned Global Address Space (PGAS) programming. The key communication facilities in UPC++ are one-sided Remote Memory Access (RMA) and Remote Procedure Call (RPC). All communication operations are syntactically explicit and default to non-blocking; asynchrony is managed through the use of futures, promises and continuation callbacks, enabling the programmer to construct a graph of operations to execute asynchronously as high-latency dependencies are satisfied. A global pointer abstraction provides system-wide addressability of shared memory, including host and accelerator memories. The parallelism model is primarily process-based, but the interface is thread-safe and designed to allow efficient and expressive use in multi-threaded applications. The interface is designed for extreme scalability throughout, and deliberately avoids design features that could inhibit scalability.

97 MATHEMATICS AND COMPUTING↗