Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “parallel systems”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 361 records · Page 20

Resilient Inverter-Driven Black Start with Collective Parallel Grid-Forming Operation: Preprint

As the modern power systems are experiencing exceptional changes with increasing penetration of inverter-based resources (IBRs), system restoration using IBRs has received attention. Using local grid-forming (GFM) assets near consumers, engineered to establish grid voltages in the absence of a stiff grid, i.e., bottom-up restoration, a distribution system would obtain high system resilience, by not relying on the bulk power system restoration requiring significant human intervention and procedure to restore. This paper studies the technical feasibility of the novel approach with detailed electromagnetic transient (EMT) simulations. To thoroughly evaluate the potential of GFM inverters and technical challenges in IBR-driven black start, a detailed three-phase inverter model is developed, including negative-sequence control for voltage balance and a phase-by-phase current limiter to sustain momentary overloading during the black start. To examine dynamic aspects of the black start process, the EMT simulation also models transformer and motor dynamics to emulate their inrush and start-up behaviors as well as network dynamics. In addition, active involvement of grid-following distributed energy resources (DER) is also studied to facilitate the black start process. It is shown that, by allowing multiple GFM inverters to collectively black start without master-slave coordination, a system can achieve high resilience even with a fraction of assets lost. Two test cases of inverter-driven black start using two and one GFM inverters, respectively, for a heavily unbalanced 2-MVA distribution feeder are demonstrated. Takeaways for further study and field deployment are provided.

grid-forming inverter↗

Modernization of the Radiation Measurements Laboratory at the Advanced Test Reactor Complex

The Advanced Test Reactor (ATR) at the Idaho National Laboratory (INL) is a unique, water-cooled, high-flux test reactor capable of performing tests prototypical of PWR operating conditions. The Radiation Measurements Laboratory (RML) was founded in the 1960s to support reactor operations and to conduct independent scientific research. For nearly six decades RML has performed four primary functions: monitoring of radioactivity by gamma-ray spectroscopy of routine reactor samples, fluence rate determinations for irradiation cycles, fission-rate measurements for the ATR-Critical (ATR-C) Facility, and independent research and development of radiation detection systems and applications. Through the decades, RML has seen technological advancements that have been integrated into each of the critical functions of the laboratory. However, many of the measurement and analysis systems employed to this day can be improved by modernization. Recent progress at RML is improving reliability and accuracy of the radiation measurements performed in support of nuclear energy research for the U.S.A. Department of Energy. The control and data collection systems supporting ATR-C have been upgraded. New High-Purity Germanium (HPGe) spectrometers have been procured with liquid nitrogen recycling capabilities to improve up-time and reduce measurement uncertainties. Fluence-rate measurement techniques are also being improved to ensure accuracy, avoid systemic errors, and identify biases. In a parallel effort, new scientific research avenues are being explored which will provide an opportunity to further enhance the utilization of ATR and ensure the sustainability of the RML as nuclear research continues to evolve. The RML is improving the effectiveness of irradiation services provided by ATR while ensuring a sustainable future for nuclear energy research.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Towards scaling community detection on distributed-memory heterogeneous systems

Distributed multi-GPU systems pose significant challenges and opportunities for efficient execution of parallel applications. Graph algorithms are generally characterized by irregular memory accesses, low computation to communication ratios, and load balancing problems that are especially hard to address on multi-GPU systems. Graph community detection is an important problem in the emerging domain of graph analytics with numerous applications. In this paper, we present our ongoing work on distributed-memory multi-GPU implementation for graph community detection. Our work parallelizes the widely used (albeit serial) Louvain method on distributed multi-GPU platforms. Supported by an extensive set of experiments on a multi-GPU enabled supercomputer (OLCF Summit) and a single compute node (Nvidia DGX-2®), we demonstrate competitive performance to existing distributed-memory CPU-based implementation, and up to 6.5 better results than Nvidia RAPIDS® CUGRAPH. To the best of our knowledge, this work represents the first effort for community detection on distributed multi-GPU systems. Our approach and related findings can be extended to numerous other iterative graph algorithms on multi-GPU systems.

97 MATHEMATICS AND COMPUTING↗

A Segregated Approach for Modeling the Electrochemistry in the 3-D Microstructure of Li-Ion Batteries and Its Acceleration Using Block Preconditioners

Abstract Battery performance is strongly correlated with electrode microstructure. Electrode materials for lithium-ion batteries have complex microstructure geometries that require millions of degrees of freedom to solve the electrochemical system at the microstructure scale. A fast-iterative solver with an appropriate preconditioner is then required to simulate large representative volume in a reasonable time. In this work, a finite element electrochemical model is developed to resolve the concentration and potential within the electrode active materials and the electrolyte domains at the microstructure scale, with an emphasis on numerical stability and scaling performances. The block Gauss-Seidel (BGS) numerical method is implemented because the system of equations within the electrodes is coupled only through the nonlinear Butler–Volmer equation, which governs the electrochemical reaction at the interface between the domains. The best solution strategy found in this work consists of splitting the system into two blocks—one for the concentration and one for the potential field—and then performing block generalized minimal residual preconditioned with algebraic multigrid, using the FEniCS and the Portable, Extensible Toolkit for Scientific Computation libraries. Significant improvements in terms of time to solution (six times faster) and memory usage (halving) are achieved compared with the MUltifrontal Massively Parallel sparse direct Solver. Additionally, BGS experiences decent strong parallel scaling within the electrode domains. Last, the system of equations is modified to specifically address numerical instability induced by electrolyte depletion, which is particularly valuable for simulating fast-charge scenarios relevant for automotive application.

25 ENERGY STORAGE↗

FPDeep: Scalable Acceleration of CNN Training on Deeply-Pipelined FPGA Clusters

In this paper, we propose a framework called FPDeep, which uses a hybrid of model and layer paral- lelism to configure distributed reconfigurable clusters to train DNNs. This approach has numerous benefits. First, the design does not suffer from batch size growth. Second, novel workload and weight partitioning leads to balanced loads of both among nodes. And third, the entire system is fine-grained pipeline. This leads to high parallelism and utilization and also minimizes the time features need to be cached while waiting for back-propagation.

Wang, Tianqi↗

Reconfigurable unitary transformations of optical beam arrays

Spatial transformations of light are ubiquitous in optics, with examples ranging from simple imaging with a lens to quantum and classical information processing in waveguide meshes. Multi-plane light converter (MPLC) systems have emerged as a platform that promises completely general spatial transformations, i.e., a universal unitary. However, until now, MPLC systems have demonstrated transformations that are far from general, e.g., converting from a Gaussian to Laguerre-Gauss mode. Here, we demonstrate the promise of an MLPC, the ability to impose an arbitrary unitary transformation that can be reconfigured dynamically. Specifically, we consider transformations on superpositions of parallel free-space beams arranged in an array, which is a common information encoding in photonics. We experimentally test the full gamut of unitary transformations for a system of two parallel beams and make a map of their fidelity. We obtain an average transformation fidelity of 0.85 ± 0.03. This high-fidelity suggests that MPLCs are a useful tool for implementing the unitary transformations that comprise quantum and classical information processing.

47 OTHER INSTRUMENTATION↗

CCM vs. CRM Design Optimization of a Boost-derived Parallel Active Power Decoupler for Microinverter Applications

Single-phase inverter or rectifier systems often make use of an auxiliary active power decoupler (APD) to balance the mismatch between steady DC power and fluctuating AC power. This paper deals with efficiency and size optimization of a parallel boost-type APD circuit for PV microinverter applications. Specifically, design of an eGaNFET-based, 400 W APD circuit, employing planar inductor and operating in either continuous conduction mode (CCM) or critical conduction mode (CRM) is considered. Available design variables including inductance value, inductor core geometry, capacitor voltage, switching frequency, and modulation scheme (CCM vs. CRM) are explored to identify Pareto-optimal configurations, which can achieve low California Energy Commission (CEC) efficiency drop while also reducing the footprint area of the inductor. The theoretical study predicts that the optimal CRM design can achieve 37% reduced inductor size, while operating with similar efficiency drop, compared to the optimal CCM design. Experimental results, obtained using two separate 40 V, 400 W hardware prototypes for CCM and CRM, are presented to verify the analyses.

14 SOLAR ENERGY↗

Phasing of seven-channel fibre laser radiation with dynamic turbulent phase distortions using a stochastic parallel gradient algorithm at a bandwidth of 450 kHz

We have demonstrated an experimental setup for the coherent phasing of a seven-channel fibre laser system ( λ = 1064 nm) in a scheme comprising a master oscillator and a set of parallel amplifiers with lithium niobate-based fibre-optic phase modulators. Using a stochastic parallel gradient algorithm, an instrumental phase modulator control unit ensures a bandwidth of the system up to 450 kHz. The effectiveness of phasing of light transmitted through a turbulent medium with a characteristic time scale τ{sub turb} has been studied experimentally as a function of phasing time τ{sub ph}. The results demonstrate that the average Strehl ratio begins to rise at τ{sub turb}/τ{sub ph} ⩾ 2 and that the effectiveness of compensation for dynamic phase distortions in the beam propagation path rises sharply at τ{sub turb}/τ{sub ph} ≈ 20. For τ{sub turb}/τ{sub ph} ⩾ 30 – 40, the average Strehl ratio remains constant at the level reached. (control of laser radiation parameters)

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

SpectrumSDT: A program for parallel calculation of coupled rotational-vibrational energies and lifetimes of bound states and scattering resonances in triatomic systems

In this work, we present SpectrumSDT – a program for calculations of energies and lifetimes of bound rotational-vibrational states below and scattering resonances above the dissociation threshold on a global potential energy surface of a triatomic system, which may include stable molecules, weekly-bound van-der-Waals complexes, and unbound atom + diatom scattering systems. Large-amplitude vibrational motion is treated explicitly using hyper-spherical coordinates. Three options for the rotational-vibrational interaction are supported: uncoupled (symmetric top rotor), partially coupled (to include interaction between several nearest states only) and full-coupled (vibrating asymmetric-top rotor). In addition to energies and lifetimes, SpectrumSDT is able to integrate ro-vibrational wave functions over the user-defined regions of potential energy surface, which helps to classify these states. In this release of the code, SpectrumSDT is limited to ABA-type molecules with wave functions that do not extend into the regions near Eckart singularities.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Performance Evaluation of Different Parallel Programming Models in SCALE-Shift Sequences for Criticality and Shielding Applications [Abstract]

The SCALE code system has been widely used for nuclear criticality safety, reactor physics, radiation shielding, source term generation, and inventory analyses by researchers, industry, and regulatory bodies. Although limited support for shared- and distributed-memory parallel processing was introduced via C++ threading, OpenMP, and MPI, a hybrid parallel programming model with both distributed- and shared-memory parallelism has not been fully supported in the SCALE code system.

Nuclear Criticality Safety Program (NCSP)↗

Distributed memory, GPU accelerated Fock construction for hybrid, Gaussian basis density functional theory

With the growing reliance of modern supercomputers on accelerator-based architecture such a graphics processing units (GPUs), the development and optimization of electronic structure methods to exploit these massively parallel resources has become a recent priority. While significant strides have been made in the development GPU accelerated, distributed memory algorithms for many modern electronic structure methods, the primary focus of GPU development for Gaussian basis atomic orbital methods has been for shared memory systems with only a handful of examples pursing massive parallelism. Here in this work, we present a set of distributed memory algorithms for the evaluation of the Coulomb and exact exchange matrices for hybrid Kohn–Sham DFT with Gaussian basis sets via direct density-fitted (DF-J-Engine) and seminumerical (sn-K) methods, respectively. The absolute performance and strong scalability of the developed methods are demonstrated on systems ranging from a few hundred to over one thousand atoms using up to 128 NVIDIA A100 GPUs on the Perlmutter supercomputer.

97 MATHEMATICS AND COMPUTING↗

DeepThermo: Deep Learning Accelerated Parallel Monte Carlo Sampling for Thermodynamics Evaluation of High Entropy Alloys

Since the introduction of Metropolis Monte Carlo (MC) sampling, it and its variants have become standard tools used for thermodynamics evaluations of physical systems. However, a long-standing problem that hinders the effectiveness and efficiency of MC sampling is the lack of a generic method (a.k.a. MC proposal) to update the system configurations. Consequently, current practices are not scalable. Here we propose a parallel MC sampling framework for thermodynamics evaluation—DeepThermo. By using deep learning–based MC proposals that can globally update the system configurations, we show that DeepThermo can effectively evaluate the phase transition behaviors of high entropy alloys, which have an astronomical configuration space. For the first time, we directly evaluate a density of states expanding over a range of ~e 10,000 for a real material. We also demonstrate DeepThermo’s performance and scalability up to 3,000 GPUs on both NVIDIA V100 and AMD MI250X-based supercomputers.

Yin, Junqi↗

Using Parameter Sweep in WaterTAP to Analyze New Water Treatment Technologies

We describe a powerful and generalized parameter sweep tool in this report that was originally developed to analyze the performance of existing and novel water treatment models being developed in WaterTAP. Since WaterTAP is built upon IDAES and Pyomo, the parameter sweep tool can be used to systematically explore and debug the behavior of most Pyomo and IDAES numerical models. In order to enable meaningful analyses, the parameter sweep tool has been designed with the following features: 1) Model flexibility: The parameter sweep tool does not enforce any restrictions on the types of models that can be used with it. As long as a Pyomo model can be solved and the parameter is active and mutable, the tool only needs functions that describe how to run the model, the sweep parameters, and the output quantities of interest. 2) Flexible sampling: The parameter sweep tool has inbuilt functions to generate samples from a random distribution or a multidimensional Euclidean space. Furthermore, the users have to ability to supply samples generated from a tool of their choice. 3) Multiple sweep types: A user can choose from one of 3 types of parameter sweeps depending on their needs. 4) Detailed outputs: Outputs generated by the parameter sweep tool can be stored in detailed H5 file or user-friendly CSV files for post processing. 5) Parallel computing: The parameter sweep supports shared and distributed memory parallel computing to enable the use of high performance computers (HPC) for large-scale analyses. 6) Modular: The parameter sweep tool is self-contained and can easily be integrated within an outer-loop analysis or as desired by the user. 7) Ease of use: The tool is well documented and a simple sweep can be easily executed by following the online documentation in a few lines of code. We demonstrate the use of the parameter sweep tool on a simple water treatment system from the WaterTAP repository and show its parallel scaling performance on an Apple laptop and NREL's Eagle HPC. The parameter sweep tool is actively being used with models currently being developed within WaterTAP and we expect its use to grow beyond it to other IDAES and Pyomo models.

97 MATHEMATICS AND COMPUTING↗

Similarity for downscaled kinetic simulations of electrostatic plasmas: Reconciling the large system size with small Debye length

A simple similarity has been proposed for kinetic (e.g., particle-in-cell) simulations of plasma transport that can effectively address the long-standing challenge of reconciling the tiny Debye length with the vast system size. This applies to both transport in unmagnetized plasma and parallel transport in magnetized plasmas, where the characteristics length scales are given by the Debye length, collisional mean free paths, and the system or gradient lengths. The controlled scaled variables are the configuration space, x/L, and an artificial Coulomb Logarithm, L ln Λ, for collisions, while the scaled time, t/L, and electric field, LE, are automatic outcomes. The similarity properties are examined, demonstrating that the macroscopic transport physics is preserved through a similarity transformation while keeping the microscopic physics at its original scale of Debye length. To showcase the utility of this approach, two examples of 1D plasma transport problems were simulated using the VPIC code: the plasma thermal quench in tokamaks [Li et al., Nuclear Fusion 63, 066030 (2023)] and the plasma sheath in the high-recycling regime [Li et al., Physics of Plasmas 30, 063505 (2023)].

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Comparison of CCM- and CRM-Based Boost Parallel Active Power Decoupler for PV Microinverter

Single-phase inverter or rectifier systems often make use of an active power decoupler (APD) to balance the mismatch between constant dc power and fluctuating ac power. This article deals with the comparison of continuous conduction mode (CCM) and critical conduction mode (CRM) operation-based design of a parallel boost-type APD for photovoltaic microinverter applications. From a design perspective, multiobjective analysis of efficiency, volume, and cost is explored within a decision space including planar inductors, gallium nitride based devices, film capacitors, switching frequency, and modulation (CCM vs. CRM). The theoretical study analyzes all possible design configurations within CCM and CRM and identifies Pareto-optimal designs, from which the selected CRM design can achieve reduced system volume and lower cost with the use of smaller inductor core, while operating with similar California Energy Commission efficiency drop as the selected CCM design. From a control perspective, a pulsewidth modulation based control strategy is proposed to implement closed-loop CRM modulation that does not rely on zero-crossing detection. Furthermore, closed-loop systems are designed for the optimal CCM and CRM realizations, and the final system characteristics are compared. Experimental results, obtained using two separate 40-V, 400-W hardware prototypes for CCM and CRM, are presented to verify the analyses.

42 ENGINEERING↗

Quantum model learning agent: characterisation of quantum systems through machine learning

Accurate models of real quantum systems are important for investigating their behaviour, yet are difficult to distil empirically. Here, we report an algorithm—the quantum model learning agent (QMLA)—to reverse engineer Hamiltonian descriptions of a target system. We test the performance of QMLA on a number of simulated experiments, demonstrating several mechanisms for the design of candidate Hamiltonian models and simultaneously entertaining numerous hypotheses about the nature of the physical interactions governing the system under study. QMLA is shown to identify the true model in the majority of instances, when provided with limited a priori information, and control of the experimental setup. Our protocol can explore Ising, Heisenberg and Hubbard families of models in parallel, reliably identifying the family which best describes the system dynamics. We demonstrate QMLA operating on large model spaces by incorporating a genetic algorithm to formulate new hypothetical models. The selection of models whose features propagate to the next generation is based upon an objective function inspired by the Elo rating scheme, typically used to rate competitors in games such as chess and football. In all instances, our protocol finds models that exhibit F 1 score ≥ 0.88 when compared with the true model, and it precisely identifies the true model in 72% of cases, whilst exploring a space of over 250 000 potential models. By testing which interactions actually occur in the target system, QMLA is a viable tool for both the exploration of fundamental physics and the characterisation and calibration of quantum devices.

97 MATHEMATICS AND COMPUTING↗

Fabricating double-sided electrodynamic screen films

An electrode film is fabricated by using a flexographic printing system to print a first electrode pattern including a first set of parallel electrodes and a second electrode pattern including a second set of parallel electrodes onto a first surface of a transparent film using a catalytic ink, wherein elements of the first electrode pattern do not cross over elements of the second electrode pattern. The flexographic printing system prints a third electrode pattern including a third set of parallel electrodes onto a second surface of the transparent film using the catalytic ink, wherein the first, second and third sets of parallel electrodes are arranged in an interlaced pattern. A conductive metallic material is electrolessly plated onto the catalytic ink such that the elements of the first, second and third electrode patterns become conductive.

36 MATERIALS SCIENCE↗

Situational Awareness of Grid Anomalies (SAGA)

The modern power industry becomes more vulnerable to cyber events due to the growing interconnectivity, interdependence, and complexity of the electric power grid. High-fidelity modeling and simulation tools that support the preventative risk analysis on potential cyber-relevant events is essential for ensuring the situational awareness of the system operator as it provides an inexpensive and risk-free environment to test the system responses under various cyber-relevant events and hereby can support research on cyber anomaly detection, optimal protective resource allocation, and mitigation measures. In this webinar, we will share NREL's cybersecurity research capabilities by highlighting the development of a scalable cyber-physical event test bed and demonstration with real hardware in the loop. The developed cyber-physical event test bed is backboned by an integrated transmission, distribution, and communication dynamic co-simulation framework and a plug-and-play cyber event generation module. It is designed to be modular and compatible with parallel computing, and thereby supports large-scale system simulations at an affordable computation cost. The test bed can capture millisecond-to-minutes dynamic frequency and voltage responses under cyber events from the bulk transmission system to the active distribution systems and distributed energy resources at the grid edge.

co-simulation↗