Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “parallel systems”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 379 records · Page 21

SiC-Based 5-kV Universal Modular Soft-Switching Solid-State Transformer (M-S4T) for Medium-Voltage DC Microgrids and Distribution Grids

Medium-voltage DC (MVDC) grids are attractive for electric aircraft and ship power systems, battery energy storage system (BESS), fast charging electric vehicle (EV), etc. Such EV or BESS applications need isolated bidirectional MVDC to LVDC or LVAC converters. However, the existing Si-based solutions cannot fulfill the requirements of a high-efficiency and robust converter for MVDC grids. This paper presents a 5 kV SiC-based universal modular solid-state transformer (SST). This universal current-source SST can interface either a LVAC or LVDC grid with a MVDC grid in single-stage power conversion, while the conventional dual active bridge (DAB) converter needs an additional inverter. The proposed SST module using 3.3 kV SiC MOSFETs and diodes is bidirectional and can serve as a building block in series or parallel for higher-voltage higher-power systems. The topology of each module is based on the soft-switching solid-state transformer (S4T) with reduced conduction loss, which features reduced EMI through controlled dv/dt, and high efficiency with full-range ZVS for main devices and ZCS for auxiliary devices. Operation principle of the modular S4T (M-S4T), capacitor voltage balancing control between the cascaded modules, design of components including a medium-voltage (MV) medium-frequency transformer (MFT) to realize a 50 kVA 5 kV DC to 600 V DC or 480 V AC M-S4T are presented. Importantly, the MV MFT prototype achieves very low leakage inductance (0.13%) and 15 kV insulation with coaxial cables and nanocrystalline cores. Here, the proposed universal modular SST is compared against the DAB solution and verified with DC-DC and DC-AC simulation and 4 kV experimental results. Significantly, the MV experimental results of a modular DC transformer with each module at MVDC are rarely covered in the literature and reported for the first time.

14 SOLAR ENERGY↗

Evaluation of PETSc on a Heterogeneous Architecture, the OLCF Summit System: Part II - Basic Communication Performance

Nearest-neighbor communication is at the heart of many high-performance parallel computations. We report on the performance of such communication on the Oak Ridge Leadership Computing Facility system Summit in the context of the PETSc communication module. The analysis in this report includes basic Ping-Pong point-to-point communication and regular and irregular nearest-neighbor communication.

97 MATHEMATICS AND COMPUTING↗

Dynamic magneto-chiral instability in photoexcited tellurium

In systems of charged chiral fermions out of equilibrium, an electric current parallel to a magnetic field can generate a dynamic instability that amplifies electromagnetic waves. Whether this mechanism also operates in chiral solid-state systems has remained uncertain. Here we observe signatures of a dynamic magneto-chiral instability in elemental tellurium, a structurally chiral crystal, using time-domain terahertz emission spectroscopy. Under transient photoexcitation in a moderate magnetic field, we observe terahertz radiation with coherent modes that grow in amplitude over time. We present a theoretical model that describes this behaviour based on a dynamic instability of electromagnetic waves interacting with infrared-active oscillators of acceptor states in tellurium, giving rise to an amplifying polariton. These results demonstrate that magneto-chiral instabilities can emerge in solid-state systems and establish a mechanism for terahertz-wave amplification in chiral materials.

42 ENGINEERING↗

Reinforcement Learning for Load-balanced Parallel Particle Tracing

We explore an online reinforcement learning (RL) paradigm to dynamically optimize parallel particle tracing performance in distributed-memory systems. Our method combines three novel components: (1) a work donation algorithm, (2) a high-order workload estimation model, and (3) a communication cost model. First, we design an RL-based work donation algorithm. Our algorithm monitors workloads of processes and creates RL agents to donate data blocks and particles from high-workload processes to low-workload processes to minimize program execution time. The agents learn the donation strategy on the fly based on reward and cost functions designed to consider processes' workload changes and data transfer costs of donation actions. Second, we propose a workload estimation model, helping RL agents estimate the workload distribution of processes in future computations. Third, we design a communication cost model that considers both block and particle data exchange costs, helping RL agents make effective decisions with minimized communication costs. We demonstrate that our algorithm adapts to different flow behaviors in large-scale fluid dynamics, ocean, and weather simulation data. Our algorithm improves parallel particle tracing performance in terms of parallel efficiency, load balance, and costs of I/O and communication for evaluations with up to 16,384 processors.

Distributed and parallel particle tracing↗

Designing the Insulation System for Motors in Electrified Aircraft: Optimization, Partial Discharge Issues and Use of Advanced Materials

Designing the insulation system for motors to be used in electrical aircraft requires efforts for maximizing specific power, but, in parallel, particular attention to achieve high reliability. As a major harm for organic insulation systems is partial discharges, design must be able to infer their likelihood during any operation stage and handle their potential inception. This paper proposes a new approach to carry out optimized or conservative insulation system designs which can provide the specified life at the chosen failure probability as well as look at the option of possibly reducing the risk of partial discharges to zero, at any altitude. Examples of designing turn, phase to ground and phase-to-phase insulation systems are reported, with cases where the design can be optimized and other cases where the optimized design does not pass IEC testing standard. Therefore, the limits for design feasibility as a function of the required level of safety and reliability are discussed, showing that the presence of partial discharges cannot be always avoided even through conservative design criteria. Therefore, the use of advanced, corona-resistant materials must be considered, in order to reach a higher, sometimes redundant, level of reliability.

Ramin, Robin (ORCID:0000000211455006)↗

Introduction: Neuromorphic Materials

The explosive growth in data collection and the need to process it efficiently, as well as the desire to automate increasingly complex tasks in transportation, medical care, manufacturing, security and many other fields have motivated a growing interest in neuromorphic computing. Unlike the binary, transistorbased ON/OFF logic gates and separate logic and memory functionalities employed in digital computing, neuromorphic computing is inspired by animal brains that use interconnected synapses and neurons to perform processing, storage and transmission of information at the same location, while only consuming ~20 W or less of power. Motivated by the brain’s efficiency, adaptability, self-learning and resiliency qualities, neuromorphic computing can be broadly defined as an approach to processing and storing information using hardware and algorithms inspired by models of biological neural systems. Present research in neuromorphic computing encompasses approaches that vary significantly in their degree of neuro-inspiration, from systems that only incorporate features such as asynchronous, event-driven operation or use crossbar arrays of non-volatile memory (NVM) elements to accelerate deep neural networks (DNNs), to designs that embrace the extreme parallelism, sparsity, reconfigurability, adaptability, complexity and stochasticity observed in nervous systems. The term ‘neuromorphic’ computing is often credited to Carver Mead, who in the 1980s investigated Si-based analog electronics to replicate functions of the animal retina. Earlier important advances in this field include the work of Frank Rosenblatt, who proposed the concept of the perceptron, Bernard Widrow, who used this concept to build one of the first analog neural networks, the Adaline and many other researchers (see ref. 6 for an historical perspective on neuromorphic computing). With the recent increase in the use of artificial intelligence and large language models, and rising concerns over the associated energy costs, interest in neuromorphic hardware has expanded rapidly. According to some estimates, driven largely by the drastic growth in the training use of artificial intelligence (AI) models using the current computing architectures, the energy cost of computing is projected to reach the energy supply worldwide by 2045. Furthermore, while this is not a realistic outcome, it means that, if more efficient computing technologies are not developed -- soon -- the world will soon become one where demand for energy and market constraints limit the continued increase of societal access to AI and cloud services from data centers. Data centers used for training and use of these models consume hundreds of terawatt hours of electricity, already past 4% of the US electricity demand.

Circuits↗

Segmented Energy Routing for a Modular AC/DC Hybrid System

This article presents a modular ac/dc system with both distributed and centralized power ports for energy router (ER) applications. In each module of the described system, photovoltaic (PV) power generation units, battery-type energy storage (ES) units, and critical loads are connected to the cascaded H-bridge (CHB)-organized medium-voltage (MV) dc links, with fully distributed low-voltage (LV) dc power ports. Copies of modules share the centralized load bus and interact with an MV ac grid in parallel. Hybrid power port (HPO) assigns flexibility to the system but makes energy routings a necessity for stable operation. In this article, a segmented energy management strategy for the HPO-ER is proposed. In terms of the grid-side power transferring, the system ratings are intentionally designed to match significant power imbalance. Focally, a segmented energy management strategy is proposed to realize fully autonomous energy routing involving MV ac grid, distributed PV generations, distributed battery storages, and LV load. The system is proven to be stable using the derived impedance-based model considering the interaction between power blocks. The feasibility of the topology and control strategy is also verified through software simulation, laboratorial hardware prototype experiment, and hardware-in-the-loop (HIL) emulation.

24 POWER TRANSMISSION AND DISTRIBUTION↗

PyOMP: Multithreaded Parallel Programming in Python

We know that Python is a widely used language in scientific computing. When the goal is high performance, however, Python lags far behind low-level languages such as C and Fortran. To support applications that stress performance, Python needs to access the full capabilities of modern CPUs. That means support for parallel multithreading. In this paper, we describe PyOMP, a system that enables OpenMP in Python. Programmers write code in Python with OpenMP, Numba generates code that compiles to LLVM, and the resulting programs run with performance that approaches that from code written with C and OpenMP. In this paper we provide an update on the PyOMP project and explain how to install it and use it to write parallel multithreaded code in Python.

97 MATHEMATICS AND COMPUTING↗

Porting Fragmentation Methods to Graphical Processing Units Using an OpenMP Application Programming Interface: Offloading the Fock Build for Low Angular Momentum Functions

Here, a framework to offload four-index two-electron repulsion integrals to graphical processing units (GPUs) using OpenMP is discussed. The method has been applied to the Fock build for low angular momentum s and p functions in both the restricted Hartree–Fock (RHF) and in the effective fragment molecular orbital (EFMO) framework. Benchmark calculations for the GPU code for the pure RHF method show an increasing speedup relative to the existing OpenMP CPU code in GAMESS from 1.04 to 52× for clusters of 70–569 water molecules. The parallel efficiency on 24 NVIDIA V100 GPU boards also increases when increasing the system size: from 75 to 94% for water clusters that contain 303–1120 molecules. In the EFMO framework, the GPU Fock build shows a high linear scalability up to 4608 V100s with a parallel efficiency of 96% for calculations on a solvated mesoporous silica nanoparticle system with ~67,000 basis functions.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Incoherent Thomson scattering system for PHAse space MApping (PHASMA) experiment

A new incoherent Thomson scattering system measures the evolution of electron velocity distribution functions perpendicular and parallel to the ambient magnetic field during kinking of a single flux rope and merging of two flux ropes through magnetic reconnection. The Thom-son scattering system provides sub-millimeter spatial resolution, sufficient to diagnose the several millimeters sized magnetic reconnection electron diffusion region in the PHAse Space MAppgin experiment. Due to the relatively modest plasma density ~ 10 19 m -3 and electron tem-perature ~1 eV, stray light suppression is critical for these measurements. Two volume Bragg gratings are used in series as a notch filter with a spectral bandwidth <0.1 nm in the collection branch. A CCD with a Gen III intensifier with peak quantum efficiency >47% is used as the detector in a 1.3 m spectrometer. Preliminary results of gun plasma electron temperature will be reported and compared with measurements obtained from a triple Langmuir probe.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Novel Solver Algorithms for Nearly Singular Linear Systems Arising in Combustion Modelling

Direct Numerical Simulations of realistic combustion devices are extremely challenging due to the wide separation of scales in the simulation, for example an internal combustion (IC) engine chamber, and the flame thickness of a high-pressure flame. The PeleLMeX solver uses adaptive mesh refinement (AMR) to evolve multi-species reacting flows in the low Mach number limit at the Exascale and relies on an embedded boundary (EB) approach to represent complex geometries. In that framework, the EB geometries often give rise to very small cut-cells along the boundary, which translate into extreme ill-conditioning of the pressure-projection, with eigenvalues that span 15-16 orders of magnitude. In this talk, we focus on the case of a typical IC piston bowl geometry for which we present on a novel approach towards solving these nearly singular linear systems with ILU-based, C-AMG smoothers on massively parallel architectures. In particular, we use scaling and equilibration algorithms to handle the non-normality of the upper triangular factors. This enables us to approximate the highly sequential triangular solve algorithm, embedded in the AMG smoothing-solve phase, with Jacobi iterations. This approximation can be written as a convergent Neumann series whose terms are composed of highly parallel sparse matrix vector multiplications. The result is an algorithm that substantially decreases setup and solve time, compared to state-of-the-art, for these challenging linear systems.

combustion modelling↗

Pretraining Billion-Scale Geospatial Foundational Models on Frontier

As AI workloads increase in scope, generalization capability becomes challenging for small task-specific models and their demand for large amounts of labeled training samples increases. On the contrary, Foundation Models (FMs) are trained with internet-scale unlabeled data via self-supervised learning and have been shown to adapt to various tasks with minimal fine-tuning. Although large FMs have demonstrated significant impact in natural language processing and computer vision, efforts toward FMs for geospatial applications have been restricted to smaller size models, as pretraining larger models requires very large computing resources equipped with state-of-the-art hardware accelerators. Current satellite constellations collect 100+TBs of data a day, resulting in images that are billions of pixels and multimodal in nature. Such geospatial data poses unique challenges opening up new opportunities to develop FMs. We investigate billion scale FMs and HPC training profiles for geospatial applications by pretraining on publicly available data. We studied from end-to-end the performance and impact in the solution by scaling the model size. Our larger 3B parameter size model achieves up to 30% improvement in top1 scene classification accuracy when comparing a 100M parameter model. Moreover, we detail performance experiments on the Frontier supercomputer, America's first exascale system, where we study different model and data parallel approaches using PyTorch's Fully Sharded Data Parallel library. Specifically, we study variants of the Vision Transformer architecture (ViT), conducting performance analysis for ViT models with size up to 15B parameters. By discussing throughput and performance bottlenecks under different parallelism configurations, we offer insights on how to leverage such leadership-class HPC resources when developing large models for geospatial imagery applications.

Tsaris, Aristeidis (aris)↗

Really Embedding Domain-Specific Languages into C++

The following topics are dealt with: program compilers; optimising compilers; parallel processing; software engineering; learning (artificial intelligence); multiprocessing systems; shared memory systems; optimisation; computational complexity; specification languages.

Finkel, Hal J.↗

An adaptive discontinuous Petrov-Galerkin method for the Grad-Shafranov equation

In this work, we propose and develop an arbitrary-order adaptive discontinuous Petrov--Galerkin (DPG) method for the nonlinear Grad--Shafranov equation. An ultraweak formulation of the DPG scheme for the equation is given based on a minimal residual method. The DPG scheme has the advantage of providing more accurate gradients compared to conventional finite element methods, which is desired for numerical solutions to the Grad--Shafranov equation. The numerical scheme is augmented with an adaptive mesh refinement approach, and a criterion based on the residual norm in the minimal residual method is developed to achieve dynamic refinement. Nonlinear solvers for the resulting system are explored and a Picard iteration with Anderson acceleration is found to be efficient to solve the system. Finally, the proposed algorithm is implemented in parallel on MFEM using a domain-decomposition approach, and our implementation is general, supporting arbitrary order of accuracy and general meshes. Furthermore, numerical results are presented to demonstrate the efficiency and accuracy of the proposed algorithm.

97 MATHEMATICS AND COMPUTING↗

Operational experience and R&D results using the Google Cloud for High-Energy Physics in the ATLAS experiment

The ATLAS experiment at CERN relies on a Worldwide Distributed Computing Grid infrastructure to support its physics program at the Large Hadron Collider. ATLAS has integrated cloud computing resources to complement its Grid infrastructure and conducted an R&D program on Google Cloud Platform. These initiatives leverage key features of commercial cloud providers: lightweight configuration and operation, elasticity and availability of diverse infrastructures. Here this paper examines the seamless integration of cloud computing services as a conventional Grid site within the ATLAS workflow management and data management systems, while also offering new setups for interactive, parallel analysis. It underscores pivotal results that enhance the on-site computing model and outlines several R&D projects that have benefited from large-scale, elastic resource provisioning models. Furthermore, this study discusses the impact of cloud-enabled R&D projects in three domains: accelerators and AI/ML, ARM CPUs and columnar data analysis techniques.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

General-Simulator-Intermediary

This application allows parallel development of simulator screens for the Human System Simulation Laboratory and connection of the backend simulators for various power plants

Lehmer, JacobP↗

GridOPTICS/GridPACK

GridPACK is a software framework consisting of a set of modules designed to simplify the development of programs that model the power grid and run on parallel, high performance computing platforms. It also contains several fully developed applications, including powerflow, dynamic simulation, state estimation, Kalman filter analysis (dynamic state estimation), contingency analysis and real time path rating. These applications can be used either standalone or as components in more complicated workflows that combine several different types of application together. The framework modules are available as a combination of libraries and software templates and consist of components for setting up and distributing power grid networks, support for modeling the behavior of individual buses and branches in the network, converting the network models to the corresponding algebraic equations, and parallel routines for manipulating and solving large algebraic systems. The framework also contains a module for distributing tasks evenly amongst computing resources, even if individual tasks vary widely in their execution times. Additional modules support input and output, basic statistical analysis of contingency based calculations, distributed data structures, as well as basic profiling and error management.

Palmer, Bruce↗

A fast particle-based approach for calibrating a 3-D model of the Antarctic ice sheet

We consider the scientifically challenging and policy-relevant task of understanding the past and projecting the future dynamics of the Antarctic ice sheet. The Antarctic ice sheet has shown a highly nonlinear threshold response to past climate forcings. Triggering such a threshold response through anthropogenic greenhouse gas emissions would drive drastic and potentially fast sea level rise with important implications for coastal flood risks. Previous studies have combined information from ice sheet models and observations to calibrate model parameters. These studies have broken important new ground but have either adopted simple ice sheet models or have limited the number of parameters to allow for the use of more complex models. These limitations are largely due to the computational challenges posed by calibration as models become more computationally intensive or when the number of parameters increases. Here, we propose a method to alleviate this problem: a fast sequential Monte Carlo method that takes advantage of the massive parallelization afforded by modern high-performance computing systems. We use simulated examples to demonstrate how our sample-based approach provides accurate approximations to the posterior distributions of the calibrated parameters. The drastic reduction in computational times enables us to provide new insights into important scientific questions, for example, the impact of Pliocene era data and prior parameter information on sea level projections. These studies would be computationally prohibitive with other computational approaches for calibration such as Markov chain Monte Carlo or emulation-based methods. We also find considerable differences in the distributions of sea level projections when we account for a larger number of uncertain parameters. For example, based on the same ice sheet model and data set, the 99th percentile of the Antarctic ice sheet contribution to sea level rise in 2300 increases from 6.5 m to 13.1 m when we increase the number of calibrated parameters from three to 11. With previous calibration methods, it would be challenging to go beyond five parameters. Here, this work provides an important next step toward improving the uncertainty quantification of complex, computationally intensive and decision-relevant models.

54 ENVIRONMENTAL SCIENCES↗