Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “parallel systems”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 415 records · Page 23

Subspace recursive Fermi-operator expansion strategies for large-scale DFT eigenvalue problems on HPC architectures

Quantum mechanical calculations for material modeling using Kohn–Sham density functional theory (DFT) involve the solution of a nonlinear eigenvalue problem for N smallest eigenvector-eigenvalue pairs, with N proportional to the number of electrons in the material system. Here, these calculations are computationally demanding and have asymptotic cubic scaling complexity with the number of electrons. Large-scale matrix eigenvalue problems arising from the discretization of the Kohn–Sham DFT equations employing a systematically convergent basis traditionally rely on iterative orthogonal projection methods, which are shown to be computationally efficient and scalable on massively parallel computing architectures. However, as the size of the material system increases, these methods are known to incur dominant computational costs through the Rayleigh–Ritz projection step of the discretized Kohn–Sham Hamiltonian matrix and the subsequent subspace diagonalization of the projected matrix. This work explores the potential of polynomial expansion approaches based on recursive Fermi-operator expansion as an alternative to the subspace diagonalization of the projected Hamiltonian matrix to reduce the computational cost. Subsequently, we perform a detailed comparison of various recursive polynomial expansion approaches to the traditional approach of explicit diagonalization on both multi-node central processing unit and graphics processing unit architectures and assess their relative performance in terms of accuracy, computational efficiency, scaling behavior, and energy efficiency.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Intelligently Partitioned Phasor-EMT Hybrid Simulations of Large-Scale, High-IBR Power Systems

As the penetration level of power electronics-interfaced renewables such as photovoltaics (PV) and wind has surged in modern electric grids, new operational risks caused by the dynamics of those inverter-based resources (IBRs) are emerging in parallel. Lessons learned from various grid events include that the impact of IBRs on system-level grid stability will become prominent along with the increase of renewables and that the short-timescale dynamic impacts of IBRs on grid stability are not fully captured by current commercial dynamic simulation tools [1] [2]. For example, IBRs can be controlled to mitigate those destabilizing interactions, but conventional phasor-domain tools (e.g. PSS/E, PSLF) often cannot capture that; likewise, the existing electromagnetic transient (EMT) simulation tools (e.g. PSCAD, EMTP) can simulate detailed IBR controls, but for large power systems with many IBRs, slow simulation speeds severely impede the ability to study dynamic events [3] [4]. Massively paralleling simulations using high-performance computing (HPC) can help address this, especially now that cloud-based HPC capability is widely available, but today s EMT tools are not HPC-compatible, and parallelization of dynamic simulation solvers is not trivial because each region can dynamically affect the others. Thus, dynamic simulation of grids with very large numbers of IBRs potentially poses a barrier to the ongoing energy transition.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Distributed Data-Driven Optimization for Voltage Regulation in Distribution Systems

Here, this paper proposes a distributed data-driven optimization framework for voltage regulation in distribution systems. The recursive kernel regression and alternating direction method of multipliers (ADMM) are selected to cover the system learning and distributed optimization tasks. The proposed distributed data-driven framework is capable of having a rapid response to system or load changes while considering the operation optimality. Besides, the distributed algorithm parallels the computation tasks and reduces the computational expense of a single agent. To validate the performance of the proposed method, a hypothetical 7-Bus system and the IEEE 123-Bus system are selected to show the effectiveness of the proposed data-driven framework. According to the numerical study results, the proposed method offers great flexibility for selecting customized kernel models for different regions and can effectively improve the system voltage profile in a distributed manner.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Additively manufactured multiplexed inertial coalescence filters

Multiphase flows often pose a significant challenge to the efficient and reliable design of thermofluidic systems. This paper describes multiplexed inertial coalescence filters composed of parallel helical pathways, designed to capture fine droplets (<40 µm) through inertial separation while maintaining a low pressure drop (<400 Pa). Filtration efficiencies for 7 µm and 30 µm droplets were characterized for varying flow conditions, with complete capture observed above a threshold flow rate. Models for filtration efficiency and pressure drop were developed and validated against experimental results to allow system design and optimization, enabled by the tunable additive manufacturing approach used to fabricate the filters. Filter quality factor was computed for varying droplet sizes, showcasing exceptional quality factor when compared to state-of-the-art filters documented in the literature. In conclusion, this multiplexed inertial coalescence filtration approach could find use in dehumidification systems, fog harvesting, chemical reactors, and microgravity droplet capture.

3D Printing↗

Smart Droplets Stabilized by Designer Surfactants: From Biomimicry to Active Motion to Materials Healing

The science and technologies of emulsion droplets have been a long‐term focus of extensive research endeavors for their practical utility across a breadth of industries, including pharmaceutical products, oil recovery processes, and the food sciences. However, with advances in materials chemistry and characterization tools, new emerging areas are arising with a focus on “smart droplets”. The versatility of emulsion droplets across is based on their ability to partition and create isolated systems with properties defined by the liquid–liquid interface, while preparative routes allow manipulation of droplet size, stability, and encapsulated contents. As described in this article, significant efforts are being devoted to creating new types of droplets by “activating” this interface through the incorporation of reactive structures that trigger droplet response to applied or environmental stimuli (e.g., pH, temperature, salt, or external fields). Moreover, parallels between droplets and live cells inspire efforts to conceive systems that resemble biological motifs or that can produce cellular behaviors that imitate biology (e.g., swarming, communication, or motion). Here, the authors highlight recent advances in smart droplets, with emphasis on organic, polymer, and/or particle surfactants that give rise to inter‐droplet communication (via aggregation, fusion, division, or mass transfer), droplet vehicles for controlled delivery, autonomous droplet motion, and tunable emulsion inversion. Especially emphasized is the macromolecular design to produce reactive and functional surfactants, which are crucial to responsive droplet behavior and their underlying mechanisms. More generally, the exquisite interplay between materials science and biology inspires the review of this research area that provides unique opportunities for insight and inspiration into the capabilities of new droplet designs.

36 MATERIALS SCIENCE↗

Ensuring statistical reproducibility of ocean model simulations in the age of hybrid computing

Novel high performance computing systems that feature hybrid architectures require large scale code refactoring to unravel underlying exploitable parallelism. Such redesign can often be accompanied with machine-precision changes as the order of computation cannot always be maintained. For chaotic systems like climate models, these round-off level differences can grow rapidly. Systematic errors may also manifest initially as machine-precision differences. Isolating genuine round off level differences from such errors remains a challenge. Here, we apply two-sample equality of distribution tests to evaluate statistical reproducibility of the ocean model component of US Department of Energy's Energy Exascale Earth System Model (E3SM). A 2-year control simulation ensemble is compared to a modified ensemble as a test case - after a known non-bit-for-bit change in a model component is introduced - to evaluate the null hypothesis that the two ensembles are statistically indistinguishable. To quantify the false negative rates of these tests, we conduct a formal power analysis using a targeted suite of short simulation ensembles. The ensemble suite contains several perturbed ensembles, each with a progressively different climate than the baseline ensemble - obtained by perturbing the magnitude of a single model tuning parameter, the Gent and McWilliams κ, in a controlled manner. The null hypothesis is evaluated for each of perturbed ensembles using these tests. The power analysis informs on the detection limits of the tests for given ensemble size allowing model developers to evaluate the impact of an introduced non-bit-for-bit change to the model.

Mahajan, Salil↗

Versatile soil gas concentration and isotope monitoring: optimization and integration of novel soil gas probes with online trace gas detection

Abstract. Gas concentrations and isotopic signatures can unveil microbial metabolisms and their responses to environmental changes in soil. Currently, few methods measure in situ soil trace gases such as the products of nitrogen and carbon cycling or volatile organic compounds (VOCs) that constrain microbial biochemical processes like nitrification, methanogenesis, respiration, and microbial communication. Versatile trace gas sampling systems that integrate soil probes with sensitive trace gas analyzers could fill this gap with in situ soil gas measurements that resolve spatial (centimeters) and temporal (minutes) patterns. We developed a system that integrates new porous and hydrophobic sintered polytetrafluoroethylene (sPTFE) diffusive soil gas probes that non-disruptively collect soil gas samples with a transfer system to direct gas from multiple probes to one or more central gas analyzer(s) such as laser and mass spectrometers. Here, we demonstrate the feasibility and versatility of this automated multiprobe system for soil gas measurements of isotopic ratios of nitrous oxide (δ18O, δ15N, and the 15N site preference of N2O), methane, carbon dioxide (δ13C), and VOCs. First, we used an inert silica matrix to challenge probe measurements under controlled gas conditions. By changing and controlling system flow parameters, including the probe flow rate, we optimized recovery of representative soil gas samples while reducing sampling artifacts on subsurface concentrations. Second, we used this system to provide a real-time window into the impact of environmental manipulation of irrigation and soil redox conditions on in situ N2O and VOC concentrations. Moreover, to reveal the dynamics in the stable isotope ratios of N2O (i.e., 14N14N16O, 14N15N16O, 15N14N16O, and 14N14N18O), we developed a new high-precision laser spectrometer with a reduced sample volume demand. Our integrated system – a tunable infrared laser direct absorption spectrometry (TILDAS) in parallel with Vocus proton transfer reaction mass spectrometry (PTR-MS), in line with sPTFE soil gas probes – successfully quantified isotopic signatures for N2O, CO2, and VOCs in real time as responses to changes in the dry–wetting cycle and redox conditions. Broadening the collection of trace gases that can be monitored in the subsurface is critical for monitoring biogeochemical cycles, ecosystem health, and management practices at scales relevant to the soil system.

54 ENVIRONMENTAL SCIENCES↗

Online data analysis and reduction: An important co-design motif for extreme-scale computers

A growing disparity between supercomputer computation speeds and I/O rates means that it is rapidly becoming infeasible to analyze supercomputer application output only after that output has been written to a file system. Instead, data-generating applications must run concurrently with data reduction and/or analysis operations, with which they exchange information via high-speed methods such as interprocess communications. The resulting parallel computing motif, online data analysis and reduction (ODAR), has important implications for both application and HPC systems design. Here we introduce the ODAR motif and its co-design concerns, describe a co-design process for identifying and addressing those concerns, present tools that assist in the co-design process, and present case studies to illustrate the use of the process and tools in practical settings.

Data Analysis↗

Parallel-in-Time Integration for Nonlinear Hyperbolic Problems (Final Report)

The work for the subcontract was situated in the area of parallel-in-time integration for hyperbolic partial differential equations (PDEs). Parallel-in-time integration is an active area of research due to its ability to enable faster numerical simulations for applications throughout many areas of science. Over the past two decades, much progress has been made in this area; however, this progress has largely been limited to diffusion-dominated PDEs, with some recent success in scalar linear hyperbolic PDEs. Given the ubiquity of numerical simulations of hyperbolic PDEs throughout the sciences, in particular, nonlinear hyperbolic systems, there is a strong need to develop efficient parallel-in-time techniques for hyperbolic problems beyond simple scalar and linear cases, which is the main topic of this subcontract. The main focus of the work was to further develop and perfect coarse-grid operators for the Multigrid Reduction-inTime (MGRIT) method applied to hyperbolic PDEs that were recently proposed in PhD thesis, based on a modified semi-Lagrangian approach.

97 MATHEMATICS AND COMPUTING↗

Online Monitoring of Medium Voltage Cable Systems with Spread Spectrum Time Domain and Frequency Domain Reflectometry

In-service failures of wave energy convertor (WEC) cable systems can have a significant cost and power availability impact. Close parallel research 2019 data showed > 1B£ and 9 Terra-Watt-Hours associated with global off-shore wind (OSW) cable failures (Strang-Moran 2020). OSW is a closely related technology but currently is significantly cheaper than WEC technology. For wave energy to compete, the problem of reliable cable transmission must be mitigated. This project develops isolation technology to allow online high frequency reflectometry testing of medium voltage cables (1 to 10 kV and higher) without arcing or damage to the test instrument. Online spread spectrum time domain reflectometry (SSTDR) testing has been established for low voltage cable systems in the aircraft and rail industry and the ability to detect and locate cable flaws of interest is well understood. Extending reflectometry testing to medium voltage systems could enable detection of cable damage before failures occur thereby allowing repair and replacement of damaged cable segments to be scheduled and managed. The seedling project succeeded to pass and receive high frequency SSTDR signals onto a cable up to 1 kV using a parallel trace isolation circuit board that can be connected onto the test cable. The approach used a novel circuit design for which an invention disclosure has been filed. A proposed sapling project would extend the technology toward the higher operating voltages used by WEC systems, thereby enabling online SSTDR cable monitoring. The goal of the seedling project was to extend the capability of the ARENA cable/motor test bed to address medium voltages and to develop a high pass filter isolation architecture to protect the reflectometry instrument from the low frequency (DC – 60 Hz) line voltage while allowing the high frequency diagnostic signal to pass to and from the test instrument to the live line. Initial efforts focused on passive LCR filter circuits to reduce 60 Hz levels below 10 volts from a 10 kV line while allowing the MHz high frequency chirps to pass onto the cables and for mV signals to be detected. We discovered that the parasitic loss behavior of real high voltage components precluded this approach from working. An alternate approach was adapted for the electric field to couple between two parallel traces on a printed circuit board much like a radio-frequency coupler. The challenge here was and is to have the parallel traces close enough to each other to effectively pass the high frequency chirp onto the live line and receive any reflected signal from any encountered impedance change along the cable. This reflected signal will be in the mV range. The traces however must be far enough apart to not allow arcing on the board. A design with 3 mm spacing was determined to allow the high frequency signal to pass onto the live line and receive the mV signal back into the instrument while reducing the 60 Hz voltage amplitude by >80 dB (more than a factor of 10,000) without allowing arcing from across the parallel traces. This was confirmed by simulation and test.

16 TIDAL AND WAVE POWER↗

Spectral quadrature for the first principles study of crystal defects: Application to magnesium

In this work, we present an accurate and efficient finite-difference formulation and parallel implementation of Kohn-Sham Density (Operator) Functional Theory (DFT) for non periodic systems embedded in a bulk environment. Specifically, employing non-local pseudopotentials, local reformulation of electrostatics, and truncation of the spatial Kohn-Sham Hamiltonian, and the Linear Scaling Spectral Quadrature method to solve for the pointwise electronic fields in real-space and the non-local component of the atomic force, we develop a parallel finite difference framework suitable for distributed memory computing architectures to simulate non-periodic systems embedded in a bulk environment. Choosing examples from magnesium-aluminum alloys, we first demonstrate the convergence of energies and forces with respect to spectral quadrature polynomial order, and the width of the spatially truncated Hamiltonian. Next, we demonstrate the parallel scaling of our framework, and show that the computation time and memory scale linearly with respect to the number of atoms. Next, we use the developed framework to simulate isolated point defects and their interactions in magnesium-aluminum alloys. Our findings conclude that the binding energies of divacancies, Al solute-vacancy and two Al solute atoms are anisotropic and are dependent on cell size. Furthermore, the binding is favorable in all three cases.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Quantum-parallel vectorized data encodings and computations on trapped-ion and transmon QPUs

Compact data representations in quantum systems are crucial for the development of quantum algorithms for data analysis. In this study, we present two innovative data encoding techniques, known as QCrank and QBArt, which exhibit significant quantum parallelism via uniformly controlled rotation gates. The QCrank method encodes a series of real-valued data as rotations on data qubits, resulting in increased storage capacity. On the other hand, QBArt directly incorporates a binary representation of the data within the computational basis, requiring fewer quantum measurements and enabling well-established arithmetic operations on binary data. We showcase various applications of the proposed encoding methods for various data types. Notably, we demonstrate quantum algorithms for tasks such as DNA pattern matching, Hamming weight computation, complex value conjugation, and the retrieval of a binary image with 384 pixels, all executed on the Quantinuum trapped-ion QPU. Furthermore, we employ several cloud-accessible QPUs, including those from IBMQ and IonQ, to conduct supplementary benchmarking experiments.

97 MATHEMATICS AND COMPUTING↗

Parallel Runtime Interface for Fortran (PRIF): A Multi-Image Solution for LLVM Flang

Fortran compilers that provide support for Fortran’s native parallel features often do so with a runtime library that depends on details of both the compiler implementation and the communication library, while others provide limited or no support at all. This paper introduces a new generalized interface that is both compiler- and runtime-library-agnostic, providing flexibility while fully supporting all of Fortran’s parallel features. The Parallel Runtime Interface for Fortran (PRIF) was developed to be portable across shared- and distributed-memory systems, with varying operating systems, toolchains and architectures. It achieves this by defining a set of Fortran procedures corresponding to each of the parallel features defined in the Fortran standard that may be invoked by a Fortran compiler and implemented by a runtime library. PRIF aims to be used as the solution for LLVM Flang to provide parallel Fortran support. This paper also briefly describes our PRIF prototype implementation: Caffeine.

Bonachea, Dan↗

Parallel three-dimensional simulations of quasi-static elastoplastic solids

Hypo-elastoplasticity is a flexible framework for modeling the mechanics of many hard materials under small elastic deformation and large plastic deformation. Under typical loading rates, most laboratory tests of these materials happen in the quasi-static limit, but there are few existing numerical methods tailor-made for this physical regime. Here, we extend to three dimensions a recent projection method for simulating quasi-static hypo-elastoplastic materials. The method is based on a mathematical correspondence to the incompressible Navier–Stokes equations, where the projection method of Chorin (1968) is an established numerical technique. We develop and utilize a three-dimensional parallel geometric multigrid solver employed to solve a linear system for the quasi-static projection. Our method is tested through simulation of three-dimensional shear band nucleation and growth, a precursor to failure in many materials. As an example system, we employ a physical model of a bulk metallic glass based on the shear transformation zone theory, but the method can be applied to any elastoplasticity model. We consider several examples of three-dimensional shear banding, and examine shear band formation in physically realistic materials with heterogeneous initial conditions under both simple shear deformation and boundary conditions inspired by friction welding.

97 MATHEMATICS AND COMPUTING↗

$\mathrm{Perturbo}$: A software package for ab initio electron–phonon interactions, charge transport and ultrafast dynamics

We report Perturbo is a software package for first-principles calculations of charge transport and ultrafast carrier dynamics in materials. The current version focuses on electron–phonon interactions and can compute phonon-limited transport properties such as the conductivity, carrier mobility and Seebeck coefficient. It can also simulate the ultrafast nonequilibrium electron dynamics in the presence of electron–phonon scattering. Perturbo uses results from density functional theory and density functional perturbation theory calculations as input, and employs Wannier interpolation to reduce the computational cost. It supports norm-conserving and ultrasoft pseudopotentials, spin–orbit coupling, and polar electron–phonon interactions for bulk and 2D materials. Hybrid MPI plus OpenMP parallelization is implemented to enable efficient calculations on large systems (up to at least 50 atoms) using high-performance computing. Taken together, Perturbo provides efficient and broadly applicable ab initio tools to investigate electron–phonon interactions and carrier dynamics quantitatively in metals, semiconductors, insulators, and 2D materials.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Energy-efficient Mott activation neuron for full-hardware implementation of neural networks

To circumvent the von Neumann bottleneck, substantial progress has been made towards in-memory computing with synaptic devices. However, compact nanodevices implementing non-linear activation functions are required for efficient full-hardware implementation of deep neural networks. Here, in this work, we present an energy-efficient and compact Mott activation neuron based on vanadium dioxide and its successful integration with a conductive bridge random access memory (CBRAM) crossbar array in hardware. The Mott activation neuron implements the rectified linear unit function in the analogue domain. The neuron devices consume substantially less energy and occupy two orders of magnitude smaller area than those of analogue complementary metal–oxide semiconductor implementations. The LeNet-5 network with Mott activation neurons achieves 98.38% accuracy on the MNIST dataset, close to the ideal software accuracy. We perform large-scale image edge detection using the Mott activation neurons integrated with a CBRAM crossbar array. Our findings provide a solution towards large-scale, highly parallel and energy-efficient in-memory computing systems for neural networks.

electrical and electronic engineering↗

Stochastic Vector Techniques in Ground-State Electronic Structure

Herein we review a suite of stochastic vector computational approaches for studying the electronic structure of extended condensed matter systems. These techniques help reduce algorithmic complexity, facilitate efficient parallelization, simplify computational tasks, accelerate calculations, and diminish memory requirements. While their scope is vast, we limit our study to ground-state and finite temperature density functional theory (DFT) and second-order many-body perturbation theory. More advanced topics, such as quasiparticle (charge) and optical (neutral) excitations and higher-order processes, are covered elsewhere. We start by explaining how to use stochastic vectors in computations, characterizing the associated statistical errors. Next, we show how to estimate the electron density in DFT and discuss effective techniques to reduce statistical errors. Finally, we review the use of stochastic vectors for calculating correlation energies within the second-order Møller-Plesset perturbation theory and its finite temperature variational form. Example calculation results are presented and used to demonstrate the efficacy of the methods.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

ddcMD.os

ddcMD is a general purpose molecular dynamics (MD) code that supports MPI parallelism. MD codes are used for simulation of particles systems and capture all the many-body effects of the underlying particle potential that defines the physical system. Though MD can be used to model systems from the subatomic to astrological length scales ddcMD is mainly focused on the atomic scale length scale, In this release of ddcMD the support will be mainly for systems using the coarse-grain Martini potential, a particle potential for biological systems.

Glosli, JamesN↗