Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “distributed parallelization”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 487 records · Page 27

Distribution function of continuously created newborn and pickup ions in outer cometary exospheres

The time evolution of the distribution function of newborn ions in the solar wind is investigated using a quasi-linear-type diffusion equation. The initial distribution is taken to be a ring beam, which is approximated by delta function in pitch angle and velocity, and it is assumed that the ions are created at a constant rate with a similar distribution. A long-time asymptotic form of ion distribution is obtained, which is a mixture of newborn ions and ions generated throughout the entire process. It is shown that the time asymptotic distribution function exists even in the presence of a continuous ionization process. The stability of the long-time asymptotic distribution was examined for the case of parallel propagation, and the results show that the distribution function can be unstable to low-frequency hydromagnetic waves. The results of the analysis were found to agree with recent satellite observations.

Gaffey, J. D., Jr.↗

Parallel Monte Carlo Simulation for control system design

The research during the 1993/94 academic year addressed the design of parallel algorithms for stochastic robustness synthesis (SRS). SRS uses Monte Carlo simulation to compute probabilities of system instability and other design-metric violations. The probabilities form a cost function which is used by a genetic algorithm (GA). The GA searches for the stochastic optimal controller. The existing sequential algorithm was analyzed and modified to execute in a distributed environment. For this, parallel approaches to Monte Carlo simulation and genetic algorithms were investigated. Initial empirical results are available for the KSR1.

Schubert, Wolfgang M.↗

Integrated Diagnostics and Prognostics for the Electrical Power System of a Planetary Rover

For electric vehicles, technology for monitoring, diagnosis, and prognosis of the electrical power system (EPS) becomes essential for safe and efficient operation. To this end, we develop a general system-level integrated diagnosis and prognosis framework, which detects, isolates, and identifies EPS faults, and predicts when the EPS will fail to deliver sufficient power. The approach takes advantage of recent work in structural model decomposition in order to distribute the global diagnosis and prognosis problems into local subproblems that can be solved in parallel, thus enabling implementation on distributed computational platforms. The framework is applied to the EPS of a planetary rover testbed, and is demonstrated using data from field experiments.

Daigle, Matthew↗

Ion Beam Heating in the Auroral Zone

Recent satellite observations at high altitudes (greater than 5000 km) in the auroral zone have shown the existence of hybrid or bimodal ion beam distributions that are evidence of both parallel and perpendicular ion acceleration. Acceleration parallel to the magnetic field is most likely due to quasi-static electric fields (double layers) which can create outflowing ion beams; since ions of different mass will have different drift speeds due to this acceleration, a plasma configuration unstable to the ion-ion two-stream acoustic mode develops. When the net drift velocity (U) between the two ion species is greater than the sound speed (C(sub 0)), the ion-ion instability has maximum growth at oblique wave propagation. To study the nonlinear effects of the ion-ion instability in terms of plasma heating, a numerical simulation parametric study has been performed. It was found that the parallel acceleration that forms the ion beams occurs on a time scale faster than ion- ion wave growth at low drifts; thus ion-ion wave growth is expected to occur primarily for higher drift speeds (U greater than C(sub 0)) which results in strong oblique heating of the ions (both hydrogen and oxygen) forming elevated ion conics (sometimes called 'bowl' distributions). Also, strong parallel electron heating in the direction of the ion beams can occur, and electrons near the top of the acceleration region may attain a net upward drift along with the elevated ion conics. Variation of the oxygen density greatly affects the ion heating due to the ion-ion instability; as the oxygen density decreases, oxygen heating increases, in agreement with observations (Collin et al., 1987). Ion-ion electrostatic wave properties and the plasma heating that results over a wide range of auroral zone parameters are included.

Schriver, David↗

ISEE 3 observations of solar wind thermal electrons with T-perpendicular greater than T-parallel

This study presents ISEE 3 observations of anomalous electron distributions for which T-perpendicular exceeds T-parallel in the solar wind near 1 AU. Twelve anomaly events were identified, lasting from 24 min to 6 hours. These events generally share the following characteristics: (1) high plasma density, (2) low solar wind speed, (3) magnetic field which is nearly transverse to the flow, and (4) low electron and ion temperatures. The processes of solar wind adiabatic expansion and isotropization via Coulomb collisions could be expected to lead to such anomalous anisotropies under conditions similar to those observed. However, these conditions actually produce T-perpendicular greater than T-parallel for only a small fraction of the time, suggesting that other mechanisms are also important in regulating solar wind electron distributions.

Phillips, J. L.↗

Parallel computation of aerodynamic influence coefficients for aeroelastic analysis on a transputer network

Aeroelastic analysis is mult-disciplinary and computationally expensive. Hence, it can greatly benefit from parallel processing. As part of an effort to develop an aeroelastic analysis capability on a distributed-memory transputer network, a parallel algorithm for the computation of aerodynamic influence coefficients is implemented on a network of 32 transputers. The aerodynamic influence coefficients are calculated using a three-dimensional unsteady aerodynamic model and a panel discretization. Efficiencies up to 85 percent are demonstrated using 32 processors. The effects of subtask ordering, problem size and network topology are presented. A comparison to results on a shared-memory computer indicates that higher speedup is achieved on the distributed-memory system.

Janetzke, D. C.↗

Using Parameter Sweep in WaterTAP to Analyze New Water Treatment Technologies

We describe a powerful and generalized parameter sweep tool in this report that was originally developed to analyze the performance of existing and novel water treatment models being developed in WaterTAP. Since WaterTAP is built upon IDAES and Pyomo, the parameter sweep tool can be used to systematically explore and debug the behavior of most Pyomo and IDAES numerical models. In order to enable meaningful analyses, the parameter sweep tool has been designed with the following features: 1) Model flexibility: The parameter sweep tool does not enforce any restrictions on the types of models that can be used with it. As long as a Pyomo model can be solved and the parameter is active and mutable, the tool only needs functions that describe how to run the model, the sweep parameters, and the output quantities of interest. 2) Flexible sampling: The parameter sweep tool has inbuilt functions to generate samples from a random distribution or a multidimensional Euclidean space. Furthermore, the users have to ability to supply samples generated from a tool of their choice. 3) Multiple sweep types: A user can choose from one of 3 types of parameter sweeps depending on their needs. 4) Detailed outputs: Outputs generated by the parameter sweep tool can be stored in detailed H5 file or user-friendly CSV files for post processing. 5) Parallel computing: The parameter sweep supports shared and distributed memory parallel computing to enable the use of high performance computers (HPC) for large-scale analyses. 6) Modular: The parameter sweep tool is self-contained and can easily be integrated within an outer-loop analysis or as desired by the user. 7) Ease of use: The tool is well documented and a simple sweep can be easily executed by following the online documentation in a few lines of code. We demonstrate the use of the parameter sweep tool on a simple water treatment system from the WaterTAP repository and show its parallel scaling performance on an Apple laptop and NREL's Eagle HPC. The parameter sweep tool is actively being used with models currently being developed within WaterTAP and we expect its use to grow beyond it to other IDAES and Pyomo models.

97 MATHEMATICS AND COMPUTING↗

Simulation of electron distributions within auroral acceleration regions

Plasma data from the Dynamics Explorer 1 satellite are used to investigate the electron distributions found within regions of parallel potential drops. A computer simulation was developed to predict major features of the observed electron distributions found within these regions. The simulation results are discussed and compared with observations. Good agreement is observed only when the potential distribution along the field line is split into separate potential drops, one above and one below the satellite. The simulated electron distribution exhibits a trapped particle region, the formation of the electron hole, and the evacuation of the loss cone.

Gurgiolo, C.↗

The electron signature of parallel electric fields

Dynamics Explorer I High-Altitude Plasma Instrument electron data are presented. The electron distribution functions have characteristics expected of a region of parallel electric fields. The data are consistent with previous test-particle simulations for observations within parallel electric field regions which indicate that typical hole, bump, and loss-cone electron distributions, which contain evidence for parallel potential differences both above and below the point of observation, are not expected to occur in regions containing actual parallel electric fields.

Burch, J. L.↗

Aerodynamic Shape Optimization Using A Combined Distributed/Shared Memory Paradigm

Current parallel computational approaches involve distributed and shared memory paradigms. In the distributed memory paradigm, each processor has its own independent memory. Message passing typically uses a function library such as MPI or PVM. In the shared memory paradigm, such as that used on the SGI Origin 2000 machine, compiler directives are used to instruct the compiler to schedule multiple threads to perform calculations. In this paradigm, it must be assured that processors (threads) do not simultaneously access regions of memory in such away that errors would occur. This paper utilizes the latest version of the SGI MPI function library to combine the two parallelization paradigms to perform aerodynamic shape optimization of a generic wing/body.

Cheung, Samson↗

Integrating Cache Performance Modeling and Tuning Support in Parallelization Tools

With the resurgence of distributed shared memory (DSM) systems based on cache-coherent Non Uniform Memory Access (ccNUMA) architectures and increasing disparity between memory and processors speeds, data locality overheads are becoming the greatest bottlenecks in the way of realizing potential high performance of these systems. While parallelization tools and compilers facilitate the users in porting their sequential applications to a DSM system, a lot of time and effort is needed to tune the memory performance of these applications to achieve reasonable speedup. In this paper, we show that integrating cache performance modeling and tuning support within a parallelization environment can alleviate this problem. The Cache Performance Modeling and Prediction Tool (CPMP), employs trace-driven simulation techniques without the overhead of generating and managing detailed address traces. CPMP predicts the cache performance impact of source code level "what-if" modifications in a program to assist a user in the tuning process. CPMP is built on top of a customized version of the Computer Aided Parallelization Tools (CAPTools) environment. Finally, we demonstrate how CPMP can be applied to tune a real Computational Fluid Dynamics (CFD) application.

Waheed, Abdul↗

Simulating energetic ions and enhanced fusion rates from ion-cyclotron resonance heating with a full-wave/Fokker–Planck model

Reproducing fast-ion enhanced fusion rates from ion-cyclotron resonance heating (ICRH) in tokamaks requires the self-consistent coupling of a full-wave solver and a Fokker–Planck solver, which evolves multiple simultaneously resonant ion species. We introduce a new self-consistent model that iterates the TORIC full-wave solver with the CQL3D Fokker–Planck solver using the integrated plasma simulator (IPS). This model evolves the bounce-averaged ion distribution functions in both parallel and perpendicular velocity-space with a quasilinear radio frequency (RF) diffusion operator valid in the ion finite Larmor radius (FLR) limit and the RF electric fields with the resultant non-Maxwellian FLR dielectric tensor. This produces non-Maxwellian ICRH simulations that are fully self-consistent, fast, and interoperable with integrated modeling frameworks, such as TRANSP/GACODE/IPS-FASTRAN. We demonstrate our model's capabilities by validating it against experimental data in Alcator C-Mod. We then perform the first RF heating simulations of SPARC using self-consistent non-Maxwellian ion distributions to investigate the potential to enhance fusion rates using ion cyclotron resonance heating generated fast ions.

Physics↗

Two-state ion heating at quasi-parallel shocks

This paper documents the alternating occurrence of two different ion states of the ion distributions downstream from several quasi-parallel shocks: a cooler (and denser) core/shoulder type and a hotter (less dense) and more Maxwellian type. Three separate lines of evidence are presented to show that the two states are not related in an evolutionary sense, but that both are produced alternately at the shock: (1) the asymptotic downstream plasma parameters (density, ion temperautre, and flow speed) are intermediate between those characterizing the two different states closer to the shock, suggesting that the asymptotic state is produced by a mixing of the two initial states; (2) examples of apparently interpenetrating distributions are found during transitions from one state to the other; and (3) examples of both types of distributions are found at actual crossings of the shock ramp.

Thomsen, M. F.↗

Use of networked workstations for parallel nonlinear structural dynamic simulations of rotating bladed-disk assemblies

The principal objective of this research is to investigate, develop and demonstrate coarse-grained, parallel-processing strategies for nonlinear dynamic simulations for rotating bladed-disk assemblies. The parallel -processing strategies addressed include numerical algorithms for parallel nonlinear solutions and techniques to effect load balancing among processors. The parallel environment employed is a distributed-memory, coarse-grained one consisting of networked workstations. A parallel explicit time integration method has been implemented for transient nonlinear solutions of rotationg bladed-disk assemblies. Automatic domain partitioning techniques have been investigated for load balancing among processors. Advanced computing environments, data structures and interactive computer graphics all contribute to an integrated parallel finite element analysis system to facilitate more efficient and powerful dynamic simulations.

Hsieh, Shang-Hsien↗

COHORT: Coordination of Heterogeneous Thermostatically Controlled Loads for Demand Flexibility

Demand flexibility is increasingly important for power grids. Careful coordination of thermostatically controlled loads (TCLs) can modulate energy demand, decrease operating costs, and increase grid resiliency. We propose a novel distributed control framework for the Coordination Of HeterOgeneous Residential Thermostatically controlled loads (COHORT). COHORT is a practical, scalable, and versatile solution that coordinates a population of TCLs to jointly optimize a grid-level objective, while satisfying each TCL’s end-use requirements and operational constraints. To achieve that, we decompose the grid-scale problem into subproblems and coordi- nate their solutions to find the global optimum using the alternating direction method of multipliers (ADMM). The TCLs’ local problems are distributed to and computed in parallel at each TCL, making COHORT highly scalable and privacy-preserving. While each TCL poses combinatorial and non-convex constraints, we characterize these constraints as a convex set through relaxation, thereby making COHORT computationally viable over long planning horizons. After coordination, each TCL is responsible for its own control and tracks the agreed-upon power trajectory with its preferred strategy. In this work, we translate continuous power back to discrete on/off actuation, using pulse width modulation. COHORT is generalizable to a wide range of grid objectives, which we demonstrate through three distinct use cases: generation following, minimizing ramping, and peak load curtailment. In a notable experiment, we validated our approach through a hardware-in-the-loop simulation, including a real-world air conditioner (AC) controlled via a smart thermostat, and simulated instances of ACs modeled after real-world data traces. During the 15-day experimental period, COHORT reduced daily peak loads by an average of 12.5% and maintained comfortable temperatures.

demand response↗

Efficient Parallel Algorithm For Direct Numerical Simulation of Turbulent Flows

A distributed algorithm for a high-order-accurate finite-difference approach to the direct numerical simulation (DNS) of transition and turbulence in compressible flows is described. This work has two major objectives. The first objective is to demonstrate that parallel and distributed-memory machines can be successfully and efficiently used to solve computationally intensive and input/output intensive algorithms of the DNS class. The second objective is to show that the computational complexity involved in solving the tridiagonal systems inherent in the DNS algorithm can be reduced by algorithm innovations that obviate the need to use a parallelized tridiagonal solver.

Moitra, Stuti↗

On the Efficient Evaluation of the Exchange Correlation Potential on Graphics Processing Unit Clusters

The predominance of Kohn–Sham density functional theory (KS-DFT) for the theoretical treatment of large experimentally relevant systems in molecular chemistry and materials science relies primarily on the existence of efficient software implementations which are capable of leveraging the latest advances in modern high-performance computing (HPC). With recent trends in HPC leading toward increasing reliance on heterogeneous accelerator-based architectures such as graphics processing units (GPU), existing code bases must embrace these architectural advances to maintain the high levels of performance that have come to be expected for these methods. In this work, we purpose a three-level parallelism scheme for the distributed numerical integration of the exchange-correlation (XC) potential in the Gaussian basis set discretization of the Kohn–Sham equations on large computing clusters consisting of multiple GPUs per compute node. In addition, we purpose and demonstrate the efficacy of the use of batched kernels, including batched level-3 BLAS operations, in achieving high levels of performance on the GPU. We demonstrate the performance and scalability of the implementation of the purposed method in the NWChemEx software package by comparing to the existing scalable CPU XC integration in NWChem.

97 MATHEMATICS AND COMPUTING↗

FTK: A Simplicial Spacetime Meshing Framework for Robust and Scalable Feature Tracking

In this work, we present the Feature Tracking Kit (FTK), a framework that simplifies, scales, and delivers various feature-tracking algorithms for scientific data. The key of FTK is our simplicial spacetime meshing scheme that generalizes both regular and unstructured spatial meshes to spacetime while tessellating spacetime mesh elements into simplices. The benefits of using simplicial spacetime meshes include (1) reducing ambiguity cases for feature extraction and tracking, (2) simplifying the handling of degeneracies using symbolic perturbations, and (3) enabling scalable and parallel processing. The use of simplicial spacetime meshing simplifies and improves the implementation of several feature-tracking algorithms for critical points, quantum vortices, and isosurfaces. As a software framework, FTK provides end users with VTK/ParaView filters, Python bindings, a command line interface, and programming interfaces for feature-tracking applications. We demonstrate use cases as well as scalability studies through both synthetic data and scientific applications including tokamak, fluid dynamics, and superconductivity simulations. We also conduct end-to-end performance studies on the Summit supercomputer. FTK is open sourced under the MIT license: https://github.com/hguo/ftk.

97 MATHEMATICS AND COMPUTING↗