Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Parallel application”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11

Parallel software support for computational structural mechanics

The application of the parallel programming methodology known as the Force was conducted. Two application issues were addressed. The first involves the efficiency of the implementation and its completeness in terms of satisfying the needs of other researchers implementing parallel algorithms. Support for, and interaction with, other Computational Structural Mechanics (CSM) researchers using the Force was the main issue, but some independent investigation of the Barrier construct, which is extremely important to overall performance, was also undertaken. Another efficiency issue which was addressed was that of relaxing the strong synchronization condition imposed on the self-scheduled parallel DO loop. The Force was extended by the addition of logical conditions to the cases of a parallel case construct and by the inclusion of a self-scheduled version of this construct. The second issue involved applying the Force to the parallelization of finite element codes such as those found in the NICE/SPAR testbed system. One of the more difficult problems encountered is the determination of what information in COMMON blocks is actually used outside of a subroutine and when a subroutine uses a COMMON block merely as scratch storage for internal temporary results.

Jordan, Harry F.↗

A comparison using APPL and PVM for a parallel implementation of an unstructured grid generation program

Efforts to parallelize the VGRIDSG unstructured surface grid generation program are described. The inherent parallel nature of the grid generation algorithm used in VGRIDSG was exploited on a cluster of Silicon Graphics IRIS 4D workstations using the message passing libraries Application Portable Parallel Library (APPL) and Parallel Virtual Machine (PVM). Comparisons of speed up are presented for generating the surface grid of a unit cube and a Mach 3.0 High Speed Civil Transport. It was concluded that for this application, both APPL and PVM give approximately the same performance, however, APPL is easier to use.

Arthur, Trey↗

Gigaflop performance on a CRAY-2: Multitasking a computational fluid dynamics application

The methodology is described for converting a large, long-running applications code that executed on a single processor of a CRAY-2 supercomputer to a version that executed efficiently on multiple processors. Although the conversion of every application is different, a discussion of the types of modification used to achieve gigaflop performance is included to assist others in the parallelization of applications for CRAY computers, especially those that were developed for other computers. An existing application, from the discipline of computational fluid dynamics, that had utilized over 2000 hrs of CPU time on CRAY-2 during the previous year was chosen as a test case to study the effectiveness of multitasking on a CRAY-2. The nature of dominant calculations within the application indicated that a sustained computational rate of 1 billion floating-point operations per second, or 1 gigaflop, might be achieved. The code was first analyzed and modified for optimal performance on a single processor in a batch environment. After optimal performance on a single CPU was achieved, the code was modified to use multiple processors in a dedicated environment. The results of these two efforts were merged into a single code that had a sustained computational rate of over 1 gigaflop on a CRAY-2. Timings and analysis of performance are given for both single- and multiple-processor runs.

Tennille, Geoffrey M.↗

The GOES-R GeoStationary Lightning Mapper (GLM)

The Geostationary Operational Environmental Satellite (GOES-R) is the next series to follow the existing GOES system currently operating over the Western Hemisphere. Superior spacecraft and instrument technology will support expanded detection of environmental phenomena, resulting in more timely and accurate forecasts and warnings. Advancements over current GOES capabilities include a new capability for total lightning detection (cloud and cloud-to-ground flashes) from the Geostationary Lightning Mapper (GLM), and improved capability for the Advanced Baseline Imager (ABI). The Geostationary Lighting Mapper (GLM) will map total lightning activity (in-cloud and cloud-to-ground lighting flashes) continuously day and night with near-uniform spatial resolution of 8 km with a product refresh rate of less than 20 sec over the Americas and adjacent oceanic regions. This will aid in forecasting severe storms and tornado activity, and convective weather impacts on aviation safety and efficiency among a number of potential applications. In parallel with the instrument development (a prototype and 4 flight models), a GOES-R Risk Reduction Team and Algorithm Working Group Lightning Applications Team have begun to develop the Level 2 algorithms (environmental data records), cal/val performance monitoring tools, and new applications using GLM alone, in combination with the ABI, merged with ground-based sensors, and decision aids augmented by numerical weather prediction model forecasts. Proxy total lightning data from the NASA Lightning Imaging Sensor on the Tropical Rainfall Measuring Mission (TRMM) satellite and regional test beds are being used to develop the pre-launch algorithms and applications, and also improve our knowledge of thunderstorm initiation and evolution. An international field campaign planned for 2011-2012 will produce concurrent observations from a VHF lightning mapping array, Meteosat multi-band imagery, Tropical Rainfall Measuring Mission (TRMM) Lightning Imaging Sensor (LIS) overpasses, and related ground and in-situ lightning and meteorological measurements in the vicinity of Sao Paulo. These data will provide a new comprehensive proxy data set for algorithm and application development.

Goodman, Steven J.↗

A study on dc fault handling of dc-dc converter based on series and parallel connected DABs in MVdc applications

With the increase of direct current (dc) based power generating stations and loads, multi-terminal medium voltage dc (MT-MVdc) systems are gaining popularity. The dc sources and loads are interfaced with the MT-MVdc systems using series and parallel connected dual-active-bridge (SP-DAB). In this paper, the dc fault ride through procedure of SP-DAB meant for MT-MVdc application is analyzed. The need for dc breaker (DCB) under the requirements of MT-MVdc is established through EMT simulation studies. It is shown that an additional dc filter capacitor is essential at the output terminals of the SP-DAB to ride through the dc faults.

Jaldanki, Sreenivasa [ORNL] (ORCID:000000028122322↗

Dynamic Phasor Modeling of Three Phase Voltage Source Inverters

With the increase in the development and implementation of distributed energy resources, application of parallel connected inverters is increasing and as a result having accurate modeling and simulation tools that can help in the design, analysis and stability assessment of the grid is of great importance. Development of fundamental methods that can achieve accurate, reliable, and computationally efficient results can be very beneficial. Application of Dynamic phasor (DP) modeling method has been limited to study of limited harmonics and small combination of interconnected converters due to the complexity associated with developing models that describe larger systems. In this paper, application of DP modeling method is expanded to model any number of parallel connected three phase voltage source inverter (VSI) with inclusion of a wider harmonic content including fundamental, subharmonics, inter-harmonics, switching frequency and their sidebands. Results achieved from this modeling method is compared with conventional average model as well as detailed switching model and the effect of inclusion of wider harmonic content on accuracy of DP modeling method is demonstrated.

Xue, Yaosuo↗

Parallel Programming Strategies for Irregular Adaptive Applications

Achieving scalable performance for dynamic irregular applications is eminently challenging. Traditional message-passing approaches have been making steady progress towards this goal; however, they suffer from complex implementation requirements. The use of a global address space greatly simplifies the programming task, but can degrade the performance for such computations. In this work, we examine two typical irregular adaptive applications, Dynamic Remeshing and N-Body, under competing programming methodologies and across various parallel architectures. The Dynamic Remeshing application simulates flow over an airfoil, and refines localized regions of the underlying unstructured mesh. The N-Body experiment models two neighboring Plummer galaxies that are about to undergo a merger. Both problems demonstrate dramatic changes in processor workloads and interprocessor communication with time; thus, dynamic load balancing is a required component.

Biswas, Rupak↗

Parallel Computation of Sensitivity Derivatives with Application to Aerodynamic Optimization of a Wing

This paper focuses on the parallel computation of aerodynamic derivatives via automatic differentiation of the Euler/Navier-Stokes solver CFL3D. The comparison with derivatives obtained by finite differences is presented and the scaling of the time required to obtain the derivatives relative to the number of processors employed for the computation is shown. Finally, the derivative computations are coupled with an optimizer and surface/volume grid deformation tools to perform an optimization to reduce the drag of a three-dimensional wing.

Biedron, Robert T.↗

Parallel adaptive mesh refinement techniques for plasticity problems

The accurate modeling of the nonlinear properties of materials can be computationally expensive. Parallel computing offers an attractive way for solving such problems; however, the efficient use of these systems requires the vertical integration of a number of very different software components, we explore the solution of two- and three-dimensional, small-strain plasticity problems. We consider a finite-element formulation of the problem with adaptive refinement of an unstructured mesh to accurately model plastic transition zones. We present a framework for the parallel implementation of such complex algorithms. This framework, using libraries from the SUMAA3d project, allows a user to build a parallel finite-element application without writing any parallel code. To demonstrate the effectiveness of this approach on widely varying parallel architectures, we present experimental results from an IBM SP parallel computer and an ATM-connected network of Sun UltraSparc workstations. The results detail the parallel performance of the computational phases of the application during the process while the material is incrementally loaded.

Barry, W. J.↗

Logically Parallel Communication for Fast MPI+Threads Applications

Supercomputing applications are increasingly adopting the MPI+threads programming model over the traditional "MPI everywhere" approach to better handle the disproportionate increase in the number of cores compared with other on-node resources. In practice, however, most applications observe a slower performance with MPI+threads primarily because of poor communication performance. Recent research efforts on MPI libraries address this bottleneck by mapping logically parallel communication, that is, operations that are not subject to MPI's ordering constraints to the underlying network parallelism. Domain scientists, however, typically do not expose such communication independence information because the existing MPI-3.1 standard's semantics can be limiting. Researchers had initially proposed user-visible endpoints to combat this issue, but such a solution requires intrusive changes to the standard (new APIs). The upcoming MPI-4.0 standard, on the other hand, allows applications to relax unneeded semantics and provides them with many opportunities to express logical communication parallelism. Here in this article, we show how MPI+threads applications can achieve high performance with logically parallel communication. Through application case studies, we compare the capabilities of the new MPI-4.0 standard with those of the existing one and user-visible endpoints (upper bound). Logical communication parallelism can boost the overall performance of an application by over 2x.

97 MATHEMATICS AND COMPUTING↗

Parallelization of a Six Degree of Freedom Entry Vehicle Trajectory Simulation Using OpenMP and OpenACC

The art and science of writing parallelized software, using methods such as Open Multi-Processing (OpenMP) and Open Accelerators (OpenACC), is dominated by computer scientists. Engineers and non-computer scientists looking to apply these techniques to their project applications face a steep learning curve, especially when looking to adapt their original single threaded software to run multi-threaded on graphics processing units (GPUs). There are significant changes in mindset that must occur; such as how to manage memory, the organization of instructions, and the use of if statements (also known as branching). The purpose of this work is twofold: 1) to demonstrate the applicability of parallelized coding methodologies, OpenMP and OpenACC, to tasks outside of the typical large scale matrix mathematics; and 2) to discuss, from an engineer’s perspective, the lessons learned from parallelizing software using these computer science techniques. This work applies OpenMP, on both multi-core central processing units (CPUs) and Intel® Xeon Phi™ 7210, and OpenACC on GPUs. These parallelization techniques are used to tackle the simulation of thousands of entry vehicle trajectories through the integration of six degree of freedom (DoF) equations of motion (EoM). The forces and moments acting on the entry vehicle, and used by the EoM, are estimated using multiple models of varying levels of complexity. Several benchmark comparisons are made on the execution of six DoF trajectory simulation: single thread Intel® Xeon® E5-2670 CPU, multi-thread CPU using OpenMP, multi-thread Xeon Phi™ 7210 using OpenMP, and multi-thread NVIDIA® Tesla® K40 GPU using OpenACC. These benchmarks are run on the Pleiades Supercomputer Cluster at the National Aeronautics and Space Administration (NASA) Ames Research Center (ARC), and a Xeon Phi™ 7210 node at NASA Langley Research Center (LaRC).

Green, Justin S.↗

Parallel Signal Processing and System Simulation using aCe

Recently, networked and cluster computation have become very popular for both signal processing and system simulation. A new language is ideally suited for parallel signal processing applications and system simulation since it allows the programmer to explicitly express the computations that can be performed concurrently. In addition, the new C based parallel language (ace C) for architecture-adaptive programming allows programmers to implement algorithms and system simulation applications on parallel architectures by providing them with the assurance that future parallel architectures will be able to run their applications with a minimum of modification. In this paper, we will focus on some fundamental features of ace C and present a signal processing application (FFT).

Dorband, John E.↗

Lime slurry treatment of soils developing on abandoned coal mine spoil: Linking contaminant transport from the micrometer to pedon-scale

Historical and abandoned coal mine spoil continues to generate acid- and metal(loid)-rich porewaters and represents a geographically large and diffuse non-point source of contamination to local watersheds. A potentially inexpensive approach to treat these materials and soils developing on them is through the application of lime slurries, to neutralize acidity and encourage the (co)precipitation of metal(loids) with Fe(III)-(oxy)hydroxides, and potentially other metal-oxides and/or Ca-bearing phases. Here, the efficacy of this approach was evaluated through parallel field application and laboratory-based flow-through column experiments. The field site is Huff Run sub-watershed 25 located in Tuscarawas County, Ohio and was chosen in part because it was previously classified as one of the most highly AMD-impacted sub-watersheds in the region. Two locations with historical spoil were chosen and suction lysimeters were installed at 25 and 75 cm depth to monitor porewater composition on two sides at the base of each pile. Half of each slope received seven lime slurry treatments from June through October of 2017. A suite of aqueous (ICP-OES, IC, and TOC-L) and solid phase geochemical and mineralogical approaches (quantitative SEM-EDS and synchrotron μ-XRF) were used to determine how composition, texture, morphology, and spatial distribution of mineral coatings differ in pre- and post-lime treated soils, and how that impacts the distribution and transport of trace metal(loid)s. Mine spoil porewater at site 1 was slightly less alkaline (pH ranging from 7.04 to 7.37) than at site 2 (ranging from 7.55 to 7.71), and average electrical conductivity values at site 1 (316–405 μS cm —1 ) were slightly lower than at site 2 (358–464 μS cm —1 ), although differences between the sites were not significant. Porewater pH and electrical conductivity in all lysimeters decreased over the course of the field season but there was no obvious response to lime treatment at either site or any depth. At site 1, both treatment and depth were significant factors affecting Ca, K, Ni, SO 4 2— , and DOC concentrations while only treatment effects were significant for dissolved Al and Cu (p < 0.05). For all soils, there were no trends in metal concentration observed over time although DOC and SO 4 2— decreased over the field season. Pedon-scale changes in metal porewater concentrations in response to treatment were linked to micrometer-scale changes in mineral surface coatings; specifically, higher concentrations of Ca, Fe, Mn, and Zn were observed in the coatings and no changes were observed in Fe redox speciation, whereas total S decreased likely due to oxidation of S in coal fragments. In contrast to the field experiment, the column experiments exhibited a much greater response in effluent composition with respect to lime treatment. The untreated columns had approximately an order of magnitude more H + leached over the course of the experiment (p < 0.001) and resulted in greater Ca, Al, Cu, and DIC leached and less Mn, Zn, and SO 4 2— . Soils treated with the lime slurry in the column experiments exhibited larger and thicker secondary Fe-coatings, including the addition of Fe-sulfates. Despite clear trends in the laboratory-based column experiment where the lime-to-soil ratio was higher, the effects were either muted or undetected in the field pilot project, suggesting that a higher application rate of lime in the field is needed to achieve a similar effect. This work provides evidence that a less alkaline lime slurry could be a practical and inexpensive method of treating coal mine spoil-impacted soils and represents an important step in linking laboratory-based remediation studies to implemented field-based studies.

58 GEOSCIENCES↗

Developing Multiphysics, Integrated, High-Fidelity, Massively Parallel Computational Capabilities for Fusion Applications Using MOOSE

As the need for fusion as a clean, sustainable, and abundant energy source grows internationally, so does the need for multiphysics, computational tools to model, study, and predict the complex interactions between plasma, materials, and engineering processes. These tools have a crucial role to play in solving scientific and engineering challenges and accelerating fusion energy deployment. To address these needs, modeling capabilities should enable massively parallel, multiphysics, fully integrated high-fidelity simulations of fusion systems. Additional attributes, such as being open source and modular while maintaining high software quality assurance standards will maximize impact by ensuring accessibility for all and wide acceptance, rapid expansion and development, as well as reliability, efficiency, and robustness. In this paper, we describe how the Multiphysics Object-Oriented Simulation Environment (MOOSE) framework, which has a track record of success in the fission space thanks to the attributes listed above, can be leveraged in the fusion energy field. We highlight key successes of the MOOSE application in the fission space and describe how MOOSE has been and is being applied to fusion applications in the United States---e.g., Tritium Migration Analysis Program, version 8 (TMAP8), MOOSE Fusion Module, Fusion ENergy Integrated multiphys-X (FENIX)---and the United Kingdom---e.g., AURORA, Achlys, Apollo. These efforts aim to establish a suite of tools that can be further extended to accelerate fusion energy deployment.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

LSPRAY-II: A Lagrangian Spray Module

LSPRAY-II is a Lagrangian spray solver developed for application with parallel computing and unstructured grids. It is designed to be massively parallel and could easily be coupled with any existing gas-phase flow and/or Monte Carlo Probability Density Function (PDF) solvers. The solver accommodates the use of an unstructured mesh with mixed elements of either triangular, quadrilateral, and/or tetrahedral type for the gas flow grid representation. It is mainly designed to predict the flow, thermal and transport properties of a rapidly vaporizing spray because of its importance in aerospace application. The manual provides the user with an understanding of various models involved in the spray formulation, its code structure and solution algorithm, and various other issues related to parallelization and its coupling with other solvers. With the development of LSPRAY-II, we have advanced the state-of-the-art in spray computations in several important ways.

Raju, M. S.↗

LSPRAY-III: A Lagrangian Spray Module

LSPRAY-III is a Lagrangian spray solver developed for application with parallel computing and unstructured grids. It is designed to be massively parallel and could easily be coupled with any existing gas-phase flow and/or Monte Carlo Probability Density Function (PDF) solvers. The solver accommodates the use of an unstructured mesh with mixed elements of either triangular, quadrilateral, and/or tetrahedral type for the gas flow grid representation. It is mainly designed to predict the flow, thermal and transport properties of a rapidly vaporizing spray because of its importance in aerospace application. The manual provides the user with an understanding of various models involved in the spray formulation, its code structure and solution algorithm, and various other issues related to parallelization and its coupling with other solvers. With the development of LSPRAY-III, we have advanced the state-of-the-art in spray computations in several important ways.

Raju, M. S.↗

Twelve Channel Optical Fiber Connector Assembly: From Commercial Off the Shelf to Space Flight Use

The commercial off the shelf (COTS) twelve channel optical fiber MTP array connector and ribbon cable assembly is being validated for space flight use and the results of this study to date are presented here. The interconnection system implemented for the Parallel Fiber Optic Data Bus (PFODB) physical layer will include a 100/140 micron diameter optical fiber in the cable configuration among other enhancements. As part of this investigation, the COTS 62.5/125 microns optical fiber cable assembly has been characterized for space environment performance as a baseline for improving the performance of the 100/140 micron diameter ribbon cable for the Parallel FODB application. Presented here are the testing and results of random vibration and thermal environmental characterization of this commercial off the shelf (COTS) MTP twelve channel ribbon cable assembly. This paper is the first in a series of papers which will characterize and document the performance of Parallel FODB's physical layer from COTS to space flight worthy.

Ott, Melaine N.↗

Bringing OpenCL to Commodity RISC-V CPUs

The importance of open-source hardware has been increasing in recent years with the introduction of the RISC-V Open ISA. This has also accelerated the push for support of the open-source software stack from compiler tools to full-blown operating systems. Parallel computing with today’s Application Programming Interfaces such as OpenCL has proven to be effective at leveraging the parallelism in commodity multi-core processors and programmable parallel accelerators. However, to the best of our knowledge, there is currently no publicly available implementation of OpenCL targeting commodity RISC-V processors that is accessible to the open-source community. Besides opening RISC-V to the existing rich variety of scientific parallel applications, OpenCL also provides access to a unique genre of benchmarks useful in computer architecture research. In this work, we extended an Open-source implementation of OpenCL to target RISC-V CPUs. Our work not only cover commodity multi-core RISC-V processors, but also plethora of low- profile embedded RISC-V CPUs that often do not support atomic instructions or multi-threading.

Tine, Blaise↗