Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “parallel machines”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 433 records · Page 24

BOXKIT

SF-23-067 BoxKit is a library that provides building blocks to parallelize and scale data science, high performance computing, and machine learning applications for block-structured datasets. Spatial data from simulations and experiments can be accessed and managed using tools available in this library when working with more data analysis oriented packages like SciKit (https://github.com/scikit-learn/scikit-learn) and FlowNet (https://github.com/NVIDIA/flownet2-pytorch)

DHRUV, AKASH↗

Method and apparatus for tensile testing of metal foil

A method for obtaining accurate and reproducible results in the tensile testing of metal foils in tensile testing machines is described. Before the test specimen are placed in the machine, foil side edges are worked until they are parallel and flaw free. The specimen are also aligned between and secured to grip end members. An aligning apparatus employed in the method is comprised of an alignment box with a longitudinal bottom wall and two upright side walls, first and second removable grip end members at each end of the box, and a means for securing the grip end members within the box.

Wade, O. W.↗

Parallel processors and nonlinear structural dynamics algorithms and software

The adaptation of a finite element program with explicit time integration to a massively parallel SIMD (single instruction multiple data) computer, the CONNECTION Machine is described. The adaptation required the development of a new algorithm, called the exchange algorithm, in which all nodal variables are allocated to the element with an exchange of nodal forces at each time step. The architectural and C* programming language features of the CONNECTION Machine are also summarized. Various alternate data structures and associated algorithms for nonlinear finite element analysis are discussed and compared. Results are presented which demonstrate that the CONNECTION Machine is capable of outperforming the CRAY XMP/14.

Belytschko, Ted↗

System and component design and test of a 10 hp, 18,000 rpm AC dynamometer utilizing a high frequency AC voltage link, part 1

Hard and soft switching test results conducted with one of the samples of first generation MOS-controlled thyristor (MCTs) and similar test results with several different samples of second generation MCT's are reported. A simple chopper circuit is used to investigate the basic switching characteristics of MCT under hard switching and various types of resonant circuits are used to determine soft switching characteristics of MCT under both zero voltage and zero current switching. Next, operation principles of a pulse density modulated converter (PDMC) for three phase (3F) to 3F two-step power conversion via parallel resonant high frequency (HF) AC link are reviewed. The details for the selection of power switches and other power components required for the construction of the power circuit for the second generation 3F to 3F converter system are discussed. The problems encountered in the first generation system are considered. Design and performance of the first generation 3F to 3F power converter system and field oriented induction moter drive based upon a 3 kVA, 20 kHz parallel resonant HF AC link are described. Low harmonic current at the input and output, unity power factor operation of input, and bidirectional flow capability of the system are shown via both computer and experimental results. The work completed on the construction and testing of the second generation converter and field oriented induction motor drive based upon specifications for a 10 hp squirrel cage dynamometer and a 20 kHz parallel resonant HF AC link is discussed. The induction machine is designed to deliver 10 hp or 7.46 kW when operated as an AC-dynamo with power fed back to the source through the converter. Results presented reveal that the proposed power level requires additional energy storage elements to overcome difficulties with a peak link voltage variation problem that limits reaching to the desired power level. The power level test of the second generation converter after the addition of extra energy storage elements to the HF link are described. The importance of the source voltage level to achieve a better current regulation for the source side PDMC is also briefly discussed. The power levels achieved in the motoring mode of operation show that the proposed power levels achieved in the generating mode of operation can also be easily achieved provided that no mechanical speed limitation were present to drive the induction machine at the proposed power level.

Lipo, Thomas A.↗

NAS Applications and Advanced Algorithms

This paper examines the applications most commonly run on the supercomputers at the Numerical Aerospace Simulation (NAS) facility. It analyzes the extent to which such applications are fundamentally oriented to vector computers, and whether or not they can be efficiently implemented on hierarchical memory machines, such as systems with cache memories and highly parallel, distributed memory systems.

Bailey, David H.↗

NAS Applications and Advanced Architectures

This paper examines the applications most commonly run on the supercomputers at the Numerical Aerospace Simulation (NAS) facility. It analyzes the extent to which such applications are fundamentally oriented to vector computers, and whether or not they can be efficiently implemented on hierarchical memory machines, such as systems with cache memories and highly parallel, distributed memory systems.

Bailey, David H.↗

NDE and SHM Simulation for CFRP Composites

Ultrasound-based nondestructive evaluation (NDE) is a common technique for damage detection in composite materials. There is a need for advanced NDE that goes beyond damage detection to damage quantification and characterization in order to enable data driven prognostics. The damage types that exist in carbon fiber-reinforced polymer (CFRP) composites include microcracking and delaminations, and can be initiated and grown via impact forces (due to ground vehicles, tool drops, bird strikes, etc), fatigue, and extreme environmental changes. X-ray microfocus computed tomography data, among other methods, have shown that these damage types often result in voids/discontinuities of a complex volumetric shape. The specific damage geometry and location within ply layers affect damage growth. Realistic threedimensional NDE and structural health monitoring (SHM) simulations can aid in the development and optimization of damage quantification and characterization techniques. This paper is an overview of ongoing work towards realistic NDE and SHM simulation tools for composites, and also discusses NASA's need for such simulation tools in aeronautics and spaceflight. The paper describes the development and implementation of a custom ultrasound simulation tool that is used to model ultrasonic wave interaction with realistic 3-dimensional damage in CFRP composites. The custom code uses elastodynamic finite integration technique and is parallelized to run efficiently on computing cluster or multicore machines.

Leckey, Cara A. C.↗

Virtual Machine Language 2.1

VML (Virtual Machine Language) is an advanced computing environment that allows spacecraft to operate using mechanisms ranging from simple, time-oriented sequencing to advanced, multicomponent reactive systems. VML has developed in four evolutionary stages. VML 0 is a core execution capability providing multi-threaded command execution, integer data types, and rudimentary branching. VML 1 added named parameterized procedures, extensive polymorphism, data typing, branching, looping issuance of commands using run-time parameters, and named global variables. VML 2 added for loops, data verification, telemetry reaction, and an open flight adaptation architecture. VML 2.1 contains major advances in control flow capabilities for executable state machines. On the resource requirements front, VML 2.1 features a reduced memory footprint in order to fit more capability into modestly sized flight processors, and endian-neutral data access for compatibility with Intel little-endian processors. Sequence packaging has been improved with object-oriented programming constructs and the use of implicit (rather than explicit) time tags on statements. Sequence event detection has been significantly enhanced with multi-variable waiting, which allows a sequence to detect and react to conditions defined by complex expressions with multiple global variables. This multi-variable waiting serves as the basis for implementing parallel rule checking, which in turn, makes possible executable state machines. The new state machine feature in VML 2.1 allows the creation of sophisticated autonomous reactive systems without the need to develop expensive flight software. Users specify named states and transitions, along with the truth conditions required, before taking transitions. Transitions with the same signal name allow separate state machines to coordinate actions: the conditions distributed across all state machines necessary to arm a particular signal are evaluated, and once found true, that signal is raised. The selected signal then causes all identically named transitions in all present state machines to be taken simultaneously. VML 2.1 has relevance to all potential space missions, both manned and unmanned. It was under consideration for use on Orion.

Riedel, Joseph E.↗

Particle simulation in a multiprocessor environment

A parallel implementation of a particle simulation method that is portable between a wide class of multiprocessor computers is presented. A fine grain spatial decomposition is utilized where several subdomains having a regular structure are computed at each processing node. This leads directly to an efficient and straightforward load balancing scheme if the number of subdomains at each processor is permitted to vary in an appropriate manner. Three dimensional simulations incorporating full thermochemical nonequilibrium are possible using the resulting code. Vectorizable algorithms are retained from earlier work allowing efficient use of deeply pipelined node processors where available. Performance results are presented from three different machine architectures demonstrating the portability of the code. On a 128-node Intel iPSC/860, performance is twice that of a single Cray-Y/MP CPU running a highly vectorized simulation code. Speedup is linear over the full range of number of processors on all target machines, indicating scalability of the method to higher degrees of parallelism.

Mcdonald, Jeffrey D.↗

To Exascale and Beyond—The Simple Cloud-Resolving E3SM Atmosphere Model (SCREAM), a Performance Portable Global Atmosphere Model for Cloud-Resolving Scales

The new generation of heterogeneous CPU/GPU computer systems offer much greater computational performance but are not yet widely used for climate modeling. One reason for this is that traditional climate models were written before GPUs were available and would require an extensive overhaul to run on these new machines. In addition, even conventional “high–resolution” simulations don't currently provide enough parallel work to keep GPUs busy, so the benefits of such overhaul would be limited for the types of simulations climate scientists are accustomed to. The vision of the Simple Cloud-Resolving Energy Exascale Earth System (E3SM) Atmosphere Model (SCREAM) project is to create a global atmospheric model with the architecture to efficiently use GPUs and horizontal resolution sufficient to fully take advantage of GPU parallelism. After 5 years of model development, SCREAM is finally ready for use. In this paper, we describe the design of this new code, its performance on both CPU and heterogeneous machines, and its ability to simulate real-world climate via a set of four 40 day simulations covering all 4 seasons of the year.

54 ENVIRONMENTAL SCIENCES↗

RISC Processors and High Performance Computing

This tutorial will discuss the top five RISC microprocessors and the parallel systems in which they are used. It will provide a unique cross-machine comparison not available elsewhere. The effective performance of these processors will be compared by citing standard benchmarks in the context of real applications. The latest NAS Parallel Benchmarks, both absolute performance and performance per dollar, will be listed. The next generation of the NPB will be described. The tutorial will conclude with a discussion of future directions in the field. Technology Transfer Considerations: All of these computer systems are commercially available internationally. Information about these processors is available in the public domain, mostly from the vendors themselves. The NAS Parallel Benchmarks and their results have been previously approved numerous times for public release, beginning back in 1991.

Bailey, David H.↗

An efficient parallel algorithm for the solution of a tridiagonal linear system of equations

Tridiagonal linear systems of equations are solved on conventional serial machines in a time proportional to N, where N is the number of equations. The conventional algorithms do not lend themselves directly to parallel computations on computers of the ILLIAC IV class, in the sense that they appear to be inherently serial. An efficient parallel algorithm is presented in which computation time grows as log sub 2 N. The algorithm is based on recursive doubling solutions of linear recurrence relations, and can be used to solve recurrence relations of all orders.

Stone, H. S.↗

Wind Power as a Virtual Synchronous Generator (WindVSG)

This project investigated the theory, implemented it in hardware, and validated the Wind as a Virtual Isochronous Generator (WindVSG) concept by combining the advantages of modern dynamic inverter technologies with static, dynamic, and transient electromechanical properties of synchronous machines. During this project we demonstrated how to control the inverters of wind turbine generators (wind alone or in parallel with other GFM sources, such battery energy storage) so that wind power behaves like a synchronous machine-based power plant with a conventional prime mover. For this purpose, testing was conducted at NLR ARIES facility with real 2.5 MW wind-turbine generator operating in GFM mode under dynamic and transient conditions. The team also developed models and conducted simulations for GFM wind power to evaluate stability impacts of GFM operation on power grid. This report describes efforts by the NLR team working in collaboration GE Vernova during 3-year project.

17 WIND ENERGY↗

Polyphony: A Workflow Orchestration Framework for Cloud Computing

Cloud Computing has delivered unprecedented compute capacity to NASA missions at affordable rates. Missions like the Mars Exploration Rovers (MER) and Mars Science Lab (MSL) are enjoying the elasticity that enables them to leverage hundreds, if not thousands, or machines for short durations without making any hardware procurements. In this paper, we describe Polyphony, a resilient, scalable, and modular framework that efficiently leverages a large set of computing resources to perform parallel computations. Polyphony can employ resources on the cloud, excess capacity on local machines, as well as spare resources on the supercomputing center, and it enables these resources to work in concert to accomplish a common goal. Polyphony is resilient to node failures, even if they occur in the middle of a transaction. We will conclude with an evaluation of a production-ready application built on top of Polyphony to perform image-processing operations of images from around the solar system, including Mars, Saturn, and Titan.

Space Exploration,↗

Local time stepping for the shallow water equations in MPAS

In this work we assess the performance of a set of local time-stepping (LTS) schemes for the shallow water equations implemented in the Model for Prediction Across Scales (MPAS). The goal of LTS is to speed up the simulation by allowing different time-steps on different regions of the computational grid. The LTS schemes considered here were originally introduced by Hoang et al. (2019) [26], who laid out the mathematical foundation of the methods. Here, the authors take on the task of presenting a fast, efficient and scalable parallel implementation of these LTS methods on high performance computing machines, with the aim to provide a recipe for other climate modeling groups that may be interested in employing LTS algorithms in their codes. As a matter of fact, even if MPAS is our framework of choice, our approach is general enough and could be of interest to other groups beyond the MPAS community. Due to their nature, LTS methods possess an inherent load imbalance that needs to be carefully addressed in order to obtain efficient scalability. Even more important is the far from trivial task of computing the right-hand side terms only on specific LTS regions during the time-stepping procedure. An inefficient handling of this task causes a drastic decay of the CPU time performance, making the LTS algorithms practically of no use. The emphasis of the present work is therefore on the computational and parallel aspects of the LTS methods, whose proper treatment is crucial to make the methods run faster against existing strategies, such as for instance high-order explicit global time-stepping schemes. This is in fact the ultimate goal of using an LTS procedure and it is the one to which we direct all our optimization efforts.

97 MATHEMATICS AND COMPUTING↗

Automated Performance Prediction of Message-Passing Parallel Programs

The increasing use of massively parallel supercomputers to solve large-scale scientific problems has generated a need for tools that can predict scalability trends of applications written for these machines. Much work has been done to create simple models that represent important characteristics of parallel programs, such as latency, network contention, and communication volume. But many of these methods still require substantial manual effort to represent an application in the model's format. The NIK toolkit described in this paper is the result of an on-going effort to automate the formation of analytic expressions of program execution time, with a minimum of programmer assistance. In this paper we demonstrate the feasibility of our approach, by extending previous work to detect and model communication patterns automatically, with and without overlapped computations. The predictions derived from these models agree, within reasonable limits, with execution times of programs measured on the Intel iPSC/860 and Paragon. Further, we demonstrate the use of MK in selecting optimal computational grain size and studying various scalability metrics.

Block, Robert J.↗

Optimal Operation of a Hybrid Hydraulic Electric Architecture (HHEA) for Off-Road Vehicles Over Discrete Operating Decisions

Many off-highway machines including construction and agriculture equipment use hydraulics for power transmission and throttling as a means for control. Trends towards better efficiency and electrification have led to the creation of a novel Hybrid Hydraulic-Electric Architecture (HHEA) which could significantly reduce energy consumption and maintain control performance, even in machines that are too large to be directly electrified. This is achieved by using a set of common pressure rails to transmit the majority of power via power dense hydraulics and modulating the power with small electric motor-drives to achieve precise control. This paper proposes a computationally efficient method for computing the optimal sequence of pressure rail selections for the HHEA over finite drive cycles. This is useful for fairly comparing the novel architecture’s energy performance to existing architectures and for use in iterative optimal design of the architecture. The optimal control technique makes use of a static model of the architecture and losses. Constraints are enforced such that the drive cycle is repeatable. The constrained optimal operation is solved using a Lagrange multiplier technique that transforms the optimization into a small set of sub-problems by considering combinations of active constraints. Each of these sub-problems can be solved efficiently because loss calculations for all time steps can be computed in parallel. A case study of an off-road construction machine demon-strates that the HHEA reduces energy consumption by 2/3 compared to the baseline load sensing architecture.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

An efficient parallel algorithm for the solution of a tridiagonal linear system of equations.

Tridiagonal linear systems of equations can be solved on conventional serial machines in a time proportional to N, where N is the number of equations. The conventional algorithms do not lend themselves directly to parallel computation on computers of the Illiac IV class, in the sense that they appear to be inherently serial. An efficient parallel algorithm is presented in which computation time grows as log(sub-2) N. The algorithm is based on recursive doubling solutions of linear recurrence relations, and can be used to solve recurrence relations of all orders.

Stone, H. S.↗