Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “parallel machines”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 397 records · Page 22

Transferring ecosystem simulation codes to supercomputers

Many ecosystem simulation computer codes have been developed in the last twenty-five years. This development took place initially on main-frame computers, then mini-computers, and more recently, on micro-computers and workstations. Supercomputing platforms (both parallel and distributed systems) have been largely unused, however, because of the perceived difficulty in accessing and using the machines. Also, significant differences in the system architectures of sequential, scalar computers and parallel and/or vector supercomputers must be considered. We have transferred a grassland simulation model (developed on a VAX) to a Cray Y-MP/C90. We describe porting the model to the Cray and the changes we made to exploit the parallelism in the application and improve code execution. The Cray executed the model 30 times faster than the VAX and 10 times faster than a Unix workstation. We achieved an additional speedup of 30 percent by using the compiler's vectoring and 'in-line' capabilities. The code runs at only about 5 percent of the Cray's peak speed because it ineffectively uses the vector and parallel processing capabilities of the Cray. We expect that by restructuring the code, it could execute an additional six to ten times faster.

Skiles, J. W.↗

On Efficient Parallel Implementation of Moving Body Overset Grid Methods

An investigation into the parallel performance of moving-body overset grid methods will be presented. Parallel versions of the OVERFLOW flow solver, DCF3D domain connectivity software, and SIXDO six-degree-of-freedom routine are coupled with an automatic load balance routine and tested for 3D Navier-Stokes calculations on the IBM SP2. The primary source of parallel inefficiency in moving and problems are the domain connectivity costs with DCF 3D. Although this algorithm constitutes a relatively low fraction of the total solution cost (e.g. 10-20%) in calculations on serial machines, the consequently cause a significant degradation in the overall parallel performance. The paper will highlight some approaches for improving the scalability of DCF3D. The paper will present results of a proposed new load balancing scheme that seeks more equal distribution of the inter-grid boundary points in order to more evenly load balance the donor search costs associated with DCF3D. Some preliminary results will also be given from a new solution-adaption algorithm coupled with OVERFLOW which incorporates overset cartesian grids with various levels of refinement. The measured parallel performance from a descending delta-wing configuration and a generic store-separation from a wing/pylon case will be presented.

Wissink, Andrew M.↗

Pairwise‐Parallel Entangling Gates on Orthogonal Modes in a Trapped‐Ion Chain

Abstract Parallel operations are important for both near‐term quantum computers and larger‐scale fault‐tolerant machines because they reduce execution time and qubit idling. This study proposes and implements a pairwise‐parallel gate scheme on a trapped‐ion quantum computer. The gates are driven simultaneously on different sets of orthogonal motional modes of a trapped‐ion chain. This work demonstrates the utility of this scheme by creating a Greenberger‐Horne‐Zeilinger (GHZ) state in one step using parallel gates with one overlapping qubit. It also shows its advantage for circuits by implementing a digital quantum simulation of the dynamics of an interacting spin system, the transverse‐field Ising model. This method effectively extends the available gate depth by up to two times with no overhead when no overlapping qubit is involved, apart from additional initial cooling. This scheme can be easily applied to different trapped‐ion qubits and gate schemes, broadly enhancing the capabilities of trapped‐ion quantum computers.

Optics↗

Scalable Second Order Optimization for Machine Learning

Many machine learning (ML) training tasks are essentially optimization processes that would at first glance appear eminently parallelizable and scalable. However, effective acceleration of these tasks with scalable parallel hardware has proven to be elusive. While standard methods for machine learning, e.g., stochastic gradient descent (SGD) for DNNs, tend to be resource efficient, they appear to be fundamentally sequential in nature.

97 MATHEMATICS AND COMPUTING↗

Parallel image compression

A parallel compression algorithm for the 16,384 processor MPP machine was developed. The serial version of the algorithm can be viewed as a combination of on-line dynamic lossless test compression techniques (which employ simple learning strategies) and vector quantization. These concepts are described. How these concepts are combined to form a new strategy for performing dynamic on-line lossy compression is discussed. Finally, the implementation of this algorithm in a massively parallel fashion on the MPP is discussed.

Reif, John H.↗

Virtual Self-Excited Induction Generator-Based Grid-Forming Inverter Control for Robust Voltage Regulation Under Nonideal Loading

This paper presents a generator-inspired control methodology for grid-forming (GFM) inverters that deliberately emulates a self-excited induction generator so that the inverter can hold its voltage and frequency under difficult loading and severe terminal disturbances across wide voltage and frequency ranges. The design integrates a Lyapunov energy function-based inner loop to provide high bandwidth and strong disturbance rejection, and it complements this with a passivity-based argument that furnishes a coherent large-signal stability guarantee beyond small-signal limits. Analytical insights are developed via the Krylov-Bogoliubov-Mitropolsky averaging method, which reveals an intrinsic resistive droop characteristic; these closed-form relations both explain the observed dynamics and yield simple, decentralized tuning rules. The methodology is validated on a controller-hardware-in-the-loop platform and exercised in real time across balanced, unbalanced, and nonlinear loads, as well as during parallel operation. Across these scenarios, the inverter maintains balanced three-phase voltages, limits harmonic content, settles quickly with well-damped transients, and remains resilient when multiple units operate in parallel. The contributions are a self-excited-machine-inspired GFM controller with enhanced dynamic performance and robustness, a single stability rationale grounded in passivity, closed-form expressions that guide tuning, and comprehensive hardware-in-the-loop validations demonstrating effectiveness and superiority under challenging operating conditions.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Graph-Learning-Assisted State and Event Tracking for Solar-Penetrated Power Grids with Heterogeneous Data Sources

Unlike transmission systems, distribution systems do not typically contain sufficient metering to enable real-time state estimation. The lack of sufficient real-time measurements prohibits accurate and timely monitoring of the state of distribution systems. As a result, control and optimal operation of distribution systems, especially those containing large numbers of renewable generation units are not possible without proper data and information about the current state of the system. The main motivation of this project is to address this shortcoming by developing an approach which provides “predicted” real-time measurements so that they can be used to execute a distribution system state estimator. Thus, the objective of the project is to make the distribution systems fully observable, such that the hosting capacity for solar generation can be accurately estimated, and unnecessary solar curtailments can be avoided. In order to accomplish this goal, the project investigated the use of a grid-model-informed machine learning (ML) tool which integrates heterogeneous data streams obtained from AMI meters, SCADA as well as PMU measurements and created synchronous measurement snapshots for the state estimator (SE); and developed a hybrid robust SE which provides not only accurate state estimates but also real-time feedback for the ML model refinement.

14 SOLAR ENERGY↗

Exago TM Users Manual: Version 1.0

The Exascale Grid Optimization (ExaGOTM) toolkit is an open source package for solving large-scale power grid optimization problems on parallel and distributed architectures, particularly targeted for exascale machines with heteregenous architectures (GPU). This manual is a guide to ExaGO's working including installation, formulation, and usage.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Grooves reduce aircraft drag

Aerodynamic drag can be reduced by many small longitudinal grooves machined in aircraft skin. Experiments show that grooves parallel to airflow reduce drag by 4 to 7 percent. Reduced drag translates into reduced engine power required to overcome drag and ultimately to lower fuel consumption.

Walsh, M. J.↗

Parallel solution of triangular systems of equations

The solution on parallel computers of systems of equations Lx = b, where L is lower triangular, is considered. Some authors have suggested that it is difficult to solve such systems in parallel on message-passing, local-memory machines when L is stored by columns. It is shown here that this is not necessarily the case if the machine can accomplish fan-in communication with reasonable efficiency.

Romine, Charles H.↗

Distributed intelligence for supervisory control

Supervisory control systems must deal with various types of intelligence distributed throughout the layers of control. Typical layers are real-time servo control, off-line planning and reasoning subsystems and finally, the human operator. Design methodologies must account for the fact that the majority of the intelligence will reside with the human operator. Hierarchical decompositions and feedback loops as conceptual building blocks that provide a common ground for man-machine interaction are discussed. Examples of types of parallelism and parallel implementation on several classes of computer architecture are also discussed.

Wolfe, W. J.↗

Distributed control architecture for real-time telerobotic operation

The emerging field of telerobotics places new demands on control system architecture to allow both autonomous operations and natural human-machine interfacing. The feasibility of multiprocessor systems performing parallel control computations is realizable. A practical distribution of control processors is presented and the issues involved in the realization of this architecture are discussed. A prototype dual axis controller based on the NOVIX computer is described, and results of its implementation are discussed. Application of this type of control system to a replicated, redundant manipulator system is also described.

Martin, H. L.↗

Motion detection in astronomical and ice floe images

Two approaches are presented for establishing correspondence between small areas in pairs of successive images for motion detection. The first one, based on local correlation, is used on a pair of successive Voyager images of the Jupiter which differ mainly in locally variable translations. This algorithm is implemented on a sequential machine (VAX 780) as well as the Massively Parallel Processor (MPP). In the case of the sequential algorithm, the pixel correspondence or match is computed on a sparse grid of points using nonoverlapping windows (typically 11 x 11) by local correlations over a predetermined search area. The displacement of the corresponding pixels in the two images is called the disparities to cubic surfaces. The disparities at points where the error between the computed values and the surface values exceeds a particular threshold are replaced by the surface values. A bilinear interpolation is then used to estimate disparities at all other pixels between the grid points. When this algorithm was applied at the red spot in the Jupiter image, the rotating velocity field of the storm was determined. The second method of motion detection is applicable to pairs of images in which corresponding areas can experience considerable translation as well as rotation.

Manohar, M.↗

The alignment-distribution graph

Implementing a data-parallel language such as Fortran 90 on a distributed-memory parallel computer requires distributing aggregate data objects (such as arrays) among the memory modules attached to the processors. The mapping of objects to the machine determines the amount of residual communication needed to bring operands of parallel operations into alignment with each other. We present a program representation called the alignment distribution graph that makes these communication requirements explicit. We describe the details of the representation, show how to model communication cost in this framework, and outline several algorithms for determining object mappings that approximately minimize residual communication.

Chatterjee, Siddhartha↗

The alignment-distribution graph

Implementing a data-parallel language such as Fortran 90 on a distributed-memory parallel computer requires distributing aggregate data objects (such as arrays) among the memory modules attached to the processors. The mapping of objects to the machine determines the amount of residual communication needed to bring operands of parallel operations into alignment with each other. We present a program representation called the alignment-distribution graph that makes these communication requirements explicit. We describe the details of the representation, show how to model communication cost in this framework, and outline several algorithms for determining object mappings that approximately minimize residual communication.

Chatterjee, Siddhartha↗

Performance Comparison of HPF and MPI Based NAS Parallel Benchmarks

Compilers supporting High Performance Form (HPF) features first appeared in late 1994 and early 1995 from Applied Parallel Research (APR), Digital Equipment Corporation, and The Portland Group (PGI). IBM introduced an HPF compiler for the IBM RS/6000 SP2 in April of 1996. Over the past two years, these implementations have shown steady improvement in terms of both features and performance. The performance of various hardware/ programming model (HPF and MPI) combinations will be compared, based on latest NAS Parallel Benchmark results, thus providing a cross-machine and cross-model comparison. Specifically, HPF based NPB results will be compared with MPI based NPB results to provide perspective on performance currently obtainable using HPF versus MPI or versus hand-tuned implementations such as those supplied by the hardware vendors. In addition, we would also present NPB, (Version 1.0) performance results for the following systems: DEC Alpha Server 8400 5/440, Fujitsu CAPP Series (VX, VPP300, and VPP700), HP/Convex Exemplar SPP2000, IBM RS/6000 SP P2SC node (120 MHz), NEC SX-4/32, SGI/CRAY T3E, and SGI Origin2000. We would also present sustained performance per dollar for Class B LU, SP and BT benchmarks.

Saini, Subhash↗

Issues on Reproducibility/Reliability of Magnetic NDE Methods

One of the critical elements related to the practicality of any NDE technique is its reproducibility under nominally the same inspection conditions. The results of certain test methodologies, however, are not always repeatable and understanding the origin of the irreproducibility is often as critical as obtaining reproducible results. One example is the characterization of residual stress in structural ferromagnets using the magnetoacoustic (MAC) method. Although it has not been widely publicized, the test results of this method are known to be time-dependent. Two distinct types of time dependencies have been observed during testing. The first type has a clearly definable relaxation time, while no such trend has been observed for the second. The purpose of the present study is to systematically investigate the time dependence of the second type, to find out the range and, if possible, the origin of the variation in the test results. For this, MAC curves were obtained under various stress levels and the tests were repeated over time. Particular attention was given to whether noise in the measuring device or a change in the laboratory environment could have been a contributing factor. The steel samples used for the study were cut from C- and U-class railroad wheels. The MAC behavior of these samples was reported previously. Each steel sample was first machined to be a cylindrical rod of 3.175 cm (1.25 in) in diameter and 26.67 cm (10.5 in) in length. The center portion of the samples were further machined to form a pair of flat and parallel surfaces for the launch and reflection of the ultrasonic pulses. Data acquisition involved two major elements; magnetic and acoustic measurements. Throughout the experiment the net magnetic induction, B, was measured by integrating the induction pickup coil output using an integrating fluxmeter. The acoustic measurements employed the phase-locked technique which will be described in the following.

Namkung, M.↗

Asynchronous distributed-memory task-parallel algorithm for compressible flows on unstructured 3D Eulerian grids

Here, we discuss the implementation of a finite element method, used to numerically solve the Euler equations of compressible flows, using an asynchronous runtime system (RTS). The algorithm is implemented for distributed-memory machines, using stationary unstructured 3D meshes, combining data-, and task-parallelism on top of the Charm++ RTS. Charm++’s execution model is asynchronous by default, allowing arbitrary overlap of computation and communication. Task-parallelism allows scheduling parts of an algorithm independently of, or dependent on, each other. Built-in automatic load balancing enables continuous redistribution of computational load by migration of work units based on real-time CPU load measurement. The RTS also features automatic checkpointing, fault tolerance, resilience against hardware failure, and supports power-, and energy-aware computation. We demonstrate scalability up to 25 x 10 9 cells at $\mathscr{O}$10 4 compute cores and the benefits of automatic load balancing for irregular workloads. The full source code with documentation is available at https://quinoacomputing.org.

42 ENGINEERING↗