Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “parallel machines”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 379 records · Page 21

Fatigue crack propagation behavior of a single crystalline superalloy

Crack propagation mechanisms occurring at various temperatures in a single crystalline Ni-base alloy, Rene N4, were investigated. The rates of crack growth at 21, 704, 927, 1038, and 1093 C were measured in specimens with 001-line and 110-line directions parallel to the load axis and the machined notch, respectively, using a pulsed dc potential drop apparatus, and the fracture surfaces at each temperature were examined using SEM. Crack growth rates (CGRs) for specimens tested at or below 927 C were similar, while at two higher temperatures, the CGRs were about an order of magnitude higher than at the lower temperatures. Results of SEM observations showed that surface morphologies depended on temperature.

Lerch, B. A.↗

A biconjugate gradient type algorithm on massively parallel architectures

The biconjugate gradient (BCG) method is the natural generalization of the classical conjugate gradient algorithm for Hermitian positive definite matrices to general non-Hermitian linear systems. Unfortunately, the original BCG algorithm is susceptible to possible breakdowns and numerical instabilities. Recently, Freund and Nachtigal have proposed a novel BCG type approach, the quasi-minimal residual method (QMR), which overcomes the problems of BCG. Here, an implementation is presented of QMR based on an s-step version of the nonsymmetric look-ahead Lanczos algorithm. The main feature of the s-step Lanczos algorithm is that, in general, all inner products, except for one, can be computed in parallel at the end of each block; this is unlike the other standard Lanczos process where inner products are generated sequentially. The resulting implementation of QMR is particularly attractive on massively parallel SIMD architectures, such as the Connection Machine.

Freund, Roland W.↗

Multi-dimensional high order essentially non-oscillatory finite difference methods in generalized coordinates

The nonlinear stability of compact schemes for shock calculations is investigated. In recent years compact schemes were used in various numerical simulations including direct numerical simulation of turbulence. However to apply them to problems containing shocks, one has to resolve the problem of spurious numerical oscillation and nonlinear instability. A framework to apply nonlinear limiting to a local mean is introduced. The resulting scheme can be proven total variation (1D) or maximum norm (multi D) stable and produces nice numerical results in the test cases. The result is summarized in the preprint entitled 'Nonlinearly Stable Compact Schemes for Shock Calculations', which was submitted to SIAM Journal on Numerical Analysis. Research was continued on issues related to two and three dimensional essentially non-oscillatory (ENO) schemes. The main research topics include: parallel implementation of ENO schemes on Connection Machines; boundary conditions; shock interaction with hydrogen bubbles, a preparation for the full combustion simulation; and direct numerical simulation of compressible sheared turbulence.

Shu, Chi-Wang↗

Unstructured grids on SIMD torus machines

Unstructured grids lead to unstructured communication on distributed memory parallel computers, a problem that has been considered difficult. Here, we consider adaptive, offline communication routing for a SIMD processor grid. Our approach is empirical. We use large data sets drawn from supercomputing applications instead of an analytic model of communication load. The chief contribution of this paper is an experimental demonstration of the effectiveness of certain routing heuristics. Our routing algorithm is adaptive, nonminimal, and is generally designed to exploit locality. We have a parallel implementation of the router, and we report on its performance.

Bjorstad, Petter E.↗

Fabrication of an Absorber-Coupled MKID Detector

Absorber-coupled microwave kinetic inductance detector (MKID) arrays were developed for submillimeter and far-infrared astronomy. These sensors comprise arrays of lambda/2 stepped microwave impedance resonators patterned on a 1.5-mm-thick silicon membrane, which is optimized for optical coupling. The detector elements are supported on a 380-mm-thick micro-machined silicon wafer. The resonators consist of parallel plate aluminum transmission lines coupled to low-impedance Nb microstrip traces of variable length, which set the resonant frequency of each resonator. This allows for multiplexed microwave readout and, consequently, good spatial discrimination between pixels in the array. The transmission lines simultaneously act to absorb optical power and employ an appropriate surface impedance and effective filling fraction. The fabrication techniques demonstrate high-fabrication yield of MKID arrays on large, single-crystal membranes and sub-micron front-to-back alignment of the micro strip circuit. An MKID is a detector that operates upon the principle that a superconducting material s kinetic inductance and surface resistance will change in response to being exposed to radiation with a power density sufficient to break its Cooper pairs. When integrated as part of a resonant circuit, the change in surface impedance will result in a shift in its resonance frequency and a decrease of its quality factor. In this approach, incident power creates quasiparticles inside a superconducting resonator, which is configured to match the impedance of free space in order to absorb the radiation being detected. For this reason MKIDs are attractive for use in large-format focal plane arrays, because they are easily multiplexed in the frequency domain and their fabrication is straightforward. The fabrication process can be summarized in seven steps: (1) Alignment marks are lithographically patterned and etched all the way through a silicon on insulator (SOI) wafer, which consists of a thin silicon membrane bonded to a thick silicon handle wafer. (2) The metal microwave circuitry on the front of the membrane is patterned and etched. (3) The wafer is then temporarily bonded with wafer wax to a Pyrex wafer, with the SOI side abutting the Pyrex. (4) The silicon handle component of the SOI wafer is subsequently etched away so as to expose the membrane backside. (5) The wafer is flipped over, and metal microwave circuitry is patterned and etched on the membrane backside. Furthermore, cuts in the membrane are made so as to define the individual detector array chips. (6) Silicon frames are micromachined and glued to the silicon membrane. (7) The membranes, which are now attached to the frames, are released from the Pyrex wafer via dissolution of the wafer wax in acetone.

Brown, Ari↗

Machine Learning Algorithm Performance on the Lucata Computer

A new parallel computing paradigm (processor in memory, or PIM) has recently become available, one that uses many lightweight threads, and where each thread migrates automatically to the memory used by that thread. Our effort focuses on understanding how suitable this architecture is for our application, and whether the hardware can sustain speedups as high as the system size permits. In particular we explore the kind of code optimizations needed, and how well optimized code scales. This paper describes some of the those optimizations, and the payoff in terms of scaling.

Kogge, Peter↗

A New Architecture for Parallelization of Complex Spacecraft Trajectory Optimization Scans

This paper describes CopScanner, a new component of the Copernicus ecosystem for spacecraft trajectory design and optimization. CopScanner is a Python library being developed at the NASA JSC which enables easy parallelization of Copernicus scans. CopScanner is currently being developed and implemented for production of Copernicus trajectory scans for upcoming Artemis Missions (Artemis II and beyond). On the backend, CopScanner utilizes Dask, an open-source Python library for parallel computing which enables parallelization over both multi-core local machines and large-scale distributed computing clusters. CopScanner abstracts the trajectory scanning process into a DAG which is constructed using a chain of individual subscans. Each node in the DAG executes a python module, called the callable, for which there are built-in defaults, or users may specify their own. Support for custom callables makes CopScanner a versatile trajectory optimization software. All output files and associated metadata from a CopScanner scan are compressed and stored in a two-file output, collectively called the FileStore, consisting of a SQLite database and a compressed JSON MessagePack file, for which CopScanner provides a Python class for interaction.

Quentin Moore↗

Generating fractal-like surfaces on general purpose mesh-connected computers

Realistic images of natural surfaces are often generated using computationally expensive stochastic modeling techniques. Here a parallel procedure to generate such models is presented. The target machines are general-purpose mesh-connected computers. The complexity of the procedure is similar to that of a proposed special-purpose parallel fractal generator.

Wainer, Michael↗

Integrating machine learning interatomic potentials with hybrid reverse Monte Carlo structure refinements in RMCProfile

Structure refinement with reverse Monte Carlo (RMC) is a powerful tool for interpreting experimental diffraction data. To ensure that the under-constrained RMC algorithm yields reasonable results, the hybrid RMC approach applies interatomic potentials to obtain solutions that are both physically sensible and in agreement with experiment. To expand the range of materials that can be studied with hybrid RMC, we have implemented a new interatomic potential constraint in RMCProfile that grants flexibility to apply potentials supported by the Large-scale Atomic/Molecular Massively Parallel Simulator ( LAMMPS ) molecular dynamics code. This includes machine learning interatomic potentials, which provide a pathway to applying hybrid RMC to materials without currently available interatomic potentials. To this end, we present a methodology to use RMC to train machine learning interatomic potentials for hybrid RMC applications.

Cuillier, Paul↗

A physics-based ensemble machine-learning approach to identifying a relationship between lightning indices and binary lightning hazard

To convert lightning indices generated by numerical weather prediction experiments into binary lightning hazard, a machine-learning tool was developed. This tool, consisting of parallel multilayer perceptron classifiers, was trained on an ensemble of planetary boundary layer schemes and microphysics parameterizations that generated four different lightning indices over 1 week. In a subsequent week, the multi-physics ensemble was applied and the machine-learning tool was used to evaluate the accuracy. Unintuitively, the machine-learning tool performed better on the testing dataset than the training dataset. Much of the error may be attributed to mischaracterizing the convection. The combination of the machine learning model and simulations could not differentiate between cloud-to-cloud lightning and cloud-to-ground lightning, despite being trained on cloud-to-ground lightning. It was found that the simulation most representative of the local operational model was the most accurate simulation tested.

54 ENVIRONMENTAL SCIENCES↗

Microprocessor arrays for large scale computation

An important new direction in computer architecture centers around the achievement of very high computational power (capacity, speed and reliability) through the use of tens of thousands of microprocessors, micromemories, and switch modules, all interconnected into a large homogeneous network using one of certain advanced connection schemes. When surrounded and supported by conventional computers and memories, such a machine holds potential for out-performing both conventional and array-based computers of the mid-1980's by one to two orders of magnitude, at least for particular classes of applications amenable to high parallelism, such as aerodynamic simulation. The homogeneous feature of this machine concept also implies size extendability, fault tolerance, and improved flexibility to handle a variety of algorithms of interest. Current work is addressing the design of technologically efficient interconnection configurations and the development of new computation algorithms that are especially efficient for highly parallel computation.

Kautz, W. H.↗

A Parallel Alternative for Energy-Efficient Neural Network Training and Inferencing

Energy efficiency of training and inferencing with large neural network models is a critical challenge facing the future of sustainable large-scale machine learning workloads. This paper introduces an alternative strategy, called phantom parallelism, to minimize the net energy consumption of traditional tensor (model) parallelism, the most energy-inefficient component of large neural network training. The approach is presented in the context of feed-forward network architectures as a preliminary, but comprehensive, proof-of-principle study of the proposed methodology. We derive new forward and backward propagation operators for phantom parallelism, implement them as custom autograd operations within an end-to-end phantom parallel training pipeline and compare its parallel performance and energy-efficiency against those of conventional tensor parallel training pipelines. Formal analyses that predict lower bandwidth and FLOP counts are presented with supporting empirical results on up to 256 GPUs that corroborate these gains. Experiments are shown to deliver ∼50% reduction in the energy consumed to train FFNs using the proposed phantom parallel approach when compared with conventional tensor parallel methods. Additionally, the proposed approach is shown to train smaller phantom models to the same model loss on smaller GPU counts as larger tensor parallel models on larger GPU counts offering the possibility for even greater energy savings.

Seal, Sudip [ORNL] (ORCID:0000000332330656)↗

Data-parallel lower-upper relaxation method for reacting flows

The implicit lower-upper symmetric Gauss-Seidel (LU-SGS) method of Yoon and Jameson is modified for use on massively parallel computers. The method has been implemented on the Thinking Machines CM-5 and the MasPar MP-1 and MP-2, where large percentages of the theoretical peak floating point performance are obtained. It is shown that the new data-parallel LU relaxation method has better convergence properties than the original method for two different inviscid compressible flow simulations. The convergence is also improved for five-species reacting air computations. The performance of the method on various partitions of the CM-5 and on the MasPar computers is discussed. The new method shows promise for the efficient simulation of very large perfect gas and reacting flows.

Candler, Graham V.↗

Did the GPU obfuscate the load imbalance in my MPI simulation?

The current proliferation of GPU-based HPC systems necessitates a method for assessing the performance of simulations on heterogeneous machines. The addition of GPUs to a system adds multiple hierarchical levels of parallelism to the node architecture. In this paper, we demonstrate that the traditional load imbalance metric is insufficient for capturing the load imbalance on GPU-based machines, since it treats the GPU as a monolithic entity and ignores the internal parallelism. We propose a new hierarchical metric that improves the correlation of measured performance and application workload by up to 20.61%. Using our metric for determining application load instead of the traditional metric as the input for the load balancing algorithm reduces the residual load imbalance by up to 4× in our application.

Eberius, David↗

Large Scale MD to Predict Epitope Regions in HIV Env [Slides]

Highly dense carbohydrates located on the surface of the HIV Env protein play a key role in immune evasion. Such evolutionary adaptation hampers any attempt to obtain a full mechanistic understanding of the role played by the glycans in protecting the virus against an effective immune response. Moreover, and due to their chemical variability, an accurate molecular understanding of the so called “glycan shield” is still limited by the lack of effective resolution of state-of-the-art experimental technics. Here, we have used extensive computational modelling in order to fill this gap, addressing the presence of a large glycan variability as observed experimentally. Based on an automated pipeline, we were able to assemble, set-up and simulate via Molecular dynamics hundreds of different glycosylated Env variants at nearly atomic resolution, recapitulating the glycosylation distributions observed experimentally. Results from these simulations were subjected to machine learning and very accurate prediction of simulation derived glycan shielding areas of each glycan as a function of static sequence features. Such predictive models of per-glycan shielding, incorporating both glycan dynamics and heterogeneity, were used to develop a novel sequence-based glycan shield mapping strategy. Parallel to these studies, we also developed an accurate machine learning approach to predict glycan heterogeneity data using sequence features and found good prediction accuracy.

59 BASIC BIOLOGICAL SCIENCES↗

SPARC-X: Quantum simulations at extreme scale - reactive dynamics from first principles

We have developed the massively parallel electronic structure code SPARC-X: a computational framework for performing Kohn-Sham Density Functional Theory (DFT) calculations that can scale linearly with the number of atoms in the system, while being able to leverage petascale and emerging exascale parallel computers to study chemical phenomena at unprecedented length and time scales. SPARC-X exploits a recent breakthrough in electronic structure methodologies: systematically improvable, strictly local, orthonormal, discontinuous real-space bases that efficiently and systematically capture the local chemistry of the system. With further adaptation using new machine-learning techniques and the use of the massively parallel Spectral Quadrature (SQ) electronic structure method, the algorithmic complexity and prefactor associated with DFT calculations involving semilocal as well as hybrid functionals are dramatically reduced. Using petascale computational resources, SPARC-X enables quantum mechanical simulations at length and time scales previously accessible only by empirical approaches, e.g., 1,000,000 atoms for a few picoseconds using semilocal functionals or 1,000 atoms for a few picoseconds using hybrid functionals. Using exascale resources, the sizes and times targeted are two orders of magnitude larger. Such a capability has applications in a wide variety of chemical sciences, including reactive interfaces where large length- and/or long time-scales are needed and traditional force fields fail. This is particularly important in dynamic catalysis, where bond breaking and formation must be understood in detail. We developed, tested, and employed the SPARC-X framework to understand the photocatalytic properties of TiO 2 nanoparticles, revealing finite size effects that cannot be captured with standard model systems or functionals. This integrated development and application strategy ensures that SPARC-X remains a robust, efficient, and scalable software package for quantum simulations on current petascale and emerging exascale computing resources.

97 MATHEMATICS AND COMPUTING↗

Programming the Navier-Stokes computer: An abstract machine model and a visual editor

The Navier-Stokes computer is a parallel computer designed to solve Computational Fluid Dynamics problems. Each processor contains several floating point units which can be configured under program control to implement a vector pipeline with several inputs and outputs. Since the development of an effective compiler for this computer appears to be very difficult, machine level programming seems necessary and support tools for this process have been studied. These support tools are organized into a graphical program editor. A programming process is described by which appropriate computations may be efficiently implemented on the Navier-Stokes computer. The graphical editor would support this programming process, verifying various programmer choices for correctness and deducing values such as pipeline delays and network configurations. Step by step details are provided and demonstrated with two example programs.

Middleton, David↗

Fast Fourier Transform algorithm design and tradeoffs

The Fast Fourier Transform (FFT) is a mainstay of certain numerical techniques for solving fluid dynamics problems. The Connection Machine CM-2 is the target for an investigation into the design of multidimensional Single Instruction Stream/Multiple Data (SIMD) parallel FFT algorithms for high performance. Critical algorithm design issues are discussed, necessary machine performance measurements are identified and made, and the performance of the developed FFT programs are measured. Fast Fourier Transform programs are compared to the currently best Cray-2 FFT program.

Kamin, Ray A., III↗