Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “concurrent computing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 325 records · Page 18

AI-enhanced Codesign for Next-Generation Neuromorphic Circuits and Systems

This report details work that was completed to address the Fiscal Year 2022 Advanced Science and Technology (AS&T) Laboratory Directed Research and Development (LDRD) call for “AI-enhanced Co-Design of Next Generation Microelectronics.” This project required concurrent contributions from the fields of 1) materials science, 2) devices and circuits, 3) physics of computing, and 4) algorithms and system architectures. During this project, we developed AI-enhanced circuit design methods that relied on reinforcement learning and evolutionary algorithms. The AI-enhanced design methods were tested on neuromorphic circuit design problems that have real-world applications related to Sandia’s mission needs. The developed methods enable the design of circuits, including circuits that are built from emerging devices, and they were also extended to enable novel device discovery. We expect that these AI-enhanced design methods will accelerate progress towards developing next-generation, high-performance neuromorphic computing systems.

42 ENGINEERING↗

Minimizing inner product data dependencies in conjugate gradient iteration

The amount of concurrency available in conjugate gradient iteration is limited by the summations required in the inner product computations. The inner product of two vectors of length N requires time c log(N), if N or more processors are available. This paper describes an algebraic restructuring of the conjugate gradient algorithm which minimizes data dependencies due to inner product calculations. After an initial start up, the new algorithm can perform a conjugate gradient iteration in time c*log(log(N)).

Vanrosendale, J.↗

Systolic multipliers for finite fields GF(2 exp m)

Two systolic architectures are developed for performing the product-sum computation AB + C in the finite field GF(2 exp m) of 2 exp m elements, where A, B, and C are arbitrary elements of GF(2 exp m). The first multiplier is a serial-in, serial-out one-dimensional systolic array, while the second multiplier is a parallel-in, parallel-out two-dimensional systolic array. The first multiplier requires a smaller number of basic cells than the second multiplier. The second multiplier needs less average time per computation than the first multiplier, if a number of computations are performed consecutively. To perform single computations both multipliers require the same computational time. In both cases the architectures are simple and regular and possess the properties of concurrency and modularity. As a consequence, they are well suited for use in VLSI systems.

Yeh, C.-S.↗

Concurrent algorithms for transient FE analysis

Information on concurrent algorithms for transient finite element analysis is given in viewgraph form. Information is given on concurrent dynamic algorithms, interprocessor communication, the performance of the BAR problem on the 32 Processor Hypercube, computational efficiency and accuracy analysis.

Ortiz, M.↗

Energy Distribution among Reaction Products. III: The Method of Measured Relaxation Applied to H + Cl2

The method of measured relaxation is described for the determination of initial vibrational energy distribution in the products of exothermic reaction. Hydrogen atoms coming from an orifice were diffused into flowing chlorine gas. Measurements were made of the resultant ir chemiluminescence at successive points along the line of flow. The concurrent processes of reaction, diffusion, flow, radiation, and deactivation were analyzed in some detail on a computer. A variety of relaxation models were used in an attempt to place limits on k(nu prime), the rate constant for reaction to form HCl in specified vibrational energy levels: H+Cl2 yields (sup K(nu prime) HCl(sub nu prime) + Cl. The set of k(?) obtained from this work is in satisfactory agreement with those obtained by another experimental method (the method of arrested relaxation described in Parts IV and V of the present series.

Pacey, P. D.↗

Grain Properties of Comet C/1995 O1 (Hale-Bopp) Deduced Through Computational Techniques

We present the computational analysis of the 7.6 - 13.2 micrometer infrared (IR) spectrophotometry (R approximately equal to 120) of comet C/1995 O1 (Hale-Bopp) in conjunction with concurrent observations which extend the spectral energy distribution from the near-infrared to far-infrared wavelengths. The observations include temporal epochs pre-perihelion, (1996 October UT and 1997 February UT), near perihelion (1997 April UT), and postperihelion (1997 June UT). Through the computational modeling of small amorphous carbon, and crystalline and amorphous silicate grains in Hale-Bopp's coma, we find that as the comet approached perihelion, the grain size distribution (the Hanner modified power law) steepened (N = 3.4 pre-perihelion, to N = 3.7 near and post-perihelion) along with an increase in the fractal porosity of larger (greater than 1 micrometer) grains. The peak of the grain size distribution remained constant (ap = 0.2 micrometer) at each epoch. We attribute the emergence of the 9.3 micrometer peak near perihelion to crystalline orthopyroxeno grains released from inside the nucleus. Crystalline silicates (olivine and orthopyroxene) make up about 30% (by mass) of the submicron sized (less than 1 micrometer) dust grains in Hale-Bopp's coma during each epoch.

Harker, David E.↗

ConnectIt: a framework for static and incremental parallel graph connectivity algorithms

Connected components is a fundamental kernel in graph applications. The fastest existing multicore algorithms for solving graph connectivity are based on some form of edge sampling and/or linking and compressing trees. However, many combinations of these design choices have been left unexplored. In this paper, we design the ConnectIt framework, which provides different sampling strategies as well as various tree linking and compression schemes. ConnectIt enables us to obtain several hundred new variants of connectivity algorithms, most of which extend to computing spanning forest. In addition to static graphs, we also extend ConnectIt to support mixes of insertions and connectivity queries in the concurrent setting. We present an experimental evaluation of ConnectIt on a 72-core machine, which we believe is the most comprehensive evaluation of parallel connectivity algorithms to date. Compared to a collection of state-of-the-art static multicore algorithms, we obtain an average speedup of 12.4x (2.36x average speedup over the fastest existing implementation for each graph). Using ConnectIt, we are able to compute connectivity on the largest publicly-available graph (with over 3.5 billion vertices and 128 billion edges) in under 10 seconds using a 72-core machine, providing a 3.1x speedup over the fastest existing connectivity result for this graph, in any computational setting. For our incremental algorithms, we show that our algorithms can ingest graph updates at up to several billion edges per second. To guide the user in selecting the best variants in ConnectIt for different situations, we provide a detailed analysis of the different strategies. Finally, we show how the techniques in ConnectIt can be used to speed up two important graph applications: approximate minimum spanning forest and SCAN clustering.

Computer Science↗

Reengineering the Project Design Process

In response to NASA's goal of working faster, better and cheaper, JPL has developed extensive plans to minimize cost, maximize customer and employee satisfaction, and implement small- and moderate-size missions. These plans include improved management structures and processes, enhanced technical design processes, the incorporation of new technology, and the development of more economical space- and ground-system designs. The Laboratory's new Flight Projects Implementation Office has been chartered to oversee these innovations and the reengineering of JPL's project design process, including establishment of the Project Design Center and the Flight System Testbed. Reengineering at JPL implies a cultural change whereby the character of its design process will change from sequential to concurrent and from hierarchical to parallel. The Project Design Center will support missions offering high science return, design to cost, demonstrations of new technology, and rapid development. Its computer-supported environment will foster high-fidelity project life-cycle development and cost estimating.

Jet Propulsion Laboratory JPL project design concu↗

FLEXAN (version 2.0) user's guide

The FLEXAN (Flexible Animation) computer program, Version 2.0 is described. FLEXAN animates 3-D wireframe structural dynamics on the Evans and Sutherland PS300 graphics workstation with a VAX/VMS host computer. Animation options include: unconstrained vibrational modes, mode time histories (multiple modes), delta time histories (modal and/or nonmodal deformations), color time histories (elements of the structure change colors through time), and rotational time histories (parts of the structure rotate through time). Concurrent color, mode, delta, and rotation, time history animations are supported. FLEXAN does not model structures or calculate the dynamics of structures; it only animates data from other computer programs. FLEXAN was developed to aid in the study of the structural dynamics of spacecraft.

Stallcup, Scott S.↗

Production Level CFD Code Acceleration for Hybrid Many-Core Architectures

In this work, a novel graphics processing unit (GPU) distributed sharing model for hybrid many-core architectures is introduced and employed in the acceleration of a production-level computational fluid dynamics (CFD) code. The latest generation graphics hardware allows multiple processor cores to simultaneously share a single GPU through concurrent kernel execution. This feature has allowed the NASA FUN3D code to be accelerated in parallel with up to four processor cores sharing a single GPU. For codes to scale and fully use resources on these and the next generation machines, codes will need to employ some type of GPU sharing model, as presented in this work. Findings include the effects of GPU sharing on overall performance. A discussion of the inherent challenges that parallel unstructured CFD codes face in accelerator-based computing environments is included, with considerations for future generation architectures. This work was completed by the author in August 2010, and reflects the analysis and results of the time.

Duffy, Austen C.↗

Method and apparatus for real time, in situ sensing and characterization of roughness, geometrical shapes, geometrical structures, composition, defects, and temperature in three-dimensional manufacturing systems

Methods and apparatuses for manufacturing are disclosed, including (a) providing an apparatus having: a laser; scanner; powder injection system; powder spreading system; dichroic filter; imager-and-processor; and computer; (b) programming the computer with specifications of a sample; (c) using the computer to set initial parameters based on the sample specifications; (d) adjusting a stage to position the sample; (e) focusing and scanning electromagnetic radiation onto the sample while powder is concurrently injected onto the sample in order to deposit a layer; (f) capturing two-dimensional images of the sample and probing the sample to determine whether the deposited layer was manufactured per the specifications; (g) use the computer to adjust the three-dimensional manufacturing parameters based on the determination made in step (f) prior to additively manufacturing a subsequent layer or making repairs; and (h) repeating steps (d), (e), (f), and (g) until the manufacture is complete. Other embodiments are described and claimed.

Liu, Jian↗

Analysis of Threading Libraries for High Performance Computing

With the appearance of multi-/many core machines, applications and runtime systems have evolved in order to exploit the new on-node concurrency brought by new software paradigms. POSIX threads (Pthreads) was widely-adopted for that purpose and it remains as the most used threading solution in current hardware. Lightweight thread (LWT) libraries emerged as an alternative offering lighter mechanisms to tackle the massive concurrency of current hardware. In this article, we analyze in detail the most representative threading libraries including Pthread- and LWT-based solutions. In addition, to examine the suitability of LWTs for different use cases, we develop a set of microbenchmarks consisting of OpenMP patterns commonly found in current parallel codes, and we compare the results using threading libraries and OpenMP implementations. Moreover, we study the semantics offered by threading libraries in order to expose the similarities among different LWT application programming interfaces and their advantages over Pthreads. This article exposes that LWT libraries outperform solutions based on operating system threads when tasks and nested parallelism are required.

GLT↗

A New Concurrent Multiscale Methodology for Coupling Molecular Dynamics and Finite Element Analyses

The coupling of molecular dynamics (MD) simulations with finite element methods (FEM) yields computationally efficient models that link fundamental material processes at the atomistic level with continuum field responses at higher length scales. The theoretical challenge involves developing a seamless connection along an interface between two inherently different simulation frameworks. Various specialized methods have been developed to solve particular classes of problems. Many of these methods link the kinematics of individual MD atoms with FEM nodes at their common interface, necessarily requiring that the finite element mesh be refined to atomic resolution. Some of these coupling approaches also require simulations to be carried out at 0 K and restrict modeling to two-dimensional material domains due to difficulties in simulating full three-dimensional material processes. In the present work, a new approach to MD-FEM coupling is developed based on a restatement of the standard boundary value problem used to define a coupled domain. The method replaces a direct linkage of individual MD atoms and finite element (FE) nodes with a statistical averaging of atomistic displacements in local atomic volumes associated with each FE node in an interface region. The FEM and MD computational systems are effectively independent and communicate only through an iterative update of their boundary conditions. With the use of statistical averages of the atomistic quantities to couple the two computational schemes, the developed approach is referred to as an embedded statistical coupling method (ESCM). ESCM provides an enhanced coupling methodology that is inherently applicable to three-dimensional domains, avoids discretization of the continuum model to atomic scale resolution, and permits finite temperature states to be applied.

Yamakov, Vesselin↗

A parallel householder tridiagonalization stratagem using scattered row decomposition

Householder's method for tridiagonalizing a real symmetric matrix, a major step in evaluating eigenvalues of the matrix, is modified into a parallel algorithm for a concurrent machine of message passing type. Each processor of the concurrent machine has its own CPU, communications control and local memory. Messages are passed through connections between processors. Although the basic algorithm is inherently serial, the computations can be spread over all processors by scattering different rows of the matrix into processors, hence the term 'Scattered Row Decomposition'. The steps in the serial and the parallel algorithms are identified. Expressions for efficiency and speedup are given in terms of problem and machine parameters. For a concurrent machine of ring type interconnection, a selected representative problem of large order exhibits efficiency approaching 66 per cent.

Chang, H. Y.↗

A distributed microprocessor system for spacecraft control and data handling

The specific requirements for spacecraft computing systems are considered. These requirements are partly related to the constraints of limited resources of power, weight, and volume. Another important factor is the requirement of extremely high reliability. These reliability requirements have led to introduction of automated redundancy techniques on board the spacecraft. The various redundant computers check each other and provide recovery procedures when a computer is found to have failed. Past and future capabilities are considered along with distributed processing requirements. System considerations are discussed, taking into account suboptimum computer throughput, sensitivity to software modifications, hierarchic timing, I/O granularity, restricted communications, synchronous functions, hierarchic control, and concurrent error detection. A description is presented of the Unified Data System (UDS), which consists of a set of standard microcomputers connected by several buses. Attention is also given to synchronization and timing, the executive control structure, the programming language, and the executive program.

Rennels, D. A.↗

Three dimensional unstructured grids for the solution of the Euler equations

The advancing front technique is being used to develop a code to generate grids around complex 3-D configurations for use in computing the invisid flow solutions by the Euler equations. By the advancing front technique points are introduced concurrently with the connectivity information so that a separate library is not required. The generation of a 3-D grid is accomplished in several steps. First the boundaries of the domain to be gridded must be described by two-, three- or four-sided surface patches. Next, a background mesh is required to control the grid spacing and stretching throughout the domain. This coarse tetrahedral grid is not required to conform to any of the boundaries. Next, each of the patches is mapped to 2-D, triangulated by the advancing front technique and mapped back to 3-D. These triangles form the initial front for the generation of the final tetrahedral mesh.

Gumbert, Clyde↗

Extend an innovative HPC-Compatible Multiple Temporal-spatial Resolution Concurrent Finite Element Modeling Approach to Guide Laser Powder Bed Fusion Additive

Laser power bed fusing (PBF) additive manufacturing is a key enabling technology to manufacture highly complex and integrated automotive structures. However, the geometric complexity of PBF-AM technique also leads to highly non-uniform heating and cooling rate in the manufactured part, which may cause flaw formation and produce excessive and nonuniform residual stresses, which increase quality uncertainties and manufacture issues, leading to increases in cost and energy consumption in the form of rejected parts. In this research project, we developed an innovative Multi-Spatial-Temporal-Resolution Finite Element (MUST-FE) method and completed the corresponding high performance computation (HPC) platform-based in-house code, which enables high accuracy prediction of temperature and residual stress fields for component-scale PBF-AM manufacture in efficient computation time. The MUST-FE model is calibrated and validated with a “2D pad” AlSi10Mg experiments by matching the melt pool shape and dimension, and with a “XY-cross” AlSi10Mg experiment by matching the thermal distortion and residual stress. The innovative multi-resolution and concurrent modeling approach adopted in this code ensures accuracy and computational efficiency, which will enable energy-efficient and high-yield, low-cost manufacturing of optimized, qualifiable automotive structures and contribute towards reaching technical targets outlined in AMO’s Program Plan to develop additive manufacturing systems that deliver consistently reliable parts with predictable properties.

36 MATERIALS SCIENCE↗