Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “near memory computing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Impacts of Hybrid Parallelism and Vectorization on the Performance of Newton-Krylov Methods in Computational Aerodynamics

Finding the numerical solution of moderate and high-fidelity aerodynamics problems on modern computer architectures involves, 1) decomposing the domain into smaller regions of nearly equal size, and 2) allocating computational resources for calculations on each domain and communication between domains. Modern computer clusters are composed from hierarchies of processing, memory, and communication resources with varying capabilities and latencies.This paper focuses on the combination of domain decomposition provided by ParMETIS [1]and Newton-Krylov Methods [2–5] for the solution of Computational Aerodynamics problems of interest to NASA. Herein, trade-offs encountered when mapping aerodynamics problems to modern computer architectures are explored through examples and discussions of trade-offs in parallelism from MPI [6], Open MP [7], and vectorization as partition sizes and computational resources are varied. An example of the impact that domain decomposition and MPI+OpenMPresource allocation can have on an adjoint calculation is presented in this abstract. The full paper will include more detailed examples, discussions of difficulties and potential methods to overcome them, and topics identified for future study.

Computational Aerodynamics, Hybrid Parallelism, Ve↗

Study on advanced information processing system

Issues related to the reliability of a redundant system with large main memory are addressed. In particular, the Fault-Tolerant Processor (FTP) for Advanced Launch System (ALS) is used as a basis for our presentation. When the system is free of latent faults, the probability of system crash due to nearly-coincident channel faults is shown to be insignificant even when the outputs of computing channels are infrequently voted on. In particular, using channel error maskers (CEMs) is shown to improve reliability more effectively than increasing the number of channels for applications with long mission times. Even without using a voter, most memory errors can be immediately corrected by CEMs implemented with conventional coding techniques. In addition to their ability to enhance system reliability, CEMs--with a low hardware overhead--can be used to reduce not only the need of memory realignment, but also the time required to realign channel memories in case, albeit rare, such a need arises. Using CEMs, we have developed two schemes, called Scheme 1 and Scheme 2, to solve the memory realignment problem. In both schemes, most errors are corrected by CEMs, and the remaining errors are masked by a voter.

Shin, Kang G.↗

Memory Network For Distributed Data Processors

Universal Memory Network (UMN) is modular, digital data-communication system enabling computers with differing bus architectures to share 32-bit-wide data between locations up to 3 km apart with less than one millisecond of latency. Makes it possible to design sophisticated real-time and near-real-time data-processing systems without data-transfer "bottlenecks". This enterprise network permits transmission of volume of data equivalent to an encyclopedia each second. Facilities benefiting from Universal Memory Network include telemetry stations, simulation facilities, power-plants, and large laboratories or any facility sharing very large volumes of data. Main hub of UMN is reflection center including smaller hubs called Shared Memory Interfaces.

Bolen, David↗

A GPU-based compressible combustion solver for applications exhibiting disparate space and time scales

High-speed chemically active flows pose significant computational challenges due to their disparate space and time scales, with stiff chemistry often dominating simulation time. While modern scientific computing programs achieve exascale performance by leveraging graphics processing units (GPUs), existing GPU-based compressible combustion solvers face critical limitations in memory management, load balancing, and handling the highly localized nature of chemical reactions. To this end, we present a high-performance compressible reacting flow solver built on the AMReX framework and optimized for multi-GPU settings. Here, our approach addresses three GPU performance bottlenecks: memory access patterns through column-major storage optimization, computational workload variability via a bulk-sparse integration strategy for chemical kinetics, and multi-GPU load distribution for adaptive mesh refinement applications. The solver adapts existing matrix-based chemical kinetics formulations to multi-grid contexts. Using representative combustion applications, including 2D and 3D detonations and a 3D jet-in-crossflow configuration, we demonstrate 1.4–5× performance improvements over initial implementations on an in-house cluster of NVIDIA H100 GPUs, and near-ideal weak scaling on the Frontier supercomputer (Oak Ridge Leadership Computing Facility) with up to 1024 AMD Instinct MI250X GPUs. Roofline analysis reveals substantial improvements in arithmetic intensity for both convection (∼ 10 ×) and chemistry (∼ 4 ×) routines, confirming efficient utilization of GPU memory bandwidth and computational resources.

42 ENGINEERING↗

Two Methods for Efficient Solution of the Hitting-Set Problem

A paper addresses much of the same subject matter as that of Fast Algorithms for Model-Based Diagnosis (NPO-30582), which appears elsewhere in this issue of NASA Tech Briefs. However, in the paper, the emphasis is more on the hitting-set problem (also known as the transversal problem), which is well known among experts in combinatorics. The authors primary interest in the hitting-set problem lies in its connection to the diagnosis problem: it is a theorem of model-based diagnosis that in the set-theory representation of the components of a system, the minimal diagnoses of a system are the minimal hitting sets of the system. In the paper, the hitting-set problem (and, hence, the diagnosis problem) is translated from a combinatorial to a computational problem by mapping it onto the Boolean satisfiability and integer- programming problems. The paper goes on to describe developments nearly identical to those summarized in the cited companion NASA Tech Briefs article, including the utilization of Boolean-satisfiability and integer- programming techniques to reduce the computation time and/or memory needed to solve the hitting-set problem.

Vatan, Farrokh↗

Force user's manual: A portable, parallel FORTRAN

The use of Force, a parallel, portable FORTRAN on shared memory parallel computers is described. Force simplifies writing code for parallel computers and, once the parallel code is written, it is easily ported to computers on which Force is installed. Although Force is nearly the same for all computers, specific details are included for the Cray-2, Cray-YMP, Convex 220, Flex/32, Encore, Sequent, Alliant computers on which it is installed.

Jordan, Harry F.↗

Computational fluid dynamics - The coming revolution

The development of aerodynamic theory is traced from the days of Aristotle to the present, with the next stage in computational fluid dynamics dependent on superspeed computers for flow calculations. Additional attention is given to the history of numerical methods inherent in writing computer codes applicable to viscous and inviscid analyses for complex configurations. The advent of the superconducting Josephson junction is noted to place configurational demands on computer design to avoid limitations imposed by the speed of light, and a Japanese projection of a computer capable of several hundred billion operations/sec is mentioned. The NASA Numerical Aerodynamic Simulator is described, showing capabilities of a billion operations/sec with a memory of 240 million words using existing technology. Near-term advances in fluid dynamics are discussed.

Graves, R. A., Jr.↗

Nonvolatile Memory Solution for Near-Term NASA Missions

Nonvolatile memory (NVM) system that could reliably function in extreme environments is one of the most critical components for many spacecrafts being developed for NASA missions to be launched in next four to seven years. NVM supports the computer system in saving and updating critical state data required for a warm restart after power cycling or in case of a power bus failure. It also provides a power independent mass storage capacity for the scientific data gathered by the instruments. In some cases the window for gathering such data is very small and occurs only once in a given mission. Commercially popular and fully developed Flash NVM technology is inappropriate for many reasons such as the limited read write cycles with slower access speeds, radiation intolerance, higher Single Event Upsets (SEU) rates, etc. It is desirable to have an NVM system based upon a robust cell technology making it immune to the SEUs and with sufficient radiation hardness. Availability of such NVM system seems to be still 5 to 10 years in the future. Meanwhile, it is possible to provide an interim hybrid solution by combining the existing rad-hard technologies. Additional information is contained in the original extended abstract.

Patel, J. U.↗

Inducer analysis/pump model development

Current design of high performance turbopumps for rocket engines requires effective and robust analytical tools to provide design information in a productive manner. The main goal of this study was to develop a robust and effective computational fluid dynamics (CFD) pump model for general turbopump design and analysis applications. A finite difference Navier-Stokes flow solver, FDNS, which includes an extended k-epsilon turbulence model and appropriate moving zonal interface boundary conditions, was developed to analyze turbulent flows in turbomachinery devices. In the present study, three key components of the turbopump, the inducer, impeller, and diffuser, were investigated by the proposed pump model, and the numerical results were benchmarked by the experimental data provided by Rocketdyne. For the numerical calculation of inducer flows with tip clearance, the turbulence model and grid spacing are very important. Meanwhile, the development of the cross-stream secondary flow, generated by curved blade passage and the flow through tip leakage, has a strong effect on the inducer flow. Hence, the prediction of the inducer performance critically depends on whether the numerical scheme of the pump model can simulate the secondary flow pattern accurately or not. The impeller and diffuser, however, are dominated by pressure-driven flows such that the effects of turbulence model and grid spacing (except near leading and trailing edges of blades) are less sensitive. The present CFD pump model has been proved to be an efficient and robust analytical tool for pump design due to its very compact numerical structure (requiring small memory), fast turnaround computing time, and versatility for different geometries.

Cheng, Gary C.↗

Simulation of 3-D shear flows around a nozzle-afterbody at high speeds

3D, compressible, unsteady, Reynolds-averaged Navier-Stokes equations are presently solved by a finite-volume and alternating-direction-implicit method in order to simulate supersonic and hypersonic turbulent shear flows. The effect of turbulence is incorporated via a modified Baldwin-Lomax eddy viscosity model which reflects the influence of high-speed compressibility, multiple walls, near-wall vortices, and turbulent memory effects, as well as local equilibrium effects. Attention is given to the simulation of the flow around the nozzle-afterbody of a generic, scramjet-propelled hypersonic vehicle; computed pressure distributions are consonant with experimental surface and off-surface flow surveys.

Baysal, Oktay↗

Computational discovery of two-dimensional rare-earth iodides: promising ferrovalley materials for valleytronics

Two-dimensional Ferrovalley materials with intrinsic valley polarization are rare but highly promising for valley-based nonvolatile random access memory and valley filter devices. These ferromagnetic materials exhibit valleys at or near the Fermi level with intrinsic magnetism. The strong coupling between magnetism and spin–orbit coupling induces intrinsic valley polarization. Using Kinetically Limited Minimization, an unconstrained crystal structure prediction algorithm, and prototype sampling based on first-principles calculations, we have discovered new Ferrovalley materials, rare-earth iodides RI 2 , where R is a rare-earth element belonging to Sc, Y, or La-Lu, and I is Iodine. The rare-earth iodides are layered and demonstrate either 2H, 1T, or 1T d phase as the ground state in bulk, analogous to transition metal dichalcogenides (TMDCs). The calculated exfoliation energy of monolayers (MLs) is comparable to that of graphene and TMDCs, suggesting possible experimental synthesis. The MLs in the 2H phase exhibit ferromagnetism due to unpaired electrons in d and f orbitals. Throughout the rare-earth series, d bands have valley polarization at K and $\overline{K}$ points in the Brillouin zone in the vicinity of the Fermi level. Large intrinsic valley polarization in the range of 15–143 meV without external stimuli is observed in these Ferrovalley materials, which can be enhanced further by applying an in-plane bi-axial strain. These valleys can selectively be probed and manipulated for information storage and processing, potentially offering superior performance beyond conventional electronics and spintronics. Here we further show that the 2H ferromagnetic phase of RI 2 MLs possesses non-zero Berry curvature and exhibits anomalous valley Hall effect with considerable anomalous Hall conductivity. Our work will incite exploratory synthesis of the predicted Ferrovalley materials and their application in valleytronics and beyond.

2D materials↗

Determining the operating characteristics of an ultraviolet interferometric spectrometer

A prototype interferometric spectrometer system is being built by NASA to explore the potential of the technique for applications involving the visible and near ultraviolet portions of the electromagnetic spectrum. The system is limited only by the frequency bandpass of the optical components used in the system, the quality of the optical components, and ultimately by the memory capacity of the computer; tradeoffs between the wavenumber resolution of the produced spectrum, the bandpass limits of the optics, and the number of samples obtained from the interferogram must be delineated explicitly. The prototype Ultraviolet Interferometric Spectrometer (UVIS) instrument is expected to be configured several different ways to ascertain its suitability for various applications. To exploit its inherent flexibility, this reference document describes these parameter tradeoffs.

Parsons, C. L.↗

The NEAT Camera Project

The NEAT (Near Earth Asteroid Tracking) camera system consists of a camera head with a 6.3 cm square 4096 x 4096 pixel CCD, fast electronics, and a Sun Sparc 20 data and control computer with dual CPUs, 256 Mbytes of memory, and 36 Gbytes of hard disk. The system was designed for optimum use with an Air Force GEODSS (Ground-based Electro-Optical Deep Space Surveillance) telescope. The GEODSS telescopes have 1 m f/2.15 objectives of the Ritchey-Chretian type, designed originally for satellite tracking. Installation of NEAT began July 25 at the Air Force Facility on Haleakala, a 3000 m peak on Maui in Hawaii.

NEAT Camera↗

Simulation of cosmic-ray induced soft errors and latchup in integrated-circuit computer memories

Soft errors have been induced in solid-state static RAM's by iron nuclei from the Lawrence Berkeley Laboratory (LBL) Bevalac, in experiments designed to prove the ability of iron-group cosmic rays to generate such errors. Subsequently, various delidded device types were tested in beams of argon and krypton ions from the LBL 88-inch Cyclotron, at energies near 2 MeV/nucleon. The latter tests showed that some devices are essentially immune to bit error while others are quite susceptible. Good agreement was obtained with model predictions in cases where the latter exist. Latchup, whose cause is attributed to individual heavy ions, was also observed in some device types.

Kolasinski, W. A.↗

Measurement of Electromagnetic Properties of Lightning with 10 Nanosecond Resolution

Electromagnetic data recorded from lightning strikes are presented. The data analysis reveals general characteristics of fast electromagnetic fields measured at the ground including rise times, amplitudes, and time patterns. A look at the electromagnetic structure of lightning shows that the shortest rise times in the vicinity of 30 ns are associated with leader leader streamers. Lightning location is based on electromagnetic field characteristics and is compared to a nearby sky camera. The fields from both leaders and return strokes were measured and are discussed. The data were obtained during 1978 and 1979 from lightning strikes occuring within 5 kilometers of an underground metal instrumentation room located on South Baldy peak near Langmuir Laboratory, New Mexico. The computer controlled instrumentation consisted of sensors previously used for measuring the nuclear electromagnetic pulse (EMP) and analog-digital recorders with 10 ns sampling, 256 levels of resolution, and 2 kilobytes of internal memory.

Baum, C. E.↗

Viscous simulation method for unsteady flows past multicomponent configurations

The present algorithm for the numerical simulation of flows about complex configurations (whose multiple and nonsimilar components have arbitraty geometries) employs a hybridization of the domain decomposition techniques for grid generation as well as to reduce computer-memory requirements. A fully vectorized, finite-volume, upwind-biased, approximately factored multigrid method is used to solve 3D Reynolds-averaged unsteady and compressible Navier-Stokes equations simulating supersonic flows past an ogive-nose-cylinder near or within a cavity. The time-averaged surface pressures obtained compare favorably with the wind tunnel data.

Fouladi, Kamran↗

Using 100G Network Technology in Support of Petascale Science

NASA in collaboration with a number of partners conducted a set of individual experiments and demonstrations during SC 10 that collectively were titled "Using 100G Network Technology in Support of Petascale Science". The partners included the iCAIR, Internet2, LAC, MAX, National LambdaRail (NLR), NOAA and SCinet Research Sandbox (SRS) as well as the vendors Ciena, Cisco, ColorChip, cPacket, Extreme Networks, Fusion-io, HP and Panduit who most generously allowed some of their leading edge 40G/100G optical transport, Ethernet switch and Internet Protocol router equipment and file server technologies to be involved. The experiments and demonstrations featured different vendor-provided 40G/100G network technology solutions for full-duplex 40G and 100G LAN data flows across SRS-deployed single-node fiber-pairs among the Exhibit Booths of NASA, the National Center for Data lining, NOAA and the SCinet Network Operations Center, as well as between the NASA Exhibit Booth in New Orleans and the Starlight Communications Exchange facility in Chicago across special SC 10- only 80- and 100-Gbps wide area network links provisioned respectively by the NLR and Internet2, then on to GSFC across a 40-Gbps link. provisioned by the Mid-Atlantic Crossroads. The networks and vendor equipment were load-stressed by sets of NASA/GSFC High End Computer Network Team-built, relatively inexpensive, net-test-workstations that are capable of demonstrating greater than 100Gbps uni-directional nuttcp-enabled memory-to-memory data transfers, greater than 80-Gbps aggregate--bidirectional memory-to-memory data transfers, and near 40-Gbps uni-directional disk-to-disk file copying. This paper will summarize the background context, key accomplishments and some significances of these experiments and demonstrations.

Gary, James P.↗

On the Development and Application of High Data Rate Architecture (HiDRA) in Future Space Networks

Historically, space missions have been severely constrained by their ability to downlink the data they have collected. These constraints are a result of relatively low link rates on the spacecraft as well as limitations on the time during which data can be sent. As part of a coherent strategy to address existing limitations and get more data to the ground more quickly, the Space Communications and Navigation (SCaN) program has been developing an architecture for a future solar system Internet. The High Data Rate Architecture (HiDRA) project is designed to fit into such a future SCaN network. HiDRA's goal is to describe a general packet-based networking capability which can be used to provide assets with efficient networking capabilities while simultaneously reducing the capital costs and operational costs of developing and flying future space systems.Along these lines, this paper begins by reviewing various characteristics of modern satellite design as well as relevant characteristics of emerging technologies (such as free-space optical links capable of working at 100+ Gbps). Next, the paper describes HiDRA's design, and how the system is able to both integrate and support the operation of not only today's high-rate systems, but also the high-rate systems likely to be found in the future. This section also explores both existing and future networking technologies, such as Delay Tolerant Networking (DTN) protocol (RFC4838 citeRFC:1, RFC5050citeRFC:2), and explains how HiDRA supports them. Additionally, this section explores how HiDRA is used for scheduling data movement through both proactive and reactive link management. After this, the paper moves on to explore a reference implementation of HiDRA. This implementation is currently being realized based on a Field Programmable Gate Array (FPGA) memory and interface controller that is itself controlled by a local computer running DTN software. Next, this paper explores HiDRA's natural evolution, which includes an integration path for software-defined networking (SDN) switches. This section also describes considerations for both near-Earth and deep-space instantiations of HiDRA, describing how differences in latencies between the environments will necessarily influence how the system is configured and the networks operate. Finally, this paper describes future work. This section includes a description of a potential ISS implementation which will allow rapid advancement through the technology readiness levels (TRL). This section also explores work being done to support HiDRA's successful implementation and operation in a heterogeneous network: such a network could include communications equipment spanning many vintages and capabilities, and one significant aspect of HiDRA's future development involves balancing compatibility with capability.

DTN↗