Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Memory systems”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11

A magnetic bubble domain flight recorder

A feasibility model of an all-electronic bubble memory system has been constructed. It uses a small 60k bit bubble recorder consisting of 6 chips of 10k bits each mounted in three separate packages operating as a FIFO at a 150 KHz bubble data rate. In addition to serving as a direct tape recorder replacement, the bubble recorder can be programmed for random access to each individual chip for ranom block access operation or for self-checking or by-passing any malfunctioning memory chip. Read and write operations can be performed asynchronously from very low frequency up to basic recording field frequency. A large 50M bit prototype is planned.

Chen, T. T.↗

The utilization of bubble memories in defense systems

The paper considers two examples of bubble memory application: the NASA solid state data recorder for spacecraft, and the POS/8 800,000 bit recorder. A matrix chart of applications delineated by capacity and chip organization, which is primarily reflected in access time, is then presented.

Mavity, W. C.↗

Memory and rejuvenation in glassy systems

Here, the memory effect in a single crystal spin glass (Cu 0.92 Mn 0.08 ) has been measured using 1 Hz ac susceptibility measurements over a reduced temperature range of 0.4 - 0.7 T g and a model of the memory effect has been developed. A double-waiting-time protocol is carried out where the spin glass is first allowed to age at a temperature below T g , T w$_{1}$ , followed by a second aging 4 K lower, T w$_{2}$ . The 4 K separation is sufficient to ensure rejuvenation has occurred. The model is based on calculating overlaps between the growth of the correlation lengths at the two temperatures. It accounts for the absolute magnitude of the memory effect as a function of both waiting times and temperatures. The data can be explained by the memory loss being a function of the relative change in the correlated volume at the first waiting temperature due to growth in the correlations at the second waiting temperature.

36 MATERIALS SCIENCE↗

Validation of CFD/Heat Transfer Software for Turbine Blade Analysis

I am an intern in the Turbine Branch of the Turbomachinery and Propulsion Systems Division. The division is primarily concerned with experimental and computational methods of calculating heat transfer effects of turbine blades during operation in jet engines and land-based power systems. These include modeling flow in internal cooling passages and film cooling, as well as calculating heat flux and peak temperatures to ensure safe and efficient operation. The branch is research-oriented, emphasizing the development of tools that may be used by gas turbine designers in industry. The branch has been developing a computational fluid dynamics (CFD) and heat transfer code called GlennHT to achieve the computational end of this analysis. The code was originally written in FORTRAN 77 and run on Silicon Graphics machines. However the code has been rewritten and compiled in FORTRAN 90 to take advantage of more modem computer memory systems. In addition the branch has made a switch in system architectures from SGI's to Linux PC's. The newly modified code therefore needs to be tested and validated. This is the primary goal of my internship. To validate the GlennHT code, it must be run using benchmark fluid mechanics and heat transfer test cases, for which there are either analytical solutions or widely accepted experimental data. From the solutions generated by the code, comparisons can be made to the correct solutions to establish the accuracy of the code. To design and create these test cases, there are many steps and programs that must be used. Before a test case can be run, pre-processing steps must be accomplished. These include generating a grid to describe the geometry, using a software package called GridPro. Also various files required by the GlennHT code must be created including a boundary condition file, a file for multi-processor computing, and a file to describe problem and algorithm parameters. A good deal of this internship will be to become familiar with these programs and the structure of the GlennHT code. Additional information is included in the original extended abstract.

Kiefer, Walter D.↗

A design study of a photorefractive page composer

A laboratory demonstration and preliminary system analysis of a page composer designed to have the dual advantages of low optical loss and small size, were reported. The current page composer is optically addressed and functions by virtue of optically induced refractive index changes in the active material. Laboratory demonstrations of the device were successfully performed using 10 x 10 bit and 128 x 128 bit data arrays. It was established that the only significant obstacle to the construction of a brass-board model working at megabit data rates is the lack of sensitivity of the photorefractive materials which were considered during the course of this study. Possible materials for future consideration are the photoplastics. While they have more than the required sensitivity, their stability and suitability for double exposure holography was not investigated. If a sufficiently sensitive material is found, then the photorefractive page composer could be built to perform in a highly efficient fashion which would result in a overall reduction of the size of the memory system and an easing of the requirements upon the sensitivity of the holographic recording material.

Source record↗

Computing material volume fractions on a superimposed mesh as applied to Monte Carlo particle transport simulations

Here, we present a newly implemented ray tracing algorithm in OpenMC for efficiently computing material volume fractions on superimposed meshes in complex geometries. By firing rays along each coordinate direction through the geometry, the approach accumulates track-length data in each mesh element, thereby determining the fractional composition of each material. Scaling studies on three different models—a random tetrahedra configuration, the Frascati Neutron Generator ITER dose rate benchmark, and a stellarator design—show excellent parallel performance, with nearly linear speedup on modern multi-threaded and distributed-memory systems. An analysis of the residual error relative to high-resolution reference solutions demonstrated that under optimal conditions it decreases as 1/R, where R is the number of rays fired, making it straightforward to achieve user-prescribed accuracy. This new functionality enables practical, mesh-based approaches for detailed nuclear analyses in production Monte Carlo workflows without resorting to expensive, fully conformal or unstructured meshing.

Monte Carlo↗

CG-Kit: Code Generation Toolkit for performant and maintainable variants of source code applied to Flash-X hydrodynamics simulations

CG-Kit is a new Code Generation tool-Kit that we have developed as a part of the solution for portability and maintainability for multiphysics computing applications. The development of CG-Kit is rooted in the urgent need created by the shifting landscape of high-performance computing platforms and the algorithmic complexities of a particular large-scale multiphysics application: Flash-X. To efficiently use computing resources on a heterogeneous node, an application must have a map of computation to resources and a mechanism to move the data and computation to the resources according to the map. Most existing performance portability solutions are focussed on abstracting the expression of computations so that a unified source code can be specialized to run on different resources. However, such an approach is insufficient for a code like Flash-X, which has a multitude of code components that can be assembled in various permutations and combinations to form different instances of applications. Similar challenges apply to any code that has composability, where a single specified way of apportioning work among devices may not be optimal. Additionally, use cases arise where the optimal control flow of computation may differ for different devices while the underlying numerics remain identical. This combination leads to unique challenges including handling an existing large code base in Fortran and/or C/C++, subdivision of code into a great variety of units supporting a wide range of physics and numerical methods, different parallelization techniques for distributed and shared memory systems and accelerator devices, and heterogeneity of computing platforms requiring coexisting variants of parallel algorithms. All of these challenges demand that scientific software developers apply existing knowledge about domain applications, algorithms, and computing platforms to determine custom abstractions and granularity for code generation. There is a critical lack of tools to tackle those problems. CG-Kit is designed to fill this gap by providing a user with the ability to express their desired control flow and computation-to-resource map in the form a pseudocode-like recipe. It consists of standalone tools that can be combined into highly specific and, we argue, highly effective portability and maintainability toolchains. Here we present the design of our new tools: parametrized source trees, control flow graphs, and recipes. The tools are implemented in Python. They are agnostic to the programming language of the source code targeted for code generation. In conclusion, we demonstrate the capabilities of the toolkit with two examples, first, multithreaded variants of the basic AXPY operation, and second, variants of parallel algorithms within a hydrodynamics solver, called Spark, from Flash-X that operates on block-structured adaptive meshes.

Algorithmic portability↗

The PetscSF Scalable Communication Layer

PetscSF, the communication component of the Portable, Extensible Toolkit for Scientific Computation (PETSc), is designed to provide PETSc's communication infrastructure suitable for exascale computers that utilize GPUs and other accelerators. PetscSF provides a simple application programming interface (API) for managing common communication patterns in scientific computations by using a star-forest graph representation. PetscSF supports several implementations based on MPI and NVSHMEM, whose selection is based on the characteristics of the application or the target architecture. An efficient and portable model for network and intra-node communication is essential for implementing large-scale applications. The Message Passing Interface, which has been the de facto standard for distributed memory systems, has developed into a large complex API that does not yet provide high performance on the emerging heterogeneous CPU-GPU-based exascale systems. Here, we discuss the design of PetscSF, how it can overcome some difficulties of working directly with MPI on GPUs, and we demonstrate its performance, scalability, and novel features.

97 MATHEMATICS AND COMPUTING↗

Reinforcement Learning for Load-balanced Parallel Particle Tracing

We explore an online reinforcement learning (RL) paradigm to dynamically optimize parallel particle tracing performance in distributed-memory systems. Our method combines three novel components: (1) a work donation algorithm, (2) a high-order workload estimation model, and (3) a communication cost model. First, we design an RL-based work donation algorithm. Our algorithm monitors workloads of processes and creates RL agents to donate data blocks and particles from high-workload processes to low-workload processes to minimize program execution time. The agents learn the donation strategy on the fly based on reward and cost functions designed to consider processes' workload changes and data transfer costs of donation actions. Second, we propose a workload estimation model, helping RL agents estimate the workload distribution of processes in future computations. Third, we design a communication cost model that considers both block and particle data exchange costs, helping RL agents make effective decisions with minimized communication costs. We demonstrate that our algorithm adapts to different flow behaviors in large-scale fluid dynamics, ocean, and weather simulation data. Our algorithm improves parallel particle tracing performance in terms of parallel efficiency, load balance, and costs of I/O and communication for evaluations with up to 16,384 processors.

Distributed and parallel particle tracing↗

Synchronization for CXL Based Memory

Compute Express Link (CXL) is an important emerging standard for disaggregated memory. While this standard provisions coherency across numerous hosts and devices, implementing hardware support for type three devices is challenging. In this work, we look at the overhead of software synchronization and using software-based coherency. Moreover, we discuss the limits of software-based coherency in fully expressing modern synchronization techniques for a CXL-based disaggregate memory system. We demonstrate our approach using a CXL hardware prototype and running a version of the famous Peterson Lock (enhanced to run with more than two threads). We analyze its performance and share how more advanced synchronization techniques might interact with software-based coherence CXL hardware and program execution models.

High Performance Computing (HPC)↗

SpecSims: A Scalable Speculative Tree-based Simulation Cloning Framework for Finite Memory Machines

Simulation cloning is a technique in which cloned simulations whose state spaces differ partially from their parent simulation due to intervening events are spawned at runtime and concurrently advanced. It is a powerful method to carry out what-if analysis by speculatively exploring and evaluating the impact of various permutations of intervening cascade of events. Due to the exponential growth in the number of possible clones even for a small number of distinct intervening events, the practical efficacy of the approach is often severely limited by the maximum available memory of the computing host. In this paper, we introduce a novel speculative simulation cloning framework that executes a simulation cloning campaign capable of efficiently exploring an exponentially large space of clone simulations created by permutation of intervening events under a finite memory constraint. We provide a theoretical analysis of the runtime characteristics of our proposed approach and highlight its novel advantages such as memory-aware and as-long-as-needed execution. Furthermore, in support of our analytical findings and to demonstrate its practical feasibility, we implement a prototype of the cloning framework on a shared memory system and report its performance characteristics in the context of a heat diffusion simulation, and a power grid simulation subject to cascading disruptions from geomagnetic disturbances.

Simulation framework↗

Variable Precision Computing (VPC) (Final Report)

The overall goal of the Variable Precision Computing project is exploring and identifying ways to reduce internodal communication cost in production physics simulation experiments running on a distributed-memory systems by compressing data, either in a lossy or lossless way, thus resulting in a significant computational performance increase. Among the several approaches proposed, the one discussed here has been proposed, designed and implemented by me. It concerns the application of information theory-derived metrics to classical molecular dynamics and computational fluid dynamics simulation experiments.

97 MATHEMATICS AND COMPUTING↗

R-Adaptivity to Enable Compression of Elementary Computations in Extreme-Scale Finite Element Simulators

Modern computing systems are capable of exascale calculations, which are revolutionizing the development and application of high-fidelity numerical models in computational science and engineering. While these systems continue to grow in processing power, the available system memory has not increased commensurately, and electrical power consumption continues to grow. A predominant approach to limit the memory usage in large-scale applications is to exploit the abundant processing power and continually recompute many low-level simulation quantities, rather than storing them. However, this approach can adversely impact the throughput of the simulation and diminish the benefits of modern computing architectures. We present three novel contributions to reduce the memory burden while maintaining, and sometimes improving, performance in simulations based on finite element discretizations. The first contribution develops dictionary-based data compression schemes that detect and exploit the structure of the discretization, due to redundancies across the finite element mesh. While these schemes are shown to reduce memory requirements by more than 99% on meshes with large numbers of identical mesh cells, there are applications where this structure does not exist. The second contribution leverages a recently developed augmented Lagrangian optimization algorithm to enable r-adaptivity for meshes with the goal of enhancing the redundancies in the mesh. The third contribution extends these methods to patch-based linear solvers and preconditioners by compressing local matrices. Numerical results demonstrate the effectiveness of the proposed methods to detect, enhance and exploit mesh structure on a suite of examples inspired by large-scale applications.

97 MATHEMATICS AND COMPUTING↗

Parallel Runtime Interface for Fortran (PRIF): A Multi-Image Solution for LLVM Flang

Fortran compilers that provide support for Fortran’s native parallel features often do so with a runtime library that depends on details of both the compiler implementation and the communication library, while others provide limited or no support at all. This paper introduces a new generalized interface that is both compiler- and runtime-library-agnostic, providing flexibility while fully supporting all of Fortran’s parallel features. The Parallel Runtime Interface for Fortran (PRIF) was developed to be portable across shared- and distributed-memory systems, with varying operating systems, toolchains and architectures. It achieves this by defining a set of Fortran procedures corresponding to each of the parallel features defined in the Fortran standard that may be invoked by a Fortran compiler and implemented by a runtime library. PRIF aims to be used as the solution for LLVM Flang to provide parallel Fortran support. This paper also briefly describes our PRIF prototype implementation: Caffeine.

Bonachea, Dan↗

Technical note: Optimizing the in situ cosmogenic 36 Cl extraction and measurement workflow for geologic applications

Abstract. In situ cosmogenic 36Cl analysis by accelerator mass spectrometry (AMS) is routinely employed to date Quaternary surfaces and assess rates of landscape evolution. However, standard laboratory preparation procedures for 36Cl dating require the addition of large amounts of isotopically enriched chlorine spike solution; these solutions are expensive and increasingly difficult to acquire from commercial sources. In addition, the typical workflow for 36Cl dating involves measuring both 35Cl/37Cl and 36Cl/Cl concurrently on the high-energy (post-accelerator) end of the AMS system, but 35Cl/37Cl determinations using this technique can be complicated by isotope fractionation and system memory during measurement. The traditional workflow also does not provide 36Cl extraction laboratories with the data needed to calculate native Cl concentrations in advance of 36Cl/Cl measurements. In light of these concerns, we present an improved workflow for extracting and measuring chlorine in geologic materials. Our initial step is to characterize 35Cl/37Cl on sample aliquots of up to ∼1 g prepared in Ag(Cl, Br) matrices, which greatly reduces the amount of isotopically enriched spike solution required to measure native Cl content in each sample. To avoid potential issues with isotope fractionation through the accelerator, 35Cl/37Cl is measured on the low-energy, pre-accelerator end of the AMS line. Then, for 36Cl/Cl measurements, we extract Cl as AgCl or Ag(Cl, Br) in analytical batches with a consistent total Cl load across all samples; this step is intended to minimize source memory effects during 36Cl/Cl measurements and allows the preparation of AMS standards that are customized to match known Cl contents in the samples. To assess the efficacy of this extraction and measurement workflow, we compare chlorine isotope ratio measurements on seven geologic samples prepared using standard procedures and the updated workflow. Measurements of 35Cl/37Cl and 36Cl/Cl are consistent between the two workflows, and 35Cl/37Cl values measured using our methods have considerably higher precision than those measured following standard protocols. The chemical preparation and measurement workflow presented here (1) reduces the amount of isotopically enriched chlorine spike used per rock sample by up to 95 %; (2) identifies rocks with high native Cl concentrations, which may be lower priority for 36Cl surface exposure dating, at an early stage of analysis; and (3) allows laboratory users to maintain control over the total chlorine content within and across analytical batches. These methods can be incorporated into existing laboratory and AMS protocols for 36Cl analyses and will increase the accessibility of 36Cl dating for geologic applications.

58 GEOSCIENCES↗

Improved solid state electron-charge-storage device

Storage device is applicable in memory systems and in high-resolution arrays for light-responsive image sensing. The device offers high yield in multiple arrays and allows charge release with light striking only the edge of a metal electrode.

Kuper, A. B.↗

An improved holographic recording medium

Solid, linear chain hydrocarbons with molecular weight ranging from about 300 to 2000 can serve as long-lived recording medium in optical memory system. Suitable recording hydrocarbons include microcrystalline waxes and low molecular weight polymers or ethylene.

Gange, R. A.↗