Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Memory systems”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 343 records · Page 19

Compiler-directed cache management in multiprocessors

The necessity of finding alternatives to hardware-based cache coherence strategies for large-scale multiprocessor systems is discussed. Three different software-based strategies sharing the same goals and general approach are presented. They consist of a simple invalidation approach, a fast selective invalidation scheme, and a version control scheme. The strategies are suitable for shared-memory multiprocessor systems with interconnection networks and a large number of processors. Results of trace-driven simulations conducted on numerical benchmark routines to compare the performance of the three schemes are presented.

Cheong, Hoichi↗

Memory access statistics monitoring

Systems, apparatuses, and methods related to memory access statistics monitoring are described. A host is configured to map pages of memory for applications to a number of memory devices coupled thereto. A first memory device comprises a monitoring component configured to monitor access statistics of pages of memory mapped to the first memory device. A second memory device does not include a monitoring component capable of monitoring access statistics of pages of memory mapped thereto. The host is configured to map a portion of pages of memory for an application to the first memory device in order to obtain access statistics corresponding to the portion of pages of memory upon execution of the application despite there being space available on the second memory device and adjust mappings of the pages of memory for the application based on the obtained access statistics corresponding to the portion of pages.

Roberts, David A.↗

Unipolar terminal-attractor based neural associative memory with adaptive threshold

A unipolar terminal-attractor based neural associative memory (TABAM) system with adaptive threshold for perfect convergence is presented. By adaptively setting the threshold values for the dynamic iteration for the unipolar binary neuron states with terminal-attractors for the purpose of reducing the spurious states in a Hopfield neural network for associative memory and using the inner product approach, perfect convergence and correct retrieval is achieved. Simulation is completed with a small number of stored states (M) and a small number of neurons (N) but a large M/N ratio. An experiment with optical exclusive-OR logic operation using LCTV SLMs shows the feasibility of optoelectronic implementation of the models. A complete inner-product TABAM is implemented using a PC for calculation of adaptive threshold values to achieve a unipolar TABAM (UIT) in the case where there is no crosstalk, and a crosstalk model (CRIT) in the case where crosstalk corrupts the desired state.

Liu, Hua-Kuang↗

Unipolar Terminal-Attractor Based Neural Associative Memory with Adaptive Threshold

A unipolar terminal-attractor based neural associative memory (TABAM) system with adaptive threshold for perfect convergence is presented. By adaptively setting the threshold values for the dynamic iteration for the unipolar binary neuron states with terminal-attractors for the purpose of reducing the spurious states in a Hopfield neural network for associative memory and using the inner-product approach, perfect convergence and correct retrieval is achieved. Simulation is completed with a small number of stored states (M) and a small number of neurons (N) but a large M/N ratio. An experiment with optical exclusive-OR logic operation using LCTV SLMs shows the feasibility of optoelectronic implementation of the models. A complete inner-product TABAM is implemented using a PC for calculation of adaptive threshold values to achieve a unipolar TABAM (UIT) in the case where there is no crosstalk, and a crosstalk model (CRIT) in the case where crosstalk corrupts the desired state.

Liu, Hua-Kuang↗

Automatic multi-banking of memory for microprocessors

A microprocessor system is provided with added memories to expand its address spaces beyond its address word length capacity by using indirect addressing instructions of a type having a detectable operations code and dedicating designated address spaces of memory to each of the added memories, one space to a memory. By decoding each operations code of instructions read from main memory into a decoder to identify indirect addressing instructions of the specified type, and then decoding the address that follows in a decoder to determine which added memory is associated therewith, the associated added memory is selectively enabled through a unit while the main memory is disabled to permit the instruction to be executed on the location to which the effective address of the indirect address instruction points, either before the indirect address is read from main memory or afterwards, depending on how the system is arranged by a switch.

Wiker, G. A.↗

Concurrent file operations in a high performance FORTRAN

Distributed memory multiprocessor systems can provide the computing power necessary for large scale scientific applications. A critical performance issue for a number of these applications is the efficient transfer of data to secondary storage. Recently several research groups have proposed FORTRAN language extensions for exploiting the data parallelism of such scientific codes on distributed memory architectures. However, few of these high performance FORTRAN's provide appropriate constructs for controlling the use of the parallel I/O capabilities of modern multiprocessing machines. In this paper, we propose constructs to specify I/O operations for distributed data structures in the context of Vienna Fortran. These operations can be used by the programmer to provide information which can help the compiler and runtime environment make the most efficient use of the I/O subsystem.

Brezany, Peter↗

Modernization efforts for the R -Matrix code SAMMY [Abstract]

The R-Matrix code SAMMY is a widely used nuclear data evaluation code focused on the resolved range, which includes corrections for experimental effects. The code is still mostly written in Fortran 77, and uses a memory management system suitable for the time of its initial writing (1984). A modernization effort is under way to bring the code in-line with modern software development practices. A continuous-integration testing framework was added, automating the large existing set of test cases. It is run on every commit. The memory management was updated to current standard practices suitable for modern software analysis tools. The code can be obtained from https://code.ornl.gov/RNSD/SAMMY. The resonance parameters and covariance information are now stored in C++ objects shared by SAMMY and AMPX, the processing code that generates nuclear data libraries for SCALE. This allows for easier maintenance and access to the resonance parameters inside and outside of SAMMY. This feature is already used by accessing and changing parameters in memory in the Bayesian Monte Carlo Evaluation Framework for Cross Sections Nuclear Data and Integral Benchmark Experiments project, Further plans include the switch to the ENDF reading and writing routines in AMPX, as these routines are more robust, easier to maintain, and support more features. Of note here is support for the new GNDS format. Previously it wasn’t easy to share the full covariance matrix for evaluations containing more than one isotope due to limitations on the ENDF format; this is now supported in GNDS. The data are currently available in a binary SAMMY format and can be exported to GNDS to make them more widely available and sharable. The next step will be to use the same resonance processing code at 0K in AMPX and SAMMY as one of the available Reich-Moore R-Matrix formalism. The first step toward this goal is to isolate the reconstruction into a module that takes resonance parameters as its input and does not depend on SAMMY global parameters. This goal has been achieved and it should now be possible to more easily change the resonance formalism and add enhancements as the Phenomenological R-Matrix parameterization of direct, doorway, and compound nuclear reactions discussed elsewhere on this conference. This concerted modernization and enhancement effort provides multiple advantages to the nuclear data community. It will allow parameter optimization using enhanced formalisms, including experimental effects, that better match complex experimental data. Then those evaluated parameters can immediately be passed off to AMPX to be reconstructed with the exact same cross section model and be put into a data library for subsequent testing using SCALE and the Valid Benchmark suite or other suitable benchmark suites.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Presentation on a Space Acceleration Measurement System (SAMS)

The primary objective of the Space Acceleration Measurement Systems (SAMS) project is to provide an acceleration measurement system capable of serving a wide variety of space experiments. The design of the system being developed under this project takes into consideration requirements for experiments located in the middeck, in the orbiter bay, and in Spacelab. In addition to measuring, conditioning, and recording accelerations, the system will be capable of performing complex calculations and interactive control. The main components consist of a remote triaxial optical storage device. In operation, the triaxial sensor head produces output signals in response to acceleration inputs. These signals are preamplified, filtered and converted into digital data which is then transferred to optical memory. The system design is modular, facilitating both software and hardware upgrading as technology advances. Two complete acceleration measurement flight systems will be build and tested under this project.

Chase, Theodore L.↗

Modular chip-integrated photonic control of artificial atoms in diamond waveguides

A central goal in creating long-distance quantum networks and distributed quantum computing is the development of interconnected and individually controlled qubit nodes. Atom-like emitters in diamond have emerged as a leading system for optically networked quantum memories, motivating the development of visible-spectrum, multi-channel photonic integrated circuit (PIC) systems for scalable atom control. However, it has remained an open challenge to realize optical programmability with a qubit layer that can achieve high optical detection probability over many optical channels. Here, we address this problem by introducing a modular architecture of piezoelectrically actuated atom-control PICs (APICs) and artificial atoms embedded in diamond nanostructures designed for high-efficiency free-space collection. The high-speed four-channel APIC is based on a splitting tree mesh with triple-phase shifter Mach–Zehnder interferometers. This design simultaneously achieves optically broadband operation at visible wavelengths, high-fidelity switching (>40dB) at low voltages, submicrosecond modulation timescales (>30MHz), and minimal channel-to-channel crosstalk for repeatable optical pulse carving. Via a reconfigurable free-space interconnect, we use the APIC to address single silicon vacancy color centers in individual diamond waveguides with inverse tapered couplers, achieving efficient single photon detection probabilities (∼15%) and second-order autocorrelation measurements g (2) (0)<0.14 for all channels. The modularity of this distributed APIC–quantum memory system simplifies the quantum control problem, potentially enabling further scaling to thousands of channels.

47 OTHER INSTRUMENTATION↗

Identification of Linear and Nonlinear Aerodynamic Impulse Responses Using Digital Filter Techniques

This paper discusses the mathematical existence and the numerically-correct identification of linear and nonlinear aerodynamic impulse response functions. Differences between continuous-time and discrete-time system theories, which permit the identification and efficient use of these functions, will be detailed. Important input/output definitions and the concept of linear and nonlinear systems with memory will also be discussed. It will be shown that indicial (step or steady) responses (such as Wagner's function), forced harmonic responses (such as Tbeodorsen's function or those from doublet lattice theory), and responses to random inputs (such as gusts) can all be obtained from an aerodynamic impulse response function. This paper establishes the aerodynamic impulse response function as the most fundamental, and, therefore, the most computationally efficient, aerodynamic function that can be extracted from any given discrete-time, aerodynamic system. The results presented in this paper help to unify the understanding of classical two-dimensional continuous-time theories with modem three-dimensional, discrete-time theories. First, the method is applied to the nonlinear viscous Burger's equation as an example. Next the method is applied to a three-dimensional aeroelastic model using the CAP-TSD (Computational Aeroelasticity Program - Transonic Small Disturbance) code and then to a two-dimensional model using the CFL3D Navier-Stokes code. Comparisons of accuracy and computational cost savings are presented. Because of its mathematical generality, an important attribute of this methodology is that it is applicable to a wide range of nonlinear, discrete-time problems.

Silva, Walter A.↗

Identification of Linear and Nonlinear Aerodynamic Impulse Responses Using Digital Filter Techniques

This paper discusses the mathematical existence and the numerically-correct identification of linear and nonlinear aerodynamic impulse response functions. Differences between continuous-time and discrete-time system theories, which permit the identification and efficient use of these functions, will be detailed. Important input/output definitions and the concept of linear and nonlinear systems with memory will also be discussed. It will be shown that indicial (step or steady) responses (such as Wagner's function), forced harmonic responses (such as Theodorsen's function or those from doublet lattice theory), and responses to random inputs (such as gusts) can all be obtained from an aerodynamic impulse response function. This paper establishes the aerodynamic impulse response function as the most fundamental, and, therefore, the most computationally efficient, aerodynamic function that can be extracted from any given discrete-time, aerodynamic system. The results presented in this paper help to unify the understanding of classical two-dimensional continuous-time theories with modern three-dimensional, discrete-time theories. First, the method is applied to the nonlinear viscous Burger's equation as an example. Next the method is applied to a three-dimensional aeroelastic model using the CAP-TSD (Computational Aeroelasticity Program - Transonic Small Disturbance) code and then to a two-dimensional model using the CFL3D Navier-Stokes code. Comparisons of accuracy and computational cost savings are presented. Because of its mathematical generality, an important attribute of this methodology is that it is applicable to a wide range of nonlinear, discrete-time problems.

Silva, Walter A.↗

A parallel dynamic load balancing algorithm for 3-D adaptive unstructured grids

Adaptive local grid refinement and coarsening results in unequal distribution of workload among the processors of a parallel system. A novel method for balancing the load in cases of dynamically changing tetrahedral grids is developed. The approach employs local exchange of cells among processors in order to redistribute the load equally. An important part of the load balancing algorithm is the method employed by a processor to determine which cells within its subdomain are to be exchanged. Two such methods are presented and compared. The strategy for load balancing is based on the Divide-and-Conquer approach which leads to an efficient parallel algorithm. This method is implemented on a distributed-memory MIMD system.

Vidwans, A.↗

An Unipolar Terminal-Attractor Based Neural Associative Memory with Adaptive Threshold and Perfect Convergence

For the first time, a unipolar terminal-attractor based neural associative memory (TABAM)system with adaptive threshold and perfect convergence is presented. By adaptively setting thethreshold values for the dynamic iteration for the unipolar binary neuron states with terminal-attractors and inner-product approach, we demonstrate via computer simulation the achievement ofperfect convergence and correct retrieval. The simulation is completed with a small number of storedstates (M) and a small number of neurons (N) but a large M/N ratio. An experiment with exclusive-or logic operation using LCTV SLMs is used to show feasibility of the optoelectronic implementationof the models.

Associative↗

Development of a Computer Architecture to Support the Optical Plume Anomaly Detection (OPAD) System

The NASA OPAD spectrometer system relies heavily on extensive software which repetitively extracts spectral information from the engine plume and reports the amounts of metals which are present in the plume. The development of this software is at a sufficiently advanced stage where it can be used in actual engine tests to provide valuable data on engine operation and health. This activity will continue and, in addition, the OPAD system is planned to be used in flight aboard space vehicles. The two implementations, test-stand and in-flight, may have some differing requirements. For example, the data stored during a test-stand experiment are much more extensive than in the in-flight case. In both cases though, the majority of the requirements are similar. New data from the spectrograph is generated at a rate of once every 0.5 sec or faster. All processing must be completed within this period of time to maintain real-time performance. Every 0.5 sec, the OPAD system must report the amounts of specific metals within the engine plume, given the spectral data. At present, the software in the OPAD system performs this function by solving the inverse problem. It uses powerful physics-based computational models (the SPECTRA code), which receive amounts of metals as inputs to produce the spectral data that would have been observed, had the same metal amounts been present in the engine plume. During the experiment, for every spectrum that is observed, an initial approximation is performed using neural networks to establish an initial metal composition which approximates as accurately as possible the real one. Then, using optimization techniques, the SPECTRA code is repetitively used to produce a fit to the data, by adjusting the metal input amounts until the produced spectrum matches the observed one to within a given level of tolerance. This iterative solution to the original problem of determining the metal composition in the plume requires a relatively long period of time to execute the software in a modern single-processor workstation, and therefore real-time operation is currently not possible. A different number of iterations may be required to perform spectral data fitting per spectral sample. Yet, the OPAD system must be designed to maintain real-time performance in all cases. Although faster single-processor workstations are available for execution of the fitting and SPECTRA software, this option is unattractive due to the excessive cost associated with very fast workstations and also due to the fact that such hardware is not easily expandable to accommodate future versions of the software which may require more processing power. Initial research has already demonstrated that the OPAD software can take advantage of a parallel computer architecture to achieve the necessary speedup. Current work has improved the software by converting it into a form which is easily parallelizable. Timing experiments have been performed to establish the computational complexity and execution speed of major components of the software. This work provides the foundation of future work which will create a fully parallel version of the software executing in a shared-memory multiprocessor system.

Katsinis, Constantine↗

Competing Interactions between Mesoscale Length-Scales, Order-Disorder, and Martensitic Transformation in Ferromagnetic Shape Memory Alloys

In the present study, the effects of composition and heat treatments and the resulting microstructural changes on the martensitic transformation and ferromagnetic transition have been investigated in NiCoMnIn metamagnetic shape memory alloys. In this shape memory alloy system, it is observed that upon heat treatments at a wide temperature range, the onset temperature of the martensitic transformation follows a non-monotonic behavior with respect to heat treatment time, and the nature of the non-monotonic behavior is also a function of composition. This behavior cannot be attributed to well-known factors such as precipitation, change in local composition due to the precipitation and/or global degree of order. In this work, a systematic investigation through synthesis, thermal processing and characterization via thermo-physical measurements, transmission electron microscopy and in-situ synchrotron x-ray diffraction experiments has been used to correlate the non-monotonic dependence of the martensitic transformation and ferromagnetic transition temperatures to the evolution of L2 1 domains arising from a order/disorder phase transition. A thermodynamic model for the magneto-structural transition is combined with classical nucleation theory to further ascertain the role of microstructural length-scales on the onset of the martensitic transformation. This work thus provides understanding of the thermodynamic and kinetic factors that can be controlled to tune the coupled magneto-structural transformations in NiCoMnIn metamagnetic shape memory alloys.

36 MATERIALS SCIENCE↗

Initial Kernel Timing Using a Simple PIM Performance Model

This presentation will describe some initial results of paper-and-pencil studies of 4 or 5 application kernels applied to a processor-in-memory (PIM) system roughly similar to the Cascade Lightweight Processor (LWP). The application kernels are: * Linked list traversal * Sun of leaf nodes on a tree * Bitonic sort * Vector sum * Gaussian elimination The intent of this work is to guide and validate work on the Cascade project in the areas of compilers, simulators, and languages. We will first discuss the generic PIM structure. Then, we will explain the concepts needed to program a parallel PIM system (locality, threads, parcels). Next, we will present a simple PIM performance model that will be used in the remainder of the presentation. For each kernel, we will then present a set of codes, including codes for a single PIM node, and codes for multiple PIM nodes that move data to threads and move threads to data. These codes are written at a fairly low level, between assembly and C, but much closer to C than to assembly. For each code, we will present some hand-drafted timing forecasts, based on the simple PIM performance model. Finally, we will conclude by discussing what we have learned from this work, including what programming styles seem to work best, from the point-of-view of both expressiveness and performance.

BRIEFING CHARTS↗

Method and apparatus for front end gather/scatter memory coalescing

A system for processing gather and scatter instructions can implement a front-end subsystem, a back-end subsystem, or both. The front-end subsystem includes a prediction unit configured to determine a predicted quantity of coalesced memory access operations required by an instruction. A decode unit converts the instruction into a plurality of access operations based on the predicted quantity, and transmits the plurality of access operations and an indication of the predicted quantity to an issue queue. The back-end subsystem includes a load-store unit that receives a plurality of access operations corresponding to an instruction, determines a subset of the plurality of access operations that can be coalesced, and forms a coalesced memory access operation from the subset. A queue stores multiple memory addresses for a given load-store entry to provide for execution of coalesced memory accesses.

97 MATHEMATICS AND COMPUTING↗

Method and apparatus for back end gather/scatter memory coalescing

A system for processing gather and scatter instructions can implement a front-end subsystem, a back-end subsystem, or both. The front-end subsystem includes a prediction unit configured to determine a predicted quantity of coalesced memory access operations required by an instruction. A decode unit converts the instruction into a plurality of access operations based on the predicted quantity, and transmits the plurality of access operations and an indication of the predicted quantity to an issue queue. The back-end subsystem includes a load-store unit that receives a plurality of access operations corresponding to an instruction, determines a subset of the plurality of access operations that can be coalesced, and forms a coalesced memory access operation from the subset. A queue stores multiple memory addresses for a given load-store entry to provide for execution of coalesced memory accesses.

97 MATHEMATICS AND COMPUTING↗