Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “limited memory”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

Redox transistors based on TiO 2 for analogue neuromorphic computing

The ability to train deep neural networks on large data sets have made significant impacts onto artificial intelligence, but consume significant amounts of energy due to the need to move information from memory to logic units. In-memory "neuromorphic" computing presents an alternative framework that processes information directly on memory elements. In-memory computing has been limited by the poor performance of the analogue information storage element, often phase-change memory or memristors. To solve this problem, we developed two types of "redox transistors" using TiO 2 (anatase) which stores analogue information states through the electrochemical concentration of dopants in the crystal. The first type of redox transistor uses lithium as the electrochemical dopant ion, and its key advantage is low operating voltage. The second uses oxygen vacancies as the dopant, which is CMOS compatible and can retain state even when scaled to nanosized dimensions. Both devices offer significant advantages in terms of predictable analogue switching over conventional filamentary-based devices, and provide a significant advance in developing materials and devices for neuromorphic computing.

36 MATERIALS SCIENCE↗

Numerical Prediction Methods (Reynolds-Averaged Navier-Stokes Simulations of Transonic Separated Flows)

During the past five years, numerous pioneering archival publications have appeared that have presented computer solutions of the mass-weighted, time-averaged Navier-Stokes equations for transonic problems pertinent to the aircraft industry. These solutions have been pathfinders of developments that could evolve into a major new technological capability, namely the computational Navier-Stokes technology, for the aircraft industry. So far these simulations have demonstrated that computational techniques, and computer capabilities have advanced to the point where it is possible to solve forms of the Navier-Stokes equations for transonic research problems. At present there are two major shortcomings of the technology: limited computer speed and memory, and difficulties in turbulence modelling and in computation of complex three-dimensional geometries. These limitations and difficulties are the pacing items of the continuing developments, although the one item that will most likely turn out to be the most crucial to the progress of this technology is turbulence modelling. The objective of this presentation is to discuss the state of the art of this technology and suggest possible future areas of research. We now discuss some of the flow conditions for which the Navier-Stokes equations appear to be required. On an airfoil there are four different types of interaction of a shock wave with a boundary layer: (1) shock-boundary-layer interaction with no separation, (2) shock-induced turbulent separation with immediate reattachment (we refer to this as a shock-induced separation bubble), (3) shock-induced turbulent separation without reattachment, and (4) shock-induced separation bubble with trailing edge separation.

Mehta, Unmeel↗

Electrically Variable Resistive Memory Devices

Nonvolatile electronic memory devices that store data in the form of electrical- resistance values, and memory circuits based on such devices, have been invented. These devices and circuits exploit an electrically-variable-resistance phenomenon that occurs in thin films of certain oxides that exhibit the colossal magnetoresistive (CMR) effect. It is worth emphasizing that, as stated in the immediately preceding article, these devices function at room temperature and do not depend on externally applied magnetic fields. A device of this type is basically a thin film resistor: it consists of a thin film of a CMR material located between, and in contact with, two electrical conductors. The application of a short-duration, low-voltage current pulse via the terminals changes the electrical resistance of the film. The amount of the change in resistance depends on the size of the pulse. The direction of change (increase or decrease of resistance) depends on the polarity of the pulse. Hence, a datum can be written (or a prior datum overwritten) in the memory device by applying a pulse of size and polarity tailored to set the resistance at a value that represents a specific numerical value. To read the datum, one applies a smaller pulse - one that is large enough to enable accurate measurement of resistance, but small enough so as not to change the resistance. In writing, the resistance can be set to any value within the dynamic range of the CMR film. Typically, the value would be one of several discrete resistance values that represent logic levels or digits. Because the number of levels can exceed 2, a memory device of this type is not limited to binary data. Like other memory devices, devices of this type can be incorporated into a memory integrated circuit by laying them out on a substrate in rows and columns, along with row and column conductors for electrically addressing them individually or collectively.

Liu, Shangqing↗

The use of computers for instruction in fluid dynamics

Applications for computers which improve instruction in fluid dynamics are examined. Computers can be used to illustrate three-dimensional flow fields and simple fluid dynamics mechanisms, to solve fluid dynamics problems, and for electronic sketching. The usefulness of computer applications is limited by computer speed, memory, and software and the clarity and field of view of the projected display. Proposed advances in personal computers which will address these limitations are discussed. Long range applications for computers in education are considered.

Watson, Val↗

GrainPaint: A multi-scale diffusion-based generative model for microstructure reconstruction of large-scale objects

Simulation-based approaches to microstructure generation can suffer from a variety of limitations, such as high memory usage, long computational times, and difficulties in generating complex geometries. Generative machine learning models present a way around these issues, but they have previously been limited by the fixed size of their generation area. Here, we present a new microstructure generation methodology leveraging advances in inpainting using denoising diffusion models to overcome this generation area limitation. We show that microstructures generated with the presented methodology are statistically similar to grain structures generated with a kinetic Monte Carlo simulator, SPPARKS.

36 MATERIALS SCIENCE↗

Higher-Order Neural Networks Applied to 2D and 3D Object Recognition

A Higher-Order Neural Network (HONN) can be designed to be invariant to geometric transformations such as scale, translation, and in-plane rotation. Invariances are built directly into the architecture of a HONN and do not need to be learned. Thus, for 2D object recognition, the network needs to be trained on just one view of each object class, not numerous scaled, translated, and rotated views. Because the 2D object recognition task is a component of the 3D object recognition task, built-in 2D invariance also decreases the size of the training set required for 3D object recognition. We present results for 2D object recognition both in simulation and within a robotic vision experiment and for 3D object recognition in simulation. We also compare our method to other approaches and show that HONNs have distinct advantages for position, scale, and rotation-invariant object recognition. The major drawback of HONNs is that the size of the input field is limited due to the memory required for the large number of interconnections in a fully connected network. We present partial connectivity strategies and a coarse-coding technique for overcoming this limitation and increasing the input field to that required by practical object recognition problems.

Spirkovska, Lilly↗

Complexity-calibrated benchmarks for machine learning reveal when prediction algorithms succeed and mislead

Abstract Recurrent neural networks are used to forecast time series in finance, climate, language, and from many other domains. Reservoir computers are a particularly easily trainable form of recurrent neural network. Recently, a “next-generation” reservoir computer was introduced in which the memory trace involves only a finite number of previous symbols. We explore the inherent limitations of finite-past memory traces in this intriguing proposal. A lower bound from Fano’s inequality shows that, on highly non-Markovian processes generated by large probabilistic state machines, next-generation reservoir computers with reasonably long memory traces have an error probability that is at least $$\sim 60\%$$ ∼ 60 % higher than the minimal attainable error probability in predicting the next observation. More generally, it appears that popular recurrent neural networks fall far short of optimally predicting such complex processes. These results highlight the need for a new generation of optimized recurrent neural network architectures. Alongside this finding, we present concentration-of-measure results for randomly-generated but complex processes. One conclusion is that large probabilistic state machines—specifically, large $$\epsilon$$ ϵ -machines—are key to generating challenging and structurally-unbiased stimuli for ground-truthing recurrent neural network architectures.

97 MATHEMATICS AND COMPUTING↗

Photonic and Opto-Electronic Applications of Polydiacetylene Films Photodeposited from Solution and Polydiacetylene Copolymer Networks

Polydiacetylenes (PDAS) are attractive materials for both electronic and photonic applications because of their highly conjugated electronic structures. They have been investigated for applications as both one-dimensional (linear chain) conductors and nonlinear optical (NLO) materials. One of the chief limitations to the use of PDAs has been the inability to readily process them into useful forms such as films and fibers. In our laboratory we have developed a novel process for obtaining amorphous films of a PDA derived from 2-methyl4-nitroaniline using photodeposition with Ultraviolet (UV) light from monomer solutions onto transparent substrates. Photodeposition from solution provides a simple technique for obtaining PDA films in any desired pattern with good optical quality. This technique has been used to produce PDA films that show potential for optical applications such as holographic memory storage and optical limiting, as well as third-order NLO applications such as all-optical refractive index modulation, phase modulation and switching. Additionally, copolymerization of diacetylenes with other monomers such as methacrylates provides a means to obtain materials with good processibility. Such copolymers can be spin cast to form films, or drawn by either melt or solution extrusion into fibers. These films or fibers can then be irradiated with UV to photopolymerize the diacetylene units to form a highly stable cross-linked PDA-copolymer network. If such films are electrically poled while being irradiated, they can achieve the asymmetry necessary for second-order NLO applications such as electro-optic switching. On Earth, formation of PDAs by the above mentioned techniques suffers from defects and inhomogeneities caused by convective flows that can arise during processing. By studying the formation of these materials in the reduced-convection, diffusion-controlled environment of space we hope to better understand the factors that affect their processing, and thereby, their nature and properties. Ultimately it may even be feasible to conduct space processing of PDAs for technological applications.

Paley, Mark S.↗

Multiplexed Holographic Data Storage in Bacteriorhodopsin

Biochrome photosensitive films in particular Bacteriorhodopsin exhibit features which make these materials an attractive recording medium for optical data storage and processing. Bacteriorhodopsin films find numerous applications in a wide range of optical data processing applications; however the short-term memory characteristics of BR limits their applications for holographic data storage. The life-time of the BR can be extended using cryogenic temperatures [1], although this method makes the system overly complicated and unstable. Longer life-times can be provided in one modification of BR - the "blue" membrane BR [2], however currently available films are characterized by both low diffraction efficiency and difficulties in providing photoreversible recording. In addition, as a dynamic recording material, the BR requires different wavelengths for recording and reconstructing of optical data in order to prevent the information erasure during its readout. This fact also put constraints on a BR-based Optical Memory, due to information loss in holographic memory systems employing the two-lambda technique for reading-writing thick multiplexed holograms.

Mehrl, David J.↗

Singleton Sieving: Overcoming the Memory/Speed Trade-Off in Exascale k-mer Analysis

Traditional filter data structures, such as Bloom filters, do not offer necessary features that modern high-performance data analytics applications need in order to efficiently perform complex data analysis tasks. For example, MetaHipMer, a de novo metagenome assembler, can use filters to weed out singleton k-mers and reduce memory usage by 30%-70%. However, the filter needs the ability to associate values with k-mers in order to perform the analysis in a single communication pass. Bloom filters do not support value associations and cause the application to perform an extra communication pass, thereby increasing the run time. Therefore, MetaHipMer faces a trade off between memory and speed due to the limited capabilities of traditional filters. In this paper, we overcome the memory and speed trade off in MetaHipMer by integrating a GPU-based feature-rich filter, the Two-Choice filter (TCF), in the MetaHipMer pipeline. The TCF uses key-value association to approximately store k-mers with extensions. This allows MetaHipMer to perform k-mer analysis on the GPUs in a single communication pass. Our empirical analysis shows a 50% reduction in memory usage in k-mer analysis on each node in MetaHipMer without any effect on the overall run time or assembly quality. The memory reduction in turn results in a 43% reduction in the number of nodes required to assemble datasets and enables MetaHipMer to scale to much larger datasets.

McCoy, Hunter↗

Filament-Free Bulk Resistive Memory Enables Deterministic Analogue Switching

Digital computing is nearing its physical limits as computing needs and energy consumption rapidly increase. Analogue-memory-based neuromorphic computing can be orders of magnitude more energy efficient at data-intensive tasks like deep neural networks, but has been limited by the inaccurate and unpredictable switching of analogue resistive memory. Filamentary resistive random access memory (RRAM) suffers from stochastic switching due to the random kinetic motion of discrete defects in the nanometer-sized filament. Here, this stochasticity is overcome by incorporating a solid electrolyte interlayer, in this case, yttria-stabilized zirconia (YSZ), toward eliminating filaments. Filament-free, bulk-RRAM cells instead store analogue states using the bulk point defect concentration, yielding predictable switching because the statistical ensemble behavior of oxygen vacancy defects is deterministic even when individual defects are stochastic. Both experiments and modeling show bulk-RRAM devices using TiO2-X switching layers and YSZ electrolytes yield deterministic and linear analogue switching for efficient inference and training. Bulk-RRAM solves many outstanding issues with memristor unpredictability that have inhibited commercialization, and can, therefore, enable unprecedented new applications for energy-efficient neuromorphic computing. Beyond RRAM, this work shows how harnessing bulk point defects in ionic materials can be used to engineer deterministic nanoelectronic materials and devices.

36 MATERIALS SCIENCE↗

Applications Performance on NAS Intel Paragon XP/S - 15#

The Numerical Aerodynamic Simulation (NAS) Systems Division received an Intel Touchstone Sigma prototype model Paragon XP/S- 15 in February, 1993. The i860 XP microprocessor with an integrated floating point unit and operating in dual -instruction mode gives peak performance of 75 million floating point operations (NIFLOPS) per second for 64 bit floating point arithmetic. It is used in the Paragon XP/S-15 which has been installed at NAS, NASA Ames Research Center. The NAS Paragon has 208 nodes and its peak performance is 15.6 GFLOPS. Here, we will report on early experience using the Paragon XP/S- 15. We have tested its performance using both kernels and applications of interest to NAS. We have measured the performance of BLAS 1, 2 and 3 both assembly-coded and Fortran coded on NAS Paragon XP/S- 15. Furthermore, we have investigated the performance of a single node one-dimensional FFT, a distributed two-dimensional FFT and a distributed three-dimensional FFT Finally, we measured the performance of NAS Parallel Benchmarks (NPB) on the Paragon and compare it with the performance obtained on other highly parallel machines, such as CM-5, CRAY T3D, IBM SP I, etc. In particular, we investigated the following issues, which can strongly affect the performance of the Paragon: a. Impact of the operating system: Intel currently uses as a default an operating system OSF/1 AD from the Open Software Foundation. The paging of Open Software Foundation (OSF) server at 22 MB to make more memory available for the application degrades the performance. We found that when the limit of 26 NIB per node out of 32 MB available is reached, the application is paged out of main memory using virtual memory. When the application starts paging, the performance is considerably reduced. We found that dynamic memory allocation can help applications performance under certain circumstances. b. Impact of data cache on the i860/XP: We measured the performance of the BLAS both assembly coded and Fortran coded. We found that the measured performance of assembly-coded BLAS is much less than what memory bandwidth limitation would predict. The influence of data cache on different sizes of vectors is also investigated using one-dimensional FFTs. c. Impact of processor layout: There are several different ways processors can be laid out within the two-dimensional grid of processors on the Paragon. We have used the FFT example to investigate performance differences based on processors layout.

Saini, Subhash↗

AENET–LAMMPS and AENET–TINKER : Interfaces for accurate and efficient molecular dynamics simulations with machine learning potentials

Machine-learning potentials (MLPs) trained on data from quantum-mechanics based first-principles methods can approach the accuracy of the reference method at a fraction of the computational cost. To facilitate efficient MLP-based molecular dynamics and Monte Carlo simulations, an integration of the MLPs with sampling software is needed. Here, we develop two interfaces that link the atomic energy network (ænet) MLP package with the popular sampling packages TINKER and LAMMPS. The three packages, ænet, TINKER, and LAMMPS, are free and open-source software that enable, in combination, accurate simulations of large and complex systems with low computational cost that scales linearly with the number of atoms. Scaling tests show that the parallel efficiency of the ænet–TINKER interface is nearly optimal but is limited to shared-memory systems. The ænet–LAMMPS interface achieves excellent parallel efficiency on highly parallel distributed memory systems and benefits from the highly optimized neighbor list implemented in LAMMPS. We demonstrate the utility of the two MLP interfaces for two relevant example applications: the investigation of diffusion phenomena in liquid water and the equilibration of nanostructured amorphous battery materials.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Trust: Triangle Counting Reloaded on GPUs

Triangle counting is a building block for a wide range of graph applications. Here, traditional wisdom suggests that i) hashing is not suitable for triangle counting, ii) edge-centric triangle counting beats vertex-centric design, and iii) communication-free and workload balanced graph partitioning is a grand challenge for triangle counting. On the contrary, we advocate that i) hashing can help the key operations for scalable triangle counting on Graphics Processing Units (GPUs), i.e., list intersection and graph partitioning, ii) vertex-centric option reduces both hash table construction cost and memory consumption, which is limited on GPUs. In addition, iii) we exploit graph and workload collaborative, and hash-based 2D partitioning to scale vertex-centric triangle counting over 1,000 GPUs with sustained scalability. In this work, we present TRUST, which performs triangle counting with the hash operation and vertex-centric paradigm. To the best of our knowledge, TRUST is the first work that achieves over one trillion Traversed Edges Per Second (TEPS) rate for triangle counting.

97 MATHEMATICS AND COMPUTING↗

Tracking 3-D body motion for docking and robot control

An advanced method of tracking three-dimensional motion of bodies has been developed. This system has the potential to dynamically characterize machine and other structural motion, even in the presence of structural flexibility, thus facilitating closed loop structural motion control. The system's operation is based on the concept that the intersection of three planes defines a point. Three rotating planes of laser light, fixed and moving photovoltaic diode targets, and a pipe-lined architecture of analog and digital electronics are used to locate multiple targets whose number is only limited by available computer memory. Data collection rates are a function of the laser scan rotation speed and are currently selectable up to 480 Hz. The tested performance on a preliminary prototype designed for 0.1 in accuracy (for tracking human motion) at a 480 Hz data rate includes a worst case resolution of 0.8 mm (0.03 inches), a repeatability of plus or minus 0.635 mm (plus or minus 0.025 inches), and an absolute accuracy of plus or minus 2.0 mm (plus or minus 0.08 inches) within an eight cubic meter volume with all results applicable at the 95 percent level of confidence along each coordinate region. The full six degrees of freedom of a body can be computed by attaching three or more target detectors to the body of interest.

Donath, M.↗

Data processing assessment for the Lunar Geoscience Observer imaging spectrometer

On the Lunar Geoscience Observer project, a Visible and Infrared Mapping Spectrometer instrument has been proposed. This instrument will have science data input rates in the hundreds of kilobits per second (kbps) and an average telemetry output data rate of 4 kbps. Techniques that can be used to reduce the throughput of the instrument are editing, summing and averaging, data compression, data preprocessing, pattern recognition and snapshot data taking. Due to instrument limitations in the buffer memory size and processing speeds, a careful selection of the available techniques must be made.

Irigoyen, R. E.↗

A high-fidelity batch simulation environment for integrated batch and piloted air combat simulation analysis

A batch air combat simulation environment known as the Tactical Maneuvering Simulator (TMS) is presented. The TMS serves as a tool for developing and evaluating tactical maneuvering logics and to evaluate the tactical implications of perturbations to aircraft performance or supporting systems. The TMS is capable of simulating air combat between any number of engagement participants, with practical limits imposed by computer memory and processing power. Aircraft are modeled using equations of motion, control laws, aerodynamics and propulsive characteristics, and databases representative of a modern high-performance aircraft with and without thrust-vectoring capability are included. A Tactical Autopilot is implemented in the aircraft simulation model to convert guidance commands issued by computerized maneuvering logics in the form of desired angle-of-attack and wind axis-bank angle into inputs to the inner-loop control augmentation system of the aircraft.

Goodrich, Kenneth H.↗