Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “memory reliability”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

Imputation of urban environmental sensor data using gated attention bidirectional long short-term memory (GA-BiLSTM): methods, performance, and implications

Urban environmental monitoring networks frequently encounter significant data gaps due to sensor malfunctions, environmental disturbances, and communication failures. Reliable approaches to address these gaps are essential for ensuring the continuity and quality of environmental data streams. In this study, we developed a gated attention bidirectional long short-term memory (GA-BiLSTM) model to impute missing data in a dense urban monitoring network. Using observations from the CROCUS network in Chicago, we evaluated GA-BiLSTM against widely used approaches (XGBoost and K-nearest neighbors) under scenarios of both short-term intermittent gaps and prolonged outages. GA-BiLSTM consistently outperformed comparative methods, particularly during extended outages of up to ten days, demonstrating its ability to capture spatiotemporal dependencies across sensor nodes. Beyond performance metrics, feature importance and spatial network analyses highlighted the unexpected but critical predictive role of peripheral rural nodes, underlining their strategic value for maintaining robust urban monitoring systems. These results emphasize that advanced imputation methods can substantially improve the reliability of environmental monitoring networks and support more resilient data infrastructures for urban sustainability.

Data imputation↗

Towards automatic Markov reliability modeling of computer architectures

The analysis and evaluation of reliability measures using time-varying Markov models is required for Processor-Memory-Switch (PMS) structures that have competing processes such as standby redundancy and repair, or renewal processes such as transient or intermittent faults. The task of generating these models is tedious and prone to human error due to the large number of states and transitions involved in any reasonable system. Therefore model formulation is a major analysis bottleneck, and model verification is a major validation problem. The general unfamiliarity of computer architects with Markov modeling techniques further increases the necessity of automating the model formulation. This paper presents an overview of the Automated Reliability Modeling (ARM) program, under development at NASA Langley Research Center. ARM will accept as input a description of the PMS interconnection graph, the behavior of the PMS components, the fault-tolerant strategies, and the operational requirements. The output of ARM will be the reliability of availability Markov model formulated for direct use by evaluation programs. The advantages of such an approach are (a) utility to a large class of users, not necessarily expert in reliability analysis, and (b) a lower probability of human error in the computation.

Liceaga, C. A.↗

Error-Free and Current-Driven Synthetic Antiferromagnetic Domain Wall Memory Enabled by Channel Meandering

We propose a new type of energy-efficient multi-bit magnetic memory based on current-driven, field-free, controlled domain wall motion. A meandering domain wall channel with precisely interspersed pinning regions provides the multi-bit capability of a magnetic tunnel junction memory. The magnetic free layer of the memory device has perpendicular magnetic anisotropy (PMA) and interfacial Dzyaloshinskii-Moriya interaction (DMI) so that spin-orbit torques (SOTs) induce efficient domain wall motion. Using micromagnetic simulations, we find two different cell designs: two-way switching and four-way switching. The memory cell design choices and the physics of pinning mechanisms are discussed in detail. Furthermore, we show that switching reliability and speed may be significantly improved by replacing the ferromagnetic free layer with a synthetic antiferromagnetic (SAF) layer. Switching behavior and material choices will be discussed for the two memory implementations.

magnetic domain wall↗

Concatenated coding systems employing a unit-memory convolutional code and a byte-oriented decoding algorithm

Concatenated coding systems utilizing a convolutional code as the inner code and a Reed-Solomon code as the outer code are considered. In order to obtain very reliable communications over a very noisy channel with relatively small coding complexity, it is proposed to concatenate a byte oriented unit memory convolutional code with an RS outer code whose symbol size is one byte. It is further proposed to utilize a real time minimal byte error probability decoding algorithm, together with feedback from the outer decoder, in the decoder for the inner convolutional code. The performance of the proposed concatenated coding system is studied, and the improvement over conventional concatenated systems due to each additional feature is isolated.

Lee, L. N.↗

Automatic specification of reliability models for fault-tolerant computers

The calculation of reliability measures using Markov models is required for life-critical processor-memory-switch structures that have standby redundancy or that are subject to transient or intermittent faults or repair. The task of specifying these models is tedious and prone to human error because of the large number of states and transitions required in any reasonable system. Therefore, model specification is a major analysis bottleneck, and model verification is a major validation problem. The general unfamiliarity of computer architects with Markov modeling techniques further increases the necessity of automating the model specification. Automation requires a general system description language (SDL). For practicality, this SDL should also provide a high level of abstraction and be easy to learn and use. The first attempt to define and implement an SDL with those characteristics is presented. A program named Automated Reliability Modeling (ARM) was constructed as a research vehicle. The ARM program uses a graphical interface as its SDL, and it outputs a Markov reliability model specification formulated for direct use by programs that generate and evaluate the model.

Liceaga, Carlos A.↗

High energy density picoliter-scale zinc-air microbatteries for colloidal robotics

The recent interest in microscopic autonomous systems, including microrobots, colloidal state machines, and smart dust, has created a need for microscale energy storage and harvesting. However, macroscopic materials for energy storage have noted incompatibilities with microfabrication techniques, creating substantial challenges to realizing microscale energy systems. Here, we photolithographically patterned a microscale zinc/platinum/SU-8 system to generate the highest energy density microbattery at the picoliter (10 −12 liter) scale. The device scavenges ambient or solution-dissolved oxygen for a zinc oxidation reaction, achieving an energy density ranging from 760 to 1070 watt-hours per liter at scales below 100 micrometers lateral and 2 micrometers thickness in size. The parallel nature of photolithography processes allows 10,000 devices per wafer to be released into solution as colloids with energy stored on board. Within a volume of only 2 picoliters each, these primary microbatteries can deliver open circuit voltages of 1.05 ± 0.12 volts, with total energies ranging from 5.5 ± 0.3 to 7.7 ± 1.0 microjoules and a maximum power near 2.7 nanowatts. We demonstrated that such systems can reliably power a micrometer-sized memristor circuit, providing access to nonvolatile memory. We also cycled power to drive the reversible bending of microscale bimorph actuators at 0.05 hertz for mechanical functions of colloidal robots. Additional capabilities, such as powering two distinct nanosensor types and a clock circuit, were also demonstrated. The high energy density, low volume, and simple configuration promise the mass fabrication and adoption of such picoliter zinc-air batteries for micrometer-scale, colloidal robotics with autonomous functions.

Robotics↗

LLaMP v0.1.0

Reducing hallucination of Large Language Models (LLMs) is imperative for use in the sciences, where reliability and reproducibility are crucial. However, LLMs inherently lack long-term memory, making it a nontrivial, ad hoc, and inevitably biased task to fine-tune them on domain-specific literature and data. LLaMP is a multimodal retrieval-augmented generation (RAG) framework of hierarchical reasoning and acting (ReAct) agents that can dynamically and recursively interact with Materials Project to ground large language models on high-fidelity materials informatics.

Riebesell, Janosh [Lawrence Berkeley National Labo↗

Neural correspondence to spectrum of environmental uncertainty in multiple-cue probability judgment system with time delay

Despite state-of-the-art technologies like artificial intelligence, human judgment is critically essential in cooperative systems, such as the multi-agent system (MAS), which collect information among agents based on multiple-cue judgment. Human agents can prevent impaired situational awareness of automated agents by confirming situations under environmental uncertainty. System error caused by uncertainty can result in an unreliable system environment, and this environment affects the human agent, resulting in non-optimal decision-making in MAS. Thus, it is necessary to know how human behavior is changed to capture system reliability under uncertainty. Another issue affecting MAS is time delay, which can delay agent information transfer, resulting in low performance and instability. However, it is difficult to find studies on the influence of time delay on human agents. This study is about understanding the human decision-making process under a specific system reliability environment by uncertainty with time delay. We used concepts of expected and unexpected uncertainty to implement reliability of the system usage environment with three types of time delay conditions: no time delay, regular time delay, and irregular time delay conditions. We used electroencephalogram (EEG) for human cognitive neural mechanisms in multiple-cue judgment systems to understand human decision-making. In the reliability of system usage environment, the unreliable system environment significantly creates less memory load by less utilization of system rules for decision-making. In terms of time delay, delayed information delivery does not significantly affect memory load for decision-making.

cognitive process↗

Feasibility study: Replacement of the inoperative decommutating buffer subsystem for the instrumentation checkout complex in the Quality and Reliability Assurance Laboratory

A general purpose computer system, that is necessary for replacement of the present inoperative signal decommutator special purpose computer subsystem is described. The present decommutator subsystem has a very poor history of reliability and since April 1970, it has become inoperative because the core memory cannot be repaired. Functions of the present signal, decommutator subsystem are to receive, demultiplex, record in real time, playback in real time, and output to the SDS-930 control computer for analysis of the telemetry data. Recommendations for replacement of the inoperative telemetry decommutator subsystem are for the purchase of a mini-computer.

Hilliard, J. W.↗

A bubble detection circuit for spacecraft applications

Low power dissipation, compactness, and reliability are attributes of a bubble detection innovation for NASA's Bubble Memory Module. Heart of the design is a detector selection multiplexer which connects a chip in an addressed cell to one of two sense hybrids. By using diode selection, expensive, bulky, and high power components such as those used for sensing are shared between cells. Eight sense amplifiers serve 64 chips.

Bohning, O. D.↗

Low-Storage, Explicit Runge-Kutta Schemes for the Compressible Navier-Stokes Equations

The derivation of storage explicit Runge-Kutta (ERK) schemes has been performed in the context of integrating the compressible Navier-Stokes equations via direct numerical simulation. Optimization of ERK methods is done across the broad range of properties, such as stability and accuracy efficiency, linear and nonlinear stability, error control reliability, step change stability, and dissipation/dispersion accuracy, subject to varying degrees of memory economization. Following van der Houwen and Wray, 16 ERK pairs are presented using from two to five registers of memory per equation, per grid point and having accuracies from third- to fifth-order. Methods have been assessed using the differential equation testing code DETEST, and with the 1D wave equation. Two of the methods have been applied to the DNS of a compressible jet as well as methane-air and hydrogen-air flames. Derived 3(2) and 4(3) pairs are competitive with existing full-storage methods. Although a substantial efficiency penalty accompanies use of two- and three-register, fifth-order methods, the best contemporary full-storage methods can be pearl), matched while still saving two to three registers of memory.

Kennedy, Chistopher A.↗

Inadvertently programmed bits in Samsung 128 Mbit flash devices: a flaky investigation

JPL's X2000 avionics design pioneers new territory by specifying a non-volatile memory (NVM) board based on flash memories. The Samsung 128Mb device chosen was found to demonstrate bit errors (mostly program disturbs) and block-erase failures that increase with cycling. Low temperature, certain pseudo- random patterns, and, probably, higher bias increase the observable bit errors. An experiment was conducted to determine the wearout dependence of the bit errors to 100k cycles at cold temperature using flight-lot devices (some pre-irradiated). The results show an exponential growth rate, a wide part-to-part variation, and some annealing behavior.

flash memory program disturb wearout erase fail re↗

On the feasibility of a spaceborne fault-tolerant hypercube

The feasibility of implementing a fault-tolerant hypercube architecture for space applications is discussed. Node-level architectures and designs are considered and a first-order reliability model is presented. It is shown how error recovery can be implemented using program rollback or roll-forward techniques. Shared memory augmentations to the message-passing structure can be used to get around the inefficiencies of multicomputers to provide efficient use of hardware to achieve the needed reliabilities while maintaining performance.

Rennels, David A.↗

Concatenated coding systems employing a unit-memory convolutional code and a byte-oriented decoding algorithm

Concatenated coding systems utilizing a convolutional code as the inner code and a Reed-Solomon code as the outer code are considered. In order to obtain very reliable communications over a very noisy channel with relatively modest coding complexity, it is proposed to concatenate a byte-oriented unit-memory convolutional code with an RS outer code whose symbol size is one byte. It is further proposed to utilize a real-time minimal-byte-error probability decoding algorithm, together with feedback from the outer decoder, in the decoder for the inner convolutional code. The performance of the proposed concatenated coding system is studied, and the improvement over conventional concatenated systems due to each additional feature is isolated.

Lee, L.-N.↗

A Study on the Impact of Temperature-Dependent Ferroelectric Switching Behavior in 3D Memory Architecture

The flourishing development of neural networks that require exponentially growing amounts of data has presented an elevated demand for memory footprint. To address this, researchers have been exploring hardware accelerators with innovative memory architectures like 3D memory. These 3D memory architectures offer enhanced storage capacity and processing capabilities, at a cost of rising on-chip temperature during operation. Hafnium Zirconium Oxide (HZO) based Ferroelectric Random Access Memory (FeRAM) is a promising nonvolatile memory candidate in neural network hardware accelerators for its outstanding write performance and reliability. However, its implementation in the architecture regarding the temperature-dependent ferroelectric switching behavior has not been well studied. In this work, we study the thermal impacts on polarization switching through experimental devices and simulation results. We conduct the circuit and architecture-level simulations to showcase that one can exploit this temperature rise to reduce FeRAM's write voltage and write energy due to its unique temperature-activated polarization switching mechanisms. As the on-chip temperature increases to 351K (ambient temperature at 300K) due to neural network workloads, the access energy per bit can be reduced by 27.6% when a dynamic write voltage is applied.

36 MATERIALS SCIENCE↗

FY12 End of Year Report for NEPP DDR2 Reliability

This document reports the status of the NASA Electronic Parts and Packaging (NEPP) Double Data Rate 2 (DDR2) Reliability effort for FY2012. The task expanded the focus of evaluating reliability effects targeted for device examination. FY11 work highlighted the need to test many more parts and to examine more operating conditions, in order to provide useful recommendations for NASA users of these devices. This year's efforts focused on development of test capabilities, particularly focusing on those that can be used to determine overall lot quality and identify outlier devices, and test methods that can be employed on components for flight use. Flight acceptance of components potentially includes considerable time for up-screening (though this time may not currently be used for much reliability testing). Manufacturers are much more knowledgeable about the relevant reliability mechanisms for each of their devices. We are not in a position to know what the appropriate reliability tests are for any given device, so although reliability testing could be focused for a given device, we are forced to perform a large campaign of reliability tests to identify devices with degraded reliability. With the available up-screening time for NASA parts, it is possible to run many device performance studies. This includes verification of basic datasheet characteristics. Furthermore, it is possible to perform significant pattern sensitivity studies. By doing these studies we can establish higher reliability of flight components. In order to develop these approaches, it is necessary to develop test capability that can identify reliability outliers. To do this we must test many devices to ensure outliers are in the sample, and we must develop characterization capability to measure many different parameters. For FY12 we increased capability for reliability characterization and sample size. We increased sample size this year by moving from loose devices to dual inline memory modules (DIMMs) with an approximate reduction of 20 to 50 times in terms of per device under test (DUT) cost. By increasing sample size we have improved our ability to characterize devices that may be considered reliability outliers. This report provides an update on the effort to improve DDR2 testing capability. Although focused on DDR2, the methods being used can be extended to DDR and DDR3 with relative ease.

Guertin, Steven M.↗

Fault-tolerant software for aircraft control systems

Concepts for software to implement real time aircraft control systems on a centralized digital computer were discussed. A fault tolerant software structure employing functionally redundant routines with concurrent error detection was proposed for critical control functions involving safety of flight and landing. A degraded recovery block concept was devised to allow collocation of critical and noncritical software modules within the same control structure. The additional computer resources required to implement the proposed software structure for a representative set of aircraft control functions were discussed. It was estimated that approximately 30 percent more memory space is required to implement the total set of control functions. A reliability model for the fault tolerant software was described and parametric estimates of failure rate were made.

Source record↗