Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “limited memory”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 271 records · Page 15

Designing shape-memory-like microstructures in intercalation materials

During the reversible insertion of ions, lattices in intercalation materials undergo structural transformations. These lattice transformations generate misfit strains and volume changes that, in turn, contribute to the structural decay of intercalation materials and limit their reversible cycling. In this paper, we draw on insights from shape-memory alloys, another class of phase transformation materials, that also undergo large lattice transformations but do so with negligible macroscopic volume changes and internal stresses. We develop a theoretical framework to predict structural transformations in intercalation compounds and identify crystallographic design rules necessary for forming shape-memory-like microstructures in intercalation materials. We use our approach to systematically screen open-source structural databases comprising $n$ > 5000 pairs of intercalation compounds. We identify candidate compounds, such as Li x Mn 2 O 4 (Spinel), Li x Ti 2 (PO 4 ) 3 (NASICON), that approximately satisfy the crystallographic design rules and can be precisely doped to form shape-memory-like microstructures. Throughout, we compare our analytical results with experimental measurements of intercalation compounds. We find a direct correlation between structural transformations, microstructures, and increased capacity retention in these materials. These results, more generally, show that crystallographic designing of intercalation materials could be a novel route to discovering compounds that do not decay with continuous usage.

36 MATERIALS SCIENCE↗

Resource-aware compression

Systems, apparatuses, and methods for implementing a multi-tiered approach to cache compression are disclosed. A cache includes a cache controller, light compressor, and heavy compressor. The decision on which compressor to use for compressing cache lines is made based on certain resource availability such as cache capacity or memory bandwidth. This allows the cache to opportunistically use complex algorithms for compression while limiting the adverse effects of high decompression latency on system performance. To address the above issue, the proposed design takes advantage of the heavy compressors for effectively reducing memory bandwidth in high bandwidth memory (HBM) interfaces as long as they do not sacrifice system performance. Accordingly, the cache combines light and heavy compressors with a decision-making unit to achieve reduced off-chip memory traffic without sacrificing system performance.

97 MATHEMATICS AND COMPUTING↗

Characterization of Ternary NiTiPt High-Temperature Shape Memory Alloys

Pt additions substituted for Ni in NiTi alloys are known to increase the transformation temperature of the alloy but only at fairly high Pt levels. However, until now only ternary compositions with a very specific stoichiometry, Ni50-xPtxTi50, have been investigated and then only to very limited extent. In order to learn about this potential high-temperature shape memory alloy system, a series of over twenty alloys along and on either side of a line of constant stoichiometry between NiTi and TiPt were arc melted, homogenized, and characterized in terms of their microstructure, transformation temperatures, and hardness. The resulting microstructures were examined by scanning electron microscopy and the phase compositions quantified by energy dispersive spectroscopy."Stoichiometric" compositions along a line of constant stoichiometry between NiTi to TiPt were essentially single phase but by any deviations from a stoichiometry of (Ni,Pt)50Ti50 resulted in the presence of at least two different intermetallic phases, depending on the overall composition of the alloy. Essentially all alloys, whether single or two-phase, still under went a martensitic transformation. It was found that the transformation temperatures were depressed with initial Pt additions but at levels greater than 10 at.% the transformation temperature increased linearly with Pt content. Also, the transformation temperatures were relatively insensitive to alloy stoichiometry within the range of alloys examined. Finally, the dependence of hardness on Pt content for a series of Ni50-xPtxTi50 alloys showed solution softening at low Pt levels, while hardening was observed in ternary alloys containing more than about 10 at.% Pt. On either side of these "stoichiometric" compositions, hardness was also found to increase significantly.

Rios, Orlando↗

Ultra-High-Density Ferroelectric Memories

Features include fast input and output via optical fibers. Memory devices of proposed type include thin ferroelectric films in which data stored in form of electric polarization. Assuming one datum stored in region as small as polarization domain, sizes of such domains impose upper limits on achievable storage densities. Limits approach 1 terabit/cm(Sup2) in all-optical versions of these ferroelectric memories and exceeds 1 gigabit/cm(Sup2) in optoelectronic versions. Memories expected to exhibit operational lives of about 10 years, input/output times of about 10 ns, and fatigue lives of about 10(Sup13) cycles.

Thakoor, Sarita↗

Benchmarking Memory Performance with the Data Cube Operator

Data movement across a computer memory hierarchy and across computational grids is known to be a limiting factor for applications processing large data sets. We use the Data Cube Operator on an Arithmetic Data Set, called ADC, to benchmark capabilities of computers and of computational grids to handle large distributed data sets. We present a prototype implementation of a parallel algorithm for computation of the operatol: The algorithm follows a known approach for computing views from the smallest parent. The ADC stresses all levels of grid memory and storage by producing some of 2d views of an Arithmetic Data Set of d-tuples described by a small number of integers. We control data intensity of the ADC by selecting the tuple parameters, the sizes of the views, and the number of realized views. Benchmarking results of memory performance of a number of computer architectures and of a small computational grid are presented.

Frumkin, Michael A.↗

LPBF Processability of NiTiHf Alloys: Systematic Modeling and Single-Track Studies

Research into the processability of NiTiHf high-temperature shape memory alloys (HTSMAs) via laser powder bed fusion (LPBF) is limited; nevertheless, these alloys show promise for applications in extreme environments. This study aims to address this limitation by investigating the printability of four NiTiHf alloys with varying Hf content (1, 2, 15, and 20 at. %) to assess their suitability for LPBF applications. Solidification cracking is one of the main limiting factors in LPBF processes, which occurs during the final stage of solidification. To investigate the effect of alloy composition on printability, this study focuses on this defect via a combination of computational modeling and experimental validation. To this end, solidification cracking susceptibility is calculated as Kou’s index and Scheil–Gulliver model, implemented in Thermo-Calc/2022a software. An innovative powder-free experimental method through laser remelting was conducted on bare NiTiHf ingots to validate the parameter impacts of the LPBF process. The result is the processability window with no cracking likelihood under diverse LPBF conditions, including laser power and scan speed. This comprehensive investigation enhances our understanding of the processability challenges and opportunities for NiTiHf HTSMAs in advanced engineering applications.

36 MATERIALS SCIENCE↗

Determining the operating characteristics of an ultraviolet interferometric spectrometer

A prototype interferometric spectrometer system is being built by NASA to explore the potential of the technique for applications involving the visible and near ultraviolet portions of the electromagnetic spectrum. The system is limited only by the frequency bandpass of the optical components used in the system, the quality of the optical components, and ultimately by the memory capacity of the computer; tradeoffs between the wavenumber resolution of the produced spectrum, the bandpass limits of the optics, and the number of samples obtained from the interferogram must be delineated explicitly. The prototype Ultraviolet Interferometric Spectrometer (UVIS) instrument is expected to be configured several different ways to ascertain its suitability for various applications. To exploit its inherent flexibility, this reference document describes these parameter tradeoffs.

Parsons, C. L.↗

Performance of an Optimized Eta Model Code on the Cray T3E and a Network of PCs

In the year 2001, NASA will launch the satellite TRIANA that will be the first Earth observing mission to provide a continuous, full disk view of the sunlit Earth. As a part of the HPCC Program at NASA GSFC, we have started a project whose objectives are to develop and implement a 3D cloud data assimilation system, by combining TRIANA measurements with model simulation, and to produce accurate statistics of global cloud coverage as an important element of the Earth's climate. For simulation of the atmosphere within this project we are using the NCEP/NOAA operational Eta model. In order to compare TRIANA and the Eta model data on approximately the same grid without significant downscaling, the Eta model will be integrated at a resolution of about 15 km. The integration domain (from -70 to +70 deg in latitude and 150 deg in longitude) will cover most of the sunlit Earth disc and will continuously rotate around the globe following TRIANA. The cloud data assimilation is supposed to run and produce 3D clouds on a near real-time basis. Such a numerical setup and integration design is very ambitious and computationally demanding. Thus, though the Eta model code has been very carefully developed and its computational efficiency has been systematically polished during the years of operational implementation at NCEP, the current MPI version may still have problems with memory and efficiency for the TRIANA simulations. Within this work, we optimize a parallel version of the Eta model code on a Cray T3E and a network of PCs (theHIVE) in order to improve its overall efficiency. Our optimization procedure consists of introducing dynamically allocated arrays to reduce the size of static memory, and optimizing on a single processor by splitting loops to limit the number of streams. All the presented results are derived using an integration domain centered at the equator, with a size of 60 x 60 deg, and with horizontal resolutions of 1/2 and 1/3 deg, respectively. In accompanying charts we report the elapsed time, the speedup and the Mflops as a function of the number of processors for the non-optimized version of the code on the T3E and theHIVE. The large amount of communication required for model integration explains its poor performance on theHIVE. Our initial implementation of the dynamic memory allocation has contributed to about 12% reduction of memory but has introduced a 3% overhead in computing time. This overhead was removed by performing loop splitting in some of the high demanding subroutines. When the Eta code is fully optimized in order to meet the memory requirement for TRIANA simulations, a non-negligeable overhead may appear that may seriously affect the efficiency of the code. To alleviate this problem, we are considering implementation of a new algorithm for the horizontal advection that is computationally less expensive, and also a new approach for marching in time.

Kouatchou, Jules↗

Adaptive Scalpel Scanning Probe Microscopy for Enhanced Volumetric Sensing in Tomographic Analysis

Controlling nanoscale tip‐induced material removal is crucial for achieving atomic‐level precision in tomographic sensing with atomic force microscopy (AFM). While advances have enabled volumetric probing of conductive features with nanometer accuracy in solid‐state devices, materials, and photovoltaics, limitations in spatial resolution and volumetric sensitivity persist. This work identifies and addresses in‐plane and vertical tip‐sample junction leakage as sources of parasitic contrast in tomographic AFM, hindering real‐space 3D reconstructions. Novel strategies are proposed to overcome these limitations. First, the contrast mechanisms analyzing nanosized conductive features are explored when confining current collection purely to in‐plane transport, thus allowing reconstruction with a reduction in the overestimation of the lateral dimensions. Furthermore, an adaptive tip‐sample biasing scheme is demonstrated for the mitigation of a class of artefacts induced by the high electric field inside the thin oxide when volumetrically reduced. This significantly enhances vertical sensitivity by approaching the intrinsic limits set by quantum tunneling processes, allowing detailed depth analysis in thin dielectrics. The effectiveness of these methods is showcased in tomographic reconstructions of conductive filaments in valence change memory, highlighting the potential for application in nanoelectronics devices and bulk materials and unlocking new limits for tomographic AFM.

36 MATERIALS SCIENCE↗

Assessing the Performance Limits of Internal Coronagraphs Through End-to-End Modeling

As part of the NASA ROSES Technology Demonstrations for Exoplanet Missions (TDEM) program, we conducted a numerical modeling study of three internal coronagraphs (PIAA, vector vortex, hybrid bandlimited) to understand their behaviors in realistically-aberrated systems with wavefront control (deformable mirrors). This investigation consisted of two milestones: (1) develop wavefront propagation codes appropriate for each coronagraph that are accurate to 1% or better (compared to a reference algorithm) but are also time and memory efficient, and (2) use these codes to determine the wavefront control limits of each architecture. We discuss here how the milestones were met and identify some of the behaviors particular to each coronagraph. The codes developed in this study are being made available for community use. We discuss here results for the HBLC and VVC systems, with PIAA having been discussed in a previous proceeding.

Inner working angle (IWA)↗

Explainable machine learning model for multi-step forecasting of reservoir inflow with uncertainty quantification

We propose an explainable machine learning (ML) model with uncertainty quantification (UQ) to improve multi-step reservoir inflow forecasting. Traditional ML methods have challenges in forecasting inflows multiple days ahead, and lack explainability and UQ. To address these limitations, we introduce an encoder–decoder long short-term memory (ED-LSTM) network for multi-step forecasting, employ the SHapley Additive exPlanation (SHAP) technique for understanding the influence of hydrometeorological factors on inflow prediction, and develop a novel UQ method for prediction trustworthiness. We apply these methods to forecast 7-day inflow in snow-dominant and rain-driven reservoirs. The results demonstrate the effectiveness of the ED-LSTM model, with high forecasting accuracy for short lead times. Our UQ method provides reliable uncertainty estimates, covering 90% of data with a 90% confidence level. The SHAP analysis reveals the importance of historical inflow and precipitation as influential factors. These findings and methods may support reservoir operators in optimizing water resources management decisions.

54 ENVIRONMENTAL SCIENCES↗

Logical Activation Functions v.1.1

SAND2024-01501O Logical Activation Functions software is a PyTorch implementation from the paper, "Logical Activation Functions for Training Arbitrary Probabilistic Boolean Logic." The activation functions approximate logit-space marginalization of probabilistic truth tables from probabilistic interpretations of inputs. They also provide a general methodology to approximate logical relationships between abstract antecedents and consequents for machine learning architectures. They do not target any specific application or use-case. By training probabilistic truth tables, these activation functions can capture more expressive relationships in a neural network than typical elementwise activation functions. This code is only designed for a single compute node with a GPU and is limited to machine learning architectures than can fit within the memory of a single GPU. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Duersch, Jed↗

Research in the design of high-performance reconfigurable systems

Computer aided design and computer aided manufacturing have the potential for greatly reducing the cost and lead time in the development of VLSI components. This potential paves the way for the design and fabrication of a wide variety of economically feasible high level functional units. It was observed that current computer systems have only a limited capacity to absorb new VLSI component types other than memory, microprocessors, and a relatively small number of other parts. The first purpose is to explore a system design which is capable of effectively incorporating a considerable number of VLSI part types and will both increase the speed of computation and reduce the attendant programming effort. A second purpose is to explore design techniques for VLSI parts which when incorporated by such a system will result in speeds and costs which are optimal. The proposed work may lay the groundwork for future efforts in the extensive simulation and measurements of the system's cost effectiveness and lead to prototype development.

Mcewan, S. D.↗

Fault injection experiments using FIAT

The results of several experiments conducted using the fault-injection-based automated testing (FIAT) system are presented. FIAT is capable of emulating a variety of distributed system architectures, and it provides the capabilities to monitor system behavior and inject faults for the purpose of experimental characterization and validation of a system's dependability. The experiments consist of exhaustively injecting three separate fault types into various locations, encompassing both the code and data portions of memory images, of two distinct applications executed with several different data values and sizes. Fault types are variations of memory bit faults. The results show that there are a limited number of system-level fault manifestations. These manifestations follow a normal distribution for each fault type. Error detection latencies are found to be normally distributed. The methodology can be used to predict the system-level fault responses during the system design stage.

Barton, James H.↗

Progress in Grid Generation: From Chimera to DRAGON Grids

Hybrid grids, composed of structured and unstructured grids, combines the best features of both. The chimera method is a major stepstone toward a hybrid grid from which the present approach is evolved. The chimera grid composes a set of overlapped structured grids which are independently generated and body-fitted, yielding a high quality grid readily accessible for efficient solution schemes. The chimera method has been shown to be efficient to generate a grid about complex geometries and has been demonstrated to deliver accurate aerodynamic prediction of complex flows. While its geometrical flexibility is attractive, interpolation of data in the overlapped regions - which in today's practice in 3D is done in a nonconservative fashion, is not. In the present paper we propose a hybrid grid scheme that maximizes the advantages of the chimera scheme and adapts the strengths of the unstructured grid while at the same time keeps its weaknesses minimal. Like the chimera method, we first divide up the physical domain by a set of structured body-fitted grids which are separately generated and overlaid throughout a complex configuration. To eliminate any pure data manipulation which does not necessarily follow governing equations, we use non-structured grids only to directly replace the region of the arbitrarily overlapped grids. This new adaptation to the chimera thinking is coined the DRAGON grid. The nonstructured grid region sandwiched between the structured grids is limited in size, resulting in only a small increase in memory and computational effort. The DRAGON method has three important advantages: (1) preserving strengths of the chimera grid; (2) eliminating difficulties sometimes encountered in the chimera scheme, such as the orphan points and bad quality of interpolation stencils; and (3) making grid communication in a fully conservative and consistent manner insofar as the governing equations are concerned. To demonstrate its use, the governing equations are discretized using the newly proposed flux scheme, AUSM+, which will be briefly described herein. Numerical tests on representative 2D inviscid flows are given for demonstration. Finally, extension to 3D is underway, only paced by the availability of the 3D unstructured grid generator.

Liou, Meng-Sing↗

A Conceptual Framework for Predicting Error in Complex Human-Machine Environments

We present a Goals, Operators, Methods, and Selection Rules-Model Human Processor (GOMS-MHP) style model-based approach to the problem of predicting human habit capture errors. Habit captures occur when the model fails to allocate limited cognitive resources to retrieve task-relevant information from memory. Lacking the unretrieved information, decision mechanisms act in accordance with implicit default assumptions, resulting in error when relied upon assumptions prove incorrect. The model helps interface designers identify situations in which such failures are especially likely.

Freed, Michael↗

Initial SEE Testing of Maestro

We have reported on initial SEE sensitivity of the full 49-core Maestro device. Supporting the low-level structures and qualitative system observation goals of phase 1 of testing. Observed sensitivities found to be consistent with Boeing predictions. Highlighted by the L1 data cache sensitivity which drives the rates on the current Maestro device. Presented details to the hardware and software setups that show where the limitations - highlighting future work. Key future work includes testing with memory and IO ports.

cache sensitivity↗

Entanglement-fidelity limits of photonically networked atomic qubits from recoil and timing

The remote entanglement of two atomic quantum memories through photonic interactions is accompanied by atomic momentum recoil. When the interactions occur at different times, such as from the random emission over the lifetime of the atomic excited state, the difference in recoil timing can expose “which-path” information and ultimately lead to decoherence. Time-bin encoded photonic qubits can be particularly sensitive to asynchronous recoil timing. In this paper we study the limits of entanglement fidelity in atomic systems due to recoil and other timing imbalances and show how these effects can be suppressed or even eliminated through proper experimental design.

Quantum communication↗