Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “limited memory”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12

Modeling of a Metal-Ferroelectric-Semiconductor Field-Effect Transistor NAND Gate

Considerable research has been performed by several organizations in the use of the Metal- Ferroelectric-Semiconductor Field-Effect Transistors (MFSFET) in memory circuits. However, research has been limited in expanding the use of the MFSFET to other electronic circuits. This research project investigates the modeling of a NAND gate constructed from MFSFETs. The NAND gate is one of the fundamental building blocks of digital electronic circuits. The first step in forming a NAND gate is to develop an inverter circuit. The inverter circuit was modeled similar to a standard CMOS inverter. A n-channel MFSFET with positive polarization was used for the n-channel transistor, and a n-channel MFSFET with negative polarization was used for the p-channel transistor. The MFSFETs were simulated by using a previously developed current model which utilized a partitioned ferroelectric layer. The inverter voltage transfer curve was obtained over a standard input of zero to five volts. Then a 2-input NAND gate was modeled similar to the inverter circuit. Voltage transfer curves were obtained for the NAND gate for various configurations of input voltages. The resultant data shows that it is feasible to construct a NAND gate with MFSFET transistors.

Phillips, Thomas A.↗

The System Complexity Metric (SCM) Explains Systems Design and is Correlated with Cost and Failure Rate

The human short term memory span and working capacity is limited to three to five items, especially if they are organized complex “chunks” of information. The impression of complexity occurs when a system is simply difficult to understand, where there is no apparent pattern to predict its behavior. Hierarchical systems design can reduce perceived complexity and increase the amount of information that can be managed. The SCM was developed to measure complexity and help compare proposed overall system architectures before detailed design information is available. The SCM is defined as the sum of the number of major nodes, N, in the system block diagram plus the number of one-way interactions, I, between the nodes. SCM = N + I. SCM’s are easily determined by direct inspection of high-level block diagrams of life support systems. Axiomatic design develops a hierarchy of subsystem requirements and designs together in a top-down, back-and-forth process. A coupling matrix is used to control the relationships between the subsystem functions and design concepts. Axiomatic design can improve system design by decoupling requirements and designs. Axiomatic design was applied to the planning of a closed life support system, similar to that used on the International Space Station. A materially open as opposed to a closed system design was created by removing the interconnections required to close the system. The open system had the same number of designed subsystems as the closed system, but it had many fewer interconnections and its SCM was lower by about half. The costs were estimated and the MTBF (Mean Time Before Failure) tabulated for open and closed space life support systems. The estimated costs were linearly proportional to SCM for the wide variations of SCM in life support, but small differences may not be significant. The flight and preflight MTBF’s both declined exponentially with increasing MTBF, faster than MTBF-2, even though the preflight estimated MTBF’s were about ten times higher than the flight MTBF’s.

System Complexity Metric (SCM)↗

Pushing the limits of NAND technology scaling with ferroelectrics

Artificial intelligence (AI) continues to drive transformative advancements across various industries. The data-intensive nature of AI training (and inferencing) has resulted in the generation of unprecedented volumes of data with machine-generated content surpassing human-generated data by more than 100-fold in 2025. Efficiently managing this data influx necessitates advanced digital storage technologies. However, traditional NAND flash memory, which is critical for supporting data flows in AI systems—alongside high-bandwidth memory, for AI training—faces fundamental scaling limitations as it approaches the 1000-layer milestone, encompassing more than 40 trillion transistors. This article delves into the potential of hafnia-based ferroelectric materials as a breakthrough solution to these challenges. Recent advancements indicate that the intrinsic limitations of ferroelectric field-effect transistors (FEFETs) can be mitigated through material and device-level engineering. These advancements enable FEFETs to meet the stringent density, reliability, and scalability requirements of future three-dimensional NAND technology. The role of ferroelectrics in addressing NAND scaling challenges and expanding storage capabilities presents a promising avenue for meeting the storage demands of the AI-driven era.

3D NAND↗

Reduced scaling extended multi-state CASPT2 (XMS-CASPT2) using supporting subspaces and tensor hyper-contraction

We present a reduced scaling formulation of the extended multi-state CASPT2 (XMS-CASPT2) method, which is based on our recently developed state-specific CASPT2 (SS-CASPT2) formulation using supporting subspaces and tensor hyper-contraction. By using these two techniques, the off-diagonal elements of the effective Hamiltonian can be computed with only O(N 3 ) operations and O(N 2 ) memory, where N is the number of basis functions. Furthermore, this limits the overall computational scaling to O(N 4 ) operations and O(N 2 ) memory. Thus, excited states can now be obtained at the same reduced (relative to previous algorithms) scaling we achieved for SS-CASPT2.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Batch Scheduling a Fresh Approach

The Network Queueing System (NQS) was designed to schedule jobs based on limits within queues. As systems obtain more memory, the number of queues increased to take advantage of the added memory resource. The problem now becomes too many queues. Having a large number of queues provides users with the capability to gain an unfair advantage over other users by tailoring their job to fit in an empty queue. Additionally, the large number of queues becomes confusing to the user community. The High Speed Processors group at the Numerical Aerodynamics Simulation (NAS) Facility at NASA Ames Research Center developed a new approach to batch job scheduling. This new method reduces the number of queues required by eliminating the need for queues based on resource limits. The scheduler examines each request for necessary resources before initiating the job. Also additional user limits at the complex level were added to provide a fairness to all users. Additional tools which include user job reordering are under development to work with the new scheduler. This paper discusses the objectives, design and implementation results of this new scheduler

Cardo, Nicholas P.↗

Scalable In Situ Computation of Lagrangian Representations via Local Flow Maps

In situ computation of Lagrangian flow maps to enable post hoc time-varying vector field analysis has recently become an active area of research. However, the current literature is largely limited to theoretical settings and lacks a solution to address scalability of the technique in distributed memory. To improve scalability, we propose and evaluate the benefits and limitations of a simple, yet novel, performance optimization. Our proposed optimization is a communication-free model resulting in local Lagrangian flow maps, requiring no message passing or synchronization between processes, intrinsically improving scalability, and thereby reducing overall execution time and alleviating the encumbrance placed on simulation codes from communication overheads. To evaluate our approach, we computed Lagrangian flow maps for four time-varying simulation vector fields and investigated how execution time and reconstruction accuracy are impacted by the number of GPUs per compute node, the total number of compute nodes, particles per rank, and storage intervals. Our study consisted of experiments computing Lagrangian flow maps with up to 67M particle trajectories over 500 cycles and used as many as 2048 GPUs across 512 compute nodes. In all, our study contributes an evaluation of a communication-free model as well as a scalability study of computing distributed Lagrangian flow maps at scale using in situ infrastructure on a modern supercomputer.

Sane, Sudhanshu↗

Spin–Phonon Coupling in Ferromagnetic Monolayer Chromium Tribromide

Novel 2D magnets exhibit intrinsic electrically tunable magnetism down to the monolayer limit, which has significant value for nonvolatile memory and emerging computing device applications. In these compounds, spin–phonon coupling (SPC) typically plays a crucial role in magnetic fluctuations, magnon dissipation, and ultimately establishing long-range ferromagnetic order. However, a systematic understanding of SPC in 2D magnets that combines theory and experiment is still lacking. Here in this work, monolayer chromium tribromide is studied to investigate SPC in 2D magnets via Raman spectroscopy and first principle calculations. The experimental Curie temperature and phonon shifts are found to be in good agreement with the numerical simulations. Specifically, it is demonstrated how magnetic exchange interactions affect phonon vibrations, which helps establish design fundamentals for 2D magnetic materials and other related devices.

36 MATERIALS SCIENCE↗

Switching of Hybrid Improper Ferroelectricity in Oxide Double Perovskites

In ABO 3 -type perovskite oxides with Pnma symmetry, rotation (Q R+ , a 0 a 0 c + ) and tilt (Q T , a – a – c 0 ) of BO 6 octahedra are the two primary order parameters. These order parameters establish an inherent trilinear coupling with anti-ferroelectric A-site displacement (Q AFE ) to form the low-symmetry phase. The symmetry is further lowered in double perovskite oxides (DPOs) due to A/A' cation ordering. It in turn makes these systems polar via hybrid improper ferroelectric mechanism, primarily driven by Q R+ and Q T . Naturally, it has been believed that functionalities such as polarization can also be switched by tuning these primary order parameters. However, mystery around finding switching mechanism still remains. Our study based on density functional theory calculations combined with finite-temperature molecular dynamics simulations shows that the polarization switching is a two-step process, driven by out-of-phase rotation (Q R– , a 0 a 0 c – when Q T = 0 or, a – a – b – when Q T ≠ 0). A series of polar DPOs such as KLnFeOsO 6 [Ln = Sm, Gd, Dy, Tm (lanthanides) and Y (rare earth)], all belonging to P2 1 symmetry, are considered in this investigation. The polarization switching P ($\overrightarrow{P}$) occurs at a very high temperature of ~1150 K through a phase transition, from a polar (P2 1 ) phase with $\overrightarrow{P}$(+) to $\overrightarrow{P}$(-) via a non-polar P4/n phase. The switching itself is metastable in nature. The switching (both polarization and spin state) is only observed for a very short period of time (~23 ps) that poses limitation on using such a mechanism in memory device realization. We demonstrate a concurrent heating–cooling procedure to overcome such shortcoming. In conclusion, simulations conducted at 600 K further imply that long lasting switching can be achieved, at least for 1.2 ns for 600 K, and ideally for an infinite time, if the material is heated just above the T c followed by rapid cooling to a temperature below T c .

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Photophysics of O-band and transition metal color centers in monolithic silicon for quantum communications

Color centers in the O-band (1260–1360 nm) are crucial for realizing long-coherence quantum network nodes in memory-assisted quantum communications. However, only a limited number of O-band color centers have been thoroughly explored in silicon hosts as spin-photon interfaces. This study explores and compares two promising O-band color centers in silicon for high-fidelity spin-photon interfaces: T and *Cu (transition metal) centers. During T center generation process, we observed the formation and dissolution of other color centers, including the copper-silver related centers with a doublet line around 1312 nm (*Cu$^0_n$), near the optical fiber zero dispersion wavelength (around 1310 nm). We then investigated the photophysics of both T and *Cu centers, focusing on their emission spectra and spin properties. The *Cu$^0_0$ line under a 0.5 T magnetic field demonstrated a 25% broadening, potentially due to spin degeneracy, suggesting that this center can be a promising alternative to T centers.

42 ENGINEERING↗

HPDR: High-Performance Portable Scientific Data Reduction Framework

The rapid growth in scientific data generation is outpacing advancements in computing systems necessary for efficient storage, transfer, and analysis, particularly in the context of exascale computing. With the deployment of first-generation exascale computing systems and next-generation experimental facilities, this gap is widening and necessitates effective data reduction techniques to manage enormous data volumes. Over the past decade, various data reduction methods, including lossless compression, error-controlled lossy compression, and data refactoring, have been developed to accelerate I/O in scientific workflows. Despite significant reductions in data volume, these methods introduce considerable computational overhead, which can become the new bottleneck in data processing. To mitigate this, GPU-accelerated data reduction algorithms have been introduced. However, challenges remain in their integration into exascale workflows, including limited portability across different GPU architectures, substantial memory transfer overhead, and reduced scalability on dense multi-GPU systems. To address these challenges, we propose HPDR, a high-performance and portable data reduction framework. HPDR is designed to enable the execution of state-of-the-art reduction algorithms across diverse processor architectures while reducing memory transfer overhead to 2.3 % of the original, resulting in up to 3.5× faster throughput compared to existing solutions. It also achieves up to 96% of the theoretical speedup in multi-GPU settings. In addition, evaluations on accelerating I/O operations at scale up to 1,024 nodes of the Frontier supercomputer demonstrate that HPDR can achieve up to 103 TB/s reduction throughput, providing up to 4× acceleration in parallel I/O performance compared to existing data reduction routines. This work highlights the potential of HPDR to significantly enhance data reduction efficiency in exascale computing environments.

Chen, Jieyang [University of Oregon]↗

Integrating Deep Learning and Hydrodynamic Modeling to Improve the Great Lakes Forecast

The Laurentian Great Lakes, one of the world’s largest surface freshwater systems, pose a modeling challenge in seasonal forecast and climate projection. While physics-based hydrodynamic modeling is a fundamental approach, improving the forecast accuracy remains critical. In recent years, machine learning (ML) has quickly emerged in geoscience applications, but its application to the Great Lakes hydrodynamic prediction is still in its early stages. This work is the first one to explore a deep learning approach to predicting spatiotemporal distributions of the lake surface temperature (LST) in the Great Lakes. Our study shows that the Long Short-Term Memory (LSTM) neural network, trained with the limited data from hypothetical monitoring networks, can provide consistent and robust performance. The LSTM prediction captured the LST spatiotemporal variabilities across the five Great Lakes well, suggesting an effective and efficient way for monitoring network design in assisting the ML-based forecast. Furthermore, we employed an explainable artificial intelligence (XAI) technique named SHapley Additive exPlanations (SHAP) to uncover how the features impact the LSTM prediction. Our XAI analysis shows air temperature is the most influential feature for predicting LST in the trained LSTM. The relatively large bias in the LSTM prediction during the spring and fall was associated with substantial heterogeneity of air temperature during the two seasons. In contrast, the physics-based hydrodynamic model performed better in spring and fall yet exhibited relatively large biases during the summer stratification period. Finally, we developed a statistical integration of the hydrodynamic modeling and deep learning results based on the Best Linear Unbiased Estimator (BLUE). The integration further enhanced prediction accuracy, suggesting its potential for next-generation Great Lakes forecast systems.

Xue, Pengfei (ORCID:000000025702421X)↗

Some recent advances in computational aerodynamics for helicopter applications

The growing application of computational aerodynamics to nonlinear helicopter problems is outlined, with particular emphasis on several recent quasi-two-dimensional examples that used the thin-layer Navier-Stokes equations and an eddy-viscosity model to approximate turbulence. Rotor blade section characteristics can now be calculated accurately over a wide range of transonic flow conditions. However, a finite-difference simulation of the complete flow field about a helicopter in forward flight is not currently feasible, despite the impressive progress that is being made in both two and three dimensions. The principal limitations are today's computer speeds and memories, algorithm and solution methods, grid generation, vortex modeling, structural and aerodynamic coupling, and a shortage of engineers who are skilled in both computational fluid dynamics and helicopter aerodynamics and dynamics.

Mccroskey, W. J.↗

Semiautomatic Design Of Zonal Computational Grids

EZGrid is knowledge-based computer program semiautomatically generating zonal computational grids for use in numerical simulations of two-dimensional flows. Zoning necessary because of limitations imposed by size of available computer memory and by topological complexity of typical flow field. Complexity and amount of required memory reduced by dividing flow field into zones, within each of which computational grid refined only to extent necessary to resolve local high gradients. Developed to speed and systematize zoning.

Vogel, Alison Andrews↗

A principled approach to the measurement of situation awareness in commercial aviation

The issue of how to support situation awareness among crews of modern commercial aircraft is becoming especially important with the introduction of automation in the form of sophisticated flight management computers and expert systems designed to assist the crew. In this paper, cognitive theories are discussed that have relevance for the definition and measurement of situation awareness. These theories suggest that comprehension of the flow of events is an active process that is limited by the modularity of attention and memory constraints, but can be enhanced by expert knowledge and strategies. Three implications of this perspective for assessing and improving situation awareness are considered: (1) Scenario variations are proposed that tax awareness by placing demands on attention; (2) Experimental tasks and probes are described for assessing the cognitive processes that underlie situation awareness; and (3) The use of computer-based human performance models to augment the measures of situation awareness derived from performance data is explored. Finally, two potential example applications of the proposed assessment techniques are described, one concerning spatial awareness using wide field of view displays and the other emphasizing fault management in aircraft systems.

Tenney, Yvette J.↗

Radiation Effects on Current Field Programmable Technologies

Manufacturers of field programmable gate arrays (FPGAS) take different technological and architectural approaches that directly affect radiation performance. Similar y technological and architectural features are used in related technologies such as programmable substrates and quick-turn application specific integrated circuits (ASICs). After analyzing current technologies and architectures and their radiation-effects implications, this paper includes extensive test data quantifying various devices total dose and single event susceptibilities, including performance degradation effects and temporary or permanent re-configuration faults. Test results will concentrate on recent technologies being used in space flight electronic systems and those being developed for use in the near term. This paper will provide the first extensive study of various configuration memories used in programmable devices. Radiation performance limits and their impacts will be discussed for each design. In addition, the interplay between device scaling, process, bias voltage, design, and architecture will be explored. Lastly, areas of ongoing research will be discussed.

Katz, R.↗

Imaging for Hypersonic Experimental Aeroheating Testing (IHEAT) Version 4.0: User Manual

The IHEAT v4.0 software is a data reduction code for global thermography data acquired in the NASA Langley Aerothermodynamics Laboratory (LAL) hypersonic wind tunnels. IHEAT uses red and green color-intensity data from two-dimensional images of wind tunnel models to compute temperatures and heat-transfer rates using a semi-infinite, one-dimensional heat transfer approximation at each image pixel. Multiple automated tools in IHEAT v4.0 decrease the time required to reduce the data from a phosphor thermography wind tunnel run. Data at one or all of the image pixel locations can be exported to computer files for further analysis. The prior version of IHEAT, v3.2, was written in PV-WAVE® (now owned by Rogue Wave® Software) in 1994 and was limited in functionality to fit within the memory constraints of the available computers at the time. IHEAT v4.0 is written in MATLAB® by MathWorks® and contains several new features that leverage the increase in available memory of the current computers. A Piecewise tool permits the user to extract data along a segmented line cut that can follow interesting features in the image better than the single, straight line cuts that were possible with the legacy Length and Profile tools. The new Load Run and Batch tools facilitate batch processing by loading in all of the input files and images for a run at the same time. Load Run permits the user to process the available run images manually, while Batch automatically saves heat transfer data from all of the images based on the analysis previously performed on a single frame. IHEAT v4.0 also can automatically calculate the temporal collapse of reference line cuts from the time history heating data for a run to indicate the appropriate frame to reduce for each run. The IHEAT v4.0 source code was compiled into a standalone executable file that can be accessed remotely from several computers with different operating systems, simultaneously. The software is run through the MATLAB® Compiler Runtime engine, and therefore, IHEAT does not require a software license to run. Any software commands executed in the IHEAT v4.0 code will not affect other similar applications running on the same machine. Similarly, changes to the parent software do not affect a compiled code. These features of IHEAT v4.0 are improvements over the legacy v3.2 code, which required regular maintenance to avoid losing functionality as the PVWAVE ® programming language was upgraded.

Mason, Michelle L.↗

Performance Evaluation of Different Parallel Programming Models in SCALE-Shift Sequences for Criticality and Shielding Applications [Abstract]

The SCALE code system has been widely used for nuclear criticality safety, reactor physics, radiation shielding, source term generation, and inventory analyses by researchers, industry, and regulatory bodies. Although limited support for shared- and distributed-memory parallel processing was introduced via C++ threading, OpenMP, and MPI, a hybrid parallel programming model with both distributed- and shared-memory parallelism has not been fully supported in the SCALE code system.

Nuclear Criticality Safety Program (NCSP)↗

Timely Reporting of Heavy Hitters Using External Memory

Given an input stream S of size N, a Φ-heavy hitter is an item that occurs at least ΦN times in S. The problem of finding heavy-hitters is extensively studied in the database literature. In this work, we study a real-time heavy-hitters variant in which an element must be reported shortly after we see its T = Φ N-th occurrence (and hence it becomes a heavy hitter). We call this the Timely Event Detection (TED) Problem. The TED problem models the needs of many real-world monitoring systems, which demand accurate (i.e., no false negatives) and timely reporting of all events from large, high-speed streams with a low reporting threshold (high sensitivity). Like the classic heavy-hitters problem, solving the TED problem without false-positives requires large space (Ω (N) words). Thus in-RAM heavy-hitters algorithms typically sacrifice accuracy (i.e., allow false positives), sensitivity, or timeliness (i.e., use multiple passes). We show how to adapt heavy-hitters algorithms to external memory to solve the TED problem on large high-speed streams while guaranteeing accuracy, sensitivity, and timeliness. Our data structures are limited only by I/O-bandwidth (not latency) and support a tunable tradeoff between reporting delay and I/O overhead. With a small bounded reporting delay, our algorithms incur only a logarithmic I/O overhead. We implement and validate our data structures empirically using the Firehose streaming benchmark. Multi-threaded versions of our structures can scale to process 11M observations per second before becoming CPU bound. In comparison, a naive adaptation of the standard heavy-hitters algorithm to external memory would be limited by the storage device’s random I/O throughput, i.e., ≈100K observations per second.

97 MATHEMATICS AND COMPUTING↗