Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “limited memory”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 307 records · Page 17

NASA's Seasonal-to-Interannual Prediction Project: In Partnership With the NCCS

Researchers with NASA's Season-to-Interannual Prediction Project (NSIPP) refer to different types of memory when running models on NCCS computers: the computer memory required for their models and the memory of the atmosphere or the ocean. Because of the atmosphere's chaotic nature, its memory is short. For weather predictions, the initial information taken from atmospheric observations has a limited useful life. Currently, there is no way to take observations, initialize an atmosphere model, integrate ahead in time, and make an accurate weather forecast beyond about 2 weeks. After that, the system becomes chaotic. What conditions could be used to make predictions beyond 2 weeks? If not conditions in the atmosphere, then the memory must be found somewhere else. That place is in the oceans. Although most changes in the atmosphere vary on a short timescale, the weather being a prime example, some important large atmospheric climate variations occur over much longer timescales-month s, years, or decades. NSIPP is interested specifically in those phenomena that occur over timescales of several months to a few years, and the El Nino Southern Oscillation (ENSO) is the most significant of these.

Source record↗

Characterizing Tradeoffs in Memory, Accuracy, and Speed for Chemistry Tabulation Techniques

Chemistry tabulation is a common approach in practical simulations of turbulent combustion at engineering scales. Linear interpolants have traditionally been used for accessing precomputed multidimensional tables but suffer from large memory requirements and discontinuous derivatives. Higher-degree interpolants address some of these restrictions but are similarly limited to relatively low-dimensional tabulation. Artificial neural networks (ANNs) can be used to overcome these limitations but cannot guarantee the same accuracy as interpolants and introduce challenges in reproducibility and reliable training. These challenges are enhanced as the physics complexity to be represented within the tabulation increases. Here, we assess the efficiency, accuracy, and memory requirements of Lagrange polynomials, tensor product B-splines, and ANNs as tabulation strategies. We analyze results in the context of nonadiabatic flamelet modeling where higher dimension counts are necessary. While ANNs do not require structuring of data, providing benefits for complex physics representation, interpolation approaches often rely on some structuring of the table. Interpolation using structured table inputs that are not directly related to the variables transported in a simulation can incur additional query costs. This is demonstrated in the present implementation of heat losses. We show that ANNs, despite being difficult to train and reproduce, can be advantageous for high-dimensional, unstructured datasets relevant to nonadiabatic flamelet models. Furthermore we demonstrate that Lagrange polynomials show significant speedup for similar accuracy compared to B-splines.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Microwave-photonic control of optically accessible spin defects in diamond

A growing variety of optically accessible spin qubits have emerged in recent years as key components for quantum sensors, computers, and memories. However, the scalability of conventional spin-based quantum architectures remains limited by direct microwave delivery, which introduces thermal noise, electromagnetic crosstalk, and design constraints for cryogenic, high-field, and distributed systems. In this work, we present a unified framework for RF-over-fiber (RFoF) control of spins accessible through optically detected magnetic resonance (ODMR) spectroscopy of nitrogen-vacancy (NV) centers in diamond. The RFoF platform relies on an intensity-modulated 1310 nm laser carrying microwave signals over fiber and a high-speed photodiode for optical-to-electrical conversion to drive NV spin transitions. We report an RFoF power-conversion efficiency of 3.42% for an RF output PRF,out=-5.5 dBm at 2.87 GHz, enabling clear resolution of Zeeman splitting in

Rahman, M Reefaz [ORNL] (ORCID:0000000323541611)↗

A PC-based hardware implementation of the maximum-likelihood classifier for the Shuttle Ice Detection System

A PC-based near-real time implementation of a two-channel maximum-likelihood classifier is described. The statistical distribution of the reflectance characteristics of ice, frost, and water formation on spray-on-foam-insulation, which covers the External Tank surface of the Space Shuttle, is acquired. The classification technique is based on these statistics. The computer, set in either a training or a classifying mode, learns the statistics of the various classes, or produces a color-coded image denoting the respective categories of classification. The classified results are memory-mapped for efficiency. The speed of the classification process is only limited by the speed of the digital frame grabber and the software that interfaces the frame grabber to the monitor. The process took 4 seconds for a 512 x 480 pixel image.

Jaggi, S.↗

A simple modern correctness condition for a space-based high-performance multiprocessor

A number of U.S. national programs, including space-based detection of ballistic missile launches, envisage putting significant computing power into space. Given sufficient progress in low-power VLSI, multichip-module packaging and liquid-cooling technologies, we will see design of high-performance multiprocessors for individual satellites. In very high speed implementations, performance depends critically on tolerating large latencies in interprocessor communication; without latency tolerance, performance is limited by the vastly differing time scales in processor and data-memory modules, including interconnect times. The modern approach to tolerating remote-communication cost in scalable, shared-memory multiprocessors is to use a multithreaded architecture, and alter the semantics of shared memory slightly, at the price of forcing the programmer either to reason about program correctness in a relaxed consistency model or to agree to program in a constrained style. The literature on multiprocessor correctness conditions has become increasingly complex, and sometimes confusing, which may hinder its practical application. We propose a simple modern correctness condition for a high-performance, shared-memory multiprocessor; the correctness condition is based on a simple interface between the multiprocessor architecture and a high-performance, shared-memory multiprocessor; the correctness condition is based on a simple interface between the multiprocessor architecture and the parallel programming system.

Probst, David K.↗

The Automatic Parallelisation of Scientific Application Codes Using a Computer Aided Parallelisation Toolkit

The shared-memory programming model is a very effective way to achieve parallelism on shared memory parallel computers. Historically, the lack of a programming standard for using directives and the rather limited performance due to scalability have affected the take-up of this programming model approach. Significant progress has been made in hardware and software technologies, as a result the performance of parallel programs with compiler directives has also made improvements. The introduction of an industrial standard for shared-memory programming with directives, OpenMP, has also addressed the issue of portability. In this study, we have extended the computer aided parallelization toolkit (developed at the University of Greenwich), to automatically generate OpenMP based parallel programs with nominal user assistance. We outline the way in which loop types are categorized and how efficient OpenMP directives can be defined and placed using the in-depth interprocedural analysis that is carried out by the toolkit. We also discuss the application of the toolkit on the NAS Parallel Benchmarks and a number of real-world application codes. This work not only demonstrates the great potential of using the toolkit to quickly parallelize serial programs but also the good performance achievable on up to 300 processors for hybrid message passing and directive-based parallelizations.

Ierotheou, C.↗

Nanoporous Dielectric Resistive Memories Using Sequential Infiltration Synthesis

Resistance switching in metal–insulator–metal structures has been extensively studied in recent years for use as synaptic elements for neuromorphic computing and as nonvolatile memory elements. However, high switching power requirements, device variabilities, and considerable trade-offs between low operating voltages, high on/off ratios, and low leakage have limited their utility. Here, we have addressed these issues by demonstrating the use of ultraporous dielectrics as a pathway for high-performance resistive memory devices. Using a modified atomic layer deposition based technique known as sequential infiltration synthesis, which was developed originally for improving polymer properties such as enhanced etch resistance of electron-beam resists and for the creation of films for filtration and oleophilic applications, we are able to create ~15 nm thick ultraporous (pore size ~5 nm) oxide dielectrics with up to 73% porosity as the medium for filament formation. We show, using the Ag/Al 2 O 3 system, that the ultraporous films result in ultrahigh on/off ratio (>10 9 ) at ultralow switching voltages (~±600 mV) that are 10× smaller than those for the bulk case. In addition, the devices demonstrate fast switching, pulsed endurance up to 1 million cycles. and high temperature (125 °C) retention up to 10 4 s, making this approach highly promising for large-scale neuromorphic and memory applications. Additionally, this synthesis methodology provides a compatible, inexpensive route that is scalable and compatible with existing semiconductor nanofabrication methods and materials.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Efficient Gradient-Based Shape Optimization Methodology Using Inviscid/Viscous CFD

The formerly developed preconditioned-biconjugate-gradient (PBCG) solvers for the analysis and the sensitivity equations had resulted in very large error reductions per iteration; quadratic convergence was achieved whenever the solution entered the domain of attraction to the root. Its memory requirement was also lower as compared to a direct inversion solver. However, this memory requirement was high enough to preclude the realistic, high grid-density design of a practical 3D geometry. This limitation served as the impetus to the first-year activity (March 9, 1995 to March 8, 1996). Therefore, the major activity for this period was the development of the low-memory methodology for the discrete-sensitivity-based shape optimization. This was accomplished by solving all the resulting sets of equations using an alternating-direction-implicit (ADI) approach. The results indicated that shape optimization problems which required large numbers of grid points could be resolved with a gradient-based approach. Therefore, to better utilize the computational resources, it was recommended that a number of coarse grid cases, using the PBCG method, should initially be conducted to better define the optimization problem and the design space, and obtain an improved initial shape. Subsequently, a fine grid shape optimization, which necessitates using the ADI method, should be conducted to accurately obtain the final optimized shape. The other activity during this period was the interaction with the members of the Aerodynamic and Aeroacoustic Methods Branch of Langley Research Center during one stage of their investigation to develop an adjoint-variable sensitivity method using the viscous flow equations. This method had algorithmic similarities to the variational sensitivity methods and the control-theory approach. However, unlike the prior studies, it was considered for the three-dimensional, viscous flow equations. The major accomplishment in the second period of this project (March 9, 1996 to March 8, 1997) was the extension of the shape optimization methodology for the Thin-Layer Navier-Stokes equations. Both the Euler-based and the TLNS-based analyses compared with the analyses obtained using the CFL3D code. The sensitivities, again from both levels of the flow equations, also compared very well with the finite-differenced sensitivities. A fairly large set of shape optimization cases were conducted to study a number of issues previously not well understood. The testbed for these cases was the shaping of an arrow wing in Mach 2.4 flow. All the final shapes, obtained either from a coarse-grid-based or a fine-grid-based optimization, using either a Euler-based or a TLNS-based analysis, were all re-analyzed using a fine-grid, TLNS solution for their function evaluations. This allowed for a more fair comparison of their relative merits. From the aerodynamic performance standpoint, the fine-grid TLNS-based optimization produced the best shape, and the fine-grid Euler-based optimization produced the lowest cruise efficiency.

Baysal, Oktay↗

Potential High-Temperature Shape-Memory Alloys Identified in the Ti(Ni,Pt) System

"Shape memory" is a unique property of certain alloys that, when deformed (within certain strain limits) at low temperatures, will remember and recover to their original predeformed shape upon heating. It occurs when an alloy is deformed in the low-temperature martensitic phase and is then heated above its transformation temperature back to an austenitic state. As the material passes through this solid-state phase transformation on heating, it also recovers its original shape. This behavior is widely exploited, near room temperature, in commercially available NiTi alloys for connectors, couplings, valves, actuators, stents, and other medical and dental devices. In addition, there are limitless applications in the aerospace, automotive, chemical processing, and many other industries for materials that exhibit this type of shape-memory behavior at higher temperatures. But for high temperatures, there are currently no commercial shape-memory alloys. Although there are significant challenges to the development of high-temperature shape-memory alloys, at the NASA Glenn Research Center we have identified a series of alloy compositions in the Ti-Ni-Pt system that show great promise as potential high-temperature shape-memory materials.

Noebe, Ronald D.↗

A Massively Parallel Implementation of the CCSD(T) Method Using the Resolution-of-the-Identity Approximation and a Hybrid Distributed/Shared Memory Parallelization Model

In this work, a parallel algorithm is described for the coupled-cluster singles and doubles method augmented with a perturbative correction for triple excitations [CCSD(T)] using the resolution-of-the-identity (RI) approximation for two-electron repulsion integrals (ERIs). The algorithm bypasses the storage of four-center ERIs by adopting an integral-direct strategy. The CCSD amplitude equations are given in a compact quasi-linear form by factorizing them in terms of amplitude-dressed three-center intermediates. A hybrid MPI/OpenMP parallelization scheme is employed, which uses the OpenMP-based shared memory model for intranode parallelization and the MPI-based distributed memory model for internode parallelization. Parallel efficiency has been optimized for all terms in the CCSD amplitude equations. Two different algorithms have been implemented for the rate-limiting terms in the CCSD amplitude equations that entail and -scaling computational costs, where N O and N V denote the number of correlated occupied and virtual orbitals, respectively. One of the algorithms assembles the four-center ERIs requiring N V 4 and N O 2 N V 2 -scaling memory costs in a distributed manner on a number of MPI ranks, while the other algorithm completely bypasses the assembling of quartic memory-scaling ERIs and thus largely reduces the memory demand. It is demonstrated that the former memory-expensive algorithm is faster on a few hundred cores, while the latter memory-economic algorithm shows a better strong scaling in the limit of a few thousand cores. The program is shown to exhibit a near-linear scaling, in particular for the compute-intensive triples correction step, on up to 8000 cores. The performance of the program is demonstrated via calculations involving molecules with 24–51 atoms and up to 1624 atomic basis functions. As the first application, the complete basis set (CBS) limit for the interaction energy of the π-stacked uracil dimer from the S66 data set has been investigated. This work reports the first calculation of the interaction energy at the CCSD(T)/aug-cc-pVQZ level without local orbital approximation. The CBS limit for the CCSD correlation contribution to the interaction energy was found to be -8.01 kcal/mol, which agrees very well with the value -7.99 kcal/mol reported by Schmitz, Hättig, and Tew [ Phys. Chem. Chem. Phys. 2014 , 16 , 22167-22178]. The CBS limit for the total interaction energy was estimated to be -9.64 kcal/mol.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Synchronization for CXL Based Memory

Compute Express Link (CXL) is an important emerging standard for disaggregated memory. While this standard provisions coherency across numerous hosts and devices, implementing hardware support for type three devices is challenging. In this work, we look at the overhead of software synchronization and using software-based coherency. Moreover, we discuss the limits of software-based coherency in fully expressing modern synchronization techniques for a CXL-based disaggregate memory system. We demonstrate our approach using a CXL hardware prototype and running a version of the famous Peterson Lock (enhanced to run with more than two threads). We analyze its performance and share how more advanced synchronization techniques might interact with software-based coherence CXL hardware and program execution models.

High Performance Computing (HPC)↗

SGD-Net: Efficient Model-Based Deep Learning with Theoretical Guarantees

Deep unfolding networks have recently gained popularity for solving imaging inverse problems. However, the computational and memory complexity of data-consistency layers within traditional deep unfolding networks scales with the number of measurements, limiting their applicability to large-scale imaging inverse problems. We propose SGD-Net as a new methodology for improving the efficiency of deep unfolding through stochastic approximations of the data-consistency layers. Our theoretical analysis shows that SGD-Net can be trained to approximate batch deep unfolding networks to an arbitrary precision. Our simulations on intensity diffraction tomography and sparse-view computed tomography show that SGD-Net can match the performance of the traditional batch network at a fraction of training and testing complexity.Deep unfolding networks have recently gained popularity for solving imaging inverse problems. However, the computational and memory complexity of data-consistency layers within traditional deep unfolding networks scales with the number of measurements, limiting their applicability to large-scale imaging inverse problems. We propose SGD-Net as a new methodology for improving the efficiency of deep unfolding through stochastic approximations of the data-consistency layers. Our theoretical analysis shows that SGD-Net can be trained to approximate batch deep unfolding networks to an arbitrary precision. Our simulations on intensity diffraction tomography and sparse-view computed tomography show that SGD-Net can match the performance of the traditional batch network at a fraction of training and testing complexity.

97 MATHEMATICS AND COMPUTING↗

High Energy Density Shape Memory Polymers Using Strain-Induced Supramolecular Nanostructures

Shape memory polymers are promising materials in many emerging applications due to their large extensibility and excellent shape recovery. However, practical application of these polymers is limited by their poor energy densities (up to ~1 MJ/m 3 ). Here, we report an approach to achieve a high energy density, one-way shape memory polymer based on the formation of strain-induced supramolecular nanostructures. As polymer chains align during strain, strong directional dynamic bonds form, creating stable supramolecular nanostructures and trapping stretched chains in a highly elongated state. Upon heating, the dynamic bonds break, and stretched chains contract to their initial disordered state. This mechanism stores large amounts of entropic energy (as high as 19.6 MJ/m 3 or 17.9 J/g), almost six times higher than the best previously reported shape memory polymers while maintaining near 100% shape recovery and fixity. The reported phenomenon of strain-induced supramolecular structures offers a new approach toward achieving high energy density shape memory polymers.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Speeding up and reducing memory usage for scientific machine learning via mixed precision

Scientific machine learning (SciML) has emerged as a versatile approach to address complex computational science and engineering problems. Within this field, physics-informed neural networks (PINNs) and deep operator networks (DeepONets) stand out as the leading techniques for solving partial differential equations by incorporating both physical equations and experimental data. However, training PINNs and DeepONets require significant computational resources, including long computational times and large amounts of memory. In search of computational efficiency, training neural networks using half precision (float16) rather than the conventional single (float32) or double (float64) precision has gained substantial interest, given the inherent benefits of reduced computational time and memory consumed. However, we find that float16 cannot be applied to SciML methods, because of gradient divergence at the start of training, weight updates going to zero, and the inability to converge to a local minima. To overcome these limitations, we explore mixed precision, which is an approach that combines the float16 and float32 numerical formats to reduce memory usage and increase computational speed. Our experiments showcase that mixed precision training not only substantially decreases training times and memory demands but also maintains model accuracy. Here, we also reinforce our empirical observations with a theoretical analysis. The research has broad implications for SciML in various computational applications.

97 MATHEMATICS AND COMPUTING↗

Domain wall-magnetic tunnel junction spin–orbit torque devices and circuits for in-memory computing

There are pressing problems with traditional computing, especially for accomplishing data-intensive and real-time tasks, that motivate the development of in-memory computing devices to both store information and perform computation. Magnetic tunnel junction memory elements can be used for computation by manipulating a domain wall, a transition region between magnetic domains, but the experimental study of such devices has been limited by high current densities and low tunnel magnetoresistance. Here, we study prototypes of three-terminal domain wall-magnetic tunnel junction in-memory computing devices that can address data processing bottlenecks and resolve these challenges by using perpendicular magnetic anisotropy, spin–orbit torque switching, and an optimized lithography process to produce average device tunnel magnetoresistance TMR = 171% and average resistance-area product RA = 29 Ω μm2, close to the RA of the unpatterned film. Device initialization variation in switching voltage is shown to be curtailed to 7%–10% by controlling the domain wall initial position, which we show corresponds to 90%–96% accuracy in a domain wall-magnetic tunnel junction full adder simulation. Repeatability of writing and resetting the device is shown. A circuit shows an inverter operation between two devices, showing that a voltage window is large enough, compared to the variation noise, to repeatably operate a domain wall-magnetic tunnel junction circuit. These results make strides in using magnetic tunnel junctions and domain walls for in-memory and neuromorphic computing applications.

Alamdar, Mahshid (ORCID:0000000221732935)↗

Atomic-Scale Modulation of Synthetic Magnetic Order in Oxide Superlattices

We report atomic-scale precision control of magnetic interactions facilitates a synthetic spin order useful for spintronics, including advanced memory and quantum logic devices. Conventional modulation of synthetic spin order has been limited to metallic heterostructures that exploit Ruderman–Kittel–Kasuya–Yosida interaction through a nonmagnetic metallic spacer; however, they face issues arising from Joule heating and/or electric breakdown. The practical realization and observation of a synthetic spin order across a nonmagnetic insulating spacer will lead to the development of spin-related devices with a completely different concept. Herein, the atomic-scale modulation of the synthetic spiral spin order in oxide superlattices composed of ferromagnetic metal and nonmagnetic insulator layers is reported. The atomically controlled superlattice exhibits an oscillatory magnetic behavior, representing the existence of a spiral spin structure. Depth-sensitive polarized neutron reflectometry evidences modulated spiral spin structures as a function of the nonmagnetic insulator layer thickness. Atomic-scale customization of the spin state can move the field one step further to actual spintronic applications.

74 ATOMIC AND MOLECULAR PHYSICS↗

Machine Learning for Slow Spill Regulation in the Fermilab Delivery Ring for Mu2e

A third-integer resonant slow extraction system is being developed for the Fermilab’s Delivery Ring to deliver protons to the Mu2e experiment. During a slow extraction process, the beam on target is liable to experience small intensity variations due to many factors. Owing to the experiment’s strict requirements in the quality of the spill, a Spill Regulation System (SRS) is currently under design. The SRS primarily consists of three components - slow regulation, fast regulation, and harmonic content tracker. In this presentation, we shall present the investigations of using Machine Learning (ML) in the fast regulation system, including further optimizations of PID controller gains for the fast regulation, prospects of an ML agent completely replacing the PID controller using supervised learning schemes such as Long Short-Term Memory (LSTM) and Gated Recurrent Unit (GRU) ML models, the simulated impact and limitation of machine response characteristics on the effectiveness of both PID and ML regulation of the spill. We also present here nascent results of Reinforcement Learning efforts, including continuous-action soft actor-critic methods, to regulate the spill rate.

43 PARTICLE ACCELERATORS↗