Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “near memory computing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Julia as a unifying end-to-end workflow language on the Frontier exascale system

We evaluate Julia as a single language and ecosystem paradigm powered by LLVM to develop workflow components for high-performance computing. We run a Gray-Scott, 2-variable diffusion-reaction application using a memory-bound, 7-point stencil kernel on Frontier, the US Department of Energy’s first exascale supercomputer. We evaluate the performance, scaling, and trade-offs of (i) the computational kernel on AMD’s MI250x GPUs, (ii) weak scaling up to 4,096 MPI processes/GPUs or 512 nodes, (iii) parallel I/O writes using the ADIOS2 library bindings, and (iv) Jupyter Notebooks for interactive analysis. Results suggest that although Julia generates a reasonable LLVM-IR, a nearly 50% performance difference exists vs. native AMD HIP stencil codes when running on the GPUs. As expected, we observed near-zero overhead when using MPI and parallel I/O bindings for system-wide installed implementations. Consequently, Julia emerges as a compelling high-performance and high-productivity workflow composition language, as measured on the fastest supercomputer in the world.

Godoy, William↗

Achieving High Efficiency in Reduced Order Modeling for Large Scale Polycrystal Plasticity Simulations

Reduced order models for the nonlinear response of heterogeneous microstructures typically require a construction (or training) stage to build the reduced order basis. In this manuscript, an efficient model construction strategy for the eigenstrain homogenization method (EHM) is presented. The proposed strategy relies on a parallel, element-by-element, conjugate gradient solver. Near linear scaling has been achieved with respect to the number of degrees of freedom used to resolve the microstructure. Linear scaling with respect to the number of pre-analyses required to construct the reduced order model (ROM) follows from the EHM formulation. Furthermore, a parallel implementation for fast evaluation of the constructed ROM has been developed using shared memory parallelization. It has been shown that for large microstructures with ≈ 10,000 grains, the total computational cost of evaluating the nonlinear response of a polycrystal could be reduced by approximately an order of magnitude using 32 cores with respect to serial ROM simulation. The present methodology has been verified using an additively manufactured polycrystalline microstructure of a nickel-based superalloy, Inconel 625. The capability of the developed framework to construct a ROM for such large microstructures, as well as the ability of the ROM to predict average and local quantities of interest has been demonstrated.

microscale↗

Computing material volume fractions on a superimposed mesh as applied to Monte Carlo particle transport simulations

Here, we present a newly implemented ray tracing algorithm in OpenMC for efficiently computing material volume fractions on superimposed meshes in complex geometries. By firing rays along each coordinate direction through the geometry, the approach accumulates track-length data in each mesh element, thereby determining the fractional composition of each material. Scaling studies on three different models—a random tetrahedra configuration, the Frascati Neutron Generator ITER dose rate benchmark, and a stellarator design—show excellent parallel performance, with nearly linear speedup on modern multi-threaded and distributed-memory systems. An analysis of the residual error relative to high-resolution reference solutions demonstrated that under optimal conditions it decreases as 1/R, where R is the number of rays fired, making it straightforward to achieve user-prescribed accuracy. This new functionality enables practical, mesh-based approaches for detailed nuclear analyses in production Monte Carlo workflows without resorting to expensive, fully conformal or unstructured meshing.

Monte Carlo↗

Electrode and Microstructure Dependence of Oxygen Diffusion in Ferroelectric Hafnium Zirconium Oxide Thin Films

Hafnia-based ferroelectrics hold promise to reduce energy demand for computing by enabling compute-in-memory and as non-volatile memories. The ferroelectric phase in this material system is, in part, stabilized by oxygen vacancies. While oxygen vacancies may be a necessity for phase stability, they limit device endurance through diffusion and accumulation into conducting channels. Herein, it is shown that oxygen diffusion is spatially variable within individual grains of ferroelectric hafnium zirconium oxide (HZO). Using 18 O tracers and finite difference modeling, it is shown that grain boundaries and regions near electrode interfaces allow for relatively rapid oxygen diffusion, with values as much as 10 4 larger than the grain cores. Further, the selection of electrode material affects the diffusion coefficients across all microstructural regions. HZO films in contact with TiN electrodes result in more oxygen-deficient HZO films and higher oxygen diffusion coefficients. Tungsten electrodes result in fewer vacancies and lower diffusion coefficients. Diffusion activation energy differences between the HZO with the two electrodes is reconciled by differing populations of charged and uncharged oxygen vacancies. This insight into the local vacancy populations and diffusion pathways provides a platform for designing hafnia-based films, deposition processes, and integration strategies to reduce vacancy gradients and improve performance.

36 MATERIALS SCIENCE↗

Developments and Validations of Fully Coupled CFD and Practical Vortex Transport Method for High-Fidelity Wake Modeling in Fixed and Rotary Wing Applications

A novel Computational Fluid Dynamics (CFD) coupling framework using a conventional Reynolds-Averaged Navier-Stokes (BANS) solver to resolve the near-body flow field and a Particle-based Vorticity Transport Method (PVTM) to predict the evolution of the far field wake is developed, refined, and evaluated for fixed and rotary wing cases. For the rotary wing case, the RANS/PVTM modules are loosely coupled to a Computational Structural Dynamics (CSD) module that provides blade motion and vehicle trim information. The PVTM module is refined by the addition of vortex diffusion, stretching, and reorientation models as well as an efficient memory model. Results from the coupled framework are compared with several experimental data sets (a fixed-wing wind tunnel test and a rotary-wing hover test).

Anusonti-Inthra, Phuriwat↗

Simulating many-engine spacecraft: Exceeding 1 quadrillion degrees of freedom via information geometric regularization

We present an optimized implementation of the recently proposed information geometric regularization (IGR) for unprecedented scale simulation of compressible fluid flows applied to multi-engine spacecraft boosters. We improve upon state-of-the-art computational fluid dynamics (CFD) techniques in terms of computational cost, memory footprint, and energy-to-solution metrics. Unified memory on coupled CPU–GPU or APU platforms increases problem size with negligible overhead. Mixed half/single-precision storage and computation are used on well-conditioned numerics. We simulate flow at 200 trillion grid points and 1 quadrillion degrees of freedom, exceeding the current record by a factor of 20. A factor of 4 wall-time speedup is achieved over optimized baselines. Ideal weak scaling is observed on OLCF Frontier, LLNL El Capitan, and CSCS Alps using the full systems. Strong scaling is near ideal at extreme conditions, including 80% efficiency on CSCS Alps with an 8 node baseline and stretching to the full system.

Wilfong, Benjamin [Georgia Institute of Technology↗

Fast Machine Learning Lidar Surrogate Simulator: Pristine Clear Sky

The simulations of lidar signals and retrievals rely on a range of optic-physical models, such as radiative transfer models, particle scattering and absorption models, along with the output data from atmospheric physical models. Integrating these different models to represent signals of a lidar system is computationally expensive, and performing backward retrievals can be complex and ambiguous. However, with the advantages of Machine Learning, there is a new potential for building effective lidar signal database linked to corresponding atmospheric profiles. For this project, we are developing a fast pre-trained neural network as the lidar surrogate simulator using simulated data for a CALIPSO-like lidar (355 nm, 532nm, and 1064nm), and a CO2 differential absorption lidar (DIAL) near 1571nm. Specifically, we utilize a long short-term memory (LSTM) model to map the relationships between atmospheric profiles (pressure, temperature, air density and CO2 mixing ratio) and lidar signals. This approach allows us to build machine learning based simulators that can reconstruct lidar signals at specific bands from MERRA reanalysis data, and perform retrievals of atmospheric profiles using lidar signals at various wavelengths. As a first step, the results show the potential of this method to establish a foundational model for sensor signals. This model offers the promise of enabling both accurate predictions and rapid retrievals, providing a more efficient approach to signal processing and analysis.

Shan Zeng↗

Sparsity-Independent Lyapunov Exponent in the Sachdev-Ye-Kitaev Model

The saturation of a recently proposed universal bound on the Lyapunov exponent has been conjectured to signal the existence of a gravity dual. This saturation occurs in the low-temperature limit of the dense Sachdev-Ye-Kitaev (SYK) model, N Majorana fermions with q body ( q > 2 ) infinite-range interactions. We calculate certain out-of-time-order correlators (OTOCs) for N ≤ 64 fermions for a highly sparse SYK model and find no significant dependence of the Lyapunov exponent on sparsity up to near the percolation limit where the Hamiltonian breaks up into blocks. This provides strong support to the saturation of the Lyapunov exponent in the low-temperature limit of the sparse SYK. A key ingredient to reaching N = 64 is the development of a novel quantum spin model simulation library that implements highly optimized matrix-free Krylov subspace methods on graphical processing units. This leads to a significantly lower simulation time as well as vastly reduced memory usage over previous approaches, while using modest computational resources. Strong sparsity-driven statistical fluctuations require both the use of a much larger number of disorder realizations with respect to the dense limit and a careful finite size scaling analysis. The saturation of the bound in the sparse SYK points to the existence of a gravity analog that would enlarge substantially the number of field theories with this feature. Published by the American Physical Society 2024

Physics↗

A High-Order Finite Spectral Volume Method for Conservation Laws on Unstructured Grids

A time accurate, high-order, conservative, yet efficient method named Finite Spectral Volume (FSV) is developed for conservation laws on unstructured grids. The concept of a 'spectral volume' is introduced to achieve high-order accuracy in an efficient manner similar to spectral element and multi-domain spectral methods. In addition, each spectral volume is further sub-divided into control volumes (CVs), and cell-averaged data from these control volumes is used to reconstruct a high-order approximation in the spectral volume. Riemann solvers are used to compute the fluxes at spectral volume boundaries. Then cell-averaged state variables in the control volumes are updated independently. Furthermore, TVD (Total Variation Diminishing) and TVB (Total Variation Bounded) limiters are introduced in the FSV method to remove/reduce spurious oscillations near discontinuities. A very desirable feature of the FSV method is that the reconstruction is carried out only once, and analytically, and is the same for all cells of the same type, and that the reconstruction stencil is always non-singular, in contrast to the memory and CPU-intensive reconstruction in a high-order finite volume (FV) method. Discussions are made concerning why the FSV method is significantly more efficient than high-order finite volume and the Discontinuous Galerkin (DG) methods. Fundamental properties of the FSV method are studied and high-order accuracy is demonstrated for several model problems with and without discontinuities.

Wang, Z. J.↗

Comparison of machine learning systems trained to detect Alfvén eigenmodes using the CO 2 interferometer on DIII-D

Abstract A Machine-Learning (ML) based detection scheme that automatically detects Alfvén Eigenmodes (AE) in a labelled DIII-D database is presented here. Controlling AEs is important for the success of planned burning plasma devices such as ITER, since resonant fast ions can drive AEs unstable and degrade the performance of the plasma or damage the first walls of the machine vessel. Artificial Intelligence could be useful for real-time detection and control of AEs in steady-state plasma scenarios by implementing ML-based models into control algorithms that drive actuators for mitigation of AE impacts. Thus, the objective is to compare differences in performance between using two different recurrent neural network systems (Reservoir Computing Network and Long Short Term Memory Network) and two different representations of the C O 2 phase data (simple and crosspower spectrograms). All C O 2 interferometer chords are used to train both models, but only one is processed during each training step. The results from the model and data comparison show higher performance for the RCN model (True Positive Rate = 90% and False Positive Rate = 14%), and that using simple magnitude spectrograms is sufficient to detect AEs. Also, the vertical C O 2 interferometer chord passing near the center is better for ML-based detection of AEs.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Supramolecular Control of Ionic Retention in Electrolyte-Gated Synaptic Transistors

Electrolyte-gated transistors with ion-trapping layers offer a promising platform for artificial synapses in neuromorphic computing, yet molecular mechanisms governing ionic retention remain poorly understood. Here, in this study, we present a supramolecular approach to modulate ion retention by incorporating a crown ether derivative-based polymer network as an ion-trapping layer on top of a semiconducting monolayer. We show that the balance between ion–host binding and ion–solvent interactions dictates the kinetics of ion capture and release, which in turn controls the memory characteristics of the device. By varying the solvent dielectric constant, we tune the ionic retention time from nearly permanent trapping to rapid relaxation. Intermediate solvent polarity enables programmable short- and long-term synaptic behaviors, including excitatory postsynaptic current, paired-pulse facilitation, and long-term potentiation and depression. These findings establish a direct link between supramolecular ion recognition and synaptic plasticity and provide a generalizable design strategy for ionic–electronic neuromorphic devices.

36 MATERIALS SCIENCE↗

A PC-based hardware implementation of the maximum-likelihood classifier for the Shuttle Ice Detection System

A PC-based near-real time implementation of a two-channel maximum-likelihood classifier is described. The statistical distribution of the reflectance characteristics of ice, frost, and water formation on spray-on-foam-insulation, which covers the External Tank surface of the Space Shuttle, is acquired. The classification technique is based on these statistics. The computer, set in either a training or a classifying mode, learns the statistics of the various classes, or produces a color-coded image denoting the respective categories of classification. The classified results are memory-mapped for efficiency. The speed of the classification process is only limited by the speed of the digital frame grabber and the software that interfaces the frame grabber to the monitor. The process took 4 seconds for a 512 x 480 pixel image.

Jaggi, S.↗

Spacecraft automated operations

Trends in automation of planetary spacecraft are examined using data from missions as far back as Mariner '67 and up to the highly sophisticated Galileo. Nine design considerations which influence the degree of automation such as protection against catastrophic failures, highly repetitive functions, loss of spacecraft communications, and the need for near-real-time adaptivity are discussed. Rapid growth of automation is shown in terms of on-board hardware by plots of number of processors on board, the average speed of processors, and total core memory. The number of commands transmitted from the ground has grown to 5 million bits in Voyager, so that increases in mission complexity have increased both in spacecraft automation and ground operations. Achieving greater automation by transferring ground operations to the spacecraft with the current means of controlling missions, are considered noting proposed changes. For the future, improved computer technology, more microprocessors and increased core storage will be used, and the number of automated functions and their complexity will grow. It is concluded that using the growing computational capability of spacecraft will achieve more autonomy thus reversing the trend of increased mission complexity and cost.

Bird, T. H.↗

Predictive Modeling in Plasma Reactor and Process Design

Research continues toward the improvement and increased understanding of high-density plasma tools. Such reactor systems are lauded for their independent control of ion flux and energy enabling high etch rates with low ion damage and for their improved ion velocity anisotropy resulting from thin collisionless sheaths and low neutral pressures. Still, with the transition to 300 mm processing, achieving etch uniformity and high etch rates concurrently may be a formidable task for such large diameter wafers for which computational modeling can play an important role in successful reactor and process design. The inductively coupled plasma (ICP) reactor is the focus of the present investigation. The present work attempts to understand the fundamental physical phenomena of such systems through computational modeling. Simulations will be presented using both computational fluid dynamics (CFD) techniques and the direct simulation Monte Carlo (DSMC) method for argon and chlorine discharges. ICP reactors generally operate at pressures on the order of 1 to 10 mTorr. At such low pressures, rarefaction can be significant to the degree that the constitutive relations used in typical CFD techniques become invalid and a particle simulation must be employed. This work will assess the extent to which CFD can be applied and evaluate the degree to which accuracy is lost in prediction of the phenomenon of interest; i.e., etch rate. If the CFD approach is found reasonably accurate and bench-marked with DSMC and experimental results, it has the potential to serve as a design tool due to the rapid time relative to DSMC. The continuum CFD simulation solves the governing equations for plasma flow using a finite difference technique with an implicit Gauss-Seidel Line Relaxation method for time marching toward a converged solution. The equation set consists of mass conservation for each species, separate energy equations for the electrons and heavy species, and momentum equations for the gas. The sheath is modeled by imposing the Bohm velocity to the ions near the walls. The DSMC method simulates each constituent of the gas as a separate species which would be analogous in CFD to employing separate species mass, momentum, and energy equations. All particles including electrons are moved and allowed to collide with one another with the stipulation that the electrons remain tied to the ions consistent with the concept of ambipolar diffusion. The velocities of the electrons are allowed to be modified during collisions and are not confined to a Maxwellian distribution. These benefits come at a price in terms of computational time and memory. The DSMC and CFD are made as consistent as possible by using similar chemistry and power deposition models. Although the comparison of CFD and DSMC is interesting, the main goal of this work is the increased understanding of high-density plasma flowfields that can then direct improvements in both techniques. This work is unique in the level of the physical models employed in both the DSMC and CFD for high-density plasma reactor applications. For example, the electrons are simulated in the present DSMC work which has not been done before for low temperature plasma processing problems. In the CFD approach, for the first time, the charged particle transport (discharge physics) has been self-consistently coupled to the gas flow and heat transfer.

Hash, D. B.↗

Nanoelectronics: Opportunities for future space applications

Further improvements in the performance of integrated electronics will eventually halt due to practical fundamental limits on our ability to downsize transistors and interconnect wiring. Avoiding these limits requires a revolutionary approach to switching device technology and computing architecture. Nanoelectronics, the technology of exploiting physics on the nanometer scale for computation and communication, attempts to avoid conventional limits by developing new approaches to switching, circuitry, and system integration. This presentation overviews the basic principles that operate on the nanometer scale that can be assembled into practical devices and circuits. Quantum resonant tunneling (RT) is used as the center-piece of the overview since RT devices already operate at high temperature (120 degrees C) and can be scaled, in principle, to a few nanometers in semiconductors. Near- and long-term applications of GaAs and silicon quantum devices are suggested for signal and information processing, memory, optoelectronics, and radio frequency (RF) communication.

Frazier, Gary↗

A structure for digital notch filters

Narrow-band digital notch filters have their poles near the unit circle. As the sampling rate is increased, the poles move towards Z = +1. Implementing such filters requires long registers to overcome the sensitivity and roundoff errors. A filter structure based on digital incremental computers is proposed which has low sensitivity and round-off errors, and simple hardware implementation. The filter structure can be directly used on differentially pulse-code modulated signals. Hardware multipliers are not required as the poles approach z = +1, and excellent results can be obtained using multipliers with very short word lengths, or with small size read-only memories.

Abu-El-haija, A. I.↗

Implementation of a fully-balanced periodic tridiagonal solver on a parallel distributed memory architecture

While parallel computers offer significant computational performance, it is generally necessary to evaluate several programming strategies. Two programming strategies for a fairly common problem - a periodic tridiagonal solver - are developed and evaluated. Simple model calculations as well as timing results are presented to evaluate the various strategies. The particular tridiagonal solver evaluated is used in many computational fluid dynamic simulation codes. The feature that makes this algorithm unique is that these simulation codes usually require simultaneous solutions for multiple right-hand-sides (RHS) of the system of equations. Each RHS solutions is independent and thus can be computed in parallel. Thus a Gaussian elimination type algorithm can be used in a parallel computation and the more complicated approaches such as cyclic reduction are not required. The two strategies are a transpose strategy and a distributed solver strategy. For the transpose strategy, the data is moved so that a subset of all the RHS problems is solved on each of the several processors. This usually requires significant data movement between processor memories across a network. The second strategy attempts to have the algorithm allow the data across processor boundaries in a chained manner. This usually requires significantly less data movement. An approach to accomplish this second strategy in a near-perfect load-balanced manner is developed. In addition, an algorithm will be shown to directly transform a sequential Gaussian elimination type algorithm into the parallel chained, load-balanced algorithm.

Eidson, T. M.↗

Beyond the supercomputer

A NASA-directed development of massively parallel processor (MPP) computers is outlined, noting intended applications for data processing for near term earth resource and environment mapping, radar, and television transmissions. The MPP is designed to perform 100 billion operations/sec to obtain satisfactory image processing, while separate processing units correct distortions, register images, calculate correlation functions, and classify multispectral characteristics. Arrays of 1s and 0s will be manipulated in analog-to-digital conversions generating separate planes corresponding to powers of binaries. Data wires are replaced by fiber-optic tubes or thousands of wires, and single logic gates are replaced by thousands of logic gates and every memory element by thousands of memory elements. Features of the interconnections and the images control processor units are detailed, along with implementation of sliders for program flexibility.

Schaefer, D. H.↗