Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “memory mapping”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

MrHyDE v.1.0

SAND2024-01324O MrHyDE, which stands for Multi-resolution Hybridized Differential Equations, is a general-purpose C++ package for the solution of coupled multiphysics and multiscale systems on massively parallel computing systems. MrHyDE is designed to enable moving beyond forward simulation for multiscale applications which includes optimization, control, uncertainty quantification, and stochastic inversion. The framework provides interfaces to several packages within the Trilinos framework and leverages automatic differentiation to enable adjoint capabilities for large-scale, gradient-based optimization. MrHyDE provides automated multiscale capabilities through a subgrid model interface and multiscale Dirichlet-to-Neumann maps. For extreme-scale applications, MrHyDE provides in situ data-compression algorithms to reduce memory requirements while maintaining performance. MrHyDE is a general-purpose, computational framework for the solution of multiscale and multiphysics applications. It uses a combination of structure-preserving, physics-compatible discretizations, fully implicit methods, multi-resolution schemes, or fully explicit methods. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

SciDAC↗

The relationship between interstellar dust and the isotopic anomalies in meteorites

Work on the ways in which the isotopic anomalies found in meteorites can be regarded as the chemical memory of even larger anomalies found in interstellar dust is outlined. This approach constitutes one theory of the isotopic anomalies, standing in contrast to the idea of a spatial inhomogeneity in the early solar system owing to inhomogeneous admixture from a neighboring supernova. The four mechanisms of isotopic chemical memory in interstellar dust are: (1) thermal condensation within expanding events of nucleosynthesis; (2) different isotopic mappings onto the grain size spectrum; (3) dust components of differing age; and (4) isotope-dependent interstellar chemistry. Specific examples of each mechanism are given to illustrate how each may have contributed to known isotopic anomalies.

Clayton, D. D.↗

Crystallographic variant mapping using precession electron diffraction data

In this work, we developed three methods to map crystallographic variants of samples at the nanoscale by analyzing precession electron diffraction data using a high-temperature shape memory alloy and a VO2 thin film on sapphire as the model systems. The three methods are (I) a user-selecting-reference pattern approach, (II) an algorithm-selecting-reference-pattern approach, and (III) a k-means approach. In the first two approaches, Euclidean distance, Cosine, and Structural Similarity (SSIM) algorithms were assessed for the diffraction pattern similarity quantification. We demonstrated that the Euclidean distance and SSIM methods outperform the Cosine algorithm. We further revealed that the random noise in the diffraction data can dramatically affect similarity quantification. Denoising processes could improve the crystallographic mapping quality. With the three methods mentioned above, we were able to map the crystallographic variants in different materials systems, thus enabling fast variant number quantification and clear variant distribution visualization. The advantages and disadvantages of each approach are also discussed. We expect these methods to benefit researchers who work on martensitic materials, in which the variant information is critical to understand their properties and functionalities.

Crystallographic variant mapping↗

CORE-BFS: Communication-Optimized REctangular-partitioned BFS Achieving 160.845 TeraTEPS on Frontier Supercomputer

Distributed Breadth-First Search (BFS) is fundamental to many large-scale graph applications, but its performance on parallel systems is often limited by high communication overhead. This paper presents CORE-BFS, an extremely scalable GPU-based BFS implementation that introduces a unique rectangular 2D partitioning-based design for Frontier supercomputer. To further improve performance, we propose four key optimizations: (1) Rectangular 2D-partition specific data formats that use two compressed row and one compressed column status array bitmaps combined with a Double Compressed Sparse Row (DCSR) format per partition, reducing memory footprint and inter-rank traffic; (2) Adaptive frontier & communication strategy that unifies top-down and bottom-up traversal on the rectangular layout, uses lazy synchronization in top-down levels, and switches variants based on frontier size to minimize communication overhead; (3) Frontier-split degree-aware update that maps frontier vertices to thread-centric, wavefront-centric, and block-centric kernels based on their degree to improve GPU utilization and memory coalescing; (4) Row-reduction pipeline that overlaps bottom-up adjacency list processing with row-wise bitmap reduction to hide inter-rank latency. Together, these techniques increase parallelism while reducing memory and communication overhead. On the Graph500 benchmark, CORE - BFS scales up to 9,248 Frontier nodes with scale-42 graphs and reaches 160.845 TTEPS, delivering a 5.42 × speedup over our previous Frontier implementation.

Yang, Haoshen [Rutgers University]↗

Registration and fusion of large-scale melt pool temperature and morphology monitoring data demonstrated for surface topography prediction in LPBF

In-situ monitoring technologies for laser powder bed fusion (LPBF) additive manufacturing often face one key challenge, extracting the ultrafast melt pool (MP) signatures for understanding the localized part properties. Further, the spatial information of each monitored MP signature is essential for correlating the MP – part property. This spatial information is often unavailable especially from commercial LPBF printers. Many MP monitoring methods have been reported and utilized. However, very few of these have the MP’s spatial information. To overcome this challenge, in this work we report a method for spatially registering the key MP signatures (MP intensity, temperature, and area) to the monitored print parts. The MP signatures are obtained from our coaxial high-speed single-camera based two-wavelength imaging pyrometry (STWIP) system and the MP spatial information is obtained from an off-axis camera system. A machine learning aided image analysis method is employed to retrieve the spatial distribution of MPs within the corresponding part’s coordinates system. Then, the MP signature maps (MPSMs) are reconstructed by mapping the STWIP measured MP signatures to the registered MP coordinates. Further, a long short-term memory (LSTM) neural network is developed for estimating the layer surface topography from the registered MPSMs. The obtained results indicate that the layer surface topography can be more accurately estimated by using MP temperature signature rather than MP intensity and/or area signatures as in common practice. Finally, our developed methods for MP monitoring, registration, and MP-surface topography prediction offer advanced capabilities for the online detection of process anomalies and part defects.

36 MATERIALS SCIENCE↗

Wavelet and Deep-Learning-Based Approach for Generation System Problematic Parameters Identification and Calibration

Accurate models of generation systems are critical for maintaining reliable and secure grid operations. In this paper, a novel and systematic approach is proposed to identify and calibrate the generation system problematic parameters using continuous wavelet transform (CWT) and advanced deep-learning technology. The phasor measurement unit (PMU) data are used through “event playback” to check whether the parameter calibration is required, and if yes, a group of suspicious parameters will be identified as the primary problematic parameter candidates (PPCs). These primary PPCs are randomly perturbed to generate the event playback simulation data, which are used by the CWT and convolutional neural networks (CNNs) to further narrow down the primary PPCs into a smaller set of candidates. Then, the identified candidates are perturbed again to generate massive event playback simulation data for training a parameter calibration neural network. Here, we designed a multi-output neural network structure to find the mappings between the perturbed parameters and the simulation data using both CNN and long short-term memory (LSTM) models. Finally, the well-trained and tested CNN-LSTM model is used to estimate the accurate value of the suspicious parameters with actual PMU measurements. The proposed CNN-LSTM network can accurately and reliably estimate the generation-system problematic parameters, and has better performance when compared to other machine-learning methods, such as the multilayer perceptron network and the conditional variational autoencoder method. The accuracy and effectiveness of the proposed approach have been validated through simulation and real-world data.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Supercomputing '91; Proceedings of the 4th Annual Conference on High Performance Computing, Albuquerque, NM, Nov. 18-22, 1991

Various papers on supercomputing are presented. The general topics addressed include: program analysis/data dependence, memory access, distributed memory code generation, numerical algorithms, supercomputer benchmarks, latency tolerance, parallel programming, applications, processor design, networks, performance tools, mapping and scheduling, characterization affecting performance, parallelism packaging, computing climate change, combinatorial algorithms, hardware and software performance issues, system issues. (No individual items are abstracted in this volume)

Source record↗

Multi-frequency electrical impedance tomography

Apparatus includes a plurality of geological subsurface electrical line sensors spaced apart from each other proximate a predetermined geological subsurface region of interest, with at least one of the electrical line sensors situated as a line source to produce a multi-frequency electrical impedance tomography source signal, and with at least one of the electrical line sensors situated as a line detector to receive the multi-frequency electrical impedance tomography response signal associated with the source signal that propagates through the predetermined geological subsurface region of interest, and a controller including a processor and a memory configured with instructions that, when executed by the processor, cause the processor to determine an electrical mapping over the predetermined geological subsurface region of interest based on the multi-frequency electrical impedance tomography source signal, response signal, and the spatial positions of the geological subsurface electrical line sensors.

Karra, Satish↗

Initial operating capability for the hypercluster parallel-processing test bed

The NASA Lewis Research Center is investigating the benefits of parallel processing to applications in computational fluid and structural mechanics. To aid this investigation, NASA Lewis is developing the Hypercluster, a multi-architecture, parallel-processing test bed. The initial operating capability (IOC) being developed for the Hypercluster is described. The IOC will provide a user with a programming/operating environment that is interactive, responsive, and easy to use. The IOC effort includes the development of the Hypercluster Operating System (HYCLOPS). HYCLOPS runs in conjunction with a vendor-supplied disk operating system on a Front-End Processor (FEP) to provide interactive, run-time operations such as program loading, execution, memory editing, and data retrieval. Run-time libraries, that augment the FEP FORTRAN libraries, are being developed to support parallel and vector processing on the Hypercluster. Special utilities are being provided to enable passage of information about application programs and their mapping to the operating system. Communications between the FEP and the Hypercluster are being handled by dedicated processors, each running a Message-Passing Kernel, (MPK). A shared-memory interface allows rapid data exchange between HYCLOPS and the communications processors. Input/output handlers are built into the HYCLOPS-MPK interface, eliminating the need for the user to supply separate I/O support programs on the FEP.

Cole, Gary L.↗

Single-pass memory system evaluation for multiprogramming workloads

Modern memory systems are composed of levels of cache memories, a virtual memory system, and a backing store. Varying more than a few design parameters and measuring the performance of such systems has traditionally be constrained by the high cost of simulation. Models of cache performance recently introduced reduce the cost simulation but at the expense of accuracy of performance prediction. Stack-based methods predict performance accurately using one pass over the trace for all cache sizes, but these techniques have been limited to fully-associative organizations. This paper presents a stack-based method of evaluating the performance of cache memories using a recurrence/conflict model for the miss ratio. Unlike previous work, the performance of realistic cache designs, such as direct-mapped caches, are predicted by the method. The method also includes a new approach to the problem of the effects of multiprogramming. This new technique separates the characteristics of the individual program from that of the workload. The recurrence/conflict method is shown to be practical, general, and powerful by comparing its performance to that of a popular traditional cache simulator. The authors expect that the availability of such a tool will have a large impact on future architectural studies of memory systems.

Conte, Thomas M.↗

Scalable In Situ Computation of Lagrangian Representations via Local Flow Maps

In situ computation of Lagrangian flow maps to enable post hoc time-varying vector field analysis has recently become an active area of research. However, the current literature is largely limited to theoretical settings and lacks a solution to address scalability of the technique in distributed memory. To improve scalability, we propose and evaluate the benefits and limitations of a simple, yet novel, performance optimization. Our proposed optimization is a communication-free model resulting in local Lagrangian flow maps, requiring no message passing or synchronization between processes, intrinsically improving scalability, and thereby reducing overall execution time and alleviating the encumbrance placed on simulation codes from communication overheads. To evaluate our approach, we computed Lagrangian flow maps for four time-varying simulation vector fields and investigated how execution time and reconstruction accuracy are impacted by the number of GPUs per compute node, the total number of compute nodes, particles per rank, and storage intervals. Our study consisted of experiments computing Lagrangian flow maps with up to 67M particle trajectories over 500 cycles and used as many as 2048 GPUs across 512 compute nodes. In all, our study contributes an evaluation of a communication-free model as well as a scalability study of computing distributed Lagrangian flow maps at scale using in situ infrastructure on a modern supercomputer.

Sane, Sudhanshu↗

Physics in the Machine: Integrating Physical Knowledge in Autonomous Phase-Mapping

Application of artificial intelligence (AI), and more specifically machine learning, to the physical sciences has expanded significantly over the past decades. In particular, science-informed AI, also known as scientific AI or inductive bias AI, has grown from a focus on data analysis to now controlling experiment design, simulation, execution and analysis in closed-loop autonomous systems. The CAMEO (closed-loop autonomous materials exploration and optimization) algorithm employs scientific AI to address two tasks: learning a material system’s composition-structure relationship and identifying materials compositions with optimal functional properties. By integrating these, accelerated materials screening across compositional phase diagrams was demonstrated, resulting in the discovery of a best-in-class phase change memory material. Key to this success is the ability to guide subsequent measurements to maximize knowledge of the composition-structure relationship, or phase map. In this work we investigate the benefits of incorporating varying levels of prior physical knowledge into CAMEO’s autonomous phase-mapping. This includes the use of ab-initio phase boundary data from the AFLOW repositories, which has been shown to optimize CAMEO’s search when used as a prior.

97 MATHEMATICS AND COMPUTING↗

A ground-based memory state tracker for satellite on-board computer memory

The TOPEX/POSEIDON satellite, currently in Earth orbit, will use radar altimetry to measure sea surface height over 90 percent of the world's ice-free oceans. In combination with a precise determination of the spacecraft orbit, the altimetry data will provide maps of ocean topography, which will be used to calculate the speed and direction of ocean currents worldwide. NASA's Jet Propulsion Laboratory (JPL) has primary responsibility for mission operations for TOPEX/POSEIDON. Software applications have been developed to automate mission operations tasks. This paper describes one of these applications, the Memory State Tracker, which allows the ground analyst to examine and track the contents of satellite on-board computer memory quickly and efficiently, in a human-readable format, without having to receive the data directly from the spacecraft. This process is accomplished by maintaining a groundbased mirror-image of spacecraft On-board Computer memory.

Quan, Alan↗

Colloidal State Machines as Smart Tracers for Chemical Reactor Analysis

A widely utilized tool in reactor analysis is passive tracers that report the residence time distribution, allowing estimation of the conversion and other properties of the system. Recently, advances in microrobotics have introduced powered and functional entities with sizes comparable to some traditional tracers. This has motivated the concept of Smart Tracers that could record the local chemical concentrations, temperature, or other conditions as they progress through reactors. Herein, the design constraints and advantages of Smart Tracers by simulating their operation in a laminar flow reactor model conducting chemical reactions of various orders are analyzed. It is noted that far fewer particles are necessary to completely map even the most complex concentration gradients compared with their conventional counterparts. Design criteria explored herein include sampling frequency, memory storage capacity, and ensemble number necessary to achieve the required accuracy to inform a reactor model. Cases of severe particle diffusion and sensor noise appear to bind the functional upper limit of such probes and require consideration for future design. The results of the study provide a starting framework for applying the new technology of microrobotics to the broad and impactful set of problems classified as chemical reactor analysis.

97 MATHEMATICS AND COMPUTING↗

Optimal strategies for optical quantum memories using long-lived noble-gas spins

Nuclear spins of noble gases exhibit exceptionally long coherence times and can potentially serve as a long-lived storage medium for quantum information. We analyze and compare the performance of two mechanisms for mapping the quantum state of light onto the collective spin state of noble gases. The first mechanism utilizes collisional exchange with the electronic spin state of metastable noble-gas atoms, while the second relies on spin-exchange collisions with ground-state alkali-metal atoms. We describe the operation of an optical quantum memory relying on these two mechanisms using a compact model and study strategies that optimize the memory storage efficiency. Through numerical simulations, we identify optimal sequences for storing optical signals with different signal bandwidths and electronic spin relaxation rates. This work highlights the qualitative difference between the two approaches for using noble gases as long-lived quantum memories at noncryogenic conditions and outlines the regimes in which they are expected to be efficient.

atomic ensemble↗

Experiences with SYCL on AMD GPUs with Kokkos

With the recent diversification of the hardware landscape in the high-performance computing (HPC) community, performance-portability solutions are becoming more and more important. One of the most popular choices is Kokkos, which recently became a Linux Foundation project. Most of its development is supported by the US Department of Energy and the French Alternative Energies and Atomic Energy Commission. Kokkos is implemented as a C++ library with multiple backends to support CPUs as well as various GPU architectures. These backends include OpenMP, CUDA, HIP, and also SCYL. This approach enables users to leverage the preferred vendor toolchain for the respective platform (e.g. CUDA, ROCm, OneAPI). The SYCL backend is used to target Intel GPUs, in particular to support the Aurora exascale supercomputer. However, SYCL itself also offers a large degree of portability, and in fact Kokkos’ CI for SYCL has been running on NVIDIA hardware due to a lack of access to Intel GPUs. In this report, we describe our experience with using Kokkos SYCL backend on AMD GPUs targeting the Frontier supercomputer at Oak Ridge National Laboratory. The two major SYCL implementations are DPC++ and AdaptiveCpp. While the Kokkos SYCL backend has been implemented using the former, the latter was the first implementation to target AMD GPUs. We will discuss the experience with both of these SYCL implementations in terms of functionality and performance. Using Kokkos to evaluate SYCL toolchains has a number of benefits. Kokkos’ use of SYCL is fairly complex, exercising features such as graphs, relocatable device functions, atomics – including for non-arithmetic types, as well as pinned and page migratable memory allocations. Kokkos also needs to implement capabilities such as Kokkos’ hierarchical parallelism that are not a straight-forward mapping to SYCL capabilities. Furthermore, a large number of libraries and applications that represent diverse use cases are implemented in Kokkos, providing readily available test cases for a toolchain evaluation. Preliminary results show that support for AMD GPUs in DPC++ is much less mature than for NVIDIA GPUs or Intel GPUs. While the situation has improved significantly over the last year, we still encounter many runtime failures, dispatching problems, and code generation issues. With AdaptiveCpp the challenges arise even earlier in the evaluation process. Since Kokkos’ SYCL implementation is largely focused on supporting Intel GPUs, we opted to leverage SYCL extensions which are available in DPC++ but not in AdaptiveCpp. Furthermore, AdaptiveCpp appears to be less conformant with the SYCL2020 standard which Kokkos relies on. In some cases, we are able to work around the lack of feature support, in other cases we have to disable certain Kokkos capabilities to evaluate the toolchain. Our evaluation will leverage Kokkos’ unit tests to establish basic functionality and feature completeness. We then use simple benchmarks for components of a CG implementation as a measure of usability and performance of the SYCL toolchains.

97 MATHEMATICS AND COMPUTING↗

A dynamic solvent chamber propagation estimation framework using RNN for warm solvent injection in heterogeneous reservoirs

Warm solvent injection (WSI), injecting low-temperature solvent into formations to reduce the viscosity of heavy oil, is a clean technology for heavy oil production through reducing greenhouse gas emissions and water usage. The success of WSI operation depends on the uniform development and propagation of solvent chambers in reservoirs. However, reservoir heterogeneity stemming from shale barriers plays a detrimental role in the conformance of solvent chamber development and oil production rate. In this work, we developed a novel recurrent neural network (RNN)-based framework with the capability of efficiently tracking and estimating the solvent chamber positions in heterogeneous reservoirs based on only production time-series data. The developed estimation model utilizes the “sequence-to-sequence" mapping methodology to correlate observed production time-series sequence and solvent chamber edge sequence via a long short-term memory (LSTM) algorithm. The trained RNN models exhibit high accuracy, evidenced by the predicted dynamic solvent chamber locations match the corresponding true locations from numerical simulation, with a high coefficient of determination (R 2 ) and a low mean squared error. Specifically, the achieved R 2 values exceed 0.98 on both the training and testing data. The developed RNN-based workflow was tested via several cases from both regularly- and irregularly-shaped shale barriers, and the results were promising. The predicted solvent chambers showed strong agreement with those obtained from numerical simulations. The major benefits of this workflow include reducing computational time and saving overall monitoring and tracking costs for conventional techniques. In conclusion, the present work would provide a good demonstration of the capability of practical integration of machine learning methods in solving engineering problems.

58 GEOSCIENCES↗

Integration of Ag-CBRAM crossbars and Mott ReLU neurons for efficient implementation of deep neural networks in hardware

In-memory computing with emerging non-volatile memory devices (eNVMs) has shown promising results in accelerating matrix-vector multiplications. However, activation function calculations are still being implemented with general processors or large and complex neuron peripheral circuits. Here, we present the integration of Ag-based conductive bridge random access memory (Ag-CBRAM) crossbar arrays with Mott rectified linear unit (ReLU) activation neurons for scalable, energy and area-efficient hardware (HW) implementation of deep neural networks. We develop Ag-CBRAM devices that can achieve a high ON/OFF ratio and multi-level programmability. Compact and energy-efficient Mott ReLU neuron devices implementing ReLU activation function are directly connected to the columns of Ag-CBRAM crossbars to compute the output from the weighted sum current. We implement convolution filters and activations for VGG-16 using our integrated HW and demonstrate the successful generation of feature maps for CIFAR-10 images in HW. Our approach paves a new way toward building a highly compact and energy-efficient eNVMs-based in-memory computing system.

Mott insulators↗