Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Limited memory method”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Competing Easy-Axis Anisotropies Impacting Magnetic Tunnel Junction-Based Molecular Spintronics Devices (MTJMSDs)

Molecular spintronics devices (MSDs) attempt to harness molecules’ quantum state, size, and configurable attributes for application in computer devices—a quest that began more than 70 years ago. In the vast number of theoretical studies and limited experimental attempts, MSDs have been found to be suitable for application in memory devices and futuristic quantum computers. MSDs have recently also exhibited intriguing spin photovoltaic-like phenomena, signaling their potential application in cost-effective and novel solar cell technologies. The molecular spintronics field’s major challenge is the lack of mass-fabrication methods producing robust magnetic molecule connections with magnetic electrodes of different anisotropies. Another main challenge is the limitations of conventional theoretical methods for understanding experimental results and designing new devices. Magnetic tunnel junction-based molecular spintronics devices (MTJMSDs) are designed by covalently connecting paramagnetic molecules across an insulating tunneling barrier. The insulating tunneling barrier serves as a mechanical spacer between two ferromagnetic (FM) electrodes of tailorable magnetic anisotropies to allow molecules to undergo many intriguing phenomena. Our experimental studies showed that the paramagnetic molecules could produce strong antiferromagnetic coupling between two FM electrodes, leading to a dramatic large-scale impact on the magnetic electrode itself. Recently, we showed that the Monte Carlo Simulation (MCS) was effective in providing plausible insights into the observation of unusual magnetic domains based on the role of single easy-axis magnetic anisotropy. Here, we experimentally show that the response of a paramagnetic molecule is dramatically different when connected to FM electrodes of different easy-axis anisotropies. Motivated by our experimental studies, here, we report on an MCS study investigating the impact of the simultaneous presence of two easy-axis anisotropies on MTJMSD equilibrium properties. In-plane easy-axis anisotropy produced multiple magnetic phases of opposite spins. The multiple magnetic phases vanished at higher thermal energy, but the MTJMSD still maintained a higher magnetic moment because of anisotropy. The out-of-plane easy-axis anisotropy caused a dominant magnetic phase in the FM electrode rather than multiple magnetic phases. The simultaneous application of equal-magnitude in-plane and out-of-plane easy-axis anisotropies on the same electrode negated the anisotropy effect. Our experimental and MCS study provides insights for designing and understanding new spintronics-based devices.

42 ENGINEERING↗

Scalable Incremental Checkpointing using GPU-Accelerated De-Duplication

Writing large amounts of data concurrently to stable storage is a typical I/O pattern of many HPC workflows. This pattern introduces high I/O overheads and results in increased storage space utilization especially for workflows that need to capture the evolution of data structures with high frequency as checkpoints. In this context, many applications, such as graph pattern matching, perform sparse updates to large data structures between checkpoints. For these applications, incremental checkpointing techniques that save only the differences from one checkpoint to another can dramatically reduce the checkpoint sizes, I/O bottlenecks, and storage space utilization. However, such techniques are not without challenges: it is non-trivial to transparently determine what data has changed since a previous checkpoint and assemble the differences in a compact fashion that does not result in excessive metadata. State-of-art data reduction techniques (e.g., compression and de-duplication) have significant limitations when applied to modern HPC applications that leverage GPUs: slow at detecting the differences, generate a large amount of metadata to keep track of the differences, and ignore crucial spatiotemporal checkpoint data redundancy. This paper addresses these challenges by proposing a Merkle tree-based incremental checkpointing method to exploit GPUs' high memory bandwidth and massive parallelism. Experimental results at scale show a significant reduction of the I/O overhead and space utilization of checkpointing compared with state-of-the-art incremental checkpointing and compression techniques.

Tan, Nigel↗

Autonomous nondestructive evaluation of resistance spot welded joints

The application of non-destructive evaluation approaches has attracted strong interests in modern automotive industries. Here, we present an autonomous deep-computing framework to analyze raw videos from infrared systems and to predict weld nugget shape and size with unprecedented accuracy and speed. In a comprehensive training and testing experiment with 90 videos (seven sets of welding material stack-ups), a new method was developed to assemble sufficient datasets for neural network training. Our framework successfully predicts all the nugget shapes with F1 scores that range from 0.84 to 0.92. The total training time on Nvidia DGX station takes less than 10 min for each set of welding material stack-up. The real inference time of an individual dataset (with 30 video frames) takes about 0.005 s. The procedure and methods developed in the study can be applied to other image-based weld property prediction, as well as other manufacturing processes. Furthermore, our well-trained neural networks take limited memory resources (2.3 MB) and are suitable for embedded microprocessors for in-situ welding quality control as edge computing within an intelligent welding framework.

42 ENGINEERING↗

SpecSims: A Scalable Speculative Tree-based Simulation Cloning Framework for Finite Memory Machines

Simulation cloning is a technique in which cloned simulations whose state spaces differ partially from their parent simulation due to intervening events are spawned at runtime and concurrently advanced. It is a powerful method to carry out what-if analysis by speculatively exploring and evaluating the impact of various permutations of intervening cascade of events. Due to the exponential growth in the number of possible clones even for a small number of distinct intervening events, the practical efficacy of the approach is often severely limited by the maximum available memory of the computing host. In this paper, we introduce a novel speculative simulation cloning framework that executes a simulation cloning campaign capable of efficiently exploring an exponentially large space of clone simulations created by permutation of intervening events under a finite memory constraint. We provide a theoretical analysis of the runtime characteristics of our proposed approach and highlight its novel advantages such as memory-aware and as-long-as-needed execution. Furthermore, in support of our analytical findings and to demonstrate its practical feasibility, we implement a prototype of the cloning framework on a shared memory system and report its performance characteristics in the context of a heat diffusion simulation, and a power grid simulation subject to cascading disruptions from geomagnetic disturbances.

Simulation framework↗

Predicting weather impacts on corn production in a data-limited region using a transfer learning approach

The stability of food supply and prices may depend more on annual changes in yields from year-to-year variability in weather than on longer-term average changes from changing climatic conditions. However, the absence of high-quality data on crop yields at fine spatial resolutions in many regions of the world makes it challenging to statistically model their response to interannual variability in weather patterns. Therefore, there is a need for empirical methods that can project annual crop yield changes even in limited data regions. Here, we propose a transfer learning algorithm that uses high spatial resolution data from one region to project yields in another region with more limited data. The goal of our work is to understand what data types can be beneficial for transferring learning from a source region to a very different target region with more limited data. We utilize Long Short-Term Memory to develop a transfer learning model that is trained on historical county-level corn yield in the United States and predicts district-level corn yield variations in India. Even using smaller amounts of data in India, simulating a data-scarce region, we achieve an average root mean square error of 0.48 bu acre−1 in predicting interannual yield variations. Using Shapley values to interpret results, we explore the contribution of the different weather parameters to interannual yield variability and find a larger influence of precipitation-related variables. Our study demonstrates the usefulness of this method for transferring models of weather impacts on crop yields trained on a data-rich country to one with more limited data. It suggests the potential of applying the transfer learning model to mitigate the need for extensive raw data globally.

Vishwakarma, Srishti [ORNL] (ORCID:000000031674419↗

Understanding the Impact of Data Staging for Coupled Scientific Workflows

We report the rate of data generated by cutting-edge experimental science facilities and large-scale simulations enabled by current high-performance computing (HPC) systems has continued to grow at a far greater pace than the development of the network and storage capabilities on which these systems rely. To cope with this challenge, scientist are moving toward the creation of autonomous experiments and HPC simulations using machine learning. However, efficiently moving, storing, and processing large amounts of data away from the point of origin presents an incredible challenge. In-memory computing, in situ analysis, data staging, and data streaming are recognized viable alternatives to traditional file-based methods for transferring data between coupled workflows. However, the performance trade-offs and limitations for these methods are not fully understood when used in HPC applications. This article presents a comprehensive performance assessment of the current solutions for data staging when applied to applications that are not necessary I/O intensive which makes them not ideal candidates for these methods. Our study is based on experiments running at scale on Oak Ridge National Laboratory's Summit supercomputer using applications and simulations that cover typical computational motifs and patterns. We investigated the usability and cost/benefit trade-offs of staging algorithms for HPC applications under different scenarios and highlight opportunities for optimizing the dataflow between coupled simulation workflows.

97 MATHEMATICS AND COMPUTING↗

Train small, model big: Scalable physics simulators via reduced order modeling and domain decomposition

Numerous cutting-edge scientific technologies originate at the laboratory scale, but transitioning them to practical industry applications is a formidable challenge. Traditional pilot projects at intermediate scales are costly and time-consuming. An alternative, the pilot-scale model, relies on high-fidelity numerical simulations, but even these simulations can be computationally prohibitive at larger scales. To overcome these limitations, we propose a scalable, physics-constrained reduced order model (ROM) method. The ROM identifies critical physics modes from small-scale unit components, projecting governing equations onto these modes to create a reduced model that retains essential physics details. We also employ Discontinuous Galerkin Domain Decomposition (DG-DD) to apply ROM to unit components and interfaces, enabling the construction of large-scale global systems without data at such large scales. Here this method is demonstrated on the Poisson and Stokes flow equations, showing that it can solve equations about 15–40 times faster with only ~1% relative error. Furthermore, ROM takes one order of magnitude less memory than the full order model, enabling larger scale predictions at a given memory limitation.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Functionality Improvements to Overaero

The functionality of the overset, static aeroelasticity, Navier-Stokes flow solver OVERAERO was increased by adding capability to the flow solver and enhancing code performance. Improvements were made to the fluids/structure interface, an MLP version of the parallel OVERAERO code was developed, and the OVERAERO-MPI code was ported to the Cray T3E. The OVERFLOW-MPI and OVERAERO-MPI codes were tested successfully on the IPG testbed and a means of reducing communication overhead within OVERFLOW-MPI was investigated. To solve an aeroelastic problem computationally, a structures grid surface definition and a fluids grid surface definition are required. Typically, the structures grid surface has a lower fidelity than the fluids grid surface. Thus, the methods developed to transfer data between the two grid systems are vital to the accuracy and efficiency of the aeroelasticity code. The fluids/structures interface developed for the OVERAERO code was improved to more accurately treat fluids surfaces that bridge between two different structural surfaces. For example, the method allowed the forward portion of a flap track fairing to deform with the wing and the aft end of the fairing to deform with the flap. A tightly-coupled version of the code based on OVERFLOW-MLP was developed to improve code performance on the SGI Origin 2000. This required a new parallelization strategy to couple the fluids and structures codes. The OVERAERO-MPI code was ported to the Cray T3E to extend the usability of the code. The port required extensive use of dynamic memory management techniques to fit large problems within the memory limitations of the T3E. The OVERFLOW-MPI and OVERAERO-MPI codes were tested on the IPG testbed being developed within NASA. For small problems with minimal data transfer between grids, there was little to no performance penalty spreading the computation across two machines. For very large problems, methods were developed to minimize intermachine communication via the grid partitioning scheme. By minimizing the intermachine communication requirements of the problem, it may still be beneficial to run a tightly-coupled flow solver across two machines within the IPG.

Gee, Ken↗

Method and system for data clustering for very large databases

Multi-dimensional data contained in very large databases is efficiently and accurately clustered to determine patterns therein and extract useful information from such patterns. Conventional computer processors may be used which have limited memory capacity and conventional operating speed, allowing massive data sets to be processed in a reasonable time and with reasonable computer resources. The clustering process is organized using a clustering feature tree structure wherein each clustering feature comprises the number of data points in the cluster, the linear sum of the data points in the cluster, and the square sum of the data points in the cluster. A dense region of data points is treated collectively as a single cluster, and points in sparsely occupied regions can be treated as outliers and removed from the clustering feature tree. The clustering can be carried out continuously with new data points being received and processed, and with the clustering feature tree being restructured as necessary to accommodate the information from the newly received data points.

Zhang, Tian↗

Two–Photon Polymerized Shape Memory Microfibers: A New Mechanical Characterization Method in Liquid

Two-photon polymerization (TPP) is widely used to create 3D micro- and nanoscale scaffolds for biological and mechanobiological studies, which often require the mechanical characterization of the TPP fabricated structures. To satisfy physiological requirements, most of the mechanical characterizations need to be conducted in liquid. However, previous characterizations of TPP fabricated structures are all conducted in air due to the limitation of conventional micro- and nanoscale mechanical testing methods. In this study, a new experimental method is reported for testing the mechanical properties of TPP-printed microfibers in liquid. The experiments show that the mechanical behaviors of the microfibers tested in liquid are significantly different from those tested in air. By controlling the TPP writing parameters, the mechanical properties of the microfibers can be tailored over a wide range to meet a variety of mechanobiology applications. In addition, it is found that, in water, the plasticly deformed microfibers can return to their predeformed shape after tensile strain is released. The shape recovery time is dependent on the size of microfibers. The experimental method represents a significant advancement in mechanical testing of TPP fabricated structures and may help release the full potential of TPP fabricated 3D tissue scaffolds for mechanobiological studies.

36 MATERIALS SCIENCE↗

Wireless Patch Antenna Characterization for Live Health Monitoring Using Machine Learning

Temperature monitoring in extreme environments, such as coal-fired power plants, was addressed by designing and testing wireless patch antennas for use in machine learning-aided temperature estimation. The sensors were designed to monitor the temperature and health of boiler systems. Wireless interrogation of the sensor was performed using a Vector Network Analyzer (VNA) and a pair of interrogation antennas to capture resonance behavior under varying thermal and spatial conditions with sensitivities ranging from 0.052 to 0.20 $\frac{𝑀𝐻𝑧}{°C}$. Sensor calibration was conducted using a Long Short-Term Memory (LSTM) model, which leveraged temporal patterns to account for hysteresis effects. The calibration method demonstrated improved performance when combined with an LSTM model, achieving up to a 76% improvement in temperature estimation error when compared with Linear Regression (LR). The experiments highlighted an innovative solution for patch antenna-based non-contact temperature measurement, which addresses limitations with conventional methods such as RFID-based systems, infrared, and thermocouples.

20 FOSSIL-FUELED POWER PLANTS↗

The persistence of memory in ionic conduction probed by nonlinear optics

Predicting practical rates of transport in condensed phases enables the rational design of materials, devices and processes. This is especially critical to developing low-carbon energy technologies such as rechargeable batteries. For ionic conduction, the collective mechanisms, variation of conductivity with timescales and confinement, and ambiguity in the phononic origin of translation, call for a direct probe of the fundamental steps of ionic diffusion: ion hops. However, such hops are rare-event large-amplitude translations, and are challenging to excite and detect. Here we use single-cycle terahertz pumps to impulsively trigger ionic hopping in battery solid electrolytes. This is visualized by an induced transient birefringence, enabling direct probing of anisotropy in ionic hopping on the picosecond timescale. The relaxation of the transient signal measures the decay of orientational memory, and the production of entropy in diffusion. We extend experimental results using in silico transient birefringence to identify vibrational attempt frequencies for ion hopping. Using nonlinear optical methods, we probe ion transport at its fastest limit, distinguish correlated conduction mechanisms from a true random walk at the atomic scale, and demonstrate the connection between activated transport and the thermodynamics of information.

25 ENERGY STORAGE↗

Efficient packing of patterns in sparse distributed memory by selective weighting of input bits

When a set of patterns is stored in a distributed memory, any given storage location participates in the storage of many patterns. From the perspective of any one stored pattern, the other patterns act as noise, and such noise limits the memory's storage capacity. The more similar the retrieval cues for two patterns are, the more the patterns interfere with each other in memory, and the harder it is to separate them on retrieval. A method is described of weighting the retrieval cues to reduce such interference and thus to improve the separability of patterns that have similar cues.

Kanerva, Pentti↗

Scalable Heterogeneous Execution of a Coupled-Cluster Model with Perturbative Triples

The CCSD(T) coupled-cluster model with perturbative triples is considered a gold standard for computational modeling of the correlated behavior of electrons in molecular systems. A fundamental constraint is the relatively small global-memory capacity in GPUs compared to the main-memory capacity on host nodes, necessitating relatively smaller tile sizes for high-dimensional tensor contractions in NWChem's GPU-accelerated implementation of the CCSD(T) method. A coordinated redesign is described to address this limitation and associated data movement overheads, including a novel fused GPU kernel for a set of tensor contractions, along with inter-node communication optimization and data caching. The new implementation of GPU-accelerated CCSD(T) improves overall performance by 3.4x. Finally, we discuss the trade-offs in using this fused algorithm on current and future supercomputing platforms.

Kim, Jinsung↗

Design, characterization and shape recovery behavior of 3D/4D printed shape memory polymers (SMPs)

Shape memory polymers (SMPs) represent a paradigm shift in material science, uniquely capable of undergoing reversible shape transformations triggered by external stimuli, positioning them as pivotal in developing next-generation biomedical devices, aerospace components, and adaptive structures. Extensive research has been done on SMPs with a major focus on high-temperature programming methods, which can limit energy efficiency and applicability with temperature-sensitive materials. Additionally, while various SMP blends have demonstrated great potential, limited work has been done on the suitability for 3D printing these materials, particularly under high-strain and ambient temperature programming conditions. In this study, a three-component optimized SMP composition was evaluated by 3D printing via the Material Extrusion (MEX) technique and investigating its ambient temperature-programming behavior at high strains. The SMP formulation studied was a tailored blend of thermoplastic polyurethane (TPU), polycaprolactone (PCL), and an octadecane diol-based copolymer (OBC) that exhibits robust shape memory behavior, high strain tolerance, and efficient force generation. Rigorous thermal, mechanical, and shape recovery analyses, along with optimized printing parameters and consistent shape recovery rates of up to 90%, were achieved under dynamic mechanical analysis (DMA), even under ambient programming conditions. This work demonstrates the SMP composition’s potential for adaptive, self-deployable systems with 4D printing characteristics ideal for bio-inspired structures and artificial muscle fibers.

Sudan, Kavish [University of Louisville, KY]↗

Finite Element Analysis in Concurrent Processing: Computational Issues

The purpose of this research is to investigate the potential application of new methods for solving large-scale static structural problems on concurrent computers. It is well known that traditional single-processor computational speed will be limited by inherent physical limits. The only path to achieve higher computational speeds lies through concurrent processing. Traditional factorization solution methods for sparse matrices are ill suited for concurrent processing because the null entries get filled, leading to high communication and memory requirements. The research reported herein investigates alternatives to factorization that promise a greater potential to achieve high concurrent computing efficiency. Two methods, and their variants, based on direct energy minimization are studied: a) minimization of the strain energy using the displacement method formulation; b) constrained minimization of the complementary strain energy using the force method formulation. Initial results indicated that in the context of the direct energy minimization the displacement formulation experienced convergence and accuracy difficulties while the force formulation showed promising potential.

Sobieszczanski-Sobieski, Jaroslaw↗

Higher-Order Neural Networks Applied to 2D and 3D Object Recognition

A Higher-Order Neural Network (HONN) can be designed to be invariant to geometric transformations such as scale, translation, and in-plane rotation. Invariances are built directly into the architecture of a HONN and do not need to be learned. Thus, for 2D object recognition, the network needs to be trained on just one view of each object class, not numerous scaled, translated, and rotated views. Because the 2D object recognition task is a component of the 3D object recognition task, built-in 2D invariance also decreases the size of the training set required for 3D object recognition. We present results for 2D object recognition both in simulation and within a robotic vision experiment and for 3D object recognition in simulation. We also compare our method to other approaches and show that HONNs have distinct advantages for position, scale, and rotation-invariant object recognition. The major drawback of HONNs is that the size of the input field is limited due to the memory required for the large number of interconnections in a fully connected network. We present partial connectivity strategies and a coarse-coding technique for overcoming this limitation and increasing the input field to that required by practical object recognition problems.

Spirkovska, Lilly↗

Scaling Resolution of Gigapixel Whole Slide Images Using Spatial Decomposition on Convolutional Neural Networks

Gigapixel images are prevalent in scientific domains ranging from remote sensing, and satellite imagery to microscopy, etc. However, training a deep learning model at the natural resolution of those images has been a challenge in terms of both, overcoming the resource limit (e.g. HBM memory constraints), as well as scaling up to a large number of GPUs. In this paper, we trained Residual neural Networks (ResNet) on 22,528 x 22,528-pixel size images using a distributed spatial decomposition method on 2,304 GPUs on the Summit Supercomputer. We applied our method on a Whole Slide Imaging (WSI) dataset from The Cancer Genome Atlas (TCGA) database. WSI images can be in the size of 100,000 x 100,000 pixels or even larger, and in this work we studied the effect of image resolution on a classification task, while achieving state-of-the-art AUC scores. Moreover, our approach doesn't need pixel-level labels, since we're avoiding patching from the WSI images completely, while adding the capability of training arbitrary large-size images. This is achieved through a distributed spatial decomposition method, by leveraging the non-block fat-tree interconnect network of the Summit architecture, which enabled GPU-to-GPU direct communication. Finally, detailed performance analysis results are shown, as well as a comparison with a data-parallel approach when possible.

Tsaris, Aristeidis (aris)↗