Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Limited memory method”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Autonomous nondestructive evaluation of resistance spot welded joints

The application of non-destructive evaluation approaches has attracted strong interests in modern automotive industries. Here, we present an autonomous deep-computing framework to analyze raw videos from infrared systems and to predict weld nugget shape and size with unprecedented accuracy and speed. In a comprehensive training and testing experiment with 90 videos (seven sets of welding material stack-ups), a new method was developed to assemble sufficient datasets for neural network training. Our framework successfully predicts all the nugget shapes with F1 scores that range from 0.84 to 0.92. The total training time on Nvidia DGX station takes less than 10 min for each set of welding material stack-up. The real inference time of an individual dataset (with 30 video frames) takes about 0.005 s. The procedure and methods developed in the study can be applied to other image-based weld property prediction, as well as other manufacturing processes. Furthermore, our well-trained neural networks take limited memory resources (2.3 MB) and are suitable for embedded microprocessors for in-situ welding quality control as edge computing within an intelligent welding framework.

42 ENGINEERING↗

SpecSims: A Scalable Speculative Tree-based Simulation Cloning Framework for Finite Memory Machines

Simulation cloning is a technique in which cloned simulations whose state spaces differ partially from their parent simulation due to intervening events are spawned at runtime and concurrently advanced. It is a powerful method to carry out what-if analysis by speculatively exploring and evaluating the impact of various permutations of intervening cascade of events. Due to the exponential growth in the number of possible clones even for a small number of distinct intervening events, the practical efficacy of the approach is often severely limited by the maximum available memory of the computing host. In this paper, we introduce a novel speculative simulation cloning framework that executes a simulation cloning campaign capable of efficiently exploring an exponentially large space of clone simulations created by permutation of intervening events under a finite memory constraint. We provide a theoretical analysis of the runtime characteristics of our proposed approach and highlight its novel advantages such as memory-aware and as-long-as-needed execution. Furthermore, in support of our analytical findings and to demonstrate its practical feasibility, we implement a prototype of the cloning framework on a shared memory system and report its performance characteristics in the context of a heat diffusion simulation, and a power grid simulation subject to cascading disruptions from geomagnetic disturbances.

Simulation framework↗

Predicting weather impacts on corn production in a data-limited region using a transfer learning approach

The stability of food supply and prices may depend more on annual changes in yields from year-to-year variability in weather than on longer-term average changes from changing climatic conditions. However, the absence of high-quality data on crop yields at fine spatial resolutions in many regions of the world makes it challenging to statistically model their response to interannual variability in weather patterns. Therefore, there is a need for empirical methods that can project annual crop yield changes even in limited data regions. Here, we propose a transfer learning algorithm that uses high spatial resolution data from one region to project yields in another region with more limited data. The goal of our work is to understand what data types can be beneficial for transferring learning from a source region to a very different target region with more limited data. We utilize Long Short-Term Memory to develop a transfer learning model that is trained on historical county-level corn yield in the United States and predicts district-level corn yield variations in India. Even using smaller amounts of data in India, simulating a data-scarce region, we achieve an average root mean square error of 0.48 bu acre−1 in predicting interannual yield variations. Using Shapley values to interpret results, we explore the contribution of the different weather parameters to interannual yield variability and find a larger influence of precipitation-related variables. Our study demonstrates the usefulness of this method for transferring models of weather impacts on crop yields trained on a data-rich country to one with more limited data. It suggests the potential of applying the transfer learning model to mitigate the need for extensive raw data globally.

Vishwakarma, Srishti [ORNL] (ORCID:000000031674419↗

Understanding the Impact of Data Staging for Coupled Scientific Workflows

We report the rate of data generated by cutting-edge experimental science facilities and large-scale simulations enabled by current high-performance computing (HPC) systems has continued to grow at a far greater pace than the development of the network and storage capabilities on which these systems rely. To cope with this challenge, scientist are moving toward the creation of autonomous experiments and HPC simulations using machine learning. However, efficiently moving, storing, and processing large amounts of data away from the point of origin presents an incredible challenge. In-memory computing, in situ analysis, data staging, and data streaming are recognized viable alternatives to traditional file-based methods for transferring data between coupled workflows. However, the performance trade-offs and limitations for these methods are not fully understood when used in HPC applications. This article presents a comprehensive performance assessment of the current solutions for data staging when applied to applications that are not necessary I/O intensive which makes them not ideal candidates for these methods. Our study is based on experiments running at scale on Oak Ridge National Laboratory's Summit supercomputer using applications and simulations that cover typical computational motifs and patterns. We investigated the usability and cost/benefit trade-offs of staging algorithms for HPC applications under different scenarios and highlight opportunities for optimizing the dataflow between coupled simulation workflows.

97 MATHEMATICS AND COMPUTING↗

NEAMS Milestone Report: M2MS–20OR030102 FW–CADIS PWR Ex-Core Analysis with Shift Through VERA

This report presents the work completed for the NEAMS milestone M2MS–20OR030102 titled "FW–CADIS PWR Ex-Core Analysis with Shift through VERA." The work completed for this milestone includes the implementation, integration, and optimization of memory and performance improvement methods in Shift for fully coupled ex-core calculations through Virtual Environment for Reactor Applications (VERA). Fully coupled in this context means the transfer of moderator boron concentration, pin-wise fission source, depleted compositions, temperatures, and moderator densities from MPACT (with COBRA-TF (CTF)) to Shift. The ability to run ex-core calculations with VERA has been enabled and used for several years by Consortium for Advanced Simulation of Light Water Reactors (CASL) partners. However, this implementation was limited and potentially computationally burdensome. This work has enabled the ability to run higher-fidelity ex-core calculations on moderate computing clusters by focusing on multithreading, domain decomposition, and Forward-Weighted CADIS (FW-CADIS) variance reduction. Tests performed on small cores, a small modular reactor (SMR), and CASL progression problems show very promising memory reduction and computational performance. Recommendations for settings when running fully coupled high-fidelity ex-core calculations with VERA are documented. Without these optimization methods, many processors on a compute node would be left unused for the entire ex-core calculation. Therefore, these methods enable the user to better use the resources available and reduce computation time.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Train small, model big: Scalable physics simulators via reduced order modeling and domain decomposition

Numerous cutting-edge scientific technologies originate at the laboratory scale, but transitioning them to practical industry applications is a formidable challenge. Traditional pilot projects at intermediate scales are costly and time-consuming. An alternative, the pilot-scale model, relies on high-fidelity numerical simulations, but even these simulations can be computationally prohibitive at larger scales. To overcome these limitations, we propose a scalable, physics-constrained reduced order model (ROM) method. The ROM identifies critical physics modes from small-scale unit components, projecting governing equations onto these modes to create a reduced model that retains essential physics details. We also employ Discontinuous Galerkin Domain Decomposition (DG-DD) to apply ROM to unit components and interfaces, enabling the construction of large-scale global systems without data at such large scales. Here this method is demonstrated on the Poisson and Stokes flow equations, showing that it can solve equations about 15–40 times faster with only ~1% relative error. Furthermore, ROM takes one order of magnitude less memory than the full order model, enabling larger scale predictions at a given memory limitation.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Two–Photon Polymerized Shape Memory Microfibers: A New Mechanical Characterization Method in Liquid

Two-photon polymerization (TPP) is widely used to create 3D micro- and nanoscale scaffolds for biological and mechanobiological studies, which often require the mechanical characterization of the TPP fabricated structures. To satisfy physiological requirements, most of the mechanical characterizations need to be conducted in liquid. However, previous characterizations of TPP fabricated structures are all conducted in air due to the limitation of conventional micro- and nanoscale mechanical testing methods. In this study, a new experimental method is reported for testing the mechanical properties of TPP-printed microfibers in liquid. The experiments show that the mechanical behaviors of the microfibers tested in liquid are significantly different from those tested in air. By controlling the TPP writing parameters, the mechanical properties of the microfibers can be tailored over a wide range to meet a variety of mechanobiology applications. In addition, it is found that, in water, the plasticly deformed microfibers can return to their predeformed shape after tensile strain is released. The shape recovery time is dependent on the size of microfibers. The experimental method represents a significant advancement in mechanical testing of TPP fabricated structures and may help release the full potential of TPP fabricated 3D tissue scaffolds for mechanobiological studies.

36 MATERIALS SCIENCE↗

Wireless Patch Antenna Characterization for Live Health Monitoring Using Machine Learning

Temperature monitoring in extreme environments, such as coal-fired power plants, was addressed by designing and testing wireless patch antennas for use in machine learning-aided temperature estimation. The sensors were designed to monitor the temperature and health of boiler systems. Wireless interrogation of the sensor was performed using a Vector Network Analyzer (VNA) and a pair of interrogation antennas to capture resonance behavior under varying thermal and spatial conditions with sensitivities ranging from 0.052 to 0.20 $\frac{𝑀𝐻𝑧}{°C}$. Sensor calibration was conducted using a Long Short-Term Memory (LSTM) model, which leveraged temporal patterns to account for hysteresis effects. The calibration method demonstrated improved performance when combined with an LSTM model, achieving up to a 76% improvement in temperature estimation error when compared with Linear Regression (LR). The experiments highlighted an innovative solution for patch antenna-based non-contact temperature measurement, which addresses limitations with conventional methods such as RFID-based systems, infrared, and thermocouples.

20 FOSSIL-FUELED POWER PLANTS↗

The persistence of memory in ionic conduction probed by nonlinear optics

Predicting practical rates of transport in condensed phases enables the rational design of materials, devices and processes. This is especially critical to developing low-carbon energy technologies such as rechargeable batteries. For ionic conduction, the collective mechanisms, variation of conductivity with timescales and confinement, and ambiguity in the phononic origin of translation, call for a direct probe of the fundamental steps of ionic diffusion: ion hops. However, such hops are rare-event large-amplitude translations, and are challenging to excite and detect. Here we use single-cycle terahertz pumps to impulsively trigger ionic hopping in battery solid electrolytes. This is visualized by an induced transient birefringence, enabling direct probing of anisotropy in ionic hopping on the picosecond timescale. The relaxation of the transient signal measures the decay of orientational memory, and the production of entropy in diffusion. We extend experimental results using in silico transient birefringence to identify vibrational attempt frequencies for ion hopping. Using nonlinear optical methods, we probe ion transport at its fastest limit, distinguish correlated conduction mechanisms from a true random walk at the atomic scale, and demonstrate the connection between activated transport and the thermodynamics of information.

25 ENERGY STORAGE↗

Scalable Heterogeneous Execution of a Coupled-Cluster Model with Perturbative Triples

The CCSD(T) coupled-cluster model with perturbative triples is considered a gold standard for computational modeling of the correlated behavior of electrons in molecular systems. A fundamental constraint is the relatively small global-memory capacity in GPUs compared to the main-memory capacity on host nodes, necessitating relatively smaller tile sizes for high-dimensional tensor contractions in NWChem's GPU-accelerated implementation of the CCSD(T) method. A coordinated redesign is described to address this limitation and associated data movement overheads, including a novel fused GPU kernel for a set of tensor contractions, along with inter-node communication optimization and data caching. The new implementation of GPU-accelerated CCSD(T) improves overall performance by 3.4x. Finally, we discuss the trade-offs in using this fused algorithm on current and future supercomputing platforms.

Kim, Jinsung↗

Design, characterization and shape recovery behavior of 3D/4D printed shape memory polymers (SMPs)

Shape memory polymers (SMPs) represent a paradigm shift in material science, uniquely capable of undergoing reversible shape transformations triggered by external stimuli, positioning them as pivotal in developing next-generation biomedical devices, aerospace components, and adaptive structures. Extensive research has been done on SMPs with a major focus on high-temperature programming methods, which can limit energy efficiency and applicability with temperature-sensitive materials. Additionally, while various SMP blends have demonstrated great potential, limited work has been done on the suitability for 3D printing these materials, particularly under high-strain and ambient temperature programming conditions. In this study, a three-component optimized SMP composition was evaluated by 3D printing via the Material Extrusion (MEX) technique and investigating its ambient temperature-programming behavior at high strains. The SMP formulation studied was a tailored blend of thermoplastic polyurethane (TPU), polycaprolactone (PCL), and an octadecane diol-based copolymer (OBC) that exhibits robust shape memory behavior, high strain tolerance, and efficient force generation. Rigorous thermal, mechanical, and shape recovery analyses, along with optimized printing parameters and consistent shape recovery rates of up to 90%, were achieved under dynamic mechanical analysis (DMA), even under ambient programming conditions. This work demonstrates the SMP composition’s potential for adaptive, self-deployable systems with 4D printing characteristics ideal for bio-inspired structures and artificial muscle fibers.

Sudan, Kavish [University of Louisville, KY]↗

Scaling Resolution of Gigapixel Whole Slide Images Using Spatial Decomposition on Convolutional Neural Networks

Gigapixel images are prevalent in scientific domains ranging from remote sensing, and satellite imagery to microscopy, etc. However, training a deep learning model at the natural resolution of those images has been a challenge in terms of both, overcoming the resource limit (e.g. HBM memory constraints), as well as scaling up to a large number of GPUs. In this paper, we trained Residual neural Networks (ResNet) on 22,528 x 22,528-pixel size images using a distributed spatial decomposition method on 2,304 GPUs on the Summit Supercomputer. We applied our method on a Whole Slide Imaging (WSI) dataset from The Cancer Genome Atlas (TCGA) database. WSI images can be in the size of 100,000 x 100,000 pixels or even larger, and in this work we studied the effect of image resolution on a classification task, while achieving state-of-the-art AUC scores. Moreover, our approach doesn't need pixel-level labels, since we're avoiding patching from the WSI images completely, while adding the capability of training arbitrary large-size images. This is achieved through a distributed spatial decomposition method, by leveraging the non-block fat-tree interconnect network of the Summit architecture, which enabled GPU-to-GPU direct communication. Finally, detailed performance analysis results are shown, as well as a comparison with a data-parallel approach when possible.

Tsaris, Aristeidis (aris)↗

Generalized quantum master equations can improve the accuracy of semiclassical predictions of multitime correlation functions

Multitime quantum correlation functions are central objects in physical science, offering a direct link between the experimental observables and the dynamics of an underlying model. While experiments such as 2D spectroscopy and quantum control can now measure such quantities, the accurate simulation of such responses remains computationally expensive and sometimes impossible, depending on the system’s complexity. A natural tool to employ is the generalized quantum master equation (GQME), which can offer computational savings by extending reference dynamics at a comparatively trivial cost. However, dynamical methods that can tackle chemical systems with atomistic resolution, such as those in the semiclassical hierarchy, often suffer from poor accuracy, limiting the credence one might lend to their results. By combining work on the accuracy-boosting formulation of semiclassical memory kernels with recent work on the multitime GQME, here we show for the first time that one can exploit a multitime semiclassical GQME to dramatically improve both the accuracy of coarse mean-field Ehrenfest dynamics and obtain orders of magnitude efficiency gains.

Chemistry↗

Reducing Memory Consumption in Calico with Shared Memory

This document details the work to reduce memory consumption in Calico. Calico is SimTools’ Constructive Solid Geometry (CSG) and geometry painting library. It is primarily used to paint material volume fractions in the Eulerian meshes of the physics codes. Calico provides point-in-body checks for the geometry supplied by an Oso model, which are then aggregated by the host codes. In addition, Calico can be used to build Oso models and is used by Ingen for that purpose. Oso models, and thus Calico, provide support for various CSG primitives such as spheres, cylinders, surfaces generated by rotating tabular curve data, and STL files as well as binary combinations of those primitives. Prior to refactoring Calico will run out of memory on CTS-1 machines when 36 MPI ranks are used per node when reading STL models on the order of 1.5 GB. This limitation is a bottleneck in designer workflow. This problem has been alleviated through the use of data structures to both reduce memory consumption and to leverage MPI-3 shared memory. This report details the data structures targeted for refactoring in Calico, the methods and implementation details for reducing memory consumption and leveraging shared memory, and results for one test problem. Results show a memory reduction when loading a 1 GB STL file by a factor of 27.5, from 93.4 to 3.4 GB.

97 MATHEMATICS AND COMPUTING↗

Ensemble Kalman filter for data assimilation coupled with low-resolution computations techniques applied in fluid dynamics

This paper presents an innovative Reduced-order model (ROM) for merging experimental and simulation data using data assimilation (DA) to estimate the "True" state of a fluid dynamics system, leading to more accurate predictions. Our methodology introduces a novel approach by implementing the ensemble Kalman filter (EnKF) within a reduced-dimensional framework, grounded in a robust theoretical foundation and applied to fluid dynamics. To address the substantial computational demands of DA, the proposed ROM employs low-resolution (LR) techniques to drastically reduce computational costs. This innovative approach involves downsampling datasets for DA computations, followed by an advanced reconstruction technique based on low-cost singular value decomposition (lcSVD). The lcSVD method, a key innovation in this paper, has never been applied to DA before and offers a highly efficient way to enhance resolution with minimal computational resources. Our results demonstrate significant reductions in both computation time and RAM usage through these LR techniques without compromising the accuracy of the estimations. For instance, in a turbulent test case, for a data compression rate of 15.9, the LR approach can achieve a speed-up of 13.7 and a RAM compression of 90.9% while maintaining a low relative root mean square error (RRMSE) of 2.6%, compared to 0.8% in the high-resolution (HR) reference. Furthermore, we highlight the effectiveness of the EnKF in estimating and predicting the state of fluid flow systems based on limited observations and given low-fidelity numerical data. This paper highlights the potential of the proposed DA method in fluid dynamics applications, particularly for improving computational efficiency in CFD and related fields. Its ability to balance accuracy with low computational and memory costs makes it especially suitable for large-scale and real-time applications, such as environmental monitoring or engineering design. This method will be incorporated into ModelFLOWs-app.

Data Assimilation↗

Addressing GPU memory limitations for Graph Neural Networks in High-Energy Physics applications

Introduction Reconstructing low-level particle tracks in neutrino physics can address some of the most fundamental questions about the universe. However, processing petabytes of raw data using deep learning techniques poses a challenging problem in the field of High Energy Physics (HEP). In the Exa.TrkX Project, an illustrative HEP application, preprocessed simulation data is fed into a state-of-art Graph Neural Network (GNN) model, accelerated by GPUs. However, limited GPU memory often leads to Out-of-Memory (OOM) exceptions during training, due to the large size of models and datasets. This problem is exacerbated when deploying models on High-Performance Computing (HPC) systems designed for large-scale applications. Methods We observe a high workload imbalance issue during GNN model training caused by the irregular sizes of input graph samples in HEP datasets, contributing to OOM exceptions. We aim to scale GNNs on HPC systems, by prioritizing workload balance in graph inputs while maintaining model accuracy. Our paper introduces diverse balancing strategies aimed at decreasing the maximum GPU memory footprint and avoiding the OOM exception, across various datasets. Results Our experiments showcase memory reduction of up to 32.14% compared to the baseline. We also demonstrate the proposed strategies can avoid OOM in application. Additionally, we create a distributed multi-GPU implementation using these samplers to demonstrate the scalability of these techniques on the HEP dataset. Discussion By assessing the performance of these strategies as data loading samplers across multiple datasets, we can gauge their effectiveness in both single-GPU and distributed environments. Our experiments, conducted on datasets of varying sizes and across multiple GPUs, broaden the applicability of our work to various GNN applications that handle input datasets with irregular graph sizes.

Lee, Claire Songhyun↗

HPDR: High-Performance Portable Scientific Data Reduction Framework

The rapid growth in scientific data generation is outpacing advancements in computing systems necessary for efficient storage, transfer, and analysis, particularly in the context of exascale computing. With the deployment of first-generation exascale computing systems and next-generation experimental facilities, this gap is widening and necessitates effective data reduction techniques to manage enormous data volumes. Over the past decade, various data reduction methods, including lossless compression, error-controlled lossy compression, and data refactoring, have been developed to accelerate I/O in scientific workflows. Despite significant reductions in data volume, these methods introduce considerable computational overhead, which can become the new bottleneck in data processing. To mitigate this, GPU-accelerated data reduction algorithms have been introduced. However, challenges remain in their integration into exascale workflows, including limited portability across different GPU architectures, substantial memory transfer overhead, and reduced scalability on dense multi-GPU systems. To address these challenges, we propose HPDR, a high-performance and portable data reduction framework. HPDR is designed to enable the execution of state-of-the-art reduction algorithms across diverse processor architectures while reducing memory transfer overhead to 2.3 % of the original, resulting in up to 3.5× faster throughput compared to existing solutions. It also achieves up to 96% of the theoretical speedup in multi-GPU settings. In addition, evaluations on accelerating I/O operations at scale up to 1,024 nodes of the Frontier supercomputer demonstrate that HPDR can achieve up to 103 TB/s reduction throughput, providing up to 4× acceleration in parallel I/O performance compared to existing data reduction routines. This work highlights the potential of HPDR to significantly enhance data reduction efficiency in exascale computing environments.

Chen, Jieyang [University of Oregon]↗

Analysis and Benchmarking of feature reduction for classification under computational constraints

Abstract Machine learning is most often expensive in terms of computational and memory costs due to training with large volumes of data. Current computational limitations of many computing systems motivate us to investigate practical approaches, such as feature selection and reduction, to reduce the time and memory costs while not sacrificing the accuracy of classification algorithms. In this work, we carefully review, analyze, and identify the feature reduction methods that have low costs/overheads in terms of time and memory. Then, we evaluate the identified reduction methods in terms of their impact on the accuracy, precision, time, and memory costs of traditional classification algorithms. Specifically, we focus on the least resource intensive feature reduction methods that are available in Scikit-Learn library. Since our goal is to identify the best performing low-cost reduction methods, we do not consider complex expensive reduction algorithms in this study. In our evaluation, we find that at quadratic-scale feature reduction, the classification algorithms achieve the best trade-off among competitive performance metrics. Results show that the overall training times are reduced 61%, the model sizes are reduced 6×, and accuracy scores increase 25% compared to the baselines on average with quadratic scale reduction.

97 MATHEMATICS AND COMPUTING↗