Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “limited memory method”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Understanding the Impact of Data Staging for Coupled Scientific Workflows

We report the rate of data generated by cutting-edge experimental science facilities and large-scale simulations enabled by current high-performance computing (HPC) systems has continued to grow at a far greater pace than the development of the network and storage capabilities on which these systems rely. To cope with this challenge, scientist are moving toward the creation of autonomous experiments and HPC simulations using machine learning. However, efficiently moving, storing, and processing large amounts of data away from the point of origin presents an incredible challenge. In-memory computing, in situ analysis, data staging, and data streaming are recognized viable alternatives to traditional file-based methods for transferring data between coupled workflows. However, the performance trade-offs and limitations for these methods are not fully understood when used in HPC applications. This article presents a comprehensive performance assessment of the current solutions for data staging when applied to applications that are not necessary I/O intensive which makes them not ideal candidates for these methods. Our study is based on experiments running at scale on Oak Ridge National Laboratory's Summit supercomputer using applications and simulations that cover typical computational motifs and patterns. We investigated the usability and cost/benefit trade-offs of staging algorithms for HPC applications under different scenarios and highlight opportunities for optimizing the dataflow between coupled simulation workflows.

97 MATHEMATICS AND COMPUTING↗

Train small, model big: Scalable physics simulators via reduced order modeling and domain decomposition

Numerous cutting-edge scientific technologies originate at the laboratory scale, but transitioning them to practical industry applications is a formidable challenge. Traditional pilot projects at intermediate scales are costly and time-consuming. An alternative, the pilot-scale model, relies on high-fidelity numerical simulations, but even these simulations can be computationally prohibitive at larger scales. To overcome these limitations, we propose a scalable, physics-constrained reduced order model (ROM) method. The ROM identifies critical physics modes from small-scale unit components, projecting governing equations onto these modes to create a reduced model that retains essential physics details. We also employ Discontinuous Galerkin Domain Decomposition (DG-DD) to apply ROM to unit components and interfaces, enabling the construction of large-scale global systems without data at such large scales. Here this method is demonstrated on the Poisson and Stokes flow equations, showing that it can solve equations about 15–40 times faster with only ~1% relative error. Furthermore, ROM takes one order of magnitude less memory than the full order model, enabling larger scale predictions at a given memory limitation.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Two–Photon Polymerized Shape Memory Microfibers: A New Mechanical Characterization Method in Liquid

Two-photon polymerization (TPP) is widely used to create 3D micro- and nanoscale scaffolds for biological and mechanobiological studies, which often require the mechanical characterization of the TPP fabricated structures. To satisfy physiological requirements, most of the mechanical characterizations need to be conducted in liquid. However, previous characterizations of TPP fabricated structures are all conducted in air due to the limitation of conventional micro- and nanoscale mechanical testing methods. In this study, a new experimental method is reported for testing the mechanical properties of TPP-printed microfibers in liquid. The experiments show that the mechanical behaviors of the microfibers tested in liquid are significantly different from those tested in air. By controlling the TPP writing parameters, the mechanical properties of the microfibers can be tailored over a wide range to meet a variety of mechanobiology applications. In addition, it is found that, in water, the plasticly deformed microfibers can return to their predeformed shape after tensile strain is released. The shape recovery time is dependent on the size of microfibers. The experimental method represents a significant advancement in mechanical testing of TPP fabricated structures and may help release the full potential of TPP fabricated 3D tissue scaffolds for mechanobiological studies.

36 MATERIALS SCIENCE↗

Wireless Patch Antenna Characterization for Live Health Monitoring Using Machine Learning

Temperature monitoring in extreme environments, such as coal-fired power plants, was addressed by designing and testing wireless patch antennas for use in machine learning-aided temperature estimation. The sensors were designed to monitor the temperature and health of boiler systems. Wireless interrogation of the sensor was performed using a Vector Network Analyzer (VNA) and a pair of interrogation antennas to capture resonance behavior under varying thermal and spatial conditions with sensitivities ranging from 0.052 to 0.20 $\frac{𝑀𝐻𝑧}{°C}$. Sensor calibration was conducted using a Long Short-Term Memory (LSTM) model, which leveraged temporal patterns to account for hysteresis effects. The calibration method demonstrated improved performance when combined with an LSTM model, achieving up to a 76% improvement in temperature estimation error when compared with Linear Regression (LR). The experiments highlighted an innovative solution for patch antenna-based non-contact temperature measurement, which addresses limitations with conventional methods such as RFID-based systems, infrared, and thermocouples.

20 FOSSIL-FUELED POWER PLANTS↗

The persistence of memory in ionic conduction probed by nonlinear optics

Predicting practical rates of transport in condensed phases enables the rational design of materials, devices and processes. This is especially critical to developing low-carbon energy technologies such as rechargeable batteries. For ionic conduction, the collective mechanisms, variation of conductivity with timescales and confinement, and ambiguity in the phononic origin of translation, call for a direct probe of the fundamental steps of ionic diffusion: ion hops. However, such hops are rare-event large-amplitude translations, and are challenging to excite and detect. Here we use single-cycle terahertz pumps to impulsively trigger ionic hopping in battery solid electrolytes. This is visualized by an induced transient birefringence, enabling direct probing of anisotropy in ionic hopping on the picosecond timescale. The relaxation of the transient signal measures the decay of orientational memory, and the production of entropy in diffusion. We extend experimental results using in silico transient birefringence to identify vibrational attempt frequencies for ion hopping. Using nonlinear optical methods, we probe ion transport at its fastest limit, distinguish correlated conduction mechanisms from a true random walk at the atomic scale, and demonstrate the connection between activated transport and the thermodynamics of information.

25 ENERGY STORAGE↗

Scalable Heterogeneous Execution of a Coupled-Cluster Model with Perturbative Triples

The CCSD(T) coupled-cluster model with perturbative triples is considered a gold standard for computational modeling of the correlated behavior of electrons in molecular systems. A fundamental constraint is the relatively small global-memory capacity in GPUs compared to the main-memory capacity on host nodes, necessitating relatively smaller tile sizes for high-dimensional tensor contractions in NWChem's GPU-accelerated implementation of the CCSD(T) method. A coordinated redesign is described to address this limitation and associated data movement overheads, including a novel fused GPU kernel for a set of tensor contractions, along with inter-node communication optimization and data caching. The new implementation of GPU-accelerated CCSD(T) improves overall performance by 3.4x. Finally, we discuss the trade-offs in using this fused algorithm on current and future supercomputing platforms.

Kim, Jinsung↗

Design, characterization and shape recovery behavior of 3D/4D printed shape memory polymers (SMPs)

Shape memory polymers (SMPs) represent a paradigm shift in material science, uniquely capable of undergoing reversible shape transformations triggered by external stimuli, positioning them as pivotal in developing next-generation biomedical devices, aerospace components, and adaptive structures. Extensive research has been done on SMPs with a major focus on high-temperature programming methods, which can limit energy efficiency and applicability with temperature-sensitive materials. Additionally, while various SMP blends have demonstrated great potential, limited work has been done on the suitability for 3D printing these materials, particularly under high-strain and ambient temperature programming conditions. In this study, a three-component optimized SMP composition was evaluated by 3D printing via the Material Extrusion (MEX) technique and investigating its ambient temperature-programming behavior at high strains. The SMP formulation studied was a tailored blend of thermoplastic polyurethane (TPU), polycaprolactone (PCL), and an octadecane diol-based copolymer (OBC) that exhibits robust shape memory behavior, high strain tolerance, and efficient force generation. Rigorous thermal, mechanical, and shape recovery analyses, along with optimized printing parameters and consistent shape recovery rates of up to 90%, were achieved under dynamic mechanical analysis (DMA), even under ambient programming conditions. This work demonstrates the SMP composition’s potential for adaptive, self-deployable systems with 4D printing characteristics ideal for bio-inspired structures and artificial muscle fibers.

Sudan, Kavish [University of Louisville, KY]↗

Scaling Resolution of Gigapixel Whole Slide Images Using Spatial Decomposition on Convolutional Neural Networks

Gigapixel images are prevalent in scientific domains ranging from remote sensing, and satellite imagery to microscopy, etc. However, training a deep learning model at the natural resolution of those images has been a challenge in terms of both, overcoming the resource limit (e.g. HBM memory constraints), as well as scaling up to a large number of GPUs. In this paper, we trained Residual neural Networks (ResNet) on 22,528 x 22,528-pixel size images using a distributed spatial decomposition method on 2,304 GPUs on the Summit Supercomputer. We applied our method on a Whole Slide Imaging (WSI) dataset from The Cancer Genome Atlas (TCGA) database. WSI images can be in the size of 100,000 x 100,000 pixels or even larger, and in this work we studied the effect of image resolution on a classification task, while achieving state-of-the-art AUC scores. Moreover, our approach doesn't need pixel-level labels, since we're avoiding patching from the WSI images completely, while adding the capability of training arbitrary large-size images. This is achieved through a distributed spatial decomposition method, by leveraging the non-block fat-tree interconnect network of the Summit architecture, which enabled GPU-to-GPU direct communication. Finally, detailed performance analysis results are shown, as well as a comparison with a data-parallel approach when possible.

Tsaris, Aristeidis (aris)↗

Imaging the magnetic nanowire cross section and magnetic ordering within a suspended 3D artificial spin-ice

Artificial spin-ice systems are patterned arrays of magnetic nanoislands arranged into frustrated geometries and provide insight into the physics of ordering and emergence. The majority of these systems have been realized in two-dimensions, mainly due to the ease of fabrication, but with recent developments in advanced nanolithography, three-dimensional artificial spin ice (ASI) structures have become possible, providing a new paradigm in their study. Such artificially engineered 3D systems provide new opportunities in realizing tunable ground states, new domain wall topologies, monopole propagation, and advanced device concepts, such as magnetic racetrack memory. Direct imaging of 3DASI structures with magnetic force microscopy has thus far been key to probing the physics of these systems but is limited in both the depth of measurement and resolution, ultimately restricting measurement to the uppermost layers of the system. In this work, a method is developed to fabricate 3DASI lattices over an aperture using two-photon lithography, thermal evaporation, and oxygen plasma exposure, allowing the probe of element-specific structural and magnetic information using soft x-ray microscopy with x-ray magnetic circular dichroism (XMCD) as magnetic contrast. The suspended polymer–permalloy lattices are found to be stable under repeated soft x-ray exposure. Analysis of the x-ray absorption signal allows the complex cross section of the magnetic nanowires to be reconstructed and demonstrates a crescent-shaped geometry. Measurement of the XMCD images after the application of an in-plane field suggests a decrease in magnetic moment on the lattice surface due to oxidation, while a measurable signal is retained on sub-lattices below the surface.

36 MATERIALS SCIENCE↗

Generalized quantum master equations can improve the accuracy of semiclassical predictions of multitime correlation functions

Multitime quantum correlation functions are central objects in physical science, offering a direct link between the experimental observables and the dynamics of an underlying model. While experiments such as 2D spectroscopy and quantum control can now measure such quantities, the accurate simulation of such responses remains computationally expensive and sometimes impossible, depending on the system’s complexity. A natural tool to employ is the generalized quantum master equation (GQME), which can offer computational savings by extending reference dynamics at a comparatively trivial cost. However, dynamical methods that can tackle chemical systems with atomistic resolution, such as those in the semiclassical hierarchy, often suffer from poor accuracy, limiting the credence one might lend to their results. By combining work on the accuracy-boosting formulation of semiclassical memory kernels with recent work on the multitime GQME, here we show for the first time that one can exploit a multitime semiclassical GQME to dramatically improve both the accuracy of coarse mean-field Ehrenfest dynamics and obtain orders of magnitude efficiency gains.

Chemistry↗

Reducing Memory Consumption in Calico with Shared Memory

This document details the work to reduce memory consumption in Calico. Calico is SimTools’ Constructive Solid Geometry (CSG) and geometry painting library. It is primarily used to paint material volume fractions in the Eulerian meshes of the physics codes. Calico provides point-in-body checks for the geometry supplied by an Oso model, which are then aggregated by the host codes. In addition, Calico can be used to build Oso models and is used by Ingen for that purpose. Oso models, and thus Calico, provide support for various CSG primitives such as spheres, cylinders, surfaces generated by rotating tabular curve data, and STL files as well as binary combinations of those primitives. Prior to refactoring Calico will run out of memory on CTS-1 machines when 36 MPI ranks are used per node when reading STL models on the order of 1.5 GB. This limitation is a bottleneck in designer workflow. This problem has been alleviated through the use of data structures to both reduce memory consumption and to leverage MPI-3 shared memory. This report details the data structures targeted for refactoring in Calico, the methods and implementation details for reducing memory consumption and leveraging shared memory, and results for one test problem. Results show a memory reduction when loading a 1 GB STL file by a factor of 27.5, from 93.4 to 3.4 GB.

97 MATHEMATICS AND COMPUTING↗

Ensemble Kalman filter for data assimilation coupled with low-resolution computations techniques applied in fluid dynamics

This paper presents an innovative Reduced-order model (ROM) for merging experimental and simulation data using data assimilation (DA) to estimate the "True" state of a fluid dynamics system, leading to more accurate predictions. Our methodology introduces a novel approach by implementing the ensemble Kalman filter (EnKF) within a reduced-dimensional framework, grounded in a robust theoretical foundation and applied to fluid dynamics. To address the substantial computational demands of DA, the proposed ROM employs low-resolution (LR) techniques to drastically reduce computational costs. This innovative approach involves downsampling datasets for DA computations, followed by an advanced reconstruction technique based on low-cost singular value decomposition (lcSVD). The lcSVD method, a key innovation in this paper, has never been applied to DA before and offers a highly efficient way to enhance resolution with minimal computational resources. Our results demonstrate significant reductions in both computation time and RAM usage through these LR techniques without compromising the accuracy of the estimations. For instance, in a turbulent test case, for a data compression rate of 15.9, the LR approach can achieve a speed-up of 13.7 and a RAM compression of 90.9% while maintaining a low relative root mean square error (RRMSE) of 2.6%, compared to 0.8% in the high-resolution (HR) reference. Furthermore, we highlight the effectiveness of the EnKF in estimating and predicting the state of fluid flow systems based on limited observations and given low-fidelity numerical data. This paper highlights the potential of the proposed DA method in fluid dynamics applications, particularly for improving computational efficiency in CFD and related fields. Its ability to balance accuracy with low computational and memory costs makes it especially suitable for large-scale and real-time applications, such as environmental monitoring or engineering design. This method will be incorporated into ModelFLOWs-app.

Data Assimilation↗

Addressing GPU memory limitations for Graph Neural Networks in High-Energy Physics applications

Introduction Reconstructing low-level particle tracks in neutrino physics can address some of the most fundamental questions about the universe. However, processing petabytes of raw data using deep learning techniques poses a challenging problem in the field of High Energy Physics (HEP). In the Exa.TrkX Project, an illustrative HEP application, preprocessed simulation data is fed into a state-of-art Graph Neural Network (GNN) model, accelerated by GPUs. However, limited GPU memory often leads to Out-of-Memory (OOM) exceptions during training, due to the large size of models and datasets. This problem is exacerbated when deploying models on High-Performance Computing (HPC) systems designed for large-scale applications. Methods We observe a high workload imbalance issue during GNN model training caused by the irregular sizes of input graph samples in HEP datasets, contributing to OOM exceptions. We aim to scale GNNs on HPC systems, by prioritizing workload balance in graph inputs while maintaining model accuracy. Our paper introduces diverse balancing strategies aimed at decreasing the maximum GPU memory footprint and avoiding the OOM exception, across various datasets. Results Our experiments showcase memory reduction of up to 32.14% compared to the baseline. We also demonstrate the proposed strategies can avoid OOM in application. Additionally, we create a distributed multi-GPU implementation using these samplers to demonstrate the scalability of these techniques on the HEP dataset. Discussion By assessing the performance of these strategies as data loading samplers across multiple datasets, we can gauge their effectiveness in both single-GPU and distributed environments. Our experiments, conducted on datasets of varying sizes and across multiple GPUs, broaden the applicability of our work to various GNN applications that handle input datasets with irregular graph sizes.

Lee, Claire Songhyun↗

HPDR: High-Performance Portable Scientific Data Reduction Framework

The rapid growth in scientific data generation is outpacing advancements in computing systems necessary for efficient storage, transfer, and analysis, particularly in the context of exascale computing. With the deployment of first-generation exascale computing systems and next-generation experimental facilities, this gap is widening and necessitates effective data reduction techniques to manage enormous data volumes. Over the past decade, various data reduction methods, including lossless compression, error-controlled lossy compression, and data refactoring, have been developed to accelerate I/O in scientific workflows. Despite significant reductions in data volume, these methods introduce considerable computational overhead, which can become the new bottleneck in data processing. To mitigate this, GPU-accelerated data reduction algorithms have been introduced. However, challenges remain in their integration into exascale workflows, including limited portability across different GPU architectures, substantial memory transfer overhead, and reduced scalability on dense multi-GPU systems. To address these challenges, we propose HPDR, a high-performance and portable data reduction framework. HPDR is designed to enable the execution of state-of-the-art reduction algorithms across diverse processor architectures while reducing memory transfer overhead to 2.3 % of the original, resulting in up to 3.5× faster throughput compared to existing solutions. It also achieves up to 96% of the theoretical speedup in multi-GPU settings. In addition, evaluations on accelerating I/O operations at scale up to 1,024 nodes of the Frontier supercomputer demonstrate that HPDR can achieve up to 103 TB/s reduction throughput, providing up to 4× acceleration in parallel I/O performance compared to existing data reduction routines. This work highlights the potential of HPDR to significantly enhance data reduction efficiency in exascale computing environments.

Chen, Jieyang [University of Oregon]↗

Analysis and Benchmarking of feature reduction for classification under computational constraints

Abstract Machine learning is most often expensive in terms of computational and memory costs due to training with large volumes of data. Current computational limitations of many computing systems motivate us to investigate practical approaches, such as feature selection and reduction, to reduce the time and memory costs while not sacrificing the accuracy of classification algorithms. In this work, we carefully review, analyze, and identify the feature reduction methods that have low costs/overheads in terms of time and memory. Then, we evaluate the identified reduction methods in terms of their impact on the accuracy, precision, time, and memory costs of traditional classification algorithms. Specifically, we focus on the least resource intensive feature reduction methods that are available in Scikit-Learn library. Since our goal is to identify the best performing low-cost reduction methods, we do not consider complex expensive reduction algorithms in this study. In our evaluation, we find that at quadratic-scale feature reduction, the classification algorithms achieve the best trade-off among competitive performance metrics. Results show that the overall training times are reduced 61%, the model sizes are reduced 6×, and accuracy scores increase 25% compared to the baselines on average with quadratic scale reduction.

97 MATHEMATICS AND COMPUTING↗

HAM: Hotspot-Aware Manager for Improving Communications with 3D-Stacked Memory

merging High-Performance Computing (HPC) workloads, such as graph analytics, machine learning, and big data science, are data-intensive. Data-intensive workloads usually present fine-grained memory accesses with limited or no data locality, and thus incur frequent cache misses and low utilization of memory bandwidth. 3D-stacked memory devices such as Hybrid Memory Cube (HMC) and High Bandwidth Memory (HBM) can provide significantly higher bandwidth than conventional memory modules. However, the traditional interfaces and optimization methods for JEDEC DDR devices do not allow to fully exploit the potential performance of 3D-stacked memory with the massive amount of irregular memory accesses of data-intensive applications. In this paper, we propose a novel Hotspot-Aware Manager (HAM) infrastructure for 3D-stacked memory devices capable of optimizing memory access streams via request aggregation, hotspot detection, and in-memory prefetching. %and an associated hotspot-aware page policy. We present the HAM design and implementation, and simulate it on a system using RISC-V embedded cores with attached HMC devices. We extensively evaluate HAM with over 12 benchmarks and applications representing diverse irregular memory access patterns. The results show that, on average, HAM reduces redundant requests by 37.51\% and increases the prefetch buffer hit rate by 4.2 times, compared to a baseline streaming prefetcher. On the selected benchmark set, HAM provides performance gains of 21.81\% in average (up to 34.28\%) and power savings of 35.07\% over a standard 3D-stacked memory.

Wang, Xi↗

Design and analysis of CXL performance models for tightly-coupled heterogeneous computing

Truly heterogeneous systems enable partitioned workloads to be mapped to the hardware that nets the best performance. However, current practice requires that inter-device communication between different vendors' hardware use host memory as an intermediary step. To date, there are no widely adopted solutions that allow accelerators to directly transfer data. A new cache-coherent protocol, CXL, aims to facilitate easier, fine-grained sharing between accelerators. In this work we analyze existing methods for designing heterogeneous applications that target GPUs and FPGAs working collaboratively, followed by an exploration to show the benefits of a CXL-enabled system. Specifically, we develop a test application that utilizes both an NVIDIA P100 GPU and a Xilinx U250 FPGA to show current communication limitations. From this application, we capture overall execution time and throughput measurements on the FPGA and GPU. We use these measurements as inputs to novel CXL performance models to show that using CXL caching instead of host memory results in a 1.31X speedup, while a more tightly-coupled pipelined implementation using CXL-enabled hardware would result in a speedup of 1.45X.

Cabrera, Anthony↗

Distributed Tomographic Reconstruction with Quantization

Conventional tomographic reconstruction typically depends on centralized servers for both data storage and computation, leading to concerns about memory limitations and data privacy. Distributed reconstruction algorithms mitigate these issues by partitioning data across multiple nodes, reducing server load and enhancing privacy. However, these algorithms often encounter challenges related to memory constraints and communication overhead between nodes. In this paper, we introduce a decentralized Alternating Directions Method of Multipliers (ADMM) with configurable quantization. By distributing local objectives across nodes, our approach is highly scalable and can efficiently reconstruct images while adapting to available resources. To overcome communication bottlenecks, we propose two quantization techniques based on K-means clustering and JPEG compression. Numerical experiments with benchmark images illustrate the tradeoffs between communication efficiency, memory use, and reconstruction accuracy.

Miao, Runxuan↗