Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “limited memory method”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Predicting Kyasanur forest disease in resource-limited settings using event-based surveillance and transfer learning

In recent years, the reports of Kyasanur forest disease (KFD) breaking endemic barriers by spreading to new regions and crossing state boundaries is alarming. Effective disease surveillance and reporting systems are lacking for this emerging zoonosis, hence hindering control and prevention efforts. We compared time-series models using weather data with and without Event-Based Surveillance (EBS) information, i.e., news media reports and internet search trends, to predict monthly KFD cases in humans. We fitted Extreme Gradient Boosting (XGB) and Long Short-Term Memory models at the national and regional levels. We utilized the rich epidemiological data from endemic regions by applying Transfer Learning (TL) techniques to predict KFD cases in new outbreak regions where disease surveillance information was scarce. Overall, the inclusion of EBS data, in addition to the weather data, substantially increased the prediction performance across all models. The XGB method produced the best predictions at the national and regional levels. The TL techniques outperformed baseline models in predicting KFD in new outbreak regions. Novel sources of data and advanced machine-learning approaches, e.g., EBS and TL, show great potential towards increasing disease prediction capabilities in data-scarce scenarios and/or resource-limited settings, for better-informed decisions in the face of emerging zoonotic threats.

60 APPLIED LIFE SCIENCES↗

Resource-aware compression

Systems, apparatuses, and methods for implementing a multi-tiered approach to cache compression are disclosed. A cache includes a cache controller, light compressor, and heavy compressor. The decision on which compressor to use for compressing cache lines is made based on certain resource availability such as cache capacity or memory bandwidth. This allows the cache to opportunistically use complex algorithms for compression while limiting the adverse effects of high decompression latency on system performance. To address the above issue, the proposed design takes advantage of the heavy compressors for effectively reducing memory bandwidth in high bandwidth memory (HBM) interfaces as long as they do not sacrifice system performance. Accordingly, the cache combines light and heavy compressors with a decision-making unit to achieve reduced off-chip memory traffic without sacrificing system performance.

97 MATHEMATICS AND COMPUTING↗

Reduced scaling extended multi-state CASPT2 (XMS-CASPT2) using supporting subspaces and tensor hyper-contraction

We present a reduced scaling formulation of the extended multi-state CASPT2 (XMS-CASPT2) method, which is based on our recently developed state-specific CASPT2 (SS-CASPT2) formulation using supporting subspaces and tensor hyper-contraction. By using these two techniques, the off-diagonal elements of the effective Hamiltonian can be computed with only O(N 3 ) operations and O(N 2 ) memory, where N is the number of basis functions. Furthermore, this limits the overall computational scaling to O(N 4 ) operations and O(N 2 ) memory. Thus, excited states can now be obtained at the same reduced (relative to previous algorithms) scaling we achieved for SS-CASPT2.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

A compute-bound formulation of Galerkin model reduction for linear time-invariant dynamical systems

This work aims to advance computational methods for projection-based reduced-order models (ROMs) of linear time-invariant (LTI) dynamical systems. For such systems, current practice relies on ROM formulations expressing the state as a rank-1 tensor (i.e., a vector), leading to computational kernels that are memory bandwidth bound and, therefore, ill-suited for scalable performance on modern architectures. This weakness can be particularly limiting when tackling many-query studies, where one needs to run a large number of simulations. This work introduces a reformulation, called rank-2 Galerkin, of the Galerkin ROM for LTI dynamical systems which converts the nature of the ROM problem from memory bandwidth to compute bound. We present the details of the formulation and its implementation, and demonstrate its utility through numerical experiments using, as a test case, the simulation of elastic seismic shear waves in an axisymmetric domain. We quantify and analyze performance and scaling results for varying numbers of threads and problem sizes. In conclusion, we present an end-to-end demonstration of using the rank-2 Galerkin ROM for a Monte Carlo sampling study. We show that the rank-2 Galerkin ROM is one order of magnitude more efficient than the rank-1 Galerkin ROM (the current practice) and about 970 times more efficient than the full-order model, while maintaining accuracy in both the mean and statistics of the field.

97 MATHEMATICS AND COMPUTING↗

StressNet - Deep learning to predict stress with fracture propagation in brittle materials

Abstract Catastrophic failure in brittle materials is often due to the rapid growth and coalescence of cracks aided by high internal stresses. Hence, accurate prediction of maximum internal stress is critical to predicting time to failure and improving the fracture resistance and reliability of materials. Existing high-fidelity methods, such as the Finite-Discrete Element Model (FDEM), are limited by their high computational cost. Therefore, to reduce computational cost while preserving accuracy, a deep learning model, StressNet, is proposed to predict the entire sequence of maximum internal stress based on fracture propagation and the initial stress data. More specifically, the Temporal Independent Convolutional Neural Network (TI-CNN) is designed to capture the spatial features of fractures like fracture path and spall regions, and the Bidirectional Long Short-term Memory (Bi-LSTM) Network is adapted to capture the temporal features. By fusing these features, the evolution in time of the maximum internal stress can be accurately predicted. Moreover, an adaptive loss function is designed by dynamically integrating the Mean Squared Error (MSE) and the Mean Absolute Percentage Error (MAPE), to reflect the fluctuations in maximum internal stress. After training, the proposed model is able to compute accurate multi-step predictions of maximum internal stress in approximately 20 seconds, as compared to the FDEM run time of 4 h, with an average MAPE of 2% relative to test data.

36 MATERIALS SCIENCE↗

LLM-Based Adaptive Distribution Voltage Regulation Under Frequent Topology Changes: An In-Context MPC Framework

This paper proposes a large language model (LLM) based adaptive inverter control for distribution voltage regulation under frequent topology changes. We leverage the ability of the LLM to perform in-context learning and create a topology-adaptive surrogate model for power flow calculation. The surrogate model is then integrated with a long short-term memory-based load forecaster and a model predictive control (MPC) scheme to achieve the optimal inverter control that adapts to frequent topology changes. Unlike many existing works that assume fixed-topology grids or require the knowledge of all possible topologies when training a model, the proposed in-context MPC method tackles the distribution voltage control problem under various topologies and adapts to unknown topologies with limited data requirement for fine-tuning. The effectiveness of our method is demonstrated on a modified IEEE 123-bus test system.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Path sampling of recurrent neural networks by incorporating known physics

Recurrent neural networks have seen widespread use in modeling dynamical systems in varied domains such as weather prediction, text prediction and several others. Often one wishes to supplement the experimentally observed dynamics with prior knowledge or intuition about the system. While the recurrent nature of these networks allows them to model arbitrarily long memories in the time series used in training, it makes it harder to impose prior knowledge or intuition through generic constraints. In this work, we present a path sampling approach based on principle of Maximum Caliber that allows us to include generic thermodynamic or kinetic constraints into recurrent neural networks. We show the method here for a widely used type of recurrent neural network known as long short-term memory network in the context of supplementing time series collected from different application domains. These include classical Molecular Dynamics of a protein and Monte Carlo simulations of an open quantum system continuously losing photons to the environment and displaying Rabi oscillations. Our method can be easily generalized to other generative artificial intelligence models and to generic time series in different areas of physical and social sciences, where one wishes to supplement limited data with intuition or theory based corrections.

59 BASIC BIOLOGICAL SCIENCES↗

Image processing tools for petabyte-scale light sheet microscopy data

Light sheet microscopy is a powerful technique for high-speed three-dimensional imaging of subcellular dynamics and large biological specimens. However, it often generates datasets ranging from hundreds of gigabytes to petabytes in size for a single experiment. Conventional computational tools process such images far slower than the time to acquire them and often fail outright due to memory limitations. To address these challenges, we present PetaKit5D, a scalable software solution for efficient petabyte-scale light sheet image processing. This software incorporates a suite of commonly used processing tools that are optimized for memory and performance. Notable advancements include rapid image readers and writers, fast and memory-efficient geometric transformations, high-performance Richardson–Lucy deconvolution and scalable Zarr-based stitching. These features outperform state-of-the-art methods by over one order of magnitude, enabling the processing of petabyte-scale image data at the full teravoxel rates of modern imaging cameras. The software opens new avenues for biological discoveries through large-scale imaging experiments.

97 MATHEMATICS AND COMPUTING↗

Hutchinson Trace Estimation for high-dimensional and high-order Physics-Informed Neural Networks

Physics-Informed Neural Networks (PINNs) have proven effective in solving partial differential equations (PDEs), especially when some data are available by seamlessly blending data and physics. However, extending PINNs to high-dimensional and even high-order PDEs encounters significant challenges due to the computational cost associated with automatic differentiation in the residual loss function calculation. Herein, we address the limitations of PINNs in handling high-dimensional and high-order PDEs by introducing the Hutchinson Trace Estimation (HTE) method. Starting with the second-order high-dimensional PDEs, which are ubiquitous in scientific computing, HTE is applied to transform the calculation of the entire Hessian matrix into a Hessian vector product (HVP). This approach not only alleviates the computational bottleneck via Taylor-mode automatic differentiation but also significantly reduces memory consumption from the Hessian matrix to an HVP’s scalar output. We further showcase HTE’s convergence to the original PINN loss and its unbiased behavior under specific conditions. Comparisons with the Stochastic Dimension Gradient Descent (SDGD) highlight the distinct advantages of HTE, particularly in scenarios with significant variability and variance among dimensions. We further extend the application of HTE to higher-order and higher-dimensional PDEs, specifically addressing the biharmonic equation. By employing tensor-vector products (TVP), HTE efficiently computes the colossal tensor associated with the fourth-order high-dimensional biharmonic equation, saving memory and enabling rapid computation. The effectiveness of HTE is illustrated through experimental setups, demonstrating comparable convergence rates with SDGD under memory and speed constraints. Additionally, HTE proves valuable in accelerating the Gradient-Enhanced PINN (gPINN) version as well as the Biharmonic equation. Overall, HTE opens up a new capability in scientific machine learning for tackling high-order and high-dimensional PDEs.

Curse of dimensionality↗

Unified many-worlds browsing of arbitrary physics-based animations

Manually tuning physics-based animation parameters to explore a simulation outcome space or achieve desired motion outcomes can be notoriously tedious. This problem has motivated many sophisticated and specialized optimization-based methods for fine-grained (keyframe) control, each of which are typically limited to specific animation phenomena, usually complicated, and, unfortunately, not widely used. In this paper, we propose Unified Many-Worlds Browsing (UMWB), a practical method for sample-level control and exploration of physics-based animations. Our approach supports browsing of large simulation ensembles of arbitrary animation phenomena by using a unified volumetric WORLDPACK representation based on spatiotemporally compressed voxel data associated with geometric occupancy and other low-fidelity animation state. Beyond memory reduction, the WORLDPACK representation also enables unified query support for interactive browsing: it provides fast evaluation of approximate spatiotemporal queries, such as occupancy tests that find ensemble samples ("worlds") where material is either IN or NOT IN a user-specified spacetime region. WORLDPACKS also support real-time hardware-accelerated voxel rendering by exploiting the spatially hierarchical and temporal RLE raster data structure. Our UMWB implementation supports interactive browsing (and offline refinement) of ensembles containing thousands of simulation samples, and fast spatiotemporal queries and ranking. We show UMWB results using a wide variety of physics-based animation phenomena---not just JELL-O ® .

Computer Science↗

Structured Adaptive Mesh Refinement Adaptations to Retain Performance Portability With Increasing Heterogeneity

Adaptive mesh refinement (AMR) is an important method that enables many mesh-based applications to run at effectively higher resolution within limited computing resources by allowing high resolution only where really needed. This advantage comes at a cost, however: greater complexity in the mesh management machinery and challenges with load distribution. With the current trend of increasing heterogeneity in hardware architecture, AMR presents an orthogonal axis of complexity. Additionally, the usual techniques, such as asynchronous communication and hierarchy management for parallelism and memory that are necessary to obtain reasonable performance are very challenging to reason about with AMR. Different groups working with AMR are bringing different approaches to this challenge. Here, we examine the design choices of several AMR codes and also the degree to which demands placed on them by their users influence these choices.

42 ENGINEERING↗

LPBF Processability of NiTiHf Alloys: Systematic Modeling and Single-Track Studies

Research into the processability of NiTiHf high-temperature shape memory alloys (HTSMAs) via laser powder bed fusion (LPBF) is limited; nevertheless, these alloys show promise for applications in extreme environments. This study aims to address this limitation by investigating the printability of four NiTiHf alloys with varying Hf content (1, 2, 15, and 20 at. %) to assess their suitability for LPBF applications. Solidification cracking is one of the main limiting factors in LPBF processes, which occurs during the final stage of solidification. To investigate the effect of alloy composition on printability, this study focuses on this defect via a combination of computational modeling and experimental validation. To this end, solidification cracking susceptibility is calculated as Kou’s index and Scheil–Gulliver model, implemented in Thermo-Calc/2022a software. An innovative powder-free experimental method through laser remelting was conducted on bare NiTiHf ingots to validate the parameter impacts of the LPBF process. The result is the processability window with no cracking likelihood under diverse LPBF conditions, including laser power and scan speed. This comprehensive investigation enhances our understanding of the processability challenges and opportunities for NiTiHf HTSMAs in advanced engineering applications.

36 MATERIALS SCIENCE↗

Knowledge Distillation for Anomaly Detection

Unsupervised deep learning techniques are widely used to identify anomalous behaviour. The performance of such methods is a product of the amount of training data and the model size. However, the size is often a limiting factor for the deployment on resource-constrained devices. Here, we present a novel procedure based on knowledge distillation for compressing an unsupervised anomaly detection model into a supervised deployable one and we suggest a set of techniques to improve the detection sensitivity. Compressed models perform comparably to their larger counterparts while significantly reducing the size and memory footprint.

Pol, Adrian Alan↗

An Integrated Framework for Memory-Centric Analysis: From Trace Collection to Co-Design

The memory wall phenomenon—where advances in processor performance significantly outpace those in memory subsystems—poses a fundamental challenge for contemporary computing systems. In memory-bound applications, memory subsystem behavior dominates performance, yet existing analysis approaches present significant limitations: detailed microarchitectural simulators require days to weeks to simulate modest workloads; hardware performance counters provide only aggregate statistics that obscure temporal and spatial access patterns; and scaled simulation approaches face challenges in capturing certain behaviors that emerge at larger scales. These limitations reflect a processor-centric design philosophy increasingly misaligned with memory-bound workloads where detailed understanding of memory access patterns, cache hierarchy interactions, and contention is critical for effective optimization. This paper presents an integrated framework for memory-centric analysis that enables effective hardware-software co-design. We describe practical trace collection techniques, including hardware-assisted processor tracing with minimal overhead and portable software-based instrumentation with statistical sampling. We present multi-perspective analysis methods that examine memory behavior from temporal, sequential, spatial, and relational viewpoints, revealing distinct optimization opportunities invisible in aggregate metrics. We detail an architectural modeling framework that uses sampled traces with temporal interpolation and confidence-based filtering to evaluate cache and memory configurations. Evaluation on representative benchmarks demonstrates that this framework achieves practical accuracy (L2 cache errors of 2.64\%, confidence-filtered L3 errors of 9.92\%, bandwidth errors of 7.33\%) while providing substantial speedup (26.8×) over cycle-accurate simulation, enabling rapid design space exploration. We demonstrate how this integrated framework enables systematic identification of both hardware optimizations (memory controller tuning, bank partitioning, NUMA configuration) and software optimizations (data layout restructuring, prefetching strategies, memory-aware scheduling). Through this comprehensive treatment of the memory-centric analysis pipeline—from trace collection through architectural modeling to co-design application—we provide researchers and practitioners with practical techniques for addressing memory bottlenecks in contemporary computing systems.

Gajaria, Dhruv Mayur↗

Georgia Tech Accelerated, Compressed, and Regularized Compute of Kinetic-based PDEs (Final Report)

This report summarizes the collaborative effort between Lawrence Livermore National Laboratory and Georgia Tech to enhance the BoBa library for tensor train computation in PDE solvers, with a target on kinetic equations and their continuum limits. We aimed to reduce computational cost and memory usage by replacing traditional array-based computations with tensor trains. We examined the compressibility of time-evolving solutions to the Euler equations with discontinuities. We also explored using the first invsicid and linear regularization of the compressible flow equations via the information geometric regularization (IGR). We explored this in a tensor train formulation. To identify that inverse terms in the IGR equations pose problems for tensor train formulations and investigate efficient methods for batched inversion of tensor trains.

97 MATHEMATICS AND COMPUTING↗

CrossLink: Advancements in Scalable Unstructured Mesh Generation [Slides]

Traditional mesh generation approaches are labor intensive and have limited robustness when applied to parametric design exploration and optimization of complex geometries. While automatic mesh generation approaches exist, they tend to generate tetrahedral or mixed-hybrid meshes which are generally unsuitable for physics applications with strong shock waves, thin boundary layers, and strong gradients. In addition, simulations sizes in the billions of cells are becoming more common with traditional mesh generation methods quickly reaching scalability limits. CrossLink offers a topology-based mesh generation approach with unstructured block-filling methods and a scalable mesh generation engine. In addition, CrossLink incorporates a python based API for seamless workflow integration and robust repeatability of the geometry handling and mesh generation process. This makes it ideal for parametric design study and optimization of complex geometries. Finally, future versions of CrossLink will offer a parametric mesh capability that optimizes a high-order mesh and enables reconstruction of the final mesh in memory by the physics solver.

97 MATHEMATICS AND COMPUTING↗

Deep learning-based spatio-temporal estimate of greenhouse gas emissions using satellite data

Accurate estimation of greenhouse gases (GHGs) emissions is very important for developing mitigation strategies to climate change by controlling and reducing GHG emissions. This project aims to develop multiple deep learning approaches to estimate anthropogenic greenhouse gas emissions using multiple types of satellite data. NO2 concentration is chosen as an example of GHGs to evaluate the proposed approach. Two sentinel satellites (sentinel-2 and sentinel-5P) provide multiscale observations of GHGs from 10-60m resolution (sentinel-2) to ~kilometer scale resolution (sentinel-5P). Among multiple deep learning (DL) architectures evaluated, two best DL models demonstrate that key features of spatio-temporal satellite data and additional information (e.g., observation times and/or coordinates of ground stations) can be extracted using convolutional neural networks and feed forward neural networks, respectively. In particular, irregular time series data from different NO 2 observation stations limit the flexibility of long short-term memory architecture, requiring zero-padding to fill in missing data. However, deep neural operator (DNO) architecture can stack time-series data as input, providing the flexibility of input structure without zero-padding. As a result, the DNO outperformed other deep learning architectures to account for time-varying features. Overall, temporal patterns with smooth seasonal variations were predicted very well, while frequent fluctuation patterns were not predicted well. In addition, uncertainty quantification using conformal inference method is performed to account for prediction ranges. Overall, this research will lead to a new groundwork for estimating greenhouse gas concentrations using multiple satellite data to enhance our capability of tracking the cause of climate change and developing mitigation strategies.

54 ENVIRONMENTAL SCIENCES↗

CrossLink: General Overview [Slides]

Problem: Traditional mesh generation approaches are labor intensive and have limited robustness when applied to parametric design exploration and optimization of complex geometries. While automatic mesh generation approaches exist, they tend to generate tetrahedral or mixed-hybrid meshes which are generally unsuitable for physics applications with strong shock waves, thin boundary layers, and strong gradients. In addition, simulations sizes in the billions of cells are becoming more common with traditional mesh generation methods quickly reaching scalability limits. Solution: CrossLink offers a topology-based mesh generation approach with unstructured block-filling methods and a scalable mesh generation engine. In addition, CrossLink incorporates a python-based API for seamless workflow integration and robust repeatability of the geometry handling and mesh generation process. This makes it ideal for parametric design study and optimization of complex geometries. Finally, future versions of CrossLink will offer a parametric mesh capability that optimizes a high-order mesh and enables reconstruction of the final mesh in memory by the physics solver.

97 MATHEMATICS AND COMPUTING↗