Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Limited memory method”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

Path sampling of recurrent neural networks by incorporating known physics

Recurrent neural networks have seen widespread use in modeling dynamical systems in varied domains such as weather prediction, text prediction and several others. Often one wishes to supplement the experimentally observed dynamics with prior knowledge or intuition about the system. While the recurrent nature of these networks allows them to model arbitrarily long memories in the time series used in training, it makes it harder to impose prior knowledge or intuition through generic constraints. In this work, we present a path sampling approach based on principle of Maximum Caliber that allows us to include generic thermodynamic or kinetic constraints into recurrent neural networks. We show the method here for a widely used type of recurrent neural network known as long short-term memory network in the context of supplementing time series collected from different application domains. These include classical Molecular Dynamics of a protein and Monte Carlo simulations of an open quantum system continuously losing photons to the environment and displaying Rabi oscillations. Our method can be easily generalized to other generative artificial intelligence models and to generic time series in different areas of physical and social sciences, where one wishes to supplement limited data with intuition or theory based corrections.

59 BASIC BIOLOGICAL SCIENCES↗

Image processing tools for petabyte-scale light sheet microscopy data

Light sheet microscopy is a powerful technique for high-speed three-dimensional imaging of subcellular dynamics and large biological specimens. However, it often generates datasets ranging from hundreds of gigabytes to petabytes in size for a single experiment. Conventional computational tools process such images far slower than the time to acquire them and often fail outright due to memory limitations. To address these challenges, we present PetaKit5D, a scalable software solution for efficient petabyte-scale light sheet image processing. This software incorporates a suite of commonly used processing tools that are optimized for memory and performance. Notable advancements include rapid image readers and writers, fast and memory-efficient geometric transformations, high-performance Richardson–Lucy deconvolution and scalable Zarr-based stitching. These features outperform state-of-the-art methods by over one order of magnitude, enabling the processing of petabyte-scale image data at the full teravoxel rates of modern imaging cameras. The software opens new avenues for biological discoveries through large-scale imaging experiments.

97 MATHEMATICS AND COMPUTING↗

Radiation and scattering by cavity-backed antennas on a circular cylinder

Conformal arrays are popular antennas for aircraft and missile platforms due to their inherent low weight and drag properties. However, to date there has been a dearth of rigorous analytical or numerical solutions to aid the designer. In fact, it has been common practice to use limited measurements and planar approximations in designing such non-planar antennas. The finite element-boundary integral method is extended to scattering and radiation by cavity-backed structures in an infinite, metallic cylinder. In particular, the formulation specifics such as weight functions, dyadic Green's function, implementation details, and particular difficulties inherent to cylindrical structures are discussed. Special care is taken to ensure that the resulting computer program has low memory demand and minimal computational requirements. Both scattering and radiation parameters are computed and validated as much as possible.

Kempel, Leo C.↗

Hutchinson Trace Estimation for high-dimensional and high-order Physics-Informed Neural Networks

Physics-Informed Neural Networks (PINNs) have proven effective in solving partial differential equations (PDEs), especially when some data are available by seamlessly blending data and physics. However, extending PINNs to high-dimensional and even high-order PDEs encounters significant challenges due to the computational cost associated with automatic differentiation in the residual loss function calculation. Herein, we address the limitations of PINNs in handling high-dimensional and high-order PDEs by introducing the Hutchinson Trace Estimation (HTE) method. Starting with the second-order high-dimensional PDEs, which are ubiquitous in scientific computing, HTE is applied to transform the calculation of the entire Hessian matrix into a Hessian vector product (HVP). This approach not only alleviates the computational bottleneck via Taylor-mode automatic differentiation but also significantly reduces memory consumption from the Hessian matrix to an HVP’s scalar output. We further showcase HTE’s convergence to the original PINN loss and its unbiased behavior under specific conditions. Comparisons with the Stochastic Dimension Gradient Descent (SDGD) highlight the distinct advantages of HTE, particularly in scenarios with significant variability and variance among dimensions. We further extend the application of HTE to higher-order and higher-dimensional PDEs, specifically addressing the biharmonic equation. By employing tensor-vector products (TVP), HTE efficiently computes the colossal tensor associated with the fourth-order high-dimensional biharmonic equation, saving memory and enabling rapid computation. The effectiveness of HTE is illustrated through experimental setups, demonstrating comparable convergence rates with SDGD under memory and speed constraints. Additionally, HTE proves valuable in accelerating the Gradient-Enhanced PINN (gPINN) version as well as the Biharmonic equation. Overall, HTE opens up a new capability in scientific machine learning for tackling high-order and high-dimensional PDEs.

Curse of dimensionality↗

Unified many-worlds browsing of arbitrary physics-based animations

Manually tuning physics-based animation parameters to explore a simulation outcome space or achieve desired motion outcomes can be notoriously tedious. This problem has motivated many sophisticated and specialized optimization-based methods for fine-grained (keyframe) control, each of which are typically limited to specific animation phenomena, usually complicated, and, unfortunately, not widely used. In this paper, we propose Unified Many-Worlds Browsing (UMWB), a practical method for sample-level control and exploration of physics-based animations. Our approach supports browsing of large simulation ensembles of arbitrary animation phenomena by using a unified volumetric WORLDPACK representation based on spatiotemporally compressed voxel data associated with geometric occupancy and other low-fidelity animation state. Beyond memory reduction, the WORLDPACK representation also enables unified query support for interactive browsing: it provides fast evaluation of approximate spatiotemporal queries, such as occupancy tests that find ensemble samples ("worlds") where material is either IN or NOT IN a user-specified spacetime region. WORLDPACKS also support real-time hardware-accelerated voxel rendering by exploiting the spatially hierarchical and temporal RLE raster data structure. Our UMWB implementation supports interactive browsing (and offline refinement) of ensembles containing thousands of simulation samples, and fast spatiotemporal queries and ranking. We show UMWB results using a wide variety of physics-based animation phenomena---not just JELL-O ® .

Computer Science↗

Structured Adaptive Mesh Refinement Adaptations to Retain Performance Portability With Increasing Heterogeneity

Adaptive mesh refinement (AMR) is an important method that enables many mesh-based applications to run at effectively higher resolution within limited computing resources by allowing high resolution only where really needed. This advantage comes at a cost, however: greater complexity in the mesh management machinery and challenges with load distribution. With the current trend of increasing heterogeneity in hardware architecture, AMR presents an orthogonal axis of complexity. Additionally, the usual techniques, such as asynchronous communication and hierarchy management for parallelism and memory that are necessary to obtain reasonable performance are very challenging to reason about with AMR. Different groups working with AMR are bringing different approaches to this challenge. Here, we examine the design choices of several AMR codes and also the degree to which demands placed on them by their users influence these choices.

42 ENGINEERING↗

LPBF Processability of NiTiHf Alloys: Systematic Modeling and Single-Track Studies

Research into the processability of NiTiHf high-temperature shape memory alloys (HTSMAs) via laser powder bed fusion (LPBF) is limited; nevertheless, these alloys show promise for applications in extreme environments. This study aims to address this limitation by investigating the printability of four NiTiHf alloys with varying Hf content (1, 2, 15, and 20 at. %) to assess their suitability for LPBF applications. Solidification cracking is one of the main limiting factors in LPBF processes, which occurs during the final stage of solidification. To investigate the effect of alloy composition on printability, this study focuses on this defect via a combination of computational modeling and experimental validation. To this end, solidification cracking susceptibility is calculated as Kou’s index and Scheil–Gulliver model, implemented in Thermo-Calc/2022a software. An innovative powder-free experimental method through laser remelting was conducted on bare NiTiHf ingots to validate the parameter impacts of the LPBF process. The result is the processability window with no cracking likelihood under diverse LPBF conditions, including laser power and scan speed. This comprehensive investigation enhances our understanding of the processability challenges and opportunities for NiTiHf HTSMAs in advanced engineering applications.

36 MATERIALS SCIENCE↗

Knowledge Distillation for Anomaly Detection

Unsupervised deep learning techniques are widely used to identify anomalous behaviour. The performance of such methods is a product of the amount of training data and the model size. However, the size is often a limiting factor for the deployment on resource-constrained devices. Here, we present a novel procedure based on knowledge distillation for compressing an unsupervised anomaly detection model into a supervised deployable one and we suggest a set of techniques to improve the detection sensitivity. Compressed models perform comparably to their larger counterparts while significantly reducing the size and memory footprint.

Pol, Adrian Alan↗

An Integrated Framework for Memory-Centric Analysis: From Trace Collection to Co-Design

The memory wall phenomenon—where advances in processor performance significantly outpace those in memory subsystems—poses a fundamental challenge for contemporary computing systems. In memory-bound applications, memory subsystem behavior dominates performance, yet existing analysis approaches present significant limitations: detailed microarchitectural simulators require days to weeks to simulate modest workloads; hardware performance counters provide only aggregate statistics that obscure temporal and spatial access patterns; and scaled simulation approaches face challenges in capturing certain behaviors that emerge at larger scales. These limitations reflect a processor-centric design philosophy increasingly misaligned with memory-bound workloads where detailed understanding of memory access patterns, cache hierarchy interactions, and contention is critical for effective optimization. This paper presents an integrated framework for memory-centric analysis that enables effective hardware-software co-design. We describe practical trace collection techniques, including hardware-assisted processor tracing with minimal overhead and portable software-based instrumentation with statistical sampling. We present multi-perspective analysis methods that examine memory behavior from temporal, sequential, spatial, and relational viewpoints, revealing distinct optimization opportunities invisible in aggregate metrics. We detail an architectural modeling framework that uses sampled traces with temporal interpolation and confidence-based filtering to evaluate cache and memory configurations. Evaluation on representative benchmarks demonstrates that this framework achieves practical accuracy (L2 cache errors of 2.64\%, confidence-filtered L3 errors of 9.92\%, bandwidth errors of 7.33\%) while providing substantial speedup (26.8×) over cycle-accurate simulation, enabling rapid design space exploration. We demonstrate how this integrated framework enables systematic identification of both hardware optimizations (memory controller tuning, bank partitioning, NUMA configuration) and software optimizations (data layout restructuring, prefetching strategies, memory-aware scheduling). Through this comprehensive treatment of the memory-centric analysis pipeline—from trace collection through architectural modeling to co-design application—we provide researchers and practitioners with practical techniques for addressing memory bottlenecks in contemporary computing systems.

Gajaria, Dhruv Mayur↗

Autonomous Information Unit for Fine-Grain Data Access Control and Information Protection in a Net-Centric System

As communication and networking technologies advance, networks will become highly complex and heterogeneous, interconnecting different network domains. There is a need to provide user authentication and data protection in order to further facilitate critical mission operations, especially in the tactical and mission-critical net-centric networking environment. The Autonomous Information Unit (AIU) technology was designed to provide the fine-grain data access and user control in a net-centric system-testing environment to meet these objectives. The AIU is a fundamental capability designed to enable fine-grain data access and user control in the cross-domain networking environments, where an AIU is composed of the mission data, metadata, and policy. An AIU provides a mechanism to establish trust among deployed AIUs based on recombining shared secrets, authentication and verify users with a username, X.509 certificate, enclave information, and classification level. AIU achieves data protection through (1) splitting data into multiple information pieces using the Shamir's secret sharing algorithm, (2) encrypting each individual information piece using military-grade AES-256 encryption, and (3) randomizing the position of the encrypted data based on the unbiased and memory efficient in-place Fisher-Yates shuffle method. Therefore, it becomes virtually impossible for attackers to compromise data since attackers need to obtain all distributed information as well as the encryption key and the random seeds to properly arrange the data. In addition, since policy can be associated with data in the AIU, different user access and data control strategies can be included. The AIU technology can greatly enhance information assurance and security management in the bandwidth-limited and ad hoc net-centric environments. In addition, AIU technology can be applicable to general complex network domains and applications where distributed user authentication and data protection are necessary. AIU achieves fine-grain data access and user control, reducing the security risk significantly, simplifying the complexity of various security operations, and providing the high information assurance across different network domains.

Chow, Edward T.↗

Spectral interpolation - Zero fill or convolution

Zero fill, or augmentation by zeros, is a method used in conjunction with fast Fourier transforms to obtain spectral spacing at intervals closer than obtainable from the original input data set. In the present paper, an interpolation technique (interpolation by repetitive convolution) is proposed which yields values accurate enough for plotting purposes and which lie within the limits of calibration accuracies. The technique is shown to operate faster than zero fill, since fewer operations are required. The major advantages of interpolation by repetitive convolution are that efficient use of memory is possible (thus avoiding the difficulties encountered in decimation in time FFTs) and that is is easy to implement.

Forman, M. L.↗

Georgia Tech Accelerated, Compressed, and Regularized Compute of Kinetic-based PDEs (Final Report)

This report summarizes the collaborative effort between Lawrence Livermore National Laboratory and Georgia Tech to enhance the BoBa library for tensor train computation in PDE solvers, with a target on kinetic equations and their continuum limits. We aimed to reduce computational cost and memory usage by replacing traditional array-based computations with tensor trains. We examined the compressibility of time-evolving solutions to the Euler equations with discontinuities. We also explored using the first invsicid and linear regularization of the compressible flow equations via the information geometric regularization (IGR). We explored this in a tensor train formulation. To identify that inverse terms in the IGR equations pose problems for tensor train formulations and investigate efficient methods for batched inversion of tensor trains.

97 MATHEMATICS AND COMPUTING↗

Practical implementation of an accurate method for multilevel design sensitivity analysis

Solution techniques for handling large scale engineering optimization problems are reviewed. Potentials for practical applications as well as their limited capabilities are discussed. A new solution algorithm for design sensitivity is proposed. The algorithm is based upon the multilevel substructuring concept to be coupled with the adjoint method of sensitivity analysis. There are no approximations involved in the present algorithm except the usual approximations introduced due to the discretization of the finite element model. Results from the six- and thirty-bar planar truss problems show that the proposed multilevel scheme for sensitivity analysis is more effective (in terms of computer incore memory and the total CPU time) than a conventional (one level) scheme even on small problems. The new algorithm is expected to perform better for larger problems and its applications on the new generation of computer hardwares with 'parallel processing' capability is very promising.

Nguyen, Duc T.↗

Improved reduced-resolution satellite imagery

The resolution of satellite imagery is often traded-off to satisfy transmission time and bandwidth, memory, and display limitations. Although there are many ways to achieve the same reduction in resolution, algorithms vary in their ability to preserve the visual quality of the original imagery. These issues are investigated in the context of the Landsat browse system, which permits the user to preview a reduced resolution version of a Landsat image. Wavelets-based techniques for resolution reduction are proposed as alternatives to subsampling used in the current system. Experts judged imagery generated by the wavelets-based methods visually superior, confirming initial quantitative results. In particular, compared to subsampling, the wavelets-based techniques were much less likely to obscure roads, transmission lines, and other linear features present in the original image, introduce artifacts and noise, and otherwise reduce the usefulness of the image. The wavelets-based techniques afford multiple levels of resolution reduction and computational speed. This study is applicable to a wide range of reduced resolution applications in satellite imaging systems, including low resolution display, spaceborne browse, emergency image transmission, and real-time video downlinking.

Ellison, James↗

CrossLink: Advancements in Scalable Unstructured Mesh Generation [Slides]

Traditional mesh generation approaches are labor intensive and have limited robustness when applied to parametric design exploration and optimization of complex geometries. While automatic mesh generation approaches exist, they tend to generate tetrahedral or mixed-hybrid meshes which are generally unsuitable for physics applications with strong shock waves, thin boundary layers, and strong gradients. In addition, simulations sizes in the billions of cells are becoming more common with traditional mesh generation methods quickly reaching scalability limits. CrossLink offers a topology-based mesh generation approach with unstructured block-filling methods and a scalable mesh generation engine. In addition, CrossLink incorporates a python based API for seamless workflow integration and robust repeatability of the geometry handling and mesh generation process. This makes it ideal for parametric design study and optimization of complex geometries. Finally, future versions of CrossLink will offer a parametric mesh capability that optimizes a high-order mesh and enables reconstruction of the final mesh in memory by the physics solver.

97 MATHEMATICS AND COMPUTING↗

A Robust Hybrid Deep Learning Model for Spatiotemporal Image Fusion

Dense time-series remote sensing data with detailed spatial information are highly desired for the monitoring of dynamic earth systems. Due to the sensor tradeoff, most remote sensing systems cannot provide images with both high spatial and temporal resolutions. Spatiotemporal image fusion models provide a feasible solution to generate such a type of satellite imagery, yet existing fusion methods are limited in predicting rapid and/or transient phenological changes. Additionally, a systematic approach to assessing and understanding how varying levels of temporal phenological changes affect fusion results is lacking in spatiotemporal fusion research. The objective of this study is to develop an innovative hybrid deep learning model that can effectively and robustly fuse the satellite imagery of various spatial and temporal resolutions. The proposed model integrates two types of network models: super-resolution convolutional neural network (SRCNN) and long short-term memory (LSTM). SRCNN can enhance the coarse images by restoring degraded spatial details, while LSTM can learn and extract the temporal changing patterns from the time-series images. To systematically assess the effects of varying levels of phenological changes, we identify image phenological transition dates and design three temporal phenological change scenarios representing rapid, moderate, and minimal phenological changes. The hybrid deep learning model, alongside three benchmark fusion models, is assessed in different scenarios of phenological changes. Results indicate the hybrid deep learning model yields significantly better results when rapid or moderate phenological changes are present. It holds great potential in generating high-quality time-series datasets of both high spatial and temporal resolutions, which can further benefit terrestrial system dynamic studies. The innovative approach to understanding phenological changes’ effect will help us better comprehend the strengths and weaknesses of current and future fusion models.

spatiotemporal fusion↗

Deep learning-based spatio-temporal estimate of greenhouse gas emissions using satellite data

Accurate estimation of greenhouse gases (GHGs) emissions is very important for developing mitigation strategies to climate change by controlling and reducing GHG emissions. This project aims to develop multiple deep learning approaches to estimate anthropogenic greenhouse gas emissions using multiple types of satellite data. NO2 concentration is chosen as an example of GHGs to evaluate the proposed approach. Two sentinel satellites (sentinel-2 and sentinel-5P) provide multiscale observations of GHGs from 10-60m resolution (sentinel-2) to ~kilometer scale resolution (sentinel-5P). Among multiple deep learning (DL) architectures evaluated, two best DL models demonstrate that key features of spatio-temporal satellite data and additional information (e.g., observation times and/or coordinates of ground stations) can be extracted using convolutional neural networks and feed forward neural networks, respectively. In particular, irregular time series data from different NO 2 observation stations limit the flexibility of long short-term memory architecture, requiring zero-padding to fill in missing data. However, deep neural operator (DNO) architecture can stack time-series data as input, providing the flexibility of input structure without zero-padding. As a result, the DNO outperformed other deep learning architectures to account for time-varying features. Overall, temporal patterns with smooth seasonal variations were predicted very well, while frequent fluctuation patterns were not predicted well. In addition, uncertainty quantification using conformal inference method is performed to account for prediction ranges. Overall, this research will lead to a new groundwork for estimating greenhouse gas concentrations using multiple satellite data to enhance our capability of tracking the cause of climate change and developing mitigation strategies.

54 ENVIRONMENTAL SCIENCES↗

CrossLink: General Overview [Slides]

Problem: Traditional mesh generation approaches are labor intensive and have limited robustness when applied to parametric design exploration and optimization of complex geometries. While automatic mesh generation approaches exist, they tend to generate tetrahedral or mixed-hybrid meshes which are generally unsuitable for physics applications with strong shock waves, thin boundary layers, and strong gradients. In addition, simulations sizes in the billions of cells are becoming more common with traditional mesh generation methods quickly reaching scalability limits. Solution: CrossLink offers a topology-based mesh generation approach with unstructured block-filling methods and a scalable mesh generation engine. In addition, CrossLink incorporates a python-based API for seamless workflow integration and robust repeatability of the geometry handling and mesh generation process. This makes it ideal for parametric design study and optimization of complex geometries. Finally, future versions of CrossLink will offer a parametric mesh capability that optimizes a high-order mesh and enables reconstruction of the final mesh in memory by the physics solver.

97 MATHEMATICS AND COMPUTING↗