Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Limited memory method”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Randomized Sketching Algorithms for Low-Memory Dynamic Optimization

This paper develops a novel limited-memory method to solve dynamic optimization problems. The memory requirements for such problems often present a major obstacle, particularly for problems with PDE constraints such as optimal flow control, full waveform inversion, and optical tomography. In these problems, PDE constraints uniquely determine the state of a physical system for a given control; the goal is to find the value of the control that minimizes an objective. While the control is often low dimensional, the state is typically more expensive to store. This paper suggests using randomized matrix approximation to compress the state as it is generated and shows how to use the compressed state to reliably solve the original dynamic optimization problem. Concretely, the compressed state is used to compute approximate gradients and to apply the Hessian to vectors. The approximation error in these quantities is controlled by the target rank of the sketch. This approximate first- and second-order information can readily be used in any optimization algorithm. As an example, we develop a sketched trust-region method that adaptively chooses the target rank using a posteriori error information and provably converges to a stationary point of the original problem. Numerical experiments with the sketched trust-region method show promising performance on challenging problems such as the optimal control of an advection-reaction-diffusion equation and the optimal control of fluid flow past a cylinder.

97 MATHEMATICS AND COMPUTING↗

Sequence length scaling in vision transformers for scientific images on frontier

Vision Transformers (ViTs) are pivotal for foundational models in scientific imagery, including Earth science applications, due to their capability to process large sequence lengths. While transformers for text have inspired scaling sequence lengths in ViTs, adapting these for ViTs introduces unique challenges. We develop distributed sequence parallelism for ViTs, enabling them to handle up to 1M tokens. Our approach, leveraging DeepSpeed-Ulysses and Long-Sequence-Segmentation with model sharding, is the first to apply sequence parallelism in ViT training, achieving a 94% batch scaling efficiency on 2,048 AMD-MI250X GPUs. Evaluating sequence parallelism in ViTs, particularly in models up to 10B parameters, highlighted substantial bottlenecks. We countered these with hybrid sequence, pipeline, and flash attention strategies, to scale beyond single GPU memory limits. Our method significantly enhances climate modeling accuracy by 20% in temperature predictions, marking the first training of a vision transformer model to convergence with a sequence length of 188K tokens, using full self-attention.

Tsaris, Aristeidis (aris) [ORNL] (ORCID:0000000277↗

Compact representations of structured BFGS matrices

For general large-scale optimization problems compact representations exist in which recursive quasi-Newton update formulas are represented as compact matrix factorizations. For problems in which the objective function contains additional structure, recent structured quasi-Newton methods exploit available second-derivative information and approximate unavailable second derivatives. Here, this article develops the compact representations of two structured Broyden-Fletcher-Goldfarb-Shanno update formulas. The compact representations enable efficient limited memory and initialization strategies. Two limited memory line search algorithms are described for which extensive numerical results demonstrate the efficacy of the algorithms, including comparisons to IPOPT on large machine learning problems, and to L-BFGS on a real world large scale ptychographic imaging application.

97 MATHEMATICS AND COMPUTING↗

Design Space Exploration of Emerging Memory Technologies for Machine Learning Applications

Memory design space exploration methods study memory systems’ performances and limitations before implementation. The computer memory design space has grown exponentially because of the enormous growth of memory types, memory controllers, and application software. Computer simulators are commonly used for memory design space exploration. However, complex memory simulations take an enormous amount of time. Hence, in this paper, we proposed a machine learning-based design space exploration method for dynamic random-access memory and non-volatile memory systems. We applied our method to the CosmoGAN and LeNet applications to predict the following six memory response parameters: (i) bandwidth, (ii) power, (iii) average latency, (iv) average total latency, (v) memory reads, and (vi) memory writes. Our experimental results show that machine learning models can predict memory response parameter values faster than simulations. We used support vector machine, random forest, and gradient boosting machine learning models. We observed that the support vector machine provides better performance for bandwidth, average latency, and average total latency. The random forest model works better for memory reads and writes. The gradient boosting model provides superior prediction performance for power. We provide a detailed discussion on learning curve characteristics, error analysis, and memory type recommendation.

Hasan, S M Shamimul↗

Advancing attenuation estimation through integration of the Hessian in multiparameter viscoacoustic full-waveform inversion

Accurate seismic attenuation models of subsurface structures not only enhance subsequent migration processes by improving fidelity, resolution, and facilitating amplitude-compliant angle gather generation but also provide valuable constraints on subsurface physical properties. Leveraging full-wavefield information, multiparameter viscoacoustic full-waveform inversion ( Q-FWI) simultaneously estimates seismic velocity and attenuation ( Q) models. However, a major challenge in Q-FWI is the contamination of crosstalk artifacts, where inaccuracies in the velocity model are mistakenly mapped to the inverted attenuation model. While incorporating the Hessian is expected to mitigate these artifacts, the explicit implementation is prohibitively expensive due to its formidable computational cost. In this study, we formulate and develop a Q-FWI algorithm via the Newton-conjugate gradient (CG) framework, where the search direction at each iteration is determined through an internal CG loop. In particular, the Hessian is integrated into each CG step in a matrix-free fashion using the second-order adjoint-state method. We find through synthetic experiments that our Newton-CG Q-FWI significantly mitigates crosstalk artifacts compared with the limited-memory Broyden-Fletcher-Goldfarb-Shanno method and the CG method, albeit with a notable computational cost. In the discussion of several key implementation details, we also determine the significance of the approximate Gauss-Newton Hessian, the second-order adjoint-state method, and the two-stage inversion strategy.

Geochemistry & Geophysics↗

Pre-conditioned BFGS-based uncertainty quantification in elastic full-waveform inversion

SUMMARY Full-waveform inversion has become an essential technique for mapping geophysical subsurface structures. However, proper uncertainty quantification is often lacking in current applications. In theory, uncertainty quantification is related to the inverse Hessian (or the posterior covariance matrix). Even for common geophysical inverse problems its calculation is beyond the computational and storage capacities of the largest high-performance computing systems. In this study, we amend the Broyden–Fletcher–Goldfarb–Shanno (BFGS) algorithm to perform uncertainty quantification for large-scale applications. For seismic inverse problems, the limited-memory BFGS (L-BFGS) method prevails as the most efficient quasi-Newton method. We aim to augment it further to obtain an approximate inverse Hessian for uncertainty quantification in FWI. To facilitate retrieval of the inverse Hessian, we combine BFGS (essentially a full-history L-BFGS) with randomized singular value decomposition to determine a low-rank approximation of the inverse Hessian. Setting the rank number equal to the number of iterations makes this solution efficient and memory-affordable even for large-scale problems. Furthermore, based on the Gauss–Newton method, we formulate different initial, diagonal Hessian matrices as pre-conditioners for the inverse scheme and compare their performances in elastic FWI applications. We highlight our approach with the elastic Marmousi benchmark model, demonstrating the applicability of pre-conditioned BFGS for large-scale FWI and uncertainty quantification.

58 GEOSCIENCES↗

L-BFGS Class Implementation in C++

This report presents a header-only C++ class implementation of the Limited-memory BroydenFletcher-Goldfarb-Shanno (L-BFGS) algorithm. The L-BFGS method is a general purpose quasi-Netwon optimization method that builds an approximation of the descent direction from consecutive iterate and gradient vectors. The limited-memory aspect stems from the replacement of the N × N approximation matrix of the original BFGS method with M vectors of length N. An example usage of the class is included along with the reference source code.

97 MATHEMATICS AND COMPUTING↗

DyG-DPCD: A Distributed Parallel Community Detection Algorithm for Large-Scale Dynamic Graphs

Dynamic (Temporal) graphs capture the valuable evolution of real-world systems, from the continuously evolving patterns of social interactions and genetic pathways to the dynamic fluctuations of economic forces. Detecting communities for such evolving networks poses unique challenges. Detecting and analyzing the evolution of communities within dynamic graphs unlocks valuable insights into the underlying structural and temporal patterns of real-world systems. However, the sheer volume of modern graph data and the inherent complexity of the temporal dimension pose significant challenges to scalable community detection algorithms. Addressing this gap, our work explores the limited landscape of scalable distributed-memory parallel methods specifically designed for dynamic network community detection. We propose a novel parallel algorithm, DyG-DPCD (Dynamic Graph Distributed Parallel Community Detection), to detect communities in dynamic networks using the Message Passing Interface (MPI) framework. We present a vertex-centric approach, allowing us to detect communities through local optimization. Furthermore, we enhance our baseline algorithm by incorporating three heuristics, which improve the algorithm’s performance significantly while maintaining the quality of the solutions. We demonstrate the efficiency of our algorithm by experimenting on several real-world large-scale networks with hundreds of millions of edges spanning diverse domains. Notably, DyG-DPCD achieves speedups between 25× and 30× for large networks that we experimented on using NERSC compute nodes. In conclusion, our algorithm outperforms the STINGER parallel re-agglomeration algorithm by 30×.

97 MATHEMATICS AND COMPUTING↗

A randomized sketching trust-region secant method for low-memory dynamic optimization

The numerical solution of dynamic optimization problems is often limited by the memory required to store the state trajectory, which is used to evaluate the objective function and its derivatives. Recently, [R. Muthukumar et al., SIAM Journal on Optimization 31(2), pp. 1242–1275 (2021)] introduced a trust-region method for dynamic optimization that employs randomized sketching to compress the state trajectory, resulting in inexact derivative computations. By adaptively learning the sketch rank, the trust-region algorithm achieves rigorous convergence guarantees. Here, we extend this approach to use secant Hessian approximations. Due to the randomness introduced by the sketch, the traditional secant update formulae can produce poor Hessian approximations. In particular, the difference of two gradients, computed from two different sketches, may be inconsistent. To overcome this, we employ a sketched approximation of the Hessian application, in lieu of computing the gradient difference. We numerically demonstrate the improved stability of this approach on an example from PDE-constrained optimization.

dynamic optimization↗

Efficiently evaluating loop integrals in the EFTofLSS using QFT integrals with massive propagators

We develop a new way to analytically calculate loop integrals in the Effective Field Theory of Large Scale-Structure. Previous available methods show severe limitations beyond the one-loop power spectrum due to analytical challenges and computational and memory costs. Our new method is based on fitting the linear power spectrum with cosmology-independent functions that resemble integer powers of quantum field theory massive propagators with complex masses. A remarkably small number of them is sufficient to reach enough accuracy. Similarly to former approaches, the cosmology dependence is encoded in the coordinate vector of the expansion of the linear power spectrum in our basis. We first produce cosmology-independent tensors where each entry is the loop integral evaluated on a given combination of basis vectors. For each cosmology, the evaluation of a loop integral amounts to contracting this tensor with the coordinate vector of the linear power spectrum. The 3-dimensional loop integrals for our basis functions can be evaluated using techniques familiar to particle physics, such as recursion relations and Feynman parametrization. We apply our formalism to evaluate the one-loop bispectrum of galaxies in redshift space. The final analytical expressions are quite simple and can be evaluated with little computational and memory cost. We show that the same expressions resolve the integration of all one-loop N-point function in the EFTofLSS. This method, which is originally presented here, has already been applied in the first one-loop bispectrum analysis of the BOSS data to constraint ΛCDM parameters and primordial non-Gaussianities [1, 2].

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Machine learning with bond information for local structure optimizations in surface science

Local optimization of adsorption systems inherently involves different scales: within the substrate, within the molecule, and between the molecule and the substrate. In this work, we show how the explicit modeling of different characteristics of the bonds in these systems improves the performance of machine learning methods for optimization. Furthermore, we introduce an anisotropic kernel in the Gaussian process regression framework that guides the search for the local minimum, and we show its overall good performance across different types of atomic systems. The method shows a speed-up of up to a factor of two compared with the fastest standard optimization methods on adsorption systems. Additionally, we show that a limited memory approach is not only beneficial in terms of overall computational resources but can also result in a further reduction of energy and force calculations.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Design of experiments to spectroscopically characterize radiation flow in stochastic media

Precise characterization of experimental radiation flow is required to validate the high energy density physics models, numerical methods, and codes that are used to simulate radiation-hydrodynamics phenomena such as thermal radiation transport in stochastic media. The Cassio code is used to simulate thermal radiation flow through inhomogeneous, stochastic-media-foam configurations containing optically thick clumps dispersed within an optically thin background aerogel. Cassio can model small inhomogeneous problems directly, but most problems require approximations to meet computer limitations on run-times and memory usage. Various examples of these approximations are methods that produce, in one calculation, an ensemble-averaged solution and associated standard deviation; reduced spatial dimensionality with approximate geometries; and full material homogenization with no geometric detail. Cassio simulations are used to design experiments at the OMEGA-60 Laser Facility that can measure the radiation flow using the spatially resolved COAX absorption spectroscopy diagnostic. The experimental platforms flow radiation through foam targets ranging from a background-only aerogel, to a single configuration of a specified stochastic medium, to a fully homogenized foam of the background and clump materials. Under constant total clump mass, larger clumps (here, larger than 10 μm diameter) will mix more slowly with the background such that the bulk radiation flow is faster than it would be in a fully homogenized material. The COAX platform can be used to infer temperature and density profiles in both the background material and clumps, simultaneously, and therefore to differentiate radiation flow in a range of stochastic and homogeneous media.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Large-Scale Optimization with Linear Equality Constraints Using Reduced Compact Representation

For optimization problems with linear equality constraints, we prove that the (1,1) block of the inverse KKT matrix remains unchanged when projected onto the nullspace of the constraint matrix. In this work, we develop reduced compact representations of the limited-memory inverse BFGS Hessian to compute search directions efficiently when the constraint Jacobian is sparse. Orthogonal projections are implemented by a sparse QR factorization or a preconditioned LSQR iteration. In numerical experiments two proposed trust-region algorithms improve in computation times, often significantly, compared to previous implementations of related algorithms and compared to IPOPT.

97 MATHEMATICS AND COMPUTING↗

Adaptive Scalpel Scanning Probe Microscopy for Enhanced Volumetric Sensing in Tomographic Analysis

Controlling nanoscale tip‐induced material removal is crucial for achieving atomic‐level precision in tomographic sensing with atomic force microscopy (AFM). While advances have enabled volumetric probing of conductive features with nanometer accuracy in solid‐state devices, materials, and photovoltaics, limitations in spatial resolution and volumetric sensitivity persist. This work identifies and addresses in‐plane and vertical tip‐sample junction leakage as sources of parasitic contrast in tomographic AFM, hindering real‐space 3D reconstructions. Novel strategies are proposed to overcome these limitations. First, the contrast mechanisms analyzing nanosized conductive features are explored when confining current collection purely to in‐plane transport, thus allowing reconstruction with a reduction in the overestimation of the lateral dimensions. Furthermore, an adaptive tip‐sample biasing scheme is demonstrated for the mitigation of a class of artefacts induced by the high electric field inside the thin oxide when volumetrically reduced. This significantly enhances vertical sensitivity by approaching the intrinsic limits set by quantum tunneling processes, allowing detailed depth analysis in thin dielectrics. The effectiveness of these methods is showcased in tomographic reconstructions of conductive filaments in valence change memory, highlighting the potential for application in nanoelectronics devices and bulk materials and unlocking new limits for tomographic AFM.

36 MATERIALS SCIENCE↗

A low-rank power iteration scheme for neutron transport criticality problems

Computing effective eigenvalues for neutron transport often requires a fine numerical resolution. Here, the main challenge of such computations is the high memory effort of classical solvers, which limits the accuracy of chosen discretizations. In this work, we derive a method for the computation of effective eigenvalues when the underlying solution has a low-rank structure. This is accomplished by utilizing dynamical low-rank approximation (DLRA), which is an efficient strategy to derive time evolution equations for low-rank solution representations. The main idea is to interpret the iterates of the classical inverse power iteration as pseudo-time steps and apply the DLRA concepts in this framework. In our numerical experiment, we demonstrate that our method significantly reduces memory requirements while achieving the desired accuracy. Analytic investigations show that the proposed iteration scheme inherits the convergence speed of the inverse power iteration, at least for a simplified setting.

97 MATHEMATICS AND COMPUTING↗

FORCE-DISPATCHES Integration - Initial Demonstration

Integrated energy systems (IES) combine, in mutually beneficial ways, power from variable renewable energy sources and nuclear power plants (NPP) to improve economic viability under uncertain market and weather conditions. The open-source Framework for Optimization of Resources and Economics (FORCE) tool suite, developed at Idaho National Laboratory (INL), has enabled comprehensive modeling and simulation of IES. The capabilities within FORCE include grid portfolio optimization through the Holistic Energy Resource Optimization Network (HERON) and the transient process model analysis library HYBRID, among others. Continuous efforts and investments from the IES programs have been made to expand and improve the versatility of the FORCE toolset in fiscal year 2022. Code-coupling and cross-tool communication have been important methods for improving this versatility. This report focuses on an additional workflow in the HERON tool for capacity and dispatch stochastic optimization through integration with the external tool Design Integration and Synthesis Platform to Advance Tightly Coupled Hybrid Energy Systems (DISPATCHES). DISPATCHES was primarily developed by the National Energy Technology Laboratory, in collaboration with other national laboratories, which included INL, universities, and industry partners. It is coupled to a library of algebraic models for specific plant components, and to a framework for stochastic optimization different from that provided in the current Risk Analysis Virtual Environment (RAVEN)-running-RAVEN algorithm in HERON. HERON currently conducts stochastic optimization via an outer-inner loop: it optimizes over variable capacity on the outer loop, and at each step within the capacity parameter space, conducts an inner optimization over scenarios (of market signals, demand, and/or weather patterns) and hourly dispatch throughout a user-specified number of years. On the other hand, DISPATCHES conducts stochastic optimization via an “all-at-once” strategy in which capacity variables are optimized at the same level as dispatch variables, as all scenarios are considered at once. The latter method works especially well for projects of limited size and project length, as the necessary computational power and memory increases with the number of variables and scenarios. The new capability to use the DISPATCHES workflow in HERON enhances standalone simulations by leveraging FORCE tools—namely, the economic metrics from the Tool for Economic Analysis (TEAL) and reduced-order model (ROM) sampling from RAVEN. The initial demonstration of the DISPATCHES workflow simulates an existing nuclear-case flowsheet within the DISPATCHES repository—this models a NPP with a secondary revenue stream for hydrogen production. Electrical output from the plant is converted to hydrogen via a proton-exchange membrane (PEM) electrolyzer, hydrogen tanks are used for storage, and an additional turbine is added for hydrogen combustion. Continued work regarding this FORCE-DISPATCHES integration will include automatic generation of DISPATCHES models from HERON inputs, offering analysts the option of using either the RAVEN-runsRAVEN or DISPATCHES workflow to solve technoeconomic optimization problems.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Accelerating structural dynamics simulations with localised phenomena through matrix compression and projection‐based model order reduction

In this work, a novel approach is introduced for accelerating the solution of structural dynamics problems in the presence of localised phenomena, such as cracks. For this category of problems, conventional projection-based Model Order Reduction (MOR) methods are either limited with respect to the range of system configurations that can be represented or require frequent solutions of the Full Order Model (FOM) to update the low-dimensional spaces, in which solutions are represented. In the proposed approach, low-dimensional spaces, constructed for the healthy structure, are enriched with appropriately selected columns of the flexibility matrix of the system. It can be shown that these spaces contain the solution to the original problem for the static case, while their dimension is much smaller. In order to allow their online construction for arbitrary localised features, the full flexibility matrix of the system should be available. To this end, a hierarchical representation is used for the matrices involved, allowing to compute the flexibility matrix efficiently and with reduced memory requirements. The resulting method offers significant speedups, without sacrificing the flexibility and accuracy of the full order model. The performance and limitations of the approach are studied through a series of examples in structural dynamics.

fracture mechanics↗

Explainable machine learning model for multi-step forecasting of reservoir inflow with uncertainty quantification

We propose an explainable machine learning (ML) model with uncertainty quantification (UQ) to improve multi-step reservoir inflow forecasting. Traditional ML methods have challenges in forecasting inflows multiple days ahead, and lack explainability and UQ. To address these limitations, we introduce an encoder–decoder long short-term memory (ED-LSTM) network for multi-step forecasting, employ the SHapley Additive exPlanation (SHAP) technique for understanding the influence of hydrometeorological factors on inflow prediction, and develop a novel UQ method for prediction trustworthiness. We apply these methods to forecast 7-day inflow in snow-dominant and rain-driven reservoirs. The results demonstrate the effectiveness of the ED-LSTM model, with high forecasting accuracy for short lead times. Our UQ method provides reliable uncertainty estimates, covering 90% of data with a 90% confidence level. The SHAP analysis reveals the importance of historical inflow and precipitation as influential factors. These findings and methods may support reservoir operators in optimizing water resources management decisions.

54 ENVIRONMENTAL SCIENCES↗