Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “limited memory”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Image processing tools for petabyte-scale light sheet microscopy data

Light sheet microscopy is a powerful technique for high-speed three-dimensional imaging of subcellular dynamics and large biological specimens. However, it often generates datasets ranging from hundreds of gigabytes to petabytes in size for a single experiment. Conventional computational tools process such images far slower than the time to acquire them and often fail outright due to memory limitations. To address these challenges, we present PetaKit5D, a scalable software solution for efficient petabyte-scale light sheet image processing. This software incorporates a suite of commonly used processing tools that are optimized for memory and performance. Notable advancements include rapid image readers and writers, fast and memory-efficient geometric transformations, high-performance Richardson–Lucy deconvolution and scalable Zarr-based stitching. These features outperform state-of-the-art methods by over one order of magnitude, enabling the processing of petabyte-scale image data at the full teravoxel rates of modern imaging cameras. The software opens new avenues for biological discoveries through large-scale imaging experiments.

97 MATHEMATICS AND COMPUTING↗

Cryogenic energy storage: Standalone design, rigorous optimization and techno-economic analysis

Energy storage allows flexible use and management of excess electricity and intermittently available renewable energy. Cryogenic energy storage (CES) is a promising storage alternative with a high technology readiness level and maturity, but the round-trip efficiency is often moderate and the Levelized Cost of Storage (LCOS) remains high. The complex flowsheets with intricate thermodynamics at cryogenic temperatures as well as the presence of multiple loops and refrigeration cycles pose considerable challenges for rigorous model-based design and optimization of CES systems. We present an optimization strategy that couples rigorous process simulation and Bayesian optimization with flowsheet decomposition and identification of hidden coupling constraints to optimally design standalone CES systems. Further refinement is done via a local search using the limited-memory Broyden–Fletcher–Goldfarb–Shanno algorithm. Here our results indicate that it is possible to achieve more than 52% round-trip efficiency and an LCOS of $153/MWh for a standalone 100 MW/400 MWh CES system limited to short-term storage with daily charging–discharging. However, a detailed techno-economic assessment reveals that the LCOS considering total capital investment may exceed $267/MWh when all direct and indirect costs of installation and operation are considered.

25 ENERGY STORAGE↗

Memory-efficient nonsmooth dynamic optimization using adaptive randomized compression

Dynamic optimization problems arise in many applications including flow control, full waveform inversion, and medical imaging. These problems are plagued by significant computational challenges. One such challenge — and the focus of this work — is the memory limitation induced by the size of the underlying dynamical system. In particular, the entire dynamic trajectory is required for derivative computation and therefore must be stored or recomputed using, e.g., checkpointing. Although recent work demonstrated the use of adaptive randomized sketching to overcome the memory challenge, that work only applies to smooth unconstrained problems, prohibiting its use for nonsmooth regularized and constrained problems. The inclusion of nonsmooth regularizers and constraints is critical as they often arise in an attempt to preserve certain physical properties or to promote sparsity. To solve these problems, we introduce a trust-region algorithm for minimizing the sum of a smooth nonconvex function and a nonsmooth convex function that leverages randomized sketching to compress the dynamical system trajectories and adaptively adjust the sketch rank to satisfy a gradient inexactness condition. We prove convergence of this algorithm and demonstrate that it achieves substantial memory reduction on three discretized PDE-constrained optimization applications.

97 MATHEMATICS AND COMPUTING↗

Train small, model big: Scalable physics simulators via reduced order modeling and domain decomposition

Numerous cutting-edge scientific technologies originate at the laboratory scale, but transitioning them to practical industry applications is a formidable challenge. Traditional pilot projects at intermediate scales are costly and time-consuming. An alternative, the pilot-scale model, relies on high-fidelity numerical simulations, but even these simulations can be computationally prohibitive at larger scales. To overcome these limitations, we propose a scalable, physics-constrained reduced order model (ROM) method. The ROM identifies critical physics modes from small-scale unit components, projecting governing equations onto these modes to create a reduced model that retains essential physics details. We also employ Discontinuous Galerkin Domain Decomposition (DG-DD) to apply ROM to unit components and interfaces, enabling the construction of large-scale global systems without data at such large scales. Here this method is demonstrated on the Poisson and Stokes flow equations, showing that it can solve equations about 15–40 times faster with only ~1% relative error. Furthermore, ROM takes one order of magnitude less memory than the full order model, enabling larger scale predictions at a given memory limitation.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

GENESPACE R Package (GENESPACE) v1.0

In short, the GENESPACE pipeline conducts analysis of orthology networks, constrained within syntenic regions. Since analyses are limited to local tests conducted within syntenic blocks, GENESPACE is agnostic to ploidy, duplicated regions, inversions or other whole-genome chromosomal complexities that are common across many evolutionary lineages. This advantage allows for evolutionary tests in polyploids (e.g. switchgrass, manuscript in review), species with ancient, but retained whole-genome duplications (e.g. pecan, manuscript in prep), high levels of tandem array proliferation (e.g. eukalypts, manuscript in review) and many other factors that can confound comparative genomic analyses. The major advances of GENESPACE are three-fold: First, this is the first R package to integrate visualization and analysis of large-scale comparative genomics. R, which offers a high-level environment for graphical and statistical exploration of data, is often speed- and memory-limited and not used for computationally intensive tasks such as comparative genomics. The highly efficient C++ scripts used in GENESPACE (via data.table) permit a much faster and computationally lightweight implementation of comparative genomics than is currently available. Second, the pipeline itself is novel. To the best of our knowledge, no other program accomplishes synteny-constrained and ploidy-agnostic comparative genomics. Since nearly all plants and many animals have a history of whole-genome duplications, this is a major and necessary advance to the field. Third, GENESPACE offers high-level and intuitive multi-genome graphical outputs. The dotplots and 'riparian' plots produced herein, which are produced entirely through original R code, are publication-ready and easily customizable.

Schmutz, Jeremy↗

Autonomous image data reduction by analysis and interpretation

Image data is a critical component of the scientific information acquired by space missions. Compression of image data is required due to the limited bandwidth of the data transmission channel and limited memory space on the acquisition vehicle. This need becomes more pressing when dealing with multispectral data where each pixel may comprise 300 or more bytes. An autonomous, real time, on-board image analysis system for an exploratory vehicle such as a Mars Rover is developed. The completed system will be capable of interpreting image data to produce reduced representations of the image, and of making decisions regarding the importance of data based on current scientific goals. Data from multiple sources, including stereo images, color images, and multispectral data, are fused into single image representations. Analysis techniques emphasize artificial neural networks. Clusters are described by their outlines and class values. These analysis and compression techniques are coupled with decision making capacity for determining importance of each image region. Areas determined to be noise or uninteresting can be discarded in favor of more important areas. Thus limited resources for data storage and transmission are allocated to the most significant images.

Eberlein, Susan↗

Autonomous image data reduction by analysis and interpretation

Image data is a critical component of the scientific information acquired by space missions. Compression of image data is required due to the limited bandwidth of the data transmission channel and limited memory space on the acquisition vehicle. This need becomes more pressing when dealing with multispectral data where each pixel may comprise 300 or more bytes. An autonomous, real time, on-board image analysis system for an exploratory vehicle such as a Mars Rover is developed. The completed system will be capable of interpreting image data to produce reduced representations of the image, and of making decisions regarding the importance of data based on current scientific goals. Data from multiple sources, including stereo images, color images, and multispectral data, are fused into single image representations. Analysis techniques emphasize artificial neural networks. Clusters are described by their outlines and class values. These analysis and compression techniques are coupled with decision-making capacity for determining importance of each image region. Areas determined to be noise or uninteresting can be discarded in favor of more important areas. Thus limited resources for data storage and transmission are allocated to the most significant images.

Eberlein, Susan↗

Tonal Emergence: An agent-based model of tonal coordination

Humans have a remarkable capacity for coordination. Our ability to interact and act jointly in groups is crucial to our success as a species. Joint Action (JA) research has often concerned itself with simplistic behaviors in highly constrained laboratory tasks. But there has been a growing interest in understanding complex coordination in more open-ended contexts. In this regard, collective music improvisation has emerged as a fascinating model domain for studying basic JA mechanisms in an unconstrained and highly sophisticated setting. A number of empirical studies have begun to elucidate coordination mechanisms underlying joint musical improvisation, but these empirical findings have yet to be cached out in a working computational model. The present work fills this gap by presenting TonalEmergence, an idealized agent-based model of improvised musical coordination. TonalEmergence models the coordination of notes played by improvisers to generate harmony (i.e., tonality), by simulating agents that stochastically generate notes biased towards maximizing harmonic consonance given their partner’s previous notes. Here, the model replicates an interesting empirical result from a previous study of professional jazz pianists: feedback loops of mutual adaptation between interacting agents support the production of consonant harmony. The model is further explored to show how complex tonal dynamics, such as the production and dissolution of stable tonal centers, are supported by agents that are characterized by (i) a tendency to strive toward consonance, (ii) stochasticity, and (iii) a limited memory for previously played notes. TonalEmergence thus provides a grounded computational model to simulate and probe the coordination mechanisms underpinning one of the more remarkable feats of human cognition: collective music improvisation.

60 APPLIED LIFE SCIENCES↗

Accelerating multigrid with streaming chiral SVD for Wilson fermions in lattice QCD

A modification to the setup algorithm for the multigrid preconditioner of Wilson fermions in lattice QCD is presented. A larger basis of test vectors than that used in regular multigrid is calculated by the smoother and truncated by singular value decomposition on the chiral components of the test vectors. The truncated basis is used to form the prolongation and restriction matrices of the multigrid hierarchy. This modification of the setup method is demonstrated to increase the convergence of linear solvers on an anisotropic lattice with m π ≈ 239 MeV from the Hadron Spectrum Collaboration and an isotropic lattice with m π ≈ 220 MeV from the MILC Collaboration. The lattice volume dependence of the method is also examined. Increasing the number of test vectors improves speedup up to a point, but storing these vectors becomes impossible in limited memory resources such as GPUs. To address storage cost, we implement a streaming singular value decomposition of the basis of test vectors on the chiral components and demonstrate a decrease in the number of fine level iterations by a factor of 1.7 for m q ≈ m crit

Iterative methods↗

Autonomous nondestructive evaluation of resistance spot welded joints

The application of non-destructive evaluation approaches has attracted strong interests in modern automotive industries. Here, we present an autonomous deep-computing framework to analyze raw videos from infrared systems and to predict weld nugget shape and size with unprecedented accuracy and speed. In a comprehensive training and testing experiment with 90 videos (seven sets of welding material stack-ups), a new method was developed to assemble sufficient datasets for neural network training. Our framework successfully predicts all the nugget shapes with F1 scores that range from 0.84 to 0.92. The total training time on Nvidia DGX station takes less than 10 min for each set of welding material stack-up. The real inference time of an individual dataset (with 30 video frames) takes about 0.005 s. The procedure and methods developed in the study can be applied to other image-based weld property prediction, as well as other manufacturing processes. Furthermore, our well-trained neural networks take limited memory resources (2.3 MB) and are suitable for embedded microprocessors for in-situ welding quality control as edge computing within an intelligent welding framework.

42 ENGINEERING↗

WUS256: An Adjoint Waveform Tomography Model of the Crust and Upper Mantle of the Western United States for Improved Waveform Simulations

Abstract We report a new model (WUS256) of radially anisotropic seismic wavespeeds of the crust and upper mantle of the western United States (WUS) obtained from adjoint waveform tomography for the purpose of improving synthetic waveform fits to observed data. WUS256 is based on inversion of over 94,000 waveforms from 72 earthquakes recorded by nearly 3,400 stations. We started with the SPiRaL global model (Simmons et al., 2021, https://doi.org/10.1093/gji/ggab277 ) and waveforms in the period band of 50–120 s. We followed a conservative multiscale inversion approach with eight stages and 256 total inversion iterations which enabled monotonic misfit reduction to 20‐s minimum‐period waves. WUS256 relied on time‐frequency (TF) phase misfits and a trust region limited memory Broyden–Fletcher–Goldfarb–Shanno (L‐BFGS) optimization. Hessian‐vector products were used to qualitatively assess model resolution. Results indicate that WUS256 has good coverage of the continental regions to depths of about 150 km and is able to resolve features on lateral scales of about 200 km. We quantify waveform fits by the reduction in TF and normalized amplitude difference misfits between WUS256 and the SPiRaL starting model. WUS256 significantly improves waveform fits with misfit reduction 64% for both inversion and validation data sets compared to the SPiRaL starting model and shows even better fits compared to other models. Waveform fits illustrate that WUS256 reproduces body‐waves, fundamental mode surface waves as well as late arriving dispersed and/or scattered short period surface waves. The improvement in waveform fit indicates that WUS256 can be used to reproduce path effects on regional complete waveforms and moment tensor inversions.

58 GEOSCIENCES↗

Ground and excited state gradients with end-to-end differentiable semiempirical quantum chemistry

Accurate and efficient gradients of molecular energy with respect to nuclear degrees of freedom are essential for geometry optimization and molecular dynamics, including simulations that go beyond the Born–Oppenheimer regime. A common approach involves deriving analytical formulas for new electronic structure methods, which is often conceptually difficult and requires tedious coding. Here, we implement analytical, semi-numerical, and automatic differentiation (AD)-based gradient pathways for semiempirical Hamiltonian models in the PYSEQM software package, leveraging both graphics processing unit (GPU) and central processing unit (CPU) architectures. We further extend these capabilities to excited states calculated using the configuration interaction singles and time-dependent Hartree–Fock ansätze. We benchmark wall time, peak memory usage, and accuracy across three molecular families of varying chemical complexity, including systems of up to a thousand atoms. For ground-state simulations, analytical and AD gradients achieve near-identical GPU runtimes, while semi-numerical gradients are slower on GPU but remain competitive on CPU. For excited states, both analytical and custom AD approaches using implicit differentiation show similar performance and low memory requirements, whereas gradients with full AD are memory-limited. AD gradients match analytical ones in accuracy across all tested systems, aided by a quaternion-based diatomic frame rotation for two-center quantities that ensures smooth energy surfaces. Overall, automatic differentiation emerges as a practical alternative to analytical gradients in semiempirical quantum chemistry, offering high accuracy while allowing seamless integration in AI-driven workflows and popular packages, such as PyTorch and JAX. Our results provide actionable guidance for selecting optimal gradient strategies in large-scale ground- and excited-state molecular dynamics simulations.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Pre-conditioned BFGS-based uncertainty quantification in elastic full-waveform inversion

SUMMARY Full-waveform inversion has become an essential technique for mapping geophysical subsurface structures. However, proper uncertainty quantification is often lacking in current applications. In theory, uncertainty quantification is related to the inverse Hessian (or the posterior covariance matrix). Even for common geophysical inverse problems its calculation is beyond the computational and storage capacities of the largest high-performance computing systems. In this study, we amend the Broyden–Fletcher–Goldfarb–Shanno (BFGS) algorithm to perform uncertainty quantification for large-scale applications. For seismic inverse problems, the limited-memory BFGS (L-BFGS) method prevails as the most efficient quasi-Newton method. We aim to augment it further to obtain an approximate inverse Hessian for uncertainty quantification in FWI. To facilitate retrieval of the inverse Hessian, we combine BFGS (essentially a full-history L-BFGS) with randomized singular value decomposition to determine a low-rank approximation of the inverse Hessian. Setting the rank number equal to the number of iterations makes this solution efficient and memory-affordable even for large-scale problems. Furthermore, based on the Gauss–Newton method, we formulate different initial, diagonal Hessian matrices as pre-conditioners for the inverse scheme and compare their performances in elastic FWI applications. We highlight our approach with the elastic Marmousi benchmark model, demonstrating the applicability of pre-conditioned BFGS for large-scale FWI and uncertainty quantification.

58 GEOSCIENCES↗

Tula: Optimizing Time, Cost, and Generalization in Distributed Large-Batch Training

Distributed training increases the number of batches processed per iteration either by scaling-out (adding more nodes) or scaling-up (increasing the batch-size). However, the largest configuration does not necessarily yield the best performance. Horizontal scaling introduces additional communication overhead, while vertical scaling is constrained by computation cost and device memory limits. Thus, simply increasing the batch-size leads to diminishing returns: training time and cost decrease initially but eventually plateaus, creating a knee-point in the time/cost vs. batch-size pareto curve. The optimal batch-size therefore depends on the underlying model, data and available compute resources. Large batches also suffer from worse model quality due to the well-known “generalization gap”. In this paper, we present Tula, an online service that automatically optimizes time, cost, and convergence quality for large-batch training of convolutional models. It combines parallel-systems modeling with statistical performance prediction to identify the optimal batchsize. Tula predicts training time and cost within 7.5−14% error across multiple models, and achieves up to 20× overall speedup and improves test accuracy by ≈9% on average over standard large-batch training on various vision tasks, thus successfully mitigating the generalization gap and accelerating training at the same time.

Tyagi, Sahil [ORNL] (ORCID:0009000783144745)↗

LATTE: Los Alamos TravelTime package based on Eikonal equation

This Fortran code focuses on traveltime computation and tomography based on eikonal equation. Specifically, the package provides three major functionalities: (1) forward modeling of traveltime from single-point or ensemble source based on factorized eikonal equation, (2) adjoint-state first-arrival traveltime tomography based on picked first arrival traveltime using steepest descent, conjugate gradient, or limited-memory BFGS inversion scheme, and (3) adjoint-state joint transmission-reflection tomography based on picked first-arrival and reflection traveltimes. The package applies to forward modeling and tomography based on traveltime in 2D and 3D isotropic regular-grid models. We name this package LATTE – Los Alamos TravelTime package based on Eikonal equation. * The code is for accompanying a journal paper under preparation. The paper will be submitted via LA-UR separately later.

Gao, Kai↗

Sequence length scaling in vision transformers for scientific images on frontier

Vision Transformers (ViTs) are pivotal for foundational models in scientific imagery, including Earth science applications, due to their capability to process large sequence lengths. While transformers for text have inspired scaling sequence lengths in ViTs, adapting these for ViTs introduces unique challenges. We develop distributed sequence parallelism for ViTs, enabling them to handle up to 1M tokens. Our approach, leveraging DeepSpeed-Ulysses and Long-Sequence-Segmentation with model sharding, is the first to apply sequence parallelism in ViT training, achieving a 94% batch scaling efficiency on 2,048 AMD-MI250X GPUs. Evaluating sequence parallelism in ViTs, particularly in models up to 10B parameters, highlighted substantial bottlenecks. We countered these with hybrid sequence, pipeline, and flash attention strategies, to scale beyond single GPU memory limits. Our method significantly enhances climate modeling accuracy by 20% in temperature predictions, marking the first training of a vision transformer model to convergence with a sequence length of 188K tokens, using full self-attention.

Tsaris, Aristeidis (aris) [ORNL] (ORCID:0000000277↗

Advancing attenuation estimation through integration of the Hessian in multiparameter viscoacoustic full-waveform inversion

Accurate seismic attenuation models of subsurface structures not only enhance subsequent migration processes by improving fidelity, resolution, and facilitating amplitude-compliant angle gather generation but also provide valuable constraints on subsurface physical properties. Leveraging full-wavefield information, multiparameter viscoacoustic full-waveform inversion ( Q-FWI) simultaneously estimates seismic velocity and attenuation ( Q) models. However, a major challenge in Q-FWI is the contamination of crosstalk artifacts, where inaccuracies in the velocity model are mistakenly mapped to the inverted attenuation model. While incorporating the Hessian is expected to mitigate these artifacts, the explicit implementation is prohibitively expensive due to its formidable computational cost. In this study, we formulate and develop a Q-FWI algorithm via the Newton-conjugate gradient (CG) framework, where the search direction at each iteration is determined through an internal CG loop. In particular, the Hessian is integrated into each CG step in a matrix-free fashion using the second-order adjoint-state method. We find through synthetic experiments that our Newton-CG Q-FWI significantly mitigates crosstalk artifacts compared with the limited-memory Broyden-Fletcher-Goldfarb-Shanno method and the CG method, albeit with a notable computational cost. In the discussion of several key implementation details, we also determine the significance of the approximate Gauss-Newton Hessian, the second-order adjoint-state method, and the two-stage inversion strategy.

Geochemistry & Geophysics↗

Multiphysics Simulations of MSRE with NEAMS Thermal Hydraulics Tools

This report documents the benchmarks being developed and simulations performed using tools and codes developed under the Nuclear Energy Advanced Modeling and Simulation (NEAMS) program, utilizing MSRE experimental data. In FY23, three main work scopes were investigated under the NEAMS MSR work package at ANL. The first scope investigated the Griffin-SAM coupling model for simulating the pump startup transient experiment of MSRE. The analyses start with a simple model (single-channel, single-lattice), gradually adding more details (multi-channel, full-core) into the model. The results show that the reactivity loss curve is very sensitive to the axial boundary conditions and the radial core discretization. The simple model can predict a similar reactivity trend as that of the more sophisticated model, which is likely due to error cancellation. Accurately modeling the axial boundary condition may further improve the reactivity trend but would require significant efforts to generate the mesh of the MSRE inlet and upper plenum. The core channel radial discretization for the Griffin-SAM coupled model also depends on the flow distribution. Given the complex geometry in the inlet plenum, the flow distribution needed to be calculated from CFD analysis, which was performed using the NekRS code. This analysis employed a MSRE CAD model developed by Copenhagen Atomics. The CAD model was disassembled to keep the inlet plenum region only, which was subsequently cleaned and modified so that the mesh generated is under the memory limit. The results are merged to a few radial regions to show that the flow rate is highest in the central region. This would be useful for future improvement of the Griffin-SAM coupling model of the MSRE core. The last task investigated is tritium transport modeling using the standalone SAM code. This task aimed to initiate the effort to demonstrate and validate the tritium transport model implemented in SAM. The preliminary investigation employed an MSRE model consisting of the primary loop. Three tritium transport pathways were examined including the retention in the graphite, the permeation through the HX tube wall, and the removal from the off-gas system. The results compare well with the MSRE data, but improvements are still needed on the initial conditions (i.e., the present state may not have reached equilibrium), the boundary conditions, the off-gas system modeling, and a better numerical strategy to reach the equilibrium state.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗