Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “limited memory method”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Evaluating the Potential and Challenges of an Uncertainty Quantification Method for Long Short–Term Memory Models for Soil Moisture Predictions

Recently, recurrent deep networks have shown promise to harness newly available satellite–sensed data for long–term soil moisture projections. However, to be useful in forecasting, deep networks must also provide uncertainty estimates. Here we evaluated Monte Carlo dropout with an input–dependent data noise term (MCD+N), an efficient uncertainty estimation framework originally developed in computer vision, for hydrologic time series predictions. MCD+N simultaneously estimates a heteroscedastic input–dependent data noise term (a trained error model attributable to observational noise) and a network weight uncertainty term (attributable to insufficiently constrained model parameters). Although MCD+N has appealing features, many heuristic approximations were employed during its derivation, and rigorous evaluations and evidence of its asserted capability to detect dissimilarity were lacking. To address this, we provided an in–depth evaluation of the scheme's potential and limitations. We showed that for reproducing soil moisture dynamics recorded by the Soil Moisture Active Passive (SMAP) mission, MCD+N indeed gave a good estimate of predictive error, provided that we tuned a hyperparameter and used a representative training data set. The input–dependent term responded strongly to observational noise, while the model term clearly acted as a detector for physiographic dissimilarity from the training data, behaving as intended. However, when the training and test data were characteristically different, the input–dependent term could be misled, undermining its reliability. Additionally, due to the data–driven nature of the model, data noise also influences network weight uncertainty, and therefore the two uncertainty terms are correlated. Altogether, this approach has promise, but care is needed to interpret the results.

54 ENVIRONMENTAL SCIENCES↗

Towards Enhancing Coding Productivity for GPU Programming Using Static Graphs

The main contribution of this work is to increase the coding productivity of GPU programming by using the concept of Static Graphs. GPU capabilities have been increasing significantly in terms of performance and memory capacity. However, there are still some problems in terms of scalability and limitations to the amount of work that a GPU can perform at a time. To minimize the overhead associated with the launch of GPU kernels, as well as to maximize the use of GPU capacity, we have combined the new CUDA Graph API with the CUDA programming model (including CUDA math libraries) and the OpenACC programming model. We use as test cases two different, well-known and widely used problems in HPC and AI: the Conjugate Gradient method and the Particle Swarm Optimization. In the first test case (Conjugate Gradient) we focus on the integration of Static Graphs with CUDA. In this case, we are able to significantly outperform the NVIDIA reference code, reaching an acceleration of up to 11x thanks to a better implementation, which can benefit from the new CUDA Graph capabilities. In the second test case (Particle Swarm Optimization), we complement the OpenACC functionality with the use of CUDA Graph, achieving again accelerations of up to one order of magnitude, with average speedups ranging from 2x to 4x, and performance very close to a reference and optimized CUDA code. Our main target is to achieve a higher coding productivity model for GPU programming by using Static Graphs, which provides, in a very transparent way, a better exploitation of the GPU capacity. The combination of using Static Graphs with two of the current most important GPU programming models (CUDA and OpenACC) is able to reduce considerably the execution time w.r.t. the use of CUDA and OpenACC only, achieving accelerations of up to more than one order of magnitude. Finally, we propose an interface to incorporate the concept of Static Graphs into the OpenACC Specifications.

58 GEOSCIENCES↗

Non-adiabatic approximations in time-dependent density functional theory: progress and prospects

Time-dependent density functional theory continues to draw a large number of users in a wide range of fields exploring myriad applications involving electronic spectra and dynamics. Although in principle exact, the predictivity of the calculations is limited by the available approximations for the exchange-correlation functional. In particular, it is known that the exact exchange-correlation functional has memory-dependence, but in practise adiabatic approximations are used which ignore this. Here we review the development of non-adiabatic functional approximations, their impact on calculations, and challenges in developing practical and accurate memory-dependent functionals for general purposes.

36 MATERIALS SCIENCE↗

Implicit fast sweeping method for hyperbolic systems of conservation laws

Implicit time-accurate methods are often used to integrate stiff problems where explicit schemes impose severe time step restrictions. This paper presents an efficient numerical framework based on the Fast Sweeping Method (FSM) for solving linear and nonlinear hyperbolic systems of conservation laws. The solution at each discrete location is computed by sweeping the numerical domain in several predetermined directions that follow the causality of the characteristic families. The use of a fractional step strategy eliminates the need for a solution selection criterion while one-sided stencils limit the number of sweeps to at most 2 d for d space dimensions. This work focuses on the first-order implicit upwind method since it constitutes the building block for high-order conservative schemes. For problems where the degree of stiffness evolves over time, implicit-explicit hybridization can be accomplished with the same algorithm by simply switching the stencil at each time level. As opposed to traditional implicit solvers, the sweeping method does not require a local time linearization of the fluxes thereby preserving the nonlinear stability properties of the original implicit scheme. It also avoids the large computational and memory requirements associated with solving large block-diagonal systems of equations. Here, a series of one- and two-dimensional test cases are presented for the inviscid Burgers' equation and the reactive Euler equations. The results indicate that the implicit FSM can allow a major reduction in the number of time steps even in the presence of discontinuous solution profiles.

74 ATOMIC AND MOLECULAR PHYSICS↗

Physics-informed latent neural operator for real-time predictions of time-dependent parametric PDEs

Deep operator network (DeepONet) has shown significant promise as surrogate models for systems governed by partial differential equations (PDEs), enabling accurate mappings between infinite-dimensional function spaces. However, when applied to systems with high-dimensional input-output mappings arising from large numbers of spatial and temporal collocation points, these models often require heavily overparameterized networks, leading to long training times. Latent DeepONet addresses some of these challenges by introducing a two-step approach: first learning a reduced latent space using a separate model, followed by operator learning within this latent space. While efficient, this method is inherently data-driven and lacks mechanisms for incorporating physical laws, limiting its robustness and generalizability in data-scarce settings. Here, in this work, we propose PI-Latent-NO, a physics-informed latent neural operator framework that integrates governing physics directly into the learning process. Our architecture features two coupled DeepONets trained end-to-end: a Latent-DeepONet that learns a low-dimensional representation of the solution, and a Reconstruction-DeepONet that maps this latent representation back to the physical space. By embedding PDE constraints into the training via automatic differentiation, our method eliminates the need for labeled training data and ensures physics-consistent predictions. The proposed framework is both memory and compute-efficient, exhibiting near-constant scaling with problem size and demonstrating significant speedups over traditional physics-informed operator models. We validate our approach on a range of parametric PDEs, showcasing its accuracy, scalability, and suitability for real-time prediction in complex physical systems.

Latent representations↗

Exploring temporal community evolution: algorithmic approaches and parallel optimization for dynamic community detection

Abstract Dynamic (temporal) graphs are a convenient mathematical abstraction for many practical complex systems including social contacts, business transactions, and computer communications. Community discovery is an extensively used graph analysis kernel with rich literature for static graphs. However, community discovery in a dynamic setting is challenging for two specific reasons. Firstly, the notion of temporal community lacks a widely accepted formalization, and only limited work exists on understanding how communities emerge over time. Secondly, the added temporal dimension along with the sheer size of modern graph data necessitates new scalable algorithms. In this paper, we investigate how communities evolve over time based on several graph metrics under a temporal formalization. We compare six different algorithmic approaches for dynamic community detection for their quality and runtime. We identify that a vertex-centric (local) optimization method works as efficiently as the classical modularity-based methods. To its advantage, such local computation allows for the efficient design of parallel algorithms without incurring a significant parallel overhead. Based on this insight, we design a shared-memory parallel algorithm DyComPar , which demonstrates between 4 and 18 fold speed-up on a multi-core machine with 20 threads, for several real-world and synthetic graphs from different domains.

97 MATHEMATICS AND COMPUTING↗

High Performance Adaptive Physics Refinement to Enable Large-Scale Tracking of Cancer Cell Trajectory

The ability to track simulated cancer cells through the circulatory system, important for developing a mechanistic understanding of metastatic spread, pushes the limits of today's supercomputers by requiring the simulation of large fluid volumes at cellular-scale resolution. To overcome this challenge, we introduce a new adaptive physics refinement (APR) method that captures cellular-scale interaction across large domains and leverages a hybrid CPU-GPU approach to maximize performance. Through algorithmic advances that integrate multi-physics and multi-resolution models, we establish a finely resolved window with explicitly modeled cells coupled to a coarsely resolved bulk fluid domain. In this work we present multiple validations of the APR framework by comparing against fully resolved fluid-structure interaction methods and employ techniques, such as latency hiding and maximizing memory bandwidth, to effectively utilize heterogeneous node architectures. Collectively, these computational developments and performance optimizations provide a robust and scalable framework to enable system-level simulations of cancer cell transport.

59 BASIC BIOLOGICAL SCIENCES↗

Sheath transitions in a cylindrical filament discharge: Axisymmetric 1D3V PIC-MCC simulations

We present the first nonplanar hot cathode discharge simulations that capture the role of the trapped-ions plasma, elucidating new phenomena unobservable in planar geometric discharges. A discharge struck between a single emitting wire filament cathode and a bounding anode is simulated in cylindrical geometry using an axisymmetric (radial) particle-in-cell Monte-Carlo collisions code. Operating the discharge near its ionization energy threshold can lead to the formation of a two plasma mode (TPM). One plasma forms in the conventional upstream region through electron impact ionization of background neutrals. A second plasma, whose global effect on the discharge was not previously well understood, forms downstream through the trapping of cold ions in the potential well of the filament’s virtual cathode, a process enabled by ion-neutral charge exchange collisions. Three space charge regions intersperse the electrode gap—an emissive sheath between the cathode filament and trapped-ions plasma, a double layer between the two plasmas, and a classical sheath between the upstream plasma and the outer anode. Simulations exhibit mode transitions and quenching instabilities that transform the discharge between the TPM and other single-plasma sheath modes that include classical (temperature-limited), space charge limited, and inverse (anode glow) modes. The transitions are explained via “aid-and-compete” dynamics wherein the growth of one plasma enhances growth in the other while concurrently exhibiting expansion dynamics antagonistic to each other. The system exhibits strong hysteresis memory during the mode transitions. Improved understanding and control of these sheath mode transitions are expected to benefit plasma applications with hot cathodes.

Electrical hysteresis↗

Fast and Flexible Inference Framework for Continuum Reverberation Mapping Using Simulation-based Inference with Deep Learning

Continuum reverberation mapping (CRM) of active galactic nuclei (AGN) monitors multiwavelength variability signatures to constrain accretion disk structure and supermassive black hole (SMBH) properties. The upcoming Vera Rubin Observatory’s Legacy Survey of Space and Time will survey tens of millions of AGN over the next decade, with thousands of AGN monitored with almost daily cadence in the deep drilling fields. However, existing CRM methodologies often require long computation time and are not designed to handle such large amounts of data. In this paper, we present a fast and flexible inference framework for CRM using simulation-based inference (SBI) with deep learning to estimate SMBH properties from AGN light curves. We use a long short-term memory summary network to reduce the high dimensionality of the light curve data and then use a neural density estimator to estimate the posterior of SMBH parameters. Using simulated light curves, we find SBI can produce more accurate SMBH parameter estimation with 10 3 –10 5 times speed up in inference efficiency compared to traditional methods. The SBI framework is particularly suitable for wide-field CRM surveys as the light curves will have identical observing patterns, which can be incorporated into the SBI simulation. We explore the performance of our SBI model on light curves with irregular-sampled, realistic observing cadence and alternative variability characteristics to demonstrate the flexibility and limitation of the SBI framework.

79 ASTRONOMY AND ASTROPHYSICS↗

Computationally Efficient Multiscale Neural Networks Applied to Fluid Flow in Complex 3D Porous Media

Abstract The permeability of complex porous materials is of interest to many engineering disciplines. This quantity can be obtained via direct flow simulation, which provides the most accurate results, but is very computationally expensive. In particular, the simulation convergence time scales poorly as the simulation domains become less porous or more heterogeneous. Semi-analytical models that rely on averaged structural properties (i.e., porosity and tortuosity) have been proposed, but these features only partly summarize the domain, resulting in limited applicability. On the other hand, data-driven machine learning approaches have shown great promise for building more general models by virtue of accounting for the spatial arrangement of the domains’ solid boundaries. However, prior approaches building on the convolutional neural network (ConvNet) literature concerning 2D image recognition problems do not scale well to the large 3D domains required to obtain a representative elementary volume (REV). As such, most prior work focused on homogeneous samples, where a small REV entails that the global nature of fluid flow could be mostly neglected, and accordingly, the memory bottleneck of addressing 3D domains with ConvNets was side-stepped. Therefore, important geometries such as fractures and vuggy domains could not be modeled properly. In this work, we address this limitation with a general multiscale deep learning model that is able to learn from porous media simulation data. By using a coupled set of neural networks that view the domain on different scales, we enable the evaluation of large ( $$>512^3$$ > 512 3 ) images in approximately one second on a single graphics processing unit. This model architecture opens up the possibility of modeling domain sizes that would not be feasible using traditional direct simulation tools on a desktop computer. We validate our method with a laminar fluid flow case using vuggy samples and fractures. As a result of viewing the entire domain at once, our model is able to perform accurate prediction on domains exhibiting a large degree of heterogeneity. We expect the methodology to be applicable to many other transport problems where complex geometries play a central role.

36 MATERIALS SCIENCE↗

Hierarchical Epoxy Structures via Tunable Polymerization-Induced Phase Separation Combined with Additive Manufacturing

Polymerization-induced phase separation (PIPS) allows for the control of thermoset morphologies and properties, enabling the tuning of domain sizes and thermomechanical response. However, its use in generating substructural features in additively manufactured materials has been limited. In this work, we combine epoxy PIPS with UV curable acrylate and rheological modifiers to print nano- to macro-phase separating materials via a two-step, dual-cure approach. This method enables direct ink write printing of hierarchical structures with both controlled morphologies through phase separation and macroscale architecture through print design. We find that formulations for phase-separating materials require judicious incorporation of additives to enable printability and to provide sufficient green strength. Atomic force microscopy-nano infrared mapping reveals tunable, reticulated nano- to micron-scale domains of the resultant multiphase materials and their morphology changes due to additives, resulting in alterations to thermomechanical and tensile properties. Shape memory behavior is also demonstrated through multimaterial additive manufacturing of epoxies with functionally graded internal morphology using active mixing techniques, highlighting this method’s ability to fabricate complex architectures with controlled morphologies and thermomechanical response.

Van Meter, Kylie E [Organic Materials Science, San↗

Dimensionality reduction of the many-body problem using coupled-cluster subsystem flow equations: classical and quantum computing perspective

We discuss reduced-scaling strategies employing recently introduced sub-system embedding sub-algebras coupled-cluster formalism (SES-CC) to describe many-body systems. These strategies utilize properties of the SES-CC formulations where the equations describing certain classes of sub- systems can be integrated into a computational flows composed coupled eigenvalue problems of reduced dimensionality. Additionally, these flows can be defined at the level of the CC Ansatz defined by selected classes of cluster amplitudes, which define the wave function ”memory” of possible partitionings of the many-body system into constituent sub-systems. One of the possible ways of solving these coupled problems is through implementing procedures, where the information is passed between the sub-systems in a self-consistent manner. As a special case, we consider local flow formulations where the so-called local character of correlation effects can be closely related to properties of sub-system embedding sub-algebras employing localized molecular basis. We also generalize flow equations to the time domain and to downfolding methods utilizing double exponential unitary CC Ansatz (DUCC), where reduced dimensionality of constituent sub-problems offer a possibility of efficient utilization of limited quantum resources in modeling realistic systems.

Electron correlation, quantum chemistry, quantum c↗

A Massively Parallel Implementation of the CCSD(T) Method Using the Resolution-of-the-Identity Approximation and a Hybrid Distributed/Shared Memory Parallelization Model

In this work, a parallel algorithm is described for the coupled-cluster singles and doubles method augmented with a perturbative correction for triple excitations [CCSD(T)] using the resolution-of-the-identity (RI) approximation for two-electron repulsion integrals (ERIs). The algorithm bypasses the storage of four-center ERIs by adopting an integral-direct strategy. The CCSD amplitude equations are given in a compact quasi-linear form by factorizing them in terms of amplitude-dressed three-center intermediates. A hybrid MPI/OpenMP parallelization scheme is employed, which uses the OpenMP-based shared memory model for intranode parallelization and the MPI-based distributed memory model for internode parallelization. Parallel efficiency has been optimized for all terms in the CCSD amplitude equations. Two different algorithms have been implemented for the rate-limiting terms in the CCSD amplitude equations that entail and -scaling computational costs, where N O and N V denote the number of correlated occupied and virtual orbitals, respectively. One of the algorithms assembles the four-center ERIs requiring N V 4 and N O 2 N V 2 -scaling memory costs in a distributed manner on a number of MPI ranks, while the other algorithm completely bypasses the assembling of quartic memory-scaling ERIs and thus largely reduces the memory demand. It is demonstrated that the former memory-expensive algorithm is faster on a few hundred cores, while the latter memory-economic algorithm shows a better strong scaling in the limit of a few thousand cores. The program is shown to exhibit a near-linear scaling, in particular for the compute-intensive triples correction step, on up to 8000 cores. The performance of the program is demonstrated via calculations involving molecules with 24–51 atoms and up to 1624 atomic basis functions. As the first application, the complete basis set (CBS) limit for the interaction energy of the π-stacked uracil dimer from the S66 data set has been investigated. This work reports the first calculation of the interaction energy at the CCSD(T)/aug-cc-pVQZ level without local orbital approximation. The CBS limit for the CCSD correlation contribution to the interaction energy was found to be -8.01 kcal/mol, which agrees very well with the value -7.99 kcal/mol reported by Schmitz, Hättig, and Tew [ Phys. Chem. Chem. Phys. 2014 , 16 , 22167-22178]. The CBS limit for the total interaction energy was estimated to be -9.64 kcal/mol.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Stochastic Vector Techniques in Ground-State Electronic Structure

Herein we review a suite of stochastic vector computational approaches for studying the electronic structure of extended condensed matter systems. These techniques help reduce algorithmic complexity, facilitate efficient parallelization, simplify computational tasks, accelerate calculations, and diminish memory requirements. While their scope is vast, we limit our study to ground-state and finite temperature density functional theory (DFT) and second-order many-body perturbation theory. More advanced topics, such as quasiparticle (charge) and optical (neutral) excitations and higher-order processes, are covered elsewhere. We start by explaining how to use stochastic vectors in computations, characterizing the associated statistical errors. Next, we show how to estimate the electron density in DFT and discuss effective techniques to reduce statistical errors. Finally, we review the use of stochastic vectors for calculating correlation energies within the second-order Møller-Plesset perturbation theory and its finite temperature variational form. Example calculation results are presented and used to demonstrate the efficacy of the methods.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Simultaneous measurement of the exchange parameter and saturation magnetization using propagating spin waves

The exchange interaction in ferromagnetic ultra thin films is a critical parameter in magnetization-based storage and logic devices, yet the accurate measurement of it remains a challenge. While a variety of approaches are currently used to determine the exchange parameter, each has its limitations, and good agreement among them has not been achieved. To date, neutron scattering, magnetometry, Brillouin light scattering, spin-torque ferromagnetic resonance spectroscopy, and Kerr microscopy have all been used to determine the exchange parameter. Here, we present a method that exploits the wavevector selectivity of Brillouin light scattering to measure the spin wave dispersion in both the backward volume and Damon–Eshbach orientations. The exchange, saturation magnetization, and magnetic thickness are then determined by a simultaneous fit of both dispersion branches with general spin wave theory without any prior knowledge of the thickness of a magnetic “dead layer.” In this study, we demonstrate the strength of this technique for ultrathin metallic films, typical of those commonly used in industrial applications for magnetic random-access memory.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Entangling Quantum Generative Adversarial Networks

Generative adversarial networks (GANs) are one of the most widely adopted machine learning methods for data generation. In this work, we propose a new type of architecture for quantum generative adversarial networks (an entangling quantum GAN, EQ-GAN) that overcomes limitations of previously proposed quantum GANs. Leveraging the entangling power of quantum circuits, the EQ-GAN converges to the Nash equilibrium by performing entangling operations between both the generator output and true quantum data. In the first multiqubit experimental demonstration of a fully quantum GAN with a provably optimal Nash equilibrium, we use the EQ-GAN on a Google Sycamore superconducting quantum processor to mitigate uncharacterized errors, and we numerically confirm successful error mitigation with simulations up to 18 qubits. Finally, we present an application of the EQ-GAN to prepare an approximate quantum random access memory and for the training of quantum neural networks via variational datasets.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Moment-based adaptive time integration for thermal radiation transport

Here, in this paper we develop a framework for moment-based adaptive time integration of deterministic multifrequency thermal radiation transpot (TRT). We generalize our recent semi-implicit-explicit (IMEX) integration framework for gray TRT to multifrequency TRT, and also introduce a semi-implicit variation that facilitates higher-order integration of TRT, where each stage is implicit in all components except opacities. To appeal to the broad literature on adaptivity with Runge–Kutta methods, we derive new embedded methods for four asymptotic preserving IMEX Runge–Kutta schemes we have found to be robust in our previous work on TRT and radiation hydrodynamics. We then use a moment-based high-order-low-order representation of the transport equations. Due to the high dimensionality, memory is always a concern in simulating TRT. We form error estimates and adaptivity in time purely based on temperature and radiation energy, for a trivial overhead in computational cost and memory usage compared with the base second order integrators. We then test the adaptivity in time on the tophat and Larsen problem, demonstrating the ability of the adaptive algorithm to naturally vary the timestep across 4–5 orders of magnitude, ranging from the dynamical timescales of the streaming regime to the thick diffusion limit.

97 MATHEMATICS AND COMPUTING↗

Probing boron vacancy defects in hBN via single spin relaxometry

Spin defects in solids offer promising platforms for quantum sensing and memory due to their long coherence times and optical addressability. Here, we integrate a single nitrogen-vacancy (NV) center in diamond with scanning probe microscopy to detect, read out, and spatially map spin-based quantum sensors at the nanoscale. Using the boron vacancy ($V$$^{–}_{B}$) center in hexagonal boron nitride—an emerging two-dimensional spin system—as a model, we detect its electron spin resonance indirectly via changes in the spin relaxation time (T 1 ) of a nearby NV center, eliminating the need for optical excitation or fluorescence detection of the $V$$^{–}_{B}$. Cross-relaxation between NV and $V$$^{–}_{B}$ ensembles significantly reduces NV T1, enabling quantitative nanoscale mapping of defect densities beyond the optical diffraction limit and clear resolution of hyperfine splitting in isotopically enriched h 10 B 15 N. Our method demonstrates interactions between spin sensors in 3D and 2D materials, establishing NV centers as versatile probes for characterizing otherwise inaccessible spin defects.

Quantum metrology↗