Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “approximate computing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 253 records · Page 14

Phase-curve Pollution of Exoplanet Transit Depths

The next generation of space telescopes will enable transformative science to understand the nature and origin of exoplanets. In particular, transit spectroscopy will reveal the chemical composition of the exoplanet atmospheres with unprecedented detail thanks to precise measurements of the visible-to-infrared transit depths down to 10 parts per million. Such a level of instrumental precision raises the challenge to obtain even more precise astrophysical models so as not to significantly influence the interpretation of the observed data. We must therefore critically revisit some of the commonly accepted assumptions that were adequate for analyzing past and current observations. A common approximation in the analysis of exoplanetary primary transits is that the planet does not contribute to the recorded flux, so-called dark planet hypothesis. In this paper, we investigate the impact of the dark planet hypothesis on the parameters obtained from the analysis of transits with particular attention to the transit depth. We develop mathematical formulae and release new software to estimate the magnitude of the potential bias. These tools will be useful in the preparation of observing proposals, as well as within the scientific consortia of the James Webb Space Telescope (JWST) and the Atmospheric Remote-sensing Infrared Exoplanet Large-survey (ARIEL) missions. We probe the accuracy of the mathematical formulae through the analysis of synthetic observations with the JWST Mid-InfraRed Instrument. We find that self-blending from nightside emission attenuates the transit depth by >3σ for some of the known exoplanet systems, in agreement with previous work. An additional unreported effect caused by the nightside rotating into view can also impart a significant effect, but in the opposite direction (increasing the transit depth); this effect can largely be removed with conventional detrending practices, at the expense of a slight increase in noise, and mixing astrophysical variations and instrumental drifts.

79 ASTRONOMY AND ASTROPHYSICS↗

Constraints on the Emission of Gamma-Rays from M31 with HAWC

Cosmic rays, along with stellar radiation and magnetic fields, are known to make up a significant fraction of the energy density of galaxies such as the Milky Way. When cosmic rays interact in the interstellar medium, they produce gamma-ray emission which provides an important indication of how the cosmic rays propagate. Gamma rays from the Andromeda Galaxy (M31), located 785 kpc away, provide a unique opportunity to study cosmic-ray acceleration and diffusion in a galaxy with a structure and evolution very similar to the Milky Way. Using 33 months of data from the High Altitude Water Cherenkov Observatory, we search for TeV gamma rays from the galactic plane of M31. We also investigate past and present evidence of galactic activity in M31 by searching for Fermi Bubble-like structures above and below the galactic nucleus. No significant gamma-ray emission is observed, so we use the null result to compute upper limits on the energy density of cosmic rays $>10$ TeV in M31. In conclusion, the computed upper limits are approximately ten times higher than expected from the extrapolation of the Fermi LAT results.

79 ASTRONOMY AND ASTROPHYSICS↗

Threshold progressions in covering and packing contexts

Here, using standard methods (due to Janson, Stein–Chen, and Talagrand) from probabilistic combinatorics, we explore the following general theme: As one progresses from each member of a family of objects $\mathcal{A}$ being “covered” by at most one object in a random collection $\mathcal{C}$, to being covered at most λ times, to being covered at least once, to being covered at least λ times, a hierarchy of thresholds emerge. We will use examples from extremal set theory, combinatorics, and additive number theory to see how these results vary according to the context, and level of dependence introduced.

97 MATHEMATICS AND COMPUTING↗

Bias-Variance Trade-Off in Physics-Informed Neural Networks with Randomized Smoothing for High-Dimensional PDEs

Physics-Informed Neural Networks (PINNs) have triggered a paradigm shift in scientific computing, leveraging mesh-free properties and robust approximation capabilities. While proving effective for low-dimensional partial differential equations (PDEs), the computational cost of PINNs remains a hurdle in high-dimensional scenarios. This is particularly pronounced when computing high-order and high-dimensional derivatives in the physics-informed loss. Randomized Smoothing PINN (RS-PINN) introduces Gaussian noise for stochastic smoothing of the original neural net model, enabling the use of Monte Carlo methods for derivative approximation, which eliminates the need for costly automatic differentiation. Despite its computational efficiency, especially in the approximation of high-dimensional derivatives, RS-PINN introduces biases in both loss and gradients, negatively impacting convergence, especially when coupled with stochastic gradient descent (SGD) algorithms. We present a comprehensive analysis of biases in RS-PINN, attributing them to the nonlinearity of the Mean Squared Error (MSE) loss as well as the intrinsic nonlinearity of the PDE itself. We propose tailored bias correction techniques, delineating their application based on the order of PDE nonlinearity. The derivation of an unbiased RS-PINN allows for a detailed examination of its advantages and disadvantages compared to the biased version. Specifically, the biased version has a lower variance and runs faster than the unbiased version, but it is less accurate due to the bias. To optimize the bias-variance trade-off, we combine the two approaches in a hybrid method that balances the rapid convergence of the biased version with the high accuracy of the unbiased version. In addition to methodological contributions, we present an enhanced implementation of RS-PINN. Extensive experiments on diverse high-dimensional PDEs, including Fokker-Planck, Hamilton-Jacobi-Bellman (HJB), viscous Burgers’, Allen-Cahn, and Sine-Gordon equations, illustrate the bias-variance trade-off and highlight the effectiveness of the hybrid RS-PINN. Empirical guidelines are provided for selecting biased, unbiased, or hybrid versions, depending on the dimensionality and nonlinearity of the specific PDE problem.

97 MATHEMATICS AND COMPUTING↗

A Perspective on Quantum Computing Applications in Quantum Chemistry Using 25-100 Logical Qubits

The intersection of quantum computing and quantum chemistry represents a promising frontier for achieving quantum utility in domains of both scientific and societal relevance. Owing to the exponential growth of classical resource requirements for simulating quantum systems, quantum chemistry has long been recognized as a natural candidate for quantum computation. This perspective focuses on identifying scientifically meaningful use cases where early fault-tolerant quantum computers, which are considered to be equipped with approximately 25-100 logical qubits, could deliver tangible impact. While recent advances in classical computing have pushed the boundaries of tractable simulations to unprecedented scales, this logical-qubit regime represents the first window where quantum devices can pursue qualitatively distinct strategies, such as polynomial-scaling phase estimation, direct simulation of quantum dynamics, and active-space embedding, that remain challenging for classical solvers, such as multireference charge-transfer and conical-intersection states central to photochemistry and materials design. We highlight near-term opportunities in algorithm and software design, discuss representative chemical problems suited for quantum acceleration, and propose strategic roadmaps and collaborative pathways for advancing practical quantum utility in quantum chemistry.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Physically regularized machine learning emulators of aerosol activation

Abstract. The activation of aerosol into cloud droplets is an important step in the formation of clouds and strongly influences the radiative budget of the Earth. Explicitly simulating aerosol activation in Earth system models is challenging due to the computational complexity required to resolve the necessary chemical and physical processes and their interactions. As such, various parameterizations have been developed to approximate these details at reduced computational cost and accuracy. Here, we explore how machine learning emulators can be used to bridge this gap in computational cost and parameterization accuracy. We evaluate a set of emulators of a detailed cloud parcel model using physically regularized machine learning regression techniques. We find that the emulators can reproduce the parcel model at higher accuracy than many existing parameterizations. Furthermore, physical regularization tends to improve emulator accuracy, most significantly when emulating very low activation fractions. This work demonstrates the value of physical constraints in machine learning model development and enables the implementation of improved hybrid physical and machine learning models of aerosol activation into next-generation Earth system models.

58 GEOSCIENCES↗

Performance of Cloud 3D Solvers in Ice Cloud Shortwave Radiation Closure Over the Equatorial Western Pacific Ocean

Abstract For retrieving cloud optical properties from satellite images or computing these properties from climate model output, computationally efficient treatments of cloud horizontal inhomogeneity include the Monte Carlo Independent Column Approximation (McICA) and the Tripleclouds method. Computationally efficient treatment of cloud horizontal radiation exchanges includes the SPeedy Algorithm for Radiative TrAnsfer through CloUd Sides (SPARTACUS). As a test to derive properties from satellite images, we collocate Moderate Resolution Imaging Spectroradiometer (MODIS) cloud retrievals with near‐nadir Cloud and the Earth's Radiant Energy System (CERES) footprints in July 2008 over an equatorial western Pacific Ocean region to compare the performance of the McICA, Tripleclouds, and SPARTACUS solvers to the conventional plane‐parallel homogeneous (PPH) treatment. PPH overestimates cloud albedo, and the three solvers effectively reduce overestimation with root mean square error of shortwave upwelling irradiance decreasing between 15.72 and 18.53 W m −2 , or about 22%–25%. Although cloud top variability does not get fed into the simulations, all three solvers also reduce the effect of cloud top variability on cloud albedo. Entrapment (energy reflected downward from clouds) and horizontal radiation transfer have opposite effects on the SPARTACUS cloud albedo simulation. The net effect depends on the cloud vertical extent, the unawareness of which limits the performance of the SPARTACUS solver.

54 ENVIRONMENTAL SCIENCES↗

An Analog Preconditioner for Solving Linear Systems [Slides]

This presentation concludes in situ computation enables new approaches to linear algebra problems which can be both more effective and more efficient as compared to conventional digital systems. Preconditioning is well-suited to analog computation due to the tolerance for approximate solutions. When combined with prior work on in situ MVM for scientific computing, analog preconditioning can enable significant speedups for important linear algebra applications.

97 MATHEMATICS AND COMPUTING↗

MAGMA: Enabling exascale performance with accelerated BLAS and LAPACK for diverse GPU architectures

MAGMA (Matrix Algebra for GPU and Multicore Architectures) is a pivotal open-source library in the landscape of GPU-enabled dense and sparse linear algebra computations. With a repertoire of approximately 750 numerical routines across four precisions, MAGMA is deeply ingrained in the DOE software stack, playing a crucial role in high-performance computing. Notable projects such as ExaConstit, HiOP, MARBL, and STRUMPACK, among others, directly harness the capabilities of MAGMA. In addition, the MAGMA development team has been acknowledged multiple times for contributing to the vendors’ numerical software stacks. Looking back over the time of the Exascale Computing Project (ECP), we highlight how MAGMA has adapted to recent changes in modern HPC systems, especially the growing gap between CPU and GPU compute capabilities, as well as the introduction of low precision arithmetic in modern GPUs. We also describe MAGMA’s direct impact on several ECP projects. Maintaining portable performance across NVIDIA and AMD GPUs, and with current efforts toward supporting Intel GPUs, MAGMA ensures its adaptability and relevance in the ever-evolving landscape of GPU architectures.

97 MATHEMATICS AND COMPUTING↗

Optimal design of acoustic metamaterial cloaks under uncertainty

In this work, we consider the problem of optimal design of an acoustic cloak under uncertainty and develop scalable approximation and optimization methods to solve this problem. The design variable is taken as an infinite-dimensional spatially-varying field that represents the material property, while an additive infinite-dimensional random field represents the variability of the material property or the manufacturing error. Discretization of this optimal design problem results in high-dimensional design variables and uncertain parameters. To solve this problem, we develop a computational approach based on a Taylor approximation and an approximate Newton method for optimization, which is based on a Hessian derived at the mean of the random field. We show our approach is scalable with respect to the dimension of both the design variables and uncertain parameters, in the sense that the necessary number of acoustic wave propagations is essentially independent of these dimensions, for numerical experiments with up to one million design variables and half a million uncertain parameters. Additionally, we demonstrate that, using our computational approach, an optimal design of the acoustic cloak that is robust to material uncertainty is achieved in a tractable manner. The optimal design under uncertainty problem is posed and solved for the classical circular obstacle surrounded by a ring-shaped cloaking region, subjected to both a single-direction single-frequency incident wave and multiple-direction multiple-frequency incident waves. Finally, we apply the method to a deterministic large-scale optimal cloaking problem with complex geometry, to demonstrate that the approximate Newton method’s Hessian computation is viable for large, complex problems.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Posterior Covariance Matrix Approximations

Here, the Davis equation of state (EOS) is commonly used to model thermodynamic relationships for high explosive (HE) reactants. Typically, the parameters in the EOS are calibrated, with uncertainty, using a Bayesian framework and Markov Chain Monte Carlo (MCMC) methods. However, MCMC methods are computationally expensive, especially for complex models with many parameters. This paper provides a comparison between MCMC and less computationally expensive Variational methods (Variational Bayesian and Hessian Variational Bayesian) for computing the posterior distribution and approximating the posterior covariance matrix based on heterogeneous experimental data. All three methods recover similar posterior distributions and posterior covariance matrices. This study demonstrates that for this EOS parameter calibration application, the assumptions made in the two Variational methods significantly reduce the computational cost but do not substantially change the results compared to MCMC.

97 MATHEMATICS AND COMPUTING↗

Integration of multiple coinflip devices for high-quality random sampling

Artificial intelligence, scientific computing, and probabilistic computing use random sampling to approximate solutions to various problems, with larger models requiring a substantial quantity of random numbers. To generate the required vast quantity of random numbers at high rates, we explore so-called “coinflip” devices, which are stochastic microelectronic devices ideally capable of independently generating random bits with a tunable weight at a high rate. However, coinflip devices are inherently analog and demonstrate nonidealities, like temperature dependence and drift, that can introduce determinism into the outputs. We present important considerations for building systems of multiple coinflip devices to produce high-quality bitstreams with low error and little dependency on previous bits. Using tunnel diodes as coinflip devices, we implement a control loop to adapt to temperature dependence and generate fair bitstreams with each device. While this can lead to dependencies between bits in a single bitstream, we demonstrate that combining results generated in parallel with individual tunnel diodes can produce fair and unpredictable bitstreams. The suitability of these bitstreams for use in probabilistic computing is then demonstrated through a Monte Carlo approximation of π.

Taylor, Brady Garland [Sandia National Laboratorie↗

PLANC: Parallel Low-rank Approximation with Nonnegativity Constraints

In this work, we consider the problem of low-rank approximation of massive dense nonnegative tensor data, for example, to discover latent patterns in video and imaging applications. As the size of data sets grows, single workstations are hitting bottlenecks in both computation time and available memory. We propose a distributed-memory parallel computing solution to handle massive data sets, loading the input data across the memories of multiple nodes, and performing efficient and scalable parallel algorithms to compute the low-rank approximation. We present a software package called Parallel Low-rank Approximation with Nonnegativity Constraints, which implements our solution and allows for extension in terms of data (dense or sparse, matrices or tensors of any order), algorithm (e.g., from multiplicative updating techniques to alternating direction method of multipliers), and architecture (we exploit GPUs to accelerate the computation in this work). We describe our parallel distributions and algorithms, which are careful to avoid unnecessary communication and computation, show how to extend the software to include new algorithms and/or constraints, and report efficiency and scalability results for both synthetic and real-world data sets.

97 MATHEMATICS AND COMPUTING↗

Multiphysics Degradation Modeling of Energy Storage Materials via RKPM with a Neural Network-Enhancement

In energy storage materials, strong electrochemical-mechanical coupling and highly anisotropic material properties contribute to the formation and propagation of micro-cracking during charge/discharge cycling, resulting in reduced performance and service life. A coupled electro-chemo-mechanical reproducing kernel particle method (RKPM) formulation is developed, and a patch-test is formulated to certify optimal convergence of the proposed RKPM method for the coupled physics system. With microstructural images supplied by the National Renewable Energy Laboratory (NREL), pixel-based model construction by RKPM is then used to represent the complex material microstructures for modeling the coupled physics of these systems. Further, a neural network-enhanced reproducing kernel particle method (NN-RKPM) [1, 2] is introduced to effectively model damage and crack propagation in the material microstructures; the location, orientation, and solution transition near a localization are automatically captured by superimposed block-level NN optimizations. This NN enrichment approach allows for effective modeling of localizations via a fixed background discretization, relieving tedious efforts for adaptive refinement in traditional mesh-based methods. Applications to the heterogeneous microstructures of Li-ion battery cathodes will be presented to demonstrate the effectiveness of the proposed methods. Reference: [1] Baek, J., Chen, J. S., Susuki, K., "Neural Network enhanced Reproducing Kernel Particle Method for Modeling Localizations," International Journal for Numerical Methods in Engineering, Vol. 123, pp 4422-4454, https://doi.org/10.1002/nme.7040, 2022. [2] Baek, J., Chen, J. S., "A Neural Network-Based Enrichment of Reproducing Kernel Approximation for Modeling Brittle Fracture", Computer Methods in Applied Mechanics and Engineering Vol. 410, 116590, 2024.

electro-chemo-mechanical coupling↗

Auger spectroscopy beyond the ultra-short core-hole relaxation time approximation

Abstract We present a time-dependent computational approach to study Auger electron spectroscopy (AES) beyond the ultra-short core-hole relaxation time approximation and, as a test case, we apply it to the paradigmatic example of a one-dimensional Mott insulator represented by a half-filled Hubbard chain. The Auger spectrum is usually calculated by assuming that, after the creation of a core-hole, the system thermalizes almost instantaneously. This leads to a relatively simple analytical expression that uses the ground-state with a core-hole as a reference state and ignores all the transient dynamics related to the screening of the core-hole. In this picture, the response of the system can be associated to the pair spectral function. On the other hand, in our numerical calculations, the core hole is created by a light pulse, allowing one to study the transient dynamics of the system in terms of the pulse duration and in the non-perturbative regime. Time-dependent density matrix renormalization group calculations reveal that the relaxation process involves the creation of a polarization cloud of doublon excitations that have an effect similar to photo-doping. As a consequence, there is a leak of spectral weight to higher energies into what otherwise would be the Mott gap. For longer pulses, these excited states, mostly comprised of doublons, can dominate the spectrum. By changing the duration of the light-pulse, the entire screening process can be resolved in time.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

Randomized Algorithms for Symmetric Nonnegative Matrix Factorization

Symmetric Nonnegative Matrix Factorization (SymNMF) is a technique in data analysis and machine learning that approximates a matrix with a product of a nonnegative, low-rank matrix and it transpose. To design faster and more scalable algorithms for SymNMF we develop two randomized algorithms for its computation. The first method uses randomized matrix sketching to compute an initial low-rank approximation to the input matrix and proceeds to uses this as a low-rank input to rapidly compute a SymNMF. The second methods uses randomized leverage score sampling to approximately solve constrained least squares problems. Many successful methods for SymNMF rely on (approximately) solving sequences of constrained least squares problems. Here, we prove theoretically that leverage score sampling can approximately solve constrained least squares problems to e-accuracy. Finally we demonstrate both methods work in practice by applying them to graph clustering tasks on large real world data sets. These experiments show that our methods approximately maintain solution quality and achieve significant speed ups for both large dense and large sparse problems.

97 MATHEMATICS AND COMPUTING↗

Efficient Neural Network Approaches for Conditional Optimal Transport with Applications in Bayesian Inference

In this work, we present two neural network approaches that approximate the solutions of static and dynamic conditional optimal transport (COT) problems. Both approaches enable conditional sampling and conditional density estimation, which are core tasks in Bayesian inference—particularly in the simulation-based (“likelihood-free”) setting. Our methods represent the target conditional distribution as a transformation of a tractable reference distribution. Obtaining such a transformation, chosen here to be an approximation of the COT map, is computationally challenging even in moderate dimensions. To improve scalability, our numerical algorithms use neural networks to parameterize candidate maps and further exploit the structure of the COT problem. Our static approach approximates the map as the gradient of a partially input convex neural network. It uses a novel numerical implementation to increase computational efficiency compared to state-of-the-art alternatives. Our dynamic approach approximates the conditional optimal transport via the flow map of a regularized neural ODE; compared to the static approach, it is slower to train but offers more modeling choices and can lead to faster sampling. We demonstrate both algorithms numerically, comparing them with competing state-of-the-art approaches, using benchmark datasets and simulation-based Bayesian inverse problems.

97 MATHEMATICS AND COMPUTING↗

Numerical solution of singular Lyapunov equations

We consider the numerical solution of large scale singular (continuous-time) Lyapunov equations of the form AX + XA T + BB T = 0, where A is semistable, that is, its spectrum is contained in the left half plane, with the exception of a few semisimple eigenvalues at zero. We also consider the case of a few semisimple eigenvalues on the imaginary axis. We assume that we know these few eigenvalues (zero or imaginary), and that we have or can compute the corresponding invariant subspaces. We use this information to build an appropriate newly proposed subspace on which to project the Lyapunov equations, and then compute a low-rank approximation to the least squares solution. Selected illustrative numerical examples are provided.

97 MATHEMATICS AND COMPUTING↗