Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “generalized algorithm”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 307 records · Page 17

LDRD Abbreviated report: High-Order General-Discrete-Ordinates Method Enabling Efficient Deterministic Transport in Hydrodynamic Simulations

Deterministic transport simulations for national-security and energy applications often operate in high-dimensional phase-space, where accuracy and cost both become major challenges. A common numerical artifact in such problems is the “ray-effect,” which appears as unphysical streaks. Beyond misinterpretation, these artifacts can contaminate tightly coupled physics, such as fluid dynamics, radiation-hydrodynamics, and laser-plasma interactions, eroding the predictive capability of entire multiphysics workflows. Our objective was to make high-dimension studies practical on modern hardware while mitigating the ray-effect without relying on prohibitively expensive sampling approaches such as Monte Carlo methods. We developed the Generic Discretization Library (GenDiL), a Graphics Processing Unit (GPU)-first framework that uses high-order Discontinuous Galerkin (DG) methods and matrix-free algorithms to reduce memory usage and improve computational efficiency, critical for phase-space simulations. GenDiL supports phase-space adaptivity in both mesh size and polynomial order (hp-adaptivity) to place resolution only where it is needed. A central capability is Local Dimensional Refinement (LDR), which couples lower-dimension continuum models to higher-dimension kinetic models through stable and conservative interfaces, so that high-fidelity physics is applied only in regions where it is essential. Building on the GenDiL framework, we developed the General SN (GSN) family of algorithms as a true generalization of the polar SN approach (discrete ordinates, often denoted SN). Rather than tying discrete ordinates to a specific polar change of coordinates, GSN formulates transport on an arbitrary change of coordinates chosen to reduce ray-effect. We studied two complementary variants: an analytic variant, where the coordinate map is prescribed in advance by a closed-form function; and a data-driven variant, where a quantity of interest, such as the net flux, guides the coordinate system. GenDiL provides the library infrastructure for efficient GPU execution, but the GSN concept is algorithmic and independent of any one library. Across representative high-dimension tests, including non-symmetric solutions, both variants delivered strong ray-effect mitigation at practical cost, moving four- to six-dimensional analysis toward repeatable, routine studies.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Quantum Computer-Aided Design: Digital Quantum Simulation of Quantum Processors

With the increasing size of quantum processors, submodules that constitute the processor hardware will become too large to accurately simulate on a classical computer. Therefore, one would soon have to fabricate and test each new design primitive and parameter choice in time-consuming coordination between design, fabrication, and experimental validation. Here we show how one can design and test the performance of next-generation quantum hardware—by using existing quantum computers. Focusing on superconducting transmon processors as a prominent hardware platform, we compute the static and dynamic properties of individual and coupled transmons. We show how the energy spectra of transmons can be obtained by variational hybrid quantum-classical algorithms that are well suited for near-term noisy quantum computers. In addition, single- and two-qubit gate simulations are demonstrated via Suzuki-Trotter decomposition. Our methods pave a promising way towards designing candidate quantum processors when the demands of calculating submodule properties exceed the capabilities of classical computing resources.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

ACCERT

The main function of ACCERT (The Algorithm for the Capital Cost Estimation of Reactor Technologies) is to provide an item-by-item estimate of the cost of a facility, at present a nuclear reactor complex, primarily nuclear power stations. The core of ACCERT is the large number of algorithms that have been developed and will continue to be developed. ACCERT is also a general methodology for identifying and organizing the individual items that are estimated using the ACCERT algorithms. ACCERT also summarize status and results; to save results and pull information from previous analyses; and to provide an interactive dynamic graphical user interface for a wide range of functions and visualizations.

Zhou, Jia [Argonne National Laboratory (ANL), Argo↗

Lanczos Algorithm, the Transfer Matrix, and the Signal-to-Noise Problem

This Letter introduces a method for determining the energy spectrum of lattice quantum chromodynamics by applying the Lanczos algorithm to the transfer matrix and using a bootstrap generalization of the Cullum-Willoughby method to filter out spurious eigenvalues. Proof-of-principle analyses of the simple harmonic oscillator and the lattice quantum chromodynamics proton mass demonstrate that this method provides faster ground-state convergence than the “effective mass,” which is related to the power-iteration algorithm. Lanczos provides more accurate energy estimates than multistate fits to correlation functions with small imaginary times while achieving comparable statistical precision. Two-sided error bounds are computed for Lanczos results and guarantee that excited-state effects cannot shift Lanczos results far outside their statistical uncertainties.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Super Resolving Unrolled Neural Networks for Remote Sensing

In remote sensing systems, the capabilities of the system are constrained by the complex interactions between size, weight, and power (SWAP) of potential designs. In electro-optical (EO) systems, examples of these critical parameters include the system’s sensitivity and resolution. Those parameters can be increased by ever larger optical apertures and focal planes but at the cost of more SWAP. Multi-image super resolution (MISR) techniques allow resolution to be enhanced via computation rather than more sophisticated optical hardware. These algorithms combine multiple images together into a single, higher resolution image, trading temporal resolution and computation for spatial resolution. Fielded MISR techniques, such as Drizzle, can require several hundred images to create a single super resolved image, implying reduced temporal resolution, increased data acquisition load, and limiting mission applications. Iterative techniques, such as model-based image reconstruction and compressive sensing, have been shown to create super resolved images using fewer images than Drizzle. They do this by posing an optimization problem that balances accuracy between a highly accurate physical model and an image model. In the case of super resolution, the physical model is defined by the relation between low resolution input images and the desired high resolution output image. The image model encodes some assumptions about the super resolved image. These assumptions are meant to suppress reconstruction artifacts that arise due to deterministic physical model error, stochastic measurement noise, and potential undersampling. In practice, the performance of iterative methods are limited by imaging models compatible with optimization. Deep learning-based methods can effectively learn image models of arbitrary complexity, but lack the theoretical explainability and robustness of iterative techniques. Consensus equilibrium (CE) generalizes the iterative techniques beyond optimization, enabling blackbox algorithms such as traditional and neural image denoisers to be used as the image model. CE-based approaches retain much of the explainability and robustness of iterative techniques while allowing the expressiveness of machine learning image models to be used. Additionally, by unrolling iterations of CE with an embedded image denoiser, the image denoiser can be further trained and specialized to the specific application with potentially higher quality reconstructions. Under this project, we demonstrated the feasibility of training an unrolled neural network based upon CE. While we didn’t train one, we showed that the CE process is differentiable and its gradient can be tractably computed. We also explored the usage of a variants of CE akin to generative neural works. Most importantly, we applied the CE framework to a number of problems including non-blind deconvolution, upsampling, single-image super resolution, MISR, event-based sensing, and saturated deconvolution. Our MISR prototype creates high quality reconstructions with an order of magnitude fewer images than previous approaches and, critically, produces these reconstructions fast enough for practical usage.

47 OTHER INSTRUMENTATION↗

RxnRover/CyRxnOpt

CyRxnOpt aims to provide a single software interface to various optimization algorithms, mainly designed for chemical process optimization applications. CyRxnOpt generalizes the optimization process into four high-level “phases”: Installation, Configuration, Training, and Prediction. This allows developers to program to a general interface for each phase of the optimization, simplifying the development of user-friendly tools to lower the barrier of entry into chemical process optimization, especially for automated laboratory workflows which can greatly benefit from access to various optimization techniques. It is also designed so researchers can easily add new or existing algorithms into existing workflows in a user-friendly manner.

Kulathunga, Dulitha Prasanna [Iowa State Universit↗

FAIR AI models in high energy physics

Abstract The findable, accessible, interoperable, and reusable (FAIR) data principles provide a framework for examining, evaluating, and improving how data is shared to facilitate scientific discovery. Generalizing these principles to research software and other digital products is an active area of research. Machine learning models—algorithms that have been trained on data without being explicitly programmed—and more generally, artificial intelligence (AI) models, are an important target for this because of the ever-increasing pace with which AI is transforming scientific domains, such as experimental high energy physics (HEP). In this paper, we propose a practical definition of FAIR principles for AI models in HEP and describe a template for the application of these principles. We demonstrate the template’s use with an example AI model applied to HEP, in which a graph neural network is used to identify Higgs bosons decaying to two bottom quarks. We report on the robustness of this FAIR AI model, its portability across hardware architectures and software frameworks, and its interpretability.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

GentenMPI: Distributed Memory Sparse Tensor Decomposition

GentenMPl is a toolkit of sparse canonical polyadic (CP) tensor decomposition algorithms that is designed to run effectively on distributed-memory high-performance computers. Its use of distributed-memory parallelism enables it to efficiently decompose tensors that are too large for a single compute node's memory. GentenMPl leverages Sandia's decades-long investment in the Trilinos solver framework for much of its parallel-computation capability. Trilinos contains numerical algorithms and linear algebra classes that have been optimized for parallel simulation of complex physical phenomena. This work applies these tools to the data science problem of sparse tensor decomposition. In this report, we describe the use of Trilinos in GentenMPl, extensions needed for sparse tensor decomposition, and implementations of the CP-ALS (CP via alternating least squares) and GCP-SGD (generalized CP via stochastic gradient descent) sparse tensor decomposition algorithms. We show that GentenMPl can decompose sparse tensors of extreme size, e.g., a 12.6-terabyte tensor on 8192 computer cores. We demonstrate that the Trilinos backbone provides good strong and weak scaling of the tensor decomposition algorithms.

97 MATHEMATICS AND COMPUTING↗

Hardware acceleration for HPS algorithms in two and three dimensions

We provide a flexible, open-source framework for hardware acceleration, namely massively-parallel execution on general-purpose graphics processing units (GPUs), applied to the hierarchical Poincaré–Steklov (HPS) family of algorithms for building fast direct solvers for linear elliptic partial differential equations. To take full advantage of the power of hardware acceleration, we propose two variants of HPS algorithms to improve performance on two- and three-dimensional problems. In the two-dimensional setting, we introduce a novel recomputation strategy that minimizes costly data transfers to and from the GPU; in three dimensions, we modify and extend the adaptive discretization technique of Geldermans and Gillman [1] to greatly reduce peak memory usage. We provide an open-source implementation of these methods written in JAX, a high-level accelerated linear algebra package, which allows for the first integration of a high-order fast direct solver with automatic differentiation tools. We conclude with extensive numerical examples showing our methods are fast and accurate on two- and three-dimensional problems.

Fast direct solvers↗

Haar-Like Wavelets on Hierarchical Trees

Here, discrete wavelet methods, originally formulated in the setting of regularly sampled signals, can be adapted to data defined on a point cloud if some multiresolution structure is imposed on the cloud. A wide variety of hierarchical clustering algorithms can be used for this purpose, and the multiresolution structure obtained can be encoded by a hierarchical tree of subsets of the cloud. Prior work introduced the use of Haar-like bases defined with respect to such trees for approximation and learning tasks on unstructured data. This paper builds on that work in two directions. First, we present an algorithm for constructing Haar-like bases on general discrete hierarchical trees. Second, with an eye towards data compression, we present thresholding techniques for data defined on a point cloud with error controlled in the $L$ $\infty$ norm and in a Hölder-type norm. In a concluding trio of numerical examples, we apply our methods to compress a point cloud dataset, study the tightness of the $L$ $\infty$ error bound, and use thresholding to identify MNIST classifiers with good generalizability.

97 MATHEMATICS AND COMPUTING↗

A Stochastic Quasi-Newton Method in the Absence of Common Random Numbers

We present Q-SASS, a quasi-Newton method for unconstrained stochastic optimization that does not rely on common random numbers. Most existing quasi-Newton approaches leverage common random numbers to construct second-order updates. However, motivated by challenges in variational quantum algorithms—where such coordination is not possible—we consider the setting in which function values and gradients are accessible only through noisy probabilistic zeroth- and first-order oracles, and no common random numbers can be exploited. We derive high-probability tail bounds on the iteration complexity of our algorithm for nonconvex, convex, and strongly convex (more generally, those satisfying the PL condition) objective functions. Finally, we demonstrate the empirical benefits of our quasi-Newton updating scheme on both synthetic and quantum chemistry problems.

Complexity bound↗

Assessing the potential of deep learning for protein–ligand docking

The effects of ligand binding on protein structures and their in vivo functions carry numerous implications for modern biomedical research and biotechnology development efforts such as drug discovery. Although several deep learning (DL) methods and benchmarks designed for protein–ligand docking have recently been introduced, so far no previous works have systematically studied the behaviour of the latest docking and structure prediction methods within the broadly applicable context of: (1) using predicted (apo) protein structures for docking (for example, for applicability to new proteins); (2) binding multiple (cofactor) ligands concurrently to a given target protein (for example, for enzyme design); and (3) having no previous knowledge of binding pockets (for example, for generalization to unknown pockets). To enable a deeper understanding of the real-world utility of docking methods, we introduce PoseBench, a comprehensive benchmark for broadly applicable protein–ligand docking. PoseBench enables researchers to rigorously and systematically evaluate DL methods for apo-to-holo protein–ligand docking and protein–ligand structure prediction using both primary ligand and multiligand benchmark datasets, the latter of which we introduce to the DL community. Empirically, using PoseBench, we find that: (1) DL cofolding methods generally outperform comparable conventional and DL docking baseline algorithms, but popular methods such as AlphaFold 3 are still challenged by prediction targets with new protein–ligand binding poses; (2) certain DL cofolding methods are highly sensitive to their input multiple sequence alignments, whereas others are not; and (3) DL methods struggle to strike a balance between structural accuracy and chemical specificity when predicting new or multiligand protein targets.

Morehead, Alex [Lawrence Berkeley National Laborat↗

Integration of scanning probe microscope with high-performance computing: Fixed-policy and reward-driven workflows implementation

The rapid development of computation power and machine learning algorithms has paved the way for automating scientific discovery with a scanning probe microscope (SPM). The key elements toward operationalization of the automated SPM are the interface to enable SPM control from Python codes, availability of high computing power, and development of workflows for scientific discovery. Here, we build a Python interface library that enables controlling an SPM from either a local computer or a remote high-performance computer, which satisfies the high computation power need of machine learning algorithms in autonomous workflows. We further introduce a general platform to abstract the operations of SPM in scientific discovery into fixed-policy or reward-driven workflows. Furthermore, our work provides a full infrastructure to build automated SPM workflows for both routine operations and autonomous scientific discovery with machine learning.

47 OTHER INSTRUMENTATION↗

Kinetic theory of particle-in-cell simulation plasma and the ensemble averaging technique

Abstract We derive the kinetic theory of fluctuations in physically and numerically stable particle-in-cell (PIC) simulations of electrostatic plasmas. The starting point is the single-time correlation at the start of the simulation between the statistical fluctuations of the weighted densities of macroparticle centers in the plasma particle phase-space. The fluctuations are associated with different initial conditions, typically due to the random initial conditions (in velocity space) of the macroparticles/simulation plasma, assigned according to their initial distribution of probability. The single-time correlations at all time steps and in each spatial grid cell are then determined from the Laplace–Fourier transforms of the discretized Klimontovich-like equation for the macroparticles and Maxwell’s equations for the fields, as computed by modern PIC codes. We recover the expressions for the electrostatic field and the plasma particle density fluctuation autocorrelation spectra as well as the kinetic equations describing the average evolution of PIC-simulated plasma particles, first derived by Langdon (1970b Proc. 4th Conf. Numerical Simulation of Plasmas ) using a test macroparticle approach perturbing a discretized Vlasovian plasma and then averaging the obtained physical quantity over the initial macroparticle velocity distribution. We generalize and extend these results to the modern algorithms in PIC codes using arbitrary macroparticle weights. Analytical estimates of statistical fluctuation amplitudes are derived as a function of the plasma simulation parameters, using the central limit theorem in the limit of a large number of macroparticles per cell. The theory is then used to analyze the ensemble averaging technique of PIC simulations where statistical averages are performed over ensembles of PIC simulations, modeling the same plasma physics problem but using different statistical realizations of the initial distribution functions of the macroparticles. This method is illustrated by linear Landau damping uncovering (from noise, which is usually considered numerical) the physical fluctuations driven by a single small amplitude electrostatic wave perturbing a PIC simulation plasma in equilibrium.

fluctuations correlations↗

Developments in SRW Code and Sirepo Framework Supporting Simulation of Time-Dependent Coherent X-ray Scattering Experiments

Physical optics simulations for beamlines and experiments are essential for the effective use of synchrotron light source facilities such as NSLS-II at BNL. The SRW software package supports such source-to-detector simulations for coherent X-ray scattering and imaging experiments through its Python interface and Sirepo browser-based graphical user interface. This allows one to define custom sample models, assess the feasibility of an experiment, and estimate most appropriate beamline settings before using valuable beamtime. We discuss the recent use of general-purpose GPU resources and coherent mode decomposition algorithms in SRW to accelerate physical optics simulations with partially coherent X-rays. To illustrate these new capabilities, we describe simulations of typical time series of partially coherent scattering images used in X-ray Photon Correlation Spectroscopy (XPCS) experiments; aiming to characterize the nanoscale dynamics of a disordered sample, representing a solution of nanoparticles undergoing Brownian diffusion.

36 MATERIALS SCIENCE↗

Towards automating structural discovery in scanning transmission electron microscopy *

Abstract Scanning transmission electron microscopy is now the primary tool for exploring functional materials on the atomic level. Often, features of interest are highly localized in specific regions in the material, such as ferroelectric domain walls, extended defects, or second phase inclusions. Selecting regions to image for structural and chemical discovery via atomically resolved imaging has traditionally proceeded via human operators making semi-informed judgements on sampling locations and parameters. Recent efforts at automation for structural and physical discovery have pointed towards the use of ‘active learning’ methods that utilize Bayesian optimization with surrogate models to quickly find relevant regions of interest. Yet despite the potential importance of this direction, there is a general lack of certainty in selecting relevant control algorithms and how to balance a priori knowledge of the material system with knowledge derived during experimentation. Here we address this gap by developing the automated experiment workflows with several combinations to both illustrate the effects of these choices and demonstrate the tradeoffs associated with each in terms of accuracy, robustness, and susceptibility to hyperparameters for structural discovery. We discuss possible methods to build descriptors using the raw image data and deep learning based semantic segmentation, as well as the implementation of variational autoencoder based representation. Furthermore, each workflow is applied to a range of feature sizes including NiO pillars within a La:SrMnO 3 matrix, ferroelectric domains in BiFeO 3 , and topological defects in graphene. The code developed in this manuscript is open sourced and will be released at github.com/nccreang/AE_Workflows .

47 OTHER INSTRUMENTATION↗

Quantum parton shower with kinematics

Parton showers, which can efficiently incorporate quantum interference effects, have been shown to be run efficiently on quantum computers. However, so far these quantum parton showers have not included the full kinematical information required to reconstruct an event, which in classical parton showers requires the use of a veto algorithm. Here, in this work, we show that adding one extra assumption about the discretization of the evolution variable allows one to construct a quantum veto algorithm, which reproduces the full quantum interference in the event, and allows one to include kinematical effects. We finally show that for certain initial states the quantum interference effects generated in this veto algorithm are classically tractable, such that an efficient classical algorithm can be devised.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Experimental Measurement of Out-of-Time-Ordered Correlators at Finite Temperature

Out-of-time-ordered correlators (OTOCs) are a key observable in a wide range of interconnected fields including many-body physics, quantum information science, and quantum gravity. Measuring OTOCs using near-term quantum simulators will extend our ability to explore fundamental aspects of these fields and the subtle connections between them. Here, we demonstrate an experimental method to measure OTOCs at finite temperatures and use the method to study their temperature dependence. These measurements are performed on a digital quantum computer running a simulation of the transverse field Ising model. Our flexible method, based on the creation of a thermofield double state, can be extended to other models and enables us to probe the OTOC’s temperature-dependent decay rate. Measuring this decay rate opens up the possibility of testing the fundamental temperature-dependent bounds on quantum information scrambling.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗