Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “multiple iterations”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

A Multilevel Approach For SolvingLarge-Scale QUBO Problems With Noisy Hybrid Quantum Approximate Optimization

Quantum approximate optimization is one ofthe promising candidates for useful quantum computation,particularly in the context of finding approximate solutionsto Quadratic Unconstrained Binary Optimization (QUBO)problems. However, the existing quantum processing units(QPUs) are of relatively small size, and canonical mappingsof QUBO via the Ising model require one qubit per vari-able, rendering direct large-scale optimization infeasible.In classical optimization, a general strategy for addressingmany large-scale problems is via multilevel/multigrid meth-ods, where the large target problem is iteratively coarsenedand the global solution is constructed from multiple small-scale optimization runs. In this work, we experimentallytest how existing QPUs perform when used as a sub-solverwithin such a multilevel strategy. To this aim, we com-bine and extend (via additional classical processing steps)the recently proposed Noise-Directed Adaptive Remapping(NDAR) and Quantum Relax&Round (QRR) algorithms.We first demonstrate the effectiveness of our heuristicextensions on Rigetti’s superconducting transmon deviceAnkaa-2. We find approximate solutions to10instances offully connected82-qubit Sherrington-Kirkpatrick graphswith random integer-valued coefficients obtaining normal-ized approximation ratios (ARs) in the range∼0.98−1.0,and the same class with real-valued coefficients (ARs∼0.94−1.0). Then, we implement the extended NDAR andQRR algorithms as subsolvers in the multilevel algorithmfor6large-scale graphs with at most∼27,000variables.In practice, the QPU (with classical post-processing steps)is used to find approximate solutions to dozens of at most82-qubit problems, which are iteratively used to constructthe global solution. We observe that quantum optimizationresults are competitive in terms of the quality of solutionswhen compared to classical heuristics used as subsolverswithin the multilevel approach.Reproducibility: source code and data are available at[TBA upon acceptance]

quantum computing↗

A Proof of the Asymptotic Variance of Path Length Estimators for Single-Collision Monte Carlo Source Iteration in the Thick Diffusion Limit

Here, we prove a theorem relating the variance of path length estimators for single-collision Monte Carlo source iteration to a parameter that becomes infinitesimally small in an important physical regime arising in radiative transfer. In our usage, “single-collision Monte Carlo source iteration” refers to Monte Carlo Boltzmann transport methods in which each Monte Carlo particle history includes no more than a single collision, and the physics of multiple scattering is modeled by lagging the scattering source term and iterating until this term converges. Our theorem can be used to construct variance reduction techniques which improve the order of the estimator variance. This enables calculations that would otherwise require impractically large sample sizes to achieve practical estimator uncertainties. We believe this is the first postulation of a theorem relating estimator variance to a limiting case parameter for single-collision Monte Carlo source iteration, and the first proof of such a theorem. We illustrate the theorem’s value with an example in which the authors of a transport method used the theorem to design a variance reduction technique that improved the uncertainty of their solution by a factor of about 500 for a proxy problem from radiative transfer that contains both optically-thick and optically-thin material.

Mathematics and Computing↗

Compensator improvement for multivariable control systems

A theory and the associated numerical technique are developed for an iterative design improvement of the compensation for linear, time-invariant control systems with multiple inputs and multiple outputs. A strict constraint algorithm is used in obtaining a solution of the specified constraints of the control design. The result of the research effort is the multiple input, multiple output Compensator Improvement Program (CIP). The objective of the Compensator Improvement Program is to modify in an iterative manner the free parameters of the dynamic compensation matrix so that the system satisfies frequency domain specifications. In this exposition, the underlying principles of the multivariable CIP algorithm are presented and the practical utility of the program is illustrated with space vehicle related examples.

Mitchell, J. R.↗

TTDFT: A GPU accelerated Tucker tensor DFT code for large-scale Kohn-Sham DFT calculations

We present the Tucker tensor DFT (TTDFT) code which uses a tensor-structured algorithm with graphic processing unit (GPU) acceleration for conducting ground-state DFT calculations on large-scale systems. The Tucker tensor DFT algorithm uses a localized Tucker tensor basis computed from an additive separable approximation to the Kohn-Sham Hamiltonian. The discrete Kohn-Sham problem is solved using Chebyshev filtered subspace iteration method that relies on matrix-matrix multiplications of a sparse symmetric Hamiltonian matrix and a dense wavefunction matrix, expressed in the localized Tucker tensor basis. These matrix-matrix multiplication operations, which constitute the most computationally intensive step of the solution procedure, are GPU accelerated providing ~8-fold GPU-CPU speedup for these operations on the largest systems studied. In conclusion, the computational performance of the TTDFT code is presented using benchmark studies on aluminum nano-particles and silicon quantum dots with system sizes ranging up to ~7,000 atoms.

97 MATHEMATICS AND COMPUTING↗

Sparse measurement medical CT reconstruction using multi-fused block matching denoising priors

A major challenge for medical X-ray CT imaging is reducing the number of X-ray projections to lower radiation dosage and reduce scan times without compromising image quality. However these under-determined inverse imaging problems rely on the formulation of an expressive prior model to constrain the solution space while remaining computationally tractable. Traditional analytical reconstruction methods like Filtered Back Projection (FBP) often fail with sparse measurements, producing artifacts due to their reliance on the Shannon-Nyquist Sampling Theorem. Consensus Equilibrium, which is a generalization of Plug and Play, is a recent advancement in Model-Based Iterative Reconstruction (MBIR), has facilitated the use of multiple denoisers are prior models in an optimization free framework to capture complex, non-linear prior information. However, 3D prior modelling in a Plug and Play approach for volumetric image reconstruction requires long processing time due to high computing requirement. Instead of directly using a 3D prior, this work proposes a BM3D Multi Slice Fusion (BM3D-MSF) prior that uses multiple 2D image denoisers fused to act as a fully 3D prior model in Plug and Play reconstruction approach. Our approach does not require training and are thus able to circumvent ethical issues related with patient training data and are readily deployable in varying noise and measurement sparsity levels. In addition, reconstruction with the BM3D-MSF prior achieves similar reconstruction image quality as fully 3D image priors, but with significantly reduced computational complexity. We test our method on clinical CT data and demonstrate that our approach improves reconstructed image quality.

Hossain, Maliha [ORNL]↗

wa-hls4ml: A Benchmark and Surrogate Models for hls4ml Resource and Latency Estimation

As machine learning (ML) is increasingly implemented in hardware to address real-time challenges in scientific applications, the development of advanced toolchains has significantly reduced the time required to iterate on various designs. These advancements have solved major obstacles, but also exposed new challenges. For example, processes that were not previously considered bottlenecks, such as hardware synthesis, are becoming limiting factors in the rapid iteration of designs. To mitigate these emerging constraints, multiple efforts have been undertaken to develop an ML-based surrogate model that estimates resource usage of ML accelerator architectures. We introduce wa-hls4ml, a benchmark for ML accelerator resource and latency estimation, and its corresponding initial dataset of over 680,000 fully connected and convolutional neural networks, all synthesized using hls4ml and targeting Xilinx FPGAs. The benchmark evaluates the performance of resource and latency predictors against several common ML model architectures, primarily originating from scientific domains, as exemplar models, and the average performance across a subset of the dataset. Additionally, we introduce GNN- and transformer-based surrogate models that predict latency and resources for ML accelerators. We present the architecture and performance of the models and find that the models generally predict latency and resources for the 75% percentile within several percent of the synthesized resources on the synthetic test dataset.

Hawks, Benjamin [Fermilab] (ORCID:0000000157000288↗

Multiqubit entanglement generation with squeezed modes

We present a hybrid continuous variable–discrete variable entanglement generation protocol using linear optics and homodyne measurements, capable of producing multiple high-fidelity Bell pairs per protocol iteration, with an approximate 0.5 success probability. The effectiveness of the protocol is determined by the squeezing strength. To increase the number of Bell pairs, approximately 3 dB of extra squeezing is needed for each additional Bell pair. The protocol also generates single Bell pairs with an approximate 0.75 probability for squeezing strengths ⪅ 15 dB , achievable with current technology.

Macridin, Alexandru [Fermilab] (ORCID:000000022228↗

Research in computer science

Synopses are given for NASA supported work in computer science at the University of Virginia. Some areas of research include: error seeding as a testing method; knowledge representation for engineering design; analysis of faults in a multi-version software experiment; implementation of a parallel programming environment; two computer graphics systems for visualization of pressure distribution and convective density particles; task decomposition for multiple robot arms; vectorized incomplete conjugate gradient; and iterative methods for solving linear equations on the Flex/32.

Ortega, J. M.↗

Accurate orbit determination strategies for the tracking and data relay satellites

The National Aeronautics and Space Administration (NASA) has developed the Tracking and Data Relay Satellite (TDRS) System (TDRSS) for tracking and communications support of low Earth-orbiting satellites. TDRSS has the operational capability of providing 85% coverage for TDRSS-user spacecraft. TDRSS currently consists of five geosynchronous spacecraft and the White Sands Complex (WSC) at White Sands, New Mexico. The Bilateration Ranging Transponder System (BRTS) provides range and Doppler measurements for each TDRS. The ground-based BRTS transponders are tracked as if they were TDRSS-user spacecraft. Since the positions of the BRTS transponders are known, their radiometric tracking measurements can be used to provide a well-determined ephemeris for the TDRS spacecraft. For high-accuracy orbit determination of a TDRSS user, such as the Ocean Topography Experiment (TOPEX)/Poseidon spacecraft, high-accuracy TDRS orbits are required. This paper reports on successive refinements in improved techniques and procedures leading to more accurate TDRS orbit determination strategies using the Goddard Trajectory Determination System (GTDS). These strategies range from the standard operational solution using only the BRTS tracking measurements to a sophisticated iterative process involving several successive simultaneous solutions for multiple TDRSs and a TDRSS-user spacecraft. Results are presented for GTDS-generated TDRS ephemerides produced in simultaneous solutions with the TOPEX/Poseidon spacecraft. Strategies with different user spacecraft, as well as schemes for recovering accurate TDRS orbits following a TDRS maneuver, are also presented. In addition, a comprehensive assessment and evaluation of alternative strategies for TDRS orbit determination, excluding BRTS tracking measurements, are presented.

Oza, D. H.↗

Gesture Based Control and EMG Decomposition

This paper presents two probabilistic developments for use with Electromyograms (EMG). First described is a new-electric interface for virtual device control based on gesture recognition. The second development is a Bayesian method for decomposing EMG into individual motor unit action potentials. This more complex technique will then allow for higher resolution in separating muscle groups for gesture recognition. All examples presented rely upon sampling EMG data from a subject's forearm. The gesture based recognition uses pattern recognition software that has been trained to identify gestures from among a given set of gestures. The pattern recognition software consists of hidden Markov models which are used to recognize the gestures as they are being performed in real-time from moving averages of EMG. Two experiments were conducted to examine the feasibility of this interface technology. The first replicated a virtual joystick interface, and the second replicated a keyboard. Moving averages of EMG do not provide easy distinction between fine muscle groups. To better distinguish between different fine motor skill muscle groups we present a Bayesian algorithm to separate surface EMG into representative motor unit action potentials. The algorithm is based upon differential Variable Component Analysis (dVCA) [l], [2] which was originally developed for Electroencephalograms. The algorithm uses a simple forward model representing a mixture of motor unit action potentials as seen across multiple channels. The parameters of this model are iteratively optimized for each component. Results are presented on both synthetic and experimental EMG data. The synthetic case has additive white noise and is compared with known components. The experimental EMG data was obtained using a custom linear electrode array designed for this study.

Wheeler, Kevin R.↗

Iterative multi-task learning and inference from seismic images

Seismic interpretation aims to extract quantitative and interpretable attributes from a seismic image produced using some migration method to inform characteristics of a subsurface reservoir or target of interest. Current paradigms for computing seismic attributes mostly rely on single-task algorithms. We develop an iterative, multi-task machine learning method to learn and infer multiple attributes from a seismic image. This method is composed of two stages: a multi-task inference stage and a multi-modal, multi-task refinement stage. The basic mechanism of this method is that we train a multi-task inference neural network (NN) to estimate a set of attributes, including a relative geological time (RGT), a denoised higher-resolution (DHR) seismic image, and multiple fault attributes (including probability, dip, and strike), from a low-resolution, noisy seismic image; then we input the inferred attributes to a multi-task refinement NN to enhance the raw inference results iteratively. The two multi-task NNs are trained separately based on synthetic seismic images and associated attributes generated by a geological modeling algorithm. The software we intend to release is a PyTorch implementation of this multi-task learning method for both 2D and 3D cases along with scripts to run the training/validation. The algorithm and software can be a useful tool for automatic seismic interpretation.

Gao, Kai↗

Macromolecular phasing using diffraction from multiple crystal forms

A phasing algorithm for macromolecular crystallography is proposed that utilizes diffraction data from multiple crystal forms – crystals of the same molecule with different unit-cell packings (different unit-cell parameters or space-group symmetries). The approach is based on the method of iterated projections, starting with no initial phase information. The practicality of the method is demonstrated by simulation using known structures that exist in multiple crystal forms, assuming some information on the molecular envelope and positional relationships between the molecules in the different unit cells. With incorporation of new or existing methods for determination of these parameters, the approach has potential as a method for ab initio phasing.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Iterative pass optimization of sequence data

The problem of determining the minimum-cost hypothetical ancestral sequences for a given cladogram is known to be NP-complete. This "tree alignment" problem has motivated the considerable effort placed in multiple sequence alignment procedures. Wheeler in 1996 proposed a heuristic method, direct optimization, to calculate cladogram costs without the intervention of multiple sequence alignment. This method, though more efficient in time and more effective in cladogram length than many alignment-based procedures, greedily optimizes nodes based on descendent information only. In their proposal of an exact multiple alignment solution, Sankoff et al. in 1976 described a heuristic procedure--the iterative improvement method--to create alignments at internal nodes by solving a series of median problems. The combination of a three-sequence direct optimization with iterative improvement and a branch-length-based cladogram cost procedure, provides an algorithm that frequently results in superior (i.e., lower) cladogram costs. This iterative pass optimization is both computation and memory intensive, but economies can be made to reduce this burden. An example in arthropod systematics is discussed. c2003 The Willi Hennig Society. Published by Elsevier Science (USA). All rights reserved.

NASA Discipline Evolutionary Biology↗

Advanced Simulation of ITER Core X-ray Crystal Spectroscopy

X-Ray Simulation Analysis (XRSA) is an analytical ray-tracing mixed code developed specifically for the ITER Core X-Ray Crystal Spectroscopy (XRCS-Core) diagnostic, which employs a dual-reflection configuration incorporating multiple pre-reflectors made of Highly Oriented Pyrolytic Graphite (HOPG) and spherically curved analyzing crystals. The ITER XRCS-Core is designed for high spectral resolution measurement in specific wavelength ranges, including narrow bands around 1.354 Å for W 64+ , 2.19 Å for Xe 51+ , and 2.555 Å for Xe 44+ and Xe 47+ , enabling diagnostic capability across a broad electron temperature range in the ITER plasma. XRSA facilitates efficient simulation of the spectral performance of this complex X-ray spectroscopic system. Recent updates to the XRSA code have incorporated two critical effects: auto-focusing, which specifically applies to HOPG, and polarization. These two effects are particularly important in the dual-reflection configuration used in the ITER XRCS-Core system to provide more accurate modeling results. Here, simulations conducted with the updated code demonstrate that polarization has a substantial impact on the performance of the dual-reflection system. Additionally, the combined influence of polarization and system layout introduces performance variations across channels through the same crystal.

Computer graphics↗

Persistent Sampling: Enhancing the Efficiency of Sequential Monte Carlo

Sequential Monte Carlo (SMC) samplers are powerful tools for Bayesian inference but suffer from high computational costs due to their reliance on large particle ensembles for accurate estimates. We introduce persistent sampling (PS), an extension of SMC that systematically retains and reuses particles from all prior iterations to construct a growing, weighted ensemble. By leveraging multiple importance sampling and resampling from a mixture of historical distributions, PS mitigates the need for excessively large particle counts, directly addressing key limitations of SMC such as particle impoverishment and mode collapse. Crucially, PS achieves this without additional likelihood evaluations-weights for persistent particles are computed using cached likelihood values. This framework not only yields more accurate posterior approximations but also produces marginal likelihood estimates with significantly lower variance, enhancing reliability in model comparison. Furthermore, the persistent ensemble enables efficient adaptation of transition kernels by leveraging a larger, decorrelated particle pool. Experiments on high-dimensional Gaussian mixtures, hierarchical models, and non-convex targets demonstrate that PS consistently outperforms standard SMC and related variants, including recycled and waste-free SMC, achieving substantial reductions in mean squared error for posterior expectations and evidence estimates, all at reduced computational cost. PS thus establishes itself as a robust, scalable, and efficient alternative for complex Bayesian inference tasks.

Karamanis, Minas↗

DART-PFLOTRAN: An ensemble-based data assimilation system for estimating subsurface flow and transport model parameters

Ensemble-based Data Assimilation (EDA), based on the Monte Carlo approach, has been effectively applied to estimate model parameters through inverse modeling in subsurface flow and transport problems. However, implementation of EDA approach involves a complicated workflow that include setting up and executing ensemble forward model simulations, processing observations and model simulation results for parameter updates, and repeat for sequential or iterative EDA. To facilitate the management of such workflow and lower the barriers for adopting EDA-based parameter estimation in subsurface science, we develop a generic software frame-work linking the Data Assimilation Research Testbed (DART) with a massively parallel subsurface FLOw and TRANsport code PFLOTRAN. The new DART-PFLOTRAN leverages both the core data assimilation engines in DART and the computational power afforded by PFLOTRAN. In addition to the standard smoother and filtering options, DART-PFLOTRAN enables an iterative EDA workflow based on the Ensemble Smoother for Multiple Data Assimilation method (ES-MDA) to improve estimation accuracy for nonlinear forward problems. Here, we verify the implementation of ES-MDA in DART-PFLOTRAN using two synthetic cases designed to estimate static permeability and dynamic exchange fluxes across the riverbed, respectively, from continuous temperature measurements made across a depth profile. One-dimensional hydro-thermal simulations are performed in both cases to relate temperature responses with the parameters of interest. In the case of estimating dynamic parameters, we demonstrate the flexibility of DART-PFLOTRAN in automating sequential ES-MDA workflow, which will significantly reduce the time researchers spend on managing complex workflows in similar applications. Both studies yield accurate estimations of the parameters compared to their synthetic truth, while ES-MDA leads to more accurate estimation when a high level of nonlinearity exist between observed responses and unknown parameters. With a code base in Python and Fortran, DART-PFLOTRAN paves the way for applications in large-scale subsurface inverse modeling by automating the complex workflow of sequential ES-MDA that can be executed on various computing platforms.

97 MATHEMATICS AND COMPUTING↗

Toward accurate measurement of electromagnetic field by retrieving and refining the center position of non-uniform diffraction disks in Lorentz 4D-STEM

Recent advancement in scanning transmission electron microscopy (STEM) allows the use of 4D-STEM, a technique that captures an electron diffraction pattern at each scan point in STEM, to measure electrostatic and magnetic potential and field in materials. However, accurate measurement, separation of the magnetic and electric signals, and removal of artifacts remain challenging, especially in the presence of complex non-uniform diffraction contrast within the disks. In this work, based on dynamic simulations of 4D-STEM patterns built upon superstructures consisting of millions of atoms to account for different sample thickness and edge geometries, we show how the shape and intensity distribution of the central disk are affected by multiple scattering. We propose a robust refinement procedure through iteration of the spin-sensitive peak position of the disk-center in the circular Hough transform filtered images from experimental Lorentz 4D-STEM dataset after minimizing the possible artifacts, such as those due to the change of thickness, dynamic scattering, and scanning process. We verify that caution must be taken as in practice the rigid-disk-shift model used to reconstruct induction maps can easily break down due to disk-protrusion when there exists a nonconstant phase gradient or thickness within the width of the probe. Through quantitative analysis and comparing experiment with calculation the effect of the non-spin-related intensity distribution inside the disk as well as that causes the disk shift due to the intensity-protrusion can be removed, and high-quality magnetic field mapping is possible.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

Discovering nuclear models from symbolic machine learning

Numerous phenomenological nuclear models have been proposed to describe specific observables within different regions of the nuclear chart. However, developing a unified model that describes the complex behavior of all nuclei remains an open challenge. Here, we explore whether symbolic Machine Learning (ML) can rediscover traditional nuclear physics models or identify alternatives with improved simplicity, fidelity, and predictive power. To address this challenge, we developed a Multi-objective Iterated Symbolic Regression approach that handles symbolic regressions over multiple target observables, accounts for experimental uncertainties and is robust against high-dimensional problems. As a proof of principle, we applied this method to describe the nuclear binding energies and charge radii of light and medium mass nuclei. Our approach identified simple analytical relationships based on the number of protons and neutrons, providing interpretable models with precision comparable to state-of-the-art nuclear models. Additionally, we integrated this ML-discovered model with an existing complementary model to estimate the limits of nuclear stability. These results highlight the potential of symbolic ML to develop accurate nuclear models and guide our description of complex many-body problems.

Nuclear structure↗