Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “multiple iterations”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Static actuator-sharing algorithm for concurrent control of multiple plasma properties

Simultaneous regulation of multiple properties in next-generation tokamaks like ITER and fusion pilot plant may require the integration of different plasma control algorithms. Such integration requires the conversion of individual controller commands into physical actuator requests while accounting for the coupling between different plasma properties. This work proposes a tokamak and scenario-agnostic actuator-sharing algorithm (ASA) to perform the above-mentioned command-request conversion and, hence, integrate multiple plasma controllers. The proposed algorithm implicitly solves a quadratic programming (QP) problem formulated to account for the saturation limits and the relation between the controller commands and physical actuator requests. Since the constraints arising in the QP program are linear, the proposed ASA is highly computationally efficient and can be implemented in the tokamak plasma control system in real time. Furthermore, the proposed algorithm is designed to handle real-time changes in the control objectives and actuators’ availability. Nonlinear simulations carried out using the Control Oriented Transport SIMulator illustrate the effectiveness of the proposed algorithm in achieving multiple control objectives simultaneously.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Solving Coupled Cluster Equations by the Newton Krylov Method

We describe using the Newton Krylov method to solve the coupled cluster equation. The method uses a Krylov iterative method to compute the Newton correction to the approximate coupled cluster amplitude. The multiplication of the Jacobian with a vector, which is required in each step of a Krylov iterative method such as the Generalized Minimum Residual (GMRES) method, is carried out through a finite difference approximation, and requires an additional residual evaluation. The overall cost of the method is determined by the sum of the inner Krylov and outer Newton iterations. We discuss the termination criterion used for the inner iteration and show how to apply pre-conditioners to accelerate convergence. We will also examine the use of regularization technique to improve the stability of convergence and compare the method with the widely used direct inversion of iterative subspace (DIIS) methods through numerical examples.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Approximate Inverse Chain Preconditioner: Iteration Count Case Study for Spectral Support Solvers

As the growing availability of computational power slows, there has been an increasing reliance on algorithmic advances. However, faster algorithms alone will not necessarily bridge the gap in allowing computational scientists to study problems at the edge of scientific discovery in the next several decades. Often, it is necessary to simplify or precondition solvers to accelerate the study of large systems of linear equations commonly seen in a number of scientific fields. Preconditioning a problem to increase efficiency is often seen as the best approach; yet, preconditioners which are fast, smart, and efficient do not always exist. Following the progress of [1], we present a new preconditioner for symmetric diagonally dominant (SDD) systems of linear equations. These systems are common in certain PDEs, network science, and supervised learning among others. Based on spectral support graph theory, this new preconditioner builds off of the work of [2], computing and applying a V-cycle chain of approximate inverse matrices. This preconditioner approach is both algebraic in nature as well as hierarchically-constrained depending on the condition number of the system to be solved. Due to its generation of an Approximate Inverse Chain of matrices, we refer to this as the AIC preconditioner. We further accelerate the AIC preconditioner by utilizing precomputations to simplify setup and multiplications in the con-text of an iterative Krylov-subspace solver. While these iterative solvers can greatly reduce solution time, the number of iterations can grow large quickly in the absence of good preconditioners. Initial results for the AIC preconditioner have shown a very large reduction in iteration counts for SDD systems as compared to standard preconditioners such as Incomplete Cholesky (ICC) and Multigrid (MG). We further show significant reduction in iteration counts against the more advanced Combinatorial Multigrid (CMG) preconditioner. We have further developed no-fill sparsification techniques to ensure that the computational cost of applying the AIC preconditioner does not grow prohibitively large as the depth of the V-cycle grows for systems with larger condition numbers. Our numerical results have shown that these sparsifiers maintain the sparsity structure of our system while also displaying significant reductions in iteration counts.1 2

97 MATHEMATICS AND COMPUTING↗

Approximate Inverse Chain Preconditioner: Iteration Count Case Study for Spectral Support Solvers

As the growing availability of computational power slows, there has been an increasing reliance on algorithmic advances. However, faster algorithms alone will not necessarily bridge the gap in allowing computational scientists to study problems at the edge of scientific discovery in the next several decades. Often, it is necessary to simplify or precondition solvers to accelerate the study of large systems of linear equations commonly seen in a number of scientific fields. Preconditioning a problem to increase efficiency is often seen as the best approach; yet, preconditioners which are fast, smart, and efficient do not always exist. Following the progress of [1], we present a new preconditioner for symmetric diagonally dominant (SDD) systems of linear equations. These systems are common in certain PDEs, network science, and supervised learning among others. Based on spectral support graph theory, this new preconditioner builds off of the work of [2], computing and applying a V-cycle chain of approximate inverse matrices. This preconditioner approach is both algebraic in nature as well as hierarchically-constrained depending on the condition number of the system to be solved. Due to its generation of an Approximate Inverse Chain of matrices, we refer to this as the AIC preconditioner. We further accelerate the AIC preconditioner by utilizing precomputations to simplify setup and multiplications in the con-text of an iterative Krylov-subspace solver. While these iterative solvers can greatly reduce solution time, the number of iterations can grow large quickly in the absence of good preconditioners. Initial results for the AIC preconditioner have shown a very large reduction in iteration counts for SDD systems as compared to standard preconditioners such as Incomplete Cholesky (ICC) and Multigrid (MG). We further show significant reduction in iteration counts against the more advanced Combinatorial Multigrid (CMG) preconditioner. We have further developed no-fill sparsification techniques to ensure that the computational cost of applying the AIC preconditioner does not grow prohibitively large as the depth of the V-cycle grows for systems with larger condition numbers. Our numerical results have shown that these sparsifiers maintain the sparsity structure of our system while also displaying significant reductions in iteration counts.1 2

97 MATHEMATICS AND COMPUTING↗

A Proof of the Asymptotic Variance of Path Length Estimators for Single-Collision Monte Carlo Source Iteration in the Thick Diffusion Limit

Here, we prove a theorem relating the variance of path length estimators for single-collision Monte Carlo source iteration to a parameter that becomes infinitesimally small in an important physical regime arising in radiative transfer. In our usage, “single-collision Monte Carlo source iteration” refers to Monte Carlo Boltzmann transport methods in which each Monte Carlo particle history includes no more than a single collision, and the physics of multiple scattering is modeled by lagging the scattering source term and iterating until this term converges. Our theorem can be used to construct variance reduction techniques which improve the order of the estimator variance. This enables calculations that would otherwise require impractically large sample sizes to achieve practical estimator uncertainties. We believe this is the first postulation of a theorem relating estimator variance to a limiting case parameter for single-collision Monte Carlo source iteration, and the first proof of such a theorem. We illustrate the theorem’s value with an example in which the authors of a transport method used the theorem to design a variance reduction technique that improved the uncertainty of their solution by a factor of about 500 for a proxy problem from radiative transfer that contains both optically-thick and optically-thin material.

Mathematics and Computing↗

TTDFT: A GPU accelerated Tucker tensor DFT code for large-scale Kohn-Sham DFT calculations

We present the Tucker tensor DFT (TTDFT) code which uses a tensor-structured algorithm with graphic processing unit (GPU) acceleration for conducting ground-state DFT calculations on large-scale systems. The Tucker tensor DFT algorithm uses a localized Tucker tensor basis computed from an additive separable approximation to the Kohn-Sham Hamiltonian. The discrete Kohn-Sham problem is solved using Chebyshev filtered subspace iteration method that relies on matrix-matrix multiplications of a sparse symmetric Hamiltonian matrix and a dense wavefunction matrix, expressed in the localized Tucker tensor basis. These matrix-matrix multiplication operations, which constitute the most computationally intensive step of the solution procedure, are GPU accelerated providing ~8-fold GPU-CPU speedup for these operations on the largest systems studied. In conclusion, the computational performance of the TTDFT code is presented using benchmark studies on aluminum nano-particles and silicon quantum dots with system sizes ranging up to ~7,000 atoms.

97 MATHEMATICS AND COMPUTING↗

Sparse measurement medical CT reconstruction using multi-fused block matching denoising priors

A major challenge for medical X-ray CT imaging is reducing the number of X-ray projections to lower radiation dosage and reduce scan times without compromising image quality. However these under-determined inverse imaging problems rely on the formulation of an expressive prior model to constrain the solution space while remaining computationally tractable. Traditional analytical reconstruction methods like Filtered Back Projection (FBP) often fail with sparse measurements, producing artifacts due to their reliance on the Shannon-Nyquist Sampling Theorem. Consensus Equilibrium, which is a generalization of Plug and Play, is a recent advancement in Model-Based Iterative Reconstruction (MBIR), has facilitated the use of multiple denoisers are prior models in an optimization free framework to capture complex, non-linear prior information. However, 3D prior modelling in a Plug and Play approach for volumetric image reconstruction requires long processing time due to high computing requirement. Instead of directly using a 3D prior, this work proposes a BM3D Multi Slice Fusion (BM3D-MSF) prior that uses multiple 2D image denoisers fused to act as a fully 3D prior model in Plug and Play reconstruction approach. Our approach does not require training and are thus able to circumvent ethical issues related with patient training data and are readily deployable in varying noise and measurement sparsity levels. In addition, reconstruction with the BM3D-MSF prior achieves similar reconstruction image quality as fully 3D image priors, but with significantly reduced computational complexity. We test our method on clinical CT data and demonstrate that our approach improves reconstructed image quality.

Hossain, Maliha [ORNL]↗

wa-hls4ml: A Benchmark and Surrogate Models for hls4ml Resource and Latency Estimation

As machine learning (ML) is increasingly implemented in hardware to address real-time challenges in scientific applications, the development of advanced toolchains has significantly reduced the time required to iterate on various designs. These advancements have solved major obstacles, but also exposed new challenges. For example, processes that were not previously considered bottlenecks, such as hardware synthesis, are becoming limiting factors in the rapid iteration of designs. To mitigate these emerging constraints, multiple efforts have been undertaken to develop an ML-based surrogate model that estimates resource usage of ML accelerator architectures. We introduce wa-hls4ml, a benchmark for ML accelerator resource and latency estimation, and its corresponding initial dataset of over 680,000 fully connected and convolutional neural networks, all synthesized using hls4ml and targeting Xilinx FPGAs. The benchmark evaluates the performance of resource and latency predictors against several common ML model architectures, primarily originating from scientific domains, as exemplar models, and the average performance across a subset of the dataset. Additionally, we introduce GNN- and transformer-based surrogate models that predict latency and resources for ML accelerators. We present the architecture and performance of the models and find that the models generally predict latency and resources for the 75% percentile within several percent of the synthesized resources on the synthetic test dataset.

Hawks, Benjamin [Fermilab] (ORCID:0000000157000288↗

Multiqubit entanglement generation with squeezed modes

We present a hybrid continuous variable–discrete variable entanglement generation protocol using linear optics and homodyne measurements, capable of producing multiple high-fidelity Bell pairs per protocol iteration, with an approximate 0.5 success probability. The effectiveness of the protocol is determined by the squeezing strength. To increase the number of Bell pairs, approximately 3 dB of extra squeezing is needed for each additional Bell pair. The protocol also generates single Bell pairs with an approximate 0.75 probability for squeezing strengths ⪅ 15 dB , achievable with current technology.

Macridin, Alexandru [Fermilab] (ORCID:000000022228↗

Iterative multi-task learning and inference from seismic images

Seismic interpretation aims to extract quantitative and interpretable attributes from a seismic image produced using some migration method to inform characteristics of a subsurface reservoir or target of interest. Current paradigms for computing seismic attributes mostly rely on single-task algorithms. We develop an iterative, multi-task machine learning method to learn and infer multiple attributes from a seismic image. This method is composed of two stages: a multi-task inference stage and a multi-modal, multi-task refinement stage. The basic mechanism of this method is that we train a multi-task inference neural network (NN) to estimate a set of attributes, including a relative geological time (RGT), a denoised higher-resolution (DHR) seismic image, and multiple fault attributes (including probability, dip, and strike), from a low-resolution, noisy seismic image; then we input the inferred attributes to a multi-task refinement NN to enhance the raw inference results iteratively. The two multi-task NNs are trained separately based on synthetic seismic images and associated attributes generated by a geological modeling algorithm. The software we intend to release is a PyTorch implementation of this multi-task learning method for both 2D and 3D cases along with scripts to run the training/validation. The algorithm and software can be a useful tool for automatic seismic interpretation.

Gao, Kai↗

Macromolecular phasing using diffraction from multiple crystal forms

A phasing algorithm for macromolecular crystallography is proposed that utilizes diffraction data from multiple crystal forms – crystals of the same molecule with different unit-cell packings (different unit-cell parameters or space-group symmetries). The approach is based on the method of iterated projections, starting with no initial phase information. The practicality of the method is demonstrated by simulation using known structures that exist in multiple crystal forms, assuming some information on the molecular envelope and positional relationships between the molecules in the different unit cells. With incorporation of new or existing methods for determination of these parameters, the approach has potential as a method for ab initio phasing.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Advanced Simulation of ITER Core X-ray Crystal Spectroscopy

X-Ray Simulation Analysis (XRSA) is an analytical ray-tracing mixed code developed specifically for the ITER Core X-Ray Crystal Spectroscopy (XRCS-Core) diagnostic, which employs a dual-reflection configuration incorporating multiple pre-reflectors made of Highly Oriented Pyrolytic Graphite (HOPG) and spherically curved analyzing crystals. The ITER XRCS-Core is designed for high spectral resolution measurement in specific wavelength ranges, including narrow bands around 1.354 Å for W 64+ , 2.19 Å for Xe 51+ , and 2.555 Å for Xe 44+ and Xe 47+ , enabling diagnostic capability across a broad electron temperature range in the ITER plasma. XRSA facilitates efficient simulation of the spectral performance of this complex X-ray spectroscopic system. Recent updates to the XRSA code have incorporated two critical effects: auto-focusing, which specifically applies to HOPG, and polarization. These two effects are particularly important in the dual-reflection configuration used in the ITER XRCS-Core system to provide more accurate modeling results. Here, simulations conducted with the updated code demonstrate that polarization has a substantial impact on the performance of the dual-reflection system. Additionally, the combined influence of polarization and system layout introduces performance variations across channels through the same crystal.

Computer graphics↗

Persistent Sampling: Enhancing the Efficiency of Sequential Monte Carlo

Sequential Monte Carlo (SMC) samplers are powerful tools for Bayesian inference but suffer from high computational costs due to their reliance on large particle ensembles for accurate estimates. We introduce persistent sampling (PS), an extension of SMC that systematically retains and reuses particles from all prior iterations to construct a growing, weighted ensemble. By leveraging multiple importance sampling and resampling from a mixture of historical distributions, PS mitigates the need for excessively large particle counts, directly addressing key limitations of SMC such as particle impoverishment and mode collapse. Crucially, PS achieves this without additional likelihood evaluations-weights for persistent particles are computed using cached likelihood values. This framework not only yields more accurate posterior approximations but also produces marginal likelihood estimates with significantly lower variance, enhancing reliability in model comparison. Furthermore, the persistent ensemble enables efficient adaptation of transition kernels by leveraging a larger, decorrelated particle pool. Experiments on high-dimensional Gaussian mixtures, hierarchical models, and non-convex targets demonstrate that PS consistently outperforms standard SMC and related variants, including recycled and waste-free SMC, achieving substantial reductions in mean squared error for posterior expectations and evidence estimates, all at reduced computational cost. PS thus establishes itself as a robust, scalable, and efficient alternative for complex Bayesian inference tasks.

Karamanis, Minas↗

DART-PFLOTRAN: An ensemble-based data assimilation system for estimating subsurface flow and transport model parameters

Ensemble-based Data Assimilation (EDA), based on the Monte Carlo approach, has been effectively applied to estimate model parameters through inverse modeling in subsurface flow and transport problems. However, implementation of EDA approach involves a complicated workflow that include setting up and executing ensemble forward model simulations, processing observations and model simulation results for parameter updates, and repeat for sequential or iterative EDA. To facilitate the management of such workflow and lower the barriers for adopting EDA-based parameter estimation in subsurface science, we develop a generic software frame-work linking the Data Assimilation Research Testbed (DART) with a massively parallel subsurface FLOw and TRANsport code PFLOTRAN. The new DART-PFLOTRAN leverages both the core data assimilation engines in DART and the computational power afforded by PFLOTRAN. In addition to the standard smoother and filtering options, DART-PFLOTRAN enables an iterative EDA workflow based on the Ensemble Smoother for Multiple Data Assimilation method (ES-MDA) to improve estimation accuracy for nonlinear forward problems. Here, we verify the implementation of ES-MDA in DART-PFLOTRAN using two synthetic cases designed to estimate static permeability and dynamic exchange fluxes across the riverbed, respectively, from continuous temperature measurements made across a depth profile. One-dimensional hydro-thermal simulations are performed in both cases to relate temperature responses with the parameters of interest. In the case of estimating dynamic parameters, we demonstrate the flexibility of DART-PFLOTRAN in automating sequential ES-MDA workflow, which will significantly reduce the time researchers spend on managing complex workflows in similar applications. Both studies yield accurate estimations of the parameters compared to their synthetic truth, while ES-MDA leads to more accurate estimation when a high level of nonlinearity exist between observed responses and unknown parameters. With a code base in Python and Fortran, DART-PFLOTRAN paves the way for applications in large-scale subsurface inverse modeling by automating the complex workflow of sequential ES-MDA that can be executed on various computing platforms.

97 MATHEMATICS AND COMPUTING↗

Toward accurate measurement of electromagnetic field by retrieving and refining the center position of non-uniform diffraction disks in Lorentz 4D-STEM

Recent advancement in scanning transmission electron microscopy (STEM) allows the use of 4D-STEM, a technique that captures an electron diffraction pattern at each scan point in STEM, to measure electrostatic and magnetic potential and field in materials. However, accurate measurement, separation of the magnetic and electric signals, and removal of artifacts remain challenging, especially in the presence of complex non-uniform diffraction contrast within the disks. In this work, based on dynamic simulations of 4D-STEM patterns built upon superstructures consisting of millions of atoms to account for different sample thickness and edge geometries, we show how the shape and intensity distribution of the central disk are affected by multiple scattering. We propose a robust refinement procedure through iteration of the spin-sensitive peak position of the disk-center in the circular Hough transform filtered images from experimental Lorentz 4D-STEM dataset after minimizing the possible artifacts, such as those due to the change of thickness, dynamic scattering, and scanning process. We verify that caution must be taken as in practice the rigid-disk-shift model used to reconstruct induction maps can easily break down due to disk-protrusion when there exists a nonconstant phase gradient or thickness within the width of the probe. Through quantitative analysis and comparing experiment with calculation the effect of the non-spin-related intensity distribution inside the disk as well as that causes the disk shift due to the intensity-protrusion can be removed, and high-quality magnetic field mapping is possible.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

Discovering nuclear models from symbolic machine learning

Numerous phenomenological nuclear models have been proposed to describe specific observables within different regions of the nuclear chart. However, developing a unified model that describes the complex behavior of all nuclei remains an open challenge. Here, we explore whether symbolic Machine Learning (ML) can rediscover traditional nuclear physics models or identify alternatives with improved simplicity, fidelity, and predictive power. To address this challenge, we developed a Multi-objective Iterated Symbolic Regression approach that handles symbolic regressions over multiple target observables, accounts for experimental uncertainties and is robust against high-dimensional problems. As a proof of principle, we applied this method to describe the nuclear binding energies and charge radii of light and medium mass nuclei. Our approach identified simple analytical relationships based on the number of protons and neutrons, providing interpretable models with precision comparable to state-of-the-art nuclear models. Additionally, we integrated this ML-discovered model with an existing complementary model to estimate the limits of nuclear stability. These results highlight the potential of symbolic ML to develop accurate nuclear models and guide our description of complex many-body problems.

Nuclear structure↗

An MILP-Based Distributed Energy Management for Coordination of Networked Microgrids

An MILP-based distributed energy management for the coordination of networked microgrids is proposed in this paper. Multiple microgrids and the utility grid are coordinated through iteratively adjusted price signals. Based on the price signals received, the microgrid controllers (MCs) and distribution management system (DMS) update their schedules separately. Then, the price signals are updated according to the generation–load mismatch and distributed to MCs and DMS for the next iteration. The iteration continues until the generation–load mismatch is small enough, i.e., the generation and load are balanced under agreed price signals. Through the proposed distributed energy management, various microgrids and the utility grid with different economic, resilient, emission and socio-economic objectives are coordinated with generation–load balance guaranteed and the microgrid customers’ privacy preserved. In particular, a piecewise linearization technique is employed to approximate the augmented Lagrange term in the alternating direction method of multipliers (ADMM) algorithm. Thus, the subproblems are transformed into mixed integer linear programming (MILP) problems and efficiently solved by open-source MILP solvers, which would accelerate the adoption and deployment of microgrids and promote clean energy. The proposed MILP-based distributed energy management is demonstrated through various case studies on a networked microgrids test system with three microgrids.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Documenting numerical experiments in support of the Coupled Model Intercomparison Project Phase 6 (CMIP6)

Abstract. Numerical simulation, and in particular simulation of the earth system, relies on contributions from diverse communities, from those who develop models to those involved in devising, executing, and analysing numerical experiments. Often these people work in different institutions and may be working with significant separation in time (particularly analysts, who may be working on data produced years earlier), and they typically communicate via published information (whether journal papers, technical notes, or websites). The complexity of the models, experiments, and methodologies, along with the diversity (and sometimes inexact nature) of information sources, can easily lead to misinterpretation of what was actually intended or done. In this paper we introduce a taxonomy of terms for more clearly defining numerical experiments, put it in the context of previous work on experimental ontologies, and describe how we have used it to document the experiments of the sixth phase for the Coupled Model Intercomparison Project (CMIP6). We describe how, through iteration with a range of CMIP6 stakeholders, we rationalized multiple sources of information and improved the clarity of experimental definitions. We demonstrate how this process has added value to CMIP6 itself by (a) helping those devising experiments to be clear about their goals and their implementation, (b) making it easier for those executing experiments to know what is intended, (c) exposing interrelationships between experiments, and (d) making it clearer for third parties (data users) to understand the CMIP6 experiments. We conclude with some lessons learnt and how these may be applied to future CMIP phases as well as other modelling campaigns.

Pascoe, Charlotte (ORCID:0000000223616778)↗