Engineering PapersSearch

SEARCH · Engineering Papers

Results for “Numerical optimization”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Portable, heterogeneous ensemble workflows at scale using libEnsemble

libEnsemble is a Python-based toolkit for running dynamic ensembles, developed as part of the DOE Exascale Computing Project. The toolkit utilizes a unique generator–simulator–allocator paradigm, where generators produce input for simulators, simulators evaluate those inputs, and allocators decide whether and when a simulator or generator should be called. The generator steers the ensemble based on simulation results. Generators may, for example, apply methods for numerical optimization, machine learning, or statistical calibration. libEnsemble communicates between a manager and workers. Flexibility is provided through multiple manager–worker communication substrates each of which has different benefits. These include Python’s multiprocessing, mpi4py, and TCP. Multisite ensembles are supported using Balsam or Globus Compute. We overview the unique characteristics of libEnsemble as well as current and potential interoperability with other packages in the workflow ecosystem. We highlight libEnsemble’s dynamic resource features: libEnsemble can detect system resources, such as available nodes, cores, and GPUs, and assign these in a portable way. These features allow users to specify the number of processors and GPUs required for each simulation; and resources will be automatically assigned on a wide range of systems, including Frontier, Aurora, and Perlmutter. Such ensembles can include multiple simulation types, some using GPUs and others using only CPUs, sharing nodes for maximum efficiency. We also describe the benefits of libEnsemble’s generator–simulator coupling, which easily exposes to the user the ability to cancel, and portably kill, running simulations based on models that are updated with intermediate simulation output. We demonstrate libEnsemble’s capabilities, scalability, and scientific impact via a Gaussian process surrogate training problem for the longitudinal density profile at the exit of a plasma accelerator stage. In conclusion, the study uses gpCAM for the surrogate model and employs either Wake-T or WarpX simulations, highlighting efficient use of resources that can easily extend to exascale.

Dynamic ensembles

Predicting ptychography probe positions using single-shot phase retrieval neural network

Ptychography is a powerful imaging technique that is used in a variety of fields, including materials science, biology, and nanotechnology. However, the accuracy of the reconstructed ptychography image is highly dependent on the accuracy of the recorded probe positions which often contain errors. These errors are typically corrected jointly with phase retrieval through numerical optimization approaches. When the error accumulates along the scan path or when the error magnitude is large, these approaches may not converge with satisfactory result. We propose a fundamentally new approach for ptychography probe position prediction for data with large position errors, where a neural network is used to make single-shot phase retrieval on individual diffraction patterns, yielding the object image at each scan point. The pairwise offsets among these images are then found using a robust image registration method, and the results are combined to yield the complete scan path by constructing and solving a linear equation. We show that our method can achieve good position prediction accuracy for data with large and accumulating errors on the order of 10 2 pixels, a magnitude that often makes optimization-based algorithms fail to converge. For ptychography instruments without sophisticated position control equipment such as interferometers, our method is of significant practical potential.

47 OTHER INSTRUMENTATION

Paradigm for universal quantum information processing with integrated acousto-optic frequency beamsplitters

Frequency-bin encoding offers tremendous potential in quantum photonic information processing, in which a single waveguide can support hundreds of lightpaths in a naturally phase-stable fashion. This stability, however, comes at a cost: arbitrary unitary operations can be realized by cascaded electro-optic phase modulators and pulse shapers, but require nontrivial numerical optimization for design and have thus far been limited to discrete tabletop components. In this article, we propose, formalize, and computationally evaluate a new paradigm for universal frequency-bin quantum information processing using acousto-optic scattering processes between distinct transverse modes. We show that controllable phase matching in intermodal processes enables 2 × 2 frequency beamsplitters and transverse-mode-dependent phase shifters, which together comprise cascadable FRequency-transverse-mODe Operations (FRODOs) that can synthesize any unitary via analytical decomposition procedures. Modeling the performance of both random gates and discrete Fourier transforms, we demonstrate the feasibility of high-fidelity quantum operations with existing integrated photonics technology, highlighting prospects of parallelizable operations achieving 100% bandwidth utilization. Our approach is realizable with CMOS technology, opening the door to scalable on-chip quantum information processing in the frequency domain.

Lukens, Joseph M. [Purdue Univ., West Lafayette, I

Exceptional points in a passive strip waveguide

Abstract Exceptional points (EPs) in non‐Hermitian systems have attracted significant interest due to their unique behaviors, including novel wave propagation and radiation. While EPs have been explored in various photonic systems, their integration into standard photonic platforms can expand their applicability to broader technological domains. In this work, we propose and experimentally demonstrate EPs in an integrated photonic strip waveguide configuration, exhibiting unique deep wave penetration and uniform‐intensity radiation profiles. By introducing the second‐order grating on one side of the waveguide, forward and backward propagating modes are coupled both directly through second‐order coupling and indirectly through first‐order coupling via a radiative intermediate mode. To describe the EP behavior in a strip configuration, we introduce modified coupled‐mode equations that account for both transverse and longitudinal components. These coupled‐mode formulas reveal the formation of EPs in bandgap closure, achieved by numerically optimizing the grating’s duty cycle to manipulate the first‐ and second‐order couplings simultaneously. Experimental observations, consistent with simulations, confirm the EP behavior, with symmetric transmission spectra and constant radiation profiles at the EP wavelength, in contrast to conventional exponential decay observed at detuned wavelengths. These results demonstrate the realization of EPs in a widely applicable strip waveguide configuration, paving the way for advanced EP applications in nonlinear and ultrafast photonics, as well as advanced sensing technologies.

Materials Science

Simulation-Based Inference for Neutrino Interaction Model Parameter Tuning

High-energy physics experiments studying neutrinos rely heavily on simulations of their interactions with atomic nuclei. Limitations in the theoretical understanding of these interactions typically necessitate ad hoc tuning of simulation model parameters to data. Traditional tuning methods for neutrino experiments have largely relied on simple algorithms for numerical optimization. While adequate for the modest goals of initial efforts, the complexity of future neutrino tuning campaigns is expected to increase substantially, and new approaches will be needed to make progress. In this paper, we examine the application of simulation-based inference (SBI) to the neutrino interaction model tuning for the first time. Using a previous tuning study performed by the MicroBooNE experiment as a test case, we find that our SBI algorithm can correctly infer the tuned parameter values when confronted with a mock data set generated according to the MicroBooNE procedure. This initial proof-of-principle illustrates a promising new technique for next-generation simulation tuning campaigns for the neutrino experimental community.

Tame-Narvaez, Karla Maria [Fermilab]

Simulation-based inference for neutrino interaction model parameter tuning

High-energy physics experiments studying neutrinos rely heavily on simulations of their interactions with atomic nuclei. Limitations in the theoretical understanding of these interactions typically necessitate ad hoc tuning of simulation model parameters to data. Traditional tuning methods for neutrino experiments have largely relied on simple algorithms for numerical optimization. While adequate for the modest goals of initial efforts, the complexity of future neutrino tuning campaigns is expected to increase substantially, and new approaches will be needed to make progress. In this paper, we examine the application of simulation-based inference (SBI) to the neutrino interaction model tuning for the first time. Using a previous tuning study performed by the MicroBooNE experiment as a test case, we find that our SBI algorithm can correctly infer the tuned parameter values when confronted with a mock data set generated according to the MicroBooNE procedure. This initial proof-of-principle illustrates a promising new technique for next-generation simulation tuning campaigns for the neutrino experimental community.

Tame-Narvaez, Karla [Fermilab] (ORCID:000000022249

Exact Fock-State Preparation with $n^{1/4}$ Circuit Depth

Efficient, deterministic, and high-fidelity preparation of large Fock states is essential for scaling bosonic quantum technologies and exploring quantum phenomena at large excitation energies. We introduce a deterministic one-parameter (D1p) protocol that maps Fock-state preparation in an infinite-dimensional Hilbert space onto two-dimensional amplitude amplification. Starting from a coherent state with $|α|\simeq\sqrt{n}$, the initial target-state population scales as $n^{-1/2}$, yielding an iteration count and circuit depth of $\mathcal{O}(n^{1/4})$. Phase matching guarantees unit fidelity in the ideal model; remarkably, preparing $|{10^6}\rangle$ requires only 39 iterations. The protocol uses only displacements and number-selective phase operations, requires no numerical optimization, and further extends to state transfer, general superpositions, finite-dimensional systems, and multipartite entangled states. In the large-amplitude regime, its multi-target form prepares $L$-legged cat states with an iteration count determined only by $L$; cats with up to ten legs require only two iterations, independent of the coherent-state amplitude. This framework provides a broadly applicable route to highly excited bosonic states on platforms supporting these elementary controls.

Roy, Tanay [Fermilab] (ORCID:000000019442862X)

A randomized sketching trust-region secant method for low-memory dynamic optimization

The numerical solution of dynamic optimization problems is often limited by the memory required to store the state trajectory, which is used to evaluate the objective function and its derivatives. Recently, [R. Muthukumar et al., SIAM Journal on Optimization 31(2), pp. 1242–1275 (2021)] introduced a trust-region method for dynamic optimization that employs randomized sketching to compress the state trajectory, resulting in inexact derivative computations. By adaptively learning the sketch rank, the trust-region algorithm achieves rigorous convergence guarantees. Here, we extend this approach to use secant Hessian approximations. Due to the randomness introduced by the sketch, the traditional secant update formulae can produce poor Hessian approximations. In particular, the difference of two gradients, computed from two different sketches, may be inconsistent. To overcome this, we employ a sketched approximation of the Hessian application, in lieu of computing the gradient difference. We numerically demonstrate the improved stability of this approach on an example from PDE-constrained optimization.

dynamic optimization

Numerical Modeling & Size Optimization of Thermal Energy Storage for Iron & Steel Production

Iron and steel production are responsible for 90 million MtCO2 per year in the United States. Hydrogen direct reduction of iron (H2DRI) is a promising pathway for a more sustainable iron production than commercially deployed technologies which rely on natural gas. The H2DRI process requires hydrogen at a temperature of up to 950 degrees C fed into a reduction furnace to produce pellets or briquettes that are used in the downstream iron and steelmaking process. In this work, we propose to use an electrical thermal energy storage (ETES) system, that can use renewable electricity to store high-temperature heat and dispatch it upon demand. Such a system can buffer the H2DRI plant from the variability of electricity prices by charging during curtailment and running the plant from storage during times of peak electricity price. We have developed heat transfer models for two different ETES systems that can be used to heat up hydrogen to the required temperatures: a particle-based ETES and a firebrick ETES. These models are used to evaluate the performance of such a system and support the sizing and preliminary cost estimation. The preliminary results using both models show that designing ETES systems for an industrial-scale H2DRI furnace is feasible. The firebrick ETES system has limited operational duration, which might limit the price buffering effect unless significantly oversized. The particle ETES system heat exchanger has industry-feasible dimensions, but its storage capacity would be decided upon the number of particle storage silos.

25 ENERGY STORAGE

Numerical Modeling and Optimization of the iProTech Pitching Inertial Pump (PIP) Wave Energy Converter (WEC) (Cooperative Research and Development Final Report, CRADA Number: CRD-22-22968)

This work generated a first-of-its-kind automated workflow to couple time-domain simulations of wave energy converters written in one software language with a set of design generation and evaluation scripts written in another software language. This automated workflow used an existing optimization package to analyze the sensitivity of different design parameters on the power output of a specific WEC, iProTech’s Pitching Inertial Pump (PIP). Geometric, inertial, and power take-off variables were all varied and optimized to find values that produced the highest amount of power generated over varying wave conditions. The findings on these parameter sensitivity studies are used to inform future design iterations of the PIP WEC. Including more design variables in the optimizations will only increase computational run time and further software development is needed to analyze a larger optimization.

16 TIDAL AND WAVE POWER

Riemannian Optimization Applied to AC Optimal Power Flow: Preprint

The nonlinear, nonconvex AC optimal power flow problem is of growing importance as the nature of the power grid evolves. This problem can be difficult to solve for interior point methods. However, the advent of optimization algorithms over smooth Riemannian manifolds presents an alternative approach. The nonlinear, nonconvex constraints in the AC power flow problem form an embedded submanifold of Euclidean space. In this paper, the authors explore the performance of Riemannian optimization algorithms for the ACOPF problem where the optimization is performed directly on the AC power flow manifold. They demonstrate that these are viable computational alternatives to interior point methods. This is done by using Julia and the packages PowerModels.jl and Manopt.jl.

manifold optimization

A Novel Manufacturing Process of Lightweight Automotive Seats (Integration of Additive Manufacturing and Reinforced Polymer Composite)

Lightweight automotive seats offer multiple benefits to original equipment manufacturers in terms of cost savings from various aspects, including less material usage, more integrated processes, and compliance with Corporate Average Fuel Economy Standards. Original equipment manufacturers have been focusing on innovative ways to produce light weight automotive seats. The commercially available automotive seats are currently made of multiple metal components combined through welding and fasteners. The use of additive manufacturing and composite structures is particularly useful for light weighting the automotive components. Additive manufacturing (AM) offers multiple advantages over traditional manufacturing processes such as freedom of design thereby enabling complex structural geometries, mass customization and waste minimization, and control over the fiber alignment through deposition in a predetermined pattern. Combining metal inserts with polymer composites through a novel manufacturing process allows design of lightweight and high-performance materials for automotive components. However, fabricating these metal polymer composite structures through traditional manufacturing processes limits their mechanical properties due to limited design freedom, lack of control over fiber orientation in composite parts, and poor interfacial bonding between the constituent materials. It is essential to develop a novel manufacturing process to enable high throughput production of lightweight automotive seats using metal and polymer composites. As such it is important to design the automotive seat suitable for manufacturing via this process and perform mechanical characterization on various subcomponents of the seat to ensure that the design and performance requirements provided by the auto manufacturer are met. The aim of this project is to develop a novel manufacturing technique to produce lightweight automotive seat by combining AM with conventional manufacturing processes. The car seat back panel will be designed via topology optimization and numerical simulations to minimize the overall weight while ensuring it meets all the performance requirements. The optimization of the seat back structure will be based on computational stress analysis to maximize the stiffness and minimize the weight. Materials currently used by Ford Motor Company will be adopted for a few subcomponents while the in-house composite materials will be used for the rest of the seat back. The composite and metallic materials will be tested to determine their mechanical properties as these are necessary for simulations. A novel manufacturing process will be developed to integrate AM metal inserts with discontinuous reinforced composite through large scale additive manufacturing and compression overmolding processes. The developed manufacturing technique will be used to fabricated various subcomponents suitable for the seat back design and mechanically tested to determine their properties. The manufacturing of the lightweight seat back design through this process involves integrated AM metal inserts with the composite structure for recliner connection. The manufacturing of the entire seat back which is lightweight through the novel manufacturing process will be discussed. The performance of the designed seat back will be investigated through numerical simulations and shown to meet all the requirements provided by the auto manufacturer. The final goal of developing a novel manufacturing process for lightweight automotive seats is met through design optimization of seat back, manufacturing of subcomponents, mechanical characterization, and validation through numerical simulations. The routes to achieve the final goal of the project and the depth in which they were investigated changed throughout the project due to personnel changes and the COVID-19 pandemic. The project resulted in the development of a novel manufacturing process to integrate metal inserts with tailored polymer composite preforms through overmolding. Leveraging this proven manufacturing process, a lightweight seat back was designed through topology optimization and numerical simulations. The designed seat back uses AM metal inserts and compression overmolding of tailored polymer composite preforms obtained via large scale additive manufacturing. The metal polymer composite structures fabricated through this process exhibited enhancement in stiffness and improved ductility upon testing. Overall, the project provided an alternative design and manufacturing technique for automotive seat back that enables weight saving while meeting the safety and performance requirements.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI

Data-Driven Kinetic Reaction Networks for Separation Chemistry

Understanding complex, multistep chemical reactions at the molecular level is a major challenge whose solution would greatly benefit the design and optimization of numerous chemical processes. The separation of rare-earth (4f) and actinide (5f) elements is an example where improving our chemical understanding is important for designing and optimizing new chemistries, even with a limited number of observations. Here, in this work, we leverage data-driven artificial intelligence and machine-learning approaches to develop kinetic reaction networks that describe the liquid–liquid extraction mechanism of uranium using N,N-di-2-ethylhexyl-isobutyramide (DEHiBA). Specifically, we compare and contrast the properties of two classes of models: (1) purely data-driven models that are regularized using chemistry-agnostic, L1 regression and (2) chemistry-informed models that are regularized using relative reaction energies provided by quantum mechanical calculations. We observe that purely data-driven models are unbiased, simple, and accurate in their predictions of experimental measurements when provided with sufficient data but are difficult to fully constrain and interpret. In contrast, chemistry-informed models exhibit significantly improved chemical interpretability and consistency, providing a detailed description of the separation process while achieving high accuracy through ensemble averaging. Overall, the dominant species predicted to be extracted into the organic phase is UO 2 (NO 3 ) 2 (DEHiBA) 2 , agreeing with experimental slope analysis, thermodynamic modeling, EXAFS, and crystal structures. This work demonstrates that leveraging the fundamental structure of the problem can lead to efficient learning schemes that provide both accurate predictions and chemical insights at a low computational cost.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH

Toward real-time optimization through model reduction and model discrepancy sensitivities

Optimization problems arise in a range of scenarios, from optimal control to model parameter estimation. In many applications, such as the development of digital twins, it is essential to solve these optimization problems within wall-clock-time limitations. However, this is often unattainable for complex systems, such as those modeled by nonlinear partial differential equations. One strategy for mitigating this issue is to construct a reduced-order model (ROM) that enables more rapid optimization. In particular, the use of nonintrusive ROMs—those that do not require access to the full-order model at evaluation time—is popular because they facilitate the computation of optimization solutions within the wall-clock time requirements. However, the optimization solution will be unreliable if the iterates move outside the ROM training data. This article proposes the use of hyper-differential sensitivity analysis with respect to model discrepancy (HDSA-MD) as a computationally efficient tool to augment ROM-constrained optimization and improve its reliability. The proposed approach consists of two phases: (i) an offline phase where several full-order model evaluations are computed to train the ROM, and (ii) an online phase where a ROM-constrained optimization problem is solved, a limited number of full-order model evaluations are computed, and HDSA-MD is used to enhance the optimization solution. Numerical results are demonstrated for two examples, atmospheric contaminant control and wildfire ignition location estimation, in which a ROM is trained offline using inaccurate atmospheric data. In conclusion, the HDSA-MD update yields a significant improvement in the ROM-constrained optimization solution using only one full-order model evaluation online with corrected atmospheric data.

PDE-constrained optimization

Stochastic Waveform Estimation at the Fundamental Quantum Limit

Although measuring the deterministic waveform of a weak classical force is a well-studied problem, estimating a random waveform, such as the spectral density of a stochastic signal field, is much less well understood despite it being a widespread task at the frontier of experimental physics. State-of-the-art precision sensors of random forces must account for the underlying quantum nature of the measurement but the optimal quantum protocol for interrogating such linear sensors is not known. We derive the fundamental precision limit: the extended-channel quantum Cramér-Rao bound. In the experimentally relevant regime in which losses dominate, we prove that non-Gaussian-state preparation and measurement are required to achieve this fundamental limit and we determine numerically the optimal non-Gaussian protocol. We discuss how this scheme could accelerate searches for signatures of quantum gravity, stochastic gravitational waves, and axionic dark matter.

Axions

Machine Learning-Driven Conservative-to-Primitive Conversion in Hybrid Piecewise Polytropic and Tabulated Equations of State

We present a novel machine learning (ML)-based method to accelerate conservative-to-primitive inversion, focusing on hybrid piecewise polytropic and tabulated equations of state. Traditional root-finding techniques are computationally expensive, particularly for large-scale relativistic hydrodynamics simulations. To address this, we employ feedforward neural networks (NNC2PS and NNC2PL), trained in PyTorch (2.0+) and optimized for GPU inference using NVIDIA TensorRT (8.4.1), achieving significant speedups with minimal accuracy loss. The NNC2PS model achieves 𝐿 1 and 𝐿 ∞ errors of 4.54 × 10 −7 and 3.44 × 10−6, respectively, while the NNC2PL model exhibits even lower error values. TensorRT optimization with mixed-precision deployment substantially accelerates performance compared to traditional root-finding methods. Specifically, the mixed-precision TensorRT engine for NNC2PS achieves inference speeds approximately 400 times faster than a traditional single-threaded CPU implementation for a dataset size of 1,000,000 points. Ideal parallelization across an entire compute node in the Delta supercomputer (dual AMD 64-core 2.45 GHz Milan processors and 8 NVIDIA A100 GPUs with 40 GB HBM2 RAM and NVLink) predicts a 25-fold speedup for TensorRT over an optimally parallelized numerical method when processing 8 million data points. Moreover, the ML method exhibits sub-linear scaling with increasing dataset sizes. We release the scientific software developed, enabling further validation and extension of our findings. By exploiting the underlying symmetries within the equation of state, these findings highlight the potential of ML, combined with GPU optimization and model quantization, to accelerate conservative-to-primitive inversion in relativistic hydrodynamics simulations.

conservative-to-primitive conversion

Extended Fayans energy density functional: optimization and analysis

The Fayans energy density functional (EDF) has been very successful in describing global nuclear properties (binding energies, charge radii, and especially differences of radii) within nuclear density functional theory. In a recent study, supervised machine learning methods were used to calibrate the Fayans EDF. Building on this experience, in this work we explore the effect of adding isovector pairing terms, which are responsible for different proton and neutron pairing fields, by comparing a 13D model without the isovector pairing term against the extended 14D model. At the heart of the calibration is a carefully selected heterogeneous dataset of experimental observables representing ground-state properties of spherical even–even nuclei. To quantify the impact of the calibration dataset on model parameters and the importance of the new terms, we carry out advanced sensitivity and correlation analysis on both models. The extension to 14D improves the overall quality of the model by about 30%. The enhanced degrees of freedom of the 14D model reduce correlations between model parameters and enhance sensitivity.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS

A GPU‐Based Ocean Dynamical Core for Routine Mesoscale‐Resolving Climate Simulations

Abstract We describe an ocean hydrostatic dynamical core implemented in Oceananigans optimized for Graphical Processing Unit (GPU) architectures. On 64 A100 GPUs, equivalent to 16 computational nodes in current state‐of‐the‐art supercomputers, our dynamical core can simulate a decade of near‐global ocean dynamics per wall‐clock day at an 8‐km horizontal resolution; a resolution adequate to resolve the ocean's mesoscale eddy field. Such efficiency, achieved with relatively modest hardware resources, suggests that climate simulations on GPUs can incorporate fully eddy‐resolving ocean models. This removes a major source of systematic bias in current IPCC coupled model projections, the parameterization of ocean eddies, and represents a major advance in climate modeling. We discuss the computational strategies, focusing on GPU‐specific optimization and numerical implementation details that enable such high performance.

Silvestri, Simone [Massachusetts Institute of Tech