Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Computational optimization”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

Computational simulations and beamline optimizations for an electron beam degrader at CEBAF

An electron beam degrader is under development with the objective of measuring the transverse and longitudinal acceptance of the Continuous Electron Beam Accelerator Facility (CEBAF) at Jefferson Lab. This project is in support of the CE+BAF positron capability. Computational simulations of beam-target interactions and particle tracking were performed integrating the GEANT4 and Elegant toolkits. A solenoid was added to the setup to control the beam's divergence. Parameter optimization of the solenoid field and magnetic quadrupoles gradient was also performed to further reduce particle loss through the rest of the injector beamline.

Lizárraga-Rubio, V.↗

aphBO-2GP-3B: a budgeted asynchronous parallel multi-acquisition functions for constrained Bayesian optimization on high-performing computing architecture

High-fidelity complex engineering simulations are often predictive, but also computationally expensive and often require substantial computational efforts. The mitigation of computational burden is usually enabled through parallelism in high-performance cluster (HPC) architecture. Optimization problems associated with these applications is a challenging problem due to the high computational cost of the high-fidelity simulations. In this paper, an asynchronous parallel constrained Bayesian optimization method is proposed to efficiently solve the computationally expensive simulation-based optimization problems on the HPC platform, with a budgeted computational resource, where the maximum number of simulations is a constant. The advantage of this method are three-fold. Firstly, the efficiency of the Bayesian optimization is improved, where multiple input locations are evaluated parallel in an asynchronous manner to accelerate the optimization convergence with respect to physical runtime. This efficiency feature is further improved so that when each of the inputs is finished, another input is queried without waiting for the whole batch to complete. Second, the proposed method can handle both known and unknown constraints. Third, the proposed method samples several acquisition functions based on their rewards using a modified GP-Hedge scheme. The proposed framework is termed aphBO-2GP-3B, which means asynchronous parallel hedge Bayesian optimization with two Gaussian processes and three batches. The numerical performance of the proposed framework aphBO-2GP-3B is comprehensively benchmarked using 16 numerical examples, compared against other 6 parallel Bayesian optimization variants and 1 parallel Monte Carlo as a baseline, and demonstrated using two real-world high-fidelity expensive industrial applications. The first engineering application is based on finite element analysis (FEA) and the second one is based on computational fluid dynamics (CFD) simulations.

97 MATHEMATICS AND COMPUTING↗

Online Convex Optimization of Programmable Quantum Computers to Simulate Time-Varying Quantum Channels

Simulating quantum channels is a fundamental primitive in quantum computing, since quantum channels define general (trace-preserving) quantum operations. An arbitrary quantum channel cannot be exactly simulated using a finite-dimensional programmable quantum processor, making it important to develop optimal approximate simulation techniques. In this paper, we study the challenging setting in which the channel to be simulated varies adversarially with time. We propose the use of matrix exponentiated gradient descent (MEGD), an online convex optimization method, and analytically show that it achieves a sublinear regret in time. Through experiments, we validate the main results for time-varying dephasing channels using a programmable generalized teleportation processor.

97 MATHEMATICS AND COMPUTING↗

Iterative methods in GPU-resident linear solvers for nonlinear constrained optimization

Linear solvers are major computational bottlenecks in a wide range of decision support and optimization computations. The challenges become even more pronounced on heterogeneous hardware, where traditional sparse numerical linear algebra methods are often inefficient. For example, methods for solving ill-conditioned linear systems have relied on conditional branching, which degrades performance on hardware accelerators such as graphical processing units (GPUs). To improve the efficiency of solving ill-conditioned systems, our computational strategy separates computations that are efficient on GPUs from those that need to run on traditional central processing units (CPUs). Our strategy maximizes the reuse of expensive CPU computations. Iterative methods, which thus far have not been broadly used for ill-conditioned linear systems, play an important role in our approach. In particular, we extend ideas from Arioli et al., (2007) to implement iterative refinement using inexact LU factors and flexible generalized minimal residual (FGMRES), with the aim of efficient performance on GPUs. In conclusion, we focus on solutions that are effective within broader application contexts, and discuss how early performance tests could be improved to be more predictive of the performance in a realistic environment.

97 MATHEMATICS AND COMPUTING↗

A Range and Performance Optimized Version of the Computer-Aided Speckle Interferometry Algorithm for Real-Time Displacement-Strain Field Monitoring

Abstract This work presents an optimized implementation of the Computer-Aided Speckle Interferometry algorithm which enables full-field determination of displacements and strains on commodity Graphics Processing Units at high resolution and frame rates. By combining careful control of the average speckle size in a laser speckle pattern with a simple sampling rate conversion scheme, a compact representation of the optical speckle is achieved. This allows for optimal use of Graphics Processing Unit architecture with robust range extension. The optimal mapping of the Computer-Aided Speckle Interferometry algorithm to Graphics Processing Unit architecture is shown in detail, and a straightforward method for disambiguating large displacements is illustrated. Lastly, this paper demonstrates a two-step subimage-tapering modification to the original algorithm that enables robust range enhancement while maintaining resolution. Results from numerical simulations on synthetic speckle patterns are shown, and runtime performance metrics are provided, with performance ranging up to 60 frames per second in some cases. The method is suitable for interactive experimental mechanics research, process and testing or any application where real-time high-resolution displacement-strain monitoring is needed. A .NET Framework class library enabling the incorporation of the algorithm into 3rd -party applications is available for download.

42 ENGINEERING↗

Communication Lower Bounds and Optimal Algorithms for Symmetric Matrix Computations

In this article, we focus on the communication costs of three symmetric matrix computations: (i) multiplying a matrix with its transpose, known as a symmetric rank-k update (SYRK) (ii) adding the result of the multiplication of a matrix with the transpose of another matrix and the transpose of that result, known as a symmetric rank-2k update (SYR2K) (iii) performing matrix multiplication with a symmetric input matrix (SYMM). All three computations appear in the Level 3 Basic Linear Algebra Subroutines (BLAS) and have wide use in applications involving symmetric matrices. We establish communication lower bounds for these kernels using sequential and distributed-memory parallel computational models, and we show that our bounds are tight by presenting communication-optimal algorithms for each setting. Our lower bound proofs rely on applying a geometric inequality for symmetric computations and analytically solving constrained nonlinear optimization problems. As a result, the symmetric matrix and its corresponding computations are accessed and performed according to a triangular block partitioning scheme in the optimal algorithms.

Al Daas, Hussam [Rutherford Appleton Laboratory, D↗

Performance Analysis of an Optimization Algorithm for Metamaterial Design on the Integrated High-Performance Computing and Quantum Systems

Optimizing metamaterials with complex geometries is a big challenge. Although an active learning algorithm, combining machine learning (ML), quantum computing, and optical simulation, has emerged as an efficient optimization tool, it still faces difficulties in optimizing complex structures that have potentially high performance. In this work, we comprehensively analyze the performance of an optimization algorithm for metamaterial design on the integrated HPC and quantum systems. We demonstrate significant time advantages through message-passing interface (MPI) parallelization on the high-performance computing (HPC) system showing approximately 54% faster ML tasks and 67 times faster optical simulation against serial workloads. Furthermore, we analyze the performance of a quantum algorithm designed for optimization, which runs with various quantum simulators on a local computer or HPC-quantum system. Results showcase ~24 times speedup when executing the optimization algorithm on the HPC-quantum hybrid system. This study paves a way to optimize complex metamaterials using the integrated HPC-quantum system.

Kim, Seongmin↗

Digital autofocusing of a coded-aperture Laue diffraction microscope

To provide optimal depth resolution with a coded-aperture Laue diffraction microscope, an accurate position of the coded-aperture and its scanning geometry need to be known. However, finding the geometry by trial and error is a time-consuming and often challenging process because of the large number of parameters involved. In this paper, we propose an optimization approach to automate the focusing process after data is collected. Here we demonstrate the robustness and efficiency of the proposed approach with experimental data taken at a synchrotron facility.

47 OTHER INSTRUMENTATION↗

A Review of Quantum Computing Technologies in Power System Optimization

As modern power grids increasingly integrate variable renewable generation, distributed energy resources, and energy storage systems, classical optimization techniques are facing unprecedented challenges. This review examines the emerging application of quantum computing to overcome these challenges in power system optimization, including optimal power flow (OPF), unit commitment (UC), economic dispatch (ED), and intelligent switching and topology optimization (IS-TO). Recent research has introduced various quantum methodologies—such as gate-based, annealing-based, variational algorithms, and quantum-inspired algorithms—to address the combinatorial complexity inherent in grid reconfiguration and energy management. The review summaries the quantum algorithms, quantum devices and the power system test cases, highlighting hybrid quantum–classical strategies that leverage the complementary strengths of both paradigms. Some quantum advantages have been observed, including theoretical speedup, accurate simulation results, scalable qubit usage, efficient QUBO mapping. In particular, the review emphasizes the importance of integrating quantum optimization techniques with classical control frameworks, these hybrid approaches demonstrate the potential to improve real-time grid management and operational reliability. A significant portion of the analysis is devoted to the practical limitations of current quantum devices. Present-day quantum hardware, operating in the noisy intermediate-scale quantum (NISQ) era, remains highly sensitive to noise and limited in qubit connectivity, which constrains the scale and accuracy of implemented algorithms. The review delves into specific challenges such as the need for qubit-efficient encoding techniques and error mitigation strategies that are critical for handling real-world grid optimization problems. In addition, the work draws attention to the performance discrepancies between theoretical quantum speedups and experimental validations, underscoring the importance of rigorous benchmark studies using representative power grid test cases. In summary, this review highlights both the promise and limitations of quantum computing for power system optimization. It provides a comprehensive overview of the state-of-the-art technologies, categorizes recent advancements in algorithm design, and discusses practical considerations for implementation, and serves as an informative resource on current research. Future research directions include developing robust hybrid frameworks, advancing qubit-efficient formulations, and scaling up experimental demonstrations to confirm the theoretical advantages of quantum methods in large-scale power system operations.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Improved Guarantees for Optimal Nash Equilibrium Seeking and Bilevel Variational Inequalities

We consider a class of hierarchical variational inequality (VI) problems that subsumes VI-constrained optimization and several other problem classes, including the optimal solution selection problem and the optimal Nash equilibrium (NE) seeking problem. Our main contribution is threefold. (i) We consider bilevel VIs with monotone and Lipschitz continuous mappings and devise a single-timescale iteratively regularized extragradient method, named IR-EG 𝚖,𝚖 . We improve the existing iteration complexity results for addressing both bilevel VI and VI-constrained convex optimization problems. (ii) Under the strong monotonicity of the outer-level mapping, we develop a method named IR-EG 𝚜,𝚖 and derive faster guarantees than those in (i). We also study the iteration complexity of this method under a constant regularization parameter. These results appear to be new for both bilevel VIs and VI-constrained optimization. (iii) To our knowledge, complexity guarantees for computing the optimal NE in nonconvex settings do not exist. Motivated by this lacuna, we consider VI-constrained nonconvex optimization problems and devise an inexactly projected gradient method, named IPR-EG, where the projection onto the unknown set of equilibria is performed using IR-EG 𝚜,𝚖 with a prescribed termination criterion and an adaptive regularization parameter. We obtain new complexity guarantees in terms of a residual map and an infeasibility metric for computing a stationary point. Here, we validate the theoretical findings using preliminary numerical experiments for computing the best and the worst NEs.

bilevel optimization↗

Bubbling Water–Treating DBD Plasma Device Optimization Using Experimental and Computational Methods

A dry air atmospheric pressure volume dielectric barrier discharge is employed to fix nitrogen in water. Producing nitrate for use as nitrogen fertilizer is the primary motivation. A 0D chemistry model is developed and informed by the electrical, and geometric characteristics of the device and the plasma gas temperature. Modeled ozone and nitrate densities are compared to those measured experimentally in the plasma effluent and treated liquid for a range of gas temperatures. Modeled and measured ozone densities are in good agreement; however, the model lacks the liquid chemistry to properly represent the measured nitrate density. A gas temperature-based shift from ozone to NO x producing regimes is observed in both experiment and model, and the reactions responsible are evaluated.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Development of a coupled experimental–computational approach for engineering optimization of spout-fluidized bed particle coating systems

The design of spout-fluidized bed (SFB) coating systems for nuclear particle fuels typically relies on trial-and-error processes, comprising iterative and time-consuming coating deposition experiments and post-deposition characterization. At an engineering scale, this approach to guided SFB system design is inefficient, highlighting the need for streamlined experimental methodologies which can correlate fluidization conditions to downstream coating outcomes. In this study, we combine time-resolved particle image velocimetry (PIV) with CFD–DEM simulations to benchmark hydrodynamic behavior in a 3D spout-fluidized bed. By exploiting easily accessible optical measurements of particle motion at the bed wall and within the spouting region, we obtain quantitative velocity fields that can be directly compared with model predictions of the occluded bed region, without resorting to complex imaging and characterization techniques such as X-ray or magnetic resonance tomography. Experimental benchmarking reveals strong agreement between CFD–DEM and PIV in the spout and annulus regions, while discrepancies near the wall highlight areas for future model development. Here, the proposed integrated experimental–numerical framework will enable a direct connection between measured variables and numerically predicted fluidization performance of dense, surrogate nuclear particle fuel feedstock such that experimental SFB component design can be rapidly evaluated, informing design decisions for nozzle geometry and operating conditions. Future work will extend this framework by correlating quantified fluidization metrics across nozzle geometries and operating conditions with the resulting coating morphology, microstructure, and uniformity. Establishing these correlations will enable predictive links between hydrodynamic performance and coating quality, providing a rational, scalable basis for optimizing SFB design prior to coating deposition.

CFD/DEM↗

Agentic AI vs ML-Based Autotuning: A Comparative Study for Loop Reordering Optimization

High Performance Computing (HPC) applications rely heavily on code optimizations to achieve good performance on modern CPU and GPU architectures. Traditional Machine Learning auto-tuning approaches have demonstrated success in exploring high-dimensional spaces, but they often require expensive compile-run evaluations and lack adaptability for large HPC applications. The recent advances in Large Language Models (LLMs) and Agentic AI systems raise intriguing questions about the potential of these approaches to address specific optimization methodologies. This work aims to answer an essential question for the HPC community: “How Agentic AI Systems Compare to Traditional ML Autotuning Techniques?” To address this question, we present a comparative analysis between a traditional ML-based optimization approach and an Agentic AI system, evaluating their respective capabilities and limitations for loop-level optimization. In addition, we introduced a new Agentic AI system named LoopGen-AI using three different Large Language Models: GPT-4.1, Claude 4.0, and Gemini 2.5. A key finding is that LoopGen-AI achieves competitive per-formance with only a few program runs, the reasoning logs from the agents revealed that their decisions rely heavily on the combination of semantic understanding of the target kernel with dynamic feedback from the environment, highlighting a promising new dimension in performance tuning. In contrast, ML-based autotuners focus on statistical exploration, and require orders of magnitude more runs to reach peak performance. Additionally, our analysis shows that prompt engineering, particularly using Persona + Context Manager patterns, significantly impacts the effectiveness of Agentic AI. Our results indicate that while Agentic AI systems are not yet a complete replacement for ML-based autotuners, it can effectively complement traditional methods.

Rosas, Miguel Romero↗

Parallelizing autotuning for HPC applications: Unveiling the potential of the speculation strategy in Bayesian optimization

In the exascale computing era, tuning High-Performance Computing (HPC) applications has become a significant computational challenge. Although Bayesian optimization (BO) has emerged as a promising tool for HPC performance tuning, the BO workflow is inherently sequential (i.e., one function evaluation at a time) and cannot leverage the huge amount of parallel resources present in modern supercomputers, resulting in a considerable underutilization of their computational capabilities. This paper explores the trade-off between search quality and parallelism in BO, investigating a diverse set of methods. Building upon both previous approaches from the literature and novel methodologies introduced in this work, our study provides a deep analysis to accelerate BO performance tuning. By examining a set of synthetic functions and practical HPC applications, our exploration analyzes the interaction among various BO methods for parallelization, the quantity of parallel resources, the runtime distribution of target HPC applications, and the costs associated with different search orchestration mechanisms that have been overlooked in previous studies. Compared to sequential BO, our novel methodology achieves comparable quality while demonstrating robust scalability in search time as the amount of parallel resources increases; it also outperforms a state-of-the-art tuner, which supports parallelization, achieving up to 3.67x faster search time. We provide high-value insights for practitioners seeking to leverage the power of parallel computing for efficient HPC application tuning. Additionally, to further assist researchers in accelerating the performance tuning of their HPC applications, we provide an extension of an existing open-source tuning framework that incorporates our methods.

Bayesian optimization↗

A Bilevel Approach for Identifying the Worst Contingencies for Nonconvex Alternating Current Power Systems

We address the bilevel optimization problem of identifying the most critical attacks to an alternating current (AC) power flow network. The upper-level binary maximization problem consists of choosing an attack that is treated as a parameter in the lower-level defender minimization problem. Instances of the lower-level global minimization problem by themselves are NP-hard due to the nonconvex AC power flow constraints, and bilevel solution approaches commonly apply a convex relaxation or approximation to allow for tractable bilevel reformulations at the cost of underestimating some power system vulnerabilities. Our main contribution is to provide an alternative branch-and-bound algorithm whose upper bounding mechanism (in a maximization context) is based on a reformulation that avoids relaxation of the AC power flow constraints in the lower-level defender problem. Lower bounding is provided with semidefinite programming (SDP) relaxed solutions to the lower-level problem. We establish finite termination with guarantees of either a globally optimal solution to the original bilevel problem, or a globally optimal solution to the SDP-relaxed bilevel problem which is included in a vetted list of upper-level attack solutions, at least one of which is a globally optimal solution to the bilevel problem. We demonstrate through computational experiments applied to IEEE case instances both the relevance of our contribution, and the effectiveness of our contributed algorithm for identifying power system vulnerabilities without resorting to convex relaxations of the lower-level problem. We conclude with a discussion of future extensions and improvements.

97 MATHEMATICS AND COMPUTING↗