Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Hyperparameter optimization”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Optimizers for stabilizing likelihood-free inference

A growing number of applications in particle physics and beyond use neural networks as unbinned likelihood ratio estimators applied to real or simulated data. Precision requirements on the inference tasks demand a high-level of stability from these networks, which are affected by the stochastic nature of training. We show how physics concepts can be used to stabilize network training through a physics-inspired optimizer. In particular, the energy conserving descent (ECD) optimization framework uses classical Hamiltonian dynamics on the space of network parameters to reduce the dependence on the initial conditions while also stabilizing the result near the minimum of the loss function. We develop a version of this optimizer known as , which has few free hyperparameters with limited ranges guided by physical reasoning. We apply to representative likelihood-ratio estimation tasks in particle physics and find on average that it out-performs the widely used Adam optimizer. We expect that ECD will be a useful tool for wide array of data-limited problems, where it is computationally expensive to exhaustively optimize hyperparameters and mitigate fluctuations with ensembling.

Monte Carlo methods↗

Reward Driven Workflows for Unsupervised Explainable Analysis of Phases and Ferroic Variants From Atomically Resolved Imaging Data

Rapid progress in aberration corrected electron microscopy necessitates development of robust methods for the identification of phases, ferroic variants, and other pertinent aspects of materials structure from imaging data. While unsupervised methods for clustering and classification are widely used for these tasks, their performance can be sensitive to hyperparameter selection in the analysis workflow. In this study, the effects of descriptors and hyperparameters are explored on the capability of unsupervised ML methods to distill local structural information, exemplified by the discovery of polarization and lattice distortion in Sm − dopped BiFeO 3 (BFO) thin films. It is demonstrated that a reward-driven approach can be used to optimize these key hyperparameters across the full workflow, where rewards are designed to reflect domain wall continuity and straightness, ensuring that the analysis aligns with the material's physical behavior. This approach allows the discovery of local descriptors that are best aligned with the specific physical behavior, providing insight into the fundamental physics of materials. The reward driven workflow is further extended to disentangle structural factors of variation via an optimized variational autoencoder (VAE). Lastly, the importance of well-defined rewards is explored as a quantifiable measure of the success of the workflow.

Barakati, Kamyar [University of Tennessee, Knoxvil↗

Optimization of Water-Alternating-CO2 Injection Field Operations Using a Machine-Learning-Assisted Workflow

Summary This paper will present a robust workflow to address multiobjective optimization (MOO) of carbon dioxide (CO2)-enhanced oil recovery (EOR)-sequestration projects with a large number of operational control parameters. Farnsworth unit (FWU) field, a mature oil reservoir undergoing CO2 alternating water injection (CO2-WAG) EOR, will be used as a field case to validate the proposed optimization protocol. The expected outcome of this work would be a repository of Pareto-optimal solutions of multiple objective functions, including oil recovery, carbon storage volume, and project economics. FWU’s numerical model is used to demonstrate the proposed optimization workflow. Because using MOO requires computationally intensive procedures, machine-learning-based proxies are introduced to substitute for the high-fidelity model, thus reducing the total computation overhead. The vector machine regression combined with the Gaussian kernel (Gaussian-SVR) is used to construct proxies. An iterative self-adjusting process prepares the training knowledge base to develop robust proxies and minimizes computational time. The proxies’ hyperparameters will be optimally designed using Bayesian optimization to achieve better generalization performance. Trained proxies will be coupled with multiobjective particle swarm Optimization (MOPSO) protocol to construct the Pareto-front solution repository. The outcomes of this workflow will be a repository containing Pareto-optimal solutions of multiple objectives considered in the CO2-WAG project. The proposed optimization workflow will be compared with another established methodology using a multilayer neural network (MLNN) to validate its feasibility in handling MOO with a large number of parameters to control. Optimization parameters used include operational variables that might be used to control the CO2-WAG process, such as the duration of the water/gas injection period, producer bottomhole pressure (BHP) control, and water injection rate of each well included in the numerical model. It is proved that the workflow coupling Gaussian-SVR proxies and the iterative self-adjusting protocol is more computationally efficient. The MOO process is made more rapid by squeezing the size of the required training knowledge base while maintaining the high accuracy of the optimized results. The outcomes of the optimization study show promising results in successfully establishing the solution repository considering multiple objective functions. Results are also verified by validating the Pareto fronts with simulation results using obtained optimized control parameters. The outcome from this work could provide field operators an opportunity to design a CO2-WAG project using as many inputs as possible from the reservoir models. The proposed work introduces a novel concept that couples Gaussian-SVR proxies with a self-adjusting protocol to increase the computational efficiency of the proposed workflow and to guarantee the high accuracy of the obtained optimized results. More importantly, the workflow can optimize a large number of control parameters used in a complex CO2-WAG process, which greatly extends its utility in solving large-scale MOO problems in various projects with similar desired outcomes.

Energy & Fuels↗

Practical Hamiltonian learning with unitary dynamics and Gibbs states

We study the problem of learning the parameters for the Hamiltonian of a quantum many-body system, given limited access to the system. In this work, we build upon recent approaches to Hamiltonian learning via derivative estimation. We propose a protocol that improves the scaling dependence of prior works, particularly with respect to parameters relating to the structure of the Hamiltonian (e.g., its locality k). Furthermore, by deriving exact bounds on the performance of our protocol, we are able to provide a precise numerical prescription for theoretically optimal settings of hyperparameters in our learning protocol, such as the maximum evolution time (when learning with unitary dynamics) or minimum temperature (when learning with Gibbs states). Thanks to these improvements, our protocol has practical scaling for large problems: we demonstrate this with a numerical simulation of our protocol on an 80-qubit system.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

From noise to information: The transfer function formalism for uncertainty quantification in reconstructing the nuclear density

The neutron distribution of neutron-rich nuclei provides critical information on the structure of finite nuclei and neutron stars. Parity violating experiments—such as PREX and CREX—provide a clean and largely model-independent determination of neutron densities. Such experiments, however, are challenging and expensive, which is why sound statistical arguments are required to maximize the information gained. We introduce a new framework, the “transfer function formalism,” aimed at uncertainty quantification, model selection, and experimental design in the context of neutron densities. The transfer functions (TFs) are built analytically by expressing the linear response of the objective function (e.g., χ2) to small perturbations of the data. Using the TF formalism, we are able to analyze the expected overall uncertainty—quantified in terms of bias and variance—of the mean square radius and interior density of 48 Ca and 208 Pb. Using relativistic mean field models as a proxy for the weak-charge density—and assuming that a total of five measurements could be performed on the weak form factor of 48 Ca and 208 Pb—we identify the optimal models and experimental locations that minimize the uncertainty in the extraction of the radius and interior density. We also explore the use of the TF formalism to understand the influence of prior distributions for the model parameters, as well as the optimization of model hyperparameters not constrained by the data. Here,we establish how the choice of experimental locations and the model that is used can have a significant impact on the final uncertainties of the extracted quantities of interest. For challenging experiments such as CREX and PREX, a proper quantification of such uncertainties is critical. We have demonstrated how the TF formalism provides several advantages for this type of analysis.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

A Machine Learning–Based Tire Life Prediction Framework for Increasing Life of Commercial Vehicle Tires

In the commercial freight industry, tire retreading decisions are often conservative due to limited knowledge of a tire’s remaining service life. This practice leads to increased costs and material waste. This paper proposes a machine learning–based approach for estimating tire casing life and retreadability, focusing on usage data rather than wear information. This approach could extend the tire’s lifespan and reduce landfill waste. Data integration from diverse tire casing measurement sources presents challenges, including imbalanced removal data. Our methodology addresses these challenges by using historical inspection, telematics, and finite element modeling (FEM) datasets. We introduce “Tire Casing Energy” as a comprehensive usage input and apply a Variance-Reduction Synthetic Minority Oversampling Technique (VR-SMOTE) for data imbalance rectification. A random forest model is used to estimate the state of the tire casing and the casing removal probability, with Bayesian optimization applied for hyperparameter tuning, enhancing model accuracy. Here, the proposed prediction framework is able to differentiate different truck fleets and tire locations based on their usage parameters. With the aid of this machine learning model, the importance and sensitivity of different tire usage parameters can be obtained, which is beneficial to maximize tire life.

Data balancing↗

Deep Learning Image Segmentation for Atmospheric Rivers

Abstract The identification of atmospheric rivers (ARs) is crucial for weather and climate predictions as they are often associated with severe storm systems and extreme precipitation, which can cause large impacts on society. This study presents a deep learning model, termed ARDetect, for image segmentation of ARs using ERA5 data from 1960 to 2020 with labels obtained from the TempestExtremes tracking algorithm. ARDetect is a convolutional neural network (CNN)-based U-Net model, with its structure having been optimized using automatic hyperparameter tuning. Inputs to ARDetect were selected to be the integrated water vapor transport (IVT) and total column water (TCW) fields, as well as the AR mask from TempestExtremes from the previous time step to the one being considered. ARDetect achieved a mean intersection-over-union (mIoU) rate of 89.04% for ARs, indicating its high accuracy in identifying these weather patterns and a superior performance than most deep learning–based models for AR detection. In addition, ARDetect can be executed faster than the TempestExtremes method (seconds vs minutes) for the same period. This provides a significant benefit for online AR detection, especially for high-resolution global models. An ensemble of 10 models, each trained on the same dataset but having different starting weights, was used to further improve on the performance produced by ARDetect, thus demonstrating the importance of model diversity in improving performance. ARDetect provides an effective and fast deep learning–based model for researchers and weather forecasters to better detect and understand ARs, which have significant impacts on weather-related events such as floods and droughts.

Galea, Daniel↗

Optimizing training trajectories in variational autoencoders via latent Bayesian optimization approach *

Unsupervised and semi-supervised ML methods such as variational autoencoders (VAE) have become widely adopted across multiple areas of physics, chemistry, and materials sciences due to their capability in disentangling representations and ability to find latent manifolds for classification and/or regression of complex experimental data. Like other ML problems, VAEs require hyperparameter tuning, e.g. balancing the Kullback–Leibler and reconstruction terms. However, the training process and resulting manifold topology and connectivity depend not only on hyperparameters, but also their evolution during training. Because of the inefficiency of exhaustive search in a high-dimensional hyperparameter space for the expensive-to-train models, here we have explored a latent Bayesian optimization (zBO) approach for the hyperparameter trajectory optimization for the unsupervised and semi-supervised ML and demonstrated for joint-VAE with rotational invariances. We have demonstrated an application of this method for finding joint discrete and continuous rotationally invariant representations for modified national institute of standards and technology database (MNIST) and experimental data of a plasmonic nanoparticles material system. The performance of the proposed approach has been discussed extensively, where it allows for any high dimensional hyperparameter trajectory optimization of other ML models.

42 ENGINEERING↗

Train Like a (Var)Pro: Efficient Training of Neural Networks with Variable Projection

Deep neural networks (DNNs) have achieved state-of-the-art performance across a variety of traditional machine learning tasks, e.g., speech recognition, image classification, and segmentation. The ability of DNNs to efficiently approximate high-dimensional functions has also motivated their use in scientific applications, e.g., to solve partial differential equations and to generate surrogate models. In this paper, we consider the supervised training of DNNs, which arises in many of the above applications. We focus on the central problem of optimizing the weights of the given DNN such that it accurately approximates the relation between observed input and target data. Devising effective solvers for this optimization problem is notoriously challenging due to the large number of weights, nonconvexity, data sparsity, and nontrivial choice of hyperparameters. To solve the optimization problem more efficiently, we propose the use of variable projection (VarPro), a method originally designed for separable nonlinear least-squares problems. Our main contribution is the Gauss--Newton VarPro method (GNvpro) that extends the reach of the VarPro idea to nonquadratic objective functions, most notably cross-entropy loss functions arising in classification. These extensions make GNvpro applicable to all training problems that involve a DNN whose last layer is an affine mapping, which is common in many state-of-the-art architectures. In our four numerical experiments from surrogate modeling, segmentation, and classification, GNvpro solves the optimization problem more efficiently than commonly used stochastic gradient descent (SGD) schemes. Finally, GNvpro finds solutions that generalize well, and in all but one example better than well-tuned SGD methods, to unseen data points.

97 MATHEMATICS AND COMPUTING↗

Robust Design Under Uncertainty in Quantum Error Mitigation

Error mitigation techniques are crucial to achieving near-term quantum advantage. Classical postprocessing of quantum computation outcomes is a popular approach for error mitigation, which includes methods, such as zero noise extrapolation, virtual distillation, and learning-based error mitigation. However, these techniques have limitations due to the propagation of uncertainty resulting from the finite shot number of a quantum measurement. In this work, we introduce general and unbiased methods for quantifying the uncertainty and error of error-mitigated observables based on the strategic sampling of error mitigation outcomes. We then extend our approach to demonstrate the optimization of performance and robustness of error mitigation under uncertainty. To illustrate our methods, we apply them to zero noise extrapolation and Clifford date regression in the ground state of the XY model simulated using depolarizing and International Business Machines Corporation (IBM) Toronto noise models, respectively. In particular, we optimize the choice of noise levels and the allocation of shots for zero noise extrapolation and the distribution of the training circuits for Clifford data regression. While our methods are readily applicable to any postprocessing-based error mitigation approach, in practice they must not be prohibitively expensive—even though they perform optimizations of the error mitigation hyperparameters requiring sampling of a statistical distribution of error mitigation outcomes. By leveraging surrogate-based optimization, we show that our methods can efficiently perform optimal design for a zero noise extrapolation implementation. We then further demonstrate the transferability of learned zero noise extrapolation hyperparameters to other similar circuits.

97 MATHEMATICS AND COMPUTING↗

Scaling and merging time-resolved pink-beam diffraction with variational inference

Time-resolved x-ray crystallography (TR-X) at synchrotrons and free electron lasers is a promising technique for recording dynamics of molecules at atomic resolution. While experimental methods for TR-X have proliferated and matured, data analysis is often difficult. Extracting small, time-dependent changes in signal is frequently a bottleneck for practitioners. Recent work demonstrated this challenge can be addressed when merging redundant observations by a statistical technique known as variational inference (VI). However, the variational approach to time-resolved data analysis requires identification of successful hyperparameters in order to optimally extract signal. In this case study, we present a successful application of VI to time-resolved changes in an enzyme, DJ-1, upon mixing with a substrate molecule, methylglyoxal. We present a strategy to extract high signal-to-noise changes in electron density from these data. Furthermore, we conduct an ablation study, in which we systematically remove one hyperparameter at a time to demonstrate the impact of each hyperparameter choice on the success of our model. We expect this case study will serve as a practical example for how others may deploy VI in order to analyze their time-resolved diffraction data.

47 OTHER INSTRUMENTATION↗

Deep learning to estimate permeability using geophysical data

Time-lapse electrical resistivity tomography (ERT) is a popular geophysical method to estimate three-dimensional (3D) permeability fields from electrical potential difference measurements. Traditional inversion and data assimilation methods are used to ingest this ERT data into hydrogeophysical models to estimate permeability. Due to ill-posedness and the curse of dimensionality, existing inversion strategies provide poor estimates and low resolution of the 3D permeability field. Recent advances in deep learning provide us with powerful algorithms to overcome this challenge. This paper presents a deep learning (DL) framework to estimate the 3D subsurface permeability from time-lapse ERT data. To test the feasibility of the proposed framework, we train DL-enabled inverse models on simulation data. Each measurement in both synthetic and field data is standardized by removing the mean and scaling the time-series to unit variance. This pre-processing step is necessary to bring simulation data closer to field observations. Subsurface process models based on hydrogeophysics are used to generate this synthetic data. Training performed on limited simulation data resulted in the DL model over-fitting. An advanced data augmentation based on mixup is implemented to generate additional training samples to overcome this issue. This mixup technique creates weakly labeled (low-fidelity) samples from strongly labeled (high-fidelity) data. The weakly labeled training data is then used to develop DL-enabled inverse models and reduce over-fitting. As both time-lapse ERT (1133048 features/realization) and 3D permeability (585453 features/realization) data samples are from a high-dimensional space, principal component analysis (PCA) is employed to reduce dimensionality. Encoded ERT and encoded permeability are generated using the trained PCA estimators. A deep neural network is then trained to map the encoded ERT to encoded permeability. This mixup training and unsupervised learning allowed us to build a fast and reasonably accurate DL-based inverse model under limited simulation data. Results show that proposed weak supervised learning can capture salient spatial features in the 3D permeability field. Quantitatively, the average mean squared error (in terms of the natural log) on the strongly labeled training, validation, and test datasets is less than 0.5. The R 2 -score (global metric) is greater than 0.75, and the percent error in each cell (local metric) is less than 10%. Finally, an added benefit in terms of computational cost is that the proposed DL-based inverse model is at least O(10 4 ) times faster than running a forward model once it is trained. Data generation, DL model training, and hyperparameter tuning to identify optimal neural network architectures utilized high-performance computing resources while the DL inference is performed on a standard laptop. Approximately, O(10 5 ) processor hours are used for generating data and DL tuning and training. We acknowledge that the data generation and DL model development are expensive. But once a DL model is trained, it can be re-used for inversion rapidly for the given system, with set physics and domain. Note that traditional inversion may require multiple forward model simulations (e.g., in the order of 10 to 1000), which are very expensive. This computational savings ≈ O(10 5 ) – O(10 7 )) makes the proposed DL-based inverse model attractive for subsurface imaging and real-time ERT monitoring applications due to fast and yet reasonably accurate estimations of permeability field.

58 GEOSCIENCES↗

The expressivity of classical and quantum neural networks on entanglement entropy

Abstract Analytically continuing the von Neumann entropy from Rényi entropies is a challenging task in quantum field theory. While then-th Rényi entropy can be computed using the replica method in the path integral representation of quantum field theory, the analytic continuation can only be achieved for some simple systems on a case-by-case basis. In this work, we propose a general framework to tackle this problem using classical and quantum neural networks with supervised learning. We begin by studying several examples with known von Neumann entropy, where the input data is generated by representing$${\text {Tr}}\rho _A^n$$ Tr ρ A n with a generating function. We adopt KerasTuner to determine the optimal network architecture and hyperparameters with limited data. In addition, we frame a similar problem in terms of quantum machine learning models, where the expressivity of the quantum models for the entanglement entropy as a partial Fourier series is established. Our proposed methods can accurately predict the von Neumann and Rényi entropies numerically, highlighting the potential of deep learning techniques for solving problems in quantum information theory.

Physics↗

Surrogate Model Based Optimization for Finding Robust Deep Learning Model Architectures

Deep Learning (DL) models are increasingly used throughout the sciences. However, their performance and usefulness depend greatly on their architecture which is defined by hyperparameters such as the number of nodes, layers, the learning rate, etc. Tuning these hyperparameters is time-consuming because evaluating their performance requires a lengthy training step. Stochastic optimizers used in training lead to performance variability and potentially prediction reliability issues. In this talk, we will describe an automated optimization method based on surrogate models and active learning strategies for tuning DL model architectures. We take into account the prediction variability with the goal to identify architectures that make reliable and robust predictions. We demonstrate our developments on an application arising in particle physics.

deep learning↗

Bayesian optimization of the beam injection process into a storage ring

We have evaluated the data-efficient Bayesian optimization method for the specific task of injection tuning in a circular accelerator. In this paper, we describe the implementation of this method at the Karlsruhe Research Accelerator with up to nine tuning parameters, including the determination of the associated hyperparameters. We show that the Bayesian optimization method outperforms manual tuning and the commonly used Nelder-Mead optimization algorithm in both simulation and experiment. The algorithm was also successfully used to ease the commissioning phase after the installation of new injection magnets and is regularly used during accelerator operations. We demonstrate that the introduction of context variables that include intrabunch scattering effects, such as the Touschek effect, further improves the control and robustness of the injection process.

43 PARTICLE ACCELERATORS↗

Gradient-Informed Design Optimization of Select Nuclear Systems

In this work, we present a gradient-informed design optimization of nuclear reactor core components based on neutronics objectives with both continuous and discrete materials. The main argument in favor of using gradient-informed design optimization is that it scales well with increasing dimensionality of the design space. First, a challenge problem with 121 free parameters is solved with a gradient-informed method and then with a genetic algorithm. Then, a challenge problem to optimize the flux profile of a simplified assembly with eight axial zones is solved. Both challenge problems are solved using directly calculated derivatives from Tools for Sensitivity and Uncertainty Analysis Methodology Implementation (TSUNAMI) in the SCALE package. Furthermore, we demonstrate how a discrete optimization problem—selection of materials for 121 voxels—can be lifted into a continuous problem with mixed materials. In the continuous space, adjoint-based gradients are well-defined, and gradient descent is applicable. Then, a forcing function is introduced that with the selection of an appropriately sized hyperparameter can be used to guide the optimized continuous solution back into a discrete solution. This paper presents an account of the challenges that were faced when applying a gradient-informed optimization algorithm using a Monte Carlo calculation to estimate the gradient information and compares a gradient descent optimization method to a genetic algorithm optimization of the same geometry. Overall, this work demonstrates the potential use of adjoint-based gradient calculations in design optimization of nuclear systems.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Optimization and supervised machine learning methods for fitting numerical physics models without derivatives

Here, we address the calibration of a computationally expensive nuclear physics model for which derivative information with respect to the fit parameters is not readily available. Of particular interest is the performance of optimization-based training algorithms when dozens, rather than millions or more, of training data are available and when the expense of the model places limitations on the number of concurrent model evaluations that can be performed. As a case study, we consider the Fayans energy density functional model, which has characteristics similar to many model fitting and calibration problems in nuclear physics. We analyze hyperparameter tuning considerations and variability associated with stochastic optimization algorithms and illustrate considerations for tuning in different computational settings.

97 MATHEMATICS AND COMPUTING↗