Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Hyperparameter optimization”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Optimization and supervised machine learning methods for fitting numerical physics models without derivatives

Here, we address the calibration of a computationally expensive nuclear physics model for which derivative information with respect to the fit parameters is not readily available. Of particular interest is the performance of optimization-based training algorithms when dozens, rather than millions or more, of training data are available and when the expense of the model places limitations on the number of concurrent model evaluations that can be performed. As a case study, we consider the Fayans energy density functional model, which has characteristics similar to many model fitting and calibration problems in nuclear physics. We analyze hyperparameter tuning considerations and variability associated with stochastic optimization algorithms and illustrate considerations for tuning in different computational settings.

97 MATHEMATICS AND COMPUTING↗

Tuning a variational autoencoder for data accountability problem in the Mars Science Laboratory ground data system

The Mars Curiosity rover is frequently sending back engineering and science data that goes through a pipeline of systems before reaching its final destination at the mission operations center making it prone to volume loss and data corruption. A ground data system analysis (GDSA) team is charged with the monitoring of this flow of information and the detection of anomalies in that data in order to request a re-transmission when necessary. This work presents ∆-MADS, a derivative-free optimization method applied for tuning the architecture and hyperparameters of a variational autoencoder trained to detect the data with missing patches in order to assist the GDSA team in their mission.

Lakhmiri, Dounia↗

Protecting and Defending against Autonomous Control Systems and Digital Twin Cyber Attacks: Response Strategy for Hyperparameter attacks of Digital Twin Machine Learning Models in Nuclear Power Plants (Final)

Navigating through the complex tapestry of technological advancements, "Response Strategy for Hyperparameter attacks of Digital Twin Machine Learning Model in Nuclear Power Plants" stands at the intersection of cybersecurity and nuclear power plant operations, embarking on a journey through the intricacies of securing digital twins against malicious cyber activities. As nuclear power plants progressively integrate digital twin technology and machine learning models to optimize operations and ensure system reliability, they inadvertently expose themselves to a new spectrum of vulnerabilities, notably in the realm of hyperparameter attacks. Hyperparameters, integral in machine learning model tuning and optimal performance of digital twins, have emerged as a target for adversaries aiming to destabilize the predictive capabilities and therefore, the operational accuracy of these digital entities within critical infrastructures like nuclear plants. This paper, therefore, meticulously threads the needle through the development of a robust response strategy, poised to shield these digital reflections against calculated hyperparameter manipulations, ensuring that the digital twin can effectively and securely function as a reliable proxy for its physical counterpart. The ensuing sections delve into the orchestrated maelstrom of multi-rate time-changing intelligent coordinated hyperparameter attacks and the implementation of event-triggered predictive control, laying down a structured, predictive, and responsive framework that safeguards the nexus where the digital and physical realms of nuclear power plants coalesce. The operational integrity of digital twins in nuclear power plants depends critically on the security of machine learning hyperparameters. This study makes two different contributions. First, a decision-based idea known as a multi-rate time changing intelligent coordinated hyperparameter attack is put forth. In this attack, many hyperparameters are repeatedly changed using both random and intelligent optimal techniques by the attacker. These assaults introduce varied rates at different attack steps, compromise various amounts of hyperparameters, and improve stealth and flexibility. Second, a technique is developed for event triggered predictive control to rapidly respond to potential hyperparameter attacks. This control integrates a sliding window framework, retaining a history of previous data points and employing linear regression to predict the next data point from the current dataset. The control gain K is determined using the Lyapunov-Krasovskii method, and subsequently, an action is developed. Finally, the outcome of the simulation demonstrates the viability of the proposed method for defending nuclear power plant digital twins from hyperparameter attacks.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Quantifying uncertainty for deep learning based forecasting and flow-reconstruction using neural architecture search ensembles

Classical problems in computational physics such as data-driven forecasting and signal reconstruction from sparse sensors have recently seen an explosion in deep neural network (DNN) based algorithmic approaches. However, most DNN models do not provide uncertainty estimates, which are crucial for establishing the trustworthiness of these techniques in downstream decision making tasks and scenarios. In recent years, ensemble-based methods have achieved significant success for the uncertainty quantification in DNNs on a number of benchmark problems. However, their performance on real-world applications remains under-explored. In this work, we present an automated approach to DNN discovery and demonstrate how this may also be utilized for ensemble-based uncertainty quantification. Specifically, we propose the use of a scalable neural and hyperparameter architecture search for discovering an ensemble of DNN models for complex dynamical systems. We highlight how the proposed method not only discovers high-performing neural network ensembles for our tasks, but also quantifies uncertainty seamlessly. This is achieved by using genetic algorithms and Bayesian optimization for sampling the search space of neural network architectures and hyperparameters. Subsequently, a model selection approach is used to identify candidate models for an ensemble set construction. Afterwards, a variance decomposition approach is used to estimate the uncertainty of the predictions from the ensemble. We demonstrate the feasibility of this framework for two tasks — forecasting from historical data and flow reconstruction from sparse sensors for the sea-surface temperature. In conclusion, we demonstrate superior performance from the ensemble in contrast with individual high-performing models and other benchmarks.

Deep ensembles↗

Development of lean, efficient, and fast physics-framed deep-learning-based proxy models for subsurface carbon storage

In this work, we present deep-learning-based surrogate models for CCUS developed with four different algorithms and a physics-framed two-phase flow problem involving displacement of water by CO 2 . The deep-learning models were trained using 3D datasets describing the pressure plume, CO 2 saturation plume, and water extraction rate generated by numerical simulation. The hyperparameters defining the architecture of the neural networks were optimized to determine the slimmest network size and training parameters that give the most efficient performance at the least training cost. To develop a robust model that closely mimics the governing physical laws, the discretized form of the two-phase fluid transport equation was used to formulate the supervised deep-learning task. The algorithms investigated in this study predicted the data to above 95% accuracy, with the multi-layer perceptron model demonstrating the best performance by balancing training speed, prediction time, and prediction accuracy with lean network capacity. Furthermore, the surrogate models simultaneously predict reservoir pressure and CO 2 saturation in every grid block, including the surface well extraction rate and bottomhole pressure, at all simulation times for a given static model realization in just a few seconds on a standard desktop computer. A key outcome of this study is that limits can be placed on network design parameters to avoid over designing neural networks, with associated efficiencies in training and prediction times. This is very useful because large volumes of data may be generated in CCUS projects and over-design of neural network architectures imposes penalties that are antithetical to the goal of near-real time forecasting.

58 GEOSCIENCES↗

RandONets: Shallow networks with random projections for learning linear and nonlinear operators

Deep neural networks have been extensively used for the solution of both the forward and the inverse problem for dynamical systems. However, their implementation necessitates optimizing a high-dimensional space of parameters and hyperparameters. This fact, along with the requirement of substantial computational resources, pose a barrier to achieving high numerical accuracy, but also interpretability. Here, to address the above challenges, we present Random Projection-based Operator Networks (RandONets): shallow networks with random projections and tailor-made numerical analysis methods that learn accurately and fast linear and nonlinear operators. Building on previous works, we prove that RandOnets are universal approximators of linear and nonlinear operators. Due to their simplicity, RandONets provide a one-step transformation of the input space, facilitating interpretability. For the evaluation of their performance, we focus on operators of PDEs. We show, that RandONets outperform by several orders of magnitude, both in terms of numerical approximation accuracy and computational cost, the “vanilla” DeepONets. Hence, we believe that our method will trigger further developments in the field of scientific machine learning, for the development of new ‘’light”schemes that will provide high accuracy while reducing dramatically the computational cost. A MATLAB toolbox for RandONets, including demos, is available on GitHub at https://github.com/GianlucaFabiani/RandONets.

Interpretable machine learning↗

Enhancing Gaussian Process Surrogates for Optimization and Posterior Approximation via Random Exploration

This paper proposes novel noise-free Bayesian optimization strategies that rely on a random exploration step to enhance the accuracy of Gaussian process surrogate models. The new algorithms retain the ease of implementation of the classical GP-UCB algorithm, but the additional random exploration step accelerates their convergence, nearly achieving the optimal convergence rate. Furthermore, to facilitate Bayesian inference with intractable likelihoods, we propose to utilize optimization iterates for maximum a posteriori estimation to build a Gaussian process surrogate model for the unnormalized log-posterior density. We provide bounds for the Hellinger distance between the true and the approximate posterior distributions in terms of the number of design points. We demonstrate the effectiveness of our Bayesian optimization algorithms in nonconvex benchmark objective functions, in a machine learning hyperparameter tuning problem, and in a black-box engineering design problem. The effectiveness of our posterior approximation approach is demonstrated in two Bayesian inference problems for parameters of dynamical systems.

Bayesian inference↗

Challenges for unsupervised anomaly detection in particle physics

Anomaly detection relies on designing a score to determine whether a particular event is uncharacteristic of a given background distribution. One way to define a score is to use autoencoders, which rely on the ability to reconstruct certain types of data (background) but not others (signals). In this paper, we study some challenges associated with variational autoencoders, such as the dependence on hyperparameters and the metric used, in the context of anomalous signal (top and W) jets in a QCD background. We find that the hyperparameter choices strongly affect the network performance and that the optimal parameters for one signal are non-optimal for another. In exploring the networks, we uncover a connection between the latent space of a variational autoencoder trained using mean-squared-error and the optimal transport distances within the dataset. We then show that optimal transport distances to representative events in the background dataset can be used directly for anomaly detection, with performance comparable to the autoencoders. Whether using autoencoders or optimal transport distances for anomaly detection, we find that the choices that best represent the background are not necessarily best for signal identification. These challenges with unsupervised anomaly detection bolster the case for additional exploration of semi-supervised or alternative approaches.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Bayes_Opt-SWMM: A Gaussian process-based Bayesian optimization tool for real-time flood modeling with SWMM

Real-time flood model plays a pivotal role in averting urban flood damage, particularly when there is minimal lead time for preparatory measures. However, urban flood modeling in real-time often contends with inherent uncertainties arising from input data uncertainty and parameter ambiguities. Here this study introduces a real-time calibration (RTC) tool called Bayes_Opt-SWMM, specifically tailored for real-time urban flood modeling and uncertainty optimization. This tool leverages the Gaussian process-based Bayesian optimization algorithm and interfaces seamlessly with the Stormwater Management Model (SWMM). It integrates real-time model forcing data and flood monitoring collected through sensors and gauges which are strategically placed within critical locations of urban drainage systems. Our approach hinges on the Surrogate Model based Uncertainty Optimization (SMUO) concept, providing an avenue for enhancing real-time flood modeling. Bayes_Opt-SWMM runs the optimization process using a surrogate model called Gaussian Process emulator with two inference methods: (1) the Gaussian Process (GP) model and (2) Markov Chain Monte Carlo (MCMC) algorithm in GP model (GP_MCMC). Furthermore, three acquisition functions, namely Expected Improvement (EI), Maximum Probability of Improvement (MPI), and Lower Confidence Bound (LCB), facilitate optimal parameter fitting within the surrogate models. The efficiency of GP-based surrogate models in learning SWMM model parameters, leads to an improved uncertainty quantification and accelerated real-time flood modeling in urban areas. Overall, Bayes_Opt-SWMM emerges as a cost-effective and valuable tool for real-time flood modeling and monitoring, with significant potential for managing intelligent storm water systems in urban environments.

54 ENVIRONMENTAL SCIENCES↗

A General Framework for Error-controlled Unstructured Scientific Data Compression

Data compression plays a key role in reducing storage and I/O costs. Traditional lossy methods primarily target data on rectilinear grids and cannot leverage the spatial coherence in unstructured mesh data, leading to suboptimal compression ratios. We present a multi-component, error-bounded compression framework designed to enhance the compression of floating-point unstructured mesh data, which is common in scientific applications. Our approach involves interpolating mesh data onto a rectilinear grid and then separately compressing the grid interpolation and the interpolation residuals. This method is general, independent of mesh types and typologies, and can be seamlessly integrated with existing lossy compressors for improved performance. We evaluated our framework across twelve variables from two synthetic datasets and two real-world simulation datasets. The results indicate that the multi-component framework consistently outperforms state-of-the-art lossy compressors on unstructured data, achieving, on average, a 2.3 − 3.5× improvement in compression ratios, with error bounds ranging from 1 × 10 the −6 to 1×10−2. We further investigate impact of hyperparameters, such as grid spacing and error allocation, to deliver optimal compression ratios in diverse datasets.

Gong, Qian↗

A general Bayesian algorithm for the autonomous alignment of beamlines

Autonomous methods to align beamlines can decrease the amount of time spent on diagnostics, and also uncover better global optima leading to better beam quality. The alignment of these beamlines is a high-dimensional expensive-to-sample optimization problem involving the simultaneous treatment of many optical elements with correlated and nonlinear dynamics. Bayesian optimization is a strategy of efficient global optimization that has proved successful in similar regimes in a wide variety of beamline alignment applications, though it has typically been implemented for particular beamlines and optimization tasks. In this paper, we present a basic formulation of Bayesian inference and Gaussian process models as they relate to multi-objective Bayesian optimization, as well as the practical challenges presented by beamline alignment. We show that the same general implementation of Bayesian optimization with special consideration for beamline alignment can quickly learn the dynamics of particular beamlines in an online fashion through hyperparameter fitting with no prior information. We present the implementation of a concise software framework for beamline alignment and test it on four different optimization problems for experiments on X-ray beamlines at the National Synchrotron Light Source II and the Advanced Light Source, and an electron beam at the Accelerator Test Facility, along with benchmarking on a simulated digital twin. We discuss new applications of the framework, and the potential for a unified approach to beamline alignment at synchrotron facilities.

47 OTHER INSTRUMENTATION↗

Emulator-Based Bayesian Calibration of the CISNET Colorectal Cancer Models

Purpose To calibrate Cancer Intervention and Surveillance Modeling Network (CISNET)'s SimCRC, MISCAN-Colon, and CRC-SPIN simulation models of the natural history colorectal cancer (CRC) with an emulator-based Bayesian algorithm and internally validate the model-predicted outcomes to calibration targets.Methods We used Latin hypercube sampling to sample up to 50,000 parameter sets for each CISNET-CRC model and generated the corresponding outputs. We trained multilayer perceptron artificial neural networks (ANNs) as emulators using the input and output samples for each CISNET-CRC model. We selected ANN structures with corresponding hyperparameters (i.e., number of hidden layers, nodes, activation functions, epochs, and optimizer) that minimize the predicted mean square error on the validation sample. We implemented the ANN emulators in a probabilistic programming language and calibrated the input parameters with Hamiltonian Monte Carlo-based algorithms to obtain the joint posterior distributions of the CISNET-CRC models' parameters. We internally validated each calibrated emulator by comparing the model-predicted posterior outputs against the calibration targets.Results The optimal ANN for SimCRC had 4 hidden layers and 360 hidden nodes, MISCAN-Colon had 4 hidden layers and 114 hidden nodes, and CRC-SPIN had 1 hidden layer and 140 hidden nodes. The total time for training and calibrating the emulators was 7.3, 4.0, and 0.66 h for SimCRC, MISCAN-Colon, and CRC-SPIN, respectively. The mean of the model-predicted outputs fell within the 95% confidence intervals of the calibration targets in 98 of 110 for SimCRC, 65 of 93 for MISCAN, and 31 of 41 targets for CRC-SPIN.Conclusions Using ANN emulators is a practical solution to reduce the computational burden and complexity for Bayesian calibration of individual-level simulation models used for policy analysis, such as the CISNET CRC models. In this work, we present a step-by-step guide to constructing emulators for calibrating 3 realistic CRC individual-level models using a Bayesian approach.

artificial neural networks↗

Automated and efficient local adaptive regression for principal component-based reduced-order modeling of turbulent reacting flows

Principal Component Analysis can be used to reduce the cost of Computational Fluid Dynamics simulations of turbulent reacting flows by reducing the dimensionality of the transported variables through projection of the thermochemical state onto a lower-dimensional manifold. However, because of the nonlinearity of the principal component source terms, nonlinear regression techniques must be utilized for the source terms in terms of the principal components. Unfortunately, widely available and utilized nonlinear regression techniques can have prohibitive computational requirements and/or accuracy that is highly dependent on user experience in ad hoc tuning of model architecture and hyperparameters. Here, in this work, a new nonlinear regression approach is proposed that is both computationally efficient and automated so does not require any user input. The approach is evaluated through a priori prediction of principal component source terms using data from a Direct Numerical Simulation of a turbulent nonpremixed n-heptane/air jet flame. In particular, the proposed framework consists of local regressions whose complexity is adapted according to the local nonlinearity of the data: local linear regression when accurate enough and local Artificial Neural Networks when nonlinear regression is required. The number of local clusters for local regression is determined automatically using the Davies-Bouldin index. In addition, Bayesian optimization is utilized for model training (i.e., to select the best architectures and hyperparameters of the nonlinear regressions in an unsupervised fashion), eliminating ad hoc hand-tuning and/or expensive grid searches. Overall, compared to a single, global neural network, the new local adaptive regression approach is shown to have comparable accuracy but 69% less training time due to the utilization of local linear regression and faster training of local neural networks.

42 ENGINEERING↗

Efficient online quantum circuit learning with no upfront training

Optimization is a promising candidate for studying the utility of variational quantum algorithms (VQAs). However, evaluating cost functions using quantum hardware introduces runtime overheads that limit exploration. Surrogate-based methods can reduce calls to a quantum computer, yet existing approaches require hyperparameter pre-training and have been tested only on small problems. Here, we show that surrogate-based methods can enable successful optimization at scale, without pre-training, by using radial basis function interpolation (RBF) to construct an adaptive, hyperparameter-free surrogate. Using the surrogate as an acquisition function drives hardware queries to the vicinity of the true optima. For 16-qubit random 3-regular Max-Cut instances with the Quantum Approximate Optimization Algorithm (QAOA), our method outperforms state-of-the-art approaches, without considering their upfront training costs. Furthermore, we successfully optimize QAOA circuits for 127-qubit random Ising models on an IBM processor using 10 4 −10 5 measurements. Strong empirical performance demonstrates the promise of automated surrogate-based learning for large-scale VQA applications.

97 MATHEMATICS AND COMPUTING↗

A Novel Spatial-Temporal Variational Quantum Circuit to Enable Deep Learning on NISQ Devices

Quantum computing presents a promising approach for machine learning with its capability for extremely parallel computation in high-dimension through superposition and entanglement. Despite its potential, existing quantum learning algorithms, such as Variational Quantum Circuits (VQCs), face challenges in handling more complex datasets, particularly those that are not linearly separable. What’s more, it encounters the deployability issue, making the learning models suffer a drastic accuracy drop after deploying them to the actual quantum devices. To overcome these limitations, this paper proposes a novel spatial-temporal design, namely “ST-VQC”, to integrate nonlinearity in quantum learning and improve the robustness of the learning model to noise. Specifically, ST-VQC can extract spatial features via a novel block-based encoding quantum sub-circuit coupled with a layer-wise computation quantum sub-circuit to enable temporal-wise deep learning. Additionally, a SWAP-Free physical circuit design is devised to improve robustness. These designs bring a number of hyperparameters. After a systematic analysis of the design space for each design component, an automated optimization framework is proposed to generate the ST-VQC quantum circuit. The proposed ST-VQC has been evaluated on two IBM quantum processors, ibm-cairo with 27 qubits and ibmq-lima with 7 qubits to assess its effectiveness. The results of the evaluation on the standard dataset for binary classification show that ST-VQC can achieve over 30% accuracy improvement compared with existing VQCs on actual quantum computers. Moreover, on a non-linear synthetic dataset, the STVQC outperforms a linear classifier by 27.9%, while the linear classifier using classical computing outperforms the existing VQC by 15.58%.

Li, Jinyang↗

ARENA: Adversary-Resistant Evolving Neural Architectures

Neural networks are becoming the cornerstone for national security prediction tasks. However, designing them requires significant research and trial/error, as they have many hyperparameters, including their computation graph (“architecture”). Neural architecture search (NAS) employs secondary optimizers to search for architectures maximizing objectives like accuracy. Evolutionary algorithms (EAs) are the most used class of optimizer for NAS. However, existing Python libraries for writing EAs limit the complexity of experiments a user can design. In this project, we built ARENA, a Python framework that encodes complex, hyper-realistic EAs. ARENA collects detailed information as it runs and is flexible enough to encode non-EA search algorithms. We tested ARENA on 4 toy optimization problems by encoding 3 search algorithms for each—random search, an EA, and simulated annealing. We also designed an EA that performs NAS on the MNIST dataset. Our experiments suggest the potential for immediate mission impact through solving lab-wide optimization problems.

97 MATHEMATICS AND COMPUTING↗

Rapid Optimization of Total Variation with Applications in Imaging, Additive Manufacturing, and Qualification

Total Variation optimization penalizes the gradient of a control variable or state. While this work focuses on image processing in particular, it has also found applications in inverse problems and topology optimization. In image processing, the goal is to maintain faithfulness to the original image while denoising and/or deblurring. Additionally, bilevel optimization over the spatially varying regularization weights can illuminate interfaces such as damage regions and other anomalies. We will address two fundamental challenges with TV-optimization: (i) the typical slow convergence of existing TV-optimization methods, and (ii) the selection of spatially varying TV parameters to promote interface detection. Additionally, we will apply such techniques to image data collected in additive manufacturing. In said context, stochasticity in build events induces flaws in the manufactured piece, compromising the integrity of said part. There is a critical need for in-situ monitoring to spot anomalies once they form, and in this setting we apply our total variation and hyperparameter solvers. We will develop a customized algorithm based on for extreme-scale TV-optimization that achieves super-linear or quadratic-convergence, a critical property for real-time, image-by-image analysis. A worst-case outcome is a preprocessing step that enhances image quality in-situ, specifically for out-of-focus and noisy images.

36 MATERIALS SCIENCE↗

Learning high-dimensional parametric maps via reduced basis adaptive residual networks

We propose a scalable framework for the learning of high-dimensional parametric maps via adaptively constructed residual network (ResNet) maps between reduced bases of the inputs and outputs. When just few training data are available, it is beneficial to have a compact parametrization in order to ameliorate the ill-posedness of the neural network training problem. By linearly restricting high-dimensional maps to informed reduced bases of the inputs, one can compress high-dimensional maps in a constructive way that can be used to detect appropriate basis ranks, equipped with rigorous error estimates. A scalable neural network learning framework is thus to learn the nonlinear compressed reduced basis mapping. Unlike the reduced basis construction, however, neural network constructions are not guaranteed to reduce errors by adding representation power, making it difficult to achieve good practical performance. Inspired by recent approximation theory that connects ResNets to sequential minimizing flows, we present an adaptive ResNet construction algorithm. This algorithm allows for depth-wise enrichment of the neural network approximation, in a manner that can achieve good practical performance by first training a shallow network and then adapting. We prove universal approximation of the associated neural network class for $L^2_v$ functions on compact sets. Our overall framework allows for constructive means to detect appropriate breadth and depth, and related compact parametrizations of neural networks, significantly reducing the need for architectural hyperparameter tuning. Numerical experiments for parametric PDE problems and a 3D CFD wing design optimization parametric map demonstrate that the proposed methodology can achieve remarkably high accuracy for limited training data, and outperformed other neural network strategies we compared against.

42 ENGINEERING↗