Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “learning rate”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Automated Pneumothorax Diagnosis using Deep Neural Networks

Thoracic ultrasound can provide information leading to rapid diagnosis of pneumothorax with improved accuracy over the standard physical examination and with higher sensitivity than anteroposterior chest radiography. However, the clinical We have Furthermore, remote environments, such as the battlefield or deep-space exploration, may lack expertise for diagnosing developed an automated image interpretation pipeline for the analysis of thoracic ultrasound data and the classification of pneumothorax events to provide decision support in such situations. Our pipeline consists of image preprocessing, data augmentation, and deep learning architectures for medical diagnosis. In this work, we demonstrate that robust, accurate interpretation of chest images and video can be achieved using deep neural networks. A number of novel image processing techniques were employed to achieve this result. Affine transformations were applied for data augmentation. Hyperparameters were optimized for learning rate, dropout regularization, batch size, and epoch iteration by a sequential model-based Bayesian approach. In addition, we utilized pretrained architecturesinterpretation of a patient medical image is highly operator dependent. certain pathologies., applying transfer learning and fine-tuning techniques to fully connected layers. Our pipeline yielded binary classification validation accuracies of 98.3% for M-mode images and 99.8% with B-mode video frames.

US Army collaboration↗

Surrogate Model Based Optimization for Finding Robust Deep Learning Model Architectures

Deep Learning (DL) models are increasingly used throughout the sciences. However, their performance and usefulness depend greatly on their architecture which is defined by hyperparameters such as the number of nodes, layers, the learning rate, etc. Tuning these hyperparameters is time-consuming because evaluating their performance requires a lengthy training step. Stochastic optimizers used in training lead to performance variability and potentially prediction reliability issues. In this talk, we will describe an automated optimization method based on surrogate models and active learning strategies for tuning DL model architectures. We take into account the prediction variability with the goal to identify architectures that make reliable and robust predictions. We demonstrate our developments on an application arising in particle physics.

deep learning↗

Unraveling the impact of initial choices and in-loop interventions on learning dynamics in autonomous scanning probe microscopy

The current focus in Autonomous Experimentation (AE) is on developing robust workflows to conduct the AE effectively. This entails the need for well-defined approaches to guide the AE process, including strategies for hyperparameter tuning and high-level human interventions within the workflow loop. This paper presents a comprehensive analysis of the influence of initial experimental conditions and in-loop interventions on the learning dynamics of Deep Kernel Learning (DKL) within the realm of AE in scanning probe microscopy. We explore the concept of the “seed effect,” where the initial experiment setup has a substantial impact on the subsequent learning trajectory. Additionally, we introduce an approach of the seed point interventions in AE allowing the operator to influence the exploration process. Using a dataset from Piezoresponse Force Microscopy on PbTiO 3 thin films, we illustrate the impact of the “seed effect” and in-loop seed interventions on the effectiveness of DKL in predicting material properties. The study highlights the importance of initial choices and adaptive interventions in optimizing learning rates and enhancing the efficiency of automated material characterization. This work offers valuable insights into designing more robust and effective AE workflows in microscopy with potential applications across various characterization techniques.

47 OTHER INSTRUMENTATION↗

Active Learning with Irrelevant Examples

An improved active learning method has been devised for training data classifiers. One example of a data classifier is the algorithm used by the United States Postal Service since the 1960s to recognize scans of handwritten digits for processing zip codes. Active learning algorithms enable rapid training with minimal investment of time on the part of human experts to provide training examples consisting of correctly classified (labeled) input data. They function by identifying which examples would be most profitable for a human expert to label. The goal is to maximize classifier accuracy while minimizing the number of examples the expert must label. Although there are several well-established methods for active learning, they may not operate well when irrelevant examples are present in the data set. That is, they may select an item for labeling that the expert simply cannot assign to any of the valid classes. In the context of classifying handwritten digits, the irrelevant items may include stray marks, smudges, and mis-scans. Querying the expert about these items results in wasted time or erroneous labels, if the expert is forced to assign the item to one of the valid classes. In contrast, the new algorithm provides a specific mechanism for avoiding querying the irrelevant items. This algorithm has two components: an active learner (which could be a conventional active learning algorithm) and a relevance classifier. The combination of these components yields a method, denoted Relevance Bias, that enables the active learner to avoid querying irrelevant data so as to increase its learning rate and efficiency when irrelevant items are present. The algorithm collects irrelevant data in a set of rejected examples, then trains the relevance classifier to distinguish between labeled (relevant) training examples and the rejected ones. The active learner combines its ranking of the items with the probability that they are relevant to yield a final decision about which item to present to the expert for labeling. Experiments on several data sets have demonstrated that the Relevance Bias approach significantly decreases the number of irrelevant items queried and also accelerates learning speed.

Wagstaff, Kiri↗

Improving Deep Neural Networks’ Training for Image Classification With Nonlinear Conjugate Gradient-Style Adaptive Momentum

Momentum is crucial in stochastic gradient-based optimization algorithms for accelerating or improving training deep neural networks (DNNs). In deep learning practice, the momentum is usually weighted by a well-calibrated constant. However, tuning the hyperparameter for momentum can be a significant computational burden. In this article, we propose a novel adaptive momentum for improving DNNs training; this adaptive momentum, with no momentum-related hyperparame- ter required, is motivated by the nonlinear conjugate gradient (NCG) method. Stochastic gradient descent (SGD) with this new adaptive momentum eliminates the need for the momentum hyperparameter calibration, allows using a significantly larger learning rate, accelerates DNN training, and improves the final accuracy and robustness of the trained DNNs. For example, SGD with this adaptive momentum reduces classification errors for training ResNet110 for CIFAR10 and CIFAR100 from 5.25% to 4.64% and 23.75% to 20.03%, respectively. Furthermore, SGD, with the new adaptive momentum, also benefits adversarial training and, hence, improves the adversarial robustness of the trained DNNs.

97 MATHEMATICS AND COMPUTING↗

Modeling injection-induced fault slip using long short-term memory networks

Stress changes due to changes in fluid pressure and temperature in a faulted formation may lead to the opening/shearing of the fault. This can be due to subsurface (geo)engineering activities such as fluid injections and geologic disposal of nuclear waste. Such activities are expected to rise in the future making it necessary to assess their short- and long-term safety. Here, a new machine learning (ML) approach to model pore pressure and fault displacements in response to high-pressure fluid injection cycles is developed. The focus is on fault behavior near the injection borehole. To capture the temporal dependencies in the data, long short-term memory (LSTM) networks are utilized. To prevent error accumulation within the forecast window, four critical measures to train a robust LSTM model for predicting fault response are highlighted: (i) setting an appropriate value of LSTM lag, (ii) calibrating the LSTM cell dimension, (iii) learning rate reduction during weight optimization, and (iv) not adopting an independent injection cycle as a validation set. Several numerical experiments were conducted, which demonstrated that the ML model can capture peaks in pressure and associated fault displacement that accompany an increase in fluid injection. The model also captured the decay in pressure and displacement during the injection shut-in period. Further, the ability of an ML model to highlight key changes in fault hydromechanical activation processes was investigated, which shows that ML can be used to monitor risk of fault activation and leakage during high pressure fluid injections.

58 GEOSCIENCES↗

CyBERT: Cybersecurity Claim Classification by Fine-Tuning the BERT Language Model

We introduce CyBERT, a cybersecurity feature claims classifier based on bidirectional encoder representations from transformers and a key component in our semi-automated cybersecurity vetting for industrial control systems (ICS). To train CyBERT, we created a corpus of labeled sequences from ICS device documentation collected across a wide range of vendors and devices. This corpus provides the foundation for fine-tuning BERT’s language model, including a prediction-guided relabeling process. We propose an approach to obtain optimal hyperparameters, including the learning rate, the number of dense layers, and their configuration, to increase the accuracy of our classifier. Fine-tuning all hyperparameters of the resulting model led to an increase in classification accuracy from 76% obtained with BertForSequenceClassification’s original architecture to 94.4% obtained with CyBERT. Furthermore, we evaluated CyBERT for the impact of randomness in the initialization, training, and data-sampling phases. CyBERT demonstrated a standard deviation of ±0.6% during validation across 100 random seed values. Finally, we also compared the performance of CyBERT to other well-established language models including GPT2, ULMFiT, and ELMo, as well as neural network models such as CNN, LSTM, and BiLSTM. The results showed that CyBERT outperforms these models on the validation accuracy and the F1 score, validating CyBERT’s robustness and accuracy as a cybersecurity feature claims classifier.

97 MATHEMATICS AND COMPUTING↗

Assessments of epistemic uncertainty using Gaussian stochastic weight averaging for fluid-flow regression

Here, we use Gaussian stochastic weight averaging (SWAG) to assess the epistemic uncertainty associated with neural-network-based function approximation relevant to fluid flows. SWAG approximates a posterior Gaussian distribution of each weight, given training data, and a constant learning rate. Having access to this distribution, it is able to create multiple models with various combinations of sampled weights, which can be used to obtain ensemble predictions. The average of such an ensemble can be regarded as the 'mean estimation', whereas its standard deviation can be used to construct 'confidence intervals', which enable us to perform uncertainty quantification (UQ) with regard to the training process of neural networks. We utilize representative neural-network-based function approximation tasks for the following cases: (i) a two-dimensional circular-cylinder wake; (ii) the DayMET dataset (maximum daily temperature in North America); (iii) a three-dimensional square-cylinder wake; and (iv) urban flow, to assess the generalizability of the present idea for a wide range of complex datasets. SWAG-based UQ can be applied regardless of the network architecture, and therefore, we demonstrate the applicability of the method for two types of neural networks: (i) global field reconstruction from sparse sensors by combining convolutional neural network (CNN) and multi-layer perceptron (MLP); and (ii) far-field state estimation from sectional data with two-dimensional CNN. We find that SWAG can obtain physically-interpretable confidence-interval estimates from the perspective of epistemic uncertainty. This capability supports its use for a wide range of problems in science and engineering.

97 MATHEMATICS AND COMPUTING↗

AI‐Driven Robot Enables Synthesis‐Property Relation Prediction for Metal Halide Perovskites in Humid Atmosphere

Materials Acceleration Platforms (MAPs) – also known as self-driving laboratories– present a new paradigm for materials science and promise an order of magnitude accelerated materials discovery compared to the traditional trial-and-error approach. Metal halide perovskites (MHPs) are an emerging class of materials for optoelectronic applications but are plagued by irreproducible optoelectronic quality, particularly for films fabricated in a humid atmosphere. Here, in this work, a machine learning (ML)-guided closed-loop platform is developed with a multimodal data fusion approach to predict synthesis–property relations for the optical quality of MHP thin films in relative humidities (RHs) ranging from 5–55%. The efficiency of this approach is confirmed by the fast-dropping learning rate to 2% after experimentally sampling less than 1% of the possible 5,000+ combinations. The prediction of synthesis–property relations is done by optical and imaging characterizations. In situ photoluminescence characterization revealed the origin of thin film quality variation at different RH. These insights provide an avenue for controlling the MHP crystallization by fine-tuning the synthesis parameters and RH for a given chemistry, thus lifting the need for stringent atmosphere control. The MAP enables an accelerated screening and understanding of the synthesis design space, facilitating rational synthesis recipe choice for a wide range of materials.

AI-driven robot↗

A phase transition for finding needles in nonlinear haystacks with LASSO artificial neural networks

To fit sparse linear associations, a LASSO sparsity inducing penalty with a single hyperparameter provably allows to recover the important features (needles) with high probability in certain regimes even if the sample size is smaller than the dimension of the input vector (haystack). More recently learners known as artificial neural networks (ANN) have shown great successes in many machine learning tasks, in particular fitting nonlinear associations. Small learning rate, stochastic gradient descent algorithm and large training set help to cope with the explosion in the number of parameters present in deep neural networks. Yet few ANN learners have been developed and studied to find needles in nonlinear haystacks. Driven by a single hyperparameter, our ANN learner, like for sparse linear associations, exhibits a phase transition in the probability of retrieving the needles, which we do not observe with other ANN learners. To select our penalty parameter, we generalize the universal threshold of Donoho and Johnstone (Biometrika 81(3):425–455, 1994) which is a better rule than the conservative (too many false detections) and expensive cross-validation. In the spirit of simulated annealing, we propose a warm-start sparsity inducing algorithm to solve the high-dimensional, non-convex and non-differentiable optimization problem. We perform simulated and real data Monte Carlo experiments to quantify the effectiveness of our approach.

97 MATHEMATICS AND COMPUTING↗

A Kaczmarz-inspired approach to accelerate the optimization of neural network wavefunctions

Neural network wavefunctions optimized using the variational Monte Carlo method have been shown to produce highly accurate results for the electronic structure of atoms and small molecules, but the high cost of optimizing such wavefunctions prevents their application to larger systems. We propose the Subsampled Projected-Increment Natural Gradient Descent (SPRING) optimizer to reduce this bottleneck. SPRING combines ideas from the recently introduced minimum-step stochastic reconfiguration optimizer (MinSR) and the classical randomized Kaczmarz method for solving linear least-squares problems. We demonstrate that SPRING outperforms both MinSR and the popular Kronecker-Factored Approximate Curvature method (KFAC) across a number of small atoms and molecules, given that the learning rates of all methods are optimally tuned. For example, on the oxygen atom, SPRING attains chemical accuracy after forty thousand training iterations, whereas both MinSR and KFAC fail to do so even after one hundred thousand iterations.

97 MATHEMATICS AND COMPUTING↗

Insights from Initial Engineering Designs of Point Source Capture at Industrial Facilities

Initial engineering design studies examining the application of state-of-the-art carbon capture technology at industrial plants contain generally overlooked real-world design considerations for near-term deployment of point source capture (PSC). The implementation of PSC across a wide range of industrial applications presents unique challenges associated with fluctuating CO<sub>2</sub> concentrations, flue gas composition, and utility and land availability. In this article, seven recent industrial retrofit PSC projects are reviewed to investigate the impact of site-specific factors on project design and cost. Across the seven projects, three capture technology classes and four industrial applications are considered, allowing insight into industry-specific opportunities for PSC technology synergy. Common challenges across projects and proposed design solutions are highlighted to propagate ideas and solutions to close the technology gaps and accelerate learning rates.

FEED Studies↗

Critically assessing sodium-ion technology roadmaps and scenarios for techno-economic competitiveness against lithium-ion batteries

Sodium-ion batteries have garnered notable attention as a potentially low-cost alternative to lithium-ion batteries, which have experienced supply shortages and price volatility for key minerals. Here we assess their techno-economic competitiveness against incumbent lithium-ion batteries using a modelling framework incorporating componential learning curves constrained by minerals prices and engineering design floors. We compare projected sodium-ion and lithium-ion price trends across over 6,000 scenarios while varying Na-ion technology development roadmaps, supply chain scenarios, market penetration and learning rates. Assuming that substantial progress can be made along technology roadmaps via targeted research and development, we identify several sodium-ion pathways that might reach cost-competitiveness with low-cost lithium-ion variants in the 2030s. In addition, we show that timelines are highly sensitive to movements in critical minerals supply chains—namely that of lithium, graphite and nickel. Our modelled outcomes suggest that being price advantageous against low-cost lithium-ion variants in the near term is challenging and increasing sodium-ion energy densities to decrease materials intensity is among the most impactful ways to improve competitiveness.

25 ENERGY STORAGE↗

Quantifying local and global mass balance errors in physics-informed neural networks

Physics-informed neural networks (PINN) have recently become attractive for solving partial differential equations (PDEs) that describe physics laws. By including PDE-based loss functions, physics laws such as mass balance are enforced softly in PINN. This paper investigates how mass balance constraints are satisfied when PINN is used to solve the resulting PDEs. We investigate PINN’s ability to solve the 1D saturated groundwater flow equations (diffusion equations) for homogeneous and heterogeneous media and evaluate the local and global mass balance errors. We compare the obtained PINN’s solution and associated mass balance errors against a two-point finite volume numerical method and the corresponding analytical solution. We also evaluate the accuracy of PINN in solving the 1D saturated groundwater flow equation with and without incorporating hydraulic heads as training data. We demonstrate that PINN’s local and global mass balance errors are significant compared to the finite volume approach. Tuning the PINN’s hyperparameters, such as the number of collocation points, training data, hidden layers, nodes, epochs, and learning rate, did not improve the solution accuracy or the mass balance errors compared to the finite volume solution. Mass balance errors could considerably challenge the utility of PINN in applications where ensuring compliance with physical and mathematical properties is crucial.

54 ENVIRONMENTAL SCIENCES↗

ZeoNet: 3D convolutional neural networks for predicting adsorption in nanoporous zeolites

Zeolites are one of the most widely used materials in the chemical industry due to their nanometer-sized pores that can adsorb and react upon molecules selectively. With hundreds of known framework topologies and hundreds of thousands of computationally predicted structures, the ability to rapidly predict zeolite performance allows researchers to prioritize their efforts on the most promising structures for a given application. Although the accuracy of forcefield-based atomistic simulations has advanced significantly in the past two decades, these simulations can be computationally expensive, especially for long-chain, complex molecules. Here, we present ZeoNet, a representation learning framework using convolutional neural networks (ConvNets) and 3D volumetric representations for predicting adsorption in zeolites. ZeoNet was trained on the task of predicting Henry's constants for adsorption, k H , of n-octadecane in more than 330 000 known and predicted zeolite materials. Employing a 3D grid based on the distances to solvent-accessible surfaces, a volumetric representation that can be generated efficiently, the best-performing ZeoNet achieved a correlation coefficient r 2 = 0.977 and a mean-squared error MSE = 3.8 in ln k H , which corresponds to an error of 9.3 kJ mol -1 in adsorption free energy. In comparison, a model based on hand-designed geometric features has values of r 2 = 0.783 and MSE = 35.7. ZeoNet is also relatively efficient and can process ≈8 structures per second on an Nvidia RTX 2080TI GPU, orders of magnitude faster than forcefield-based simulations. A systematic analysis was conducted to investigate how the choice of ConvNet architectures, the linear dimension (L) and spatial resolution (Δd) of the distance grids, batch size, optimizer, and learning rate impact the model performance. We found that ConvNets based on the ResNet architecture offer the best tradeoff between expressiveness and efficiency. The performance for all models reaches a plateau at L = 30–45 Å and depends less sensitively on grid resolution, with a small benefit around Δd = 0.30–0.45 Å. Finally, saliency maps were visualized to identify which regions of the materials contributed the most to model predictions. It was found, interestingly, that the predictions are driven primarily by the accessible pore volume rather than the region occupied by the framework atoms.

36 MATERIALS SCIENCE↗

Machine learning predictions of high-Curie-temperature materials

Technologies that function at room temperature often require magnets with a high Curie temperature, $T$ C , and can be improved with better materials. Discovering magnetic materials with a substantial $T$ C is challenging because of the large number of candidates and the cost of fabricating and testing them. Using the two largest known datasets of experimental Curie temperatures, we develop machine-learning models to make rapid $T$ C predictions solely based on the chemical composition of a material. We train a random-forest model and a k -NN one and predict on an initial dataset of over 2500 materials and then validate the model on a new dataset containing over 3000 entries. The accuracy is compared for multiple compounds' representations (“descriptors”) and regression approaches. A random-forest model provides the most accurate predictions and is not improved by dimensionality reduction or by using more complex descriptors based on atomic properties. Further, a random-forest model trained on a combination of both datasets shows that cobalt-rich and iron-rich materials have the highest Curie temperatures for all binary and ternary compounds. An analysis of the model reveals systematic error that causes the model to over-predict low-$T$ C materials and under-predict high-$T$ C materials. For exhaustive searches to find new high-$T$ C materials, analysis of the learning rate suggests either that much more data is needed or that more efficient descriptors are necessary.

36 MATERIALS SCIENCE↗

Grad–Shafranov equilibria via data-free physics informed neural networks

A large number of magnetohydrodynamic (MHD) equilibrium calculations are often required for uncertainty quantification, optimization, and real-time diagnostic information, making MHD equilibrium codes vital to the field of plasma physics. In this paper, we explore a method for solving the Grad–Shafranov equation by using physics-informed neural networks (PINNs). For PINNs, we optimize neural networks by directly minimizing the residual of the partial differential equation as a loss function. We show that PINNs can accurately and effectively solve the Grad–Shafranov equation with several different boundary conditions, making it more flexible than traditional solvers. This method is flexible as it does not require any mesh and basis choice, thereby streamlining the computational process. We also explore the parameter space by varying the size of the model, the learning rate, and boundary conditions to map various tradeoffs such as between reconstruction error and computational speed. Additionally, we introduce a parameterized PINN framework, expanding the input space to include variables such as pressure, aspect ratio, elongation, and triangularity in order to handle a broader range of plasma scenarios within a single network. Parameterized PINNs could be used in future work to solve inverse problems such as shape optimization.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Group-equivariant autoencoder for identifying spontaneously broken symmetries

We introduce the group-equivariant autoencoder (GE autoencoder), a deep neural network (DNN) method that locates phase boundaries by determining which symmetries of the Hamiltonian have spontaneously broken at each temperature. We use group theory to deduce which symmetries of the system remain intact in all phases, and then use this information to constrain the parameters of the GE autoencoder such that the encoder learns an order parameter invariant to these “never-broken” symmetries. This procedure produces a dramatic reduction in the number of free parameters such that the GE-autoencoder size is independent of the system size. We include symmetry regularization terms in the loss function of the GE autoencoder so that the learned order parameter is also equivariant to the remaining symmetries of the system. By examining the group representation by which the learned order parameter transforms, we are then able to extract information about the associated spontaneous symmetry breaking. We test the GE autoencoder on the 2D classical ferromagnetic and antiferromagnetic Ising models, finding that the GE autoencoder (1) accurately determines which symmetries have spontaneously broken at each temperature; (2) estimates the critical temperature in the thermodynamic limit with greater accuracy, robustness, and time efficiency than a symmetry-agnostic baseline autoencoder; and (3) detects the presence of an external symmetry-breaking magnetic field with greater sensitivity than the baseline method. Lastly, we describe various key implementation details, including a quadratic-programming-based method for extracting the critical temperature estimate from trained autoencoders and calculations of the DNN initialization and learning rate settings required for fair model comparisons.

42 ENGINEERING↗