Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Stochastic Gradient Descent”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

A phase transition for finding needles in nonlinear haystacks with LASSO artificial neural networks

To fit sparse linear associations, a LASSO sparsity inducing penalty with a single hyperparameter provably allows to recover the important features (needles) with high probability in certain regimes even if the sample size is smaller than the dimension of the input vector (haystack). More recently learners known as artificial neural networks (ANN) have shown great successes in many machine learning tasks, in particular fitting nonlinear associations. Small learning rate, stochastic gradient descent algorithm and large training set help to cope with the explosion in the number of parameters present in deep neural networks. Yet few ANN learners have been developed and studied to find needles in nonlinear haystacks. Driven by a single hyperparameter, our ANN learner, like for sparse linear associations, exhibits a phase transition in the probability of retrieving the needles, which we do not observe with other ANN learners. To select our penalty parameter, we generalize the universal threshold of Donoho and Johnstone (Biometrika 81(3):425–455, 1994) which is a better rule than the conservative (too many false detections) and expensive cross-validation. In the spirit of simulated annealing, we propose a warm-start sparsity inducing algorithm to solve the high-dimensional, non-convex and non-differentiable optimization problem. We perform simulated and real data Monte Carlo experiments to quantify the effectiveness of our approach.

97 MATHEMATICS AND COMPUTING↗

Mode connectivity in the loss landscape of parameterized quantum circuits

Variational training of parameterized quantum circuits (PQCs) underpins many workflows employed on near-term noisy intermediate scale quantum (NISQ) devices. It is a hybrid quantum-classical approach that minimizes an associated cost function in order to train a parameterized ansatz. In this work we adapt the qualitative loss landscape characterization for neural networks introduced in Goodfellow et al. (2014); Li et al. (2017) and tests for connectivity used in Draxler et al. (2018) to study the loss landscape features in PQC training. We present results for PQCs trained on a simple regression task, using the bilayer circuit ansatz, which consists of alternating layers of parameterized rotation gates and entangling gates. Multiple circuits are trained with 3 different batch gradient optimizers: stochastic gradient descent, the quantum natural gradient, and Adam. We identify large features in the landscape that can lead to faster convergence in training workflows.

97 MATHEMATICS AND COMPUTING↗

Concurrent multi-peak Bragg coherent x-ray diffraction imaging of 3D nanocrystal lattice displacement via global optimization

Abstract In this paper we demonstrated a method to reconstruct vector-valued lattice distortion fields within nanoscale crystals by optimization of a forward model of multi-reflection Bragg coherent diffraction imaging (MR-BCDI) data. The method flexibly accounts for geometric factors that arise when making BCDI measurements, is amenable to efficient inversion with modern optimization toolkits, and allows for globally constraining a single image reconstruction to multiple Bragg peak measurements. This is enabled by a forward model that emulates the multiple Bragg peaks of a MR-BCDI experiment from a single estimate of the 3D crystal sample. We present this forward model, we implement it within the stochastic gradient descent optimization framework, and we demonstrate it with simulated and experimental data of nanocrystals with inhomogeneous internal lattice displacement. We find that utilizing a global optimization approach to MR-BCDI affords a reliable path to convergence of data which is otherwise challenging to reconstruct.

36 MATERIALS SCIENCE↗

Stable diagonal stripes in the t-J model at $\bar{n}$ h = 1/8 doping from fPEPS calculations

We investigate the two-dimensional t–J model at a hole doping of $\bar{n}$ h= 1 / 8 using recently developed high accuracy fermionic projected entangled pair states method. By applying the stochastic gradient descent method combined with the Monte Carlo sampling technique, we obtain the ground state hole energy Ehole = -1.621 for J/t = 0.4. We show that the ground state has stable diagonal stripes instead of vertical stripes with a width of 4 unit cells, and stripe filling ρ l = 0.5. We further show that the long-range superconductivity order is suppressed at this point.

36 MATERIALS SCIENCE↗

Criticality analysis of nuclear binding energy neural networks

Machine learning methods, in particular deep learning methods such as artificial neural networks (ANNs) with many layers, have become widespread and useful tools in nuclear physics. However, these ANNs are typically treated as ‘black boxes’, with their architecture (width, depth, and weight/bias initialization) and the training algorithm and parameters chosen empirically by optimizing learning based on limited exploration. We test a non-empirical approach to understanding and optimizing nuclear physics ANNs by adapting a criticality analysis based on renormalization group flows in terms of the hyperparameters for weight/bias initialization, training rates, and the ratio of depth to width. This treatment utilizes the statistical properties of neural network initialization to find a generating functional for network outputs at any layer, allowing for a path integral formulation of the ANN outputs as a Euclidean statistical field theory. We use a prototypical example to test the applicability of this approach: a simple ANN for nuclear binding energies. We find that with training using a stochastic gradient descent optimizer, the predicted criticality behavior is realized, and optimal performance is found with critical tuning. However, the use of an adaptive learning algorithm leads to somewhat superior results without concern for tuning and thus obscures the analysis. Nevertheless, the criticality analysis offers a way to look within the black box of ANNs, which is a first step towards potential improvements in network performance beyond using adaptive optimizers.

artificial neural network↗

Embedded training of neural-network subgrid-scale turbulence models

We report that the weights of a deep neural-network model are optimized in conjunction with the governing flow equations to provide a model for subgrid-scale stresses in a temporally developing plane turbulent jet at Reynolds number Re 0 = 6000 . The objective function for training is first based on the instantaneous filtered velocity fields from a corresponding direct numerical simulation, and the training is by a stochastic gradient descent method, which uses the adjoint Navier-Stokes equations to provide the end-to-end sensitivities of the model weights to the velocity fields. In-sample and out-of-sample testing on multiple dual-jet configurations show that its required mesh density in each coordinate direction for prediction of mean flow, Reynolds stresses, and spectra is half that needed by the dynamic Smagorinsky model for comparable accuracy. The same neural-network model trained directly to match filtered subgrid-scale stresses, without the constraint of being embedded within the flow equations during the training, fails to provide a qualitatively correct prediction. The coupled formulation is generalized to train based only on mean-flow and Reynolds stresses, which are more readily available in experiments. The mean-flow training provides a robust model, which is important, though a somewhat less accurate prediction for the same coarse meshes, as might be anticipated due to the reduced information available for training in this case. The anticipated advantage of the formulation is that the inclusion of resolved physics in the training increases its capacity to extrapolate. This is assessed for the case of passive scalar transport, for which it outperforms established models due to improved mixing predictions.

42 ENGINEERING↗

AlignOT: An Optimal Transport Based Algorithm for Fast 3D Alignment With Applications to Cryogenic Electron Microscopy Density Maps

Aligning electron density maps from Cryogenic electron microscopy (cryo-EM) is a first key step for studying multiple conformations of a biomolecule. As this step remains costly and challenging, with standard alignment tools being potentially stuck in local minima, we propose here a new procedure, called AlignOT, which relies on the use of computational optimal transport (OT) to align EM maps in 3D space. By embedding a fast estimation of OT maps within a stochastic gradient descent algorithm, our method searches for a rotation that minimizes the Wasserstein distance between two maps, represented as point clouds. Here, we quantify the impact of various parameters on the precision and accuracy of the alignment, and show that AlignOT can outperform the standard local alignment methods, with an increased range of rotation angles leading to proper alignment. We further benchmark AlignOT on various pairs of experimental maps, which account for different types of conformational heterogeneities and geometric properties. As our experiments show good performance, we anticipate that our method can be broadly applied to align 3D EM maps.

3D alignment↗

Improving Deep Neural Networks’ Training for Image Classification With Nonlinear Conjugate Gradient-Style Adaptive Momentum

Momentum is crucial in stochastic gradient-based optimization algorithms for accelerating or improving training deep neural networks (DNNs). In deep learning practice, the momentum is usually weighted by a well-calibrated constant. However, tuning the hyperparameter for momentum can be a significant computational burden. In this article, we propose a novel adaptive momentum for improving DNNs training; this adaptive momentum, with no momentum-related hyperparame- ter required, is motivated by the nonlinear conjugate gradient (NCG) method. Stochastic gradient descent (SGD) with this new adaptive momentum eliminates the need for the momentum hyperparameter calibration, allows using a significantly larger learning rate, accelerates DNN training, and improves the final accuracy and robustness of the trained DNNs. For example, SGD with this adaptive momentum reduces classification errors for training ResNet110 for CIFAR10 and CIFAR100 from 5.25% to 4.64% and 23.75% to 20.03%, respectively. Furthermore, SGD, with the new adaptive momentum, also benefits adversarial training and, hence, improves the adversarial robustness of the trained DNNs.

97 MATHEMATICS AND COMPUTING↗

Machine learning and atomic layer deposition: Predicting saturation times from reactor growth profiles using artificial neural networks

In this work, we explore the application of deep neural networks to the optimization of atomic layer deposition (ALD) processes. In particular, we focus on a one-shot optimization problem, where we try to predict the optimal dose time that leads to saturation everywhere in the reactor based on thickness values measured at different points of an ALD reactor after a single trial growth. In order to tackle this problem, we introduce a dataset designed to train neural networks to predict saturation times based on these inputs for a cross-flow ALD reactor. Here, we then explore the predictive ability of artificial neural networks of different depths and sizes using a separate testing dataset to evaluate their accuracies. The results obtained show that networks trained using stochastic gradient descent methods can accurately predict saturation times without requiring any additional information on the surface kinetics. This provides a viable approach to minimize the number of experiments required to optimize new ALD processes in a known reactor, and it highlights the way machine learning can be leveraged for thin film growth and manufacturing. While the datasets and training procedure depend on the reactor geometry, the trained neural networks provide a general surrogate model connecting thickness values and trial dose times with optimal saturation times that can be reused for different ALD processes within the same reactor.

36 MATERIALS SCIENCE↗

Train Like a (Var)Pro: Efficient Training of Neural Networks with Variable Projection

Deep neural networks (DNNs) have achieved state-of-the-art performance across a variety of traditional machine learning tasks, e.g., speech recognition, image classification, and segmentation. The ability of DNNs to efficiently approximate high-dimensional functions has also motivated their use in scientific applications, e.g., to solve partial differential equations and to generate surrogate models. In this paper, we consider the supervised training of DNNs, which arises in many of the above applications. We focus on the central problem of optimizing the weights of the given DNN such that it accurately approximates the relation between observed input and target data. Devising effective solvers for this optimization problem is notoriously challenging due to the large number of weights, nonconvexity, data sparsity, and nontrivial choice of hyperparameters. To solve the optimization problem more efficiently, we propose the use of variable projection (VarPro), a method originally designed for separable nonlinear least-squares problems. Our main contribution is the Gauss--Newton VarPro method (GNvpro) that extends the reach of the VarPro idea to nonquadratic objective functions, most notably cross-entropy loss functions arising in classification. These extensions make GNvpro applicable to all training problems that involve a DNN whose last layer is an affine mapping, which is common in many state-of-the-art architectures. In our four numerical experiments from surrogate modeling, segmentation, and classification, GNvpro solves the optimization problem more efficiently than commonly used stochastic gradient descent (SGD) schemes. Finally, GNvpro finds solutions that generalize well, and in all but one example better than well-tuned SGD methods, to unseen data points.

97 MATHEMATICS AND COMPUTING↗

Bias-Variance Trade-Off in Physics-Informed Neural Networks with Randomized Smoothing for High-Dimensional PDEs

Physics-Informed Neural Networks (PINNs) have triggered a paradigm shift in scientific computing, leveraging mesh-free properties and robust approximation capabilities. While proving effective for low-dimensional partial differential equations (PDEs), the computational cost of PINNs remains a hurdle in high-dimensional scenarios. This is particularly pronounced when computing high-order and high-dimensional derivatives in the physics-informed loss. Randomized Smoothing PINN (RS-PINN) introduces Gaussian noise for stochastic smoothing of the original neural net model, enabling the use of Monte Carlo methods for derivative approximation, which eliminates the need for costly automatic differentiation. Despite its computational efficiency, especially in the approximation of high-dimensional derivatives, RS-PINN introduces biases in both loss and gradients, negatively impacting convergence, especially when coupled with stochastic gradient descent (SGD) algorithms. We present a comprehensive analysis of biases in RS-PINN, attributing them to the nonlinearity of the Mean Squared Error (MSE) loss as well as the intrinsic nonlinearity of the PDE itself. We propose tailored bias correction techniques, delineating their application based on the order of PDE nonlinearity. The derivation of an unbiased RS-PINN allows for a detailed examination of its advantages and disadvantages compared to the biased version. Specifically, the biased version has a lower variance and runs faster than the unbiased version, but it is less accurate due to the bias. To optimize the bias-variance trade-off, we combine the two approaches in a hybrid method that balances the rapid convergence of the biased version with the high accuracy of the unbiased version. In addition to methodological contributions, we present an enhanced implementation of RS-PINN. Extensive experiments on diverse high-dimensional PDEs, including Fokker-Planck, Hamilton-Jacobi-Bellman (HJB), viscous Burgers’, Allen-Cahn, and Sine-Gordon equations, illustrate the bias-variance trade-off and highlight the effectiveness of the hybrid RS-PINN. Empirical guidelines are provided for selecting biased, unbiased, or hybrid versions, depending on the dimensionality and nonlinearity of the specific PDE problem.

97 MATHEMATICS AND COMPUTING↗

Efficient Training of Deep Neural Operator Networks via Randomized Sampling

Neural operators (NOs) employ deep neural networks to learn the mappings between infinitedimensional function spaces. Deep operator network (DeepONet), a popular NO architecture, has demonstrated success in the real-time prediction of complex dynamics across various scientific and engineering applications. In this work, we introduce a random sampling technique to be adopted during the training of DeepONet, aimed at improving the generalization ability of the model, while significantly reducing the computational time. The proposed approach targets the trunk network of the DeepONet model that outputs the basis functions corresponding to the spatiotemporal locations of the bounded domain on which the physical system is defined. While constructing the loss function, DeepONet training traditionally considers a uniform grid of spatiotemporal points at which all the output functions are evaluated for each iteration. This approach leads to a larger batch size, resulting in poor generalization and increased memory demands, due to the limitations of the stochastic gradient descent (SGD) optimizer. The proposed random sampling over the inputs of the trunk net mitigates these challenges, improving generalization and reducing the memory requirements during training, resulting in significant computational gains. We validate our hypothesis through three benchmark examples, demonstrating substantial reductions in training time while achieving comparable or lower overall test errors relative to the traditional training approach. Our results indicate that incorporating randomization in the trunk network inputs during training enhances the efficiency and robustness of DeepONet, offering a promising avenue for improving the framework’s performance in modeling complex physical systems.

Karumuri, Sharmila [Department of Civil & Systems ↗

Finding MIDDLE Ground: Scalable and Secure Distributed Learning

Edge computing methods allow devices to efficiently train a high-performing, robust, and personalized model for predictive tasks. However, these methods succumb to privacy and scalability concerns such as adversarial data recovery and expensive model communication. Furthermore, edge computing methods unrealistically assume that all devices train an identical model. In practice, edge devices have varying computational and memory constraints which may not allow certain devices to have the space or speed to train a specific model. To overcome these issues, we propose MIDDLE: a model independent distributed learning algorithm which allows heterogeneous edge devices to assist each other’s training while communicating only non-sensitive information. MIDDLE unlocks the ability for edge devices, regardless of computational or memory constraints, to assist each other even with completely different model architectures. Furthermore, MIDDLE does not require model or gradient communication which greatly reduces communication size and time. We prove that MIDDLE attains the optimal convergence rate O(1/sqrt(TM)) of stochastic gradient descent for convex and non-convex smooth optimization (for total iterations T and batch size M). Finally, our experimental results demonstrate that MIDDLE (even in non-IID data settings) attains robust and high-performing models without model or gradient communication.

Bornstein, Marc I.↗

Deep Neural Network Algorithm for CMC Microstructure Characterization and Variability Quantification

Microstructure characterization and variability quantification are crucial for understanding ceramic matrix composites (CMCs) mechanical behavior and deformation mechanisms across length scales. Traditionally, analyses of the micrographs obtained from microscopy are labor-intensive. However, with the vast improvement in computer vision (CV) and deep learning (DL), an automated algorithm can be designed to extract essential microstructure variability from micrographs which can then be used to construct a statistically representative volume element (SRVE). The DL-based algorithm spans the taxonomy of microstructure analyses, including semantic segmentation of microstructure constituents, secondary phases, matrix/fiber interface, and defects, and quantifying the microstructure variability in terms of probability distributions. In this work, C/SiNC and SiC/SiNC CMCs microstructures are semantically segmented through a deep convolutional neural network, followed by variability quantification through the implementation of a fully connected regression layer, hence forming a deep regression network. The deep regression network operates in a feedforward regime, in which the neuron output signal traverses through the network in a unidirectional manner. The weight tensor associated with each layer is updated through a backpropagation stochastic gradient descent approach. The input gray-scale image obtained through in-house scanning electron microscope and confocal microscope micrographs is augmented through affine transformations to increase the training set size, which is then processed through four strided convolutional layers. This compresses the image resolution by half at each layer while increasing the image depth by applying different filters (image encoding). The class activation maps (CAMs) corresponding to the applied filters highlight the key architectural features and assist with the semantic segmentation of the microstructure.

Hamza, Mohamed H.↗

Imaging extended single crystal lattice distortion fields with multi-peak Bragg ptychography

Recent advances in phase-retrieval-based x-ray imaging methods have demonstrated the ability to reconstruct 3D distortion vector fields within a nanocrystal by using coherent diffraction information from multiple crystal Bragg reflections. However, these works do not provide a solution to the challenges encountered in imaging lattice distortions in crystals with significant defect content that result in phase wrapping. Moreover, these methods only apply to isolated crystals smaller than the x-ray illumination, and therefore cannot be used for imaging of distortions in extended crystals. We introduce multi-peak Bragg ptychography which addresses both challenges via an optimization framework that combines stochastic gradient descent and phase unwrapping methods for robust image reconstruction of lattice distortions and defects in extended crystals. Our work uses modern automatic differentiation toolsets so that the method is easy to extend to other settings and easy to implement in high-performance computers. This work is particularly timely given the broad interest in using the increased coherent flux in fourth-generation synchrotrons for innovative material research.

36 MATERIALS SCIENCE↗

Stochastic minibatch approach to the ptychographic iterative engine

The ptychographic iterative engine (PIE) is a widely used algorithm that enables phase retrieval at nanometer-scale resolution over a wide range of imaging experiment configurations. By analyzing diffraction intensities from multiple scanning locations where a probing wavefield interacts with a sample, the algorithm solves a difficult optimization problem with constraints derived from the experimental geometry as well as sample properties. The effectiveness at which this optimization problem is solved is highly dependent on the ordering in which we use the measured diffraction intensities in the algorithm, and random ordering is widely used due to the limited ability to escape from stagnation in poor-quality local solutions. In this study, we introduce an extension to the PIE algorithm that uses ideas popularized in recent machine learning training methods, in this case minibatch stochastic gradient descent. Our results demonstrate that these new techniques significantly improve the convergence properties of the PIE numerical optimization problem.

47 OTHER INSTRUMENTATION↗

GentenMPI: Distributed Memory Sparse Tensor Decomposition

GentenMPl is a toolkit of sparse canonical polyadic (CP) tensor decomposition algorithms that is designed to run effectively on distributed-memory high-performance computers. Its use of distributed-memory parallelism enables it to efficiently decompose tensors that are too large for a single compute node's memory. GentenMPl leverages Sandia's decades-long investment in the Trilinos solver framework for much of its parallel-computation capability. Trilinos contains numerical algorithms and linear algebra classes that have been optimized for parallel simulation of complex physical phenomena. This work applies these tools to the data science problem of sparse tensor decomposition. In this report, we describe the use of Trilinos in GentenMPl, extensions needed for sparse tensor decomposition, and implementations of the CP-ALS (CP via alternating least squares) and GCP-SGD (generalized CP via stochastic gradient descent) sparse tensor decomposition algorithms. We show that GentenMPl can decompose sparse tensors of extreme size, e.g., a 12.6-terabyte tensor on 8192 computer cores. We demonstrate that the Trilinos backbone provides good strong and weak scaling of the tensor decomposition algorithms.

97 MATHEMATICS AND COMPUTING↗

End-to-End Differentiable Modeling and Management of the Environment

Focal Area: (2) Data acquisition and assimilation enabled by machine learning, AI, and advanced methods including experimental/network design/optimization. We emphasize the importance of leveraging optimization techniques from AI/machine learning (ML) to solve challenging problems in Earth system modeling. Science Challenge: Automatic differentiation has had a transformative effect on ML by allowing the calculation of gradients of arbitrary functions in an incredibly large class of models. We can potentially realize similar improvements in parameter estimation and control for Earth system models (ESMs) by reimplementing them in computational frameworks from ML. Practitioners working with large (>10 7 parameters) models in ML and AI can obtain good predictive performance in a range of spatiotemporal tasks by making use of optimization via stochastic gradient descent and incorporating prior knowledge at multiple levels. We propose writing ESMs in open-source computational frameworks such as Torch, Tensorflow, and JAX to greatly expand the scope of environmental forecasting and management challenges, which can be addressed by leveraging automatic differentiation and gradient descent-like algorithms. We do not call for a wholesale replacement of physical models with data-driven surrogates, but rather advocate for interleaving physical and empirical equations in a manner that is most faithful to the extent of our scientific knowledge and observational data. Central to this topic is the merging of differentiable physical simulations with differentiable optimization layers, which are now both beginning to come to the forefront.

54 ENVIRONMENTAL SCIENCES↗