Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Network parameter error”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Skip-Connected Self-Recurrent Spiking Neural Networks with Joint Intrinsic Parameter and Synaptic Weight Training

As an important class of spiking neural networks (SNNs), recurrent spiking neural networks (RSNNs) possess great computational power and have been widely used for processing sequential data like audio and text. However, most RSNNs suffer from two problems. 1. Due to the lack of architectural guidance, random recurrent connectivity is often adopted, which does not guarantee good performance. 2. Training of RSNNs is in general challenging, bottlenecking achievable model accuracy. To address these problems, we propose a new type of RSNNs called Skip-Connected Self-Recurrent SNNs (ScSr-SNNs). Recurrence in ScSr-SNNs is introduced in a stereotyped manner by adding self-recurrent connections to spiking neurons. The SNNs with self-recurrent connections can realize recurrent behaviors similar to those of more complex RSNNs while the error gradients can be more straightforwardly calculated due to the mostly feedforward nature of the network. The network dynamics is enriched by skip connections between nonadjacent layers. Moreover, we propose a new backpropagation (BP) method called backpropagated intrinsic plasticity (BIP) to further boost the performance of ScSr-SNNs by training intrinsic model parameters. Unlike standard intrinsic plasticity rules that adjust the neuron's intrinsic parameters according to neuronal activity, the proposed BIP method optimizes intrinsic parameters based on the backpropagated error gradient of a well-defined global loss function in addition to synaptic weight training. Here, based on challenging speech, neuromorphic speech, and neuromorphic image datasets, the proposed ScSr-SNNs can boost performance by up to 2.85% compared with other types of RSNNs trained by state-of-the-art BP methods.

97 MATHEMATICS AND COMPUTING↗

Invertible neural networks for real-time control of extrusion additive manufacturing

Material extrusion additive manufacturing (AM) has enabled an elegant fabrication pathway for a vast material library. Nonetheless, each material requires optimization of printing parameters generally determined through significant trial-and-error testing. To eliminate arduous, iteration-based optimization approaches, many researchers have used machine learning (ML) algorithms which provide opportunities for automated process optimization. Here, in this work, we demonstrate the use of an ML-driven approach for real-time material extrusion print-parameter optimization through in-situ monitoring of printed line geometry. To do this, we use deep invertible neural networks (INNs) which can solve both forward and inverse, or optimization, problems using a single network. By combining in-situ computer vision and deep INNs, the printing parameters can be autonomously optimized to print a target line width in 1.2 s. Furthermore, defects that occur during printing can be rapidly identified and corrected autonomously. The methods developed and presented in this work eliminate user-intensive, time-consuming, and iterative parameter discovery approaches that currently limit accelerated implementation of extrusion-based AM processes. Furthermore, the presented approach can be generalized to provide real-time monitoring and optimization pathways for increasingly complex AM environments.

36 MATERIALS SCIENCE↗

Inverse Design of Photonic Surfaces via High throughput Femtosecond Laser Processing and Tandem Neural Networks

Abstract This work demonstrates a method to design photonic surfaces by combining femtosecond laser processing with the inverse design capabilities of tandem neural networks that directly link laser fabrication parameters to their resulting textured substrate optical properties. High throughput fabrication and characterization platforms are developed that generate a dataset comprising 35280 unique microtextured surfaces on stainless steel with corresponding measured spectral emissivities. The trained model utilizes the nonlinear one‐to‐many mapping between spectral emissivity and laser parameters. Consequently, it generates predominantly novel designs, which reproduce the full range of spectral emissivities (average root‐mean‐squared‐error < 2.5%) using only a compact region of laser parameter space 25 times smaller than what is represented in the training data. Finally, the inverse design model is experimentally validated on a thermophotovoltaic emitter design application. By synergizing laser‐matter interactions with neural network capabilities, the approach offers insights into accelerating the discovery of photonic surfaces, advancing energy harvesting technologies.

36 MATERIALS SCIENCE↗

Application of machine learning for optical emission spectroscopy data in NAGDIS-II

In this study, we applied machine learning to optical emission spectroscopy (OES) data and device parameters from the linear plasma device NAGDIS-II to explore the potential application of machine learning for predicting electron density, $n$ e , and temperature, $T$ e . The covered ranges of $n$ e and $T$ e , which were measured by an electrostatic probe, are 3.6 × 10 17 –2.4 × 10 19 m -3 and 0.3–7.1 eV, respectively. A three hidden layer neural network (NN) is introduced to model the relationship between $n$ e /$T$ e and the combination of line intensities, radial position, and device parameters. It is shown that the errors in $n$ e and $T$ e become 18.0 and 18.8%, respectively, which were almost the same level for the electrostatic probe, using all available data. Lasso regression and greedy algorithm are used to select the necessary line emissions. In conclusion, it is shown that four- or five-line intensities are sufficient to obtain almost the same quality as the one with all the other lines.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Data-Efficient Strategies for Probabilistic Voltage Envelopes under Network Contingencies

This work presents an efficient data-driven method to construct probabilistic voltage envelopes (PVE) using power flow learning in grids with network contingencies. First, a network-aware Gaussian process (GP) termed Vertex-Degree Kernel (VDK-GP), developed in prior work, is used to estimate voltage–power functions for a few network configurations. The paper introduces a novel multi-task vertex degree kernel (MT-VDK) that amalgamates the learned VDK-GPs to determine power flows for unseen networks, with a significant reduction in the computational complexity and hyperparameter requirements compared to alternate approaches. Simulations on the IEEE 30-Bus network demonstrate the retention and transfer of power flow knowledge in both N-1 and N-2 contingency scenarios. The MT-VDK-GP approach achieves over 50 % reduction in mean prediction error for novel N-1 contingency network configurations in low training data regimes (50–250 samples) over VDK-GP. Additionally, MT-VDK-GP outperforms a hyper-parameter based transfer learning approach in over 75 % of N-2 contingency network structures, even without historical N-2 outage data. Furthermore, the proposed method demonstrates the ability to achieve PVEs using sixteen times fewer power flow solutions compared to Monte-Carlo sampling-based methods.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Machine learning enables national assessment of wind plant controls with implications for land use

Summary As deployment of wind energy continues to expand, computationally efficient tools for predicting wind plant performance over a wide range of layout designs, technology innovations, and spatial locations are increasingly important for policy and investment decisions. We demonstrate two approaches to training a surrogate model to predict annual energy production (AEP) of parameterized wind plant layouts: one using a Gaussian process (GP) and the other using a fully convolutional neural network (FCNN). We leverage the powerful FCNN architecture by encoding wind plant design parameters and output response surface as an image. The FCNN produces more accurate results than the GP with mean absolute errors equivalent to 1% and 1.9% of plant rated power, respectively, although the GP performs well under limited training data and provides useful uncertainty information. We also evaluate a surrogate model for wake steering, enabling a nationwide assessment of the impact of plant control strategies and plant layout decisions. Across two million locations, we find that wake steering strategies boost AEP with relative gains upwards of 3%. Gains are most pronounced at sites without a dominant wind direction and where layout optimization is less fruitful. Additionally, we perform a nationwide sensitivity analysis showing that wake steering can mitigate wake losses from higher density plant layouts. Our results suggest that regions which have not been previously viable for wind deployment due to moderate wind resources are especially well enhanced by wake steering strategies that could help overcome land constraints and inflexible layout options, potentially identifying new deployment opportunities.

17 WIND ENERGY↗

JANUS: Resilient and Adaptive Data Transmission for Enabling Timely and Efficient Cross-Facility Scientific Workflows

In modern science, the growing complexity of large-scale scientific projects has led to an increasing reliance on cross-facility scientific workflows, where resources and expertise from multiple institutions and geographic locations are leveraged to accelerate scientific discovery. These workflows often require transmitting huge amounts of scientific data through wide-area networks. Although high-speed networks like ESnet and transfer services such as Globus have improved data mobility, several challenges remain. The sheer volume of data can overwhelm network bandwidth, widely used transport protocols such as TCP suffer from inefficiencies due to retransmissions triggered by packet loss, and existing fault-tolerance mechanisms like erasure coding introduce substantial overhead. In this paper, we propose Janus, a resilient and adaptable data transmission approach designed for cross-facility scientific workflows. Unlike traditional TCP-based methods, Janus leverages UDP, integrates erasure coding for fault tolerance, and combines it with error-bounded lossy compression to reduce overhead. This novel design allows users to balance data transmission time and accuracy, optimizing transfer performance based on specific scientific requirements. Additionally, Janus dynamically adjusts erasure coding parameters in response to real-time network conditions, ensuring efficient data transfers even in fluctuating environments. We develop optimization models for determining ideal configurations and implement adaptive data transfer protocols to enhance reliability. Through extensive simulations and real-network experiments, we demonstrate that Janus significantly improves transfer efficiency while maintaining data fidelity.

Esaulov, Vladislav [Georgia State University, Atla↗

Residual-based error correction for neural operator accelerated infinite-dimensional Bayesian inverse problems

We explore using neural operators, or neural network representations of nonlinear maps between function spaces, to accelerate infinite-dimensional Bayesian inverse problems (BIPs) with models governed by nonlinear parametric partial differential equations (PDEs). Neural operators have gained significant attention in recent years for their ability to approximate the parameter-to-solution maps defined by PDEs using as training data solutions of PDEs at a limited number of parameter samples. The computational cost of BIPs can be drastically reduced if the large number of PDE solves required for posterior characterization are replaced with evaluations of trained neural operators. However, reducing error in the resulting BIP solutions via reducing the approximation error of the neural operators in training can be challenging and unreliable. We provide an a priori error bound result that implies certain BIPs can be ill-conditioned to the approximation error of neural operators, thus leading to inaccessible accuracy requirements in training. To reliably deploy neural operators in BIPs, we consider a strategy for enhancing the performance of neural operators: correcting the prediction of a trained neural operator by solving a linear variational problem based on the PDE residual. We show that a trained neural operator with error correction can achieve a quadratic reduction of its approximation error, all while retaining substantial computational speedups of posterior sampling when models are governed by highly nonlinear PDEs. The strategy is applied to two numerical examples of BIPs based on a nonlinear reaction–diffusion problem and deformation of hyperelastic materials. We demonstrate that posterior representations of the two BIPs produced using trained neural operators are greatly and consistently enhanced by error correction.

97 MATHEMATICS AND COMPUTING↗

Scalable deep learning for watershed model calibration

Watershed models such as the Soil and Water Assessment Tool (SWAT) consist of high-dimensional physical and empirical parameters. These parameters often need to be estimated/calibrated through inverse modeling to produce reliable predictions on hydrological fluxes and states. Existing parameter estimation methods can be time consuming, inefficient, and computationally expensive for high-dimensional problems. In this paper, we present an accurate and robust method to calibrate the SWAT model (i.e., 20 parameters) using scalable deep learning (DL). We developed inverse models based on convolutional neural networks (CNN) to assimilate observed streamflow data and estimate the SWAT model parameters. Scalable hyperparameter tuning is performed using high-performance computing resources to identify the top 50 optimal neural network architectures. We used ensemble SWAT simulations to train, validate, and test the CNN models. We estimated the parameters of the SWAT model using observed streamflow data and assessed the impact of measurement errors on SWAT model calibration. We tested and validated the proposed scalable DL methodology on the American River Watershed, located in the Pacific Northwest-based Yakima River basin. Our results show that the CNN-based calibration is better than two popular parameter estimation methods (i.e., the generalized likelihood uncertainty estimation [GLUE] and the dynamically dimensioned search [DDS], which is a global optimization algorithm). For the set of parameters that are sensitive to the observations, our proposed method yields narrower ranges than the GLUE method but broader ranges than values produced using the DDS method within the sampling range even under high relative observational errors. The SWAT model calibration performance using the CNNs, GLUE, and DDS methods are compared using R 2 and a set of efficiency metrics, including Nash-Sutcliffe, logarithmic Nash-Sutcliffe, Kling-Gupta, modified Kling-Gupta, and non-parametric Kling-Gupta scores, computed on the observed and simulated watershed responses. The best CNN-based calibrated set has scores of 0.71, 0.75, 0.85, 0.85, 0.86, and 0.91. The best DDS-based calibrated set has scores of 0.62, 0.69, 0.8, 0.77, 0.79, and 0.82. The best GLUE-based calibrated set has scores of 0.56, 0.58, 0.71, 0.7, 0.71, and 0.8. The scores above show that the CNN-based calibration leads to more accurate low and high streamflow predictions than the GLUE and DDS sets. Our research demonstrates that the proposed method has high potential to improve our current practice in calibrating large-scale integrated hydrologic models.

54 ENVIRONMENTAL SCIENCES↗

Finite-Time Analysis of Whittle Index based Q-Learning for Restless Multi-Armed Bandits with Neural Network Function Approximation

Whittle index policy is a heuristic to the intractable restless multi-armed bandits (RMAB) problem. Although it is provably asymptotically optimal, finding Whittle indices remains difficult. In this paper, we present Neural-Q-Whittle, a Whittle index based Q-learning algorithm for RMAB with neural network function approximation, which is an example of nonlinear two-timescale stochastic approximation with Q-function values updated on a faster timescale and Whittle indices on a slower timescale. Despite the empirical success of deep Q-learning, the non-asymptotic convergence rate of Neural-Q-Whittle, which couples neural networks with two-timescale Q-learning largely remains unclear. This paper provides a finite-time analysis of Neural-Q-Whittle, where data are generated from a Markov chain, and Q-function is approximated by a ReLU neural network. Our analysis leverages a Lyapunov drift approach to capture the evolution of two coupled parameters, and the nonlinearity in value function approximation further requires us to characterize the approximation error. Combing these provide Neural-Q-Whittle with convergence rate, where is the number of iterations.

reinforcement learning, structured learning, conve↗

Challenges for unsupervised anomaly detection in particle physics

Anomaly detection relies on designing a score to determine whether a particular event is uncharacteristic of a given background distribution. One way to define a score is to use autoencoders, which rely on the ability to reconstruct certain types of data (background) but not others (signals). In this paper, we study some challenges associated with variational autoencoders, such as the dependence on hyperparameters and the metric used, in the context of anomalous signal (top and W) jets in a QCD background. We find that the hyperparameter choices strongly affect the network performance and that the optimal parameters for one signal are non-optimal for another. In exploring the networks, we uncover a connection between the latent space of a variational autoencoder trained using mean-squared-error and the optimal transport distances within the dataset. We then show that optimal transport distances to representative events in the background dataset can be used directly for anomaly detection, with performance comparable to the autoencoders. Whether using autoencoders or optimal transport distances for anomaly detection, we find that the choices that best represent the background are not necessarily best for signal identification. These challenges with unsupervised anomaly detection bolster the case for additional exploration of semi-supervised or alternative approaches.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Development of a hybrid neural network and transfer learning model for optimized ICP-MS/MS operation

Correct function and calibration of instrumentation is a crucial assumption for any scientific experiment. One such instrument, tandem inductively coupled plasma mass spectrometer (ICP-MS/MS), has in-depth calibration settings that range across 30+ different parameters, making it difficult to determine optimal conditions without expertise and some degree of trial and error. Often, these settings are hand-tuned, a time-intensive process prone to local maxima and human error. While some automation is available, the automation also may favor local optimizations over a global optimum. In addition to these difficulties, day to day instrument variability can further complicate the calibration process. We propose a solution to this problem as a machine learning (ML) algorithm that learns how each parameter helps determine the calibration sensitivity across several elements, and re-weights parameters over time as instrument variability changes (e.g., a global neural network (NN) with a time-dependent transfer learning (TL) component). This model would be able to generate a surface of predicted calibration sensitivities and their respective parameters, and a simple multivariate algorithm would be able to pull out the optimum results with the settings associated with them. Here-in, we describe our initial findings in working towards this goal, including data extraction from historical files, exploratory data analysis, and some initial model building to better describe the data and the feasibility of our goal.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Lossy compression of statistical data using quantum annealer

Abstract We present a new lossy compression algorithm for statistical floating-point data through a representation learning with binary variables. The algorithm finds a set of basis vectors and their binary coefficients that precisely reconstruct the original data. The optimization for the basis vectors is performed classically, while binary coefficients are retrieved through both simulated and quantum annealing for comparison. A bias correction procedure is also presented to estimate and eliminate the error and bias introduced from the inexact reconstruction of the lossy compression for statistical data analyses. The compression algorithm is demonstrated on two different datasets of lattice quantum chromodynamics simulations. The results obtained using simulated annealing show 3–3.5 times better compression performance than the algorithm based on neural-network autoencoder. Calculations using quantum annealing also show promising results, but performance is limited by the integrated control error of the quantum processing unit, which yields large uncertainties in the biases and coupling parameters. Hardware comparison is further studied between the previous generation D-Wave 2000Q and the current D-Wave Advantage system. Our study shows that the Advantage system is more likely to obtain low-energy solutions for the problems than the 2000Q.

97 MATHEMATICS AND COMPUTING↗

Lightweight and effective tensor sensitivity for atomistic neural networks

Atomistic machine learning focuses on the creation of models that obey fundamental symmetries of atomistic configurations, such as permutation, translation, and rotation invariances. In many of these schemes, translation and rotation invariance are achieved by building on scalar invariants, e.g., distances between atom pairs. There is growing interest in molecular representations that work internally with higher rank rotational tensors, e.g., vector displacements between atoms, and tensor products thereof. Here, we present a framework for extending the Hierarchically Interacting Particle Neural Network (HIP-NN) with Tensor Sensitivity information (HIP-NN-TS) from each local atomic environment. Crucially, the method employs a weight tying strategy that allows direct incorporation of many-body information while adding very few model parameters. We show that HIP-NN-TS is more accurate than HIP-NN, with negligible increase in parameter count, for several datasets and network sizes. As the dataset becomes more complex, tensor sensitivities provide greater improvements to model accuracy. In particular, HIP-NN-TS achieves a record mean absolute error of 0.927 $\frac{\text{kcal}}{\text{mol}}$ for conformational energy variation on the challenging COMP6 benchmark, which includes a broad set of organic molecules. We also compare the computational performance of HIP-NN-TS to HIP-NN and other models in the literature.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Deep Learning without Global Optimization by Random Fourier Neural Networks

Here we introduce a new training algorithm for deep neural networks that utilize random complex exponential activation functions. Our approach employs a Markov chain Monte Carlo sampling procedure to iteratively train network layers, avoiding global and gradient-based optimization while maintaining error control. It consistently attains the theoretical approximation rate for residual networks with complex exponential activation functions, determined by network complexity. Additionally, it enables efficient learning of multiscale and high-frequency features, producing interpretable parameter distributions. Despite using sinusoidal basis functions, we do not observe Gibbs phenomena in approximating discontinuous target functions.

97 MATHEMATICS AND COMPUTING↗

Comprehensive Material Characterization and Simultaneous Model Calibration for Improved Computational Simulation Credibility

Computational simulation is increasingly relied upon for high-consequence engineering decisions, and a foundational element to solid mechanics simulations is a credible material model. Our ultimate vision is to interlace material characterization and model calibration in a real-time feedback loop, where the current model calibration results will drive the experiment to load regimes that add the most useful information to reduce parameter uncertainty. The current work investigated one key step to this Interlaced Characterization and Calibration (ICC) paradigm, using a finite load-path tree to incorporate history/path dependency of nonlinear material models into a network of surrogate models that replace computationally-expensive finite-element analyses. Our reference simulation was an elastoplastic material point subject to biaxial deformation with a Hill anisotropic yield criterion. Training data was generated using either a space-filling or adaptive sampling method, and surrogates were built using either Gaussian process or polynomial chaos expansion methods. Surrogate error was evaluated to be on the order of 10 ⁻5 and 10 ⁻3 percent for the space-filling and adaptive sampling training data, respectively. Direct Bayesian inference was performed with the surrogate network and with the reference material point simulator, and results agreed to within 3 significant figures for the mean parameter values, with a reduction in computational cost over 5 orders of magnitude. These results bought down risk regarding the surrogate network and facilitated a successful FY22-24 full LDRD proposal to research and develop the complete ICC paradigm.

36 MATERIALS SCIENCE↗

Reconstructing Richtmyer–Meshkov instabilities from noisy radiographs using low dimensional features and attention-based neural networks

We develop an ML-based approach for density reconstruction based on transformer neural networks. This approach is demonstrated in the setting of ICF-like double shell hydrodynamic simulations wherein the parameters related to material properties and initial conditions are varied. The new method can robustly recover the complex topologies given by the Richtmyer-Meshkoff instability (RMI) from a sequence of hydrodynamic features derived from radiographic images corrupted with blur, scatter, and noise. A noise model is developed to characterize errors in extracting features from synthetic radiographs of the simulated density field. The key component of the network is a transformer encoder that acts on a sequence of features extracted from noisy radiographs. This encoder includes numerous self-attention layers that act to learn temporal dependencies in the input sequences and increase the expressiveness of the model. This approach is shown to exhibit an excellent ability to accurately recover the RMI growth rates, despite the gas-metal interface being greatly obscured by radiographic noise. Our approach can be applied in a broad array of fields involving shock physics and material science.

47 OTHER INSTRUMENTATION↗

Enhanced analysis of experimental x-ray spectra through deep learning

X-ray spectroscopic data from high-energy-density laser-produced plasmas has long required thorough, time-consuming analysis to extract meaningful source conditions. There are often confounding factors due to rapidly evolving states and finite spatial gradients (e.g., the existence of multi-temperature, multi-density, multi-ionization states, etc.) that make spectral measurements and analysis difficult. Here, in this paper, we demonstrate how deep learning can be applied to enhance x-ray spectral data analysis in both speed and intricacy. Neural networks (NNs) are trained on ensemble atomic physics simulations so that they can subsequently construct a model capable of extracting plasma parameters directly from experimental spectra. Through deep learning, the models can extract temperature distributions as opposed to single or dual temperature/density fits from standard trial-and-error atomic modeling at a significantly reduced computational cost compared to traditional trial-and-error methods. These NNs are envisioned to be deployed with high repetition rate x-ray spectrometers in order to provide detailed real-time analysis of experimental spectra.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗