Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Network parameter error”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

What do physics-informed DeepONets learn? Understanding and improving training for scientific computing applications

Physics-informed deep operator networks (DeepONets) have emerged as a promising approach toward numerically approximating the solution of partial differential equations (PDEs). In this work, we aim to develop further understanding of what is being learned by physics-informed DeepONets by assessing the universality of the extracted basis functions and demonstrating their potential toward model reduction with spectral methods. Results provide clarity about measuring the performance of a physics-informed DeepONet through the decays of singular values and expansion coefficients. In addition, we propose a transfer learning approach for improving training for physics-informed DeepONets between parameters of the same PDE as well as across different, but related, PDEs where these models struggle to train well. This approach results in significant error reduction and learned basis functions that are more effective in representing the solution of a PDE.

Deep operator networks↗

A hybrid Penman-Monteith and machine learning model for simulating evapotranspiration and its components

Integrating physical processes with machine learning has advanced evapotranspiration (ET) simulation, yet most hybrid models fail to partition total ET into its components: soil evaporation (E) and vegetation transpiration (T). This study introduces Residual Neural Network–Penman–Monteith (RNN-PM), a novel hybrid dual-source ET model designed to overcome this limitation. The model synergizes the physically-based Penman–Monteith framework with three specialized residual neural networks trained to estimate key conductance parameters (canopy conductance, soil surface conductance, and aerodynamic conductance). Furthermore this explicit parameterization allows for the direct partitioning of total ET. Validation at National Ecological Observatory Network (NEON) flux sites using high-frequency partitioned E and T shows that RNN-PM reliably reproduces ET and the transpiration fraction (T/ET). For ET, the model achieves an average Kling–Gupta efficiency (KGE) of 0.89 and a root-mean-square error (RMSE) of 0.55 mm/day; for T/ET, the KGE is 0.87 with an RMSE of 0.06. Furthermore, RNN-PM demonstrates robust generalization, accurately simulating ET and its components well beyond the initial training dataset, even under extreme climatic conditions. This study extended the analysis by comparing the RNN-PM model with seven established dual-source ET models. The results indicate that RNN-PM outperforms both conventional machine learning models and purely physical process-based models in simulating ET components in most cases. Among the purely physical process-based dual-source models, those based on surface temperature decomposition showed improved performance as the leaf area index (LAI) decreased when evaluated against high-frequency ET component datasets. In contrast, the performance of conductance-based dual-source models declined with decreasing LAI. Although purely machine learning-based models can produce relatively accurate simulations of ET components, they often exhibit limited generalization capability, an issue that the RNN-PM model effectively overcomes. Ultimately, the RNN-PM model represents a significant advance in simulating ET components, offering a novel and scalable approach for improving the representation of land–atmosphere interactions in Earth system models.

54 ENVIRONMENTAL SCIENCES↗

Hamiltonian learning using machine-learning models trained with continuous measurements

Here, we build upon recent work on the use of machine-learning models to estimate Hamiltonian parameters using continuous weak measurement of qubits as input. We consider two settings for the training of our model: (1) supervised learning, where the weak-measurement training record can be labeled with known Hamiltonian parameters, and (2) unsupervised learning, where no labels are available. The first has the advantage of not requiring an explicit representation of the quantum state, thus potentially scaling very favorably to a larger number of qubits. The second requires the implementation of a physical model to map the Hamiltonian parameters to a measurement record, which we implement using an integrator of the physical model with a recurrent neural network to provide a model-free correction at every time step to account for small effects not captured by the physical model. We test our construction on a system of two qubits and demonstrate accurate prediction of multiple physical parameters in both the supervised context and the unsupervised context. We demonstrate that the model benefits from larger training sets, establishing that it is “learning,” and we show robustness regarding errors in the assumed physical model by achieving accurate parameter estimation in the presence of unanticipated single-particle relaxation.

97 MATHEMATICS AND COMPUTING↗

Spread Spectrum Time Domain Reflectometry (SSTDR) Digital Twin Simulation of Photovoltaic Systems for Fault Detection and Location

Utilizing spread spectrum time domain reflectometry (SSTDR) to detect, locate, and characterize faults in photovoltaic (PV) systems is examined in this paper. We present a method to obtain the model parameters that are needed to produce digital twin SSTDR responses for PV systems. The digital twin SSTDR responses could be used to predict faults within the PV systems. Here, the model parameters are the reflection and transmission coefficients at each impedance discontinuity in the PV system along with the propagation coefficients across each PV cable segment. We obtain model parameter by applying inverse modeling techniques to experimental SSTDR data associated with PV systems. Our model parameters can be used in any digital twin simulation method for modeling reflectometry in frequency-dependent and complex loads. For validation, we used the model parameters in a graph network simulation engine and adapted it to be used for SSTDR digital twin simulations in PV systems. We produced simulations for 0 to 10 PV modules connected in series. We also simulated SSTDR responses for open circuit disconnections in a PV setup containing 10 PV modules in series. Results show that all but one simulated disconnect locations match experimental disconnection locations of the same setup with an error of less than 5%.

14 SOLAR ENERGY↗

Machine Learning-Based Process Control for Injection Molding of Recycled Polypropylene

The increased interest in artificial intelligence in manufacturing has driven the adoption of machine learning to optimize processes and improve efficiency. A key challenge in injection molding is the variability of recycled materials, which affects part quality and processing stability. This study presents a novel closed-loop process control approach for injection molding, leveraging machine learning to adaptively predict processing inputs and quality outcomes. The methodology was tested on five blends of recycled polypropylene (rPP), using artificial neural networks (ANNs), linear regression, and polynomial regression to model the relationships between material properties and process parameters. The dataset was split 80/20 into training and testing sets. The ANN model was implemented using TensorFlow and Keras, with six hidden layers of 32 neurons per layer, ReLU activation, and an Adam optimizer. Empirical tuning and early stopping were used to optimize performance and prevent overfitting. Predictions were evaluated based on mean absolute error (MAE), mean squared error (MSE), and percentage error. The results showed that yield stress, ultimate elongation, and part weight were accurately predicted within a 5% error for linear and polynomial regression models and within a 10% error for the ANN. However, modulus predictions were less reliable, with errors of ~11% for ANN and linear regression and ~40% for polynomial regression, reflecting the inherent variability of this property in rPP blends. Predictions of processing inputs had errors ranging from 3% to 25%, depending on the model and response variable. No single modeling approach was consistently superior across all responses, highlighting the complexity of the relationship between material properties, process parameters, and quality metrics. Overall, the work demonstrates that closed-loop process control, powered by machine learning, can effectively predict key quality parameters in injection molding of recycled materials. The proposed approach can improve process stability and material utilization, facilitating increased adoption of sustainable materials.

Krantz, Joshua↗

Component-Level Inverse Design of Transmon Qubits Using Neural Networks

Designing a superconducting qubit to realize specific Hamiltonian parameters typically requires iterating through a time and compute-intensive forward loop in which the designer chooses a layout geometry, simulates it, extracts circuit parameters such as capacitances, and refines the geometry. We study the inverse version of this task using a neural-network workflow that maps target Hamiltonian parameters directly to component-level layout parameters, which we subsequently demonstrate on a planar transmon layout. During training, we pair the inverse model with a frozen forward surrogate model and evaluate the loss in Hamiltonian space rather than in layout-parameter space. In validation against a conventional EM solver, 97% of generated designs produce usable geometries, and the inverse-plus-surrogate pipeline reaches mean percent errors of 0.73% for qubit frequency and 1.58% for anharmonicity, comparable to or below the fabrication and simulation-to-measurement uncertainty expected for academic-process transmon devices of this type. A single pipeline query takes ~60 ms on CPU, versus ~2 min for a conventional EM capacitance extraction on the same hardware, a speedup of approximately 2,000x. Batching minimizes the AI model inference overhead, reducing the runtime to 3.1 microseconds per sample on CPU and 2.6 microseconds per sample on GPU at a batch size of 2048, resulting in speedups of 3.9 x 10^7 and 4.6 x 10^7, respectively, relative to a single conventional CPU EM extraction. Our results indicate that component-level inverse design usefully extends and complements conventional EM simulation, including for small datasets on the order of 1,000 samples.

Seidel, Olivia [Fermilab; Texas U., Arlington]↗

PINN surrogate of Li-ion battery models for parameter inference, Part II: Regularization and application of the pseudo-2D model

Bayesian parameter inference is useful to improve Li-ion battery diagnostics and can help formulate battery aging models. However, it is computationally intensive and cannot be easily repeated for multiple cycles, multiple operating conditions, or multiple replicate cells. To reduce the computational cost of Bayesian calibration, numerical solvers for physics-based models can be replaced with faster surrogates. A physics-informed neural network (PINN) is developed as a surrogate for the pseudo-2D (P2D) battery model calibration. For the P2D surrogate, additional training regularization was needed as compared to the PINN single-particle model (SPM) developed in Part I. Both the PINN SPM and P2D surrogate models are exercised for parameter inference and compared to data obtained from a direct numerical solution of the governing equations. A parameter inference study highlights the ability to use these PINNs to calibrate scaling parameters for the cathode Li diffusion and the anode exchange current density. By realizing computational speed-ups of ~2250x for the P2D model, as compared to using standard integrating methods, the PINN surrogates enable rapid state-of-health diagnostics. Finally, in the low-data availability scenario, the testing error was estimated to ~2 mV for the SPM surrogate and ~10 mV for the P2D surrogate which could be mitigated with additional data.

25 ENERGY STORAGE↗

Efficient Bayesian inference with latent Hamiltonian neural networks in No-U-Turn Sampling

When sampling for Bayesian inference, one popular approach in the computational field is to use Hamiltonian Monte Carlo (HMC) and specifically the No-U-Turn Sampler (NUTS), which automatically decides the end time of the Hamiltonian trajectory. However, HMC and NUTS can require numerous numerical gradients of the target density and can prove slow in practice when relying on computationally expensive forward models. We propose Latent Hamiltonian neural networks (L-HNNs) with HMC and NUTS for solving Bayesian inference problems. Once trained, L-HNNs do not require numerical gradients of the target density during sampling, and hence numerous evaluations of the forward computational model. Moreover, L-HNNs satisfy important properties such as perfect time reversibility and Hamiltonian conservation, making them well-suited for use within HMC and NUTS because stationarity can be shown. We also propose the integration of L-HNNs in an online error monitoring scheme, in which numerical gradients of the target density are used for a few samples whenever the L-HNNs prediction errors are large. This online error monitor scheme prevents sample degeneracy in regions of low probability density and ensures robust uncertainty quantification. We demonstrate L-HNNs in NUTS with online error monitoring on several analytical examples involving complex, heavy-tailed, and high-local-curvature probability densities. We then demonstrate the applicability of L-HNNs in NUTS to two computational case studies, namely the Allen-Cahn stochastic partial differential equation and an elliptic partial differential equation with 25 and 50 inference parameters, respectively. Overall, the L-HNNs in NUTS with online error monitoring satisfactorily inferred these probability densities. In conclusion, compared to traditional NUTS, L-HNNs in NUTS with online error monitoring required 1–2 orders of magnitude fewer numerical gradients of the target density and improved the effective sample size (ESS) per gradient (which is a measure of both the sampling quality and the computational expense) by an order of magnitude.

97 MATHEMATICS AND COMPUTING↗

A framework for data-driven solution and parameter estimation of PDEs using conditional generative adversarial networks

We employ and adapt the image-to-image translation concept based on conditional generative adversarial networks (cGAN) for learning a forward and an inverse solution operator of partial differential equations (PDEs). We focus on steady-state solutions of coupled hydromechanical processes in heterogeneous porous media and present the parameterization of the spatially heterogeneous coefficients, which is exceedingly difficult using standard reduced-order modeling techniques. We show that our framework provides a speed-up of at least 2,000 times compared to a finite-element solver and achieves a relative root-mean-square error (r.m.s.e.) of less than 2% for forward modeling. For inverse modeling, the framework estimates the heterogeneous coefficients, given an input of pressure and/or displacement fields, with a relative r.m.s.e. of less than 7%, even for cases where the input data are incomplete and contaminated by noise. The framework also provides a speed-up of 120,000 times compared to a Gaussian prior-based inverse modeling approach while also delivering more accurate results.

97 MATHEMATICS AND COMPUTING↗

Towards multi-fidelity deep learning of wind turbine wakes

We report engineering wake models that accurately predict wake in a computationally efficient manner are very important for tasks such as layout optimization and control of wind farms. In this paper, we explore an application of deep learning (DL) to learn the wake model from hierarchies of physics-based approaches ranging from analytical models to an approximate form of the Reynolds-averaged Navier-Stokes equations. We first illustrate the application of principal component analysis to obtain a lower-dimensional representation that allows a computationally tractable training and deployment of DL models. Then, the DL model is trained to learn the mapping from input parameter space to the principal components, which are then used to reconstruct the three-dimensional flow field. Additionally, we investigate a composite framework consisting of two neural networks to learn the correlation between low- and high-fidelity data with Gauss and curl models treated as proxies for low- and high-fidelity models, respectively. The prediction from both DL models matches well with the high-fidelity data with a maximum relative percentage error for the kinetic energy flux of <1%. This work opens up possibilities for data-efficient construction of surrogate models for wake prediction that can be used to study the influence of wind speed and yaw angles on wind farm power production.

17 WIND ENERGY↗

Fault-Tolerant Decentralized Control for Large-Scale Inverter-Based Resources for Active Power Tracking

Integration of inverter-based resources (IBRs) which lack the intrinsic characteristics such as the inertial response of the traditional synchronous-generator (SG)-based sources presents a new challenge in the form of analyzing the grid stability under their presence. While the dynamic composition of IBRs differs from that of the SGs, the control objective remains similar in terms of tracking the desired active power. This letter presents a decentralized primal-dual-based fault-tolerant control framework for the power allocation in IBRs. Overall, a hierarchical control algorithm is developed with a lower level addressing the current control and the parameter estimation for the IBRs and the higher level acting as the reference power generator to the low level based on the desired active power profile. The decentralized network-based algorithm adaptively splits the desired power between the IBRs taking into consideration the health of the IBRs transmission lines. The proposed framework is tested through a simulation on the network of IBRs and the high-level controller performance is compared against the existing framework in the literature. The proposed algorithm shows significant performance improvement in the magnitude of power deviation and settling time to the nominal value under faulty conditions as compared to the algorithm in the literature.

24 POWER TRANSMISSION AND DISTRIBUTION↗

ROM-Based Surrogate Systems Modeling of EBR-II

We report the System Analysis Module (SAM), developed and maintained by Argonne National Laboratory, is designed to provide whole-plant transient safety analysis capabilities for a number of advanced non-light water reactors, including sodium-cooled fast reactor (SFR), lead-cooled fast reactor (LFR), and molten salt reactor (MSR)/fluoride-salt-cooled high-temperature reactor (FHR) designs. SAM is primarily constructed as a systems-level analysis tool, with the potential to incorporate reduced order models from three-dimensional computational fluid dynamics (CFD) simulations to improve characterization of complex, multidimensional physics. It is recognized that the computational expense associated with CFD can be intractable for various engineering analyses, such as uncertainty quantification, inference, and design optimization. This paper explores the reducibility of a SAM model using recent advances in randomized linear algebra techniques, which attempt to find recurring patterns in the various realizations generated by a model after randomly perturbing all its input parameters. The reduction is described in terms of fewer degrees of freedom (DOFs), referred to as the active DOFs, for the model variables such as input model parameters and model responses. The results indicate that there is significant room for additional reduction that may be leveraged for additional computational gains when employing SAM for engineering-intensive analyses that require repeated model executions. Different from physics-based reduction approaches, the proposed approach allows one to estimate upper bounds on the reduction errors, which are rigorously developed in this work. Finally, different methods for surrogate model construction, such as regression and neural network-based training, are employed to correlate the input and output active DOFs, which are related back to the original variables using matrix-based linear transformations.

42 ENGINEERING↗

Uncertainty propagation in feed-forward neural network models

We develop new uncertainty propagation methods for feed-forward neural network architectures with leaky ReLU activation functions subject to random perturbations in the input vectors. In particular, we derive analytical expressions for the probability density function (PDF) of the neural network output and its statistical moments as a function of the input uncertainty and the parameters of the network, i.e., weights and biases. A key finding is that an appropriate linearization of the leaky ReLU activation function yields accurate statistical results even for large perturbations in the input vectors. This can be attributed to the way information propagates through the network. We also propose new analytically tractable Gaussian copula surrogate models to approximate the full joint PDF of the neural network output. To validate our theoretical results, we conduct Monte Carlo simulations and a thorough error analysis on a multi-layer neural network representing a nonlinear integro-differential operator between two polynomial function spaces. Our findings demonstrate excellent agreement between the theoretical predictions and Monte Carlo simulations.

MLP networks↗

Application of machine learning in the determination of impact parameter in the 132 Sn+ 124 Sn system

Here, 132 Sn + 124 Sn collisions at a beam energy of 270 MeV/nucleon were performed at the Radioactive Isotope Beam Factory (RIBF) in RIKEN to investigate the nuclear equation of state. Reconstructing the impact parameter is one of the important tasks in the experiment as it relates to many observable. In this work, we employ three commonly used algorithms in machine learning, the artificial neural network (ANN), the convolutional neural network (CNN), and the light gradient boosting machine (LightGBM), to determine the impact parameter by analyzing either the charged particle spectra or several features simulated with events from the ultrarelativistic quantum molecular dynamics (UrQMD) model. To closely imitate experimental data and investigate the generalizability of the trained machine learning algorithms, incompressibility of nuclear equation of state and the in-medium nucleon-nucleon cross sections are varied in the UrQMD model to generate the training data. The mean absolute error Δb between the true and the predicted impact parameter is smaller than 0.45 fm if training and testing sets are sampled from the UrQMD model with the same parameter set. However, if training and testing sets are sampled with different parameter sets, Δb would increase to 0.8 fm. The generalizability of the trained machine learning algorithms suggests that these machine learning algorithms can be used reliably to reconstruct the impact parameter in experiment.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

GAINN: The Galaxy Assembly and Interaction Neural Networks for High-redshift JWST Observations

We present the Galaxy Assembly and Interaction Neural Networks (Gainn), a series of artificial neural networks for predicting the redshift, stellar mass, halo mass, and mass-weighted age of simulated galaxies based on James Webb Space Telescope (JWST) photometry. Our goal is to determine the best neural network for predicting these variables at 11 < z < 15. The parameters of the optimal neural network can then be used to estimate these variables for real, observed galaxies. The inputs of the neural networks are JWST filter magnitudes of a subset of five broadband filters (F150W, F200W, F277W, F356W, and F444W) and two medium-band filters (F162M and F182M). We compare the performance of the neural networks using different combinations of these filters, as well as different activation functions and numbers of layers. The best neural network predicted redshift with a normalized rms error of $0.010^{+0.003}_{-0.001}$, stellar mass with rms = $0.089^{+0.044}_{-0.022}$, halo mass with a mean-squared error of $0.022^{+0.014}_{-0.008}$, and mass-weighted age with rms = $12.466^{+5.065}_{-2.408}$. We also test the performance of Gainn on real data from MACS0647JD, an object observed by JWST. Predictions from Gainn for the first projection of the object (JD1) have normalized bias $\langle$Δz$\rangle$ < 0.00228, which is significantly smaller than found with template-fitting methods. We find that the optimal filter combination is F277W, F356W, F162M, and F200W when considering both theoretical accuracy and observational resources from JWST.

97 MATHEMATICS AND COMPUTING↗

Smart Pixels: towards on-sensor inference of charged particle track parameters and uncertainties

The combinatorics of track seeding has long been a computational bottleneck for triggering and offline computing in High Energy Physics (HEP), and remains so for the HL-LHC. Next-generation pixel sensors will be sufficiently fine-grained to determine angular information of the charged particle passing through from pixel-cluster properties. This detector technology immediately improves the situation for offline tracking, but any major improvements in physics reach are unrealized since they are dominated by lowest-level hardware trigger acceptance. We will demonstrate track angle and hit position prediction, including errors, using a mixture density network within a single layer of silicon as well as the progress towards and status of implementing the neural network in hardware on both FPGAs and ASICs.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Smart Pixels: towards on-sensor inference of charged particle track parameters and uncertainties

The combinatorics of track seeding has long been a computational bottleneck for triggering and offline computing in High Energy Physics (HEP), and remains so for the HL-LHC. Next-generation pixel sensors will be sufficiently fine-grained to determine angular information of the charged particle passing through from pixel-cluster properties. This detector technology immediately improves the situation for offline tracking, but any major improvements in physics reach are unrealized since they are dominated by lowest-level hardware trigger acceptance. We will demonstrate track angle and hit position prediction, including errors, using a mixture density network within a single layer of silicon as well as the progress towards and status of implementing the neural network in hardware on both FPGAs and ASICs.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Smartpixels: Towards on-sensor inference of charged particle track parameters and uncertainties

The combinatorics of track seeding has long been a computational bottleneck for triggering and offline computing in High Energy Physics (HEP), and remains so for the HL-LHC. Next-generation pixel sensors will be sufficiently fine-grained to determine angular information of the charged particle passing through from pixel-cluster properties. This detector technology immediately improves the situation for offline tracking, but any major improvements in physics reach are unrealized since they are dominated by lowest-level hardware trigger acceptance. We will demonstrate track angle and hit position prediction, including errors, using a mixture density network within a single layer of silicon as well as the progress towards and status of implementing the neural network in hardware on both FPGAs and ASICs.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗