Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Statistical accuracy”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Quantitative Insight to Fission Gas Pores Distribution in Irradiated Annular U-10Zr Metallic Fuel Using Machine Learning

Metallic fuels, particularly U-10Zr and its performance in reactor irradiation conditions, have been thoroughly investigated and are a promising candidate for next-generation sodium-cooled fast spectrum nuclear reactors. Irradiation in reactors can lead to the formation of fission gas and increased pore formation which can significantly impact fuel performance. Due to the large number of pores and various phases formed in metallic fuel during irradiation, a quantitative description of fission gas pores as a function of irradiation conditions is not yet available, undermining the fidelity of fuel performance modeling to support fuel qualification. It has been difficult to clearly detect pore boundaries and distinguish matrix phases from fission gas pores using optical microscopy by using simple threshold methods working with low magnification images. The pre-trained deep learning model for fission gas pore detection was applied to ~10,260 high magnification scanning electron microscopy images. The model increased the accuracy of fission gas pore segmentation to obtain statistical features, which cannot be processed manually. A pre-trained decision tree model was used to classify pores as isolated or connected pores, providing new insight into the correlation between the movement of lanthanides, solid fission products, and the radial temperature gradient developed in fuel irradiation conditions. This paper emphasizes the potential that artificial intelligence-based machine learning models have to accelerate qualification and support nuclear fuel development.

36 MATERIALS SCIENCE↗

Tree-Based Ensemble Learning Models for Wall Temperature Predictions in Post-Critical Heat Flux Flow Regimes at Subcooled and Low-Quality Conditions

Accurately predicting post-critical heat flux (CHF) heat transfer is an important but challenging task in water-cooled reactor design and safety analysis. Although numerous heat transfer correlations have been developed to predict post-CHF heat transfer, these correlations are only applicable to relatively narrow ranges of flow conditions due to the complex physical nature of the post-CHF heat transfer regimes. In this paper, a large quantity of experimental data is collected and summarized from the literature for steady-state subcooled and low-quality film boiling regimes with water as the working fluid in vertical tubular test sections. In addition, a low-quality water film boiling (LWFB) database is consolidated with a total of 22,813 experimental data points, which cover a wide flow range of the system pressure from 0.1 to 9.0 MPa, mass flux from 25 to 2750 kg/m 2 s, and inlet subcooling from 1 to 70 °C. Two machine learning (ML) models, based on random forest (RF) and gradient boosted decision tree (GBDT), are trained and validated to predict wall temperatures in post-CHF flow regimes. The trained ML models demonstrate significantly improved accuracies compared to conventional empirical correlations. To further evaluate the performance of these two ML models from a statistical perspective, three criteria are investigated and three metrics are calculated to quantitatively assess the accuracy of these two ML models. For the full LWFB database, the root-mean-square errors between the measured and predicted wall temperatures by the GBDT and RF models are 5.7% and 6.2%, respectively, confirming the accuracy of the two ML models.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Full-stack Quantification of Variability in Predicting Ion Transport Properties using Machine-learned Interatomic Potentials

Machine-learned interatomic potentials (MLIPs) have become the state-of-the-art for performing accurate, scalable molecular dynamics (MD) simulations. It is therefore crucial to understand and quantify the reliability of MLIPs for downstream property predictions. Uncertainty in predicted properties can arise from limitations in first-principles training data, intrinsic MLIP model errors in representing the data, and the statistical noise introduced during subsequent MD simulations. Using ion transport in Li7P3S11 as a case study, we systematically assess the impact of training set size and selection, neural network stochasticity, and MD sampling statistics on predicted diffusivity and activation energy. We find that when using equivariant MLIP architectures with standard MD protocols, uncertainty arising from MD sampling dominates over model-induced errors. In contrast, MLIP errors relative to the underlying first-principles data are consistently minor. Given this, there are two main routes to improving the accuracy of predictions based on MLIP potentials: adopting higher accuracy reference data generation methods, and improving the MD sampling statistics.

36 MATERIALS SCIENCE↗

Evaluation of artificial neural network performance for classification of potato plants infected with potato virus Y using spectral data on multiple varieties and genotypes

Potato virus Y (Potyviridae, PVY) is a plant virus that poses a significant threat to potato producers on a global basis. The pathogen has disrupted seed potato supplies and negatively impacted yield and quality of commercial potato crops. The potato industry currently manages PVY infection levels via insecticide applications, regional seed certification programs that rely on field scouting to visually assess individual plants for infection status, and destructive and costly tissue sampling coupled with laboratory assays. Despite these efforts, PVY continues to confound potato industry stakeholders resulting in economic harm. Remote sensing and machine learning provide for the development of new tools to more accurately detect and spatially quantify PVY-infected plants versus the current state of the art. However, there is a need to understand how the occurrence of many different potato varieties impact the dynamics of developing models to detect potato plants impacted with PVY and their potential effectiveness. This study evaluates classification modelling outcomes using spectral datasets collected in different temporal and spatial environments (greenhouse and a production field) on multiple potato varieties consisting of labelled instances of plants infected with PVY and those not infected with the virus. A modelling framework was developed to support iterative modelling runs using artificial neural network (ANN) architectures configured as binary classifiers to develop sample populations to support statistical analysis on model performance using specific spectral subsets. When using spectral data to detect PVY-infected plants, ANN models achieved the highest mean accuracy of 0.894 on a single variety. Conversely, the same ANN model architecture only achieved a mean accuracy of 0.575 on a spectral data set representing 29 potato breeding lines. Additionally, statistical analysis indicates spectral regions including the red edge, near infrared and shortwave infrared contain more important spectral features for the ANN classifier introduced in this research.

60 APPLIED LIFE SCIENCES↗

Modeling the Behavior of Complex Aqueous Electrolytes Using Machine Learning Interatomic Potentials: The Case of Sodium Sulfate

Understanding the structure and thermodynamics of solvated ions is essential for advancing applications in electrochemistry, water treatment, and energy storage. While ab initio molecular dynamics methods are highly accurate, they are limited by short accessible time and length scales whereas classical force fields struggle with accuracy. Herein, we explore the structure and thermodynamics of complex monovalent-divalent ion pairs using Na 2 SO 4 (aq) as a case study by applying a machine learning interatomic potential (MLIP) trained on density functional theory (DFT) data. Our MLIP-based approach reproduces key bulk properties such as density and radial distribution functions of water. We provide the hydration structure of the sodium and sulfate ions in the 0.1–2 M concentration range and the one-dimensional and two-dimensional potentials of mean force for the sodium–sulfate ion pairing at the low concentration limit (0.1 M), which are inaccessible to DFT. At low concentrations, the sulfate ion is strongly solvated, leading to the stabilization of solvent-separated ion pairs over contact ion pairs. Minimum energy pathway analysis revealed that coordinating two sodium ions with a sulfate ion is a multistep process whereby the sodium ions coordinate to the sulfate ion sequentially. Finally, we demonstrate that MLIPs allow the study of solvated ions beyond simple monovalent pairs with DFT-level accuracy in their low concentration limit (0.1 M) via statistically converged properties from ns-long simulations.

anions↗

Extracting forces from noisy dynamics in dusty plasmas

Extracting environmental forces from noisy data is a common yet challenging task in complex physical systems. Machine learning (ML) represents a robust approach to this problem, yet is mostly tested on simulated data with known parameters. Here we use supervised ML to extract the electrostatic, dissipative, and stochastic forces acting on micron-sized charged particles levitated in an argon plasma (dusty plasma). By tracking the sub-pixel motion of particles in subsequent images, we successfully estimated these forces from their random motion. The experiments contained important sources of non-Gaussian noise, such as drift and pixel-locking, representing a data mismatch from methods used to analyze simulated data with purely Gaussian noise. Our model was trained on simulated particle trajectories that included all of these artifacts, and used more than 100 dynamical and statistical features, resulting in a prediction with 50\% better accuracy than conventional methods. Lastly, in systems with two interacting particles, the model provided non-contact measurements of the particle charge and Debye length in the plasma environment.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

A Hybrid Anomaly Detection Approach for Obfuscated Malware

With the rapid evolution of malicious software, cyber threats have become increasingly sophisticated, employing advanced obfuscation techniques to evade traditional detection methods. This study presents a hybrid anomaly detection approach applied to obfuscated malware. Even though there is a large body of research in this field, existing malware detection techniques have some drawbacks, such as requiring large amounts of data, trustworthiness (imprecise results) of algorithms, and advanced obfuscation. To overcome these challenges, there is a need to employ solid and efficient techniques for malware detection. This paper proposes a hybrid approach, combining an autoencoder with traditional machine-learning methods to create an efficient malware detection framework. We used the malware memory dataset (MalMemAnalysis-2022) to evaluate this framework. The results indicate that our proposed approach can detect obfuscated malware when a deep autoencoder used for feature learning is combined with logistic regression, and it is extremely fast with an Accuracy, Detection Rate (DR), Matthew Correlation Coefficient(MCC), and Statistical Parity Difference

malware detection, Hybrid Anomly Detection, Obfusc↗

Topological Interpretability for Deep Learning

With the growing adoption of AI-based systems across everyday life, the need to understand their decision-making mechanisms is correspondingly increasing. The level at which we can trust the statistical inferences made from AI-based decision systems is an increasing concern, especially in high-risk systems such as criminal justice or medical diagnosis, where incorrect inferences may have tragic consequences. Despite their successes in providing solutions to problems involving real-world data, deep learning (DL) models cannot quantify the certainty of their predictions. These models are frequently quite confident, even when their solutions are incorrect. This work presents a method to infer prominent features in two DL classification models trained on clinical and non-clinical text by employing techniques from topological and geometric data analysis. We create a graph of a model's feature space and cluster the inputs into the graph's vertices by the similarity of features and prediction statistics. We then extract subgraphs demonstrating high-predictive accuracy for a given label. These subgraphs contain a wealth of information about features that the DL model has recognized as relevant to its decisions. We infer these features for a given label using a distance metric between probability measures, and demonstrate the stability of our method compared to the LIME and SHAP interpretability methods. This work establishes that we may gain insights into the decision mechanism of a DL model. This method allows us to ascertain if the model is making its decisions based on information germane to the problem or identifies extraneous patterns within the data.

Spannaus, Adam↗

Error-Level-Controlled Synthetic Forecasts for Renewable Generation

Renewable energy resources, including solar and wind energy, play a significant role in sustainable energy systems. However, the inherent uncertainty and intermittency of renewable generation pose challenges to the safe and efficient operation of power systems. Recognizing the importance of short-term (hours ahead) renewable generation forecasting in power systems operation, it becomes crucial to address the potential inaccuracies in these forecasts. To systematically evaluate the performance of controllers in the presence of imperfect forecasts, we generate synthetic forecasts using actual renewable generation profiles (one from solar and one from wind). These synthetic forecasts incorporate different levels of statistical error, allowing us to control and manipulate the accuracy of the predictions. The primary objective is to employ synthetic forecasts with controlled yet realistic error levels to systematically investigate how controllers adapt to variations in forecast accuracy, providing valuable insights into their robustness and effectiveness under real-world conditions.

Array↗

Assessment of software methods for estimating protein-protein relative binding affinities

A growing number of computational tools have been developed to accurately and rapidly predict the impact of amino acid mutations on protein-protein relative binding affinities. Such tools have many applications, for example, designing new drugs and studying evolutionary mechanisms. In the search for accuracy, many of these methods employ expensive yet rigorous molecular dynamics simulations. By contrast, non-rigorous methods use less exhaustive statistical mechanics, allowing for more efficient calculations. However, it is unclear if such methods retain enough accuracy to replace rigorous methods in binding affinity calculations. This trade-off between accuracy and computational expense makes it difficult to determine the best method for a particular system or study. Here, eight non-rigorous computational methods were assessed using eight antibody-antigen and eight non-antibody-antigen complexes for their ability to accurately predict relative binding affinities (ΔΔG) for 654 single mutations. In addition to assessing accuracy, we analyzed the CPU cost and performance for each method using a variety of physico-chemical structural features. This allowed us to posit scenarios in which each method may be best utilized. Most methods performed worse when applied to antibody-antigen complexes compared to non-antibody-antigen complexes. Rosetta-based JayZ and EasyE methods classified mutations as destabilizing (ΔΔG < -0.5 kcal/mol) with high (83–98%) accuracy and a relatively low computational cost for non-antibody-antigen complexes. Some of the most accurate results for antibody-antigen systems came from combining molecular dynamics with FoldX with a correlation coefficient (r) of 0.46, but this was also the most computationally expensive method. Overall, our results suggest these methods can be used to quickly and accurately predict stabilizing versus destabilizing mutations but are less accurate at predicting actual binding affinities. This study highlights the need for continued development of reliable, accessible, and reproducible methods for predicting binding affinities in antibody-antigen proteins and provides a recipe for using current methods.

59 BASIC BIOLOGICAL SCIENCES↗

Large Scale Study of Ligand–Protein Relative Binding Free Energy Calculations: Actionable Predictions from Statistically Robust Protocols

The accurate and reliable prediction of protein–ligand binding affinities can play a central role in the drug discovery process as well as in personalized medicine. Of considerable importance during lead optimization are the alchemical free energy methods that furnish an estimation of relative binding free energies (RBFE) of similar molecules. Recent advances in these methods have increased their speed, accuracy, and precision. This is evident from the increasing number of retrospective as well as prospective studies employing them. However, such methods still have limited applicability in real-world scenarios due to a number of important yet unresolved issues. Here, we report the findings from a large data set comprising over 500 ligand transformations spanning over 300 ligands binding to a diverse set of 14 different protein targets which furnish statistically robust results on the accuracy, precision, and reproducibility of RBFE calculations. We use ensemble-based methods which are the only way to provide reliable uncertainty quantification given that the underlying molecular dynamics is chaotic. These are implemented using TIES (Thermodynamic Integration with Enhanced Sampling). Results achieve chemical accuracy in all cases. Ensemble simulations also furnish information on the statistical distributions of the free energy calculations which exhibit non-normal behavior. We find that the “enhanced sampling” method known as replica exchange with solute tempering degrades RBFE predictions. We also report definitively on numerous associated alchemical factors including the choice of ligand charge method, flexibility in ligand structure, and the size of the alchemical region including the number of atoms involved in transforming one ligand into another. Our findings provide a key set of recommendations that should be adopted for the reliable application of RBFE methods.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Stochastic modeling and statistical calibration with model error and scarce data

This paper introduces a procedure to assess the predictive accuracy of stochastic models subject to model error and sparse data. Model error is introduced as uncertainty on the coefficients of appropriate polynomial chaos expansions (PCE). The error associated with finite sample size allows us to conceive of these coefficients as statistics of the data that we describe as random variables whose influence on output quantities of interest is evaluated through the extended polynomial chaos expansion (EPCE). A Bayesian data assimilation scheme is introduced to update these expansions by considering the resulting nested chaos expansion as a hierarchical probabilistic model. Stochastic models of quantities of interest (QoI) are thus constructed and efficiently evaluated. Here, the Metropolis–Hastings Markov chain Monte Carlo procedure is used to sample the posterior. Two illustrative analytical and numerical problems are used to demonstrate the proposed approach.

Bayesian inference↗

Universal energy-speed-accuracy trade-offs in driven nonequilibrium systems

The connection between measure theoretic optimal transport and dissipative nonequilibrium dynamics provides a language for quantifying nonequilibrium control costs, leading to a collection of thermodynamic speed limits, which rely on the assumption that the target probability distribution is perfectly realized. This is almost never the case in experiments or numerical simulations, so here we address the situation in which the external controller is imperfect. We obtain a lower bound for the dissipated work in generic nonequilibrium control problems that (1) is asymptotically tight and (2) matches the thermodynamic speed limit in the case of optimal driving. Along with analytically solvable examples, we refine this imperfect driving notion to systems in which the controlled degrees of freedom are slow relative to the nonequilibrium relaxation rate, and identify independent energy contributions from fast and slow degrees of freedom. Furthermore, we develop a strategy for optimizing minimally dissipative protocols based on optimal transport flow matching, a generative machine learning technique. Furthermore, this latter approach ensures the scalability of both the theoretical and computational framework we put forth. Crucially, we demonstrate that we can compute the terms in our bound numerically using efficient algorithms from the computational optimal transport literature and that the protocols we learn saturate the bound.

59 BASIC BIOLOGICAL SCIENCES↗

Impact of classical statistics on thermal conductivity predictions of BAs and diamond using machine learning molecular dynamics

Machine learning interatomic potentials (MLIPs) have greatly enhanced molecular dynamics (MD) simulations, achieving near-first-principles accuracy in thermal conductivity studies. In this work, we reveal that this accuracy, observed in BAs and diamond at sub-Debye temperatures, stems from an accidental error cancelation: classical statistics overestimates specific heat while underestimating phonon lifetimes, balancing out in thermal conductivity predictions. However, this balance is disrupted when isotopes are introduced, leading MLIP-based MD to significantly underpredict thermal conductivity compared to experiments and quantum statistics-based Boltzmann transport equation. This discrepancy arises not from classical statistics affecting phonon–isotope scattering rates but from its impact on the interplay between phonon–isotope and phonon–phonon scattering in the normal scattering-dominated BAs and diamond. In conclusion, this work underscores the limitations of MLIP-based MD for thermal conductivity studies at sub-Debye temperatures.

36 MATERIALS SCIENCE↗

Examination of probability distribution of mixture fraction in LES/FDF modelling of a turbulent partially premixed jet flame

An accurate prediction of the probability density function (PDF) of the mixture fraction is crucial to the prediction of combustion since mixing plays an important role in turbulent non-premixed and partially premixed flames. This work provides an assessment of the large-eddy simulation (LES)/filtered density function (FDF) method for the prediction of the PDF of the mixture fraction. The advantage of the LES/FDF method is that it provides the full predictions of the statistical distribution of scalars including the mixture fraction. The predictive accuracy of the method for the PDF is yet to be fully validated. Assessing the prediction of the PDF of the mixture fraction, a conserved scalar, is an important starting point. The Sydney/Sandia inhomogeneous inlet jet flame is used as a test case. A quick comparison shows that the LES/FDF predicted PDF shapes of the mixture fraction deviate significantly from the commonly presumed Beta-PDF as well as from the experimental data in the flame. Here, to examine the source of the discrepancy, we clarify the different PDF definitions used in the comparison among the predictions, measurements, and the presumed shape PDFs. The discrepancy observed from the comparison is largely reconciled by clarifying the difference between the PDFs that are examined. The PDF of the resolved mixture fraction is shown to be close to the Beta-PDF in both the measurements and predictions, while the PDF directly deduced from the LES/FDF particles deviates significantly from the Beta-PDF. A multimodal PDF analysis and a pseudo convergence analysis are conducted to provide plausible evidence to support the predicted multimodal PDF shapes. The sub-filter scale FDF is shown to be close to the Beta-PDF too through the construction of a synthesized PDF, which supports the common presumed Beta-PDF assumption used in the presumed PDF methods when combined with LES.

42 ENGINEERING↗

DCTRGAN: improving the precision of generative models with reweighting

Significant advances in deep learning have led to more widely used and precise neural network-based generative models such as Generative Adversarial Networks (Gans). We introduce a post-hoc correction to deep generative models to further improve their fidelity, based on the Deep neural networks using the Classification for Tuning and Reweighting (Dctr) protocol. The correction takes the form of a reweighting function that can be applied to generated examples when making predictions from the simulation. We illustrate this approach using Gans trained on standard multimodal probability densities as well as calorimeter simulations from high energy physics. We show that the weighted Gan examples significantly improve the accuracy of the generated samples without a large loss in statistical power. This approach could be applied to any generative model and is a promising refinement method for high energy physics applications and beyond.

47 OTHER INSTRUMENTATION↗

Data assimilation empowered neural network parametrizations for subgrid processes in geophysical flows

In the past couple of years, there has been a proliferation in the use of machine learning approaches to represent subgrid-scale processes in geophysical flows with an aim to improve the forecasting capability and to accelerate numerical simulations of these flows. Despite its success for different types of flow, the online deployment of a data-driven closure model can cause instabilities and biases in modeling the overall effect of subgrid-scale processes, which in turn leads to inaccurate prediction. To tackle this issue, we exploit the data assimilation technique to correct the physics-based model coupled with the neural network as a surrogate for unresolved flow dynamics in multiscale systems. In particular, we use a set of neural network architectures to learn the correlation between resolved flow variables and the parametrizations of unresolved flow dynamics and formulate a data assimilation approach to correct the hybrid model during their online deployment. We illustrate our framework in a set of applications of the multiscale Lorenz 96 system for which the parametrization model for unresolved scales is exactly known, and the two-dimensional Kraichnan turbulence system for which the parametrization model for unresolved scales is not known a priori. Our analysis, therefore, comprises a predictive dynamical core empowered by (i) a data-driven closure model for subgrid-scale processes, (ii) a data assimilation approach for forecast error correction, and (iii) both data-driven closure and data assimilation procedures. We show significant improvement in the long-term prediction of the underlying chaotic dynamics with our framework compared to using only neural network parametrizations for future prediction. Moreover, we demonstrate that these data-driven parametrization models can handle the non-Gaussian statistics of subgrid-scale processes, and effectively improve the accuracy of outer data assimilation workflow loops in a modular nonintrusive way.

42 ENGINEERING↗

ELG×LRG Distribution through Dark Matter Halo Dynamics

We investigate the clustering and halo occupation distribution (HOD) of DESI Y1 emission-line (ELGs) and luminous red (LRGs) galaxies at 0.8 < z < 1.1, including their cross-correlation (ELG×LRG), using the A BACUS S UMMIT suite and a new Halo Occupation Model (H OME ) for galaxy multitracers. This integrates intrahalo dynamics, halo exclusion, and quenching, bridging insights from hydrodynamical, HOD, abundance-matching, and semianalytic studies. Leveraging full phase-space information from the Uchuu N-body simulation, and sampling satellites from dark-matter particle positions via physically motivated prescriptions, Home reproduces the anisotropic clustering down to s = 200 h −1 kpc with unprecedented accuracy. Model parameters are inferred solely from two-point statistics using a two-level Bayesian framework, yielding high-fidelity ELG, LRG, and cross-reference catalogs. We find that satellite ELGs behave as incoherent flows within their parent halos, dominating the clustering below 4 h −1 Mpc. The HOD from the best-fit Home has the following properties: (i) 90.50% (85.91%) of ELGs (LRGs) are central galaxies without satellites, residing in halos of M vir ∼ 6.6 × 10 11 (1.2 × 10 13 ) h −1 M ⊙ ; (ii) the ELG×LRG cross-correlation is governed by central-central pairs and shaped by halo exclusion on 2–5 h −1 Mpc scales; (iii) 9.50% (14.09%) of ELGs (LRGs) are satellites, of which 1.09% (3.52%) inhabit halos with a central galaxy of the same species in a maximally conformal configuration, 7.02% (0.005%) orbit complementary hosts in a minimally conformal state, and 0.58% (10.57%) are orphans. The high sensitivity of Home precisely captures the dynamics of satellites in different host environments, opening a promising avenue for understanding systematics and the dynamical nature of dark matter, potentially distinguishing gravity models.

Favole, Ginevra [Universidad de La Laguna (Spain);↗