Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “statistical learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 289 records · Page 16

Cluster characterization in atom probe tomography: Machine learning using multiple summary functions

In this work, we develop a machine learning-based method to characterize intracluster concentration (ρ c ), background concentration (ρ b ), clustering radius (r̄), and radius dispersity (δ r ) in simulated atom probe tomography data using multiple spatial statistics summary functions to train a Bayesian regularized neural network. Here, we build upon previous work that utilized Ripley’s K-function by incorporating additional features from nearest-neighbor spatial statistics summary functions to better characterize concentration-based metrics. The addition of nearest-neighbor based features allows for highly accurate estimates of ρ c and ρ b , both with 90% of the predictions within 4.0% of the real value; the root-mean-square errors are reduced by 81.5% and 92.8% from predictions using only K-function based features, respectively. Additionally, including these nearest-neighbor based features improves the ability to differentiate between r̄ and δ r .

36 MATERIALS SCIENCE↗

Deep Learning Parameterization of Vertical Wind Velocity Variability via Constrained Adversarial Training

Atmospheric models with typical resolution in the tenths of kilometers cannot resolve the dynamics of air parcel ascent, which varies on scales ranging from tens to hundreds of meters. Small-scale wind fluctuations are thus characterized by a subgrid distribution of vertical wind velocity W with standard deviation σ W . The parameterization of σ W is fundamental to the representation of aerosol–cloud interactions, yet it is poorly constrained. Using a novel deep learning technique, this work develops a new parameterization for σ W merging data from global storm-resolving model simulations, high-frequency retrievals of W , and climate reanalysis products. The parameterization reproduces the observed statistics of σ W and leverages learned physical relations from the model simulations to guide extrapolation beyond the observed domain. Incorporating observational data during the training phase was found to be critical for its performance. The parameterization can be applied online within large-scale atmospheric models, or offline using output from weather forecasting and reanalysis products.

54 ENVIRONMENTAL SCIENCES↗

Machine Learning-Based Anomaly Detection for PMT Data Quality Monitoring in the SBN and DUNE

Maintaining high-quality detector data is essential for achieving the scientific objectives of the Short-Baseline Neutrino (SBN) Program at Fermilab. Current data quality monitoring (DQM) procedures rely primarily on threshold-based metrics and manual inspection of detector monitoring plots, making the detection of subtle or gradually developing anomalies both time-consuming and dependent on expert interpretation. This project developed and evaluated a machine-learning workflow for automatically identifying anomalous photomultiplier tube (PMT) channels in the Short-Baseline Near Detector (SBND) using optical-hit amplitude data. A Python-based analysis program was developed to process ROOT files, extract statistical features describing individual PMT amplitude distributions, and generate feature vectors for anomaly detection. These features were used to train an Isolation Forest model using data representing normal detector operation. The trained model was subsequently applied to independent detector runs to identify channels exhibiting statistically unusual behavior relative to the learned reference response. To support expert interpretation, the workflow generated complementary diagnostic products, including anomaly score distributions, normalized amplitude comparisons, decision-tree visualizations, and principal component analysis (PCA) projections. This project demonstrated the feasibility of integrating unsupervised machine learning into detector data-quality monitoring and developed a complete workflow for automated PMT performance assessment to aid expert-driven review. Beyond its technical contributions, the VFP appointment fostered a research collaboration between Aurora University and Fermilab and provided direct workforce development benefits by training the visiting faculty member in detector-scale machine-learning methods that are now being incorporated into undergraduate coursework and research. The methodology developed here provides a foundation for future applications to ProtoDUNE and other liquid argon time projection chamber (LArTPC) detectors, contributing to ongoing efforts to improve detector reliability, reduce manual monitoring requirements, and enable scalable data quality monitoring for future large-scale neutrino experiments, including the Deep Underground Neutrino Experiment (DUNE).

Colón Santana, Juan A. [Unlisted, US, IL]↗

Machine Learning to Select Experiments Driven by Fundamental Science and Applications for Targeted Nuclear Data Improvement

This work describes a blueprint for a process that accelerates progress in science by quantitatively answering the following question: What is the optimal combination of fundamental-science and application-driven experiments to maximally reduce pertinent data uncertainties? Answering this question entails solving a high-dimensional and complex optimization problem that is best solved with advanced statistic techniques often classified as machine learning. We apply this process within the framework of nuclear data with the aim to select an experiment combination that will reduce uncertainties in 239 Pu nuclear data for neutron energies between 1 and 600 keV. In this field, fundamental-physics driven data, called differential, look at one nuclear physics observable at a time. They are contrasted to application-driven, integral, data where one or few resulting values inform a broad set of nuclear data across several nuclides and energies. The candidates for integral experiments are criticality measurements that were refined by a genetic algorithm to be maximally sensitive to 239 Pu fission cross sections in the desired energy range. Twenty-three candidate differential experiments were investigated and span multiple nuclear physics observables (e.g., total, capture cross sections) for isotopes appearing in the integral experiments. The optimal combination among these candidate experiments was investigated via generalized least squares fitting, augmented with Gaussian processes to ameliorate statistical irregularities in data, and the D-optimality criterion. The latter evaluates for each pair of candidates the joint reduction in uncertainties of all 12200 nuclear data appearing in the integral experiments compared to the knowledge we have from 168 past experiments, theory, and nuclear data. We chose as differential measurements those that investigate 63 Cu and 239 Pu total cross sections, based on D-optimality rank and feasibility constraints. Two integral (criticality) experiments were selected: An experiment with Al 2 ⁢O 3 and graphite interleaved with Pu and a thick Cu reflector explores 1–30 keV, while we target the 30–600 keV range with an experiment that swaps boron in place of graphite with a different geometry.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

General-Purpose Unsupervised Cyber Anomaly Detection via Non-Negative Tensor Factorization

Distinguishing malicious anomalous activities from unusual but benign activities is a fundamental challenge for cyber defenders. Prior studies have shown that statistical user behavior analysis yields accurate detections by learning behavior profiles from observed user activity. These unsupervised models are able to generalize to unseen types of attacks by detecting deviations from normal behavior, without knowledge of specific attack signatures. However, approaches proposed to date based on probabilistic matrix factorization are limited by the information conveyed in a two-dimensional space. Non-negative tensor factorization, on the other hand, is a powerful unsupervised machine learning method that naturally models multi-dimensional data, capturing complex and multi-faceted details of behavior profiles. Herein, our new unsupervised statistical anomaly detection methodology matches or surpasses state-of-the-art supervised learning baselines across several challenging and diverse cyber application areas, including detection of compromised user credentials, botnets, spam e-mails, and fraudulent credit card transactions.

97 MATHEMATICS AND COMPUTING↗

Applications of flow models to the generation of correlated lattice QCD ensembles

Machine-learned normalizing flows can be used in the context of lattice quantum field theory to generate statistically correlated ensembles of lattice gauge fields at different action parameters. This work demonstrates how these correlations can be exploited for variance reduction in the computation of observables. Three different proof-of-concept applications are demonstrated using a novel residual flow architecture: continuum limits of gauge theories, the mass dependence of QCD observables, and hadronic matrix elements based on the Feynman–Hellmann approach. In all three cases, it is shown that statistical uncertainties are significantly reduced when machine-learned flows are incorporated as compared with the same calculations performed with uncorrelated ensembles or direct reweighting. Published by the American Physical Society 2024

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Fouling modeling and prediction approach for heat exchangers using deep learning

In this article, we develop a generalized and scalable statistical model for accurate prediction of fouling resistance using commonly measured parameters of industrial heat exchangers. This prediction model is based on deep learning where a scalable algorithmic architecture learns non-linear functional relationships between a set of target and predictor variables from large number of training samples. Here, the efficacy of this modeling approach is demonstrated for predicting fouling in an analytically modeled cross-flow heat exchanger, designed for waste heat recovery from flue-gas using room temperature water. The performance results of the trained models demonstrate that the mean absolute prediction errors are under 10 –4 KW –1 for flue-gas side, water side and overall fouling resistances. The coefficients of determination (R 2 ), which characterize the goodness of fit between the predictions and observed data, are over 99%. Even under varying levels of measurement noise in the inputs, we demonstrate that predictions over an ensemble of multiple neural networks achieves better accuracy and robustness to noise. We find that the proposed deep-learning fouling prediction framework learns to follow heat exchanger flow and heat transfer physics, which we confirm using locally interpretable model agnostic explanations around randomly selected operating points. Overall, we provide a robust algorithmic framework for fouling prediction that can be generalized and scaled to various types of industrial heat exchangers.

42 ENGINEERING↗

Leveraging interpolation models and error bounds for verifiable scientific machine learning

Effective verification and validation techniques for modern scientific machine learning workflows are challenging to devise. Statistical methods are abundant and easily deployed, but often rely on speculative assumptions about the data and methods involved. Error bounds for classical interpolation techniques can provide mathematically rigorous estimates of accuracy, but often are difficult or impractical to determine computationally. Here, in this work, we present a best-of-both-worlds approach to verifiable scientific machine learning by demonstrating that (1) multiple standard interpolation techniques have informative error bounds that can be computed or estimated efficiently; (2) comparative performance among distinct interpolants can aid in validation goals; (3) deploying interpolation methods on latent spaces generated by deep learning techniques enables some interpretability for black-box models. We present a detailed case study of our approach for predicting lift-drag ratios from airfoil images. Code developed for this work is available in a public Github repository.

97 MATHEMATICS AND COMPUTING↗

Joint Estimation of Topology and Injection Statistics in Distribution Grids with Missing Nodes

Optimal operation of distribution grid resources relies on accurate estimation of its state and topology. Practical estimation of such quantities is complicated by the limited presence of real-time meters. This article discusses a theoretical framework to jointly estimate the operational topology and statistics of injections in radial distribution grids under limited availability of nodal voltage measurements. In particular, we show that our proposed algorithms are able to provably learn the exact grid topology and injection statistics at all unobserved nodes as long as they are not adjacent. The algorithm design is based on novel ordered trends in voltage magnitude fluctuations at node groups, that are independently of interest for radial physical flow networks. The complexity of the designed algorithms is theoretically analyzed and their performance is validated using both linearized and nonlinear ac power flow samples in test distribution grids.

97 MATHEMATICS AND COMPUTING↗

Machine learning BPS spectra and the gap conjecture

We explore statistical properties of Bogomol’nyi-Prasad-Sommerfield q-series for strongly coupled supersymmetric theories that correspond to a particular family of three-manifolds. We discover that gaps between exponents in the -series are statistically more significant at the beginning of the -series compared to gaps that appear in higher powers of. Our observations are obtained by calculating saliencies of -series features used as input data for principal component analysis, which is a standard example of an explainable machine learning technique that allows for a direct calculation and a better analysis of feature saliencies.

97 MATHEMATICS AND COMPUTING↗

Implementation of disruptive designs for gas turbine components using direct energy deposition additive manufacturing

This research aims to develop a framework for establishing the correlation between in-situ monitoring data, process parameters, and microstructure evolution in blown-powder laser-directed energy deposition (DED) additive manufacturing (AM). To achieve this, a comprehensive manufacturing framework has been developed, spanning from in-situ data acquisition, melt-pool simulation, microstructure modeling, and statistical microstructure quantification. A machine learning-based surrogate model is constructed to predict melt pool geometry directly from in-situ coaxial camera data. The surrogate model is trained using outputs from a high-fidelity melt pool simulation, which provides accurate melt pool dimension data under varying process conditions. The predicted melt pool geometry is then used as input to a microstructure model to predict microstructural features. To rigorously compare and analyze microstructures, the project introduces statistical metrics that quantify differences based on key features such as morphology and texture. Microstructures are represented using advanced statistical descriptors including angular chord length distribution, two-point spatial statistics, orientation distribution function, and global spherical harmonic. These representations are used to compute four distinct “dissimilarity scores” that quantitatively capture differences in texture and morphology. This framework is demonstrated to enable automated calibration of simulation parameters by minimizing discrepancies between simulated and target microstructures. The technology developed in this project enables direct correlation between in-situ monitoring data and resulting microstructure, paving the way for adaptive microstructure control in metal AM. This capability strengthens the connection between process parameters and final material properties, facilitating more precise and reliable material design.

36 MATERIALS SCIENCE↗

Inverse methods for design of soft materials

Functional soft materials, comprising colloidal and molecular building blocks that self-organize into complex structures as a result of their tunable interactions, enable a wide array of technological applications. Inverse methods provide a systematic means for navigating their inherently high-dimensional design spaces to create materials with targeted properties. Furthermore, while multiple physically motivated inverse strategies have been successfully implemented in silico, their translation to guiding experimental materials discovery has thus far been limited to a handful of proof-of-concept studies. In this perspective, we discuss recent advances in inverse methods for design of soft materials that address two challenges: (1) methodological limitations that prevent such approaches from satisfying design constraints and (2) computational challenges that limit the size and complexity of systems that can be addressed. Strategies that leverage machine learning have proven particularly effective, including methods to discover order parameters that characterize complex structural motifs and schemes to efficiently compute macroscopic properties from the underlying structure. We also highlight promising opportunities to improve the experimental realizability of materials designed computationally, including discovery of materials with functionality at multiple thermodynamic states, design of externally directed assembly protocols that are simple to implement in experiments, and strategies to improve the accuracy and computational efficiency of experimentally relevant models.

36 MATERIALS SCIENCE↗

Automated detector simulation and reconstruction parametrization using machine learning

Rapidly applying the effects of detector response to physics objects (e.g. electrons, muons, showers of particles) is essential in high energy physics. Presently available tools for the transformation from truth-level physics objects to reconstructed detector-level physics objects involve manually defining resolution functions. These resolution functions are typically derived in bins of variables that are correlated with the resolution (e.g. pseudorapidity and transverse momentum). This process is time consuming, requires manual updates when detector conditions change, and can miss important correlations. Machine learning offers a way to automate the process of building these truth-to-reconstructed object transformations and can capture complex correlation for any given set of input variables. Such machine learning algorithms, with sufficient optimization, could have a wide range of applications: improving phenomenological studies by using a better detector representation, allowing for more efficient production of Geant4 simulation by only simulating events within an interesting part of phase space, and studies on future experimental sensitivity to new physics.

47 OTHER INSTRUMENTATION↗

DESI mock challenge: Halo and galaxy catalogues with the bias assignment method

We present a novel approach to the construction of mock galaxy catalogues for large-scale structure analysis based on the distribution of dark matter halos obtained with effective bias models at the field level. We aim to produce mock galaxy catalogues capable of generating accurate covariance matrices for a number of cosmological probes that are expected to be measured in current and forthcoming galaxy redshift surveys (e.g. two- and three-point statistics). The construction of the catalogues shown in this paper is part of a mock-comparison project within the Dark Energy Spectroscopic Instrument (DESI) collaboration. We use the bias assignment method ( BAM ) to model the statistics of halo distribution through a learning algorithm using a few detailed N-body simulations, and approximated gravity solvers based on Lagrangian perturbation theory. We introduce cosmic-web-dependent corrections to modelling redshift-space distortions at the N-body level – both in the halo and galaxy distributions –, as well as a multi-scale approach for accurate assignment of halo properties. Using specific models of halo occupation distributions to populate halos, we generate galaxy mocks with the expected number density and central-satellite fraction of emission-line galaxies, which are a key target of the DESI experiment. BAM generates mock catalogues with per cent accuracy in a number of summary statistics, such as the abundance, the two- and three-point statistics of halo distributions, both in real and redshift space. In particular, the mock galaxy catalogues display ~3%-10% accuracy in the multipoles of the power spectrum up to scales of k ~ 0.4 h -1 Mpc. We show that covariance matrices of two- and three-point statistics obtained with BAM display a similar structure to the reference simulation. BAM offers an efficient way to produce mock halo catalogues with accurate two- and three-point statistics and is able to generate a variety of multi-tracer catalogues with precise covariance matrices of several cosmological probes. We discuss future developments of the algorithm towards mock production in DESI and other galaxy-redshift surveys.

79 ASTRONOMY AND ASTROPHYSICS↗

Stream Temperature Prediction in a Shifting Environment: Explaining the Influence of Deep Learning Architecture

Stream temperature is a fundamental control on ecosystem health. Recent efforts incorporating process guidance into deep learning models for predicting stream temperature have been shown to outperform existing statistical and physical models. This performance is in part because deep learning architectures can actively learn spatiotemporal relationships that govern how water and energy propagate through a river network. However, exploration of how spatiotemporal awareness and process guidance influence a model's generalizability under shifting environmental conditions such as climate change is limited. Here, we use Explainable Artificial Intelligence (XAI) to interrogate how differing deep learning architectures affect a model's learned spatial and temporal dependencies, and how those learned dependencies affect a model's ability to maintain high accuracy when applied to unseen environmental conditions. Using the Delaware River Basin in the northeastern United States as a test case, we compare two spatiotemporally aware process–guided deep learning models for predicting stream temperature (a recurrent graph convolution network—RGCN, and a temporal convolution graph model—Graph WaveNet). Both models achieve equally high predictive performance when testing data are well represented in the training data (test root mean squared errors of 1.64°C and 1.65°C); however, Graph WaveNet significantly outperforms RGCN in 4 out of 5 experiments where test partitions represent different types of unseen environmental conditions. XAI results show that the architecture of Graph WaveNet leads to learned spatial relationships with greater fidelity to physical processes, and that this fidelity improves the generalizability of the model when applied to shifting and/or unseen environmental conditions.

54 ENVIRONMENTAL SCIENCES↗

Supervised and unsupervised machine learning of structural phases of polymers adsorbed to nanowires

Here, we identify configurational phases and structural transitions in a polymer nanotube composite by means of machine learning. We employ various unsupervised dimensionality reduction methods, conventional neural networks, as well as the confusion method, an unsupervised neural-network-based approach. We find neural networks are able to reliably recognize all configurational phases that have been found previously in experiment and simulation. Furthermore, we locate the boundaries between configurational phases in a way that removes human intuition or bias. This could be done before only by relying on preconceived, ad hoc order parameters.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Adaptive nonequilibrium design of actin-based metamaterials: Fundamental and practical limits of control

The adaptive and surprising emergent properties of biological materials self-assembled in far-from-equilibrium environments serve as an inspiration for efforts to design nanomaterials. In particular, controlling the conditions of self-assembly can modulate material properties, but there is no systematic understanding of either how to parameterize external control or how controllable a given material can be. Here, we demonstrate that branched actin networks can be encoded with metamaterial properties by dynamically controlling the applied force under which they grow and that the protocols can be selected using multi-task reinforcement learning. These actin networks have tunable responses over a large dynamic range depending on the chosen external protocol, providing a pathway to encoding “memory” within these structures. Interestingly, we obtain a bound that relates the dissipation rate and the rate of “encoding” that gives insight into the constraints on control—both physical and information theoretical. Taken together, these results emphasize the utility and necessity of nonequilibrium control for designing self-assembled nanostructures.

59 BASIC BIOLOGICAL SCIENCES↗

Smart quantum statistical imaging beyond the Abbe-Rayleigh criterion

The wave nature of light imposes limits on the resolution of optical imaging systems. For over a century, the Abbe-Rayleigh criterion has been utilized to assess the spatial resolution limits of imaging instruments. Recently, there has been interest in using spatial projective measurements to enhance the resolution of imaging systems. Unfortunately, these schemes require a priori information regarding the coherence properties of “unknown” light beams and impose stringent alignment conditions. Here, we introduce a smart quantum camera for superresolving imaging that exploits the self-learning features of artificial intelligence to identify the statistical fluctuations of unknown mixtures of light sources at each pixel. This is achieved through a universal quantum model that enables the design of artificial neural networks for the identification of photon fluctuations. Our protocol overcomes limitations of existing superresolution schemes based on spatial mode projections, and consequently provides alternative methods for microscopy, remote sensing, and astronomy.

77 NANOSCIENCE AND NANOTECHNOLOGY↗