Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “statistical learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12

Near-Real-Time Statistical Analysis and Visualization of Streamflow from a Deep-Learning Rainfall-Runoff Model

Near-real-time (NRT) streamflow data are critical importance for timely water resources management. Here, we developed an open-source tool, FlowStats, for NRT streamflow analysis and visualization in Germany, based on NRT meteorological data from the German Weather Service and simulated streamflow from a long short-term memory neural network (LSTM). The LSTM model achieved very good overall performance, median NSE of 0.80 for the test period across 1,479 catchments. FlowStats provides options for deriving various streamflow statistics, from normal and abnormal streamflow detection to drought and flood analyses. An example analysis from FlowStats revealed widespread below-normal to extreme low-flow conditions across Germany from March to May 2025, which weakened from June to September 2025. Drought analysis for September 2025 highlighted severe to extreme drought conditions in northwestern Germany, while flood classifications indicated that high-flow events occurred in southwestern Germany. FlowStats can be used for various hydrological assessments to support water resources management.

Hydrological modeling↗

Machine Learning–Based Condition Monitoring of a Circulating Water System of a Canadian Nuclear Plant

With the need to maintain long-term reliable energy using nuclear power plants, there is an underlying demand to ensure that the maintenance of plant components and systems is also done in an efficient and cost-effective manner. One way to achieve this is by moving from time-based maintenance to condition-based maintenance. The research presented in this paper focuses on applying statistical and machine-learning-based methods to capture anomalies within data for fault detection to further develop into condition monitoring. This paper focuses on system data for a circulating water system (CWS) of a pressurized heavy-water reactor for detecting anomalies. The different methodologies used for detecting and capturing anomalies in the CWS data are matrix profile, density-based spatial clustering of applications with noise (DBSCAN), and support vector machines (SVMs). Matrix profile and DBSCAN are used to distinguish between normal data and anomalous data. This paper presents a hybrid method using DBSCAN and SVM when a portion of the data is used for DBSCAN to generate clusters. This portion of data is then used to train the SVM along with the clusters generated by DBSCAN as output. SVM is then tested on unseen data as a predictive tool, which can work in real time to categorize data points as either normal or anomalous. This paper presents results that show the high accuracies of DBSCAN and SVM in capturing anomalies within the data for a CWS for fault detection. Thus, the maintenance plan would be focused on component condition rather than a time-based schedule by switching to an automated system to identify and predict faults within a CWS.

21 SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLANTS↗

Emerging Technologies for Privacy Preservation in Energy Systems

This study explores the intersection of digitalization and privacy within the energy sector, focusing on the emerging challenges and opportunities presented by integrating Distributed Energy Resources (DERs) and advanced metering infrastructure. The need for robust digital privacy measures has become crucial as the energy industry evolves towards a more decentralized, digitalized, and decarbonized future. This study delves into four cutting-edge privacy-preserving technologies—Homomorphic Encryption (HE), Secure Multiparty Computation (SMPC), Differential Privacy (DP), and Federated Learning (FL)—each offering unique solutions to safeguard consumer data by increasing digital connectivity and data exchange. Through a detailed examination of these methods, the study explains how each technology operates, its applications within the energy sector, and the specific privacy challenges it addresses. Homomorphic Encryption allows for secure computations on encrypted data, enabling data analysis without compromising privacy. Secure Multiparty Computation enables collaborative data analysis across different entities while protecting the confidentiality of the inputs. Differential Privacy introduces randomness into the assembled data set, preventing the identification of individual records in statistical databases. Lastly, Federated Learning offers a paradigm shift in data analysis, where machine learning models are trained at the edge, minimizing the centralization of sensitive data. The research underscores the significance of implementing these privacy-enhancing technologies to comply with strict data protection regulations, foster consumer trust, and enhance the security of the energy infrastructure. By providing a comprehensive overview of these methodologies and their practical implications for the energy sector, this study aims to contribute to the ongoing discourse on digital privacy, offering insights into how the energy industry can navigate the complexities of data privacy in the digital age.

Cali, Umit↗

Multitask graph neural networks for elastoplastic response prediction in dual-phase polycrystals

Microstructure-sensitive prediction of elastoplastic response remains a recurring bottleneck in multiscale damage and fatigue modeling, where large ensembles of statistically distinct polycrystals are required to quantify variability and extreme-value behavior. In this work, we develop a multitask graph neural network (GNN) surrogate that maps dual-phase ferrite–martensite polycrystal microstructures to Statistical Volume Element (SVE)-level elastoplastic Quantities of Interest (QoIs). Each SVE is represented as a grain-adjacency graph, with node features encoding phase, geometry, and crystallographic orientation, and edge features encoding relative misorientation. A message-passing graph convolution generates node embeddings, which are pooled into a graph representation and passed to a multitask regression head that jointly predicts 10 scalar QoIs and vector-valued stress–strain responses in orthogonal loading directions across multiple martensite volume fractions and SVE sizes. Results show high accuracy for scalar QoIs and strong agreement for full stress–strain trajectories, with population envelopes reproducing both median behavior and finite-SVE variability across compositions and partition scales. A unified model trained on pooled volume-fraction data preserves most within-regime accuracy relative to regime-specific models while also capturing the broader cross-regime variation reflected in the pooled test set. Distributional comparisons further demonstrate that the surrogate preserves heterogeneity under SVE partitioning, enabling statistically consistent block-wise random-field construction for mesoscale analyses. Overall, the proposed grain-graph surrogate provides a practical pathway to accelerate ensemble-based studies of SVE-level constitutive variability in dual-phase polycrystals.

Crystal plasticity↗

Small angle scattering of diblock copolymers profiled by machine learning

We outline a machine learning strategy for quantitively determining the conformation of AB-type diblock copolymers with excluded volume effects using small angle scattering. Complemented by computer simulations, a correlation matrix connecting conformations of different copolymers according to their scattering features is established on the mathematical framework of a Gaussian process, a multivariate extension of the familiar univariate Gaussian distribution. We show that the relevant conformational characteristics of copolymers can be probabilistically inferred from their coherent scattering cross sections without any restriction imposed by model assumptions. This work not only facilitates the quantitative structural analysis of copolymer solutions but also provides the reliable benchmarking for the related theoretical development of scattering functions.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Prediction of hydration energies of adsorbates at Pt(111) and liquid water interfaces using machine learning

Aqueous phase heterogeneous catalysis is important to various industrial processes, including biomass conversion, Fischer–Tropsch synthesis, and electrocatalysis. Accurate calculation of solvation thermodynamic properties is essential for modeling the performance of catalysts for these processes. Explicit solvation methods employing multiscale modeling, e.g., involving density functional theory and molecular dynamics have emerged for this purpose. Although accurate, these methods are computationally intensive. This study introduces machine learning (ML) models to predict solvation thermodynamics for adsorbates on a Pt(111) surface, aiming to enhance computational efficiency without compromising accuracy. In particular, ML models are developed using a combination of molecular descriptors and fingerprints and trained on previously published water–adsorbate interaction energies, energies of solvation, and free energies of solvation of adsorbates bound to Pt(111). These models achieve root mean square error values of 0.09 eV for interaction energies, 0.04 eV for energies of solvation, and 0.06 eV for free energies of solvation, demonstrating accuracy within the standard error of multiscale modeling. Feature importance analysis reveals that hydrogen bonding, van der Waals interactions, and solvent density, together with the properties of the adsorbate, are critical factors influencing solvation thermodynamics. Furthermore, these findings suggest that ML models can provide rapid and reliable predictions of solvation properties. This approach not only reduces computational costs but also offers insights into the solvation characteristics of adsorbates at Pt(111)–water interfaces.

Adsorption↗

Impact of classical statistics on thermal conductivity predictions of BAs and diamond using machine learning molecular dynamics

Machine learning interatomic potentials (MLIPs) have greatly enhanced molecular dynamics (MD) simulations, achieving near-first-principles accuracy in thermal conductivity studies. In this work, we reveal that this accuracy, observed in BAs and diamond at sub-Debye temperatures, stems from an accidental error cancelation: classical statistics overestimates specific heat while underestimating phonon lifetimes, balancing out in thermal conductivity predictions. However, this balance is disrupted when isotopes are introduced, leading MLIP-based MD to significantly underpredict thermal conductivity compared to experiments and quantum statistics-based Boltzmann transport equation. This discrepancy arises not from classical statistics affecting phonon–isotope scattering rates but from its impact on the interplay between phonon–isotope and phonon–phonon scattering in the normal scattering-dominated BAs and diamond. In conclusion, this work underscores the limitations of MLIP-based MD for thermal conductivity studies at sub-Debye temperatures.

36 MATERIALS SCIENCE↗

Evaluation of Extreme Weather Impacts on Utility-scale Photovoltaic Plant Performance in the United States

The global energy system is undergoing significant changes, including a shift in energy generating technologies to more renewable energy sources. However, the dependence of renewable energy sources on local environmental conditions could also increase disruptions in service through exposures to compound, extreme weather events. By fusing three diverse datasets (operations and maintenance tickets, weather data, and production data), this analysis presents a novel methodology to identify and evaluate performance impacts arising from extreme weather events across diverse geographical regions. Text analysis of maintenance tickets identified snow, hurricanes, and storms as the leading extreme weather events affecting photovoltaic plants in the United States. Statistical techniques and machine learning were then implemented to identify the magnitude and variability of these extreme weather impacts on site performance. Impacts varied between event and non-event days, with snow events causing the greatest reductions in performance (54.5%), followed by hurricanes (12.6%) and storms (1.1%). Machine learning analysis identified key features in determining if a day is categorized as low performing, such as low irradiance, geographic location, weather features, and site size. The analysis improves our understanding of compound, extreme weather event impacts on photovoltaic systems, which can inform planning activities, especially as the industry continues to expand into new geographic and climatic regions around the world.

14 SOLAR ENERGY↗

On the transferability of residence time distributions in two 10-km long river sections with similar hydromorphic units

Quantifying hydrologic exchange fluxes (HEFs) at the stream-groundwater interface and their residence time distributions (RTDs) in the subsurface are important for managing the water quality and ecosystem health in dynamic river corridors. However, direct simulating high-spatial resolution HEFs and RTDs can be time-consuming, especially for watershed-scale modeling. Efficient surrogate models linking RTDs to hydromorphic units (HUs) can be alternatives for simulating RTDs in large-scale models. A common concern of these surrogate models, though, is the transferability of the relationship between the RTDs and HUs from one river corridor to another. To address this issue, this work evaluates the HEFs and resulting RTD-HU relationships for two 10-km long river corridors along the Columbia River leveraging a one-way coupled three-dimensional transient surface-subsurface water transport modeling framework we previously developed. Applying such a framework at the two river corridors with similar HUs allows for quantitative comparisons of HEFs and RTDs using both statistical tests and machine learning classification models. Finally, our comparison shows that the similarity and transferability of the RTD-HU relationship is very low for the two investigated river sections, which suggests that devising a general algorithm to estimate RTDs based solely on surface water hydrodynamics and short-distance river channel topography data, as well as HU classification, might be nearly impossible.

54 ENVIRONMENTAL SCIENCES↗

Deep neural network improves the estimation of polygenic risk scores for breast cancer

Polygenic risk scores (PRS) estimate the genetic risk of an individual for a complex disease based on many genetic variants across the whole genome. Here, we compared a series of computational models for estimation of breast cancer PRS. A deep neural network (DNN) was found to outperform alternative machine learning techniques and established statistical algorithms, including BLUP, BayesA, and LDpred. In the test cohort with 50% prevalence, the Area Under the receiver operating characteristic Curve (AUC) were 67.4% for DNN, 64.2% for BLUP, 64.5% for BayesA, and 62.4% for LDpred. BLUP, BayesA, and LPpred all generated PRS that followed a normal distribution in the case population. However, the PRS generated by DNN in the case population followed a bimodal distribution composed of two normal distributions with distinctly different means. This suggests that DNN was able to separate the case population into a high-genetic-risk case subpopulation with an average PRS significantly higher than the control population and a normal-genetic-risk case subpopulation with an average PRS similar to the control population. This allowed DNN to achieve 18.8% recall at 90% precision in the test cohort with 50% prevalence, which can be extrapolated to 65.4% recall at 20% precision in a general population with 12% prevalence. Interpretation of the DNN model identified salient variants that were assigned insignificant p values by association studies, but were important for DNN prediction. These variants may be associated with the phenotype through nonlinear relationships.

59 BASIC BIOLOGICAL SCIENCES↗

Multidimensional quantitative phenotypic and molecular analysis reveals neomorphic behaviors of p53 missense mutants

Abstract Mutations in the TP53 tumor suppressor gene occur in >80% of the triple-negative or basal-like breast cancer. To test whether neomorphic functions of specific TP53 missense mutations contribute to phenotypic heterogeneity, we characterized phenotypes of non-transformed MCF10A-derived cell lines expressing the ten most common missense mutant p53 proteins and observed a wide spectrum of phenotypic changes in cell survival, resistance to apoptosis and anoikis, cell migration, invasion and 3D mammosphere architecture. The p53 mutants R248W, R273C, R248Q, and Y220C are the most aggressive while G245S and Y234C are the least, which correlates with survival rates of basal-like breast cancer patients. Interestingly, a crucial amino acid difference at one position—R273C vs. R273H—has drastic changes on cellular phenotype. RNA-Seq and ChIP-Seq analyses show distinct DNA binding properties of different p53 mutants, yielding heterogeneous transcriptomics profiles, and MD simulation provided structural basis of differential DNA binding of different p53 mutants. Integrative statistical and machine-learning-based pathway analysis on gene expression profiles with phenotype vectors across the mutant cell lines identifies quantitative association of multiple pathways including the Hippo/YAP/TAZ pathway with phenotypic aggressiveness. Further, comparative analyses of large transcriptomics datasets on breast cancer cell lines and tumors suggest that dysregulation of the Hippo/YAP/TAZ pathway plays a key role in driving the cellular phenotypes towards basal-like in the presence of more aggressive p53 mutants. Overall, our study describes distinct gain-of-function impacts on protein functions, transcriptional profiles, and cellular behaviors of different p53 missense mutants, which contribute to clinical phenotypic heterogeneity of triple-negative breast tumors.

60 APPLIED LIFE SCIENCES↗

Scaling and merging time-resolved pink-beam diffraction with variational inference

Time-resolved x-ray crystallography (TR-X) at synchrotrons and free electron lasers is a promising technique for recording dynamics of molecules at atomic resolution. While experimental methods for TR-X have proliferated and matured, data analysis is often difficult. Extracting small, time-dependent changes in signal is frequently a bottleneck for practitioners. Recent work demonstrated this challenge can be addressed when merging redundant observations by a statistical technique known as variational inference (VI). However, the variational approach to time-resolved data analysis requires identification of successful hyperparameters in order to optimally extract signal. In this case study, we present a successful application of VI to time-resolved changes in an enzyme, DJ-1, upon mixing with a substrate molecule, methylglyoxal. We present a strategy to extract high signal-to-noise changes in electron density from these data. Furthermore, we conduct an ablation study, in which we systematically remove one hyperparameter at a time to demonstrate the impact of each hyperparameter choice on the success of our model. We expect this case study will serve as a practical example for how others may deploy VI in order to analyze their time-resolved diffraction data.

47 OTHER INSTRUMENTATION↗

Signatures of a liquid–liquid transition in an ab initio deep neural network model for water

Significance Water is central across much of the physical and biological sciences and exhibits physical properties that are qualitatively distinct from those of most other liquids. Understanding the microscopic basis of water’s peculiar properties remains an active area of research. One intriguing hypothesis is that liquid water can separate into metastable high- and low-density liquid phases at low temperatures and high pressures, and the existence of this liquid–liquid transition could explain many of water’s anomalous properties. We used state-of-the-art approaches in computational quantum chemistry, statistical mechanics, and machine learning and obtained evidence consistent with a liquid–liquid transition, supporting the argument for the existence of this phenomenon in real water.

36 MATERIALS SCIENCE↗

Noisy intermediate-scale quantum algorithms

A universal fault-tolerant quantum computer that can efficiently solve problems such as integer factorization and unstructured database search requires millions of qubits with low error rates and long coherence times. While the experimental advancement toward realizing such devices will potentially take decades of research, noisy intermediate-scale quantum (NISQ) computers already exist. These computers are composed of hundreds of noisy qubits, i.e., qubits that are not error corrected, and therefore perform imperfect operations within a limited coherence time. In the search for achieving quantum advantage with these devices, algorithms have been proposed for applications in various disciplines spanning physics, machine learning, quantum chemistry, and combinatorial optimization. The overarching goal of such algorithms is to leverage the limited available resources to perform classically challenging tasks. In this review, a thorough summary of NISQ computational paradigms and algorithms is provided. The key structure of these algorithms and their limitations and advantages are discussed. Finally, a comprehensive overview of various benchmarking and software tools useful for programming and testing NISQ devices is additionally provided.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Crash Risks Evaluation of Urban Expressways: A Case Study in Shanghai

We report that proactive traffic safety management systems can reduce crashes by identifying crash precursors, evaluating real-time crash risks, and implementing suitable interventions. The basic prerequisite for developing such a system is to propose a reliable crash risk evaluation model that takes real-time traffic flow data as input. Previous studies have primarily focused on real-time crash prediction using some statistical or machine-learning methods. However, further quantitative evaluation and classification of crash risks have been ignored. In this study, we conduct a systematic crash risk evaluation workflow, including crash risk prediction, crash risk quantification, and crash risk classification. Specifically, the crash risk prediction using an extended logit model is proposed, from which CAS, CSD, UAS, DAS, DTV are identified to be contributing factors of crash risks. Then a crash risk quantification model based on the parameter evaluation of the extended logit model is developed. The crash risks of urban expressways and their spatial-temporal evolution trends are quantified. Finally, the crash risks are classified into high crash risk level, moderate crash risk level, and low crash risk level by the k-means cluster algorithm. Then the threshold boundaries of different crash risk levels are determined. The research results provide a proactive guidance for traffic safety management of urban expressways.

33 ADVANCED PROPULSION SYSTEMS↗

Rapid Electrochemical Diagnosis of Battery Health and Safety from Cells to Modules

Rapid electrochemical diagnosis of battery health and failure is critical for ensuring reliable battery performance and battery safety. Traditional battery health diagnostics such as capacity measurements and DC pulse tests are reliable and well-understood, however, these measurements of battery capacity and resistance do not capture all aspects of battery degradation. Other aspects of degradation, such as electrolyte decomposition, lithium-plating, and particle cracking are difficult to detect electrochemically but are crucial to measure to get a full picture of battery safety and flag out potential failures. In this work, lab- and field-aged commercial lithium-ion batteries and modules of various chemistries and formats are tested using a variety of traditional electrochemical characterization methods as well as using 2-minute pseudo-random DC pulse sequences at rest and during charge/discharge. The electrochemical measurements are compared to physical cell measurements, cell efficiency, drive cycle performance, physical and thermal heterogeneity, and qualitative safety metrics using statistical and machine-learning methods to discover if a comprehensive "battery health map" can be accurately identified using only rapid DC measurements.

ADVANCED PROPULSION SYSTEMS,ENERGY STORAGE↗