Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Statistical Learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

A Framework for the Analysis of Deep Neural Networks in Autonomous Aerospace Applications using Bayesian Statistics

Deep Neural Networks (DNNs) are considered to be key components in many autonomous systems. Applications range from vision-based obstacle avoidance to intelligent/learning control and planning. Safety-critical applications as found in the aerospace domain require that the behavior of the DNN is validated and tested rigorously for safety of the autonomous system (AUS). In this paper, we present a framework to support testing of DNNs and the analysis of the network structure. Our framework employs techniques from statistical modeling and active learning to effectively generate test cases for DNN safety testing and performance analysis. We will present results of a case study on a physics-based Deep recurrent residual neural network (DR-RNN), which has been trained to emulate the aerodynamics behavior of a fixed-wing aircraft.

Deep Neural networks↗

Emerging Technologies for Privacy Preservation in Energy Systems

This study explores the intersection of digitalization and privacy within the energy sector, focusing on the emerging challenges and opportunities presented by integrating Distributed Energy Resources (DERs) and advanced metering infrastructure. The need for robust digital privacy measures has become crucial as the energy industry evolves towards a more decentralized, digitalized, and decarbonized future. This study delves into four cutting-edge privacy-preserving technologies—Homomorphic Encryption (HE), Secure Multiparty Computation (SMPC), Differential Privacy (DP), and Federated Learning (FL)—each offering unique solutions to safeguard consumer data by increasing digital connectivity and data exchange. Through a detailed examination of these methods, the study explains how each technology operates, its applications within the energy sector, and the specific privacy challenges it addresses. Homomorphic Encryption allows for secure computations on encrypted data, enabling data analysis without compromising privacy. Secure Multiparty Computation enables collaborative data analysis across different entities while protecting the confidentiality of the inputs. Differential Privacy introduces randomness into the assembled data set, preventing the identification of individual records in statistical databases. Lastly, Federated Learning offers a paradigm shift in data analysis, where machine learning models are trained at the edge, minimizing the centralization of sensitive data. The research underscores the significance of implementing these privacy-enhancing technologies to comply with strict data protection regulations, foster consumer trust, and enhance the security of the energy infrastructure. By providing a comprehensive overview of these methodologies and their practical implications for the energy sector, this study aims to contribute to the ongoing discourse on digital privacy, offering insights into how the energy industry can navigate the complexities of data privacy in the digital age.

Cali, Umit↗

Multitask graph neural networks for elastoplastic response prediction in dual-phase polycrystals

Microstructure-sensitive prediction of elastoplastic response remains a recurring bottleneck in multiscale damage and fatigue modeling, where large ensembles of statistically distinct polycrystals are required to quantify variability and extreme-value behavior. In this work, we develop a multitask graph neural network (GNN) surrogate that maps dual-phase ferrite–martensite polycrystal microstructures to Statistical Volume Element (SVE)-level elastoplastic Quantities of Interest (QoIs). Each SVE is represented as a grain-adjacency graph, with node features encoding phase, geometry, and crystallographic orientation, and edge features encoding relative misorientation. A message-passing graph convolution generates node embeddings, which are pooled into a graph representation and passed to a multitask regression head that jointly predicts 10 scalar QoIs and vector-valued stress–strain responses in orthogonal loading directions across multiple martensite volume fractions and SVE sizes. Results show high accuracy for scalar QoIs and strong agreement for full stress–strain trajectories, with population envelopes reproducing both median behavior and finite-SVE variability across compositions and partition scales. A unified model trained on pooled volume-fraction data preserves most within-regime accuracy relative to regime-specific models while also capturing the broader cross-regime variation reflected in the pooled test set. Distributional comparisons further demonstrate that the surrogate preserves heterogeneity under SVE partitioning, enabling statistically consistent block-wise random-field construction for mesoscale analyses. Overall, the proposed grain-graph surrogate provides a practical pathway to accelerate ensemble-based studies of SVE-level constitutive variability in dual-phase polycrystals.

Crystal plasticity↗

Cloud Vertical and Horizontal Structure from ICESat/GLAS and MODIS

To accurately model radiative fluxes at the surface and within the atmosphere, we need to know both vertical and horizontal structures of cloudiness. While MODIS provides accurate information on cloud horizontal structure, it has limited ability to estimate cloud vertical structure. ICESat/GLAS on the other hand, provides the vertical distribution and internal structure of clouds as deep as the laser beam can penetrate and return a signal. Having different orbits, MODIS and GLAS provide few collocated measurements; hence a statistical approach is needed to learn about 3D cloud structures from the two instruments. In the presentation, we show the results of the statistical analysis of vertical and horizontal structure of cloudiness using GLAS and MODIS cloud top(s) data acquired in October-November 2003. We revisit the (H1, C1) plot, previously used for analyzing cloud liquid water data, and illustrate cloud structure for single and multiple-layer clouds.

Marshak, Alexander↗

Prediction of hydration energies of adsorbates at Pt(111) and liquid water interfaces using machine learning

Aqueous phase heterogeneous catalysis is important to various industrial processes, including biomass conversion, Fischer–Tropsch synthesis, and electrocatalysis. Accurate calculation of solvation thermodynamic properties is essential for modeling the performance of catalysts for these processes. Explicit solvation methods employing multiscale modeling, e.g., involving density functional theory and molecular dynamics have emerged for this purpose. Although accurate, these methods are computationally intensive. This study introduces machine learning (ML) models to predict solvation thermodynamics for adsorbates on a Pt(111) surface, aiming to enhance computational efficiency without compromising accuracy. In particular, ML models are developed using a combination of molecular descriptors and fingerprints and trained on previously published water–adsorbate interaction energies, energies of solvation, and free energies of solvation of adsorbates bound to Pt(111). These models achieve root mean square error values of 0.09 eV for interaction energies, 0.04 eV for energies of solvation, and 0.06 eV for free energies of solvation, demonstrating accuracy within the standard error of multiscale modeling. Feature importance analysis reveals that hydrogen bonding, van der Waals interactions, and solvent density, together with the properties of the adsorbate, are critical factors influencing solvation thermodynamics. Furthermore, these findings suggest that ML models can provide rapid and reliable predictions of solvation properties. This approach not only reduces computational costs but also offers insights into the solvation characteristics of adsorbates at Pt(111)–water interfaces.

Adsorption↗

Impact of classical statistics on thermal conductivity predictions of BAs and diamond using machine learning molecular dynamics

Machine learning interatomic potentials (MLIPs) have greatly enhanced molecular dynamics (MD) simulations, achieving near-first-principles accuracy in thermal conductivity studies. In this work, we reveal that this accuracy, observed in BAs and diamond at sub-Debye temperatures, stems from an accidental error cancelation: classical statistics overestimates specific heat while underestimating phonon lifetimes, balancing out in thermal conductivity predictions. However, this balance is disrupted when isotopes are introduced, leading MLIP-based MD to significantly underpredict thermal conductivity compared to experiments and quantum statistics-based Boltzmann transport equation. This discrepancy arises not from classical statistics affecting phonon–isotope scattering rates but from its impact on the interplay between phonon–isotope and phonon–phonon scattering in the normal scattering-dominated BAs and diamond. In conclusion, this work underscores the limitations of MLIP-based MD for thermal conductivity studies at sub-Debye temperatures.

36 MATERIALS SCIENCE↗

On the transferability of residence time distributions in two 10-km long river sections with similar hydromorphic units

Quantifying hydrologic exchange fluxes (HEFs) at the stream-groundwater interface and their residence time distributions (RTDs) in the subsurface are important for managing the water quality and ecosystem health in dynamic river corridors. However, direct simulating high-spatial resolution HEFs and RTDs can be time-consuming, especially for watershed-scale modeling. Efficient surrogate models linking RTDs to hydromorphic units (HUs) can be alternatives for simulating RTDs in large-scale models. A common concern of these surrogate models, though, is the transferability of the relationship between the RTDs and HUs from one river corridor to another. To address this issue, this work evaluates the HEFs and resulting RTD-HU relationships for two 10-km long river corridors along the Columbia River leveraging a one-way coupled three-dimensional transient surface-subsurface water transport modeling framework we previously developed. Applying such a framework at the two river corridors with similar HUs allows for quantitative comparisons of HEFs and RTDs using both statistical tests and machine learning classification models. Finally, our comparison shows that the similarity and transferability of the RTD-HU relationship is very low for the two investigated river sections, which suggests that devising a general algorithm to estimate RTDs based solely on surface water hydrodynamics and short-distance river channel topography data, as well as HU classification, might be nearly impossible.

54 ENVIRONMENTAL SCIENCES↗

Scaling and merging time-resolved pink-beam diffraction with variational inference

Time-resolved x-ray crystallography (TR-X) at synchrotrons and free electron lasers is a promising technique for recording dynamics of molecules at atomic resolution. While experimental methods for TR-X have proliferated and matured, data analysis is often difficult. Extracting small, time-dependent changes in signal is frequently a bottleneck for practitioners. Recent work demonstrated this challenge can be addressed when merging redundant observations by a statistical technique known as variational inference (VI). However, the variational approach to time-resolved data analysis requires identification of successful hyperparameters in order to optimally extract signal. In this case study, we present a successful application of VI to time-resolved changes in an enzyme, DJ-1, upon mixing with a substrate molecule, methylglyoxal. We present a strategy to extract high signal-to-noise changes in electron density from these data. Furthermore, we conduct an ablation study, in which we systematically remove one hyperparameter at a time to demonstrate the impact of each hyperparameter choice on the success of our model. We expect this case study will serve as a practical example for how others may deploy VI in order to analyze their time-resolved diffraction data.

47 OTHER INSTRUMENTATION↗

A Computational Review of Privacy-Preserving Mechanisms for the Smart Grid

Smart grid technologies have rapidly become one of the largest and most comprehensive sources of data for the modern utility. For the most part, data streams are seen as an essential tool that enable utilities to carry their day-to-day business operations, but they also create the need for efficient and secure data management strategies. In the context of the smart grid, ensuring data privacy is becoming an increasing concern due to a combination of factors that range from shifts in operational paradigms and rapid technology evolution to changes in legislation. Furthermore, researchers have highlighted the risks associated with improperly protected energy records. For example, energy consumption data from homes could be used to infer the behaviors and habits of home occupants through activity recognition or user profiling (Fan, 2017), which may lead to unfair service pricing, targeted advertising, or other personal security violations. Similarly, Electric Vehicles’ (EVs) charging metadata could be used to reveal private information about the owner such as their payment methods, preferred charging stations, and other locational and timing information that could be used to reconstruct the vehicle owner’s behaviors. The privacy of user data, even when used for statistical analysis or machine learning training processes, also needs to be carefully considered, as an individual’s private traits may still be vulnerable if their inclusion/exclusion greatly impacts the result or could be linked to a public dataset through cross-reference. The breach of user privacy also has severe impacts for organizations that store, transmit, or work on the data in the form of diminishing the public’s trust in them while potentially incurring legal consequences (e.g., fines and suspensions under the European Union General Data Protection Regulation, Health Insurance Portability and Accountability Act, etc.). Because of these risks, several privacy-preserving mechanisms are available to help organizations comply with privacy legislations and prevent the unauthorized and malicious use of user data. In light of these concerns, this report focuses on performing a computational review of privacy-preserving mechanisms that have received a significant amount of interest in literature. It specifically focuses on 1) homomorphic encryption, 2) zero-knowledge proofs, 3) differential privacy, and 4) federated learning. It is worth noting that although many of the methods presented in this document rely on cryptographic primitives, their intent is not to provide perfect secrecy, but rather to enable users to maintain privacy, and thus they shall not be compared or equated to other constructs that are aimed to address cybersecurity constructs.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Uncertainty in Synthetic Tropical Cyclone Hazard and Risk Estimates: Insights from RAFT, CHAZ, MIT, STORM, and CLIMADA

We synthesize five complementary tropical cyclone (TC) hazard frameworks—RAFT (physics-based machine learning), CHAZ and MIT (statistical–dynamical), STORM (fully statistical), and CLIMADA (observation-driven resampling)—to characterize uncertainty in wind-related TC metrics relevant to energy applications. All datasets and the IBTrACS observational record are harmonized to a common 6-hourly, 2.5° grid. We compare basin-wide and coastal properties using consistent definitions for TC frequency, mean and maximum intensity, 24-hour intensification, and 6-hour translation speed, and quantify agreement with Pearson r, RMSE, and Kling–Gupta efficiency (KGE) alongside resampling-based confidence intervals. CLIMADA is included for basin context but excluded from coastal skill scoring because it resamples historical IBTrACS; if supplied with projected future tracks from an external hazard model, CLIMADA can be used to simulate future TC scenarios. Results show robust, cross-model signals: (i) a corridor of activity from the tropical Atlantic through the Caribbean into the Bahamas and western subtropical Atlantic; (ii) a meridional dipole in 24-hour intensification (low-latitude strengthening, subtropical weakening); and (iii) a transition from slower tropical motion to faster midlatitude translation. Coastal winds (mean and maximum) consistently cluster from the eastern Gulf into the Bahamas–western Atlantic transition. The largest structural spread occurs in the amplitude and footprint of lifetime maximum intensity and, secondarily, in translation speed; intensification exhibits similar central behavior across frameworks with variability in extremes. Translation speed shows the most uniform coastal agreement. These findings provide a decision envelope for wind-focused risk screening and clarify where uncertainty should be carried forward; wind-only results represent a lower bound on total hazard, motivating integration of surge and rainfall modules and a companion, asset-level damage analysis.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Market Survey 2020: Commercial Clinical Decision Support Systems and Wellness Tools

For long-duration, deep space exploration missions, current methods for managing and supporting crew health and medical conditions will be unsuitable. Communication and data transmission lags will necessitate the use of a sophisticated clinical decision support system (CDSS) that will tailor diagnosis and treatment guidance that is context-sensitive for anticipated astronaut health, wellness, and medical conditions. A variety of clinical decision support (CDS) and wellness tools (WT) are currently available in the commercial market and a broad-brush survey of this market can provide an initial impression of the current state of the art which, in turn, can inform the roadmap of NASA deep space CDSS development and associated requirements. Such a survey was undertaken during the first six months of 2020 using directed convenience sampling to obtain information provided by vendors on their websites; both commercially available CDS and WT (such as those used to track and monitor nutrition, exercise, and sleep) were included. Areas assessed were item type (e.g., software/application, device); primary purpose of the item (e.g., diagnostic support, nutrition tracking); additional purposes (if any); reported features, capabilities, and functionality; setting of use (e.g., inpatient, outpatient); intended user (e.g., clinician, patient); location and sources of data/information used or produced by the item; integration with patient electronic health record (EHR); compliance with interoperability ontologies and standards (e.g., Health Level 7 [HL7], Systematized Nomenclature of Medicine – Clinical Terminology [SNOMED-CT]); and whether the item is knowledge-based (derived from research findings) or non-knowledge-based (derived through artificial intelligence, machine learning, advanced probability and statistics), among others. Ninety-seven (97) vendor websites describing 196 CDS and 73 WT (269 total) were reviewed and coded. The primary purpose of the majority of CDS reviewed is diagnosis or diagnosis/treatment/drug decision support—targeted for clinician use— and the primary purpose of the majority of WT reviewed is the monitoring of different health metrics, most often through the use of a biosensor device (e.g., blood pressure)—targeted for patient use. Very few CDS or WT appear to comply with major international interoperability standards or can be integrated with a patient’s EHR data. None consider contextual factors, such as conditions of the physical environment (e.g., CO2 levels). The majority of CDS and WT reviewed are non-knowledge, cloud- or web-based applications or software. Forty-three (43) major findings were identified and the implications those findings have for NASA will be discussed. Example major findings include: CDS-WT capabilities range from diagnosis to treatment applications, CDS-WT may be wearable or non-wearable and are technologically advanced and only a few CDS-WT tools referenced compliance to ensure interoperability, among other findings. Recommendations will also be offered that will help to address ExMC Gap, Medical-701: Enhance medical capabilities within an exploration medical system.

market survey↗

Discovery and Analysis of Rare High-Impact Failure Modes using Adversarial RL-Informed Sampling

Adaptive learning agents have tremendous potential to handle critical tasks currently performed by humans. Unfortunately, due to their complexity, it can be difficult to verify that these learning agents do not have critical failure modes. Standard verification and validation methods often do not apply directly to learning agents and Monte Carlo methods have difficulty covering even a small fraction of the state space, especially in multiagent systems or over long time horizons. To overcome this difficulty, we demonstrate an adaptive stress-testing method based on reinforcement learning of correlations that raise the probability of failure. This approach has three key properties: (1) it is able to find rare failure modes with far greater sample efficiency than Monte Carlo methods, (2) it can estimate the true probability of a failure mode despite the inherent bias in the learning method, and (3) it is capable of learning and resampling compact representations of multimodal failure spaces. These properties are important in practice as we need to find disparate failure modes while accounting for their actual relevance. This is a significant advantage over traditional adaptive stress testing methods that give abstract likelihoods of particular failure instances, but cannot estimate the probability of a broader failure mode. We test our algorithm on a simple problem from the aviation domain where an autonomous aircraft lands in gusty wind conditions. The results suggest that we can find failure modes with far fewer samples than the Monte Carlo approach and simultaneously estimate the probability of failure.

reinforcement learning↗

Operations on Graphical Models with Plates

This paper explains how graphical models, for instance Bayesian or Markov networks, can be extended to model problems in data analysis and learning. This provides a unified framework that combines lessons learned from the artificial intelligence, statistical and connectionist communities. This also offers a set of principles for developing a software generator for data analysis, whereby a learning or discovery system can be compiled from specifications. Many of the popular learning algorithms can be compiled in this way from graphical specifications. While in a sense this paper is a multidisciplinary review of learning, the main contribution here is the presentation of the material within the unifying framework of graphical models, and the observation that, as a result, the process of developing learning algorithms can be partly automated.

Buntine, Wray L.↗

Improved Subseasonal Forecasting of Extreme Polar Vortices Using Machine Learning

Our research was focused on forecasting the position and shape of the winter stratospheric polar vortex at a subseasonal timescale of 15 days in advance. To achieve this, we employed both statistical and neural network machine learning techniques. The analysis was performed on 42 winter seasons of reanalysis data provided by NASA giving us a total of 6,342 days of data. The state of the polar vortex for determined by using geometric moments to calculate the centroid latitude and the aspect ratio of an ellipse fit onto the vortex. Timeseries for thirty additional precursors were calculated to help improve the predictive capabilities of the algorithm. Feature importance of these precursors was performed using random forest to measure the predictive importance and the ideal number of precursors. Then, using the precursors identified as important, various statistical methods were tested for predictive accuracy with random forest and nearest neighbor performing the best. An echo state network, a type of recurrent neural network that features sparsely connected hidden layer and a reduced number of trainable parameters that allows for rapid training and testing, was also implemented for the forecasting problem. Hyperparameter tuning was performed for each methods using a subset of the training data. The algorithms were trained and tuned on the first 41 years of data, then tested for accuracy on the final year. In general, the centroid latitude of the polar vortex proved easier to predict than the aspect ratio across all algorithms. Random forest outperformed other statistical forecasting algorithms overall but struggled to predict extreme values. Forecasting from echo state network suggested a strong predictive capability past 15 days, but further work is required to fully realize the potential of recurrent neural network approaches.

54 ENVIRONMENTAL SCIENCES↗

Learning classification trees

Algorithms for learning classification trees have had successes in artificial intelligence and statistics over many years. How a tree learning algorithm can be derived from Bayesian decision theory is outlined. This introduces Bayesian techniques for splitting, smoothing, and tree averaging. The splitting rule turns out to be similar to Quinlan's information gain splitting rule, while smoothing and averaging replace pruning. Comparative experiments with reimplementations of a minimum encoding approach, Quinlan's C4 and Breiman et al. Cart show the full Bayesian algorithm is consistently as good, or more accurate than these other approaches though at a computational price.

Buntine, Wray↗

Statistical Engineering

This webinar provides an overview of the International Statistical Engineering Association (ISEA), and it illustrates the practice of statistical engineering at NASA. ISEA was formed to promote the study of how successful data-based problem-solving methods are leveraged to realize innovative opportunities and solve problems sustainably. ISEA is comprised of statisticians, engineers, scientists, and other professionals that exchange ideas and experiences in the development and application of statistical engineering theories. ISEA is building the body of knowledge of the statistical engineering discipline with a particular focus on improving academic preparation for tackling complex problems. Over the past 15 years, the practice of statistical engineering has gained recognition within NASA by spurring innovation and efficiency, and it has demonstrated significant impact. Aerospace research and development benefits from an application-focused statistical engineering perspective to accelerate learning, maximize knowledge, ensure strategic resource investment, and inform data-driven decisions. The second portion of this presentation provides an overview of infusing statistical engineering at NASA through pioneering case studies in aeronautics, space exploration, and atmospheric science.

Peter A Parker↗

Emerging anomaly detection techniques for electronic health records: A survey

Background Anomaly detection in electronic health records (EHRs) is a cornerstone of biomedical informatics, with direct implications for patient safety, clinical decision-making, and the prevention of healthcare fraud. Once guided primarily by simple rule-based methods, the field has advanced rapidly, driven by increased computing power, richer and more detailed health data, and the rise of machine learning and deep learning techniques. The objective of this paper is to provide a comprehensive overview of modern approaches to detecting anomalies in EHRs, outlining their strengths, limitations, and relevance to key healthcare challenges. We review traditional statistical methods alongside newer ML- and DL-based strategies and hybrid models, with particular attention to how these techniques support transparency and build clinical trust. Methods This paper presents a thorough and critical survey through systematic review (PRISMA-based) of the latest anomaly detection strategies in time-sequence data domains within electronic health record systems. Results We explore a broad spectrum of methodologies, including statistical models, supervised and unsupervised learning approaches, hybrid frameworks, and state-of-the-art ML-based techniques that collectively advance the precision and scalability of detecting anomalies in complex clinical datasets. In addition to mapping current capabilities, we address the enduring challenges that hinder widespread implementation and provide a forward-looking perspective on the future of anomaly detection in the data-rich landscape of modern healthcare. Summary The advancement in AI-based approaches is reported along with the basic principles of the individual approaches and their applicability. The increased availability of high-quality data, advancements in DL approaches, and enhanced computation power are leading to more frequent adaptation of DL-based approaches. Emerging DL-based approaches that have been adapted in other domains or recently applied in the EHR domain are also discussed in detail. Although DL-based approaches can improve model predictions by incorporating comorbidities, their application is limited in low-frequency data domains (e.g., when the total available data remains in the single digits). Therefore, the user must carefully consider the application based on data availability.

Anomaly detection↗