Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “statistical learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 289 records · Page 16

On the transferability of residence time distributions in two 10-km long river sections with similar hydromorphic units

Quantifying hydrologic exchange fluxes (HEFs) at the stream-groundwater interface and their residence time distributions (RTDs) in the subsurface are important for managing the water quality and ecosystem health in dynamic river corridors. However, direct simulating high-spatial resolution HEFs and RTDs can be time-consuming, especially for watershed-scale modeling. Efficient surrogate models linking RTDs to hydromorphic units (HUs) can be alternatives for simulating RTDs in large-scale models. A common concern of these surrogate models, though, is the transferability of the relationship between the RTDs and HUs from one river corridor to another. To address this issue, this work evaluates the HEFs and resulting RTD-HU relationships for two 10-km long river corridors along the Columbia River leveraging a one-way coupled three-dimensional transient surface-subsurface water transport modeling framework we previously developed. Applying such a framework at the two river corridors with similar HUs allows for quantitative comparisons of HEFs and RTDs using both statistical tests and machine learning classification models. Finally, our comparison shows that the similarity and transferability of the RTD-HU relationship is very low for the two investigated river sections, which suggests that devising a general algorithm to estimate RTDs based solely on surface water hydrodynamics and short-distance river channel topography data, as well as HU classification, might be nearly impossible.

54 ENVIRONMENTAL SCIENCES↗

Deep neural network improves the estimation of polygenic risk scores for breast cancer

Polygenic risk scores (PRS) estimate the genetic risk of an individual for a complex disease based on many genetic variants across the whole genome. Here, we compared a series of computational models for estimation of breast cancer PRS. A deep neural network (DNN) was found to outperform alternative machine learning techniques and established statistical algorithms, including BLUP, BayesA, and LDpred. In the test cohort with 50% prevalence, the Area Under the receiver operating characteristic Curve (AUC) were 67.4% for DNN, 64.2% for BLUP, 64.5% for BayesA, and 62.4% for LDpred. BLUP, BayesA, and LPpred all generated PRS that followed a normal distribution in the case population. However, the PRS generated by DNN in the case population followed a bimodal distribution composed of two normal distributions with distinctly different means. This suggests that DNN was able to separate the case population into a high-genetic-risk case subpopulation with an average PRS significantly higher than the control population and a normal-genetic-risk case subpopulation with an average PRS similar to the control population. This allowed DNN to achieve 18.8% recall at 90% precision in the test cohort with 50% prevalence, which can be extrapolated to 65.4% recall at 20% precision in a general population with 12% prevalence. Interpretation of the DNN model identified salient variants that were assigned insignificant p values by association studies, but were important for DNN prediction. These variants may be associated with the phenotype through nonlinear relationships.

59 BASIC BIOLOGICAL SCIENCES↗

Multidimensional quantitative phenotypic and molecular analysis reveals neomorphic behaviors of p53 missense mutants

Abstract Mutations in the TP53 tumor suppressor gene occur in >80% of the triple-negative or basal-like breast cancer. To test whether neomorphic functions of specific TP53 missense mutations contribute to phenotypic heterogeneity, we characterized phenotypes of non-transformed MCF10A-derived cell lines expressing the ten most common missense mutant p53 proteins and observed a wide spectrum of phenotypic changes in cell survival, resistance to apoptosis and anoikis, cell migration, invasion and 3D mammosphere architecture. The p53 mutants R248W, R273C, R248Q, and Y220C are the most aggressive while G245S and Y234C are the least, which correlates with survival rates of basal-like breast cancer patients. Interestingly, a crucial amino acid difference at one position—R273C vs. R273H—has drastic changes on cellular phenotype. RNA-Seq and ChIP-Seq analyses show distinct DNA binding properties of different p53 mutants, yielding heterogeneous transcriptomics profiles, and MD simulation provided structural basis of differential DNA binding of different p53 mutants. Integrative statistical and machine-learning-based pathway analysis on gene expression profiles with phenotype vectors across the mutant cell lines identifies quantitative association of multiple pathways including the Hippo/YAP/TAZ pathway with phenotypic aggressiveness. Further, comparative analyses of large transcriptomics datasets on breast cancer cell lines and tumors suggest that dysregulation of the Hippo/YAP/TAZ pathway plays a key role in driving the cellular phenotypes towards basal-like in the presence of more aggressive p53 mutants. Overall, our study describes distinct gain-of-function impacts on protein functions, transcriptional profiles, and cellular behaviors of different p53 missense mutants, which contribute to clinical phenotypic heterogeneity of triple-negative breast tumors.

60 APPLIED LIFE SCIENCES↗

Scaling and merging time-resolved pink-beam diffraction with variational inference

Time-resolved x-ray crystallography (TR-X) at synchrotrons and free electron lasers is a promising technique for recording dynamics of molecules at atomic resolution. While experimental methods for TR-X have proliferated and matured, data analysis is often difficult. Extracting small, time-dependent changes in signal is frequently a bottleneck for practitioners. Recent work demonstrated this challenge can be addressed when merging redundant observations by a statistical technique known as variational inference (VI). However, the variational approach to time-resolved data analysis requires identification of successful hyperparameters in order to optimally extract signal. In this case study, we present a successful application of VI to time-resolved changes in an enzyme, DJ-1, upon mixing with a substrate molecule, methylglyoxal. We present a strategy to extract high signal-to-noise changes in electron density from these data. Furthermore, we conduct an ablation study, in which we systematically remove one hyperparameter at a time to demonstrate the impact of each hyperparameter choice on the success of our model. We expect this case study will serve as a practical example for how others may deploy VI in order to analyze their time-resolved diffraction data.

47 OTHER INSTRUMENTATION↗

Signatures of a liquid–liquid transition in an ab initio deep neural network model for water

Significance Water is central across much of the physical and biological sciences and exhibits physical properties that are qualitatively distinct from those of most other liquids. Understanding the microscopic basis of water’s peculiar properties remains an active area of research. One intriguing hypothesis is that liquid water can separate into metastable high- and low-density liquid phases at low temperatures and high pressures, and the existence of this liquid–liquid transition could explain many of water’s anomalous properties. We used state-of-the-art approaches in computational quantum chemistry, statistical mechanics, and machine learning and obtained evidence consistent with a liquid–liquid transition, supporting the argument for the existence of this phenomenon in real water.

36 MATERIALS SCIENCE↗

Noisy intermediate-scale quantum algorithms

A universal fault-tolerant quantum computer that can efficiently solve problems such as integer factorization and unstructured database search requires millions of qubits with low error rates and long coherence times. While the experimental advancement toward realizing such devices will potentially take decades of research, noisy intermediate-scale quantum (NISQ) computers already exist. These computers are composed of hundreds of noisy qubits, i.e., qubits that are not error corrected, and therefore perform imperfect operations within a limited coherence time. In the search for achieving quantum advantage with these devices, algorithms have been proposed for applications in various disciplines spanning physics, machine learning, quantum chemistry, and combinatorial optimization. The overarching goal of such algorithms is to leverage the limited available resources to perform classically challenging tasks. In this review, a thorough summary of NISQ computational paradigms and algorithms is provided. The key structure of these algorithms and their limitations and advantages are discussed. Finally, a comprehensive overview of various benchmarking and software tools useful for programming and testing NISQ devices is additionally provided.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Crash Risks Evaluation of Urban Expressways: A Case Study in Shanghai

We report that proactive traffic safety management systems can reduce crashes by identifying crash precursors, evaluating real-time crash risks, and implementing suitable interventions. The basic prerequisite for developing such a system is to propose a reliable crash risk evaluation model that takes real-time traffic flow data as input. Previous studies have primarily focused on real-time crash prediction using some statistical or machine-learning methods. However, further quantitative evaluation and classification of crash risks have been ignored. In this study, we conduct a systematic crash risk evaluation workflow, including crash risk prediction, crash risk quantification, and crash risk classification. Specifically, the crash risk prediction using an extended logit model is proposed, from which CAS, CSD, UAS, DAS, DTV are identified to be contributing factors of crash risks. Then a crash risk quantification model based on the parameter evaluation of the extended logit model is developed. The crash risks of urban expressways and their spatial-temporal evolution trends are quantified. Finally, the crash risks are classified into high crash risk level, moderate crash risk level, and low crash risk level by the k-means cluster algorithm. Then the threshold boundaries of different crash risk levels are determined. The research results provide a proactive guidance for traffic safety management of urban expressways.

33 ADVANCED PROPULSION SYSTEMS↗

Rapid Electrochemical Diagnosis of Battery Health and Safety from Cells to Modules

Rapid electrochemical diagnosis of battery health and failure is critical for ensuring reliable battery performance and battery safety. Traditional battery health diagnostics such as capacity measurements and DC pulse tests are reliable and well-understood, however, these measurements of battery capacity and resistance do not capture all aspects of battery degradation. Other aspects of degradation, such as electrolyte decomposition, lithium-plating, and particle cracking are difficult to detect electrochemically but are crucial to measure to get a full picture of battery safety and flag out potential failures. In this work, lab- and field-aged commercial lithium-ion batteries and modules of various chemistries and formats are tested using a variety of traditional electrochemical characterization methods as well as using 2-minute pseudo-random DC pulse sequences at rest and during charge/discharge. The electrochemical measurements are compared to physical cell measurements, cell efficiency, drive cycle performance, physical and thermal heterogeneity, and qualitative safety metrics using statistical and machine-learning methods to discover if a comprehensive "battery health map" can be accurately identified using only rapid DC measurements.

ADVANCED PROPULSION SYSTEMS,ENERGY STORAGE↗

Data and scripts associated with a manuscript on residence time distribution simulation in two 10-kilometer long river sections

This data package is associated with the publication “On the Transferability of Residence Time Distributions in Two 10-km Long River Sections with Similar Hydromorphic Units” submitted to the Journal of Hydrology (Bao et al. 2024).Quantifying hydrologic exchange fluxes (HEFs) at the stream-groundwater interface, along with their residence time distributions (RTDs) in the subsurface, is crucial for managing water quality and ecosystem health in dynamic river corridors. However, directly simulating high-spatial resolution HEFs and RTDs can be a time-consuming process, particularly for watershed-scale modeling. Efficient surrogate models that link RTDs to hydromorphic units (HUs) may serve as alternatives for simulating RTDs in large-scale models. One common concern with these surrogate models, however, is the transferability of the relationship between the RTDs and HUs from one river corridor to another. To address this, we evaluated the HEFs and the resulting RTD-HU relationships for two 10-kilometer-long river corridors along the Columbia River, using a one-way coupled three-dimensional transient surface-subsurface water transport modeling framework that we previously developed. Applying this framework to the two river corridors with similar HUs allows for quantitative comparisons of HEFs and RTDs using both statistical tests and machine learning classification models. This data package includes the model inputs files and the simulation results data. This data package contains 10 folders. The modeling simulation results data are in the folders 100H_pt_data and 300area_pt_data, for the study domain Hanford 100H and 300 area respectively. The remaining eight folders contain the scripts and data to generate the manuscript figures. The file-level metadata file (Bao_2024_Residence_Time_Distribution _flmd.csv) includes a list of all files contained in this data package and descriptions for each. The data dictionary file (Bao_2024_Residence_Time_Distribution _dd.csv) includes column header definitions and units of all tabular files.

54 ENVIRONMENTAL SCIENCES↗

Feature Detection

Focal Area(s): This proposal aims to develop and evaluate statistical models and machine learning algorithms for detecting and tracking features in spatiotemporal remotely sensed data with uncertainty quantification. We focus a particular application on the detection of sea ice leads and ridges in the Arctic and use these key sea ice features for model calibration and to gain insight into the physics of sea ice thermodynamics and deformation.

54 ENVIRONMENTAL SCIENCES↗

A Computational Review of Privacy-Preserving Mechanisms for the Smart Grid

Smart grid technologies have rapidly become one of the largest and most comprehensive sources of data for the modern utility. For the most part, data streams are seen as an essential tool that enable utilities to carry their day-to-day business operations, but they also create the need for efficient and secure data management strategies. In the context of the smart grid, ensuring data privacy is becoming an increasing concern due to a combination of factors that range from shifts in operational paradigms and rapid technology evolution to changes in legislation. Furthermore, researchers have highlighted the risks associated with improperly protected energy records. For example, energy consumption data from homes could be used to infer the behaviors and habits of home occupants through activity recognition or user profiling (Fan, 2017), which may lead to unfair service pricing, targeted advertising, or other personal security violations. Similarly, Electric Vehicles’ (EVs) charging metadata could be used to reveal private information about the owner such as their payment methods, preferred charging stations, and other locational and timing information that could be used to reconstruct the vehicle owner’s behaviors. The privacy of user data, even when used for statistical analysis or machine learning training processes, also needs to be carefully considered, as an individual’s private traits may still be vulnerable if their inclusion/exclusion greatly impacts the result or could be linked to a public dataset through cross-reference. The breach of user privacy also has severe impacts for organizations that store, transmit, or work on the data in the form of diminishing the public’s trust in them while potentially incurring legal consequences (e.g., fines and suspensions under the European Union General Data Protection Regulation, Health Insurance Portability and Accountability Act, etc.). Because of these risks, several privacy-preserving mechanisms are available to help organizations comply with privacy legislations and prevent the unauthorized and malicious use of user data. In light of these concerns, this report focuses on performing a computational review of privacy-preserving mechanisms that have received a significant amount of interest in literature. It specifically focuses on 1) homomorphic encryption, 2) zero-knowledge proofs, 3) differential privacy, and 4) federated learning. It is worth noting that although many of the methods presented in this document rely on cryptographic primitives, their intent is not to provide perfect secrecy, but rather to enable users to maintain privacy, and thus they shall not be compared or equated to other constructs that are aimed to address cybersecurity constructs.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Uncertainty in Synthetic Tropical Cyclone Hazard and Risk Estimates: Insights from RAFT, CHAZ, MIT, STORM, and CLIMADA

We synthesize five complementary tropical cyclone (TC) hazard frameworks—RAFT (physics-based machine learning), CHAZ and MIT (statistical–dynamical), STORM (fully statistical), and CLIMADA (observation-driven resampling)—to characterize uncertainty in wind-related TC metrics relevant to energy applications. All datasets and the IBTrACS observational record are harmonized to a common 6-hourly, 2.5° grid. We compare basin-wide and coastal properties using consistent definitions for TC frequency, mean and maximum intensity, 24-hour intensification, and 6-hour translation speed, and quantify agreement with Pearson r, RMSE, and Kling–Gupta efficiency (KGE) alongside resampling-based confidence intervals. CLIMADA is included for basin context but excluded from coastal skill scoring because it resamples historical IBTrACS; if supplied with projected future tracks from an external hazard model, CLIMADA can be used to simulate future TC scenarios. Results show robust, cross-model signals: (i) a corridor of activity from the tropical Atlantic through the Caribbean into the Bahamas and western subtropical Atlantic; (ii) a meridional dipole in 24-hour intensification (low-latitude strengthening, subtropical weakening); and (iii) a transition from slower tropical motion to faster midlatitude translation. Coastal winds (mean and maximum) consistently cluster from the eastern Gulf into the Bahamas–western Atlantic transition. The largest structural spread occurs in the amplitude and footprint of lifetime maximum intensity and, secondarily, in translation speed; intensification exhibits similar central behavior across frameworks with variability in extremes. Translation speed shows the most uniform coastal agreement. These findings provide a decision envelope for wind-focused risk screening and clarify where uncertainty should be carried forward; wind-only results represent a lower bound on total hazard, motivating integration of surge and rainfall modules and a companion, asset-level damage analysis.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Market Survey 2020: Commercial Clinical Decision Support Systems and Wellness Tools

For long-duration, deep space exploration missions, current methods for managing and supporting crew health and medical conditions will be unsuitable. Communication and data transmission lags will necessitate the use of a sophisticated clinical decision support system (CDSS) that will tailor diagnosis and treatment guidance that is context-sensitive for anticipated astronaut health, wellness, and medical conditions. A variety of clinical decision support (CDS) and wellness tools (WT) are currently available in the commercial market and a broad-brush survey of this market can provide an initial impression of the current state of the art which, in turn, can inform the roadmap of NASA deep space CDSS development and associated requirements. Such a survey was undertaken during the first six months of 2020 using directed convenience sampling to obtain information provided by vendors on their websites; both commercially available CDS and WT (such as those used to track and monitor nutrition, exercise, and sleep) were included. Areas assessed were item type (e.g., software/application, device); primary purpose of the item (e.g., diagnostic support, nutrition tracking); additional purposes (if any); reported features, capabilities, and functionality; setting of use (e.g., inpatient, outpatient); intended user (e.g., clinician, patient); location and sources of data/information used or produced by the item; integration with patient electronic health record (EHR); compliance with interoperability ontologies and standards (e.g., Health Level 7 [HL7], Systematized Nomenclature of Medicine – Clinical Terminology [SNOMED-CT]); and whether the item is knowledge-based (derived from research findings) or non-knowledge-based (derived through artificial intelligence, machine learning, advanced probability and statistics), among others. Ninety-seven (97) vendor websites describing 196 CDS and 73 WT (269 total) were reviewed and coded. The primary purpose of the majority of CDS reviewed is diagnosis or diagnosis/treatment/drug decision support—targeted for clinician use— and the primary purpose of the majority of WT reviewed is the monitoring of different health metrics, most often through the use of a biosensor device (e.g., blood pressure)—targeted for patient use. Very few CDS or WT appear to comply with major international interoperability standards or can be integrated with a patient’s EHR data. None consider contextual factors, such as conditions of the physical environment (e.g., CO2 levels). The majority of CDS and WT reviewed are non-knowledge, cloud- or web-based applications or software. Forty-three (43) major findings were identified and the implications those findings have for NASA will be discussed. Example major findings include: CDS-WT capabilities range from diagnosis to treatment applications, CDS-WT may be wearable or non-wearable and are technologically advanced and only a few CDS-WT tools referenced compliance to ensure interoperability, among other findings. Recommendations will also be offered that will help to address ExMC Gap, Medical-701: Enhance medical capabilities within an exploration medical system.

market survey↗

Discovery and Analysis of Rare High-Impact Failure Modes using Adversarial RL-Informed Sampling

Adaptive learning agents have tremendous potential to handle critical tasks currently performed by humans. Unfortunately, due to their complexity, it can be difficult to verify that these learning agents do not have critical failure modes. Standard verification and validation methods often do not apply directly to learning agents and Monte Carlo methods have difficulty covering even a small fraction of the state space, especially in multiagent systems or over long time horizons. To overcome this difficulty, we demonstrate an adaptive stress-testing method based on reinforcement learning of correlations that raise the probability of failure. This approach has three key properties: (1) it is able to find rare failure modes with far greater sample efficiency than Monte Carlo methods, (2) it can estimate the true probability of a failure mode despite the inherent bias in the learning method, and (3) it is capable of learning and resampling compact representations of multimodal failure spaces. These properties are important in practice as we need to find disparate failure modes while accounting for their actual relevance. This is a significant advantage over traditional adaptive stress testing methods that give abstract likelihoods of particular failure instances, but cannot estimate the probability of a broader failure mode. We test our algorithm on a simple problem from the aviation domain where an autonomous aircraft lands in gusty wind conditions. The results suggest that we can find failure modes with far fewer samples than the Monte Carlo approach and simultaneously estimate the probability of failure.

reinforcement learning↗

The Data Synergy Effects of Time-Series Deep Learning Models in Hydrology

When fitting statistical models to variables in geoscientific disciplines such as hydrology, it is a customary practice to stratify a large domain into multiple regions (or regimes) and study each region separately. Traditional wisdom suggests that models built for each region separately will have higher performance because of homogeneity within each region. However, each stratified model has access to fewer and less diverse data points. Here, through two hydrologic examples (soil moisture and streamflow), we show that conventional wisdom may no longer hold in the era of big data and deep learning (DL). We systematically examined an effect we call data synergy, where the results of the DL models improved when data were pooled together from characteristically different regions. The performance of the DL models benefited from modest diversity in the training data compared to a homogeneous training set, even with similar data quantity. Moreover, allowing heterogeneous training data makes eligible much larger training datasets, which is an inherent advantage of DL. A large, diverse data set is advantageous in terms of representing extreme events and future scenarios, which has strong implications for climate change impact assessment. The results here suggest the research community should place greater emphasis on data sharing.

54 ENVIRONMENTAL SCIENCES↗