Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “hypothesis learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Physics-Guided Machine Learning (PGML) for Improved Aerostructure Manufacturing

The Physics-Guided Machine Learning (PGML) for Improved Aerostructure Manufacturing project objective is to automatically tune machining parameter predictions from physics-based models using process data and Bayesian machine learning. The intent is to enable a step change in aerospace manufacturing by combining machine learning, physics-based process models, and sensors/data in a comprehensive digital environment that simultaneously considers the computer numerically controlled (CNC) machining center capabilities, the workpiece material and geometry, and the workpiece support (fixturing). The project hypothesis is that this combination will enable improved performance in machining operations.

42 ENGINEERING↗

Natural Language Processing for Text Based Event Extraction: Identifying Events of Interest Related to Worldwide State-Sponsored Civil Nuclear Power

Beginning in FY20, SRNL was funded by the National Nuclear Security Administration’s Office of Defense Nuclear Non-Proliferation Research and Development to develop a prototype natural language processing/natural language understating machine learning-based modeling and analysis pipeline to extract and forecast events of interest from massive open data sources. The working hypothesis within the approach is that contextual shifts in key words and phrases act as indicators of events of interest over time. Therefore, by identifying points in time where contextual shifts occur, events of interest can be extracted along with explicit and implicit connections of entities and activities. The development of the preliminary prototype pipeline proved successful, meriting further testing of the pipeline on more broad topical domains and in a worldwide data environment. Therefore, SRNL, in collaboration with the Sanghani Center for Artificial Intelligence and Data Analytics at Virginia Tech, have continued development with a test case of identifying events of interest related to worldwide state-sponsored civil nuclear power in open data sources. In the first year of this follow-on effort, the team has curated domain-specific data corpuses using an automated scheme and applied the modeling and analysis pipeline. This robust, focused, and efficient approach consists of an ensemble of analyses applied to time dependent word embedding models that are trained on the data corpuses. In this report, the team has demonstrated the capability of the existing pipeline (as development has continued in parallel) by exploring several specific case-studies centered around Rosatom’s international activities regarding the planning, construction, operation, and/or shutdown of nuclear reactors. A basic timeline events has been generated by manually cataloging known “milestone” events that have occurred at reactors in Turkey, Finland, Hungary, and Egypt and compared with the output of the modeling pipeline. In this approach, the team has characterized the lead time using the prototype pipeline, as well as the ability to capture relevant information, which proved 100% successful. A deep dive example of the Akkuyu reactor (Turkey) is presented that shows the breadth of information that can be captured using the approach. In this case study, events were extracted pertaining to the planning/construction of Akkuyu including protests from the population, information campaigns in response to the protests, forged regulatory documents and lawsuits, budgetary/shareholder information, geopolitical tensions, and the various construction milestones. This has demonstrated the pipeline’s utility as a research aid or real-time event extraction tool, where summary-level information and detailed text extractions from millions of articles or Tweets across long time periods can be generated with significantly less effort than current techniques.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Causal machine learning uncovers conditions for convective intensification driven by organic and sulfate aerosols

Aerosols are often hypothesized to invigorate deep convective clouds (DCCs), but observational evidence remains limited and inconclusive. Clarifying this hypothesis is critical for regions vulnerable to thunderstorms and flooding, particularly highly polluted coastal cities. Leveraging a novel causal discovery–inference pipeline and high-resolution observations near Houston, TX, we identify multiple causal pathways among aerosols (mostly organic and sulfate), DCCs, and meteorological factors. However, a direct causal link from aerosols to DCCs is found to be uncommon, occurring in less than 35% of analyzed scenarios, and is characterized by strong conditionality and nonlinearity. When aerosol impacts on DCCs do occur, they can be substantial, enhancing DCC core heights by approximately 1.7 km, with 92% of this effect concentrated in warmer-phase cloud regions. Notably, the presence of sea breezes and the inclusion of all measured aerosol particles each enhance DCCs in over 95% of aerosol-sensitive cases.

54 ENVIRONMENTAL SCIENCES↗

SNM Radiation Signature Classification Using Different Semi-Supervised Machine Learning Models

The timely detection of special nuclear material (SNM) transfers between nuclear facilities is an important monitoring objective in nuclear nonproliferation. Persistent monitoring enabled by successful detection and characterization of radiological material movements could greatly enhance the nuclear nonproliferation mission in a range of applications. Supervised machine learning can be used to signal detections when material is present if a model is trained on sufficient volumes of labeled measurements. However, the nuclear monitoring data needed to train robust machine learning models can be costly to label since radiation spectra may require strict scrutiny for characterization. Therefore, this work investigates the application of semi-supervised learning to utilize both labeled and unlabeled data. As a demonstration experiment, radiation measurements from sodium iodide (NaI) detectors are provided by the Multi-Informatics for Nuclear Operating Scenarios (MINOS) venture at Oak Ridge National Laboratory (ORNL) as sample data. Anomalous measurements are identified using a method of statistical hypothesis testing. After background estimation, an energy-dependent spectroscopic analysis is used to characterize an anomaly based on its radiation signatures. In the absence of ground-truth information, a labeling heuristic provides data necessary for training and testing machine learning models. Supervised logistic regression serves as a baseline to compare three semi-supervised machine learning models: co-training, label propagation, and a convolutional neural network (CNN). In each case, the semi-supervised models outperform logistic regression, suggesting that unlabeled data can be valuable when training and demonstrating value in semi-supervised nonproliferation implementations.

98 NUCLEAR DISARMAMENT, SAFEGUARDS, AND PHYSICAL P↗

Observation of high-energy neutrinos from the Galactic plane

The origin of high-energy cosmic rays, atomic nuclei that continuously impact Earth’s atmosphere, is unknown. Because of deflection by interstellar magnetic fields, cosmic rays produced within the Milky Way arrive at Earth from random directions. However, cosmic rays interact with matter near their sources and during propagation, which produces high-energy neutrinos. We searched for neutrino emission using machine learning techniques applied to 10 years of data from the IceCube Neutrino Observatory. By comparing diffuse emission models to a background-only hypothesis, we identified neutrino emission from the Galactic plane at the 4.5σ level of significance. The signal is consistent with diffuse emission of neutrinos from the Milky Way but could also arise from a population of unresolved point sources.

Science & Technology - Other Topics↗

Synthesis of motif and symmetry for accelerated learning, discovery, and design of electronic structures for energy conversion applications (Final Technical Report)

The overall goal of the projects is to develop a framework to incorporate structure motifs and crystal/orbital symmetries into the data-driven materials discovery infrastructure. The PI proposed to develop structure-motif- and symmetry-based graph convolutional networks for effective learning and efficient predictions of electronic structures and related properties. Fundamental understanding of the roles of structure motif and symmetry will establish new hypothesis and design rules, which will be combined with high-throughput computations based on density functional theory to discover novel light absorbers, transparent conductors, as well as 2D light emitting materials and heterojunctions for optoelectronics.

36 MATERIALS SCIENCE↗

NREM sleep as a novel protective cognitive reserve factor in the face of Alzheimer's disease pathology

Alzheimer’s disease (AD) pathology impairs cognitive function. Yet some individuals with high amounts of AD pathology suffer marked memory impairment, while others with the same degree of pathology burden show little impairment. Why is this? One proposed explanation is cognitive reserve i.e., factors that confer resilience against, or compensation for the effects of AD pathology. Deep NREM slow wave sleep (SWS) is recognized to enhance functions of learning and memory in healthy older adults. However, that the quality of NREM SWS (NREM slow wave activity, SWA) represents a novel cognitive reserve factor in older adults with AD pathology, thereby providing compensation against memory dysfunction otherwise caused by high AD pathology burden, remains unknown. Here, we tested this hypothesis in cognitively normal older adults (N = 62) by combining 11 C-PiB (Pittsburgh compound B) positron emission tomography (PET) scanning for the quantification of β-amyloid (Aβ) with sleep electroencephalography (EEG) recordings to quantify NREM SWA and a hippocampal-dependent face-name learning task. We demonstrated that NREM SWA significantly moderates the effect of Aβ status on memory function. Specifically, NREM SWA selectively supported superior memory function in individuals suffering high Aβ burden, i.e., those most in need of cognitive reserve (B = 2.694, p = 0.019). In contrast, those without significant Aβ pathological burden, and thus without the same need for cognitive reserve, did not similarly benefit from the presence of NREM SWA (B = -0.115, p = 0.876). This interaction between NREM SWA and Aβ status predicting memory function was significant after correcting for age, sex, Body Mass Index, gray matter atrophy, and previously identified cognitive reserve factors, such as education and physical activity (p = 0.042). These findings indicate that NREM SWA is a novel cognitive reserve factor providing resilience against the memory impairment otherwise caused by high AD pathology burden. Furthermore, this cognitive reserve function of NREM SWA remained significant when accounting both for covariates, and factors previously linked to resilience, suggesting that sleep might be an independent cognitive reserve resource. Beyond such mechanistic insights are potential therapeutic implications. Unlike many other cognitive reserve factors (e.g., years of education, prior job complexity), sleep is a modifiable factor. As such, it represents an intervention possibility that may aid the preservation of cognitive function in the face of AD pathology, both present moment and longitudinally.

60 APPLIED LIFE SCIENCES↗

Downhole Sensing and Event-Driven Sensor Fusion for Depth-of-Cut Based Autonomous Fault Response and Drilling Optimization

Achieving robust and efficient drilling is a critical part of reducing the cost of geothermal energy exploration and extraction. Drilling performance is often evaluated using one or more of three key metrics: depth of cut (DOC), rate of penetration (ROP), and mechanical specific energy (MSE). All three of these quantities are related to each other. DOC refers to the depth a bit penetrates into rock during drilling. This is an important quantity for estimating bit behavior. ROP is the simply the DOC multiplied by the rotational rate, and represents how quickly the drill bit is advancing through the ground. ROP is often the parameter used for drilling control and optimization. Finally, MSE provides insight into drilling efficiency and rock type. MSE calculations rely on ROP, drilling force, and drilling torque. Surface-based sensors at the top of the drill are often used to measure all these quantities. However, top-hole measurements can deviate substantially from the behavior at the bit due to lag, vibrations, and friction. Therefore, relying only on top-hole information can lead to suboptimal drilling control. In this work, we describe recent progress towards estimating ROP, DOC, and MSE using down-hole sensing. We assume down-hole measurements of torque, weight-on-bit (WOB). Our hypothesis is that these measurements can provide more rapid and accurate measures of drilling performance. We show how a multi-layer perceptron (MLP) machine learning algorithm can provide rapid and accurate performance when evaluated on experimental data taken from Sandia’s Hard Rock Drilling Facility. In addition, we implement our algorithms on an embedded system intended to emulate a bottom-hole-assembly for sensing and estimation. Our experimental results show that DOC can be estimated accurately and in real-time. These estimates when combined with measurements for rotary speed, torque, and force can provide improved estimates for ROP and MSE. These results have the potential to enable better drilling assessment, improved control, and extended component lifetimes.

15 GEOTHERMAL ENERGY↗

Accelerating Materials Discovery: Artificial Intelligence for Sustainable, High-Performance Polymers

PolyID enables the discovery of polymers with advanced performance and greater sustainability while reducing material development timelines. The material design space is immense and cannot be reasonably probed using an Edisionian approach. High-throughput property prediction, enabled by artificial intelligence provides a hypothesis driven approach for down selection of candidate polymers to pursue experimentally. To aid experimentalists in the down selection of material targets this high-throughput, machine learning-based tool is capable of predicting polymer properties simply from molecular structures. Currently, transport, thermal, and mechanical properties across 7 polymer class (polyamides, polyesters, polycarbonates, polyimides, polyolefins, polyacrylates, and polyurethanes) can be predicted, and the PolyID platform has been flexibly designed so new materials and properties can be added.

artificial intelligence↗

Regional-scale soil carbon predictions can be enhanced by transferring global-scale soil–environment relationships

Accurate modelling and mapping soil organic carbon are crucial for supporting soil health restoration and climate change mitigation at both regional and global scales. However, regional soil predictions often suffer from data scarcity and high prediction uncertainty. Utilizing a pre-trained global-to-regional soil carbon predictive model can be a potential solution to address this challenge. Despite its promise, how to construct and apply the global-scale model to enhance regional-scale soil carbon mapping remains largely unexplored. Here, we propose the Global Soil Carbon Pre-trained Model (GSoilCPM), a deep-learning-based domain adaptative model, to enhance regional-scale soil carbon predictions. Based on large amount of environmental covariate data and 106,167 soil samples across the globe, we verify our hypothesis of the effectiveness of this 'global-to-regional' modelling strategy. The pre-trained model can be then transferred and fine-tuned to bridge the regional- and global-scale soil–environment relationships. We applied and validated this modelling strategy in four regional-scale study areas, three in the Northern Hemisphere and one in the Southern Hemisphere, each with distinct environmental background. Compared to traditional modelling approaches as a baseline, four case studies all demonstrated significant improvement in prediction accuracy across diverse environments and varying data availabilities. The average percentage improvement across all regions is 10.93% (absolute values decreased by 1.20 g kg−1 averagely) in MAE and 29.04% (absolute values increased by 0.10 averagely) in CCC. The applicability and future horizons of using GSoilCPM were further discussed. We further reveal that regions with fewer soil samples or lower baseline accuracy benefit more from the pre-trained global model. Our findings highlight the advantages of leveraging the generalized knowledge from global models to enhance specifically localized soil modelling, positioning a potential paradigm shift in digital soil mapping, and far-reaching implications for soil monitoring and land management.

Deep learning↗

Automated Framework for Groundwater Monitoring Using DWT with LSTM and Transformers

Environmental monitoring is critical for safeguarding public health and ecological well-being. Traditional data structuring and workflow monitoring methods consume significant time and effort, hindering timely insights and effective decision-making. Our study addresses this challenge by presenting an AI framework that automates data cleaning, structuring, and modeling processes, specifically targeting applications in groundwater monitoring. By leveraging automation for data processing and model training, our framework establishes a novel and efficient paradigm for environmental monitoring, with its potential application to the vast network of over a hundred Department of Energy Environmental Management (DoE-EM) cleanup sites across the country. It analyzes data streams from a network of groundwater Internet-of-Things (IoT) sensors deployed at the Savannah River Site (SRS) for prediction modeling. This allows human experts to focus on analysis and decision-making, ultimately leading to better environmental outcomes.The framework employs multivariate time-series forecasting methods to study and model the behavior of varying chemical analytes. The continuous learning process is enabled by utilizing deep learning techniques. It allows the framework to become more nuanced in its analysis over time, adapting to the specific characteristics of the environmental site and the evolving nature of contaminant behavior. Deep learning models known for sequence modeling, LSTM, and Transformers are employed for time series forecasting. Data processing and structuring are essential components significantly impacting the final model's performance. This hypothesis was proven by presenting a comparative analysis of model performance with processed and unprocessed data. The feature engineering approach utilized was the Discrete Wavelet Transform, which works well with time series data.

Discrete Wavelet Transform (DWT)↗

Data-driven causal model discovery and personalized prediction in Alzheimer's disease

Abstract With the explosive growth of biomarker data in Alzheimer’s disease (AD) clinical trials, numerous mathematical models have been developed to characterize disease-relevant biomarker trajectories over time. While some of these models are purely empiric, others are causal, built upon various hypotheses of AD pathophysiology, a complex and incompletely understood area of research. One of the most challenging problems in computational causal modeling is using a purely data-driven approach to derive the model’s parameters and the mathematical model itself, without any prior hypothesis bias. In this paper, we develop an innovative data-driven modeling approach to build and parameterize a causal model to characterize the trajectories of AD biomarkers. This approach integrates causal model learning, population parameterization, parameter sensitivity analysis, and personalized prediction. By applying this integrated approach to a large multicenter database of AD biomarkers, the Alzheimer’s Disease Neuroimaging Initiative, several causal models for different AD stages are revealed. In addition, personalized models for each subject are calibrated and provide accurate predictions of future cognitive status.

Zheng, Haoyang (ORCID:0000000168358242)↗

Spatiotemporal features of traffic help reduce automatic accident detection time

Quick and reliable automatic detection of traffic accidents is of paramount importance to save human lives in transportation systems. However, automatically detecting when accidents occur has proven challenging, and minimizing the time to detect accidents (TTDA) by using traditional features in machine learning (ML) classifiers has plateaued. We hypothesize that accidents affect traffic farther from the accident location than previously reported. Therefore, leveraging traffic signatures from neighboring sensors that are adjacent to accidents should help improve their detection. We confirm this hypothesis by using verified ground-truth accident data, traffic data from radar detection system sensors, and light and weather conditions and show that we can minimize the TTDA while maximizing classification performance by considering spatiotemporal features of traffic. Specifically, we compare the performance of different ML classifiers (i.e, logistic regression, random forest, and XGBoost) when controlling for different numbers of neighboring sensors and TTDA horizons. We use data from interstates 75 and 24 in the metropolitan area that surrounds Chattanooga, TN. Our results show that the XGBoost classifier produces the best results by detecting accidents as quickly as 1.0 min after their occurrence with an area under the receiver operating characteristic curve of up to 83% and an average precision of up to 49%. We describe limitations, open challenges, and how the proposed framework can be used for quicker operational accident detection.

33 ADVANCED PROPULSION SYSTEMS↗

Comparing the effect of virtual and in-person instruction on students’ performance in a design for additive manufacturing learning activity

The goal of this work is to compare the outcome of a design for additive manufacturing (DfAM) heuristics lesson conducted in a virtual learning environment to the same in an in-person learning environment. Prior work revealed that receiving DfAM heuristics at different points in the design process impacts the quality and novelty of designs produced afterward, but this work may have been limited by the solely virtual format. In this work, an identical experiment was performed in a face-to-face learning environment. Results indicate that neither learning format presents an advantage over the other when it comes to the quality of designs produced during the intervention. Participants across all experimental groups reported an increase in self-efficacy after the intervention, with improved performance on quiz-type questions. Furthermore, the novelty and variety of the designs produced by the in-person experimental groups were significantly lower than that of the virtual experimental groups. In addition to validating the effectiveness of virtual instruction as a teaching method, these results also support the authors’ hypothesis that the priming effect is stronger in an in-person classroom than in a virtual classroom.

Design for additive manufacturing↗

Integration of Waveform Simulation Methods

The generation of synthetic seismograms through simulation is a fundamental tool of seismology required to run quantitative hypothesis tests. A variety of approaches have been developed throughout the seismological community and each has their own specific user interface based on their implementation. This causes a challenge to researchers who will need to learn new interfaces with each new software they wish to use and create substantial challenges when attempting to compare results from different tools. Here we provide a unified interface that facilitates interoperability amongst several simulation tools through a modern containerized Python package. Further, this package includes post-processing analysis modules designed to facilitate end-to-end analysis of synthetic seismograms. In this report we present the conceptual guidance and an example implementation of the new Waveform Simulation Framework.

58 GEOSCIENCES↗

Structure of the divergent human astrovirus MLB capsid spike

Despite their worldwide prevalence and association with human disease, the molecular bases of human astrovirus (HAstV) infection and evolution remain poorly characterized. Here, we report the structure of the capsid protein spike of the divergent HAstV MLB clade (HAstV MLB). While the structure shares a similar folding topology with that of classical-clade HAstV spikes, it is otherwise strikingly different. We find no evidence of a conserved receptor-binding site between the MLB and classical HAstV spikes, suggesting that MLB and classical HAstVs utilize different receptors for host-cell attachment. We provide evidence for this hypothesis using a novel HAstV infection competition assay. Comparisons of the HAstV MLB spike structure with structures predicted from its sequence reveal poor matches, but template-based predictions were surprisingly accurate relative to machine-learning-based predictions. Our data provide a foundation for understanding the mechanisms of infection by diverse HAstVs and can support structure determination in similarly unstudied systems.

59 BASIC BIOLOGICAL SCIENCES↗

Aging matrix visualizes complexity of battery aging across hundreds of cycling protocols

To reliably deploy lithium-ion batteries, a fundamental understanding of cycling aging behavior is critical. Battery aging consists of complex and highly coupled phenomena, making it challenging to develop a holistic interpretation. In this work, we generate a diverse battery cycling dataset with a broad range of degradation trajectories, consisting of 359 high energy density commercial Li(Ni,Co,Al)O 2 /graphite + SiO x cylindrical 21 700 cells cycled across 207 unique cycling protocols. We consolidate aging via 16 mechanistic state-of-health (SOH) metrics, including cell-level performance metrics, electrode-specific capacities/state-of-charges (SOCs), and aging trajectory metrics. We develop a framework using interpretable machine learning and explainable features to generate an aging matrix that visually deconvolutes the complex battery degradation behavior. This generalizable data-driven mechanistic framework simplifies the complex interplay between cycling conditions, degradation modes, and SOH, acting as a hypothesis-generation tool to aid battery users in identifying key degradation regimes for further study and experimentation.

25 ENERGY STORAGE↗

Machine learning enables identification of an alternative yeast galactose utilization pathway

How genomic differences contribute to phenotypic differences is a major question in biology. The recently characterized genomes, isolation environments, and qualitative patterns of growth on 122 sources and conditions of 1,154 strains from 1,049 fungal species (nearly all known) in the yeast subphylum Saccharomycotina provide a powerful, yet complex, dataset for addressing this question. We used a random forest algorithm trained on these genomic, metabolic, and environmental data to predict growth on several carbon sources with high accuracy. Known structural genes involved in assimilation of these sources and presence/absence patterns of growth in other sources were important features contributing to prediction accuracy. By further examining growth on galactose, we found that it can be predicted with high accuracy from either genomic (92.2%) or growth data (82.6%) but not from isolation environment data (65.6%). Prediction accuracy was even higher (93.3%) when we combined genomic and growth data. After the GALactose utilization genes, the most important feature for predicting growth on galactose was growth on galactitol, raising the hypothesis that several species in two orders, Serinales and Pichiales (containing the emerging pathogen Candida auris and the genus Ogataea, respectively), have an alternative galactose utilization pathway because they lack the GAL genes. Growth and biochemical assays confirmed that several of these species utilize galactose through an alternative oxidoreductive D-galactose pathway, rather than the canonical GAL pathway. Machine learning approaches are powerful for investigating the evolution of the yeast genotype–phenotype map, and their application will uncover novel biology, even in well-studied traits.

59 BASIC BIOLOGICAL SCIENCES↗