Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “hypothesis learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Accelerating Materials Discovery: Artificial Intelligence for Sustainable, High-Performance Polymers

PolyID enables the discovery of polymers with advanced performance and greater sustainability while reducing material development timelines. The material design space is immense and cannot be reasonably probed using an Edisionian approach. High-throughput property prediction, enabled by artificial intelligence provides a hypothesis driven approach for down selection of candidate polymers to pursue experimentally. To aid experimentalists in the down selection of material targets this high-throughput, machine learning-based tool is capable of predicting polymer properties simply from molecular structures. Currently, transport, thermal, and mechanical properties across 7 polymer class (polyamides, polyesters, polycarbonates, polyimides, polyolefins, polyacrylates, and polyurethanes) can be predicted, and the PolyID platform has been flexibly designed so new materials and properties can be added.

artificial intelligence↗

Regional-scale soil carbon predictions can be enhanced by transferring global-scale soil–environment relationships

Accurate modelling and mapping soil organic carbon are crucial for supporting soil health restoration and climate change mitigation at both regional and global scales. However, regional soil predictions often suffer from data scarcity and high prediction uncertainty. Utilizing a pre-trained global-to-regional soil carbon predictive model can be a potential solution to address this challenge. Despite its promise, how to construct and apply the global-scale model to enhance regional-scale soil carbon mapping remains largely unexplored. Here, we propose the Global Soil Carbon Pre-trained Model (GSoilCPM), a deep-learning-based domain adaptative model, to enhance regional-scale soil carbon predictions. Based on large amount of environmental covariate data and 106,167 soil samples across the globe, we verify our hypothesis of the effectiveness of this 'global-to-regional' modelling strategy. The pre-trained model can be then transferred and fine-tuned to bridge the regional- and global-scale soil–environment relationships. We applied and validated this modelling strategy in four regional-scale study areas, three in the Northern Hemisphere and one in the Southern Hemisphere, each with distinct environmental background. Compared to traditional modelling approaches as a baseline, four case studies all demonstrated significant improvement in prediction accuracy across diverse environments and varying data availabilities. The average percentage improvement across all regions is 10.93% (absolute values decreased by 1.20 g kg−1 averagely) in MAE and 29.04% (absolute values increased by 0.10 averagely) in CCC. The applicability and future horizons of using GSoilCPM were further discussed. We further reveal that regions with fewer soil samples or lower baseline accuracy benefit more from the pre-trained global model. Our findings highlight the advantages of leveraging the generalized knowledge from global models to enhance specifically localized soil modelling, positioning a potential paradigm shift in digital soil mapping, and far-reaching implications for soil monitoring and land management.

Deep learning↗

Working and Learning with Knowledge in the Lobes of a Humanoid's Mind

Humanoid class robots must have sufficient dexterity to assist people and work in an environment designed for human comfort and productivity. This dexterity, in particular the ability to use tools, requires a cognitive understanding of self and the world that exceeds contemporary robotics. Our hypothesis is that the sense-think-act paradigm that has proven so successful for autonomous robots is missing one or more key elements that will be needed for humanoids to meet their full potential as autonomous human assistants. This key ingredient is knowledge. The presented work includes experiments conducted on the Robonaut system, a NASA and the Defense Advanced research Projects Agency (DARPA) joint project, and includes collaborative efforts with a DARPA Mobile Autonomous Robot Software technical program team of researchers at NASA, MIT, USC, NRL, UMass and Vanderbilt. The paper reports on results in the areas of human-robot interaction (human tracking, gesture recognition, natural language, supervised control), perception (stereo vision, object identification, object pose estimation), autonomous grasping (tactile sensing, grasp reflex, grasp stability) and learning (human instruction, task level sequences, and sensorimotor association).

Ambrose, Robert↗

Automated Framework for Groundwater Monitoring Using DWT with LSTM and Transformers

Environmental monitoring is critical for safeguarding public health and ecological well-being. Traditional data structuring and workflow monitoring methods consume significant time and effort, hindering timely insights and effective decision-making. Our study addresses this challenge by presenting an AI framework that automates data cleaning, structuring, and modeling processes, specifically targeting applications in groundwater monitoring. By leveraging automation for data processing and model training, our framework establishes a novel and efficient paradigm for environmental monitoring, with its potential application to the vast network of over a hundred Department of Energy Environmental Management (DoE-EM) cleanup sites across the country. It analyzes data streams from a network of groundwater Internet-of-Things (IoT) sensors deployed at the Savannah River Site (SRS) for prediction modeling. This allows human experts to focus on analysis and decision-making, ultimately leading to better environmental outcomes.The framework employs multivariate time-series forecasting methods to study and model the behavior of varying chemical analytes. The continuous learning process is enabled by utilizing deep learning techniques. It allows the framework to become more nuanced in its analysis over time, adapting to the specific characteristics of the environmental site and the evolving nature of contaminant behavior. Deep learning models known for sequence modeling, LSTM, and Transformers are employed for time series forecasting. Data processing and structuring are essential components significantly impacting the final model's performance. This hypothesis was proven by presenting a comparative analysis of model performance with processed and unprocessed data. The feature engineering approach utilized was the Discrete Wavelet Transform, which works well with time series data.

Discrete Wavelet Transform (DWT)↗

Preventing Failures By Dataset Shift Detection in Safety-Critical Graph Applications

Dataset shift refers to the problem where the input data distribution may change over time (e.g., between training and test stages). Since this can be a critical bottleneck in several safety-critical applications such as healthcare, drug-discovery, etc., dataset shift detection has become an important research issue in machine learning. Though several existing efforts have focused on image/video data, applications with graph-structured data have not received sufficient attention. Therefore, in this paper, we investigate the problem of detecting shifts in graph structured data through the lens of statistical hypothesis testing. Specifically, we propose a practical two-sample test based approach for shift detection in large-scale graph structured data. Our approach is very flexible in that it is suitable for both undirected and directed graphs, and eliminates the need for equal sample sizes. Using empirical studies, we demonstrate the effectiveness of the proposed test in detecting dataset shifts. We also corroborate these findings using real-world datasets, characterized by directed graphs and a large number of nodes.

97 MATHEMATICS AND COMPUTING↗

Data-driven causal model discovery and personalized prediction in Alzheimer's disease

Abstract With the explosive growth of biomarker data in Alzheimer’s disease (AD) clinical trials, numerous mathematical models have been developed to characterize disease-relevant biomarker trajectories over time. While some of these models are purely empiric, others are causal, built upon various hypotheses of AD pathophysiology, a complex and incompletely understood area of research. One of the most challenging problems in computational causal modeling is using a purely data-driven approach to derive the model’s parameters and the mathematical model itself, without any prior hypothesis bias. In this paper, we develop an innovative data-driven modeling approach to build and parameterize a causal model to characterize the trajectories of AD biomarkers. This approach integrates causal model learning, population parameterization, parameter sensitivity analysis, and personalized prediction. By applying this integrated approach to a large multicenter database of AD biomarkers, the Alzheimer’s Disease Neuroimaging Initiative, several causal models for different AD stages are revealed. In addition, personalized models for each subject are calibrated and provide accurate predictions of future cognitive status.

Zheng, Haoyang (ORCID:0000000168358242)↗

Spatiotemporal features of traffic help reduce automatic accident detection time

Quick and reliable automatic detection of traffic accidents is of paramount importance to save human lives in transportation systems. However, automatically detecting when accidents occur has proven challenging, and minimizing the time to detect accidents (TTDA) by using traditional features in machine learning (ML) classifiers has plateaued. We hypothesize that accidents affect traffic farther from the accident location than previously reported. Therefore, leveraging traffic signatures from neighboring sensors that are adjacent to accidents should help improve their detection. We confirm this hypothesis by using verified ground-truth accident data, traffic data from radar detection system sensors, and light and weather conditions and show that we can minimize the TTDA while maximizing classification performance by considering spatiotemporal features of traffic. Specifically, we compare the performance of different ML classifiers (i.e, logistic regression, random forest, and XGBoost) when controlling for different numbers of neighboring sensors and TTDA horizons. We use data from interstates 75 and 24 in the metropolitan area that surrounds Chattanooga, TN. Our results show that the XGBoost classifier produces the best results by detecting accidents as quickly as 1.0 min after their occurrence with an area under the receiver operating characteristic curve of up to 83% and an average precision of up to 49%. We describe limitations, open challenges, and how the proposed framework can be used for quicker operational accident detection.

33 ADVANCED PROPULSION SYSTEMS↗

Comparing the effect of virtual and in-person instruction on students’ performance in a design for additive manufacturing learning activity

The goal of this work is to compare the outcome of a design for additive manufacturing (DfAM) heuristics lesson conducted in a virtual learning environment to the same in an in-person learning environment. Prior work revealed that receiving DfAM heuristics at different points in the design process impacts the quality and novelty of designs produced afterward, but this work may have been limited by the solely virtual format. In this work, an identical experiment was performed in a face-to-face learning environment. Results indicate that neither learning format presents an advantage over the other when it comes to the quality of designs produced during the intervention. Participants across all experimental groups reported an increase in self-efficacy after the intervention, with improved performance on quiz-type questions. Furthermore, the novelty and variety of the designs produced by the in-person experimental groups were significantly lower than that of the virtual experimental groups. In addition to validating the effectiveness of virtual instruction as a teaching method, these results also support the authors’ hypothesis that the priming effect is stronger in an in-person classroom than in a virtual classroom.

Design for additive manufacturing↗

Minimal Experimental Bias on the Hydrogen Bond Greatly Improves Ab Initio Molecular Dynamics Simulations of Water

Experiment Directed Simulations (EDS) is a method within a class of techniques seeking to improve molecular simulations by minimally biasing the system Hamiltonian to reproduce certain experimental observables. In a previous application of EDS to ab initio molecular dynamics (AIMD) simulation based on electronic density functional theory (DFT), the AIMD simulations of water were biased to reproduce its experimentally derived solvation structure. In particular, by solely biasing the O-O pair correlation functions, other structural and dynamical properties that were not biased were improved. In this work, the hypothesis is tested that directly biasing the OH pair correlation (and hence the H-O∙∙∙H hydrogen bonding), will provide an even better improvement of DFT-based water properties in AIMD simulations. The logic behind this hypothesis is that for most electronic DFT descriptions of water the hydrogen bonding is known to be deficient due to anomalous charge transfer and over polarization in the DFT. Using recent advances to the EDS learning algorithm, we thus train a minimal bias on AIMD water that reproduces the O-H radial distribution function derived from the highly accurate MB-pol model of water. Finally, it is then confirmed that biasing the O-H pair correlation alone can lead to improved AIMD water properties, with structural and dynamical properties in even closer to experiment than the previous EDS-AIMD model.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Integration of Waveform Simulation Methods

The generation of synthetic seismograms through simulation is a fundamental tool of seismology required to run quantitative hypothesis tests. A variety of approaches have been developed throughout the seismological community and each has their own specific user interface based on their implementation. This causes a challenge to researchers who will need to learn new interfaces with each new software they wish to use and create substantial challenges when attempting to compare results from different tools. Here we provide a unified interface that facilitates interoperability amongst several simulation tools through a modern containerized Python package. Further, this package includes post-processing analysis modules designed to facilitate end-to-end analysis of synthetic seismograms. In this report we present the conceptual guidance and an example implementation of the new Waveform Simulation Framework.

58 GEOSCIENCES↗

Report on the international Cambodian crater expedition, 1992

It has been proposed that Tonle Sap, a lake in Cambodia, 100 km long and 30 km wide, marks the location of an elongate basin formed by the oblique impact of a comet or asteroid. The impact is considered to have produced melted ejecta found now as tektites over much of southeast Asia and Australia. The location of the lake, its approximate age, its size, and the orientation of its long axis (toward Australia) are consistent with this hypothesis. Our scientific objectives were to find impact or shock metamorphosed rocks unambiguously related to the Tonle Sap basin, to collect samples of rocks that may represent those melted to produce Australasian tektites, and to learn as much as possible about Cambodian geology. Using 1:200,000-scale geologic maps with fairly detailed descriptions of the rock units, we selected a number of acceptable 'phnoms' (hills that rise abruptly out of the surrounding plain) that may contain rocks affected by the postulated Tonle Sap impact. A map of central Cambodia is shown, and the locations of sites where samples were collected are indicated. A list of those sites, together with a description of the rocks reported to be present at each site, is given. No obviously shock-metamorphosed or suevite-like rocks were observed. Recent alluvium surrounding Tonle Sap is judged to be lake sediment deposited when the lake surface was at a higher elevation.

Hartung, J.↗

Structure of the divergent human astrovirus MLB capsid spike

Despite their worldwide prevalence and association with human disease, the molecular bases of human astrovirus (HAstV) infection and evolution remain poorly characterized. Here, we report the structure of the capsid protein spike of the divergent HAstV MLB clade (HAstV MLB). While the structure shares a similar folding topology with that of classical-clade HAstV spikes, it is otherwise strikingly different. We find no evidence of a conserved receptor-binding site between the MLB and classical HAstV spikes, suggesting that MLB and classical HAstVs utilize different receptors for host-cell attachment. We provide evidence for this hypothesis using a novel HAstV infection competition assay. Comparisons of the HAstV MLB spike structure with structures predicted from its sequence reveal poor matches, but template-based predictions were surprisingly accurate relative to machine-learning-based predictions. Our data provide a foundation for understanding the mechanisms of infection by diverse HAstVs and can support structure determination in similarly unstudied systems.

59 BASIC BIOLOGICAL SCIENCES↗

The Regolith of 4 Vesta - Inferences from Howardites

Asteroid 4 Vesta is quite likely the parent asteroid of the howardite, eucrite and diogenite meteorites - the HED clan. Eucrites and diogenites are the products of igneous processes; the former are basaltic composition rocks from flows, and shallow and deep intrusive bodies, whilst the latter are cumulate orthopyroxenites thought to have formed deep in the crust. Impact processes have excavated these materials and mixed them into a suite of polymict breccias. Howardites are polymict breccias composed mostly of clasts and mineral fragments of eucritic and diogenitic parentage, with neither end-member comprising more than 90% of the rock. Early work interpreted howardites as representing the lithified regolith of their parent asteroid. Recently, howardites have been divided into two subtypes; fragmental howardites, being a type of non-regolithic polymict breccia, and regolithic howardites, being lithified remnants of the active regolith of 4 Vesta. We are in the thralls of a collaborative investigation of the record of impact mixing contained within howardites, which includes studies of their mineralogy, petrology, bulk rock compositions, and bulk rock and clast noble gas contents. One goal of our investigation is to test the hypothesis that some howardites represent breccias formed from an ancient, well-mixed regolith on Vesta. Another is to use our results to further understand regolith processing on differentiated asteroids as compared to what has been learned from the Moon. We have made petrographic observations and electron microprobe analyses on 21 howardites and 3 polymict eucrites. We have done bulk rock analyses using X-ray fluorescence spectrometry and are completing inductively coupled plasma mass spectrometry analyses. Here, we discuss our petrologic and bulk compositional results in the context of regolith formation. Companion presentations describe the noble gas results and compositional studies of low-Ca pyroxene clasts.

Mittlefehldt, D. W.↗

Data-driven materials research enabled by natural language processing and information extraction

Given the emergence of data science and machine learning throughout all aspects of society, but particularly in the scientific domain, there is increased importance placed on obtaining data. Data in materials science are particularly heterogeneous, based on the significant range in materials classes that are explored and the variety of materials properties that are of interest. This leads to data that range many orders of magnitude, and these data may manifest as numerical text or image-based information, which requires quantitative interpretation. The ability to automatically consume and codify the scientific literature across domains - enabled by techniques adapted from the field of natural language processing - therefore has immense potential to unlock and generate the rich datasets necessary for data science and machine learning. This review focuses on the progress and practices of natural language processing and text mining of materials science literature and highlights opportunities for extracting additional information beyond text contained in figures and tables in articles. Here, we discuss and provide examples for several reasons for the pursuit of natural language processing for materials, including data compilation, hypothesis development, and understanding the trends within and across fields. Current and emerging natural language processing methods along with their applications to materials science are detailed. We, then, discuss natural language processing and data challenges within the materials science domain where future directions may prove valuable.

36 MATERIALS SCIENCE↗

Aging matrix visualizes complexity of battery aging across hundreds of cycling protocols

To reliably deploy lithium-ion batteries, a fundamental understanding of cycling aging behavior is critical. Battery aging consists of complex and highly coupled phenomena, making it challenging to develop a holistic interpretation. In this work, we generate a diverse battery cycling dataset with a broad range of degradation trajectories, consisting of 359 high energy density commercial Li(Ni,Co,Al)O 2 /graphite + SiO x cylindrical 21 700 cells cycled across 207 unique cycling protocols. We consolidate aging via 16 mechanistic state-of-health (SOH) metrics, including cell-level performance metrics, electrode-specific capacities/state-of-charges (SOCs), and aging trajectory metrics. We develop a framework using interpretable machine learning and explainable features to generate an aging matrix that visually deconvolutes the complex battery degradation behavior. This generalizable data-driven mechanistic framework simplifies the complex interplay between cycling conditions, degradation modes, and SOH, acting as a hypothesis-generation tool to aid battery users in identifying key degradation regimes for further study and experimentation.

25 ENERGY STORAGE↗

Enzyme activities predicted by metabolite concentrations and solvent capacity in the cell

Experimental measurements or computational model predictions of the post-translational regulation of enzymes needed in a metabolic pathway is a difficult problem. Consequently, regulation is mostly known only for well-studied reactions of central metabolism in various model organisms. In this study, we use two approaches to predict enzyme regulation policies and investigate the hypothesis that regulation is driven by the need to maintain the solvent capacity in the cell. The first predictive method uses a statistical thermodynamics and metabolic control theory framework while the second method is performed using a hybrid optimization–reinforcement learning approach. Efficient regulation schemes were learned from experimental data that either agree with theoretical calculations or result in a higher cell fitness using maximum useful work as a metric. As previously hypothesized, regulation is herein shown to control the concentrations of both immediate and downstream product concentrations at physiological levels. Model predictions provide the following two novel general principles: (1) the regulation itself causes the reactions to be much further from equilibrium instead of the common assumption that highly non-equilibrium reactions are the targets for regulation; and (2) the minimal regulation needed to maintain metabolite levels at physiological concentrations maximizes the free energy dissipation rate instead of preserving a specific energy charge. The resulting energy dissipation rate is an emergent property of regulation which may be represented by a high value of the adenylate energy charge. In addition, the predictions demonstrate that the amount of regulation needed can be minimized if it is applied at the beginning or branch point of a pathway, in agreement with common notions. The approach is demonstrated for three pathways in the central metabolism of E. coli (gluconeogenesis, glycolysis-tricarboxylic acid (TCA) and pentose phosphate-TCA) that each require different regulation schemes. It is shown quantitatively that hexokinase, glucose 6-phosphate dehydrogenase and glyceraldehyde phosphate dehydrogenase, all branch points of pathways, play the largest roles in regulating central metabolism.

59 BASIC BIOLOGICAL SCIENCES↗

Machine learning enables identification of an alternative yeast galactose utilization pathway

How genomic differences contribute to phenotypic differences is a major question in biology. The recently characterized genomes, isolation environments, and qualitative patterns of growth on 122 sources and conditions of 1,154 strains from 1,049 fungal species (nearly all known) in the yeast subphylum Saccharomycotina provide a powerful, yet complex, dataset for addressing this question. We used a random forest algorithm trained on these genomic, metabolic, and environmental data to predict growth on several carbon sources with high accuracy. Known structural genes involved in assimilation of these sources and presence/absence patterns of growth in other sources were important features contributing to prediction accuracy. By further examining growth on galactose, we found that it can be predicted with high accuracy from either genomic (92.2%) or growth data (82.6%) but not from isolation environment data (65.6%). Prediction accuracy was even higher (93.3%) when we combined genomic and growth data. After the GALactose utilization genes, the most important feature for predicting growth on galactose was growth on galactitol, raising the hypothesis that several species in two orders, Serinales and Pichiales (containing the emerging pathogen Candida auris and the genus Ogataea, respectively), have an alternative galactose utilization pathway because they lack the GAL genes. Growth and biochemical assays confirmed that several of these species utilize galactose through an alternative oxidoreductive D-galactose pathway, rather than the canonical GAL pathway. Machine learning approaches are powerful for investigating the evolution of the yeast genotype–phenotype map, and their application will uncover novel biology, even in well-studied traits.

59 BASIC BIOLOGICAL SCIENCES↗

Heavy particle irradiation, neurochemistry and behavior: thresholds, dose-response curves and recovery of function

Exposure to heavy particles can affect the functioning of the central nervous system (CNS), particularly the dopaminergic system. In turn, the radiation-induced disruption of dopaminergic function affects a variety of behaviors that are dependent upon the integrity of this system, including motor behavior (upper body strength), amphetamine (dopamine)-mediated taste aversion learning, and operant conditioning (fixed-ratio bar pressing). Although the relationships between heavy particle irradiation and the effects of exposure depend, to some extent, upon the specific behavioral or neurochemical endpoint under consideration, a review of the available research leads to the hypothesis that the endpoints mediated by the CNS have certain characteristics in common. These include: (1) a threshold, below which there is no apparent effect; (2) the lack of a dose-response relationship, or an extremely steep dose-response curve, depending on the particular endpoint; and (3) the absence of recovery of function, such that the heavy particle-induced behavioral and neural changes are present when tested up to one year following exposure. The current report reviews the data relevant to the degree to which these characteristics are common to neurochemical and behavioral endpoints that are mediated by the effects of exposure to heavy particles on CNS activity. c2004 COSPAR. Published by Elsevier Ltd. All rights reserved.

Review↗