Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “hypothesis learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

A Systems Biology Approach to Identify Essential Epigenetic Regulators for Specific Biological Processes in Plants

Upon sensing developmental or environmental cues, epigenetic regulators transform the chromatin landscape of a network of genes to modulate their expression and dictate adequate cellular and organismal responses. Knowledge of the specific biological processes and genomic loci controlled by each epigenetic regulator will greatly advance our understanding of epigenetic regulation in plants. To facilitate hypothesis generation and testing in this domain, we present EpiNet, an extensive gene regulatory network (GRN) featuring epigenetic regulators. EpiNet was enabled by (i) curated knowledge of epigenetic regulators involved in DNA methylation, histone modification, chromatin remodeling, and siRNA pathways; and (ii) a machine-learning network inference approach powered by a wealth of public transcriptome datasets. We applied GENIE3, a machine-learning network inference approach, to mine public Arabidopsis transcriptomes and construct tissue-specific GRNs with both epigenetic regulators and transcription factors as predictors. The resultant GRNs, named EpiNet, can now be intersected with individual transcriptomic studies on biological processes of interest to identify the most influential epigenetic regulators, as well as predicted gene targets of the epigenetic regulators. We demonstrate the validity of this approach using case studies of shoot and root apical meristem development.

root apical meristem↗

From sequence to protein structure and conformational dynamics with artificial intelligence/machine learning

The 2024 Nobel Prize in Chemistry was awarded in part for de novo protein structure prediction using AlphaFold2, an artificial intelligence/machine learning (AI/ML) model trained on vast amounts of sequence and three-dimensional structure data. AlphaFold2 and related models, including RoseTTAFold and ESMFold, employ specialized neural network architectures driven by attention mechanisms to infer relationships between sequence and structure. At a fundamental level, these AI/ML models operate on the long-standing hypothesis that the structure of a protein is determined by its amino acid sequence. More recently, AlphaFold2 has been adapted for the prediction of multiple protein conformations by subsampling multiple sequence alignments. Herein, we provide an overview of the deterministic relationship between sequence and structure, which was hypothesized over half a century ago with profound implications for the biological sciences ever since. We postulate that protein conformational dynamics are also determined, at least in part, by amino acid sequence and that this relationship may be leveraged for construction of AI/ML models dedicated to predicting protein conformational ensembles. Accordingly, we describe a conceptual model architecture, which may be trained on sequence data in combination with conformationally sensitive structural information, coming primarily from nuclear magnetic resonance (NMR) spectroscopy. Notwithstanding certain limitations in this context, NMR offers abundant structural heterogeneity conducive to conformational ensemble prediction. As NMR and other data continue to accumulate, sequence-informed prediction of protein structural dynamics with AI/ML has the potential to emerge as a transformative capability across the biological sciences.

Artificial intelligence↗

Quantitation of chlorpromazine-bound calmodulin during chlorpromazine inhibition of gravitropism

The regulatory protein, calmodulin (CaM), controls the activity of a plasma membrane localized ATPase in plants which serves to pump calcium out of cells. Recent data are consistent with the hypothesis that activation of this pump is one of the early steps necessary for gravitropism. Chlorpromazine (CPZ), a CaM antagonist, reversibly inhibits gravitropism in oat coleoptiles at concentrations which permit normal growth rates. C-14-labeled CPZ was used to photo-affinity label endogenous CaM in vivo to learn whether the drug is actually binding to some portion of endogenous CaM when it inhibits gravitropism. Under conditions in which CPZ inhibits gravitropism for over an hour, at least 11% of the CaM in gravitropically stimulated coleoptiles is bound to CPZ. In a given CPZ experiment the degree of inhibition of gravitropism correlates well with the amount of CaM bound to CPZ.

Roux, S. J.↗

Preprocessing for Unintended Conducted Emissions Classification with ResNet

Characterization of Unintended Conducted Emissions (UCE) from electronic devices is important when diagnosing electromagnetic interference, performing nonintrusive load monitoring (NILM) of power systems, and monitoring electronic device health, among other applications. Prior work has demonstrated that UCE analysis can serve as a diagnostic tool for energy efficiency investigations and detailed load analysis. While explaining the feature selection of deep networks with certainty is often not fully comprehensive, or in other applications, quite lacking, additional tools/methods for further corroboration and confirmation can help further the understanding of the researcher. This is true especially in the subject application of the study in this paper. Often the focus of such efforts is the selected features themselves, and there is not as much understanding gained about the noise in the collected data. If selected feature and noise characteristics are known, it can be used to further shape the design of the deep network or associated preprocessing. This is additionally difficult when the available data are limited, as in the case which the authors investigated in this study. Here, the authors present a novel work (which is a proposed complementary portion of the overall solution to the deep network classification explainability problem for this application) by applying a systematic progression of preprocessing and a deep neural network (ResNet architecture) to classify UCE data obtained via current transformers. By using a methodical application of preprocessing techniques prior to a deep classifier, hypotheses can be produced concerning what features the deep network deems important relative to what it perceives as noise. For instance, it is hypothesized in this particular study as a result of execution of the proposed method and periodic inspection of the classifier output that the UCE spectral features are relatively close to each other or to the interferers, as systematically reducing the beta parameter of the Kaiser window produced progressively better classification performance, but only to a point, as going below the Beta of eight produced decreased classifier performance, as well as the hypothesis that further spectral feature resolution was not as important to the classifier as rejection of the leakage from a spectrally distant interference. This can be very important in unpredictable low-FNR applications, where knowing the difference between features and noise is difficult. As a side-benefit, much was learned regarding the best preprocessing to use with the selected deep network for the UCE collected from these low power consumer devices obtained via current transformers. Baseline rectangular windowed FFT preprocessing provided a 62% classification increase versus using raw samples. After performing a more optimal preprocessing, more than 90% classification accuracy was achieved across 18 low-power consumer devices for scenarios in which the in-band features-to-noise ratio (FNR) was very poor.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

A Mass‐Conserving‐Perceptron for Machine‐Learning‐Based Modeling of Geoscientific Systems

Although decades of effort have been devoted to building Physical-Conceptual (PC) models for predicting the time-series evolution of geoscientific systems, recent work shows that Machine Learning (ML) based Gated Recurrent Neural Network technology can be used to develop models that are much more accurate. However, the difficulty of extracting physical understanding from ML-based models complicates their utility for enhancing scientific knowledge regarding system structure and function. Here, we propose a physically interpretable Mass-Conserving-Perceptron (MCP) as a way to bridge the gap between PC-based and ML-based modeling approaches. The MCP exploits the inherent isomorphism between the directed graph structures underlying both PC models and GRNNs to explicitly represent the mass-conserving nature of physical processes while enabling the functional nature of such processes to be directly learned (in an interpretable manner) from available data using off-the-shelf ML technology. As a proof of concept, we investigate the functional expressivity (capacity) of the MCP, explore its ability to parsimoniously represent the rainfall-runoff (RR) dynamics of the Leaf River Basin, and demonstrate its utility for scientific hypothesis testing. To conclude, we discuss extensions of the concept to enable ML-based physical-conceptual representation of the coupled nature of mass-energy-information flows through geoscientific systems.

58 GEOSCIENCES↗

A physics‐informed learning technique for fault location of DC microgrids using traveling waves

Abstract Fast and accurate fault location in DC power systems is of particular importance to ensure their reliable operation. One of the approaches for implementing a fast‐tripping protection scheme is to use Traveling waves (TW) initiated by a fault scenario. This paper proposes a physics‐informed machine learning approach that utilizes TWs for fault location in DC microgrids. TWs are extracted by the so‐called multiresolution analysis which identifies the TW's wavelet coefficients for multiple frequency ranges. This paper deploys Parseval's theorem to find the energy of wavelet coefficients as a quantitative metric for describing TWs. The hypothesis of this paper is that once the Parseval energy curves for a specific cable are extracted, they can be utilized to locate faults along with that cable regardless of the DC system in which the cable is deployed. The fault location algorithm uses Parseval energy curves to train a Gaussian Process (GP) estimator. With the Parseval energy values of measured current at the protection device location, the GP estimator is able to estimate fault locations with high accuracy. The effectiveness of the proposed algorithm is verified by simulating a DC microgrid system in PSCAD/EMTDC.

Paruthiyil, Sajay Krishnan↗

Evolution of the Earth and Origin of Life: The Role of Gas/Fluid Interactions with Rocks

The work under the Cooperative Agreement will be centered on questions of the evolution of Life on the early Earth and possibly on Mars. It is still hotly debated whether the essential organic molecules were delivered to the early Earth from space (by comets, meteorites or interplanetary dust particles) or were generated in situ on Earth. Prior work that has shown that the matrix of igneous minerals is a medium in which progenitors of organic molecules assemble from H2O, C02 and N2 incorporated as minority "impurities" in minerals of igneous rocks during crystallization from H2O/CO2/N2-laden magmas. The underlying processes involve a redox. conversion whereby C, H, and N become chemically reduced, while 0 becomes oxidized to the peroxy state. During Year 02 the work will be divided into three tasks. Task 1: After carboxylic (fatty) acids and N-bearing compounds have been identified, other extractable organic molecules including lipids, oily substances and amino acids will be studied. Dedicated lipid analysis will be combined with gas chromatographic-mass spectroscopic (GCMS) analysis of organic compounds extracted from minerals and rocks. Task 2: Using infrared (IR) spectroscopy, C-H entities that are indicators for the organic progenitors in mineral matrices will be studied. A preliminary heating experiment with MgO single crystals has shown that the C-H entities can be pyrolyzed, causing the IR bands to disappear, but at room temperature the IR bands reappear in a matter of days to weeks. This work will be expanded, both by studying synthetic MgO crystals and olivine crystals from the Earth's upper mantle. The C-H bands will be compared to the published "organic" IR feature of dust in the interstellar medium (ISM) and interplanetary dust particles (IDP). Task 3: A paradox marks the evolution of early Life: Oxygen is highly toxic to primitive life, yet early organisms "learned" to detoxify reactive oxygen species, to utilize oxygen, and even produce it. Why would organisms on the early anaerobic Earth be under evolutionary pressure to evolve defenses against reactive oxygen species? Minerals in igneous rocks are now known to contain peroxy. When such minerals weather, the peroxy hydrolyzes to H2O2. The hypothesis will be tested whether organisms living in intimate contact with rock surfaces are subjected to a constant trickle of H202 and thus under stress to develop strategies to either detoxify the reactive oxygen species or repair the molecular damage that they cause. Understanding these processes is central to the Astrobiology mission. It opens new avenues toward understanding the evolution of early life on Earth, and the potential for aerobic life elsewhere. This Cooperative Agreement also has a strong educational and public outreach component involving high school, undergraduate students, and high school teachers.

Freund, Friedemann↗

Signatures of a liquid–liquid transition in an ab initio deep neural network model for water

Significance Water is central across much of the physical and biological sciences and exhibits physical properties that are qualitatively distinct from those of most other liquids. Understanding the microscopic basis of water’s peculiar properties remains an active area of research. One intriguing hypothesis is that liquid water can separate into metastable high- and low-density liquid phases at low temperatures and high pressures, and the existence of this liquid–liquid transition could explain many of water’s anomalous properties. We used state-of-the-art approaches in computational quantum chemistry, statistical mechanics, and machine learning and obtained evidence consistent with a liquid–liquid transition, supporting the argument for the existence of this phenomenon in real water.

36 MATERIALS SCIENCE↗

It takes a village: using a crowdsourced approach to investigate organic matter composition in global rivers through the lens of ecological theory

Though community-based scientific approaches are becoming more common, many scientific efforts are conducted by small groups of researchers that together develop a concept, analyze data, and interpret results that ultimately translate into a publication. Here, we present a community effort that breaks these traditional boundaries of the publication process by engaging the scientific community from initial hypothesis generation to final publication. We leverage community-generated data from the Worldwide Hydrobiogeochemistry Observation Network for Dynamic River Systems (WHONDRS) consortium to study organic matter composition through the lens of ecological theory. This community endeavor will use a suite of paired physical and chemical datasets collected from 97 river corridors across the globe. With our first step aimed at ideation, we engaged a community of scientists from 20 countries and 60 institutions, spanning disciplines and career stages by holding a virtual workshop (April 2021). In the workshop, participants generated content for questions, hypotheses, and proposed analyses based on the WHONDRS dataset. These ideation efforts resulted in several narratives investigating different questions led by different teams, which will be the basis for research articles in a Frontiers in Water collection. Currently, the community is collectively analyzing, interpreting, and synthesizing these data that will result in seven crowdsourced articles using a single, existing WHONDRS dataset. The use of a shared dataset across articles not only lowers barriers for broad participation by not requiring generation of new data, but also provides unique opportunities for emergent learning by connecting outcomes across studies. Here we will explain methods used to enable this community endeavor aimed to promote a greater diversity of thinking on river corridor biogeochemistry through community science.

Borton, Mikayla A.↗

Advancing river corridor science beyond disciplinary boundaries with an inductive approach to catalyse hypothesis generation

Abstract A unified conceptual framework for river corridors requires synthesis of diverse site‐, method‐ and discipline‐specific findings. The river research community has developed a substantial body of observations and process‐specific interpretations, but we are still lacking a comprehensive model to distill this knowledge into fundamental transferable concepts. We confront the challenge of how a discipline classically organized around the deductive model of systematically collecting of site‐, scale‐, and mechanism‐specific observations begins the process of synthesis. Machine learning is particularly well‐suited to inductive generation of hypotheses. In this study, we prototype an inductive approach to holistic synthesis of river corridor observations, using support vector machine regression to identify potential couplings or feedbacks that would not necessarily arise from classical approaches. This approach generated 672 relationships linking a suite of 157 variables each measured at 62 locations in a fifth order river network. Eighty four percent of these relationships have not been previously investigated, and representing potential (hypothetical) process connections. We document relationships consistent with current understanding including hydrologic exchange processes, microbial ecology, and the River Continuum Concept, supporting that the approach can identify meaningful relationships in the data. Moreover, we highlight examples of two novel research questions that stem from interpretation of inductively‐generated relationships. This study demonstrates the implementation of machine learning to sieve complex data sets and identify a small set of candidate relationships that warrant further study, including data types not commonly measured together. This structured approach complements traditional modes of inquiry, which are often limited by disciplinary perspectives and favour the careful pursuit of parsimony. Finally, we emphasize that this approach should be viewed as a complement to, rather than in place of, more traditional, deductive approaches to scientific discovery.

54 ENVIRONMENTAL SCIENCES↗

Improving Slip Prediction on Mars Using Thermal Inertia Measurements

Rovers operating on Mars have been delayed, diverted, and trapped by loose granular materials. Vision-based mobility prediction approaches often fail because hazardous sand is difficult to distinguish from safe sand based on surface appearance alone. Unlike surface appearance, the thermal inertia of terrain is directly correlated to the same geophysical properties that control slip. This paper presents a quantitative analysis showing that considering thermal inertia improves rover slip prediction on Mars using in-situ data from the Curiosity rover. Thermal inertia is estimated for each slip measurement in sand using both on-board and orbital instruments. Slip models are learned using a mixture of experts approach where the experts are identified using thermal inertia. Two-expert models are compared to a single-expert, vision-only model to show that slip predictions are improved by separating high-slip, low thermal inertia sand from low-slip, high thermal inertia sand. These results support the hypothesis that the consideration of thermal inertia improves mobility estimates for rovers on Mars.

Nesnas, Issa A.↗

A Mass Conservation Relaxed (MCR) LSTM Model for Streamflow Simulation Across CONUS

The recent development of the physics-aware Mass-Conserving Long Short-Term Memory network (MC-LSTM) provides an alternative to other data-driven Deep Learning (DL) models in hydrology. Mass-Conserving Long Short-Term Memory incorporates mass conservation directly into the LSTM architecture. Despite the theoretical advancements, studies have reported a surprisingly limited performance of the MC-LSTM in streamflow simulation. We hypothesize that such a limitation is due to the unrealistic mass conservation scheme in MC-LSTM, which overlooks unobserved incoming water fluxes beyond precipitation. As an attempt to verify this hypothesis, we propose a Mass Conservation Relaxed LSTM (MCR-LSTM), which incorporates a bi-directional mass relaxation (MR) component to account for potential incoming water fluxes beyond precipitation. We train and test the proposed MCR-LSTM model across 531 watersheds in the contiguous United States (CONUS) against three baseline models: the Sacramento Soil Moisture Accounting, LSTM, and MC-LSTM. Our results show that MCR-LSTM outperforms MC-LSTM despite its underperformance compared to LSTM. Specifically, MCR-LSTM's advantage over MC-LSTM is mainly seen in the Plains and Western U.S., where the newly incorporated MR component better simulates water loss and suggests the likely existence of additional incoming water fluxes beyond precipitation, respectively. The novelty and contribution of this study are twofold: firstly, it introduces an alternative physics-aware DL tool (i.e., MCR-LSTM) in hydrology with higher accuracy in specific regions compared to MC-LSTM. Secondly, it provides a diagnosis of regions where strict, precipitation-based mass conservation constraints may be unrealistic in streamflow simulation.

deep learning↗

A Framework for Deep Learning Emulation of Numerical Models With a Case Study in Satellite Remote Sensing

Numerical models based on physics represent the state of the art in Earth system modeling and comprise our best tools for generating insights and predictions. Despite rapid growth in computational power, the perceived need for higher model resolutions overwhelms the latest generation computers, reducing the ability of modelers to generate simulations for understanding parameter sensitivities and characterizing variability and uncertainty. Thus, surrogate models are often developed to capture the essential attributes of the full-blown numerical models. Recent successes of machine learning methods, especially deep learning (DL), across many disciplines offer the possibility that complex nonlinear connectionist representations may be able to capture the underlying complex structures and nonlinear processes in Earth systems. A difficult test for DL-based emulation, which refers to function approximation of numerical models, is to understand whether they can be comparable to traditional forms of surrogate models in terms of computational efficiency while simultaneously reproducing model results in a credible manner. A DL emulation that passes this test may be expected to perform even better than simple models with respect to capturing complex processes and spatiotemporal dependencies. Here, we examine, with a case study in satellite-based remote sensing, the hypothesis that DL approaches can credibly represent the simulations from a surrogate model with comparable computational efficiency. Our results are encouraging in that the DL emulation reproduces the results with acceptable accuracy and often even faster performance. We discuss the broader implications of our results in light of the pace of improvements in high-performance implementations of DL and the growing desire for higher resolution simulations in the Earth sciences.

Bayesian Deep Learning↗

Microglia are implicated in the development of paclitaxel chemotherapy-associated cognitive impairment in female mice

Chemotherapy remains a mainstay in the treatment of many types of cancer even though it is associated with debilitating behavioral side effects referred to as “chemobrain,” including difficulty concentrating and memory impairment. The predominant hypothesis in the field is that systemic inflammation drives these cognitive impairments, although the brain mechanisms by which this occurs remain poorly understood. Here, we hypothesized that microglia are activated by chemotherapy and drive chemotherapy-associated cognitive impairments. To test this hypothesis, we treated female C57BL/6 mice with a clinically-relevant regimen of a common chemotherapeutic, paclitaxel (6 i.p. doses at 30 mg/kg), which impairs memory of an aversive stimulus as assessed via a contextual fear conditioning (CFC) paradigm. In this work, paclitaxel increased the percent area of IBA1 staining in the dentate gyrus of the hippocampus. Moreover, using a machine learning random forest classifier we identified immunohistochemical features of reactive microglia in multiple hippocampal subregions that were distinct between vehicle- and paclitaxel-treated mice. Paclitaxel treatment also increased gene expression of inflammatory cytokines in a microglia-enriched population of cells from mice. Lastly, a selective inhibitor of colony stimulating factor 1 receptor, PLX5622, was employed to deplete microglia and then assess CFC performance following paclitaxel treatment. PLX5622 significantly reduced hippocampal gene expression of paclitaxel-induced proinflammatory cytokines and restored memory, suggesting that microglia play a critical role in the development of chemotherapy-associated neuroinflammation and cognitive impairments. This work provides critical evidence that microglia drive paclitaxel-associated cognitive impairments, a key mechanistic detail for determining preventative and intervention strategies for these burdensome side effects.

60 APPLIED LIFE SCIENCES↗

Graph-based machine learning improves just-in-time defect prediction

The increasing complexity of today’s software requires the contribution of thousands of developers. This complex collaboration structure makes developers more likely to introduce defect-prone changes that lead to software faults. Determining when these defect-prone changes are introduced has proven challenging, and using traditional machine learning (ML) methods to make these determinations seems to have reached a plateau. In this work, we build contribution graphs consisting of developers and source files to capture the nuanced complexity of changes required to build software. By leveraging these contribution graphs, our research shows the potential of using graph-based ML to improve Just-In-Time (JIT) defect prediction. We hypothesize that features extracted from the contribution graphs may be better predictors of defect-prone changes than intrinsic features derived from software characteristics. We corroborate our hypothesis using graph-based ML for classifying edges that represent defect-prone changes. This new framing of the JIT defect prediction problem leads to remarkably better results. We test our approach on 14 open-source projects and show that our best model can predict whether or not a code change will lead to a defect with an F1 score as high as 77.55% and a Matthews correlation coefficient (MCC) as high as 53.16%. This represents a 152% higher F1 score and a 3% higher MCC over the state-of-the-art JIT defect prediction. We describe limitations, open challenges, and how this method can be used for operational JIT defect prediction.

59 BASIC BIOLOGICAL SCIENCES↗

Training data selection for event classification in a highly variable environment

A problem of interest for nuclear nonproliferation is monitoring activities at nuclear facilities, where proliferation events may only take place a few times and often under variable conditions. Machine learning has revolutionized data analytics by enabling the use of measurable signatures to generate predictive models of facility operations. However, traditional methods for training these models require large, reliable data sets with labeled observations, a challenge for nonproliferation. Highly variable conditions further complicate this as events from training data may have occurred in conditions quite different from the event of interest. Our hypothesis is that when events occur in a highly variable environment, careful training data selection for each test event could outperform the standard approach of using all available training data. We developed a method to optimize training data selection for the given test event and applied it to predicting the power level of the High Flux Isotope Reactor (HFIR) at Oak Ridge National Laboratory. In this study, the reactor startup exhibits variability between occurrences due to natural variability in environmental conditions and operational procedures. Using a combination of analysis techniques, a similitude assessment was performed on data collected from HFIR to isolate clusters that were optimal for training a predictive model. Concepts such as dynamic time warping and Jaccard similarity were used in conjunction with clustering analysis. In order to validate this approach, the model was trained on every combination of unique training events and the predictive performance was compared to the performance using a subset of the training data selected by isolated clusters found through the similitude assessment.

Iyer, A↗

Bayesian learning

In 1983 and 1984, the Infrared Astronomical Satellite (IRAS) detected 5,425 stellar objects and measured their infrared spectra. In 1987 a program called AUTOCLASS used Bayesian inference methods to discover the classes present in these data and determine the most probable class of each object, revealing unknown phenomena in astronomy. AUTOCLASS has rekindled the old debate on the suitability of Bayesian methods, which are computationally intensive, interpret probabilities as plausibility measures rather than frequencies, and appear to depend on a subjective assessment of the probability of a hypothesis before the data were collected. Modern statistical methods have, however, recently been shown to also depend on subjective elements. These debates bring into question the whole tradition of scientific objectivity and offer scientists a new way to take responsibility for their findings and conclusions.

Denning, Peter J.↗

Observation of 𝑡⁢𝑊⁢𝑍 Production at the CMS Experiment

The first observation of single top quark production in association with a 𝑊 and a 𝑍 boson in proton-proton collisions is reported. The analysis uses data at center-of-mass energies of 13 and 13.6 TeV recorded with the CMS detector at the CERN LHC, corresponding to a total integrated luminosity of 200 fb −1 . Events with three or four charged leptons, which can be electrons or muons, are selected. Advanced machine-learning algorithms and improved reconstruction methods, compared to an earlier analysis, result in an unprecedented sensitivity to 𝑡⁢𝑊⁢𝑍 production. The measured cross sections for 𝑡⁢𝑊⁢𝑍 production are 248 ± 52 fb and 242 ± 77 fb for $\sqrt{s}$ =13 and 13.6 TeV, respectively. The signal is established with a statistical significance of 5.8 standard deviations, with 3.5 expected, compared to the background-only hypothesis.

Hayrapetyan, Aram [Yerevan Physics Institute]↗