Engineering PapersSearch

SEARCH · Engineering Papers

Results for “Causal inference”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Biological Research and Space Health Enabled by Machine Learning to Support Deep Space Missions

A key science goal of the NASA “Moon to Mars” campaign is to understand how biology responds to the Lunar, Martian, and deep space environments in order to advance fundamental knowledge, reduce risk, and support safe, productive human space missions. Through the powerful emerging computer science approaches of artificial intelligence (AI) and machine learning (ML), a paradigm shift has begun in biomedical science and engineered astronaut health systems, to enable Earth-independence and autonomy of mission operations. We present a decadal view of AI/ML architecture to support deep space mission goals, developed in concert with leaders in the field. We describe current AI/ML methods to support 1) fundamental biology, 2) in situ analytics, 3) high performance computing hardware, 4) automated science, 5) self-driving labs, 6) remote data management, 7) integrated real-time mission biomonitoring, and 8) a Precision Space Health system. Cutting-edge AI/ML approaches that can be integrated to support these domains include active learning, explainable AI, adaptive learning, causal inference, knowledge graphs, federated learning, transfer learning, and large language models. Finally, we present results from several current ML projects that are underway in the field to address key challenges of small sample n, high feature count, heterogeneity, and sparse data. These include 1) connecting omics data to phenotypic data using an ensemble model to infer causality of spaceflight rodent liver health disruption, 2) usage of explainable ML to interrogate the muscular underpinnings of spaceflight muscle atrophy, 3) ML models analyzing and determining directed acyclic graphs of human space health risk leveraging rodent bone datasets, 4) usage of large pre-trained models connecting biomedical knowledgebases with small spaceflight datasets to understand gene-to-gene interaction networks, and 5) a suite of benchmarked open science datasets (spaceflight mouse liver; radiation DNA damage) enabling programmers to identify the best ML algorithms to answer space biological science questions.

space biology

Biological Research and Space Health Enabled by Machine Learning to Support Deep Space Missions

A key science goal of the NASA “Moon to Mars” campaign is to understand how biology responds to the Lunar, Martian, and deep space environments in order to advance fundamental knowledge and support human space missions. Through artificial intelligence (AI) and machine learning (ML), a paradigm shift has begun in space biosciences and engineered astronaut health systems, to enable Earth-independence and mission operations autonomy. We describe current AI/ML methods to support 1) fundamental biology, 2) in situ analytics, 3) high performance computing, 4) automated science, 5) self-driving labs, 6) remote data management, 7) integrated mission biomonitoring, and 8) a Precision Space Health system. AI/ML approaches that can be integrated to support these domains include active learning, explainable AI, adaptive learning, causal inference, knowledge graphs, federated learning, transfer learning, and large language models. Finally, we present results from several current ML projects that are underway in the space biology field to address key challenges of small sample n, high feature count, heterogeneity, and sparse data. These include 1) connecting omics to phenotypic data using an ensemble model to infer causality of rodent liver health disruption, 2) usage of explainable ML to interrogate muscular underpinnings of muscle atrophy, 3) ML models analyzing and determining directed acyclic graphs of human health risk leveraging rodent bone datasets, 4) usage of large pre-trained models connecting biomedical knowledgebases with small spaceflight datasets to understand gene-to-gene interactions, and 5) a suite of benchmarked open science datasets enabling programmers to identify best algorithms to answer space biology questions.

space biology

Biological Research and Space Health Enabled by Machine Learning to Support Deep Space Missions

A key science goal of the NASA “Moon to Mars” campaign is to understand how biology responds to the Lunar, Martian, and deep space environments in order to advance fundamental knowledge, reduce risk, and support safe, productive human space missions. Through the powerful emerging computer science approaches of artificial intelligence (AI) and machine learning (ML), a paradigm shift has begun in biomedical science and engineered astronaut health systems, to enable Earth-independence and autonomy of mission operations. We present a decadal view of AI/ML architecture to support deep space mission goals, developed in concert with leaders in the field. We describe current AI/ML methods to support 1) fundamental biology, 2) in situ analytics, 3) high performance computing hardware, 4) automated science, 5) self-driving labs, 6) remote data management, 7) integrated real-time mission biomonitoring, and 8) a Precision Space Health system. Cutting-edge AI/ML approaches that can be integrated to support these domains include active learning, explainable AI, adaptive learning, causal inference, knowledge graphs, federated learning, transfer learning, and large language models. Finally, we present results from several current ML projects that are underway in the field to address key challenges of small sample n, high feature count, heterogeneity, and sparse data. These include 1) connecting omics data to phenotypic data using an ensemble model to infer causality of spaceflight rodent liver health disruption, 2) usage of explainable ML to interrogate the muscular underpinnings of spaceflight muscle atrophy, 3) ML models analyzing and determining directed acyclic graphs of human space health risk leveraging rodent bone datasets, 4) usage of large pre-trained models connecting biomedical knowledgebases with small spaceflight datasets to understand gene-to-gene interaction networks, and 5) a suite of benchmarked open science datasets (spaceflight mouse liver; radiation DNA damage) enabling programmers to identify the best ML algorithms to answer space biological science questions.

space biology

Incorporating Biological Knowledge into Evaluation of Casual Regulatory Hypothesis

Biological data can be scarce and costly to obtain. The small number of samples available typically limits statistical power and makes reliable inference of causal relations extremely difficult. However, we argue that statistical power can be increased substantially by incorporating prior knowledge and data from diverse sources. We present a Bayesian framework that combines information from different sources and we show empirically that this lets one make correct causal inferences with small sample sizes that otherwise would be impossible.

Chrisman, Lonnie

Using Federated Learning to Overcome Data Gravity in Space

Humans intend to take longer missions to outer space. Understanding the impact that space has on human health is paramount to the success of these missions. Controlled experiments with model organisms are run to infer the impact of space conditions on human health, but the data these experiments generate are too large to transfer to Earth for building models. The same is true for space-relevant data generated on Earth. Ideally, these datasets should be combined to improve statistical power and model accuracy without having to transfer data. Federated learning is such a method which trains an algorithm across decentralized computing systems, each of which has their own local copy of training and testing data. In this research, made possible by NASA@Work, the AI for Life in Space group at NASA demonstrates the use of federated learning to train an ensemble of causality inference models on a combination of data residing on the International Space Station (ISS) and in the cloud. Our work leverages CRISP, a causal inference platform developed during the 2020 Frontier Development Lab’s “Astronaut Health Challenge.” We also leverage the OpenFL federated learning library which was collaboratively developed at Intel and UPenn. We used publicly available data from the NASA Ames Life Sciences Data Archive to identify features in ionizing radiation experiments as causal of changes in cardiac blood velocity. This research demonstrates, for the first time, the possibility of running machine learning algorithms on datasets separated by astronomical distances. In this experiment, all the data were generated in terra, half of which were transferred to the ISS and analyzed on the Spaceborne Computer. In the future, our research will leverage federated learning on data generated in situ on the ISS with data generated terrestrially to predict the impact of spaceflight on mammalian female reproductive capacity.

James Casaletto

Opportunistic Experiments to Constrain Aerosol Effective Radiative Forcing

Aerosol-cloud interactions (ACI) are considered to be the most uncertain driver of present-day radiative forcing due to human activities. The non-linearity of cloud-state changes to aerosol perturbations make it challenging to attribute causality in observed relationships of aerosol radiative forcing. Using correlations to infer causality can be challenging when meteorological variability also drives both aerosol and cloud changes independently. Natural and anthropogenic aerosol perturbations from well defined sources provide ‘opportunistic experiments’ (also known as natural experiments) to investigate ACI in cases where causality may be more confidently inferred. These perturbations cover a wide range of locations and spatio-temporal scales, including point sources such as volcanic eruptions or industrial sources, plumes from biomass burning or forest fires, and tracks from individual ships or shipping corridors.We review the different experimental conditions and conduct a synthesis of the available satellite data sets and field campaigns to place these opportunistic experiments on a common footing, facilitating new insights and a clearer understanding of key uncertainties in aerosol radiative forcing. Cloud albedo perturbations are strongly sensitive to background meteorological conditions. Strong liquid water path increases due to aerosol perturbations are largely ruled out by averaging across experiments. Opportunistic experiments have significantly improved process-level understanding of ACI, but it remains unclear how reliably the relationships found can be scaled to the global level, thus demonstrating a need for deeper investigation in order to improve assessments of aerosol radiative forcing and climate change.

Matthew W Christensen

Prediction and causal reasoning in planning

Nonlinear planners are often touted as having an efficiency advantage over linear planners. The reason usually given is that nonlinear planners, unlike their linear counterparts, are not forced to make arbitrary commitments to the order in which actions are to be performed. This ability to delay commitment enables nonlinear planners to solve certain problems with far less effort than would be required of linear planners. Here, it is argued that this advantage is bought with a significant reduction in the ability of a nonlinear planner to accurately predict the consequences of actions. Unfortunately, the general problem of predicting the consequences of a partially ordered set of actions is intractable. In gaining the predictive power of linear planners, nonlinear planners sacrifice their efficiency advantage. There are, however, other advantages to nonlinear planning (e.g., the ability to reason about partial orders and incomplete information) that make it well worth the effort needed to extend nonlinear methods. A framework is supplied for causal inference that supports reasoning about partially ordered events and actions whose effects depend upon the context in which they are executed. As an alternative to a complete but potentially exponential-time algorithm, researchers provide a provably sound polynomial-time algorithm for predicting the consequences of partially ordered events.

Dean, T.

Genetic Network Inference: From Co-Expression Clustering to Reverse Engineering

Advances in molecular biological, analytical, and computational technologies are enabling us to systematically investigate the complex molecular processes underlying biological systems. In particular, using high-throughput gene expression assays, we are able to measure the output of the gene regulatory network. We aim here to review datamining and modeling approaches for conceptualizing and unraveling the functional relationships implicit in these datasets. Clustering of co-expression profiles allows us to infer shared regulatory inputs and functional pathways. We discuss various aspects of clustering, ranging from distance measures to clustering algorithms and multiple-duster memberships. More advanced analysis aims to infer causal connections between genes directly, i.e., who is regulating whom and how. We discuss several approaches to the problem of reverse engineering of genetic networks, from discrete Boolean networks, to continuous linear and non-linear models. We conclude that the combination of predictive modeling with systematic experimental verification will be required to gain a deeper insight into living organisms, therapeutic targeting, and bioengineering.

Dhaeseleer, Patrik

InvestigationOrganizer: The Development and Testing of a Web-based Tool to Support Mishap Investigations

InvestigationOrganizer (IO) is a collaborative web-based system designed to support the conduct of mishap investigations. IO provides a common repository for a wide range of mishap related information, and allows investigators to make explicit, shared, and meaningful links between evidence, causal models, findings and recommendations. It integrates the functionality of a database, a common document repository, a semantic knowledge network, a rule-based inference engine, and causal modeling and visualization. Thus far, IO has been used to support four mishap investigations within NASA, ranging from a small property damage case to the loss of the Space Shuttle Columbia. This paper describes how the functionality of IO supports mishap investigations and the lessons learned from the experience of supporting two of the NASA mishap investigations: the Columbia Accident Investigation and the CONTOUR Loss Investigation.

Carvalho, Robert F.

Robust Strategy for Rocket Engine Health Monitoring

Monitoring the health of rocket engine systems is essentially a two-phase process. The acquisition phase involves sensing physical conditions at selected locations, converting physical inputs to electrical signals, conditioning the signals as appropriate to establish scale or filter interference, and recording results in a form that is easy to interpret. The inference phase involves analysis of results from the acquisition phase, comparison of analysis results to established health measures, and assessment of health indications. A variety of analytical tools may be employed in the inference phase of health monitoring. These tools can be separated into three broad categories: statistical, rule based, and model based. Statistical methods can provide excellent comparative measures of engine operating health. They require well-characterized data from an ensemble of "typical" engines, or "golden" data from a specific test assumed to define the operating norm in order to establish reliable comparative measures. Statistical methods are generally suitable for real-time health monitoring because they do not deal with the physical complexities of engine operation. The utility of statistical methods in rocket engine health monitoring is hindered by practical limits on the quantity and quality of available data. This is due to the difficulty and high cost of data acquisition, the limited number of available test engines, and the problem of simulating flight conditions in ground test facilities. In addition, statistical methods incur a penalty for disregarding flow complexity and are therefore limited in their ability to define performance shift causality. Rule based methods infer the health state of the engine system based on comparison of individual measurements or combinations of measurements with defined health norms or rules. This does not mean that rule based methods are necessarily simple. Although binary yes-no health assessment can sometimes be established by relatively simple rules, the causality assignment needed for refined health monitoring often requires an exceptionally complex rule base involving complicated logical maps. Structuring the rule system to be clear and unambiguous can be difficult, and the expert input required to maintain a large logic network and associated rule base can be prohibitive.

Santi, L. Michael

Phenomena Associated with EIT Waves

We discuss phenomena associated with 'EIT Wave' transients. These phenomena include coronal mass ejections, flares, EUV/SXR dimmings, chromospheric waves, Moreton waves, solar energetic particle events, energetic electron events, and radio signatures. Although the occurrence of many phenomena correlate with the appearance of EIT waves, it is difficult to infer which associations are causal. The presentation will include a discussion of correlation surveys of these phenomena.

Thompson, B. J.

Log-Based Recovery in Asynchronous Distributed Systems

A log-based mechanism is described for restoring consistent states to replicated data objects after failures. Preserving a causal form of consistency based on the notion of virtual time is focused upon in this report. Causal consistency has been shown to apply to a variety of applications, including distributed simulation, task decomposition, and mail delivery systems. Several mechanisms have been proposed for implementing causally consistent recovery, most notably those of Strom and Yemini, and Johnson and Zwaenepoel. The mechanism proposed here differs from these in two major respects. First, a roll-forward style of recovery is implemented. A functioning process is never required to roll-back its state in order to achieve consistency with a recovering process. Second, the mechanism does not require any explicit information about the causal dependencies between updates. Instead, all necessary dependency information is inferred from the orders in which updates are logged by the object servers. This basic recovery technique appears to be applicable to forms of consistency other than causal consistency. In particular, it is shown how the recovery technique can be modified to support an atomic form of consistency (grouping consistency). By combining grouping consistency with casual consistency, it may even be possible to implement serializable consistency within this mechanism.

Kane, Kenneth Paul

Untangling causality in midlatitude aerosol–cloud adjustments

Aerosol–cloud interactions represent the leading uncertainty in our ability to infer climate sensitivity from the observational record. The forcing from changes in cloud albedo driven by increases in cloud droplet number (Nd) (the first indirect effect) is confidently negative and has narrowed its probable range in the last decade, but the sign and strength of forcing associated with changes in cloud macrophysics in response to aerosol (aerosol–cloud adjustments) remain uncertain. This uncertainty reflects our inability to accurately quantify variability not associated with a causal link flowing from the cloud microphysical state to the cloud macrophysical state. Once variability associated with meteorology has been removed, covariance between the liquid water path (LWP) averaged across cloudy and clear regions (here characterizing the macrophysical state) and Nd (characterizing the microphysical) is the sum of two causal pathways linking Nd to LWP: Nd altering LWP (adjustments) and precipitation scavenging aerosol and thus depleting Nd. Only the former term is relevant to constraining adjustments, but disentangling these terms in observations is challenging. We hypothesize that the diversity of constraints on aerosol–cloud adjustments in the literature may be partly due to not explicitly characterizing covariance flowing from cloud to aerosol and aerosol to cloud. Here, we restrict our analysis to the regime of extratropical clouds outside of low-pressure centers associated with cyclonic activity. Observations from MAC-LWP (Multisensor Advanced Climatology of Liquid Water Path) and MODIS are compared to simulations in the Met Office Unified Model (UM) GA7.1 (the atmosphere model of HadGEM3-GC3.1 and UKESM1). The meteorological predictors of LWP are found to be similar between the model and observations. There is also agreement with previous literature on cloud-controlling factors finding that increasing stability, moisture, and sensible heat flux enhance LWP, while increasing subsidence and sea surface temperature decrease it. A simulation where cloud microphysics are insensitive to changes in Nd is used to characterize covariance between Nd and LWP that is induced by factors other than aerosol–cloud adjustments. By removing variability associated with meteorology and scavenging, we infer the sensitivity of LWP to changes in Nd. Application of this technique to UM GA7.1 simulations reproduces the true model adjustment strength. Observational constraints developed using simulated covariability not induced by adjustments and observed covariability between Nd and LWP predict a 25 %–30 % overestimate by the UM GA7.1 in LWP change and a 30 %–35 % overestimate in associated radiative forcing.

Daniel T. McCoy

EBIC investigation of hydrogenation of crystal defects in EFG solar silicon ribbons

Changes in the contrast and resolution of defect structures in 205 Ohm-cm EFG polysilicon ribbon subjected to annealing and hydrogenation treatments were observed in a JEOL 733 Superprobe scanning electron microscope, using electron beam induced current (EBIC) collected at an A1 Schottky barrier. The Schottky barrier was formed by evaporation of A1 onto the cleaned and polished surface of the ribbon material. Measurement of beam energy, beam current, and the current induced in the Schottky diode enabled observations to be quantified. Exposure to hydrogen plasma increased charge collection efficiency. However, no simple causal relationship between the hydrogenation and charge collection efficiency could be inferred, because the collection efficiency also displayed an unexpected thermal dependence. Good quality intermediate-magnification (1000X-5400X) EBIC micrographs of several specific defect structures were obtained. Comparison of grown-in and stress-induced dislocations after annealing in vacuum at 500 C revealed that stress-induced dislocations are hydrogenated to a much greater degree than grown-in dislocations. The theoretical approximations used to predict EBIC contrast and resolution may not be entirely adequate to describe them under high beam energy and low beam current conditions.

Sullivan, T.

Inference of precipitation through thermal infrared measurements of soil moisture

The physics of microwave radiative transfer is well understood so that causal models can be assembled which relate the observed brightness temperatures to assumed distributions of hydrometeors (both liquid and ice), non-precipitating clouds, water vapor oxygen, and surface conditions. Present models assume a Marshall Palmer size distribution of liquid hydrometers from the surface to the freezing level (near the 0 C isotherm) and a variable thickness of frozen hydrometeors above that with various reasonable distribution of the other relevant constituents. The validity of such models is discussed. All uncertainties in the rain rate retrieval algorithms can be expressed in terms of specific model uncertainties which can be addressed through appropriate measurements. Those factors which must be known to achieve umambiguous results can be identified so that rainfall measuring algorithms can be developed and improved. The emissivity of the underlying surface significantly affects the contrast that may be measured between areas covered by rain and those which are dry. Sensing strategies for measuring rain over the ocean and rain over land are reviewed.

Wetzel, P. J.

Systems Modeling to Implement Integrated System Health Management Capability

ISHM capability includes: detection of anomalies, diagnosis of causes of anomalies, prediction of future anomalies, and user interfaces that enable integrated awareness (past, present, and future) by users. This is achieved by focused management of data, information and knowledge (DIaK) that will likely be distributed across networks. Management of DIaK implies storage, sharing (timely availability), maintaining, evolving, and processing. Processing of DIaK encapsulates strategies, methodologies, algorithms, etc. focused on achieving high ISHM Functional Capability Level (FCL). High FCL means a high degree of success in detecting anomalies, diagnosing causes, predicting future anomalies, and enabling health integrated awareness by the user. A model that enables ISHM capability, and hence, DIaK management, is denominated the ISHM Model of the System (IMS). We describe aspects of the IMS that focus on processing of DIaK. Strategies, methodologies, and algorithms require proper context. We describe an approach to define and use contexts, implementation in an object-oriented software environment (G2), and validation using actual test data from a methane thruster test program at NASA SSC. Context is linked to existence of relationships among elements of a system. For example, the context to use a strategy to detect leak is to identify closed subsystems (e.g. bounded by closed valves and by tanks) that include pressure sensors, and check if the pressure is changing. We call these subsystems Pressurizable Subsystems. If pressure changes are detected, then all members of the closed subsystem become suspect of leakage. In this case, the context is defined by identifying a subsystem that is suitable for applying a strategy. Contexts are defined in many ways. Often, a context is defined by relationships of function (e.g. liquid flow, maintaining pressure, etc.), form (e.g. part of the same component, connected to other components, etc.), or space (e.g. physically close, touching the same common element, etc.). The context might be defined dynamically (if conditions for the context appear and disappear dynamically) or statically. Although this approach is akin to case-based reasoning, we are implementing it using a software environment that embodies tools to define and manage relationships (of any nature) among objects in a very intuitive manner. Context for higher level inferences (that use detected anomalies or events), primarily for diagnosis and prognosis, are related to causal relationships. This is useful to develop root-cause analysis trees showing an event linked to its possible causes and effects. The innovation pertaining to RCA trees encompasses use of previously defined subsystems as well as individual elements in the tree. This approach allows more powerful implementations of RCA capability in object-oriented environments. For example, if a pressurizable subsystem is leaking, its root-cause representation within an RCA tree will show that the cause is that all elements of that subsystem are suspect of leak. Such a tree would apply to all instances of leak-events detected and all elements in all pressurizable subsystems in the system. Example subsystems in our environment to build IMS include: Pressurizable Subsystem, Fluid-Fill Subsystem, Flow-Thru-Valve Subsystem, and Fluid Supply Subsystem. The software environment for IMS is designed to potentially allow definition of any relationship suitable to create a context to achieve ISHM capability.

Figueroa, Jorge F.

Accident/Mishap Investigation System

InvestigationOrganizer (IO) is a Web-based collaborative information system that integrates the generic functionality of a database, a document repository, a semantic hypermedia browser, and a rule-based inference system with specialized modeling and visualization functionality to support accident/mishap investigation teams. This accessible, online structure is designed to support investigators by allowing them to make explicit, shared, and meaningful links among evidence, causal models, findings, and recommendations.

Keller, Richard

Temperature Dependence of Factors Controlling Isoprene Emissions

We investigated the relationship of variability in the formaldehyde (HCHO) columns measured by the Aura Ozone Monitoring Instrument (OMI) to isoprene emissions in the southeastern United States for 2005-2007. The data show that the inferred, regional-average isoprene emissions varied by about 22% during summer and are well correlated with temperature, which is known to influence emissions. Part of the correlation with temperature is likely associated with other causal factors that are temperature-dependent. We show that the variations in HCHO are convolved with the temperature dependence of surface ozone, which influences isoprene emissions, and the dependence of the HCHO column to mixed layer height as OMI's sensitivity to HCHO increases with altitude. Furthermore, we show that while there is an association of drought with the variation in HCHO, drought in the southeastern U.S. is convolved with temperature.

Duncan, Bryan N.