Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “causality”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Causal Directions Matter: How Environmental Factors Drive Convective Cloud Detrainment Heights

This study investigates how environmental factors influence the level of maximum detrainment (LMD) in deep convective clouds. Through a novel application of the Linear Non‐Gaussian Acyclic Model (LiNGAM), we discover causal structures between environmental variables and LMD, observed at six tropical sites operated by the Atmospheric Radiation Measurement (ARM) user facility. LiNGAM effectively identifies causal directions among variables of interest, revealing robust relationships such as those among the lifting condensation level (LCL), level of free convection (LFC), and convective inhibition (CIN), aligning with prior knowledge. Relative humidity is shown to directly influence LMD; however, this relationship exhibits strong nonlinearity and becomes difficult to detect when the contrast between oceanic and continental environments is excluded from the analysis. This study highlights the importance of establishing causal relationships before performing statistical inference.

54 ENVIRONMENTAL SCIENCES↗

VALIDATION, VERIFICATION, AND CALIBRATION THROUGH A CAUSAL LENS

This paper presents an alternative method based on causal inference to perform validation, verification, and calibration of simulation models. While classical validation and verification approaches focus on the identification of the associations between data elements using statistical and machine learning methods, the novel methods in this paper focus instead on the identification of causal relationships between data elements. Statistical and machine learning-based approaches are strictly data-driven, meaning that they provide quantitative comparison measures between datasets without explicitly considering the hypotheses behind them. This can lead to the erroneous conclusion that, if two data sets are close enough, then the models that generated them are similar. In addition, when experimental and simulated data differ to an extent that fails to meet the acceptance criteria, calibration techniques are used to tweak simulation model parameters to reduce the gap between the two types of data. This produces the false expectation that a simulation model will match reality. The methods presented in this paper move away from these strictly data-driven methods for validation and calibration toward more robust, model-driven methods based on causal inference. Causal inference aims to identify the possible mechanisms that might have generated data. Thus, this analysis targets the prediction of the effects when one (or more) of the identified mechanisms are altered. There are many approaches to identify, quantify and illustrate causal relationships. For the scope of this paper, directed graphs are employed as causal models. If the directed graph lacks cycles it is known as a directed acyclic graph (DAG). A node in such a graph represents an observed data element while a directed edge connecting two nodes represents a causal relationship between two variables. The developed causal methods are designed to extract causal models from simulation models and from experimental data. Causal models capture the causal relationships between data elements (e.g., simulated and experimental data). In this context, validation and verification are performed by comparing causal models. The proposed approach does not only inform system analysts on how a simulation model matches real-world data, but also identifies elements of the simulation model that should be revised when discrepancies between simulation and experimental data are observed. Through these causal methods, analysts have a means to identify the portion of the model equation(s) that are behind an edge connecting two variables. Hence, once the structural differences between causal models have been determined, model calibration can occur by changing only those model parameters that impact the identified causal relationships.

97 MATHEMATICS AND COMPUTING↗

Verification, Validation, and Calibration Through a Causal Lens

While typical validation and verification approaches focus on identifying the associations between data elements using statistical and machine learning methods, the novel methods in this paper focus instead on identifying causal relationships between data elements. Statistical and machine-learning-based approaches are strictly data-driven, meaning that they provide quantitative comparison measures between data sets without explicitly considering the hypotheses behind them. This can lead to the erroneous conclusion that, if two data sets are close enough, the models that generated them are similar. In addition, when experimental and simulated data differ to an extent that fails to meet the acceptance criteria, calibration techniques are used to tweak simulation model parameters to reduce the gap between the two types of data. This produces the false expectation that a simulation model will match reality. The methods presented in this paper move away from these strictly data-driven methods for validation and calibration toward more robust, model-driven methods based on causal inference. Causal inference aims to identify the possible mechanisms that might have generated data. Thus, this analysis targets the prediction of the effects when one (or more) of the identified mechanisms are altered. There are many approaches to identify, quantify, and illustrate causal relationships. For the scope of this paper, directed graphs are employed as causal models. If the directed graph lacks cycles, it is known as a directed acyclic graph. A node in such a graph represents an observed data element while a directed edge connecting two nodes represents a causal relationship between two variables. The developed causal methods are designed to extract causal models from simulation models and experimental data. Causal models capture the causal relationships between data elements (e.g., simulated and experimental data). In this context, validation and verification are performed by comparing causal models. The proposed approach does not only inform system analysts on how a simulation model matches real-world data, but also identifies elements of the simulation model that should be revised when discrepancies between simulation and experimental data are observed. Through these causal methods, analysts can identify the portion of the model equation(s) that are behind an edge connecting two variables. Hence, once the structural differences between causal models have been determined, model calibration can occur by changing only those model parameters that impact the identified causal relationships.

97 MATHEMATICS AND COMPUTING↗

Explaining and predicting human behavior and social dynamics in simulated virtual worlds: reproducibility, generalizability, and robustness of causal discovery methods

Ground Truth program was designed to evaluate social science modeling approaches using simulation test beds with ground truth intentionally and systematically embedded to understand and model complex Human Domain systems and their dynamics Lazer et al. (Science 369:1060–1062, 2020). Our multidisciplinary team of data scientists, statisticians, experts in Artificial Intelligence (AI) and visual analytics had a unique role on the program to investigate accuracy, reproducibility, generalizability, and robustness of the state-of-the-art (SOTA) causal structure learning approaches applied to fully observed and sampled simulated data across virtual worlds. In addition, we analyzed the feasibility of using machine learning models to predict future social behavior with and without causal knowledge explicitly embedded. In this paper, we first present our causal modeling approach to discover the causal structure of four virtual worlds produced by the simulation teams—Urban Life, Financial Governance, Disaster and Geopolitical Conflict. Our approach adapts the state-of-the-art causal discovery (including ensemble models), machine learning, data analytics, and visualization techniques to allow a human-machine team to reverse-engineer the true causal relations from sampled and fully observed data. We next present our reproducibility analysis of two research methods team’s performance using a range of causal discovery models applied to both sampled and fully observed data, and analyze their effectiveness and limitations. We further investigate the generalizability and robustness to sampling of the SOTA causal discovery approaches on additional simulated datasets with known ground truth. Our results reveal the limitations of existing causal modeling approaches when applied to large-scale, noisy, high-dimensional data with unobserved variables and unknown relationships between them. We show that the SOTA causal models explored in our experiments are not designed to take advantage from vasts amounts of data and have difficulty recovering ground truth when latent confounders are present; they do not generalize well across simulation scenarios and are not robust to sampling; they are vulnerable to data and modeling assumptions, and therefore, the results are hard to reproduce. Finally, when we outline lessons learned and provide recommendations to improve models for causal discovery and prediction of human social behavior from observational data, we highlight the importance of learning data to knowledge representations or transformations to improve causal discovery and describe the benefit of causal feature selection for predictive and prescriptive modeling.

97 MATHEMATICS AND COMPUTING↗

A Method for Validating Causal Diagrams of Human Health Risk in Space Flight

The complexity of cause-and-effect relationships between spaceflight hazards and resulting health conditions clouds understanding of the totality of human system risk in space. In response, NASA has introduced Directed Acyclic Graphs (causal diagrams) into the human systems risk management process. These diagrams allow for a common understanding of the mechanisms that lead from unique hazards of spaceflight to the health outcomes important to agencies and astronauts. However, the paucity of available biomedical data from spaceflight creates a need for methods of validating causal models that can accommodate data from spaceflight model analogs. Here we outline one approach utilizing open-access rodent bone datasets from the Ames Life Sciences Data Archive. The properties of directed acyclic graphs themselves can provide an epistemological and statistical framework for validation of a priori causal representations of human system risk in space flight. The assumed causal connections on the graph creates sets of logical implications: variables that – if the causal diagram is correct – should be correlated, as well as sets that should be conditionally independent. By testing these implied correlations and conditional independencies both statistically and heuristically, we can provide evidence for or against specific causal pathways on the causal diagram. In addition to validation of expert-generated causal diagrams, machine learning techniques can learn the most likely structure of a causal diagram from a given dataset. Comparison with and reconciliation between machine-learned causal diagrams and expert-generated diagrams is another technique for challenging assumptions and improving our understanding of causal mechanisms. Accurately representing complex causation is essential to systemic understanding of human health risks in space travel. Having a robust system of validating causal diagrams helps us arrive at more accurate representations of causal systems. This process will be integral to developing the countermeasures necessary for extended exploration of the moon and Mars.

Robert Reynolds↗

Space‐Time Causal Discovery in Earth System Science: A Local Stencil Learning Approach

Causal discovery tools enable scientists to infer meaningful relationships from observational data, spurring advances in fields as diverse as biology, economics, and climate science. Despite these successes, the application of causal discovery to space-time systems remains immensely challenging due to the high-dimensional nature of the data. For example, in climate sciences, modern observational temperature records over the past few decades regularly measure thousands of locations around the globe. To address these challenges, we introduce Causal Space-Time Stencil Learning (CaStLe), a novel meta-algorithm for discovering causal structures in complex space-time systems. CaStLe leverages regularities in local space-time dependencies to learn governing global dynamics. This local perspective eliminates spurious confounding and drastically reduces sample complexity, making space-time causal discovery practical and effective. For causal discovery, CaStLe flexibly accepts any appropriately adapted time series causal discovery algorithm to recover local causal structures. These advances enable causal discovery of geophysical phenomena that were previously unapproachable, including non-periodic, transient phenomena such as volcanic eruption plumes. Regularities in local space-time dependencies are transformed into informative spatial replicates, which actually improve CaStLe's performance when applied to ever-larger spatial grids. We successfully apply CaStLe to discover the atmospheric dynamics governing the climate response to the 1991 Mount Pinatubo volcanic eruption. We provide validation experiments to demonstrate the effectiveness of CaStLe over existing causal-discovery frameworks on a range of geophysics-inspired benchmarks while identifying the method's limitations and domains where its assumptions may not hold.

Nichol, J. Jake [Univ. of New Mexico, Albuquerque,↗

Granger causal inference for climate change attribution

Abstract Climate change detection and attribution (D&A) is concerned with determining the extent to which anthropogenic activities have influenced specific aspects of the global climate system. D&A fits within the broader field of causal inference, the collection of statistical methods that identify cause and effect relationships. There are a wide variety of methods for making attribution statements, each of which require different types of input data and focus on different types of weather and climate events and each of which are conditional to varying extents. Some methods are based on Pearl causality (direct experimental interference) while others leverage Granger (predictive) causality, and the causal framing provides important context for how the resulting attribution conclusion should be interpreted. However, while Granger-causal attribution analyses have become more common, there is no clear statement of their strengths and weaknesses relative to Pearl-causal attribution and no clear consensus on where and when Granger-causal perspectives are appropriate. In this prospective paper, we provide a formal definition for Granger-based approaches to trend and event attribution and a clear comparison with more traditional methods for assessing the human influence on extreme weather and climate events. Broadly speaking, Granger-causal attribution statements can be constructed quickly from observations and do not require computationally-intesive dynamical experiments. These analyses also enable rapid attribution, which is useful in the aftermath of a severe weather event, and provide multiple lines of evidence for anthropogenic climate change when paired with Pearl-causal attribution. Confidence in attribution statements is increased when different methodologies arrive at similar conclusions. Moving forward, we encourage the D&A community to embrace hybrid approaches to climate change attribution that leverage the strengths of both Granger and Pearl causality.

Risser, Mark D. (ORCID:0000000319561783)↗

Application of advanced causal analyses to identify processes governing secondary organic aerosols

Abstract Understanding how different physical and chemical atmospheric processes affect the formation of fine particles has been a persistent challenge. Inferring causal relations between the various measured features affecting the formation of secondary organic aerosol (SOA) particles is complicated since correlations between variables do not necessarily imply causality. Here, we apply a state-of-the-art information transfer measure coupled with the Koopman operator framework to infer causal relations between isoprene epoxydiol SOA (IEPOX-SOA) and different chemistry and meteorological variables derived from detailed regional model predictions over the Amazon rainforest. IEPOX-SOA represents one of the most complex SOA formation pathways and is formed by the interactions between natural biogenic isoprene emissions and anthropogenic emissions affecting sulfate, acidity and particle water. Since the regional model captures the known relations of IEPOX-SOA with different chemistry and meteorological features, their simulated time series implicitly include their causal relations. We show that our causal model successfully infers the known major causal relations between total particle phase 2-methyl tetrols (the dominant component of IEPOX-SOA over the Amazon) and input features. We provide the first proof of concept that the application of our causal model better identifies causal relations compared to correlation and random forest analyses performed over the same dataset. Our work has tremendous implications, as our methodology of causal discovery could be used to identify unknown processes and features affecting fine particles and atmospheric chemistry in the Earth’s atmosphere.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Benchmarking the PCMCI Causal Discovery Algorithm for Spatiotemporal Systems

Causal discovery algorithms construct hypothesized causal graphs that depict causal dependencies among variables in observational data. While powerful, the accuracy of these algorithms is highly sensitive to the underlying dynamics of the system in ways that have not been fully characterized in the literature. In this report, we benchmark the PCMCI causal discovery algorithm in its application to gridded spatiotemporal systems. Effectively computing grid-level causal graphs on large grids will enable analysis of the causal impacts of transient and mobile spatial phenomena in large systems, such as the Earth’s climate. We evaluate the performance of PCMCI with a set of structural causal models, using simulated spatial vector autoregressive processes in one- and two-dimensions. We develop computational and analytical tools for characterizing these processes and their associated causal graphs. Our findings suggest that direct application of PCMCI is not suitable for the analysis of dynamical spatiotemporal gridded systems, such as climatological data, without significant preprocessing and downscaling of the data. PCMCI requires unrealistic sample sizes to achieve acceptable performance on even modestly sized problems and suffers from a notable curse of dimensionality. This work suggests that, even under generous structural assumptions, significant additional algorithmic improvements are needed before causal discovery algorithms can be reliably applied to grid-level outputs of earth system models.

54 ENVIRONMENTAL SCIENCES↗

Causally‐Informed Deep Learning to Improve Climate Models and Projections

Abstract Climate models are essential to understand and project climate change, yet long‐standing biases and uncertainties in their projections remain. This is largely associated with the representation of subgrid‐scale processes, particularly clouds and convection. Deep learning can learn these subgrid‐scale processes from computationally expensive storm‐resolving models while retaining many features at a fraction of computational cost. Yet, climate simulations with embedded neural network parameterizations are still challenging and highly depend on the deep learning solution. This is likely associated with spurious non‐physical correlations learned by the neural networks due to the complexity of the physical dynamical system. Here, we show that the combination of causality with deep learning helps removing spurious correlations and optimizing the neural network algorithm. To resolve this, we apply a causal discovery method to unveil causal drivers in the set of input predictors of atmospheric subgrid‐scale processes of a superparameterized climate model in which deep convection is explicitly resolved. The resulting causally‐informed neural networks are coupled to the climate model, hence, replacing the superparameterization and radiation scheme. We show that the climate simulations with causally‐informed neural network parameterizations retain many convection‐related properties and accurately generate the climate of the original high‐resolution climate model, while retaining similar generalization capabilities to unseen climates compared to the non‐causal approach. The combination of causal discovery and deep learning is a new and promising approach that leads to stable and more trustworthy climate simulations and paves the way toward more physically‐based causal deep learning approaches also in other scientific disciplines.

Meteorology & Atmospheric Sciences↗

Deep Koopman operators for causal discovery

Causal discovery aims to identify cause-effect mechanisms for better scientific understanding, explainable decision-making, and more accurate modeling. Standard statistical frameworks, such as Granger causality, lack the ability to quantify causal relationships in nonlinear dynamics due to the presence of complex feedback mechanisms, timescale mixing, and nonstationarity. Thus, applying these methods to study causal dynamics in real-world systems, such as the Earth, is a major challenge. Addressing this shortcoming, we leverage deep learning and a Koopman operator-theoretic formalism to present a class of causal discovery algorithms. Kausal uses deep Koopman operator methods to approximate nonlinear dynamics in a linearized vector space in which traditional causal inference methods such as Granger causality can be more easily applied. Our idealized experiments demonstrate Kausal’s superior ability in discovering and characterizing causal signals compared to existing deep learning and non-deep learning state-of-the-art approaches. Finally, the successful identification of major El Niño and La Niña events in observations showcases Kausal’s skill to handle real-world applications.

54 ENVIRONMENTAL SCIENCES↗

Unraveling complex causal processes that affect sustainability requires more integration between empirical and modeling approaches

Scientists seek to understand the causal processes that generate sustainability problems and determine effective solutions. Yet, causal inquiry in nature–society systems is hampered by conceptual and methodological challenges that arise from nature–society interdependencies and the complex dynamics they create. Here, we demonstrate how sustainability scientists can address these challenges and make more robust causal claims through better integration between empirical analyses and process- or agent-based modeling. To illustrate how these different epistemological traditions can be integrated, we present four studies of air pollution regulation, natural resource management, and the spread of COVID-19. The studies show how integration can improve empirical estimates of causal effects, inform future research designs and data collection, enhance understanding of the complex dynamics that underlie observed temporal patterns, and elucidate causal mechanisms and the contexts in which they operate. These advances in causal understanding can help sustainability scientists develop better theories of phenomena where social and ecological processes are dynamically intertwined and prior causal knowledge and data are limited. The improved causal understanding also enhances governance by helping scientists and practitioners choose among potential interventions, decide when and how the timing of an intervention matters, and anticipate unexpected outcomes. Methodological integration, however, requires skills and efforts of all involved to learn how members of the respective other tradition think and analyze nature–society systems.

42 ENGINEERING↗

A causal data fusion method for the general exposure and outcome

Abstract With the advent of the big data era, the need to combine multiple individual data sets to draw causal effects arises naturally in many medical and biological applications. Especially each data set cannot measure enough confounders to infer the causal effect of an exposure on an outcome. In this article, we extend the method proposed by a previous study to causal data fusion of more than two data sets without external validation and to a more general (continuous or discrete) exposure and outcome. Theoretically, we obtain the condition for identifiability of exposure effects using multiple individual data sources for the continuous or discrete exposure and outcome. The simulation results show that our proposed causal data fusion method has unbiased causal effect estimate and higher precision than traditional regression, meta‐analysis and statistical matching methods. We further apply our method to study the causal effect of BMI on glucose level in individuals with diabetes by combining two data sets. Our method is essential for causal data fusion and provides important insights into the ongoing discourse on the empirical analysis of merging multiple individual data sources.

Li, Hongkai↗

Discovering causal structure with reproducing-kernel Hilbert space ε -machines

We merge computational mechanics’ definition of causal states (predictively equivalent histories) with reproducing-kernel Hilbert space (RKHS) representation inference. The result is a widely applicable method that infers causal structure directly from observations of a system’s behaviors whether they are over discrete or continuous events or time. A structural representation—a finite- or infinite-state kernel ϵ-machine—is extracted by a reduced-dimension transform that gives an efficient representation of causal states and their topology. In this way, the system dynamics are represented by a stochastic (ordinary or partial) differential equation that acts on causal states. We introduce an algorithm to estimate the associated evolution operator. Paralleling the Fokker–Planck equation, it efficiently evolves causal-state distributions and makes predictions in the original data space via an RKHS functional mapping. We demonstrate these techniques, together with their predictive abilities, on discrete-time, discrete-value infinite Markov-order processes generated by finite-state hidden Markov models with (i) finite or (ii) uncountably infinite causal states and (iii) continuous-time, continuous-value processes generated by thermally driven chaotic flows. The method robustly estimates causal structure in the presence of varying external and measurement noise levels and for very high-dimensional data.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Path integral suppression of badly behaved causal sets

Abstract Causal set theory is a discrete model of spacetime that retains a notion of causal structure. We understand how to construct causal sets that approximate a given spacetime, but most causal sets are not at all manifold-like, and must be dynamically excluded if something like our Universe is to emerge from the theory. Here we show that the most common of these ‘bad’ causal sets, the Kleitman–Rothschild orders, are strongly suppressed in the gravitational path integral, and we provide evidence that a large class of other ‘bad’ causal sets are similarly suppressed. It thus becomes plausible that continuum behavior could emerge naturally from causal set quantum theory.

Astronomy & Astrophysics↗

Do-calculus enables estimation of causal effects in partially observed biomolecular pathways

Abstract Motivation Estimating causal queries, such as changes in protein abundance in response to a perturbation, is a fundamental task in the analysis of biomolecular pathways. The estimation requires experimental measurements on the pathway components. However, in practice many pathway components are left unobserved (latent) because they are either unknown, or difficult to measure. Latent variable models (LVMs) are well-suited for such estimation. Unfortunately, LVM-based estimation of causal queries can be inaccurate when parameters of the latent variables are not uniquely identified, or when the number of latent variables is misspecified. This has limited the use of LVMs for causal inference in biomolecular pathways. Results In this article, we propose a general and practical approach for LVM-based estimation of causal queries. We prove that, despite the challenges above, LVM-based estimators of causal queries are accurate if the queries are identifiable according to Pearl’s do-calculus and describe an algorithm for its estimation. We illustrate the breadth and the practical utility of this approach for estimating causal queries in four synthetic and two experimental case studies, where structures of biomolecular pathways challenge the existing methods for causal query estimation. Availability and implementation The code and the data documenting all the case studies are available at https://github.com/srtaheri/LVMwithDoCalculus. Supplementary information Supplementary data are available at Bioinformatics online.

59 BASIC BIOLOGICAL SCIENCES↗

Regime-oriented causal model evaluation of Atlantic–Pacific teleconnections in CMIP6

The climate system and its spatio-temporal changes are strongly affected by modes of long-term internal variability, like the Pacific decadal variability (PDV) and the Atlantic multidecadal variability (AMV). As they alternate between warm and cold phases, the interplay between PDV and AMV varies over decadal to multidecadal timescales. Here, we use a causal discovery method to derive fingerprints in the Atlantic–Pacific interactions and to investigate their phase-dependent changes. Dependent on the phases of PDV and AMV, different regimes with characteristic causal fingerprints are identified in reanalyses in a first step. In a second step, a regime-oriented causal model evaluation is performed to evaluate the ability of models participating in the Coupled Model Intercomparison Project Phase 6 (CMIP6) in representing the observed changing interactions between PDV, AMV and their extra-tropical teleconnections. The causal graphs obtained from reanalyses detect a direct opposite-sign response from AMV to PDV when analyzing the complete 1900–2014 period and during several defined regimes within that period, for example, when AMV is going through its negative (cold) phase. Reanalyses also demonstrate a same-sign response from PDV to AMV during the cold phase of PDV. Historical CMIP6 simulations exhibit varying skill in simulating the observed causal patterns. Generally, large-ensemble (LE) simulations showed better network similarity when PDV and AMV were out of phase compared to other regimes. Also, the two largest ensembles (in terms of number of members) were found to contain realizations with similar causal fingerprints to observations. For most regimes, these same models showed higher network similarity when compared to each other. This work shows how causal discovery on LEs complements the available diagnostics and statistical metrics of climate variability to provide a powerful tool for climate model evaluation.

54 ENVIRONMENTAL SCIENCES↗

Observational process data analytics using causal inference

Voluminous process data are available with the paradigm shift toward smart manufacturing. However, most historical data are observational, containing noncausal correlations due to confounders and mediators. Estimating causal effects from observational data remains a bottleneck in leveraging them for active applications such as optimization and control. Further, this work aims to introduce a causal modeling framework for analyzing observational process data and extracting quantitative causal information. We demonstrate a real-world application in steel manufacturing where causal inference is used to analyze observational production data and improve the steelmaking process. Additionally, we propose a novel formulation for identifying critical process parameters from observational data, where causal inference is combined with variance-based methods to estimate corresponding risks of interventions to the manufacturing system. The proposed methods are compared with statistical ones to illustrate that causally interpreting statistical correlation leads to problematic results, while the provided workflow generates satisfactory strategies for process improvement.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗