Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “causality”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 325 records · Page 18

VAINE: Visualization and AI for Natural Experiments

Natural experiments are observational studies where the assignment of treatment conditions to different populations occur by chance ``in the wild''. Researchers from fields such as economics, healthcare, and the social sciences leverage natural experiments to conduct hypothesis testing and causal effect estimation for treatment and outcome variables that would otherwise be costly, infeasible, or unethical. In this paper, we introduce VAINE (Visualization and AI for Natural Experiments), a visual analytics tool for identifying and understanding natural experiments from observational data. We then demonstrate how VAINE can be used to validate causal relationships, estimate average treatment effects, and identify statistical phenomena such as Simpson’s paradox through two use cases.

Guo, Grace↗

Sustainable Development Tool Using Meta‐Analysis and DPSIR Framework — Application to Savannah River Basin, U.S.

Abstract The Savannah River Basin (SRB), a highly stressed southeastern river in United States is a conservation priority for State, Federal government, and nongovernment organizations. A four‐stage sustainable development tool was developed in this study using meta‐analysis and the drivers–pressures–state–impacts–responses (DPSIR) framework. Through the synthesis of ~150 references in the SRB this study addressed three research questions: (1) What were the drivers, pressures, state, impacts, and responses (components of DRSIR framework) in SRB (2) Can these components be grouped together from various studies in SRB (3) Can causal chain/loops be developed, and will they be useful for policy and decision making? First in the Stage 1, the state of the SRB was represented (S component of DPSIR), in Stage 2, the drivers–pressures–impacts–responses (DPIR components of DPSIR) were represented, in the third stage (Stage 3) the common units characterizing each DPSIR component were identified. Finally, in Stage 4, the causal chains/loops were developed and organized into scientific research at a level appropriate for building better understanding about SRB and helping stakeholders and policy makers in managing basin sustainability challenges. Although the tool was applied to SRB, the methodology is applicable to other river basins and ecosystems.

Pagan, Janeesa↗

Salinity–induced limits to mangrove canopy height

Aim: Mangrove canopy height is a key metric to assess tidal forests' resilience in the face of climate change. In terrestrial forests, tree height is primarily determined by water availability, plant hydraulic design, and disturbance regime. However, the role of water stress remains elusive in tidal environments, where saturated soils are prevalent, and salinity can substantially affect the soil water potential. Location: Global. Time Period: The canopy height dataset provides a global snapshot of the maximum mangrove height geographical distribution for the year 2000. Climate and environmental variables extend over the period 1970–2018. Major Taxa Studied: Mangroves. Methods: We use global observations of maximum canopy height, species richness, air temperature, and seawater salinity—a proxy of soil water salt concentration—to explore the causal link between salinity and mangrove stature. Results: Our findings suggest that salt stress limits mangrove height. High salinity favours more salt-tolerant species, narrowing the spectrum of viable traits. Highly salt-tolerant mangroves have evolved to cope with high salt concentrations in the soil, but this adaptation comes at a cost. They typically have lower rates of photosynthesis and growth, resulting in reduced productivity and smaller stature compared to more salt-sensitive mangrove species. This suggests a causal link between salinity, biodiversity, and tree height, where high salinity selects for more salt-tolerant species that tend to be less productive and shorter. Conclusions: We hypothesize that the salinity-induced limit to mangrove canopy height is the direct result of a reduction of primary productivity, an increment in the risk of xylem cavitation, and an indirect consequence of the decrease in biodiversity. As sea-level rise enhances coastal salinisation, failure to account for these effects can lead to incorrect estimates of future carbon stocks in Tropical coastal ecosystems and endanger preservation efforts.

54 ENVIRONMENTAL SCIENCES↗

A noncoding single-nucleotide polymorphism at 8q24 drives IDH1 -mutant glioma formation

Establishing causal links between inherited polymorphisms and cancer risk is challenging. Here, we focus on the single-nucleotide polymorphism rs55705857, which confers a sixfold greater risk of isocitrate dehydrogenase (IDH)–mutant low-grade glioma (LGG). Here, we reveal that rs55705857 itself is the causal variant and is associated with molecular pathways that drive LGG. Mechanistically, we show that rs55705857 resides within a brain-specific enhancer, where the risk allele disrupts OCT2/4 binding, allowing increased interaction with the Myc promoter and increased Myc expression. Mutating the orthologous mouse rs55705857 locus accelerated tumor development in an Idh1 R132H -driven LGG mouse model from 472 to 172 days and increased penetrance from 30% to 75%. Our work reveals mechanisms of the heritable predisposition to lethal glioma in ~40% of LGG patients.

59 BASIC BIOLOGICAL SCIENCES↗

Diversity and scale: Genetic architecture of 2068 traits in the VA Million Veteran Program

One of the justifiable criticisms of human genetic studies is the underrepresentation of participants from diverse populations. Lack of inclusion must be addressed at-scale to identify causal disease factors and understand the genetic causes of health disparities. We present genome-wide associations for 2068 traits from 635,969 participants in the Department of Veterans Affairs Million Veteran Program, a longitudinal study of diverse United States Veterans. Systematic analysis revealed 13,672 genomic risk loci; 1608 were only significant after including non-European populations. Fine-mapping identified causal variants at 6318 signals across 613 traits. One-third (n = 2069) were identified in participants from non-European populations. This reveals a broadly similar genetic architecture across populations, highlights genetic insights gained from underrepresented groups, and presents an extensive atlas of genetic associations.

59 BASIC BIOLOGICAL SCIENCES↗

GNATFinder

SAND2023-05510O GNATFinder is a software tool for extracting graphical neural activity threads (GNATs) from spiking neural networks. GNATFinder efficiently extracts GNATs from spiking neural network data. A GNAT is a graph that represents the causal relationships between individual spikes in a spiking network using a combination of the network connectivity structure and the timing of individual spikes. The software allows for identifying similar, repeating threads that often reappear in a dataset. Using quadtree data structures to recursively subdivide the time interval over which spikes were produced by the spiking network, GNATFinder allows for efficient pairwise comparisons of spike times to determine the degree of causal relatedness among all pairs of spikes in a dataset. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

SciDAC↗

CRISPR-CARB/nocap

Network Optimization and Causal Analysis of Perturb-seq (NOCAP) is a software package for causal inference of gene regulation networks using data from perturb-seq.

George, August↗

Simplified Approximations of Direct Cumulus Entrainment and Detrainment

Abstract In recent years, direct calculations of simulated cumulus entrainment and detrainment have facilitated new physical insights into these highly elusive but critically important processes. However, these calculations require substantial computational resources that may limit their widespread usage. To facilitate such calculations, two simplified approximations of direct cumulus entrainment and detrainment are examined herein. The first approximation, termed the “semidirect” method, follows a standard bulk approach but makes more realistic assumptions about the sources of entrained and detrained air near the cloud edges. In contrast, the second approximation (the “projection” method) uses the governing equations of motion to project whether grid points near the cloud edge will entrain or detrain as the mean cloud ascends by one grid point. Verification exercises using large-eddy simulations reveal that both methods generally agree better with corresponding direct entrainment/detrainment estimates than the traditional bulk formulation, with the projection method outperforming the semidirect method. The two methods can be used in a synergistic fashion, with the semidirect method helping to optimize the projection method, to suit a wide range of applications. Because the latter incorporates the essential dynamics of entrainment and detrainment at the local scale, it can be used to gain physical insight into the causal mechanisms regulating these complex processes.

Meteorology & Atmospheric Sciences↗

A systematic decision-making methodology to formalize the selection of degree of realism in screening analysis of probabilistic risk assessment

In the nuclear power domain, Probabilistic Risk Assessment (PRA) is used to inform decision-making for Nuclear Power Plants (NPPs). Recently, there has been an increase in the utilization of modeling and simulation (M&S) to support the estimation of PRA inputs. Risk analysts should carefully select the PRA items that require M&S and their degree of realism (DoR) with consideration of the required resources. To support this selection, this article formulates a systematic decision-making approach for the DoR selection. The DoR selection is made based on two predictive decision-making attributes: the predicted differences in safety risk estimate (ΔSaRi) and the cost of analysis (ΔCAN). This research also develops and quantifies causal models to estimate ΔSaRi and ΔCAN. The causal model-based prediction of ΔSaRi and ΔCAN helps reduce the trial-and-error nature of the DoR selection in the PRA screening analysis and provides insights for DoR selection and the gradual refinements of PRA realism. This approach is demonstrated for a case study on fire PRA of NPPs, where an adequate DoR is selected from two fire models: an engineering correlation and a zone model.

Alkhatib, Sari [Department of Nuclear, Plasma, and↗

Decision-making based on Markov decision process in integrated artificial reasoning framework—Part I: Theory

This paper presents a decision-making framework based on an integrated artificial reasoning framework and Markov decision process (MDP). The integrated artificial reasoning framework provides a physics-based approach that converts system information into state transition models, and the analysis result will be represented by the transition probabilities that can be used with an MDP to find a traceable and explainable optimal pathway. A dynamic Bayesian network (DBN) is well suited for representing the structure of an MDP. The causality information among process variables (or among subsystems) is mathematically represented in a DBN by the conditional probabilities of the node’s states provided different probabilities of the parent node’s states. To define node states in a physically understandable manner, we used multilevel flow modeling (MFM). An MFM follows the fundamental energy and mass conservation laws and supports the selection of process variables that represent the system of interest so that causal relations among process variables are properly captured. An MFM-based DBN supports developing state transition models in an MDP to capture the effect of process variables of system having physical relations. The operators of the target system can capture stochastic system dynamics as multiple subsystem state transitions based on their physical relations and uncertainties coming from component degradation or random failures. We analyzed a simplified exemplary system to illustrate an optimal operational policy using the suggested approach.

Markov decision process↗

Training Population Optimization for Genomic Selection in Miscanthus

Miscanthus is a perennial grass with potential for lignocellulosic ethanol production. To ensure its utility for this purpose, breeding efforts should focus on increasing genetic diversity of the nothospecies Miscanthus × giganteus (M×g) beyond the single clone used in many programs. Germplasm from the corresponding parental species M. sinensis (Msi) and M. sacchariflorus (Msa) could theoretically be used as training sets for genomic prediction of M×g clones with optimal genomic estimated breeding values for biofuel traits. To this end, we first showed that subpopulation structure makes a substantial contribution to the genomic selection (GS) prediction accuracies within a 538-member diversity panel of predominately Msi individuals and a 598-member diversity panels of Msa individuals. We then assessed the ability of these two diversity panels to train GS models that predict breeding values in an interspecific diploid 216-member M×g F2 panel. Low and negative prediction accuracies were observed when various subsets of the two diversity panels were used to train these GS models. To overcome the drawback of having only one interspecific M×g F2 panel available, we also evaluated prediction accuracies for traits simulated in 50 simulated interspecific M×g F2 panels derived from different sets of Msi and diploid Msa parents. The results revealed that genetic architectures with common causal mutations across Msi and Msa yielded the highest prediction accuracies. Ultimately, these results suggest that the ideal training set should contain the same causal mutations segregating within interspecific M×g populations, and thus efforts should be undertaken to ensure that individuals in the training and validation sets are as closely related as possible.

59 BASIC BIOLOGICAL SCIENCES↗

Constraining braneworlds with entanglement entropy

We propose swampland criteria for braneworlds viewed as effective field theories of defects coupled to semiclassical gravity. We do this by exploiting their holographic interpretation. We focus on general features of entanglement entropies and their holographic calculations. Entropies have to be positive. Furthermore, causality imposes certain constraints on the surfaces that are used holographically to compute them, most notably a property known as causal wedge inclusion. As a test case, we explicitly constrain the Dvali-Gabadadze-Porrati term as a second-order-in-derivatives correction to the Randall-Sundrum action. We conclude by discussing the implications of these criteria for the question on whether entanglement islands in theories with massless gravitons are possible in Karch-Randall braneworlds.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

IC Project: w19_hossfault “Modelling of stick-slip behavior in sheared granular fault gouge & nonlinear elasticity behavior in cracked solid”

Low frequency earthquakes, non-volcanic tremor, and acoustic emissions are examples of weak seismic signals that may help detect major seismic events, i.e., earthquakes. In a laboratory setting, in which stick-slip events are simulated, acoustic emissions are detected far from the stick-slip events. In both, field and laboratory scale cases, the acoustic emission signals are sourced or detected in a volume remote from the volume that spawns the earthquake. Therefore, it is relevant to establish the causal relationship between signals detected on passive, remote monitors and the dynamics of the elastic structures that launch important seismic events. The work conducted under this IC allocation allowed us to utilize a numerical model that let us follow this causality, i.e., examine and connect the dynamics in a granular system (fault gouge), to signals detected on passive remote monitors. It was demonstrated that stress chains are key in the dynamics of granular systems and that their evolution is the source of the acoustic emission.

58 GEOSCIENCES↗

Development of Integrated Safety and Security Models for Comprehensive Reliability and Resiliency Evaluation

The security of the electric grid and supporting energy systems is crucial to national security. One of the complexities in analyzing the security of energy systems is the safety consequences that may result from accidents. For energy systems, the goal is to ensure that they operate as intended and that any consequences are mitigated or prevented. The integration of safety and security is paramount to protecting these systems from attacks and ensuring that large consequences are prevented. This report describes an integrated safety and security methodology to evaluate cybersecurity events that can lead to large consequences. This novel approach first describes how Systems-Theoretic Process Analysis (STPA) provides a digital causal analysis for Bayesian Networks (BNs). The use of STPA causal analysis provides a systematic approach to constructing BNs that adequately model cyber scenarios that result in consequences. When combined with the technical principles described in Risk-Informed Management of Enterprise Systems (RIMES), a comprehensive risk-informed cybersecurity analysis results that allows decision-makers to prioritize systems that most impact risk.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Phasor-Measurement-Unit-Based Data Analytics Using Digital Twin and PhasorAnalytics Software

A major objective of this project was to apply GE’s commercial machine learning and data analytics toolsets to large-scale, real-world, anonymized Phasor Measurement Unit (PMU) datasets in order to extract signatures, correlated and/or causal factors, and precursor patterns associated with significant power system phenomena. The project had a particular emphasis on extraction of insights relevant to asset health monitoring, real-time load modeling and cybersecurity monitoring. Additionally, the team was directed to undertake a comprehensive data quality analysis for the provided datasets and encouraged to estimate the ‘machine-learning readiness’ of the datasets by documenting any major obstacles to the application of commercial machine learning algorithms. To accomplish the aforementioned objectives, the project team’s work centered around the identification of key event signatures and application of the identified event signatures for event detection and event classification. The industry-validated, semi-supervised machine learning strategy employed for event signature identification involved several major tasks, including data-preprocessing, generation of an overabundance of features, normal data identification, normality modeling, and event signature identification through a methodical, quantitative ranking of features in order of relevance to each studied event type. Throughout the project, data quality issues and mitigation techniques were investigated. In this report, insights are provided regarding the readiness of the provided synchrophasor datasets for application of machine learning and data analytics. The methodologies employed for this technical strategy are summarized in this report. With regards to data preprocessing and feature generation, the provided Training and Test Datasets were ingested into GE’s big data environment. Subsequently, the team applied bad data cleansing and data imputation scripts, event detection scripts, and application programming interfaces (APIs) to the datasets for convenient data access. The project team completed development and validation of dozens of physics-based, statistics-based and transformation-based feature functions used for the extraction of over 60 synchrophasor features. Using a new parallel feature generation technology developed on this project, over 60 features have been rapidly generated for the full two years’ worth of Training and Test Dataset data associated with both the Eastern and Western interconnects. Even accommodating for temporal down-sampling inherent to the feature extraction procedure, this parallel feature generation activity resulted in a massive feature set with a storage requirement approximately equal to that of the raw training dataset itself. With regards to normal data identification and normality modeling, a normality model was built using the feature data extracted from the Training Dataset and iteratively refined subsequent to incremental adjustments and expansions of the Training Dataset feature data. With respect to event characterization and signature identification, an event signature identification pipeline was developed and used in conjunction with the normality model to identify over 15 event signatures for key event categories within the Training Dataset. The identified event signatures were used to characterize hundreds of key events in terms of relative severity, duration, and location of the event. An investigation was undertaken to identify correlated and causal factors involved in transformer events. A separate investigation into temporal trends in ring-down analysis results was undertaken to determine possible associations between system dynamics and various other factors such as loading, season or year. To validate the identified event signatures, additional work was undertaken to develop signature-based anomaly detection and classification tools suitable for convenient application to the synchrophasor datasets. The anomaly detection and classification tools, suitable for online application, were then applied to the entirety of the Eastern Interconnect Training and Test Datasets. Performance of the event detection and classification tools was evaluated upon receipt of the Test Dataset event logs (i.e., the labels for events contained in the Test Dataset), and promising results were obtained despite several challenges (documented herein) associated with application of supervised or semi-supervised machine learning methods to large-scale, anonymized datasets. Finally, the detection and classification tools were used to detect, classify, and characterize thousands of new events not included in the original event logs provided by the DOE within both the Training and Test Datasets.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Investigation of Surface and Marine-Cloud Coupling and its Impact on Cloud Droplet Number Concentrations and Cloud Cover Over the Southern Ocean

The proposed work involves characterizing and quantifying both the thermodynamic and dynamical coupling of marine low clouds (MLC) with the sea surface using Atmospheric Radiation Measurement (ARM) observations during MARCUS field campaign and an LES model. ARM data from multiple sensors will be used (e.g., Doppler cloud radar, ceilometer, and microwave radiometer) to characterize the MLC-surface coupling by virtue of the vertical structure and integrated quantities of boundary-layer clouds, aerosols, as well as atmospheric profiles and surface meteorology. We have used a combination of case studies and statistics-based composite analyses to find any linkages between large-scale dynamics, MLC-surface coupling, and cloud and boundary layer properties, the sequence of which reflects the chain of causality. An LES model with explicit aerosol physics, which includes the cycle of aerosols by being consumed as CCN and be regenerated following cloud droplet evaporation, has been run to determine the sources and sinks of Nd and their dependence on the degree of MLC-surface coupling. We have examined the systematic differences in both Nd and cloud occurrence between the two clusters under different meteorological conditions and further examine their respective roles, as well as causal relationships by means of LES modeling. This study helped improve our understanding of the ACI by differentiating the dynamic role of the coupling and cloud physics denoted by Nd, bridging the linkage in the chain toward understanding mechanisms governing the persistence of MLC over the Southern Ocean, solving the long-lasting problem of the cloud cover underestimation over the SO by GCMs. Ample ARM data and LES model have been employed to achieve the objectives of the study.

54 ENVIRONMENTAL SCIENCES↗

AI-Assisted Conceptual Development of a Pre-Geometric Cosmological Model - An Exercise in AI-Assisted Conceptual Framework Generation, Paper II: Local Geometry and Metric Structure

This paper develops the geometric sector of the replication-driven cosmogenesis framework introduced in Paper I. Starting from a pre-geometric spectral substrate and a minimal set of replication axioms, we show how coherent self-replicating units generate a spatial adjacency graph whose continuum limit acquires an effective Riemannian structure. The replication dynamics determines a characteristic correlation length that seeds the local metric, while overlap relations among coherent units produce an isotropic neighborhood geometry with an emergent dimensionality $d_{\rm eff}\simeq 3$ across a broad range of replication factors. As replication slows and causal order stabilizes, a limiting signal speed $c_\ast$ appears, providing the basis for the Lorentzian structure of spacetime without assuming a pre-existing light cone. We derive conditions under which the adjacency graph converges to a smooth three-dimensional manifold, describe the transition from Euclidean to Lorentzian propagation, and identify geometric invariants controlled by the replication parameters. This work establishes the geometric and causal layer of the replication cosmogenesis program, bridging the spectral axioms of Paper I to the cosmological dynamics explored in Paper III.

79 ASTRONOMY AND ASTROPHYSICS↗

AI-Assisted Conceptual Development of a Pre-Geometric Cosmological Model - An Exercise in AI-Assisted Conceptual Framework Generation, Paper I: Foundations and Replication Dynamics

We develop a pre-geometric cosmological framework in which existence is identified with a finite amount of unstructured energy possessing vibration as its only intrinsic property. This vibrational substrate occupies an open, bounded spectral interval $(\omega_{\min},\omega_{\max})$, ensuring finiteness of total energy and excluding infinitely stable configurations. The substrate evolves under two fundamental and competing tendencies---excitation, which amplifies coherence, and randomization, which scrambles it. Their balance produces a metastable unstructured regime in which rare fluctuations may form long-lived self-consistent spectral configurations. Because the substrate is finite and subject to competing order--disorder dynamics, no coherent configuration can be perpetually stable. We show that the only mechanism capable of sustaining long-lived organization is a replication instability: a coherent unit may reproduce into multiple offspring according to a general $1\!\to n$ rule. Replication consumes energy from the finite substrate, breaks the metastable symmetry, and induces a discrete notion of event time through the replication tick $\Delta\tau$. Temporal succession is defined through correlation ordering of spectral microstates, producing an intrinsic pre-causal structure. The compactness of the spectral domain imposes minimal and maximal timescales, bounds the internal coherence of emergent units, and limits their proliferation. These spectral constraints serve as precursors for the emergence of geometry, adjacency, and a limiting propagation speed, developed in subsequent papers of this series. Paper I provides the foundational axioms (PG1--PG9) governing the spectral substrate, its metastable dynamics, the formation of coherent units, and the necessity of replication, establishing a fully pre-geometric stage from which causal and geometric structure naturally emerge.

79 ASTRONOMY AND ASTROPHYSICS↗