Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “mutual information”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Demonstration of Optimal Benchmark Selection Website and Validation of the q c Coverage Metric Using HEU-SOL-THERM-013-003 Experiment

In the work documented in this interim report, the experiment selection toolkit web site was demonstrated and q C coverage metric methodology was validated for IEU-MET-FAST-002-001, MIX-COMP-THERM 004-004, and HEU-SOL-THERM-013-003 experiments. 𝑞 𝐶 is an information-theoretic measure based on mutual information that quantifies the ability of candidate benchmark experiments to reduce the bias and uncertainty of a target criticality safety application. The metric and an accompanying open-source Python toolkit with a web-based interface were tested against a benchmark set of 425 experiments drawn from the International Criticality Safety Benchmark Evaluation Project Handbook. The interface is hosted at https://edim.covdef.com. It accepts sensitivity data files produced by the TSUNAMI-IP module of the SCALE code system and supports both (i) deterministic analysis using the ENDF/B-VII.0 covariance library and (ii) stochastic analysis based on user-supplied keff samples. Demonstrations on representative applications across a range of material composition, spectrum, and form show that q C -guided benchmark selection achieves greater uncertainty reduction with fewer experiments and yields more stable posterior bias and uncertainty estimates than traditional similarity coefficient ( c k )–based selection, while also capturing valuable low-ck experiments that one-to-one metrics overlook.

Abdel-khalik, Hany S. [Indiana Univ.-Purdue Univ. ↗

Information-Theoretic Exploration of Multivariate Time-Varying Image Databases

Modern scientific simulations produce very large datasets, making interactive exploration of such data computationally prohibitive. An increasingly common data reduction technique is to store visualizations and other data extracts in a database. The Cinema project is one such approach, storing visualizations in an image database for post hoc exploration and interactive image-based analysis. This work focuses on developing efficient algorithms that can quantify various types of multivariate dependencies existing within multi-variable datasets. It applies specific mutual information measures for the quantification of salient regions from multivariate image data. Here, using such information measures, the opacity of the images is modulated so that the salient regions are automatically highlighted and the domain scientists can interactively explore the most relevant regions for scientific discovery.

97 MATHEMATICS AND COMPUTING↗

A holistic cyber-physical security protocol for authenticating the provenance and integrity of structural health monitoring imagery data

Modern infrastructure systems, such as bridges, dams, power generation stations, and buildings increasingly have an intrinsic cyber-physical nature to them. Infrastructure now commonly, includes actuators, network connections, sensors, control systems, and computational resources. It is of increasing concern that modern infrastructure is vulnerable to cyber-attacks that can damage both the cyber and physical nature of the infrastructure. To date, the physical and cyber health of infrastructure has been considered separately. However, the increasing concerns associated with the cyber-physical security of infrastructure coupled with the emergence of 5G networks made using components that are not universally considered trustworthy, and the emergence of techniques for creating deepfakes and adversarial examples suggests the time has come to begin considering cyber health and structural health with a more holistic approach. In this work, a protocol is developed for ensuring the imagery data captured by a structural health monitoring system can be unambiguously attributed to legitimate sensors associated with the structural health monitoring system. A computer vision approach based on the idea of mutual information is then presented to detect damage in an image. This work presents the protocol for authenticating the provenance of imager data and demonstrates that this protocol does not have overly adverse effects when used with the mutual information-based technique for detecting damage in the resulting imagery data.

Jung, HweeKwon↗

Identifying dominant environmental predictors of freshwater wetland methane fluxes across diurnal to seasonal time scales

While wetlands are the largest natural source of methane (CH 4 ) to the atmosphere, they represent a large source of uncertainty in the global CH 4 budget due to the complex biogeochemical controls on CH 4 dynamics. Here we present, to our knowledge, the first multi-site synthesis of how predictors of CH 4 fluxes (FCH4) in freshwater wetlands vary across wetland types at diel, multiday (synoptic), and seasonal time scales. In this work we used several statistical approaches (correlation analysis, generalized additive modeling, mutual information, and random forests) in a wavelet-based multi-resolution framework to assess the importance of environmental predictors, nonlinearities and lags on FCH4 across 23 eddy covariance sites. Seasonally, soil and air temperature were dominant predictors of FCH4 at sites with smaller seasonal variation in water table depth (WTD). In contrast, WTD was the dominant predictor for wetlands with smaller variations in temperature (e.g., seasonal tropical/subtropical wetlands). Changes in seasonal FCH4 lagged fluctuations in WTD by ~17 ± 11 days, and lagged air and soil temperature by median values of 8 ± 16 and 5 ± 15 days, respectively. Temperature and WTD were also dominant predictors at the multiday scale. Atmospheric pressure (PA) was another important multiday scale predictor for peat-dominated sites, with drops in PA coinciding with synchronous releases of CH 4 . At the diel scale, synchronous relationships with latent heat flux and vapor pressure deficit suggest that physical processes controlling evaporation and boundary layer mixing exert similar controls on CH 4 volatilization, and suggest the influence of pressurized ventilation in aerenchymatous vegetation. In addition, 1- to 4-h lagged relationships with ecosystem photosynthesis indicate recent carbon substrates, such as root exudates, may also control FCH4. By addressing issues of scale, asynchrony, and nonlinearity, this work improves understanding of the predictors and timing of wetland FCH4 that can inform future studies and models, and help constrain wetland CH 4 emissions.

59 BASIC BIOLOGICAL SCIENCES↗

Data and scripts associated with “Allometric scaling of hyporheic respiration across basins in the Pacific Northwest USA"

This data package is associated with the publication “Allometric scaling of hyporheic respiration across basins in the Pacific Northwest USA” submitted to JGR-Biogeosciences (Regier et al. 2025).This study used reach-scale modeled estimates of hyporheic aerobic respiration made by the River Corridor Model (Fang et al. 2020) and watershed characteristics across the Willamette and Yakima River basins to explore potential allometric scaling (i.e., power-law relationships between size and function) of cumulative hyporheic respiration across catchment-to-basin scales. Scaling was explored quantitatively via the R2, slope, and y-intercept of relationships between cumulative hyporheic respiration and watershed area, divided into hyporheic exchange flux (HEF) quantiles. We also explored relationships between allometric scaling and other watershed characteristics through linear regression, spatial patterns, and mutual information analyses. Our results also suggest variability of hyporheic respiration allometry for middle exchange flux quantiles, and in relation to land-cover. Our findings provide initial evidence that allometric scaling may be useful for predicting hyporheic biogeochemical dynamics across watersheds from reach to basin scales. This data package is associated with the GitHub repository found at https://github.com/peterregier/rc_wrb_yrb_scaling. The data package is organized into several key directories. The “data” folder contains multiple CSV files, including landscape heterogeneity, scaling analysis, and watershed boundary data. The “figures” folder has all figure files in both PDF and PNG formats. Core analysis scripts and figure generation scripts are in the “scripts” directory, systematically numbered for sequential execution. The root directory includes essential project files; please see the file ending in “flmd.csv” for a list and description of all files contained in this data package and the file ending in “dd.csv” for data dictionaries used to describe tabular column headers.

54 ENVIRONMENTAL SCIENCES↗

A Contextually Supervised Optimization-Based HVAC Load Disaggregation Methodology

This paper presents a novel contextually supervised optimization-based approach for disaggregating heating, ventilation, and air-conditioning (HVAC) loads using smart meter or Supervisory Control and Data Acquisition data. To disaggregate the load into HVAC loads, large and infrequently used loads (LIUL), and base loads, we formulate an optimization problem to minimize a set of five loss terms, consisting of the reconstruction errors of the overall load profile, the ramp rate losses, and three distinct loss functions linked with the HVAC load, base load, and LIUL, respectively. To enhance accuracy, we incorporate two forms of contextual information into the problem formulation. First, we utilize mutual information to estimate HVAC energy consumption. Second, we employ a base load dictionary to constrain HVAC load estimation errors. The obtained HVAC load profiles are fine-tuned by abnormal ramp detection followed by binary hypothesis testing. Here, the proposed method is developed and tested using sub-metered residential and commercial building data. Simulation results show that the proposed method outperforms existing methods across various data resolutions and load aggregation levels, showing excellent transferability and generalizability.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Measurement of Volatile Compounds for Real-Time Analysis of Soil Microbial Metabolic Response to Simulated Snowmelt

Snowmelt dynamics are a significant determinant of microbial metabolism in soil and regulate global biogeochemical cycles of carbon and nutrients by creating seasonal variations in soil redox and nutrient pools. With an increasing concern that climate change accelerates both snowmelt timing and rate, obtaining an accurate characterization of microbial response to snowmelt is important for understanding biogeochemical cycles intertwined with soil. However, observing microbial metabolism and its dynamics non-destructively remains a major challenge for systems such as soil. Microbial volatile compounds (mVCs) emitted from soil represent information-dense signatures and when assayed non-destructively using state-of-the-art instrumentation such as Proton Transfer Reaction-Time of Flight-Mass Spectrometry (PTR-TOF-MS) provide time resolved insights into the metabolism of active microbiomes. In this study, we used PTR-TOF-MS to investigate the metabolic trajectory of microbiomes from a subalpine forest soil, and their response to a simulated wet-up event akin to snowmelt. Using an information theory approach based on the partitioning of mutual information, we identified mVC metabolite pairs with robust interactions, including those that were non-linear and with time lags. The biological context for these mVC interactions was evaluated by projecting the connections onto the Kyoto Encyclopedia of Genes and Genomes (KEGG) network of known metabolic pathways. Simulated snowmelt resulted in a rapid increase in the production of trimethylamine (TMA) suggesting that anaerobic degradation of quaternary amine osmo/cryoprotectants, such as glycine betaine, may be important contributors to this resource pulse. Unique and synergistic connections between intermediates of methylotrophic pathways such as dimethylamine, formaldehyde and methanol were observed upon wet-up and indicate that the initial pulse of TMA was likely transformed into these intermediates by methylotrophs. Increases in ammonia oxidation signatures (transformation of hydroxylamine to nitrite) were observed in parallel, and while the relative role of nitrifiers or methylotrophs cannot be confirmed, the inferred connection to TMA oxidation suggests either a direct or indirect coupling between these processes. Overall, it appears that such mVC time-series from PTR-TOF-MS combined with causal inference represents an attractive approach to non-destructively observe soil microbial metabolism and its response to environmental perturbation.

59 BASIC BIOLOGICAL SCIENCES↗

Preventing Reverse Engineering of Critical Industrial Data with DIOD

Business analytics augmented by artificial intelligence and machine learning (AI/ML) have revolutionized the role of data in the modern world. In recent years, businesses have incorporated data into their decision-making process for better prediction, risk-assessment, content creation, etc. While such businesses often seek to leverage the full use of their data through third-party AI/ML services, they are often hampered by the risks of data leaks, reverse-engineering, stolen technology, etc. that often have disastrous consequences for businesses and their stakeholders alike. Thus, there arises a need for data masking prior to its transmission that obfuscates proprietary information while preserving the information relevant for AI/ML applications. In order to meet the needs of industrial data which are significantly different from those of data warehouses, previous work proposed an efficient time and space-scalable data masking paradigm known as the deceptive infusion of data (DIOD) methodology. The present work expands upon this work by leveraging existing reverse-engineering capabilities to facilitate the decomposition of industrial data into its proprietary and AI/ML-relevant parts, referred to as fundamental and inference metadata respectively. Both sets of metadata are further obfuscated in accordance with the DIOD methodology to create the DIOD rendition of the industrial data, which is rendered immune to reverse-engineering by discarding proprietary information and only preserving AI/ML-relevant information. Additionally, constraints of the original DIOD manuscript are relaxed using mutual information by configuring the methodology to the target AI/ML application to unlock the full potential of the DIOD methodology. As an example, data from a nuclear reactor is transformed into that from a nonlinear spring-mass system with different levels of data masking as required by the generic system and the target application.

97 MATHEMATICS AND COMPUTING↗

On the origin of red spirals: does assembly bias play a role?

The formation of the red spirals is a puzzling issue in the standard picture of galaxy formation and evolution. Most studies attribute the colour of the red spirals to different environmental effects. We analyze a volume limited sample from the SDSS to study the roles of small-scale and large-scale environments on the colour of spiral galaxies. We compare the star formation rate, stellar age and stellar mass distributions of the red and blue spirals and find statistically significant differences between them at 99.9% confidence level. The red spirals inhabit significantly denser regions than the blue spirals, explaining some of the observed differences in their physical properties. However, the differences persist in all types of environments, indicating that the local density alone is not sufficient to explain the origin of the red spirals. Using an information theoretic framework, we find a small but non-zero mutual information between the colour of spiral galaxies and their large-scale environment that are statistically significant (99.9% confidence level) throughout the entire length scale probed. Such correlations between the colour and the large-scale environment of spiral galaxies may result from the assembly bias. Thus both the local environment and the assembly bias may play essential roles in forming the red spirals. The spiral galaxies may have different assembly history across all types of environments. We propose a picture where the differences in the assembly history may produce spiral galaxies with different cold gas content. Such a difference would make some spirals more susceptible to quenching. Finally, in all environments, the spirals with high cold gas content could delay the quenching and maintain a blue colour, whereas the spirals with low cold gas fractions would be easily quenched and become red.

79 ASTRONOMY AND ASTROPHYSICS↗

Security Analysis of a Class of Spread Spectrum Systems Presentation

A method of adding physical layer security to a class of spread spectrum systems has been recently proposed. In this paper, we look into the rate at which an eavesdropper may gain information about the system to decipher the data symbols. The Shannon mutual information is used to measure the rate of information that may be gained by an eavesdropper. The k-nearest neighbors (k-NN) method is used to obtain estimates of relevant entropy values, which will then be used to quantify the rate of information recovery as more data is transmitted. It turns out that such information recovery requires the adoption of special methods that avoid any destructive bias in the estimates. Details of these methods are also presented.

97 - MATHEMATICS AND COMPUTING↗

Security Analysis of a Class of Secured Spread Spectrum Systems

Abstract—A method of adding physical layer security to a class of spread spectrum systems has been recently proposed. In this paper, we look into the rate at which an eavesdropper may gain information about the system to decipher the data symbols. The Shannon mutual information is used to measure the rate of information that may be gained by an eavesdropper. The k-nearest neighbors (k-NN) method is used to obtain the estimates of relevant entropy values which will be then used to quantify the rate of information recovery as more data are being transmitted. It turns out that such information recovery requires adoption of special methods that avoid any destructive bias in the estimates. Details of these methods are also presented.

97 - MATHEMATICS AND COMPUTING↗

A Hybrid Gradient Method to Designing Bayesian Experiments for Implicit Models

Bayesian experimental design (BED) aims at designing an experiment to maximize the information gathering from the collected data. The optimal design is usually achieved by maximizing the mutual information (MI) between the data and the model parameters. When the analytical expression of the MI is unavailable, e.g.,having implicit models with intractable data distributions, a neural network-based lower bound of the MI was recently proposed and a gradient ascent method was used to maximize the lower bound [1]. However, the approach in [1] requires a pathwise sampling path to compute the gradient of the MI lower bound with respect to the design variables, and such a pathwise sampling path is usually inaccessible for implicit models. In this work, we propose a hybrid gradient approach that leverages recent advances in variational MI estimator and evolution strategies (ES)combined with black-box stochastic gradient ascent (SGA) to maximize the MI lower bound. This allows the design process to be achieved through a unified scalable procedure for implicit models without sampling path gradients. Several experiments demonstrate that our approach significantly improves the scalability of BED for implicit models in high-dimensional design space.

Zhang, Jiaxin↗

Variational Information Planning for Sequential Decision Making

We consider the setting of sequential decision making where, at each stage, potential actions are evaluated based on expected reduction in posterior uncertainty, given by mutual information (MI). As MI typically lacks a closed form, we propose an approach which maintains variational approximations of, both, the posterior and MI utility. Our planning objective extends an established variational bound on MI to the setting of sequential planning. The result, variational information planning (VIP), is an efficient method for sequential decision making. We further establish convexity of the variational planning objective and, under conditional exponential family approximations, we show that the optimal MI bound arises from a relaxation of the well-known exponential family moment matching property. Here, we demonstrate VIP for sensor selection, experiment design, and active learning, where it meets or exceeds methods requiring more computation, or those specialized to the task.

Pacheco, Jason↗

Replica wormhole and information retrieval in the SYK model coupled to Majorana chains

Motivated by recent studies of the information paradox in (1+1)-D anti-de Sitter spacetime with a bath described by a (1+1)-D conformal field theory, we study the dynamics of second Ŕenyi entropy of the Sachdev-Ye-Kitaev (SYK) model (χ) coupled to a Majorana chain bath (ψ). The system is prepared in the thermofield double (TFD) state and then evolved by HL + HR. For small system-bath coupling, we find that the second Rényi entropy $S^(2)_{χL,χR}$ of the SYK model undergoes a first order transition during the evolution. In the sense of holographic duality, the long-time solution corresponds to a “replica wormhole”. The transition time corresponds to the Page time of a black hole coupled to a thermal bath. We further study the information scrambling and retrieval by introducing a classical control bit, which controls whether or not we add a perturbation in the SYK system. The mutual information between the bath and the control bit shows a positive jump at the Page time, indicating that the entanglement wedge of the bath includes an island in the holographic bulk.

1/N Expansion↗

Redundantly Amplified Information Suppresses Quantum Correlations in Many-Body Systems

We establish bounds on quantum correlations in many-body systems. They reveal what sort of information about a quantum system can be simultaneously recorded in different parts of its environment. Specifically, independent agents who monitor environment fragments can eavesdrop only on amplified and redundantly disseminated—hence, effectively classical—information about the decoherence-resistant pointer observable. We also show that the emergence of classical objectivity is signaled by a distinctive scaling of the conditional mutual information, bypassing hard numerical optimizations. Our results validate the core idea of quantum Darwinism: objective classical reality does not need to be postulated and is not accidental, but rather a compelling emergent feature of quantum theory that otherwise—in the absence of decoherence and amplification—leads to “quantum weirdness.” In particular, a lack of consensus between agents that access environment fragments is bounded by the information deficit, a measure of the incompleteness of the information about the system.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

A Scalable Gradient Free Method for Bayesian Experimental Design with Implicit Models

Bayesian experimental design (BED) is to answer the question that how to choose designs that maximize the information gathering. For implicit models, where the likelihood is intractable but sampling is possible, conventional BED methods have difficulties in efficiently estimating the posterior distribution and maximizing the mutual information (MI) between data and parameters. Recent work proposed the use of gradient ascent to maximize a lower bound on MI to deal with these issues. However, the approach requires a sampling path to compute the pathwise gradient of the MI lower bound with respect to the design variables, and such a pathwise gradient is usually inaccessible for implicit models. In this paper, we propose a novel approach that leverages recent advances in stochastic approximate gradient ascent incorporated with a smoothed variational MI estimator for efficient and robust BED. Without the necessity of pathwise gradients, our approach allows the design process to be achieved through a unified procedure with an approximate gradient for implicit models. Several experiments show that our approach outperforms baseline methods, and significantly improves the scalability of BED in high-dimensional problems.

Zhang, Jiaxin↗

CAD-VAE: Leveraging Correlation-Aware Latents for Comprehensive Fair Disentanglement

While deep generative models have significantly advanced representation learning, they may inherit or amplify biases and fairness issues by encoding sensitive attributes alongside predictive features. Enforcing strict independence in disentanglement is often unrealistic when target and sensitive factors are naturally correlated. To address this challenge, we propose CAD-VAE(Correlation-Aware Disentangled VAE), which introduces a correlated latent code to capture the information shared between the target and sensitive attributes. Given this correlated latent, our method effectively separates over-lapping factors without extra domain knowledge by directly minimizing the conditional mutual information between target and sensitive codes. A relevance-driven optimization strategy refines the correlated code by efficiently capturing essential correlated features and eliminating redundancy. Extensive experiments on benchmark datasets demonstrate that CAD-VAE produces fairer representations, realistic counterfactuals, and improved fairness-aware image editing.

Ma, Chenrui [University of California Irvine]↗

Improvement and generalization of ABCD method with Bayesian inference

To find New Physics or to refine our knowledge of the Standard Model at the LHC is an enterprise that involves many factors, such as the capabilities and the performance of the accelerator and detectors, the use and exploitation of the available information, the design of search strategies and observables, as well as the proposal of new models. We focus on the use of the information and pour our effort in re-thinking the usual data-driven ABCD method to improve it and to generalize it using Bayesian Machine Learning techniques and tools. We propose that a dataset consisting of a signal and many backgrounds is well described through a mixture model. Signal, backgrounds and their relative fractions in the sample can be well extracted by exploiting the prior knowledge and the dependence between the different observables at the event-by-event level with Bayesian tools. We show how, in contrast to the ABCD method, one can take advantage of understanding some properties of the different backgrounds and of having more than two independent observables to measure in each event. In addition, instead of regions defined through hard cuts, the Bayesian framework uses the information of continuous distribution to obtain soft-assignments of the events which are statistically more robust. To compare both methods we use a toy problem inspired by pp\to hh\to b\bar b b \bar b p p → h h → b b ‾ b b ‾ , selecting a reduced and simplified number of processes and analysing the flavor of the four jets and the invariant mass of the jet-pairs, modeled with simplified distributions. Taking advantage of all this information, and starting from a combination of biased and agnostic priors, leads us to a very good posterior once we use the Bayesian framework to exploit the data and the mutual information of the observables at the event-by-event level. We show how, in this simplified model, the Bayesian framework outperforms the ABCD method sensitivity in obtaining the signal fraction in scenarios with 1% and 0.5% true signal fractions in the dataset. We also show that the method is robust against the absence of signal. We discuss potential prospects for taking this Bayesian data-driven paradigm into more realistic scenarios.

Alvarez, Ezequiel↗