Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “mutual information”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Mitigating Scattering in a Quantum System Using Only an Integrating Sphere

Strong quantum correlated sources are essential but delicate resources for quantum information science and engineering protocols. Decoherence and loss are the two main disruptive processes that lead to the loss of nonclassical behavior in quantum correlations. In quantum systems, scattering can contribute to both decoherence and loss. In this work, we present an experimental scheme capable of significantly mitigating the adverse impact of scattering in quantum systems. Our quantum system is composed of a two-mode squeezed light generated with the four-wave-mixing process in hot rubidium vapor and a scatterer is introduced to one of the two modes. An integrating sphere is then placed after the scatterer to recollect the scattered photons. We use mutual information between the two modes as the measure of quantum correlations and demonstrate a 47.5% mutual information recovery from scattering, despite an enormous photon loss of greater than 85%. Our scheme is the very first step toward recovering quantum correlations from disruptive random processes and thus has the potential to bridge the gap between proof-of-principle demonstrations and practical real-world implementations of quantum protocols. Published by the American Physical Society 2024

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

News Media Framing of Suicide Circumstances and Gender: Mixed Methods Analysis

Background: Suicide is a leading cause of death worldwide. Journalistic reporting guidelines were created to curb the impact of unsafe reporting; however, how suicide is framed in news reports may differ by important characteristics such as the circumstances and the decedent’s gender. Objective: This study aimed to examine the degree to which news media reports of suicides are framed using stigmatized or glorified language and differences in such framing by gender and circumstance of suicide. Methods: We analyzed 200 news articles regarding suicides and applied the validated Stigma of Suicide Scale to identify stigmatized and glorified language. We assessed linguistic similarity with 2 widely used metrics, cosine similarity and mutual information scores, using a machine learning–based large language model. Results: News reports of male suicides were framed more similarly to stigmatizing (P<.001) and glorifying (P=.005) language than reports of female suicides. Considering the circumstances of suicide, mutual information scores indicated that differences in the use of stigmatizing or glorifying language by gender were most pronounced for articles attributing legal (0.155), relationship (0.268), or mental health problems (0.251) as the cause.

59 BASIC BIOLOGICAL SCIENCES↗

Quantifying the Long‐Range Coupling of Electronic Properties in Proteins with ab initio Molecular Dynamics**

Abstract The delicate interplay of covalent and non‐covalent interactions in proteins is inherently quantum mechanical and highly dynamic in nature. To directly interrogate the evolving nature of the electronic structure of proteins, we carry out 100‐ps‐scale ab initio molecular dynamics simulations of three representative small proteins with range‐separated hybrid density functional theory. We quantify the nature and length‐scale of the coupling of residue‐specific charge probability distributions in these proteins. While some nonpolar residues exhibit expectedly narrow charge distributions, most polar and charged residues exhibit broad, multimodal distributions. Even for nonpolar residues, we observe sequence‐specific deviations corresponding to charge accumulation or depletion that would be challenging to capture in a fixed charge force field. We quantify the effect of residue‐residue interactions on charge distributions first with linear cross‐correlations. We then show how additional insight can be gained from evaluating the mutual information of charge distributions. We show that a significant number of residues couple most strongly with residues that are distant in both sequence and space over a range of secondary structures including α‐helical, β‐sheet, disulfide bridging, and lasso motifs. The mutual information analysis is necessary to capture coupling between some polar and charged residues that would be otherwise missed.

Yang, Zhongyue↗

Multi-body entanglement and information rearrangement in nuclear many-body systems: a study of the Lipkin–Meshkov–Glick model

Here, we examine how effective-model-space (EMS) calculations of nuclear many-body systems rearrange and converge multi-particle entanglement. The generalized Lipkin-Meshkov-Glick (LMG) model is used to motivate and provide insight for future developments of entanglement-driven descriptions of nuclei. The effective approach is based on a truncation of the Hilbert space together with a variational rotation of the qubits (spins), which constitute the relevant elementary degrees of freedom. The non-commutivity of the rotation and truncation allows for an exponential improvement of the energy convergence throughout much of the model space. Our analysis examines measures of correlations and entanglement, and quantifies their convergence with increasing cut-off. We focus on one- and two-spin entanglement entropies, mutual information, and $n$-tangles for $n=2,4$ to estimate multi-body entanglement. The effective description strongly suppresses entropies and mutual information of the rotated spins, while being able to recover the exact results to a large extent with low cut-offs. Naive truncations of the bare Hamiltonian, on the other hand, artificially underestimate these measures. The $n$-tangles in the present model provide a basis-independent measures of $n$-particle entanglement. While these are more difficult to capture with the EMS description, the improvement in convergence, compared to truncations of the bare Hamiltonian, is significantly more dramatic. We conclude that the low-energy EMS techniques, that successfully provide predictive capabilities for low-lying observables in many-body systems, exhibit analogous efficacy for quantum correlations and multi-body entanglement in the LMG model, motivating future studies in nuclear many-body systems and effective field theories relevant to high-energy physics and nuclear physics.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Allosteric prediction via convolutional neural networks and protein structural and dynamical features

Allostery is the phenomenon whereby a binding event or covalent modification at one site in a protein modulates function at a distal site, thus changing a protein’s functional state. As such, it is a ubiquitous aspect of protein functional regulation. Computationally predicting allosteric states is important as part of the broader challenge of functional annotation, but it also has practical implications for drug development, as targeting an allosteric site often affords greater specificity compared with targeting an orthosteric site. This study introduces a machine learning approach to predict the allosteric functional state using the small G-protein KRas as the model system, due to its implication in many types of cancer and being well studied as a result with many x-ray crystallographic structures of KRas available with different mutations and ligands bound. Using structural and dynamical features that can be cast as images, namely interatomic distances, contact maps, covariance, and mutual information, supervised learning was performed using convolutional neural networks. Two pretrained convolutional neural network architectures, GoogLeNet and ResNet18, were fine-tuned to classify KRas into active or inactive states based on these features. Across training regimes, atomic contact maps emerged as the most effective structural feature, whereas linearized mutual information outperformed covariance in capturing dynamical correlations relevant to allostery. Models achieved significant validation accuracy, with atomic contact maps yielding up to 90% accuracy. In conclusion, the findings suggest that integrating global structural rearrangements and correlated motion patterns with deep learning can reliably predict protein allosteric states, offering a promising framework for understanding allosteric regulation and developing targeted therapeutics.

Rajeshwar T., Rajitha [Oak Ridge National Laborato↗

Data and scripts associated with the manuscript "Encoding Diel Hysteresis and the Birch Effect in Dryland Soil Respiration Models through Knowledge-Guided Deep Learning"

This package contains the data and scripts used in "Encoding Diel Hysteresis and the Birch Effect in Dryland Soil Respiration Models through Knowledge-Guided Deep Learning" (Jiang et al., 2022). The data.zip file contains the flux tower and automated chamber observations used for developing the deep learning model for modeling soil respiration. The scripts.zip file contains the Jupyter notebooks and python scripts for preprocessing the data, training the deep learning models, and postprocessing the results. The src.zip contains the source code for training the deep learning model, performing mutual information analysis, and plotting functions. The trained_models.zip contains multiple folders used for hosting the trained deep-learning models and the associated soil respiration predictions. The whole process is performed using python. We include the REAMD.md to document the python package requirements.Soil respiration in dryland ecosystems is challenging to model due to its complex interactions with environmental drivers. Knowledge-guided deep learning provides a much more effective means of accurately representing these complex interactions than traditional Q10-based models. Mutual information analysis revealed that future soil temperature shares more information with soil respiration than past soil temperature, consistent with their clockwise diel hysteresis. We explicitly encoded diel hysteresis, soil drying, and soil rewetting effects on soil respiration dynamics in a newly designed Long Short Term Memory (LSTM) model. The model takes both past and future environmental drivers as inputs to predict soil respiration. The new LSTM model substantially outperformed three Q10-based models and the Community Land Model when reproducing the observed soil respiration dynamics in a semi-arid ecosystem. The new LSTM model clearly demonstrated its superiority for temporally extrapolating soil respiration dynamics, such that the resulting correlation with observational data is up to 0.7 while the correlations of both Q10-based models and the Community Land Model (CLM) are less than 0.4. Our results underscore the high potential for knowledge-guided deep learning to replace Q10-based soil respiration modules in Earth system models.

54 ENVIRONMENTAL SCIENCES↗

Landslide Likelihood Prediction using Machine Learning Algorithms

The supply of electricity via power plants is criticalto the operation of many critical infrastructure systems in mod-ern society. Natural hazards can disrupt the power supply, causepower outages that can halt economic growth, and impede emer-gency response until power is restored. The proposed work aimsto predict the landslides likelihood in these critical infrastructurelocations in the Northeastern USA using integrated databases ofexplanatory variables and machine learning algorithms. First,data related to landslides are obtained and merged, includingtopographic, soil moisture, and precipitation-related data. Fiveregression algorithms, namely: Random Forest, Extreme Gradi-ent Boosting (XGBoost), K-Nearest Neighbor regression (KNN),Linear Support Vector Regressor (SVR), and Linear regression,are utilized to predict the landslide probability and evaluatedon the dataset. The accuracy of the models is assessed by usingstatistical metrics such as mean absolute error (MAE), meansquared error (MSE), and root mean squared error (RMSE).The study results show that Random Forest outperformed othermodels with the mutual information feature selection method.It achieved an MSE of 0.0011 with mutual information-basedfeature selection and an MSE of 0.00157 without feature selection.KNN regressor outperformed the other models with an MSEof 0.00139 with correlation-based information selection. Theproposed landslide identification model with Random Forestalgorithm shows outstanding robustness and great potential intackling the landslide likelihood prediction by employing MLalgorithms.

Vasundhara Acharya↗

Demonstration of Optimal Benchmark Selection Website and Validation of the q c Coverage Metric Using HEU-SOL-THERM-013-003 Experiment

In the work documented in this interim report, the experiment selection toolkit web site was demonstrated and q C coverage metric methodology was validated for IEU-MET-FAST-002-001, MIX-COMP-THERM 004-004, and HEU-SOL-THERM-013-003 experiments. 𝑞 𝐶 is an information-theoretic measure based on mutual information that quantifies the ability of candidate benchmark experiments to reduce the bias and uncertainty of a target criticality safety application. The metric and an accompanying open-source Python toolkit with a web-based interface were tested against a benchmark set of 425 experiments drawn from the International Criticality Safety Benchmark Evaluation Project Handbook. The interface is hosted at https://edim.covdef.com. It accepts sensitivity data files produced by the TSUNAMI-IP module of the SCALE code system and supports both (i) deterministic analysis using the ENDF/B-VII.0 covariance library and (ii) stochastic analysis based on user-supplied keff samples. Demonstrations on representative applications across a range of material composition, spectrum, and form show that q C -guided benchmark selection achieves greater uncertainty reduction with fewer experiments and yields more stable posterior bias and uncertainty estimates than traditional similarity coefficient ( c k )–based selection, while also capturing valuable low-ck experiments that one-to-one metrics overlook.

Abdel-khalik, Hany S. [Indiana Univ.-Purdue Univ. ↗

Information theoretic comparisons of original and transformed data from Landsat MSS and TM

The dispersion and concentration of signal values in transformed data from the Landsat-4 MSS and TM instruments are analyzed using a communications theory approach. The definition of entropy of Shannon was used to quantify information, and the concept of mutual information was employed to develop a measure of information contained in several subsets of variables. Several comparisons of information content are made on the basis of the information content measure, including: system design capacities; data volume occupied by agricultural data; and the information content of original bands and Tasseled Cap variables. A method for analyzing noise effects in MSS and TM data is proposed.

Malila, W. A.↗

Information-Theoretic Exploration of Multivariate Time-Varying Image Databases

Modern scientific simulations produce very large datasets, making interactive exploration of such data computationally prohibitive. An increasingly common data reduction technique is to store visualizations and other data extracts in a database. The Cinema project is one such approach, storing visualizations in an image database for post hoc exploration and interactive image-based analysis. This work focuses on developing efficient algorithms that can quantify various types of multivariate dependencies existing within multi-variable datasets. It applies specific mutual information measures for the quantification of salient regions from multivariate image data. Here, using such information measures, the opacity of the images is modulated so that the salient regions are automatically highlighted and the domain scientists can interactively explore the most relevant regions for scientific discovery.

97 MATHEMATICS AND COMPUTING↗

Optimal Experimental Design With Fast Neural Network Surrogate Models

Designing optimal experiments minimizes the uncertainty of results and maximizes the efficient use of resources. Herein, machine learning surrogate models and the approximate coordinate exchange (ACE) algorithm are used to determine optimum experimental designs over large or arbitrarily restrictive design spaces. Optimal experimental design is particularly salient in materials science where experiments are expensive and material properties must often be inferred indirectly. The proposed framework is demonstrated by finding optimal experiments with which the hidden constituent properties of composite materials can be most efficiently inferred from observable experimental outcomes. The optimum experimental design is given by an information-theoretic criteria, which maximizes the conditional mutual information between the hidden properties and the expected experimental outcomes. To perform tractable optimization a neural network is trained as a surrogate model to mimic a physics based simulation, which can calculate the expected experimental outcome based on a candidate experimental design and sampled constituent properties. The ACE algorithm is used to optimize over large design spaces with many tests and controlled parameters where an exhaustive search would be intractable even with the surrogate model. Using this approach, optimal experimental designs that are consistent with those produced by heuristic knowledge and established best practices are found; then optimal designs in larger design spaces where heuristic knowledge is unavailable are examined.

machine learning↗

Data and scripts associated with “Allometric scaling of hyporheic respiration across basins in the Pacific Northwest USA"

This data package is associated with the publication “Allometric scaling of hyporheic respiration across basins in the Pacific Northwest USA” submitted to JGR-Biogeosciences (Regier et al. 2025).This study used reach-scale modeled estimates of hyporheic aerobic respiration made by the River Corridor Model (Fang et al. 2020) and watershed characteristics across the Willamette and Yakima River basins to explore potential allometric scaling (i.e., power-law relationships between size and function) of cumulative hyporheic respiration across catchment-to-basin scales. Scaling was explored quantitatively via the R2, slope, and y-intercept of relationships between cumulative hyporheic respiration and watershed area, divided into hyporheic exchange flux (HEF) quantiles. We also explored relationships between allometric scaling and other watershed characteristics through linear regression, spatial patterns, and mutual information analyses. Our results also suggest variability of hyporheic respiration allometry for middle exchange flux quantiles, and in relation to land-cover. Our findings provide initial evidence that allometric scaling may be useful for predicting hyporheic biogeochemical dynamics across watersheds from reach to basin scales. This data package is associated with the GitHub repository found at https://github.com/peterregier/rc_wrb_yrb_scaling. The data package is organized into several key directories. The “data” folder contains multiple CSV files, including landscape heterogeneity, scaling analysis, and watershed boundary data. The “figures” folder has all figure files in both PDF and PNG formats. Core analysis scripts and figure generation scripts are in the “scripts” directory, systematically numbered for sequential execution. The root directory includes essential project files; please see the file ending in “flmd.csv” for a list and description of all files contained in this data package and the file ending in “dd.csv” for data dictionaries used to describe tabular column headers.

54 ENVIRONMENTAL SCIENCES↗

A Contextually Supervised Optimization-Based HVAC Load Disaggregation Methodology

This paper presents a novel contextually supervised optimization-based approach for disaggregating heating, ventilation, and air-conditioning (HVAC) loads using smart meter or Supervisory Control and Data Acquisition data. To disaggregate the load into HVAC loads, large and infrequently used loads (LIUL), and base loads, we formulate an optimization problem to minimize a set of five loss terms, consisting of the reconstruction errors of the overall load profile, the ramp rate losses, and three distinct loss functions linked with the HVAC load, base load, and LIUL, respectively. To enhance accuracy, we incorporate two forms of contextual information into the problem formulation. First, we utilize mutual information to estimate HVAC energy consumption. Second, we employ a base load dictionary to constrain HVAC load estimation errors. The obtained HVAC load profiles are fine-tuned by abnormal ramp detection followed by binary hypothesis testing. Here, the proposed method is developed and tested using sub-metered residential and commercial building data. Simulation results show that the proposed method outperforms existing methods across various data resolutions and load aggregation levels, showing excellent transferability and generalizability.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Measurement of Volatile Compounds for Real-Time Analysis of Soil Microbial Metabolic Response to Simulated Snowmelt

Snowmelt dynamics are a significant determinant of microbial metabolism in soil and regulate global biogeochemical cycles of carbon and nutrients by creating seasonal variations in soil redox and nutrient pools. With an increasing concern that climate change accelerates both snowmelt timing and rate, obtaining an accurate characterization of microbial response to snowmelt is important for understanding biogeochemical cycles intertwined with soil. However, observing microbial metabolism and its dynamics non-destructively remains a major challenge for systems such as soil. Microbial volatile compounds (mVCs) emitted from soil represent information-dense signatures and when assayed non-destructively using state-of-the-art instrumentation such as Proton Transfer Reaction-Time of Flight-Mass Spectrometry (PTR-TOF-MS) provide time resolved insights into the metabolism of active microbiomes. In this study, we used PTR-TOF-MS to investigate the metabolic trajectory of microbiomes from a subalpine forest soil, and their response to a simulated wet-up event akin to snowmelt. Using an information theory approach based on the partitioning of mutual information, we identified mVC metabolite pairs with robust interactions, including those that were non-linear and with time lags. The biological context for these mVC interactions was evaluated by projecting the connections onto the Kyoto Encyclopedia of Genes and Genomes (KEGG) network of known metabolic pathways. Simulated snowmelt resulted in a rapid increase in the production of trimethylamine (TMA) suggesting that anaerobic degradation of quaternary amine osmo/cryoprotectants, such as glycine betaine, may be important contributors to this resource pulse. Unique and synergistic connections between intermediates of methylotrophic pathways such as dimethylamine, formaldehyde and methanol were observed upon wet-up and indicate that the initial pulse of TMA was likely transformed into these intermediates by methylotrophs. Increases in ammonia oxidation signatures (transformation of hydroxylamine to nitrite) were observed in parallel, and while the relative role of nitrifiers or methylotrophs cannot be confirmed, the inferred connection to TMA oxidation suggests either a direct or indirect coupling between these processes. Overall, it appears that such mVC time-series from PTR-TOF-MS combined with causal inference represents an attractive approach to non-destructively observe soil microbial metabolism and its response to environmental perturbation.

59 BASIC BIOLOGICAL SCIENCES↗

Preventing Reverse Engineering of Critical Industrial Data with DIOD

Business analytics augmented by artificial intelligence and machine learning (AI/ML) have revolutionized the role of data in the modern world. In recent years, businesses have incorporated data into their decision-making process for better prediction, risk-assessment, content creation, etc. While such businesses often seek to leverage the full use of their data through third-party AI/ML services, they are often hampered by the risks of data leaks, reverse-engineering, stolen technology, etc. that often have disastrous consequences for businesses and their stakeholders alike. Thus, there arises a need for data masking prior to its transmission that obfuscates proprietary information while preserving the information relevant for AI/ML applications. In order to meet the needs of industrial data which are significantly different from those of data warehouses, previous work proposed an efficient time and space-scalable data masking paradigm known as the deceptive infusion of data (DIOD) methodology. The present work expands upon this work by leveraging existing reverse-engineering capabilities to facilitate the decomposition of industrial data into its proprietary and AI/ML-relevant parts, referred to as fundamental and inference metadata respectively. Both sets of metadata are further obfuscated in accordance with the DIOD methodology to create the DIOD rendition of the industrial data, which is rendered immune to reverse-engineering by discarding proprietary information and only preserving AI/ML-relevant information. Additionally, constraints of the original DIOD manuscript are relaxed using mutual information by configuring the methodology to the target AI/ML application to unlock the full potential of the DIOD methodology. As an example, data from a nuclear reactor is transformed into that from a nonlinear spring-mass system with different levels of data masking as required by the generic system and the target application.

97 MATHEMATICS AND COMPUTING↗

On the origin of red spirals: does assembly bias play a role?

The formation of the red spirals is a puzzling issue in the standard picture of galaxy formation and evolution. Most studies attribute the colour of the red spirals to different environmental effects. We analyze a volume limited sample from the SDSS to study the roles of small-scale and large-scale environments on the colour of spiral galaxies. We compare the star formation rate, stellar age and stellar mass distributions of the red and blue spirals and find statistically significant differences between them at 99.9% confidence level. The red spirals inhabit significantly denser regions than the blue spirals, explaining some of the observed differences in their physical properties. However, the differences persist in all types of environments, indicating that the local density alone is not sufficient to explain the origin of the red spirals. Using an information theoretic framework, we find a small but non-zero mutual information between the colour of spiral galaxies and their large-scale environment that are statistically significant (99.9% confidence level) throughout the entire length scale probed. Such correlations between the colour and the large-scale environment of spiral galaxies may result from the assembly bias. Thus both the local environment and the assembly bias may play essential roles in forming the red spirals. The spiral galaxies may have different assembly history across all types of environments. We propose a picture where the differences in the assembly history may produce spiral galaxies with different cold gas content. Such a difference would make some spirals more susceptible to quenching. Finally, in all environments, the spirals with high cold gas content could delay the quenching and maintain a blue colour, whereas the spirals with low cold gas fractions would be easily quenched and become red.

79 ASTRONOMY AND ASTROPHYSICS↗

Security Analysis of a Class of Spread Spectrum Systems Presentation

A method of adding physical layer security to a class of spread spectrum systems has been recently proposed. In this paper, we look into the rate at which an eavesdropper may gain information about the system to decipher the data symbols. The Shannon mutual information is used to measure the rate of information that may be gained by an eavesdropper. The k-nearest neighbors (k-NN) method is used to obtain estimates of relevant entropy values, which will then be used to quantify the rate of information recovery as more data is transmitted. It turns out that such information recovery requires the adoption of special methods that avoid any destructive bias in the estimates. Details of these methods are also presented.

97 - MATHEMATICS AND COMPUTING↗