Engineering PapersSearch

SEARCH · Engineering Papers

Results for “Information theory entropy”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Discrimination of coherent features in turbulent boundary layers by the entropy method

Entropy in information theory is defined as the expected or mean value of the measure of the amount of self-information contained in the ith point of a distribution series x sub i, based on its probability of occurrence p(x sub i). If p(x sub i) is the probability of the ith state of the system in probability space, then the entropy, E(X) = - sigma p(x sub i) logp (x sub i), is a measure of the disorder in the system. Based on this concept, a method was devised which sought to minimize the entropy in a time series in order to construct the signature of the most coherent motions. The constrained minimization was performed using a Lagrange multiplier approach which resulted in the solution of a simultaneous set of non-linear coupled equations to obtain the coherent time series. The application of the method to space-time data taken by a rake of sensors in the near-wall region of a turbulent boundary layer was presented. The results yielded coherent velocity motions made up of locally decelerated or accelerated fluid having a streamwise scale of approximately 100 nu/u(tau), which is in qualitative agreement with the results from other less objective discrimination methods.

Corke, T. C.

Complexity analysis of a CT injection experiment on BRB

In this work, we use Jensen–Shannon complexity and permutation entropy to analyze the magnetic field fluctuations of an astrophysically scaled plasma experiment. The experiment was intended to emulate an interplanetary coronal mass ejection event in the lab, recreating the major sections seen in satellite data. We also use a technique called “delay,” in which we use select elements, skipping one or more data points at a time, in our time series data to obtain Jensen–Shannon complexity as a function of frequency and investigate the frequency of maximized complexity. We then compare the delay frequencies to other frequencies in the plasma. We found that the frequencies for maximum complexity do not correspond to the frequencies investigated, implying that other physical mechanisms lead to an increase in complexity at these frequencies.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY

Structural and compositional complexities of hierarchical self-assembly: A hypergraph approach

Programmable self-assembly enables the construction of complex molecular, supramolecular, and crystalline architectures from well-designed building blocks. In this work, we introduce a hypergraph-based formalism, Blocks & Bonds (B&B), which generalizes classical chemical graph theory by incorporating directed and multicolored interactions, internal symmetries, and hierarchical organization. Within this framework, we develop the Structure Code (SC), a compact and versatile language for describing self-assembled architectures. We define a Kolmogorov-style structural complexity as the total information content of SC, obtained through its tokenization and Shannon information assignment. Complementing this encoding-based measure, we introduce a much simpler quantity, the compositional complexity, which depends only on the number and cumulative usage of block and bond types in the construction set. A central result of this work is a strong empirical correlation between the token-based structural complexity and the compositional complexity across all examined systems. Owing to this agreement, the compositional complexity emerges as the most practical and broadly applicable measure: it is easy to compute, requires no explicit encoding, and yet closely tracks the actual information content of structurally diverse architectures. Applications to molecular systems (ethylene glycol and glucose), DNA-origami lattices, and crystalline assemblies show that B&B hypergraphs provide a unified, scalable, and information-efficient representation of structural organization, naturally capturing symmetry, modularity, and stereochemistry. This framework establishes a quantitative foundation for complexity-aware classification and inverse design of programmable matter.

36 MATERIALS SCIENCE

Maximizing efficiency of dataset compression for machine learning potentials with information theory

Machine learning interatomic potentials (MLIPs) balance high accuracy and lower costs compared to density functional theory calculations, but their performance often depends on the size and diversity of training datasets. Large datasets improve model accuracy and generalization but are computationally expensive to produce and train on, while smaller datasets risk discarding rare but important atomic environments and compromising MLIP accuracy/reliability. Here, we develop an information-theoretical framework to quantify the efficiency of dataset compression methods and propose an algorithm that maximizes this efficiency. By framing atomistic dataset compression as an instance of the minimum set cover (MSC) problem over atom-centered environments, our method identifies the smallest subset of structures that contains as much information as possible from the original dataset while pruning redundant information. The approach is extensively demonstrated on the GAP-20 and TM23 datasets and validated on 64 varied datasets from the ColabFit repository. Across all cases, MSC consistently retains outliers, preserves dataset diversity, and reproduces the long-tail distributions of forces even at high compression rates, outperforming other subsampling methods. Furthermore, MLIPs trained on MSC-compressed datasets exhibit reduced error for out-of-distribution data even in low-data regimes. We explain these results using an outlier analysis and show that such quantitative conclusions could not be achieved with conventional dimensionality reduction methods. The algorithm is implemented in the open-source QUESTS package and can be used for several tasks in atomistic modeling, from data subsampling, outlier detection, and training improved MLIPs at a lower cost.

36 MATERIALS SCIENCE

Entropy of the Quantum–Classical Interface: A Potential Metric for Security

Hybrid quantum–classical systems are emerging as key platforms in quantum computing, sensing, and communication technologies, but the quantum–classical interface (QCI)—the boundary enabling these systems—introduces unique and largely unexplored security vulnerabilities. This position paper proposes using entropy-based metrics to monitor and enhance security, specifically at the QCI. We present a theoretical security outline that leverages well-established information-theoretic entropy measures, such as Shannon entropy, von Neumann entropy, and quantum relative entropy, to detect anomalous behaviors and potential breaches at the QCI. By linking entropy fluctuations to scenarios of practical relevance—including quantum key distribution, quantum sensing, and hybrid control systems—we promote the potential value and applicability of entropy-based security monitoring. While explicitly acknowledging practical limitations and theoretical assumptions, we argue that entropy-based metrics provide a complementary approach to existing security methods, inviting further empirical studies and theoretical refinements that can strengthen future quantum technologies.

97 MATHEMATICS AND COMPUTING

Kullback-Leibler information function and the sequential selection of experiments to discriminate among several linear models

A sequential adaptive experimental design procedure for a related problem is studied. It is assumed that a finite set of potential linear models relating certain controlled variables to an observed variable is postulated, and that exactly one of these models is correct. The problem is to sequentially design most informative experiments so that the correct model equation can be determined with as little experimentation as possible. Discussion includes: structure of the linear models; prerequisite distribution theory; entropy functions and the Kullback-Leibler information function; the sequential decision procedure; and computer simulation results. An example of application is given.

Sidik, S. M.

Information theory lateral density distribution for Earth inferred from global gravity field

Information Theory Inference, better known as the Maximum Entropy Method, was used to infer the lateral density distribution inside the Earth. The approach assumed that the Earth consists of indistinguishable Maxwell-Boltzmann particles populating infinitesimal volume elements, and followed the standard methods of statistical mechanics (maximizing the entropy function). The GEM 10B spherical harmonic gravity field coefficients, complete to degree and order 36, were used as constraints on the lateral density distribution. The spherically symmetric part of the density distribution was assumed to be known. The lateral density variation was assumed to be small compared to the spherically symmetric part. The resulting information theory density distribution for the cases of no crust removed, 30 km of compensated crust removed, and 30 km of uncompensated crust removed all gave broad density anomalies extending deep into the mantle, but with the density contrasts being the greatest towards the surface (typically + or 0.004 g cm 3 in the first two cases and + or - 0.04 g cm 3 in the third). None of the density distributions resemble classical organized convection cells. The information theory approach may have use in choosing Standard Earth Models, but, the inclusion of seismic data into the approach appears difficult.

Rubincam, D. P.

Entropy, instrument scan and pilot workload

Correlation and information theory which analyze the relationships between mental loading and visual scanpath of aircraft pilots are described. The relationship between skill, performance, mental workload, and visual scanning behavior are investigated. The experimental method required pilots to maintain a general aviation flight simulator on a straight and level, constant sensitivity, Instrument Landing System (ILS) course with a low level of turbulence. An additional periodic verbal task whose difficulty increased with frequency was used to increment the subject's mental workload. The subject's looppoint on the instrument panel during each ten minute run was computed via a TV oculometer and stored. Several pilots ranging in skill from novices to test pilots took part in the experiment. Analysis of the periodicity of the subject's instrument scan was accomplished by means of correlation techniques. For skilled pilots, the autocorrelation of instrument/dwell times sequences showed the same periodicity as the verbal task. The ability to multiplex simultaneous tasks increases with skill. Thus autocorrelation provides a way of evaluating the operator's skill level.

Tole, J. R.

Information theoretic comparisons of original and transformed data from Landsat MSS and TM

The dispersion and concentration of signal values in transformed data from the Landsat-4 MSS and TM instruments are analyzed using a communications theory approach. The definition of entropy of Shannon was used to quantify information, and the concept of mutual information was employed to develop a measure of information contained in several subsets of variables. Several comparisons of information content are made on the basis of the information content measure, including: system design capacities; data volume occupied by agricultural data; and the information content of original bands and Tasseled Cap variables. A method for analyzing noise effects in MSS and TM data is proposed.

Malila, W. A.

Measuring Questions: Relevance and its Relation to Entropy

The Boolean lattice of logical statements induces the free distributive lattice of questions. Inclusion on this lattice is based on whether one question answers another. Generalizing the zeta function of the question lattice leads to a valuation called relevance or bearing, which is a measure of the degree to which one question answers another. Richard Cox conjectured that this degree can be expressed as a generalized entropy. With the assistance of yet another important result from Janos Acz6l, I show that this is indeed the case; and that the resulting inquiry calculus is a natural generalization of information theory. This approach provides a new perspective of the Principle of Maximum Entropy.

Knuth, Kevin H.

A safety-based decision making architecture for autonomous systems

Engineering systems designed specifically for space applications often exhibit a high level of autonomy in the control and decision-making architecture. As the level of autonomy increases, more emphasis must be placed on assimilating the safety functions normally executed at the hardware level or by human supervisors into the control architecture of the system. The development of a decision-making structure which utilizes information on system safety is detailed. A quantitative measure of system safety, called the safety self-information, is defined. This measure is analogous to the reliability self-information defined by McInroy and Saridis, but includes weighting of task constraints to provide a measure of both reliability and cost. An example is presented in which the safety self-information is used as a decision criterion in a mobile robot controller. The safety self-information is shown to be consistent with the entropy-based Theory of Intelligent Machines defined by Saridis.

Musto, Joseph C.

Fuzzy geometry, entropy, and image information

Presented here are various uncertainty measures arising from grayness ambiguity and spatial ambiguity in an image, and their possible applications as image information measures. Definitions are given of an image in the light of fuzzy set theory, and of information measures and tools relevant for processing/analysis e.g., fuzzy geometrical properties, correlation, bound functions and entropy measures. Also given is a formulation of algorithms along with management of uncertainties for segmentation and object extraction, and edge detection. The output obtained here is both fuzzy and nonfuzzy. Ambiguity in evaluation and assessment of membership function are also described.

Pal, Sankar K.

Understanding Local Structure Globally in Earth Science Remote Sensing Data Sets

Empirical probability distributions derived from the data are the signatures of physical processes generating the data. Distributions defined on different space-time windows can be compared and differences or changes can be attributed to physical processes. This presentation discusses on ways to reduce remote sensing data in a way that preserves information, focusing on the rate-distortion theory and using the entropy-constrained vector quantization algorithm.

massive data sets

Environmental Controls on Water Vapor Deuterium Excess in the Coastal Boundary Layer: An Information Theory Perspective

We use information theory to quantify the environmental controls on water vapor deuterium excess (D-excess) in coastal Southern California from June 2023 through February 2024. Using Shannon entropy, mutual information (MI), and joint mutual information, metrics that capture both linear and nonlinear relationships, we identify the most informative variables and variable combinations governing D-excess across contrasting marine and continental regimes. Relative humidity with respect to sea surface temperature (RHS) is consistently the strongest individual predictor, explaining up to 27% of D-excess variability during marine conditions but only 10% in continental air masses. The Relative humidity(RHS) + sea surface temperature (SST) combination demonstrates synergistic effects, where their joint influence (explaining up to 36% of D-excess variability) exceeds what either variable achieves individually, confirming their coupled influence on deuterium excess. Wind direction complements RHS most effectively during continental conditions. The best three-variable combination (RHS + SST + Planetary Boundary Layer height) explains 38% of D-excess variability in marine air, while no combination exceeds 20% explanatory power during continental periods. Information theory shows that heteroscedasticity in D-excess relationships indicates regime shifts in controlling processes and quantifies fundamental constraints on predictor variables: some environmental factors like surface pressure or water vapor flux contain insufficient information content to explain D-excess variability regardless of their physical relevance. These results highlight the different predictability limits between marine and continental regimes, challenging the adequacy of linear models and providing a rigorous framework for quantifying the information content of isotope-climate relationships with implications for both modern and paleoclimate applications.

information theory

Comparison of the information contents of Landsat TM and MSS data

A communications-theory approach is taken to analyze the dispersion and concentration of signal values in various data spaces, irrespective of specific class membership. Entropy is used to quantify information, and mutual information is used to measure the information represented by subsets of spectral variables. Several different comparisons of information content are made. These include comparisons of system design capacities, of data volumes occupied by agricultural data in the spaces defined by original bands and by transformed spectral (Tasseled Cap) variables, of the information contents of original bands and Tasseled Cap variables, and of the information contents of TM and MSS for the given agricultural data sets. Also, the effects of sample size, scene content, and quantization level are examined.

Malila, W. A.

Comparison of the information contents of LANDSAT TM and MSS data

A communications-theory approach is taken to analyze the dispersion and concentration of signal values in various data spaces, irrespective of specific class membership. Entropy is used to quantify information, and mutual information is used to measure the information represented by subsets of spectral variables. Several different comparisons of information content are made. These include comparisons of system design capacities, of data volumes occupied by agricultural data in the spaces defined by original bands and by transformed spectral (Tasseled Cap) variables, of the information contents of original bands and Tasseled Cap variables, and of the information contents of TM and MSS for the given agricultural data sets. Also, the effects of sample size, scene content, and quantization level are examined.

Malila, W. A.

A Study of Feature Extraction Using Divergence Analysis of Texture Features

An empirical study of texture analysis for feature extraction and classification of high spatial resolution remotely sensed imagery (10 meters) is presented in terms of specific land cover types. The principal method examined is the use of spatial gray tone dependence (SGTD). The SGTD method reduces the gray levels within a moving window into a two-dimensional spatial gray tone dependence matrix which can be interpreted as a probability matrix of gray tone pairs. Haralick et al (1973) used a number of information theory measures to extract texture features from these matrices, including angular second moment (inertia), correlation, entropy, homogeneity, and energy. The derivation of the SGTD matrix is a function of: (1) the number of gray tones in an image; (2) the angle along which the frequency of SGTD is calculated; (3) the size of the moving window; and (4) the distance between gray tone pairs. The first three parameters were varied and tested on a 10 meter resolution panchromatic image of Maryville, Tennessee using the five SGTD measures. A transformed divergence measure was used to determine the statistical separability between four land cover categories forest, new residential, old residential, and industrial for each variation in texture parameters.

Hallada, W. A.