Engineering PapersSearch

SEARCH · Engineering Papers

Results for “Abstract Interpretation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Knowledge-guided learning with curated prior genetic biomarkers for robust model interpretation

Abstract Motivation Knowledge-guided learning offers effective and robust model training strategies in data-scarce settings by incorporating established domain knowledge, thereby enhancing generalization, robustness, and interpretability. By contrast, conventional deep learning approaches rely purely on data-driven learning, which can limit robust model interpretability, particularly in high-dimensional settings with limited size samples. In computational biology, knowledge-guided learning has primarily leveraged network- and structural-based knowledge, leading to biologically interpretable representations and enhanced predictive performance compared to conventional approaches. However, curated biomarkers, one of the most accessible forms of biological knowledge, remain largely unexplored within knowledge-guided paradigms. Results In this study, we propose a model-agnostic training paradigm, Biomarker-driven Explainable Prior-guided Learning (BioExPL), that can be applied to any neural networks that incorporates curated prior knowledge. BioExPL enforces neural networks to reflect curated biomarker priors in their latent representations through a novel knowledge-alignment loss. BioExPL consistently demonstrated significantly improved predictive performance and enhanced model interpretability with minimized computational overhead in simulation studies and intensive experiments on multiple cancer datasets. BioExPL not only integrates prior curated knowledge into the model but also accurately identifies unknown associated signals additionally. BioExPL is model-agnostic and domain-independent, enabling its integration into diverse neural network architectures. Availability and implementation The open-source is publicly available at: https://github.com/datax-lab/BioExPL.

Baek, Beomsu [Department of Computer Science, Univ

Integrating adaptive learning with post hoc model explanation and symbolic regression to build interpretable surrogate models

Abstract We develop a materials informatics workflow to build an interpretable surrogate model for micromagnetic simulations. Our goal is to predict the energy barrier of a moving isolated skyrmion in rare-earth-free $$\hbox {Mn}_4$$ Mn 4 N. Our approach integrates adaptive learning with post hoc model explanation and symbolic regression methods. We discuss an unexplored acquisition function (information condensing active learning) within the adaptive learning loop and compare it with the known standard deviation function for efficient navigation of the search space. Model-agnostic post hoc explanation techniques then uncover trends learned by the trained model, which we then leverage to constrain the expressions used for symbolic regression. Graphical abstract

Biswas, Ankita

Poster Abstract: Leveraging Large Language Models to Reveal Interpretable Cooling Behaviors from Smart Thermostat Data

Frequent heatwaves and hot summers increasingly challenge occupant comfort, health, and energy grid stability. Addressing these challenges requires a detailed understanding of household cooling behaviors, such as thermostat adjustments and adaptive responses to extreme conditions. Traditional analyses often rely on aggregated numerical metrics that overlook subtle but important household-specific variations. In this study, we introduce a generalizable methodology that integrates large language models (LLMs) with vision capabilities to enable scalable and detailed analysis of residential thermostat data. Using Ecobee's Donate Your Data (DYD) dataset—which provides five-minute records of indoor temperatures, thermostat setpoints, and HVAC runtimes—we focus on two U.S. cities with contrasting summer climates : Austin (TX) and Phoenix (AZ). Because raw time-series data are not well suited for direct LLM analysis, we transform them into visual representations, such as daily indoor temperature trajectories and weekly runtime histograms, to better capture behavioral variations. Leveraging LLMs' visual interpretation, we extract descriptive behavioral features, including temperature preferences, time-of-day cooling orientation, anticipatory versus reactive heatwave responses, and behavioral consistency. These semantic features support unsupervised clustering to identify distinct occupant archetypes at scale, revealing differences—such as morning-centric anticipatory coolers versus households that shift toward warmer setpoints during heatwaves—that can inform demand response, resilience planning, and health-aware interventions. By converting raw numerical data into interpretable behavioral patterns, this methodology enables scalable and practical analysis of occupant behavior, supporting actionable insights for comfort, resilience, and energy management.

Nihar, Kopal

A Unified Interpretation of Variability in Precipitation Isotope Ratios

Abstract Several mechanisms have been proposed to explain why the isotope ratios of precipitation vary in space and time and why they correlate with other climate variables like temperature and precipitation. Here, we argue that this behavior is best understood through the lens of radiative transfer, which treats the depletion of atmospheric vapor transport by precipitation as analogous to the attenuation of light by absorption or scattering. Building on earlier work by Siler et al., we introduce a simple model that uses the equations of radiative transfer to approximate the two-dimensional pattern of the oxygen isotope composition of precipitation ( δ p ) from monthly mean hydrologic variables. The model accurately simulates the spatial and seasonal variability in δ p within a state-of-the-art climate model and permits a simple decomposition of δ p variability into contributions from gradients in evaporation and the length scale of vapor transport. Outside the tropics, δ p is mostly controlled by gradients in evaporation, whose dependence on temperature explains the positive correlation between δ p and temperature (i.e., the temperature effect). At low latitudes, δ p is mostly controlled by gradients in the transport length scale, whose inverse relationship with precipitation explains the negative correlation between δ p and precipitation (i.e., the amount effect). This suggests that the temperature and amount effects are both mostly explained by the variability in upstream rainout, but they reflect distinct mechanisms governing rainout at different latitudes. Significance Statement The isotopic composition of precipitation has long been used to make inferences about past climates based on its observed relationship with precipitation in the tropics and with temperature at higher latitudes. These relationships—known as the “amount effect” and “temperature effect,” respectively—have been attributed to many different mechanisms, most of which are thought to operate at either high or low latitudes but not both. Here, we present a unified framework for interpreting the isotope variability that can explain the latitude dependence of the temperature and amount effects despite making no distinction between high and low latitudes. Although our results are generally consistent with certain interpretations of the amount effect, they suggest that the temperature effect is widely misunderstood.

54 ENVIRONMENTAL SCIENCES

eDNAjoint: An R package for interpreting paired or semi‐paired environmental DNA and traditional survey data in a Bayesian framework

Abstract Environmental DNA (eDNA) sampling is increasingly used in surveys of species distribution as a potentially sensitive and efficient monitoring method. Yet access to modelling tools designed specifically for interpreting this new data type lags behind its ubiquity. While occupancy modelling software has dominated the analytical landscape for eDNA data analysis of single species, this type of model may not always be the most appropriate. The rate of eDNA detection often corresponds to species density, rather than just occupancy, and researchers often have access to observations from non‐genetic sampling methods at the same sites. To provide users access to a modelling framework designed to maximize the use of all available data, we developed an R package, eDNAjoint . The package provides an easy‐to‐use interface for fitting a ‘joint’ model that integrates data from paired or semi‐paired eDNA and traditional surveys in a Bayesian framework. The model can be used to estimate parameters like the probability of a false positive eDNA detection and mean catch rate at a site, and the package allows access to multiple model variations and Bayesian prior customization. Additional functionality can be used for model selection, summarising posteriors and comparing the relative sensitivities of the two survey methods. We demonstrate the use of eDNAjoint by fitting a variation of the model with site‐level covariates that scale the sensitivity of eDNA sampling relative to traditional sampling. The example workflow uses binary eDNA and seine count data for the endangered tidewater goby ( Eucyclogobius newberryi ) from a study by Schmelzle and Kinziger (2016). This use case includes a prior sensitivity analysis and an evaluation of the relationship between detection rates and environmental variables. eDNAjoint has the potential to greatly increase the range of users who will be able to rigorously analyse eDNA and traditional survey data in a Bayesian framework, understand if and how eDNA can improve monitoring practices, and gain confidence in the interpretability of eDNA data.

Keller, Abigail G. [Department of Environment Scie

Interpreting Mass and Radius Measurements of Neutron Stars with Dark Matter Halos

Abstract The high densities of neutron stars (NSs) could provide astrophysical locations for dark matter (DM) to accumulate. Depending on the DM model, these DM admixed NSs (DANSs) could have significantly different properties than pure baryonic NSs, accessible through X-ray observations of rotation-powered pulsars. We adopt the two-fluid formalism in general relativity to numerically simulate stable configurations of DANSs, assuming a fermionic equation of state (EOS) for the DM with repulsive self-interaction. The distribution of DM in the DANS as a halo affects the path of X-rays emitted from hot spots on the visible baryonic surface, causing notable changes in the pulse profile observed by telescopes such as NICER, compared to pure baryonic NSs. We explore how various DM models affect the DM mass distribution, leading to different types of dark halos. We quantify the deviation in observed X-ray flux from stars with each of these halos. We identify the pitfalls in interpreting mass and radius measurements of NSs inferred from electromagnetic radiation and constraining the baryonic matter EOS if these dark halos exist.

Shawqi, Shafayat (ORCID:0000000210956183)

Commutative Algebra Modeling in Materials Science – A Case Study on Metal–Organic Frameworks (MOFs)

Metal-organic frameworks (MOFs) are a class of important crystalline and highly porous materials whose hierarchical geometry and chemistry hinder interpretable predictions in materials properties. Commutative algebra is a branch of abstract algebra that has been rarely applied in data and material sciences. We introduce the first ever commutative algebra modeling and prediction in materials science. Specifically, category-specific commutative algebra (CSCA) is proposed as a new framework for MOF representation and learning. It integrates element-based categorization with multiscale algebraic invariants to encode both local coordination motifs and global network organization of MOFs. These algebraically consistent, chemically aware representations enable compact, interpretable, and data efficient modeling of MOF properties such as Henry’s constants and uptake capacities for common gases. Compared to traditional geometric and graph-based approaches, CSCA achieves comparable or superior predictive accuracy while substantially improving interpretability and stability across data sets. By aligning commutative algebra with the chemical hierarchy, the CSCA establishes a rigorous and generalizable paradigm for understanding structure and property relationships in porous materials and provides a nonlinear algebra-based framework for data-driven material discovery.

Khaemba, Caleb S.

The Analysis Description Language Ecosystem: Latest developments and physics applications

We present latest developments in Analysis Description Language (ADL), a declarative domain-specific language describing the physics algorithm of a HEP data analysis decoupled from software frameworks. Analyses written in ADL can be integrated into any framework for various tasks. ADL is a multipurpose construct with uses ranging from analysis design to preservation, reinterpretation, queries, visualisation, combination, etc. The most advanced infrastructure to execute ADL on events is the CutLang runtime interpreter. Recent technical developments include an automated interface with different data types, generation of the abstract syntax tree, a visualization tool that that auto-converts analysis flows to graphs, incorporation of trained machine learning models and a Jupyter-based plotting tool. We also report physics implications including a large scale LHC analysis implementation and validation effort for beyond the standard model reinterpretation purposes and studies with ATLAS and CMS open data.

Sekmen, Sezen [Kyungpook National Univ., Daegu (Ko

Element Formation in Radiation-hydrodynamics Simulations of Kilonovae

Abstract Understanding the details of r -process nucleosynthesis in binary neutron star merger (BNSM) ejecta is key to interpreting kilonova observations and identifying the role of BNSMs in the origin of heavy elements. We present a self-consistent, two-dimensional, ray-by-ray radiation-hydrodynamic evolution of BNSM ejecta with an online nuclear network (NN) up to a timescale of days. For the first time, an initial numerical relativity ejecta profile composed of the dynamical component and spiral-wave and disk winds is evolved including detailed r -process reactions and nuclear heating effects. A simple model for the jet energy deposition is also included. Our simulation highlights that the common approach of relating in postprocessing the final nucleosynthesis yields to the initial thermodynamic profile of the ejecta can lead to inaccurate predictions. Moreover, we find that neglecting the details of the radiation-hydrodynamic evolution of the ejecta in nuclear calculations can introduce deviations of up to 1 order of magnitude in the final abundances of several elements, including very light and second r -process peak elements. The presence of a jet affects element production only in the innermost part of the polar ejecta, and it does not alter the global nucleosynthesis results. Overall, our analysis shows that employing an online NN improves the reliability of nucleosynthesis and kilonova light-curve predictions.

Magistrelli, Fabio (ORCID:0009000509767851)

The importance of electron scattering in the analysis of actinide X-ray spectroscopy

Abstract Manifestations of electron scattering in X-ray spectroscopy have been evident for decades. Here, it will be shown that the proper interpretation of variants of X-ray Absorption Spectroscopy (XAS) of actinide materials must include an accurate treatment of features caused by electron scattering, i.e., EXAFS or Extended X-ray Absorption Fine Structure. These EXAFS features can be of such low energy that they are within ten to twenty electron volts of the Unoccupied Density of States (UDOS), immediately above the Fermi Energy or Band Gap. The adaption of simple models using the FEFF simulation program will be presented, including the demonstration of the robust nature of the results from different models. Graphical abstract

Tobin, J. G. (ORCID:0000000322943301)

AI-powered topic modeling: comparing LDA and BERTopic in analyzing opioid-related cardiovascular risks in women

Topic modeling is a crucial technique in natural language processing (NLP), enabling the extraction of latent themes from large text corpora. Traditional topic modeling, such as Latent Dirichlet Allocation (LDA), faces limitations in capturing the semantic relationships in the text document although it has been widely applied in text mining. BERTopic, created in 2022, leveraged advances in deep learning and can capture the contextual relationships between words. In this work, we integrated Artificial Intelligence (AI) modules to LDA and BERTopic and provided a comprehensive comparison on the analysis of prescription opioid-related cardiovascular risks in women. Opioid use can increase the risk of cardiovascular problems in women such as arrhythmia, hypotension etc. 1,837 abstracts were retrieved and downloaded from PubMed as of April 2024 using three Medical Subject Headings (MeSH) words: “opioid,” “cardiovascular,” and “women.” Machine Learning of Language Toolkit (MALLET) was employed for the implementation of LDA. BioBERT was used for document embedding in BERTopic. Eighteen was selected as the optimal topic number for MALLET and 23 for BERTopic. ChatGPT-4-Turbo was integrated to interpret and compare the results. The short descriptions created by ChatGPT for each topic from LDA and BERTopic were highly correlated, and the performance accuracies of LDA and BERTopic were similar as determined by expert manual reviews of the abstracts grouped by their predominant topics. The results of the t-SNE (t-distributed Stochastic Neighbor Embedding) plots showed that the clusters created from BERTopic were more compact and well-separated, representing improved coherence and distinctiveness between the topics. Our findings indicated that AI algorithms could augment both traditional and contemporary topic modeling techniques. In addition, BERTopic has the connection port for ChatGPT-4-Turbo or other large language models in its algorithm for automatic interpretation, while with LDA interpretation must be manually, and needs special procedures for data pre-processing and stop words exclusion. Therefore, while LDA remains valuable for large-scale text analysis with resource constraints, AI-assisted BERTopic offers significant advantages in providing the enhanced interpretability and the improved semantic coherence for extracting valuable insights from textual data.

Research & Experimental Medicine

Replacing non-biomedical concepts improves embedding of biomedical concepts

Embeddings are semantically meaningful representations of words in a vector space, commonly used to enhance downstream machine learning applications. Traditional biomedical embedding techniques often replace all synonymous words representing biological or medical concepts with a unique token, ensuring consistent representation and improving embedding quality. However, the potential impact of replacing non-biomedical concept synonyms has received less attention. Embedding approaches often employ concept replacement to replace concepts that span multiple words, such as non-small-cell lung carcinoma, with a single concept identifier (e.g., D002289). Also, all synonyms of each concept are merged into the same identifier. Here, we additionally leveraged WordNet to identify and replace sets of non-biomedical synonyms with their most common representatives. This combined approach aimed to reduce embedding noise from non-biomedical terms while preserving the integrity of biomedical concept representations. We applied this method to 1,055 biomedical concept sets representing molecular signatures or medical categories and assessed the mean pairwise distance of embeddings with and without non-biomedical synonym replacement. A smaller mean pairwise distance was interpreted as greater intra-cluster coherence and higher embedding quality. Embeddings were generated using the Word2Vec algorithm applied to a corpus of 10 million PubMed abstracts. Our results demonstrate that the addition of non-biomedical synonym replacement reduced the mean intra-cluster distance by an average of 8%, suggesting that this complementary approach enhances embedding quality. Future work will assess its applicability to other embedding techniques and downstream tasks. Python code implementing this method is provided under an open-source license.

algorithms

Dynamic allostery in the peptide/MHC complex enables TCR neoantigen selectivity

Abstract The inherent antigen cross-reactivity of the T cell receptor (TCR) is balanced by high specificity. Surprisingly, TCR specificity often manifests in ways not easily interpreted from static structures. Here we show that TCR discrimination between an HLA-A*03:01 (HLA-A3)-restricted public neoantigen and its wild-type (WT) counterpart emerges from distinct motions within the HLA-A3 peptide binding groove that vary with the identity of the peptide’s first primary anchor. These motions create a dynamic gate that, in the presence of the WT peptide, impedes a large conformational change required for TCR binding. The neoantigen is insusceptible to this limiting dynamic, and, with the gate open, upon TCR binding the central tryptophan can transit underneath the peptide backbone to the opposing side of the HLA-A3 peptide binding groove. Our findings thus reveal a novel mechanism driving TCR specificity for a cancer neoantigen that is rooted in the dynamic and allosteric nature of peptide/MHC-I binding grooves, with implications for resolving long-standing and often confounding questions about T cell specificity.

Science & Technology - Other Topics

Chemical ionization mass spectrometry utilizing benzene cations for measurements of volatile organic compounds and nitric oxide

We evaluate the capability of chemical ionization mass spectrometry (CIMS) using benzene cations as reagent ions (benzene CIMS) for detecting atmospheric trace gases. We characterize the ionization pathways and product ion distributions for 27 analytes spanning diverse chemical classes. To interpret the complex ion chemistry involving two reagent ions (C 6 H$^{+}_{6}$ and (C 6 H 6 )$^{+}_{2}$) and multiple ionization pathways (charge transfer, proton transfer, adduct formation, and hydride abstraction), we introduce a thermodynamics-based framework that classifies analytes into three categories based on their ionization energy (IE), relative to those of benzene monomer (9.24 eV) and dimer (8.69 eV). Each class exhibits distinct ionization mechanisms and product ions. Analytes with IE smaller than 8.69 eV (low IE) undergo charge transfer with both reagent ions; analytes with IE between 8.69 and 9.24 eV (mid IE) undergo charge transfer with C 6 H$^{+}_{6}$ and potential adduct formation with (C 6 H 6 )$^{+}_{2}$; analytes with IE larger than 9.24 eV (high IE) could undergo adduct formation, proton transfer, or hydride abstraction. Analytes within each class also show similar sensitivity, enabling sensitivity estimation for compounds lacking calibration standards. In addition to volatile organic compounds (VOCs), benzene CIMS detects nitric oxide (NO) with a detection limit of 5 pptv for 1 min integration time, exceeding the performance of most commercial NOx analyzers. Field deployments in Chicago and St. Louis demonstrate good agreement with reference NO measurements. Isoprene measurements show good agreement with a co-located gas chromatography–photoionization detector (GC-PID) in St. Louis, but exhibit substantial positive bias in Chicago, likely due to interferences from anthropogenic VOCs in the polluted urban environment. These results highlight the potential of benzene CIMS for concurrent measurements of NO, VOCs, and their oxidation products using a single instrument, while also underscoring challenges in complex atmospheric conditions.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH

Chiral-odd generalized parton distributions in the large-𝑁 𝑐 limit of QCD: Spin-flavor structure, polynomiality, and sum rules

We study the nonperturbative properties of the nucleon’s chiral-odd generalized parton distributions (transversity GPDs) in the large-𝑁 𝑐 limit of QCD. This includes the parametric ordering of the spin-flavor components, the polynomiality property of the moments, and the sum rules connecting the GPDs with the tensor form factors. A multipole expansion in the transverse momentum transfer is used to enumerate and interpret the structures in the nucleon matrix element of the chiral-odd partonic operator, including monopole, dipole and quadrupole terms. The 1/𝑁 𝑐 expansion of the GPDs is performed using the abstract mean-field picture of baryons in the large-𝑁 𝑐 limit and its symmetries. We derive a large-𝑁 𝑐 relation between the flavor-nonsinglet GPDs 𝐸$^{𝑢−𝑑}_𝑇$ and $\tilde{𝐻}^{𝑢−𝑑}_𝑇$ and test it with recent lattice QCD results. We show that the polynomiality property and sum rules of the GPDs are fulfilled with the restricted realization of translational and rotational invariance in the mean-field picture. The results provide a basis for the phenomenological analysis of chiral-odd GPDs and hard exclusive processes in the large-𝑁 𝑐 limit, and for calculations in specific dynamical models.

generalized parton distributions

New horizon symmetries, hydrodynamics, and quantum chaos

Abstract We generalize the formulation of horizon symmetries presented in previous literature to include diffeomorphisms that can shift the location of the horizon. In the context of the AdS/CFT duality, we show that horizon symmetries can be interpreted on the boundary as emergent low-energy gauge symmetries. In particular, we identify a new class of horizon symmetries that extend the so-called shift symmetry, which was previously postulated for effective field theories of maximally chaotic systems. Additionally, we comment on the connections of horizon symmetries with bulk calculations of out-of-time-ordered correlation functions and the phenomenon of pole-skipping.

Physics

Modern chemical graph theory

Abstract Graph theory has a long history in chemistry. Yet as the breadth and variety of chemical data is rapidly changing, so too do graph encoding methods and analyses that yield qualitative and quantitative insights. Using illustrative cases within a basic mathematical framework, we showcase modern chemical graph theory's utility in Chemists' analysis and model development toolkit. The encoding of both experimental and simulation data is discussed at various levels of granularity of information. This is followed by a discussion of the two major classes of graph theoretical analyses: identifying connectivity patterns and partitioning methods. Measures, metrics, descriptors, and topological indices are then introduced with an emphasis upon enhancing interpretability and incorporation into physical models. Challenging data cases are described that include strategies for studying time dependence. Throughout, we incorporate recent advancements in computer science and applied mathematics that are propelling chemical graph theory into new domains of chemical study. This article is categorized under: Molecular and Statistical Mechanics > Molecular Dynamics and Monte‐Carlo Methods Structure and Mechanism > Computational Materials Science Structure and Mechanism > Molecular Structures

Leite, Leonardo S. G.

Fracture Characterization Via AI‐Assisted Analysis of Temperature Logs

Abstract Fractures control fluid flow, mass transport, and heat transfer in a geothermal reservoir. This makes accurate characterization of fracture networks a prerequisite for optimal design and control of a reservoir's exploitation. We develop a deep‐learning procedure to identify fracture locations via interpretation of temporally and spatially continuous downhole temperature measurements. A long short‐term memory fully convolutional network (LSTM‐FCN) is used both to capture long‐term dependencies in sequential temperature data and to distill local features around fractures. A wellbore and fractured‐reservoir thermal model is established to generate temperature data for network training. The trained LSTM‐FCN exhibits a unique ability to detect multiple fractures intersecting a borehole. We use the LSTM‐FCN algorithm to evaluate the effectiveness of different‐stage wellbore temperature measurements on fracture detection in a complex fractured system. Our experiments reveal that the use of various‐stage temperature information as an input feature set improves the robustness of fracture detection to noise interference. This study indicates the practical feasibility of obtaining accurate fracture‐network reconstructions from temperature signals, at reasonable computational cost.

Yang, Xiaoyu