Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “knowledge representation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Molecular dynamics study on the interfacial properties of mixtures of monomers of polyvinylpyrrolidone (PVP)-based battery binders on graphene and graphite surfaces

This study investigates the behavior of two different mixtures of monomers of polyvinylpyrrolidone (PVP)-based battery binders, polyvinylpyrrolidone:polyvinylidene difluoride (PVP:PVDF) and polyvinylpyrrolidone:polyacrylic acid (PVP:PAA), at graphene and graphite interfaces using classical molecular dynamics simulations. The aim is to identify the best performing monomer binder blend and carbon-based material for the design of battery-optimized energy devices. The PVP:PAA monomer binder blend and graphite are found to have the best interaction energies, densification upon adsorption, and more ordered structure. The adsorption of both monomer binder blends is strongly guided by the higher affinity of PVP and PAA monomeric molecules for the surfaces compared to PVDF. The structure of adsorbed layers of PVP:PVDF monomer binder blend on graphene and graphite develops more quickly than PVP:PAA, indicating faster kinetics. This study complements a previous density functional theory study recently reported by our group and contributes to a better understanding of the nanoscopic features of relevant interfacial regions involving mixtures of monomers of PVP-based battery binders and different carbon-based materials. In conclusion, the effect of a blend of commonly used monomer binders on carbon-based materials is essential for obtaining tightly bound anode and cathode active materials in lithium-ion batteries, which is crucial for designing battery-optimized energy devices.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Minimal implicit-solvent coarse-grained simulation of Pluronic block copolymers with ionic liquids

Pluronic block copolymers, composed of poly(ethylene oxide) (PEO) and poly(propylene oxide) (PPO) in a triblock structure (PEO–PPO–PEO), are well known for their amphiphilic character and ability to self‐assemble into micelles in aqueous solution. The addition of ionic liquids (ILs) can further modulate the core–shell structures of these copolymers, influencing their stability, critical micellization temperature, and size. However, fully atomistic simulations often become prohibitively expensive due to the size and complexity of these systems. In this work, coarse‐grained simulations using a minimal implicit‐solvent model were performed to examine how two classes of ILs, namely, 1‐alkyl‐3‐methylimidazolium ([C n C 1 im]) and 1‐alkyl‐3‐methylpyrrolidinium ([C n C 1 pyrr]), change the micellization of Pluronic block copolymers in aqueous solution. The effects of IL concentration and alkyl group length were investigated, and the model greatly improved the efficiency of simulating large‐scale micelle systems. Furthermore, the numerical simulations are qualitatively compared with experimental investigations. Our results show that adding ILs expands the micelle core by embedding IL tails among the PPO blocks, thereby increasing overall micelle size. Less polar ILs generally induce more pronounced micellar growth. However, the effect of IL tail length on conformation and micellar packing is non‐monotonic. Up to moderate chain lengths (around C8–C10), the IL tails can extend sufficiently to increase local separation within the micelle; at longer tail lengths, enhanced hydrophobic clustering and steric hindrance cause the tails to bend or fold, capping further expansion. In addition, although block copolymer chains tend to pack more closely in the presence of longer‐tailed ILs, the random coil size of an individual polymer chain does not necessarily shrink. Meanwhile, these insights provide a deeper understanding of how Pluronic/IL systems interact, informing applications in drug delivery, cosmetics, food, and environmental engineering. Finally, our minimal implicit‐solvent model can be applied to larger systems and longer timescales, substantially reducing computational cost while reproducing key structural trends observed experimentally.

Atomistic simulations↗

Facet-dependent structure and dissociation of water at pristine IrO 2 /water interfaces

Understanding the microscopic structure of water at metal oxide interfaces is crucial for advancing electrocatalysis. IrO 2 , specifically, has shown exceptional activity for electrochemical water oxidation, but we currently lack a fundamental understanding of how the surface structure of IrO 2 impacts water reactivity. In this work, we developed a machine learning potential trained to first-principles accuracy for modeling IrO 2 /water interfaces across different facets: (110), (100), (101), and (001). Using extensive machine learning molecular dynamics simulations, we investigated the spontaneous dissociation of water molecules at these interfaces. Our results reveal a distinct dissociation probability trend: (110) > (100) ≈ (101) > (001), which we attribute primarily to the reaction thermodynamics of surface water dissociation. A strong correlation is observed between the surface Ir–O bond distances and the dissociation probabilities, highlighting the role of surface geometry in modulating reactivity. As a consequence, the interfacial solvation structures and hydrogen bonding environments are dynamically tuned by the varying water dissociation capabilities across facets. This work elucidates how water dissociation energetics depend on surface orientation and interfacial structure, offering atomistic insights into manipulating reaction chemistry at electrocatalytic interfaces.

organic↗

Physical discovery in representation learning via conditioning on prior knowledge

Recent advances in electron, scanning probe, optical, and chemical imaging and spectroscopy yield bespoke data sets containing the information of structure and functionality of complex systems. In many cases, the resulting data sets are underpinned by low-dimensional simple representations encoding the factors of variability within the data. The representation learning methods seek to discover these factors of variability, ideally further connecting them with relevant physical mechanisms. However, generally, the task of identifying the latent variables corresponding to actual physical mechanisms is extremely complex. Here, we present an empirical study of an approach based on conditioning the data on the known (continuous) physical parameters and systematically compare it with the previously introduced approach based on the invariant variational autoencoders. The conditional variational autoencoder (cVAE) approach does not rely on the existence of the invariant transforms and hence allows for much greater flexibility and applicability. Interestingly, cVAE allows for limited extrapolation outside of the original domain of the conditional variable. However, this extrapolation is limited compared to the cases when true physical mechanisms are known, and the physical factor of variability can be disentangled in full. We further show that introducing the known conditioning results in the simplification of the latent distribution if the conditioning vector is correlated with the factor of variability in the data, thus allowing us to separate relevant physical factors. We initially demonstrate this approach using 1D and 2D examples on a synthetic data set and then extend it to the analysis of experimental data on ferroelectric domain dynamics visualized via piezoresponse force microscopy.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Semantic Property Graph for Scalable Knowledge Graph Analytics

Graphs are a natural and fundamental representation to describe entities, relationships, activities, and evolution of complex systems. Many domains such as communication, citation, procurement, biology, social media, and transportation can be modeled as a set of entities and their relationships. Resource Description Framework (RDF) and Labeled Property Graph (LPG) are two of the most used data models to encode information in a graph. Both models are similar in terms of using basic graph elements such as nodes and edges but differ in terms of the modeling approach, expressibility, serialization, and target applications. RDF is a flexible data exchange model for expressing information about entities but it tends to a have high memory footprint and inefficient storage, which does not make it a natural choice to perform scalable graph analytics. In contrast, LPG has gained traction as a reliable model to perform scalable graph analytic tasks such as sub-graph matching, network alignment, and real-time knowledge graph query. It provides efficient storage, fast traversal, and flexibility to model various real-world domains. At the same time, the LPG lacks the support of a formal knowledge representation such as an ontology to provide automated knowledge inference. We propose Semantic Property Graph (SPG) as a logical projection of reified RDF into the LPG model. SPG continues to use RDF ontology to define the type hierarchy of the projected graph and validate it against a given ontology. We present a framework to convert reified RDF graphs into SPG using two different computing environments. We also present cloud-based graph migration capabilities using Amazon Web Services.

Purohit, Sumit↗

Bridging 20 Years of Soil Organic Matter Frameworks: Empirical Support, Model Representation, and Next Steps

Abstract In the past few decades, there has been an evolution in our understanding of soil organic matter (SOM) dynamics from one of inherent biochemical recalcitrance to one deriving from plant‐microbe‐mineral interactions. This shift in understanding has been driven, in part, by influential conceptual frameworks which put forth hypotheses about SOM dynamics. Here, we summarize several focal conceptual frameworks and derive from them six controls related to SOM formation, (de)stabilization, and loss. These include: (a) physical inaccessibility; (b) organo‐mineral and ‐metal stabilization; (c) biodegradability of plant inputs; (d) abiotic environmental factors; (e) biochemical reactivity and diversity; and (f) microbial physiology and morphology. We then review the empirical evidence for these controls, their model representation, and outstanding knowledge gaps. We find relatively strong empirical support and model representation of abiotic environmental factors but disparities between data and models for biochemical reactivity and diversity, organo‐mineral and ‐metal stabilization, and biodegradability of plant inputs, particularly with respect to SOM destabilization for the latter two controls. More empirical research on physical inaccessibility and microbial physiology and morphology is needed to deepen our understanding of these critical SOM controls and improve their model representation. The SOM controls are highly interactive and also present some inconsistencies which may be reconciled by considering methodological limitations or temporal and spatial variation. Future conceptual frameworks must simultaneously refine our understanding of these six SOM controls at various spatial and temporal scales and within a hierarchical structure, while incorporating emerging insights. This will advance our ability to accurately predict SOM dynamics.

54 ENVIRONMENTAL SCIENCES↗

Reducing Model Uncertainty of Climate Change Impacts on High Latitude Carbon Assimilation

The Arctic Boreal Region (ABR) has a large impact on global vegetation-atmosphere interactions and is experiencing markedly greater warming than the rest of the planet, a trend that is projected to continue with anticipated future emissions of CO 2 . The ABR is a significant source of uncertainty in estimates of carbon uptake in terrestrial biosphere models (TBMs) such that reducing this uncertainty is critical for more accurately estimating global carbon cycling and understanding the response of the region to global change. Process representation and parameterization associated with gross primary productivity (GPP) drives a large amount of this model uncertainty, particularly within the next 50 years, where the response of existing vegetation to climate change will dominate estimates of GPP for the region. Furthermore, we review our current understanding and model representation of GPP in northern latitudes, focusing on vegetation composition, phenology and physiology, and consider how climate change alters these three components. We highlight challenges in the ABR for predicting GPP, but also focus on the unique opportunities for advancing knowledge and model representation, particularly through the combination of remote sensing and traditional boots-on-the-ground science.

54 ENVIRONMENTAL SCIENCES↗

Artificial Reasoning System for Symptom-Based Conditional Failure Probability Estimation Using Bayesian Network

Advances in nuclear power technologies require enhanced capabilities for operator advice and autonomous control. One of the first tasks in the development of such capabilities is the formulation of symptom-based conditional failure probabilities for structures, systems, and components (SSCs) of interest, for which the primary goal is to aid plant personnel in deducing the probabilistic performance status of the monitored SSCs and in detecting impending faults/failure. The task of conditional failure probability estimation is a bidirectional inference problem and shall be logically tackled by the Bayesian network (BN) approach. As a knowledge-based artificial intelligence tool and a probabilistic graphical model, BN offers the capability of reasoning under uncertainty and graphical representation emulating the physical behavior of the target SSC. This paper provides a systematic overview of the BN technique and the software tools for handling implementation of BN models, along with the associated knowledge representation and reasoning paradigm. Both operational data and expert judgement can be readily incorporated into the knowledge base of a BN model. The challenges with data availability are highlighted, and the general approach to target SSC identification is presented. Our focus is upon failure-prone and risk-important balance of plant assets, especially cases having strong operator involvement. An exemplary case study on the failure of a motor-driven centrifugal pump is also conducted to demonstrate the usefulness and technical feasibility of the proposed artificial reasoning system using an expert system shell.

Zhao, Xingang↗

Quantifying microbial control of soil organic matter dynamics at macrosystem scales

Soil organic matter (SOM) stocks, decomposition and persistence are largely the product of controls that act locally. Yet the controls are shaped and interact at multiple spatiotemporal scales, from which macrosystem patterns in SOM emerge. Theory on SOM turnover recognizes the resulting spatial and temporal conditionality in the effect sizes of controls that play out across macrosystems, and couples them through evolutionary and community assembly processes. For example, climate history shapes plant functional traits, which in turn interact with contemporary climate to influence SOM dynamics. Selection and assembly also shape the functional traits of soil decomposer communities, but it is less clear how in turn these traits influence temporal macrosystem patterns in SOM turnover. Here, we review evidence that establishes the expectation that selection and assembly should generate decomposer communities across macrosystems that have distinct functional effects on SOM dynamics. Representation of this knowledge in soil biogeochemical models affects the magnitude and direction of projected SOM responses under global change. Yet there is high uncertainty and low confidence in these projections. To address these issues, we make the case that a coordinated set of empirical practices are required which necessitate (1) greater use of statistical approaches in biogeochemistry that are suited to causative inference; (2) long-term, macrosystem-scale, observational and experimental networks to reveal conditionality in effect sizes, and embedded correlation, in controls on SOM turnover; and (3) use of multiple measurement grains to capture local- and macroscale variation in controls and outcomes, to avoid obscuring causative understanding through data aggregation. Here, when employed together, along with process-based models to synthesize knowledge and guide further empirical work, we believe these practices will rapidly advance understanding of microbial controls on SOM and improve carbon cycle projections that guide policies on climate adaptation and mitigation.

59 BASIC BIOLOGICAL SCIENCES↗

Improving the Quasi‐Biennial Oscillation via a Surrogate‐Accelerated Multi‐Objective Optimization

Accurate simulation of the quasi-biennial oscillation (QBO) is challenging due to uncertainties in representing convectively generated gravity waves. We develop an end-to-end uncertainty quantification workflow that calibrates these gravity wave processes in E3SM for a realistic QBO. Central to our approach is a domain knowledge-informed, compressed representation of high-dimensional spatio-temporal wind fields. By employing a parsimonious statistical model that learns the fundamental frequency from complex observations, we extract interpretable and physically meaningful quantities capturing key attributes. Building on this, we train a probabilistic surrogate model that approximates the fundamental characteristics of the QBO as functions of critical physics parameters governing gravity wave generation. Leveraging the Karhunen–Loève decomposition, our surrogate efficiently represents these characteristics as a set of orthogonal features, capturing cross-correlations among multiple physics quantities evaluated at different pressure levels and enabling rapid surrogate-based inference at a fraction of the computational cost of full-scale simulations. Finally, we analyze the inverse problem using a multi-objective approach. Our study reveals a tension between amplitude and period that constrains the QBO representation, precluding a single optimal solution. To navigate this, we quantify the bi-criteria trade-off and generate a set of Pareto optimal parameter values that balance the conflicting objectives. This integrated workflow improves the fidelity of QBO simulations and offers a versatile template for uncertainty quantification in complex geophysical models.

54 ENVIRONMENTAL SCIENCES↗

Ontologizing health systems data at scale: making translational discovery a reality

Common data models solve many challenges of standardizing electronic health record (EHR) data but are unable to semantically integrate all of the resources needed for deep phenotyping. Open Biological and Biomedical Ontology (OBO) Foundry ontologies provide computable representations of biological knowledge and enable the integration of heterogeneous data. However, mapping EHR data to OBO ontologies requires significant manual curation and domain expertise. We introduce OMOP2OBO, an algorithm for mapping Observational Medical Outcomes Partnership (OMOP) vocabularies to OBO ontologies. Using OMOP2OBO, we produced mappings for 92,367 conditions, 8611 drug ingredients, and 10,673 measurement results, which covered 68–99% of concepts used in clinical practice when examined across 24 hospitals. When used to phenotype rare disease patients, the mappings helped systematically identify undiagnosed patients who might benefit from genetic testing. By aligning OMOP vocabularies to OBO ontologies our algorithm presents new opportunities to advance EHR-based deep phenotyping.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

BuildingQA: A Benchmark for Natural Language Question Answering over Building Knowledge Graphs

Graph-based representations of building metadata using ontologies like Brick are vital for smart building applications, but querying them remains a challenge for practitioners. Knowledge Graph Question Answering (KGQA) systems, meant to retrieve answers from natural language questions, traditionally require large-scale training data, making them ill-suited for the specialized and data-scarce building domain. The advent of Large Language Models (LLMs) offers a paradigm shift, enabling zero-shot natural language querying without building/domain-specific training. Yet, there is no standardized benchmark for building-specific KGQA which can guide and validate research in this area. To address this gap, our work makes three primary contributions. First, we introduce the BuildingQA Benchmark Dataset, constructed through a multi-stage process of collecting practitioner data, augmenting it with LLMs for linguistic diversity, and curating a final set of 188 questions across 4 buildings. Second, we characterize the benchmark's complexity and ambiguity, introducing a novel method to quantify its "lexical gap" and providing a four-stage diagnostic framework for analyzing how systems fail. Third, we benchmark zero-shot LLM-powered KGQA systems to establish baseline performance and analyze their failure modes. Our evaluation reveals that top-performing systems achieve a maximum F1 score of only 0.38. This result does not indicate a failure of these powerful systems, but rather underscores the unique challenges posed by our benchmark. It demonstrates a critical performance gap, showing that current methods successful on general KGs struggle with the specific lexical and structural nuances of the building domain. BuildingQA1 thus provides the benchmark dataset and foundational analysis needed to drive the development of novel, domain-aware methods required to unlock the use of semantic data in buildings.

Mulayim, Ozan Baris↗

Knowledge-guided learning with curated prior genetic biomarkers for robust model interpretation

Abstract Motivation Knowledge-guided learning offers effective and robust model training strategies in data-scarce settings by incorporating established domain knowledge, thereby enhancing generalization, robustness, and interpretability. By contrast, conventional deep learning approaches rely purely on data-driven learning, which can limit robust model interpretability, particularly in high-dimensional settings with limited size samples. In computational biology, knowledge-guided learning has primarily leveraged network- and structural-based knowledge, leading to biologically interpretable representations and enhanced predictive performance compared to conventional approaches. However, curated biomarkers, one of the most accessible forms of biological knowledge, remain largely unexplored within knowledge-guided paradigms. Results In this study, we propose a model-agnostic training paradigm, Biomarker-driven Explainable Prior-guided Learning (BioExPL), that can be applied to any neural networks that incorporates curated prior knowledge. BioExPL enforces neural networks to reflect curated biomarker priors in their latent representations through a novel knowledge-alignment loss. BioExPL consistently demonstrated significantly improved predictive performance and enhanced model interpretability with minimized computational overhead in simulation studies and intensive experiments on multiple cancer datasets. BioExPL not only integrates prior curated knowledge into the model but also accurately identifies unknown associated signals additionally. BioExPL is model-agnostic and domain-independent, enabling its integration into diverse neural network architectures. Availability and implementation The open-source is publicly available at: https://github.com/datax-lab/BioExPL.

Baek, Beomsu [Department of Computer Science, Univ↗

Towards the next generation of Geospatial Artificial Intelligence

Geospatial Artificial Intelligence (GeoAI), as the integration of geospatial studies and AI, has become one of the fastest-developing research directions in spatial data science and geography. This rapid change in the field calls for a deeper understanding of the recent developments and envision where the field is going in the near future. In this work, we provide a quantitative analysis of the GeoAI literature from the spatial, temporal, and semantic aspects. We briefly discuss the history of AI and GeoAI by highlighting some pioneering work. Then we discuss the current landscape of GeoAI by selecting five representative subdomains including remote sensing, urban computing, Earth system science, cartography, and geospatial semantics. Finally, we highlight several unique future research directions of GeoAI which are classified into two groups: GeoAI method development challenges and GeoAI Ethics challenges. Topics include heterogeneity-aware GeoAI, knowledge-guided GeoAI, spatial representation learning, geo-foundation models, fairness-aware GeoAI, privacy-aware GeoAI, as well as interpretable and explainable GeoAI. We hope our review of GeoAI’s past, present, and future is comprehensive and can enlighten the next generation of GeoAI research.

58 GEOSCIENCES↗

Detecting Anomalies in Time Series Using Kernel Density Approaches

This paper introduces a novel anomaly detection approach tailored for time series data with exclusive reliance on normal events during training. Our key innovation lies in the application of kernel-density estimation (KDE) to scrutinize reconstruction errors, providing an empirically derived probability distribution for normal events post-reconstruction. This non-parametric density estimation technique offers a nuanced understanding of anomaly detection, differentiating it from prevalent threshold-based mechanisms in existing methodologies. In post-training, events are encoded, decoded, and evaluated against the estimated density, providing a comprehensive notion of normality. In addition, we propose a data augmentation strategy involving variational autoencoder-generated events and a smoothing step for enhanced model robustness. The significance of our autoencoder-based approach is evident in its capacity to learn normal representation without prior anomaly knowledge. Through the KDE step on reconstruction errors, our method addresses the versatility of anomalies, departing from assumptions tied to larger reconstruction errors for anomalous events. Our proposed likelihood measure then distinguishes normal from anomalous events, providing a concise yet comprehensive anomaly detection solution. The extensive experimental results support the feasibility of our proposed method, yielding significantly improved classification performance by nearly 10% on the UCR benchmark data.

Frehner, Robin↗

EXCHANGE Campaign Degradation (ECD): Understanding Decomposition Dynamics Across Mid-Atlantic and Great Lakes Coastal Ecosystems

The EXploration of Coastal Hydrobiogeochemistry Across a Network of Gradients and Experiments (EXCHANGE) Degradation Experiment (EXCHANGE-D) is an in situ experiment designed to assess organic matter decomposition rates across coastal terrestrial-aquatic interfaces (TAIs), from coastal uplands through transition zones to wetlands. Through a network of partner scientists and coastal sites, we are testing how environmental gradients shape decomposition and carbon dynamics across terrestrial-aquatic interfaces. Using standardized tea bag substrates deployed across a network of diverse coastal sites, we compare decomposition rates at different fresh- and salt-water TAIs to develop transferable knowledge that improves the representation of organic matter degradation in coastal ecosystem models. For more information, please see https://compass.pnnl.gov/FME/EXCHANGE. This is Version 1 of the data package, which includes: ecd_README.pdf flmd.csv dd.csv ecd_soil_weom_L2.csv ecd_soil_ph_conductivity_L2.csv ecd_soil_gwc_L2.csv ecd_soil_teabag_degradation_L2.csv ecd_readme.pdf

coastal soils↗