Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “knowledge representation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 307 records · Page 17

Unifying Combinatorial and Graphical Methods in Artificial Intelligence

Recently, a new graph Laplacian, called the inner product Laplacian, was introduced which generalizes many existing Laplacians, including the normalized and combinatorial Laplacian and their weighted variants. The key observation behind the inner product Laplacian is that by defining appropriate inner product spaces on the vertices and edges, the standard Laplacians can be recovered as Hodge Laplacians over the simplicial complex formed by the edges and vertices. These inner product spaces form a natural way to incorporate non-combinatorial information into the definition of a domain-specific Laplacian. In particular, in contrast to current domain-specific weighting schemes which rely solely on edge weights, information regarding the similarity of non-adjacent vertices and arbitrary pairs of edges can be effectively incorporated into the Laplacian. In order to illustrate this approach we consider the problem of calculating the potential energy of an atomistic configuration using Graph Neural Networks. In comparison with start-of-the-art approaches, such as SchNet, our approach replaces a learned (via auto-encoder) representation of the atom types with an inner product space on atoms based on scientific knowledge (e.g., electronegativity). We will illustrate how this approach captures key chemical properties of the molecules and compare the energy calculations with state-of-the-art neural network approaches. However, to compute the resulting Laplacian involves a mixture of sparse and dense matrix computation and yields a dense matrix as the basis for the graph convolution. This dense convolutional kernel necessitates moving away from the standard message passing framework for graph neural networks and increases the computational cost of applying the kernel. In order to mitigate these costs we investigate means of leveraging the mixed sparse and dense computations to reduce the overall computational cost and how these approaches can be automatically transferred to energy efficient hardware (e.g., field programmable gate arrays (FPGAs)).

97 MATHEMATICS AND COMPUTING↗

Exaflops Biomedical Knowledge Graph Analytics

We are motivated by newly proposed methods for mining large-scale corpora of scholarly publications (e.g., full biomedical literature), which consists of tens of millions of papers spanning decades of research. In this setting, analysts seek to discover relationships among concepts. They construct graph representations from annotated text databases and then formulate the relationship-mining problem as an all-pairs shortest paths (APSP) and validate connective paths against curated biomedical knowledge graphs (e.g., Spoke). In this context, we present Coast (Exascale Communication-Optimized All-Pairs Shortest Path) and demonstrate 1.004 EF/s on 9,200 Frontier nodes (73,600 GCDs). We develop hyperbolic performance models (HYPERMOD), which guide optimizations and parametric tuning. The proposed Coast algorithm achieved the memory constant parallel efficiency of 99% in the single-precision tropical semiring. Looking forward, Coast will enable the integration of scholarly corpora like PubMed into the Spoke biomedical knowledge graph.

Kannan, Ramakrishnan {ramki}↗

Knowledge Oriented Graph Unified Transformer (KOGUT) v0.1

KOGUT — Knowledge Oriented Graph Unified Transformer KOGUT implements the Relational Graph Transformer (RelGT) architecture for knowledge graph link prediction in biological domains, with a primary focus on microbial growth media prediction. While the original RelGT (arXiv:2505.10960) targets relational tables, time series, and multi-table databases, KOGUT adapts this architecture for heterogeneous biological knowledge graphs, providing first-in-class AI predictive models for microbial cultivation. Key Adaptations Beyond Original RelGT: - Knowledge Graph Focus: Applied to biological KGs with semantic node types (taxa, chemicals, media, phenotypes, environments) versus generic relational database tables, trained on the KG-Microbe knowledge graph (1.3M entities, 2.9M edges, 24 relation types). - Multimodal Node Encoding: Integrates node labels, categories, descriptions, and synonyms from KG metadata through learned embedding layers—adapting relational column features to graph node attributes with textual semantics. - Extended K-Hop Subgraph Strategy: Optimized neighborhood sampling (3-hop default, configurable up to 200 nodes) tuned for sparse biological networks, building on the original local-global attention framework with biological relation preservation. - Biolink Predicate Preservation: Type-specific transformations for 24 biological edge semantics (occurs_in, consumes, produces, has_phenotype, subclass_of) beyond standard relational foreign keys, enabling multi-relation link prediction. - Inductive Learning Support: Enables zero-shot predictions for novel taxa through feature-based embeddings (temperature, oxygen requirements, gram stain, cell shape), extending the original transductive relational benchmark scope to uncultured microorganisms. CheapSOTA Performance Optimizations (This Distribution): - VQ-EMA Centroid Attention: Vector quantization with exponential moving average for improved global context modeling (+5-10% MRR improvement). - HDF5 Precomputed Data Loading: One-time preprocessing of k-hop subgraphs to eliminate redundant graph traversals (2-5× training speedup). - Distributed Data Parallel Training: Multi-GPU support for scaling to larger knowledge graphs (tested on 4× NVIDIA A100 GPUs at NERSC Perlmutter). - Mixed Precision Training: Automatic mixed precision (AMP) for memory efficiency and faster training. Advantages Over Standard Knowledge Graph Embedding Models: Combines RelGT's proven multi-element tokenization (features, type, hop, structure) with graph-native biological representations, enabling interpretable link prediction across heterogeneous entities that standard embedding models (TransE, RotatE, ComplEx) and table-based transformers cannot directly model. Achieves near-perfect performance on microbial growth media prediction (MRR: 0.9966, Precision@1: 0.9932, Hit@10: 1.0000) while maintaining explainability through attention-based reasoning over biological pathways. Training Data: - KG-Microbe merged knowledge graph: 1,379,337 nodes, 2,960,472 edges - 24 biological relation types including taxonomic hierarchies, metabolic interactions, phenotype associations, and environmental relationships - Primary prediction task: Growth media suitability for microbial taxa (biolink:occurs_in, 50K edges) - Multi-relation capability: Predicts links for any of the 24 relation types, including chemical consumption/production, phenotype associations, and taxonomic classification Citation: Original RelGT Architecture: Dwivedi et al., "Relational Graph Transformer", arXiv:2505.10960, 2025 KOGUT Implementation: Knowledge Oriented Graph Unified Transformer for Microbial Growth Media Prediction Developed at Lawrence Berkeley National Laboratory (LBNL) Trained on NERSC Perlmutter supercomputer

Joachimiak, Marcin [Lawrence Berkeley National Lab↗

Language model-accelerated deep symbolic optimization

Symbolic optimization methods have been used to solve varied challenging and relevant problems such as symbolic regression and neural architecture search. However, the current state of the art typically learns each problem from scratch and is unable to leverage pre-existing knowledge and datasets that are available for many applications. Here, inspired by the similarity between sequence representations learned in natural language processing and the formulation of symbolic optimization as a discrete sequence optimization problem, we propose language model-accelerated deep symbolic optimization (LA-DSO), a method that leverages language models to learn symbolic optimization solutions more efficiently. We demonstrate LA-DSO in two tasks: symbolic regression, which allows us to perform extensive experimentation due to its low computation requirements, and computational antibody optimization, which shows that our proposal accelerates learning in challenging real-world problems.

97 MATHEMATICS AND COMPUTING↗

Progress toward a universal biomedical data translator

Clinical, biomedical, and translational science has reached an inflection point in the breadth and diversity of available data and the potential impact of such data to improve human health and well-being. However, the data are often siloed, disorganized, and not broadly accessible due to discipline-specific differences in terminology and representation. To address these challenges, the Biomedical Data Translator Consortium has developed and tested a pilot knowledge graph-based “Translator” system capable of integrating existing biomedical data sets and “translating” those data into insights intended to augment human reasoning and accelerate translational science. Having demonstrated feasibility of the Translator system, the Translator program has since moved into development, and the Translator Consortium has made significant progress in the research, design, and implementation of an operational system. Herein, we describe the current system’s architecture, performance, and quality of results. We apply Translator to several real-world use cases developed in collaboration with subject-matter experts. Finally, we discuss the scientific and technical features of Translator and compare those features to other state-of-the-art, biomedical graph-based question-answering systems.

60 APPLIED LIFE SCIENCES↗

Analysis of Interpretable Data Representations for 4D-STEM Using Unsupervised Learning

Abstract Understanding the structure of materials is crucial for engineering devices and materials with enhanced performance. Four-dimensional scanning transmission electron microscopy (4D-STEM) is capable of mapping nanometer-scale local crystallographic structure over micron-scale field of views. However, 4D-STEM datasets can contain tens of thousands of images from a wide variety of material structures, making it difficult to automate detection and classification of structures. Traditional automated analysis pipelines for 4D-STEM focus on supervised approaches, which require prior knowledge of the material structure and cannot describe anomalous or deviant structures. In this article, a pipeline for engineering 4D-STEM feature representations for unsupervised clustering using non-negative matrix factorization (NMF) is introduced. Each feature is evaluated using NMF and results are presented for both simulated and experimental data. It is shown that some data representations more reliably identify overlapping grains. Additionally, real space refinement is applied to identify spatially distinct sample regions, allowing for size and shape analysis to be performed. This work lays the foundation for improved analysis of nanoscale structural features in materials that deviate from expected crystallographic arrangement using 4D-STEM.

Bruefach, Alexandra (ORCID:0000000209323477)↗

Use of machine learning to analyze chemistry card sort tasks

Education researchers are deeply interested in understanding the way students organize their knowledge. Card sort tasks, which require students to group concepts, are one mechanism to infer a student’s organizational strategy. However, the limited resolution of card sort tasks means they necessarily miss some of the nuance in a student’s strategy. Here in this work, we propose new machine learning strategies that leverage a potentially richer source of student thinking: free-form written language justifications associated with student sorts. Using data from a university chemistry card sort task, we use vectorized representations of language and unsupervised learning techniques to generate qualitatively interpretable clusters, which can provide unique insight in how students organize their knowledge. We compared these to machine learning analysis of the students’ sorts themselves. Machine learning-generated clusters revealed different organizational strategies than those built into the task; for example, sorts by difficulty or even discipline. There were also many more categories generated by machine learning for what we would identify as more novice-like sorts and justifications than originally built into the task, suggesting students’ organizational strategies converge when they become more expert-like. Finally, we learned that categories generated by machine learning for students’ justifications did not always match the categories for their sorts, and these cases highlight the need for future research on students’ organizational strategies, both manually and aided by machine learning. In sum, the use of machine learning to analyze results from a card sort task has helped us gain a more nuanced understanding of students’ expertise, and demonstrates a promising tool to add to existing analytic methods for card sorts.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Integrative Modeling and Analysis of Fungal Central Carbon Metabolism

Over a thousand fungal genomes have been sequenced, yet manually curated genome-scale metabolic models (GEMs) are available for only a limited number of species. Moreover, these models have often been developed independently, leading to inconsistencies in namespaces, compartment definitions, and pathway representations that hinder comparative analysis, the systematic reuse of prior curation efforts, and the integration of consolidated metabolic knowledge. Here, we present the Consolidated Fungal Core Metabolism Model (CFCMM), constructed by integrating thirteen published fungal models spanning Ascomycota, Mucoromycota, and both Crabtree-positive and Crabtree-negative yeasts. We harmonized metabolites and reactions into a non-redundant shared ModelSEED ontological space, standardized compartmentalization, and refined gene–protein–reaction (GPR) rules. Using pathway-level visualization and systematic gap detection, we further improved the integrated network through literature-guided curation to correct stoichiometry, stereospecificity, and pathway architecture. Orthologous protein family reconstruction and functional annotation workflows were used to validate and inform GPR associations, with particular emphasis on ambiguous enzyme superfamilies and membrane-associated components. Using the resulting CFCMM, we built high-quality central carbon core models for each fungus and performed flux balance analysis to quantify ATP-yield variation under aerobic and anaerobic conditions, explicitly evaluating scenarios driven by differences in electron transport chain (ETC) composition. Simulations reproduced the expected fermentative yield of approximately 2 mmol ATP per mmol glucose under anaerobic conditions and separated the thirteen fungi into two bioenergetic groups under aerobic respiration based on Complex I status, with predicted yields of approximately 30 versus 22 mmol ATP per mmol glucose. Forcing flux through the alternative oxidase bypass further reduced ATP yields to approximately 12 and 4 mmol ATP per mmol glucose in Complex I-containing and Complex I-lacking fungi, respectively. Collectively, this work provides a manually curated, ModelSEED-consistent, and extensible fungal core metabolic template, deployed in DOE KBase as a resource for automated reconstruction of central carbon core models from any sequenced fungal genome. In addition, the CFCMM provides modular components for developing GEMs with more accurate energy predictions and enables robust comparative analyses of fungal bioenergetics and core metabolic diversity

59 BASIC BIOLOGICAL SCIENCES↗

Strategic roadmap to assess forest vulnerability under air pollution and climate change

Abstract Although it is an integral part of global change, most of the research addressing the effects of climate change on forests have overlooked the role of environmental pollution. Similarly, most studies investigating the effects of air pollutants on forests have generally neglected the impacts of climate change. We review the current knowledge on combined air pollution and climate change effects on global forest ecosystems and identify several key research priorities as a roadmap for the future. Specifically, we recommend (1) the establishment of much denser array of monitoring sites, particularly in the South Hemisphere; (2) further integration of ground and satellite monitoring; (3) generation of flux‐based standards and critical levels taking into account the sensitivity of dominant forest tree species; (4) long‐term monitoring of N, S, P cycles and base cations deposition together at global scale; (5) intensification of experimental studies, addressing the combined effects of different abiotic factors on forests by assuring a better representation of taxonomic and functional diversity across the ~73,000 tree species on Earth; (6) more experimental focus on phenomics and genomics; (7) improved knowledge on key processes regulating the dynamics of radionuclides in forest systems; and (8) development of models integrating air pollution and climate change data from long‐term monitoring programs.

54 ENVIRONMENTAL SCIENCES↗

Evaporation Sub-model Development for Volume of Fluid (eVOF) Method Applicable to Spray-Wall Interaction Including Film Characteristics with Validation at High Pressure and Temperature Conditions (Final Report)

Internal combustion engines have seen a great evolution over the last several decades through application of high pressure direct injection, multiple injections, and other technologies to reduced fuel consumption, NOx, and PM. Although combustion systems with advanced injection strategies have been studied extensively, there exists a significant fundamental knowledge gap on the fuel-spray interactions with the piston surface and chamber walls. Advanced computational codes validated with experimental techniques have to be developed for accurate representation of the drop impingement, fuel film formation, and vaporization. Current engine CFD spray models utilize a Lagrangian framework for modeling which lacks critical considerations of the physics pertaining to these interactions and thus requiring extensive parameterization, tuning and validation. The team from Michigan Technological University, University of Massachusetts Dartmouth, and Argonne National Laboratory is composed of experts in sprays, combustion, engines and CFD with a wide spectrum of knowledge including specific expertise in the area under consideration. In the proposed work, a VOF modeling approach has been adopted for the spray-wall interaction, film formation and spreading, and vaporization. With the inclusion of a vaporization submodel, a more predictive and accurate simulation of the spray-film was performed without extensive need of parameterization and tuning. Extensive experimentation of the spray-wall interaction under the range of conditions matching the thermodynamic and surface temperatures that occur in diesel and gasoline engines were conducted to validate the SWI submodels and for development of the evaporation sub-model, which has been implemented in the Converge software.

42 ENGINEERING↗

High temporal resolution generation expansion planning for the clean energy transition

As power systems integrate increasing quantities of wind, solar and energy storage resources, it is important to revisit power system capacity expansion modeling methods and assumptions that have been utilized in thermal- dominated systems. We conduct a series of case study analyses using a simplified representation of the Electric Reliability Council of Texas (ERCOT) system to demonstrate how least-cost capacity expansion outcomes are impacted by changes in model resolution across two temporal dimensions: 1) the number of considered representative periods, and 2) the system dispatch interval. First, we find that the least-cost generation portfolio can differ significantly for small changes in the number of representative days, but largely converges to the 365-day result once 104 representative days are considered. Furthermore, systems with wind, solar and storage resources were more sensitive to changes in the number of representative days than a thermal-dominated system. Second, we find that considering five-minute dispatch resolution consistently results in least-cost generation portfolios with less solar capacity and more energy storage capacity than corresponding scenarios with hourly dispatch intervals. This suggests that hourly dispatch representation fails to capture the intra-hour volatility of solar generation, and therefore also overlooks opportunities for storage resources to provide system value by balancing this volatility. Collectively these results indicate that capacity expansion modelers should revisit conventional approaches to temporal representation when conducting analyses of deeply decarbonized power systems to ensure that such analyses are robust and actionable. To our knowledge, this is the first study to analyze capacity expansion outcomes with five-minute dispatch resolution in this manner.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Thermodynamics of order and randomness in dopant distributions inferred from atomically resolved imaging

Abstract Exploration of structure-property relationships as a function of dopant concentration is commonly based on mean field theories for solid solutions. However, such theories that work well for semiconductors tend to fail in materials with strong correlations, either in electronic behavior or chemical segregation. In these cases, the details of atomic arrangements are generally not explored and analyzed. The knowledge of the generative physics and chemistry of the material can obviate this problem, since defect configuration libraries as stochastic representation of atomic level structures can be generated, or parameters of mesoscopic thermodynamic models can be derived. To obtain such information for improved predictions, we use data from atomically resolved microscopic images that visualize complex structural correlations within the system and translate them into statistical mechanical models of structure formation. Given the significant uncertainties about the microscopic aspects of the material’s processing history along with the limited number of available images, we combine model optimization techniques with the principles of statistical hypothesis testing. We demonstrate the approach on data from a series of atomically-resolved scanning transmission electron microscopy images of Mo x Re 1- x S 2 at varying ratios of Mo/Re stoichiometries, for which we propose an effective interaction model that is then used to generate atomic configurations and make testable predictions at a range of concentrations and formation temperatures.

25 ENERGY STORAGE↗

Metagenomic clustering links specific metabolic functions to globally relevant ecosystems

ABSTRACT Metagenomic sequencing has advanced our understanding of biogeochemical processes by providing an unprecedented view into the microbial composition of different ecosystems. While the amount of metagenomic data has grown rapidly, simple-to-use methods to analyze and compare across studies have lagged behind. Thus, tools expressing the metabolic traits of a community are needed to broaden the utility of existing data. Gene abundance profiles are a relatively low-dimensional embedding of a metagenome’s functional potential and are, thus, tractable for comparison across many samples. Here, we compare the abundance of KEGG Ortholog Groups (KOs) from 6,539 metagenomes from the Joint Genome Institute’s Integrated Microbial Genomes and Metagenomes (JGI IMG/M) database. We find that samples cluster into terrestrial, aquatic, and anaerobic ecosystems with marker KOs reflecting adaptations to these environments. For instance, functional clusters were differentiated by the metabolism of antibiotics, photosynthesis, methanogenesis, and surprisingly GC content. Using this functional gene approach, we reveal the broad-scale patterns shaping microbial communities and demonstrate the utility of ortholog abundance profiles for representing a rapidly expanding body of metagenomic data. IMPORTANCE Metagenomics, or the sequencing of DNA from complex microbiomes, provides a view into the microbial composition of different environments. Metagenome databases were created to compile sequencing data across studies, but it remains challenging to compare and gain insight from these large data sets. Consequently, there is a need to develop accessible approaches to extract knowledge across metagenomes. The abundance of different orthologs (i.e., genes that perform a similar function across species) provides a simplified representation of a metagenome’s metabolic potential that can easily be compared with others. In this study, we cluster the ortholog abundance profiles of thousands of metagenomes from diverse environments and uncover the traits that distinguish them. This work provides a simple to use framework for functional comparison and advances our understanding of how the environment shapes microbial communities.

54 ENVIRONMENTAL SCIENCES↗

Evaporation Submodel Development for Volume of Fluid (eVOF) Method Applicable to Spray-Wall Interaction Including Film Characteristics with Validation at High Pressure and Temperature Conditions

Internal combustion engines have seen a great evolution over the last several decades through application of high pressure direct injection, multiple injections, and other technologies to reduced fuel consumption, NOx, and PM. Although combustion systems with advanced injection strategies have been studied extensively, there exists a significant fundamental knowledge gap on the fuel-spray interactions with the piston surface and chamber walls. Advanced computational codes validated with experimental techniques have to be developed for accurate representation of the drop impingement, fuel film formation, and vaporization. Current engine CFD (Computational Fluid Dynamics) spray models utilize a Lagrangian framework for modeling which lacks critical considerations of the physics pertaining to these interactions and thus requiring extensive parameterization, tuning and validation. The team from Michigan Technological University, University of Massachusetts Dartmouth, and Argonne National Laboratory is composed of experts in sprays, combustion, engines and CFD with a wide spectrum of knowledge including specific expertise in the area under consideration. In the proposed work, a VOF (Volume of Fluid) modeling approach has been adopted for the spray-wall interaction, film formation and spreading, and vaporization. With the inclusion of a vaporization submodel, a more predictive and accurate simulation of the spray-film was performed without extensive need of parameterization and tuning. Extensive experimentation of the spray-wall interaction under the range of conditions matching the thermodynamic and surface temperatures that occur in diesel and gasoline engines were conducted to validate the SWI submodels and for development of the evaporation submodel, which has been implemented in the flow solver.

42 ENGINEERING↗

The Influence of Visual Provenance Representations on Strategies in a Collaborative Hand-off Data Analysis Scenario

Conducting data analysis tasks rarely occur in isolation. Especially in intelligence analysis scenarios where different experts contribute knowledge to a shared understanding, members must communicate how insights develop to establish common ground among collaborators. The use of provenance to communicate analytic sensemaking carries promise by describing the interactions and summarizing the steps taken to reach insights. Yet, no universal guidelines exist for communicating provenance in different settings. Our work here focuses on the presentation of provenance information and the resulting conclusions reached and strategies used by new analysts. In an open-ended, 30-minute, textual exploration scenario, we qualitatively compare how adding different types of provenance information (specifically data coverage and interaction history) affects analysts' confidence in conclusions developed, propensity to repeat work, filtering of data, identification of relevant information, and typical investigation strategies. We see that data coverage (i.e., what was interacted with) provides provenance information without limiting individual investigation freedom. On the other hand, while interaction history (i.e., when something was interacted with) does not significantly encourage more mimicry, it does take more time to comfortably understand, as represented by less confident conclusions and less relevant information-gathering behaviors. In conclusion, our results contribute empirical data towards understanding how provenance summarizations can influence analysis behaviors.

97 MATHEMATICS AND COMPUTING↗

A Study on Efficient Reinforcement Learning Through Knowledge Transfer

Although Reinforcement Learning (RL) algorithms have made impressive progress in learning complex tasks over the past years, there are still prevailing short-comings and challenges. Specifically, the sample-inefficiency and limited adaptation across tasks often make classic RL techniques impractical for real-world applications despite the gained representational power when combining deep neural networks with RL, known as Deep Reinforcement Learning (DRL). Recently, a number of approaches to address those issues have emerged. Many of those solutions are based on smart DRL architectures that enhance single task algorithms with the capability to share knowledge between agents and across tasks by introducing Transfer Learning (TL) capabilities. Here this survey addresses strategies of knowledge transfer from simple parameter sharing to privacy preserving federated learning and aims at providing a general overview of the field of TL in the DRL domain, establishes a classification framework, and briefly describes representative works in the area.

97 MATHEMATICS AND COMPUTING↗

Machine learning on neutron and x-ray scattering and spectroscopies

Neutron and x-ray scattering represent two classes of state-of-the-art materials characterization techniques that measure materials structural and dynamical properties with high precision. These techniques play critical roles in understanding a wide variety of materials systems from catalysts to polymers, nanomaterials to macromolecules, and energy materials to quantum materials. In recent years, neutron and x-ray scattering have received a significant boost due to the development and increased application of machine learning to materials problems. This article reviews the recent progress in applying machine learning techniques to augment various neutron and x-ray techniques, including neutron scattering, x-ray absorption, x-ray scattering, and photoemission. We highlight the integration of machine learning methods into the typical workflow of scattering experiments, focusing on problems that challenge traditional analysis approaches but are addressable through machine learning, including leveraging the knowledge of simple materials to model more complicated systems, learning with limited data or incomplete labels, identifying meaningful spectra and materials representations, mitigating spectral noise, and others. We present an outlook on a few emerging roles machine learning may play in broad types of scattering and spectroscopic problems in the foreseeable future.

Chen, Zhantao↗

Enhancing occupant behavior representation for interoperability between building information modeling and building energy modeling

Building Performance Simulation (BPS) has been adopted as an essential tool for designing, operating, and retrofitting buildings to optimize energy efficiency throughout the building life cycle. The Green Building XML (gbXML) schema facilitates seamless data exchange between Building Information Modeling (BIM) and Building Energy Modeling (BEM) software tools. However, limited occupant behavior (OB) representation in BIM often leads to inconsistent and inaccurate energy simulation in BEM software. This paper presents 154 systematic enhancements to the existing occupant behavior XML (obXML) schema v1.3.4, initially developed for standardizing OB representation for BEM, to address existing limitations and improve interoperability with BIM models. The enhancements encompass improved integration with BIM models through extended building representations and system operations, expanded support for advanced OB models with additional environmental parameters and mathematical capabilities, and implementation of a standardized model documentation framework. To facilitate seamless data transformation between gbXML and obXML schemas, we developed a publicly available gb-obXML Schema Converter. Three case studies demonstrate the enhanced schema’s capabilities: representation of building information using a two-story office building model, documentation of a window operation behavior model, and validation of the schema converter’s functionality. The enhanced obXML schema v1.4 enables sophisticated modeling of occupant-building interactions while maintaining consistency with industry-standard BIM schemas. The standardized documentation framework facilitates reproducibility and knowledge sharing in the OB research community, while the schema converter automates the integration of building information into OB simulation workflows. These enhancements establish a foundation for more accurate building performance simulation by supporting sophisticated representation of occupant behavior within the BIM-to-BEM simulation workflows.

Chung, Jihoon↗