Engineering PapersSearch

SEARCH · Engineering Papers

Results for “Data structures”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Custom Accessors: Enabling Scalable Data Ingestion, (Re-)Organization, and Analysis on Distributed Systems

The emerging class of high velocity and high volume data analytic workflows comprise interwoven data ingestion, organization, and processing stages, with ingestion and organization steps often contributing comparable or even higher computational costs than actual processing steps. Since complex workflows consist of a variety of phases that view and use data differently, being able to construct efficient, scalable, distributed data structures (arrays, vectors, sets, maps, and multi-maps) is essential and requires custom methods to extend and shrink containers, analyze and position data, and, maintain globallyconsistent meta-data. In this paper, we propose a novel datastructure access paradigm based on the concept of Accessors. At a high level, accessors are customizable callable objects that can modify the behavior of insert, read, update, and delete operations for distributed containers while preserving atomicity guarantees. Accessors provide a very clean and natural way to implement a variety of programming patterns, e.g., conditional insertion/deletion and cascading computations, which would be otherwise hard (or even impossible) to express in parallel and distributed settings without using locks. We demonstrate the practicality and usefulness of our approach with two representative use cases and study the performance of these applications on a distributed High-Performance Computing system. Our analysis highlights that our proposed abstraction allows for an effective overlapping and concurrent execution of different workflow steps (e.g., data ingestion and analysis), which in a conventional analytics pipeline would execute sequentially, contributing cumulatively to the overall latency.

Castellana, Vito G. [BATTELLE (PACIFIC NW LAB)] (O

Nuclear Data Sheets for A=154

The experimental results published before Aug 2022 from the various reaction and decay studies leading to nuclides of Z=56 to Z=72, 154 Ba, 154 La, 154 Ce, 154 Pr, 154 Nd, 154 Pm, 154 Sm, 154 Eu, 154 Gd, 154 Tb, 154 Dy, 154 Ho, 154 Er, 154 Tm, 154 Yb, 154 Lu, 154 Hf, in the A=154 mass chain have been reviewed. These data are collected and presented in decay or reaction datasets, together with Adopted Levels and gammas datasets that are the most extensive collections of nuclear structure data for each nuclide. Furthermore this work is intended to supersede the previous evaluation of the A=154 nuclides by C.W. Reich (2009Re14), which was published in Nuclear Data Sheets 110, 2257 (2009).

Nica, N. [Texas A&M University, College Station, T

Low‐dimensional manifold learning for uncertainty quantification in complex multi‐scale stochastic systems

Broadly speaking, the goals of the project are to develop techniques to use manifold learning to develop reduced‐order and surrogate models for "hyper‐reduction" of very high‐dimensional complex multi‐scale systems. This is being achieved by employing a newly proposed form of manifold projection and learning that leverages recent advancements in computational geometry and data‐driven modeling. In particular, we are applying a manifold projection technique to project the solutions of very high‐dimensional systems onto the so‐called Grassmannmanifold, a Reimannian manifold comprised of orthonormal matrices. We then apply data‐driven machine learning techniques to classify the solutions on the manifold (e.g. clustering techniques) according to their proximity on the manifold and leverage a further nonlinear dimension reduction to organize the structured data on the manifold. Finally, we are developing novel techniques that enable us to directly interpolate the hyper‐reduced data such that we can predict the solution of the complex, high‐ dimensional system without need to call the full expensive computational model. Given their adherence to the underlying structure of the solution of the physical system, it is expected that these approximate solutions will be sufficiently constrained so as to (approximately) adhere to physical principles.

97 MATHEMATICS AND COMPUTING

Universal Workflow Language and Software Enable Geometric Learning and FAIR Scientific Protocol Reporting

Written language and conventional data structures for representing scientific procedures suffer from low process detail, often fail to accurately represent protocols, and lack universality. New strategies for the handling of experimental data are needed to provide viable process information for both humans and machines. In this work, we present the universal workflow language (UWL) and interface (UWLi). UWL is a findable, accessible, interoperable, and reusable (FAIR)-compatible, graph-based data architecture that can capture arbitrary scientific procedures through workflow representation, and UWLi is an accompanying software package for building, manipulating, and interpreting UWL entries. The UWL format was found to be highly effective in identifying deficiencies in the reported process details of high-impact, peer-reviewed scientific journals, and in simulated scenarios, the graph format was shown to be more effective than conventional methods in predictively modeling the outcome of diverse scientific protocols. Implementation of UWL could enable more accurate scientific communication and more impactful process datasets.

14 SOLAR ENERGY

Development and validation of a software for simulating γ-γ coincidence emission and detection probabilities

Gamma-gamma coincidence spectrometers have the potential to significantly enhance detection sensitivity for ultra-trace radionuclide measurements. The implementation of these spectrometers, however, is limited by the complexity of acquisition hardware, data processing and quantification. This work reports development of a novel radionuclide quantification software for γ-γ coincidence measurements. For any radionuclide, the software parses the Evaluated Nuclear Structure Data File (ENSDF) database, recursively simulating all possible γ-γ coincidence signatures and their respective emission and detection probabilities. Implemented using Python programming language, the software employs several strategies to boost overall computational performance. Since coincidence-based spectrometers are of notable interest in monitoring compliance for the Comprehensive Nuclear-Test-Ban Treaty (CTBT), the software’s execution was tested for 84 CTBT-relevant radionuclides. To date, the software has been experimentally validated for 15 radionuclides using the Advanced Radionuclide Gamma spectrOmeter (ARGO) at Pacific Northwest National Laboratory, USA (PNNL). Notably, the software can be operated in convergence mode, whereby coincidence detection efficiency’s convergence behavior can help avoid unreliable radionuclide activity estimates. With growing number of coincidence spectrometers worldwide, this paper aims to assist the radiation metrology community in developing similar software for their system.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND

MBX V1.2: Accelerating Data-Driven Many-Body Molecular Dynamics Simulations

The MBX software provides an advanced platform for molecular dynamics simulations, leveraging state-of-the-art MB-pol and MB-nrg data-driven many-body potential energy functions. Developed over the past decade, these potential energy functions integrate physics-based and machine-learned many-body terms trained on electronic structure data calculated at the "gold standard" coupled-cluster level of theory. Recent advancements in MBX have focused on optimizing its performance, resulting in the release of MBX v1.2. While the inherently many-body nature of MB-pol and MB-nrg ensures high accuracy, it poses computational challenges. MBX v1.2 addresses these challenges with significant performance improvements, including enhanced parallelism that fully harnesses the power of modern multicore CPUs. In conclusion, these advancements enable simulations on nanosecond time scales for condensed-phase systems, significantly expanding the scope of high-accuracy, predictive simulations of complex molecular systems powered by data-driven many-body potential energy functions.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH

Cell-free bioelectrocatalytic platform for carbon dioxide reduction (Final Technical Report)

The University of Minnesota (UMN) EcoSynBio Team aimed to develop a cell-free, enzyme-based platform for the electro- biocatalytic conversion of CO2 into formate as a platform chemical for further upgrading. This type of bio electrocatalytic process delivers a clean product stream without the need for extensive separation from the electrolyte as in electrochemical synthesis and microbial processes. The reduction reaction is catalyzed by metal-dependent formate dehydrogenases (mFDHs) that are capable of efficient electrocatalytic CO2 reduction without the need of costly co-factors. The development of an efficient, scalable electrobiocatalytic process with high total turnover numbers and viable space time yields, however, was not without its challenges. The UM team has developed a protein-based scaffolding system that facilitates enzyme stabilization and attachment to electrodes along with electron transfer. Yet, although FDHs are highly promising enzymes for cell-free, electrobiochemical CO2 reduction, they are also greatly understudied and especially for applications in electrocatalysis. The UM team used the best described mFDH from Clostridium as its benchmark system and spent significant time and effort in attempting to replicate published data and finally, redesigned a recombinant production system for proper metal co-factor incorporation. The UM team has also identified a small set of new enzyme homologs from extreme microorganisms with superior stabilities that have yielded initial structural data for further engineering. In addition, a new bioelectrocatalytic reactor system has been developed that can be 3D printed and used for enzyme attachment to electrodes. In summary the project has generated critical basic information for the further development of this class of enzymes for the electricity driven reduction of CO2 into formate as platform chemical for upgrading into various other chemicals, including fuels.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH

Local lattice distortions drive the transition of BaIrO 3 into a ferromagnetic insulator state

Using variable temperature total and resonant x-ray scattering at the K edge of Ir species, we study the “bad metal” to insulator transition in BaIrO 3 , a canonical third transition series oxide. The usage of advanced experimental techniques and large-scale computer modeling helps us show that, contrary to the widely accepted view, charge disproportionation leading to the formation of Ir-trimers with a different number of 5d valence electrons already exists at room temperature. The charge disbalance between the trimers does not evolve much with decreasing temperature while local lattice distortions do, suggesting that the latter and not the former make a key contribution to the emergence of the enigmatic ferromagnetic insulator state of BaIrO 3 . The conclusion is supported by DFT calculations based on unmodified experimental structure data. Our work calls for a reconsideration of the role of lattice distortions in determining the electronic properties of third transition series oxides. It also charts a path to assessing these properties on a realistic and not assumed crystal structure basis.

36 MATERIALS SCIENCE

Observational Data for Next-Generation Climate Model Evaluation: Requirements, Considerations, and Best Practices

Climate model simulations are an important source of information about our planet’s climate system and also enable informed decision-making under different future scenarios. As a new archive of results from the next generation of climate models is anticipated to become available with the Coupled Model Intercomparison Project phase 7 (CMIP7), the need to develop efficient and robust methods to evaluate models is paramount. Observations are an integral part of model evaluation, providing a means to quantify and understand the degree to which climate models can faithfully reproduce Earth system processes. Such analysis is critical for constraining climate projections, identifying areas of focus for model development, and assisting analysts in deciphering the utility of models for specific applications. Observations of Earth system come from a diversity of sources, span different space–time domains, and are produced by different communities, and each dataset features different data structures and formats, metadata standards, and its own unique uncertainties. Uncertainties in an observational dataset may stem from gaps in temporal and spatial coverage, instrumentation errors, or assumptions in retrieval and processing methods. How then does one ensure that observational data are ready for use and utilized in the most appropriate way for robust, rapid, and routine climate model evaluation? The CMIP7 Model Benchmarking Task Team with input from the broader climate modeling, model evaluation, and observational data communities present a vision and considerations for best practices toward the optimal and appropriate use of observational data to support next-generation climate model evaluation.

Climate models

Cyote-attack Chain Estimator

Attack Chain Estimator (ACE) Application Overview The Attack Chain Estimator (ACE) Application is a sophisticated tool designed for the ingestion, classification, sequencing, and enrichment of cybersecurity threat reports. This application leverages advanced machine learning models and extensive historical data to provide comprehensive insights into cyber threats, specifically targeting Industrial Control Systems (ICS). Purpose The primary functions of the ACE Application include: Ingestion of Cybersecurity Threat Reporting: Capable of ingesting text-based threat reports in markdown or text file format. Supports ingestion of structured data from other sources in STIX/JSON format. Classification of Report’s Text-Based Events: Utilizes a DeBERTa classifier, specifically trained on cybersecurity data, to map the events to MITRE ATT&CK for ICS Tactics and Techniques. Classification is performed using multiple Jupyter notebooks and machine learning workflows hosted as FastAPI microservices: regex_data deberta_base_35_train_hft_classifier_mlflow.ipynb hft_regex_classifier_mlflow.ipynb param_train_hft_classifier_mlflow.ipynb regex_tactic_tech.ipynb Ordering of Tactics, Techniques, and Observable Events: Sequences the identified tactics, techniques, and events to form a coherent attack chain. Enrichment with Historical Attack Chain Details: Enhances the attack chain with details from historical attacks using a Markov model developed from CyOTE Precursor Analysis Report data. The Markov model is available as a FastAPI endpoint for seamless integration. Enrichment with Adversary Emulation Capabilities Data: Integrates adversary emulation capabilities data using MITRE Caldera for OT adversary abilities UUIDs. Export of Output Files: Provides options to export the enriched attack chain in JSON or CSV formats. Routing of Output to Other Applications: Facilitates routing of output to various platforms and applications, including: Threat Intelligence Platforms COREII Scout for Threat Intelligence Analysis COREII Modeling and Simulation for Adversary Emulation Technical Description The ACE Application is an advanced cybersecurity tool designed to provide detailed threat analysis and sequence generation. It is built on a robust architecture that integrates natural language processing, machine learning, and historical data modeling. Key Components: Data Ingestion Module: Handles the input of threat reports and data from various formats, ensuring flexibility in data sources. Classification Engine: Employs DeBERTa-based classifiers hosted as FastAPI microservices to analyze and classify threat report events in accordance with the MITRE ATT&CK framework for ICS. Sequence Generator: Orders the classified events into a logical attack chain, providing clear insight into the sequence of tactics and techniques used in the threat. Enrichment Engine: Integrates historical data and adversary emulation capabilities to enhance the attack chain with valuable context and additional details. The historical data enrichment is powered by a Markov model, which is available as a FastAPI endpoint. Export and Routing Module: Facilitates the export of the enriched attack chain in multiple formats and routes the output to designated applications for further analysis or emulation.

Paul, Tony [Idaho National Laboratory (INL), Idaho

Hydrazinoacetic acid is a biosynthetic precursor of the bacterially produced nitramine, N -nitroglycine

Nitramines [R(R′)N–NO 2 ; R,R′=H or alkyl] are valuable synthetic products, but knowledge of the biosynthetic processes that generate these compounds is limited. This work sought to elucidate the biosynthesis of a nitramine natural product, N-nitroglycine (NNG) by Streptomyces noursei . Stable isotope studies showed that S. noursei cells supplemented with L-(ε- 15 N)lysine, ( 15 N)glycine, or ( 13 C)hydrazinoacetic acid (HAA) incorporated 67%, 88%, and 67% of the isotope label into NNG, respectively, indicating that these compounds are biosynthetic precursors of NNG. Liquid chromatography coupled tandem mass spectrometry (LC-MS/MS) of 15 N-Lys-labeled NNG confirmed that the nitro nitrogen of NNG originates from Lys. Bioinformatics analysis of the S. noursei genome showed evidence for a biosynthetic gene cluster (BGC) that contained machinery for HAA biosynthesis ( nngKLM ), consistent with the results of the isotope labeling. In vitro reconstitution of the gene products produced HAA. The borders of this BGC were defined by cross-referencing the predicted BGC with previously published differential proteomics data. Furthermore, we show that azaserine is produced alongside NNG in S. noursei cultures, linking the two biosynthetic pathways via a proposed nitrosamine biosynthetic intermediate. Finally, the oxygen balance for NNG is −20.2% for the formation of carbon dioxide (CO 2 ), which is comparable to that of hexahydro-1,3,5- trinitro-1,3,5-triazene (common name: RDX; −21.6%). Crystal structure data of NNG indicate that the unit crystalizes as a pure material, not a hydrate, suggesting a favorable energetic crystallization phase. The combined results suggest a route that, with further development, could lead to sustainable production of energetic nitramines via synthetic biology or biocatalytic approaches.

biosynthesis

Hybrid Quantum–Classical Graph Transformers for Efficient Sentiment Analysis

Quantum Machine Learning (QML) offers a promising paradigm that leverages quantum computing principles to develop efficient and expressive models for learning from complex and structured data. Recent advances in natural language processing (NLP) and artificial intelligence (AI) have demonstrated capabilities in understanding, generating, and reasoning over linguistic and multimodal information. In this work, we present the Quantum Graph Transformer (QGT), a hybrid quantum–classical architecture that extends graph transformer capabilities through quantum self-attention. The QGT models variable-length sentences as token graphs, where both the embedding encoding and the self-attention mechanisms are implemented using parameterized quantum circuits (PQCs), enabling efficient contextual learning with significantly fewer trainable parameters. We train QGT using both fully connected and 𝑘 -nearest-neighbor graph structures and evaluate it on five benchmark sentiment-classification datasets. Experimental results show that QGT consistently achieves higher or comparable accuracy to existing quantum NLP models and outperforms a Classical Graph Transformer (CGT) baseline with identical architecture, achieving 29.4 × fewer parameters while requiring 3–5 × fewer samples to reach comparable performance. These findings highlight the potential of graph-based quantum models as scalable and data-efficient architectures for natural language understanding.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC

EV-ELM (Electric Vehicle Policies with the Energy Language Model) [SWR-25-156]

Electric Vehicle Policies with the Energy Language Model (EV-ELM) leverages previous work using Large Language Models (LLMs) to find, download, and parse policy information related to energy infrastructure. In this application, we use LLMs to find policy documents related to the permitting and installation of electric vehicle charging infrastructure. This software contains the code to find, download, and parse these documents, while a related data record in the Open Energy Data Initiative (OEDI) will include the resulting output dataset that can be used for downstream analysis. The EV-ELM repository contains code for the EV-ELM project, which focuses on retrieving and processing EV permitting processes using large language models. The project is composed of two pipelines: (1) a web scraping pipeline for discovering and downloading EV permitting documents, and (2) a document parsing and extraction pipeline that processes the downloaded files to produce structured data. The web scraping pipeline is designed to extract relevant information from various websites, while the document parsing pipeline processes and analyzes the extracted documents to derive meaningful insights. Both pipelines depend on the NLR elm repository, which provides essential tools and functionalities for handling and processing the data. The web scraping pipeline is a modified version of the ordinance_gpt example within the elm repository. It has been adapted to fit the specific requirements of the EV-ELM project, ensuring that it effectively captures and processes the necessary information related to EV permitting.

Olson, Reid [National Laboratory of the Rockies (N

The ECP SICM project: Managing complex memory hierarchies for exascale applications

The Exascale Computing Project (ECP)’s Simplified Interface to Complex Memories (SICM) effort focuses on developing universal interfaces for discovering, managing, and sharing data across complex memory hierarchies. These facilitate the exploitation of emerging memory technologies and support precise control over their various trade-offs such as high-bandwidth versus low-latency, persistent versus ephemeral, high-capacity versus low-capacity, and near-CPU versus near-GPU. SICM comprises three interrelated components: a low-level interface, a high-level interface, and a persistent-heap interface. The low-level SICM interface is intended for system and run-time developers as well as expert application developers who prefer full control of the memory objects used within their application. The high-level SICM interface builds upon the low-level interface, employing application-level profiling and analysis to optimize data management for complex memory hierarchies. The persistent-heap interface provides applications with a persistent memory allocator that can allocate custom C++ data structures in both block-storage and byte-addressable persistent memories.

97 MATHEMATICS AND COMPUTING

Vulcan-Forge: Architecture and Design of a Multi-Modal Forensic Analysis Plugin for CALDERA

Forge and VULCAN together describe an open-architecture cybersecurity analysis ecosystem that unifies forensic artifact processing, detection engineering, and vulnerability intelligence within integrated platforms. Forge operates as a plugin for MITRE CALDERA, ingesting diverse evidence formats—including EVTX, PCAP/PCAPNG, CSV, JSON, YAML, XML, binaries, and archives—to construct a unified artifact graph enriched with severity scoring, TLP classification, and audit trails. It provides subsystems for artifact parsing, streaming structured-data visualization, NetworkMiner-based packet inspection, PE/.NET binary analysis, and LLM-assisted triage and rule generation, with outputs validated against CCCS-YARA and pySigma schemas. VULCAN complements this by serving as a cybersecurity analyst platform that integrates a Neo4j knowledge graph, Qdrant vector retrieval, SSVC-based triage, and a local LLM to deliver CVE intelligence and forensic analysis through a multi-source ingest pipeline drawing from NVD, CISA KEV, EPSS, MITRE ATT&CK, and CAPEC. Together, they bridge structured threat intelligence with automated forensic analysis and detection workflows.

97 MATHEMATICS AND COMPUTING

Inputs to GCAM-USA: IM3 Phase 2 Experiments

Overview This dataset contains XML input files for the IM3 Phase 2 version of GCAM-USA. The files are organized into two categories: Scenario-specific inputs represent hydroclimate and socioeconomic effects on water availability, heating and cooling degree-hours, and agricultural productivity. They support eight IM3 canonical scenarios: rcp45cooler_ssp3 rcp45cooler_ssp5 rcp45hotter_ssp3 rcp45hotter_ssp5 rcp85cooler_ssp3 rcp85cooler_ssp5 rcp85hotter_ssp3 rcp85hotter_ssp5 Model-improvement inputs extend GCAM-USA v5.3 with updated representations of coal and nuclear power plant retirements, electricity trade among U.S. interconnections, offshore carbon storage costs, and groundwater depletion constraints. Data structure Scenario-specific inputs rcp45_runoff/ and rcp85_runoff/XML files describing water availability by HUC2 basin under the RCP 4.5 and RCP 8.5 scenarios. rcp45_hdcd/ and rcp85_hdcd/XML files containing monthly-day and monthly-night heating and cooling degree-hours at the U.S. state level for different RCP-SSP combinations. rcp45_agyields/ and rcp85_agyields/XML files describing changes in agricultural productivity at the intersection of GCAM regions and HUC2 water basins for different RCP-SSP combinations. rcp45_emissions_pathway/The emissions-constraint XML file used to represent the RCP 4.5 pathway. Model-improvement inputs core_retire/Updates coal-fired power plant retirement schedules based on New England ISO. GCAMUSA_IM3_elec_trade_interconnect.xmlRestricts electricity trade to occur within the ERCOT, WECC, and IE interconnections. nuclear_USA.xmlUpdates the retirement schedules of the Diablo Canyon and Palisades nuclear power plants. high_cost_offshore_carbon.xmlUpdates the assumed cost of offshore carbon storage. water_supply_constrained_gleeson_5pct.xmlReplaces WaterGAP historical groundwater-depletion estimates with data from the Gleeson dataset and limits groundwater extraction to 5% of the available groundwater in each Superwell grid cell. How to use the data This dataset is designed for use with the IM3 version of GCAM-USA. Download or clone GCAM-USA from the IM3 GCAM GitHub repository at https://github.com/IMMM-SFA/gcam-core and check out the gcam-usa-im3 branch. Place the downloaded folder im3scenarios in the gcam-core/input directory while preserving the provided folder structure.

Energy

Consistent performance of large language models in rare disease diagnosis across ten languages and 4917 cases

Background Large language models (LLMs) are increasingly used medicine for diverse applications including differential diagnostic support. The training data used to create LLMs such as the Generative Pretrained Transformer (GPT) predominantly consist of English-language texts, but LLMs could be used across the globe to support diagnostics if language barriers could be overcome. Initial pilot studies on the utility of LLMs for differential diagnosis in languages other than English have shown promise, but a large-scale assessment on the relative performance of these models in a variety of European and non-European languages on a comprehensive corpus of challenging rare-disease cases is lacking. Methods We created 4917 clinical vignettes using structured data captured with Human Phenotype Ontology (HPO) terms with the Global Alliance for Genomics and Health (GA4GH) Phenopacket Schema. These clinical vignettes span a total of 360 distinct genetic diseases with 2525 associated phenotypic features. We used translations of the Human Phenotype Ontology together with language-specific templates to generate prompts in English, Chinese, Czech, Dutch, French, German, Italian, Japanese, Spanish, and Turkish. We applied GPT-4o, version gpt-4o-2024-08-06, and the medically fine-tuned Meditron3-70B to the task of delivering a ranked differential diagnosis using a zero-shot prompt. An ontology-based approach with the Mondo disease ontology was used to map synonyms and to map disease subtypes to clinical diagnoses in order to automate evaluation of LLM responses. Findings For English, GPT-4o placed the correct diagnosis at the first rank 19.9% and within the top-3 ranks 27.0% of the time. In comparison, for the nine non-English languages tested here the correct diagnosis was placed at rank 1 between 16.9% and 20.6%, within top-3 between 25.4% and 28.6% of cases. The Meditron3 model placed the correct diagnosis within the first 3 ranks for 20.9% of cases in English and between 19.9% and 24.0% for the other nine languages. Interpretation The differential diagnostic performance of LLMs across a comprehensive corpus of rare-disease cases was largely consistent across the ten languages tested. This suggests that the utility of LLMs in clinical settings may extend to non-English clinical settings.

Artificial intelligence

Decoding substrate specificity determining factors in glycosyltransferase-B enzymes – insights from machine learning models

Substrate specificity is an essential characteristic of any enzyme's function and an understanding of the factors that determine this specificity is crucial for enzyme engineering. Unlike the structure of an enzyme which is directly impacted by its sequence, substrate specificity as an enzyme attribute involves a rather indirect relationship with sequence as it also depends on structural aspects that dictate substrate accessibility and active site dynamics. In this study, we explore the performance of classifier-based machine learning models trained on curated sequence and structural data for a class of glycosyltransferases (GTs), namely GT-Bs, to understand their substrate specificity determining factors. GTs enable the transfer of sugar moieties to other biomolecules such as oligosaccharides or proteins and are found in all kingdoms of life. In plants, GTs participate in the biosynthesis of plant cell wall biopolymers (e.g.: hemicelluloses and pectins) and are an integral part of the enzymatic machinery that enables the storage of carbon and energy as plant biomass. To elucidate the substrate specificity of uncharacterized GT-Bs, we constructed multi-label machine learning models (Support Vector Classifier, K-Nearest Neighbors, Gaussian Naïve-Bayes, Random Forest) that incorporate both sequence and structural features. These models achieve good predictive accuracies on test datasets. However, despite our use of structural information, we highlight that there is further scope for improvement in training these models to draw interpretable relationships between sequence, structure and substrate specificity determining motifs in GT-Bs.

97 MATHEMATICS AND COMPUTING