Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “graph databases”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

90 records · Page 5

Long plasma duration operation analyses with an international multi-machine (tokamaks and stellarators) database

Combined high-fusion performance and long-pulse operation is one of the key integration challenges for fusion energy development in magnetic devices. Addressing these challenges requires an integrated vision of physics and engineering aspects with the purpose of simultaneously increasing time duration and fusion performance. Significant progress has been made in tokamaks and stellarators, including very recent achievement in duration and/or performance. This progress is reviewed by analyzing the experimental data (109 plasma pulses with a total of 3200 data points, i.e. on average 29 data per pulse) provided by ten tokamaks (in alphabetical order: Axially Symmetric Divertor Experiment Upgrade, DIII-D, Experimental Advanced Superconducting Tokamak, Joint European Torus, JT-60 Upgrade, Korea Superconducting Tokamak Advanced Research, tokamak à configuration variable, Tokamak Fusion Test Reactor, Tore Supra, W Environment in Steady-State Tokamak) and two stellarators (Large Helical Device and Wendelstein 7-X) expanding the pioneering work of Kikuchi (Kikuchi M. and Azumi M. 2015 Frontiers in Fusion Research II: Introduction to Modern Tokamak Physics (Springer)). Data have been gathered up to January 2022 and coordination has been provided by the recently created International Energy Agency-International Atomic Energy Agency international Coordination on International Challenges on Long duration OPeration group. By exploiting the multi-machine international database, recent progress in terms of injected energies (e.g. 1730 MJ in L-mode, 425 MJ in H-mode), durations (1056 s in L-mode, 101 s in H-mode), injected powers, and sustained performance will be reviewed. Progress has been made to sustain long-pulse operation in tokamaks and stellarators with superconducting coils, actively cooled components, and/or with metallic walls. The graph of the fusion triple products as a function of duration shows a dramatic reduction of, at least two orders of magnitude when increasing the plasma duration from less than 1 s to 100 s. Indeed, long-pulse operation is usually reached in dominant electron-heating modes at reduced density (current drive optimization) but with low ion temperatures ranging from 1 to 3 keV for discharges above 100 s. Difficulties in extending the duration may arise from coupling high heating powers over long durations and the evolving plasma-wall interaction towards an unstable operational domain. Possible causes limiting the duration and critical issues to be addressed prior to ITER operation and DEMO design are reported and analyzed.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

An automated workflow that generates atom mappings for large‐scale metabolic models and its application to Arabidopsis thaliana

SUMMARY Quantification of reaction fluxes of metabolic networks can help us understand how the integration of different metabolic pathways determines cellular functions. Yet, intracellular fluxes cannot be measured directly but are estimated with metabolic flux analysis (MFA), which relies on the patterns of isotope labeling of metabolites in the network. The application of MFA also requires a stoichiometric model with atom mappings that are currently not available for the majority of large‐scale metabolic network models, particularly of plants. While automated approaches such as the Reaction Decoder Toolkit (RDT) can produce atom mappings for individual reactions, tracing the flow of individual atoms of the entire reactions across a metabolic model remains challenging. Here we establish an automated workflow to obtain reliable atom mappings for large‐scale metabolic models by refining the outcome of RDT, and apply the workflow to metabolic models of Arabidopsis thaliana . We demonstrate the accuracy of RDT through a comparative analysis with atom mappings from a large database of biochemical reactions, MetaCyc. We further show the utility of our automated workflow by simulating 15 N isotope enrichment and identifying nitrogen (N)‐containing metabolites which show enrichment patterns that are informative for flux estimation in future 15 N‐MFA studies of A. thaliana . The automated workflow established in this study can be readily expanded to other species for which metabolic models have been established and the resulting atom mappings will facilitate MFA and graph‐theoretic structural analyses with large‐scale metabolic networks.

59 BASIC BIOLOGICAL SCIENCES↗

Defect Diffusion Graph Neural Networks for Materials Discovery in High-Temperature Energy Applications

Here, the migration of crystallographic defects dictates material properties and performance for a plethora of technological applications. Density functional theory (DFT)-based nudged elastic band (NEB) calculations are a powerful computational technique for predicting defect migration activation energy barriers, yet they become prohibitively expensive for high-throughput screening of defect diffusivities. Without introducing hand-crafted (i.e., chemistry- or structure-specific) descriptors, we propose a generalized deep learning approach to train surrogate models for NEB energies of vacancy migration by hybridizing graph neural networks with transformer encoders and simply using pristine host structures as input. With sufficient training data, computationally efficient and simultaneous inference of vacancy defect thermodynamics and migration activation energies can be obtained to compute temperature-dependent vacancy diffusivities and to down-select candidates for more thorough DFT analysis or experiments. Thus, as we specifically demonstrate for potential water-splitting materials, candidates with desired defect thermodynamics, kinetics, and host stability properties can be more rapidly targeted from open-source databases of experimentally validated or hypothetical materials.

14 SOLAR ENERGY↗

Retaining Systems Engineering Model Meaning Through Transformation: Demo 2

Digital engineering strategies typically assume that digital engineering models interoperate seamlessly across the multiple different engineering modeling software applications involved, such as model- based systems engineering (MBSE), mechanical computer-aided design (MCAD), electrical computer-aided design (ECAD), and other engineering modeling applications. The presumption is that the data schema in these modeling software applications are structured in the familiar flat- tabular schema like any other software application. Engineering domain-specific applications (e.g., systems, mechanical, electrical, simulation) are typically designed to solve domain-specific problems, necessarily excluding explicit representations of non-domain information to help the engineer focus on the domain problems (system definition, design, simulation). Such exclusions become problematic in inter-domain information exchange. The obvious assumptions of one domain might not be so obvious to experts in another domain. Ambiguity in domain-specific language can erode the ability to enable different domain modeling applications to interoperate, unless the underlying language is understood and used as the basis for translation from one application to another. The engineering modeling software application industry has struggled for decades to enable these applications to interoperate. Industry standards have been developed, but they have not unified the industry. Why is this? The authors assert that the industry has relied on traditional database integration methods. The basic issue prohibiting successful application integration then is that traditional database-driven integration does not consider the distinct languages of each domain. An engineering models meaning is expressed through the underlying language of that engineering domain. In essence, traditional integration methods do not retain the semantic context (meaning) of the model. The basis of this research stems from the widely held assumption that systems engineering models are (or can be) structured according to the underlying semantic ontology of the model. This assumption can be imagined from two thoughts. 1) Digital systems engineering models are often represented using graph theory (the graph of a complex systems model can contain millions of nodes and edges). When examining the nodes one at a time and following the outbound edges of each node one by one, one can end up with rudimentary statements about the model (i.e., node A relates to node B), as in a semantic graph. 2) Likewise, from the study of natural languages, a sentence can be structured into unambiguous triples of subject-predicate-object within formal and highly expressive semantic ontologies. The rudimentary statements about a systems model discerned with graph theory closely mimic the triples used in the ontologies that try to structure natural languages. In other words, a systems models semantic graph can be (or is) structured into an ontology. Additionally, it is well established in industry that through natural language processing (NLP), which provides the means to create language structures, that computers can interpret ontological graphs. Therefore, the authors hypothesized that if the integrity of the underlying semantic structure of a systems model is retained, the contextual meaning of the model is retained. By structuring system models into the triples of the underlying ontology during the transformation from one MBSE application to another, the authors have provided a proof of the concept that the meaning of a system model can be retained during transformation. The authors assert that this is the missing ingredient in effective systems model-to-model interoperability. ACKNOWLEDGEMENTS The authors would like to thank the FY19 Model Interoperability team members who provided a solid foundation for the FY20 team to leverage: John McCloud, for the work he did to guide us toward the right use of technology that will appropriately discover and manipulate ontologies. Carlos Tafoya, for the work he did to develop an application programming interface (API)/Adapter that would export ontology-based data from GENESYS. Peter Chandler, for the work he did to architect our overall integration solution, with an eye toward the future that would influence a large-scale federated production-level systems engineering digital model ecosystem.

42 ENGINEERING↗

A framework to evaluate machine learning crystal stability predictions

The rapid adoption of machine learning in various scientific domains calls for the development of best practices and community agreed-upon benchmarking tasks and metrics. We present Matbench Discovery as an example evaluation framework for machine learning energy models, here applied as pre-filters to first-principles computed data in a high-throughput search for stable inorganic crystals. We address the disconnect between (1) thermodynamic stability and formation energy and (2) retrospective and prospective benchmarking for materials discovery. Alongside this paper, we publish a Python package to aid with future model submissions and a growing online leaderboard with adaptive user-defined weighting of various performance metrics allowing researchers to prioritize the metrics they value most. To answer the question of which machine learning methodology performs best at materials discovery, our initial release includes random forests, graph neural networks, one-shot predictors, iterative Bayesian optimizers and universal interatomic potentials. We highlight a misalignment between commonly used regression metrics and more task-relevant classification metrics for materials discovery. Accurate regressors are susceptible to unexpectedly high false-positive rates if those accurate predictions lie close to the decision boundary at 0 eV per atom above the convex hull. The benchmark results demonstrate that universal interatomic potentials have advanced sufficiently to effectively and cheaply pre-screen thermodynamic stable hypothetical materials in future expansions of high-throughput materials databases.

Riebesell, Janosh↗

Database Performance Monitoring for DUNE

This report presents the research, design, and implementation of improved PostgreSQL monitoring for DUNE Rucio database services using Checkmk. The project began with a request to improve dashboard visibility for database performance metrics, including connection usage, configured connection limits, lock activity, wait behavior, storage trends, query performance, and saturation alerts. The initial implementation focused on the dune_rucio_prod database on the rucio_prod PostgreSQL instance because connection saturation and lock contention are direct reliability risks for database-backed services. Existing Checkmk PostgreSQL monitoring was investigated, and several gaps were identified. Built-in connection monitoring did not clearly separate active, idle, idle-in-transaction, total, and usage-percent metrics, while the built-in lock monitoring simplified PostgreSQL lock modes into shared and exclusive categories. To address these gaps, two DSG-specific Checkmk local checks were created: one for connection-state monitoring and one for lock-state monitoring. These checks supplement the built-in PostgreSQL checks and provide additional performance data for dashboard graphs, service states, and alerts.

Bowers, Elliot [Cabrillo Coll.]↗

Rapid Computational Identification of Therapeutic Targets for Pathogens

Biological threats continue to persist and evolve as an important challenge to national security. There are multiple ways in which novel viral pathogens could emerge to pose a serious threat to human health. This project developed a pathogen target identification tool that can rapidly respond to a novel or emerging viral biological threat. A set of computational tools were developed that provide detailed information on the newly sequenced genes, their protein products and the drug target sites for the proteins that are best suited for biological countermeasure development. Three key innovations were developed in the project. 1) Development of a new extensive database of protein pocket structures with structure-based search algorithms to rapidly link novel protein targets with the complete collection of previously experimentally solved protein structures. 2) A novel clustering pipeline was introduced to group matching structures and associated small-molecule binding ligands into a consensus protein pocket with the associated small-molecule chemotypes predicted to fit in the pocket site. The matching experimentally solved structures were used to inform the value of different target sites. 3) Where there are viral protein targets with pockets structurally matched to similar human proteins, a biological knowledge graph, which links molecular interactions with human disease, was used to further assess the potential negative impact of a viral protein target with similarities to human proteins that could have important off target side effects. In total, the project produced a new resource for rapid and detailed assessment of promising targets for countermeasures, reflecting the ongoing wet lab, clinical, and computational data being collected. These capabilities will improve the ability to respond to a biological threat in multiple domains.

59 BASIC BIOLOGICAL SCIENCES↗

Database Performance Monitoring for DUNE

This project improves Checkmk monitoring for DUNE Rucio PostgreSQL database services by adding clearer dashboard visibility for connection and lock behavior. The work began with a request to monitor database performance metrics such as connection usage, configured limits, lock activity, wait behavior, query performance, storage trends, and saturation alerts. Existing Checkmk PostgreSQL checks were reviewed, and gaps were identified in how connection states and lock modes were displayed. To address these gaps, two DSG-specific local checks were added for dune_rucio_prod: one for connection-state monitoring and one for lock-state monitoring. These checks report active, idle, idle-in-transaction, total, usage-percent, lock-mode, waiting-lock, and wait-age metrics. The added metrics supplement built-in Checkmk monitoring and provide DUNE application developers with clearer service states, history graphs, dashboard widgets, and alerts.

Bowers, Elliot [Cabrillo Coll.]↗

cuTS: Scaling Subgraph Isomorphism on Distributed Multi-GPUSystems Using Trie Based Data Structure

Subgraph isomorphism is a pattern-matching algorithm widely used in many domains such as chem-informatics, bioinformatics, databases, and social network analysis. It is computationally expensive and is a proven NP-hard problem. The massive parallelism offered by the GPU hardware is well suited for solving the subgraph isomorphism. However, current GPU implementations are far from the achievable performance. Moreover, the enormous memory requirement of current approaches limits the problem size that can be handled. This work analyzes the fundamental challenges associated with processing the subgraph isomorphism on GPUs and develops an efficient GPU hardware-aware implementation. We also develop a new GPU-friendly trie-based data structure to drastically reduce the intermediate storage space requirement. Hence, our approach runs larger benchmarks than the competitors. We also develop the first distributed sub-graph isomorphism algorithm for GPUs. Our experimental evaluation section demonstrates the efficacy of our approach by comparing the execution time and number of cases that we can handle against the state-of-the-art GPU implementations.

Xiang, Lizhi↗

ChemoGraph: Interactive Visual Exploration of the Chemical Space

Exploratory analysis of the chemical space is an important task in the field of cheminformatics. For example, in drug discovery research, chemists investigate sets of thousands of chemical compounds in order to identify novel yet structurally similar synthetic compounds to replace natural products. Manually exploring the chemical space inhabited by all possible molecules and chemical compounds is impractical, and therefore presents a challenge. To fill this gap, we present ChemoGraph, a novel visual analytics technique for interactively exploring related chemicals. In ChemoGraph, we formalize a chemical space as a hypergraph and apply novel machine learning models to compute related chemical compounds. It uses a database to find related compounds from a known space and a machine learning model to generate new ones, which helps enlarge the known space. Moreover, ChemoGraph highlights interactive features that support users in viewing, comparing, and organizing computationally identified related chemicals. With a drug discovery usage scenario and initial expert feedback from a case study, we demonstrate the usefulness of ChemoGraph.

chemical space exploration↗

Machine learning in materials research: Developments over the last decade and challenges for the future

The number of studies that apply machine learning (ML) to materials science has been growing at a rate of approximately 1.67 times per year over the past decade. In this review, I examine this growth in various contexts. First, I present an analysis of the most commonly used tools (software, databases, materials science methods, and ML methods) used within papers that apply ML to materials science. The analysis demonstrates that despite the growth of deep learning techniques, the use of classical machine learning is still dominant as a whole. It also demonstrates how new research can effectively build upon past research, particular in the domain of ML models trained on density functional theory calculation data. Next, I present the progression of best scores as a function of time on the matbench materials science benchmark for formation enthalpy prediction. In particular, a dramatic improvement of 7 times reduction in error is obtained when progressing from feature-based methods that use conventional ML (random forest, support vector regression, etc.) to the use of graph neural network techniques. Finally, I provide views on future challenges and opportunities, focusing on data size and complexity, extrapolation, interpretation, access, and relevance.

36 MATERIALS SCIENCE↗

Can a deep-learning model make fast predictions of vacancy formation in diverse materials?

The presence of point defects, such as vacancies, plays an important role in materials design. Here, we explore the extrapolative power of a graph neural network (GNN) to predict vacancy formation energies. We show that a model trained only on perfect materials can also be used to predict vacancy formation energies (E vac ) of defect structures without the need for additional training data. Such GNN-based predictions are considerably faster than density functional theory (DFT) calculations and show potential as a quick pre-screening tool for defect systems. To test this strategy, we developed a DFT dataset of 530 E vac consisting of 3D elemental solids, alloys, oxides, semiconductors, and 2D monolayer materials. We analyzed and discussed the applicability of such direct and fast predictions. We applied the model to predict 192 494 E vac for 55 723 materials in the JARVIS-DFT database. Our work demonstrates how a GNN-model performs on unseen data.

2D materials↗

Optimizing FPGA-based Accelerator Design for Large-Scale Molecular Similarity Search (Special Session Paper)

Molecular similarity search has been widely used in drug discovery to rapidly identify structurally similar compounds from large molecular databases. With the increasing size of chemical libraries, there is growing interest in the efficient ac- celeration of large-scale similarity search. Existing works mainly focus on CPU and GPU to accelerate the computation of Tatimoto coefficient in measuring the pairwise similarity between different molecular fingerprints. In this paper, we propose and optimize an FPGA-based accelerator design on exhaustive and approximate search algorithms. On exhaustive search using BitBound & fold- ing, we analyze the similarity cutoff and folding level relationship with search speedup and accuracy, and propose a scalable on- the-fly query engine on FPGAs to reduce the resource utilization and pipeline interval. We achieve a 450 million compounds-per- second processing throughput for a single query engine. On approximate search using hierarchical navigable small world (HNSW), a popular algorithm with high recall and query speed, we propose an FPGA-based graph traversal engine to utilize high throughput register array based priority queue and fine- grained distance calculation engine to increase the processing capability. Experimental results show that the proposed FPGA- based HNSW implementation achieves a 35× speedup than existing works on CPU. To the best of our knowledge, our FPGA- based implementation is the first attempt to accelerate molecular similarity search on FPGA and has the highest performance among existing approaches.

Peng, Hongwu↗

Optimizing Management of Persistent Data Structures in High-Performance Analytics

Large-scale data analytics workflows ingest massive input data into various data structures, including graphs and key-value datastores. These data structures undergo multiple transformations and computations and are typically reused in incremental and iterative analytics workflows. Persisting in-memory views of these data structures enables reusing them beyond the scope of a single program run while avoiding repetitive raw data ingestion overheads. Memory-mapped I/O enables persisting in-memory data structures without data serialization and deserialization overheads. However, memory-mapped I/O lacks the key feature of persisting consistent snapshots of these data structures for incremental ingestion and processing. The obstacles to efficient virtual memory snapshots using memory-mapped I/O include background writebacks outside the application’s control, and the significantly high storage footprint of such snapshots. To address these limitations, we present Privateer, a memory and storage management tool that enables storage-efficient virtual memory snapshotting while also optimizing snapshot I/O performance. Here, we integrated Privateer into Metall, a state-of-the-art persistent memory allocator for C++, and the Lightning Memory-Mapped Database (LMDB), a widely-used key-value datastore in data analytics and machine learning. Privateer optimized application performance by 1.22× when storing data structure snapshots to node-local storage, and up to 16.7× when storing snapshots to a parallel file system. Privateer also optimizes storage efficiency of incremental data structure snapshots by up to 11× using data deduplication and compression.

Computer science↗

Active learning of ternary alloy structures and energies

Abstract Machine learning models with uncertainty quantification have recently emerged as attractive tools to accelerate the navigation of catalyst design spaces in a data-efficient manner. Here, we combine active learning with a dropout graph convolutional network (dGCN) as a surrogate model to explore the complex materials space of high-entropy alloys (HEAs). We train the dGCN on the formation energies of disordered binary alloy structures in the Pd-Pt-Sn ternary alloy system and improve predictions on ternary structures by performing reduced optimization of the formation free energy, the target property that determines HEA stability, over ensembles of ternary structures constructed based on two coordinate systems: (a) a physics-informed ternary composition space, and (b) data-driven coordinates discovered by the Diffusion Maps manifold learning scheme. Both reduced optimization techniques improve predictions of the formation free energy in the ternary alloy space with a significantly reduced number of DFT calculations compared to a high-fidelity model. The physics-based scheme converges to the target property in a manner akin to a depth-first strategy, whereas the data-driven scheme appears more akin to a breadth-first approach. Both sampling schemes, coupled with our acquisition function, successfully exploit a database of DFT-calculated binary alloy structures and energies, augmented with a relatively small number of ternary alloy calculations, to identify stable ternary HEA compositions and structures. This generalized framework can be extended to incorporate more complex bulk and surface structural motifs, and the results demonstrate that significant dimensionality reduction is possible in thermodynamic sampling problems when suitable active learning schemes are employed.

Chemistry↗

GROWdb US River Systems - Samples

GROW Overview We developed the Genome Resolved Open Watersheds database (GROWdb), which aims to increase genomic sampling and understanding of global river microbiomes. An emphasis of GROWdb is to create a publicly available and ever-expanding microbial genome database that is focused on rivers while being interoperable with databases from other ecosystems. GROWdb is based on a network-of-networks approach to move beyond a small collection of well-studied rivers, towards a spatially distributed, global network of systematic observations. GROWdb represents the first microbial, river-focused resource parsed at various scales from genes to MAGs to community level including expression and potential based measurements that will be of interest to microbiologists, ecologists, geochemists, hydrologists, and modelers. Dataset Acknowledgement GROWdb contains data from various research campaigns, please acknowledge the following data generators, as appropriate: WHONDRS derived genomes or samples - include this statement in your acknowledgements: “This study used data from the Worldwide Hydrobiogeochemistry Observation Network for Dynamic River Systems (WHONDRS) under the River Corridor Science Focus Area (SFA) at the Pacific Northwest National Laboratory (PNNL) that was generated at the U.S. Department of Energy (DOE) Joint Genome Institute User Facility. PNNL is operated by Battelle Memorial Institute for the U.S. DOE under Contract No. DE-AC05-76RL01830. The SFA is supported by the U.S. DOE, Office of Biological and Environmental Research (BER), Environmental System Science (ESS) Program.” Total Samples loaded onto this Narrative: 178 Note: Not all GROW samples may be loaded into KBase Data Availability The data underlying GROWdb are accessible across various platforms to ensure all levels of data structure are widely available. First, all reads and MAGs are publicly hosted on National Center for Biotechnology (NCBI) under Bioproject PRJNA946291. Second, all data related data presented here including MAG annotations, extended data tables, phylogenetic tree files, antibiotic resistance gene database files, and MAG abundance tables are available in Zenodo (link). Beyond the flat database files listed above, our aim for GROWdb was to maximize data use by making the data available in searchable and interactive platforms including the National Microbiome Data Collaborative (NMDC) data portal, the Department of Energy’s Systems Biology Knowledgebase (KBase), and a GROW specific user interface released here, GROWdb Explorer. Each platform provides different ways to interact with GROWdb: NMDC GROWdb formed a pilot project for the NMDC. Specifically, individual GROWdb datasets (metagenomes, metatranscriptomes, etc) are easily accessible and searchable through the NMDC data portal, where they are systematically connected to each other and to a rich suite of sample information and standard analysis results, following Findable, Accessible, Interoperable, and Reusable (FAIR) data practices. KBase GROWdb is publicly available within KBase, including samples (this Narrative), MAGs, and corresponding genome scale metabolic models. Access within KBase allows for immediate access and reuse of data, including comparison to private data using KBase’s 500+ analysis tools. Other linked narratives in KBase: GROW Metagenome Assembled Genomes (MAGs) GROW Metabolic Models GROWdb Explorer GROWdb data is also explorable through a graphical user interface built through the Colorado State University Geospatial Centroid (https://geocentroid.shinyapps.io/GROWdatabase/), allowing users to search and graph microbial and spatial data simultaneously. In summary, this microbial genome resource represents the first publicly available genome collection from rivers and offers data that can be leveraged across microbiome studies. GROWdb is an expanding repository to incorporate and unify global river multi-omic data for the future.

59 BASIC BIOLOGICAL SCIENCES↗

Verification Testing of OLI Systems Mixed Solvent Electrolyte Model for the Na-K-Mg-Ca-H-Cl-SO 4 -OH-HCO 3 -CO 3 -CO 2 -H 2 ) System to High Ionic Strength at 25°C

This technical report summarizes model verification results and summary statistics for 41 evaporite mineral solubility cases evaluated by Savannah River National Laboratory using OLI Systems’ aqueous electrolyte thermodynamic modeling software. The 41 verification cases containing a total of 60 solubility curves comprise mineral solubility data from low to high ionic strength at 25°C for the eight-component system Na-K-Mg-Ca-H-Cl-SO 4 -OH-HCO 3 -CO 3 -CO 2 -H 2 O as reported by Harvie et al. (1984). Thermodynamic calculations were executed using OLI Systems’ Stream Analyzer computation module within the OLI Studio software platform (Ver. 11.0, Rev. 11.0.1.9). The Mixed Solvent Electrolyte (MSE) thermodynamic framework was chosen for this investigation because of its superiority in modeling high ionic-strength inorganic salt solutions and actinide redox chemistry and solubility, both of which are relevant to the geological repository conditions at the Waste Isolation Pilot Plant in Carlsbad, New Mexico. Mineral solubility data in various inorganic salt solutions were digitized and extracted from figures generated by Harvie et al. (1984). For each of the 60 solubility curves, a case-specific chemistry model and input file were generated in OLI Studio using OLI Stream Analyzer and the MSE (H 3 O + ion) public databank provided by OLI Systems. Model simulation results were exported to Microsoft Excel to calculate summary statistics and to generate graphs comparing the OLI model predictions to the solubility data. Summary statistics include residuals (model – data) and concordance (accuracy × precision, where precision is indicated by the Pearson correlation coefficient and accuracy accounts for bias and scale differential). Private databanks were not developed, and activity coefficient model regressions were not performed to improve OLI model fits to the data. Of the 41 model verification plots, 83% have a mean of the percent residuals less than or equal to 25%. Similarly, 75% display a concordance greater than or equal to 0.75. Only seven of the 41 verification plots fail to show good agreement between the model and data. Of these seven, three are relevant to the WIPP repository because they involve the Mg-OH-Cl-SO 4 -CO 3 aqueous system. The remaining four address salt solubilities at the pH extremes (strong acid and strong base). It should be noted that in two of the three Mg-OH-Cl-SO 4 -CO 3 system cases, the regressed Harvie et al. (1984) solubility curve also deviated from the data. Lack of agreement between the OLI model-predicted solubility curves and the data is attributable to one or more of the following: specific solid species are not included in the OLI MSE databank; there is significant variation among the different solubility datasets chosen by Harvie et al. (1984); the OLI MSE model’s thermodynamic parameters were determined using different solubility datasets; and the activity coefficient parameters for certain relevant ion-ion and ion-molecule pairs have not been optimized via data regression. Two recommendations for future work are to (1) evaluate solubility data for the Mg-OH-Cl-SO 4 -CO 3 system at high ionic strength and, if necessary, develop a private OLI MSE database that includes missing species and, where necessary, regressed standard state properties and interaction parameters; (2) perform similar verification testing of the OLI model for actinide solubility data.

12 MANAGEMENT OF RADIOACTIVE AND NON-RADIOACTIVE W↗

Ca X ML: Chemistry‐informed machine learning explains mutual changes between protein conformations and calcium ions in calcium‐binding proteins using structural and topological features

Proteins' flexibility is a feature in communicating changes in cell signaling instigated by binding with secondary messengers, such as calcium ions, associated with the coordination of muscle contraction, neurotransmitter release, and gene expression. When binding with the disordered parts of a protein, calcium ions must balance their charge states with the shape of calcium-binding proteins and their versatile pool of partners depending on the circumstances they transmit. Accurately determining the ionic charges of those ions is essential for understanding their role in such processes. However, it is unclear whether the limited experimental data available can be effectively used to train models to accurately predict the charges of calcium-binding protein variants. Here, we developed a chemistry-informed, machine-learning algorithm that implements a game theoretic approach to explain the output of a machine-learning model without the prerequisite of an excessively large database for high-performance prediction of atomic charges. We used the ab initio electronic structure data representing calcium ions and the structures of the disordered segments of calcium-binding peptides with surrounding water molecules to train several explainable models. Network theory was used to extract the topological features of atomic interactions in the structurally complex data dictated by the coordination chemistry of a calcium ion, a potent indicator of its charge state in protein. Our design created a computational tool of Ca X ML, which provided a framework of explainable machine learning model to annotate ionic charges of calcium ions in calcium-binding proteins in response to the chemical changes in an environment. Our framework will provide new insights into protein design for engineering functionality based on the limited size of scientific data in a genome space.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗