Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “graph databases”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

From Data to Knowledge: A Graph-Based Reliability Approach to Assess System Health

With the goal of maximizing plant reliability and availability, complex systems such as nuclear power plants continuously monitor and record the performance and the health status of many components, assets, and systems. Such data may take the form of online monitoring data, condition reports, and maintenance reports and it carries the potential to provide system engineers with insights into anomalous behaviors or degradation trends as well as the possible causes behind them and to predict their direct consequences. The analysis of such data poses however few challenges. While some of these challenges are technical in nature (i.e., data are often distributed over several physical servers or databases), others are conceptual in nature (i.e., data elements come in different formats, numeric or textual), and measured values have different scales (e.g., vibration spectra and oil temperature). This paper directly tackles these challenges, and it focuses on the integration of all these data elements in order to assist plant system engineers in analyzing component, assets, and systems performances and optimize maintenance activities. This is performed by 1) extracting knowledge from textual data via technical language processing methods, and 2) quantifying system, asset, and component health from numeric condition-based data. We rely on model-based system engineering (MBSE) models of systems and assets to identify their architecture and functional (i.e., cause and effect) relations. Numeric and textual data elements are then associated with an MBSE graph element, based on their nature. This bonding of MBSE models and data elements constitutes a first-of-its-kind knowledge graph of a nuclear power plants system, with data elements being organized in a structured manner that enables system engineers to identify cause-effect trends in data elements and carry out appropriate actions in response.

97 MATHEMATICS AND COMPUTING↗

RHSEG and Subdue: Background and Preliminary Approach for Combining these Technologies for Enhanced Image Data Analysis, Mining and Knowledge Discovery

Under a project recently selected for funding by NASA's Science Mission Directorate under the Applied Information Systems Research (AISR) program, Tilton and Cook will design and implement the integration of the Subdue graph based knowledge discovery system, developed at the University of Texas Arlington and Washington State University, with image segmentation hierarchies produced by the RHSEG software, developed at NASA GSFC, and perform pilot demonstration studies of data analysis, mining and knowledge discovery on NASA data. Subdue represents a method for discovering substructures in structural databases. Subdue is devised for general-purpose automated discovery, concept learning, and hierarchical clustering, with or without domain knowledge. Subdue was developed by Cook and her colleague, Lawrence B. Holder. For Subdue to be effective in finding patterns in imagery data, the data must be abstracted up from the pixel domain. An appropriate abstraction of imagery data is a segmentation hierarchy: a set of several segmentations of the same image at different levels of detail in which the segmentations at coarser levels of detail can be produced from simple merges of regions at finer levels of detail. The RHSEG program, a recursive approximation to a Hierarchical Segmentation approach (HSEG), can produce segmentation hierarchies quickly and effectively for a wide variety of images. RHSEG and HSEG were developed at NASA GSFC by Tilton. In this presentation we provide background on the RHSEG and Subdue technologies and present a preliminary analysis on how RHSEG and Subdue may be combined to enhance image data analysis, mining and knowledge discovery.

Tilton, James C.↗

Orbital Debris Ontology, Terminology, and Knowledge Modeling

The looming threat orbital debris poses to assets in orbit demands solutions. As the orbital population grows, so does this hazard, but so does the sea of data. The problem is also an opportunity for interdisciplinary innovation and cooperation. This paper focuses on the data and information management aspect of developing solutions for a sustainable and safe orbital space environment. The corresponding author’s in-progress work to develop an orbital debris domain ontology is summarized in order to discuss knowledge modeling for this domain. Methodological approaches of this effort can also contribute to standards efforts and address terminological and policy questions. Leveraging the growing volumes of orbital debris and space situational awareness (SSA) data will create a more complete picture of the orbital space environment. Part of the solution will be: consistent and correct data interpretation, sharing orbital debris and SSA data in one form or another, terminology development & harmonization, and knowledge or domain modeling. To facilitate this, [Rovetto, 2015/16] discussed ontology development for the orbital debris domain. This paper lists concepts from that paper, and subsequently developed concepts [2-9]. Ontology engineering is an interdisciplinary field related to knowledge representation and reasoning in artificial intelligence, semantic technologies and the so-called semantic web. An ontology is effectively a computable and semantically rich terminology that presents a knowledge or domain model for a topic area. Expressions of knowledge or assertions are stored using formally defined term. This knowledge base is reasoned over to yield answers to queries, among other things. Ontologies have been developed in knowledge-based projects across various disciplines, and used for such things as search engines, chatbots, enterprise knowledge graphs, etc. Ontologies support: interoperability, automated reasoning, data sharing and integration, data search and retrieval, and communicating the meaning of data. The Orbital Debris Ontology (ODO), and related ontologies [Rovetto & Kelso 2016] [Rovetto 2016, 2017], were proposed to help achieve this. ODO, for instance, is intended as a domain ontology that can be used across federated databases, offering an explicitly specified set of concepts describing the orbital debris domain. Its meaning-rich taxonomy will provide a sharable semantics for orbital debris data to, in part, consistently communicate the meaning of data to both humans and machines, and tag data elements in space object catalogs to help afford inference tasks, decision support, knowledge discovery, and information integration. ODO and the SSA ontology (SSAO) is part of the overall Orbital Space Domain Ontology concept, which is conceived as a broader domain reference ontology. It aims to provide a knowledge representation structure of the orbital space environment, a common semantic model, and develop a sharable terminology. Collectively this will provide common meaning for datasets, a high-level taxonomy or classification for orbital space objects, and thus means to characterize space objects. Ongoing efforts have included using visualizations, R, JSON-LD, and contemporary semantic technologies. Potential applications and interdisciplinary partnerships include web-based platforms, web apps, visualizations, and academia projects. Community input and participation may yield a more widely understood domain model as well as facilitate terminological standards. For example, the proposed conceptual, terminological and ontological analysis may contribute to such efforts as the Space Debris Mitigation Requirements in the International Standards Organization by developing more precise, consistent and coherent terms and definitions. Projects that seek to develop in-house ontologies can use ODO and related ontologies as domain reference ontologies. This paper was developed independent of author affiliations. Readers are encouraged to contact corresponding author(1) with general interest and potential opportunities to support or realize the described project.

Robert J. Rovetto↗

A Data Processing Pipeline for Adversarial Socio-Technical Network Analysis

With the rapid adoption of emerging technologies, there is a need to catalog and model sociotechnical interdependencies that have been historically used to influence the operation of Critical Infrastructure networks including the impacts of mergers and acquisitions, hostile takeovers, and foreign investment. Our research intends to address this need with two primary contributions. First, we have developed a data curation and processing pipeline to generate sociotechnical networks extracted from a variety of data sources including SEC filings and infrastructure asset databases. The pipeline, implemented in Apache Airflow, extracts and normalizes the representation of entities and relations, specified within ontologies. Our intent is to provide an extensible, machine-actionable approach to quickly communicate such models, reproduce previous results, and adapt them to new, unanticipated situations. Second, networks produced by our pipeline enable the development of graph-theoretic metrics that consider the properties of network components in addition to its topology. Metadata associated with network components---whether semantic, temporal, or geospatial---affects the alignment of generated networks with assumptions underlying complexity metrics. Validation of generated networks relative to component types defined by an ontology, may allow the research community to adapt metrics to the semantics of the domains being studied. Generated networks may be processed as knowledge, dynamic, or spatial graphs and enables a variety of analyses including automated reasoning and measures of network complexity. Automated reasoning views extracted entities and relations as a knowledge graph; this enables application of inference rules that represent historically-attested adversarial business methods and applies that behavior to a specific geographic context. Measures of network complexity, including degree distribution, reachability analyses, temporal analysis, and community detection can be adapted to indicate adversarial organizational influence.

97 MATHEMATICS AND COMPUTING↗

Accurate Prediction of Voltage of Battery Electrode Materials Using Attention-Based Graph Neural Networks

Performing first-principles calculations to discover electrodes’ properties in the large chemical space is a challenging task. While machine learning (ML) has been applied to effectively accelerate those discoveries, most of the applied methods ignore the materials’ spatial information and only use predefined features: based only on chemical compositions. Here, we propose two attention-based graph convolutional neural network techniques to learn the average voltage of electrodes. Our proposed methods, which combine both atomic composition and atomic coordinates in 3D-space, improve the accuracy in voltage prediction significantly when compared to composition-based ML models. The first model directly learns the chemical reaction of electrodes and metal ions to predict their average voltage, whereas the second model combines electrodes’ ML predicted formation energy (E form ) to compute their average voltage. Our E form -based model demonstrates improved accuracy in transferability from our subset of learned Li ions to Na ions. Moreover, we predicted the theoretical voltage of 10 Na x MPO 4 F (M = Ti, Cr, Fe, Cu, Mn, Co, and Ni) fluorophosphate battery frameworks, which are unavailable in the Material Project database. It could be shown that we can expect average voltages higher than 3.1 V from those Na battery frameworks except from the NaTiPO 4 F and TiPO 4 F pair of electrodes, which offer an average voltage of 1.32 V.

25 ENERGY STORAGE↗

Customer Avionics Interface Development and Analysis Development Activity Tracking System

The Customer Avionics Interface Development and Analysis (CAIDA) Development Activity (DA)Tracking System is a Microsoft Access Database that tracks, organizes, and analyzes data about DAs from the work management tool. The CAIDA DA Tracking System takes data imported from the work management system. Once in the system, the data is filtered to generate each DA’s Asset. From there, many different queries are run on the data and their results are imported into forms to create graphs to get and display a wide range of metrics. These graphs are automatically updated with each new data import and over time. They can be easily exported for use in presentations, documents, etc. Additionally, the system is highly customizable and can be added upon to include more data members, generate new graphs, and much more.

Tara Conti↗

EnviroNET: On-line information for LDEF

EnviroNET is an on-line, free-form database intended to provide a centralized repository for a wide range of technical information on environmentally induced interactions of use to Space Shuttle customers and spacecraft designers. It provides a user-friendly, menu-driven format on networks that are connected globally and is available twenty-four hours a day - every day. The information, updated regularly, includes expository text, tabular numerical data, charts and graphs, and models. The system pools space data collected over the years by NASA, USAF, other government research facilities, industry, universities, and the European Space Agency. The models accept parameter input from the user, then calculate and display the derived values corresponding to that input. In addition to the archive, interactive graphics programs are also available on space debris, the neutral atmosphere, radiation, magnetic fields, and the ionosphere. A user-friendly, informative interface is standard for all the models and includes a pop-up help window with information on inputs, outputs, and caveats. The system will eventually simplify mission analysis with analytical tools and deliver solutions for computationally intense graphical applications to do 'What if...' scenarios. A proposed plan for developing a repository of information from the Long Duration Exposure Facility (LDEF) for a user group is presented.

Lauriente, Michael↗

Long plasma duration operation analyses with an international multi-machine (tokamaks and stellarators) database

Combined high-fusion performance and long-pulse operation is one of the key integration challenges for fusion energy development in magnetic devices. Addressing these challenges requires an integrated vision of physics and engineering aspects with the purpose of simultaneously increasing time duration and fusion performance. Significant progress has been made in tokamaks and stellarators, including very recent achievement in duration and/or performance. This progress is reviewed by analyzing the experimental data (109 plasma pulses with a total of 3200 data points, i.e. on average 29 data per pulse) provided by ten tokamaks (in alphabetical order: Axially Symmetric Divertor Experiment Upgrade, DIII-D, Experimental Advanced Superconducting Tokamak, Joint European Torus, JT-60 Upgrade, Korea Superconducting Tokamak Advanced Research, tokamak à configuration variable, Tokamak Fusion Test Reactor, Tore Supra, W Environment in Steady-State Tokamak) and two stellarators (Large Helical Device and Wendelstein 7-X) expanding the pioneering work of Kikuchi (Kikuchi M. and Azumi M. 2015 Frontiers in Fusion Research II: Introduction to Modern Tokamak Physics (Springer)). Data have been gathered up to January 2022 and coordination has been provided by the recently created International Energy Agency-International Atomic Energy Agency international Coordination on International Challenges on Long duration OPeration group. By exploiting the multi-machine international database, recent progress in terms of injected energies (e.g. 1730 MJ in L-mode, 425 MJ in H-mode), durations (1056 s in L-mode, 101 s in H-mode), injected powers, and sustained performance will be reviewed. Progress has been made to sustain long-pulse operation in tokamaks and stellarators with superconducting coils, actively cooled components, and/or with metallic walls. The graph of the fusion triple products as a function of duration shows a dramatic reduction of, at least two orders of magnitude when increasing the plasma duration from less than 1 s to 100 s. Indeed, long-pulse operation is usually reached in dominant electron-heating modes at reduced density (current drive optimization) but with low ion temperatures ranging from 1 to 3 keV for discharges above 100 s. Difficulties in extending the duration may arise from coupling high heating powers over long durations and the evolving plasma-wall interaction towards an unstable operational domain. Possible causes limiting the duration and critical issues to be addressed prior to ITER operation and DEMO design are reported and analyzed.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Telemetry and Science Data Software System

The Telemetry and Science Data Software System (TSDSS) was designed to validate the operational health of a spacecraft, ease test verification, assist in debugging system anomalies, and provide trending data and advanced science analysis. In doing so, the system parses, processes, and organizes raw data from the Aquarius instrument both on the ground and while in space. In addition, it provides a user-friendly telemetry viewer, and an instant pushbutton test report generator. Existing ground data systems can parse and provide simple data processing, but have limitations in advanced science analysis and instant report generation. The TSDSS functions as an offline data analysis system during I&T (integration and test) and mission operations phases. After raw data are downloaded from an instrument, TSDSS ingests the data files, parses, converts telemetry to engineering units, and applies advanced algorithms to produce science level 0, 1, and 2 data products. Meanwhile, it automatically schedules upload of the raw data to a remote server and archives all intermediate and final values in a MySQL database in time order. All data saved in the system can be straightforwardly retrieved, exported, and migrated. Using TSDSS s interactive data visualization tool, a user can conveniently choose any combination and mathematical computation of interesting telemetry points from a large range of time periods (life cycle of mission ground data and mission operations testing), and display a graphical and statistical view of the data. With this graphical user interface (GUI), the data queried graphs can be exported and saved in multiple formats. This GUI is especially useful in trending data analysis, debugging anomalies, and advanced data analysis. At the request of the user, mission-specific instrument performance assessment reports can be generated with a simple click of a button on the GUI. From instrument level to observatory level, the TSDSS has been operating supporting functional and performance tests and refining system calibration algorithms and coefficients, in sync with the Aquarius/SAC-D spacecraft. At the time of this reporting, it was prepared and set up to perform anomaly investigation for mission operations preceding the Aquarius/SAC-D spacecraft launch on June 10, 2011.

Bates, Lakesha↗

An automated workflow that generates atom mappings for large‐scale metabolic models and its application to Arabidopsis thaliana

SUMMARY Quantification of reaction fluxes of metabolic networks can help us understand how the integration of different metabolic pathways determines cellular functions. Yet, intracellular fluxes cannot be measured directly but are estimated with metabolic flux analysis (MFA), which relies on the patterns of isotope labeling of metabolites in the network. The application of MFA also requires a stoichiometric model with atom mappings that are currently not available for the majority of large‐scale metabolic network models, particularly of plants. While automated approaches such as the Reaction Decoder Toolkit (RDT) can produce atom mappings for individual reactions, tracing the flow of individual atoms of the entire reactions across a metabolic model remains challenging. Here we establish an automated workflow to obtain reliable atom mappings for large‐scale metabolic models by refining the outcome of RDT, and apply the workflow to metabolic models of Arabidopsis thaliana . We demonstrate the accuracy of RDT through a comparative analysis with atom mappings from a large database of biochemical reactions, MetaCyc. We further show the utility of our automated workflow by simulating 15 N isotope enrichment and identifying nitrogen (N)‐containing metabolites which show enrichment patterns that are informative for flux estimation in future 15 N‐MFA studies of A. thaliana . The automated workflow established in this study can be readily expanded to other species for which metabolic models have been established and the resulting atom mappings will facilitate MFA and graph‐theoretic structural analyses with large‐scale metabolic networks.

59 BASIC BIOLOGICAL SCIENCES↗

Defect Diffusion Graph Neural Networks for Materials Discovery in High-Temperature Energy Applications

Here, the migration of crystallographic defects dictates material properties and performance for a plethora of technological applications. Density functional theory (DFT)-based nudged elastic band (NEB) calculations are a powerful computational technique for predicting defect migration activation energy barriers, yet they become prohibitively expensive for high-throughput screening of defect diffusivities. Without introducing hand-crafted (i.e., chemistry- or structure-specific) descriptors, we propose a generalized deep learning approach to train surrogate models for NEB energies of vacancy migration by hybridizing graph neural networks with transformer encoders and simply using pristine host structures as input. With sufficient training data, computationally efficient and simultaneous inference of vacancy defect thermodynamics and migration activation energies can be obtained to compute temperature-dependent vacancy diffusivities and to down-select candidates for more thorough DFT analysis or experiments. Thus, as we specifically demonstrate for potential water-splitting materials, candidates with desired defect thermodynamics, kinetics, and host stability properties can be more rapidly targeted from open-source databases of experimentally validated or hypothetical materials.

14 SOLAR ENERGY↗

Retaining Systems Engineering Model Meaning Through Transformation: Demo 2

Digital engineering strategies typically assume that digital engineering models interoperate seamlessly across the multiple different engineering modeling software applications involved, such as model- based systems engineering (MBSE), mechanical computer-aided design (MCAD), electrical computer-aided design (ECAD), and other engineering modeling applications. The presumption is that the data schema in these modeling software applications are structured in the familiar flat- tabular schema like any other software application. Engineering domain-specific applications (e.g., systems, mechanical, electrical, simulation) are typically designed to solve domain-specific problems, necessarily excluding explicit representations of non-domain information to help the engineer focus on the domain problems (system definition, design, simulation). Such exclusions become problematic in inter-domain information exchange. The obvious assumptions of one domain might not be so obvious to experts in another domain. Ambiguity in domain-specific language can erode the ability to enable different domain modeling applications to interoperate, unless the underlying language is understood and used as the basis for translation from one application to another. The engineering modeling software application industry has struggled for decades to enable these applications to interoperate. Industry standards have been developed, but they have not unified the industry. Why is this? The authors assert that the industry has relied on traditional database integration methods. The basic issue prohibiting successful application integration then is that traditional database-driven integration does not consider the distinct languages of each domain. An engineering models meaning is expressed through the underlying language of that engineering domain. In essence, traditional integration methods do not retain the semantic context (meaning) of the model. The basis of this research stems from the widely held assumption that systems engineering models are (or can be) structured according to the underlying semantic ontology of the model. This assumption can be imagined from two thoughts. 1) Digital systems engineering models are often represented using graph theory (the graph of a complex systems model can contain millions of nodes and edges). When examining the nodes one at a time and following the outbound edges of each node one by one, one can end up with rudimentary statements about the model (i.e., node A relates to node B), as in a semantic graph. 2) Likewise, from the study of natural languages, a sentence can be structured into unambiguous triples of subject-predicate-object within formal and highly expressive semantic ontologies. The rudimentary statements about a systems model discerned with graph theory closely mimic the triples used in the ontologies that try to structure natural languages. In other words, a systems models semantic graph can be (or is) structured into an ontology. Additionally, it is well established in industry that through natural language processing (NLP), which provides the means to create language structures, that computers can interpret ontological graphs. Therefore, the authors hypothesized that if the integrity of the underlying semantic structure of a systems model is retained, the contextual meaning of the model is retained. By structuring system models into the triples of the underlying ontology during the transformation from one MBSE application to another, the authors have provided a proof of the concept that the meaning of a system model can be retained during transformation. The authors assert that this is the missing ingredient in effective systems model-to-model interoperability. ACKNOWLEDGEMENTS The authors would like to thank the FY19 Model Interoperability team members who provided a solid foundation for the FY20 team to leverage: John McCloud, for the work he did to guide us toward the right use of technology that will appropriately discover and manipulate ontologies. Carlos Tafoya, for the work he did to develop an application programming interface (API)/Adapter that would export ontology-based data from GENESYS. Peter Chandler, for the work he did to architect our overall integration solution, with an eye toward the future that would influence a large-scale federated production-level systems engineering digital model ecosystem.

42 ENGINEERING↗

A framework to evaluate machine learning crystal stability predictions

The rapid adoption of machine learning in various scientific domains calls for the development of best practices and community agreed-upon benchmarking tasks and metrics. We present Matbench Discovery as an example evaluation framework for machine learning energy models, here applied as pre-filters to first-principles computed data in a high-throughput search for stable inorganic crystals. We address the disconnect between (1) thermodynamic stability and formation energy and (2) retrospective and prospective benchmarking for materials discovery. Alongside this paper, we publish a Python package to aid with future model submissions and a growing online leaderboard with adaptive user-defined weighting of various performance metrics allowing researchers to prioritize the metrics they value most. To answer the question of which machine learning methodology performs best at materials discovery, our initial release includes random forests, graph neural networks, one-shot predictors, iterative Bayesian optimizers and universal interatomic potentials. We highlight a misalignment between commonly used regression metrics and more task-relevant classification metrics for materials discovery. Accurate regressors are susceptible to unexpectedly high false-positive rates if those accurate predictions lie close to the decision boundary at 0 eV per atom above the convex hull. The benchmark results demonstrate that universal interatomic potentials have advanced sufficiently to effectively and cheaply pre-screen thermodynamic stable hypothetical materials in future expansions of high-throughput materials databases.

Riebesell, Janosh↗

The Open Data Repositorys Data Publisher

Data management and data publication are becoming increasingly important components of researcher's workflows. The complexity of managing data, publishing data online, and archiving data has not decreased significantly even as computing access and power has greatly increased. The Open Data Repository's Data Publisher software strives to make data archiving, management, and publication a standard part of a researcher's workflow using simple, web-based tools and commodity server hardware. The publication engine allows for uploading, searching, and display of data with graphing capabilities and downloadable files. Access is controlled through a robust permissions system that can control publication at the field level and can be granted to the general public or protected so that only registered users at various permission levels receive access. Data Publisher also allows researchers to subscribe to meta-data standards through a plugin system, embargo data publication at their discretion, and collaborate with other researchers through various levels of data sharing. As the software matures, semantic data standards will be implemented to facilitate machine reading of data and each database will provide a REST application programming interface for programmatic access. Additionally, a citation system will allow snapshots of any data set to be archived and cited for publication while the data itself can remain living and continuously evolve beyond the snapshot date. The software runs on a traditional LAMP (Linux, Apache, MySQL, PHP) server and is available on GitHub (http://github.com/opendatarepository) under a GPLv2 open source license. The goal of the Open Data Repository is to lower the cost and training barrier to entry so that any researcher can easily publish their data and ensure it is archived for posterity.

Astrobiology data↗

Database Performance Monitoring for DUNE

This report presents the research, design, and implementation of improved PostgreSQL monitoring for DUNE Rucio database services using Checkmk. The project began with a request to improve dashboard visibility for database performance metrics, including connection usage, configured connection limits, lock activity, wait behavior, storage trends, query performance, and saturation alerts. The initial implementation focused on the dune_rucio_prod database on the rucio_prod PostgreSQL instance because connection saturation and lock contention are direct reliability risks for database-backed services. Existing Checkmk PostgreSQL monitoring was investigated, and several gaps were identified. Built-in connection monitoring did not clearly separate active, idle, idle-in-transaction, total, and usage-percent metrics, while the built-in lock monitoring simplified PostgreSQL lock modes into shared and exclusive categories. To address these gaps, two DSG-specific Checkmk local checks were created: one for connection-state monitoring and one for lock-state monitoring. These checks supplement the built-in PostgreSQL checks and provide additional performance data for dashboard graphs, service states, and alerts.

Bowers, Elliot [Cabrillo Coll.]↗

Rapid Computational Identification of Therapeutic Targets for Pathogens

Biological threats continue to persist and evolve as an important challenge to national security. There are multiple ways in which novel viral pathogens could emerge to pose a serious threat to human health. This project developed a pathogen target identification tool that can rapidly respond to a novel or emerging viral biological threat. A set of computational tools were developed that provide detailed information on the newly sequenced genes, their protein products and the drug target sites for the proteins that are best suited for biological countermeasure development. Three key innovations were developed in the project. 1) Development of a new extensive database of protein pocket structures with structure-based search algorithms to rapidly link novel protein targets with the complete collection of previously experimentally solved protein structures. 2) A novel clustering pipeline was introduced to group matching structures and associated small-molecule binding ligands into a consensus protein pocket with the associated small-molecule chemotypes predicted to fit in the pocket site. The matching experimentally solved structures were used to inform the value of different target sites. 3) Where there are viral protein targets with pockets structurally matched to similar human proteins, a biological knowledge graph, which links molecular interactions with human disease, was used to further assess the potential negative impact of a viral protein target with similarities to human proteins that could have important off target side effects. In total, the project produced a new resource for rapid and detailed assessment of promising targets for countermeasures, reflecting the ongoing wet lab, clinical, and computational data being collected. These capabilities will improve the ability to respond to a biological threat in multiple domains.

59 BASIC BIOLOGICAL SCIENCES↗

Database Performance Monitoring for DUNE

This project improves Checkmk monitoring for DUNE Rucio PostgreSQL database services by adding clearer dashboard visibility for connection and lock behavior. The work began with a request to monitor database performance metrics such as connection usage, configured limits, lock activity, wait behavior, query performance, storage trends, and saturation alerts. Existing Checkmk PostgreSQL checks were reviewed, and gaps were identified in how connection states and lock modes were displayed. To address these gaps, two DSG-specific local checks were added for dune_rucio_prod: one for connection-state monitoring and one for lock-state monitoring. These checks report active, idle, idle-in-transaction, total, usage-percent, lock-mode, waiting-lock, and wait-age metrics. The added metrics supplement built-in Checkmk monitoring and provide DUNE application developers with clearer service states, history graphs, dashboard widgets, and alerts.

Bowers, Elliot [Cabrillo Coll.]↗

Noiseless compression using non-Markov models

Adaptive data compression techniques can be viewed as consisting of a model specified by a database common to the encoder and decoder, an encoding rule and a rule for updating the model to ensure that the encoder and decoder always agree on the interpretation of the next transmission. The techniques which fit this framework range from run-length coding, to adaptive Huffman and arithmetic coding, to the string-matching techniques of Lempel and Ziv. The compression obtained by arithmetic coding is dependent on the generality of the source model. For many sources, an independent-letter model is clearly insufficient. Unfortunately, a straightforward implementation of a Markov model requires an amount of space exponential in the number of letters remembered. The Directed Acyclic Word Graph (DAWG) can be constructed in time and space proportional to the text encoded, and can be used to estimate the probabilities required for arithmetic coding based on an amount of memory which varies naturally depending on the encoded text. The tail of that portion of the text which was encoded is the longest suffix that has occurred previously. The frequencies of letters following these previous occurrences can be used to estimate the probability distribution of the next letter. Experimental results indicate that compression is often far better than that obtained using independent-letter models, and sometimes also significantly better than other non-independent techniques.

Blumer, Anselm↗