Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Domain knowledge”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12

Mining for Gluon Saturation at Colliders

Quantum chromodynamics (QCD) is the theory of strong interactions of quarks and gluons collectively called partons, the basic constituents of all nuclear matter. Its non-abelian character manifests in nature in the form of two remarkable properties: color confinement and asymptotic freedom. At high energies, perturbation theory can result in the growth and dominance of very gluon densities at small-x. If left uncontrolled, this growth can result in gluons eternally growing violating a number of mathematical bounds. The resolution to this problem lies by balancing gluon emissions by recombinating gluons at high energies: phenomena of gluon saturation. High energy nuclear and particle physics experiments have spent the past decades quantifying the structure of protons and nuclei in terms of their fundamental constituents confirming predicted extraordinary behavior of matter at extreme density and pressure conditions. In the process they have also measured seemingly unexpected phenomena. We will give a state of the art review of the underlying theoretical and experimental tools and measurements pertinent to gluon saturation physics. We will argue for the need of high energy electron-proton/ion colliders such as the proposed EIC (USA) and LHeC (Europe) to consolidate our knowledge of QCD knowledge in the small x kinematic domains.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Spatial Transferability of Machine Learning Based Volume Estimation Models

High-quality traffic volume data is essential for efficient transportation planning and operations. However, such high-quality data is expensive to collect, owing primarily to the high capital cost of installing and maintaining continuous counting stations (CCSs). Recent availability of probe-based vehicle data offers a cost-effective solution for increasing the observability of traffic volumes. However, having ample ground truth traffic data is a prerequisite for developing robust volume estimation models. Though this might not be a big issue in many states, states with scarce CCS data might be able to benefit from robust volume estimation models developed in (adjacent) data-rich states. While there is a reasonable amount of spatial transferability research in the transportation domain, there is a dearth of knowledge on the spatial transferability of probe-based volume estimation models. To address this gap, this paper explores spatial transferability of volume estimation models developed from data in three states (Colorado, North Carolina, and Pennsylvania). Results indicate that it is extremely important to maintain temporal consistency when attempting spatial transferability of volume estimation models. It was also found that models trained on regions with lower peak traffic volumes will limit the performance of models transferred to states with higher peak hourly traffic volumes. Corroborating findings from existing spatial transferability research on other topics, it was found that a meta-model (developed using data from multiple states) performs better than volume estimation models developed within any one of the states.

ADVANCED PROPULSION SYSTEMS↗

An adaptive knowledge-based data-driven approach for turbulence modeling using ensemble learning technique under complex flow configuration: 3D PWR sub-channel with DNS data

This work describes a new approach to increase the accuracy of Reynolds-averaged Navier–Stokes (RANS) in modeling turbulence flow leveraging the machine learning technique. Traditionally, different turbulence models for Reynolds stress are developed for different flow patterns based on human knowledge. Each turbulence model has a certain application domain and prediction uncertainty. In recent years, with the rapid improvements of machine learning techniques, researchers start to develop an approach to compensate for the prediction discrepancy of traditional turbulence models with statistical models and data. However, the approach has deficiencies in several aspects. For example, the amount of human knowledge introduced to the statistical model couldn’t be controlled, which makes the statistical model learn from a very naïve stage and limits its application. In this work, a new approach is developed to address those deficiencies. Here, the new approach uses the “ensemble learning” technique to control the amount of human knowledge introduced into the statistical model. Therefore, the new approach could be adaptive to the multiple application domains. In conclusion, according to the results of case study, the new approach shows higher accuracy than both traditional turbulence models and the previous machine learning approach.

42 ENGINEERING↗

Dynamic Retrieval Augmented Generation of Ontologies using Artificial Intelligence (DRAGON-AI)

Ontologies are fundamental components of informatics infrastructure in domains such as biomedical, environmental, and food sciences, representing consensus knowledge in an accurate and computable form. However, their construction and maintenance demand substantial resources and necessitate substantial collaboration between domain experts, curators, and ontology experts. We present Dynamic Retrieval Augmented Generation of Ontologies using AI (DRAGON-AI), an ontology generation method employing Large Language Models (LLMs) and Retrieval Augmented Generation (RAG). DRAGON-AI can generate textual and logical ontology components, drawing from existing knowledge in multiple ontologies and unstructured text sources.We assessed performance of DRAGON-AI on de novo term construction across ten diverse ontologies, making use of extensive manual evaluation of results. Our method has high precision for relationship generation, but has slightly lower precision than from logic-based reasoning. Our method is also able to generate definitions deemed acceptable by expert evaluators, but these scored worse than human-authored definitions. Notably, evaluators with the highest level of confidence in a domain were better able to discern flaws in AI-generated definitions. We also demonstrated the ability of DRAGON-AI to incorporate natural language instructions in the form of GitHub issues.These findings suggest DRAGON-AI's potential to substantially aid the manual ontology construction process. However, our results also underscore the importance of having expert curators and ontology editors drive the ontology generation process.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

What Are Ontologies and When Should They Be Used?

Data without description is at best unusable, and at worst, misused. If we do not understand the assumptions and meaning of our data, we are unable to confidently use it. Data today is largely described within a database’s schema, detailing structure and primitive datatypes as part of a relational model, but if we require assurance some data value can be correctly evaluated alongside others beyond the immediate systems in which they are defined, a more portable, richer semantics is needed. Ontologies define knowledge unambiguously across systems and establish the means to reason upon said knowledge using logical inference. They model neutral domains of information rather than data definitions from software or databases that would only serve to enrich a single system’s idiosyncrasies. In this paper, we take a casual stance to explore what ontologies are, how they are built, why they are useful, and when they should be used.

97 MATHEMATICS AND COMPUTING↗

A change language for ontologies and knowledge graphs

Ontologies and knowledge graphs (KGs) are general-purpose computable representations of some domain, such as human anatomy, and are frequently a crucial part of modern information systems. Most of these structures change over time, incorporating new knowledge or information that was previously missing. Managing these changes is a challenge, both in terms of communicating changes to users and providing mechanisms to make it easier for multiple stakeholders to contribute. To fill that need, we have created KGCL, the Knowledge Graph Change Language (https://github.com/INCATools/kgcl), a standard data model for describing changes to KGs and ontologies at a high level, and an accompanying human-readable Controlled Natural Language (CNL). This language serves two purposes: a curator can use it to request desired changes, and it can also be used to describe changes that have already happened, corresponding to the concepts of “apply patch” and “diff” commonly used for managing changes in text documents and computer programs. Another key feature of KGCL is that descriptions are at a high enough level to be useful and understood by a variety of stakeholders—e.g. ontology edits can be specified by commands like “add synonym ‘arm’ to ‘forelimb’” or “move ‘Parkinson disease’ under ‘neurodegenerative disease’.” We have also built a suite of tools for managing ontology changes. These include an automated agent that integrates with and monitors GitHub ontology repositories and applies any requested changes and a new component in the BioPortal ontology resource that allows users to make change requests directly from within the BioPortal user interface. Overall, the KGCL data model, its CNL, and associated tooling allow for easier management and processing of changes associated with the development of ontologies and KGs.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Technology Transfer from Fermi Research Alliance to Itasca Plastics for the purpose of Commercializing Scintillator Material

Researchers at the Fermi National Accelerator Laboratory (Fermilab) developed extruded plastic scintillator in the late 1990s, which was first used in the D-Zero experiment. Extruded plastic scintillator is currently produced at Fermilab and is used in particle detectors worldwide. The purpose of this CRADA is to transfer the knowledge related to the Fermilab extrusion process to Itasca Plastics, Inc. (Itasca Plastics). Much of this knowledge is contained in documentation that is in the public domain, although it is distributed over several communications (papers, conference records, etc.) and over several years. Under this CRADA Fermilab will assemble the information, provide it to Itasca Plastics and provide limited consulting to complete the knowledge transfer. If the transfer is successful, Itasca Plastics will be able to establish a U.S. commercial manufacturing capability for extruded scintillator material that can be used for high energy physics and commercial applications.

36 MATERIALS SCIENCE↗

Outdoor annual algae productivity improvements at the pre-pilot scale through crop rotation and pond operational management strategies

The Development of Integrated Screening, Cultivar Optimization, and Verification Research (DISCOVR) collaborative consortium operated pre-pilot scale outdoor ponds to deliver much-needed multi-year, long-term and consistent, algae cultivation data relevant to understanding the current state of technology in terms of expected seasonal algae biomass productivity. Over the course of four years from 2018 to 2021, twelve identical 4.2 m 2 mini-ponds were run in triplicate sets to test strains and operational strategies demonstrated in small-, indoor photobioreactors, in pursuit of increasing overall algae areal productivity and projected farm yield. Fourteen different cultivars derived from a strain screening pipeline were tested. Through deliberate seasonal crop rotation and improvements in operational strategies, annual biomass productivity increased from 11.6 to 17.6 g m -2 day -1 , a > 50% increase over the 2018 baseline. Both brackish and marine strains were included and four out of the fourteen strains consistently yielded high productivity across multiple years; brackish strains Monoraphidium minutum (26BAM) and Scenedesmus obliquus (UTEX393), and marine strains Tetraselmis striata (LANL1001) and Picochlorum celeri (TG2). These freely available datasets, which represent nearly complete annual daily coverage of cultivation metrics including weather, pond temperature and pH, nutrients, and productivity, are unique in the public domain and seek to fill agronomic and operational knowledge gaps to help in the eventual commercialization of algal biofuels and bioproducts.

09 BIOMASS FUELS↗

Local Strain and Polarization Mapping in Ferrielectric Materials

CuInP 2 S 6 (CIPS) is a van der Waals material that has attracted attention because of its unusual properties. Recently, a combination of density functional theory (DFT) calculations and piezoresponse force microscopy (PFM) showed that CIPS is a uniaxial quadruple-well ferrielectric featuring two polar phases and a total of four polarization states that can be controlled by external strain. In this study, we combine DFT and PFM to investigate the stress-dependent piezoelectric properties of CIPS, which have so far remained unexplored. The two different polarization phases are predicted to differ in their mechanical properties and the stress sensitivity of their piezoelectric constants. This knowledge is applied to the interpretation of ferroelectric domain images, which enables investigation of local strain and stress distributions. The interplay of theory and experiment produces polarization maps and layer spacings which we compare to macroscopic X-ray measurements. We found that the sample contains only the low-polarization phase and that domains of one polarization orientation are strained, whereas domains of the opposite polarization direction are fully relaxed. The described nanoscale imaging methodology is applicable to any material for which the relationship between electromechanical and mechanical characteristics is known, providing insight on structural, mechanical, and electromechanical properties down to ~10 nm length scales.

36 MATERIALS SCIENCE↗

Assessment of the Distributed Ledger Technology for Energy Sector Industrial and Operational Applications Using the MITRE ATT&CK® ICS Matrix

In recent times, Distributed Ledger Technology (DLT) has gained significant attention for its potential application in the energy sector. Utilizing blockchain and DLT has demonstrated the ability to enhance the resilience of the electric infrastructure, which will support a more flexible infrastructure and advance grid modernization. However, the deployment of these technologies increases the overall attack surface. The MITRE ATT&CK® matrices have been developed to document an adversary’s tactics and techniques based on real-world observations. The MITRE ATT&CK® matrices provide a common taxonomy for offense and defense and have become a valuable conceptual tool across multiple cybersecurity disciplines for conveying threat intelligence, performing testing through red teaming or adversary emulation, and enhancing network and system defenses against intrusions. The MITRE ATT&CK® for Industrial Control Systems (ICS) matrix was created to provide knowledge about adversary behavior in the ICS technology domain. This study analyzes the relevance of various tactics and techniques across a seven-layer DLT engineering and cybersecurity stack, known as the DLT stack, designed by the Cybersecurity Taskforce under IEEE P2418.5 - Standard for Blockchain in Energy working group sponsored by Power and Energy Systems - Smart Buildings, Loads and Customer Systems (PES/SBLC) Technical Committee. Additionally, this paper identifies specific mitigation strategies tailored to the energy ICS environment

42 ENGINEERING↗

Leveraging Structured Biological Knowledge for Counterfactual Inference: A Case Study of Viral Pathogenesis

Counterfactual inference is a useful tool for comparing outcomes of interventions on complex systems. It requires us to represent the system in form of a structural causal model, complete with a causal diagram, probabilistic assumptions on exogenous variables, and functional assignments. Specifying such models can be extremely difficult in practice. The process requires substantial domain expertise, and does not scale easily to large systems, multiple systems, or novel system modifications. At the same time, many application domains, such as molecular biology, are rich in structured causal knowledge that is qualitative in nature. This manuscript proposes a general approach for querying a causal knowledge graph with a causal question and converting the qualitative result into a quantitative structural causal model that can learn from data to answer the question. Here, we demonstrate the feasibility, accuracy and versatility of this approach using two case studies in systems biology. The first demonstrates the appropriateness of the underlying assumptions and the accuracy of the results. The second demonstrates the versatility of the approach by querying a knowledge base for the molecular determinants of a severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2)-induced cytokine storm and performing counterfactual inference to predict the causal effect of medical countermeasures for severely ill COVID-19 patients.

60 APPLIED LIFE SCIENCES↗

Staphylococcus aureus FtsZ and PBP4 bind to the conformationally dynamic N-terminal domain of GpsB

In the Firmicutes phylum, GpsB is a membrane associated protein that coordinates peptidoglycan synthesis with cell growth and division. Although GpsB has been studied in several bacteria, the structure, function, and interactome of Staphylococcus aureus GpsB is largely uncharacterized. To address this knowledge gap, we solved the crystal structure of the N-terminal domain of S. aureus GpsB, which adopts an atypical, asymmetric dimer, and demonstrates major conformational flexibility that can be mapped to a hinge region formed by a three-residue insertion exclusive to Staphylococci. When this three-residue insertion is excised, its thermal stability increases, and the mutant no longer produces a previously reported lethal phenotype when overexpressed in Bacillus subtilis. In S. aureus, we show that these hinge mutants are less functional and speculate that the conformational flexibility imparted by the hinge region may serve as a dynamic switch to fine-tune the function of the GpsB complex and/or to promote interaction with its various partners. Furthermore, we provide the first biochemical, biophysical, and crystallographic evidence that the N-terminal domain of GpsB binds not only PBP4, but also FtsZ, through a conserved recognition motif located on their C-termini, thus coupling peptidoglycan synthesis to cell division. Taken together, the unique structure of S. aureus GpsB and its direct interaction with FtsZ/PBP4 provide deeper insight into the central role of GpsB in S. aureus cell division.

59 BASIC BIOLOGICAL SCIENCES↗

Predictive Indicators of the Performance of Large Language Models

In several mission contexts, it is desirable to estimate the performance of large language models (LLMs) on tasks that we cannot run directly. In light of published “scaling laws” our hypothesis is that some tasks should be consistently more challenging than others based on characteristics of the task. The goal of this project was to begin quantifying how much information about LLM performance can be gained from the features of a model and a task. Two of our statistical models struggled to converge. Pass/fail test results may provide limited information for inference beyond model quality and task difficulty, but we see no evidence at this time for significant feature interaction effect sizes, arguing for simple models. Future work extending the models to capitalize on perplexity of ground truth answers is suggested. This project also introduces “Depth of Knowledge Variant Testing” as a strategy for more finely assessing language models on open domain question and answer tasks. We developed sets of questions that ask a language model to produce similar information while demonstrating increasing depth of knowledge, and also relabeled existing Q&A test questions with their depth of knowledge. Our results suggest further consideration of Bloom’s taxonomy and further refinement of prompts to properly elicit information at varying depths. In the course of this work, we set up a basic infrastructure for standardizing tasks and testing many language models on these tasks. In addition to testing the predictive quality of model features and performance across test suites, with this project we have introduced two new task features to contextualize each test question: the Dewey Classification main category of information covered, and the Bloom’s taxonomy level that corresponds to the depth of knowledge probed by the question. Splits across these and other features produced over five hundred task subtypes with distinct feature vectors, which we tested on half a dozen models.

97 MATHEMATICS AND COMPUTING↗

A Data Processing Pipeline for Adversarial Socio-Technical Network Analysis

With the rapid adoption of emerging technologies, there is a need to catalog and model sociotechnical interdependencies that have been historically used to influence the operation of Critical Infrastructure networks including the impacts of mergers and acquisitions, hostile takeovers, and foreign investment. Our research intends to address this need with two primary contributions. First, we have developed a data curation and processing pipeline to generate sociotechnical networks extracted from a variety of data sources including SEC filings and infrastructure asset databases. The pipeline, implemented in Apache Airflow, extracts and normalizes the representation of entities and relations, specified within ontologies. Our intent is to provide an extensible, machine-actionable approach to quickly communicate such models, reproduce previous results, and adapt them to new, unanticipated situations. Second, networks produced by our pipeline enable the development of graph-theoretic metrics that consider the properties of network components in addition to its topology. Metadata associated with network components---whether semantic, temporal, or geospatial---affects the alignment of generated networks with assumptions underlying complexity metrics. Validation of generated networks relative to component types defined by an ontology, may allow the research community to adapt metrics to the semantics of the domains being studied. Generated networks may be processed as knowledge, dynamic, or spatial graphs and enables a variety of analyses including automated reasoning and measures of network complexity. Automated reasoning views extracted entities and relations as a knowledge graph; this enables application of inference rules that represent historically-attested adversarial business methods and applies that behavior to a specific geographic context. Measures of network complexity, including degree distribution, reachability analyses, temporal analysis, and community detection can be adapted to indicate adversarial organizational influence.

97 MATHEMATICS AND COMPUTING↗

Semantic Property Graph for Scalable Knowledge Graph Analytics

Graphs are a natural and fundamental representation to describe entities, relationships, activities, and evolution of complex systems. Many domains such as communication, citation, procurement, biology, social media, and transportation can be modeled as a set of entities and their relationships. Resource Description Framework (RDF) and Labeled Property Graph (LPG) are two of the most used data models to encode information in a graph. Both models are similar in terms of using basic graph elements such as nodes and edges but differ in terms of the modeling approach, expressibility, serialization, and target applications. RDF is a flexible data exchange model for expressing information about entities but it tends to a have high memory footprint and inefficient storage, which does not make it a natural choice to perform scalable graph analytics. In contrast, LPG has gained traction as a reliable model to perform scalable graph analytic tasks such as sub-graph matching, network alignment, and real-time knowledge graph query. It provides efficient storage, fast traversal, and flexibility to model various real-world domains. At the same time, the LPG lacks the support of a formal knowledge representation such as an ontology to provide automated knowledge inference. We propose Semantic Property Graph (SPG) as a logical projection of reified RDF into the LPG model. SPG continues to use RDF ontology to define the type hierarchy of the projected graph and validate it against a given ontology. We present a framework to convert reified RDF graphs into SPG using two different computing environments. We also present cloud-based graph migration capabilities using Amazon Web Services.

Purohit, Sumit↗

Structural basis for antibody binding to adenylate cyclase toxin reveals RTX linkers as neutralization-sensitive epitopes

RTX leukotoxins are a diverse family of prokaryotic virulence factors that are secreted by the type 1 secretion system (T1SS) and target leukocytes to subvert host defenses. T1SS substrates all contain a C-terminal RTX domain that mediates recruitment to the T1SS and drives secretion via a Brownian ratchet mechanism. Neutralizing antibodies against the Bordetella pertussis adenylate cyclase toxin, an RTX leukotoxin essential for B . pertussis colonization, have been shown to target the RTX domain and prevent binding to the α M β 2 integrin receptor. Knowledge of the mechanisms by which antibodies bind and neutralize RTX leukotoxins is required to inform structure-based design of bacterial vaccines, however, no structural data are available for antibody binding to any T1SS substrate. Here, we determine the crystal structure of an engineered RTX domain fragment containing the α M β 2 -binding site bound to two neutralizing antibodies. Notably, the receptor-blocking antibodies bind to the linker regions of RTX blocks I–III, suggesting they are key neutralization-sensitive sites within the RTX domain and are likely involved in binding the α M β 2 receptor. As the engineered RTX fragment contained these key epitopes, we assessed its immunogenicity in mice and showed that it elicits similar neutralizing antibody titers to the full RTX domain. The results from these studies will support the development of bacterial vaccines targeting RTX leukotoxins, as well as next-generation B . pertussis vaccines.

59 BASIC BIOLOGICAL SCIENCES↗

Towards Geospatial Knowledge Graph Infused Neuro-Symbolic AI for Remote Sensing Scene Understanding

Deep learning has proven its effectiveness in numerous tasks for remote sensing scene understanding. However there is an increasing interest to explore fusion of domain-specific background information to the deep neural network to further improve its performance. Remote sensing researchers are also working towards developing models that generalize and adapt to multiple applications. Generalization challenges coupled with the scarcity of large corpora of high-quality noise-free labelled data, have together fueled an interest for leveraging background information. Knowledge graphs serve as excellent choice to represent domain-specific information in a structured, standardized and extensible manner. Integrating symbolic knowledge representations in the form of Knowledge Graph Embedding (KGE) to perform neuro-symbolic reasoning is an emerging research direction promising significant impacts. This vision paper seeks to position ideas and provoke early thoughts toward advancing neuro-symbolic artificial intelligence in the context of geospatial challenges. Specifically, it conceptualizes and elaborates on an architecture for infusing geospatial knowledge from knowledge graph in a deep neural network pipeline. As guiding case studies - land-use land-cover classification, object detection and instance segmentation can benefit from infusing spatio-contextual information with remote sensing imagery. The discussion further reflects on and articulates the challenges and explainable AI opportunities anticipated when scaling and maintaining large-scale geospatial knowledge graphs.

Potnis, Abhishek↗

Data-Driven Cyber-Attack Detection for PV Farms via Time-Frequency Domain Features

The internetworking of grid-connected power electronics converters (PECs) in photovoltaic (PV) farms has inevitably expanded the cyber-attack surfaces. Here this paper presents a comprehensive study on cyber-attack detection and diagnosis for PEC-enabled PV farms via single waveform sensor to distinguish between normal conditions, open-circuit faults, short-circuit faults, and cyber-attacks. To our knowledge, this has not been attempted before. Firstly, we propose frequency-domain magnitude-based residuals to identify short-circuit faults and a time-domain mean current vector-based feature to distinguish open-circuit faults from other threats. These features can fully reflect the specific physical characteristics of PV farms during threat duration. Secondly, unlike micro phasor measurement units (µPMU) and raw electric waveform-based methods, the proposed innovative features can address novel cyber-attacks that are excluded from the training process. Thirdly, an online hardware-in-the-loop (HIL) testbed using the OPAL-RT real-time digital simulator has verified the effectiveness. The monitoring system runs in real-time while using HIL as an operational solar farm and a National Instruments (NI) data acquisition card as the electric waveform sensor at the point of coupling.

42 ENGINEERING↗