Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Domain knowledge”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13

Foundation Models of Scientific Knowledge for Chemistry: Opportunities, Challenges and Lessons Learned

Foundation models pre-trained on large corpora demonstrate significant gains across many natural language processing tasks and domains e.g., law, healthcare, education, etc. However, only limited efforts have investigated the opportunities and limitations of applying these powerful models to science and security applications. In this work we develop foundation models of scientific knowledge for chemistry to augment scientists with the advanced ability to perceive and reason at scale previously unimagined. Specifically, we build large-scale (1.47B parameter) general-purpose models for chemistry that can be effectively used to perform a wide range of in-domain and out-of-domain tasks. Evaluating these models in a zero-shot setting, we analyze the effect of model and data scaling, knowledge depth, and temporality on model performance in context of model training efficiency. Our novel findings demonstrate that (1) model size significantly contributes to the task performance when evaluated in a zero-shot setting; (2) data quality (aka diversity) affects model performance more than data quantity; (3) similarly, unlike previous work (Luu et al., 2021) temporal order of the documents in the corpus boosts model performance only for specific tasks, e.g., SciQ; and (4) models pre-trained from scratch perform better on in-domain tasks than those tuned from general-purpose models like Open AI’s GPT-2.

Foundation Models, Chemistry↗

Domain-aware Control-oriented Neural Models for Autonomous Underwater Vehicles

Conventional physics-based modeling is a time-consuming bottleneck in control design for complex nonlinear systems like autonomous underwater vehicles (AUVs). In contrast, purely data-driven models, require a large number of observations and lack operational guarantees for safety-critical systems. Data-driven models leveraging available partially characterized dynamics have potential to provide reliable systems models in a typical data-limited scenario for high value complex systems, thereby avoiding months of expensive expert modeling time. In this work we explore this middle-ground between expert-modeled and pure data-driven modeling. We present control-oriented parametric models with varying levels of domain-awareness that exploit known system structure and prior physics knowledge to create constrained deep neural dynamical system models. We employ universal differential equations to construct data-driven blackbox and graybox representations of the AUV dynamics. In addition, we explore a hybrid formulation that explicitly models the residual error related to imperfect graybox models. We compare the prediction performance of the learned models for different distributions of initial conditions and control inputs to assess their suitability for control.

Shaw Cortez, Wenceslao E.↗

Bayesian model-data comparison incorporating theoretical uncertainties

Accurate comparisons between theoretical models and experimental data are critical for scientific progress. However, inferred physical model parameters can vary significantly with the chosen physics model, highlighting the importance of properly accounting for theoretical uncertainties. In this Letter, we present a Bayesian framework that explicitly quantifies these uncertainties by statistically modeling theory errors, guided by qualitative knowledge of a theory’s varying reliability across the input domain. We demonstrate the effectiveness of this approach using two systems: a simple ball drop experiment and multi-stage heavy-ion simulations. In both cases incorporating model discrepancy leads to improved parameter estimates, with systematic improvements observed as additional experimental observables are integrated.

Bayesian methods↗

Pulse shaping in strong-field ionization: Theory and experiments

Intense ultrafast pulses cause dissociative ionization and shaping the pulses may allow control of both electronic and nuclear dynamics that determine ion yields. We report on a combined experimental and theoretical effort to determine how shaped laser pulses affect tunnel ionization, the process that precedes many strong-field phenomena. We carried out experiments on Ar, N 2 , H 2 O, and O 2 using a phase-step function of amplitude 3/4π that is scanned across the spectrum of the pulse. In addition, we changed the amount of chirp in the pulses. Semiclassical as well as fully quantum mechanical time-dependent Schrödinger equation calculations are found to be in excellent agreement with experimental results. We find that precise knowledge of the field parameters in the time and frequency domains is essential to afford reproducible results and quantitative theory and experiment comparisons.

74 ATOMIC AND MOLECULAR PHYSICS↗

Localization and coherent imaging of hidden moving objects using laser speckle

Imaging and sensing of moving objects through opaque scattering media is a challenging but important problem in a variety of applications, including environmental sensing, biomedical imaging, and material inspection. We have previously demonstrated a technique to coherently image a moving object through thick, heavily scattering random media using correlations of speckle images as a function of the object’s spatial translation. Here, we demonstrate that this technique can be combined with localization to achieve imaging without prior knowledge of the object’s motion, greatly extending the application domain. This method is effective beyond the thin or weakly scattering regime and, rather than motion being deleterious, exploits the information available when the hidden object is moving, as could be the case in a cluttered terrestrial environment or through substantial levels of biological tissue scatter.

Hastings, Ryan L. (ORCID:0009000095977807)↗

Large Language Models for the Creation and Use of Semantic Ontologies in Buildings: Requirements and Challenges

Semantic ontologies offer a formalized, machine-readable framework for representing knowledge, enabling the structured description of complex systems. In the building domain, the adoption of ontologies like the Brick schema has transformed how buildings and their systems are modeled by providing a standardized, interoperable language. However, the complexity and the steep learning curve involved in developing and querying semantic models present substantial challenges, often requiring a workforce with specialized expertise. This paper builds on our experience in investigating how Large Language Models (LLMs) can help address these challenges, focusing on their role in constructing and querying of semantic models, particularly using the Brick Schema. Our study outlines the requirements and metrics for evaluating the scalability and effectiveness of LLM-based tools, while also discussing the current challenges and limitations in developing such tools. Ultimately, this paper aims to orient research efforts as various groups experiment with diverse techniques, while enabling more effective comparison of emerging solutions and fostering collaboration across the field.

Mulayim, Ozan Baris↗

Application of Accelerator Technology to Quantum Information Science

The intersection of accelerator and quantum information science (QIS) offers a unique platform to advance both fields through shared technology and infrastructure. This talk will discuss the synergies which exist between these two vastly different but complementary domains. We demonstrate how we leverage pre-existing infrastructure and knowledge to perform research and development which helps to realize dramatic improvements in both 10 km long accelerators and 10 cm large quantum processors. We will explore niobium superconducting radio-frequency (SRF) cavities, a highly advanced technology that excels in efficiently storing electromagnetic energy, enabling ultra-long photon lifetimes critical for quantum processors and facilitating the characterization of quantum materials with parts-per-billion precision. We will also discuss how advancements in superconducting materials, cryogenic systems, and control techniques help to reduce cost and improve performance for both quantum systems and particle accelerators. Moreover, we will discuss cross-disciplinary applications such as dark-matter searches and demonstrate the convergence of these fields in addressing fundamental scientific questions.

Bafia, Daniel [Fermilab]↗

Grid Architecture Mapping to Understand Transformation (GAMUT): Methods and Framework Architecture

Grid architecture (GA) is a concept that was developed to address the need for a comprehensive view of power grid challenges. GA can be viewed as a relatively consistent and fixed high-level approach; however, for any instantiation of grid structures, a combinatorial explosion results from each lower-layer expansion. This constitutes the main challenge with GA—it is a grid architect’s view of the system, which might not be very informative at the implementation level. Grid Architecture Mapping to Understand Transformation (GAMUT project) seeks to bridge that gap by integrating subject matter expertise across GA structures, providing users who lack expertise in GA approaches with valuable insights and informational materials. GAMUT seeks to answer feasibility questions for the approach. System-level expectations are that a GA baseline needs to be established in order for GA to be the common framework to which any lower layer approach is tied. This report explores a potential information ingestion and documentation framework to support GAMUT. The main concepts that enable the solution domain of GAMUT are discussed, and examples are provided. The solution domain leverages already-existing technology and concepts related to GA, knowledge management, and other relevant areas. To assess GAMUT building blocks and the overall approach, a feasibility assessment is proposed, rooted in systems engineering and GA architecture evaluation concepts.

24 POWER TRANSMISSION AND DISTRIBUTION↗

The Impact of Individual Traits on Domain Task Performance: Exploring the Dunning-Kruger Effect

Research shows that individuals often overestimate their knowledge and performance without realizing they have done so, which can lead to faulty technical outcomes. This phenomenon is known as the Dunning-Kruger effect (Kruger & Dunning, 1999). This research sought to determine if some individuals were more prone to overestimating their performance due to underlying personality and cognitive characteristics. To test our hypothesis, we first collected individual difference measures. Next, we asked participants to estimate their performance on three performance tasks to assess the likelihood of overestimation. We found that some individuals may be more prone to overestimating their performance than others, and that faulty problem-solving abilities and low skill may be to blame. Encouraging individuals to think critically through all options and to consult with others before making a high-consequence decision may reduce overestimation.

42 ENGINEERING↗

NukeLM: Pre-Trained and Fine-Tuned Language Models for the Nuclear and Energy Domains

Natural language processing (NLP) tasks (text classification, named entity recognition, etc.) have seen amazing improvements over the last few years. This is due to models such as BERT that achieve deep knowledge transfer by using a large pre-trained model, then fine-tuning the model on specific tasks. The BERT architecture has shown even better performance on domain-specific tasks when the model is pre-trained using domain-relevant texts. Here, inspired by these recent advancements, we have developed NukeLM, a nuclear-domain BERT model pre-trained on 1.5 million abstracts from the DOE Office of Scientific and Technical Information (OSTI) database. This NukeLM model is then fine-tuned for the classification of research articles into either binary classes (related to the nuclear fuel cycle (NFC) or not) or multiple categories related to the subject of the article. We show that continued pre-training of a BERT-style architecture prior to fine-tuning results in greater performance in both article classification tasks. This information is critical for properly triaging manuscripts, a necessary task for better understanding citation networks that publish in the nuclear space and uncovering new areas of research in the nuclear (or nuclear relevant) domain.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

The first DEP domain of the RhoGEF P-Rex1 autoinhibits activity and contributes to membrane binding

Phosphatidylinositol (3,4,5)-trisphosphate (PIP 3 )-dependent Rac exchanger 1 (P-Rex1) catalyzes the exchange of GDP for GTP on Rac GTPases, thereby triggering changes in the actin cytoskeleton and in transcription. Its overexpression is highly correlated with the metastasis of certain cancers. P-Rex1 recruitment to the plasma membrane and its activity are regulated via interactions with heterotrimeric Gβγ subunits, PIP 3 , and protein kinase A (PKA). Deletion analysis has further shown that domains C-terminal to its catalytic Dbl homology (DH) domain confer autoinhibition. Among these, the first dishevelled, Egl-10, and pleckstrin domain (DEP1) remains to be structurally characterized. DEP1 also harbors the primary PKA phosphorylation site, suggesting that an improved understanding of this region could substantially increase our knowledge of P-Rex1 signaling and open the door to new selective chemotherapeutics. Here we show that the DEP1 domain alone can autoinhibit activity in context of the DH/PH-DEP1 fragment of P-Rex1 and interacts with the DH/PH domains in solution. The 3.1 Å crystal structure of DEP1 features a domain swap, similar to that observed previously in the Dvl2 DEP domain, involving an exposed basic loop that contains the PKA site. Using purified proteins, we show that although DEP1 phosphorylation has no effect on the activity or solution conformation of the DH/PH-DEP1 fragment, it inhibits binding of the DEP1 domain to liposomes containing phosphatidic acid. Thus, we propose that PKA phosphorylation of the DEP1 domain hampers P-Rex1 binding to negatively charged membranes in cells, freeing the DEP1 domain to associate with and inhibit the DH/PH module.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Geometrically Matched Multi-source Microscopic Image Synthesis Using Bidirectional Adversarial Networks

Microscopic images from multiple modalities can produce plentiful experimental information. In practice, biological or physical constraints under a given observation period may prevent researchers from acquiring enough microscopic scanning. Recent studies demonstrate that image synthesis is one of the popular approaches to release such constraints. Nonetheless, most existing synthesis approaches only translate images from the source domain to the target domain without solid geometric associations. To embrace this challenge, we propose an innovative model architecture, BANIS, to synthesize diversified microscopic images from multi-source domains with distinct geometric features. The experimental outcomes indicate that BANIS successfully synthesizes favorable image pairs on C. elegans microscopy embryonic images. To the best of our knowledge, BANIS is the first application to synthesize microscopic images that associate distinct spatial geometric features from multi-source domains.

Wang, Dali↗

Decoding the protein–ligand interactions using parallel graph neural networks

Abstract Protein–ligand interactions (PLIs) are essential for biochemical functionality and their identification is crucial for estimating biophysical properties for rational therapeutic design. Currently, experimental characterization of these properties is the most accurate method, however, this is very time-consuming and labor-intensive. A number of computational methods have been developed in this context but most of the existing PLI prediction heavily depends on 2D protein sequence data. Here, we present a novel parallel graph neural network (GNN) to integrate knowledge representation and reasoning for PLI prediction to perform deep learning guided by expert knowledge and informed by 3D structural data. We develop two distinct GNN architectures: $$\hbox {GNN}_{\mathrm{F}}$$ GNN F is the base implementation that employs distinct featurization to enhance domain-awareness, while $$\hbox {GNN}_{\mathrm{P}}$$ GNN P is a novel implementation that can predict with no prior knowledge of the intermolecular interactions. The comprehensive evaluation demonstrated that GNN can successfully capture the binary interactions between ligand and protein’s 3D structure with 0.979 test accuracy for $$\hbox {GNN}_{\mathrm{F}}$$ GNN F and 0.958 for $$\hbox {GNN}_{\mathrm{P}}$$ GNN P for predicting activity of a protein–ligand complex. These models are further adapted for regression tasks to predict experimental binding affinities and $$\hbox {pIC}_{\mathrm{50}}$$ pIC 50 crucial for compound’s potency and efficacy. We achieve a Pearson correlation coefficient of 0.66 and 0.65 on experimental affinity and 0.50 and 0.51 on $$\hbox {pIC}_{\mathrm{50}}$$ pIC 50 with $$\hbox {GNN}_{\mathrm{F}}$$ GNN F and $$\hbox {GNN}_{\mathrm{P}}$$ GNN P , respectively, outperforming similar 2D sequence based models. Our method can serve as an interpretable and explainable artificial intelligence (AI) tool for predicted activity, potency, and biophysical properties of lead candidates. To this end, we show the utility of $$\hbox {GNN}_{\mathrm{P}}$$ GNN P on SARS-Cov-2 protein targets by screening a large compound library and comparing the prediction with the experimentally measured data.

59 BASIC BIOLOGICAL SCIENCES↗

Clinical knowledge extraction via sparse embedding regression (KESER) with multi-center large scale electronic health record data

The increasing availability of electronic health record (EHR) systems has created enormous potential for translational research. However, it is difficult to know all the relevant codes related to a phenotype due to the large number of codes available. Traditional data mining approaches often require the use of patient-level data, which hinders the ability to share data across institutions. In this project, we demonstrate that multi-center large-scale code embeddings can be used to efficiently identify relevant features related to a disease of interest. We constructed large-scale code embeddings for a wide range of codified concepts from EHRs from two large medical centers. We developed knowledge extraction via sparse embedding regression (KESER) for feature selection and integrative network analysis. We evaluated the quality of the code embeddings and assessed the performance of KESER in feature selection for eight diseases. Besides, we developed an integrated clinical knowledge map combining embedding data from both institutions. The features selected by KESER were comprehensive compared to lists of codified data generated by domain experts. Features identified via KESER resulted in comparable performance to those built upon features selected manually or with patient-level data. The knowledge map created using an integrative analysis identified disease-disease and disease-drug pairs more accurately compared to those identified using single institution data. Analysis of code embeddings via KESER can effectively reveal clinical knowledge and infer relatedness among codified concepts. KESER bypasses the need for patient-level data in individual analyses providing a significant advance in enabling multi-center studies using EHR data.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

A Systems Biology Approach to Identify Essential Epigenetic Regulators for Specific Biological Processes in Plants

Upon sensing developmental or environmental cues, epigenetic regulators transform the chromatin landscape of a network of genes to modulate their expression and dictate adequate cellular and organismal responses. Knowledge of the specific biological processes and genomic loci controlled by each epigenetic regulator will greatly advance our understanding of epigenetic regulation in plants. To facilitate hypothesis generation and testing in this domain, we present EpiNet, an extensive gene regulatory network (GRN) featuring epigenetic regulators. EpiNet was enabled by (i) curated knowledge of epigenetic regulators involved in DNA methylation, histone modification, chromatin remodeling, and siRNA pathways; and (ii) a machine-learning network inference approach powered by a wealth of public transcriptome datasets. We applied GENIE3, a machine-learning network inference approach, to mine public Arabidopsis transcriptomes and construct tissue-specific GRNs with both epigenetic regulators and transcription factors as predictors. The resultant GRNs, named EpiNet, can now be intersected with individual transcriptomic studies on biological processes of interest to identify the most influential epigenetic regulators, as well as predicted gene targets of the epigenetic regulators. We demonstrate the validity of this approach using case studies of shoot and root apical meristem development.

root apical meristem↗

Designing workflows for materials characterization

Experimental science is enabled by the combination of synthesis, imaging, and functional characterization organized into evolving discovery loop. Synthesis of new material is typically followed by a set of characterization steps aiming to provide feedback for optimization or discover fundamental mechanisms. However, the sequence of synthesis and characterization methods and their interpretation, or research workflow, has traditionally been driven by human intuition and is highly domain specific. Here, we explore concepts of scientific workflows that emerge at the interface between theory, characterization, and imaging. In this study, we discuss the criteria by which these workflows can be constructed for special cases of multiresolution structural imaging and functional characterization, as a part of more general material synthesis workflows. Some considerations for theory–experiment workflows are provided. We further pose that the emergence of user facilities and cloud labs disrupts the classical progression from ideation, orchestration, and execution stages of workflow development. To accelerate this transition, we propose the framework for workflow design, including universal hyperlanguages describing laboratory operation, ontological domain matching, reward functions and their integration between domains, and policy development for workflow optimization. These tools will enable knowledge-based workflow optimization; enable lateral instrumental networks, sequential and parallel orchestration of characterization between dissimilar facilities; and empower distributed research.

36 MATERIALS SCIENCE↗

Function, Structure, and Regulation of Nitrogen Fixation-like Metalloproteins for Nitrogen, Energy, Carbon, and Sulfur Metabolism

Nitrogenases (N 2 ases) and nitrogen fixation-like (NFL) systems play distinct roles in nitrogen, carbon, sulfur, and energy metabolism based on their fundamental differences in structure and metallocofactor identity. As new NFL systems have recently been identified and characterized, striking parallels and differences compared to N 2 ase structure, catalysis, and regulation have emerged. NFL systems use metallocofactors that span from simple [4Fe-4S] clusters to complex clusters akin to FeMo-co, previously only thought to occur in N 2 ase. This review describes the present state of knowledge on the function, structure, catalytic mechanisms, and regulation of NFL systems that perform distinct biological roles across all three domains of life. Recent advancements in N 2 ase spectroscopic techniques for probing metallocofactor structure and electronic states guide current and future work on how each NFL system catalyzes its specific biological reaction(s). Key knowledge gaps and needed areas of research for uncovering the specific metallocofactors and structural motifs that are at the heart of NFL system reaction specificity, along with how these systems are regulated, are discussed.

Bacteria↗