Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “model queries”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

Active Learning for Anomaly Detection in Environmental data

Due to the growing amount of data from in-situ sensors in environmental monitoring, it becomes necessary to automatically detect anomalous data points. Nowadays, this is mainly performed using supervised machine learning models, which need a fully labelled data set for their training process. However, the process of labelling data is typically cumbersome and, as a result, a hindrance to the adoption of machine learning methods for automated anomaly detection. In this work, we propose to address this challenge by means of active learning. This method consists of querying the domain expert for the labels of only a selected subset of the full data set. We show that this reduces the time and costs associated to labelling while delivering the same or similar anomaly detection performances. Finally, we also show that machine learning models providing a nonlinear classification boundary are to be recommended for anomaly detection in complex environmental data sets.

54 ENVIRONMENTAL SCIENCES↗

An ontology-based knowledge graph for representing interactions involving RNA molecules

The "RNA world" represents a novel frontier for the study of fundamental biological processes and human diseases and is paving the way for the development of new drugs tailored to each patient's biomolecular characteristics. Although scientific data about coding and non-coding RNA molecules are constantly produced and available from public repositories, they are scattered across different databases and a centralized, uniform, and semantically consistent representation of the "RNA world" is still lacking. We propose RNA-KG, a knowledge graph (KG) encompassing biological knowledge about RNAs gathered from more than 60 public databases, integrating functional relationships with genes, proteins, and chemicals and ontologically grounded biomedical concepts. To develop RNA-KG, we first identified, pre-processed, and characterized each data source; next, we built a meta-graph that provides an ontological description of the KG by representing all the bio-molecular entities and medical concepts of interest in this domain, as well as the types of interactions connecting them. Finally, we leveraged an instance-based semantically abstracted knowledge model to specify the ontological alignment according to which RNA-KG was generated. RNA-KG can be downloaded in different formats and also queried by a SPARQL endpoint. A thorough topological analysis of the resulting heterogeneous graph provides further insights into the characteristics of the "RNA world". RNA-KG can be both directly explored and visualized, and/or analyzed by applying computational methods to infer bio-medical knowledge from its heterogeneous nodes and edges. The resource can be easily updated with new experimental data, and specific views of the overall KG can be extracted according to the bio-medical problem to be studied.

59 BASIC BIOLOGICAL SCIENCES↗

RAG for FLAG: AI Assistance for a Physics Code

Artificial intelligence (AI) has quickly become an important tool in scientific research, where significant efforts are underway to develop tools that will expedite the research process. One area of particular impact is scientific software, which can be particularly complex, and therefore time consuming to learn and use effectively. AI assistants are increasingly helping to streamline the process by performing tasks such as interactively answering user questions or suggesting solutions. Los Alamos National Laboratory (LANL) develops several advanced scientific codes, such as FLAG, which can be used to run multiphysics simulations. With this study, our goal was to develop an AI assistant for FLAG that could help make the process of understanding the software and running physics simulations more efficient. To develop an AI assistant for FLAG, we used a method called retrieval-augmented generation (RAG), which is a technique that uses information from relevant data sources to enhance the accuracy of large language models (LLMs). We used the FLAG user manual and other FLAG documentation as the knowledge base for the RAG system. When a user provides a query, RAG retrieves relevant sections from the knowledge base in response, then uses those excerpts to generate grounded and contextually rich answers. We found that our AI assistant was able to provide context aware answers and source references to user queries. To evaluate performance, we developed a set of 40 benchmark questions and compared the accuracy of the responses to those of two standard LLMs without retrieval. Our AI assistant significantly outperformed the standard LLMs at answering FLAG-related questions, with an 82.5% accuracy rate, compared to 47.5% for both of the standard LLMs. This has the potential to make the process of learning and using FLAG much easier, especially for new users. Ultimately, it supports LANL’s broader mission by empowering scientists and engineers to focus more on discovery and analysis rather than on navigating complex software systems.

97 MATHEMATICS AND COMPUTING↗

A Semi-Automated Approach for Curating a Glossary of Key Terms for Open-Source Data Queries

In FY20, the Savannah River National Laboratory (SRNL) was funded by the National Nuclear Security Administration’s Office of Defense Nuclear Non-Proliferation Research and Development (NA-22) to build a machine learning based modeling pipeline that could extract proliferation events of interest from open text-based data sources. As a test case, the research team targeted the identification/fusion of events and indicators that fissile core fabrication would be executed at the Savannah River Site prior to its official announcement in May of 2018. The demonstration prototype proved successful by applying natural language processing and graph theoretical techniques to identify contextual shifts in key words and phrases that acted as indicators that pit production would be carried out at the Savannah River Site up to two years prior to the official announcement.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Retrieval Augmented Generation for Robust Cyber Defense

In cybersecurity, the ability to efficiently analyze and respond to vulnerabilities, weaknesses, attack patterns, and threat tactics is critical for effective defense strategies. With the increasing complexity and volume of cybersecurity data, traditional methods of querying and retrieving information are often inadequate. To address this challenge, we implemented Retrieval-Augmented Generation (RAG) systems—CyRAG and GraphCyRAG—that integrate large language models (LLMs) with both structured data from relational databases and knowledge graphs such as Neo4j. CyRAG is designed to handle structured data, focusing on CVE (Common Vulnerabilities and Exposures) and CWE (Common Weakness Enumeration) entities to generate accurate and context-rich responses. In contrast, GraphCyRAG leverages Neo4j knowledge graphs to retrieve interconnected information from CVE, CWE, CAPEC (Common Attack Pattern Enumeration and Classification), and ATT&CK (Adversarial Tactics, Techniques, and Common Knowledge) datasets. By utilizing Neo4j’s graph-based framework, GraphCyRAG enables deeper traversal of relationships between vulnerabilities and attack patterns, providing cybersecurity analysts with more comprehensive insights into potential attack vectors and mitigation strategies. Our preliminary results demonstrate that integrating knowledge graphs with RAG significantly enhances both the accuracy and depth of threat analysis, allowing for the retrieval of dynamic, real-time data and the generation of contextually aware responses. This approach helps analysts uncover hidden relationships between cyber entities, predict exploit paths, and prioritize mitigation efforts effectively. The integration of RAG with cybersecurity knowledge graphs represents a significant advancement in cybersecurity threat intelligence, enabling more informed decision-making and stronger defense strategies.

97 MATHEMATICS AND COMPUTING↗

Inferring adversarial behaviour in cyber‐physical power systems using a Bayesian attack graph approach

Abstract Highly connected smart power systems are subject to increasing vulnerabilities and adversarial threats. Defenders need to proactively identify and defend new high‐risk access paths of cyber intruders that target grid resilience. However, cyber‐physical risk analysis and defense in power systems often requires making assumptions on adversary behaviour, and these assumptions can be wrong. Thus, this work examines the problem of inferring adversary behaviour in power systems to improve risk‐based defense and detection. To achieve this, a Bayesian approach for inference of the Cyber‐Adversarial Power System (Bayes‐CAPS) is proposed that uses Bayesian networks (BNs) to define and solve the inference problem of adversarial movement in the grid infrastructure towards targets of physical impact. Specifically, BNs are used to compute conditional probabilities to queries, such as the probability of observing an event given a set of alerts. Bayes‐CAPS builds initial Bayesian attack graphs for realistic power system cyber‐physical models. These models are adaptable using collected data from the system under study. Then, Bayes‐CAPS computes the posterior probabilities of the occurrence of a security breach event in power systems. Experiments are conducted that evaluate algorithms based on time complexity, accuracy and impact of evidence for different scales and densities of network. The performance is evaluated and compared for five realistic cyber‐physical power system models of increasing size and complexities ranging from 8 to 300 substations based on computation and accuracy impacts.

Sahu, Abhijeet↗

Aromatic amino acid metabolism and active transport regulation are implicated in microbial persistence in fractured shale reservoirs

Abstract Hydraulic fracturing has unlocked vast amounts of hydrocarbons trapped within unconventional shale formations. This large-scale engineering approach inadvertently introduces microorganisms into the hydrocarbon reservoir, allowing them to inhabit a new physical space and thrive in the unique biogeochemical resources present in the environment. Advancing our fundamental understanding of microbial growth and physiology in this extreme subsurface environment is critical to improving biofouling control efficacy and maximizing opportunities for beneficial natural resource exploitation. Here, we used metaproteomics and exometabolomics to investigate the biochemical mechanisms underpinning the adaptation of model bacterium Halanaerobium congolense WG10 and mixed microbial consortia enriched from shale-produced fluids to hypersalinity and very low reservoir flow rates (metabolic stress). We also queried the metabolic foundation for biofilm formation in this system, a major impediment to subsurface energy exploration. For the first time, we report that H. congolense WG10 accumulates tyrosine for osmoprotection, an indication of the flexible robustness of stress tolerance that enables its long-term persistence in fractured shale environments. We also identified aromatic amino acid synthesis and cell wall maintenance as critical to biofilm formation. Finally, regulation of transmembrane transport is key to metabolic stress adaptation in shale bacteria under very low well flow rates. These results provide unique insights that enable better management of hydraulically fractured shale systems, for more efficient and sustainable energy extraction.

04 OIL SHALES AND TAR SANDS↗

WA-IsoC_MSC1.1.0

The soil microbiome is central to the cycling of carbon and other nutrients and to the promotion of plant growth. Despite its importance, analysis of the soil microbiome is difficult due to its sheer complexity, with thousands of interacting species. Here, we reduced this complexity by developing model soil microbial consortia that are simpler and more amenable to experimental analysis but still represent important microbial functions of the native soil ecosystem. Samples were collected from an arid grassland soil and microbial communities (consisting mainly of bacterial species) were enriched on agar plates containing chitin as the main carbon source. Chitin was chosen because it is an abundant carbon and nitrogen polymer in soil that often requires the coordinated action of several microorganisms for complete metabolic degradation. Several soil consortia were derived that had tractable richness (30-50 OTUs) with diverse phyla representative of the native soil, including Actinobacteria, Bacteroidetes, Firmicutes, Proteobacteria and Verrucomicrobia. The resulting consortia can be stored as glycerol or lyophilized stocks at -80° C and revived while retaining community composition, greatly increasing their use as tools for the research community at large. One of the consortia that was particularly stable was chosen as a model soil consortium (MSC-1) for further analysis. MSC-1 species interactions were studied using both pairwise co-cultivation in liquid media and during growth in soil under several perturbations. Co-abundance analyses highlighted interspecies interactions and helped to define keystone species, including Mycobacterium, Rhodococcus, and Rhizobiales taxa. These experiments demonstrate the success of an approach based on naturally enriching a community of interacting species that can be stored, revived, and shared. The knowledge gained from querying these communities and their interactions will enable better understanding of the soil microbiome and the roles these interactions play in this environment.

54 ENVIRONMENTAL SCIENCES↗

Large Language Models (LLMs) for Energy Systems Research

The integration of Large Language Models (LLMs) in energy systems research promises transformative results, as demonstrated in this work, particularly in the realms of information retrieval and legal document analysis. We have developed a chat-based interface, specifically designed to query an extensive corpus of technical reports from the National Renewable Energy Laboratory (NREL). This interface capitalizes on the natural language processing capabilities of LLMs, providing future consumers of NREL research with a user-friendly platform to access and extract valuable information from technical documents, thus enhancing the dissemination of research to the public. In addition to information retrieval, we have employed LLMs to extract renewable energy siting ordinances from a variety of legal documents, a task traditionally driven by significant human labor. This automated extraction not only supports the ongoing development of the high-impact NREL siting ordinance database but also ensures the database's accuracy and comprehensiveness. Crucially, we have augmented the performance of LLMs through the integration of a decision tree framework, resulting in a substantial improvement in extraction accuracy. Comparative analysis with manual efforts has shown that this approach not only rivals but also significantly surpasses human accuracy, heralding increased reliability in legal document analysis for energy systems research. To democratize access to these advancements and foster collaborative research, we introduce the "Energy Language Model" (ELM), an open-source software package. ELM encapsulates the methodologies and tools developed in this work, providing researchers and practitioners with a robust toolkit to conduct similar analyses within their respective domains. Through these contributions, this work underscores the immense potential of LLMs in revolutionizing energy systems research, improving accuracy, efficiency, and accessibility in the field.

automation↗

EI_MS_ML

The unambiguous identification of compounds from their electron ionization mass (EI-MS) spectra remains a significant unsolved problem in the field of metabolomics and analytical chemistry as a whole. Typically EI-MS spectra are compared using various mathematical operations that convert the spectral similarity or differences into a distance-like metric that roughly approximates the similarity of any two spectra. A commonly used metric for this is the cosine similarity metric which has values close to one for very similar spectra and a value of zero for very dissimilar spectra; however, no metric is perfect. Due to the prevalence of structurally-similar compounds such as isomers and the prevalence of certain fragmentation patterns across structurally-dissimilar compounds, the unambiguous assignment of EI-MS spectra compounds remains difficult. Frequently, querying an observed EI-MS spectrum against a large database such as the NIST17 library yields multiple possible assignments requiring the end user to distinguish between multiple high scoring hits, or multiple low scoring hits while keeping in mind that the correct hit may not be in the database at all. Although techniques such as orthogonal information from techniques such as chromatography can greatly aid in unambiguous assignment, this also requires more complicated experimental designs and access to more complicated analytical instrumentation. Substructures can be trivially detected and represented as strings using a previously published technique called node coloring from a known chemical structure. However, for experimentally-derived EI-MS spectra this information must be derived from the spectra itself (i.e., because we do not know what compound it represents). To achieve this, the software uses techniques from the field of machine learning and a large training dataset of EI-MS spectra corresponding to known structures annotated with substructure strings, to build models that can predict the presence of a given chemical substructure from an EI-MS spectrum directly.If these predictions are of high-quality (i.e., are unlikely to be false positives), the presence of one or more predicted substructures can be used to constrain the number of possible hits for a query spectrum. Mathematically, this restriction could be expressed in many forms, but the most straight-forward implementation is to weight the cosine similarity of a query spectrum and a plausible database match with a Tanimoto-like coefficient based on the ratio of the number of substructures predicted to the number of substructures present in the potential database hit. Determining which combination of models best reduces assignment ambiguity will be achieved using a combination of manual curation and optimization techniques such as genetic algorithms. This software will perform all the steps necessary to construct said models from a training dataset and evaluate them using a holdout dataset. Various statistical analyses can be performed to determine if this approach does decrease assignment ambiguity. For example, if this approach works, on average, the rank-order of the correct assignment for the holdout set of EI-MS spectra should decrease and the weighted cosine similarities for most of the possible matches in the database should be better than the unweighted cosine similarities. Furthermore, this same pipeline can be used on real experimental data to generate less ambiguous assignments.

Mitchell, Joshua↗

Track Seeding and Labelling with Embedded-space Graph Neural Networks

To address the unprecedented scale of HL-LHC data, the Exa.TrkX project is investigating a variety of machine learning approaches to particle track reconstruction. The most promising of these solutions, graph neural networks (GNN), process the event as a graph that connects track measurements (detector hits corresponding to nodes) with candidate line segments between the hits (corresponding to edges). Detector information can be associated with nodes and edges, enabling a GNN to propagate the embedded parameters around the graph and predict node-, edge- and graph-level observables. Previously, message-passing GNNs have shown success in predicting doublet likelihood, and we here report updates on the state-of-the-art architectures for this task. In addition, the Exa.TrkX project has investigated innovations in both graph construction, and embedded representations, in an effort to achieve fully learned end-to-end track finding. Hence, we present a suite of extensions to the original model, with encouraging results for hitgraph classification. In addition, we explore increased performance by constructing graphs from learned representations which contain non-linear metric structure, allowing for efficient clustering and neighborhood queries of data points. We demonstrate how this framework fits in with both traditional clustering pipelines, and GNN approaches. The embedded graphs feed into high-accuracy doublet and triplet classifiers, or can be used as an end-to-end track classifier by clustering in an embedded space. A set of post-processing methods improve performance with knowledge of the detector physics. Finally, we present numerical results on the TrackML particle tracking challenge dataset, where our framework shows favorable results in both seeding and track finding.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Active Learning A Neural Network Model For Gold Clusters & Bulk From Sparse First Principles Training Data

Small metal clusters are of fundamental scientific interest and of tremendous significance in catalysis. These nanoscale clusters display diverse geometries and structural motifs depending on the cluster size; a knowledge of this size-dependent structural motifs and their dynamical evolution has been of longstanding interest. Given the high computational cost of first-principles calculations, molecular modeling and atomistic simulations such as molecular dynamics (MD) has proven to be an important complementary tool to aid this understanding. Classical MD typically employ predefined functional forms which limits their ability to capture such complex size-dependent structural and dynamical transformation. Neural Network (NN) based potentials represent flexible alternatives and in principle, well-trained NN potentials can provide high level of flexibility, transferability and accuracy on-par with the reference model used for training. A major challenge, however, is that NN models are interpolative and requires large quantities (similar to 10 4 or greater) of training data to ensure that the model adequately samples the energy landscape both near and far-from-equilibrium. A highly desirable goal is minimize the number of training data, especially if the underlying reference model is first-principles based and hence expensive. In this work, we introduce an active learning (AL) scheme that trains a NN model on-the-fly with minimal amount of first-principles based training data. Our AL workflow is initiated with a sparse training dataset (similar to 1 to 5 data points) and is updated on-the-fly via a Nested Ensemble Monte Carlo scheme that iteratively queries the energy landscape in regions of failure and updates the training pool to improve the network performance. Using a representative system of gold clusters, we demonstrate that our AL workflow can train a NN with similar to 500 total reference calculations. Using an extensive DFT test set of similar to 1100 configurations, we show that our AL-NN is able to accurately predict both the DFT energies and the forces for clusters of a myriad of different sizes. Our NN predictions are within 30 meV/atom and 40 meV/angstrom of the reference DFT calculations. Moreover, our AL-NN model also adequately captures the various size-dependent structural and dynamical properties of gold clusters in excellent agreement with DFT calculations and available experiments. We finally show that our AL-NN model also captures bulk properties reasonably well, even though they were not included in the training data.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Gradient-based constrained optimization using a database of linear reduced-order models

A methodology grounded in model reduction is presented for accelerating the gradient-based solution of a family of linear or nonlinear constrained optimization problems where the constraints include at least one linear Partial Differential Equation (PDE). A key component of this methodology is the construction, during an offline phase, of a database of pointwise, linear, Projection-based Reduced-Order Models (PROM)s associated with a design parameter space and the linear PDE(s). A parameter sampling procedure based on an appropriate saturation assumption is proposed to maximize the efficiency of such a database of PROMs. A real-time method is also presented for interpolating at any queried but unsampled parameter vector in the design parameter space the relevant sensitivities of a PROM. The practical feasibility, computational advantages, and performance of the proposed methodology are demonstrated for several realistic, nonlinear, aerodynamic shape optimization problems governed by linear aeroelastic constraints.

97 MATHEMATICS AND COMPUTING↗

dsgrid (Demand-Side Grid Toolkit) [SWR-21-52]

The demand-side grid model harnesses decades of sector-specific energy modeling expertise to understand current and future U.S. electricity load for power systems analyses. The API provides tools for submitting datasets to dsgrid projects, validating individual datasets and project data as a whole, querying, analysis, and visualization.

Hale, Elaine↗

Swap Path Network for Robust Person Search Pre-training

This code corresponds to the WACV25 conference paper, "Swap Path Network for Robust Person Search Pre-training". In that paper, we introduce a new model for the person search task called the Swap Path Net (SPNet). The person search task is a problem in computer vision, where we locate and rank matches to an image of a query person in a set of other images where we want to find them. We also introduce a novel pre-training algorithm specific to the Swap Path Net architecture. The code implements pre-training and fine-tuning of the Swap Path Net (SPNet). This includes ingesting image datasets and updating the weights of the SPNet neural network to train it for the person search task. The repository contains code, configs, and instructions to reproduce all results from the paper.

Jaffe, LucasW [Lawrence Livermore National Laborat↗

Updated resources for exploring experimentally-determined PDB structures and Computed Structure Models at the RCSB Protein Data Bank

The Research Collaboratory for Structural Bioinformatics Protein Data Bank (RCSB PDB, RCSB.org), the US Worldwide Protein Data Bank (wwPDB, wwPDB.org) data center for the global PDB archive, provides access to the PDB data via its RCSB.org research-focused web portal. We report substantial additions to the tools and visualization features available at RCSB.org, which now delivers more than 227000 experimentally determined atomic-level three-dimensional (3D) biostructures stored in the global PDB archive alongside more than 1 million Computed Structure Models (CSMs) of proteins (including models for human, model organisms, select human pathogens, crop plants and organisms important for addressing climate change). In addition to providing support for 3D structure motif searches with user-provided coordinates, new features highlighted herein include query results organized by redundancy-reduced Groups and summary pages that facilitate exploration of groups of similar proteins. Newly released programmatic tools are also described, as are enhanced training opportunities.

Burley, Stephen K.↗

Optimal adjustment sets for causal query estimation in partially observed biomolecular networks

Abstract Causal query estimation in biomolecular networks commonly selects a ‘valid adjustment set’, i.e. a subset of network variables that eliminates the bias of the estimator. A same query may have multiple valid adjustment sets, each with a different variance. When networks are partially observed, current methods use graph-based criteria to find an adjustment set that minimizes asymptotic variance. Unfortunately, many models that share the same graph topology, and therefore same functional dependencies, may differ in the processes that generate the observational data. In these cases, the topology-based criteria fail to distinguish the variances of the adjustment sets. This deficiency can lead to sub-optimal adjustment sets, and to miss-characterization of the effect of the intervention. We propose an approach for deriving ‘optimal adjustment sets’ that takes into account the nature of the data, bias and finite-sample variance of the estimator, and cost. It empirically learns the data generating processes from historical experimental data, and characterizes the properties of the estimators by simulation. We demonstrate the utility of the proposed approach in four biomolecular Case studies with different topologies and different data generation processes. The implementation and reproducible Case studies are at https://github.com/srtaheri/OptimalAdjustmentSet.

59 BASIC BIOLOGICAL SCIENCES↗

Text Mining for Process–Structure–Properties Relationships in Metals

With the advent of large language models (LLMs), the vast unstructured text within millions of academic papers is increasingly accessible for materials discovery—although significant challenges remain. While LLMs offer promising few- and zero-shot learning capabilities, particularly valuable in the materials domain where expert annotations are scarce, general-purpose LLMs often fail to address key materials-specific queries without further adaptation. To bridge this gap, fine-tuning LLMs on human-labeled data is essential for effective structured knowledge extraction (Liu in The Importance of Human-Labeled Data in the Era of LLMs, 2023). Here, in this study, we introduce a novel annotation schema designed to extract generic process–structure–properties relationships from scientific literature. We demonstrate the utility of this approach using a dataset of 128 abstracts, with annotations drawn from two distinct domains: high-temperature materials (Domain I) and uncertainty quantification in simulating materials microstructure (Domain II). Initially, we developed a conditional random field (CRF) model based on MatBERT—a domain-specific BERT variant—and evaluated its performance on Domain I. Subsequently, we compared this model with a fine-tuned LLM (GPT-4o from OpenAI) under identical conditions. Our results indicate that fine-tuning LLMs can significantly improve entity extraction performance over the BERT-CRF baseline on Domain I. However, when additional examples from Domain II were incorporated, the performance of the BERT-CRF model became comparable to that of the GPT-4o model. These findings underscore the potential of our schema for structured knowledge extraction and highlight the complementary strengths of both modeling approaches.

Materials science↗