Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “model queries”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

Graph Convolutional Network-Strengthened Topic Modeling for Scientific Papers

Machine learning has been woven into statistics to modernize topic modeling over textual documents written in natural language, and scientific paper search and recommendation can consequently offer higher accuracy instead of counting on traditional keyword-based search. However, topic distribution of a paper resulted from existing topic modeling techniques only relies on the statistics of words contained in the paper itself. We argue that community users’ views of a paper may also provide insights at the time of recommendation. For example, if a paper on fake image detection has been cited heavily by machine learning papers, such a feature should be absorbed in the embedding of this paper, so that it can be recommended for future query on machine learning. In this paper, we present a Graph Convolutional Network-strengthened Topic Modeling (GCN-TM) method, which employs GCN technique to refine topic modeling of scientific papers. A citation-oriented knowledge graph is constructed, and topic modeling is mapped to feature embedding of the comprising papers. On top of its own topics carried in its content, each paper learns topics from its neighbors and revise its embedding accordingly. Our empirical studies over real-life scientific literature has proved the necessity and effectiveness of our proposed approach.

Jia Zhang↗

Consist v0.1.0

A Python library for provenance tracking, intelligent caching, and data virtualization in scientific simulation workflows. It automatically records code, configuration, and input data to skip redundant computations and enables querying results across many runs without manual bookkeeping. Designed to support multi-model simulation workflows like the BEAM CORE toolset at LBL, but designed to be extensible to a wide range of research workflows. Combines lineage tracking features as provided by OpenLineage with deterministic hashing like SnakeMake, and adds powerful analysis tools on model outputs.

Needell, Zachary [Lawrence Berkeley National Labor↗

An evidential approach to problem solving when a large number of knowledge systems is available

Some recent problems are no longer formulated in terms of imprecise facts, missing data or inadequate measuring devices. Instead, questions pertaining to knowledge and information itself arise and can be phrased independently of any particular area of knowledge. The problem considered in the present work is how to model a problem solver that is trying to find the answer to some query. The problem solver has access to a large number of knowledge systems that specialize in diverse features. In this context, feature means an indicator of what the possibilities for the answer are. The knowledge systems should not be accessed more than once, in order to have truly independent sources of information. Moreover, these systems are allowed to run in parallel. Since access might be expensive, it is necessary to construct a management policy for accessing these knowledge systems. To help in the access policy, some control knowledge systems are available. Control knowledge systems have knowledge about the performance parameters status of the knowledge systems. In order to carry out the double goal of estimating what units to access and to answer the given query, diverse pieces of evidence must be fused. The Dempster-Shafer Theory of Evidence is used to pool the knowledge bases.

Dekorvin, Andre↗

TAT-C: A Trade-Space Analysis Tool for Constellations

Under a changing technological and economic environment, there is growing interest in implementing future NASA Earth Science missions as Distributed Spacecraft Missions (DSM). The objective of our project is to provide a framework that facilitates DSM Pre-Phase A investigations and optimizes DSM designs with respect to a-priori Science goals. In this first version of our Trade-space Analysis Tool for Constellations (TAT-C), we are investigating questions such as: Which type of constellations should be chosen? How many spacecraft should be included in the constellation? Which design has the best costrisk value? This paper describes the overall architecture of TAT-C including: a User Interface (UI) interacting with multiple users - scientists, missions designers or program managers; an Executive Driver gathering requirements from UI and formulating Trade-space Search Requests for the Trade-space Search Iterator, which in collaboration with the Orbit Coverage, Reduction Metrics, and Cost Risk modules generates multiple potential architectures and their associated characteristics. UI will include Graphical, Command Line and Application Programmer Interfaces to respond to the demands of various levels of users expertise. Science inputs are grouped into various mission concepts, satellite specifications, and payload specifications, while science outputs are grouped into several types of metrics - spatial, temporal, angular and radiometric. Orbit Coverage leverages the use of the Goddard Mission Analysis Tool (GMAT) to compute coverage and ancillary data that are passed to Reduction Metrics. Then, for each architecture design, Cost Risk will provide estimates of the cost and life cycle cost as well as technical and cost risk of the proposed mission. Additionally, the Knowledge Base module is a centralized store of structured data readable by humans and machines. It will support both TAT-C analysis when composing new mission concepts from existing model inputs, and TAT-C exploration when discovering new mission concepts by querying previous results.

Science Data Processing↗

Large Language Model Integration for Knowledge Retrieval and Interaction for the DUNE Experiment

The Deep Underground Neutrino Experiment (DUNE) is a next-generation neutrino experiment that will generate an unprecedented volume of heterogeneous information-from documentation and technical notes to experimental data and reconstruction pipelines. Efficient knowledge retrieval and contextual understanding are increasingly critical for collaboration-wide productivity and onboarding. In this work, we present DUNE-GPT, a prototype framework that leverages large language models (LLMs) and retrieval-augmented generation (RAG) to enable natural-language querying of DUNE's internal documentation and technical resources. The system provides an intelligent interface for DUNE collaborators to interact with experiment-specific knowledge while maintaining data privacy and infrastructure compliance within Fermilab computing resources.

Rafique, A. [Argonne (main)]↗

The Schwarz Alternating Method for the Seamless Coupling of Nonlinear Reduced Order Models and Full Order Models

Projection-based model order reduction allows for the parsimonious representation of full order models (FOMs), typically obtained through the discretization of a set of partial differential equations (PDEs) using conventional techniques (e.g., finite element, finite volume, finite difference methods) where the discretization may contain a very large number of degrees of freedom. As a result of this more compact representation, the resulting projection-based reduced order models (ROMs) can achieve considerable computational speedups, which are especially useful in real-time or multi-query analyses. One known deficiency of projection-based ROMs is that they can suffer from a lack of robustness, stability and accuracy, especially in the predictive regime, which ultimately limits their useful application. Another research gap that has prevented the widespread adoption of ROMs within the modeling and simulation community is the lack of theoretical and algorithmic foundations necessary for the “plug-and-play” integration of these models into existing multi-scale and multi-physics frameworks. This paper describes a new methodology that has the potential to address both of the aforementioned deficiencies by coupling projection-based ROMs with each other as well as with conventional FOMs by means of the Schwarz alternating method [41]. Leveraging recent work that adapted the Schwarz alternating method to enable consistent and concurrent multiscale coupling of finite element FOMs in solid mechanics [35, 36], we present a new extension of the Schwarz framework that enables FOM-ROM and ROM-ROM coupling, following a domain decomposition of the physical geometry on which a PDE is posed. In order to maintain efficiency and achieve computation speed-ups, we employ hyper-reduction via the Energy-Conserving Sampling and Weighting (ECSW) approach [13]. We evaluate the proposed coupling approach in the reproductive as well as in the predictive regime on a canonical test case that involves the dynamic propagation of a traveling wave in a nonlinear hyper-elastic material.

97 MATHEMATICS AND COMPUTING↗

A User-Centered Approach to Adaptive Hypertext Based on an Information Relevance Model

Rapid and effective to information in large electronic documentation systems can be facilitated if information relevant in an individual user's content can be automatically supplied to this user. However most of this knowledge on contextual relevance is not found within the contents of documents, it is rather established incrementally by users during information access. We propose a new model for interactively learning contextual relevance during information retrieval, and incrementally adapting retrieved information to individual user profiles. The model, called a relevance network, records the relevance of references based on user feedback for specific queries and user profiles. It also generalizes such knowledge to later derive relevant references for similar queries and profiles. The relevance network lets users filter information by context of relevance. Compared to other approaches, it does not require any prior knowledge nor training. More importantly, our approach to adaptivity is user-centered. It facilitates acceptance and understanding by users by giving them shared control over the adaptation without disturbing their primary task. Users easily control when to adapt and when to use the adapted system. Lastly, the model is independent of the particular application used to access information, and supports sharing of adaptations among users.

Mathe, Nathalie↗

Biolink Model: A universal schema for knowledge graphs in clinical, biomedical, and translational science

Abstract Within clinical, biomedical, and translational science, an increasing number of projects are adopting graphs for knowledge representation. Graph‐based data models elucidate the interconnectedness among core biomedical concepts, enable data structures to be easily updated, and support intuitive queries, visualizations, and inference algorithms. However, knowledge discovery across these “knowledge graphs” (KGs) has remained difficult. Data set heterogeneity and complexity; the proliferation of ad hoc data formats; poor compliance with guidelines on findability, accessibility, interoperability, and reusability; and, in particular, the lack of a universally accepted, open‐access model for standardization across biomedical KGs has left the task of reconciling data sources to downstream consumers. Biolink Model is an open‐source data model that can be used to formalize the relationships between data structures in translational science. It incorporates object‐oriented classification and graph‐oriented features. The core of the model is a set of hierarchical, interconnected classes (or categories) and relationships between them (or predicates) representing biomedical entities such as gene, disease, chemical, anatomic structure, and phenotype. The model provides class and edge attributes and associations that guide how entities should relate to one another. Here, we highlight the need for a standardized data model for KGs, describe Biolink Model, and compare it with other models. We demonstrate the utility of Biolink Model in various initiatives, including the Biomedical Data Translator Consortium and the Monarch Initiative, and show how it has supported easier integration and interoperability of biomedical KGs, bringing together knowledge from multiple sources and helping to realize the goals of translational science.

60 APPLIED LIFE SCIENCES↗

Data-driven wind turbine wake modeling via probabilistic machine learning

Wind farm design primarily depends on the variability of the wind turbine wake flows to the atmospheric wind conditions and the interaction between wakes. Physics-based models that capture the wake flow field with high-fidelity are computationally very expensive to perform layout optimization of wind farms, and, thus, data-driven reduced-order models can represent an efficient alternative for simulating wind farms. In this work, we use real-world light detection and ranging (LiDAR) measurements of wind-turbine wakes to construct predictive surrogate models using machine learning. Specifically, we first demonstrate the use of deep autoencoders to find a low-dimensional latent space that gives a computationally tractable approximation of the wake LiDAR measurements. Then, we learn the mapping between the parameter space and the (latent space) wake flow fields using a deep neural network. Additionally, we also demonstrate the use of a probabilistic machine learning technique, namely, Gaussian process modeling, to learn the parameter-space-latent-space mapping in addition to the epistemic and aleatoric uncertainty in the data. Finally, to cope with training large datasets, we demonstrate the use of variational Gaussian process models that provide a tractable alternative to the conventional Gaussian process models for large datasets. Furthermore, we introduce the use of active learning to adaptively build and improve a conventional Gaussian process model predictive capability. Overall, we find that our approach provides accurate approximations of the wind-turbine wake flow field that can be queried at an orders-of-magnitude cheaper cost than those generated with high-fidelity physics-based simulations.

Deep neural networks↗

Hierarchical Gaussian process-based Bayesian optimization for materials discovery in high entropy alloy spaces

Bayesian optimization (BO) is a powerful and data-efficient method for iterative materials discovery and design, particularly valuable when prior knowledge is limited, underlying functional relationships are complex or unknown, and the cost of querying the materials space is significant. Traditional BO methodologies typically utilize conventional Gaussian Processes (cGPs) to model the relationships between material inputs and properties, as well as correlations within the input space. However, cGP-BO approaches often fall short in multi-objective optimization scenarios, where they are unable to fully exploit correlations between distinct material properties. Leveraging these correlations can significantly enhance the discovery process, as information about one property can inform and improve predictions about others. Here, this study addresses this limitation by employing advanced kernel structures to capture and model multi-dimensional property correlations through multi-task (MTGPs) or deep Gaussian Processes (DGPs), thus accelerating the discovery process. We demonstrate the effectiveness of MTGP-BO and DGP-BO in rapidly and robustly solving complex materials design challenges that occur within the context of complex multi-objective optimization over FCC FeCrNiCoCu high entropy alloy (HEA) spaces, where traditional cGP-BO approaches fail. Furthermore, we highlight how the differential costs associated with querying various material properties can be strategically leveraged to make the materials discovery process more cost-efficient.

36 MATERIALS SCIENCE↗

Leveraging Structured Biological Knowledge for Counterfactual Inference: A Case Study of Viral Pathogenesis

Counterfactual inference is a useful tool for comparing outcomes of interventions on complex systems. It requires us to represent the system in form of a structural causal model, complete with a causal diagram, probabilistic assumptions on exogenous variables, and functional assignments. Specifying such models can be extremely difficult in practice. The process requires substantial domain expertise, and does not scale easily to large systems, multiple systems, or novel system modifications. At the same time, many application domains, such as molecular biology, are rich in structured causal knowledge that is qualitative in nature. This manuscript proposes a general approach for querying a causal knowledge graph with a causal question and converting the qualitative result into a quantitative structural causal model that can learn from data to answer the question. Here, we demonstrate the feasibility, accuracy and versatility of this approach using two case studies in systems biology. The first demonstrates the appropriateness of the underlying assumptions and the accuracy of the results. The second demonstrates the versatility of the approach by querying a knowledge base for the molecular determinants of a severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2)-induced cytokine storm and performing counterfactual inference to predict the causal effect of medical countermeasures for severely ill COVID-19 patients.

60 APPLIED LIFE SCIENCES↗

Discovery of multi-functional polyimides through high-throughput screening using explainable machine learning

Polyimides have been widely used in modern industries because of their excellent mechanical and thermal properties, e.g., high-temperature fuel cells, displays, and aerospace composites. However, it usually takes decades of experimental efforts to develop a successful product. Aiming to expedite the discovery of high-performance polyimides, we utilize computational methods of machine learning (ML) and molecular dynamics (MD) simulations. Our study provides compelling evidence for the effectiveness of a data-driven approach in discovering novel polyimides. We first build a comprehensive library of more than 8 million hypothetical polyimides based on the polycondensation of existing dianhydride and diamine/diisocyanate molecules. Then we establish multiple ML models for the thermal and mechanical properties of polyimides based on their experimentally reported values, including glass transition temperature, Young’s modulus, and tensile yield strength. The obtained ML models demonstrate excellent predictive performance in identifying the key chemical substructures influencing the thermal and mechanical properties of polyimides. The use of explainable machine learning describes the effect of chemical substructures on individual properties, from which human experts can understand the cause of the ML model decision. Applying the well-trained ML models, we obtain property predictions of the 8 million hypothetical polyimides. Then, we screen the whole hypothetical dataset and identify three (3) best-performing novel polyimides that have better-combined properties than existing ones through Pareto frontier analysis. For an easy query of the discovered high-performing polyimides, we also create an online platform https://polyimide-explorer.herokuapp.com/ that embeds the developed ML model with interactive visualization. Furthermore, we validate the ML predictions through all-atom MD simulations and examine their synthesizability. The MD simulations are in good agreement with the ML predictions and the three novel polyimides are predicted to be easy to synthesize via Schuffenhauer’s synthetic accessibility score. Following the proposed ML guidance, we successfully synthesized a novel polyimide and the experimentally obtained high glass transition/thermal decomposition temperature demonstrated its excellent thermal stability. Here our study demonstrates an efficient way to expedite the discovery of novel polymers using ML prediction and MD validation. The high-throughput screening of a large computational dataset can serve as a general approach for new material discovery in other polymeric material exploration problems, such as organic photovoltaics, polymer membranes, and dielectrics.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Predicting Metabolic Reaction Networks with Perturbation-Theory Machine Learning (PTML) Models

Background: Checking the connectivity (structure) of complex Metabolic Reaction Networks(MRNs) models proposed for new microorganisms with promising properties is an importantgoal for chemical biology. Objective: In principle, we can perform a hand-on checking (Manual Curation). However, this is achallenging task due to the high number of combinations of pairs of nodes (possible metabolic reactions). Results: The CPTML linear model obtained using the LDA algorithm is able to discriminate nodes(metabolites) with the correct assignation of reactions from incorrect nodes with values of accuracy,specificity, and sensitivity in the range of 85-100% in both training and external validation dataseries. Methods: In this work, we used Combinatorial Perturbation Theory and Machine Learning techniquesto seek a CPTML model for MRNs >40 organisms compiled by Barabasis’ group. First, wequantified the local structure of a very large set of nodes in each MRN using a new class of node indexcalled Markov linear indices fk. Next, we calculated CPT operators for 150000 combinationsof query and reference nodes of MRNs. Last, we used these CPT operators as inputs of differentML algorithms. Conclusion: Meanwhile, PTML models based on Bayesian network, J48-Decision Tree and RandomForest algorithms were identified as the three best non-linear models with accuracy greaterthan 97.5%. The present work opens the door to the study of MRNs of multiple organisms usingPTML models.

Pharmacology & Pharmacy↗

Conversational Grid Storage: Bridging Rucio and LLMs with Model Context Protocol

Experiments at Fermilab use Rucio to handle datasets that can be up to exabyte scale. However, navigating through Rucio’s syntax-heavy Command Line Interface (CLI) is a major workflow obstruction for researchers who just want to check quotas, track data identifiers (DIDs), or locate data sets. This project introduces a natural language interface. By building a containerized Model Context Protocol (MCP) server, an AI agent is created that translates plain English queries into data operations.

Akella, Kashyap [Fermilab; Illinois U., Urbana (ma↗

Formal Provenance Representation of the Data and Information Supporting the National Climate Assessment

The Global Change Information System (GCIS) provides a framework for the formal representation of structured metadata about data and information about global change. The pilot deployment of the system supports the National Climate Assessment (NCA), a major report of the U.S. Global Change Research Program (USGCRP). A consumer of that report can use the system to browse and explore that supporting information. Additionally, capturing that information into a structured data model and presenting it in standard formats through well defined open inter- faces, including query interfaces suitable for data mining and linking with other databases, the information becomes valuable for other analytic uses as well.

Provenance↗

A Spatiotemporal Indexing Approach for Efficient Processing of Big Array-Based Climate Data with MapReduce

Climate observations and model simulations are producing vast amounts of array-based spatiotemporal data. Efficient processing of these data is essential for assessing global challenges such as climate change, natural disasters, and diseases. This is challenging not only because of the large data volume, but also because of the intrinsic high-dimensional nature of geoscience data. To tackle this challenge, we propose a spatiotemporal indexing approach to efficiently manage and process big climate data with MapReduce in a highly scalable environment. Using this approach, big climate data are directly stored in a Hadoop Distributed File System in its original, native file format. A spatiotemporal index is built to bridge the logical array-based data model and the physical data layout, which enables fast data retrieval when performing spatiotemporal queries. Based on the index, a data-partitioning algorithm is applied to enable MapReduce to achieve high data locality, as well as balancing the workload. The proposed indexing approach is evaluated using the National Aeronautics and Space Administration (NASA) Modern-Era Retrospective Analysis for Research and Applications (MERRA) climate reanalysis dataset. The experimental results show that the index can significantly accelerate querying and processing (10 speedup compared to the baseline test using the same computing cluster), while keeping the index-to-data ratio small (0.0328). The applicability of the indexing approach is demonstrated by a climate anomaly detection deployed on a NASA Hadoop cluster. This approach is also able to support efficient processing of general array-based spatiotemporal data in various geoscience domains without special configuration on a Hadoop cluster.

big data↗

Latent space dynamics identification for interface tracking with application to shock-induced pore collapse

Capturing sharp, evolving interfaces remains a central challenge in reduced-order modeling, especially when data is limited and the system exhibits localized nonlinearities or discontinuities. Here, we propose LaSDI-IT (Latent Space Dynamics Identification for Interface Tracking), a data-driven framework that combines low-dimensional latent dynamics learning with explicit interface-aware encoding to enable accurate and efficient modeling of physical systems involving moving material boundaries. At the core of LaSDI-IT is a revised autoencoder architecture that jointly reconstructs the physical field and an indicator function representing material regions or phases, allowing the model to track complex interface evolution without requiring detailed physical models or mesh adaptation. The latent dynamics are learned through linear regression in the encoded space and generalized across parameter regimes using Gaussian process interpolation with greedy sampling. We demonstrate LaSDI-IT on the problem of shock-induced pore collapse in high explosives, a process characterized by sharp temperature gradients and dynamically deforming pore geometries. The method achieves relative prediction errors below 9% across the parameter space, accurately recovers key quantities of interest such as pore area and hot spot formation, and matches the performance of dense training with only half the data. This latent dynamics prediction was 10 6 times faster than the conventional high-fidelity simulation, proving its utility for multi-query applications. These results highlight LaSDI-IT as a general, data-efficient framework for modeling discontinuity-rich systems in computational physics, with potential applications in multiphase flows, fracture mechanics, and phase change problems.

Gaussian process↗

An information-matching approach to optimal experimental design and active learning

The efficacy of mathematical models heavily depends on the quality of the training data, yet collecting sufficient data is often expensive and challenging. Many modeling applications require inferring parameters only as a means to predict other quantities of interest (QoI). Because models often contain many unidentifiable (sloppy) parameters, QoIs often depend on a relatively small number of parameter combinations. Therefore, we introduce an information-matching criterion based on the Fisher information matrix to select the most informative training data from a candidate pool. This method ensures that the selected data contain sufficient information to learn only those parameters that are needed to constrain downstream QoIs. It is formulated as a convex optimization problem, making it scalable to large models and datasets. Here, we demonstrate the effectiveness of this approach across various modeling problems in diverse scientific fields, including power systems and underwater acoustics. Finally, we use information-matching as a query function within an active learning (AL) loop for materials science applications. In all these applications, we find that a relatively small set of optimal training data can provide the necessary information for achieving precise predictions. These results are encouraging for diverse future applications, particularly AL in large machine-learning models.

Materials science↗