Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “model queries”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Optimization on Manifolds via Graph Gaussian Processes

This paper integrates manifold learning techniques within a Gaussian process upper confidence bound algorithm to optimize an objective function on a manifold. Our approach is motivated by applications where a full representation of the manifold is not available and querying the objective is expensive. We rely on a point cloud of manifold samples to define a graph Gaussian process surrogate model for the objective. Query points are sequentially chosen using the posterior distribution of the surrogate model given all previous queries. We establish regret bounds in terms of the number of queries and the size of the point cloud. Several numerical examples complement the theory and illustrate the performance of our method.

Bayesian optimization↗

Dimensional Control over Metal Halide Perovskite Crystallization Guided by Active Learning

Metal halide perovskite (MHP) derivatives, a promising class of optoelectronic materials, have been synthesized with a range of dimensionalities that govern their optoelectronic properties and determine their applications. We demonstrate a data-driven approach combining active learning and high-throughput experimentation to discover, control, and understand the formation of phases with different dimensionalities in the morpholinium (morph) lead iodide system. Using a robot-assisted workflow, we synthesized and characterized two novel MHP derivatives that have distinct optical properties: a one-dimensional (1D) morphPbI 3 phase ([C 4 H 10 NO][PbI 3 ]) and a two-dimensional (2D) (morph) 2 PbI 4 phase ([C 4 H 10 NO] 2 [PbI 4 ]). To efficiently acquire the data needed to construct a machine learning (ML) model of the reaction conditions where the 1D and 2D phases are formed, data acquisition was guided by a diverse-mini-batch-sampling active learning algorithm, using prediction confidence as a stopping criterion. Querying the ML model uncovered the reaction parameters that have the most significant effects on dimensionality control. Based on these insights, we discuss possible reaction schemes that may selectively promote the formation of morph-Pb-I phases with different dimensionalities. The data-driven approach presented here, including the use of additives to manipulate dimensionality, will be valuable for controlling the crystallization of a range of materials over large reaction-composition spaces.

36 MATERIALS SCIENCE↗

Role of Uncertainty Quantification in the Explainability of Large Language Models for the Nuclear Industry

The meteoric rise of generative artificial intelligence (AI) large language models (LLMs) has created an opportunity to utilize them to increase efficiencies in a multitude of industries. While LLMs carry great potential to revolutionize the manner in which work is performed, numerous known deficiencies limit their utility, including the black box nature of the models, the stochastic nature of the response (i.e., presenting the same prompt multiple times results in different responses), and the potential for hallucination. Widespread adoption of LLMs in safety-critical industries such as nuclear will require some form of explainability to assure end users that the LLM’s response to a given query is valid. Model uncertainty is inherently linked to the concepts of trust and explainability, and can be used to identify situations in which the model is insufficiently certain about its answer. Although uncertainty is not enough in and of itself to determine the suitability of an answer—a model can be very certain of an inaccurate answer—it still provides valuable supporting information. Practical methodologies for gauging or quantifying the uncertainty in LLM outputs are presented herein, along with examples based on nuclear-specific prompts.

22 - GENERAL STUDIES OF NUCLEAR REACTORS↗

Exploring scenarios for enhanced fuel compression and performance on the National Ignition Facility with machine-learning-aided design techniques

Recent fusion experiments on the National Ignition Facility (NIF) have achieved ignition, producing multi-MJ fusion yields for input laser energies of roughly 2 MJ [Abu-Shawareb et al., Phys. Rev. Lett. 132, 065102 (2024)]. Building on the success of the target designs that have achieved ignition, we explore new implosion scenarios predicted to generate significantly more compression of the dense DT ice layer and correspondingly higher yields while preserving many of the key physics characteristics of present-day ignition designs. Our main result is a novel 3-shock implosion scheme that effectively minimizes the shock-induced entropy in the dense, accelerating DT shell and maximizes the resulting fuel compression subject to a fixed leading shock strength consistent with present-day ignition experiments, which is necessary to melt the crystalline high-density carbon ablator. Compared to the first NIF experiment to fulfill Lawson's ignition criterion, shot N210808 [Abu-Shawareb et al., Phys. Rev. Lett. 129, 075001 (2022)], our design exhibits a 40% increase in simulated peak areal density (ρR) and a 5× increase in 1D fusion yield using a 4% lighter ablator and identical DT payloads. We also present a complete integrated 2D hohlraum design and laser pulse specifications capable of generating the desired 3-shock drive and maintaining control of the low-mode capsule implosion symmetry, where the increase in simulated 2D yield relative to N210808 is > 10×. This new implosion regime was discovered with help from a machine-learning-enabled capsule design optimization framework. We outline the workflow this automated tool uses to identify improved design candidates by running several rounds of capsule simulations, constructing a surrogate model mapping input variations to key physics output quantities, and querying the resulting statistical model to propose adjustments to the x-ray drive and capsule to reach a set of physics objectives prescribed by the designer.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Differentially Private Map Matching (DPMM) v1.0

Human mobility trajectories provide valuable information for developing mobility applications, as they contain diverse and rich information about the users. User mobility data is valuable for various applications such as intelligent transportation systems (ITS), commercial business models, and disease-spread models. However, such spatio-temporal traces may pose a threat to user privacy. GPS trajectories in their raw form are not suitable for transportation studies, as they require matching locations with nearest road links — a process called map-matching. This software implements a differential privacy (DP)-based map-matching algorithm, called DPMM, that generates link-level location trajectories in a privacy-preserving manner to protect users' origin destinations (OD) and travel paths. OD privacy is achieved by injecting Planar Laplace noise to the user OD GPS points. Travel-path privacy is provided with randomized travel path construction using exponential DP mechanism. The injected noise level is selected adaptively, by considering the link density of the location and the functional category of the localized links. For path privacy, our mechanism samples waypoints and selects candidate paths between waypoints. DPMM provides privacy effectively with respect to link density instead of other trajectory samples in the database compared to other privacy mechanisms. Compared to the different baseline models our DP-based privacy model offers closer query responses to the raw data in terms of individual and aggregate trajectory-level statistics with an average at absolute deviation from the baseline for individual statistics on ϵ = 1.0. Beyond individual trajectory statistics, the DPMM outperforms the other benchmark DP-based mechanisms on different aggregate statistics with up to 8x improvement in utility.

Peisert, Sean [Lawrence Berkeley National Laborato↗

Scalable and Energy-Efficient Methods for Interactive Exploration of Scientific Data

The main scientific contributions of this project are the following novel concepts for multidimensional arrays: shape-based similarity join (SIGMOD 2016), incremental view maintenance (SIGMOD 2017), user-defined stencil functions (HPDC 2017), and distributed caching for in-situ processing (SSDBM 2018). Building on our collaboration with the astrophysics group at LBNL, we applied these techniques to the data generated in the Palomar Transient Factory (PTF) astronomical survey. They played a pivotal role in the first-ever observation of a neutron star merger, which produces gravitational waves and turns out to be the origin of heavy elements, including gold. This has lead to a Science magazine article that has received extensive media coverage on ACM TechNews, Slashdot, FiveThirtyEight, and Quanta Magazine, among others. Additionally, two other articles detailing related aspects of the same discovery have been published in the Astrophysical Journal Letters journal. These publications have more than 3,000 citations according to Google Scholar (as of February 2022). This cross-disciplinary collaboration provided very good opportunities to apply database techniques to real-life scientific problems. The fact that they facilitated major discoveries in astrophysics proves the importance of our research. In addition to the work on multidimensional array databases, this project has also developed stochastic gradient descent (SGD) optimization algorithms for training large scale machine learning models, methods for querying in-situ data, and a database query optimizer based on sketch synopses.

79 ASTRONOMY AND ASTROPHYSICS↗

I Can’t Read All That! Improving the Usability of Semantic Models Using Concise, Ontology-Agnostic, Building-Specific Schemas

Semantic ontologies have enabled the creation of formalized, machine-readable descriptions of heterogenous building systems by providing dictionaries of well defined concepts that can be applied to model them. Within a semantic model of a particular building, a subset of an ontology's concepts may be applied in different ways to represent a particular perspective of the building's systems. How the concepts were applied can only be understood by examining the large amount of instance data within a semantic model, which leads to usability challenges. We propose a concise, ontology-agnostic method for defining building-specific schema (b-schema) graphs that summarize the structure and content of a semantic model. This approach provides a queryable and concise representation of the model's contents, separate from the instance data within a model, that can mitigate the challenges posed by the size and complexity of semantic models in processes such as visualization, querying, validation, and the use of large language models (LLMs). We validate our approach on semantic models based on the Brick and ASHRAE S223 ontologies. Results demonstrate that b-schemas significantly reduce the complexity of visual interpretation, accelerate SPARQL queries and SHACL validation, and improve LLM-based knowledge graph question answering.

Paul, Lazlo [Lawrence Berkeley National Laboratory↗

Large Language Model for Validation, Optical Calibration, and Learning (VOCAL) Distributed Temperature Sensing Interface

Distributed temperature sensing (DTS) using fiber optic sensors (FOS) offers a promising method for temperature measurements in advanced reactors, such as sodium fast reactors and molten salt cooled reactors. To support the calibration and validation of DTS measurements, Argonne National Laboratory developed the Validation, Optical Calibration, and Learning (VOCAL) software package. This report describes the integration of a local large language model (LLM) with a retrieval-augmented generation (RAG) system into the VOCAL interface to serve as an interactive user assistant. The LLM framework enhances the VOCAL platform’s accessibility to users by explaining interface components, clarifying inputs and outputs, and answering user queries dynamically in real-time. The accuracy of the LLM assistant performance was evaluated with 20 queries regarding the interface and its parameters using experimental data from the Thermal Hydraulic Experimental Test Article (THETA) facility. Results demonstrate that the LLM achieved a 95% accuracy rate, with a BERTScore of 0.8816 and SBERT value of 0.7417. Furthermore, validation of the RAG system within the LLM framework showed optimal accuracy with k-values between 1 and 2 using the k-refinement convergence test. The prompt perturbation analysis demonstrated good initial consistency for the RAG system, exhibiting the highest accuracy under punctuation variations and the greatest sensitivity under query reordering. Notably, the model’s errors were limited to data retrieval failures rather than factual hallucinations, reinforcing its baseline reliability. The integration of LLM provides a highly accurate, userfriendly enhancement to the VOCAL platform without disrupting its core computational capabilities for FOS calibration and validation.

Hong, Evan↗

Bayesian Optimization of Catalysis with In-Context Learning

Large language models (LLMs) can perform accurate classification with zero or few examples through in-context learning (ICL), allowing the model to observe query-relevant examples at inference time and eliminating the need for additional weight updates to generalize beyond its original training data. We extend this capability to regression with uncertainty estimation using frozen LLMs (e.g., GPT-4o, Gemini), enabling Bayesian optimization (BO) in natural language without explicit model training or feature engineering. We apply this to materials discovery by representing materials as synthesis and testing procedures for use in natural language prompts. This Bayesian, design-first approach prioritizes optimization toward target material properties before detailed characterization, in contrast to conventional experimental workflows that often emphasize characterization of suboptimal materials. On benchmarks like aqueous solubility and oxidative coupling of methane (OCM), BO-ICL matches or outperforms Gaussian processes. In live experiments on the reverse water–gas shift (RWGS) reaction, BO-ICL identifies multimetallic catalysts that approach equilibrium CO yield within 6 and 10 iterations from a pool of 3,700 and 360,000 candidates, respectively. Our method redefines materials representation and accelerates discovery, with broad applications across catalysis, materials science, and AI.

Calibration↗

What Is the Agent Doing? Visualizing Agentic AI Querying Workflows

We explore how visualizations can help users understand what an AI agent is doing as it builds and runs queries over data. As part of the LinkQ system, a natural language interface for querying knowledge graphs with a large language model (LLM), we designed two complementary views: A State Diagram that shows where the agent is within a larger workflow, and a Live Action Display that gives real-time updates about the agent's current task. In a study with 14 practitioners, we found that these visuals helped participants build stronger mental models of the agent's behavior while also increasing their confidence in the system. However, we also observed that users sometimes trusted incorrect outputs simply because the agent appeared to be doing the "right" thing. Our findings point to both the value and risk of visualizing agent behavior in interactive AI systems.

97 MATHEMATICS AND COMPUTING↗

CrossLink: Geometry API [Slides]

The mesh generation process is very challenging and time consuming when working with complex CAD models. The process of creating and sorting geometric entities into groups appropriate for meshing is labor intensive and prone to error. In addition, the common data exchange formats such as STEP and IGES do not propagate information such as entity names that may be defined in the original model. Finally, entity counts change frequently with parameter variation as a result of tolerance-based geometry operations. Thus, sorting by index does not provide a robust and repeatable means for grouping. xGeom is a geometry library that enables the creation of NURBS curves and surfaces via a python scripting interface. xGeom is ideal for studying relatively simple models and is fully integrated with CrossLink’s mesh generation capabilities. For more complex models, xCAD is a python-based Creo Parametric CAD model driver that enables the model to be generated, queried, parametrically modified, regenerated, and exported without data loss and in a fully repeatable manner.

97 MATHEMATICS AND COMPUTING↗

Extraction and Analysis of Time Series Data from Building Automation Systems Using Large Language Models

Semantic schemas like Haystack 4, Brick and ASHRAE standard 223 enable the structured, standardized, and machine-readable representation of building data, facilitating interoperability, data integration, and advanced analytics. However, extracting information from these models requires specialized expertise in SPARQL and other programming languages, skills that are not commonly found among building professionals. Recent advancements in Large Language Models (LLMs), such as ChatGPT, enable the construction of queries using natural language, making it easier for individuals to interact with these systems in a manner that resembles everyday speech. However, these methods have not yet been tested on building semantic ontologies. This paper introduces a novel workflow and tool for enabling users to ask questions about a specific building's data, using natural language and receive answers automatically generated by GPT-4o. Our approach integrates semantic ontologies with advanced LLM capabilities to automate three critical steps: (1) generating SPARQL queries to retrieve time series references from ontological models, (2) extracting the corresponding time series data from the Building Automation System, and (3) performing computations and visualizations tailored to the user's query. The proposed method simplifies access to BAS data, allowing both domain experts and non-specialists to conduct sophisticated analyses without needing extensive technical knowledge of semantic web technologies. By demonstrating this pipeline, we facilitate more accessible and scalable data-driven decision-making in building operations and management.

Mulayim, Ozan Baris↗

Extraction and Analysis of Time Series Data from Building Automation Systems Using Large Language Models

Semantic schemas like Haystack 4, Brick and ASHRAE standard 223 enable the structured, standardized, and machine-readable representation of building data, facilitating interoperability, data integration, and advanced analytics. However, extracting information from these models requires specialized expertise in SPARQL and other programming languages, skills that are not commonly found among building professionals. Recent advancements in Large Language Models (LLMs), such as ChatGPT, enable the construction of queries using natural language, making it easier for individuals to interact with these systems in a manner that resembles everyday speech. However, these methods have not yet been tested on building semantic ontologies. This paper introduces a novel workflow and tool for enabling users to ask questions about a specific building's data, using natural language and receive answers automatically generated by GPT-4o. Our approach integrates semantic ontologies with advanced LLM capabilities to automate three critical steps: (1) generating SPARQL queries to retrieve time series references from ontological models, (2) extracting the corresponding time series data from the Building Automation System, and (3) performing computations and visualizations tailored to the user's query. The proposed method simplifies access to BAS data, allowing both domain experts and non-specialists to conduct sophisticated analyses without needing extensive technical knowledge of semantic web technologies. By demonstrating this pipeline, we facilitate more accessible and scalable data-driven decision-making in building operations and management.

Mulayim, Ozan Baris↗

QLiG: Query Like a Graph For Subgraph Matching

A graph is a natural and flexible modeling approach to represent entities and relationships between them in real-world. A Knowledge Graphs (KG) is a specialized graph with formal and structured representation of facts, relationships, annotated with semantic descriptions. Subgraph matching is one of the fundamental graph problems to identify relationships, interactions and activities of interest within a large graph. A query specification is a collection of abstract components, operations, and constraints to express a pattern. The specification can be implemented in different ways based on underlying data model. Various graph query specifications have been developed over the years and have led to the development of different open-sourced and vendor-specific query languages. Such specification are modeled as an extension of relational algebra used to develop relational query languages such as SQL. Such relational concepts do not inherently support graph queries. There is a need to represent graph queries in terms on graph-based components to expedite query construction by non-database experts. We present a graph-based query approach QLiG (pronounced cleeg), to perform subgraph matching in Labeled Property Graph. We present the query specifications, salient features, and a use case to show functional examples.

Purohit, Sumit↗

Training calibration-based counterfactual explainers for deep learning models in medical image analysis

The rapid adoption of artificial intelligence methods in healthcare is coupled with the critical need for techniques to rigorously introspect models and thereby ensure that they behave reliably. This has led to the design of explainable AI techniques that uncover the relationships between discernible data signatures and model predictions. In this context, counterfactual explanations that synthesize small, interpretable changes to a given query while producing desired changes in model predictions have become popular. This under-constrained, inverse problem is vulnerable to introducing irrelevant feature manipulations, particularly when the model’s predictions are not well-calibrated. Hence, in this paper, we propose the TraCE (training calibration-based explainers) technique, which utilizes a novel uncertainty-based interval calibration strategy for reliably synthesizing counterfactuals. Given the wide-spread adoption of machine-learned solutions in radiology, our study focuses on deep models used for identifying anomalies in chest X-ray images. Using rigorous empirical studies, we demonstrate the superiority of TraCE explanations over several state-of-the-art baseline approaches, in terms of several widely adopted evaluation metrics. Our findings show that TraCE can be used to obtain a holistic understanding of deep models by enabling progressive exploration of decision boundaries, to detect shortcuts, and to infer relationships between patient attributes and disease severity.

62 RADIOLOGY AND NUCLEAR MEDICINE↗

Contextual Active Online Model Selection with Expert Advice

How can we collect the most useful labels to learn a model selection policy, when presented with arbitrary heterogeneous data streams? In this paper, we formulate this task as a contextual active model selection problem, where at each round the learner receives an unlabeled data point along with a context. The goal is to output the best model for any given context without obtaining an excessive amount of labels. In particular, we focus on the task of selecting pre-trained classifiers, and propose a contextual active model selection algorithm (CAMS), which relies on a novel uncertainty sampling query criterion defined on a given policy class for adaptive model selection. In comparison to prior art, our algorithm does not assume a globally optimal model. We provide rigorous theoretical analysis for the regret and query complexity under both adversarial and stochastic settings. Our experiments on several benchmark classification datasets demonstrate the algorithm’s effectiveness in terms of both regret and query complexity. Notably, to achieve the same accuracy, CAMS incurs less than 10% of the label cost when compared to the best online model selection baselines on CIFAR10.

Liu, Xuefeng↗

Calibrating hypersonic turbulence flow models with the HIFiRE-1 experiment using data-driven machine-learned models.

In this paper we study the efficacy of combining machine-learning methods with projection-based model reduction techniques for creating data-driven surrogate models of computationally expensive, high-fidelity physics models. Such surrogate models are essential for many-query applications e.g., engineering design optimization and parameter estimation, where it is necessary to invoke the high-fidelity model sequentially, many times. Surrogate models are usually constructed for individual scalar quantities. However there are scenarios where a spatially varying field needs to be modeled as a function of the model’s input parameters. Here we develop a method to do so, using projections to represent spatial variability while a machine-learned model captures the dependence of the model’s response on the inputs. The method is demonstrated on modeling the heat flux and pressure on the surface of the HIFiRE-1 geometry in a Mach 7.16 turbulent flow. The surrogate model is then used to perform Bayesian estimation of freestream conditions and parameters of the SST (Shear Stress Transport) turbulence model embedded in the high-fidelity (Reynolds-Averaged Navier–Stokes) flow simulator, using shock-tunnel data. The paper provides the first-ever Bayesian calibration of a turbulence model for complex hypersonic turbulent flows. We find that the primary issues in estimating the SST model parameters are the limited information content of the heat flux and pressure measurements and the large model-form error encountered in a certain part of the flow.

42 ENGINEERING↗

a priori uncertainty quantification of reacting turbulence closure models using Bayesian neural networks

While many physics-based closure model forms have been posited for the sub-filter scale (SFS) in large eddy simulation (LES), vast amounts of data available from direct numerical simulations (DNS) create opportunities to leverage data-driven modeling techniques. Albeit flexible, data-driven models still depend on the dataset and the functional form of the model chosen. Increased adoption of such models requires reliable uncertainty estimates both in the data-informed and out-of-distribution regimes. Here, in this work, we employ Bayesian neural networks (BNNs) to capture both epistemic and aleatoric uncertainties in a reacting flow model. In particular, we model the filtered progress variable scalar dissipation rate which plays a key role in the dynamics of turbulent premixed flames. We demonstrate that BNN models can provide unique insights about the structure of uncertainty of the data-driven closure models. We also propose a method for the incorporation of out-of-distribution information in a BNN, which can be used for out-of-distribution query detection. The efficacy of the model is demonstrated by a priori evaluation on a dataset consisting of a variety of flame conditions and fuels.

97 MATHEMATICS AND COMPUTING↗