Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “model queries”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Monte Carlo Thought Search: Large Language Model Querying for Complex Scientific Reasoning in Catalyst Design

Discovering novel catalysts requires complex reasoning involving multiple chemical properties and resultant trade-offs, leading to a combinatorial growth in the search space. While large language models (LLM) have demonstrated novel capabilities for chemistry through complex instruction following capabilities and high quality reasoning, a goal-driven combinatorial search using LLMs has not been explored in detail. In this work, we present a Monte Carlo Tree Search-based approach that improves beyond state-of-the-art chain-of-thought prompting variants to augment scientific reasoning. We introduce two new reasoning datasets: 1) a curation of computational chemistry simulations, and 2) diverse questions written by catalysis researchers for reasoning about novel chemical conversion processes. We improve over the best baseline by 25.8\% and find that our approach can augment scientist's reasoning and discovery process with novel insights.\footnote{All resources will be publicly available upon publication.

Sprueill, Henry W.↗

Keep CALM and CRDT On

Despite decades of research and practical experience, developers have few tools for programming reliable distributed applications without resorting to expensive coordination techniques. Conflict-free replicated datatypes (CRDTs) are a promising line of work that enable coordination-free replication and offer certain eventual consistency guarantees in a relatively simple object-oriented API. Yet CRDT guarantees extend only to data updates; observations of CRDT state are unconstrained and unsafe. We propose an agenda that embraces the simplicity of CRDTs, but provides richer, more uniform guarantees. We extend CRDTs with a query model that reasons about which queries are safe without coordination by applying monotonicity results from the CALM Theorem, and lay out a larger agenda for developing CRDT data stores that let developers safely and efficiently interact with replicated application state.

Computer Science↗

PCalc User's Manual

PCalc is a software tool that computes travel-time predictions, ray path geometry and model queries. This software has a rich set of features, including the ability to use custom 3D velocity models to compute predictions using a variety of geometries. The PCalc software is especially useful for research related to seismic monitoring applications.

58 GEOSCIENCES↗

PNNL-CIM-Tools/EASY-CIM

Extracting Attributes in a Simple DictionarY of CIM (EASY-CIM) is a python library to reduce the time and burden of model queries using simplified data representations of CIM power system models.

Anderson, Alex↗

Methods and system for siting advanced nuclear reactors and evaluating energy policy concerns

There is a growing sociopolitical desire to develop cleaner energy sources in the United States and maintain energy security. Regardless of politics, many coal-fired electric plants have already been shut down and many utilities are vowing to retire their current coal-fired assets within the next two decades. Replacement power assets require consideration of appropriate siting. A geographic information system (GIS)-based multicriteria decision analysis approach is useful to assist utility and energy companies, as well as policymakers, to evaluate potential areas for siting new plants in the contiguous United States. A GIS-based framework is simply a database of location information that allows for mapping, querying, modeling, and analyzing data based on location. The spatial output can be structured to be visual, allowing for easier analysis of location data. The need to site additional power assets, including renewable resources and clean power sources, such as nuclear, led to the development of the Oak Ridge Siting Analysis for power Generation Expansion (OR-SAGE) tool discussed in this paper. The tool takes inputs such as population growth, water availability, environmental indicators, and tectonic and geological hazards to provide an in-depth visual analysis for siting options. Energy companies and other stakeholders can use OR-SAGE to procure feedback quickly and effectively on land suitability based on technology specific inputs. Policymakers can use OR-SAGE to analyze the impacts of future energy technology decisions, while balancing competing resource use. Overall, this paper discusses the recent use of OR-SAGE for these purposes and plans for future development.

21 SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLANTS↗

Large Language Models for the Creation and Use of Semantic Ontologies in Buildings: Requirements and Challenges

Semantic ontologies offer a formalized, machine-readable framework for representing knowledge, enabling the structured description of complex systems. In the building domain, the adoption of ontologies like the Brick schema has transformed how buildings and their systems are modeled by providing a standardized, interoperable language. However, the complexity and the steep learning curve involved in developing and querying semantic models present substantial challenges, often requiring a workforce with specialized expertise. This paper builds on our experience in investigating how Large Language Models (LLMs) can help address these challenges, focusing on their role in constructing and querying of semantic models, particularly using the Brick Schema. Our study outlines the requirements and metrics for evaluating the scalability and effectiveness of LLM-based tools, while also discussing the current challenges and limitations in developing such tools. Ultimately, this paper aims to orient research efforts as various groups experiment with diverse techniques, while enabling more effective comparison of emerging solutions and fostering collaboration across the field.

Mulayim, Ozan Baris↗

QED: A Powerful Query Equivalence Decider for SQL

Checking query equivalence is of great significance in database systems. Prior work in automated query equivalence checking sets the first steps in formally modeling and reasoning about query optimization rules, but only supports a limited number of query features. In this paper, we present Qed, a new framework for query equivalence checking based on bag semantics. Qed uses a new formalism called Q-expressions that models queries using different normal forms for efficient equivalence checking, and models features such as integrity constraints and NULLs in a principled way unlike prior work. Our formalism also allows us to define a new query fragment that encompasses many real-world queries with a complete equivalence checking algorithm, assuming a complete first-order theory solver. Empirically, Qed can verify 299 out of 444 query pairs extracted from the Calcite framework and 979 out of 1287 query pairs extracted from CockroachDB, which is more than 2× the number of cases proven by prior state-of-the-art solver.

Computer Science↗

Multipoint-BAX: a new approach for efficiently tuning particle accelerator emittance via virtual objectives

Abstract Although beam emittance is critical for the performance of high-brightness accelerators, optimization is often time limited as emittance calculations, commonly done via quadrupole scans, are typically slow. Such calculations are a type of multipoint query , i.e. each query requires multiple secondary measurements. Traditional black-box optimizers such as Bayesian optimization are slow and inefficient when dealing with such objectives as they must acquire the full series of measurements, but return only the emittance, with each query. We propose a new information-theoretic algorithm, Multipoint-BAX , for black-box optimization on multipoint queries, which queries and models individual beam-size measurements using techniques from Bayesian Algorithm Execution (BAX). Our method avoids the slow multipoint query on the accelerator by acquiring points through a virtual objective , i.e. calculating the emittance objective from a fast learned model rather than directly from the accelerator. We use Multipoint-BAX to minimize emittance at the Linac Coherent Light Source (LCLS) and the Facility for Advanced Accelerator Experimental Tests II (FACET-II). In simulation, our method is 20× faster and more robust to noise compared to existing methods. In live tests, it matched the hand-tuned emittance at FACET-II and achieved a 24% lower emittance than hand-tuning at LCLS. Our method represents a conceptual shift for optimizing multipoint queries, and we anticipate that it can be readily adapted to similar problems in particle accelerators and other scientific instruments.

43 PARTICLE ACCELERATORS↗

Differentially Private Adaptive Noise Injection (DP-ANI) v1.0

Location data is collected from users continuously to understand their mobility patterns. Releasing the user trajectories may compromise user privacy. Therefore, the general practice is to release aggregated location datasets. However, private information may still be inferred from an aggregated version of location trajectories. Differential privacy (DP) protects the query output against inference attacks regardless of background knowledge. This software implements a differential privacy-based privacy model that protects the user's origins and destinations from being inferred from aggregated mobility datasets. This is achieved by injecting Planar Laplace noise to the user origin and destination GPS points. The noisy GPS points are then transformed into a link representation using a link-matching algorithm. Finally, the link trajectories form an aggregated mobility network. The injected noise level is selected using the Sparse Vector Mechanism. This DP selection mechanism considers the link density of the location and the functional category of the localized links. Compared to the different baseline models, including a k-anonymity method, our differential privacy-based aggregation model offers query responses that are close to the raw data in terms of aggregate statistics at both the network and trajectory-levels with maximum 9% deviation from the baseline in terms of network length.

Peisert, Sean [Lawrence Berkeley National Laborato↗

Query relaxation for portable brick-based applications

Semantic metadata standards pave the way for interoperability by providing building operators and application developers with common schemes to describe building resources. Applications can query building metadata models to retrieve the set of entities and relationships they need to operate, instead of hard-coding references to specific points and objects from the underlying data sources. Currently, querying such models requires the developer to be very specific when formulating queries in order to obtain meaningful answers (or any answer). The developer is inevitably expected to be familiar with the systems and components of the buildings being queried, as well as the schema used to represent them. The variety of buildings - both in the composition of their subsystems and in how they happen to be modeled - means that the developer will need to use multiple queries in order to retrieve necessary results. This is complex, time-consuming and error-prone. To address this limitation, we investigate query relaxation as a technique to facilitate discovery of meaningful building resources in a collection of ontology-based buildings data. We evaluate our query relaxation approach over a set of Brick models and demonstrate its use in the context of real-world building applications.

Brick↗

Efficient mapping between void shapes and stress fields using Deep Convolutional Neural Networks with sparse data

Establishing fast and accurate structure-to-property relationships is an important component in the design and discovery of advanced materials. Physics-based simulation models like the finite element method (FEM) are often used to predict deformation, stress, and strain fields as a function of material microstructure in material and structural systems. Such models may be computationally expensive and time intensive if the underlying physics of the system is complex. This limits their application to solve inverse design problems and identify structures that maximize performance. In such scenarios, surrogate models are employed to make the forward mapping computationally efficient to evaluate. However, the high dimensionality of the input microstructure and the output field of interest often renders such surrogate models inefficient, especially when dealing with sparse data. Deep convolutional neural network (CNN) based surrogate models have shown great promise in handling such high-dimensional problems. In this paper, a single ellipsoidal void structure under a uniaxial tensile load represented by a linear elastic, high-dimensional and expensive-to-query, FEM model. We consider two deep CNN architectures, a modified convolutional autoencoder framework with a fully connected bottleneck and a UNet CNN, and compare their accuracy in predicting the von Mises stress field for any given input void shape in the FEM model. Additionally, a sensitivity analysis study is performed using the two approaches, where the variation in the prediction accuracy on unseen test data is studied through numerical experiments by varying the number of training samples from 20 to 100.

surrogate modeling; convolutional neural networks;↗

A portable application framework for energy management and information systems (EMIS) solutions using Brick semantic schema

This paper introduces a portable framework for developing, scaling and maintaining energy management and information systems (EMIS) applications using an ontology-based approach. Key contributions include an interoperable layer based on Brick schema, the formalization of application constraints pertaining metadata and data requirements, and a field demonstration. The framework allows for querying metadata models, fetching data, preprocessing, and analyzing data, thereby offering a modular and flexible workflow for application development. Its effectiveness is demonstrated through a case study involving the development and implementation of a data-driven anomaly detection tool for the photovoltaic systems installed at the Politecnico di Torino, Italy. During eight months of testing, the framework was used to tackle practical challenges including: (i) developing a machine learning-based anomaly detection pipeline, (ii) replacing data-driven models during operation, (iii) optimizing model deployment and retraining, (iv) handling critical changes in variable naming conventions and sensor availability (v) extending the pipeline from one system to additional ones.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

NeuralCubes: Deep Representations for Visual Data Exploration

Visual exploration of large multi-dimensional datasets has seen tremendous progress in recent years, allowing users to express rich data queries that produce informative visual summaries, all in real time. Techniques based on data cubes are some of the most promising approaches. However, these techniques usually require a large memory footprint for large datasets. To tackle this problem, we present NeuralCubes: neural networks that predict results for aggregate queries, similar to data cubes. NeuralCubes learns a function that takes as input a given query, for instance, a geographic region and temporal interval, and outputs the result of the query. The learned function serves as a real-time, low-memory approximator for aggregation queries. Our models are small enough to be sent to the client side (e.g. the web browser for a web-based application) for evaluation, enabling data exploration of large datasets without database/network connection. Here, we demonstrate the effectiveness of NeuralCubes through extensive experiments on a variety of datasets and discuss how NeuralCubes opens up opportunities for new types of visualization and interaction.

97 MATHEMATICS AND COMPUTING↗

Pareto Optimization of Analog Circuits Using Reinforcement Learning

Analog circuit optimization and design presents a unique set of challenges in the IC design process. Many applications require the designer to optimize for multiple competing objectives, which poses a crucial challenge. Motivated by these practical aspects, we propose a novel method to tackle multi-objective optimization for analog circuit design in continuous action spaces. In particular, we propose to (i) extrapolate current techniques in Multi-Objective Reinforcement Learning to continuous state and action spaces and (ii) provide for a dynamically tunable trained model to query user defined preferences in multi-objective optimization in the analog circuit design context.

99 GENERAL AND MISCELLANEOUS↗

A dataset of cyber-induced mechanical faults on buildings with network and buildings data

We have collected data of cyber-induced mechanical faults on buildings using a simulation platform. A DOE reference building model was used for running the simulation under a Rogue device attack and collected the network data as well as the physical buildings data to better understand the impacts of cyber attacks on the building and help identify the source of the mechanical fault with the network data. Alfalfa is the tool used for simulating the DOE reference buildings and acts as an interface to the model for querying the status and providing input externally. The Building Automation System (BAS) is the centralized controller providing control commands to other BACnet devices on the network based on the building status received from Alfalfa. The BACnet devices like damper will listen for the control commands from BAS on the BACnet network and implement it. The attacker is the malicious actor on the network creating disruptions by placing cyber-attacks.

24 POWER TRANSMISSION AND DISTRIBUTION↗

msr-genie

MSR-Genie is a vendor-neutral tool for fast and efficient queries about model-specific registers (MSRs). The tool allows end users to query bi-directionally across MSR lists as well as processor families and models, and provides them with guidance on appropriate bitmasks. The msr-genie tool is open-source and easily extensible, and currently supports 30 Intel processor models and two-thousand unique MSRs.

Fan, Kyle↗

Robust Multi-fidelity Bayesian Optimization with Deep Kernel and Partition

Multi-fidelity Bayesian optimization (MFBO) is a powerful approach that utilizes lowfidelity, cost-effective sources to expedite the exploration and exploitation of a high-fidelity objective function. Existing MFBO methods with theoretical foundations either lack justification for performance improvements over single-fidelity optimization or rely on strong assumptions about the relationships between fidelity sources to construct surrogate models and direct queries to low-fidelity sources. To mitigate the dependency on cross-fidelity assumptions while maintaining the advantages of low-fidelity queries, we introduce a random sampling and partition-based MFBO framework with deep kernel learning. This framework is robust to cross-fidelity model misspecification and explicitly illustrates the benefits of low-fidelity queries. Our results demonstrate that the proposed algorithm effectively manages complex cross-fidelity relationships and efficiently optimizes the target fidelity function.

Zhang, Fengxue [University of Chicago, Illinois, U↗

GraphAide: Advanced Graph-Assisted Query and Reasoning System

Curating knowledge from multiple siloed sources that contain both structured and unstructured data is a major challenge in many real-world applications. Pattern matching and querying represent fundamental tasks in modern data analytics that leverage this curated knowledge. The development of such applications necessitates overcoming several research challenges, including data extraction, named entity recognition, data modeling, and designing query interfaces. Moreover, the explainability of these functionalities is critical for their broader adoption. The emergence of Large Language Models (LLMs) has accelerated the development lifecycle of new capabilities. Nonetheless, there is an ongoing need for domain-specific tools tailored to user activities. The creation of digital assistants has gained considerable traction in recent years, with LLMs offering a promising avenue to develop such assistants utilizing domain-specific knowledge and assumptions. In this context, we introduce an advanced query and reasoning system, GraphAide, which constructs a knowledge graph (KG) from diverse sources and allows to query and reason over the resulting KG. GraphAide harnesses both the KG and LLMs to rapidly develop domain-specific digital assistants. It integrates design patterns from retrieval augmented generation (RAG) and the semantic web to create an agentic LLM application. GraphAide underscores the potential for streamlined and efficient development of specialized digital assistants, thereby enhancing their applicability across various domains.

Purohit, Sumit [BATTELLE (PACIFIC NW LAB)] (ORCID:↗