Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “sql”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

On-Demand Column Joining for High Energy Physics

As the Large Hadron Collider (LHC) transitions into the High-Luminosity LHC (HL-LHC) era, the volume of data to be processed is expected to increase significantly. The CMS Experiment currently utilizes various data formats, including AOD, MiniAOD, and NanoAOD, each with different levels of detail and storage requirements. This paper addresses the challenges of data duplication and storage inefficiencies in high-energy physics (HEP) analyses by proposing an on-demand column-joining solution. This approach aims to reduce data duplication by enabling the dynamic combination of NanoAOD data with auxiliary information from larger data tiers, such as MiniAOD. The proposed solution leverages Trino, a high-performance distributed SQL query engine, to perform efficient and scalable data joins. Benchmarks using CMS OpenData demonstrate the feasibility of this approach, showing that it can handle large datasets with low latency. Integration with the scikit-hep ecosystem and the coffea analysis framework is also discussed, highlighting the potential for seamless end-to-end data processing and analysis. Ongoing and future work focuses on expanding benchmarks, integrating ServiceX for data transformation, and exploring the use of native object storage solutions.

Manganelli, Nicholas [Northeastern U.]↗

Quantum metrology of low-frequency electromagnetic modes with frequency upconverters

We present the RF Quantum Upconverter (RQU) and describe its application to quantum metrology of electromagnetic modes between dc and the very high frequency band (VHF) ( ≲ 300 MHz). The RQU uses a Josephson interferometer made up of superconducting loops and Josephson junctions to implement a parametric interaction between a low-frequency electromagnetic mode (between dc and VHF) and a mode in the microwave C Band ( ∼ 5 GHz), analogous to the radiation pressure interaction between electromagnetic and mechanical modes in cavity optomechanics. We analyze RQU performance with quantum amplifier theory and show that the RQU can operate as a quantum-limited op-amp in this frequency range. It can also use nonclassical measurement protocols equivalent to those used in cavity optomechanics, including back-action evading (BAE) measurements, sideband cooling, and two-mode squeezing. These protocols enable experiments using dc VHF electromagnetic modes as quantum sensors with sensitivity better than the standard quantum limit (SQL). We demonstrate signal upconversion from low frequencies to the microwave C band using an RQU and show a phase-sensitive gain (extinction ratio) of 46.9 dB , which is a necessary step towards the realization of full BAE. Published by the American Physical Society 2025

Kuenstner, Stephen E. (ORCID:0000000346128846)↗

Quantum Frequency Combs with Path Identity for Quantum Remote Sensing

Quantum sensing promises to revolutionize sensing applications by employing quantum states of light or matter as sensing probes. Photons are the clear choice as quantum probes for remote sensing because they can travel to and interact with a distant target. Existing schemes are mainly based on the quantum illumination framework, which requires quantum memory to store a single photon of an initially entangled pair until its twin reflects off a target and returns for final correlation measurements. Existing demonstrations are limited to tabletop experiments, and expanding the sensing range faces various roadblocks, including long-time quantum storage and photon loss and noise when transmitting quantum signals over long distances. We propose a novel quantum sensing framework that addresses these challenges using quantum frequency combs with path identity for remote sensing of signatures (“qCOMBPASS”). The combination of one key quantum phenomenon and two quantum resources—namely, quantum-induced coherence by path identity, quantum frequency combs, and two-mode squeezed light—allows for quantum remote sensing without requiring quantum memory. The proposed scheme is akin to a quantum radar based on entangled frequency-comb pairs that uses path identity to detect, range, or sense a remote target of interest by measuring pulses of one comb in the pair that never traveled to the target but that contains target information “teleported” by quantum-induced coherence by path identity from the other comb in the pair that traveled to the target but is not detected. We develop the basic qCOMBPASS theory, analyze the properties of the qCOMBPASS transceiver, and introduce the qCOMBPASS equation—a quantum analog of the well-known LIDAR equation in classical remote sensing. We also describe an experimental scheme to demonstrate the concept using two-mode squeezed quantum combs. qCOMBPASS can strongly impact various applications in remote quantum sensing, imaging, metrology, and communications. These applications include detection and ranging of low-reflectivity objects, measurement of small displacements of a remote target with precision beyond the standard quantum limit (SQL), standoff hyperspectral quantum imaging, discreet surveillance from space with low detection probability (detect without being detected), very-long-baseline interferometry, quantum Doppler sensing, quantum clock synchronization, and networks of distributed quantum sensors. Published by the American Physical Society 2024

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

QLiG: Query Like a Graph For Subgraph Matching

A graph is a natural and flexible modeling approach to represent entities and relationships between them in real-world. A Knowledge Graphs (KG) is a specialized graph with formal and structured representation of facts, relationships, annotated with semantic descriptions. Subgraph matching is one of the fundamental graph problems to identify relationships, interactions and activities of interest within a large graph. A query specification is a collection of abstract components, operations, and constraints to express a pattern. The specification can be implemented in different ways based on underlying data model. Various graph query specifications have been developed over the years and have led to the development of different open-sourced and vendor-specific query languages. Such specification are modeled as an extension of relational algebra used to develop relational query languages such as SQL. Such relational concepts do not inherently support graph queries. There is a need to represent graph queries in terms on graph-based components to expedite query construction by non-database experts. We present a graph-based query approach QLiG (pronounced cleeg), to perform subgraph matching in Labeled Property Graph. We present the query specifications, salient features, and a use case to show functional examples.

Purohit, Sumit↗

App2Net: Moving App Functions to Network & a Case Study on Low-latency Feedback

Recent advances in programmable networks enable custom processing of data at hundreds of gigabits per second. These advances can boost the performance of many distributed applications. Yet the high-level languages used by application developers are different from the data plane programming languages (such as P4 and NPL) used by network equipment. This language barrier slows innovation. Our hourglass-shaped architectural solution aims to lower this language barrier. This enables the application developer community to leverage programmable networks for achieving better performance. In this paper we propose a JSON-based intermediate representation to bridge the gap between applications and in-network computing. We demonstrate an instance of the solution in the context of a low-latency feedback application that enables SQL-based data filtering in a P4-based programmable environment. We also present a prototype compiler to convert an intermediate representation in JSON to P4 source.

Sankaran, Ganesh↗

Performance Analysis of Data Processing in Distributed File Systems with Near Data Processing

In the era of big data, the escalating volume and velocity of data generation pose significant challenges in data processing. Traditional systems like Spark and Hadoop manage the increasing amount and velocity of data by improving data placement and processing speeds. However, they face inherent limitations due to the essential data movement required for processing. In this paper, we explore the Skyhook framework, a novel extension of the Ceph distributed system, which significantly reduces the need for data movement. We present an extensive case study using the Skyhook framework, applying it with the TPC-H and K-means clustering algorithms. More specifically, we leverage the TPC-H benchmark to distinguish between CPU-intensive and I/O-intensive tasks. We explore the integration of K-means clustering into SQL, coupled with a near-data processing system to offload the computational burden of the K-means clustering algorithm to storage nodes. We conduct a comprehensive performance evaluation of distributed data processing applications across three processing approaches: traditional layout (baseline), optimized layout, and near-data processing. Additionally, we introduce the use of the FIO tool to simulate real-world system workloads, enabling the measurement of performance metrics such as average latency and CPU utilization. Our research is a significant advance in understanding how to optimize data processing systems to meet the demands of the modern data landscape.

Hou, Shiyue↗

mkite

mkite is a distributed computing platform for materials simulation. mkite is built with the server-client pattern, decoupling production databases from client runners. When used in combination with message brokers, mkite enables any available client to perform calculations without prior hardware specification on the server side. Furthermore, the software enables the creation of complex workflows with multiple inputs and branches, facilitating the exploration of combinatorial chemical spaces. The mkite suite provides recipes and tools to interact with package such as VASP, but is extensible to any other simulation package. Finally, mkite helps keeping the provenance of calculations in a SQL database. A complete description of the software is available at https://arxiv.org/abs/2301.08841.

Schwalbe Koda, Daniel↗

Active Learning Framework

Machine learning (ML) of interatomic potentials show great promise to accelerate scientific simulation, e.g., by emulating expensive computations at a high accuracy but much reduced computational cost. Training datasets are calculated from computationally expensive ab initio quantum mechanics methods, density functional theory (DFT). Trained on this data, an ML model can be very successful in predicting energy and forces for new atomic configurations. A critical factor is the quality and diversity of the training dataset. Thus, a highly automated approach to dataset construction based on active learning framework is designed suitable for material physics. The active learning scheme begins with fully randomized atomic configurations. Then, many Molecular Dynamics (MD) trajectories are simulated using current ML potentials, where each MD trajectory is initialized to a random disordered configuration. The temperature is varied in order to diversify the sampled configuration during these simulations. The variance of predictions for eight neural networks within an ensemble is analyzed to determine whether the model is operating as expected. This helps in determining whether collecting more data would be helpful to the model by checking the ensemble variance is greater than the threshold. In this case, the MD trajectory is terminated and the final atomic configuration is placed on a queue (SQL database) for DFT calculations and added to training dataset. Periodically, ML model is retrained to the updated training model. This Active Learning loop is iterated until the cost of MD simulations becomes prohibitively expensive. The MD simulations will hopefully be sufficiently robust to support nucleation after many active learning iterations. In this sense, active learning scheme must automatically discover the important low energy and nonequilibrium physics.

Nebgen, Benjamin↗

BuildStockQuery [SWR-23-58]

BuildStockQuery is a python library designed to simplify and streamline the process of querying massive, terabyte-scale datasets generated by ResStock(TM). ResStock (SWR-19-15) is a U.S. DOE-supported, NREL-built, national residential building energy stock model that enables a new approach to large-scale residential energy analysis across the U.S. by combining large public and private data sources, statistical sampling, detailed sub-hourly building simulations, and high-performance computing. BuildStockQuery offers an intuitive Object-Oriented Programming (OOP) interface to the ResStock output dataset allowing users to easily perform common queries and receive results in familiar pandas DataFrame format, abstracting away the need for complex SQL query. By initializing a query object with the pertinent Athena database and table names, users can easily query for various kinds of insights, for example, timeseries electricity for an end use for a given state grouped by building types.

Adhikari, Rajendra↗

Scribe Network API v.1.0.0

SAND2023-07864O Scribe Network API is an add-on tool for Scribe3D software. Scribe Network API documents tabletop exercises in trainings and plays back simulated videos of the scenarios and responses. Users can apply the software to compiled projects, or to a simple visual studio project solution package. The software runs a web application that relays information between computers using Scribe3D through a web application, or web app. The web app involves a representational state transfer (REST) application programming interface that handles sending and receiving Scribe save files and an SQL server that stores the save files. This follow-on package allows users to facilitate a networked tabletop exercise. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Noel, Todd↗

Libra

Libra is a Python package that extends support of its parent package, SQLAlchemy. Libra’s primary functionality supports the dynamic creation of object-oriented analogs of SQL tables from a variety of user-defined, text-based schema definition formats and provides quality control and analysis tools and methods

Spears, Brady↗

OpenStudio®-MCP [SWR-26-035]

OpenStudio®-MCP is a Model Context Protocol (MCP) server that lets AI assistants perform building energy modeling through natural language. Rather than requiring users to learn the OpenStudio® SDK, EnergyPlus® scripting, or Ruby/Python automation, the server translates conversational requests into sequences of tool calls that create models, design HVAC systems, run simulations, and extract results — all within a single chat session. The server's 124 tools are organized into a skills architecture where each skill encapsulates a domain of building energy modeling (envelope, HVAC, loads, weather, simulation, results) behind typed, LLM-friendly interfaces. High-leverage operations like applying ASHRAE 90.1 baseline systems or generating standards-compliant typical buildings are exposed as single tool calls that internally wire dozens of OpenStudio® objects. Bundled measures from ComStock™ and Openstudio® -common-measures-gem are wrapped with dedicated tools and typed arguments rather than exposed through a generic measure interface, so AI models get consistent, error-resistant recipes without needing to discover measure arguments at runtime. A key design decision is structured results extraction: six SQL-based tools return surgical ~300–1,000 token responses (end-use breakdowns, envelope summaries, HVAC sizing, timeseries data) instead of requiring the AI to parse ~100K-token raw HTML reports, making iterative design exploration practical within context window limits. The codebase is designed as a reference implementation — explicit, well-commented, and modular — so that other simulation engines (EnergyPlus® standalone, TRNSYS, DOE-2) can use it as a template for building their own MCP servers.

Ball, Brian [National Laboratory of the Rockies (N↗

Accessing Microsoft Access Databases Using ODBC and RODBC

ODBC (Open Database Connectivity) is an industry-standard API (Application Program Interface) that provides a standard interface, based on SQL (Structured Query Language) between applications and databases. This insulates applications from specific details of different database management systems (DBMS). Microsoft Windows TM provides an implementation of ODBC, which, along with drivers for various databases, supports the API. Windows also provides a driver for Access databases.

97 MATHEMATICS AND COMPUTING↗

Evaluation of Apex Alpha with LabWare-LIMS Data Reduction for Alpha Analyses

Savannah River Site's Environmental Bioassay Laboratory is migrating its laboratory information management system (LIMS) from SQL-LIMS to Oracle Labware-LIMS (LW-LIMS) systems to align with the Department of Energy's cyber security policies. Concurrently with the LIMS upgrade, the current VMS based Alpha Measurement System (AMS) software is being replaced with Apex Alpha software to reduce other cyber vulnerabilities, streamline procedural production aspects of data handling, and improve outdated reduction processes to align the program with requirements outlined in ANSI N13.30. The primary technical changes which are being implemented are: - Reagent contributions are added in data reduction, - Moving average background equivalent activity are used for gross signal corrections, - Measured uncertainty of both are included in the decision level and propagated uncertainty, - Tracer levels are increased for precision and accuracy improvements, and - Reports are modified for compliance.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Modeling of spin squeezing in OPMs and performance improvements

We propose to develop an optically pumped magnetometer (OPM) that implements spin squeezing to achieve better than standard quantum limit (SQL) performance while operating in mission relevant environments. This performance enhancement would allow us to increase the useable bandwidth by a factor > 2 while maintaining world class magnetic sensitivity, an order of magnitude better than commercially available options.

42 ENGINEERING↗

HBET V3.0 Installation Manual

The Hydropower Biological Evaluation Toolset (HBET) V3.0 now requires Python v3.11.0 to be installed, following the addition of the absolute fish injury rate prediction feature. This version introduces two new strike metrics—based on velocity and pressure—to provide a more precise understanding of the biological effects of fish collisions with rigid structures within the fish passage system. Additionally, SQL Server 2019 is the supported database for this release. This installation guide will walk users through the process of installing HBET V3.0 along with all necessary dependencies.

13 HYDRO ENERGY↗

Evaluation of Graph Analytics Frameworks Using the GAP Benchmark Suite

The analysis of connected data is an increasingly important application in high-performance computing. Such analyses can reveal fraudulent patterns in financial transactions, optimize telecommunications networks, predict information flow in social networks, etc. However, the landscape of graph analytics is highly diverse. Graph algorithms stress processor architectures differently, and no one graph can represent all topologies. Consequently, no single approach or framework is expected to be optimal for all graph analytics problems. To help make sense of this diverse landscape, we evaluated four approaches to graph analytics: GraphBLAS, Galois, BGL17, GraphIt; and compare them against hand-tuned implementations that take advantage of hardware features on our test platform. Graph- BLAS formulates graph analytics as sparse linear algebra. Galois provides syntactic constructs for data parallelism over irregular data structures. BGL17 is a generic C++ template library for implementing graph algorithms. GraphIt provides a domain- specific language to describe and optimize graph algorithms. We use the GAP Benchmark Suite to establish baseline performance and guide the side-by-side evaluation of each framework. GAP consists of 30 tests: six graph analytics algorithms (breadth- first search, single-source shortest path, PageRank, betweenness centrality, connected components, and triangle counting) run on five graphs, each with different topological characteristics (e.g., high diameter, skewed degree distribution, high average degree). High-performance reference implementations are included for each benchmark algorithm. Because a graph can be loaded into memory a number of ways (e.g., flat file on disk, compressed sparse format, data frames, retrieved from SQL or NoSQL databases), our evaluation focused on computational performance rather than I/O. Our results show the relative strengths of each framework.

Graph algorithms, Benchmarking, shared-memory prog↗

GBCGE Subsurface Database Explorer and APIs

This submission defines a DOI for the Great Basin Center for Geothermal Energy's (GBCGE) Subsurface Database Explorer web application and underlying data services, and acknowledges the INGENIOUS project as a major source of funding for data compilation and quality assurance. The GBCGE Subsurface Database Explorer is an interactive web mapping application that provides public access to the GBCGE Subsurface Database, and its collection of datasets pertinent to geothermal exploration, oil and gas exploration, critical mineral exploration, and other subsurface characterization for the Great Basin Region, western US. This is a living database, and will be continuously updated with new data and datasets as funding and motivations allow. The underlying database views that populate the web application are on an automated refresh schedule. Data sources and acknowledgements: We thank our partners with the Nevada Division of Minerals (NDOM), the Southern Methodist University (SMU), and Great Basin State Geological Surveys for their active efforts in data curation, schema design, and quality assurance. We also thank contributors among the USGS, Oregon Institute of Technology, State Divisions of Water Resources, State Divisions of Oil, Gas, and Minerals, and State Geological Surveys for open data availability and direct contributions made under the National Geothermal Data System (NGDS).

15 GEOTHERMAL ENERGY↗