Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Data Systems”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

Data Acquisition and Control for Marine Energy Devices: Cost Considerations

This document discusses the process involved with developing a data acquisition system specifically in the context of applications for Marine Renewable Energy (MRE) technologies however, much of what is presented is applicable to applications requiring data acquisition in general. The detail on the process is provided to highlight the critical steps and needs for a successful measurement campaign and to understand what can impact the overall outcome, cost, and schedule. The process presented is an amalgamation of best practices, lessons learned, recommendations, and prudent technical project planning and management. Data acquisition systems may be tightly integrated with or into the device under measurement and it often has its own dependencies that must be met. Therefore, early consideration and planning for the data acquisition system are stressed throughout this document.

13 HYDRO ENERGY↗

Integration of Solid Oxide Fuel Cell Systems Into Artificial Intelligence Data Centers

This report presents the results of a techno-economic analysis (TEA) that evaluates the economic benefits of integrating solid oxide fuel cell (SOFC) systems with artificial intelligence (AI) data centers. The analysis was completed in two phases: a scoping-level analysis was performed to identify impactful integration opportunities, followed by a more detailed TEA. Results show that, due to their modularity, SOFC can meet the 99.999% availability requirement of data centers with minimal additional costs. Heat integration via absorption chillers decreases data center electricity consumption at the tradeoff of increased water consumption. Higher SOFC exhaust temperatures are important for achieving larger electricity savings. Finally, power electronics integration with SOFC direct current electricity can reduce electricity consumption by 9 percent and reduce water consumption by 6.4 percent.

20 FOSSIL-FUELED POWER PLANTS↗

Oscilloscope Data Push Program

This paper details the development of a Python program designed to automate the data acquisition and conversion for an oscilloscope for the purposes of a one-off/temporary data acquisition system for users that readily need data, and do not have the option of obtaining a Data Acquisition (DAQ) solution. Creating DAQ systems for analyzing a system requires expensive electronics and a dedicated team of engineers for support. Traditionally, manual data collection and processing are time consuming and prone to error. By automating these processes, the cost, efficiency and accuracy of data handling are improved upon. This project involves the creation of a program that interacts with the oscilloscope. During this interaction, there are various functions being performed such as the acquisition of waveform data via floating points, generating plots with the acquired wave points, and storing of floating points in a CSV file format for future reference and plotting purposes. While the initial aim of the project included continuous logging to a cloud database, this was deferred due to time constraints. The results portrayed an almost-instant rate of data collection with a buffer time, showcasing the potential for further integration and real-time data processing.

Osei-Tutu, Jason↗

Extraction and Analysis of Time Series Data from Building Automation Systems Using Large Language Models

Semantic schemas like Haystack 4, Brick and ASHRAE standard 223 enable the structured, standardized, and machine-readable representation of building data, facilitating interoperability, data integration, and advanced analytics. However, extracting information from these models requires specialized expertise in SPARQL and other programming languages, skills that are not commonly found among building professionals. Recent advancements in Large Language Models (LLMs), such as ChatGPT, enable the construction of queries using natural language, making it easier for individuals to interact with these systems in a manner that resembles everyday speech. However, these methods have not yet been tested on building semantic ontologies. This paper introduces a novel workflow and tool for enabling users to ask questions about a specific building's data, using natural language and receive answers automatically generated by GPT-4o. Our approach integrates semantic ontologies with advanced LLM capabilities to automate three critical steps: (1) generating SPARQL queries to retrieve time series references from ontological models, (2) extracting the corresponding time series data from the Building Automation System, and (3) performing computations and visualizations tailored to the user's query. The proposed method simplifies access to BAS data, allowing both domain experts and non-specialists to conduct sophisticated analyses without needing extensive technical knowledge of semantic web technologies. By demonstrating this pipeline, we facilitate more accessible and scalable data-driven decision-making in building operations and management.

Mulayim, Ozan Baris↗

Extraction and Analysis of Time Series Data from Building Automation Systems Using Large Language Models

Semantic schemas like Haystack 4, Brick and ASHRAE standard 223 enable the structured, standardized, and machine-readable representation of building data, facilitating interoperability, data integration, and advanced analytics. However, extracting information from these models requires specialized expertise in SPARQL and other programming languages, skills that are not commonly found among building professionals. Recent advancements in Large Language Models (LLMs), such as ChatGPT, enable the construction of queries using natural language, making it easier for individuals to interact with these systems in a manner that resembles everyday speech. However, these methods have not yet been tested on building semantic ontologies. This paper introduces a novel workflow and tool for enabling users to ask questions about a specific building's data, using natural language and receive answers automatically generated by GPT-4o. Our approach integrates semantic ontologies with advanced LLM capabilities to automate three critical steps: (1) generating SPARQL queries to retrieve time series references from ontological models, (2) extracting the corresponding time series data from the Building Automation System, and (3) performing computations and visualizations tailored to the user's query. The proposed method simplifies access to BAS data, allowing both domain experts and non-specialists to conduct sophisticated analyses without needing extensive technical knowledge of semantic web technologies. By demonstrating this pipeline, we facilitate more accessible and scalable data-driven decision-making in building operations and management.

Mulayim, Ozan Baris↗

Cataloging Legacy Data from the Tritium Systems Test Assembly Program

The Tritium Systems Test Assembly (TSTA) at Los Alamos National Laboratory, operational from 1984 to 2001, was critical in advancing fusion fuel cycle technologies, including tritium storage, gas separation, and pumping. TSTA’s contributions, particularly in safe tritium operations, have influenced subsequent fusion projects. This paper discusses the ongoing effort to digitize and catalog TSTA’s historical data to create a searchable resource for the fusion research community. While the long-term objective is to develop a relational database for structured data management, the project remains in the early phase, with current efforts focused on scanning and indexing physical documents. Initial plans for database implementations are also presented, outlining key considerations for structure, query indexing, and standardization. As digitization progresses, future discussions will refine these implantation details to ensure an efficient and comprehensive system. This initiative aims to preserve critical legacy data, enhance the design of tritium system facilities, and support the next generation of fusion energy research.

42 ENGINEERING↗

A Graph-Net with Node Embeddings to Detect False Data Injection Attacks in Photovoltaic Systems

Distributed energy resources (DER) contribute to the operational stability of the larger power grid both at utility-scale as well as commercial and residential scales in aggregated forms. These DER in-turn are susceptible to increasing cyber threats. An adversary can plug into the same local network that a field photovoltaic (PV) system uses to interconnect its data loggers and inverters and manipulate certain measurements collected from the network or trick existing irradiance and inverter readings through false data injection attacks (FDIA). Control routines that rely on these measurements can propagate the false data, impacting critical decisions that result in a suboptimal operation or even cause intentional harm leading to inverter-tripping or unscheduled loads that need to be shed. To detect FDIA in PV systems, the paper introduces an attention-based graph neural network with node embeddings and applied it to a simple prototypical DC-coupled microgrid with PV, energy storage, and load. The algorithm shows a detection accuracy of up to 98.95%. The proposed FDIA detection technique will provide micro-grid operators with an effective method to safeguard their systems, guaranteeing the secure and reliable operation.

Parvez, Imtiaz [Utah Valley University]↗

A Practical Comparison of Data-Driven Prognostics Methods for Energy Systems

This study explores data-driven prognostics for nuclear power plant (NPP) condensers, focusing on tube fouling. We utilized the Asherah nuclear power plant simulator (ANS) to compare four methods: Random Forest (RF), Support Vector Regressor (SVR), Fully Connected Neural Network (FCNN), and Long Short-Term Memory Neural Network (LSTM). By simulating various fouling scenarios in the ANS, we generated data with different degradation rates under transient operations. The models were trained and tested on these data, with performance evaluated visually and numerically including uncertainty assessment. The LSTM model excelled, exhibiting minimal prediction noise and the most accurate remaining useful life estimates across all degradation levels. Its ability to capture long-term dependencies and produce cleaner outputs makes it a strong candidate, although accurate training data across the entire component lifespan are crucial. The RF model emerged as a robust alternative, providing reliable predictions with high confidence. The FCNN and SVR models, while less effective overall, showed potential under specific conditions. FCNN offers a less complex alternative to LSTM and might benefit from larger datasets. SVR excels in precision when the quality of the training data is high. Furthermore, this study highlights the operational benefits of advanced prognostics in the energy sector and emphasizes the need for further research in NPP condenser health management through real-life experiments.

21 SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLANTS↗

NN-OpInf

SAND2026-18878O The NN-OpInf tool is a PyTorch-based approach to operator inference that uses composable, structure-preserving neural networks to represent nonlinear operators. Operator inference is a machine learning method for inferring low-dimensional systems from data and polynomial models for system dynamics. However, many systems do not conform to polynomial structures, which NN-OpInf addresses by parameterizing operators with neural networks. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy's National Nuclear Security Administration under contract DE-NA0003525.

SciDAC↗

Quality-Controlled Meteorological Data from the Flood Control District of Maricopa County (FCDMC) Network, Phoenix, Arizona (1987-2024)

This dataset contains 15- or 30-minute interval meteorological data from the Flood Control District of Maricopa County (FCDMC), Arizona, USA, covering eight key variables across multiple sensor stations between 1987 and 2024. Each variable is stored as a separate CSV file, containing time-series data that have undergone rigorous quality control (QC) procedures and, where appropriate, short-gap interpolation for consistency. The quality control (QC) pipeline consisted of four sequential tests: (1) a range test to ensure all values fall within physically realistic limits, (2) a step test to identify abrupt and implausible changes between consecutive records, (3) a proximity test that validates flagged values from step test using data from nearby stations and exceedance probability thresholds, and (4) a persistence test to detect and remove periods of unrealistically constant readings. These thresholds were calibrated to Arizona’s environmental conditions and sensor specifications. After QC, short gaps (≤2 hours) were linearly interpolated to ensure consistent temporal resolution, except for wind variables. Due to a major upgrade in FCDMC’s data transmission system, only ALERT-2 protocol data (2016–2024) for wind variables are included; earlier ALERT-1 data were excluded because of irregular sampling and high missing rates. This dataset supports regional climate and infrastructure resilience studies by providing standardized, high-resolution meteorological data for the greater Phoenix metropolitan area.

54 ENVIRONMENTAL SCIENCES↗

Calculating beam extinction in a pulsed proton beam using FPGA-based peak detection

The Mu2e experiment at Fermilab imposes stringent requirements on the elimination of out-of-time beam in its pulsed proton beam, a requirement known as “extinction”. Utilizing a new μTCA-based FPGA data acquisition system, we recorded live particle data from scattered particles incident on an array of quartz Cherenkov radiators and photomultiplier tubes to measure the extinction in the inter-pulse gaps in the pulsed proton beam. Minuscule errors in the derived signal period can make a measurement of the extinction impossible, so after taking a Fourier transform, further optimizations on the period were done based on the assumption that the signal period is stable over the full time of the beam spill while it is being resonantly extracted. After these optimizations, the beam extinction was shown to be on the level of 10^3.

Hensley, Ryan [UC, Davis]↗

Two-terminator RF adapter for background/environment noise measurement

A two-terminator RF adapter for background noise measurement in a test environment comprises a system test port comprising a system test port termination and a system test port connector to connect to a system under test; and a data acquisition port, comprising a data acquisition port termination and a data acquisition port connector to connect to a data acquisition system.

Kolski, Jeffrey↗

RE-INTEGRATE EMT Simulation Tool: Input Data Processing Layer for Bulk Power System

This paper introduces an advanced input data processing layer for EMT simulations of large-scale bulk power systems. The paper proposes two versions of the RE-INTEGRATE EMT simulation tool, RE-INTEGRATE Gen-0 and RE-INTEGRATE Gen-1, which are developed to enhance simulation generalizability, scalability, and accuracy. The framework leverages a generic class design for components to incorporate linear equations, which are generated by discretizing the Differential-Algebraic Equations (DAEs) that represent the dynamics of the components. In addition, the framework employs a parsing algorithm that parses a power system’s raw and dyr files to generate a connectivity graph which is then traversed to form the overall system’s dynamics. The proposed input data processing layer is used to simulate the IEEE 39-bus test system. The obtained results demonstrate the framework’s capability to achieve simulation scalability and accuracy. Further, the results indicate that EMT simulations performed using the proposed automations can effectively handle complex grid configurations.

Mishra, Rahul [ORNL] (ORCID:0000000328205932)↗

Community Geothermal: Soil Conductivity, Borehole Design, Energy Models, and Load Data for a Residential System Development - Hinesburg, VT

This dataset contains materials from the Coalition for Community-Supported Affordable Geothermal Energy Systems (C2SAGES) project, which evaluated the techno-economic feasibility of a community geothermal system for a residential development in Hinesburg, VT. The dataset includes detailed soil conductivity test reports, energy models, borehole design reports, hourly energy loads for heating, cooling, and hot water, and design layouts. EnergyPlus was used to model building energy loads, and Modelica software was applied for geothermal loop sizing based on these loads and soil conductivity results. Python scripts for network design further refined the models. Key files include PDF reports on borehole design (with projections for 1-year, 15-year, and 30-year systems), soil conductivity test results, EnergyPlus modeling outputs, and 2D/3D design drawings in PDF, DWG, and DXF formats. Python notebooks for network design and OnePipe model files are also provided, with Modelica required for viewing certain files. Outputs and modeling data are in various formats including CSV, JPG, HTML, and IDF, with units and data clearly labeled to support understanding of system design and performance for the proposed geothermal solution.

15 GEOTHERMAL ENERGY↗

NLR HPC Facility Power Usage Effectiveness (PUE) Data

Timeseries of Energy Systems Integration Facility (ESIF) Data Center Power Usage Effectiveness (PUE) Data provided in Parquet and compressed CSV formats Power Metrics Timeseries Fields: ts: Timestamp cooling_kw: Cooling (kilowatts) - Captures the power used by fans and pipe trace heaters associated with outdoor cooling equipment. The dedicated tower filter pump power is also captured as cooling load. energy_reuse: Energy Reuse Effectiveness hvac_kw: Heating, ventilation, and air conditioning (kilowatts) - Captures fan walls, fan coils that support the data center electrical rooms, and the make-up air unit. it_power_kw: IT equipment (kilowatts) - Captures power used by the IT equipment on the data center floor. plug_and_light_kw: Lights and utility plugs (kilowatts) - Captures power associated with the data center and dedicated mechanical room. The crank-case heater for the emergency standby generator is also captured as light and plug load. pue: Power Usage Effectiveness pump_kw: Pumps (kilowatts) - Captures power from pumps that move water in the data center Energy Recover Water loop and the Tower Water loops, and also captures power used by the boost pumps that circulate water through the fan walls. Note: The tower filter pump runs constantly to filter water from the data center cooling tower system, so 2.67 kilowatts are attributed to this pump and that is not reflected in this data field. day: Day of month Outside Weather Station Timeseries Fields: ts: Timestamp outside_air_humidity: Outside air humidity - Relative humidity percent outside_air_temp: Outside air temperature - Degrees Fahrenheit day: Day of month More detail: High-Performance Computing Data Center Power Usage Effectiveness

97 MATHEMATICS AND COMPUTING↗

Detecting anomalous SRF cavity behavior with unsupervised learning

We present an unsupervised learning framework for detecting anomalous superconducting radio-frequency (SRF) cavity behavior at the Continuous Electron Beam Accelerator Facility (CEBAF), emphasizing its initial performance and effectiveness. Key to the system’s success was the development of data acquisition systems (DAQs) that capture fast-sampled, information-rich signals, essential for detecting transient effects. The approach involves creating daily cavity-specific models using principal component analysis to handle variations in rf signal behavior and mitigate performance degradation from data drift. This unsupervised method eliminates the need for expensive labeling by continuously updating models with recent data. Deployed and operational for 3 months before a scheduled shutdown, the system successfully identified several issues with DAQ signals, confirming its effectiveness. Despite access to only a fraction of CEBAF’s SRF cavity signals, the framework efficiently detected several instances requiring intervention, demonstrating a significant improvement over traditional, labor-intensive methods of manual plot inspection. Published by the American Physical Society 2025

43 PARTICLE ACCELERATORS↗

Bayesian Physics Informed Spatio-Temporal Network for Streamflow Data Imputation

Reliable reconstruction of incomplete streamflow records is critical for improving hydrological forecasting, flood preparedness, and water resource management. However, large observational gaps and uncertainties in governing physical parameters limit the accuracy of traditional statistical and machinelearning imputation frameworks. To address these challenges, we develop a Bayesian Physics-Informed Spatio-Temporal Network (BPI-STNet) that jointly captures spatial and temporal dependencies while enforcing hydrologic consistency through embedded physical constraints. The framework integrates a GraphSAGE-LSTM architecture to model spatial connectivity across gauges and temporal flow dynamics, coupled with a Bayesian update mechanism to estimate uncertain parameters in a simplified water-balance framework. Unlike conventional physics-informed networks that rely on sampling-based posterior estimation, BPI-STNet derives an analytic solution to the inverse problem, allowing closed-form Bayesian updates of uncertain parameters Λ={α,β,k} using Gaussian priors and likelihoods. Applied to daily observations from the Susquehanna River Basin (1980-2022), BPI-STNet achieves substantial improvements over a purely data-driven RGNN baseline, which reduced RMSE by 23 % and MAE by 9 %, and achieving an average NSE values up to 0.96. The results demonstrate that coupling Bayesian inference with physics-informed learning yields physically consistent, uncertainty-aware reconstructions that preserve the temporal persistence and statistical distribution of observed flows. The proposed framework establishes a generalizable paradigm for data-sparse hydrologic systems where both data fidelity and physical interpretability are essential.

Krishnan Kutty Ambika, Anukesh [ORNL] (ORCID:00000↗