Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “data processing and logging”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Utah FORGE: 16B(78)-32 RFS DSS Strain Change Rate vs. Depth During 16A(78)-32 Stimulation

This dataset contains strain change rate versus depth data acquired using a Rayleigh frequency shift (RFS) distributed strain sensing (DSS) system during hydraulic stimulation of well 16A(78)-32 at the Utah FORGE site in April 2024. The data were collected from an optical fiber installed in the annulus of production well 16B(78)-32, approximately 300 feet from the injection well. The dataset includes tabulated strain data and an explanation of the methodology used to generate the frac log, which integrates strain change rate signals over selected time windows to identify fracture events.

15 GEOTHERMAL ENERGY↗

A semi–automatic analytical methodology for characterizing the energy consumption of MRI systems using load duration curves

Background and purpose: Magnetic resonance imaging (MRI) scanners are a major contributor to greenhouse gas emissions from the healthcare sector, and efforts to improve energy efficiency and reduce energy consumption rely on quantification of the characteristics of energy consumption. The purpose of this work was to develop a semi-automatic analytical methodology for the characterization of the energy consumption of MRI systems using only the load duration curve (LDC). LDCs are a fundamental tool used across various fields to analyze and understand the behavior of loads over time. Methods: An electric current transformer sensor and data logger were installed on two 3T MRI scanners from two vendors, termed M1 (outpatient scanner) and M2 (inpatient/emergency scanner). Data was collected for 1 month (7/11/2023 to 8/11/2023). Active power was calculated, assuming a balanced three-phase system, using the average current measured across all three phases, a 480 V reference voltage for both machines, and vendor-provided power factors. An LDC was constructed for each system by sorting the active power values in descending order and computing the cumulative time (in units of percentage) for each data point. The first derivative of the LDC was then computed (LDC’), smoothed by convolution with a window function (sLDC’), and used to detect transitions between different system modes including (in descending power levels): scan, prepared-to-scan, idle, low-power, and off. The final, segmented LDC was used to measure time (% total time), total energy (kWh), and mean power (kW) for each system mode on both scanners. The method was validated by comparing mean power values, computed using the segmented 1-month LDC, for each nonproductive system mode (i.e., prepared-to-scan, idle, lower-power, and off) against power levels measured after a deliberate system shutdown was performed for each scanner (1 day worth of data). Results: The validation revealed differences in mean power values <1.4% for all nonproductive modes and both scanners. In the scan system mode, the mean power values ranged from 29.8 to 37.2 kW and the total energy consumed for 1 month ranged from 11 106 to 14 466 kWh depending on the scanner. Over the course of 1 month, the portion of time the scanners were in nonproductive modes ranged from 76% to 80% across scanners and the nonproductive energy consumption ranged from 8010 to 6722 kWh depending on the scanner. The M1 (outpatient) scanner consumed 99.9 and 183.9 kWh/day in idle mode for weekdays and weekends, respectively, because the scanner spent 23% more time proportionally in idle mode on the weekends. Conclusions: A semi-automatic method for quantifying energy consumption characteristics of MRI scanners was introduced and validated. This method is relatively simple to implement as it requires only power data from the scanners and avoids the technical challenges associated with extracting and processing scanner log files. Finally, the methodology enables quantitative evaluation of the power, time, and energy characteristics of MRI scanners in scan and nonproductive system modes, providing baseline data and the capability of identifying potential opportunities for enhancing the energy efficiency of MRI scanners.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

TEAMER: Twin Ocean Power Wave Energy Converter Comprehensive Overview

These files collectively provide a comprehensive overview of the testing process, data analysis, and validation for the Twin Ocean Power device tested at the O.H. Hinsdale Wave Research Laboratory, supported by TEAMER funding. This resource includes an overview of power results for a series of 7 trials. The files included in this comprehensive overview include a comprehensive log sheet for each trial, a summary of all trials, and processing scripts for the raw data. It includes all raw data in .tsv and MATLAB compatible formats, an average power chart, angular velocity charts for each trial, trial metrics, and power output files. This resource includes images of the Twin Ocean Power Wave Energy Converter device components and movement during testing and video recordings of each trial.

16 TIDAL AND WAVE POWER↗

Leveraging Natural Language Processing and Generative Models in Molecular Chemistry: Property Prediction and Novel Compound Generation

The accurate prediction of molecular properties is important for the rational design and the advancement of green chemistry and sustainable materials research. However, the predictive power of traditional computational chemistry methods is limited due to computational restrictions. Here, in this study, we examine an alternative approach to the accurate prediction of properties of organic compounds: natural language processing (NLP)-based molecular embedding. Using viscosity, partition coefficient (log P), and enthalpy of vaporization as test properties through a survey of comprehensive datasets comprising 5695 data points for viscosity, 25 870 data points for log P, and 2296 data points for enthalpy of vaporization. These are important properties for the design of greener, safer, and sustainable chemical processes. Models were trained using NLP methods such as Mol2vec and fine-tuned ChemBERTa, and results were compared with traditional input featurization techniques such as Morgan fingerprints and quantum chemistry derived sigma profiles and DFT features. Among the various machine learning models, Mol2vec demonstrated superior predictive capabilities, achieving the highest correlation coefficient (R 2 = 0.945) and lowest RMSE (0.106 mPa s) for viscosity, as well as high accuracy for log P and enthalpy of vaporization predictions. These findings establish the Mol2vec featurization technique, graph-convolutional neural networks (GCNN), and fine-tuned ChemBERTa model as powerful tools for predictive modeling of organic compounds properties, offering a significant improvement over previously used featurization techniques and opening up strategies for very-high-throughput computational screening. Finally, we integrated ML models with hybrid language-model-based generative adversarial networks (LM-GAN) to generate novel molecular sequences with desirable properties for different research applications. The ability to computationally design solvents with lower viscosity, lower log P, and lower enthalpy of vaporization offers a data-driven route to accelerating the discovery of sustainable alternatives to traditionally toxic solvents.

ChemBERTa↗

Informing Plant Asset Reliability and Availability Through AI-Driven Analysis of Operator Logs

The availability and reliability of nuclear power plant (NPP) structures, systems, and components (SSCs) are critical parameters for NPP safety. Tracking these parameters is necessary but costly and labor-intensive, requiring the collection and evaluation of SSC event data such as shutdowns, startups, and failures. To show how these events are needed for the parameters an example is given: one measure of reliability is based on the number of equipment failure events and the number of run hours (i.e., the time from a startup event to a shutdown event). Here, this work investigates using artificial intelligence (AI) to mine NPP operator log entry texts for SSC event data. Four AI approaches were explored for identifying these events, including natural language processing (NLP) methods, generative AI, generative AI combined with NLP, and topic modeling. A key challenge addressed with all four approaches is the brevity of operator log entries. Among these four a neural network–based NLP method was shown to be the most promising for this application, achieving F1 scores of 86.0% for shutdowns, 92.2% for startups, and 80.4% for failures on a subject-matter-expert-curated dataset from NPP operator logs, compared to a baseline of 66.6% for a random classifier. This shows that NLP methods can perform better than generative AI. Additionally, the NLP methods combined with generative AI were shown to perform better than generative AI alone. Generative AI was most successful at providing the background information for the NLP methods to use. This work demonstrates the potential to use AI to automate parameter collection from NPP operator log entries and other records.

97 - MATHEMATICS AND COMPUTING↗

Summit Darshan Archival Dataset

Summit Darshan Archival Dataset contains 2021 Summit Darshan log data for 25 applications and is grouped into science domains. The dataset is processed, and all the propriety fields are anonymized. The resultant data is converted into a tabular structure and saved in parquet file format. In this notebook, we demonstrate how to access the data. Data Organization: The data is organized into two directories: Darshan total (`darshan_total`): List all the high levels generated by the `darshan-parser --total` command on `.darshan` files. There is one parquet file for each application. Note: `uid` and `exe` field are masked Darshan detail (`darshan_detail`): This data contains detailed job level log information extracted by command `darshan-parser` on the raw `.darshan` files. The data is sorted by directory hierarchy in the order of `year/month/day (2021/12/07)`. For instance, to get the data for a `job_id` 3819766 of application `App11`, which was executed on `2021-12-07`can be accessed as follows. Note:`uid` and `filename` fields are masked

97 MATHEMATICS AND COMPUTING↗

Are System Baselines within OT Environments Feasible?

Critical infrastructure stakeholders need to baseline their systems to understand expected protocol communications.Baseline behaviors may vary based on operational context.Expected operations during a maintenance window, for example, may be different from normal operations.Furthermore, constructing system baselines for Industrial Control Systems (ICS) is difficult and time-consuming.ICS processes generate artifacts expressed across heterogeneous data sources such as network and device logs. There needs to be a corpus of data in order to develop and compare methods that evaluate the feasibility, performance, and generality of approaches to construct baselines for ICS events. Standalone repositories of network packet captures are insufficient to develop methods to classify or recognize operational events expressed across multiple data sources. Moreover, static data corpora do not enable researchers to compare the impact of changing the underlying system for which a baseline is being constructed and this limits the ability to evaluate the performance of system baselines given system changes (e.g. patches, configuration, maintenance events). In order to address these limitations within the community, this talk intends to promote discussion about the state of the practice of constructing baselines. In this manner, we can continue to understand requirements within industry that are not being met by current approaches to baseline construction. This talk builds on two previous talks on the topic of system baselines for OT environments. First, Weaver co-presented at the RSA Conference ICS Sandbox with Dan Gunter. The talk confirmed the need within industry to construct baselines across multiple types of data sources relative to the semantics of specific business processes. Second, Weaver presented at IEEE Security and Privacy Workshop on Language-Theoretic Security.

02 PETROLEUM↗

Convergence of Emerging Technologies - EAGL Test Information

The Emergency Automatic Gunshot Detection and Lockdown (EAGL) system provides automatic, autonomous, and timely gunshot detection in both indoor and outdoor environments. This system uses both wired and wireless devices. Self-contained wireless EAGL sensors passively “listen” for gunshot events. These devices also perform a single, daily supervisory heartbeat (HB) function to include a device self-check with reporting capability. Transmissions are received by an assigned EAGL Gateway, which translates the RF sensor data to a PoE network format solely for use by the EAGL system server. The server then performs additional processes after data receipt, which include but are not limited to: event validation and logging, GUI presentation, notifications, and other independent operations.

47 OTHER INSTRUMENTATION↗

Language-Theoretic Data Analysis to Support ICS Protocol Baselining

Critical infrastructure stakeholders need to baseline their systems to understand expected protocol communications. Baseline behaviors may vary based on operational context. Expected operations during a maintenance window, for example, may be different from normal operations. Furthermore, constructing system baselines for Industrial Control Systems (ICS) is difficult and time-consuming. ICS processes generate artifacts expressed across heterogeneous data sources such as network traffic and device logs. This paper explores the hypothesis that such ICS artifacts form a language in the language-theoretic sense. From a theoretical perspective, the variety of implementations of ICS protocols and constrained environment of OT networks provide a rich application domain for language-theoretic approaches. We present several use cases related to the practical construction of system baselines: grammars for data fusion, language dialects for device fingerprinting, and security automata for system baselining

24 POWER TRANSMISSION AND DISTRIBUTION↗

HERO WEC Belt Test Data

The following submission includes raw and processed data from the 2024 Hydraulic and Electric Reverse Osmosis Wave Energy Converter (HERO WEC) belt tests conducted using NREL's Large Amplitude Motion Platform (LAMP). A description of the motion profiles run during testing can be found in the run log document. Data was collected using NREL's Modular Ocean Data AcQuisition (MODAQ) system in the form of TDMS files. Data was then processed using Python and MATLAB and converted to MATLAB workspace, parquet, and csv file formats. During Data processing, a low pass filter was applied to each array and the arrays were then resampled to common 10Hz timestamps. A MATLAB data viewer script is provided to quickly visualize these data sets. The following arrays are contained in each test data file: - Time: Unix seconds timestamp - Test_Time: Time in seconds since beginning of test - POS_OS_1001: Encoder position in degrees (the encoder is located on the secondary shaft of the spring return and is driven by the winch after a 4.5:1 gear reduction) - LC_ST_1001: Anchor load cell data in lbf - PRESS_OS_2002: Air spring pressure in psi This data set has been developed by the National Renewable Energy Laboratory, operated by Alliance for Sustainable Energy, LLC, for the U.S. Department of Energy (DOE) under Contract No. DE-AC36-08GO28308. Funding provided by the U.S. Department of Energy Office of Energy Efficiency and Renewable Energy Water Power Technologies Office.

16 TIDAL AND WAVE POWER↗

Condition-Based Maintenance of a Circulating Water System of a Canadian Nuclear Power Plant using Machine Learning and Statistical Tools

Canada Deuterium Uranium pressurized-heavy-water reactors (PHWR) are a type of nuclear power plant that generate clean and reliable energy. The scope of this work is to automate data analysis methodologies to inform a condition-based maintenance strategy of a circulating water system (CWS) of a PHWR. The multiunit CWS provides a continuous supply of water to cool steam condensers, even during transient scenarios, thereby improving the thermal efficiency. This work aims to develop a machine learning (ML) based approach to detect anomalies in heterogeneous data of a CWS in a PHWR to help inform a predictive maintenance strategy. The heterogeneous data include textual and numeric time series data for a PHWR. Natural-language-processing (NLP)-based models are used to analyze textual data contained in work orders and operator logs and an event-timeseries correlation detection method is applied to assist anomalies diagnoses for CWS. An ML model Robust Linear Model (RLM) is also used to remove the seasonal variations in the system variable distributions based on distributions of environmental variables. A machine learning model, Density-Based Spatial Clustering of Applications with Noise (DBSCAN), trained on both original data and data without any seasonal variations will then be used to detect if an anomaly exists. Thus, by moving to an automated methodology to detect, classify, and forecast anomalies, the maintenance strategy would be based on component condition instead of a time-based schedule.

97 - MATHEMATICS AND COMPUTING↗

Equipment Testing Environment (ETE) Process Specification

This document is intended to be utilized with the Equipment Test Environment being developed to provide a standard process by which the ETE can be validated. The ETE is developed with the intent of establishing cyber intrusion, data collection and through automation provide objective goals that provide repeatability. This testing process is being developed to interface with the Technical Area V physical protection system. The document will overview the testing structure, interfaces, device and network logging and data capture. Additionally, it will cover the testing procedure, criteria and constraints necessary to properly capture data and logs and record them for experimental data capture and analysis.

97 MATHEMATICS AND COMPUTING↗

CSEM Fluid Monitoring Methodology Using Real Data Examples

Conference presentation at International Meeting for Applied Geoscience & Energy (IMAGE), Houston, Texas, August 28 – September 1, 2023. Using field data from hydrocarbon and CO 2 applications, we illustrate the importance of a workflow and adaption to the target on hand. Verifying the geophysical acquisition and processing steps with 3D modeling and checking them against a 3D anisotropic log-derived model maintains confidence in the workflow and minimizes the influence on the data. This allows us to predict data validity and to certify the data with respect to the borehole logs.

20 FOSSIL-FUELED POWER PLANTS↗

Zero Resistance Ammetry (ZRA) Measurements of Sediment Electrochemical Gradients, Old Woman Creek, Ohio, USA, June–November 2022

This dataset contains raw and processed zero resistance ammetry (ZRA) measurements collected from wetland sediments at Old Woman Creek, a freshwater estuary on Lake Erie, Ohio, USA, between June and November 2022. Measurements were obtained using a vertically deployed electrode array positioned at multiple depths within the sediment profile to capture electrochemical gradients associated with microbial activity and sediment geochemistry. The raw dataset consists of parsed instrument log files containing timestamps, electrode pair identifiers, and measured electrical potential (mV). The processed dataset includes standardized and quality-controlled values with instrument saturation limits removed and timestamps converted to ISO 8601 format. Electrode line identifiers were mapped to physical depths, enabling interpretation of depth-resolved electrochemical gradients. Instrument saturation values (−2048, −2047, 2047, and 2048 mV) were identified as measurement limits and excluded from quantitative analyses. All data parsing, processing, and quality control steps are documented in an accompanying R Markdown script, ensuring full reproducibility from raw instrument logs to final datasets.

EARTH SCIENCE > AGRICULTURE > SOILS > ELECTRICAL C↗

MASK4 Test Campaign for Sandia WaveBot Device

This data and report details the findings from a wave tank test focused on production of useful work of a wave energy converter (WEC) device. The experimental system and test were specifically designed to validate models for power transmission throughout the WEC system. Additionally, the validity of co-design informed changes to the power take-off (PTO) were assessed and shown to provide the expected improvements in system performance. These data describe the "MASK4" wave tank test of the Sandia WaveBot device. The WaveBot device has been tested a number of times in different permutations at the US Navy's Maneuvering and Sea Keeping (MASK) basin. Each test in this series is referred to as MASK1, MASK2, etc. The WaveBot device was first tested in one degree of freedom (heave) in 2016. This MASK1 test focused primarily on system identification and modeling. After MASK1, major modifications were performed to improve the overall real-time control and measurement system, improve the heave drive train, and add surge and pitch degrees of freedom. The second set of testing, which was broken up in to two stages: MASK2A and MASK2B, focused on bench testing and closed-loop control performance as well as nonlinear modeling. MASK3 then focused on multi-input, multi-output modeling and control for maximization of electrical power. The attached report presents the results from MASK4, which focuses on detailed modeling of the power conversion chain and validation co-design principles by way of the introduction of a magnetic spring. The test log, report, and data from the MASK4 test of the WaveBot augmented with a tunable magnetic spring. Processing codes can be found at the Github link below.

16 TIDAL AND WAVE POWER↗

Darshan for HEP applications

Modern HEP workflows must manage increasingly large and complex data collections. HPC facilities may be employed to help meet these workflows’ growing data processing needs. However, a better understanding of the I/O patterns and underlying bottlenecks of these workflows is necessary to meet the performance expectations of HPC systems.Darshan is a lightweight I/O characterization tool that captures concise views of HPC application I/O behavior. It intercepts application I/O calls at runtime, records file access statistics for each process, and generates log files detailing application I/O access patterns.Typical HEP workflows include event generation, detector simulation, event reconstruction, and subsequent analysis stages. A study of the I/O behavior of the ATLAS simulation and filtering stage, and the CMS simulation workflow using Darshan is presented, including insights into the I/O operations and data access size.

Wang, Rui↗

A total of 19 months of daily weather logging on the US east coast: the WFIP3 event log

The Third Wind Forecast Improvement Project (WFIP3) is a multi-institutional field campaign designed to advance the understanding and prediction of the offshore atmospheric boundary layer along the US east coast. Extending from February 2024 through August 2025, WFIP3 combines long-term coastal and offshore measurements with targeted modeling and forecasting efforts. This data paper presents the WFIP3 event log, a curated record of 578 d of meteorological phenomena and field observations that complements the campaign's extensive high-frequency datasets. The event log provides both manually documented daily weather discussions and automatically derived indicators of atmospheric processes – including low-level jets, wind ramps, extreme wind veer, and weak wind conditions – based on observations from scanning lidars deployed at three coastal and offshore sites. The dataset offers structured metadata, standardized time and site identifiers, and consistent terminology to facilitate its integration with WFIP3's observational and modeling data products. The log supports diverse applications, from model evaluation and forecast verification to the selection of case studies on offshore boundary-layer dynamics. The WFIP3 event log is publicly available through the US Department of Energy's Wind Data Hub, providing the research community with a transparent and enduring contextual reference for the interpretation and use of WFIP3 measurements.

17 WIND ENERGY↗

Mass Spectrometry Sample Submission Portal

Each step in the scientific process generates contextual information about the data that is important to consider when performing data integration, developing models of biological process, or training AI models. We will develop a flexible, template-driven tool that will log biological samples, capture metadata about those samples, and track the type(s) of analysis being performed by researchers providing samples for analysis by mass spectrometry.

97 MATHEMATICS AND COMPUTING↗