Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “data processing and logging”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Skylab Medical Data Center and Archives

The founding of the Skylab medical data center and archives as a central area to house medical data from space flights is described. Skylab program strip charts, various daily reports and summaries, experiment reports and logs, status report on Skylab data quality, raw data digital tapes, processed data microfilm, and other Skylab documents are housed in the data center. In addition, this memorandum describes how the data center acted as a central point for the coordination of preflight and postflight baseline data and how it served as coordinator for all data processing through computation and analysis. Also described is a catalog identifying Skylab medical experiments and all related data currently archived in the data center.

Spross, F. R.↗

Multichannel Networked Phasemeter Readout and Analysis

Netmeter software reads a data stream from up to 250 networked phasemeters, synchronizes the data, saves the reduced data to disk (after applying a low-pass filter), and provides a Web server interface for remote control. Unlike older phasemeter software that requires a special, real-time operating system, this program can run on any general-purpose computer. It needs about five percent of the CPU (central processing unit) to process 20 channels because it adds built-in data logging and network-based GUIs (graphical user interfaces) that are implemented in Scalable Vector Graphics (SVG). Netmeter runs on Linux and Windows. It displays the instantaneous displacements measured by several phasemeters at a user-selectable rate, up to 1 kHz. The program monitors the measure and reference channel frequencies. For ease of use, levels of status in Netmeter are color coded: green for normal operation, yellow for network errors, and red for optical misalignment problems. Netmeter includes user-selectable filters up to 4 k samples, and user-selectable averaging windows (after filtering). Before filtering, the program saves raw data to disk using a burst-write technique.

Edmonds, Karina↗

A semi–automatic analytical methodology for characterizing the energy consumption of MRI systems using load duration curves

Background and purpose: Magnetic resonance imaging (MRI) scanners are a major contributor to greenhouse gas emissions from the healthcare sector, and efforts to improve energy efficiency and reduce energy consumption rely on quantification of the characteristics of energy consumption. The purpose of this work was to develop a semi-automatic analytical methodology for the characterization of the energy consumption of MRI systems using only the load duration curve (LDC). LDCs are a fundamental tool used across various fields to analyze and understand the behavior of loads over time. Methods: An electric current transformer sensor and data logger were installed on two 3T MRI scanners from two vendors, termed M1 (outpatient scanner) and M2 (inpatient/emergency scanner). Data was collected for 1 month (7/11/2023 to 8/11/2023). Active power was calculated, assuming a balanced three-phase system, using the average current measured across all three phases, a 480 V reference voltage for both machines, and vendor-provided power factors. An LDC was constructed for each system by sorting the active power values in descending order and computing the cumulative time (in units of percentage) for each data point. The first derivative of the LDC was then computed (LDC’), smoothed by convolution with a window function (sLDC’), and used to detect transitions between different system modes including (in descending power levels): scan, prepared-to-scan, idle, low-power, and off. The final, segmented LDC was used to measure time (% total time), total energy (kWh), and mean power (kW) for each system mode on both scanners. The method was validated by comparing mean power values, computed using the segmented 1-month LDC, for each nonproductive system mode (i.e., prepared-to-scan, idle, lower-power, and off) against power levels measured after a deliberate system shutdown was performed for each scanner (1 day worth of data). Results: The validation revealed differences in mean power values <1.4% for all nonproductive modes and both scanners. In the scan system mode, the mean power values ranged from 29.8 to 37.2 kW and the total energy consumed for 1 month ranged from 11 106 to 14 466 kWh depending on the scanner. Over the course of 1 month, the portion of time the scanners were in nonproductive modes ranged from 76% to 80% across scanners and the nonproductive energy consumption ranged from 8010 to 6722 kWh depending on the scanner. The M1 (outpatient) scanner consumed 99.9 and 183.9 kWh/day in idle mode for weekdays and weekends, respectively, because the scanner spent 23% more time proportionally in idle mode on the weekends. Conclusions: A semi-automatic method for quantifying energy consumption characteristics of MRI scanners was introduced and validated. This method is relatively simple to implement as it requires only power data from the scanners and avoids the technical challenges associated with extracting and processing scanner log files. Finally, the methodology enables quantitative evaluation of the power, time, and energy characteristics of MRI scanners in scan and nonproductive system modes, providing baseline data and the capability of identifying potential opportunities for enhancing the energy efficiency of MRI scanners.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Millimeter-Wave Radar Field Measurements and Inversion of Cloud Parameters for the 1999 Mt. Washington Icing Sensors Project

The Mount Washington Icing Sensors Project (MWISP) was a multi-investigator experiment with participants from Quadrant Engineering, NOAA Environmental Technology Laboratory (NOAA/ETL), the Microwave Remote Sensing Laboratory (MIRSL) of the University of Massachusetts (UMass), and others. Radar systems from UMass and NOAA/ETL were used to measure X-, Ka-, and W-band backscatter data from the base of Mt. Washington, while simultaneous in-situ particle measurements were made from aircraft and from the observatory at the summit. This report presents range and time profiles of liquid water content and particle size parameters derived from range profiles of radar reflectivity as measured at X-, Ka-, and W-band (9.3, 33.1, and 94.9 GHz) using an artificial neural network inversion algorithm. In this report, we provide a brief description of the experiment configuration, radar systems, and a review of the artificial neural network used to extract cloud parameters from the radar data. Time histories of liquid water content (LWC), mean volume diameter (MVD) and mean Z diameter (MZD) are plotted at 300 m range intervals for slant ranges between 1.1 and 4 km. Appendix A provides details on the extraction of radar reflectivity from measured radar power, and Appendix B provides summary logs of the weather conditions for each day in which we processed data.

Pazmany, Andrew L.↗

TEAMER: Twin Ocean Power Wave Energy Converter Comprehensive Overview

These files collectively provide a comprehensive overview of the testing process, data analysis, and validation for the Twin Ocean Power device tested at the O.H. Hinsdale Wave Research Laboratory, supported by TEAMER funding. This resource includes an overview of power results for a series of 7 trials. The files included in this comprehensive overview include a comprehensive log sheet for each trial, a summary of all trials, and processing scripts for the raw data. It includes all raw data in .tsv and MATLAB compatible formats, an average power chart, angular velocity charts for each trial, trial metrics, and power output files. This resource includes images of the Twin Ocean Power Wave Energy Converter device components and movement during testing and video recordings of each trial.

16 TIDAL AND WAVE POWER↗

Leveraging Natural Language Processing and Generative Models in Molecular Chemistry: Property Prediction and Novel Compound Generation

The accurate prediction of molecular properties is important for the rational design and the advancement of green chemistry and sustainable materials research. However, the predictive power of traditional computational chemistry methods is limited due to computational restrictions. Here, in this study, we examine an alternative approach to the accurate prediction of properties of organic compounds: natural language processing (NLP)-based molecular embedding. Using viscosity, partition coefficient (log P), and enthalpy of vaporization as test properties through a survey of comprehensive datasets comprising 5695 data points for viscosity, 25 870 data points for log P, and 2296 data points for enthalpy of vaporization. These are important properties for the design of greener, safer, and sustainable chemical processes. Models were trained using NLP methods such as Mol2vec and fine-tuned ChemBERTa, and results were compared with traditional input featurization techniques such as Morgan fingerprints and quantum chemistry derived sigma profiles and DFT features. Among the various machine learning models, Mol2vec demonstrated superior predictive capabilities, achieving the highest correlation coefficient (R 2 = 0.945) and lowest RMSE (0.106 mPa s) for viscosity, as well as high accuracy for log P and enthalpy of vaporization predictions. These findings establish the Mol2vec featurization technique, graph-convolutional neural networks (GCNN), and fine-tuned ChemBERTa model as powerful tools for predictive modeling of organic compounds properties, offering a significant improvement over previously used featurization techniques and opening up strategies for very-high-throughput computational screening. Finally, we integrated ML models with hybrid language-model-based generative adversarial networks (LM-GAN) to generate novel molecular sequences with desirable properties for different research applications. The ability to computationally design solvents with lower viscosity, lower log P, and lower enthalpy of vaporization offers a data-driven route to accelerating the discovery of sustainable alternatives to traditionally toxic solvents.

ChemBERTa↗

The Mobile Agents Integrated Field Test: Mars Desert Research Station April 2003

The Mobile Agents model-based, distributed architecture, which integrates diverse components in a system for lunar and planetary surface operations, was extensively tested in a two-week field "technology retreat" at the Mars Society s Desert Research Station (MDRS) during April 2003. More than twenty scientists and engineers from three NASA centers and two universities refined and tested the system through a series of incremental scenarios. Agent software, implemented in runtime Brahms, processed GPS, health data, and voice commands-monitoring, controlling and logging science data throughout simulated EVAs with two geologists. Predefined EVA plans, modified on the fly by voice command, enabled the Mobile Agents system to provide navigation and timing advice. Communications were maintained over five wireless nodes distributed over hills and into canyons for 5 km; data, including photographs and status was transmitted automatically to the desktop at mission control in Houston. This paper describes the system configurations, communication protocols, scenarios, and test results.

Clancey, William J.↗

Informing Plant Asset Reliability and Availability Through AI-Driven Analysis of Operator Logs

The availability and reliability of nuclear power plant (NPP) structures, systems, and components (SSCs) are critical parameters for NPP safety. Tracking these parameters is necessary but costly and labor-intensive, requiring the collection and evaluation of SSC event data such as shutdowns, startups, and failures. To show how these events are needed for the parameters an example is given: one measure of reliability is based on the number of equipment failure events and the number of run hours (i.e., the time from a startup event to a shutdown event). Here, this work investigates using artificial intelligence (AI) to mine NPP operator log entry texts for SSC event data. Four AI approaches were explored for identifying these events, including natural language processing (NLP) methods, generative AI, generative AI combined with NLP, and topic modeling. A key challenge addressed with all four approaches is the brevity of operator log entries. Among these four a neural network–based NLP method was shown to be the most promising for this application, achieving F1 scores of 86.0% for shutdowns, 92.2% for startups, and 80.4% for failures on a subject-matter-expert-curated dataset from NPP operator logs, compared to a baseline of 66.6% for a random classifier. This shows that NLP methods can perform better than generative AI. Additionally, the NLP methods combined with generative AI were shown to perform better than generative AI alone. Generative AI was most successful at providing the background information for the NLP methods to use. This work demonstrates the potential to use AI to automate parameter collection from NPP operator log entries and other records.

97 - MATHEMATICS AND COMPUTING↗

Contrasting Patterns of Damage and Recovery in Logged Amazon Forests From Small Footprint LiDAR Data

Tropical forests ecosystems respond dynamically to climate variability and disturbances on time scales of minutes to millennia. To date, our knowledge of disturbance and recovery processes in tropical forests is derived almost exclusively from networks of forest inventory plots. These plots typically sample small areas (less than or equal to 1 ha) in conservation units that are protected from logging and fire. Amazon forests with frequent disturbances from human activity remain under-studied. Ongoing negotiations on REDD+ (Reducing Emissions from Deforestation and Forest Degradation plus enhancing forest carbon stocks) have placed additional emphasis on identifying degraded forests and quantifying changing carbon stocks in both degraded and intact tropical forests. We evaluated patterns of forest disturbance and recovery at four -1000 ha sites in the Brazilian Amazon using small footprint LiDAR data and coincident field measurements. Large area coverage with airborne LiDAR data in 2011-2012 included logged and unmanaged areas in Cotriguacu (Mato Grosso), Fiona do Jamari (Rondonia), and Floresta Estadual do Antimary (Acre), and unmanaged forest within Reserva Ducke (Amazonas). Logging infrastructure (skid trails, log decks, and roads) was identified using LiDAR returns from understory vegetation and validated based on field data. At each logged site, canopy gaps from logging activity and LiDAR metrics of canopy heights were used to quantify differences in forest structure between logged and unlogged areas. Contrasting patterns of harvesting operations and canopy damages at the three logged sites reflect different levels of pre-harvest planning (i.e., informal logging compared to state or national logging concessions), harvest intensity, and site conditions. Finally, we used multi-temporal LiDAR data from two sites, Reserva Ducke (2009, 2012) and Antimary (2010, 2011), to evaluate gap phase dynamics in unmanaged forest areas. The rates and patterns of canopy gap formation at these sites illustrate potential issues for separating logging damages from natural forest disturbances over longer time scales. Multi-temporal airborne LiDAR data and coincident field measurements provide complementary perspectives on disturbance and recovery processes in intact and degraded Amazon forests. Compared to forest inventory plots, the large size of each individual site permitted analyses of landscape-scale processes that would require extremely high investments to study using traditional forest inventory methods.

Morton, D. C.↗

Summit Darshan Archival Dataset

Summit Darshan Archival Dataset contains 2021 Summit Darshan log data for 25 applications and is grouped into science domains. The dataset is processed, and all the propriety fields are anonymized. The resultant data is converted into a tabular structure and saved in parquet file format. In this notebook, we demonstrate how to access the data. Data Organization: The data is organized into two directories: Darshan total (`darshan_total`): List all the high levels generated by the `darshan-parser --total` command on `.darshan` files. There is one parquet file for each application. Note: `uid` and `exe` field are masked Darshan detail (`darshan_detail`): This data contains detailed job level log information extracted by command `darshan-parser` on the raw `.darshan` files. The data is sorted by directory hierarchy in the order of `year/month/day (2021/12/07)`. For instance, to get the data for a `job_id` 3819766 of application `App11`, which was executed on `2021-12-07`can be accessed as follows. Note:`uid` and `filename` fields are masked

97 MATHEMATICS AND COMPUTING↗

Are System Baselines within OT Environments Feasible?

Critical infrastructure stakeholders need to baseline their systems to understand expected protocol communications.Baseline behaviors may vary based on operational context.Expected operations during a maintenance window, for example, may be different from normal operations.Furthermore, constructing system baselines for Industrial Control Systems (ICS) is difficult and time-consuming.ICS processes generate artifacts expressed across heterogeneous data sources such as network and device logs. There needs to be a corpus of data in order to develop and compare methods that evaluate the feasibility, performance, and generality of approaches to construct baselines for ICS events. Standalone repositories of network packet captures are insufficient to develop methods to classify or recognize operational events expressed across multiple data sources. Moreover, static data corpora do not enable researchers to compare the impact of changing the underlying system for which a baseline is being constructed and this limits the ability to evaluate the performance of system baselines given system changes (e.g. patches, configuration, maintenance events). In order to address these limitations within the community, this talk intends to promote discussion about the state of the practice of constructing baselines. In this manner, we can continue to understand requirements within industry that are not being met by current approaches to baseline construction. This talk builds on two previous talks on the topic of system baselines for OT environments. First, Weaver co-presented at the RSA Conference ICS Sandbox with Dan Gunter. The talk confirmed the need within industry to construct baselines across multiple types of data sources relative to the semantics of specific business processes. Second, Weaver presented at IEEE Security and Privacy Workshop on Language-Theoretic Security.

02 PETROLEUM↗

On The Processing of Log Files for Monitoring Antenna Health

In order to improve the quality of geodetic results, we have developed an infrastructure for timely processing of telemetry from IVS observing stations. We check every hour for new log files with telemetry from both VLBI observing sessions, single dish experiments, and stow-in data collection and automatically process them. The telemetry data we use is the system temperature, phase calibration phases and amplitudes, system equivalent flux density, and the differences between formatter clock and GPS clock. For the system temperature and phase calibration, processing includes filtering out outliers and computing averages and rms of the scatter in each scan. Furthermore, for the phase calibration we also compute the group delay and detect spurious signals. Cleaned and post-processed telemetry is archived. Our process detects abnormalities, such as, anomalously high system temperature, unstable phase calibration phases, jumps in the GPS and formatter clock differences, and others. With our procedure, the latency of detection of station abnormalities is reduced to less than two hours. Early detection of abnormalities reduces the amount of affected data since station personnel get early alerts. We discuss our experience of running this system since 2022.

Phase Calibration↗

On The Processing of Log Files for Monitoring Antenna Health

In order to improve the quality of geodetic results, we have developed an infrastructure for timely processing of telemetry from IVS observing stations. We check every hour for new log files with telemetry from both VLBI observing sessions, single dish experiments, and stow-in data collection and automatically process them. The telemetry data we use is the system temperature, phase calibration phases and amplitudes, system equivalent flux density, and the differences between formatter clock and GPS clock. For the system temperature and phase calibration, processing includes filtering out outliers and computing averages and rms of the scatter in each scan. Furthermore, for the phase calibration we also compute the group delay and detect spurious signals. Cleaned and post-processed telemetry is archived. Our process detects abnormalities, such as, anomalously high system temperature, unstable phase calibration phases, jumps in the GPS and formatter clock differences, and others. With our procedure, the latency of detection of station abnormalities is reduced to less than two hours. Early detection of abnormalities reduces the amount of affected data since station personnel get early alerts. We discuss our experience of running this system since 2022.

VLBI↗

The CCD/Transit Instrument (CTI) data-analysis system

The automated software system for archiving, analyzing, and interrogating data from the CCD/Transit Instrument (CTI) is described. The CTI collects up to 450 Mbytes of image-data each clear night in the form of a narrow strip of sky observed in two colors. The large data-volumes and the scientific aims of the project make it imperative that the data are analyzed within the 24-hour period following the observations. To this end a fully automatic and self evaluating software system has been developed. The data are collected from the telescope in real-time and then transported to Tucson for analysis. Verification is performed by visual inspection of random subsets of the data and obvious cosmic rays are detected and removed before permanent archival is made to the optical disc. The analysis phase is performed by a pair of linked algorithms, one operating on the absolute pixel-values and the other on the spatial derivative of the data. In this way both isolated and merged images are reliably detected in a single pass. In order to isolate the latter algorithm from the effects of noise spikes a 3x3 Hanning filter is applied to the raw data before the analysis is run. The algorithms reduce the input pixel-data to a database of measured parameters for each image which has been found. A contrast filter is applied in order to assign a detection-probability to each image and then x-y calibration and intensity calibration are performed using known reference stars in the strip. These are added to as necessary by secondary standards boot-strapped from the CTI data itself. The final stages involve merging the new data into the CTI Master-list and History-list and the automatic comparison of each new detection with a set of pre-defined templates in parameter-space to find interesting objects such as supernovae, quasars and variable stars. Each stage of the processing from verification to interesting image selection is performed under a data-logging system which both controls the pipe-lining of data through the system and records key performance monitor parameters which are built into the software. Furthermore, the data from each stage are stored in databases to facilitate evaluation, and all stages offer the facility to enter keyword-indexed free-format text into the data-logging system. In this way a large measure of certification is built into the system to provide the necessary confidence in the end results.

Cawson, M. G. M.↗

Convergence of Emerging Technologies - EAGL Test Information

The Emergency Automatic Gunshot Detection and Lockdown (EAGL) system provides automatic, autonomous, and timely gunshot detection in both indoor and outdoor environments. This system uses both wired and wireless devices. Self-contained wireless EAGL sensors passively “listen” for gunshot events. These devices also perform a single, daily supervisory heartbeat (HB) function to include a device self-check with reporting capability. Transmissions are received by an assigned EAGL Gateway, which translates the RF sensor data to a PoE network format solely for use by the EAGL system server. The server then performs additional processes after data receipt, which include but are not limited to: event validation and logging, GUI presentation, notifications, and other independent operations.

47 OTHER INSTRUMENTATION↗

Improved Discrete Approximation of Laplacian of Gaussian

An improved method of computing a discrete approximation of the Laplacian of a Gaussian convolution of an image has been devised. The primary advantage of the method is that without substantially degrading the accuracy of the end result, it reduces the amount of information that must be processed and thus reduces the amount of circuitry needed to perform the Laplacian-of- Gaussian (LOG) operation. Some background information is necessary to place the method in context. The method is intended for application to the LOG part of a process of real-time digital filtering of digitized video data that represent brightnesses in pixels in a square array. The particular filtering process of interest is one that converts pixel brightnesses to binary form, thereby reducing the amount of information that must be performed in subsequent correlation processing (e.g., correlations between images in a stereoscopic pair for determining distances or correlations between successive frames of the same image for detecting motions). The Laplacian is often included in the filtering process because it emphasizes edges and textures, while the Gaussian is often included because it smooths out noise that might not be consistent between left and right images or between successive frames of the same image.

Shuler, Robert L., Jr.↗

Language-Theoretic Data Analysis to Support ICS Protocol Baselining

Critical infrastructure stakeholders need to baseline their systems to understand expected protocol communications. Baseline behaviors may vary based on operational context. Expected operations during a maintenance window, for example, may be different from normal operations. Furthermore, constructing system baselines for Industrial Control Systems (ICS) is difficult and time-consuming. ICS processes generate artifacts expressed across heterogeneous data sources such as network traffic and device logs. This paper explores the hypothesis that such ICS artifacts form a language in the language-theoretic sense. From a theoretical perspective, the variety of implementations of ICS protocols and constrained environment of OT networks provide a rich application domain for language-theoretic approaches. We present several use cases related to the practical construction of system baselines: grammars for data fusion, language dialects for device fingerprinting, and security automata for system baselining

24 POWER TRANSMISSION AND DISTRIBUTION↗

HERO WEC Belt Test Data

The following submission includes raw and processed data from the 2024 Hydraulic and Electric Reverse Osmosis Wave Energy Converter (HERO WEC) belt tests conducted using NREL's Large Amplitude Motion Platform (LAMP). A description of the motion profiles run during testing can be found in the run log document. Data was collected using NREL's Modular Ocean Data AcQuisition (MODAQ) system in the form of TDMS files. Data was then processed using Python and MATLAB and converted to MATLAB workspace, parquet, and csv file formats. During Data processing, a low pass filter was applied to each array and the arrays were then resampled to common 10Hz timestamps. A MATLAB data viewer script is provided to quickly visualize these data sets. The following arrays are contained in each test data file: - Time: Unix seconds timestamp - Test_Time: Time in seconds since beginning of test - POS_OS_1001: Encoder position in degrees (the encoder is located on the secondary shaft of the spring return and is driven by the winch after a 4.5:1 gear reduction) - LC_ST_1001: Anchor load cell data in lbf - PRESS_OS_2002: Air spring pressure in psi This data set has been developed by the National Renewable Energy Laboratory, operated by Alliance for Sustainable Energy, LLC, for the U.S. Department of Energy (DOE) under Contract No. DE-AC36-08GO28308. Funding provided by the U.S. Department of Energy Office of Energy Efficiency and Renewable Energy Water Power Technologies Office.

16 TIDAL AND WAVE POWER↗