Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “data processing and logging”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Oscilloscope Data Push Program

This paper details the development of a Python program designed to automate the data acquisition and conversion for an oscilloscope for the purposes of a one-off/temporary data acquisition system for users that readily need data, and do not have the option of obtaining a Data Acquisition (DAQ) solution. Creating DAQ systems for analyzing a system requires expensive electronics and a dedicated team of engineers for support. Traditionally, manual data collection and processing are time consuming and prone to error. By automating these processes, the cost, efficiency and accuracy of data handling are improved upon. This project involves the creation of a program that interacts with the oscilloscope. During this interaction, there are various functions being performed such as the acquisition of waveform data via floating points, generating plots with the acquired wave points, and storing of floating points in a CSV file format for future reference and plotting purposes. While the initial aim of the project included continuous logging to a cloud database, this was deferred due to time constraints. The results portrayed an almost-instant rate of data collection with a buffer time, showcasing the potential for further integration and real-time data processing.

Osei-Tutu, Jason↗

Access Patterns and Performance Behaviors of Multi-layer Supercomputer I/O Subsystems under Production Load

Scientific computing workloads at HPC facilities have been shifting from traditional numerical simulations to AI/ML applications for training and inference while processing and producing ever-increasing amounts of scientific data. To address the growing need for increased storage capacity, lower access latency, and higher bandwidth, emerging technologies such as non-volatile memory are integrated into supercomputer I/O subsystems. With these emerging trends, we need a better understanding of the multilayer supercomputer I/O systems and ways to use these subsystems efficiently. In this work, we study the I/O access patterns and performance characteristics of two representative supercomputer I/O subsystems. Through an extensive analysis of year-long I/O logs on each system, we report new observations in I/O reads and writes, unbalanced use of storage system layers, and new trends in user behaviors at the HPC I/O middleware stack.

Bez, JL↗

Pseudonymization at Scale: OLCF’s Summit Usage Data Case Study

The analysis of vast amounts of data and the processing of complex computational jobs have traditionally relied upon high performance computing (HPC) systems, which offer reliable and efficient management of large-scale computational and data resources. Understanding these analyses’ needs is paramount for designing solutions that can lead to better science, and similarly, understanding the characteristics of the user behavior on those systems is important for improving user experiences on HPC systems. A common approach to gathering data about user behavior is to extract workload characteristics from system log data available only to system administrators. Recently at Oak Ridge Leadership Computing Facility (OLCF), however, we unveiled user behavior about the Summit supercomputer by collecting data from a user’s point of view with ordinary Unix commands.In this paper, we discuss the process, challenges, and lessons learned while preparing this dataset for publication and submission to an open data challenge. The original dataset contains personal identifiable information (PII) about the users of OLCF which needed be masked prior to publication, and we determined that anonymization, which scrubs PII completely, destroyed too much of the structure of the data to be interesting for the data challenge. We instead chose to pseudonymize the dataset, which reduced the linkability of the dataset to the users’ identities. Pseudonymization is significantly more computationally expensive than anonymization, and the size of our dataset, which is approximately 175 million lines of raw text, necessitated the development of a parallelized workflow that could be reused on different HPC machines. We demonstrate the scaling behavior of the workflow on two leadership class HPC systems at OLCF, and we show that we were able to bring the overall makespan time from an impractical 20+ hours on a single node down to around 2 hours. As a result of this work, we release the entire pseudonymized dataset and make the workflows and source code publicly available.

Maheshwari, Ketan↗

pvOps: a Python package for empirical analysis of photovoltaic field data

The purpose of pvOps is to support empirical evaluations of data collected in the field related to the operations and maintenance (O&M) of photovoltaic (PV) power plants. pvOps presently contains modules that address the diversity of field data, including text-based maintenance logs, current-voltage (IV) curves, and timeseries of production information. The package functions leverage machine learning, visualization, and other techniques to enable cleaning, processing, and fusion of these datasets. These capabilities are intended to facilitate easier evaluation of field patterns and extraction of relevant insights to support reliability-related decision-making for PV sites. The open-source code, examples, and instructions for installing the package through PyPI can be accessed through the GitHub repository.

14 SOLAR ENERGY↗

G-LiHT Campaign Leaf Carbon and Nitrogen Content, Mar2017: Puerto Rico

Measurements of leaf carbon and nitrogen content collected from 68 tropical tree species. Data includes leaves collected from fully sunlit and shaded canopy strata as well as leaves for young, mature, old and senescent leaf ages. Data for each sample includes the relative age estimate, leaf canopy position and sample number. This data was collected as part of the 2017 NGEE-Tropics / NASA G-LiHT airborne campaign. This data package includes processed data for leaf carbon and nitrogen content (*.csv). Metadata files include data description (_dd.csv) for tabular data, site information (*.csv), sampling protocol (*.pdf) and the NGEE-Tropics FRAMES e-field log and file submission metadata (*.xlsx). See related datasets for sample details including photographs, leaf-level reflectance and transmittance spectra, leaf mass per area (LMA) and water content.

54 ENVIRONMENTAL SCIENCES↗

Event Log / Raw Data

The WFIP3 event log is a curated record spanning 578 days of meteorological phenomena and field observations that complements the campaign’s high-frequency measurements. The log combines manually documented daily weather discussions with automatically derived indicators of key atmospheric processes, providing standardized, publicly available context to support model evaluation, forecast verification, and case-study selection for offshore boundary-layer research.

17 WIND ENERGY↗

Utah FORGE: Optimization of a Plug-and-Perf Stimulation (Fervo Energy)

Information around the plug-and-perf treatment design at Utah FORGE by Fervo Energy. Objective and Purpose: - Develop a multistage hydraulic stimulation approach designed specifically to target the top three factors that control the technical and commercial viability of an EGS system: i) Achieving sufficient injectivity to support high cross-well flow rates ii) Distributing flow evenly across the wellbore and reservoir to maximize heat mining efficiency, ensure sustained heat transfer, and mitigate thermal breakthrough iii) Overcoming the effects of stress heterogeneity, stress shadowing, and variations in natural fracture properties during the stimulation treatment, leading to a more predictable stimulated reservoir volume and offset well placement - The following activities will be performed: i) Design, plan, and execute a multistage plug-and-perf stimulation treatment at a Fervo site with data acquisition and well testing activities aimed at addressing key technical aspects of the issues above ii) Perform data processing and interpretation of field results to translate the results form the Fervo site to a site-specific design at the Utah FORGE site iii) Design, plan, and execute a multistage plug-and-perf stimulation treatment design at the Utah FORGE site Methods and Approach: - Design a detailed data acquisition plan to maximize learning around: i) DFIT testing ii) Petrophysical logging, image logging iii) Permanent DAS/DTS fiber optic monitoring iv) Deep borehole microseismic monitoring v) Shallow borehole induced seismicity monitoring vi) Injection/production testing (RTA analysis, tracer testing) vii) Integrated numerical modeling and production forecasting

15 GEOTHERMAL ENERGY↗

tbsinp (00)

The ice nucleation spectrometer (INS) is an offline analytical measurement system used to process filter samples for freezing temperature spectra of immersion-mode ice-nucleating particle (INP) number concentrations. It is almost identical to the Colorado State University (CSU) ice spectrometer design. The INS-AIR is specific to filter samples collected on aerial platforms, such as the DOE ARM tethered balloon system (TBS) operated by Sandia National Laboratories, using miniaturized aerosol filter samplers. Users are referred to the INS instrument handbook for INS-AIR sampler details. INS-AIR filter samples are collected during intensive operational periods at various ARM sites and then processed on the INS at CSU. This filter log contains the detailed metadata at all ARM sites where INP filter sample collection has occurred or is currently ongoing, including both ground-based routine INP sampling (INS) and TBS INP sampling (INS-AIR). Metadata include start and end times, vacuum line pressures and temperatures, and flow rates; total accumulated flow through each filter; and notes on collection issues or weather conditions. Users can also keep up to date with the status of filter and data processing, even before data are available on ARM’s Data Discovery. Users can contact INP mentors Jessie Creamean or Thomas Hill with any questions.

54 ENVIRONMENTAL SCIENCES↗

TEAMER - AquaHarmonics High Fidelity WEC Sim PTO and Control Model Validation, Test Logs and Results

Collaborative effort between AquaHarmonics, Sandia National Laboratories (SNL), and the National Renewable Energy Laboratory (NREL) to revise and validate Aquaharmonics' full wave to wire model, allowing for reduced uncertainty and increased understanding of design requirements of a utility scale wave energy converter (WEC). SNL and NREL in collaboration with AquaHarmonics, will set up and run WEC Simulator (WEC-Sim) models of the AquaHarmonics WEC, building off past model developments for inclusion of custom PTO (power take-off) dynamics. The intent is to review, update, and verify or validate a new WEC-Sim model against wave tank experimental data. Furthermore, the WEC-Sim model will be coupled to an energy storage system model to better understand the wave-to-wire functionality. This data set is described in the "Test Log" excel file. Please refer to that document for details on each specific test date/time, constraint parameters and model hardware setup details. Sim model can be found in the associated MHKDR link below.

16 TIDAL AND WAVE POWER↗

Application of Principal Component Analysis to Electrochemical Reprocessing PM and NMAC

In this report, data from an electrorefiner (ER) for nuclear fuel reprocessing is evaluated for process monitoring (PM) conclusions. This data comes from tests performed at the Idaho National Laboratory in 2022. Multivariate approaches utilizing methods of Principal Component Analysis (PCA) is applied. This is based off established work in process monitoring for fault detection in industrial facilities. This report will discuss the background, methods, and results of the application and some of the conclusions and applications that can be drawn from them. PCA is applied to two different electrorefiner (ER) operations that occurred at Idaho National Laboratory between August and October 2022. The first operation occurred with little incident while the second had several noted faults in the equipment in operational logs. The data from the first run was used to train the data for “normal” operations and applied to both sets of data to determine when operations were in an “off-normal” condition and identify where the fault occurs through PCA. PCA was able to identify off-normal events and identify the cause for off-normal operations. These identified off-normal events matched with the events and their causes in the operational logs. However, small amounts of variance in the data led to false detection of “off-normal” events. Thus, careful selection of training data and a-posteriori conclusions based off operator assessments will both be required for application of PCA to PM applications. This work demonstrated that multivariate approaches and latent variables are applicable to pyroprocessing PM applications and can be further expanded in future work as quality variables such as salt concentration from sensors and sampling become available.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Continuous precipitation‐filtration process for initial capture of a monoclonal antibody product using a four‐stage countercurrent hollow fiber membrane washing step

The significant increase in product titers, coupled with the growing focus on continuous bioprocessing, has renewed interest in using precipitation as a low‐cost alternative to Protein A chromatography for the primary capture of monoclonal antibody (mAb) products. In this work, a commercially relevant mAb was purified from clarified cell culture fluid using a tubular flow precipitation reactor with dewatering and washing provided by tangential flow microfiltration. The particle morphology was evaluated using an inline high‐resolution optical probe, providing quantitative data on the particle size distribution throughout the precipitation process. Data were obtained in both a lab‐built 2‐stage countercurrent washing system and a commercial countercurrent contacting skid that provided 4 stages of continuous washing. The processes were operated continuously for 2 h with overall mAb yield of 92 ± 3% and DNA removal of nearly 3 logs in the 4‐stage system. The high DNA clearance was achieved by selective redissolution of the mAb using a low pH acetate buffer. Host cell protein clearance was 0.59 ± 0.08 logs, comparable to that based on model predictions. The process mass intensity was slightly better than typical Protein A processes and could be significantly improved by preconcentration of the antibody feed material.

59 BASIC BIOLOGICAL SCIENCES↗

Zero-truncated Poisson regression for sparse multiway count data corrupted by false zeros

Abstract We propose a novel statistical inference methodology for multiway count data that is corrupted by false zeros that are indistinguishable from true zero counts. Our approach consists of zero-truncating the Poisson distribution to neglect all zero values. This simple truncated approach dispenses with the need to distinguish between true and false zero counts and reduces the amount of data to be processed. Inference is accomplished via tensor completion that imposes low-rank tensor structure on the Poisson parameter space. Our main result shows that an $N$-way rank-$R$ parametric tensor $\boldsymbol{\mathscr{M}}\in (0,\infty )^{I\times \cdots \times I}$ generating Poisson observations can be accurately estimated by zero-truncated Poisson regression from approximately $IR^2\log _2^2(I)$ non-zero counts under the nonnegative canonical polyadic decomposition. Our result also quantifies the error made by zero-truncating the Poisson distribution when the parameter is uniformly bounded from below. Therefore, under a low-rank multiparameter model, we propose an implementable approach guaranteed to achieve accurate regression in under-determined scenarios with substantial corruption by false zeros. Several numerical experiments are presented to explore the theoretical results.

97 MATHEMATICS AND COMPUTING↗

Combinatorial Evaluation of Physical Feature Engineering, Classical Machine Learning, and Deep Learning Models for Synchrophasor Data at Scale

A major objective of the project was to train and evaluate the effectiveness of multiple event and anomaly detection, identification and classification deep temporal learning models for processing of real-time phasor measurement unit (PMU) data streams. A vast dataset, consisting of two years of phasor measurements from all three U.S. Interconnections, was curated and released by the Department of Energy (DOE) through Pacific Northwest National Laboratory (PNNL). The dataset also included an event log that provided event times and types (e.g. generator trips, line trips, planned service events, transformer operations, etc.). Our analysis of this dataset addressed six (6) of the eleven (11) research priorities identified in Funding Opportunity Announcement (FOA) DE-FOA-0001861 “Big Data Analysis of Synchrophasor Data” (FOA 1861). Rather than being limited to pre-determined specific algorithms, this project relied on the uniquely structured, highly performant underlying time series database capabilities of the PredictiveGrid platform to assess the vast dataset utilizing a wide variety of algorithms.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Regionalized Life Cycle Greenhouse Gas Emissions of Forest Biomass Use for Electricity Generation in the United States

This study presents a cradle-to-grave life cycle analysis (LCA) of the greenhouse gas (GHG) emissions of the electricity generated from forest biomass in different regions of the United States (U.S.), taking into consideration regional variations in biomass availabilities and logistics. The regional biomass supply for a 20 MW bioelectricity facility is estimated using the Land Use and Resource Allocation (LURA) model. Results from LURA and data on regional forest management, harvesting, and processing are incorporated into the GHGs, Regulated Emissions, and Energy Use in Technologies (GREET) model for LCA. The results suggest that GHG emissions of mill residues-based pathways can be 15-52% lower than those of pulpwood-based pathways, with logging residues falling in between. Nonetheless, our analysis suggests that screening bioenergy projects on specific feedstock types alone is not sufficient because GHG emissions of a pulpwood-based pathway in one state can be lower than those of a mill residue-based pathway in another state. Furthermore, the available biomass supply often consists of several woody feedstocks, and its composition is region-dependent. Forest biomass-derived electricity is associated with 86-93% lower life-cycle GHG emissions than the emissions of the average grid electricity in the U.S. Key factors driving bioelectricity GHG emissions include electricity generation efficiency, transportation distance, and energy use for biomass harvesting and processing.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Utah FORGE: Well 16B(78)-32 Logs from Schlumberger Technologies

This dataset is a collection of well logs provided by Schlumberger Technologies from the Utah FORGE well 16B(78)-32 drilling project. Information here includes critical borehole information collected by an ultrasonic borehole imager (UBI) and a fullbore formation microimager (FMI). Well 16B(78)-32 serves as the production well for reservoir creation, fluid circulation, and demonstration of heat extraction for the FORGE project. It has been drilled as a doublet approximately 300 feet parallel to and above the injection well 16A(78)-32. The total depth measured 10,947 feet and the vertical depth measured 8,357 feet.

11.5 inch intermediate casing logs↗

Probabilistic Assessment and Uncertainty Analysis of CO2 Storage Capacity of the Morrow B Sandstone—Farnsworth Field Unit

This paper presents probabilistic methods to estimate the quantity of carbon dioxide (CO2) that can be stored in a mature oil reservoir and analyzes the uncertainties associated with the estimation. This work uses data from the Farnsworth Field Unit (FWU), Ochiltree County, Texas, which is currently undergoing a tertiary recovery process. The input parameters are determined from seismic, core, and fluid analyses. The results of the estimation of the CO2 storage capacity of the reservoir are presented with both expectation curve and log probability plot. The expectation curve provides a range of possible outcomes such as the P90, P50, and P10. The deterministic value is calculated as the statistical mean of the storage capacity. The coefficient of variation and the uncertainty index, P10/P90, is used to analyze the overall uncertainty of the estimations. A relative impact plot is developed to analyze the sensitivity of the input parameters towards the total uncertainty and compared with Monte Carlo. In comparison to the Monte Carlo method, the results are practically the same. The probabilistic technique presented in this paper can be applied in different geological settings as well as other engineering applications.

03 NATURAL GAS↗

Geothermal Play-Fairway Analysis of Washington State Prospects: Final Report

The Washington State Geothermal Play-Fairway Analysis overcomes the exploration challenges posed by dense vegetation, glacial deposits, and extreme precipitation. The geothermal play-fairways we target are locations where heat, permeability, and saturated porosity are present in sufficient volume to provide adequate heat exchange at depths accessible by modern drilling technology. The three study areas lie along the Cascade Range magmatic arc and are near Mount Baker, Mount St. Helens, and the Wind River Valley. The seven-year project is divided into three phases. In Phase 1 we build on a previous statewide assessment of geothermal resources and develop an initial modeling approach. The results are a series of favorability, uncertainty, and risk maps for three targeted study areas. Based on these initial results, we collect new geologic and geophysical data to further refine our modeling and reduce exploration uncertainty in Phase 2. We improve the modeling method to handle the new data and update the favorability, uncertainty, and risk maps. We also update the conceptual geothermal resource models. In Phase 3 we validate our modeling approach by drilling two temperature-gradient holes and collecting and analyzing core, image logs, and new geochemistry. Our modeling approach improves on an earlier statewide method through a more-rigorous and detailed assessment of heat and permeability. Permeability potential is assessed through geomechanical modeling of the deformation that can generate and maintain reservoir porosity and permeability. Metrics to inform heat potential include temperature-gradient wells, which are sparse in Washington; proximity of Quaternary volcanic vents and young intrusive rock; spring temperature; and reservoir temperature inferred from geothermometry. We weight the individual components using an expert-guided approach known as the Analytical Hierarchy Process. During Phase 2 we also develop a fluid-filled fracture model, and an infrastructure model that helps to delineate areas which are more favorable for geothermal development based on proximity to transmission lines, elevation, land ownership and use restrictions, and availability of process water. New geologic and geophysical data is collected during Phase 2 in each of our three main study areas. At Mount Baker and north of Mount St. Helens we conduct 1:24,000-scale geologic mapping and lidar analysis to better constrain the location and character of surface faults; detailed mapping in the Wind River Valley was completed just prior to the start of this project. Ages of intrusive rocks are determined with 40 Ar/ 39 Ar geochronology, though all of our samples are Miocene or older. We collect ground based gravity observations (a total of 1,580 new stations) in all of our study areas and ground-based magnetic lines (a total of 93 km) at Mount Baker. These data are combined with existing gravity and aeromagnetic data and used to constrain fault locations and geometry. Two to three cross sections are constructed at each study area using the mapped surface geology and forward-modeling of the gravity and magnetic data; these cross sections form the basis for our updated conceptual models. We collect magnetotelluric surveys at Mount Baker and Mount St. Helens and these data are inverted to form a resistivity model from the surface to about 10 km depth; each model shows conductive zones that can be interpreted as upwelling geothermal fluids. At Mount St. Helens we deploy a passive seismic array and use the newly detected events to refine the location of the Saint Helens seismic zone. We also employ ambient-noise tomography to develop a detailed seismic-velocity model for the study area and use this model to help constrain our cross sections and conceptual model. Based on the new data collected during Phase 2—and our updated models—we develop a campaign of temperature-gradient holes and core analysis to validate our modeling in Phase 3. Drill hole MB76-31 is located near Little Park Creek, 11 km west-southwest of the summit of Mount Baker, and is 1,471 ft deep. About 410 ft of core from the lower portion of the hole—and image logs from ~175 ft below ground surface to the bottom—are collected and analyzed. Water samples are collected and processed for geothermometry. Drill hole MSH17-24 is located along upper Schultz Creek, 16 km north-northeast of Mount St. Helens and has core from 470 ft to the bottom at 1,053 ft. We did not collect image logs due to borehole stability concerns, but water samples are collected and analyzed for geothermometry. Repeat temperature-gradient measurements are made at both sites and thermal conductivity is measured from core samples. At MB76-31, the equilibrated temperature gradient of 64°C/km and calculated heat flow of 141–159 mW/m 2 is more than twice the regional average. Detailed mapping and analysis of the core, coupled with correlation to the image logs, indicates a history of permeability generation consistent with our predictions of high permeability. Because the site has high favorability in the Phase 2 model, we consider the results a positive validation of the modeling. At site MSH17-24, the equilibrated temperature gradient of ~15°C/km and calculated heat flow of 41–43 mW/m 2 are similar to regional. Geochemical analysis of the water samples indicates a meteoric source without any geothermal component. Detailed outcrop-based mapping of fault exposures near the drill site and analysis of image logs from nearby boreholes indicates a history of permeability generation consistent with our predictions. Because the site has low favorability in the Phase 2 model, we consider the results a positive validation of the modeling. Together, the two sites provide a reasonably positive validation of the Phase 2 modeling and should encourage future use of this modeling approach.

15 GEOTHERMAL ENERGY↗

Data-driven organic solubility prediction at the limit of aleatoric uncertainty

Abstract Small molecule solubility is a critically important property which affects the efficiency, environmental impact, and phase behavior of synthetic processes. Experimental determination of solubility is a time- and resource-intensive process and existing methods for in silico estimation of solubility are limited by their generality, speed, and accuracy. This work presents two models derived from the FASTPROP and CHEMPROP architectures and trained on BigSolDB which are capable of predicting solubility at arbitrary temperatures for a wide range of small molecules in organic solvent. Both extrapolate to unseen solutes 2–3 times more accurately than the current state-of-the-art model and we demonstrate that they are approaching the aleatoric limit (0.5–1$$\log S$$ log S ) of available test data, suggesting that further improvements in prediction accuracy require more accurate datasets. The FASTPROP-derived model (called FASTSOLV) and the CHEMPROP-based model are open source, freely accessible via a Python package and web interface, highly reproducible, and up to 2 orders of magnitude faster than current alternatives.

Science & Technology - Other Topics↗