Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “data processing and logging”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

A total of 19 months of daily weather logging on the US east coast: the WFIP3 event log

The Third Wind Forecast Improvement Project (WFIP3) is a multi-institutional field campaign designed to advance the understanding and prediction of the offshore atmospheric boundary layer along the US east coast. Extending from February 2024 through August 2025, WFIP3 combines long-term coastal and offshore measurements with targeted modeling and forecasting efforts. This data paper presents the WFIP3 event log, a curated record of 578 d of meteorological phenomena and field observations that complements the campaign's extensive high-frequency datasets. The event log provides both manually documented daily weather discussions and automatically derived indicators of atmospheric processes – including low-level jets, wind ramps, extreme wind veer, and weak wind conditions – based on observations from scanning lidars deployed at three coastal and offshore sites. The dataset offers structured metadata, standardized time and site identifiers, and consistent terminology to facilitate its integration with WFIP3's observational and modeling data products. The log supports diverse applications, from model evaluation and forecast verification to the selection of case studies on offshore boundary-layer dynamics. The WFIP3 event log is publicly available through the US Department of Energy's Wind Data Hub, providing the research community with a transparent and enduring contextual reference for the interpretation and use of WFIP3 measurements.

17 WIND ENERGY↗

Mass Spectrometry Sample Submission Portal

Each step in the scientific process generates contextual information about the data that is important to consider when performing data integration, developing models of biological process, or training AI models. We will develop a flexible, template-driven tool that will log biological samples, capture metadata about those samples, and track the type(s) of analysis being performed by researchers providing samples for analysis by mass spectrometry.

97 MATHEMATICS AND COMPUTING↗

Toward Drilling the Perfect Geothermal Well: An International Research Coordination Network for Geothermal Drilling Optimization Supported by Deep Machine Learning and Cloud Based Data Aggregation

The EDGE project, supported by the U.S. Department of Energy Geothermal Technologies Office under award DE-EE0008793, established a data-driven framework for improving the efficiency, cost-effectiveness, and reliability of geothermal well drilling. The project focused on developing scalable data infrastructure, advanced machine learning and probabilistic models, and integrated analytics tools to support continuous drilling optimization. A central objective was to reduce geothermal drilling costs by up to seventy percent while minimizing the risk of well failure through predictive diagnostics and adaptive planning. Over the project period, a comprehensive data repository was designed and deployed, incorporating records from over one hundred geothermal wells across varied geological settings. This repository supported both structured and unstructured data and adhered to FAIR data principles, enabling provenance tracking, quality control, and standardized metadata. The project introduced automated ingestion pipelines and a cloud-hosted platform that facilitated access to raw, processed, and derived datasets. This infrastructure served as the foundation for model development and analysis. Machine learning workflows were developed to predict key drilling metrics including rate of penetration, non-productive time, and total drilling costs. Self-organizing maps and dimensionality reduction methods were used to uncover operational patterns and outliers, while supervised learning algorithms such as random forests and deep neural networks were applied to forecast performance outcomes. The models were validated on heterogeneous datasets from both U.S. and Icelandic fields, demonstrating variable but significant predictive accuracy. The results indicated that finer temporal resolution, inclusion of lithological data, and consistency in operational annotations could substantially improve model performance. The project also implemented process mining techniques to reconstruct state-transition models from drilling event logs. These models enabled the identification of deviations from optimal workflows and provided insights into recurring failure modes. Analysis of non-productive time highlighted the impact of equipment failures, geological challenges, and human factors, offering opportunities for targeted mitigation strategies. The EDGE Dashboard was developed as a web-based expert system integrating data visualization, model outputs, and user-driven queries. It provided an accessible interface for operators to explore historical data, evaluate predicted outcomes, and compare drilling scenarios. Initial feedback from project partners suggested that the dashboard could serve as a foundation for more advanced advisory and optimization tools. Overall, the EDGE project demonstrated the feasibility and value of applying modern data science techniques to geothermal drilling. It delivered a set of interoperable tools and models that can support more efficient, lower-risk well development. The findings point toward a viable path for transitioning from advisory analytics to semi-autonomous drilling systems, contingent on continued collaboration, expanded datasets, and field validation. The project results have immediate relevance for drilling operations, data management practices, and future geothermal R&D efforts aimed at achieving reliable, cost-competitive geothermal energy at scale.

15 GEOTHERMAL ENERGY↗

Prediction of stability constants of metal–ligand complexes by machine learning for the design of ligands with optimal metal ion selectivity

The new LOGKPREDICT program integrates HostDesigner molecular design software with the machine learning (ML) program Chemprop. By supplying HostDesigner with predicted log K values, LOGKPREDICT enhances the computer-aided molecular design process by ranking ligands directly by metal–ligand binding strength. Harnessing reliable experimental data from a historic National Institute of Standards and Technology (NIST) database and data from the International Union of Pure and Applied Chemistry (IUPAC), we train message passing neural net algorithms. The multi-metal NIST-based ML model has a root mean square error (RMSE) of 0.629 ± 0.044 (R 2 of 0.960 ± 0.006), while two versions of lanthanide-only IUPAC-based ML models have, respectively, RMSE of 0.764 ± 0.073 (R 2 of 0.976 ± 0.005) and 0.757 ± 0.071 (R 2 of 0.959 ± 0.007). For relative log K predictions on an out-of-sample set of six ligands, demonstrating metal ion selectivity, the RMSE value reaches a commendably low 0.25. Here we showcase the use of LOGKPREDICT in identifying ligands with high selectivity for lanthanides in aqueous solutions, a finding supported by recent experimental evidence. We also predict new ligands yet to be verified experimentally. Therefore, our ML models implemented through LOGKPREDICT and interfaced with the ligand design software HostDesigner pave the way for designing new ligands with predetermined selectivity for competing metal ions in an aqueous solution.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

MARSAME Radiological Release Report for Metal Items from Technical Area 53, Set 28

Environmental Protection and Compliance, Environmental Stewardship Group (EPC-ES) has evaluated the survey results for metal items from the Los Alamos Neutron Science Center (LANSCE) at Technical Area 53 (TA-53) and found that the metal items described in Table 1 of this report (identified by Radiation Protection [RP] Tracking Numbers) meet the criteria for unrestricted release under Department of Energy (DOE) Order 458.1 Chg 4, Radiation Protection of the Public and the Environment (DOE 2020) and can be recycled. This conclusion is based on the known history of the metal items and radiation survey data (see the completed RP-Form-031 LANSCE Metals Clearance Log [LANL 2021a] for each item in Attachment 1). Process knowledge indicates that items were released from radiological areas, including radiation areas, prior to the implementation of the 2000 metals moratorium. Therefore, the items are considered unencumbered and are not subject to the moratorium suspension on metal recycling from DOE facilities. Additionally, Los Alamos National Laboratory (LANL) has determined that there is no practical opportunity for internal DOE reuse of this metal. Process knowledge indicates that these metal items were unlikely to ever be in direct contact with the beam and thus are unlikely to have become activated. Surface contamination measurements (both total and removable) showed either no detectable radioactivity or activity levels within the range of background. All measurements for volumetric contamination were indistinguishable from background based on calculated decision limits. Additionally, all gamma isotopic surveys conducted for defense-in depth showed no identifiable gamma radiation from beam activation.

61 RADIATION PROTECTION AND DOSIMETRY↗

Predicting Li-Ion Battery Capacity Fade Using Early-Life Data and a Hybrid Data-Driven Gaussian Process-Bayesian Regression Approach

Accurately predicting Li-ion battery capacity trajectories using early-life data can dramatically improve battery-life understandings and be used to rapidly evaluate design/cost/performance trade-offs when developing new battery materials. Accurate early-life predictions enable researchers to quickly iterate over cell designs and material precursor properties without consistently cycling cells to failure. To this end, we present a toolbox that uses a combined Gaussian Process and Bayesian regression approach that capitalizes on signals other than just capacity (e.g., dQ/dV, voltage drops) to rapidly predict capacity-fade trajectories. The prediction tool uses Bayesian regression to fit functional forms, e.g., power law, sigmoids, etc., to predict capacity-fade dynamics. By fitting functional forms, the capacity fade can be interrogated at any point in the future, allowing for early cell-failure prediction. Additionally, Bayesian regression allows for accurate uncertainty estimates that account for cell-to-cell variability (aleatoric uncertainty) and the lack of observation data (epistemic uncertainty). By only using early cycle data to predict the capacity fade trajectory, uncertainty bounds at end-of-life can be extremely large. The large uncertainty bounds are further exacerbated because there is no systematic way to define the prior distribution of the functional forms' parameters. We improve our the predicted trajectory confidence interval of our predicted trajectory using two methods. First, we shows that a small amount of held-out cycling data is sufficientuse some train cells, that have been cycled to failure to derive information regarding the appropriate prior distributions for the functional forms' parameters of the functional form, effectively leading to data-driven priors.. We propose constructing the data-driven priors by first running a Bayesian regression starting with uninformed priors to generate intermediate cell-specific posterior parameter distributions. These posterior distributions are combined using a Ggaussian mixture model for each parameter to create the data-driven priors. These mixture models serve as the data-driven prior distributions for the parameters for. Second, we derive multiple features, e.g., C_dchg 0.5 DoD 0.5, log (|mean(dQ/dV_(w_3-w_0 ) (V)|), etc., from the train cellsheld-out cycling data, identify which the features are that best predicting capacity at early/mid-life cycles, and then create Ggaussian process regression models that are used for predicting capacity at early/mid-life cycles for the test cells (see blue dots with error bars in Fig 1b). Finally, these predicted data-points are used in addition to the actual early cycle data capacity fade to construct the Bayesian regression trajectory for the test cell s. Notably. We note that these two methods are complementary and can be combined with each other. We evaluate the performance of our proposed method on an testing open-source dataset from Iowa State University and Iowa Lakes Community College (ISU-ILCC). This dataset comprises of 251 nickel-manganese-cobalt/graphite Lithium-ion cells that are cycled under 63 different conditions. We compute the mean average percentage error (MAPE) and negative log predictive density (NLPD) to quantify the efficacy of our method. Our initial findings suggest that, when only few observations are available, for test cells, when using only Bayesian regression with uninformed priors, a power law functional provides the most accurate predictions. with very few data points. However, asHowever, a the number of data points increases, a twin sigmoidal function becomes more accurate as the number of observations further increases. We also find that using as little as 10% of the data set towards generating data-driven priors can lead to significant improvement in prediction accuracy when using early cycle data. Lastly, we found that augmenting early-cycle data with Gaussian process-predicted capacity data for Bayesian regression greatly improves the prediction accuracy. We will present a comprehensive comparison of our methods to other methods available in the literature and apply this method to additional battery datasets.

42 ENGINEERING↗

Malicious Cyber Activity Detection using Zigzag Persistence

In this study we synthesize zigzag persistence from topological data analysis with autoencoder-based approaches to detect malicious cyber activity, and derive analytic insights. Cybersecurity aims to safeguard computers, networks, and servers from various forms of malicious attacks, including network damage, data theft, and activity monitoring. We focus on the cybersecurity domain and investigate the detection of malicious activity using log data. We consider the dynamics of the log data and explore the changing topology of a hypergraph representation of this data to gain insights into the underlying activity. These hypergraphs capture complex interactions between processes, together with their temporal information. To study the changing topology we use zigzag persistence, which captures how topological features persist at multiple dimensions over time. We observe that this detects malicious activity in a cyber data set. To automate this detection we implement an autoencoder trained on a vectorization of the resulting zigzag persistence barcodes. Our experimental results demonstrate the effectiveness of the autoencoder in detecting malicious activity. Overall, this study highlights the potential of zigzag persistence and its combination with temporal hypergraphs for analyzing cybersecurity log data and detecting malicious behavior.

hypergraphs, temporal hypergraph, topological data↗

SlimIO: Lightweight I/O Path Design for Write Isolation in FDP-backed In-Memory Databases

In-Memory Databases (IMDBs) are widely used with HPC applications to manage transient data, often using snapshot-based persistence for backups. Redis, a representative IMDB, employs both snapshot and Write-Ahead Log (WAL) mechanisms, storing data on persistent devices via the traditional kernel I/O path. This method incurs syscall overhead, I/O contention between processes, and SSD garbage collection (GC) delays. To address these issues, we propose SlimIO, which adopts I/O passthru to minimize syscall overhead and inter-process I/O interference. Additionally, it leverages Flexible Data Placement (FDP) SSDs as backup storage to avoid performance degradation from SSD GC. Experimental results show that SlimIO reduces snapshot time by up to 25%, increases query throughput by up to 30% during non-snapshot periods, and lowers 99.9%-ile latency by up to 50%. Furthermore, it achieves a write amplification factor (WAF) of 1.00, indicating no redundant internal writes, thus extending SSD lifespan.

Lee, Sangyun [Sogang University]↗

Terrestrial laser scanning data (Levels 0 and 1) for Pasoh, Malaysia, Sep 2024

This data package contains data from terrestrial laser scanning (TLS) at the Pasoh Forest Reserve, Malaysia. The Pasoh Forest Reserve is a facility of the Forest Research Institute Malaysia, and contains evergreen lowland dipterocarp forest. The Next-Generation Ecosystem Experiments Tropics (NGEE-Tropics) study areas at Pasoh were established to study how different species respond to climatic variation and soil water availability. Two study areas were chosen representing different topography and species. The TLS data archived here were collected to provide detailed, three-dimensional information about forest structure. Specifically, data were collected to allow tree-level characterization of woody structure and leaf area for 12 focal trees with FloraPulse and sap flux sensors, facilitating estimation of woody biomass and leaf area to allow upscaling of water content and transpiration data to the tree-level. Scan positions were not selected to provide consistent data for non-focal trees with the study areas. This data package contains the following data: - High-level files document further details of the campaign and data package: 1_CampaignSummary.csv provides details about the campaign and study site, 2_ScanAreasDetail.csv provides details about each separate scan area (groups of scans post-processed into a single point cloud), 3_TerrestrialLidarSensor.csv provides further technical details about the Riegl VZ-400i TLS sensor, TLS_CSV_dd.csv is a CSV Data Dictionary providing information about the fields in CSV files following the ESS-DIVE CSV File Formatting Guidelines Reporting Format, TLS_flmd.csv is a File Level Metadata file providing information about each file in the data package following the ESS-DIVE File Level Metadata Reporting Format, and README.txt is a text file describing the overall project and file structure. - Level 0 data are the raw data (.PROJ folders) as recorded by the Riegl VZ-400i TLS instrument before scan co-registration and post-processing with the Riegl's proprietary RiSCAN PRO software, which requires a license. - Level 1 data contain post-processed, co-registered data from each scan area. The "PointClouds" folder for each scan area contains a .las file with 1 cm resolution point cloud data exported from RiSCAN PRO. These are the main files likely to be of interest to most users and can be further processed with any software capable of manipulating .las files (e.g. Python, R CloudCompare). The "Project Information" folder contains log files from post-processing in RiSCAN PRO that may be of interest to users who want to see detailed records of post-processing, including all PDF reports generated by RiSCAN PRO. The "ScanPositions" folder contains information about the final position of all TLS scans, after post-processing, in multiple formats. The file ScanPositions_*.csv provides final geo-referenced scan positions, and the file SOP_backup_*.csv can be used in RiSCAN PRO to restore the co-registered scan positions if users wish to re-process raw data (Level 0 .PROJ folders) with RiSCAN PRO software (e.g., subsample to a different resolution, exclude a certain scan position, or apply different filters on reflectance or deviation values) without redoing time-consuming co-registration steps.

54 ENVIRONMENTAL SCIENCES↗

Terrestrial laser scanning data (Levels 0 and 1) from Urban Biogeochemistry Pilot Project sites, Knoxville, Tennessee, Jul 2024 - Jul 2025

This data package contains data from terrestrial laser scanning (TLS) at five urban park sites in Knoxville, Tennessee, USA. All parks include open-grown and/or closed-canopy trees and mixed nearby land use. These study sites were established as part of the Urban Biogeochemistry Pilot Project, which has an overall goal of better understanding how hydrobiogeochemical cycling is altered within the human environment. These five sites represent a gradient of urbanization, and were instrumented to understand hydrological and biogeochemical cycling (e.g., soil moisture, soil physical properties and biogeochemistry, tree transpiration, species type). The TLS data archived here were collected to provide detailed, three-dimensional information about forest structure. Specifically, data were collected to allow tree- and stand-level characterization of woody structure and leaf area. TLS scans were placed to capture the area around trees with sap flow sensors, and as much of a 50 m radius area around the meteorological station as possible given site property limits. Derived products will allow upscaling of water content and transpiration data. This data package contains the following data: - High-level files document further details of the campaign and data package: 1_CampaignSummary.csv provides details about the campaign and study site, 2_ScanAreasDetail.csv provides details about each separate scan area (groups of scans post-processed into a single point cloud), 3_TerrestrialLidarSensor.csv provides further technical details about the Riegl VZ-400i TLS sensor, TLS_CSV_dd.csv is a CSV Data Dictionary providing information about the fields in CSV files following the ESS-DIVE CSV File Formatting Guidelines Reporting Format, TLS_flmd.csv is a File Level Metadata file providing information about each file in the data package following the ESS-DIVE File Level Metadata Reporting Format, and README.txt is a text file describing the overall project and file structure. - Level 0 data are the raw data (.PROJ folders) as recorded by the Riegl VZ-400i TLS instrument before scan co-registration and post-processing with the Riegl's proprietary RiSCAN PRO software, which requires a license. - Level 1 data contain post-processed, co-registered data from each scan area. The "PointClouds" folder for each scan area contains a .las file with 1 cm resolution point cloud data exported from RiSCAN PRO. These are the main files likely to be of interest to most users and can be further processed with any software capable of manipulating .las files (e.g. Python, R CloudCompare). The "Project Information" folder contains log files from post-processing in RiSCAN PRO that may be of interest to users who want to see detailed records of post-processing, including all PDF reports generated by RiSCAN PRO. The "ScanPositions" folder contains information about the final position of all TLS scans, after post-processing, in multiple formats. The file ScanPositions_*.csv provides final geo-referenced scan positions, and the file SOP_backup_*.csv can be used in RiSCAN PRO to restore the co-registered scan positions if users wish to re-process raw data (Level 0 .PROJ folders) with RiSCAN PRO software (e.g., subsample to a different resolution, exclude a certain scan position, or apply different filters on reflectance or deviation values) without redoing time-consuming co-registration steps.

54 ENVIRONMENTAL SCIENCES↗

UBW (USLCI-Brightway2) [SWR-25-169]

Life cycle inventory (LCI) data are critical for robust life cycle assessment (LCA), yet many widely used datasets such as the U.S. Life Cycle Inventory (USLCI) are not natively compatible with advanced modeling frameworks like Brightway2. This work presents an automated pipeline to transform USLCI data into a fully functional Brightway2 project. The workflow performs systematic data cleaning, resolves duplicate process and exchange identifiers, and applies allocation to multi-output processes. Technosphere and biosphere flows are harmonized through unit conversions and a bridge mapping to the biosphere3 database, with comprehensive logging of missing flows and cutoff issues. The resulting Brightway2 database is validated using matrix diagnostics to ensure consistency of the technosphere, and is benchmarked via life cycle impact assessment (LCIA) methods such as ReCiPe and IPCC GWP. Outputs include reproducible CSV exports of corrected processes, elementary flows, characterization factors, and LCIA results, alongside backup utilities for project sharing. This pipeline lowers barriers for integrating USLCI data into open-source LCA workflows, enabling reproducible, validated LCA inventories within the Brightway 2 framework.

Ghosh, Tapajyoti [National Laboratory of the Rocki↗

Roughrider Carbon Storage Hub (Final Report)

The Roughrider Carbon Storage Hub was a 2-year project (October 2023 – September 2025) conducted by the Energy & Environmental Research Center (EERC) focused on advancing the feasibility of a commercial-scale carbon dioxide (CO 2 ) geologic storage hub in McKenzie County, North Dakota. The project’s objective was to investigate the potential that stacked storage complexes (multiple deep saline formations) can safely and economically store at least 50 million tonnes of CO 2 within 30 years. The captured CO 2 would be sourced from industrial emitters including project partner ONEOK, Inc.’s gas-processing plants and a planned gas-to-liquids facility. Drilling of the Roughrider 1 stratigraphic test well (14,979-ft total depth) was completed in November 2024. The wellbore intersected four candidate storage formations: Inyan Kara, Broom Creek, Mission Canyon, and Black Island–Deadwood. Operational challenges, including a stuck drill string, were resolved without long-term impact. A comprehensive logging and coring program was conducted, followed by successful well abandonment and site reclamation. Over 660 ft of 4-in. whole core was retrieved. Core plug samples were processed and analyzed for petrophysical and geochemical properties. Results confirmed promising porosity and permeability in the Inyan Kara and Broom Creek Formations and removal of the Mission Canyon and Black Island–Deadwood horizons from further investigation. Data derived from the logging and coring program were used to improve initial geologic models built from legacy data. CO 2 injection simulations showed that the Inyan Kara alone can feasibly store the target mass of CO 2 . Because of subtle differences in geologic structure and porosity trends between the formations, a stacked storage scenario using the Broom Creek and Inyan Kara Formations resulted in a larger overall plume area than using the Inyan Kara alone. Preliminary CO 2 pipeline routes from the industrial sources were mapped utilizing existing rights of way and evaluated for capacity and cost using U.S. Department of Energy Office of Fossil Energy and Carbon Management/National Energy Technology Laboratory models and U.S. Environmental Protection Agency emissions data. Integrating capture, transport, and storage cost estimates with policy incentives (e.g., 45Q credits) provided a total cost-per-ton analysis. Results indicate that the small scale of the volumes to be transported over the cumulative large distances does not support the project’s financial viability. However, the groundwork laid during this project from geological, regulatory, and social perspectives positions the Roughrider hub site as a promising candidate for commercial carbon storage in North Dakota, especially if the economy of scale is introduced for CO 2 transportation to the hub site.

01 COAL, LIGNITE, AND PEAT↗

Oscilloscope Data Push Program

This paper details the development of a Python program designed to automate the data acquisition and conversion for an oscilloscope for the purposes of a one-off/temporary data acquisition system for users that readily need data, and do not have the option of obtaining a Data Acquisition (DAQ) solution. Creating DAQ systems for analyzing a system requires expensive electronics and a dedicated team of engineers for support. Traditionally, manual data collection and processing are time consuming and prone to error. By automating these processes, the cost, efficiency and accuracy of data handling are improved upon. This project involves the creation of a program that interacts with the oscilloscope. During this interaction, there are various functions being performed such as the acquisition of waveform data via floating points, generating plots with the acquired wave points, and storing of floating points in a CSV file format for future reference and plotting purposes. While the initial aim of the project included continuous logging to a cloud database, this was deferred due to time constraints. The results portrayed an almost-instant rate of data collection with a buffer time, showcasing the potential for further integration and real-time data processing.

Osei-Tutu, Jason↗

Access Patterns and Performance Behaviors of Multi-layer Supercomputer I/O Subsystems under Production Load

Scientific computing workloads at HPC facilities have been shifting from traditional numerical simulations to AI/ML applications for training and inference while processing and producing ever-increasing amounts of scientific data. To address the growing need for increased storage capacity, lower access latency, and higher bandwidth, emerging technologies such as non-volatile memory are integrated into supercomputer I/O subsystems. With these emerging trends, we need a better understanding of the multilayer supercomputer I/O systems and ways to use these subsystems efficiently. In this work, we study the I/O access patterns and performance characteristics of two representative supercomputer I/O subsystems. Through an extensive analysis of year-long I/O logs on each system, we report new observations in I/O reads and writes, unbalanced use of storage system layers, and new trends in user behaviors at the HPC I/O middleware stack.

Bez, JL↗

Pseudonymization at Scale: OLCF’s Summit Usage Data Case Study

The analysis of vast amounts of data and the processing of complex computational jobs have traditionally relied upon high performance computing (HPC) systems, which offer reliable and efficient management of large-scale computational and data resources. Understanding these analyses’ needs is paramount for designing solutions that can lead to better science, and similarly, understanding the characteristics of the user behavior on those systems is important for improving user experiences on HPC systems. A common approach to gathering data about user behavior is to extract workload characteristics from system log data available only to system administrators. Recently at Oak Ridge Leadership Computing Facility (OLCF), however, we unveiled user behavior about the Summit supercomputer by collecting data from a user’s point of view with ordinary Unix commands.In this paper, we discuss the process, challenges, and lessons learned while preparing this dataset for publication and submission to an open data challenge. The original dataset contains personal identifiable information (PII) about the users of OLCF which needed be masked prior to publication, and we determined that anonymization, which scrubs PII completely, destroyed too much of the structure of the data to be interesting for the data challenge. We instead chose to pseudonymize the dataset, which reduced the linkability of the dataset to the users’ identities. Pseudonymization is significantly more computationally expensive than anonymization, and the size of our dataset, which is approximately 175 million lines of raw text, necessitated the development of a parallelized workflow that could be reused on different HPC machines. We demonstrate the scaling behavior of the workflow on two leadership class HPC systems at OLCF, and we show that we were able to bring the overall makespan time from an impractical 20+ hours on a single node down to around 2 hours. As a result of this work, we release the entire pseudonymized dataset and make the workflows and source code publicly available.

Maheshwari, Ketan↗

pvOps: a Python package for empirical analysis of photovoltaic field data

The purpose of pvOps is to support empirical evaluations of data collected in the field related to the operations and maintenance (O&M) of photovoltaic (PV) power plants. pvOps presently contains modules that address the diversity of field data, including text-based maintenance logs, current-voltage (IV) curves, and timeseries of production information. The package functions leverage machine learning, visualization, and other techniques to enable cleaning, processing, and fusion of these datasets. These capabilities are intended to facilitate easier evaluation of field patterns and extraction of relevant insights to support reliability-related decision-making for PV sites. The open-source code, examples, and instructions for installing the package through PyPI can be accessed through the GitHub repository.

14 SOLAR ENERGY↗

Event Log / Raw Data

The WFIP3 event log is a curated record spanning 578 days of meteorological phenomena and field observations that complements the campaign’s high-frequency measurements. The log combines manually documented daily weather discussions with automatically derived indicators of key atmospheric processes, providing standardized, publicly available context to support model evaluation, forecast verification, and case-study selection for offshore boundary-layer research.

17 WIND ENERGY↗

Utah FORGE: Optimization of a Plug-and-Perf Stimulation (Fervo Energy)

Information around the plug-and-perf treatment design at Utah FORGE by Fervo Energy. Objective and Purpose: - Develop a multistage hydraulic stimulation approach designed specifically to target the top three factors that control the technical and commercial viability of an EGS system: i) Achieving sufficient injectivity to support high cross-well flow rates ii) Distributing flow evenly across the wellbore and reservoir to maximize heat mining efficiency, ensure sustained heat transfer, and mitigate thermal breakthrough iii) Overcoming the effects of stress heterogeneity, stress shadowing, and variations in natural fracture properties during the stimulation treatment, leading to a more predictable stimulated reservoir volume and offset well placement - The following activities will be performed: i) Design, plan, and execute a multistage plug-and-perf stimulation treatment at a Fervo site with data acquisition and well testing activities aimed at addressing key technical aspects of the issues above ii) Perform data processing and interpretation of field results to translate the results form the Fervo site to a site-specific design at the Utah FORGE site iii) Design, plan, and execute a multistage plug-and-perf stimulation treatment design at the Utah FORGE site Methods and Approach: - Design a detailed data acquisition plan to maximize learning around: i) DFIT testing ii) Petrophysical logging, image logging iii) Permanent DAS/DTS fiber optic monitoring iv) Deep borehole microseismic monitoring v) Shallow borehole induced seismicity monitoring vi) Injection/production testing (RTA analysis, tracer testing) vii) Integrated numerical modeling and production forecasting

15 GEOTHERMAL ENERGY↗