Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “data transfer pipeline”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

PhytoOracle: Scalable, modular phenomics data processing pipelines

As phenomics data volume and dimensionality increase due to advancements in sensor technology, there is an urgent need to develop and implement scalable data processing pipelines. Current phenomics data processing pipelines lack modularity, extensibility, and processing distribution across sensor modalities and phenotyping platforms. To address these challenges, we developed PhytoOracle (PO), a suite of modular, scalable pipelines for processing large volumes of field phenomics RGB, thermal, PSII chlorophyll fluorescence 2D images, and 3D point clouds. PhytoOracle aims to ( i ) improve data processing efficiency; ( ii ) provide an extensible, reproducible computing framework; and ( iii ) enable data fusion of multi-modal phenomics data. PhytoOracle integrates open-source distributed computing frameworks for parallel processing on high-performance computing, cloud, and local computing environments. Each pipeline component is available as a standalone container, providing transferability, extensibility, and reproducibility. The PO pipeline extracts and associates individual plant traits across sensor modalities and collection time points, representing a unique multi-system approach to addressing the genotype-phenotype gap. To date, PO supports lettuce and sorghum phenotypic trait extraction, with a goal of widening the range of supported species in the future. At the maximum number of cores tested in this study (1,024 cores), PO processing times were: 235 minutes for 9,270 RGB images (140.7 GB), 235 minutes for 9,270 thermal images (5.4 GB), and 13 minutes for 39,678 PSII images (86.2 GB). These processing times represent end-to-end processing, from raw data to fully processed numerical phenotypic trait data. Repeatability values of 0.39-0.95 (bounding area), 0.81-0.95 (axis-aligned bounding volume), 0.79-0.94 (oriented bounding volume), 0.83-0.95 (plant height), and 0.81-0.95 (number of points) were observed in Field Scanalyzer data. We also show the ability of PO to process drone data with a repeatability of 0.55-0.95 (bounding area).

59 BASIC BIOLOGICAL SCIENCES↗

Artificial Intelligence for Event Reconstruction and Higgs Physics at CMS and Future Colliders

This dissertation charts a trajectory in which advances in artificial intelligence (AI) play a central role in pushing the high-energy physics frontier, complementing progress driven by higher collision energies and larger colliders. The discovery potential of the LHC and future colliders relies on accurate reconstruction of increasingly complex particle collision events. In the CMS experiment, this task is performed by the particle-flow (PF) algorithm. This dissertation presents the first implementation of a machine-learning-based particle-flow (MLPF) reconstruction in the CMS detector based on transformer architectures. In simulated top quark--antiquark pair (ttbar) events under LHC Run~3 (2023--2024) conditions, MLPF improves jet energy resolution by 10--20\% compared to standard PF for jets with transverse momentum between 30--100\GeV. Runtime performance is evaluated using simulated multijet events, with a median inference time of 20\unit{ms} per event on an NVIDIA L4 GPU, compa red to approximately 110\unit{ms} for standard PF. The MLPF algorithm is also validated on Run~3 collision data, representing the first data-validated ML-based reconstruction pipeline at any LHC experiment. We then extend MLPF toward future electron--positron colliders and introduce the first full-simulation cross-detector transfer learning workflow for PF reconstruction. The model is pre-trained on simulated events from the Compact Linear Collider detector (CLICdet) and fine-tuned on the CLIC-like detector (CLD) proposed for the Future Circular Collider (FCC). This approach achieves up to a 40\% improvement in jet energy resolution over rule-based reconstruction while reducing the required training dataset size by an order of magnitude, demonstrating the potential of AI to accelerate detector development and optimization. This dissertation also demonstrates how modern AI techniques enhance the sensitivity of LHC physics analyses. A CMS search for highly Lorentz-boosted Higgs bosons decaying to \textrm{W} boson pairs is presented, focusing on the single-lepton final state. A dedicated fine-tuning strategy for \ParT yields an approximately 70\% increase in expected sensitivity relative to the baseline model. The analysis uses proton--proton collision data at a center-of-mass energy of \ensuremath{\sqrt{s}=13\TeV} collected by CMS between 2016 and 2018, corresponding to an integrated luminosity of 138\ensuremath{\ \mathrm{fb}^{-1}}. The expected significance of the search is $1.86\sigma$, with an observed signal strength of $-0.19^{+0.48}_{-0.46}$. Finally, explainable AI techniques are applied to the MLPF and \ParticleNet algorithms using layerwise relevance propagation, showing that both models base their predictions on physically meaningful features consistent with our physics intuition. Together, these results demonstrate how advanced AI methods can enhance reconstruction, analysis sensitivity, and interpretability, shaping the next era of experimental parti cle physics.

Mokhtar, Farouk [UC, San Diego]↗

Enhancing transfer learning in angle-resolved photoemission spectroscopy (ARPES) with spatially-aware representations via graph convolution

A recent application of machine learning has been to spatially-resolved angle-resolved photoemission spectroscopy (ARPES). Here we advance the state-of-the-art by applying representational learning to transform ARPES data into an embedding space of a pre-trained self-supervised learning model, thus enhancing the pipeline that improves the bandstructure classification and domain assignment/segmentation performance compared to a k-means clustering method. In the current iteration, the real-space information is entered into the domain assignment through the graph convolution method, which improves the transfer learning performance of the original self-supervised model. Lastly, an unsupervised automated tool is developed that incorporates these techniques to enable automatic domain assignment.

ARPES↗

An FPGA-based hardware accelerator supporting sensitive sequence homology filtering with profile hidden Markov models

Abstract Background Sequence alignment lies at the heart of genome sequence annotation. While the BLAST suite of alignment tools has long held an important role in alignment-based sequence database search, greater sensitivity is achieved through the use of profile hidden Markov models (pHMMs). Here, we describe an FPGA hardware accelerator, called HAVAC, that targets a key bottleneck step (SSV) in the analysis pipeline of the popular pHMM alignment tool, HMMER. Results The HAVAC kernel calculates the SSV matrix at 1739 GCUPS on a $$\sim$$ ∼ $3000 Xilinx Alveo U50 FPGA accelerator card, $$\sim$$ ∼ 227× faster than the optimized SSV implementation in nhmmer . Accounting for PCI-e data transfer data processing, HAVAC is 65× faster than nhmmer’s SSV with one thread and 35× faster than nhmmer with four threads, and uses $$\sim$$ ∼ 31% the energy of a traditional high end Intel CPU. Conclusions HAVAC demonstrates the potential offered by FPGA hardware accelerators to produce dramatic speed gains in sequence annotation and related bioinformatics applications. Because these computations are performed on a co-processor, the host CPU remains free to simultaneously compute other aspects of the analysis pipeline.

59 BASIC BIOLOGICAL SCIENCES↗

Streaming Large-Scale Microscopy Data to a Supercomputing Facility

Data management is a critical component of modern experimental workflows. As data generation rates increase, transferring data from acquisition servers to processing servers via conventional file-based methods is becoming increasingly impractical. The 4D Camera at the National Center for Electron Microscopy generates data at a nominal rate of 480 Gbit s -1 (87,000 frames s -1 ⁠), producing a 700 GB dataset in 15 s. To address the challenges associated with storing and processing such quantities of data, we developed a streaming workflow that utilizes a high-speed network to connect the 4D Camera’s data acquisition system to supercomputing nodes at the National Energy Research Scientific Computing Center, bypassing intermediate file storage entirely. In this work, we demonstrate the effectiveness of our streaming pipeline in a production setting through an hour-long experiment that generated over 10 TB of raw data, yielding high-quality datasets suitable for advanced analyses. Additionally, we compare the efficacy of this streaming workflow against the conventional file-transfer workflow by conducting a postmortem analysis on historical data from experiments performed by real users. Our findings show that the streaming workflow significantly improves data turnaround time, enables real-time decision-making, and minimizes the potential for human error by eliminating manual user interactions.

4D-STEM↗

Complete and Correct Transfer of Information (CACTI)

Many distributed systems, file transfer mechanisms, and message passing systems offer reliability mechanisms such as acknowledgements, retries, and durability. While these tools may be “good enough” for their typical use cases, they may not offer sufficient coverage for the wide range of faults that impact data transfers and communication. A gap in the reliability measures may lead to some small amount of data loss. Some high-consequence systems cannot tolerate the loss or corruption of even a single record. We present seven principles that will counter a wide range of faults and protect against data loss and corruption. These principles bring together lessons learned from a wide range of technologies and can inform appropriate system design and application usage. These principles will help readers reason on how prevent data loss in a multi-hop pipeline and how to properly use tools that may have a deficiency in reliability.

97 MATHEMATICS AND COMPUTING↗

A computational pipeline to generate a synthetic dataset of metal ion sorption to oxides for AI/ML exploration

The charged mineral/electrolyte interfaces are ubiquitous in the surface and subsurface–including the surroundings of the geological disposal sites for radioactive waste. Therefore, understanding how ions interact with charged surfaces is critically important for predicting radionuclide mobility in the case of waste leakage. At present, the Surface Complexation Models (SCMs) are the most successful thermodynamic frameworks to describe ion retention by mineral surfaces. SCMs are interfacial speciation models that account for the effect of the electric field generated by charged surfaces on sorption equilibria. These models have been successfully used to analyze and interpret a broad range of experimental observations including potentiometric and electrokinetic titrations or spectroscopy. Unfortunately, many of the current procedures to solve and fit SCM to experimental data are not optimal, which leads to a non-transferable or non-unique description of interfacial electrostatics and consequently of the strength and extent of ion retention by mineral surfaces. Recent developments in Artificial Intelligence (AI) offer a new avenue to replace SCM solvers and fitting algorithms with trained AI surrogates. Unfortunately, there is a lack of a standardized dataset covering a wide range of SCM parameter values available for AI exploration and training–a gap filled by this study. Here, we described the computational pipeline to generate synthetic SCM data and discussed approaches to transform this dataset into AI-learnable input. First, we used this pipeline to generate a synthetic dataset of electrostatic properties for a broad range of the prototypical oxide/electrolyte interfaces. The next step is to extend this dataset to include complex radionuclide sorption and complexation, and finally, to provide trained AI architectures able to infer SCMs parameter values rapidly from experimental data. Here, we illustrated the AI-surrogate development using the ensemble learning algorithms, such as Random Forest and Gradient Boosting. These surrogate models allow a rapid prediction of the SCM model parameters, do not rely on an initial guess, and guarantee convergence in all cases.

Li, Chunhui↗

Workflow for High-throughput Screening of Enzyme Mutant Libraries Using Matrix-assisted Laser Desorption/Ionization Mass Spectrometry Analysis of Escherichia coli Colonies

High-throughput molecular screening of microbial colonies and DNA libraries are critical procedures that enable applications such as directed evolution, functional genomics, microbial identification, and creation of engineered microbial strains to produce high-value molecules. A promising chemical screening approach is the measurement of products directly from microbial colonies via optically guided matrix-assisted laser desorption/ionization mass spectrometry (MALDI-MS). Measuring the compounds from microbial colonies bypasses liquid culture with a screen that takes approximately 5 s per sample. We describe a protocol combining a dedicated informatics pipeline and sample preparation method that can prepare up to 3,000 colonies in under 3 h. The screening protocol starts from colonies grown on Petri dishes and then transferred onto MALDI plates via imprinting. The target plate with the colonies is imaged by a flatbed scanner and the colonies are located via custom software. The target plate is coated with MALDI matrix, MALDI-MS analyzes the colony locations, and data analysis enables the determination of colonies with the desired biochemical properties. This workflow screens thousands of colonies per day without requiring additional automation. The wide chemical coverage and the high sensitivity of MALDI-MS enable diverse screening projects such as modifying enzymes and functional genomics surveys of gene activation/inhibition libraries.

Choe, Kisurb↗

The Geothermal Data Repository: Ten Years of Supporting the Geothermal Industry with Open Access to Geothermal Data: Preprint

The Department of Energy's (DOE) Geothermal Data Repository (GDR) is celebrating its tenth anniversary! Over the last decade it has grown from the simple idea of storing public data in a centralized location to a valuable tool at the center of the US geothermal scientific community and an integral part of the DOE Geothermal Technologies Office (DOE GTO) project management strategy. Researchers funded by the DOE GTO have contributed over 1,300 data submissions to the GDR. These data have been used to further advancements in geothermal science, economic analysis, exploration, research, development, and operational efficiency. The adoption of open data methodologies and a data management strategy that prioritizes universal open access and standardized, interoperable data have further increased the value of GDR data, making them available across a distributed network of data sharing partners and improving their utility to other industries and related fields, including material science and space exploration. Incorporating feedback from users has been critical to the GDRs success, allowing it to grow over the years to meet the evolving needs of the geothermal community. This paper will explore some of many changes that occurred throughout the GDRs tenure and the lessons learned along the way, as well as highlight some of the new features and recent improvements that been implemented to support innovation, reduce duplication of effort, and advance the geothermal industry as a whole.

accessibility↗

Effect of Drying on Corrosion Mitigation of Hanford Transfer Lines

Radioactive waste is stored in underground, carbon-steel double-shell tanks at the Department of Energy Hanford site. The waste is transferred between the tanks and other assets using the transfer lines spanning throughout the various tank farms at Hanford. The transfer lines consist of a pipe-in-pipe design, small diameter pipes, and are not piggable. Recent inspection data of the transfer lines have shown areas with corrosion on both interior of the encasements and exterior of the primary pipes, with nearly 50 percent wall loss on the primary pipes and nearly 25% wall loss on the encasement pipes due to pitting corrosion. The visual inspections of the transfer lines have shown presence of corrosion products near the pipeline risers and beyond. It has been hypothesized that the corrosion is predominantly due to the high humidity conditions and in some cases is driven by the presence of residual hydrotest water in the encasement and the associated contact with the safety significant primary pipe. Therefore, drying of the transfer lines could lead to corrosion mitigation. Experimental studies are being conducted to understand the effect of environmental conditions, especially, relative humidity and temperature, on transfer line grade carbon steel corrosion and on mitigating corrosion. The experimental conditions are selected based on the seasonal temperature changes, and relative humidity conditions ranging from 30 to 100 percent. The experimental data will be used as guidance for maintaining a dry environment that will help mitigate the transfer-line corrosion caused by the high humidity conditions.

Shukla, Pavan K.↗

Evaluation of High Level Waste Sludge Processing Behavior

The U.S. Department of Energy’s (DOE) Hanford Site has 177 underground storage tanks that contain wastes from past nuclear fuel reprocessing and waste-management operations. Over 20% of this waste is in the form of an insoluble sludge that will require slurry modification before its transfer to the Waste Treatment and Immobilization Plant (WTP). Specific WTP acceptance criteria for waste feed delivery describe the physical and chemical characteristics of the waste that must be met before the waste is transferred to the WTP. One challenging requirement relates to the undissolved solids (UDS) composition in a waste feed because the waste contains solid particles that settle, and their concentration and relative proportion can change during the transfer of the waste in individual batches. A key uncertainty is the ability to transfer and mix wastes with large variations in UDS concentrations and resulting settling rates. To address this uncertainty, a number of small scale mixing and settling tests have been conducted to determine the mobilization performance of variable chemistry simulants. Comparison of the size and density of the particulate for each simulant to that of southeast area Hanford sludge was made using metrics for particle mobilization, suspension, settling, and pipeline transfer where dependance on particle size and density may be different, including: 1. Settling velocity, 2. Critical shear stress for erosion, 3. Just-suspended impeller speed, and 4. Pipeline critical transport velocity. Existing high-level waste sludge data has shown the effect that increasing Al concentration has on resulting settled solids. This differential settling of particles in the sludge has the possibility of resulting in solids segregation during feed preparation and uneven particle distribution during pipeline transportation or mixer jet pump operations. Understanding the predictive capabilities of HLW solids settling and transport as well as potential remedies for addressing disparate sludge behaviors can help provide technical guidance during HLW flowsheet planning.

12 MANAGEMENT OF RADIOACTIVE AND NON-RADIOACTIVE W↗

Expanding the Domain of Applicability of Machine Learning Models with Limited Data for Drug Property Prediction

Accurate machine learning models for predicting small molecule interactions with biological targets are essential for therapeutic discovery, biothreat response, and computational drug design, but their performance is often limited for understudied targets with sparse experimental data. To address this challenge, we developed and evaluated methods to improve molecular property prediction under low-data conditions, using the NimA-related kinase (NEK) family as a proof-of-concept. This work focused on two complementary goals within the ATOM Modeling PipeLine (AMPL) and the Generative Molecular Design (GMD) loop: expanding model applicability through transfer learning, representation learning, feature scaling, sampling strategies, and active-learning-inspired compound selection; and enabling efficient virtual screening to prioritize compounds that balance predicted activity, design objectives, and synthetic accessibility.

organic↗

Data Curation for Machine Learning Applied to Geothermal Power Plant Operational Data for GOOML: Geothermal Operational Optimization with Machine Learning: Preprint

Geothermal Operational Optimization with Machine Learning (GOOML) is a transferable and extensible component-based geothermal asset modeling framework that considers complex steamfield relationships and identifies optimization prospects using a data-driven approach to physics-guided, data-centric machine learning. This framework has been used to develop digital twins that provide steamfield operators with operational environments to analyze and understand historical and forecasted power production, explore new steamfield configuration possibilities, and seek optimal asset management in real world applications. To create, test, and apply the GOOML framework, diverse time-series datasets spanning multiple years were sourced from various geothermal power plant components within several complex real-world geothermal operations. These operations are based in the United States and New Zealand and include a variety of technologies, end-uses and configurations, collectively covering nearly all relevant operating conditions for modern geothermal fields. Datasets were acquired from multiple sources to ensure that machine learning experiments generalized properly to various operating conditions. It was found that the data varied in quality, format, and completeness. To ensure consistency between the various datasets, a standardized data curation process was developed to reliably streamline data preparation. This paper will discuss best practices as learned from the GOOML data curation process which takes the following steps: 1) acquisition of large quantities of data from power plant operators, 2) digestion of data to gain an initial understanding of what is included, 3) data transformation, which includes converting the data into a standardized machine-readable format so that they can be visualized, quality checked, and cleaned, 4) quality assurance and quality control, involving identification of significant data gaps and apparent anomalies through mapping of data features to real world componentry via the GOOML historical model, followed by discussion with modelers and power plant operators to identify additional data needs and to resolve issues, 5) use in machine learning algorithms, and 6) repetition of steps one through five until all data needs are met and data are deemed suitable for producing trustworthy modeling results which may be disseminated, ideally along with the curated dataset. This iterative process is focused on improving the quality of the data rather than tuning machine learning model parameters and supports a shift towards data-centric AI as a means to improving real-world applicability of geothermal machine learning projects.

access↗

Improving the Accessibility and Usability of Geothermal Information with Data Lakes and Data Pipelines on the Geothermal Data Repository: Preprint

The Geothermal Data Repository (GDR) provides universal access to data and information resulting from research and development activities funded by the Department of Energy (DOE). The GDR has extended this universal access to big data through integration with data lakes developed by the Open Energy Data Initiative (OEDI). Previously, large datasets such as seismic waveform or distributed acoustic sensing (DAS) data could only be accessed by institutions with high performance data storage and compute capabilities, effectively limiting the accessibility of big data to national labs, larger universities, and major corporations. Moreover, the time and resources needed to transport big data and configure them can produce additional barriers to use. Many of the standard formats used for structured data models (also known as content models) are incapable of handling big data and can introduce additional usability problems, often requiring data to be reformatted prior to use. This paper will explore how recent integrations between the GDR and the OEDI data lake have improved the accessibility and usability of geothermal data in a big way, making the data available to a broader audience, and enabling collaborative analysis and innovation across the greater geothermal industry.

access↗

A Centralized AI Lakehouse Framework for Brain Tumor MRI Classification and Segmentation, University KPI Forecasting, and Water Potability Prediction

In many university and healthcare projects, models are built for very different data types such as tables, institutional time series, and medical images, but they are deployed as separate applications. In this work, that separation made testing and maintenance difficult because each module had its own pipeline and runtime requirements. This paper presents an integrated AI lakehouse-style implementation that runs three model pipelines inside one containerized backend. For medical imaging, we used MRI datasets from IEEE DataPort: a four-class classification set with 7012 images (5708 train/1304 test) and a segmentation set with 3063 image–mask pairs. The classification model (ResNet50 transfer learning) is evaluated using a proper train–validation–test protocol across multiple splits (80/10/10, 70/10/20, 60/10/30, and 10/30/60), achieving a test accuracy of 99.00% under the standard 80/10/10 split. Additionally, a patient-level evaluation is conducted using an external glioma dataset to provide a more realistic assessment without data leakage. The segmentation model (DeepLabV3-ResNet50) achieved 83.09% validation mIoU and 88.79% Dice score. For university KPI forecasting, we used annual IPEDS and NSF HERD data from 2010 to 2023 for three universities (BSU, EOU, and UAB). To examine the effect of preprocessing on forecasting performance, two case studies are conducted. In the first case, linear interpolation is applied to generate semester-level data. In the second case, the original annual data is used directly without interpolation. Random Forest regression and ARIMA models are evaluated using MAE, RMSE, MAPE, and R 2 . The results showed that interpolation improved apparent forecasting performance due to smoothing, while evaluation on the original annual data provided a more realistic assessment of model behavior. To further validate the framework on a larger dataset, an additional case study is conducted using a student dropout dataset. For water potability, we trained and compared multiple tabular classifiers on a large dataset (1,048,575 samples). A Random Forest model (100 trees, max depth 10) achieved 85.86% test accuracy and high recall for unsafe samples (0.8447). All modules are served via FastAPI and deployed together using Docker, with workflow automation routing requests to the correct endpoint. System-level benchmarking indicates that the backend maintains stable throughput and latency under concurrent requests.

97 MATHEMATICS AND COMPUTING↗

The Geothermal Data Repository: Ten Years of Supporting the Geothermal Industry with Open Access to Geothermal Data

The Department of Energy's (DOE) Geothermal Data Repository (GDR) is celebrating its tenth anniversary! Over the last decade it has grown from the simple idea of storing public data in a centralized location to a valuable tool at the center of the US geothermal scientific community and an integral part of the DOE Geothermal Technologies Office (DOE GTO) project management strategy. Researchers funded by the DOE GTO have contributed over 1,300 data submissions to the GDR. These data have been used to further advancements in geothermal science, economic analysis, exploration, research, development, and operational efficiency. The adoption of open data methodologies and a data management strategy that prioritizes universal open access and standardized, interoperable data have further increased the value of GDR data, making them available across a distributed network of data sharing partners and improving their utility to other industries and related fields, including material science and space exploration. Incorporating feedback from users has been critical to the GDR's success, allowing it to grow over the years to meet the evolving needs of the geothermal community. This paper will explore some of many changes that occurred throughout the GDRs tenure and the lessons learned along the way, as well as highlight some of the new features and recent improvements that been implemented to support innovation, reduce duplication of effort, and advance the geothermal industry as a whole.

access↗

GOOML - Finding Optimization Opportunities for Geothermal Operations: Preprint

Geothermal Operational Optimization with Machine Learning (GOOML) is a transferable and extensible component-based geothermal asset modeling framework that considers complex steamfield relationships and identifies optimization prospects using a data-driven approach. We have used this framework to develop digital twins that provide steamfield operators with an operational environment to analyze and understand historical and forecasted power production, explore new steamfield configuration possibilities, and seek optimal asset management for real world applications. The GOOML modeling software is built on a generic component-based systems framework that allows for both historical and forecast analysis. A GOOML model can perform historical data-assimilation using first-principal thermodynamics to create a meaningful data model. Historical production data can then be coupled with a forecast framework to train machine-learning models of steamfield components to predict future outputs. This modeling environment enables digital exploration of steamfield design configurations and operational scenarios. GOOML digital twins have been developed for steamfields in New Zealand and the United States representing differing power generation and field conditions. These digital twins have been validated by comparing hindcast predictions against historical production data. Reinforcement learning experiments were conducted to demonstrate the ability to programmatically explore the operations space using machine learning agents. Our initial results are compelling; two to five percent increases in annual energy production were demonstrated by the GOOML models with no additional infrastructure build required. GOOML offers a new approach to geothermal operations by applying state-of-the-art machine learning algorithms, comprehensive data analytics, and interaction with digital twins. Through application of these tools, operators will realize greater availability and higher net generation which will increase the cost effectiveness of geothermal energy projects.

access↗

Improving the Accessibility and Usability of Geothermal Information with Data Lakes and Data Pipelines on the Geothermal Data Repository

The Geothermal Data Repository (GDR) provides universal access to data and information resulting from research and development activities funded by the Department of Energy (DOE). The GDR has extended this universal access to big data through integration with data lakes developed by the Open Energy Data Initiative (OEDI). Previously, large datasets such as seismic waveform or distributed acoustic sensing (DAS) data could only be accessed by institutions with high performance data storage and compute capabilities, effectively limiting the accessibility of big data to national labs, larger universities, and major corporations. Moreover, the time and resources needed to transport big data and configure them can produce additional barriers to use. Many of the standard formats used for structured data models (also known as content models) are incapable of handling big data and can introduce additional usability problems, often requiring data to be reformatted prior to use. This paper will explore how recent integrations between the GDR and the OEDI data lake have improved the accessibility and usability of geothermal data in a big way, making the data available to a broader audience, and enabling collaborative analysis and innovation across the greater geothermal industry.

access↗