Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “data repositories”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

The EGS Collab Project – Stimulations at Two Depths

The EGS Collab project, supported by the US Department of Energy, is performing intensively monitored rock stimulation and flow tests at the 10-m scale in an underground research laboratory to address challenges in implementing enhanced geothermal systems (EGS). Data and observations from the field tests are compared to simulations to understand processes and build confidence in numerical modeling of the processes. We have completed Experiment 1 (of 3), which examined hydraulic fracturing in a well-characterized underground fractured phyllite test bed at a depth of approximately 1.5 km at the Sanford Underground Research Facility (SURF) in Lead, South Dakota. Testbed characterization included fracture mapping, borehole acoustic and optical televiewers, full waveform sonic, conductivity, resistivity, temperature, campaign p- and s-wave investigations and electrical resistance tomography. Borehole geophysical techniques including passive seismic, continuous active source seismic monitoring, electrical resistance tomography, fiber-based distributed strain, distributed temperature, and distributed acoustic monitoring, were used to carefully monitor stimulation events and flow tests. More than a dozen stimulations and nearly one year of flow tests were performed. Quality data and detailed observations were collected and analyzed during stimulation and water flow tests using ambient temperature and chilled water. We achieved adaptive control of the tests using real-time monitoring and rapid dissemination of data and near-real-time simulation. More detailed numerical simulation was performed to answer key experimental design questions, forecast fracture propagation trajectories and extents, and analyze and evaluate results. Data are freely available from the Geothermal Data Repository. Experiment 2 examines the potential for hydraulic shearing in amphibolite at a depth of about 1.25 km at SURF. This site has a different set of stress and fracture conditions than Experiment 1. The Experiment 2 testbed consists of nine subhorizontal boreholes configured in two fans of two boreholes which surround the testbed and contain grouted-in electrical resistance tomography, seismic sensors, active seismic sources and distributed fiber sensors. A “five-spot” set of test wells that extends from a custom mined alcove includes an injection well and four production/monitoring wells. The testbed was characterized geophysically and hydrologically, and three stimulations have been performed using the Step-Rate Injection Method for Fracture In-Situ Properties (SIMFIP) tool to measure strains, and a new strain quantifying tool (downhole robotic strain analysis tool -DORSA) was deployed in a monitoring hole during stimulation. Real-time data were broadcast during stimulations to allow real-time response to arising issues.

EGS Collab, Enhanced Geothermal Systems, EGS, fiel↗

Observation of spin‑wave altermagnetic splitting in MnF2

-Contents of the Data Repository: - Polarized neutron diffraction data acquired in the (HK0) scattering plane. - Inelastic neutron scattering (INS) data from both unpolarized and polarized measurements with an incident neutron energy of Ei =9 meV.- - Reduced multidimensional single-crystal datasets (MDE) and corresponding S(Q,ω) slices used to generate all figures presented in the manuscript. - Julia source code and supporting input files used for spin-wave calculations, model fitting, and simulation of neutron scattering intensity maps.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

Secure Data Logging and Processing with Blockchain and Machine Learning (Final Report)

Secure Data Logging and Processing with Blockchain and Machine Learning (ML) research is focused on the development of a platform to securely log and process sensor data in fossil power plants. The platform integrates two emerging technologies, blockchain and ML, and incorporates several innovative mechanisms to ensure the integrity, reliability, and resiliency of power systems. The goal is to protect the power plant from various cyberattacks such as false data injection and denial of service attacks using these technologies. The research goal was enabled by the following Research Project Objectives: 1) Secure authentication and identity verification of sensor nodes, actuators, and other equipment within a network. 2) Development of mechanisms that ensure only data sent by legitimate sensors are accepted and stored in the data repository. 3) Development of data aggregation methodologies using ML / Deep Learning (DL) algorithms to minimize noise / faulty data. 4) Implementation of the blockchain technologies to provide data security using secured IOTA framework & nodes.

20 FOSSIL-FUELED POWER PLANTS↗

IT Challenges for Space Medicine

This viewgraph presentation reviews the various Information Technology challenges for aerospace medicine. The contents include: 1) Space Medicine Activities; 2) Private Medical Information; 3) Lifetime Surveillance of Astronaut Health; 4) Mission Medical Support; 5) Data Repositories for Research; 6) Data Input and Output; 7) Finding Data/Information; 8) Summary of Challenges; and 9) Solutions and questions.

Johnson-Throop, Kathy↗

Elevating the Quality of Space Omics Sequencing Data: Innovations and Methodologies from NASA GeneLab Sample Processing Laboratory

NASA’s GeneLab, part of the NASA Open Science Data Repository, is a space-related database that hosts a diverse range of transcriptomics, proteomics, epigenomics and genomics data. The NASA GeneLab Sample Processing Laboratory (SPL) generates omics data from biological experiments conducted aboard the International Space Station, Space Shuttle and space related ground experiments, this omics data then hosted on the GeneLab repository. Samples generated such experiments pose numerous technical challenges such as small experimental sample size, variance in dissection times, limited tissue preservation methods, prolonged storage time, and more. GeneLab SPL team had developed specialized expertise in nucleic acid extraction, library preparation and sequencing of such biological samples via extensive training and years of experience. In order to ensure data accuracy and consistency across experiments, SPL has developed standardized protocols for each species and tissue type. These protocols in conjunction with quality control metrics and data standards are crucial in generating of high-quality data. SPL protocols and standards have been developed in collaboration with the scientific community and had been made publicly available on the GeneLab portal, guaranteeing comparability of datasets across spaceflight experiments. To ensure reliability of data generation, SPL leverages cutting-edge innovations in laboratory automation for sample processing. By leveraging these state-of-the-art platforms, SPL achieves high levels of data reproducibility while significantly minimizing sources of bias and variability, especially across experiments with large numbers of samples. Over the past few years, the space biology investigator community has accessed SPL-generated data from the Open Science Data Repository for a myriad of data re-analysis and re-use studies. We observe a trend that in-house SPL-generated data consistently outperforms outsourced sequencing data in terms of technical standards, quality control metrics, timeliness of data delivery, and sequencing and reagent efficiency. Superior data generation has and will continue to enable discoveries in disease, diagnostic tools, and the biological effects of long duration spaceflight.

GeneLab↗

Elevating the Quality of Space Omics Sequencing Data: Innovations and Methodologies from NASA GeneLab Sample Processing Laboratory

NASA’s GeneLab, part of the NASA Open Science Data Repository, is a space-related database that hosts a diverse range of transcriptomics, proteomics, epigenomics and genomics data. The NASA GeneLab Sample Processing Laboratory (SPL) generates omics data from biological experiments conducted aboard the International Space Station, Space Shuttle and space related ground experiments, this omics data then hosted on the GeneLab repository. Samples generated such experiments pose numerous technical challenges such as small experimental sample size, variance in dissection times, limited tissue preservation methods, prolonged storage time, and more. GeneLab SPL team had developed specialized expertise in nucleic acid extraction, library preparation and sequencing of such biological samples via extensive training and years of experience. In order to ensure data accuracy and consistency across experiments, SPL has developed standardized protocols for each species and tissue type. These protocols in conjunction with quality control metrics and data standards are crucial in generating of high-quality data. SPL protocols and standards have been developed in collaboration with the scientific community and had been made publicly available on the GeneLab portal, guaranteeing comparability of datasets across spaceflight experiments. To ensure reliability of data generation, SPL leverages cutting-edge innovations in laboratory automation for sample processing. By leveraging these state-of-the-art platforms, SPL achieves high levels of data reproducibility while significantly minimizing sources of bias and variability, especially across experiments with large numbers of samples. Over the past few years, the space biology investigator community has accessed SPL-generated data from the Open Science Data Repository for a myriad of data re-analysis and re-use studies. We observe a trend that in-house SPL-generated data consistently outperforms outsourced sequencing data in terms of technical standards, quality control metrics, timeliness of data delivery, and sequencing and reagent efficiency. Superior data generation has and will continue to enable discoveries in disease, diagnostic tools, and the biological effects of long duration spaceflight.

GeneLab↗

Enabling pan-repository reanalysis for big data science of public metabolomics data

Public untargeted metabolomics data is a growing resource for metabolite and phenotype discovery; however, accessing and utilizing these data across repositories pose significant challenges. Therefore, here we develop pan-repository universal identifiers and harmonized cross-repository metadata. This ecosystem facilitates discovery by integrating diverse data sources from public repositories including MetaboLights, Metabolomics Workbench, and GNPS/MassIVE. Our approach simplified data handling and unlocks previously inaccessible reanalysis workflows, fostering unmatched research opportunities.

El Abiead, Yasin↗

University Data Management Pilot Utilizing the Nuclear Research Data System

Background In 2022, the Office of Science and Technology Policy (OSTP) issued a memo that significantly reshaped the landscape of access to federally funded research. The memo mandated that all taxpayer-funded research be made available to the public without delay upon publication, without an embargo period, superseding the 2013 OSTP public access policy. This public access policy promotes transparency and the democratization of knowledge, ensuring that the fruits of scientific endeavors funded by federal agencies could be immediately accessed and built upon by scientists, educators, students, and the public at large. To implement the requirements of the OSTP guidance and DOE Public Access Plan, the Office of Nuclear Energy (NE) has implemented public access plan guidance and has identified several areas where better data management practices would further expand public access to important nuclear energy related scientific data, reports, and other technical products. Significant NE supported efforts are already underway for data management and public access to important nuclear energy related data.1 2 To address gaps in data management practices, and improve retention and accessibility of data, NE is actively exploring enhanced data management options utilizing its high-performance computing resources administered by its Nuclear Scientific User Facility Program. A newly piloted system, the Nuclear Research Data System (NRDS) acts as a portal for data collection and dissemination. Nuclear Energy University Program Research and Development Portfolio According to Web of Science, NEUP has produced 2,345 journal publication that have been cited more than 61,000 times3 and countless conference proceedings. These publications are publicly available through OSTI.gov and in the open literature. Additional scientific and technical products including project milestones that are not publications and NEUP project final reports are vetted through OSTI.gov and released once reviewed and approved by DOE. Since 2009, NEUP has awarded close to 1,000 different R&D projects in technical areas across the NE research programs. As of June 2023, 512 NEUP reports are publicly available on OSTI. The underlying data for projects is still held at universities, and data transfer, co-location, and dissemination has not occurred in a systematic way. NEUP data is currently accessible through myriad university-based data repositories, or through direct requests to PIs. The program identified this patchwork of repositories, or often lack of publicly available data, as a significant barrier to an organized, accessible, and comprehensive solution to sharing data with the larger nuclear energy community. Approach The goal of this pilot project is to establish a pathway to a consolidated long-term repository for NEUP project data. To accomplish this goal, the pilot strives to accomplish the following objectives: Establish data collection standards, including a standard set of required supplementary information to contextualize and support raw data files. Work with the HPC group collect and upload information and to modify the NRDS system, as needed, to support a standardized approach. Resolve potential barriers to successful roll out of an expanded data collection strategy, including modifying data management plan guidelines and establishing a document and data release process that accounts for potential intellectual property and/or export control concerns. Results Overall, the pilot was successful in collecting 8,982 raw and processes data files, 220 reports, 56 calibration files, and 5,931 other supplementary documents. Supplementary documents included experimental plans, methods, journal publications and conference proceedings, milestone reports, and final reports. Figure 2 shows the number of data sets and supplementary project information provided by each project. Projects has significantly different input, depending on experimental data produced and completeness of the datasets provided.

Data collection↗

Hydropower Capacity Factor Trends & Analytics for the United States

This data repository contains all code, input data, and data generated for Turner et al. (2024)—“Hydropower capacity factors trending down in the United States”. File descriptions: – hydro-cf-trends-inputs.zip: Full set of input data used in this study, organized for direct entry into “/data” directory of hydro-cf-trends data processing pipeline. – hydro-cf-trends.zip: Full data processing pipeline, coded using the R {targets} framework. This is a snapshot release (v1.0) of the code repository stored at https://code.ornl.gov/turnersw/hydro-cf-trends/. – hydro-cf-trends-results.zip: Provides all dam level results required to reproduce results and graphics in Turner et al. (2024). Dams are identified by the “complxID” (root of the hydropower plant ID in the Existing Hydropower Assets Database, inherited from HILARRI). Results include: • dam_CF_trends.csv: Table of long-term trends in annualized capacity factors for 610 dams and modeled annualized capacity factors for 362 modeled dams (naturalized and assimilated flows). • dam_annualized_CF_gen.csv: Annualized time series of the following variables for each of 610 hydropower dams with nameplate > 5MW – Reported nameplate capacity (MW) – Implied maximum annual generation (MWh) – Reported net generation (MWh) – Computed annual capacity factor – Modeled annual capacity factor (362 modeled plants only)

13 HYDRO ENERGY↗

Pathways to a Sustainable Aviation Ecosystem: Flight DNA: An Anonymized Aviation Data Tool and Repository

The National Renewable Energy Laboratory (NREL) has deep experience developing secure, national data repositories, which it augments with analysis, technology and market expertise, high-performance computing, and innovative data visualization. By adapting the architecture of existing mobility databases (i.e., Fleet DNA, the Transportation Secure Data Center, the National Fuel Cell Technology Evaluation Center), NREL can build a powerful aviation data clearinghouse - Flight DNA - that helps stakeholders navigate the web of pitfalls and possibilities generated by new aviation technologies.

aviation↗

AmeriFlux BASE data pipeline to support network growth and data sharing

Abstract AmeriFlux is a network of research sites that measure carbon, water, and energy fluxes between ecosystems and the atmosphere using the eddy covariance technique to study a variety of Earth science questions. AmeriFlux’s diversity of ecosystems, instruments, and data-processing routines create challenges for data standardization, quality assurance, and sharing across the network. To address these challenges, the AmeriFlux Management Project (AMP) designed and implemented the BASE data-processing pipeline. The pipeline begins with data uploaded by the site teams, followed by the AMP team’s quality assurance and quality control (QA/QC), ingestion of site metadata, and publication of the BASE data product. The semi-automated pipeline enables us to keep pace with the rapid growth of the network. As of 2022, the AmeriFlux BASE data product contains 3,130 site years of data from 444 sites, with standardized units and variable names of more than 60 common variables, representing the largest long-term data repository for flux-met data in the world. The standardized, quality-ensured data product facilitates multisite comparisons, model evaluations, and data syntheses.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Designing and Implementing a Distributed System Architecture for the Mars Rover Mission Planning Software (Maestro)

Distributed systems allow scientists from around the world to plan missions concurrently, while being updated on the revisions of their colleagues in real time. However, permitting multiple clients to simultaneously modify a single data repository can quickly lead to data corruption or inconsistent states between users. Since our message broker, the Java Message Service, does not ensure that messages will be received in the order they were published, we must implement our own numbering scheme to guarantee that changes to mission plans are performed in the correct sequence. Furthermore, distributed architectures must ensure that as new users connect to the system, they synchronize with the database without missing any messages or falling into an inconsistent state. Robust systems must also guarantee that all clients will remain synchronized with the database even in the case of multiple client failure, which can occur at any time due to lost network connections or a user's own system instability. The final design for the distributed system behind the Mars rover mission planning software fulfills all of these requirements and upon completion will be deployed to MER at the end of 2005 as well as Phoenix (2007) and MSL (2009).

Goldgof, Gregory M.↗

Biological Data for Deep Space Mission Support

Increased biomedical risks and challenges associated with deep space missions (cis-Lunar, Mars transit, Mars surface) require new knowledge discovery and development of novel ecosystem and biomedical support capabilities. This paradigm shift supporting distant and long-duration missions requires biological data to be findable, accessible, interoperable, reusable (FAIR), and maximally open-access (i.e., there is a data governance continuum from closed to mediated to embargoed to open). The NASA “Open Science Data Repositories” (OSDR) aims to meet scientific, technical, and operational spaceflight needs, and offers the ability to upload, download, search, share, analyze, and visualize data across physiological, behavioral, ‘omics, and environmental monitoring telemetry datasets. OSDR includes NASA GeneLab, NASA Ames Life Sciences Data Archive (ALSDA), and NASA Biological Institutional Scientific Collection (NBISC). In the past year, ALSDA has undergone a transformation in its data collection, curation, and architecture methods. Standardizing non-genomic (phenotypic) datasets was, and will continue to be, a challenge because of their diverse nature (e.g., molecular, cellular, tissue, whole organism behavior; micro-computed tomography, intraocular pressure, fluorescence microscopy, western blot, ultrasonography; tabular, images, video). This year ALSDA, alongside GeneLab, introduced the Biological Data Management Environment (BDME) with the purpose to accept submission of data from space relevant experiments including spaceflight, radiation, simulated gravity, gravitropism, isolation and confinement, hostile closed environments and/or distance from Earth. In addition to bringing together omics, phenotypic, physiological, bioimaging, and behavioral data into one repository. By integrating with GeneLab a multi-project submission portal aims to reduce the burden on PIs submitting data and enabling the discovery of both omics and phenotypic data. The purpose of ALSDA is to collect, curate, and make all non-human space-relevant biological data maximally findable, accessible, interoperable, and reusable (FAIR). These scope of ALSDA data collected and submitted by PIs include study design metadata, subject metadata, assay metadata (parameters), raw and processed assay data, assay imagery/video, and subject-experienced mission data telemetry (radiation, temperature, humidity, acoustics, vibrations, etc.). In 2021, a community of researchers rallied to form the ALSDA Analysis Working Group (AWG) and provided scientific consensus on dataset sample and assay metadata. The community and excitement around the ALSDA/OSDR system has already led to several data reuse studies, demonstrating value using machine learning (ML), knowledge graphs, and meta-analysis approaches.

space biology↗

ESS-DIVE guidelines for archiving terrestrial model data

This dataset contains supporting documents and images for ESS-DIVE terrestrial model data archiving guidelines.Terrestrial models are broadly defined as numerical models that couple both land dynamics and energy, water, carbon, or nutrient fluxes. We created these guidelines based on input from the U.S. Department of Energy’s Biological and Environmental Research land modeling community. The guidelines are intended to help modelers determine which components of their terrestrial model data associated with publication should be archived. Based on input from the land modeling community, the guidelines recommend archiving both model input and testing data, as well as code, script, and metadata. The guidelines also recommend archiving model data output, depending on the limitations set by data repositories. Lastly, we provide recommendations for bundling data files for publication as well as a discussion about tools that can facilitate model data archiving and reuse.This dataset is an archive of the associated GitHub repository for our model archiving guidelines (https://github.com/ess-dive-community/essdive-model-data-archiving-guidelines). The ‘README.pdf’ file gives a general introduction to the guidelines, and the ‘instructions.pdf’ file provides more detailed steps for following the guidelines. We also provide 2 figures in this data package: 1) a decision tree (model_data_guidelines_decision_tree.png) that can help users determine which components of their model data to archive. and 2) the ‘model_data_guidelines_flmd.png’ file depicts the different files that can be archived in addition to the model data itself. Lastly, we include 3 digitized tables from our associated manuscript and 3 CSV files with anonymized input from DOE scientists about the importance of different aspects of model data archiving from which we developed the guidelines.Dataset updates for v1.1.0: We updated this data package on 2021-11-22 in response to review comments on our related manuscript. In this update we removed one figure so that the model archiving guidelines are conveyed in text rather than an image. We updated the file-level metadata (FLMD) figure to be in accord with the most recent FLMD recommendations. We made minor edits to the README file to update the recommended citation and added two co-authors. We also added 6 new data files (3 are anonymized input from DOE scientists that helped to inform guidelines, and 3 are digitized tables from our manuscript.

54 ENVIRONMENTAL SCIENCES↗

The Colorado East River Community Observatory Data Collection

Abstract The U.S. Department of Energy's (DOE) Colorado East River Community Observatory (ER) in the Upper Colorado River Basin was established in 2015 as a representative mountainous, snow‐dominated watershed to study hydrobiogeochemical responses to hydrological perturbations in headwater systems. The ER is characterized by steep elevation, geologic, hydrologic and vegetation gradients along floodplain, montane, subalpine, and alpine life zones, which makes it an ideal location for researchers to understand how different mountain subsystems contribute to overall watershed behaviour. The ER has both long‐term and spatially‐extensive observations and experimental campaigns carried out by the Watershed Function Scientific Focus Area (SFA), led by Lawrence Berkeley National Laboratory, and researchers from over 30 organizations who conduct cross‐disciplinary process‐based investigations and modelling of watershed behaviour. The heterogeneous data generated at the ER include hydrological, genomic, biogeochemical, climate, vegetation, geological, and remote sensing data, which combined with model inputs and outputs comprise a collection of datasets and value‐added products within a mountainous watershed that span multiple spatiotemporal scales, compartments, and life zones. Within 5 years of collection, these datasets have revealed insights into numerous aspects of watershed function such as factors influencing snow accumulation and melt timing, water balance partitioning, and impacts of floodplain biogeochemistry and hillslope ecohydrology on riverine geochemical exports. Data generated by the SFA are managed and curated through its Data Management Framework. The SFA has an open data policy, and over 70 ER datasets are publicly available through relevant data repositories. A public interactive map of data collection sites run by the SFA is available to inform the broader community about SFA field activities. Here, we describe the ER and the SFA measurement network, present the public data collection generated by the SFA and partner institutions, and highlight the value of collecting multidisciplinary multiscale measurements in representative catchment observatories.

54 ENVIRONMENTAL SCIENCES↗

The disCO2ver Platform: Curating Data and Tools for Geologic Carbon Sequestration and Deep Subsurface Research Systems

The U.S. DOE National Energy Technology Laboratory has invested 12+ years of development into the data repository and digital laboratory, the Energy Data eXchange (EDX, edx.netl.doe.gov). Supporting a variety of research areas across the DOE Office of Fossil Energy and Carbon Management, the platform has successfully curated and preserved thousands of data products from DOE research. The Carbon Storage Program has successfully supported data curation, upload, and publishing of data products on EDX for many years, demonstrating a success story of how resources like EDX can effectively help with long term preservation and publishing of DOE data products. EDX continues to shift towards cloud-supported infrastructure, taking a hybrid approach combining on-premises compute and storage integrated with cloud-hosted services. The integration of cloud compute and hybrid architecture enables the development of EDX-hosted platforms that tailor the data and tools hosted on them to a specific community, enables implementation of machine learning tools for data discovery and filtering, and enables the hosting of virtual (online user interface) tools. Geologic carbon sequestration (GCS) research continues to scale up in response to the current administration goals to reduce greenhouse gas emissions and transition the energy economy. Over the last year, EDX’s disCO2ver platform has been developed in response to the need for access to data products and tools to support the scaling up of GCS research. disCO2ver provides access to data resources and tools, produced by DOE and outside authoritative external resources. The platform also provides a user-access control component for the virtualization and cloud hosting of tools. Tools that need to be virtualized, to eliminate the need for users to download the tool and use local compute resources, is essential to supporting big-data analysis and machine learning that is becoming common place in carbon storage modeling, risk analysis, and data publishing practices. This talk will review the EDX’s disCO2ver platform and the current work ongoing to curate data and tools to support GCS and deep subsurface systems research.

Morkner, Paige↗

Toward Drilling the Perfect Geothermal Well: An International Research Coordination Network for Geothermal Drilling Optimization Supported by Deep Machine Learning and Cloud Based Data Aggregation

The EDGE project, supported by the U.S. Department of Energy Geothermal Technologies Office under award DE-EE0008793, established a data-driven framework for improving the efficiency, cost-effectiveness, and reliability of geothermal well drilling. The project focused on developing scalable data infrastructure, advanced machine learning and probabilistic models, and integrated analytics tools to support continuous drilling optimization. A central objective was to reduce geothermal drilling costs by up to seventy percent while minimizing the risk of well failure through predictive diagnostics and adaptive planning. Over the project period, a comprehensive data repository was designed and deployed, incorporating records from over one hundred geothermal wells across varied geological settings. This repository supported both structured and unstructured data and adhered to FAIR data principles, enabling provenance tracking, quality control, and standardized metadata. The project introduced automated ingestion pipelines and a cloud-hosted platform that facilitated access to raw, processed, and derived datasets. This infrastructure served as the foundation for model development and analysis. Machine learning workflows were developed to predict key drilling metrics including rate of penetration, non-productive time, and total drilling costs. Self-organizing maps and dimensionality reduction methods were used to uncover operational patterns and outliers, while supervised learning algorithms such as random forests and deep neural networks were applied to forecast performance outcomes. The models were validated on heterogeneous datasets from both U.S. and Icelandic fields, demonstrating variable but significant predictive accuracy. The results indicated that finer temporal resolution, inclusion of lithological data, and consistency in operational annotations could substantially improve model performance. The project also implemented process mining techniques to reconstruct state-transition models from drilling event logs. These models enabled the identification of deviations from optimal workflows and provided insights into recurring failure modes. Analysis of non-productive time highlighted the impact of equipment failures, geological challenges, and human factors, offering opportunities for targeted mitigation strategies. The EDGE Dashboard was developed as a web-based expert system integrating data visualization, model outputs, and user-driven queries. It provided an accessible interface for operators to explore historical data, evaluate predicted outcomes, and compare drilling scenarios. Initial feedback from project partners suggested that the dashboard could serve as a foundation for more advanced advisory and optimization tools. Overall, the EDGE project demonstrated the feasibility and value of applying modern data science techniques to geothermal drilling. It delivered a set of interoperable tools and models that can support more efficient, lower-risk well development. The findings point toward a viable path for transitioning from advisory analytics to semi-autonomous drilling systems, contingent on continued collaboration, expanded datasets, and field validation. The project results have immediate relevance for drilling operations, data management practices, and future geothermal R&D efforts aimed at achieving reliable, cost-competitive geothermal energy at scale.

15 GEOTHERMAL ENERGY↗

Enabling Open and Interoperable Science: Multi-Omics Data Processing Platform with NASA GeneLab Standardized Bioinformatics Workflows for Space and Earth Research

Multi-omics biological data continues to be generated at an astounding pace. Genomics, transcriptomics, metabolomics, and proteomics, or collectively known as multi-omics data, are used to assess biological functions, and provide invaluable insights into human, animal, plant, and environmental health both on Earth and in Space. Despite the abundance of these valuable data, the need for bioinformatics expertise, particularly as it relates to the niche filed of space biology, and a lack of accessible resources for processing these data limit their usefulness in deriving biological insights. The NASA Open Science Data Repository (OSDR) provides access to omics data from various spaceflight and analog studies. To enhance the accessibility and reusability of these data, GeneLab (part of OSDR) designs and implements standardized, community-driven, open-source bioinformatics workflows to transform raw omics data into standardized processed data. Currently, GeneLab-processed data from hundreds of space studies have been reused for meta-analyses. This has led to new insights and scientific publications that extend beyond the initial research, thereby enriching our understanding of molecular-scale biological responses to the space environment. To make these bioinformatics workflows open and accessible, GeneLab teamed up with DOE-funded initiatives, including the National Microbiome Data Collaborative (NMDC), to create the NASA EDGE [Empowering the Development of Genomics Expertise] Bioinformatics web-based platform. NASA EDGE utilizes shared compute resources to run the GeneLab standardized bioinformatics workflows, which eliminates the need for researchers to have their own high performance computing cluster. The web-based platform makes complicated biological analyses incredibly easy to perform, thus expanding the reach of these analyses to bioinformatics novices, students, and even citizen scientists enabling them to contribute to scientific discoveries and progress. The authors will demonstrate how the NASA EDGE platform can be used to process microbial omics data hosted on OSDR as well as user-generated omics datasets using GeneLab’s standard workflows.

Amanda M. Saravia-Butler↗