Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “metadata extraction”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

SPASE: The Connection Among Solar and Space Physics Data Centers

The Space Physics Archive Search and Extract (SPASE) project is an international collaboration among Heliophysics (solar and space physics) groups concerned with data acquisition and archiving. Within this community there are a variety of old and new data centers, resident archives, "virtual observatories", etc. acquiring, holding, and distributing data. A researcher interested in finding data of value for his or her study faces a complex data environment. The SPASE group has simplified the search for data through the development of the SPASE Data Model as a common method to describe data sets in the various archives. The data model is an XML-based schema and is now in operational use. There are both positives and negatives to this approach. The advantage is the common metadata language enabling wide-ranging searches across the archives, but it is difficult to inspire the data holders to spend the time necessary to describe their data using the Model. Software tools have helped, but the main motivational factor is wide-ranging use of the standard by the community. The use is expanding, but there are still other groups who could benefit from adopting SPASE. The SPASE Data Model is also being expanded in the sense of providing the means for more detailed description of data sets with the aim of enabling more automated ingestion and use of the data through detailed format descriptions. We will discuss the present state of SPASE usage and how we foresee development in the future. The evolution is based on a number of lessons learned - some unique to Heliophysics, but many common to the various data disciplines.

Thieman, James R.↗

Comprehensive Database of Environmental Mitigations Extracted from FERC-Licensed Hydropower Projects Using Artificial Intelligence Techniques, 1998-2023

This dataset provides a comprehensive inventory of environmental mitigation measures required by Federal Energy Regulatory Commission (FERC) licensed hydropower facilities from 461 licenses that were issued from 1998 to 2023. These licenses constitute 446 of the 1015 FERC projects that were active at the end of 2023. 17,612 mentions of environmental mitigations were identified and categorized in 128 unique categories. Mitigations were identified using a Natural Language Processing (NLP) approach, specifically with a Bidirectional Encoder Representations from Transformer (BERT) model. Model-derived results were then reviewed and updated by a subject matter expert as needed. This dataset introduces important enhancements to previous efforts to inventory environmental mitigations, such as including associated license text for each mitigation, tracking the number of instances a mitigation was identified within a license, and providing improved location information. These enhancements significantly expand the dataset's utility, offering greater analytical capabilities and ensuring reproducibility. The dataset is downloadable as a zip file containing the metadata and dataset files.

Ruggles, Thomas [Oak Ridge National Laboratory (OR↗

Fe(III) reducing bacterial activities in Old Woman Creek wetland sediments, June 2023

To evaluate the Fe(III) reducing microbiological activities in Old Woman Creek Nature Preserve (OWC) wetland sediments, we incubated OWC sediments under anoxic and oxic conditions and with or without Fe(III) amendment [as hydrous ferric oxide (HFO)]. No Fe(III) reduction was observed in heat-deactivated incubations. In non-sterile anoxic incubations, measurement of 0.5 M HCl-extractable Fe(II) indicated that Fe(III) reduction occurred in both Fe(III)-amended and -unamended incubations, indicating that abundant Fe(III) is associated with the OWC sediments. Little Fe(II) accumulated in solution, indicating that the most biogenic Fe(II) adsorbs to the sediments. When air was added to the headspace of non-sterile incubations, Fe(III) reduction was halted and any biogenic Fe(II) that accumulated was oxidized. These experiments were used to guide preparation and analyses of incubations to determine if electrochemical measuements can be used to detect microbiological activities in contrasting terminal electron accepting regimes (i.e., aerobic and Fe(III) reducing conditions). Data package includes methods and data from experiments, including dissolved anion concentrations, dissolved Fe(II) concentrations, and 0.5 M HCl-extractable Fe(II) concentrations. All files are either .txt or .csv and can be opened by any plain text editor application.

EARTH SCIENCE↗

The catalog-to-cosmology framework for weak lensing and galaxy clustering for LSST

We present TXPipe, a modular, automated and reproducible pipeline for ingesting catalog data and performing all the calculations required to obtain quality-assured two-point measurements of lensing and clustering, and their covariances, with the metadata necessary for parameter estimation. The pipeline is developed within the Rubin Observatory Legacy Survey of Space and Time (LSST) Dark Energy Science Collaboration (DESC), and designed for cosmology analyses using LSST data. In this paper, we present the pipeline for the so-called 3x2pt analysis -- a combination of three two-point functions that measure the auto- and cross-correlation between galaxy density and shapes. We perform the analysis both in real and harmonic space using TXPipe and other LSST-DESC tools. We validate the pipeline using Gaussian simulations and show that it accurately measures data vectors and recovers the input cosmology to the accuracy level required for the first year of LSST data under this simplified scenario. We also apply the pipeline to a realistic mock galaxy sample extracted from the CosmoDC2 simulation suite (Korytov et al. 2019). TXPipe establishes a baseline framework that can be built upon as the LSST survey proceeds. Furthermore, the pipeline is designed to be easily extended to science probes beyond the 3x2pt analysis.

79 ASTRONOMY AND ASTROPHYSICS↗

Metagenome-assembled genomes from topsoils collected during NEON campaign in East River, CO (06/14/2018-06/28/2018)

The Watershed Function Science Focus Area (WF SFA) at Lawrence Berkeley National Lab is working to build a mechanistic understanding of the distribution and dynamics of biogeochemical processes in mountainous watersheds and their response to perturbation. In June 2018, the NEON (National Ecological Observatory Network) Airborne Observatory Platform (AOP) performed a taskable airborne imaging campaign to collect visible to shortwave infrared (VSWIR) imaging spectroscopy and LiDAR data across 330 km2 in the Upper East River at Crested Butte, CO. We conducted a parallel ground sampling campaign to sample vegetation traits, as well as soil physical, chemical, and microbiological characteristics. We collected these samples from 438 sites across 12 locations spanning much of the elevation, topographic, and geologic variability across the study area. A subset of 250 samples were used for soil metagenomics which is presented here. In addition, at each site, vegetation samples were collected to measure species-specific leaf water content and leaf mass area, foliar elemental composition and foliar CN stable isotope ratios. Soil samples were collected to measure soil physical properties which include bulk density and soil texture analysis. A suite of soil chemical properties was measured from the samples collected at each site, including pH, organic matter, concentrations exchangeable cations, total elemental composition, and the concentrations of extractable N pools (e.g. total free amino acids, ammonium, nitrate, dissolved organic N, and total dissolved N). Additionally, we have measured soil microbial biomass CN stoichiometry. Here, we present 1982 metagenome-assembled genomes (MAGs) for the bacterial and archaeal community from topsoil collected from during NEON 2018 campaign. All metagenomes were sequenced at JGI (Joint Genome Institute) (GOLD Study ID: Gs0149986). Metagenomes were assembled using JGI Metagenome Workflow (10.1128/mSystems.00804-20). The dataset includes (1) zip files for 1982 MAG fasta files (neon_genomes1-5.tar.gz, split into 5 tarballs to keep tarballs under 0.5 GB), (2) neon_Gs0149986_samples_soilproperties_metagenomes.csv: the sample information together with the accession numbers for the underlying metagenomes and the associated soil physical and chemical measurements in NMDC (National Microbiome Data Collaborative) compliant format, (3) neon_Gs0149986.kml: location bounding box file for the sampled locations, (4) samples.csv: sample metadata file used to register Internationall Generic Sample Numbers (IGSNs), (5) flmd.csv: file level metadata file, and (6) dd.csv: data dictionary file. This work was supported by the Watershed Function Science Focus Area at Lawrence Berkeley National Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-05CH11231.

2018 NEON and 2025 CHESS Campaigns↗

SPRUCE FT-ICR MS, Bulk Chemistry, and Mass Loss from Litter Decomposition Study in Experimental Plots, Marcell Experimental Forest, Minnesota, 2015-2017

This dataset contains molecular, bulk chemical, and mass loss measurements from a litter decomposition study at the Spruce and Peatland Responses Under Changing Environments (SPRUCE) experimental site within the Marcell Experimental Forest in northern Minnesota, USA. This site is in a Sphagnum spp. ombrotrophic bog forest. Litterbags were deployed into the peat in September 2015 across three warming levels (+0, +4.5, and +9°C) under ambient and elevated carbon dioxide (CO₂ - +500 ppm) and retrieved after roughly 0.5, 1, and 2 years of field incubation (2015-09-23 to 2017-08-02). Litterbags containing six peatland litter types: black spruce needles (Picea mariana - SPL), spruce fine roots (SPR), Sphagnum angustifolium (ANG), Sphagnum magellanicum (MAG), Labrador tea leaves (Rhododendron groenlandicum - LTL), and Labrador tea roots (LTR). Molecular composition of water-soluble organic matter extracts was characterized using Fourier Transform Ion Cyclotron Resonance Mass Spectrometry (FT-ICR MS) at 9.4 Tesla, operated in negative ion mode with electrospray ionization, providing molecular formula assignments and compound-class distributions across the decomposition time series. Bulk chemical characterization included elemental analysis (percent carbon, nitrogen, and phosphorus) and Fourier Transform Infrared Spectroscopy (FTIR) to quantify functional group composition. Litter mass loss was tracked gravimetrically at each retrieval interval, expressed as percent mass remaining relative to initial dry mass for each litter type and treatment combination. These data are valuable for understanding how vegetation shifts driven by increased atmospheric CO2 and temperature in peatlands alter litter inputs and organic matter stabilization trajectories, with implications for projecting and modeling peatland carbon cycling. This dataset contains two data files in comma-separated value (.csv) format. Additional metadata are provided: two data dictionaries and a file-level metadata file in comma separate (.csv) format and a user guide in PDF (*.pdf) format.

decomposition↗

CHESS 2025: Waveform LiDAR data from NEON AOP surveys

This dataset provides Level 1 (L1) full-waveform light detection and ranging (LiDAR) data collected for the 2025 Colorado Headwaters Ecological Spectroscopy Study (CHESS). These data were acquired to enable characterization of vegetation structure and other three-dimensional features of the land surface, and to evaluate structural changes that may have occurred between a prior LiDAR acquisition in 2018 and the 2025 overflight. Waveform LiDAR data can provide more detailed information about objects on the ground than discrete point clouds typically do, and they are often used for granular target segmentation and characterization of subcanopy vegetation. The data were acquired over three study domains in the Upper Gunnison river basin: the upper East River watershed (CRBU); Almont Triangle and Taylor Canyon (ALMO); and Upper Taylor River watershed (UPTA) between 2025-06-13 and 2025-07-15. LiDAR data were acquired using the Optech Galaxy Prime Airborne LiDAR Terrain Mapper onboard the National Ecological Observatory Network (NEON) Airborne Observation Platform (AOP). These are the primary waveform LiDAR data delivered by NEON and are provided per flightline in compressed Pulsewaves format, an open-source binary file standard. A Pulsewaves object comprises a two files: a pulse (.pls) file, which stores the geographic origin, outgoing vector, and metadata for every laser pulse emitted by the scanner, and a wave file (.wvs), which stores the sequential amplitude samples of the outgoing pulse and the returning signals. The files are published here in their compressed forms (.plz, .wvz). All waveform data were processed following the theoretical workflow described in the NEON L0-to-L1 Waveform LiDAR Algorithm Theoretical Basis Document (Krause and Goulden 2022a); however, the Pulsewaves output format differs from a legacy format described in that document. Waveform amplitude samples are recorded at 1 nanosecond intervals. All coordinates are provided in meters. Horizontal coordinates are referenced in Universal Transverse Mercator (UTM) zone 13N and the World Geodetic System (WGS) 1984 ensemble datum. Elevations are referenced to Geoid12A. Waveform data for the UPTA survey area were collected without incident and the published records are complete. However, both the ALMO and CRBU collections experienced issues that resulted in incomplete data for those areas. On collection day 2018-06-16 a hardware failure caused the waveform digitizer to lose data from the eastern edge of the ALMO site (Figure 22). The waveform data for flightlines 2–20 could not be extracted from the digitizer, and the data proved unrecoverable. As a result, a portion of the site does not have coverage with waveform data. Although no hardware failure was observed during collection over the CRBU area, final waveform files generated by vendor software contained only ~25% of the expected number of return pulses. After discovery, NEON initiated troubleshooting with the vendor. The root cause of the data ablation had not been identified at the time of publication. Additional data will be published in an update to this package if further recovery proves successful. CHESS Project Description: The Colorado Headwaters Ecological Spectroscopy Study (CHESS) comprised a multi-week airborne remote sensing and field observation campaign in the Upper Gunnison Basin, Colorado, conducted in June and July of 2025. Airborne remote sensing was conducted by the National Ecological Observatory Network Airborne Observation Platform (NEON AOP), concurrent with a field campaign run by the Rocky Mountain Biological Laboratory (RMBL), the Lawrence Berkeley National Laboratory (LBNL) and SLAC National Accelerator Laboratory Watershed Function Science Focus Area (SFA), and NASA-JPL (Jet Propulsion Laboratory) Earth Surface Mineral Dust Source Investigation (EMIT) program. Between June 10 and July 18, 2025, the NEON AOP flight team collected high-resolution aerial imaging spectroscopy and Light Detection and Ranging (LiDAR) data over three domains: the Upper East River (CRBU), Almont Triangle (ALMO), and the Upper Taylor Basin (UPTA). In coordination with the flights, a field campaign acquired ground-truth observations, including observations of vegetation composition, foliar traits, forest demography, and subsurface properties in 18 core sampling areas within the domains. Additional surface water observations were taken at over 380 point locations. All CHESS campaign datasets can be found within the CHESS ESS-DIVE data portal: https://data.ess-dive.lbl.gov/portals/chess. Funding Acknowledgement: Field and remote-sensing data acquisition was performed under a grant from the National Aeronautics and Space Administration (80NSSC24K1005). This work was also supported by the Watershed Function Science Focus Area at Lawrence Berkeley National Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-05CH11231.

2018 NEON and 2025 CHESS Campaigns↗

A High-Throughput Computing Infrastructure to Generate Custom, Open Community Geothermal Datasets

The most significant challenge facing geothermal research, development, and deployment is a lack of comprehensive datasets describing the geological and economical properties of North America. Automated knowledge base construction, the process of designing algorithms to analyze text and images to programmatically build new datasets, is one possible solution to this problem. The xDD library of full-text scientific articles (https://xdd.wisc.edu) is one of the largest collections of open and controlled-access scientific documents available for knowledge base construction in the world, but it has been underutilized by experts in geothermal research. The xDD development team attributed the lack of engagement by software developers and geothermal researchers to two perceived shortcomings of the system. First, the workflow for obtaining data from xDD for local development and testing of data mining applications was unnecessarily abstruse and required significant manual intervention by xDD systems administrators. Second, although xDD already held articles from a broad cross-section of scientific literature with an emphasis on the geosciences, it did not have an explicit set of geothermal research documents that could serve as the nucleus of a geothermal data mining application. To address these issues, the Automated Data Extraction PlaTform (ADEPT) was proposed to extend the data distribution capabilities of the xDD system. The ADEPT extension added the following four key features to xDD: 1) integration of National Geothermal Data System (NGDS) documents into the xDD library to provide an explicitly geothermally-themed collection; 2) improved RESTful (i.e., https-protocol driven) web services for external partners to access xDD data for machine learning application development; 3) a web platform for end-users and xDD administrators to coordinate the development of data mining applications from the initial step of browsing available documents to the final stage of deploying a production-quality machine learning application on high-throughput computing infrastructure; and 4) the development of demonstration data mining applications to illustrate the new workflow to potential collaborators. A total of 21,674 geothermal documents from NGDS were fully ingested into the xDD library and the associated metadata is publicly available through the xDD web services; furthermore, the ADEPT web platform is now publicly accessible and fully live at https://xdd.wisc.edu/adept/.

15 GEOTHERMAL ENERGY↗

Technical Report on Subsurface Monitoring of the Brady Hot Spring Geothermal Site, Nevada, based upon Full Waveform Inversion

Abilities to accurately characterize the subsurface in a geothermal setting is key to assess and support production. An important element of geothermal reservoir monitoring is also the ability to investigate fluid transport within fracture network. This report focuses on improving subsurface imaging and monitoring in geothermal settings using full waveform inversion based on the adjoint method and time-lapse imaging. To assess our method, we rely on a dense seismic dataset collected in 2016 at the Brady Hot Springs geothermal site in Nevada for the DOE-funded project Poroelastic Tomography by Adjoint Inverse Modeling of Data from Seismology, Geodesy, and Hydrology. This dataset captures subsurface changes across four stages of geothermal power plant operations, which involve varying rates of fluid injection and extraction. Two velocity models were previously derived from this dataset using different methods: one based on travel times and another on sweep interferometry. Our first step is to refine these models using adjoint tomography, which has been applied successfully at global and regional-scales but is less common at the reservoir-scale. Two approaches are then explored for time-lapse analysis: directly comparing refined tomographic models from different stages or backpropagating waveform differences relative to a baseline tomographic model. The main take away is that both approaches highlight similar reservoir behaviors, but the latter approach is more computationally effective in capturing small-scale changes in subsurface properties. For this work, we leverage the use of Salvus (www.mondaic.com), an end-to-end seismic imaging solution, relying on the spectral element method to compute forward and adjoint simulations, and developed by Mondaic Ltd. It includes integrated workflow management that handles waveform and metadata, launches simulations, computes waveform misfits and adjoint sources, and iterates for model updates by nonlinear optimization.

15 GEOTHERMAL ENERGY↗

Framework for Processing Citizens Science Data for Applications to NASA Earth Science Missions

Citizen science (or crowdsourcing) has drawn much high-level recent and ongoing interest and support. It is poised to be applied, beyond the by-now fairly familiar use of, e.g., Twitter for natural hazards monitoring, to science research, such as augmenting the validation of NASA earth science mission data. This interest and support is seen in the 2014 National Plan for Civil Earth Observations, the 2015 White House forum on citizen science and crowdsourcing, the ongoing Senate Bill 2013 (Crowdsourcing and Citizen Science Act of 2015), the recent (August 2016) Open Geospatial Consortium (OGC) call for public participation in its newly-established Citizen Science Domain Working Group, and NASA's initiation of a new Citizen Science for Earth Systems Program (along with its first citizen science-focused solicitation for proposals). Over the past several years, we have been exploring the feasibility of extracting from the Twitter data stream useful information for application to NASA precipitation research, with both "passive" and "active" participation by the twitterers. The Twitter database, which recently passed its tenth anniversary, is potentially a rich source of real-time and historical global information for science applications. The time-varying set of "precipitation" tweets can be thought of as an organic network of rain gauges, potentially providing a widespread view of precipitation occurrence. The validation of satellite precipitation estimates is challenging, because many regions lack data or access to data, especially outside of the U.S. and in remote and developing areas. Mining the Twitter stream could augment these validation programs and, potentially, help tune existing algorithms. Our ongoing work, though exploratory, has resulted in key components for processing and managing tweets, including the capabilities to filter the Twitter stream in real time, to extract location information, to filter for exact phrases, and to plot tweet distributions. The key step is to process the "precipitation" tweets to be compatible with satellite-retrieved precipitation data. These key components for processing and managing "precipitation" tweets (and additional ones to be developed) are not limited to precipitation, nor are they limited to the Twitter social medium. Indeed, to maximize the value of our work for NASA earth science programs, these components should be generalized and be part of an overall framework for processing citizen science data for science research. In this paper, we outline such a framework.

earth science satellite data↗

Unlocking the Mysteries of the Moon’s Shadowed Regions

The Moon poles host large quantities of water-ice deposits in the permanently shadowed regions (PSRs), which are vital for enabling sustainable human space exploration, making these regions high-priority targets for upcoming Artemis missions [1]. Unfortunately, today, the best available orbital lunar imagery [2, 3] lacks the meter-scale resolution and signal needed to understand the geomorphology and trafficability of PSRs, complicating the planning and execution of future missions seeking to explore PSRs. We have developed an image enhancement tool called HORUS (Hyper-effective nOise Removal Unet Software) [4, 5], designed to enhance LRO Narrow-Angle Camera (NAC) optical low-light imagery of permanently shadowed regions by effectively removing the CCD-related, photon, and other residual noises that corrupt the images. The tool is composed of two deep learning neural networks trained on environmental metadata and real and synthetic imagery, the latter generated by a physical noise model (LROC). We demonstrated that HORUS effectively produces low-noise, high-resolution images (~1.5m/px), achieving a 5 to 10x improvement over existing long-exposure images of PSRs. HORUS allows scientists and engineers to identify geomorphic features (e.g., craters and boulders) in shadowed regions as small as 3 meters across as well as to peek inside of small shadowed regions, for the first time. The tool was deployed and thoroughly validated for NASA's VIPER mission [6], where it was applied to 20 candidate target regions across the lunar South Pole. Additionally, we conducted different approaches to validate the resulting HORUS-processed images. With HORUS denoised images, VIPER scientists can increase their confidence on what surface features (previously unseen) exist in the shadowed regions, helping them plan rover traverses more safely and efficiently (e.g., Fig. 1) In this manuscript, we will describe how VIPER scientists are utilizing HORUS denoised images to extract new information from the terrain and increase their confidence in what surface features exist in the shadowed regions. In combination with other high-resolution images and digital elevation maps, HORUS images are helping the team analyze potential lading and science sites, as well as planning traverses more safely and efficiently (e.g., Fig. 1). Additionally, we will describe how HORUS tool unlocks a broad range of scientific and exploration applications to other Artemis and CPLS missions to the lunar poles, including (but not limited to) geomorphic analysis, change detection, surface hazard detection, and terrain relative navigation.

artificial intelligence↗

Chloroform Fumigation Extraction for Microbial Biomass and Dissolved Organic Carbon from SPRUCE, Marcell Experimental Forest, Minnesota, 2021, 2022, and 2024

This data set provides the results for chloroform fumigation extraction (CFE) of peat samples collected from ambient and experimental plots in the Spruce and Peatland Responses Under Environmental Change (SPRUCE) Experiment site in June and August of 2021, June of 2022, and June, August, and October of 2024. The SPRUCE Experiment site is in the Marcell Experimental Forest in northern Minnesota, USA. The data set includes values for microbial biomass carbon (MBC), microbial biomass nitrogen (MBN), dissolved organic carbon (DOC), dissolved nitrogen (DN), moisture content (MC, available for 2021 and 2022 only) and gravimetric water content (GWC) at 11 depth increments of two-meter peat cores taken from 12 sampling sites at SPRUCE (10 temperature treatment enclosures, 2 ambient temperature treatment enclosures). The sample analysis followed standard methods. The samples were analyzed using a Shimadzu Total Organic Carbon/Nitrogen (TOC/N) analyzer (TOC-V and TOC-L; 2021-2022) or an Elementar vario TOC Cube (2024), liquid catalytic oxidation combustion analyzers for total carbon and nitrogen analysis. This dataset contains two data files in comma separate (.csv) format. Additional metadata are provided: two data dictionaries and a file-level metadata file in comma separate (.csv) format and a user guide in PDF (*.pdf) format.

dissolved nitrogen↗

Dataset for 'Ombadi, M. & Varadharajan, C. (2022). Urbanization and aridity mediate distinct salinity response to floods in rivers and streams across the Contiguous United States, Water Research'

This package contains data sets and code used to obtain the results in Ombadi, M., & Varadharajan, C. (2022). Urbanization and aridity mediate distinct salinity response to floods in rivers and streams across the Contiguous United States. Water Research, 118664. The folder "data" contains 259 .csv files, each of which has daily time series of concurrent streamflow (Q) and specific conductance (SC) for each of the sites used in this study originally downloaded from the USGS National Water Information System (NWIS; USGS, 2016). The number of data points in each of the files is at least 3650 (i.e. 10 years of daily measurements). The folder "RF_single_data" contains 259 .csv files, each of which include data used to train and test the Random Forest models at individual sites for predicting SC during days of floods. The folder "RF_regional_data" contains 3 .csv files, each of which include scaled data compiled from all sites within each climate zone (arid, temperate and wet). "metadata.csv" contains the physical properties of the 259 catchments corresponding to the sites used in this study; this data was extracted from GAGES-II dataset (Falcone et al., 2010). "RF_implementation.ipynb" is a Jupyter notebook with the code needed to implement the analysis using Random Forest models either for individual sites or for the regional models (for each climate zone). The code utilizes the data in the two folders: "RF_single_data" and "RF_regional_data" and the metadata.csv file.

54 ENVIRONMENTAL SCIENCES↗

Evalution of a DE-Identification Process for Ocular Imaging

Medical privacy of NASA astronauts requires an organized and comprehensive approach when data are being made available outside NASA systems. A combination of factors, including the uniquely small patient population, the extensive medical testing done on these individuals, and the relative cultural popularity of the astronauts puts them at a far greater risk to potential exposure of personal information than the general public. Therefore, care must be taken to ensure that the astronauts' identities are concealed. Magnetic Resonance Imaging (MRI) medical data is a recent source of interest to researchers concerned with the development of Visual Impairment due to Intracranial Pressure (VIIP) in the astronaut population. Each vision MRI scan of an astronaut includes 176 separate sagittal images that are saved as an "image series" for clinical use. In addition to the medical information these image sets provide, they also inherently contain a substantial amount of non-medical personally identifiable information (PII) such as-name, date of birth, and date of exam. We have shown that an image set of this type can be rendered, using free software, to give an accurate representation of the patient's face. This currently restricts NASA from dispensing MRI data to researchers in a deidentified format. Automated software programs, such as the Brain Extraction Tool, are available to researchers who wish to de-identify MRI sagittal brain images by "erasing" identifying characteristics such as the nose and jaw on the image sets. However, this software is not useful to NASA for vision research because it removes the portion of the images around the eye orbits, which is the main area of interest to researchers studying the VIIP syndrome. The Lifetime Surveillance of Astronaut Health program has resolved this issue by developing a protocol to de-identify MRI sagittal brain images using Showcase Premier, a DICOM (Digital Imaging and Communications in Medicine) software package. The software allows manual editing of one image from a patient's image set to be automatically applied to the entire image series. This new approach would allow a new level of access to untapped medical imaging data relating to VIIP that can be utilized by researchers while protecting the privacy of the astronauts. In the next step toward finalizing this technique, NASA clinical radiology consultants will test the images to verify removal of all metadata and PII.

LaPelusa, Michael B.↗

Data for "Depth of nutrient uptake by deep-rooted plants is regulated by water availability"

The data set consists of strontium (Sr) isotope ratios (87Sr/86Sr), water isotopes, soil cation concentrations, soil water potential sensor data, and results of 87Sr/86Sr mixing model. The plant canopy size files include the dataset of canopy dimension of sagebrush, lupine, and sunflower. The soil and plant ICPMS (Inductively Coupled Plasma Mass Spectrometry) data file includes both of 87Sr/86Sr, and cation concentration dataset from soil exchangeable pool, apatite pool, silicate extract, atmospheric rain deposition, and plant leaf and stem tissues. The plant dendrochronology file includes the dendrochronogical ring width of several sagebrush, and dendrochemical sample data includes the 87Sr/86Sr for each separated growth ring. The modeling result gives the proportion of nutrient sources of each plants (based on their 87Sr/86Sr in leaf tissues and growth rings) from atmospheric deposition and mineral weathering. Soil water potential data includes continuous collection of soil water potential dataset at 2 depths (30 cm and 60 cm, from Nov 24 - Jun 25) of the sampling site. All the samples were collected from 2 sampling campaign June and July 2023, and rain water is a separate sampling from Aug - Sept 2023, at north-facing hillslope near pumphouse site. The data showed that the depth of cation nutrient acquisition is thus tightly coupled with, and likely determined by, water availability in soil, saprolite and bedrock. The enhanced uptake of cations and water from regions of mineral weathering could confer plant and ecosystem resilience during low water years and may impact the rate of bedrock weathering and watershed chemistry during drought. This dataset includes a file-level metadata (flmd.csv) file that lists each file contained in the dataset with associated metadata; a data dictionary (dd.csv) file that contains column/row headers used throughout the files along with a definition, units, and data type; a location metadata file (locations.csv); and a samples metadata file (samples.csv). All files are provided as comma-separated values (CSV) files (.csv). This work was supported by the Watershed Function Science Focus Area at Lawrence Berkeley National Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-05CH11231.

54 ENVIRONMENTAL SCIENCES↗

Power System Waveform Datasets for Machine Learning

The desire for increased visibility across the electricity grid will necessarily increase the deployment of sensing and measurement devices and associated data management needs to unprecedented levels. For the existing sensing and measurement infrastructure, there remains a great amount of “value” yet to be extracted through advanced data management and analytics. Availability of more data will not, by itself, lead to changes in grid visibility, security, and resiliency. To create the predictive and prescriptive environment required to enable new markets and transactions for customer revenue and a reliable grid, the data must be collected, organized, evaluated, and analyzed using sophisticated algorithms to provide actionable information allowing operators and customers to reliably manage an increasingly complex grid. Progress in artificial intelligence (AI) has been largely driven by large, publicly available datasets that can be used to train AI algorithms such as MNIST, a database of handwritten images of digits, and ImageNet, an image database of everyday objects. These types of publicly available databases of real-world training datasets have been largely credited for advancement of image processing, computer vision, and deep learning algorithms that these use cases deploy. However, in the power systems industry to date, there are few databases with proper event labeling, and data access to a publicly available collection of power system event waveforms that will allow users to interact with grid signature data. Publicly available datasets of power system event waveforms, such as the DOE/EPRI dataset, often lack critical metadata or contain limited examples of each event type, and data formats vary widely across these datasets.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

BEAST: Expanding Sustainable Data Infrastructure for High-Enthalpy Facilities

Reproducible, data-driven thermal protection system (TPS) research requires that experimental records from high-enthalpy testing be consistently structured, traceable, and accessible across campaigns and institutions. In practice, however, arcjet and plasma facilities data remain largely fragmented: raw diagnostics are stored in ad hoc formats, material sample histories are disconnected from test conditions, and metadata standards are absent, precluding systematic cross-campaign analysis and long-term reuse. BEAST (Backend for Experiment Analysis, Storage, and Traceability) is an open-source, web-based platform that addresses these limitations by providing a unified, queryable infrastructure for high-enthalpy ground-test data [1]. First presented at the 15th Ablation Workshop [2], BEAST has since undergone significant development. The platform ingests and structures multi-channel time-series diagnostics, facility configurations, and material property records within a common provenance model, ensuring end-to-end traceability from raw sensor acquisition to reduced experimental quantities. A versioned material library links specimen identity and processing history to the specific runs in which each sample was tested. An integrated modeling workbench enables training and evaluation of regression models directly on archived experimental data, supporting condition interpolation and the construction of empirical material response databases. Beyond its original deployment at NASA Ames Research Center, BEAST has been designed to be facility-agnostic, with ongoing efforts to extend its adoption to other facilities. Its modular architecture accommodates heterogeneous diagnostic setups and facility types, and its future open-source distribution allows institutions to build on a common data standard rather than maintaining isolated, bespoke solutions. BEAST is further integrated within a broader ecosystem of companion tools: arcjetCV [3] extracts recession rates and shock standoff distances from high-speed video using computer vision, and miniSTARscan [4] provides sub-minute, portable photogrammetric surface reconstruction of test articles before and after exposure. All tools share a common data schema, enabling seamless ingestion of surface geometry, imagery, and time-series data into a single, coherent experimental record.

Database↗

BEAST: Expanding Sustainable Data Infrastructure for High-Enthalpy Facilities

Reproducible, data-driven thermal protection system (TPS) research requires that experimental records from high-enthalpy testing be consistently structured, traceable, and accessible across campaigns and institutions. In practice, however, arcjet and plasma facilities data remain largely fragmented: raw diagnostics are stored in ad hoc formats, material sample histories are disconnected from test conditions, and metadata standards are absent, precluding systematic cross-campaign analysis and long-term reuse. BEAST (Backend for Experiment Analysis, Storage, and Traceability) is an open-source, web-based platform that addresses these limitations by providing a unified, queryable infrastructure for high-enthalpy ground-test data [1]. First presented at the 15th Ablation Workshop [2], BEAST has since undergone significant development. The platform ingests and structures multi-channel time-series diagnostics, facility configurations, and material property records within a common provenance model, ensuring end-to-end traceability from raw sensor acquisition to reduced experimental quantities. A versioned material library links specimen identity and processing history to the specific runs in which each sample was tested. An integrated modeling workbench enables training and evaluation of regression models directly on archived experimental data, supporting condition interpolation and the construction of empirical material response databases. Beyond its original deployment at NASA Ames Research Center, BEAST has been designed to be facility-agnostic, with ongoing efforts to extend its adoption to other facilities. Its modular architecture accommodates heterogeneous diagnostic setups and facility types, and its future open-source distribution allows institutions to build on a common data standard rather than maintaining isolated, bespoke solutions. BEAST is further integrated within a broader ecosystem of companion tools: arcjetCV [3] extracts recession rates and shock standoff distances from high-speed video using computer vision, and miniSTARscan [4] provides sub-minute, portable photogrammetric surface reconstruction of test articles before and after exposure. All tools share a common data schema, enabling seamless ingestion of surface geometry, imagery, and time-series data into a single, coherent experimental record.

Database↗