Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “metadata extraction”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

History and Status of ALSEP and the Apollo Lunar Data Project

A suite of automated scientific instruments (the Apollo Lunar Surface Experiment Package, or ALSEP) was installed at each of the landing sites of Apollo 12, 14, 15, 16, and 17 from 1969 to 1972. They operated from deployment until decommissioning on 30 September 1977. These data were continuously transmitted to Earth and saved on the Range Tapes, which were recorded at the Manned Space Flight Network stations. These data were also broken out by experiment and sent to the experiment Principal Investigators on what were called the P.I. Tapes. Starting in April 1973 the Range Tape data were stored in digital format on 7-track magnetic tapes, the ARCSAV Tapes. In February 1976, the handling of the Range Tapes was transferred to UT Galveston. They produced 9-track tapes referred to as the Work Tapes. Following the Apollo program the Range and ARCSAV tapes, which were never archived, were lost. The Work Tapes were archived at the National Space Science Data Center (NSSDC). Some investigators archived their individual experiment data with NSSDC as well, but much of the data had minimal documentation, were not in digital form, or were stored in difficult to translate formats. Data from many experiments were never delivered to the NSSDC. The Lunar Data Project was started to address the problem of both missing and not readily usable data. Our effort has resulted in recovery of some of the ARCSAV tapes, recovery and digitization of a large volume of Apollo scientific and technical documentation, and restoration of many ALSEP and other Apollo data collections. Restoration involves deciphering formats, assembling necessary ancillary data (metadata), and packaging data in digital format to be archived with the Planetary Data System (PDS). Recovery of the data from the ARCSAV tapes involved having the tapes read on special equipment and extracting the individual experiment data out of the integrated data stream. We will report on the history and status of the various recovery efforts.

Work Tapes↗

Availability of previously lost data and metadata from the Apollo Lunar Surface Experiments Package (ALSEP)

Fourteen types of geophysical instruments deployed at the Apollo 12, 14, 15, 16, and 17 sites by the astronauts for long-term observation were collectively called the Apollo Lunar Surface Experiments Package (ALSEP). These instruments were active from the times of their deployment (November 1969–December 1972) to September 1977. At the conclusion of the experiments, the raw instrument data received from the Moon prior to March 1976 were left unarchived. Portions of the data processed by the principal investigators (PIs) of these experiments had been archived at the NASA Space Science Data Coordinated Archive (NSSDCA) in various formats. The unarchived data, residing then on open-reel magnetic tapes, became lost in the decades since, along with much of the metadata (the supporting documents for these data). We have recently recovered 440 of the previously lost tapes, containing raw ALSEP instrument data from April through June of 1975. Here we describe the data extracted from these tapes and summarize the data products generated for archiving at the NASA Planetary Data System (PDS) and NSSDCA, along with their historical narrative. In addition, we have reformatted many of the datasets delivered to NSSDCA by the PIs in the 1970s for archiving at the PDS. Finally, we have compiled an online searchable repository of ALSEP-related documents by optically scanning tens of thousands of pages of them kept at the Lunar and Planetary Institute in Texas.

S. Nagihara↗

A Data Processing Pipeline for Adversarial Socio-Technical Network Analysis

With the rapid adoption of emerging technologies, there is a need to catalog and model sociotechnical interdependencies that have been historically used to influence the operation of Critical Infrastructure networks including the impacts of mergers and acquisitions, hostile takeovers, and foreign investment. Our research intends to address this need with two primary contributions. First, we have developed a data curation and processing pipeline to generate sociotechnical networks extracted from a variety of data sources including SEC filings and infrastructure asset databases. The pipeline, implemented in Apache Airflow, extracts and normalizes the representation of entities and relations, specified within ontologies. Our intent is to provide an extensible, machine-actionable approach to quickly communicate such models, reproduce previous results, and adapt them to new, unanticipated situations. Second, networks produced by our pipeline enable the development of graph-theoretic metrics that consider the properties of network components in addition to its topology. Metadata associated with network components---whether semantic, temporal, or geospatial---affects the alignment of generated networks with assumptions underlying complexity metrics. Validation of generated networks relative to component types defined by an ontology, may allow the research community to adapt metrics to the semantics of the domains being studied. Generated networks may be processed as knowledge, dynamic, or spatial graphs and enables a variety of analyses including automated reasoning and measures of network complexity. Automated reasoning views extracted entities and relations as a knowledge graph; this enables application of inference rules that represent historically-attested adversarial business methods and applies that behavior to a specific geographic context. Measures of network complexity, including degree distribution, reachability analyses, temporal analysis, and community detection can be adapted to indicate adversarial organizational influence.

97 MATHEMATICS AND COMPUTING↗

Human Host Cellular Response to HCoV-229E Infection Proteomics (ACS-JM-DP2)

The purpose of this experiment was to evaluate the human host cellular response to wild-type Human coronavirus strain 229E (HCoV-229E) infection. Sample data was obtained for mock and infected immortalized human lung epithelial cells (A549) (MOI 5) nuclear extracts, immortalized human lung fibroblasts cells (MRC5) (MOI5) nuclear extracts, and primary human airway epithelial (HAE) (MOI 3) cells from lung tissue and processed for proteome analysis. Processed datasets are openly accessible from the download button and contain secondary processed proteomic results files and supporting metadata materials. Experimental proteomics samples were prepared using Limited Proteolysis (LiP) methods for Label-free quantification (LFQ) and global proteomic evaluation. Sample data was acquired using a Q-Exactive HF-X mass spectrometer and was processed and compiled using MaxQuant software (v.1.6.17.0). Processed proteomic data downloads include a sample naming key, processed MaxQuant results/parameters, and protein annotated relative abundance files. See corresponding primary data accessions below and Viral Experiment LiP Analysis source code supporting data transparency and reuse. Experimental transcriptomics samples were collected in parallel and processed for RNA sequencing (RNA-Seq) as summarized under ACS-DP1 (https://data.pnnl.gov/group/nodes/dataset/34069).

59 BASIC BIOLOGICAL SCIENCES↗

EXCHANGE Campaign 1: A Community-Driven Baseline Characterization of Soils, Sediments, and Water Across Coastal Gradients

The EXploration of Coastal Hydrobiogeochemistry Across a Network of Gradients and Experiments (EXCHANGE) program is a consortium of scientists working together to improve our understanding of how the two-way exchange of water between estuaries or large lake lacustuaries and the terrestrial landscape influence the state and function of ecosystems across the coastal interface. EXCHANGE Campaign 1 (EC1) focuses on the spatial variation in biogeochemical structure and function at the coastal terrestrial-aquatic interface (TAI). In the Fall of 2021, the EXCHANGE Consortium gathered samples from 52 TAIs. Samples collected from EC1 were analyzed for bulk geochemical parameters, bulk physicochemical parameters, organic matter characteristics, and redox-sensitive elements.Please download ec1_README.pdf for a complete list of available data in each .zip folder, package version history, and detailed information about the project. This README will serve as the central place for EC1 Data Package updates. Experimental setup and v1 methods are documented in Myers-Pigg and Pennington et al., 2023 (https://doi.org/10.1038/s41597-023-02548-7).EC1 Data Package Structure:ec1_README.pdfec1_methods.pdfec1_metadata_v3.zip...ec1_dd.csv...ec1_flmd.csv...ec1_sample_catalog.csv...ec1_metadata_kitlevel.csv...ec1_metadata_collectionlevel.csv...ec1_data_collectionlevel.csv...ec1_igsn_metadata.csvec1_soil_v3.zipec1_sediment_v3.zipec1_water_v3.zipec1_processingscripts_v3.zipThis data package is on v3 and was originally published May 2023 (v1). Subsequent updates will be published here with new version numbers. Please see the Change History section in ec1_README.pdf for detailed changes.---Acknowledging EXCHANGE: General Support and Data Product UseWe ask that users of EXCHANGE data add the following acknowledgement when publishing data in scholarly articles and data repositories:"This research is based on work supported by COMPASS-FME, a multi-institutional project supported by the U.S. Department of Energy, Office of Science, Biological and Environmental Research as part of the Environmental System Science Program."

54 ENVIRONMENTAL SCIENCES↗

Machine Learning Enabled Quantitative Risk Assessment of Aerial Wildfire Response

Aerial wildfire operations are high risk and account for a large number of firefighter deaths. Increasing intensity of wildfires is driving a surge in aerial operations, while simultaneously there is growing interest in improving system safety and performance. In this work, wildfire aviation mishaps documented using the SAFECOM system are analyzed using a previously developed framework for hazard extraction and analysis of trends (HEAT). Hazards and specific failure modes are extracted from the narrative data in SAFECOM forms using natural language processing techniques. Metrics for each hazard are calculated, including frequency, rate, and severity. We examine whether these metrics change over time, and whether they are related to metadata, such as region and aircraft type. The results of the hazard analysis are presented in a risk matrix, identifying the highest and lowest risk hazards based on rate of occurrence and average severity. Results identify jumper operations hazards as high-risk, in addition to bucket drop failures, cargo let down failures, and severe weather as medium risk.

machine learning↗

Machine Learning Enabled Quantitative Risk Assessment of Aerial Wildfire Response

Aerial wildfire operations are high risk and account for a large number of firefighter deaths. Increasing intensity of wildfires is driving a surge in aerial operations, while simultaneously there is growing interest in improving system safety and performance. In this work, wildfire aviation mishaps documented using the SAFECOM system are analyzed using a previously developed framework for hazard extraction and analysis of trends (HEAT). Hazards and specific failure modes are extracted from the narrative data in SAFECOM forms using natural language processing techniques. Metrics for each hazard are calculated, including frequency, rate, and severity. We examine whether these metrics change over time, and whether they are related to metadata, such as region and aircraft type. The results of the hazard analysis are presented in a risk matrix, identifying the highest and lowest risk hazards based on rate of occurrence and average severity. Results identify jumper operations hazards as high-risk, in addition to bucket drop failures, cargo let down failures, and severe weather as medium risk.

machine learning↗

Textural-Contextual Labeling and Metadata Generation for Remote Sensing Applications

Despite the extensive research and the advent of several new information technologies in the last three decades, machine labeling of ground categories using remotely sensed data has not become a routine process. Considerable amount of human intervention is needed to achieve a level of acceptable labeling accuracy. A number of fundamental reasons may explain why machine labeling has not become automatic. In addition, there may be shortcomings in the methodology for labeling ground categories. The spatial information of a pixel, whether textural or contextual, relates a pixel to its surroundings. This information should be utilized to improve the performance of machine labeling of ground categories. Landsat-4 Thematic Mapper (TM) data taken in July 1982 over an area in the vicinity of Washington, D.C. are used in this study. On-line texture extraction by neural networks may not be the most efficient way to incorporate textural information into the labeling process. Texture features are pre-computed from cooccurrence matrices and then combined with a pixel's spectral and contextual information as the input to a neural network. The improvement in labeling accuracy with spatial information included is significant. The prospect of automatic generation of metadata consisting of ground categories, textural and contextual information is discussed.

Kiang, Richard K.↗

In-situ electrochemical and water quality data; Slate River and East River floodplains, Crested Butte, CO; May 2022-September 2022

This data package includes a time-series of field measurements from May to September 2022 in groundwater and surface water from the Slate River and East River floodplains in Crested Butte, CO, a focus field site for the SLAC Floodplain Hydro-Biogeochemistry SFA. The data was generated as part of the work targeting the overarching research question for the SLAC SFA: How do ubiquitous subsurface interfaces mediate molecular-scale biogeochemical processes and groundwater quality in floodplains and watersheds? The data package includes 5 data files, one for each measured variable: specific conductivity, pH, dissolved oxygen, water temperature and alkalinity. All measurements were recorded in the field immediately after water sampling. Groundwater samples were extracted from a network of installed rhizons (Rhizosphere Research Products, part no. 19.60.21F, 0.6 micrometer mesh size) and piezometer wells within the floodplain. In addition to the data files, there is a terminology file explaining the terms used, a file level metadata file, and a sensor file with metadata about the sensors used.All files are in csv format.

54 ENVIRONMENTAL SCIENCES↗

Characterizing DebriSat Fragments: So Many Fragments, So Much Data, and So Little Time

To improve prediction accuracy, the DebriSat project was conceived by NASA and DoD to update existing standard break-up models. Updating standard break-up models require detailed fragment characteristics such as physical size, material properties, bulk density, and ballistic coefficient. For the DebriSat project, a representative modern LEO spacecraft was developed and subjected to a laboratory hypervelocity impact test and all generated fragments with at least one dimension greater than 2 mm are collected, characterized and archived. Since the beginning of the characterization phase of the DebriSat project, over 130,000 fragments have been collected and approximately 250,000 fragments are expected to be collected in total, a three-fold increase over the 85,000 fragments predicted by the current break-up model. The challenge throughout the project has been to ensure the integrity and accuracy of the characteristics of each fragment. To this end, the post hypervelocity-impact test activities, which include fragment collection, extraction, and characterization, have been designed to minimize handling of the fragments. The procedures for fragment collection, extraction, and characterization were painstakingly designed and implemented to maintain the post-impact state of the fragments, thus ensuring the integrity and accuracy of the characterization data. Each process is designed to expedite the accumulation of data, however, the need for speed is restrained by the need to protect the fragments. Methods to expedite the process such as parallel processing have been explored and implemented while continuing to maintain the highest integrity and value of the data. To minimize fragment handling, automated systems have been developed and implemented. Errors due to human inputs are also minimized by the use of these automated systems. This paper discusses the processes and challenges involved in the collection, extraction, and characterization of the fragments as well as the time required to complete the processes. The objective is to provide the orbital debris community an understanding of the scale of the effort required to generate and archive high quality data and metadata for each debris fragment 2 mm or larger generated by the DebriSat project.

Shiotani, B.↗

Long-Lasting Science Returns from the Apollo Heat Flow Experiments

The Apollo astronauts deployed geothermal heat flow instruments at landing sites 15 and 17 as part of the Apollo Lunar Surface Experiments Packages (ALSEP) in July 1971 and December 1972, respectively. These instruments continuously transmitted data to the Earth until September 1977. Four decades later, the data from the two Apollo sites remain the only set of in-situ heat flow measurements obtained on an extra-terrestrial body. Researchers continue to extract additional knowledge from this dataset by utilizing new analytical techniques and by synthesizing it with data from more recent lunar orbital missions such as the Lunar Reconnaissance Orbiter. In addition, lessons learned from the Apollo experiments help contemporary researchers in designing heat flow instruments for future missions to the Moon and other planetary bodies. For example, the data from both Apollo sites showed gradual warming trends in the subsurface from 1971 to 1977. The cause of this warming has been debated in recent years. It may have resulted from fluctuation in insolation associated with the 18.6-year-cycle precession of the Moon, or sudden changes in surface thermal environment/properties resulting from the installation of the instruments and the astronauts' activities. These types of reanalyses of the Apollo data have lead a panel of scientists to recommend that a heat flow probe carried on a future lunar mission reach 3 m into the subsurface, approx 0.6 m deeper than the depths reached by the Apollo 17 experiment. This presentation describes the authors current efforts for (1) restoring a part of the Apollo heat flow data that were left unprocessed by the original investigators and (2) designing a compact heat flow instrument for future robotic missions to the Moon. First, at the conclusion of the ALSEP program in 1977, heat flow data obtained at the two Apollo sites after December 1974 were left unprocessed and not properly archived through NASA. In the following decades, heat flow data from January 1975 through February 1976, as well as the metadata necessary for processing the data (the data reduction algorithm, instrument calibration data, etc.), were somehow lost. In 2010, we located 450 original master archival tapes of unprocessed data from all the ALSEP instruments for a period of April through June 1975 at the Washington National Records Center. We are currently extracting the heat flow data packets from these tapes and processing them. Second, on future lunar missions, heat flow probes will likely be deployed by a network of small robotic landers, as recommended by the latest Decadal Survey of the National Academy of Science. In such a scenario, the heat flow probe must be a compact system, and that precludes use of heavy excavation equipment such as a rotary drill for reaching the 3-m target depth. The new heat flow system under development uses a pneumatically driven penetrator. It utilizes a stem that winds out of a reel and pushes its conical tip into the regolith. Simultaneously, gas jets, emitted from the cone tip, loosen and blow away the soil. Lab experiments have demonstrated its effectiveness in lunar vacuum.

Nagihara, S.↗

Lessons Learned from AskGDR: Usage and Impact Analysis of the Geothermal Data Repository's AI Research Assistant: Preprint

In October of 2024, the Department of Energy's (DOE) Geothermal Data Repository (GDR) team officially launched AskGDR, an AI research assistant resulting from the integration of a Large Language Model (LLM) with the metadata and supporting documents associated with GDR datasets. AskGDR allows GDR users to ask deeper questions about the origin of datasets, the methods used to collect them, and the findings they help support. Using Retrieval Augmented Generation (RAG), AskGDR can be used to summarize findings spread across dozens of papers and technical reports or to extract relevant information describing a single data field. However, generative AI is experimental. The National Renewable Energy Laboratory (NREL) has been collecting metrics on AskGDR and documenting lessons learned during its deployment. This paper will outline the efficacy and impact of AskGDR through analysis of its use, operating costs, number and types of questions asked, and the quality of answers provided.

15 GEOTHERMAL ENERGY↗

NEPATEC2.0: NEPA Text Corpus v2.0

The National Environmental Policy Act of 1969, as amended (NEPA), is a major environmental law in the United States, requiring Federal agencies to consider and document potential environmental impacts before deciding on a proposed action. Modernization of NEPA and permitting processes faces significant challenges due to the lack of standardized formats and interoperable systems for organizing and sharing NEPA-related information across agencies. Much of the information gathered during NEPA reviews is written into documents such as categorical exclusions, environmental assessments, and environmental impact statements, then filed in predominately independent agency file stores that may or may not be publicly accessible. The application of metadata and data standards, such as those recommended by the Council on Environmental Quality (CEQ), to NEPA documents offers a shared vocabulary and structure for key entities like projects, processes, and documents that can streamline information exchange and enhance collaboration across systems. In this work, we publicly release NEPATEC2.0, an expanded corpus of NEPA documents with associated metadata. NEPATEC2.0 encompasses approximately 120,000 documents from 60,000 projects prepared by more than 60 different agencies. Modeled to align with CEQ metadata standards, NEPATEC2.0 promotes consistency in environmental reviews and supports the ongoing effort to modernize permitting technologies by facilitating more transparent, efficient, and data-driven decision-making. Importantly, NEPATEC2.0 demonstrates the possibilities and limitations of large language model-based prompting to extract information from NEPA documents at scale.

environmental review↗

NEPATEC v2.0: Standardized Metadata and Text Corpus of National Environmental Policy Act Documents

The National Environmental Policy Act of 1969, as amended (NEPA), is a major environmental law in the United States, requiring Federal agencies to consider and document potential environmental impacts before deciding on a proposed action. Modernization of NEPA and permitting processes faces significant challenges due to the lack of standardized formats and interoperable systems for organizing and sharing NEPA-related information across agencies. Much of the information gathered during NEPA reviews is written into documents such as categorical exclusions, environmental assessments, and environmental impact statements, then filed in predominately independent agency file stores that may or may not be publicly accessible. The application of metadata and data standards, such as those recommended by the Council on Environmental Quality (CEQ), to NEPA documents offers a shared vocabulary and structure for key entities like projects, processes, and documents that can streamline information exchange and enhance collaboration across systems. In this work, we publicly release NEPATEC2.0, an expanded corpus of NEPA documents with associated metadata. NEPATEC2.0 encompasses approximately 120,000 documents from 60,000 projects prepared by more than 60 different agencies. Modeled to align with CEQ metadata standards, NEPATEC2.0 promotes consistency in environmental reviews and supports the ongoing effort to modernize permitting technologies by facilitating more transparent, efficient, and data-driven decision-making. Importantly, NEPATEC2.0 demonstrates the possibilities and limitations of large language model-based prompting to extract information from NEPA documents at scale.

54 ENVIRONMENTAL SCIENCES↗

Interrelationships among methods of estimating microbial biomass across multiple soil orders and biomes: Supporting data

This dataset contains environmental and soil measurements from 18 different locations across the globe including the SPRUCE experiment site and multiple sampling depths, with 17 of these locations having samples processed between 2012-2013 and one location (SPRUCE) collected in 2021 and processed in 2022. Environmental measurements include: mean annual temperature, mean annual precipitation, and 30-day presampling temperature. Soil physicochemical measurements include: particle size analysis (PSA), pH, gravimetric moisture content (GMC), bulk soil carbon (C) and nitrogen (N), total organic C and N, C:N ratio, and dissolved organic carbon (DOC). Soil biological measurements include: microbial biomass carbon (MBC) measured through chloroform fumigation extraction (CFE), gene copy numbers (GCN) of bacteria, fungi, and archaea measured through quantitative polymerase chain reaction (qPCR), DNA yield measured through Nanodrop spectrophotometry, and phospholipid fatty acids (PLFA) of bacteria and fungi measured through PLFA analysis. This data set contains one file in comma separate (*.csv) format.

archaea gene copy number↗

An Automated Approach to Labelling Datasets in Earth Science Publications

NASA Data Active Archive Centers, orDAACs, ingest, store, and distribute dataacquired from satellites, ground systems as well asreanalysis models. Many authors use this datain their research. However, most of the datasets usedin Earth Science Publications are not citedcorrectly or not cited at all. Thus, there is no directlink between the datasets used and thescientific publications which reference them. Thisleads to issues with reproducibility of theresults, attribution of the research results, anddiscovery of new datasets. This project began byexploring various methods of automatically labellingGoddard Earth Sciences Data andInformation Services Center (GES DISC) datasets usingSupervised Machine Learning and EarthData Search Common Metadata Repository (CMR) queries.The ultimate goal was to create alibrary of citations that utilized automated citationlabeling to directly link the researchpublications to the data they use. Supervised MachineLearning approaches struggled due to thelimited amount of labelled training data to learnfrom. Increasing the volume of training data isdifficult as it requires subject matter experts todevote time to manually reviewing journalarticles and determining the datasets used. The CMRqueries were inconsistent because theunderlying metadata is continuously being updated.Thus, it is hard to generalize theeffectiveness of the CMR results as they are dependenton the internal state of CMR. Theseapproaches helped inform the decision to transitionthe project into using a Knowledge Graph.Another key aspect of this project focused on theautomated extraction of features (platform,instrument, variables, etc) and explicit citationsfrom within Earth Science Publications. Theseautomated extractions were used to classify researchpapers based on their platform/instrumentcouples. This information was input into the CitationManagement System for GES DISC. Theseplatform/instrument couples also provide an additionalfacet that can be searched on the GESDISC website.

Edward Jahoda↗

Changuinola peat soil characteristics and gas emission raw data October 2019

This dataset comprises radiocarbon and geochemical measurements from peat and porewater samples collected across various depths at a site in Bocas del Toro, Panama. The study focuses on carbon cycling dynamics in tropical peatlands by examining carbon isotopic signatures (¹⁴C and ¹³C) and elemental compositions of bulk peat, dissolved organic carbon (DOC), carbon dioxide (CO₂), and methane (CH₄). Key parameters include radiocarbon ages and isotopic ratios (δ¹³C) of bulk peat, concentrations of carbon (%C) and nitrogen (%N), and radiocarbon content of porewater gases and dissolved organic carbon (DOC). The data provide insights into the vertical and spatial distribution of carbon sources and possible preservation and decomposition processes within tropical peat profiles, offering critical information for understanding carbon storage and greenhouse gas emissions in these ecosystems.This dataset is comprised of one main data folder containing (1) file-level metadata; (2) data dictionary; (3) field metadata; (4) carbon isotopic signatures (¹⁴C and ¹³C); (5) concentrations of carbon (%C) and nitrogen (%N); (6) radiocarbon content of porewater carbon dioxide (CO₂), and methane (CH₄) ; (7) porewater DOC; (8) bulk peat sampling protocol; (9) porewater sampling protocol; (10) porewater gas collection methods; and (11) gas extraction methods. All files are in .csv format and can be opened with any software that supports this file types.

54 ENVIRONMENTAL SCIENCES↗

1H-NMR characterization of soil dissolved organic matter from soil samples in control and warming plots in Blodgett Forest, CA (2014 and 2018)

The pathways of carbon transport and loss through and from soils—soil organic matter (SOM) depolymerization to dissolved organic carbon and mineralization to carbon dioxide (CO2)—are fundamentally driven by microbial activity, which is strongly regulated by environmental conditions. As part of Lawrence Berkeley National Laboratory Terrestrial Ecosystem Science Belowground Biogeochemistry Science Focus Area (SFA), we have established a novel whole-soil long-term warming experiment at the University of California (UC) Blodgett Forest Research Station (Sierra Nevada) in 2014, where we study the role of biogeochemical, microbial and geochemical process interactions in SOM (soil organic matter) decomposition and stabilization. This package contains metabolite data obtained through 1H nuclear magnetic resonance (NMR) spectroscopy on water-extracted soils. Soil samples were collected in 2014/06/03 and 2018/06/04 from 3 replicated paired plots that had been subjected to experimental warming since June 2014 to simulate a predicted climate change scenario for northern California. The following files are included: (1) nmr_h2o_data_raw.csv: raw data, (2) nmr_h2o_data_processed.csv: computed compound concentrations and metadata, (3) nmr_h2o_compound_metadata.csv: compound metadata, (4) nmr_h2o_sample_metadata.csv: sample metadata

1H-NMR (nucleic magnetic resonance) spectroscopy↗