Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “metadata extraction”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

139 records · Page 8

Power System Waveform Datasets for Machine Learning

The desire for increased visibility across the electricity grid will necessarily increase the deployment of sensing and measurement devices and associated data management needs to unprecedented levels. For the existing sensing and measurement infrastructure, there remains a great amount of “value” yet to be extracted through advanced data management and analytics. Availability of more data will not, by itself, lead to changes in grid visibility, security, and resiliency. To create the predictive and prescriptive environment required to enable new markets and transactions for customer revenue and a reliable grid, the data must be collected, organized, evaluated, and analyzed using sophisticated algorithms to provide actionable information allowing operators and customers to reliably manage an increasingly complex grid. Progress in artificial intelligence (AI) has been largely driven by large, publicly available datasets that can be used to train AI algorithms such as MNIST, a database of handwritten images of digits, and ImageNet, an image database of everyday objects. These types of publicly available databases of real-world training datasets have been largely credited for advancement of image processing, computer vision, and deep learning algorithms that these use cases deploy. However, in the power systems industry to date, there are few databases with proper event labeling, and data access to a publicly available collection of power system event waveforms that will allow users to interact with grid signature data. Publicly available datasets of power system event waveforms, such as the DOE/EPRI dataset, often lack critical metadata or contain limited examples of each event type, and data formats vary widely across these datasets.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

BEAST: Expanding Sustainable Data Infrastructure for High-Enthalpy Facilities

Reproducible, data-driven thermal protection system (TPS) research requires that experimental records from high-enthalpy testing be consistently structured, traceable, and accessible across campaigns and institutions. In practice, however, arcjet and plasma facilities data remain largely fragmented: raw diagnostics are stored in ad hoc formats, material sample histories are disconnected from test conditions, and metadata standards are absent, precluding systematic cross-campaign analysis and long-term reuse. BEAST (Backend for Experiment Analysis, Storage, and Traceability) is an open-source, web-based platform that addresses these limitations by providing a unified, queryable infrastructure for high-enthalpy ground-test data [1]. First presented at the 15th Ablation Workshop [2], BEAST has since undergone significant development. The platform ingests and structures multi-channel time-series diagnostics, facility configurations, and material property records within a common provenance model, ensuring end-to-end traceability from raw sensor acquisition to reduced experimental quantities. A versioned material library links specimen identity and processing history to the specific runs in which each sample was tested. An integrated modeling workbench enables training and evaluation of regression models directly on archived experimental data, supporting condition interpolation and the construction of empirical material response databases. Beyond its original deployment at NASA Ames Research Center, BEAST has been designed to be facility-agnostic, with ongoing efforts to extend its adoption to other facilities. Its modular architecture accommodates heterogeneous diagnostic setups and facility types, and its future open-source distribution allows institutions to build on a common data standard rather than maintaining isolated, bespoke solutions. BEAST is further integrated within a broader ecosystem of companion tools: arcjetCV [3] extracts recession rates and shock standoff distances from high-speed video using computer vision, and miniSTARscan [4] provides sub-minute, portable photogrammetric surface reconstruction of test articles before and after exposure. All tools share a common data schema, enabling seamless ingestion of surface geometry, imagery, and time-series data into a single, coherent experimental record.

Database↗

BEAST: Expanding Sustainable Data Infrastructure for High-Enthalpy Facilities

Reproducible, data-driven thermal protection system (TPS) research requires that experimental records from high-enthalpy testing be consistently structured, traceable, and accessible across campaigns and institutions. In practice, however, arcjet and plasma facilities data remain largely fragmented: raw diagnostics are stored in ad hoc formats, material sample histories are disconnected from test conditions, and metadata standards are absent, precluding systematic cross-campaign analysis and long-term reuse. BEAST (Backend for Experiment Analysis, Storage, and Traceability) is an open-source, web-based platform that addresses these limitations by providing a unified, queryable infrastructure for high-enthalpy ground-test data [1]. First presented at the 15th Ablation Workshop [2], BEAST has since undergone significant development. The platform ingests and structures multi-channel time-series diagnostics, facility configurations, and material property records within a common provenance model, ensuring end-to-end traceability from raw sensor acquisition to reduced experimental quantities. A versioned material library links specimen identity and processing history to the specific runs in which each sample was tested. An integrated modeling workbench enables training and evaluation of regression models directly on archived experimental data, supporting condition interpolation and the construction of empirical material response databases. Beyond its original deployment at NASA Ames Research Center, BEAST has been designed to be facility-agnostic, with ongoing efforts to extend its adoption to other facilities. Its modular architecture accommodates heterogeneous diagnostic setups and facility types, and its future open-source distribution allows institutions to build on a common data standard rather than maintaining isolated, bespoke solutions. BEAST is further integrated within a broader ecosystem of companion tools: arcjetCV [3] extracts recession rates and shock standoff distances from high-speed video using computer vision, and miniSTARscan [4] provides sub-minute, portable photogrammetric surface reconstruction of test articles before and after exposure. All tools share a common data schema, enabling seamless ingestion of surface geometry, imagery, and time-series data into a single, coherent experimental record.

Database↗

Distributed Real-time Plume Monitoring for Deep Sea Mineral Extraction​

In the emerging industry of deep-sea mining for minerals and deposits (e.g. polymetallic nodules for nickel, cobalt, copper, and manganese), more data is required to understand the effects of sediment plume generation and predict the distribution of disturbed sediment. There are two main sources of plume generation, the first being at the active mining site where the “collector” directly removes the top layer of the sea floor. The other is the “midwater plume” consisting of unwanted sediment that was collected during extraction that is pumped back into the aphotic zone. The vast majority of plume generation is caused by the collector, causing detrimental and long-lasting impacts on seafloor ecosystems due to the lack of wave activity or strong currents at the sea floor. Therefore, it is crucial to invest in the infrastructure to support the study and constant monitoring over a large area of the sea floor where plume generation is present. Due to the limited number of usable channels and power requirements, current subsea wireless communications technologies are not well suited to instrumenting the large areas of the sea floor needed to monitor plume migration. The scope of this effort is to transition experimental demonstrations of high-bandwidth, full-duplex scalable underwater laser communications to the seafloor in an open ocean environment. Specifically tackling challenges associated with the dynamic nature of the subsea world, including but not limited to, deployment logistics, sustainability, and range. The goal is to enable the internet of underwater things for deep sea industries by broadening the capabilities of subsea communications. By using high-precision laser transmitters, many of the challenges current subsea optical systems face can be circumvented, such as power consumption, interference, and bandwidth limitations. This approach lends itself to wireless interlinking multi-node networks, in series or parallel, facilitating the implementation of a wide array of sensor types. This interlinking allows all the data gathered from the network to be processed through a single hardline uplink to the surface, lowering the complexity required for near real-time data processing. Additionally, the laser control systems produce metadata that can be used to help characterize the water column between the nodes. Combining data from various sensors such as turbidity, temperature, current velocity with metadata such as beam attenuation and deflection can produce a high-resolution model of sea floor conditions around an active mining zone. The resulting near real-time model can be used to optimize location and flow rate of the mining operation to minimize and quantify the environmental impact.

Mons, Ishan↗

NLR HPC Eagle Jobs Data and Additional Energy Metrics

Overview: Anonymized job-level records from the Eagle high-performance computing (HPC) system at the National Laboratory of the Rockies (NLR). Each record represents a Slurm batch job with scheduling metadata, resource requests, resource utilization, CPU/GPU energy consumption, and efficiency metrics. Sensitive fields (user, account, job name) are replaced with cryptographic hashes. System & Timeframe: Eagle was a 2,000-node, 8-petaflop system operated at NLR from 2019–2024. Data covers the full operational lifetime of the system. Slurm data was processed nightly; timestamps are in Mountain Time. Funding provided by the U.S. Department of Energy, EERE. Files: esif.hpc.eagle.job-anon.zip — Core anonymized job records (Hive-partitioned Parquet) esif.hpc.eagle.job-anon-energy-metrics.zip — Same records with additional iLO and Ganglia energy metrics datacard.md — Full dataset documentation ~13.8 million rows, 62 variables. Readable with PyArrow, pandas, DuckDB, Apache Spark, or any Parquet-compatible tool. Data Collection: Jobs collected via sacct through a pipeline: Eagle Jobs API → Redpanda → StreamSets → HPCMON API → PostgreSQL. Node-level power from iLO (HP Integrated Lights-Out); GPU power from Ganglia monitoring, joined to jobs via node lists and time ranges. Preprocessing: Anonymization of name, user, and account fields via cryptographic hashing Derived columns: queue_wait, cpu_eff, max_mem_eff Simplified job state mapping (e.g., "CANCELLED BY 12345" → "CANCELLED") QoS accounting rules (buy-in, standby, or Slurm QoS value) CPU energy estimated from TDP (200W, Intel Xeon Gold 6154, 18 cores) Timezone-aware columns (_tz) sourced from LEX accounting database to correctly handle DST transitions Key Variables: Scheduling: job_id, partition, state_simple, submit_time_tz, start_time_tz, end_time_tz, queue_waitResources: nodes_req/used, processors_req/used, memory_req, wallclock_req/used, gpus_requested Efficiency: cpu_eff, max_mem_eff Energy: cpu_energy_tdp_estimated_max/used_watt_hours, node_energy_total_watt_hours (iLO), gpu0/1_energy_total_watt_hours (Ganglia) Partitions: bigmem, bigmem-8600, bigscratch, csc, dav, ddn, debug, gpu, haswell, long, mono, short, standard Job States: CANCELLED, COMPLETED, FAILED, NODE_FAIL, OUT_OF_MEMORY, PENDING, RUNNING, TIMEOUT QoS Levels: Unknown, normal, buy-in, debug, penalty, high, standby Important Notes: Non-_tz timestamp columns may be off by one hour across DST boundaries; use _tz columns for time difference calculations Energy fields are null for jobs without monitoring coverage Job step records and raw Slurm JSONB fields are excluded from this extract Do not attempt to re-identify individuals from hashed fields

97 MATHEMATICS AND COMPUTING↗

Aqueous Organic Matter from Kougarok Fire Complex, Alaska, 2023

Chemical analyses of aqueous organic matter extracted by filtration from a small set of organic layer samples collected from burned and unburned tussock tundra sites in the Kougarok Fire Complex, near Nome, Alaska. There are five files in *.csv format with one data file and four data description files including data dictionary, methods, terminology, and file-level metadata.The Next-Generation Ecosystem Experiments: Arctic (NGEE Arctic), was a research effort to reduce uncertainty in Earth System Models by developing a predictive understanding of carbon-rich Arctic ecosystems and feedbacks to climate. NGEE Arctic was supported by the Department of Energy's Office of Biological and Environmental Research.The NGEE Arctic project had two field research sites: 1) located within the Arctic polygonal tundra coastal region on the Barrow Environmental Observatory (BEO) and the North Slope near Utqiagvik (Barrow), Alaska and 2) multiple areas on the discontinuous permafrost region of the Seward Peninsula north of Nome, Alaska.Through observations, experiments, and synthesis with existing datasets, NGEE Arctic provided an enhanced knowledge base for multi-scale modeling and contributed to improved process representation at global pan-Arctic scales within the Department of Energy's Earth system Model (the Energy Exascale Earth System Model, or E3SM), and specifically within the E3SM Land Model component (ELM).

54 ENVIRONMENTAL SCIENCES↗

Soil Core Chemistry of Wetland, Old Woman Creek National Estuarine Research Reserve, Huron, OH, 2023-10-24 to 2024-05-20

This dataset contains the chemistry data of soil samples collected from a wetland, referred to as The Cove, at Old Woman Creek Estuarine Research Reserve in Huron, OH. Soil core extractions were performed to analyze what nutrient and/or metal constituents were present at different depths and what biogeochemical activity this could indicate. Three soil cores were collected at three locations within The Cove. The soil cores were removed from their core tubing and were cut into 4 segments down the length (or depth) of the core: top to 1-inch deep, from 1 inch to 5 inches, 5 inches to 7 inches, and 7 inches to the bottom of the core (approximately 10 inches). These soils segments were each homogenized and sub-sampled for various chemical analyses. Soil chemistry measurements are reported in the SoilChem_DataTable.csv file. Collection information about the samples can be found within the SoilChem_SampleMetdata.csv file. All files associated with this dataset are listed in the SoilChem_FLMD.csv file.

54 ENVIRONMENTAL SCIENCES↗

Total metals, sulfur and organic carbon data; Slate River floodplain, Crested Butte, CO; March 2021-October 2021

This data package includes processed and undiluted measurements for metal, sulfur and organic carbon concentrations from pore water (groundwater) samples from the Slate River floodplain of Crested Butte, CO, a focus field site for the SLAC Floodplain Hydro-Biogeochemistry SFA. The data was generated as part of the work targeting the overarching research question for the SLAC SFA: How do ubiquitous subsurface interfaces mediate molecular-scale biogeochemical processes and groundwater quality in floodplains and watersheds? Samples were collected between March and October of 2021. These measurements were all recorded at the Environmental Measurements Facility (EM-1) at Stanford University in Stanford, CA. Groundwater samples were extracted from a network of installed rhizon (Rhizosphere Research Products, part no. 19.60.21F, 0.6 micrometer mesh size) and piezometer wells within the river floodplain. All water samples were shaded from sun exposure during extraction from the subsurface and preserved at 2C until measured at EM-1.Analysis by ICP-MS:Measurements for total metals were performed on an inductively coupled plasma optical emission spectrometer (ICP-OES; iCAP 6300, Thermo Scientific, Cambridge, U.K.). Calibration standards were prepared from the mulit element stock solution (Sigma-Aldrich Multielement standard for ICP, St. Louis, MO) using matrix matched to sample solutions (2% HNO3). Calibration curves included 5 points with correlation coefficients of >0.99. The QC protocol includes a continuing calibration blank and quality control samples that are analyzed just after calibration and again every 20 samples and at the completion of the run. Acceptable QC responses must be between 90-110% of the certified value. Analysis by Total Organic Carbon:Dissolved organic carbon concentrations were quantified on a total organic carbon (TOC) analyzer (TOC-L, Shimadzu, Kyoto, Japan) running the NPOC method. Standard curves were developed using an Organic Carbon Standard from RICCA Chemical Company (Arlington,TX). A series of 4 to 6 standards were automatically diluted by the instrument in a concentration range that spans that of the samples. A blank sample was run just after the calibration curve and at the end of the run. A QC sample was run every 25 samples. Acceptable QC responses must be between 90-110% of the certified value.All files are in csv format.

54 ENVIRONMENTAL SCIENCES↗

Co-located temperature and electrical resistivity measurements for permafrost mapping, Teller 27, Teller 47 and Kougarok 64, Seward Peninsula, Alaska, late summers of 2018, 2019, 2021, and 2022

This data set contains co-located shallow soil (0.8m below ground level) and electrical resistivity at various depth extracted from Electrical Resistivity Tomography (ERT) measurements. Data were acquired at three watersheds on the southern Seward Peninsula during the late summers of 2018, 2019, 2021, and 2022. The three watersheds are located along the Nome-Teller Highway at mile markers 27 and 47, and along Kougarok Road mile marker 64. The co-located temperature and electrical resistivity data were used to (1) map the spatial extend of near surface permafrost, and (2) for supervised classification of permafrost bodies. This dataset contains .csv, .txt, .srv files and flmd reporting format with data dictionary. Find information about the .srv files at "https://e4d-userguide.pnnl.gov/e4d_guide/elec/e4d_e4d-survey.html". The Next-Generation Ecosystem Experiments: Arctic (NGEE Arctic), was a research effort to reduce uncertainty in Earth System Models by developing a predictive understanding of carbon-rich Arctic ecosystems and feedbacks to climate. NGEE Arctic was supported by the Department of Energy's Office of Biological and Environmental Research. The NGEE Arctic project had two field research sites: 1) located within the Arctic polygonal tundra coastal region on the Barrow Environmental Observatory (BEO) and the North Slope near Utqiagvik (Barrow), Alaska and 2) multiple areas on the discontinuous permafrost region of the Seward Peninsula north of Nome, Alaska. Through observations, experiments, and synthesis with existing datasets, NGEE Arctic provided an enhanced knowledge base for multi-scale modeling and contributed to improved process representation at global pan-Arctic scales within the Department of Energy's Earth system Model (the Energy Exascale Earth System Model, or E3SM), and specifically within the E3SM Land Model component (ELM).

54 ENVIRONMENTAL SCIENCES↗

Depth-resolved sagebrush root metabolomics, rhizosphere microbial communities, and geochemistry at the East River Watershed

This data set consists of results from soil nutrient profile, untargeted metabolomics, mass spec imaging, and amplicon sequencing. Data for soil nutrient profile includes common cations (Ca, Mg, Na, and K etc.) extracted from 3 digesting steps – ammonia acetate (for exchangeable cations), nitric acid (for acid dissolved fraction), and hydrofluoric acid/perchloric acid (HF/HClO4) for whole soil digestion. It also includes concentration of organic carbon, inorganic nitrogen (ammonia and nitrate) and phosphorus (Bray-1 P and nitric acid extract), and total nitrogen and phosphorus. Data for untargeted metabolomics includes metabolomic profile for root exudate/tissues and soil extracts from depths at surface soil to saprolite, that were measured using gas chromatography – mass spectrometry (GC-MS), and liquid chromatography – tandem mass spectrometry (LC-MS/MS). Data for mass spec imaging includes spatial distribution of metabolites that were detected and annotated with Fourier transformation ion cyclotron resonance mass spectrometer (FTICR-MS). Data for amplicon sequencing includes the base paired 16S and ITS ribosomal RNA sequences from Miseq Illumina sequencing. All samples were collected from 2 sampling campaign October 2022 and June 2023. Collectively, these datasets enable a mechanistic evaluation of how nutrient acquisition, especially nitrogen and phosphorus, differs between shallow roots operating in soil and deep roots functioning within the fractured bedrock zone. All files are provided as comma-separated values (CSV) fies (.csv) and (GZIP) file (.gz). The compressed .gz FASTQ files can be read directly in R using the dada2 package as part of the amplicon sequence analysis workflow. This work was supported by the Watershed Function Science Focus Area at Lawrence Berkeley National Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-05CH11231. This research was performed on a project award 60563 (https://dx.doi.org/10.46936/expl.proj.2022.60563/60008727) from the Environmental Molecular Sciences Laboratory, a DOE Office of Science User Facility sponsored by the Biological and Environmental Research program under Contract No. DE-AC05-76RL01830.

EARTH SCIENCE > AGRICULTURE > SOILS > CARBON↗

Frost Tables, NGEE Areas B, C and D, Utqiagvik (Barrow), Alaska, 2012-2014

This dataset represents spatially intensive thaw depth surveys with individual point measurements spaced ~0.5 m apart. The three ~10x10m grids cover an ice wedge and a portion of its two neighboring polygons. The data contain thaw depth, frost table elevation, ground surface elevation, active layer depth and surface water inundation across three seasons (2012, 2013 and 2014) at Utqiagvik (Barrow) NGEE Areas B, C and D. There are a total of 69 extracted files separated by month, year, and area. The data were originally provided in excel workbooks, these files have been preserved within the zip "Original_Files." The files were transformed for archiving with images pulled out into JPEGs, data preserved as CSV files and some methodological information preserved as TXT files. Some functions, features, and formatting were lost during the transformation. Missing data were defined with blank cells. Additional documentation can be found in files labelled with "information.txt" file endings.The Next-Generation Ecosystem Experiments: Arctic (NGEE Arctic), was a research effort to reduce uncertainty in Earth System Models by developing a predictive understanding of carbon-rich Arctic ecosystems and feedbacks to climate. NGEE Arctic was supported by the Department of Energy's Office of Biological and Environmental Research.The NGEE Arctic project had two field research sites: 1) located within the Arctic polygonal tundra coastal region on the Barrow Environmental Observatory (BEO) and the North Slope near Utqiagvik (Barrow), Alaska and 2) multiple areas on the discontinuous permafrost region of the Seward Peninsula north of Nome, Alaska.Through observations, experiments, and synthesis with existing datasets, NGEE Arctic provided an enhanced knowledge base for multi-scale modeling and contributed to improved process representation at global pan-Arctic scales within the Department of Energy's Earth system Model (the Energy Exascale Earth System Model, or E3SM), and specifically within the E3SM Land Model component (ELM).[2024-07-15] Renamed BEO_Thaw_Depth_Frost_Table_Area_*_*_*.csv, BEO_Thaw_Depth_Frost_Table_Area_*_*_*_Plot_1.JPEG, and BEO_Thaw_Depth_Frost_Table_Area_*_*_*_Plot_2.JPEG names that had E or L to be Early or Late. Renamed BEO_Thaw_Depth_Frost_Table_Area_*_*_*_Plot_1.JPEG, BEO_Thaw_Depth_Frost_Table_Area_*_*_*_Plot_2.JPEG, and BEO_Thaw_Depth_Frost_Table_Area_*_information_Plot_1.JPEG to BEO_Thaw_Depth_Frost_Table_Area_*_*_*_Gridpoints.JPEG, BEO_Thaw_Depth_Frost_Table_Area_*_*_*_3Dgraphic.JPEG, and BEO_Thaw_Depth_Frost_Table_Area_*_information_LiDAR_DEM.JPEG, respectively. Updated BEO_Thaw_Depth_Frost_Table_flmd.csv file to reflect above updates.

54 ENVIRONMENTAL SCIENCES↗

Total metals, carbon, nitrogen & anion concentration data; Slate River & East River floodplains, Crested Butte, CO; May 2022-October 2022

This data package includes processed and undiluted measurements for metal, total carbon, total nitrogen, and anion concentrations from pore water (groundwater) and surface water samples from the Slate River and East River floodplains of Crested Butte, CO, focus field sites for the SLAC Floodplain Hydro-Biogeochemistry SFA. The data was generated as part of the work targeting the overarching research question for the SLAC SFA: How do ubiquitous subsurface interfaces mediate molecular-scale biogeochemical processes and groundwater quality in floodplains and watersheds? Samples were collected between May and October of 2022. These measurements were all recorded at the Arizona Laboratory for Emerging Contaminants (ALEC) at the University of Arizona located in Tucson, AZ. Groundwater samples were extracted from a network of installed rhizon (Rhizosphere Research Products, part no. 19.60.21F, 0.6 micrometer mesh size) and piezometer wells within the river floodplain. All water samples were shaded from sun exposure during extraction from the subsurface and preserved at 4C until measured at ALEC.Analysis by ICP-MS (metals):Measurements for total metals were made on the Agilent 7700x ICP-MS (for total metals) – Agilent Technologies, Santa Clara, CA. The analytical QA/QC protocol was adapted from US EPA Method 200.8 for analysis by ICP-MS. Calibration standards were prepared from multi-element stock solutions (SPEX Certiprep, Metuchen, NJ). Calibration curves include at least 7 points with correlation coefficients > 0.995. The QC protocol includes a continuing calibration blank (CCB), a continuing calibration verification (CCV) solution and at least one quality control sample (QCS) to be analyzed just after calibration and again after every 12 samples and at the completion of the run. The QCS solutions are from an independent source, such as NIST SRM 1643e - Trace elements in water, or QCS solutions from High Purity Standards (Charleston, SC). Acceptable QC responses must be between 90 and 110% of the certified value. Lastly, a suitable internal standard (usually Rh, In, Ga or Ge) is added using on-line addition into the sample line and mixing tee.Analysis by Shimadzu TOC-L (TOC/TN):The TOC-L system is a combustion technique where liquid samples are injected and combusted into CO2 for carbon detection by non-dispersive infrared (NDIR) and NO for detection by chemiluminescence. A calibration curve using five standard solutions between 0.1 and 7 ppm for carbon and 0.05 and 3.5 ppm for nitrogen is made for each type of measurement with a linearity >0.99. All samples, standards, and QC’s are prepared in 24mL scintillation vials that have been baked for 4hrs at 475 Cº and made using RO water (18.2mΩ). QC’s include a calibration blank check (CCB), continuing calibration check (CCC), and a certified reference material check (CRM). All QC’s are within ±10% error and are run before and after each batch of samples. Samples are diluted and rerun if any measurement concentrations are above the highest standard.Analysis by Ion Chromatography (Anions):The instrument used is the Thermo Scientific Dionex ICS-6000 using AS+AG22 column set for anion analysis with sodium carbonate eluent. A calibration curve using five standard solutions between 5 and 250 umol/L is made with a linearity >0.99. Standards and QC’s are prepared in 15mL polypropylene conical tubes, pipetted along with the samples into 1.5mL polypropylene vials. Dilutions are made using RO water (18.2mΩ). QC’s include a calibration blank check (CCB), continuing calibration check (CCC), and a certified reference material check (CRM). All QC’s are within ±10% error and are run before and after each batch of samples. Samples are diluted and rerun if any measurement concentrations are above the highest standard.All files are in csv format.

54 ENVIRONMENTAL SCIENCES↗

Data from TropiRoot 1.0 database: tropical root characteristics across environments

TropiRoot 1.0 is a new tropical root database with root characteristics across environment gradients. It has data extracted from 104 new sources, resulting in more than 8000 rows of data (either species or community data). Most of the data in TropiRoot 1.0 includes root characteristics such as root biomass, morphology, root dynamics, mass fraction, architecture, anatomy, physiology and root chemistry. This initiative represents an approximately 30% increase in the currently available data for tropical roots in the Fine Root Ecology Database (FRED). TropiRoot 1.0, contains root characteristics from 25 different countries where seven are located in Asia, six in South America, five in Central America and the Caribbean, four in Africa, two in North America, and 1 in Oceania. Due to the volume of data, when ancillary data was available, including soil data, these data was either extracted and included in the database or their availability was recorded in an additional column. Multiple contributors checked the entries for outliers during the collation process to ensure data quality. For text-based observations, we examined all cells to ensure that their content relates to their specific categories. For numerical observations, we ordered each numerical value from least to greatest and plotted the values, checking apparent outliers against the data in their respective sources and correcting or removing incorrect or impossible values. Some data (soil and aboveground) have different columns for the same variable presented in different units, including originally published units, but root characteristics data had units converted to match the ones reported in FRED. By filling a gap from global databases, TropiRoot 1.0 expands our knowledge of otherwise so far underrepresented regions, and our ability to assess global trends. This advancement can be used to improve tropical forest representation in vegetation models.

54 ENVIRONMENTAL SCIENCES↗