Engineering PapersSearch

SEARCH · Engineering Papers

Results for “missing values”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Stable Water Isotope Data for the East River Watershed, Colorado (2014-2025)

The stable water isotope data for the East River Watershed, Colorado, consists of delta2H (hydrogen) and delta18O (oxygen) values from samples collected at multiple, long-term monitoring sites including streams, groundwater wells, springs, and a precipitation collector used to establish a local meteoric water line (LMWL) for the watershed. These locations represent important and/or unique end-member locations for which stable isotope values can be diagnostic of the connection between precipitation inputs as snow and rain and riverine export. Such locations include drainages underline entirely or largely by shale bedrock, land covered dominated by conifers, aspens, or meadows, and drainages impacted by historic mining activity and the presence of naturally mineralized rock. Developing a long-term record of water isotope values from a diversity of environments is a critical component of quantifying the impacts of both climate change and discrete climate perturbations, such as drought, forest mortality, and wildfire, on water export. Such data may be combined with stream gaging stations co-located at each surface water monitoring site to relate seasonal variations in water export to their stable isotopic signature. Data for liquid water delta2H and delta18O values are reported in units of parts per thousand (per-mil; ‰). This data package contains (1) a zip file (isotope_data_2014-2025.zip) containing a total of 95 files: 96 data files of isotope data from across the Lawrence Berkeley National Laboratory (LBNL) Watershed Function Scientific Focus Area (SFA) which is reported in .csv files per location and a locations.csv (1 file) with latitude and longitude for each location; (2) a file-level metadata (v6_20260901_flmd.csv) file that lists each file contained in the dataset with associated metadata; and (3) a data dictionary (v6_20260901_dd.csv) file that contains terms/column_headers used throughout the files along with a definition, units, and data type. Missing values within the anion data files are noted as either "-9999" or "0.0" for not detectable (N.D.) data. There are a total of 43 locations containing isotope data. Update on 2022-06-10: versioned updates to this dataset was made along with these changes: (1) updated isotope data for all locations up to 2021-12-31 and (2) the addition of the file-level metadata (flmd.csv) and data dictionary (dd.csv) were added to comply with the File-Level Metadata Reporting Format. Update on 2022-09-09: Updates were made to reporting format specific files (file-level metadata and data dictionary) to correct swapped file names, add additional details on metadata descriptions on both files, add a header_row column to enable parsing, and add version number and date to file names (v2_20220909_flmd.csv and v2_20220909_dd.csv). Update on 2023-08-08: Updates were made to both the data files and reporting format specific files. New available anion data was added, up until 2023-03-13. The file level metadata and data dictionary files were updated to reflect the additional data added. Update on 2024-03-11: Updates were made to both the data files and reporting format specific files. New available anion data was added, up until 2024-02-19. Further, revisions to the data files were made to remove incorrect data points (from 1970 and 2001). The reporting format specific files were updated to reflect the additional data added. Update on 2025-05-15: Updates were made to both the data files and reporting format specific files. New available isotope data was added, up until the end of WY2024 (September 30, 2024). International Generic Sample Numbers (IGSNs), when registered, were added to the data files. The reporting format specific files were updated to reflect the additional data added. Update on 2026-09-01: Updates were made to both the data files and reporting format specific files. New available isotope data was added, up until the end of WY2025 (September 30, 2025).

54 ENVIRONMENTAL SCIENCES

Dissolved Inorganic Carbon and Dissolved Organic Carbon Data for the East River Watershed, Colorado (2015-2025)

This data package contains mean values for dissolved organic carbon (DOC) and dissolved inorganic carbon (DIC) for water samples taken from the East River Watershed in Colorado. The East River is part of the Watershed Function Scientific Focus Area (WFSFA) located in the Upper Colorado River Basin, United States. DOC and DIC concentrations in water samples were determined using a TOC-VCPH analyzer (Shimadzu Corporation, Japan). DOC was analyzed as non-purgeable organic carbon (NPOC) by purging HCl-acidified samples with carbon-free air to remove DIC prior to measurement. After the acidified sample has been sparged, it is injected into a combustion tube filled with oxidation catalyst heated to 680 oC. The DOC in samples is combusted to CO2 and measured by a non-dispersive infrared (NDIR) detector. The peak area of the analog signal produced by the NDIR detector is proportional to the DOC concentration of the sample. DIC was determined by acidifying the samples with HCl first, and then purging with carbon-free air to release CO2 for analysis by NDIR detector. Total dissolved nitrogen (TDN) was analyzed using a Shimadzu Total Nitrogen Module (TNM-L) combined with the TOC-L analyzer (Shimadzu Corporation, Japan). TNM-L is a non-specific measurement of total nitrogen (TN). All nitrogen species in samples are combusted to nitrogen monoxide and nitrogen dioxide, then reacted with ozone to form an excited state of nitrogen dioxide. Upon returning to ground state, light energy is emitted. Then, TDN is measured using a chemiluminescence detector. All data reported are the mean values upon minimum of three replicate measurements, with a relative standard deviation < 3%. All samples were analyzed under a rigorous quality assurance and quality control (QA/QC) process. This data package contains (1) a zip file (dic_npoc_data_2014-2025.zip) containing a total of 337 files: 336 data files of DIC and NPOC data from across the Lawrence Berkeley National Laboratory (LBNL) Watershed Function Scientific Focus Area (SFA) which is reported in .csv files per location and a locations.csv (1 file) with latitude and longitude for each location; (2) a file-level metadata (v6_20250901_flmd.csv) file that lists each file contained in the dataset with associated metadata; (3) a data dictionary (v6_20250901_dd.csv) file that contains terms/column_headers used throughout the files along with a definition, units, and data type; and (4) PDF and docx files for the determiniation of Method Detection Limits (MDLs) for DIC and NPOC data, which has been updated in 2026-08. Missing values within the anion data files are noted as either "-9999" or "0.0" for not detectable (N.D.) data. There are a total of 113 locations containing DIC/NPOC data. Update on 2020-10-07: Updated the data files to remove times from the timestamps, so that only dates remain. The data values have not changed. Update on 2021-04-11: Added Determination of Method Detection Limits (MDLs) for DIC, NPOC and TDN Analyses document, which can be accessed as a PDF or with Microsoft Word.Update on 6/10/2022: versioned updates to this dataset was made along with these changes: (1) updated dissolved inorganic carbon and dissolved organic carbon data for all locations up to 2021-12-31, (2) removal of units from column headers in datafiles, (3) added row underneath headers to contain units of variables, (4) restructure of units to comply with CSV reporting format requirements, (5) added -9999 for empty numerical cells, and (6) the addition of the file-level metadata (flmd.csv) and data dictionary (dd.csv) were added to comply with the File-Level Metadata Reporting Format. Update on 2022-09-09: Updates were made to reporting format specific files (file-level metadata and data dictionary) to correct swapped file names, add additional details on metadata descriptions on both files, add a header_row column to enable parsing, and add version number and date to file names (v2_20220909_flmd.csv and v2_20220909_dd.csv). Update on 2023-08-08: Updates were made to both the data files and reporting format specific files. New available anion data was added, up until 2023-01-05. The file level metadata and data dictionary files were updated to reflect the additional data added. Update on 2024-03-11: Updates were made to both the data files and reporting format specific files. New available anion data was added, up until 2023-11-21. Further, revisions to the data files were made to remove incorrect data points (from 1970 and 2001). The reporting format specific files were updated to reflect the additional data added. Revised versions of the PDF and docx files for determination of MDLs for DIC and NPOC were added to replace previous versions. Update on 2025-05-15: Updates were made to both the data files and reporting format specific files. New available DIC and NPOC data was added, up until the end of WY2024 (September 30, 2024). International Generic Sample Numbers (IGSNs), when registered, were added to the data files. The reporting format specific files were updated to reflect the additional data added. Update on 2026-09-01: Updates were made to both the data files and reporting format specific files. New available DIC and NPOC data was added, up until the end of WY2025 (September 30, 2025). Updated versions, as of 2026-08-10, of the PDF and docx files for determination of MDLs for DIC and NPOC data were added to this dataset.

54 ENVIRONMENTAL SCIENCES

Total Dissolved Nitrogen and Ammonia Data for the East River Watershed, Colorado (2015-2025)

This data package contains mean values for total dissolved nitrogen (TDN) and ammonia concentrations for water samples taken from the East River Watershed in Colorado. The East River is part of the Watershed Function Scientific Focus Area (WFSFA) located in the Upper Colorado River Basin, United States. TDN was analyzed using a Shimadzu Total Nitrogen Module (TNM-1) combined with the TOC-VCSH analyzer (Shimadzu Corporation, Japan). TNM-1 is a non-specific measurement of total nitrogen (TN). All nitrogen species in samples are combusted to nitrogen monoxide and nitrogen dioxide, then reacted with ozone to form an excited state of nitrogen dioxide. Upon returning to ground state, light energy is emitted. Then, TDN is measured using a chemiluminescence detector. Ammonia was determined using a Lachat's QuikChem 8500 Series 2 Flow Injection Analysis System (LACHAT Instruments, QuckChem 8500 series 2, Automated Ion Analyzer, Loveland, Colorado). When ammonia in water samples is heated (60 degrees C) with salicylate and hypochlorite in an alkaline phosphate buffer, an emerald green color is produced which is proportional to the ammonia concentration. The color is intensified by the addition of nitroprusside. Ethylenediaminetetraacetic acid (EDTA) is added to the buffer to prevent the interference of metal ions (Ca, Mg, and Fe etc.). Ammonia-N is then determined by LACHAT flow injection and a colorimetric assay at an absorbance wavelength 660 nm. (Reference: LACHAT Instruments: QuickChem Method 90-107-06-3-A, Determination of Ammonia by Flow Injection Analysis (High Throughput, Salicylate Method/DCIC) (Multi Matrix method). Written by Lynn Egan (Application group), February 08, 2011.) All files are labeled by location and variable, and data reported are the mean values upon replicate measurements. All samples were analyzed under a rigorous quality assurance and quality control (QA/QC) process as detailed in the methods. This data package contains (1) a zip file (tdn_ammonia_data_2015-2025.zip) containing a total of 299 files: 298 data files of ammonia and TDN data from across the Lawrence Berkeley National Laboratory (LBNL) Watershed Function Scientific Focus Area (SFA) which is reported in .csv files per location and a locations.csv (1 file) with latitude and longitude for each location; (2) a file-level metadata (v7_20260901_flmd.csv) file that lists each file contained in the dataset with associated metadata; (3) a data dictionary (v7_20260901_dd.csv) file that contains terms/column_headers used throughout the files along with a definition, units, and data type; (4) PDF and docx files for the determination of Method Detection Limits (MDLs) for TDN data, which has been updated in 2026-08; and (5) PDF and docx files for the detemination of Method Detection Limits (MDLs) for Ammonia and the Interferences by LACHAT Flow Injection Analysis. Missing values within the anion data files are noted as either "-9999" or "0.0" for not detectable (N.D.) data. There are a total of 105 locations containing TDN and Ammonia-N data. Update 2020-10-07: Updated the data files to remove times from the timestamps, so that only dates remain. The data values have not changed. Update 2021-04-11: Added Determination of Method Detection Limits (MDLs) for DIC, NPOC and TDN Analyses and Determination of Method Detection Limit for Ammonia and the Interferences by LACHAT Flow Injection Analysis documents, which can be accessed as PDFs or with Microsoft Word.Update on 6/10/2022: versioned updates to this dataset was made along with these changes: (1) updated total dissolved nitrogen and ammonia data for all locations up to 2021-12-31, (2) removal of units from column headers in datafiles, (3) added row underneath headers to contain units of variables, (4) restructure of units to comply with CSV reporting format requirements, (5) added -9999 for empty numerical cells, and (6) the addition of the file-level metadata (flmd.csv) and data dictionary (dd.csv) were added to comply with the File-Level Metadata Reporting Format. Update on 2022-09-09: Updates were made to reporting format specific files (file-level metadata and data dictionary) to correct swapped file names, add additional details on metadata descriptions on both files, add a header_row column to enable parsing, and add version number and date to file names (v2_20220909_flmd.csv and v2_20220909_dd.csv). Update on 2022-12-20: Updates were made to both the data files and reporting format specific files. Units were listed incorrectly, but have been fixed to reflect correct units (ug/L). File level metadata (flmd) and data dictionary (dd) files were updated to reflect the updated versions of these files. Available data was added up until 2022-06-01. Update on 2023-08-08: Updates were made to both the data files and reporting format specific files. New available anion data was added, up until 2023-01-05. The file level metadata and data dictionary files were updated to reflect the additional data added. Update on 2024-03-11: Updates were made to both the data files and reporting format specific files. New available anion data was added, up until 2023-10-27. Further, revisions to the data files were made to remove incorrect data points (from 1970 and 2001). The reporting format specific files were updated to reflect the additional data added. Revised versions of the PDF and docx files for determination of MDLs for TDN were added to replace previous versions. Update on 2025-05-15: Updates were made to both the data files and reporting format specific files. New available TDN and Ammonia-N data was added, up until the end of WY2024 (September 30, 2024). International Generic Sample Numbers (IGSNs), when registered, were added to the data files. The reporting format specific files were updated to reflect the additional data added. Update on 2026-09-01: Updates were made to both the data files and reporting format specific files. New available TDN and Ammonia-N data was added, up until the end of WY2025 (September 30, 2025). Updated versions, as of 2026-08-10, of the PDF and docx files for determination of MDLs for TDN data were added to this dataset.

54 ENVIRONMENTAL SCIENCES

Cation Data for the East River Watershed, Colorado (2014-2025)

This data package contains mean values for cation concentration for water samples taken from the East River Watershed in Colorado. Inductively coupled plasma mass spectrometry (ICP-MS) has been used to measure the concentrations of elements of interest simultaneously for the East River Watershed, Colorado groundwater and surface water samples to inform insights on the biogeochemistry processes within the watershed. The East River is part of the Watershed Function Scientific Focus Area (WFSFA) located in the Upper Colorado River Basin, United States. For samples collected prior to 06-16-2021, the instrumentation, Elan DRC II, PerkinElmer SCIEX, automatically switches among the three models necessary to analyze all 37 elements. These 37 elements include: (1) Lithium (Li), Beryllium (Be), Boron (B), Sodium (Na), Magnesium (Mg), Aluminium (Al), Silicon (Si), Phosphorus (P), Titanium (Ti), Cobalt (Co), Nickel (Ni), Copper (Cu), Zinc (Zn), Germanium (Ge), Arsenic (As), Rubidium (Rb), Strontium (Sr), Zirconium (Zr), Molybdenum (Mo), Silver (Ag), Cadmium (Cd), Tin (Sn), Antimony (Sb), Caesium (Cs), Barium (Ba), Europium (Eu), Lead (Pb), Thorium (Th), Uranium (U) using standard model, argon Ar as reaction gas, (2) Potassium (K), Calcium (Ca), Vanadium (V), Chromium (Cr), Manganese (Mn), Iron (Fe) using dynamic reaction cell (DRC) model, ammonia NH3 as reaction gas, and (3) Phosphorus (P) and Selenium (Se) using DRC model, oxygen O2 as reaction gas. Note for the samples with higher concentrations of chloride (Cl-), asenic (As) concentrations were analysed with DRC model (oxygen O2 as reaction gas) to avoid the interference of chloride. For samples collected on and after 06-16-2021, an advanced Agilent 8900 triple quadrupole inductively coupled plasma mass spectrometry system (Agilent 8900 QQQ ICP-MS, Agilent Technologies) has been used to measure the concentrations of interested 36 elements simultaneously for environmental samples, including (1) Lithium (Li), Beryllium (Be) and Boron (B) using standard no gas mode, (2) Sodium (Na), Magnesium (Mg), Aluminium (Al) Phosphorus (P), Potassium (K), Chromium (Cr), Manganese (Mn), Iron (Fe), Cobalt (Co), Nickel (Ni), Copper (Cu), Zinc (Zn), Germanium (Ge), Arsenic (As), Rubidium (Rb), Strontium (Sr), Zirconium (Zr), Molybdenum (Mo), Silver (Ag), Cadmium (Cd), Tin (Sn), Antimony (Sb), Cesium (Cs), Barium (Ba), Europium (Eu), Lead (Pb), Thorium (Th) and Uranium (U) using standard helium (He) collision mode, (3) Titanium (Ti) and Vanadium (V) using high Energy (HEHe) helium (He) collision mode, and (4) Silicon (Si), Calcium (Ca) and Selenium (Se) using standard H2 reaction mode. All samples were prepared/diluted with 2% (v/v) ultrapure nitric acid in Milli-Q water (18.2 mega ohm-cm), and analyzed under a rigorous quality assurance and quality control (QA/QC) process. This data package contains (1) a zip file (cation_data_2014_2025.zip) containing a total of 5,849 files: 5.848 data files of cation data from across the Lawrence Berkeley National Laboratory (LBNL) Watershed Function Scientific Focus Area (SFA) which is reported in .csv files per location and a locations.csv (1 file) with latitude and longitude for each location; (2) a file-level metadata (v6_20260901_flmd.csv) file that lists each file contained in the dataset with associated metadata; (3) a data dictionary (v6_20260901_dd.csv) file that contains terms/column_headers used throughout the files along with a definition, units, and data type; (4) PDF and docx files for the detemination of Method Detection Limits (MDLs) for ICP-MS PerkinElmer DRC II instrumentation (Detemination_of_Method_Detection_Limits__MDLs__for_ICP_MS__PerkinElmer_Elan_DRC_II__LBL_Bldg74_Lab214D) for samples before November 2021; (5) PDF and docx files for the determination of MDLs for ICP-MS Agilent 8900 QQQ instrumentation (ICP_MS_Analysis_detection_limits_and_QA_QC_WenmingDong_updated_2026-08-06) for samples November 2021 and onward. Missing values within the anion data files are noted as either "-9999" or "0.0" for not detectable (N.D.) data. There are a total of 113 locations containing cation data. Update on 2021-04-11: Added Detemination of Method Detection Limits (MDLs) for ICP-MS document, which can be accessed as a PDF or with Microsoft Word. Update on 2022-06-10: versioned updates to this dataset was made along with these changes: (1) updated cation data for all locations up to 2021-12-31, (2) removal of units from column headers in datafiles, (3) added row underneath headers to contain units of variables, (4) removed suffix and prefix on two variables (“aqberylliumion_asberyllium” and “aqlithiumion_aslithium”), (5) added -9999 for empty numerical cells, and (6) the addition of the file-level metadata (flmd.csv) and data dictionary (dd.csv) were added to comply with the File-Level Metadata Reporting Format. Update on 2022-09-09: Updates were made to reporting format specific files (file-level metadata and data dictionary) to correct swapped file names, add additional details on metadata descriptions on both files, add a header_row column to enable parsing, and add version number and date to file names (v2_20220909_flmd.csv and v2_20220909_dd.csv). Update on 2023-08-08: Updates were made to both the data files and reporting format specific files. New available anion data was added, up until 2023-01-05. The file level metadata and data dictionary files were updated to reflect the additional data added. Update on 2024-03-11: Updates were made to both the data files and reporting format specific files. New available anion data was added, up until 2023-10-16. Further, revisions to the data files were made to remove incorrect data points (from 1970 and 2001). The reporting format specific files were updated to reflect the additional data added. Updated versions of the PDF and docx files for determination of MDLs for ICP-MS data were added to this dataset for samples starting in November 2021. Update on 2025-05-15: Updates were made to both the data files and reporting format specific files. New available cation data was added, up until the end of WY2024 (September 30, 2024). International Generic Sample Numbers (IGSNs), when registered, were added to the data files. The reporting format specific files were updated to reflect the additional data added. Update on 2026-09-01: Updates were made to both the data files and reporting format specific files. New available cation data was added, up until the end of WY2025 (September 30, 2025). Updated versions, as of 2026-08-06, of the PDF and docx files for determination of MDLs for ICP-MS data were added to this dataset for samples starting in November 2021.

54 ENVIRONMENTAL SCIENCES

Predictive analytics of selections of russet potatoes

We explore the application of machine learning algorithms specifically to enhance the selection process of Russet potato (Solanum tuberosum L.) clones in breeding trials by predicting their suitability for advancement. This study addresses the challenge of efficiently identifying high-yield, disease-resistant, and climate-resilient potato varieties that meet processing industry standards. Leveraging manually collected data from trials in the state of Oregon, we investigate the potential of a wide variety of state-of-the-art binary classification models. The dataset includes 1086 clones, with data on 38 attributes recorded for each clone, focusing on yield, size, appearance, and frying characteristics, with several control varieties planted consistently across four Oregon regions from 2013 to 2021. We conduct a comprehensive analysis of the dataset that includes preprocessing, feature engineering, and imputation to address missing values. We focus on several key metrics such as accuracy, F1-score, and Matthews correlation coefficient (MCC) for model evaluation. The top-performing models, namely a feedforward neural network classifier (Neural Net), a histogram-based gradient boosting classifier (HGBC), and a support vector machine classifier (SVM), demonstrate consistent and significant results. To further validate our findings, we conducted a simulation study using the aims, data-generating mechanisms, estimands, methods, and performance measures (ADEMP) framework, simulating different data-generating scenarios to assess model robustness and performance through true positive, true negative, false positive, and false negative distributions, area under the receiver operating characteristic curve (AUC-ROC) and MCC. The simulation results highlight that non-linear models like SVM and HGBC consistently show higher AUC-ROC and MCC than logistic regression, thus outperforming the traditional linear model across various distributions, and emphasizing the importance of model selection and tuning in agricultural trials. Variable selection further enhances model performance and identifies influential features in predicting trial outcomes. The findings emphasize the potential of machine learning in streamlining the selection process for potato varieties, offering benefits such as increased efficiency, substantial cost savings, and judicious resource utilization. Our study contributes insights into precision agriculture and showcases the relevance of advanced technologies for informed decision-making in breeding programs.

60 APPLIED LIFE SCIENCES

Anion Data for the East River Watershed, Colorado (2014-2025)

The anion data for the East River Watershed, Colorado, consist of fluoride, chloride, sulfate, nitrate, and phosphate concentrations collected at multiple, long-term monitoring sites that include stream, groundwater, and spring sampling locations. These locations represent important and/or unique end-member locations for which solute concentrations can be diagnostic of the connection between terrestrial and aquatic systems. Such locations include drainages underlined entirely or largely by shale bedrock, land covered dominated by conifers, aspens, or meadows, and drainages impacted by historic mining activity and the presence of naturally mineralized rock. Developing a long-term record of solute concentrations from a diversity of environments is a critical component of quantifying the impacts of both climate change and discrete climate perturbations, such as drought, forest mortality, and wildfire, on the riverine export of multiple anionic species. Such data may be combined with stream gauging stations co-located at each monitoring site to directly quantify the seasonal and annual mass flux of these anionic species out of the watershed. This data package contains (1) a zip file (anion_data_2014_2025.zip) containing a total of 386 files: 387 data files of anion data from across the Lawrence Berkeley National Laboratory (LBNL) Watershed Function Scientific Focus Area (SFA) which is reported in .csv files per location and a locations.csv (1 file) with latitude and longitude for each location; (2) a file-level metadata (v7_20260901_flmd.csv) file that lists each file contained in the dataset with associated metadata; (3) a data dictionary (v7_20260901_dd.csv) file that contains terms/column_headers used throughout the files along with a definition, units, and data type; and (4) a anion MDL fact sheet (anion_MDLs_202608 in PDF and docx formats). Missing values within the anion data files are noted as either "-9999" or "0.0" for not detectable (N.D.) data. There are a total of 47 locations containing anion data. Update on 2022-06-10: versioned updates to this dataset was made along with these changes: (1) updated anion data for all locations up to 2021-12-31, (2) removal of units from column headers in datafiles, (3) added row underneath headers to contain units of variables, (4) restructure of units to comply with CSV reporting format requirements, and (5) the addition of the file-level metadata (flmd.csv) and data dictionary (dd.csv) were added to comply with the File-Level Metadata Reporting Format. Update on 2022-09-09: Updates were made to reporting format specific files (file-level metadata and data dictionary) to correct swapped file names, add additional details on metadata descriptions on both files, add a header_row column to enable parsing, and add version number and date to file names (v2_20220909_flmd.csv and v2_20220909_dd.csv). Update on 2022-12-20: Updates were made to both the data files and reporting format specific files. Conversion issues affecting ER-PLM locations for anion data was resolved for the data files. Additionally, the flmd and dd files were updated to reflect the updated versions of these files. Available data was added up until 2022-03-14. Update on 2023-08-08: Updates were made to both the data files and reporting format specific files. New available anion data was added, up until 2023-05-19. The file level metadata and data dictionary files were updated to reflect the additional data added. Update on 2024-03-11: Updates were made to both the data files and reporting format specific files. New available anion data was added, up until 2023-09-11. Further, revisions to the data files were made to remove incorrect data points (from 1970 and 2001). The reporting format specific files were updated to reflect the additional data added. Update on 2025-05-15: Updates were made to both the data files and reporting format specific files. New available anion data was added, up until the end of WY2024 (September 30, 2024). International Generic Sample Numbers (IGSNs), when registered, were added to the data files. The reporting format specific files were updated to reflect the additional data added. Update on 2026-09-01: Updates were made to both the data files and reporting format specific files. New available anion data was added, up until the end of WY2025 (September 30, 2025). An anion MDL document was included in this update.

54 ENVIRONMENTAL SCIENCES

Gap-filled methane and carbon dioxide fluxes across two ecosystem states at the US-OWC AmeriFlux site (2015−2016, 2020−2022)

This dataset contains gap-filled measurements of methane flux (FCH4), net ecosystem CO2 exchange (NEE) partitioned into gross primary productivity (GPP) and ecosystem respiration (RE), as well as latent heat flux (LE) from a Great Lakes coastal freshwater wetland at the US-OWC AmeriFlux site. The dataset covers the peak growing seasons (June−September) of 2015−2016, dominated by Typha spp., and 2020−2022, characterized by floating-leaved species (lotus and water lily). These data were generated to investigate how rising water levels and vegetation shifts influence CH4 and CO2 fluxes across two distinct ecosystem states in this wetland. The dataset, provided in CSV format, includes half-hourly gap-filled flux data from June to September for 2015, 2016, 2020, 2021, and 2022. The gap-filled data refers to measurements where missing values due to instrument issues or quality control were filled using artificial neural networks (ANNs).

54 ENVIRONMENTAL SCIENCES

Data from a throughfall exclusion experiment: Fine root dynamics, morphology, chemistry, and AMF colonization across four lowland Panamanian forests

Fine roots regulate forest nutrient, carbon, and water cycling, yet their variation within and among tropical forests remains under-characterized. We quantified root productivity, disappearance, and stocks to 1 m using minirhizotron imaging, and we measured morphology, elemental composition [root carbon (C), root nitrogen (N), root phosphorus (P)], and arbuscular mycorrhizal fungi (AMF) colonization to 20 cm using ingrowth cores and sequential coring. Sampling took place in four distinct lowland Panamanian forests (32 plots; 8 per forest) from 2018 through 2022 under control and throughfall-exclusion (drought) treatments in the Panama Rainforest Changes with Experimental Drying (PARCHED) experiment.The dataset is presented as an Excel workbook with six tabs. The first tab is the data dictionary. Tab S1 contains ingrowth-core production and mortality, morphology and soil moisture. Tab S2 contains sequential-coring standing stocks with associated morphology and soil moisture. Tab S3 contains minirhizotron row data records to 1 m depth, including per-frame root length and diameter, normalized length metrics, and session timing. Tab S4 contains AMF colonization. Tab S5 contains fine-root chemistry at 0–10 cm, reporting %P, %C, %N, and C:N for samples collected via ingrowth cores and sequential-coring standing stocks. CSV mirrors for each tab are provided, and a KML file supplies coordinates for all 32 plots.Key variables span live and dead fine-root biomass (and coarse fractions where applicable), specific root length (SRL) and area (SRA), diameter, root tissue density (RTD), soil moisture, AMF colonization, root %N, %C, %P, and C:N, along with minirhizotron root length and diameter. Depth, season, treatment, and plot/site identifiers are included to support cross-tab integration and analysis from 0–100 cm (minirhizotron) and 0–20 cm (cores).Units are reported in-column and missing values are coded as NA. No special software is required to open or use the files (Excel, CSV, and KML compatible).

54 ENVIRONMENTAL SCIENCES

A Photochemical Phosphorus-Hydrogen-Oxygen Network for Hydrogen-dominated Exoplanet Atmospheres

Due to the detection of phosphine (PH 3 ) in the solar system gas giants Jupiter and Saturn, PH 3 has long been suggested to be detectable in exosolar substellar atmospheres too. However, to date, direct detection of phosphine has proven to be elusive in exoplanet atmosphere surveys. We construct an updated phosphorus-hydrogen-oxygen (PHO) photochemical network suitable for the simulation of gas giant hydrogen-dominated atmospheres. Using this network, we examine PHO photochemistry in hot Jupiter and warm Neptune exoplanet atmospheres at solar and enriched metallicities. Our results show for HD 189733b-like hot Jupiters that HOPO, PO, and P 2 are typically the dominant P carriers at pressures important for transit and emission spectra, rather than PH 3 . For GJ1214b-like warm Neptune atmospheres our results suggest that at solar metallicity PH 3 is dominant in the absence of photochemistry, but is generally not in high abundance for all other chemical environments. At 10 and 100 times solar, small oxygenated phosphorus molecules such as HOPO and PO dominate for both thermochemical and photochemical simulations. The network is able to reproduce well the observed PH 3 abundances on Jupiter and Saturn. Despite progress in improving the accuracy of the PHO network, large portions of the reaction rate data remain with approximate, uncertain, or missing values, which could change the conclusions of the current study significantly. Improving understanding of the kinetics of phosphorus-bearing chemical reactions will be a key undertaking for astronomers aiming to detect phosphine and other phosphorus species in both rocky and gaseous exoplanetary atmospheres in the near future.

atmospheric composition

Confidence-Based Feature Acquisition

Confidence-based Feature Acquisition (CFA) is a novel, supervised learning method for acquiring missing feature values when there is missing data at both training (learning) and test (deployment) time. To train a machine learning classifier, data is encoded with a series of input features describing each item. In some applications, the training data may have missing values for some of the features, which can be acquired at a given cost. A relevant JPL example is that of the Mars rover exploration in which the features are obtained from a variety of different instruments, with different power consumption and integration time costs. The challenge is to decide which features will lead to increased classification performance and are therefore worth acquiring (paying the cost). To solve this problem, CFA, which is made up of two algorithms (CFA-train and CFA-predict), has been designed to greedily minimize total acquisition cost (during training and testing) while aiming for a specific accuracy level (specified as a confidence threshold). With this method, it is assumed that there is a nonempty subset of features that are free; that is, every instance in the data set includes these features initially for zero cost. It is also assumed that the feature acquisition (FA) cost associated with each feature is known in advance, and that the FA cost for a given feature is the same for all instances. Finally, CFA requires that the base-level classifiers produce not only a classification, but also a confidence (or posterior probability).

Wagstaff, Kiri L.

Land Surface Phenology from MODIS: Characterization of the Collection 5 Global Land Cover Dynamics Product

Information related to land surface phenology is important for a variety of applications. For example, phenology is widely used as a diagnostic of ecosystem response to global change. In addition, phenology influences seasonal scale fluxes of water, energy, and carbon between the land surface and atmosphere. Increasingly, the importance of phenology for studies of habitat and biodiversity is also being recognized. While many data sets related to plant phenology have been collected at specific sites or in networks focused on individual plants or plant species, remote sensing provides the only way to observe and monitor phenology over large scales and at regular intervals. The MODIS Global Land Cover Dynamics Product was developed to support investigations that require regional to global scale information related to spatiotemporal dynamics in land surface phenology. Here we describe the Collection 5 version of this product, which represents a substantial refinement relative to the Collection 4 product. This new version provides information related to land surface phenology at higher spatial resolution than Collection 4 (500-m vs. 1-km), and is based on 8-day instead of 16-day input data. The paper presents a brief overview of the algorithm, followed by an assessment of the product. To this end, we present (1) a comparison of results from Collection 5 versus Collection 4 for selected MODIS tiles that span a range of climate and ecological conditions, (2) a characterization of interannual variation in Collections 4 and 5 data for North America from 2001 to 2006, and (3) a comparison of Collection 5 results against ground observations for two forest sites in the northeastern United States. Results show that the Collection 5 product is qualitatively similar to Collection 4. However, Collection 5 has fewer missing values outside of regions with persistent cloud cover and atmospheric aerosols. Interannual variability in Collection 5 is consistent with expected ranges of variance suggesting that the algorithm is reliable and robust, except in the tropics where some systematic differences are observed. Finally, comparisons with ground data suggest that the algorithm is performing well, but that end of season metrics associated with vegetation senescence and dormancy have higher uncertainties than start of season metrics.

Land cover dynamics

MODIS Collection 6 MAIAC Algorithm

This paper describes the latest version of the algorithm MAIAC (Multi-Angle Implementation of Atmospheric Correction) used for processing the MODIS (Moderate-resolution Imaging Spectroradiometer) Collection6 data record. Since initial publication in 2011-2012, MAIAC has changed considerably to adapt to global processing and improve cloud/snow detection, aerosol retrievals and atmospheric correction of MODIS data. The main changes include (1) transition from a 25 to 1 km scale for retrieval of the spectral regression coefficient (SRC) which helped to remove occasional blockiness at 25 km scale in the aerosol optical depth (AOD) and in the surface reflectance, (2) continuous improvements of cloud detection, (3) introduction of smoke and dust tests to discriminate absorbing fine- and coarse mode aerosols, (4) adding over-water processing, (5) general optimization of the LUT (LookUp-Table)-based radiative transfer for the global processing, and others. MAIAC provides an interdisciplinary suite of atmospheric and land products, including cloud mask (CM), column water vapor (CWV), AOD at 0.47 and 0.55 m, aerosol type (background, smoke or dust) and fine-mode fraction over water; spectral bidirectional reflectance factors (BRF), parameters of Ross-thick Lisparse (RTLS) bidirectional reflectance distribution function (BRDF) model and instantaneous albedo. For snow-covered surfaces, we provide subpixel snow fraction and snow grain size. All products come in standard HDF4 (software library) format at 1 km resolution, except for BRF, which is also provided at 500 m resolution on a sinusoidal grid adopted by the MODIS Land team. All products are provided on per-observation basis in daily files except for the BRDF/Albedo product, which is reported every 8 days. Because MAIAC uses a time series approach, BRDF/Albedo is naturally gap-filled over land where missing values are filled-in with results from the previous retrieval. While the BRDF model is reported for MODIS Land bands 1-7 and ocean band 8, BRF is reported for both land and ocean bands 1-12. This paper focuses on MAIAC cloud detection, aerosol retrievals and atmospheric correction and describes MCD19 data products and quality assurance (QA) flags.

MAIAC Algorithm

MLtool: Universal Supervised Machine Learning Tool to Model Tabulated Data

Machine Learning (ML) is a subfield of Artificial Intelligence that gives computers the ability to learn from past data without being explicitly programmed. The predictive capabilities of ML models have already been used to facilitate several scientific breakthroughs. However, the practical application of ML is often limited due to the gaps in technical knowledge of its users. The common issue faced by many scientific researchers is the inability to choose the appropriate ML pipelines that are needed to treat real-world data, which is often sparse and noisy. To solve this problem, we have developed an automated Machine Learning tool (MLtool) that includes a set of ML algorithms and approaches to aid scientific researchers. The current version of MLtool is implemented as an object-oriented Python code that is easily extensible. It includes 44 different regression algorithms used to model data. MLtool helps users select the best model for their data, based on the scoring metrics used. Besides regression algorithms, MLtool also includes a suite of pre- and post-processing techniques such as missing value imputation, categorical variable encoding, input feature normalization, uncertainty quantification, exploratory data analysis (EDA), etc. MLtool was tested on several publicly available multi-dimensional data sets and was found capable of making accurate predictions.

Machine learning

MLtool Python Code

Machine Learning (ML) is a subfield of Artificial Intelligence that gives computers the ability to learn from past data without being explicitly programmed. The predictive capabilities of ML models have already been used to facilitate several scientific breakthroughs. However, the practical application of ML is often limited due to the gaps in technical knowledge of its users. The common issue faced by many scientific researchers is the inability to choose the appropriate ML pipelines that are needed to treat real-world data, which is often sparse and noisy. To solve this problem, we have developed an automated Machine Learning tool (MLtool) that includes a set of ML algorithms and approaches to aid scientific researchers. The current version of MLtool is implemented as an object-oriented Python code that is easily extensible. It includes 44 different regression algorithms used to model data. MLtool helps users select the best model for their data, based on the scoring metrics used. Besides regression algorithms, MLtool also includes a suite of pre- and post-processing techniques such as missing value imputation, categorical variable encoding, input feature normalization, uncertainty quantification, exploratory data analysis (EDA), etc. MLtool was tested on several publicly available multi-dimensional data sets and was found capable of making accurate predictions.

Machine Learning

Merging of OMI and AIRS Ozone Data

The OMI Instrument measures ozone using the backscattered light in the UV part of the spectrum. In polar night there are no OMI measurements so we hope to incorporate the AIRS ozone data to fill in these missing regions. AIRS is on the Aqua platform and has been operating since May 2002. AIRS is a multi-detector array grating spectrometer containing 2378 IR channels between 650 per centimeter and 2760 per centimeter which measures atmospheric temperature, precipitable water, water vapor, CO, CH4, CO2 and ozone profiles and column amount. It can also measure effective cloud fraction and cloud top pressure for up to two cloud layers and sea-land skin temperature. Since 2008, OMI has had part of its aperture occulted with a piece of the thermal blanket resulting in several scan positions being unusable. We hope to use the AIRS data to fill in the missing ozone values for those missing scan positions.

OMI

Observed and Imputed Volumetric Soil Water Content Timeseries for the New Mexico Elevation Gradient

Reliable soil water content (SWC) data are essential for understanding dryland ecosystem dynamics, but high-frequency SWC sensors often fail, creating gaps in critical datasets. To address this, we developed a Bayesian mixture model that imputes missing SWC using both linear interpolation and an ecosystem water balance model (SOILWAT2), tested across six AmeriFlux eddy covariance tower sites in the New Mexico Elevation Gradient, demonstrating its effectiveness in reconstructing SWC patterns while providing insights into the factors driving SWC variability. Daily volumetric soil water content (SWC) data are provided as csv-formatted spreadsheets for the six AmeriFlux sites (US-Seg, US-Ses, US-Wjs, US-Mpi, US-Vcp, and US-Vcs). For each site there is an observed SWC file (site_SWC_gapfill.csv) and a file that contains imputed SWC (imputed_SWC_site.csv). The observed SWC files contain temperature corrected sensor values, tower precipitation data, as well as outputs from SOILWAT2 simulations that were used to impute SWC. The imputed files contain the original observed SWC values and the imputed missing SWC values. When SWC was missing from the original data, the missing value was imputed based on the Bayesian imputation mixture model. The posterior mean of all imputed values is reported as "mean_X". When the observed SWC was NOT missing, mean_X = observed SWC value (original data). The standard deviation, 2.5th percentile and the 97.5th percentile for the imputed values are also reported in the imputed files. There are readme text files for each file type explaining the contents of each column.

54 ENVIRONMENTAL SCIENCES