Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “oregon”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 253 records · Page 14

Measurements and Model Improvement: Insight into NWP Model Error Using Doppler Lidar and Other WFIP2 Measurement Systems

Abstract Doppler-lidar wind-profile measurements at three sites were used to evaluate NWP model errors from two versions of NOAA’s 3-km-grid HRRR model, to see whether updates in the latest version 4 reduced errors when compared against the original version 1. Nested (750-m grid) versions of each were also tested to see how grid spacing affected forecast skill. The measurements were part of the field phase of the Second Wind Forecasting Improvement Project (WFIP2), an 18-month deployment into central Oregon–Washington, a major wind-energy-producing region. This study focuses on errors in simulating marine intrusions, a summertime, 600–800-m-deep, regional sea-breeze flow found to generate large errors. HRRR errors proved to be complex and site dependent. The most prominent error resulted from a premature drop in modeled marine-intrusion wind speeds after local midnight, when lidar-measured winds of greater than 8 m s −1 persisted through the next morning. These large negative errors were offset at low levels by positive errors due to excessive mixing, complicating the interpretation of model “improvement,” such that the updates to the full-scale versions produced mixed results, sometimes enhancing but sometimes degrading model skill. Nesting consistently improved model performance, with version 1’s nest producing the smallest errors overall. HRRR’s ability to represent the stages of sea-breeze forcing was evaluated using radiation budget, surface-energy balance, and near-surface temperature measurements available during WFIP2. The significant site-to-site differences in model error and the complex nature of these errors mean that field-measurement campaigns having dense arrays of profiling sensors are necessary to properly diagnose and characterize model errors, as part of a systematic approach to NWP model improvement. Significance Statement Dramatic increases in NWP model skill will be required over the coming decades. This paper describes the role of major deployments of accurate profiling sensors in achieving that goal and presents an example from the Second Wind Forecast Improvement Program (WFIP2). Wind-profile data from scanning Doppler lidars were used to evaluate two versions of HRRR, the original and an updated version, and nested versions of each. This study focuses on the ability of updated HRRR versions to improve upon predicting a regional sea-breeze flow, which was found to generate large errors by the original HRRR. Updates to the full-scale HRRR versions produced mixed results, but the finer-mesh versions consistently reduced model errors.

Meteorology & Atmospheric Sciences↗

Doppler-Lidar Evaluation of HRRR-Model Skill at Simulating Summertime Wind Regimes in the Columbia River Basin during WFIP2

Complex-terrain locations often have repeatable near-surface wind patterns, such as synoptic gap flows and local thermally forced flows. An example is the Columbia River Valley in east-central Oregon-Washington, a significant wind-energy-generation region and the site of the Second Wind-Forecast Improvement Project (WFIP2). Data from three Doppler lidars deployed during WFIP2 define and characterize summertime wind regimes and their large-scale contexts, and provide insight into NWP model errors by examining differences in the ability of a model [NOAA’s High-Resolution Rapid-Refresh (HRRR-version1)] to forecast wind-speed profiles for different regimes. Seven regimes were identified based on daily time series of the lidar-measured rotor-layer winds, which then suggested two broad categories. First, in three regimes the primary dynamic forcing was the large-scale pressure gradient. Second, in two regimes the dominant forcing was the diurnal heating-cooling cycle (regional sea-breeze-type dynamics), including the marine intrusion previously described, which generates strong nocturnal winds over the region. The other two included a hybrid regime and a non-conforming regime. For the large-scale pressure-gradient regimes, HRRR had wind-speed biases of ~1 m s -1 and RMSEs of 2-3 m s -1 . Errors were much larger for the thermally forced regimes, owing to the premature demise of the strong nocturnal flow in HRRR. Thus, the more dominant the role of surface heating in generating the flow, the larger the errors. Major errors could result from surface heating of the atmosphere, boundary-layer responses to that heating, and associated terrain interactions. Finally, measurement/modeling research programs should be aimed at determining which modeled processes produce the largest errors, so those processes can be improved and errors reduced.

17 WIND ENERGY↗

Optimization of a Mixed Fleet of Aerial Drones for Medical Supplies: A Case Study of Blood Delivery Logistics

Aerial drones have emerged as an innovative solution for faster transportation of time-sensitive items (e.g., emergency medical supplies), potentially reducing the transmission of contagious diseases and enhancing healthcare availability through contactless autonomous delivery. We study fleet sizing and efficient scheduling of a mixed fleet of drones for delivering time-sensitive medical items having distinct release and due times to minimize the required fleet size and fleet composition, the required number of additional batteries, and the total energy consumption. We continuously track the remaining battery energy of drones to determine the optimal timing for battery replacement, rather than replacing the battery at each node. Using actual drone flight test data, we employed a machine learning (ML) method to estimate the energy consumption of different drone types during flight segments for different operating parameters. We present a novel mixed-integer programming model to efficiently formulate the problem that integrates the estimated energy consumption functions from ML. We propose a new greedy heuristic (GH) algorithm and a customized genetic algorithm (GA) for solving large-scale instances of this problem faster. Results demonstrate that the GH algorithm is substantially faster than the accelerated CPLEX and the GA, while sacrificing the solution quality by a small amount. Results based on an actual blood sample delivery case study from Pendleton, Oregon, United States, show that using a mixed fleet of drones reduces the total cost and total energy consumption up to 18.18% and 28.7%, respectively, compared to using a homogeneous fleet.

29 - ENERGY PLANNING, POLICY AND ECONOMY↗

Characterizing post-fire delayed tree mortality with remote sensing: sizing up the elephant in the room

Abstract Background Despite recent advances in understanding the drivers of tree-level delayed mortality, we lack a method for mapping delayed mortality at landscape and regional scales. Consequently, the extent, magnitude, and effects of delayed mortality on post-fire landscape patterns of burn severity are unknown. We introduce a remote sensing approach for mapping delayed mortality based on post-fire decline in the normalized burn ratio (NBR). NBR decline is defined as the change in NBR between the first post-fire measurement and the minimum NBR value up to 5 years post-fire for each pixel. We validate the method with high-resolution aerial photography from six wildfires in California, Oregon, and Washington, USA, and then compare the extent, magnitude, and effects of delayed mortality on landscape patterns of burn severity among fires and forest types. Results NBR decline was significantly correlated with post-fire canopy mortality (r 2 = 0.50) and predicted the presence of delayed mortality with 83% accuracy based on a threshold of 105 NBR decline. Plots with NBR decline greater than 105 were 23 times more likely to experience delayed mortality than those below the threshold (p < 0.001). Delayed mortality occurred across 6–38% of fire perimeters not affected by stand-replacing fire, generally affecting more areas in cold (22–41%) and wet (30%) forest types than in dry (1.7–19%) types. The total area initially mapped as unburned/very low-severity declined an average of 38.1% and generally persisted in smaller, more fragmented patches when considering delayed mortality. The total area initially mapped as high-severity increased an average of 16.2% and shifted towards larger, more contiguous patches. Conclusions Differences between 1- and 5-year post-fire burn severity maps depict dynamic post-fire mosaics resulting from delayed mortality, with variability among fires reflecting a range of potential drivers. We demonstrate that tree-level delayed mortality scales up to alter higher-level landscape patterns of burn severity with important implications for forest resilience and a range of fire-driven ecological outcomes. Our method can complement existing tree-level studies on drivers of delayed mortality, refine mapping of fire refugia, inform estimates of habitat and carbon losses, and provide a more comprehensive assessment of landscape and regional scale fire effects and trends.

Environmental Sciences & Ecology↗

Phylogenetic diversity of 200+ isolates of the ectomycorrhizal fungus Cenococcum geophilum associated with Populus trichocarpa soils in the Pacific Northwest, USA and comparison to globally distributed representatives

The ectomycorrhizal fungal symbiont Cenococcum geophilum is of high interest as it is globally distributed, associates with many plant species, and has resistance to multiple environmental stressors. C . geophilum is only known from asexual states but is often considered a cryptic species complex, since extreme phylogenetic divergence is often observed within nearly morphologically identical strains. Alternatively, C . geophilum may represent a highly diverse single species, which would suggest cryptic but frequent recombination. Here we describe a new isolate collection of 229 C . geophilum isolates from soils under Populus trichocarpa at 123 collection sites spanning a ~283 mile north-south transect in Western Washington and Oregon, USA (PNW). To further understanding of the phylogenetic relationships within C . geophilum , we performed maximum likelihood and Bayesian phylogenetic analyses to assess divergence within the PNW isolate collection, as well as a global phylogenetic analysis of 789 isolates with publicly available data from the United States, Japan, and European countries. Phylogenetic analyses of the PNW isolates revealed three distinct phylogenetic groups, with 15 clades that strongly resolved at >80% bootstrap support based on a GAPDH phylogeny and one clade segregating strongly in two principle component analyses. The abundance and representation of PNW isolate clades varied greatly across the North-South range, including a monophyletic group of isolates that spanned nearly the entire gradient at ~250 miles. A direct comparison between the GAPDH and ITS rRNA gene region phylogenies, combined with additional analyses revealed stark incongruence between the ITS and GAPDH gene regions, consistent with intra-species recombination between PNW isolates. In the global isolate collection phylogeny, 34 clades were strongly resolved using Maximum Likelihood and Bayesian approaches (at >80% MLBS and >0.90 BPP respectively), with some clades having intra- and intercontinental distributions. Together these data are highly suggestive of divergence within multiple cryptic species, however additional analyses such as higher resolution genotype-by-sequencing approaches are needed to distinguish potential species boundaries and the mode and tempo of recombination patterns.

59 BASIC BIOLOGICAL SCIENCES↗

Characterizing juvenile salmon predation risk during early marine residence

Predation mortality can influence the distribution and abundance of fish populations. While predation is often assessed using direct observations of prey consumption, potential predation can be predicted from co-occurring predator and prey densities under varying environmental conditions. Juvenile Pacific salmon Oncorhynchus spp. (i.e., smolts) from the Columbia River Basin experience elevated mortality during the transition from estuarine to ocean habitat, but a thorough understanding of the role of predation remains incomplete. We used a Holling type II functional response to estimate smolt predation risk based on observations of piscivorous seabirds (sooty shearwater [ Ardenna griseus ] and common murre [ Uria aalge ]) and local densities of alternative prey fish including northern anchovy ( Engraulis mordax ) in Oregon and Washington coastal waters during May and June 2010–2012. We evaluated predation risk relative to the availability of alternative prey and physical factors including turbidity and Columbia River plume area, and compared risk to returns of adult salmon. Seabirds and smolts consistently co-occurred at sampling stations throughout most of the study area (mean = 0.79 ± 0.41, SD), indicating that juvenile salmon are regularly exposed to avian predators during early marine residence. Predation risk for juvenile coho ( Oncorhynchus kisutch ), yearling Chinook salmon ( O . tshawytscha ), and subyearling Chinook salmon was on average 70% lower when alternative prey were present. Predation risk was greater in turbid waters, and decreased as water clarity increased. Juvenile coho and yearling Chinook salmon predation risk was lower when river plume surface areas were greater than 15,000 km 2 , while the opposite was estimated for subyearling Chinook salmon. These results suggest that plume area, turbidity, and forage fish abundance near the mouth of the Columbia River, all of which are influenced by river discharge, are useful indicators of potential juvenile salmon mortality that could inform salmonid management.

Phillips, Elizabeth M. (ORCID:0000000327752563)↗

Forest carbon sequestration on the west coast, USA: Role of species, productivity, and stockability

Forest ecosystems store large amounts of carbon and can be important sources, or sinks, of the atmospheric carbon dioxide that is contributing to global warming. Understanding the carbon storage potential of different forests and their response to management and disturbance events are fundamental to developing policies and scenarios to partially offset greenhouse gas emissions. Projections of live tree carbon accumulation are handled differently in different models, with inconsistent results. We developed growth-and-yield style models to predict stand-level live tree carbon density as a function of stand age in all vegetation types of the coastal Pacific region, US (California, Oregon, and Washington), from 7,523 national forest inventory plots. We incorporated site productivity and stockability within the Chapman-Richards equation and tested whether intensively managed private forests behaved differently from less managed public forests. We found that the best models incorporated stockability in the equation term controlling stand carrying capacity, and site productivity in the equation terms controlling the growth rate and shape of the curve. RMSEs ranged from 10 to 137 Mg C/ha for different vegetation types. There was not a significant effect of ownership over the standard industrial rotation length (~50 yrs) for the productive Douglas-fir/western hemlock zone, indicating that differences in stockability and productivity captured much of the variation attributed to management intensity. Our models suggest that doubling the rotation length on these intensively managed lands from 35 to 70 years would result in 2.35 times more live tree carbon stored on the landscape. These findings are at odds with some studies that have projected higher carbon densities with stand age for the same vegetation types, and have not found an increase in yields (on an annual basis) with longer rotations. We suspect that differences are primarily due to the application of yield curves developed from fully-stocked, undisturbed, single-species, “normal” stands without accounting for the substantial proportion of forests that don’t meet those assumptions. The carbon accumulation curves developed here can be applied directly in growth-and-yield style projection models, and used to validate the predictions of ecophysiological, cohort, or single-tree style models being used to project carbon futures for forests in the region. Our approach may prove useful for developing robust models in other forest types.

Chisholm, Paul J. (ORCID:0000000238784707)↗

DEEPEN Global Standardized Categorical Exploration Datasets for Magmatic Plays

DEEPEN stands for DE-risking Exploration of geothermal Plays in magmatic ENvironments. As part of the development of the DEEPEN 3D play fairway analysis (PFA) methodology for magmatic plays (conventional hydrothermal, superhot EGS, and supercritical), weights needed to be developed for use in the weighted sum of the different favorability index models produced from geoscientific exploration datasets. This was done using two different approaches: one based on expert opinions, and one based on statistical learning. This GDR submission includes the datasets used to produce the statistical learning-based weights. While expert opinions allow us to include more nuanced information in the weights, expert opinions are subject to human bias. Data-centric or statistical approaches help to overcome these potential human biases by focusing on and drawing conclusions from the data alone. The drawback is that, to apply these types of approaches, a dataset is needed. Therefore, we attempted to build comprehensive standardized datasets mapping anomalies in each exploration dataset to each component of each play. This data was gathered through a literature review focused on magmatic hydrothermal plays along with well-characterized areas where superhot or supercritical conditions are thought to exist. Datasets were assembled for all three play types, but the hydrothermal dataset is the least complete due to its relatively low priority. For each known or assumed resource, the dataset states what anomaly in each exploration dataset is associated with each component of the system. The data is only a semi-quantitative, where values are either high, medium, or low, relative to background levels. In addition, the dataset has significant gaps, as not every possible exploration dataset has been collected and analyzed at every known or suspected geothermal resource area, in the context of all possible play types. The following training sites were used to assemble this dataset: - Conventional magmatic hydrothermal: Akutan (from AK PFA), Oregon Cascades PFA, Glass Buttes OR, Mauna Kea (from HI PFA), Lanai (from HI PFA), Mt St Helens Shear Zone (from WA PFA), Wind River Valley (From WA PFA), Mount Baker (from WA PFA). - Superhot EGS: Newberry (EGS demonstration project), Coso (EGS demonstration project), Geysers (EGS demonstration project), Eastern Snake River Plain (EGS demonstration project), Utah FORGE, Larderello, Kakkonda, Taupo Volcanic Zone, Acoculco, Krafla. - Supercritical: Coso, Geysers, Salton Sea, Larderello, Los Humeros, Taupo Volcanic Zone, Krafla, Reyjanes, Hengill. **Disclaimer: Treat the supercritical fluid anomalies with skepticism. They are based on assumptions due to the general lack of confirmed supercritical fluid encounters and samples at the sites included in this dataset, at the time of assembling the dataset. The main assumption was that the supercritical fluid in a given geothermal system has shared properties with the hydrothermal fluid, which may not be the case in reality. Once the datasets were assembled, principal component analysis (PCA) was applied to each. PCA is an unsupervised statistical learning technique, meaning that labels are not required on the data, that summarized the directions of variance in the data. This approach was chosen because our labels are not certain, i.e., we do not know with 100% confidence that superhot resources exist at all the assumed positive areas. We also do not have data for any known non-geothermal areas, meaning that it would be challenging to apply a supervised learning technique. In order to generate weights from the PCA, an analysis of the PCA loading values was conducted. PCA loading values represent how much a feature is contributing to each principal component, and therefore the overall variance in the data.

15 GEOTHERMAL ENERGY↗

Hawaii Wave Surge Energy Converter (HAWSEC) OSU O.H. Hinsdale Basin

The following information and metadata applies to both the Phase I (Hydrodynamics) and Phase II (Full System Power Take-Off) zip folders which contain testing data from the OSU (Oregon State University) O.H. Hinsdale Wave Research Laboratory, from both OSU and the University of Hawaii at Manoa (UH). See zip folders provided further below in the downloads section. For experimental data of the full system, including PTO, see Phase II dataset. There are two main directories in each Phases's zip folder: "OSU_data" and "UH_data". The "OSU_data" directory contains data collected from their DAQ (data acquisition system), which includes all wave gauge observations, as well as body motions derived from their Qualisys motion tracking system. The organization of the directory follows OSU's convention. Detailed information on the instrument setup can be found under "OSU_data/docs/setup/instm_locations". The experiments conducted are documented in the "OSU_data/docs/daq_logs", which provides the trial number to the corresponding data located under "OSU_data/data" in several formats (e.g., ".mat" and ".txt"). Inside the trial directory, data is provided for each of the instruments defined in "OSU_data/docs/setup/instm_locations". The "UH_data" directory contains data collected from their DAQ. The data is stored in a ".tdms" file format. There are free plug-ins for Microsoft Excel and MathWorks MATLAB to read the ".tdms" format. Below are a few links providing methods to read in the data, but a Google search should identify alternatives sources if these no longer exist (valid as of January 2024): Excel: http://www.ni.com/example/27944/en/ MATLAB: https://www.mathworks.com/matlabcentral/fileexchange/30023-tdms-reader The Excel plugin is recommend to get a quick overview of the data. The UH data is organized by directory name, in which the sub-directories for each experiment contains a directory whose name defines the wave height and period for the experimental data within. For example, a directory name "H02_T0275" corresponds to an experiment with wave height 0.1m and a period of 2.75s. For random wave data, the gamma value is also included in the directory name. For example, a directory name "H02_T0225_G18" corresponds to an experiment with a significant wave height of 0.2m, a peak period of 2.25s, and a gamma value of 1.8, with each spectra being a TMA spectrum. For the free decay experiments, the directory name is defined by the initial angular displacement. For example, a directory name "ang05_run01" corresponds to an experiment with an initial angular displacement of 5 degrees. There is a dataset in the UH data for each corresponding experiment defined in the OSU DAQ logs. The ".tdms" data is output from the DAQ at fixed intervals. Therefore, if multiple files are contained within the folder, the data will need to be stitched together. Within the UH dataset, there are two input channels from the OSU DAQ providing a random square wave signal for time synchronization ("ENV-WHT-0010") and a high/low signal ("ENV-WHT-0012") to identify when the wave maker is active (+5V). The UH data is logged as a collection of channel outputs. Channels not in use for the OSU testing (either Phase I or Phase II) are marked "nan" below. If the sensor is disconnected, it will record noise throughout the experiment. Below are the channel definitions in terms of what they measure: GPS Time = time CYL-POS-0001 = position between flap and fixed reference CYL-LCA-0001 = force between flap and hydraulic cylinder REC-LPT-0001 = nan REC-HPT-0001 = nan REC-HPT-0002 = nan REC-HPT-0003 = nan HHT-HPT-0001 = pressure at exhaust ("head" only) REC-FQC-0001 = nan REC-FQC-0002 = nan HHT-FQC-0001 = flow at exhaust ("head" only) ENV-WHT-0001 = nan ENV-WHT-0002 = nan ENV-WHT-0003 = nan ENV-WHT-0010 = random signal from OSU DAQ ENV-WHT-0012 = high/low signal from OSU DAQ Also included is a calibration curve to convert the string pot data to flap pi...

16 TIDAL AND WAVE POWER↗

TEAMER: Drifting Hydrophone System - Block Diagram and Pre-Amplifier Calibrations

This data release is part of TEAMER RFTS 2, where the Cooperative Institute for Marine Resources Studies (CIMRS) at Oregon State University is performing hardware and software development and integration of four newly designed drifting hydrophone systems for underwater noise measurements at marine renewable energy projects. These new acoustic systems will provide advanced technology support available for use at both tidal and wave energy deployments adding additional resources to the limited amount of drifting hydrophone technologies that are available to the marine energy community The submission data contains 2 files plus a Post-Access Report: One file containing the calibration parameters for four (4) pre-amplifier boards used on the drifting hydrophones. Pre-amplifier boards are model "WB PREAMP REV6," with serial numbers; 28, 29, 30, and 31. One image file illustrating the system functional block diagram of the drifting hydrophone system configuration.

16 TIDAL AND WAVE POWER↗

1994 Portland Household Travel Survey

The Portland Household Travel Survey was conducted under the auspices of the Oregon Department of Transportation to provide information suitable for gaining an in-depth understanding of the activity and travel behavior of both households in metropolitan areas and the individuals within those households. The first round of surveys collected household activity data from the four metropolitan planning organization areas around Portland, plus three extra counties (Marion, Polk, and Yamhill). A total of 11,762 households participated in this study. Each household member was asked to record any activity that lasted 30 minutes or longer, or any activity that required travel for the specified 48-hour period.

1Hz data↗

FTICR-MS, Sensor, and Environmental Data from 5 Streams Impacted by the 2020 Holiday Farm Fire Associated with: "Spatiotemporal controls on the delivery of dissolved organic matter to streams following a wildfire"

This data package is associated with the publication "Spatiotemporal Controls on the Delivery of Dissolved Organic Matter to Streams Following a Wildfire" submitted to Geophysical Research Letters (Roebuck et al., 2022). The study aims to understand storm induced transport of pyrogenic materials to streams impacted by varying degrees of burn severity. Time series samples (24 samples in 1-hour intervals) were collected at 5 sites within the McKenzie River Watershed (Oregon, USA) whose catchment were each completely engulfed by the 2020 Holiday Farm Fire. The samples were collected in November 2020 during the first major storm pulse following the conclusion of the wildfire. Samples were characterized for dissolved organic carbon, total dissolved nitrogen, and by ultra-high resolution mass spectrometry. In situ turbidity data also collected.This data package contains 4 primary folders that include the following: 1) Metadata, 2) EnvData (Environmental Data), 3) SensorData, and 4) FTICR_SupportingData. The package contains a single file-level metadata (flmd) file. Each primary folder also contains individual data dictionaries (dd) to define and provide descriptors of column/row headers and data flags. The FTICR_SupportingData, folder 4, contains raw, unprocessed FTICR-MS Data files in addition to a csv containing processed FTICR-MS data. This package contains the following file types: csv, xml, pdf.

54 ENVIRONMENTAL SCIENCES↗

WHONDRS Surface Water Dissolved Organic Carbon and FTICR-MS across Stream Orders in Four United States Watersheds in 2019 and 2020 (v3)

This dataset supports the broader Watershed Rules of Life (WROL) study examining spatial and temporal biogeochemical and microbial relationships across the Connecticut River, Deschutes River, Gunnison River, and Willamette River watersheds. The dataset provides geochemistry and organic matter characterization data generated from surface water collected from August 2019 to September 2020 in Colorado, Connecticut, Oregon, and Vermont. Related data were collected and will be published separately in collaboration with P. A. Raymond and B. C. Crump as part of WROL. This dataset is comprised of one folder containing (1) file-level metadata; (2) data dictionary; (3) readme; (4) field metadata; (5) dissolved organic carbon (DOC, measured as non-purgeable organic carbon, NPOC) data and averages; (6) surface water sampling protocol; (7) methods codes; (8) international generic sample number (IGSN) mapping file; and (9) a folder of high resolution characterization of organic matter via 12 Tesla Fourier transform ion cyclotron resonance mass spectrometry (FTICR-MS) generated through the Environmental Molecular Sciences Laboratory (EMSL; https://www.pnnl.gov/environmental-molecular-sciences-laboratory). The FTICR folder contains two subfolders, one containing the .xml data files and the other containing instructions for using Formularity (https://omics.pnl.gov/software/formularity) and an R script to process the data based on the user's specific needs. All files are .csv, .pdf, .R, .ref, or .xml. This data package was originally published in October 2022. It was updated in July 2023 (v2) and again in November 2023 (v3). See the change history section in the readme for more details.

54 ENVIRONMENTAL SCIENCES↗

WHONDRS River Corridor Sediment and Water Geochemistry and In Situ Sensor Data from Machine-Learning-Informed Sites across the Contiguous United States (v6)

This dataset supports a broader study examining hyporheic zone respiration rates to improve predictive models at a contiguous United States (CONUS) scale. The CONUS-Scale Model-Sample Study (CM) was designed following ICON (integrated, coordinated, open, and networked) principles to facilitate a model-experiment (ModEx) iteration approach, leveraging crowdsourced sampling across the CONUS. New machine learning models were created every month to guide sampling locations. Data from the resulting samples were used to test and rebuild the machine learning models for the next round of sampling guidance. Sampling began in April 2022 and ended in October 2023. In addition to the widely distributed CONUS sites, a more spatially focused sampling occurred in the Yakima River Basin, WA in summer 2022. Data from this more spatially intensive sampling occurred under the label “Second Spatial Study (SSS)” and were also included in the machine learning models. Other data types collected from SSS that were not part of CM were published in a separate data package (https://data.ess-dive.lbl.gov/view/doi:10.15485/1969566). This data package was originally published in February 2023. It was updated in June 2023 (v2; new and modified files); December 2023 (v3; new and modified files); June 2024 (v4; new and modified files); April 2024 (v5; new and modified files); and September 2025 (v6; modified files). See the change history section in the readme for more details. For details on how to navigate data packages generated by this project, see https://data.ess-dive.lbl.gov/portals/PNNLRiverCorridorSFA/About. This dataset is comprised of two folders of field photos and videos, one folder of raw Fourier transform ion cyclotron resonance mass spectrometry (FTICR-MS) data and one main data folder containing (1) file-level metadata; (2) data dictionary; (3) field metadata; (4) readme; (5) international generic sample number (IGSN) mapping file; (6) field protocols; (7) a subfolder with sample data; and (8) a subfolder with sensor data. The sample data subfolder contains (1) surface water and sediment dissolved organic carbon (DOC, measured as non-purgeable organic carbon, NPOC) data and averages; (2) surface water and sediment total nitrogen data and averages; (3) surface water major cations and anions and averages; (4) sediment grain size data; (5) sediment iron (II) data and averages; (6) wet sediment mass, dry sediment mass, water mass, and wet sediment volume in incubation and sediment ICR vials; (7) sediment incubation respiration rate data and averages; (8) normalized respiration rate data and averages; (9) methods codes; (10) sediment specific surface area; (11) sediment percent carbon and nitrogen; (12) sediment gravimetric moisture and averages; (15) sediment X-ray diffraction (XRD) data; (16) sediment adenosine triphosphate (ATP) and averages; (17) a subfolder with sediment incubation respiration data, scripts, and plots; (18) surface water and sediment FTICR methods; and (19) a subfolder of 9.4 Tesla (9.4T) FTICR-MS data. This folder contains five subfolders, one containing the sediment .xml data files, one containing the water .xml files, one containing the sediment CoreMS output files, one containing the water CoreMS output files, and the other containing instructions and scripts for processing the files in CoreMS (https://github.com/EMSL-Computing/CoreMS).The sensor data subfolder contains (1) a subfolder with miniDOT dissolved oxygen and temperature data and plots; (2) miniDOT dissolved oxygen and temperature summary data; and (3) miniDOT installation methods. All files are .csv, .pdf, .R, .xml, .d, .html, .Rmd, .py, .cal, .json, .jpg, .jpeg, .png, .mov, or .mp4. CORRECTION: Carbon and nitrogen content are reported as percentages. The current column headers "01395_C_percent_per_mg" and "01397_N_percent_per_mg" are incorrect. These should read "01395_C_percent" and "01397_N_percent" and will be corrected in the next version of this data package. We thank the United States Forest Service, Washington Department of Fish and Wildlife, Washington Department of Natural Resources, Cowiche Canyon Conservatory, Washington State Parks and Recreation Commission (Scientific Research Permit #210901), and the Confederated Tribes and Bands of the Yakama Nation for access to field locations where the samples labeled “SSS” were collected. We also thank the Yakama Nation Tribal Council and Yakama Nation Fisheries for working with us to facilitate sample collection and optimization of data usage according to their values and worldview. WHONDRS consortium members were asked to provide any acknowledgments for the collection of samples labeled “CM” and the following is a list of acknowledgments that were submitted with their corresponding Site IDs: (MART) Research activities were conducted in part on the Wind River Experimental Forest within the Gifford Pinchot National Forest; (MP- 100379) Philadelphia is part of Lenapehoking, the ancestral homelands of the Lenape peoples; (MP-102398) Land surveyed is the ancestral homelands of the Nookhose'iinenno (Arapaho), Tsis tsis'tas (Cheyenne), and Nuuchu (Ute); (MP-100749 and MP- 100747) Georgia Coastal Ecosystem LTER, OCE-1832178; (SP-70 and SP-72) Eastern Shoshone, Shoshone-Bannock; (MP- 102944) Funded by Oregon Watershed Enhancement Board. On the traditional lands of the Confederated Tribes of the Siletz, Confederated Tribes of the Grand Rhonde, and the Clatsop-Nehalem Confederated Tribe; (MP- 100607) Holiday Creek is located on the traditional territory of the Monacan Indian Nation; (SP-45) Lafayette Blue Springs State Park; (MP-102420) NSF DEB-2016749; (MP-100019) New Hampshire Agriculture Experiment Station; (SP-35) Rayonier (land owner; https://www.rayonier.com/); (MP- 101276) US Department of Energy, Office of Science, Biological and Environmental Research, Subsurface Biogeochemical Research, Watershed Dynamics and Evolution SFA at ORNL; (MP- 103224) Watershed Dynamics and Evolution SFA at ORNL; (MP- 101584) Traditional lands of the Oceti Sakowin (Dakota, Lakota, Nakoda) and Anishinaabe Peoples.

54 ENVIRONMENTAL SCIENCES↗

Meteorological and hydrological parameters for 17 locations of meteorological stations of the East River Watershed

The Data Package includes a set of csv files of input meteorological parameters for locations of 17 meteorological stations within the East River watershed. These meteorological datasets were downloaded from the (1) PRISM database--monthly precipitation, air temperature (minimum, mean, and maximum), vapor pressure deficit (minimum and maximum), and dewpoint temperature, and (2) NCEP/NCAR Reanalysis database – wind database. The datasets were used for calculations of the Potential Evapotranspiration (ETo), Actual Evapotranspiration (ET), Standard Precipitation Index (SPI) , Standard Evapotranspiration-Precipitation Index (SPEI) for the period from 1966 to 2021. The main research questions addressed are: the evaluation of the long-term temporal trends of climatic parameters, hierarchical clustering, and areal mapping/zonation of the East River watershed. Calculations were conducted in the Rstudio environment. The dataset additionally includes a file-level metadata (flmd.csv) file that lists each file contained in the dataset with associated metadata; and a data dictionary (dd.csv) file that contains column/row headers used throughout the files along with a definition, units, and data type.The input datasets were downloaded from (a) PRISM database (the Northwest Alliance for Computational Science and Engineering at the Oregon State University), and (b) NCEP/NCAR Reanalysis database.

54 ENVIRONMENTAL SCIENCES↗

Organic Matter Concentration and Composition in November 2021 and April 2022 from 12 Streams Impacted by the 2020 Holiday Farm Fire (v2)

This dataset represents results from a field study aiming to understand storm induced transport of pyrogenic materials to streams impacted by varying degrees of burn severity. Time series samples were collected at 5 sites within the McKenzie River Watershed (Oregon, USA) whose catchment were each completely engulfed by the 2020 Holiday Farm Fire. An additional 7 sites were sampled once during the storm. The samples were collected during storm events in November 2020, January 2021, November 2021, and April 2022. Samples were characterized for benezenepolycarboxylic acids (BPCA), ultra-high resolution mass spectrometry, dissolved organic carbon and optics (absorbance and fluorescence). Fourier-transform ion cyclotron resonance mass spectrometry (FTICR) and dissolved organic carbon data from the November 2020 (referred to as “EWEB_2020”) sampling can be found in a separate data package (doi: 10.15485/1869708). NOTE: The 2020 samples were run on FTICR-MS in two unique instances. The first run can be found in the previous data package (EWEB_2020). The second run is included in this data package. These samples were run for a second time so that the data were more directly interoperable with the other samples in this data package. We have not done any investigation into the differences/similarities between these datasets and the previously ran/published data in the other data package. This data package was originally published in November 2024. It was updated in April 2025 (v2; new and modified files). See the change history section below for more details. For details on how to navigate data packages generated by this project, see https://data.ess-dive.lbl.gov/portals/PNNLRiverCorridorSFA/About. This dataset contains (1) file-level metadata; (2) data dictionary; (3) data package readme; (4) metadata; (5) methods information; (6) dissolved organic carbon (DOC, measured as non-purgeable organic carbon, NPOC) data; (7) excitation emission matrix (EEM) methods; and (8) a sub-folder with processed EEM data (9) benzene polycarboxylic acid (BPCA) concentration data; (10) Fourier transform ion cyclotron resonance mass spectrometry (FTICR-MS) methods; and (11) folder of high-resolution characterization of organic matter via 12 Tesla FTICR-MS generated through the Environmental Molecular Sciences Laboratory (EMSL; https://www.pnnl.gov/environmental-molecular-sciences-laboratory). The EEMs sub-folder contains two additional folders; the Absorbance and Fluorescence folders which contain the processed EEMs absorbance and fluorescence data respectively. This package contains the following file types: csv, xml, pdf.

54 ENVIRONMENTAL SCIENCES↗

Machine learning model inputs, outputs, and scripts associated with “Artificial intelligence-guided iterations between observations and modeling significantly improve environmental predictions”

NOTE: The manuscript associated with this data package is currently in review. The data may be revised based on reviewer feedback. Upon manuscript acceptance, this data package will be updated with the final dataset and additional metadata. This data package is associated with the manuscript “Artificial intelligence-guided iterations between observations and modeling significantly improve environmental predictions” (Malhotra et al., in prep). This effort was designed following ICON (integrated, coordinated, open, and networked) principles to facilitate a model-experiment (ModEx) iteration approach, leveraging crowdsourced sampling across the contiguous United States (CONUS). New machine learning models were created every month to guide sampling locations. Data from the resulting samples were used to test and rebuild the machine learning models for the next round of sampling guidance. Associated sediment and water geochemistry and in situ sensor data can be found at https://data.ess-dive.lbl.gov/datasets/doi:10.15485/1923689, https://data.ess-dive.lbl.gov/datasets/doi:10.15485/1729719, and https://data.ess-dive.lbl.gov/datasets/doi:10.15485/1603775. This data package is associated with two GitHub repositories found at https://github.com/parallelworks/dynamic-learning-rivers and https://github.com/WHONDRS-Hub/ICON-ModEx_Open_Manuscript. In addition to this readme, this data package also includes two file-level metadata (FLMD) files that describes each file and two data dictionaries (DD) that describe all column/row headers and variable definitions. This data package consists of two main folders (1) dynamic-learning-rivers and (2) ICON-ModEx_Open_Manuscript which contain snapshots of the associated GitHub repositories. The input data, output data, and machine learning models used to guide sampling locations are within dynamic-learning-rivers. The folder is organized into five top-level directories: (1) “input_data” holds the training data for the ML models; (2) “ml_models” holds machine learning (ML) models trained on the data in “input_data”; (3) “examples” contains files for direct experimentation with the machine learning model, including scripts for setting up “hindcast” run; (4) “scripts” contains data preprocessing and postprocessing scripts and intermediate results specific to this data set that bookend the ML workflow; and (5) “output_data” holds the overall results of the ML model on that branch. Each trained ML model resides on its own branch in the repository; this means that inputs and outputs can be different branch-to-branch. There is also one hidden directory “.github/workflows”. This hidden directory contains information for how to run the ML workflow as an end-to-end automated GitHub Action but it is not needed for reusing the ML models archived here. Please see the top-level README.md in the GitHub repository for more details on the automation. The scripts and data used to create figures in the manuscript are within ICON-ModEx_Open_Manuscript. The folder is organized into four folders which contain the scripts, data, and pdf for each figure. Within the “fig-model-score-evolution” folder, there is a folder called “intermediate_branch_data” which contains some intermediate files pulled from dynamic-learning-rivers and reorganized to easily integrate into the workflows. NOTE: THIS FOLDER INCLUDES THE FILES AT THE POINT OF PAPER SUBMISSION. IT WILL BE UPDATED ONCE THE PAPER IS ACCEPTED WITH ANY REVISIONS AND WILL INCLUDE A DD/FLMD AT THAT POINT. We thank the United States Forest Service, Washington Department of Fish and Wildlife, Washington Department of Natural Resources, Cowiche Canyon Conservatory, Washington State Parks and Recreation Commission (Scientific Research Permit #210901), and the Confederated Tribes and Bands of the Yakama Nation for access to field locations where the samples labeled “SSS” were collected. We also thank the Yakama Nation Tribal Council and Yakama Nation Fisheries for working with us to facilitate sample collection and optimization of data usage according to their values and worldview. WHONDRS consortium members were asked to provide any acknowledgments for the collection of samples labeled “CM” and the following is a list of acknowledgments that were submitted with their corresponding Site IDs: (MART) Research activities were conducted in part on the Wind River Experimental Forest within the Gifford Pinchot National Forest; (MP- 100379) Philadelphia is part of Lenapehoking, the ancestral homelands of the Lenape peoples; (MP-102398) Land surveyed is the ancestral homelands of the Nookhose'iinenno (Arapaho), Tsis tsis'tas (Cheyenne), and Nuuchu (Ute); (MP-100749 and MP- 100747) Georgia Coastal Ecosystem LTER, OCE-1832178; (SP-70 and SP-72) Eastern Shoshone, Shoshone-Bannock; (MP- 102944) Funded by Oregon Watershed Enhancement Board. On the traditional lands of the Confederated Tribes of the Siletz, Confederated Tribes of the Grand Rhonde, and the Clatsop-Nehalem Confederated Tribe; (MP- 100607) Holiday Creek is located on the traditional territory of the Monacan Indian Nation; (SP-45) Lafayette Blue Springs State Park; (MP-102420) NSF DEB-2016749; (MP-100019) New Hampshire Agriculture Experiment Station; (SP-35) Rayonier (land owner; https://www.rayonier.com/); (MP- 101276) US Department of Energy, Office of Science, Biological and Environmental Research, Subsurface Biogeochemical Research, Watershed Dynamics and Evolution SFA at ORNL; (MP- 103224) Watershed Dynamics and Evolution SFA at ORNL; (MP- 101584) Traditional lands of the Oceti Sakowin (Dakota, Lakota, Nakoda) and Anishinaabe Peoples.

54 ENVIRONMENTAL SCIENCES↗

AmeriFlux FLUXNET-1F US-Me6 Metolius Young Pine Burn

This is the AmeriFlux Management Project (AMP) created FLUXNET-1F version of the carbon flux data for the site US-Me6 Metolius Young Pine Burn. This is the FLUXNET version of the carbon flux data for the site US-Me6 Metolius Young Pine Burn produced by applying the standard ONEFlux (1F) software. Site Description - The study site is located east of the Cascade mountains, near Sisters, Central Oregon and is part of the Metolius cluster sites with different age and disturbance classes within the AmeriFlux network. After a severe fire in 1979, the site was salvage logged, was acquired by the US Forest Service land and re-forested in 1990. The dominant overstory vegetation are 20-year old ponderosa pine trees with an average height of 5.2 +/- 1.1 m. The season maximum overstory half-sided LAI was 0.6 m2 m-2 in 2010. Tree density is low, with ca. 162 trees ha-1.

Hanson, Chad↗