Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “STEP file”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

CrossLink: Meshing a 3D part from a STEP file [Slides]

This example shows how to use CrossLink to create a mesh for a 3D geometry part read in from a STEP file. A STEP file (Standard for the Exchange of Product data) is a common file format used for 3D modelling that can be written out from a CAD program (e.g., Creo Parametric). Being able to read and mesh STEP file geometries is essential for meshing parts from engineering models. Creating the mesh typically involves: (1) Importing the geometry; (2) Creating geometry groups; (3) Assigning curves and surfaces to the geometry groups; (4) Building a topology; (5) Applying geometric and mesh constraints; and (6) Generating the final mesh. Current issues with this process in CrossLink include: (1) The GUI does not display trimmed surfaces; (2) The user must manually create and assign geometry groups; and (3) The user must be aware of duplicate curves and surfaces from the CAD model.

97 MATHEMATICS AND COMPUTING↗

A universal method to compare parts from STEP files

Abstract Model Based Definition (MBD) captures the complete specification of a part in digital form and leverages (at least) the universal “Standard for the Exchange of Product” (STEP) file format. MBD has revolutionized manufacturing due to time and cost savings associated with containing all engineering data within a single digital source. This work presents a novel method to transform digital definitions in any given STEP file into a tensor-like structure that is unique for each part and can be used to regenerate the original STEP file completely. Resulting STEP tensors are amenable to part comparison based on various part specifications in a general and straightforward manner. Here, part similarity is evaluated among sets of parts according to specific geometry, material composition, and design intent. Importantly, specification similarity can be quantified using only the tensors’ structure. As such, this approach is not limited to families of geometric shapes, part types, or fabrication methods; nor does it require any prior knowledge about the parts being compared.

36 MATERIALS SCIENCE↗

Laboratory Upgrade Point Absorber (LUPA) CAD Files

The Laboratory Upgrade Point Absorber (LUPA) is an open-source wave energy converter designed and tested by Oregon State University. The computer-aided design (CAD) files are provided here in two forms: the original SOLIDWORKS (2021) model as "LUPA SOLIDWORKS.zip" and as a STEP file "LUPA-A1000.step". The bill of materials is provided as an Excel file with assemblies (LUPA-Axxx), part numbers (LUPA-Axxx-Pyyy), part descriptions, manufacturers, and manufacturer part numbers. This comprehensive CAD model represents LUPA as it was deployed in Fall 2022 testing at the O.H. Hinsdale Wave Research Laboratory. The mass properties including mass, center of gravity, and moments of inertia have been overridden for some parts and assemblies to match the physical device properties as determined from experiments. This appears as "overridden by user" when viewing mass properties in SOLIDWORKS. The LUPA-A1000.SLDASM file from the LUPA SOLIDWORKS.zip folder is the topmost assembly, open this file to see the entire model as one assembly. See "PMEC Page", "OpenEI Wiki Page", and the "Signature Project Page" resources below for more information on LUPA.

16 TIDAL AND WAVE POWER↗

Baltimore Social-Environmental Collaborative (BSEC) Doppler Lidar & Derived Products

This repository contains all processed Doppler‐lidar outputs from the PSU lidar deployed for the Baltimore Social‐Environmental Collaborative (BSEC) project. Vertical Stare Scans (fixed‐beam, vertical profiling): 1 Hz backscatter intensity (m⁻¹ sr⁻¹), signal‐to‐noise ratio (unitless), and Doppler vertical‐velocity (m s⁻¹) on ~30 m range gates, stored as CF-compliant NetCDF. Wind Profiles (horizontal‐wind retrieval): daily NetCDF outputs of retrieved horizontal wind speed (m s⁻¹) and direction (degrees), computed from the angled‐scan returns. Profile Statistics (summary statistics on the vertical velocity): 15 min windows (default) of mean, variance, skewness, kurtosis, high-frequency variance, etc., as a function of height; saved as CF-compliant NetCDF files. Boundary Layer Height (BLH) (fuzzy-logic output): 15 min BLH estimates (m), with lower/upper fuzzy bounds (m) and a quality flag (0–4) indicating data status (e.g., no data, good, below range, ran out of signal, cloud-topped). Cloud Base Height (Haar-gradient detection): 15 min estimates of cloud-base height (m) with a cloud-detection quality flag (0–3: none, low, moderate, high). All five product streams are organized by year and date under their own top-level folders (01_Vertical_Stare_Scans/ through 05_Cloud_Height/). Each folder contains a data_ /YYYY/ subdirectory with daily CF-compliant NetCDF outputs (96 windows per day at 15 min intervals). Global attributes in each file include creation history, version (2.0.0), institution, and source. Instrument & MeasurementsThe PSU Doppler Lidar samples aerosol backscatter (m⁻¹ sr⁻¹), signal-to-noise ratio, and radial velocity at ~1 Hz. Vertical stare scans point the beam straight up; after collecting angled scans through multiple elevation angles, the "Wind Profiles" product contains the fully retrieved horizontal wind speed and direction. Data were collected continuously at ~30 m range resolution, with a typical height ceiling of ~12 km. How to Use Open any NetCDF with Python's xarray, MATLAB, or similar CF-compliant tools. Stare scans and angled-scan retrievals (Wind Profiles) are CF-compliant daily NetCDF files. Profile-Statistics, BLH, and Cloud Height files are daily 15 min summaries (96 time steps per file). Inspect the included variables (e.g., vertical_velocity_variance, wind_speed, BLH, cloud_base_height) for your analyses. Use the quality flags (BLH_flag, cloud_flag) to filter out poor-quality retrievals. For more information or questions about processing methods, please contact:Nicholas E. Prince ⟨nec5299@psu.edu⟩Penn State Department of Meteorology & Atmospheric Science

Air Quality↗

Study of reference burnup steps optimization in fuel segment data file generation for NEXUS/ANC9 code system

For any two-step core design code system, the cross-section files are the primary factor to determine the accuracy of the system prediction. For a once-through cross-section system to cover all potential applicable conditions, the system may need to perform tens of thousands state points of lattice calculations. These calculations take a significant amount of CPU times and computing power, especially when a fine or ultra-fine energy group library is used. As computer power increases, so does the calculational complexity. Therefore, it is important to make the lattice calculations effective and efficient. Based on the Westinghouse core design code system NEXUS/ANC9, this work focuses on the reference burnup steps used by the lattice code when performing the calculations. Following the fundamental cross-section methodology and analyzing the contribution of each individual terms, this study provides an applicable solution of the optimized reference burnup steps. The test results show that, with the existing cross-section methodology, it is possible to significantly reduce the lattice calculation cases and cross-section file generation time without sacrificing the accuracy of system prediction. (authors)

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Baltimore Social-Environmental Collaborative (BSEC) Doppler Lidar & Derived Products

This repository contains all processed Doppler‐lidar outputs from the PSU lidar deployed for the Baltimore Social‐Environmental Collaborative (BSEC) project. Vertical Stare Scans (fixed‐beam, vertical profiling): 1 Hz backscatter intensity (m⁻¹ sr⁻¹), signal‐to‐noise ratio (unitless), and Doppler vertical‐velocity (m s⁻¹) on 30 m range gates (and 3 m range gates), stored as CF-compliant NetCDF. Wind Profiles (horizontal‐wind retrieval): daily NetCDF outputs of retrieved horizontal wind speed (m s⁻¹) and direction (degrees), computed from the angled‐scan returns. Profile Statistics (summary statistics on the vertical velocity): 15 min windows (default) of mean, variance, skewness, kurtosis, high-frequency variance, etc., as a function of height; saved as CF-compliant NetCDF files. Boundary Layer Height (BLH) (fuzzy-logic output): 15 min BLH estimates (m), with lower/upper fuzzy bounds (m) and a quality flag (0–4) indicating data status (e.g., no data, good, ran out of signal, below range, cloud-topped). Cloud Base Height (Haar-gradient detection): 10 min estimates of cloud-base height (m). All five product streams are organized by year and date under their own top-level folders (01_Vertical_Stare_Scans/ through 05_Cloud_Height/). Each folder contains a data_ /YYYY/ subdirectory with daily CF-compliant NetCDF outputs (96 windows per day at 15 min intervals). Global attributes in each file include creation history, version (3.0.0), institution, and source. Instrument & MeasurementsThe PSU Doppler Lidar samples aerosol backscatter (m⁻¹ sr⁻¹), signal-to-noise ratio, and radial velocity at ~1 Hz. Vertical stare scans point the beam straight up; after collecting angled scans through multiple elevation angles, the "Wind Profiles" product contains the fully retrieved horizontal wind speed and direction. Data were collected continuously at ~30 m range resolution (and 3 m for the year of 2025), with a typical height ceiling of ~12 km. How to Use Open any NetCDF with Python's xarray, MATLAB, or similar CF-compliant tools. Stare scans and angled-scan retrievals (Wind Profiles) are CF-compliant daily NetCDF files. Profile-Statistics, BLH, and Cloud Height files are daily 15 min (10 min for Cloud Heights) summaries (96 time steps per file). Inspect the included variables (e.g., vertical_velocity_variance, wind_speed, BLH, cloud_base_height) for your analyses. Use the quality flags (BLH_flag, cloud_flag) to filter out poor-quality retrievals. For more information or questions about processing methods, please contact:Nicholas E. Prince ⟨nec5299@psu.edu⟩Penn State Department of Meteorology & Atmospheric Science

Air Quality↗

Baltimore Social-Environmental Collaborative (BSEC) Doppler Lidar & Derived Products

This repository contains all processed Doppler‐lidar outputs from the PSU lidar deployed for the Baltimore Social‐Environmental Collaborative (BSEC) project. Vertical Stare Scans (fixed‐beam, vertical profiling): 1 Hz backscatter intensity (m⁻¹ sr⁻¹), signal‐to‐noise ratio (unitless), and Doppler vertical‐velocity (m s⁻¹) on 30 m range gates (and 3 m range gates), stored as CF-compliant NetCDF. Wind Profiles (horizontal‐wind retrieval): daily NetCDF outputs of retrieved horizontal wind speed (m s⁻¹) and direction (degrees), computed from the angled‐scan returns. Profile Statistics (summary statistics on the vertical velocity): 15 min windows (default) of mean, variance, skewness, kurtosis, high-frequency variance, etc., as a function of height; saved as CF-compliant NetCDF files. Boundary Layer Height (BLH) (fuzzy-logic output): 15 min BLH estimates (m), with lower/upper fuzzy bounds (m) and a quality flag (0–4) indicating data status (e.g., no data, good, ran out of signal, below range, cloud-topped). Cloud Base Height (Haar-gradient detection): 10 min estimates of cloud-base height (m). All five product streams are organized by year and date under their own top-level folders (01_Vertical_Stare_Scans/ through 05_Cloud_Height/). Each folder contains a data_ /YYYY/ subdirectory with daily CF-compliant NetCDF outputs (96 windows per day at 15 min intervals). Global attributes in each file include creation history, version (3.0.1), institution, and source. Instrument & MeasurementsThe PSU Doppler Lidar samples aerosol backscatter (m⁻¹ sr⁻¹), signal-to-noise ratio, and radial velocity at ~1 Hz. Vertical stare scans point the beam straight up; after collecting angled scans through multiple elevation angles, the "Wind Profiles" product contains the fully retrieved horizontal wind speed and direction. Data were collected continuously at ~30 m range resolution (and 3 m for the year of 2025), with a typical height ceiling of ~12 km. How to Use Open any NetCDF with Python's xarray, MATLAB, or similar CF-compliant tools. Stare scans and angled-scan retrievals (Wind Profiles) are CF-compliant daily NetCDF files. Profile-Statistics, BLH, and Cloud Height files are daily 15 min (10 min for Cloud Heights) summaries (96 time steps per file). Inspect the included variables (e.g., vertical_velocity_variance, wind_speed, BLH, cloud_base_height) for your analyses. Use the quality flags (BLH_flag, cloud_flag) to filter out poor-quality retrievals. For more information or questions about processing methods, please contact:Nicholas E. Prince ⟨nec5299@psu.edu⟩Penn State Department of Meteorology & Atmospheric Science

Air Quality↗

Model files for estimating snow dynamics and stable water isotopes across the East River, CO.

A coupled hydrologic and snowpack stable water isotope model is used to assesses controls on isotopic inputs across the East River, Colorado, a large, mountainous basin. The hydrologic model uses the semi-empirical, spatially distributed and publicly available U.S. Geological Survey numerical code Precipitation-Modelling Runoff System (PRMS). Water and energy are tracked daily through the atmosphere, canopy and subsurface at a 100-m grid resolution. The isotope mass balance model follows previous work by Ala-aho et al. (2017) to track stable isotopes entering the soil system as snowmelt or rain. Water stores and fluxes needed for the isotope model use hydrologic model output for each timestep and model modeled grid location. This data package provides all files related to the hydrologic model (executable, input and output files), the isotopic model calibration and the historic isotopic model (water years 2015-2020). The isotope model source code and executable are provided. The file "readme.txt" includes information on the file structure, steps to run the model, and description of included folders.

54 ENVIRONMENTAL SCIENCES↗

SPRUCE Vegetation Phenology in Experimental Plots from Phenocam Imagery, 2015-2022

This data set consists of PhenoCam data from the SPRUCE experiment from the beginning of whole ecosystem warming (Hanson et al. 2017) in August 2015 through March 31 of 2023, with start- and end-of-season phenological transition dates derived through the end of autumn 2022. Digital cameras, or phenocams, installed in each SPRUCE enclosure track seasonal variation in vegetation “greenness”, a proxy for vegetation phenology and associated physiological activity. Three separate regions of interest (ROIs) were defined for each camera field of view, corresponding to different vegetation types and demarcating (1) Picea trees (vegetation type EN, for evergreen needleleaf); (2) Larix trees (vegetation type DN, for deciduous needleleaf); and (3) the mixed shrub layer (vegetation type SH). This data set consists of three sets of data files: 3-day summary product files: One file for each camera and each ROI (i.e. vegetation type), characterizing vegetation color at a 3-day time step Contains 35 files in *.csv format inside a compressed (*.zip) file. Transition date file: Estimates “greenness rising” (spring) and “greenness falling” (autumn) transition dates derived from the smoothed daily green chromatic coordinate (GCC) values, for each camera and each ROI (i.e. vegetation type) Contains one file in *.csv format Snow flag files: Indicate days with snow on trees or snow on ground for each experimental enclosure Contains two files in *.csv format, one for snow on trees and one for snow on ground This data set consists of two sets of companion files: Accompanying HTML files show the 90th quantiles of the mean GCC plotted together with transition dates for each vegetation type and plot. Contains three files in HTML format, one for each vegetation type One additional file in HTML format with the transition dates plotted for each vegetation type, by year R files for processing Phenocam files and flags. Contains five files in R file (*.R) format in one compressed (*.zip) file User Note: All imagery is posted in near-real time to the PhenoCam Project web page (http://phenocam.sr.unh.edu/), where it is publicly available. Scroll to “spruce” in the Gallery or link directly to the 29 SPRUCE cameras at https://tinyurl.com/y7z5mau7. The data reported here are based on the complete camera record from SPRUCE and supersedes the previously released phenocam datasets (see Related Data Sets). The estimated transition dates for previously released datasets may differ slightly (in most cases, by ±3 days or less), because following standard PhenoCam processing protocols (Richardson et al. 2018, Scientific Data), smoothing and interpolation, outlier removal, and transition date estimation are always conducted using the full data record.

SPRUCE Experiment, Marcell Experimental Forest, Sp↗

SPRUCE Vegetation Phenology in Experimental Plots from Phenocam Imagery, 2015-2023

This data set consists of PhenoCam data from the SPRUCE experiment from the beginning of whole ecosystem warming (Hanson et al. 2017) in August 2015 through March 31 of 2024, with start- and end-of-season phenological transition dates derived through the end of autumn 2023. Digital cameras, or phenocams, installed in each SPRUCE enclosure track seasonal variation in vegetation “greenness”, a proxy for vegetation phenology and associated physiological activity. Three separate regions of interest (ROIs) were defined for each camera field of view, corresponding to different vegetation types and demarcating (1) Picea trees (vegetation type EN, for evergreen needleleaf); (2) Larix trees (vegetation type DN, for deciduous needleleaf); and (3) the mixed shrub layer (vegetation type SH). This data set consists of three sets of data files: (1) 3-day summary product files: One file for each camera and each ROI (i.e. vegetation type), characterizing vegetation color at a 3-day time step • Contains 36 files in *.csv format inside a compressed (*.zip) file. (2) Transition date file: Estimates “greenness rising” (spring) and “greenness falling” (autumn) transition dates derived from the smoothed daily green chromatic coordinate (GCC) values, for each camera and each ROI (i.e. vegetation type) • Contains one file in *.csv format (3) Snow flag files: Indicate days with snow on trees or snow on ground for each experimental enclosure • Contains two files in *.csv format, one for snow on trees and one for snow on ground This data set consists of two sets of companion files: (1) Accompanying HTML files show the 90th quantiles of the mean GCC plotted together with transition dates for each vegetation type and plot. • Contains three files in HTML format, one for each vegetation type • One additional file in HTML format with the transition dates plotted for each vegetation type, by year (2) R files for processing Phenocam files and flags. • Contains five files in R file (*.R) format in one compressed (*.zip) file User Note: All imagery is posted in near-real time to the PhenoCam Project web page (http://phenocam.sr.unh.edu/), where it is publicly available. Scroll to “spruce” in the Gallery or link directly to the 29 SPRUCE cameras at https://tinyurl.com/sprucecams. The data reported here are based on the complete camera record from SPRUCE and supersedes the previously released data inclusive of the 2015-2022 data (see Related Data Sets). The estimated transition dates for previously released datasets may differ slightly (in most cases, by ±3 days or less), because following standard PhenoCam processing protocols (Richardson et al. 2018, Scientific Data), smoothing and interpolation, outlier removal, and transition date estimation are always conducted using the full data record.

Spruce and Peatland Responses Under Changing Envir↗

SPRUCE Vegetation Phenology in Experimental Plots from PhenoCam Imagery, 2015-2024

This data set consists of PhenoCam data from the SPRUCE experiment from the beginning of whole ecosystem warming (Hanson et al. 2017) in August 2015 through March 31 of 2025 (2015-08-24 to 2025-03-31), with start- and end-of-season phenological transition dates derived through the end of autumn 2024. Digital cameras, or phenocams, installed in each SPRUCE enclosure track seasonal variation in vegetation “greenness”, a proxy for vegetation phenology and associated physiological activity. Three separate regions of interest (ROIs) were defined for each camera field of view, corresponding to different vegetation types and demarcating (1) Picea trees (vegetation type EN, for evergreen needleleaf); (2) Larix trees (vegetation type DN, for deciduous needleleaf); and (3) the mixed shrub layer (vegetation type SH). This data set consists of three sets of data files: (1) 3-day summary product files: One file for each camera and each ROI (i.e. vegetation type), characterizing vegetation color at a 3-day time step. • Contains 36 files in *.csv format inside a compressed (*.zip) file. (2) Transition date file: Estimates “greenness rising” (spring) and “greenness falling” (autumn) transition dates derived from the smoothed daily green chromatic coordinate (GCC) values, for each camera and each ROI (i.e., vegetation type). • Contains one file in *.csv format. (3) Snow flag files: Indicate days with snow on trees or snow on ground for each experimental enclosure. • Contains two files in *.csv format, one for snow on trees and one for snow on ground. This data set consists of two sets of companion files: (1) Accompanying HTML files show the 90th quantiles of the mean GCC plotted together with transition dates for each vegetation type and plot. • Contains three files in HTML format, one for each vegetation type. • One additional file in HTML format with the transition dates plotted for each vegetation type, by year. (2) R files for processing PhenoCam files and flags. • Contains five files in R file(*.R) format and the components of the phenocamr package (Version 1.1.4) used for calculating transition dates for 2015-2024. These are contained in a compressed (*.zip) file. User Note: All imagery is posted in near-real time to the PhenoCam Project web page (https://phenocam.nau.edu), where it is publicly available. Scroll to “spruce” in the Gallery or link directly to the 29 SPRUCE cameras at https://tinyurl.com/sprucecams. This data set is based on the complete camera record from SPRUCE and supersedes all previously released PhenoCam datasets (see Related Data Sets). The estimated transition dates for previously released datasets may differ slightly (in most cases, by ±3 days or less), because following standard PhenoCam processing protocols (Richardson et al. 2018, Scientific Data), smoothing and interpolation, outlier removal, and transition date estimation are always conducted using the full data record.

54 ENVIRONMENTAL SCIENCES↗

CORPSE model with litter decomposition parameters derived from the LIDET dataset

This is a version of the CORPSE model (Carbon, Organisms, Rhizosphere and Protection in the Soil Environment, Sulman et al. 2014) that uses litter decomposition parameters derived from a modified Monte Carlo simulation using the LIDET litter decomposition dataset (Long-term Intersite Decomposition Experiment Team, Harmon 2013). The code also includes the Baseline parameters, and the eight other best parameter sets identified in a modified Monte Carlo simulation. Related publication:Juice, S.M., Ridgeway, J.R., Hartman, M.D., Parton, W.J., Berardi, D.M., Sulman, B.N., Allen, K.E., & Brzostek, E.R. Reparameterizing litter decomposition using a simplified Monte Carlo method improves litter decay simulated by a microbial model and alters bioenergy soil carbon estimates. Description of files:The folder "Input Files" contains one folder for each LIDET site with data necessary to run the model. Note that "(site)" in the filenames below indicates where the LIDET site code appears (see Table 1 for site codes). Data streams include: CORPSE_full_spinup_litter.csv, CORPSE_full_spinup_rhizo.csv, CORPSE_full_spinup_bulk.csv, litterbag_init_100g_6spp.csv: initial C and N (kg C or N/m2) pool values for each soil layer, the litterbag_init_100_6spp.csv file is for the litterbag layer and is the same file for all sites. All initial C and N files have the same columns (Column - Description - Units) uFastC - Unprotected fast decomposing carbon - kg carbon/m2 uSlowC - Unprotected slow decomposing carbon - kg carbon/m2 uNecroC - Unprotected necromass carbon - kg carbon/m2 pFastC - Protected fast decomposing carbon - kg carbon/m2 pSlowC - Protected slow decomposing carbon - kg carbon/m2 pNecroC - Protected necromass carbon - kg carbon/m2 livingMicrobeC - Carbon in living microbial biomass - kg carbon/m2 uFastN - Unprotected fast decomposing nitrogen - kg nitrogen/m2 uSlowN - Unprotected slow decomposing nitrogen - kg nitrogen/m2 uNecroN - Unprotected necromass nitrogen - kg nitrogen/m2 pFastN - Protected fast decomposing nitrogen - kg nitrogen/m2 pSlowN - Protected slow decomposing nitrogen - kg nitrogen/m2 pNecroN - Protected necromass nitrogen - kg nitrogen/m2 inorganicN - Inorganic nitrogen - kg nitrogen/m2 CO2 - Carbon in carbon dioxide - kg carbon/m2 livingMicrobeN - Nitrogen in living microbial biomass - kg nitrogen/m2 soilT (site) DOY274start.csv: Average daily soil temperature (oC) interpolated from previously calculated monthly values used in DayCent LIDET simulations (Bonan et al., 2013). soilT (site) DOY274start.csv: Average daily soil volumetric water content (VWC) scalar interpolated from previously calculated monthly values used in DayCent LIDET simulations (Bonan et al., 2013). litter production.csv: Average daily litter production values for each site, data sources listed in Table S3 of related publication. litter (site) CN.csv: C:N ratio for each species from LIDET dataset (Table 2, Harmon 2013). (site).csv: Table indicating number of observations for each species decomposed at each site. Instructions: Save the model code ("CORPSE_LIDET.R") and "Input Files" folder in the same folder. Also make a folder for the model output (e.g., "results_Baseline") in the same folder. Set the working directory (setwd) in the model code to the folder with the files saved in step #1. Select the parameter set to use for the litter and litterbag compartments, comment out all other parameter sets. Run code. Output will be saved in the folder made in step 1. Output destination can be changed as necessary in code section called "Running the model." Table 1 LIDET sites and site codes used in model files. Site Code - Site AND - H.J. Andrews Experimental Forest BNZ - Bonanza Creek Experimental Forest BSF - Blodgett Research Forest CDR - Cedar Creek Natural History Area CPR - Central Plains Experimental Range HBR - Hubbard Brook Experimental Forest HFR - Harvard Forest JUN - Juneau KBS - Kellogg Biological Station KNZ - Konza Prairie Research Natural Area NWT - Niwot Ridge/Green Lakes Valley OLY - Olympic National Park OLY Conifer forest SEV - Sevilleta National Wildlife Refuge SMR - Santa Margarita Ecological Reserve UFL - University of Florida VCR - Virginia Coast Reserve Table 2 LIDET species and species codes used in model files (6 common species). Species - Species Code Sugar maple (Acer saccharum) - ACSA Drypetes (Drypetes glauca) - DRGL Red pine (Pinus resinosa) - PIRE Chestnut oak (Quercus prinus) - QUPR Western redcedar (Thuja plicata) - THPL Wheat (Triticum aestivum) - TRAE References:Bonan, G. B., Hartman, M. D., Parton, W. J., & Wieder, W. R. (2013). Evaluating litter decomposition in earth system models with long-term litterbag experiments: an example using the Community Land Model version 4 (CLM4). Global Change Biology, 19(3), 957-974. https://doi.org/https://doi.org/10.1111/gcb.12031 Harmon, M. (2013). LTER Intersite Fine Litter Decomposition Experiment (LIDET), 1990 to 2002. Long-Term Ecological Research. Forest Science Data Bank, Corvallis, OR. [Data set]. Accessed http://andlter.forestry.oregonstate.edu/data/abstract.aspx?dbcode=TD023. https://doi.org/10.6073/pasta/f35f56bea52d78b6a1ecf1952b4889c5. Sulman, B. N., Phillips, R. P., Oishi, A. C., Shevliakova, E., & Pacala, S. W. (2014). Microbe-driven turnover offsets mineral-mediated storage of soil carbon under elevated CO2. Nature Climate Change, 4, 1099 - 1102. https://doi.org/10.1038/nclimate2436

Juice, Stephanie↗

FUN-BioCROP model with litter decomposition parameters derived from the LIDET dataset

This repository contains the code and data necessary to run the FUN-BioCROP (Fixation and Uptake of Nitrogen-Bioenergy Carbon, Rhizosphere, Organisms, and Protection) model with litter decomposition parameters derived from a modified Monte Carlo simulation that used the Long-term Intersite Decomposition Experiment Team dataset. Related publication:Juice, S.M., Ridgeway, J.R., Hartman, M.D., Parton, W.J., Berardi, D.M., Sulman, B.N., Allen, K.E., & Brzostek, E.R. Reparameterizing litter decomposition using a simplified Monte Carlo method improves litter decay simulated by a microbial model and alters bioenergy soil carbon estimates. Description of Files: FUNBioCROP_LIDET Study.Rmd R code with FUN-BioCROP model that can be run with 10 different sets of parameters for litter decomposition (Baseline, LIDET, or eight other best parameter sets identified in the modified Monte Carlo simulation. CORPSE Functions_Bioenergy_V2.R Code with CORPSE model functions, called by FUNBioCROP_LIDET Study.Rmd Model Input Data: bulk.csv, bulk_till.csv, rhizo.csv, rhizo_till.csv, litter.csv Initial C and N (kg C or N/m2) pool values for each soil compartment, final values from spin up. All five files have the same columns: (Column - Description - Units) uFastC - Unprotected fast decomposing carbon - kg carbon/m2 uSlowC - Unprotected slow decomposing carbon - kg carbon/m2 uNecroC - Unprotected necromass carbon - kg carbon/m2 pFastC - Protected fast decomposing carbon - kg carbon/m2 pSlowC - Protected slow decomposing carbon - kg carbon/m2 pNecroC - Protected necromass carbon - kg carbon/m2 livingMicrobeC - Carbon in living microbial biomass - kg carbon/m2 uFastN - Unprotected fast decomposing nitrogen - kg nitrogen/m2 uSlowN - Unprotected slow decomposing nitrogen - kg nitrogen/m2 uNecroN - Unprotected necromass nitrogen - kg nitrogen/m2 pFastN - Protected fast decomposing nitrogen - kg nitrogen/m2 pSlowN - Protected slow decomposing nitrogen - kg nitrogen/m2 pNecroN - Protected necromass nitrogen - kg nitrogen/m2 inorganicN - Inorganic nitrogen - kg nitrogen/m2 CO2 - Carbon in carbon dioxide - kg carbon/m2 livingMicrobeN - Nitrogen in living microbial biomass - kg nitrogen/m2 Model Input Data: FluxTower_AvgSoilT.csv: Average daily soil temperature (oC) at 10 cm depth at University of Illinois Urbana-Champaign (UIUC) Energy Farm flux tower from 7/2008-3/2016. (One year of averaged data) Model Input Data: FluxTower_AvgSoilVWC.csv: Average daily soil volumetric water content (VWC) at 10 cm depth at UIUC Energy Farm flux tower from 7/2008-3/2016. (One year of averaged data) Model Input Data: input_CCS_LIDET Study.csv: This file has daily data to run FUN-BioCROP (Column - Description - Units): yr - calendar year - year doy - day of year (1 to 365) (no leap year) - day anpp - aboveground NPP (DayCent) - kg C/m2/day bnpp - belowground NPP (DayCent) - kg C/m2/day aglivc - live aboveground biomass carbon (DayCent) - kg C/m2 bglivcj - live juvenile fine root biomass carbon (DayCent) - kg C/m2 bglivcm - live mature fine root biomass carbon (DayCent) - kg C/m2 aglivn - live aboveground biomass nitrogen (DayCent) - kg N/m2 bglivnj - live juvenile fine root biomass nitrogen (DayCent) - kg N/m2 bglivnm - live mature fine root biomass nitrogen (DayCent) - kg N/m2 nyr - simulation year - year cult - indicates a cultivation event (0 or 1) crop - indicates a new crop (0 or 1) fert - indicates a fertilizer event (0 or 1) frst - indicates the first day of the growing season (0 or 1) harv - indicates a harvest event (0 or 1) last - indicates the end of the growing season (0 or 1) croptype - crop type (0=none; 1=alfalfa; 2=corn; 3=grass clover pasture; 4=soybean; 5=wheat) cropsrl - crop specific root length - mm/g root cultrhizmix - fraction of rhizosphere mixed with bulk soil during cultivation (0.0-1.0) - fraction cultlitmix - fraction of litter mixed with bulk soil during cultivation (0.0-1.0) - fraction harvremov - fraction of above ground biomass removed during harvest (0.0-1.0) - fraction fertamt - fertilization amount - g N/m2 lifehist - plant life history (0 = annual, 1 = perennial) froot_turnover_c - amount of C in fine root turnover - kg C/m2 froot_turnover_n - amount of N in fine root turnover - kg N/m2 agrd_turnover_c - amount of C in aboveground biomass turnover - kg C/m2 agrd_turnover_n - amount of N in aboveground biomass turnover - kg N/m2 leaf_litter_fastfrac - Fast decomposing fraction of leaf litter (0.0-1.0) - fraction root_litter_fastfrac - Fast decomposing fraction of root litter (0.0-1.0) - fraction root_diameter - root diameter - mm root_length - root length - mm root/m2 rhizo_frac - fraction of total soil volume that is rhizosphere (0.0 - 1.0) - fraction date - date in format YYYY-MM-DD Instructions: Save the model code ("FUN-BioCROP_LIDET Study.Rmd") and accompanything files (data streams and CORPSE function code) in the same folder. In model code "Chunk 3: Load CORPSE Data Streams" set the working directory (setwd) to the folder with the files saved in step #1. In "Chunk 5: Define LIDET parameter sets" select the litter decomposition parameter set to be used in the run, and comment out all other sets. If changing any parameter values, edit them in "Chunk 6: Load parameters." Run all chunks up to and including "Chunk 10: Prepare Data for Export." In "Chunk 11: Export Output Data" edit data frames for export and filenames, as necessary. "Chunk 12: Graph Total Soil C" makes a figure of C remaining over the model run period. Description of each model chunk (in file FUN-BioCROP_LIDET Study.Rmd): Chunk 1: Remove all functions, clear memory. Removes all functions from R environment, clears the memory. Chunk 2: Load Packages. Loads packages necessary to run the code. Chunk 3: Load CORPSE Data Streams. Sets the working directory and loads the data files necessary to run CORPSE. Chunk 4: Load CORPSE Functions. Loads the R script with CORPSE functions from the working directory, "CORPSE Functions_Bioenergy_V2.R". Chunk 5: Define LIDET parameter sets. Has ten different parameter sets for litter decomposition tested in this study: Baseline parameters, LIDET parameters, and the other 8 best performing parameter sets identified in the modified Monte Carlo. To run the model, all but one parameter set must be commented out. Chunk 6: Load Parameters. Loads all fixed parameters to run the model. Data frame with definitions of parameters is in the CORPSE function script "CORPSE Functions_Bioenergy_V2.R" Chunk 7: Prepare Data Streams. Takes data streams loaded in Chunk 3 and puts them in the format necessary to run the model. The model is coded to run at least two sites at a time, so if only one site is being run it must be run in duplicate. Individual data tables of daily values are created in this chunk from the input data file. Chunk 8: Set Initial Conditions. Creates data tables of soil C and N pools for each soil compartment (rhizo_till, rhizo, bulk_till, bulk, litter) and loads initial values into the data tables. Creates lists for each soil compartment to hold model output. Chunk 9: Load FUN Data and Set Up Matrices. Uses DayCent data to calculate FUN input data: root and leaf N demand, total N demand, plant CN, leaf N available for retranslocation, and litter production. Creates matrices for FUN model outputs. Chunk 10: Run Model. Runs the model. Chunk 11: Prepare Data for Export. Combines data from each day saved as lists into data frames for each soil compartment. Adds values from all soil compartments together to calculate total soil values, creates separate data frames for each soil C and N pool (e.g., protected slow C) for the total soil value. Adds different C and N pools together to calculate total soil C and N for all layers. Creates data frame of ratio of protected to unprotected SOC. Organizes FUN data for export. Chunk 12: Export Results. Exports CSV files of model results to the working directory. Chunk 13: Graph Total Soil C. Makes figure of C remaining over time. Related Links: Original FUN-BioCROP model: https://github.com/BrzostekEcologyLab/FUN-BioCROP LIDET dataset: https://andlter.forestry.oregonstate.edu/data/abstract.aspx?dbcode=TD023

Juice, Stephanie↗

Supporting information for Few-Shot Learning Enables Population-Scale Analysis of Leaf Traits in Populus trichocarpa

In this work, we use few-shot learning to segment the body and vein architecture of P. trichocarpa leaves from high-resolution scans obtained in the UC Davis common garden. Leaf and vein segmentation are formulated as separate tasks, in which convolutional neural networks (CNNs) are used to iteratively expand partial segmentations until reaching stopping criteria. Our leaf and vein segmentation approaches use just 50 and 8 manually traced images for training, respectively, and are applied to a set of 2,634 top and bottom leaf scans. We show that both methods achieve high segmentation accuracy, in some cases exceeding even human-level segmentation. The leaf and vein segmentations are subsequently used to extract 68 morphological traits using traditional open-source image processing tools, which are validated using real-world physical measurements. For a biological perspective, we perform a genome-wide association study using the vein density trait to discover novel genetic architectures associated with multiple physiological processes relating to leaf development and function. In addition to sharing all of the few-shot learning code (see https://github.com/jlager/few-shot-leaf-segmentation), we are releasing all images, manual segmentations, model predictions, 68 extracted leaf phenotypes, and a new set of SNPs called against the v4 P. trichocarpa genome for 1,419 genotypes. The data folder includes all images, ground truth segmentations, predicted segmentations, and extracted leaf traits. All images encode the sample ID in the file name by indicating the treatment, block, row, position, and leaf side, respectively. For example, the file, C_1_1_2_bot.jpeg, indicates the control treatment, block 1, row 1, position 2, and the bottom side of the leaf. Tabulated results include position IDs as well as the corresponding genotype IDs. The images folder includes the 2,906 high-resolution leaf scans taken in the field. The leaf_masks folder includes 50 ground truth segmentations used for training the leaf tracing algorithm. The leaf_preds folder includes the 2,906 predicted segmentations from the leaf tracing algorithm. The vein_masks folder includes 8 ground truth segmentations used for training the vein growing algorithm. The vein_preds folder includes the 1,453 predicted segmentations from the vein growing algorithm. The vein_probs folder includes the 1,453 predicted probability maps from the vein growing algorithm before thresholding. The genomes folder includes the set of SNPs called against the v4 P. trichocarpa genome for 1,419 genotypes with a README file detailing the steps taken. The results folder includes: (i) raw values of the 68 predicted leaf traits in digital_traits.tsv, (ii) manually measured values of petiole length and width in manual_traits.tsv, (iii) thin plate spline (TPS) adjusted values of the vein density trait in vein_density_tps_adj.tsv, (iv) best linear unbiased prediction (BLUP) adjusted values of the vein density trait in vein_density_blups.tsv, and (v) GWAS results for the vein density trait, including chromosome positions and corresponding P values, in gwas_results.csv.

09 BIOMASS FUELS↗

KBase Narrative - Viral Analysis End-to-End

This pipeline processes a single metagenomic dataset (downloaded from NCBI SRA) from the Global Ocean Viromes dataset, identifies the viral sequences using VirSorter, and subsequently classifies them with vConTACT2. Along the way there are several intermediary steps to convert files from one KBase object to another - and though some of these can be skipped for this particular dataset - we're explicitly doing them here as to be more easily generalized to other datasets.

Bolduc, Benjamin↗

Spatio-Temporal Surrogates for Interaction of a Jet with High Explosives: Part II - Clustering Extremely High-Dimensional Grid-Based Data

Building an accurate surrogate model for the spatio-temporal outputs of a computer simulation is a challenging task. A simple approach to improve the accuracy of the surrogate is to cluster the outputs based on similarity and build a separate surrogate model for each cluster. This clustering is relatively straightforward when the output at each time step is of moderate size. However, when the spatial domain is represented by a large number of grid points, numbering in the millions, the clustering of the data becomes more challenging. In this report, we consider output data from simulations of a jet interacting with high explosives. These data are available on spatial domains of different sizes, at grid points that vary in their spatial coordinates, and in a format that distributes the output across multiple files at each time step of the simulation. We first describe how we bring these data into a consistent format prior to clustering. Borrowing the idea of random projections from data mining, we reduce the dimension of our data by a factor of thousand, making it possible to use the iterative k-means method for clustering. We show how we can use the randomness of both the random projections, and the choice of initial centroids in k-means clustering, to determine the number of clusters in our data set. Our approach makes clustering of extremely high dimensional data tractable, generating meaningful cluster assignments for our problem, despite the approximation introduced in the random projections.

97 MATHEMATICS AND COMPUTING↗

Topsoil bulk geochemical compositions - An updated harmonized global dataset

Mineral weathering is a key biogeochemical process because of the capacity of minerals to stabilize organic matter. However, predicting soil weathering status across large spatial areas still isn’t possible due to a lack of global data and theoretical frameworks. To address this knowledge gap, multiple global datasets of bulk topsoil geochemical compositions have been harmonized using R. These datasets document topsoil bulk geochemical compositions across five continents (n = ~16,000 observations). Source data for these observations include the EuroGEOSurveys Geochemical Baseline Database (FOREGS), the US Geological Survey National Geochemical Database (NASGLP), the Geochemical Atlas of Australia (GAA), the US Geological Survey Alaska Geochemical Database (AGD84), the National Cooperative Soil Survey (NCSS), the European Geochemical Mapping of Agricultural Soil (GEMAS), Ecorespira-Amazon (ERA), the New Zealand Geochemical Baseline Survey (NZ_GBS), and the African Soil Information Service (AFSIS). Major elements observed include Aluminum (Al), Calcium (Ca), Iron (Fe), Potassium (K), Magnesium (Mg), Sodium (Na), Titanium (Ti), Manganese (Mn), Phosphorus (P), Carbon (C), and Sulfur (S). This data package includes the harmonized dataset itself, and the R scripts necessary to harmonize these datasets, in addition to metadata that describes all columns, files, and databases used in this project. Methods & Sampling Step 1 – Databases of geochemical data identified This study aimed to leverage existing measurements of topsoil geochemical data. Databases were first identified and deemed appropriate for inclusion if they were measuring soils and performed these measurements on the <2mm soil fraction. Databases such as NCSS and AGD84 needed more post processing to include in the database and this was done using the NCSS_datamerge_031626 R file and Alaska_USGSmerge_031626 R file, respectively. Step 2 – Database harmonization Once appropriate databases were identified, they were harmonized for ease of analysis using the R script Database_Harmonization_031826. This included removing columns from original datasets that would not be used in analysis (removed columns are noted in the code). Then, data cleaning procedures specific to each dataset were undertaken. This includes standardizing columns to include units and adding metadata columns regarding procedures for analyzing specific elements. Functions for standardizing measurements and units are outline in R files: calculate element_mg_kg_031626, calculate_oxide_wt_perc_031626, change_oxide_caps_031626, and conv_2_numeric_031626. This also included adding a unique identifier for each sample to identify it with its respective database (see CD_ID in data dictionary). Geographic information: Data reflect a compilation of datasets collected globally. Geographic areas covered by each of the datasets include: - EuroGEOSurveys Geochemical Baseline Database (FOREGS) - European continent - North American Soil Geochemical Landscapes (NASGLP) - continental United States and limited parts of Canada (see database key for more details) - National Geochemical Survey of Australia (GAA) - Australia - Alaska geochemical database (AGDB4) - Alaska - National Cooperative Soil Survey (NCSS) - Global measurements, but concentrated in the continental United States - Geochemical data for arable land and land under permanent grass cover in continental Europe (GEMAS) - continental Europe - Ecorespira-Amazon (ERA) - Geochemical data from the Amazon basin - Geochemical baseline data for New Zealand (NZGBS) - New Zealand - Geochemical data collected across continental Africa (AfSIS) - Measurements across Africa

EARTH SCIENCE > LAND SURFACE > SOILS↗

Code for BALDR Study 07.05 v.1.0.0

SAND2024-01293O The Code for BALDR Study 07.05 software reproduces results from the BALDR study concerning "A Semi-Supervised Learning Method to Produce Explainable Radioisotope Proportion Estimates for NaI-based Synthetic and Measured Gamma Spectra." The code for BALDR Study 07.05 can reproduce a scientific study following the step numbers within the file names. Using synthetic and measured data to find the best model for the radioisotope proportion estimation task of interest, the study generated results for inclusion in a paper. The high-level methods involved are neural networks, semi-supervised learning, and out-of-distribution detection. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Morrow, Tyler↗