Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Data Variables”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Soil temperature and soil moisture raw data, permafrost table depths, and accompanying environmental variable data, Kenai Wildlife Refuge, 2019-2022

Data package purpose: This data package was created to contain all data used in an upcoming article, "Canopy Cover and Microtopography Control Precipitation-Enhanced Thaw of Ecosystem-Protected Permafrost." In review.This data package includes: Raw output from 19 distributed temperature profilers with a thermistor every 10 cm along a 160 cm length at a measurement interval of 15 minutes (.CSV). Raw output from two soil moisture and temperature profilers (90 cm length and 120 cm length) that took composite soil moisture readings every 15 cm along the sensor length at a measurement interval of 30 minutes (.CSV). Permafrost depths were measured annually in mid-September at DTP sensor locations (.CSV) and along an across-site transect (.CSV). Environmental variables (snow depth, canopy closure, moss depth, and elevation) for all sensor locations. Real-time kinetic (RTK) GPS points showing site microtopography (.CSV).Analysis software: Our analysis was done in Matlab. File types can be used with any software.

54 ENVIRONMENTAL SCIENCES↗

Strength and ductility of additively manufactured 316L stainless steel: Impact of neutron irradiation and data variability

Here, this article presents the mechanical properties of additively manufactured (AM) 316L stainless steel processed via the laser powder bed fusion (LPBF) method, focusing on the effects of neutron irradiation on mechanical properties and the variability in strength and ductility data. The rapid melting-solidification process and multiple heating-cooling cycles inherent in LPBF typically result in a fine, metastable microstructure with significant local variability. AM 316L builds of varying thicknesses were fabricated, and SS-J3 miniature tensile specimens were machined from six different locations. These specimens were irradiated with fast neutrons to doses of 2 and 10 dpa at target temperatures of 300 °C and 600 °C. Post-irradiation tensile tests were conducted at room temperature, 300 °C, and 600 °C. Compared to conventional 316L stainless steel, AM 316L exhibited higher initial strength but lower ductility. Irradiation at 300 °C caused significant hardening and prompt necking at yield, with limited uniform ductility, although embrittlement was not observed up to 10 dpa. While neutron irradiation, particularly at 600 °C, increased the variability in strength and ductility data, no clear dependence of mechanical properties on build thickness or sampling location was found—contrary to the conventional perception that AM materials may exhibit high property variability. Furthermore, we observed that the variability in property data for LPBF-processed 316L was relatively low compared to that of wrought 316L stainless steel. This reduced variability in AM 316L steel may be attributed to its highly metastable, stress-containing microstructure, which is discussed in the context of general tensile property variations.

Additively manufactured 316L stainless steel↗

Tencoder: tensor-product encoder-decoder architecture for predicting solutions of PDEs with variable boundary data

It is widely hoped that artificial intelligence will boost data-driven surrogate models in science and engineering. However, fundamental spatial aspects of AI surrogate models remain under-studied. We investigate the ability of neural-network surrogate models to predict solutions to PDEs under variable boundary values. We do not wish to retrain the model when the boundary values change but to make them inputs to the model and infer the solution of the PDE under those boundary conditions. Such a capability is essential to making AI-based surrogate models practically useful. While simple feedforward networks are used for one-dimensional (1D) Poisson equation, an encoder-decoder architecture with a tensor-product layer is developed for the two-dimensional Poisson equation posed on a rectangular domain. We show that it is indeed possible to infer solutions to PDEs from variable boundary data using neural networks in this relatively simple setting, and point to future directions.

Kashi, Aditya↗

Dataset from ORNL Flexible Research Platform (FRP)

The data comes from a two-story light commercial building which can be used to physically simulate light commercial buildings common in the nation's existing building stock. Based on a data collection plan, these measurements were performed for five different building operations. In the data set for each period, 98 data variables were selected and collected, in which 7 variables are weather data and the rest are all building and system operation data.

Cui, Borui↗

Timeseries Unlabeled and Labeled Photos, Modeled Stream Elevation, and (Meta)Data of Variably Inundated Streams Across The Yakima River Basin, Washington, United States (v2)

This dataset is associated with the “River Monitoring Photos” (RMP) study and subsequent manuscript (Bao et al. 2025. Monitoring river flow status using low-cost wildlife camera and image segmentation artificial intelligence doi: 10.1016/j.envsoft.2025.106715). Game camera timeseries photos were collected to evaluate stream variable inundation via changes in width. A subset of photos was labeled for training the YOLOv8 and Mask2Former models and used to segment water surface fractions from all the game camera photos.This data package was originally published in March 2024. It was updated in October 2025 (v2) to add additional photos and files associated with the manuscript (i.e., processed data, labeled photos, and trained models). For details on how to navigate data packages generated by this project, see https://data.ess-dive.lbl.gov/portals/PNNLRiverCorridorSFA/About.In addition to a readme, this data package also includes two file-level metadata (FLMD) files that describes each file and two data dictionaries (DD) that describe all column/row headers and variable definitions. This dataset is comprised of (1) file-level metadata; (2) data dictionary; (3) readme; (4) field metadata; and (5) folders containing game camera photos and manuscript-associated files. Each Yakima River Basin site has a folder that contains subfolders for each month photos were collected. There is also a folder for files associated with the manuscript which has subfolders for labeled data, trained models, Yakima River Basin site water surface fractions, and USGS site water surface fractions. All files are .csv, .json, .txt, .yaml, .pth, .pt, or .pdf. We thank the Confederated Tribes and Bands of the Yakama Nation Tribal Council and Yakama Nation Fisheries for working with us to facilitate sample collection and optimization of data usage according to their values and worldview.

54 ENVIRONMENTAL SCIENCES↗

Disentangling the Impacts of Microtopography and Shrub Distribution on Snow Depth in a Subarctic Watershed: Toward a Predictive Understanding of Snow Spatial Variability: Supporting Data and Code

This repository contains R code and associated datasets for reproducing the analysis described in the manuscript titled “Disentangling the Impacts of Microtopography and Shrub Distribution on Snow Depth in a Subarctic Watershed: Toward a Predictive Understanding of Snow Spatial Variability” (DOI: 10.1029/2024JG008604). The provided scripts facilitate a comprehensive analysis of snow depth variability influenced by microtopography and vegetation distribution in a subarctic watershed. Included datasets are high-resolution spatial maps of snow depth, terrain elevation, vegetation height, and distance from shrubs taller than 1 meter, all formatted as text files (.txt). These data are fully describe in doi:10.15485/2316038. Users can adapt the provided R scripts to accommodate different data formats or larger spatial domains, noting that some output files may require modification due to their size.The code includes implementations for boosted regression tree analysis adapted from methods outlined in Elith et al. (2008). Users interested in understanding or modeling landscape-scale snow distribution patterns, particularly in Arctic or subarctic ecosystems, will find this package useful. The Next-Generation Ecosystem Experiments: Arctic (NGEE Arctic), was a research effort to reduce uncertainty in Earth System Models by developing a predictive understanding of carbon-rich Arctic ecosystems and feedbacks to climate. NGEE Arctic was supported by the Department of Energy's Office of Biological and Environmental Research.The NGEE Arctic project had two field research sites: 1) located within the Arctic polygonal tundra coastal region on the Barrow Environmental Observatory (BEO) and the North Slope near Utqiagvik (Barrow), Alaska and 2) multiple areas on the discontinuous permafrost region of the Seward Peninsula north of Nome, Alaska.Through observations, experiments, and synthesis with existing datasets, NGEE Arctic provided an enhanced knowledge base for multi-scale modeling and contributed to improved process representation at global pan-Arctic scales within the Department of Energy's Earth system Model (the Energy Exascale Earth System Model, or E3SM), and specifically within the E3SM Land Model component (ELM).

54 ENVIRONMENTAL SCIENCES↗

Arctic Impact Identification with Less Data Using Variable Relationships: An Exploratory Express LDRD project.

Regional impacts from sea ice loss can be challenging to separate from internal climate variability, potentially requiring thousands of ensemble members. East Asian wintertime cooling has been linked to sea ice loss from present day conditions in the Polar Amplification Model Intercomparison Project with these large ensemble counts. This cooling is theorized to arise from a strengthened Siberian High and East Asian Jet response. The strengthened Siberian High can be detected with one fifth the ensemble members needed for the East Asian wintertime cooling in a single model. We thus hypothesize that leveraging relationships between multiple variables in a conditional pathways-based approach would reduce the number of required ensemble members to conclusively attribute East Asian wintertime cooling to future sea ice concentrations. In all analyzed cases, confidence was increased when evaluating sea ice loss’s responsibility for the joint effects of East Asian cooling, East Asian Jet strengthening, and Siberian High strengthening over just East Asian cooling. However, we were not able to confidently attribute future East Asian wintertime cooling to sea ice loss in a single model. We found that significant intra-ensemble variability within single Earth System Models (ESMs) produced highly uncertain forcing response models upon which attribution results were undermined. We were able to show that ensemble mean seasonally averaged metrics from multiple ESMs greatly improved the accuracy of the forcing response linear models and exposed the necessity of all three steps in the pathway (sea ice area, Siberian High pressure, and East Asian Jet speed) for accurate prediction of East Asian wintertime cooling. Although all three steps were necessary, East Asian wintertime cooling possesses a large dependence on the Siberian High pressure, which weakens the confidence associated with overall strong joint-attribution comparing present day and future scenarios. We believe transitioning the pathway nodes to relative changes between the Siberian High and Aleutian Low as well as between the midlatitude westerlies and subtropical jet in the East Asianj Jet region may be able to produce significant attribution more fully dependent upon all three steps. Ultimately, this research demonstrates the simple extensibility of conditional pathways-based attribution to sea ice loss forcing on the Earth system.

54 ENVIRONMENTAL SCIENCES↗

Hourly gap-filled meteorological data from PIE LTER measurements (2004-2023) used as drivers to run ELM PFLOTRAN simulations

This dataset contains continuous gap-filled precipitation, solar radiation, photosynthetically active radiation (PAR), air temperature, relative humidity, wind speed, and barometric pressure data recorded primarily at the Marshview Farm weather station within the Plum Island Long Term Ecosystems Research (PIE LTER) in Newbury Massachusetts (MA) from 2004 to 2023. We compiled the data set from published annual data packages in 15min resolution available on DataOne. Gaps were filled using different statistical techniques or available observations from the vicinity, e.g. the US-PLo and the US-PHM Ameriflux sites, also located within the PIE LTER. Flags are included in this dataset to indicate the origin of each data point. Metadata files ELMPFLOTRAN_met_dd.csv and ELMPFLOTRAN_met_flmd.csv contain more information on site locations, gap filling protocols, data variables, flags, and QA/QC methods. The data set was used in the spin up and simulations of a land surface model coupled to a biogeochemical reaction network (ELM PFLOTRAN) assessing impacts of hydrology and salinity input on methane fluxes in 2022 and 2023 (Sulman et al., 2024).

54 ENVIRONMENTAL SCIENCES↗

Metrics and Methods for Radiation Detection Algorithm Characterization for Nuclear/Radiological Source Search

This report presents a series of recommendations for data to train and evaluate radiation detection algorithms and performance metrics to evaluate these algorithms. These recommendations were formed through a community consensus approach through the Detection Radiation Algorithms Group (DRAG), a multi-institution collaboration spanning eight Department of Energy laboratories and John Hopkins Applied Physics Laboratory. This report includes recommendations on background data variability, and metrics to quantify variability, sources and shielding configurations to include in data collection campaigns and detector response variability. In addition, this report describes several anomaly detection and identification algorithms and recommends metrics to report their performance. Finally, this report ends with a discussion on machine learning algorithms.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Decoding Golden Eagle Movement Behavior from High-Resolution, Variable-Rate Telemetry Data Through Bayesian Filtering

The recent advances in animal tracking technology have enabled the collection of a vast amount of in situ data regarding the movement of wildlife at high spatiotemporal resolution. These data are usually available at variable time resolutions and contains noise (error) originating from GPS fixes. Decoding movement characteristics, particularly of flying animals, from telemetry data while handling these factors is a challenging yet important task for conservation purposes. Typically, this task is broken into two subtasks: resampling, and model calibration. The resampling subtask converts the variable rate positional data into a constant time interval data, while the model calibration subtask uses the resampled data to tune time-invariant parameters of the proposed models. For telemetry data at high temporal resolutions (order of 1 second), it is very challenging to decouple noise from actual movements using interpolation-based resampling techniques. Any errors introduced during resampling can significantly alter the the calibration and prediction attributes of the movement model. We address this problem through a unified Bayesian state-space framework that can handle both the resampling and calibration tasks in a single step. In addition, we use the speed and heading of the bird from telemetry data to regularize the position information of the bird. We use a Kalman filtering approach to include these nonlinearly related motion parameters within the state space framework. We cross-validated to quantify how this inclusion affects the model performance in estimating true bird movements. The relationship between the true state of the bird and environmental and topographical covariates is then represented parametrically. These parameters are then tuned using stochastic sampling strategies like Markov Chain Monte Carlo (MCMC). We use the telemetry data collected from golden eagles in the western USA to demonstrate the applicability of this approach to build a predictive, probabilistic movement model. Our preliminary results show that this approach provides improved predictive performance in terms of capturing higher-order motion parameters such as angular and horizontal accelerations, which may have simpler and more direct relationships with environmental covariates than corresponding speeds. In this talk, we will demonstrate how this state-space approach benefits the prediction capabilities of a movement model in simulating golden eagle paths through a wind power plant in Wyoming given certain atmospheric conditions. The model outcomes are aimed at informing mitigation strategies that can minimize the potential for collisions of golden eagles with wind turbines.

Bayesian methods↗

Challenges of open data in aquatic sciences: issues faced by data users and data providers

Free use and redistribution of data (i.e., Open Data) increases the reproducibility, transparency, and pace of aquatic sciences research. However, barriers to both data users and data providers may limit the adoption of Open Data practices. Here, we describe common Open Data challenges faced by data users and data providers within the aquatic sciences community (i.e., oceanography, limnology, hydrology, and others). These challenges were synthesized from literature, authors’ experiences, and a broad survey of 174 data users and data providers across academia, government agencies, industry, and other sectors. Through this work, we identified seven main challenges: 1) metadata shortcomings, 2) variable data quality and reusability, 3) open data inaccessibility, 4) lack of standardization, 5) authorship and acknowledgement issues 6) lack of funding, and 7) unequal barriers around the globe. Our key recommendation is to improve resources to advance Open Data practices. This includes dedicated funds for capacity building, hiring and maintaining of skilled personnel, and robust digital infrastructures for preparation, storage, and long-term maintenance of Open Data. Further, to incentivize data sharing we reinforce the need for standardized best practices to handle data acknowledgement and citations for both data users and data providers. We also highlight and discuss regional disparities in resources and research practices within a global perspective.

54 ENVIRONMENTAL SCIENCES↗

Evapotranspiration partitioning estimates from 8 methods from 47 NEON sites, 2019-2021

This dataset provides daily estimates of evapotranspiration (ET) and the transpiration-to-evapotranspiration ratio (T/ET) across 47 terrestrial National Ecological Observatory Network (NEON) sites spanning diverse environmental and biome conditions in the United States across three years of data (2019-2021). Daily ET is reported in both energy units (MJ m⁻² day⁻¹) and equivalent water depth (mm day⁻¹), assuming a constant latent heat of vaporization of 2.45 MJ/kg. The primary method uses a hybrid recurrent neural network–Penman–Monteith framework (RNN-PM), which integrates physically based surface energy balance constraints with data-driven learning to partition ET into transpiration and evaporation components. Model inputs include in situ meteorological observations (air temperature, vapor pressure deficit, wind speed, and radiation) combined with satellite-derived land surface temperature, leaf area index, and soil moisture. For benchmarking and uncertainty assessment, T/ET estimates from seven additional models are included: Priestley-Taylor Jet Propulsion Laboratory (PT-JPL), Penman-Monteith (P-M), Two-Source Energy Balance (TSEB), Support Vector Regression (SVR), and Categorical Boosting (CatBoost), among others—spanning empirical, machine-learning, and process-based approaches (see methods section or linked publication for detailed descriptions). Data Package Contents: The dataset a csv files containing daily ET and T/ET estimates for each site and model, along with associated metadata files these variables. Data can be accessed using common spreadsheet software (e.g., Microsoft Excel, LibreOffice) or programming environments such as R or Python. Together, these data support cross-site comparisons of ecosystem water use, evaluation of ET partitioning methods, and development of improved land–atmosphere exchange models.

EARTH SCIENCE > ATMOSPHERE↗

OpenCRUMS USA: An Open Machine Learning Framework for Characterizing Variability in Aerosol Reanalysis Data

Advances in artificial intelligence (AI) have called for exploring how these techniques can be used for exploring patterns in large climate datasets. To that regard, the U.S. Department of Energy AI for Earth System Predictability (AI4ESP) supported a pilot initiative called the Open Classification of Regimes in the Southeast USA (OpenCRUMS USA) project to explore how AI can be used to characterize modes of spatial variability in large climate datasets. For this study, we focus on comparing two methods for characterizing the modes of spatial variability of surface aerosol concentration over the Houston region: empirical orthogonal functions (EOFs) and layerwise relevance propagation (LRP) applied to a convolutional neural network (CNN) classifier. We show that EOF analysis typically attributes spatial variability modes that span all of southeast Texas, prohibiting the attribution of spatial variability to localized regions. However, using LRP on the CNN classifier resolves the explanatory parameters at a finer spatial resolution than EOFs. This allows for the attribution of the spatial variability of surface aerosols to local regions of organic carbon which was not possible using EOFs. In addition, the LRP analysis also suggests that synoptic-scale transport of dust is most prevalent during anticyclonic and pretrough synoptic conditions as categorized by self-organizing maps.

54 ENVIRONMENTAL SCIENCES↗

Prediction of Distributed River Sediment Respiration Rates Using Community-Generated Data and Machine Learning

River sediment microbial respiration is a key indicator of ecosystem functioning and the biogeochemical fluxes across this critical zone link surface and subsurface waters. As such, there is tremendous interest in measuring and mapping these respiration rates. Respiration observations are expensive and labor intensive; there is limited data available to the community. An open science, collaborative initiative is collecting samples for respiration rate analysis and multi-scale metadata; this evolving data set is being used for making machine learning (ML) predictions at unsampled sites to help inform continued community engagement. However, it is a challenge to find an optimum configuration for ML models to work with this feature-rich (i.e., 100+ possible input variables) data set. Here, we present results from a two-tiered approach to managing the analysis of this complex data set: (a) a stacked ensemble of models that automatically optimizes hyperparameters and manages the training of many models and (b) feature permutation importance to detect the most important features in the models. The major elements of this workflow are modular, portable, open, and cloud-based thus making this implementation a potential template for other applications. The models developed here predict that sediment organic matter chemistry is one of the most important features for predicting sediment respiration rate. Other larger-scale, important features fall into the categories of climatic, ecological, geological, and fluvial settings. Leveraging these larger-scale features to generate data-driven estimates of river sediment respiration rates reveals spatially consistent but heterogeneous patterns across the river network of the Columbia River Basin.

54 ENVIRONMENTAL SCIENCES↗

A lightweight, user-configurable detector ASIC digital architecture with on-chip data compression for MHz X-ray coherent diffraction imaging

Today, most X-ray pixel detectors used at light sources transmit raw pixel data off the detector ASIC. With the availability of more advanced ASIC technology nodes for scientific application, more digital functionalities from the computing domains (e.g., compression) can be integrated directly into a detector ASIC to increase data velocity. In this paper, we describe a lightweight, user-configurable detector ASIC digital architecture with on-chip compression which can be implemented in 130 nm technologies in a reasonable area on the ASIC periphery. In addition, we present a design to efficiently handle the variable data from the stream of parallel compressors. The architecture includes user-selectable lossy and lossless compression blocks. The impact of lossy compression algorithms is evaluated on simulated and experimental X-ray ptychography datasets. This architecture is a practical approach to increase pixel detector frame rates towards the continuous 1 MHz regime for not only coherent imaging techniques such as ptychography, but also for other diffraction techniques at X-ray light sources.

47 OTHER INSTRUMENTATION↗

Forcing data (CESM2/CMIP6) for projection of drought impacts (2015-2100) at the K34 site in Manaus, Brazil

Historical and projected output data variables extracted and derived from Community Earth System Model 2 (CESM2) runs from the Coupled Model Intercomparison Project Phase 6 (CMIP6) archive. CESM2 is a fully coupled Earth system model used in simulations of Earth's past, present, and future climates (Danabasoglu et al., 2020). Variables in this dataset are in six-hourly resolution, and include air temperature (both in K and ℃), specific humidity (kg kg-1), air pressure (Pa), relative humidity (%), and vapor pressure deficit (kPa). CESM2 was the only CMIP6 model that provided VPD at the high temporal resolution required for this analysis. Data are included in .csv files, and the text file CESM2-CMIP6_forcing_K34-Manaus_headers.txt provides descriptions of data file headers.

54 ENVIRONMENTAL SCIENCES↗

Models, data, and scripts associated with “Prediction of Distributed River Sediment Respiration Rates using Community-Generated Data and Machine Learning”

This data package is associated with the publication “Prediction of Distributed River Sediment Respiration Rates using Community-Generated Data and Machine Learning’’ submitted to the Journal of Geophysical Research: Machine Learning and Computation (Scheibe et al. 2024). River sediment respiration observations are expensive and labor intensive to obtain and there is no physical model for predicting this quantity. The Worldwide Hydrobiogeochemisty Observation Network for Dynamic River Systems (WHONDRS) observational data set (Goldman et al.; 2020) is used to train machine learning (ML) models to predict respiration rates at unsampled sites. This repository archives training data, ML models, predictions, and model evaluation results for the purposes of reproducibility of the results in the associated manuscript and community reuse of the ML models trained in this project. One of the key challenges in this work was to find an optimum configuration for machine learning models to work with this feature-rich (i.e. 100+ possible input variables) data set. Here, we used a two-tiered approach to managing the analysis of this complex data set: 1) a stacked ensemble of ML models that can automatically optimize hyperparameters to accelerate the process of model selection and tuning and 2) feature permutation importance to iteratively select the most important features (i.e. inputs) to the ML models. The major elements of this ML workflow are modular, portable, open, and cloud-based, thus making this implementation a potential template for other applications. This data package is associated with the GitHub repository found at Please see the file level metadata (flmd; “sl-archive-whondrs_flmd.csv”) for a list of all files contained in this data package and descriptions for each. Please see the data dictionary (dd; “sl-archive-whondrs_dd.csv”) for a list of all column headers contained within comma separated value (csv) files in this data package and descriptions for each. The GitHub repository is organized into five top-level directories: (1) “input_data” holds the training data for the ML models; (2) “ml_models” holds machine learning models trained on the data in “input_data”; (3) “scripts” contains data preprocessing and postprocessing scripts and intermediate results specific to this data set that bookend the ML workflow; (4) “examples” contains the visualization of the results in this repository including plotting scripts for the manuscript (e.g., model evaluation, FPI results) and scripts for running predictions with the ML models (i.e., reusing the trained ML models); (5) “output_data” holds the overall results of the ML model on that branch. Each trained ML model resides on its own branch in the repository; this means that inputs and outputs can be different branch-to-branch. Furthermore, depending on the number of features used to train the ML models, the preprocessing and postprocessing scripts, and their intermediate results, can also be different branch-to-branch. The “main-*” branches are meant to be starting points (i.e. trunks) for each model branch (i.e. sprouts). Please see the Branch Navigation section in the top-level README.md in the GitHub repository for more details. There is also one hidden directory “.github/workflows”. This hidden directory contains information for how to run the ML workflow as an end-to-end automated GitHub Action but it is not needed for reusing the ML models archived here. Please the top-level README.md in the GitHub repository for more details on the automation.

13C↗