Engineering PapersSearch

SEARCH · Engineering Papers

Results for “data set”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Transcriptomic data sets for Novosphingobium aromaticivorans DSM12444 and a ΔSARO_RS14285 mutant grown in the presence of glucose and either protocatechuic, vanillic, syringic, or 4-coumaric acid

The SARO_RS14285 gene, encoding a transcription factor, was deleted in Novosphingobium aromaticivorans DSM12444. The transcriptomes of the parent and ΔSARO_RS14285 strains were determined when grown in medium containing glucose with or without protocatechuic, vanillic, syringic, or 4-coumaric acid. We present the raw RNA sequencing data obtained from these cultures.

Novosphingobium aromaticivorans

Transcriptomic data sets for Novosphingobium aromaticivorans grown with the β-5-linked aromatic dimer dehydrodiconiferyl alcohol and the related G-aromatic monomers vanillin and ferulic acid

ABSTRACT The transcriptomes of a 2-pyrone-4,6-dicarboxylic acid-producing strain of Novosphingobium aromaticivorans DSM12444 were determined when grown in minimal medium containing glucose alone or glucose plus vanillin, ferulic acid, or the β-5-linked aromatic dimer dehydrodiconiferyl alcohol as carbon sources. Here, we present the RNA-sequencing data we obtained.

Metz, Fletcher

Exploring Data Set Bias and Decision Support with Predictive Uncertainty Through Bayesian Approximations and Convolutional Neural Networks

Individual seismic catalogs can contain multiscale observations from fault level to global scales and associated waveforms from discrete events reflect crustal structure across many different scales and locations. Seismic network aperture, geographic location, and observation distance may not provide informative guidance or intuition on how different catalogs will behave across models trained under different conditions. We rely on uncertainty to provide guardrails for when to trust model decisions, but understanding when our uncertainty is trustworthy is an open challenge. Here, in this work, we explore Bayesian approximation methods for assigning predictive uncertainty in seismic event classification problems. We find that computationally expensive Bayesian approximations do not outperform simple ensemble methods. We also find that when exploiting multiple seismic event catalogs, joint training with data from all the catalogs combined with Bayesian approximations and supervised training for classification can obscure bias and result in less robust uncertainty while also not providing substantial performance benefits compared to training individual models for each catalog.

58 GEOSCIENCES

A Cloud-Tracking Data Set for the CSAPR2 Adaptive Scanning during TRACER

The U.S. Department of Energy (DOE) Atmospheric Radiation Measurement (ARM) User Facility (Mather and Voyles 2013) deployed the first ARM Mobile Facility (AMF1; Miller et al. 2016) near LaPorte, Texas to support the Tracking Aerosol Convection Interactions Experiment (TRACER) (Jensen et al. 2025) near Houston, Texas. From October 2021 to September 2022, AMF1 was deployed to 29.67° N, 95.06° W near LaPorte, Texas and the 2nd Generation C-band Scanning ARM Precipitation Radar (CSAPR2) was deployed to a supplementary site at 29.53° N, 95.28° W (Figure 1). During an intensive operational period (IOP) from 1 June to 30 September 2022, the CSAPR2 sampled precipitation echoes in an adaptive scanning mode following the Multisensor Agile Adaptive Scanning (MAAS) framework (Kollias et al. 2020). MAAS helped optimize the CSAPR2 scan strategy to perform frequent plan position indicator (PPI) and range height indicator (RHI) scans (Lamer et al. 2023). Details of the CSAPR2 scanning, data processing, and calibration procedures used by the principal investigator (PI), and the PI data files are described by Oue et al. (2023). Details of the CSAPR2 operational performance, ARM data processing and correction procedures, and data quality masks are described by Feng et al. (2024a).

54 ENVIRONMENTAL SCIENCES

Integrase-On-Demand-Pipeline Data Set

Files needed to run the Integrase-On-Demand-Pipeline, a program designed to provide users with a list of putative attachment site and integrase pairs for a prokaryotic genome of interest. isles.pkl: Serialized python-object file, containing a dictionary of attachment site sequences and reference genomic island information extracted from the Genomic island database ints.gff: Gene format file containing annotations for all integrases referenced in isles.pkl. The source genome, gene coordinates, integrase name, protein IDs and amino acid sequence included. reps.msh: Binary file containing 1000 128-bit MurmurHash3 hashes for >80,000 genomes

McClain, Hannah Marie [Sandia National Laboratorie

Advanced Materials & Manufacturing Technology (AMMT): Development of Additive Manufacturing Agnostic Process Parameter Procedure, 316H Stainless Steel Readiness Level Data Sets, and Machine Maintenance Plan

The University of California, Davis is involved in a project to deploy and enhance an artificial intelligence (AI) system for predicting and preventing plasma disruptions on the DIII D tokamak, under the funding from Department of Energy DE-SC0023500 (title: AI/Deep Learning FRNN Software for Prediction & Real-Time Control of DIII-D Plasma Control System (PCS)). The overarching goal is to demonstrate that real-time, AI-guided intervention can proactively modify the plasma state to avoid or mitigate disruptions—a critical challenge for the future of fusion energy.

36 MATERIALS SCIENCE

The ARM Precipitation Best Estimate (PrecipBE) Value-Added Product Report

The U.S. Department of Energy Atmospheric Radiation Measurement (ARM) User Facility’s Precipitation Best-Estimate (PrecipBE) Value-Added Product integrates multiple precipitation datastreams, accounting for data quality and instrument limitations, to deliver comprehensive per-precipitation event properties alongside ancillary ARM data set data. PrecipBE bundles all valid surface rainfall samples into artificial intelligence (AI)-ready tabular and time-series formats, reporting bundle means and uncertainty ranges. This per-event structure provides an insightful and easy-to-use resource for researchers analyzing precipitation characteristics.

54 ENVIRONMENTAL SCIENCES

SPRUCE Vegetation Phenology in Experimental Plots from PhenoCam Imagery, 2015-2024

This data set consists of PhenoCam data from the SPRUCE experiment from the beginning of whole ecosystem warming (Hanson et al. 2017) in August 2015 through March 31 of 2025 (2015-08-24 to 2025-03-31), with start- and end-of-season phenological transition dates derived through the end of autumn 2024. Digital cameras, or phenocams, installed in each SPRUCE enclosure track seasonal variation in vegetation “greenness”, a proxy for vegetation phenology and associated physiological activity. Three separate regions of interest (ROIs) were defined for each camera field of view, corresponding to different vegetation types and demarcating (1) Picea trees (vegetation type EN, for evergreen needleleaf); (2) Larix trees (vegetation type DN, for deciduous needleleaf); and (3) the mixed shrub layer (vegetation type SH). This data set consists of three sets of data files: (1) 3-day summary product files: One file for each camera and each ROI (i.e. vegetation type), characterizing vegetation color at a 3-day time step. • Contains 36 files in *.csv format inside a compressed (*.zip) file. (2) Transition date file: Estimates “greenness rising” (spring) and “greenness falling” (autumn) transition dates derived from the smoothed daily green chromatic coordinate (GCC) values, for each camera and each ROI (i.e., vegetation type). • Contains one file in *.csv format. (3) Snow flag files: Indicate days with snow on trees or snow on ground for each experimental enclosure. • Contains two files in *.csv format, one for snow on trees and one for snow on ground. This data set consists of two sets of companion files: (1) Accompanying HTML files show the 90th quantiles of the mean GCC plotted together with transition dates for each vegetation type and plot. • Contains three files in HTML format, one for each vegetation type. • One additional file in HTML format with the transition dates plotted for each vegetation type, by year. (2) R files for processing PhenoCam files and flags. • Contains five files in R file(*.R) format and the components of the phenocamr package (Version 1.1.4) used for calculating transition dates for 2015-2024. These are contained in a compressed (*.zip) file. User Note: All imagery is posted in near-real time to the PhenoCam Project web page (https://phenocam.nau.edu), where it is publicly available. Scroll to “spruce” in the Gallery or link directly to the 29 SPRUCE cameras at https://tinyurl.com/sprucecams. This data set is based on the complete camera record from SPRUCE and supersedes all previously released PhenoCam datasets (see Related Data Sets). The estimated transition dates for previously released datasets may differ slightly (in most cases, by ±3 days or less), because following standard PhenoCam processing protocols (Richardson et al. 2018, Scientific Data), smoothing and interpolation, outlier removal, and transition date estimation are always conducted using the full data record.

54 ENVIRONMENTAL SCIENCES

Non-linear relationships between daily temperature extremes and US agricultural yields uncovered by global gridded meteorological datasets

Global agricultural commodity markets are highly integrated among major producers. Prices are driven by aggregate supply rather than what happens in individual countries in isolation. Furthermore, estimating the effects of weather-induced shocks on production, trade patterns and prices hence requires a globally representative weather data set. Recently, two data sets that provide daily or hourly records, GMFD and ERA5-Land, became available. Starting with the US, a data rich region, we formally test whether these global data sets are as good as more fine-scaled country-specific data in explaining yields and whether they estimate similar response functions. While GMFD and ERA5-Land have lower predictive skill for US corn and soybeans yields than the fine-scaled PRISM data, they still correctly uncover the underlying non-linear temperature relationship. All specifications using daily temperature extremes under any of the weather data sets outperform models that use a quadratic in average temperature. Correctly capturing the effect of daily extremes has a larger effect than the choice of weather data. In a second step, focusing on Sub Saharan Africa, a data sparse region, we confirm that GMFD and ERA5-Land have superior predictive power to CRU, a global weather data set previously employed for modeling climate effects in the region.

54 ENVIRONMENTAL SCIENCES

Ab Initio-Based Bond Order Potential for Arsenene Polymorphs Developed via Hierarchical Reinforcement Learning

Arsenene, a less-explored two-dimensional material, holds the potential for applications in wearable electronics, memory devices, and quantum systems. This study introduces a bond-order potential model with Tersoff formalism, the ML-Tersoff, which leverages multireward hierarchical reinforcement learning (RL), trained on an ab initio data set. This data set covers a spectrum of properties for arsenene polymorphs, enhancing our understanding of its mechanical and thermal behaviors without the complexities of traditional models requiring multiple parameter sets. Our RL strategy utilizes decision trees coupled with a hierarchical reward strategy to accelerate convergence in high-dimensional continuous search spaces. Unlike the Stillinger-Weber approach, which demands separate formalisms for buckled and puckered forms, the ML-Tersoff model concurrently captures multiple properties of the two polymorphs by effectively representing the local environment, thereby avoiding the need for different atomic types. Here, we apply the ML model to understand the mechanical and thermal properties of the arsenene polymorphs and nanostructures. We observe an inverse relationship between the critical strain and temperature in arsenene. Thermal conductivity calculations in nanosheets show good agreement with ab initio data, reflecting a decrease in thermal conductivity attributable to increased anharmonic effects at higher temperatures. We also apply the model to predict the thermal behavior of arsenene nanotubes.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH

Ligand-Based Compound Activity Prediction via Few-Shot Learning

Predicting the activities of new compounds against biophysical or phenotypic assays based on the known activities of one or a few existing compounds is a common goal in early stage drug discovery. This problem can be cast as a “few-shot learning” challenge, and prior studies have developed few-shot learning methods to classify compounds as active versus inactive. However, the ability to go beyond classification and rank compounds by expected affinity is more valuable. We describe Few-Shot Compound Activity Prediction (FS-CAP), a novel neural architecture trained on a large bioactivity data set to predict compound activities against an assay outside the training set, based on only the activities of a few known compounds against the same assay. Our model aggregates encodings generated from the known compounds and their activities to capture assay information and uses a separate encoder for the new compound whose activity is to be predicted. The new method provides encouraging results relative to traditional chemical-similarity-based techniques as well as other state-of-the-art few-shot learning methods in tests on a variety of ligand-based drug discovery settings and data sets.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH

Cambium 2024 Scenario Descriptions and Documentation

The National Renewable Energy Laboratory's (NREL's) Cambium data sets are annually released sets of simulated hourly data for a range of modeled futures of the U.S. electric sector with metrics designed to be useful for long-term decision- making. The 2024 Cambium data set is the fifth annual release. The data sets are a companion product to NREL's Standard Scenarios, which are likewise released annually and are a set of projections of how the U.S. electric sector could evolve across a suite of different potential futures, but covering more scenarios with less temporal granularity. Information about Cambium and related publications can be found at https://www.nrel.gov/analysis/cambium.html, and the Cambium data sets can be viewed and downloaded at https://scenarioviewer.nrel.gov/. In this documentation, we describe Cambium 2024's scenarios, define the metrics, and document the Cambium-specific methods for calculating those metrics.

24 POWER TRANSMISSION AND DISTRIBUTION