Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “evaluation datasets”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 253 records · Page 14

CausalFS

SAND2021-4723 O CausalFS is a Python package that evaluates causal relationships between dataset features and the metric to be predicted. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Nguyen, Bernard↗

HydraGNN_GFM_FineTuning4Materials v1.0

This repository enables fine-tuning of the HydraGNN Predictive GFM 2026 — an open-source ensemble of pre-trained graph foundation models for atomistic materials modeling, developed at Oak Ridge National Laboratory. The GFM 2026 is freely available and downloadable via Globus from the OLCF Data Constellation (DOI: 10.13139/OLCF/2562660). Starting from these pre-trained weights, this repository provides a complete transfer learning pipeline for adapting the GFM ensemble to domain-specific molecular and materials property prediction tasks. It includes: 1) Utilities for ensemble fine-tuning with task-specific output heads 2) Example pipelines for eight widely-used materials and molecular datasets 3) Tools for model adaptation and head configuration 4) Data preprocessing utilities for each supported dataset 5) Benchmarking and evaluation scripts

Ungerboeck, Linda↗

VECTOR Phase 1 Dataset: CAV Trajectory and Energy Consumption Records

This dataset contains benchmark experimental data from Phase 1 of the VECTOR project, focusing on the energy impact of CAV hardware components. The dataset includes vehicle trajectory data (speed and position) and corresponding energy consumption records collected from a CAV platform equipped with lidar, cameras, onboard computation units, and communication modules. The primary objective is to quantify the baseline energy consumption attributable to sensing and computing systems, independent of any advanced cooperative control strategies. During experiments, the leading vehicle followed a predetermined velocity profile, and the following CAV mirrored this trajectory using a basic car-following control to ensure consistent driving behavior. This setup enables a reliable benchmark for assessing the energy cost introduced by onboard CDA hardware (e.g., lidar and GPU-based processing). The dataset is essential for evaluating energy baselines and supports future comparative studies involving additional cooperative strategies. ![system img](system.png) ![vector img](vector.png)

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Depth-resolved sagebrush root metabolomics, rhizosphere microbial communities, and geochemistry at the East River Watershed

This data set consists of results from soil nutrient profile, untargeted metabolomics, mass spec imaging, and amplicon sequencing. Data for soil nutrient profile includes common cations (Ca, Mg, Na, and K etc.) extracted from 3 digesting steps – ammonia acetate (for exchangeable cations), nitric acid (for acid dissolved fraction), and hydrofluoric acid/perchloric acid (HF/HClO4) for whole soil digestion. It also includes concentration of organic carbon, inorganic nitrogen (ammonia and nitrate) and phosphorus (Bray-1 P and nitric acid extract), and total nitrogen and phosphorus. Data for untargeted metabolomics includes metabolomic profile for root exudate/tissues and soil extracts from depths at surface soil to saprolite, that were measured using gas chromatography – mass spectrometry (GC-MS), and liquid chromatography – tandem mass spectrometry (LC-MS/MS). Data for mass spec imaging includes spatial distribution of metabolites that were detected and annotated with Fourier transformation ion cyclotron resonance mass spectrometer (FTICR-MS). Data for amplicon sequencing includes the base paired 16S and ITS ribosomal RNA sequences from Miseq Illumina sequencing. All samples were collected from 2 sampling campaign October 2022 and June 2023. Collectively, these datasets enable a mechanistic evaluation of how nutrient acquisition, especially nitrogen and phosphorus, differs between shallow roots operating in soil and deep roots functioning within the fractured bedrock zone. All files are provided as comma-separated values (CSV) fies (.csv) and (GZIP) file (.gz). The compressed .gz FASTQ files can be read directly in R using the dada2 package as part of the amplicon sequence analysis workflow. This work was supported by the Watershed Function Science Focus Area at Lawrence Berkeley National Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-05CH11231. This research was performed on a project award 60563 (https://dx.doi.org/10.46936/expl.proj.2022.60563/60008727) from the Environmental Molecular Sciences Laboratory, a DOE Office of Science User Facility sponsored by the Biological and Environmental Research program under Contract No. DE-AC05-76RL01830.

EARTH SCIENCE > AGRICULTURE > SOILS > CARBON↗

Towards Query-Efficient Black-Box Adversary with Zeroth-Order Natural Gradient Descent

Despite the great achievements of the modern deep neural networks (DNNs), the vulnerability/robustness of state-of-the-art DNNs raises security concerns in many application domains requiring high reliability. Various adversarial attacks are proposed to sabotage the learning performance of DNN models. Among those, the black-box adversarial attack methods have received special attentions owing to their practicality and simplicity. Black-box attacks usually prefer less queries in order to maintain stealthy and low costs. However, most of the current black-box attack methods adopt the first-order gradient descent method, which may come with certain deficiencies such as relatively slow convergence and high sensitivity to hyper-parameter settings. In this paper, we propose a zeroth-order natural gradient descent (ZO-NGD) method to design the adversarial attacks, which incorporates the zeroth-order gradient estimation technique catering to the black-box attack scenario and the second-order natural gradient descent to achieve higher query efficiency. The empirical evaluations on image classification datasets demonstrate that ZO-NGD can obtain significantly lower model query complexities compared with state-of-the-art attack methods.

Zhao, Pu↗

JetNet: A Python package for accessing open datasets and benchmarking machine learning methods in high energy physics

JetNet is a Python package that aims to increase accessibility and reproducibility for machinelearning (ML) research in high energy physics (HEP), primarily related to particle jets. Basedon the popular PyTorch ML framework, it provides easy-to-access and standardized interfacesfor multiple heterogeneous HEP datasets and implementations of evaluation metrics, lossfunctions, and more general utilities relevant to HEP.

97 MATHEMATICS AND COMPUTING↗

A library of AI-assisted FAIR water cycle and related disturbance datasets to enable model training, parameterization and validation

This whitepaper is responsive to focal area Data acquisition and assimilation enabled by machine learning, AI, and advanced methods. Here we describe how FAIR (Findable, Accessible, Reusable, Interoperable) datasets related to water cycle extremes are essential for successful implementation of ML in Earth System and other models. We also describe how AI can be used to acquire and integrate water cycle data related to extreme events to create a library of FAIR datasets for training and evaluating algorithms.

58 GEOSCIENCES↗

Seismic Signal Detection on International Monitoring System 3-Component Stations using PhaseNet

In this report we discuss training a deep learning seismic signal detection model on 3-component stations from the International Monitoring System (IMS) using the PhaseNet architecture. Using 14 years of associated signals from the International Data Centre’s (IDC) Late Event Bulletin (LEB), we auto-curated training data consisting of signal windows containing associated arrivals, and noise windows that contain no LEB-associated signals. We trained several models using different waveform window durations (30 seconds and 100 seconds), with and without bandpass filtering. We evaluated the effectiveness of our models using associated signals from the Unconstrained Global Event Bulletin (UGEB) and found that several of our models outperformed the signal detections from the IDC’s Selected Event List 3 (SEL3) arrival table. The SEL3 bulletin evaluated on the UGEB dataset with 100-second waveform windows registered a precision and recall of .15 and .48, respectively, versus .19 and .59 for our filtered-data model. For the 30-second waveform window dataset, the SEL3 bulletin achieved a precision and recall of .31 and .47, respectively, versus .32 and .60 for our filtered-data model. Finally, our models detected signals from all source-to-receiver distances, suggesting it is feasible to use a single PhaseNet model for the IMS network.

58 GEOSCIENCES↗

Machine learning guided selection of broad-spectrum epitope-specific functional antibodies for "Disease X"

Our project established and demonstrated a transfer learning framework that enables prediction of antibody–antigen interactions across related viruses. The approach focused on three major activities: 1. Conserved region and epitope identification – We compared viral protein structures and sequences to identify shared receptor-binding domains and neutralizing epitope regions across variants and related viruses. These conserved features formed the foundation for discovering broadly functional antibodies. 2. Machine learning model development – We built neural network–based models that integrate epitope features with antibody sequence information. Instead of relying solely on structural or physical properties, the models learned transferable patterns that describe antibody binding potential across different viral families. 3. Transfer learning and validation – Using SARS-CoV-2 and Ebola as source systems, we successfully transferred learned epitope features to predict antibody interactions for SARS CoV-1 and Marburg virus. Iterative cycles of dataset generation, retraining, and evaluation improved generalization and predictive power, ensuring the framework can adapt to new threats.

59 BASIC BIOLOGICAL SCIENCES↗

Comparison of metagenomes from fermentation of various agroindustrial residues suggests a common model of community organization

The liquid residue resulting from various agroindustrial processes is both rich in organic material and an attractive source to produce a variety of chemicals. Using microbial communities to produce chemicals from these liquid residues is an active area of research, but it is unclear how to deploy microbial communities to produce specific products from the different agroindustrial residues. To address this, we fed anaerobic bioreactors one of several agroindustrial residues (carbohydrate-rich lignocellulosic fermentation conversion residue, xylose, dairy manure hydrolysate, ultra-filtered milk permeate, and thin stillage from a starch bioethanol plant) and inoculated them with a microbial community from an acid-phase digester operated at the wastewater treatment plant in Madison, WI, United States. The bioreactors were monitored over a period of months and sampled to assess microbial community composition and extracellular fermentation products. We obtained metagenome assembled genomes (MAGs) from the microbial communities in each bioreactor and performed comparative genomic analyses to identify common microorganisms, as well as any community members that were unique to each reactor. Collectively, we obtained a dataset of 217 non-redundant MAGs from these bioreactors. This metagenome assembled genome dataset was used to evaluate whether a specific microbial ecology model in which medium chain fatty acids (MCFAs) are simultaneously produced from intermediate products (e.g., lactic acid) and carbohydrates could be applicable to all fermentation systems, regardless of the feedstock. MAGs were classified using a multiclass classification machine learning algorithm into three groups, organisms fermenting the carbohydrates to intermediate products, organisms utilizing the intermediate products to produce MCFAs, and organisms producing MCFAs directly from carbohydrates. This analysis revealed common biological functions among the microbial communities in different bioreactors, and although different microorganisms were enriched depending on the agroindustrial residue tested, the results supported the conclusion that the microbial ecology model tested was appropriate to explain the MCFA production potential from all agricultural residues.

60 APPLIED LIFE SCIENCES↗

Characterization of organic aerosol across the global remote troposphere: a comparison of ATom measurements and global chemistry models

The spatial distribution and properties of submicron organic aerosol (OA) are among the key sources of uncertainty in our understanding of aerosol effects on climate. Uncertainties are particularly large over remote regions of the free troposphere and Southern Ocean, where very few data have been available and where OA predictions from AeroCom Phase II global models span 2 to 3 orders of magnitude, greatly exceeding the model spread over source regions. The (nearly) pole-to-pole vertical distribution of non-refractory aerosols was measured with an aerosol mass spectrometer onboard the NASA DC-8 aircraft as part of the Atmospheric Tomography (ATom) mission during the Northern Hemisphere summer (August 2016) and winter (February 2017). This study presents the first extensive characterization of OA mass concentrations and their level of oxidation in the remote atmosphere. OA and sulfate are the major contributors by mass to submicron aerosols in the remote troposphere, together with sea salt in the marine boundary layer. Sulfate was dominant in the lower stratosphere. OA concentrations have a strong seasonal and zonal variability, with the highest levels measured in the lower troposphere in the summer and over the regions influenced by biomass burning from Africa (up to 10 µg sm -3 ). Lower concentrations (~0.1–0.3 µg sm -3 ) are observed in the northern middle and high latitudes and very low concentrations (<0.1 µg sm -3 ) in the southern middle and high latitudes. The ATom dataset is used to evaluate predictions of eight current global chemistry models that implement a variety of commonly used representations of OA sources and chemistry, as well as of the AeroCom-II ensemble. The current model ensemble captures the average vertical and spatial distribution of measured OA concentrations, and the spread of the individual models remains within a factor of 5. These results are significantly improved over the AeroCom-II model ensemble, which shows large overestimations over these regions. However, some of the improved agreement with observations occurs for the wrong reasons, as models have the tendency to greatly overestimate the primary OA fraction and underestimate the secondary fraction. Measured OA in the remote free troposphere is highly oxygenated, with organic aerosol to organic carbon (OA/OC) ratios of ~2.2–2.8, and is 30 %–60 % more oxygenated than in current models, which can lead to significant errors in OA concentrations. The model–measurement comparisons presented here support the concept of a more dynamic OA system as proposed by Hodzic et al. (2016), with enhanced removal of primary OA and a stronger production of secondary OA in global models needed to provide better agreement with observations.

54 ENVIRONMENTAL SCIENCES↗

Climate Model Output Rewriter

The Climate Model Output Rewriter (CMOR) software was first developed by LLNL’s PCMDI program in early 2000s and was formally released with v1.0 (July 2006), v2.0 (January 2011), and v3.1(June 2016). CMOR is used to produce Climate and Forecast Convention (http://cfconventions.org/) CF-compliant netCDF files, in the standard format required to satisfy the World Climate Research Program (WCRP) Coupled Model Intercomparison Project (CMIP). The software has been used across multiple phases of the Earth System Modeling (ESM) project CMIP (CMIP3, CMIP5, CMIP6, and planned use in CMIP7) along with numerous parallel projects focused on preparation observations for use in model evaluation (obs4MIPs) and forcing datasets (input4MIPs) to guide ESM simulations to meet strict experimental protocols. More information can be obtained from the CMOR website and code repositories: https://cmor.llnl.gov/; https://github.com/pcmdi/cmor; https://github.com/PCMDI/cmor3_documentation The ESM variable definitions used as input for CMOR can also be viewed in code repositories: https://github.com/PCMDI/cmip3-cmor-tables/; https://github.com/PCMDI/cmip5-cmor-tables/; https://github.com/PCMDI/cmip6-cmor-tables/

Mauzey, ChristopherF↗

Synthesis of Multispectral Bands from Hyperspectral Data: Validation Based on Images Acquired by AVIRIS, Hyperion, ALI, and ETM+

Multispectral data requirements for Earth science applications are not always studied rigorously studied before a new remote sensing system is designed. A study of the spatial resolution, spectral bandpasses, and radiometric sensitivity requirements of real-world applications would focus the design onto providing maximum benefits to the end-user community. To support systematic studies of multispectral data requirements, the Applications Research Toolbox (ART) has been developed at NASA's Stennis Space Center. The ART software allows users to create and assess simulated datasets while varying a wide range of system parameters. The simulations are based on data acquired by existing multispectral and hyperspectral instruments. The produced datasets can be further evaluated for specific end-user applications. Spectral synthesis of multispectral images from hyperspectral data is a key part of the ART software. In this process, hyperspectral image cubes are transformed into multispectral imagery without changes in spatial sampling and resolution. The transformation algorithm takes into account spectral responses of both the synthesized, broad, multispectral bands and the utilized, narrow, hyperspectral bands. To validate the spectral synthesis algorithm, simulated multispectral images are compared with images collected near-coincidentally by the Landsat 7 ETM+ and the EO-1 ALI instruments. Hyperspectral images acquired with the airborne AVIRIS instrument and with the Hyperion instrument onboard the EO-1 satellite were used as input data to the presented simulations.

Blonksi, Slawomir↗

Toward a Physical Characterization of Raindrop Collision Outcome Regimes

A comprehensive raindrop collision outcome regime diagram that delineates the physical conditions associated with the outcome regimes (i.e., bounce, coalescence, and different breakup types) of binary raindrop collisions is proposed. The proposed diagram builds on a theoretical regime diagram defined in the phase space of collision Weber numbers We and the drop diameter ratio p by including critical angle of impact considerations. In this study, the theoretical regime diagram is first evaluated against a comprehensive dataset for drop collision experiments representative of raindrop collisions in nature. Subsequently, the theoretical regime diagram is modified to explicitly describe the dominant regimes of raindrop interactions in (We, p) by delineating the physical conditions necessary for the occurrence of distinct types of collision-induced breakup (neck/filament, sheet, disk, and crown breakups) based on critical angle of impact consideration. Crown breakup is a subtype of disk breakup for lower collision kinetic energy that presents distinctive morphology. Finally, the experimental results are analyzed in the context of the comprehensive collision regime diagram, and conditional probabilities that can be used in the parameterization of breakup kernels in stochastic models of raindrop dynamics are provided.

Testik, F. Y.↗

Influence of Precipitation Forcing Uncertainty on Hydrological Simulations with the NASA South Asia Land Data Assimilation System

Accurate meteorological estimates are critical for process-based hydrological simulationand prediction. This presents a significant challenge in mountainous Asia where in situmeteorological stations are limited and major river basins cross international borders. In thiscontext, remotely sensed and model-derived meteorological estimates are often necessary inputsfor distributed hydrological analysis. However, these datasets are difficult to evaluate on accountof limited access to ground data. In this case, the implications of uncertainty associated withprecipitation forcing for hydrological simulations is explored by driving the South Asia Land DataAssimilation System (South Asia LDAS) using a range of meteorological forcing products.MERRA2, GDAS, and CHIRPS produce a wide range of estimates for rainfall, which causes awidespread simulated streamflow and evapotranspiration. A combination of satellite-derived andlimited in situ data are applied to evaluate model simulations and, by extension, to constrain theestimates of precipitation. The results show that available gridded precipitation estimates based onin situ data may systematically underestimate precipitation in mountainous regions and thatperformance of gridded satellite-derived or modeled precipitation estimates varies systematicallyacross the region. Since no station-based data or product including station data is satisfactoryeverywhere, our results suggest that the evaluation of the hydrological simulation of streamflowand ET can be used as an indirect evaluation of precipitation forcing based on ground-basedproducts or in-situ data. South Asia LDAS produces reasonable evapotranspiration and streamflowwhen forced with appropriate meteorological forcing and the choice of meteorological forcingshould be made based on the geographical location as well as on the purpose of the simulations.

South Asia land data assimilation system (South As↗

Precipitation over the U.S. Coastal Land/Water Using Gauge-Corrected Multi-Radar/Multi-Sensor System and Three Satellite Products

The weather and climate over the coastal regions have received increasing attention because of substantial population growth, the rising sea level, and extreme weather. Satellite remote sensing provides global precipitation estimates (including coastal land/ocean). While these datasets have been extensively evaluated over land, they have rarely been assessed over coastal ocean. As precipitation radars cover both coastal land and ocean, we used the Multi-Radar/Multi-Sensor System (MRMS) gauge-corrected precipitation product from 2018 to 2020 to evaluate three widely used satellite-based precipitation products over the U.S. coastal land versus the ocean (and the water over the Great Lakes). These products included the Integrated Multi-satellite Retrievals for GPM (IMERG), Precipitation Estimation from Remotely Sensed Information using Artificial Neural Networks (PERSIANN), and Climate Prediction Center Morphing technique (CMORPH). The MRMS data showed a precipitation climatology difference between the coastal land and the ocean that was higher in the winter and lower in the summer and autumn. IMERG and CMORPH performed best over land and water, respectively, while PERSIANN was the most consistent in its performance over land versus water. Heavy precipitation was overestimated by the three products, with larger overestimates over water than over land. These results were not affected by the MRMS uncertainties due to the gauge correction or by the use of different versions.

coastal precipitation↗

GOLEM: GOld standard for Learning and Evaluation of Motifs

Motifs are distinctive, recurring, widely used idiom-like words or phrases, often originating from folklore, whose meaning is anchored in a narrative and have a significance as communicative devices across a wide range of media, including news, literature, and propaganda. Many motifs concisely imply a large constellation of culturally relevant information, and their broad usage suggests their cognitive importance as touchstones of cultural knowledge. As such, their detection is a step towards culturally aware natural language processing. We present GOLEM (GOld standard for Learning and Evaluation of Motifs) a dataset of English news articles, opinion pieces, and broadcast transcripts annotated for motific information. The dataset identifies 25,737 motif candidates across 34 motif types drawn from three cultural or national groups: Jewish, Irish, and Puerto Rican. The dataset contains 2,024,141 words split into 25,737 text snippets drawn from 8,073 articles. Each motif candidate is labeled according to a scheme which identifies the type of usage (motific, referential, eponymic, or unrelated), resulting in 1,743 actual motific instances in the data. Annotation was performed by individuals identifying as members of each group and achieved a Fleiss’ kappa (?) of > 0.55. In addition to the data, we demonstrate that classification of the candidate type is a challenging task for Large Language Models (LLMs) using a few-shot approach; recent models such as T5, FLAN-T5, GPT-2, and Llama 2 (7B) achieved a performance of 41% accuracy at best, where the majority class accuracy is 41% and the average chance accuracy is 27%. These data will support development of new models and approaches for detecting (and reasoning about) motific information in text.

motif, culture, natural language, artificial intel↗

On the feasibility of using physics-informed machine learning for underground reservoir pressure management

In this work, we evaluate the feasibility of using physics-informed machine learning (PIML) for underground energy-related pressure management. To this end, we develop a PIML framework to manage underground reservoir pressures by training neural networks to determine fluid extraction rates for dedicated extraction wells during fluid injection operations given a range of reservoir conditions (e.g., transmissivity and storativity). We implement an automatically-differentiable analytical physics model of fluid flow in porous media within the PIML framework as a proxy for more complicated models. This allows us to execute a sufficient number of training scenarios to fully evaluate the feasibility of using PIML to support pressure management activities. We quantify the number of physics-model parameters required for automatic differentiation to become more efficient than finite-difference gradient calculations. We use a simple scenario with a single injector, extractor, and critical location for our feasibility analysis. We evaluate the effect of the size of the training dataset (i.e., the number of reservoir condition samples) on the accuracy and efficiency of the PIML framework. For an equivalent number of model evaluations, the larger training dataset took less time to train and produced a neural network that was able to more accurately manage reservoir pressures. We also evaluate the effect of the training dataset batch size (i.e., number of reservoir condition samples used to update the neural network coefficients during training; i.e., how the training dataset is partitioned). While training ran faster with larger batch sizes, they produced neural networks that managed pressures less accurately. We demonstrate the approach on a more complex scenario involving 10 injectors, 10 extractors, and 4 critical locations (a relatively high well density of 20 wells/km2). We provide the number of forward and adjoint model evaluations required in each case as an indication of the feasibility of using PIML for pressure management when more complicated physics models with longer execution times are used.

54 ENVIRONMENTAL SCIENCES↗