Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “evaluation datasets”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Re-evaluating probable maximum precipitation estimates: sensitivity to transposition domains and storm rotation using modern datasets

This study examines the sensitivity of Probable Maximum Precipitation (PMP) estimates to key methodological decisions embedded in the legacy approach adopted in the U.S. National Weather Service Hydrometeorological Reports No. 51 and No. 52. Although widely used for infrastructure design and risk regulation, fundamental aspects of PMP estimation—such as storm sample size, transposition domain, maximization procedures, and storm rotation—remain poorly constrained and lack formal guidance. Using the Red Rock watershed in Iowa as a case study, and leveraging the 2002–2023 NOAA Analysis of Record for Calibration (AORC) precipitation dataset, we systematically evaluate how each methodological choice, individually and in combination, influences PMP estimates. Our findings demonstrate that PMP is not a fixed physical upper bound but rather a modeling construct shaped heavily by user-defined assumptions. Notably, PMP values derived from modern gridded rainfall datasets can be substantially higher than the legacy estimate used in the original spillway design for Red Rock Dam. Decisions regarding storm sample size, domain extent, climatological window, and particularly storm rotation all contributed to higher PMP estimates. Storm rotation alone—a loosely constrained element in the current PMP practice—can amplify PMP by more than 25%. These results reveal the lack of standardized bounds in current PMP workflows and the need for systematic sensitivity and uncertainty analysis. As PMP estimation shifts toward probabilistic approaches, incorporating physically meaningful storm attributes will be key to developing more transparent, defensible methods for dam safety and climate-resilient infrastructure.

Probable maximum precipitation↗

Leveraging large language models to address data scarcity in machine learning for graphene synthesis

Machine learning in experimental materials science faces significant challenges due to the scarcity of data, which are costly and time-consuming to generate, particularly when relying on in-house experiments. Literature data mining offers a potential solution but introduces issues like mixed data quality, inconsistent formats, and non-uniform reporting of synthesis parameters, resulting in partially missing and heterogeneous features across the dataset. Here, we propose data imputation and feature engineering methods that employ pre-trained large language models (LLMs) to enhance machine learning performance on scarce, heterogeneous datasets, demonstrated on graphene CVD synthesis data and the ML-HydPARK hydrogen storage dataset. GPT models perform data imputation via tailored prompting and semantic normalization of inconsistently reported features through embeddings, for example, to harmonize the complex nomenclature of CVD substrates. Beyond yielding more diverse and richer feature representations than traditional methods such as K-nearest neighbors (KNN) and Multivariate Imputation by Chained Equations (MICE), LLM-based data imputation is evaluated against dataset characteristics and prompting strategies. We vary the level of autonomy granted to the LLM, from generic prompting that leverages pre-trained knowledge for autonomous data generation to data-informed prompting that constrains outputs using target-specific information, and demonstrate which level of autonomy yields superior imputation performance across datasets and feature types. The proposed data engineering methods markedly improve downstream performance; for example, in graphene layer number classification using a support vector machine (SVM), binary accuracy increases from 39% to 65% and ternary accuracy from 52% to 72%. Fine-tuning experiments on both datasets show that combining our proposed LLM-based data imputation and feature encoding methods with numerical machine learning predictors outperforms standalone fine-tuned LLM predictors in data-scarce settings. The proposed strategies emphasize data enhancement techniques rather than refining learning architectures or regularizing loss functions, offering a broadly applicable framework for improving machine learning performance on scarce, inhomogeneous datasets.

Chemical vapor deposition↗

High bias machine learning for antineutrino-based safeguards for small reactors

The statistical methods used for antineutrino detection will need to be improved to effectively monitor the inventory of next-generation nuclear reactors. In this sensitivity study, we evaluate machine learning models compared to previously used statistical approaches to identify diversion scenarios in a simulated Advanced Fast Reactor (AFR)-100. A chi-square goodness-of-fit technique, which individually compares the simulated antineutrino yields to the expected antineutrino yield, resulted in precise but low diversion detection probability. Various support vector machine (SVM) models were applied with diverse training datasets to evaluate the robustness of the method towards unexpected or “unseen” diversion scenarios. Furthermore, our results indicate that while the SVM models significantly improved the detection probability of near-field antineutrino-based safeguards, up to a probability of ~0.04, for the simulated small reactor, the detection system still needs improvements to reach the 0.2 detection limit established by the International Atomic Energy Agency.

98 NUCLEAR DISARMAMENT, SAFEGUARDS, AND PHYSICAL P↗

Machine learning and atomic layer deposition: Predicting saturation times from reactor growth profiles using artificial neural networks

In this work, we explore the application of deep neural networks to the optimization of atomic layer deposition (ALD) processes. In particular, we focus on a one-shot optimization problem, where we try to predict the optimal dose time that leads to saturation everywhere in the reactor based on thickness values measured at different points of an ALD reactor after a single trial growth. In order to tackle this problem, we introduce a dataset designed to train neural networks to predict saturation times based on these inputs for a cross-flow ALD reactor. Here, we then explore the predictive ability of artificial neural networks of different depths and sizes using a separate testing dataset to evaluate their accuracies. The results obtained show that networks trained using stochastic gradient descent methods can accurately predict saturation times without requiring any additional information on the surface kinetics. This provides a viable approach to minimize the number of experiments required to optimize new ALD processes in a known reactor, and it highlights the way machine learning can be leveraged for thin film growth and manufacturing. While the datasets and training procedure depend on the reactor geometry, the trained neural networks provide a general surrogate model connecting thickness values and trial dose times with optimal saturation times that can be reused for different ALD processes within the same reactor.

36 MATERIALS SCIENCE↗

Low Latency Flux and Concentration Datasets in Support of Greenhouse Gas Monitoring Based on NASA's GEOS Modeling and Data Assimilation System

We present efforts to develop space-based greenhouse gas monitoring systems that can provide low latency information and traceability to independent observations. Through support from its Carbon Monitoring System program, NASA has developed the capability to assimilate XCO2 retrievals from the Orbiting Carbon Observatory, 2 (OCO-2) into the Goddard Earth Observing System (GEOS) Constituent Data Assimilation System (CoDAS) to create gap-filled, three-dimensional (3D) estimates of CO2 mixing ratio. When OCO-2 data are not available, concentration fields are further informed by a bottom-up flux package based on remotely sensed fire radiative power, nighttime lights, and vegetation reflectance combined with estimates of atmospheric growth rate based on surface in situ data. The 3D nature of this dataset supports evaluation with independent aircraft data, helping to ensure transparency of remotely sensed data products. These quasi-operational data are currently produced 2-3 months behind real time and are distributed via NASA and international dashboard services to a variety of end users. In this presentation, we provide an overview of the system as well as remaining data gaps and modeling challenges. We also highlight the application of this dataset for detecting emissions anomalies associated with COVID-19 and comparing against independent emissions estimates. Finally, we highlight a new NASA initiative called the Earth Information System (EIS), which aims to support open science and applications by leveraging emerging cloud computing capabilities to increase access to NASA’s greenhouse gas datasets, opportunities for co-development, and transparency in methods for analysis and flux attribution.

Lesley Ott↗

The SPoRT-WRF: Evaluating the Impact of NASA Datasets on Convective Forecasts

Short-term Prediction Research and Transition (SPoRT) seeks to improve short-term, regional weather forecasts using unique NASA products and capabilities SPoRT has developed a unique, real-time configuration of the NASA Unified Weather Research and Forecasting (WRF)WRF (ARW) that integrates all SPoRT modeling research data: (1) 2-km SPoRT Sea Surface Temperature (SST) Composite, (2) 3-km LIS with 1-km Greenness Vegetation Fraction (GVFs) (3) 45-km AIRS retrieved profiles. Transitioned this real-time forecast to NOAA's Hazardous Weather Testbed (HWT) as deterministic model at Experimental Forecast Program (EFP). Feedback from forecasters/participants and internal evaluation of SPoRT-WRF shows a cool, dry bias that appears to suppress convection likely related to methodology for assimilation of AIRS profiles Version 2 of the SPoRT-WRF will premier at the 2012 EFP and include NASA physics, cycling data assimilation methodology, better coverage of precipitation forcing, and new GVFs

Zavodsky, Bradley↗

Application of artificial intelligence methods in the international roughness index prediction of rigid and composite pavements: a systematic review

The International Roughness Index (IRI) is a widely adopted metric for quantifying pavement roughness, directly influencing vehicle safety, ride comfort, and overall roadway performance. In recent years, the use of Machine Learning (ML) models for IRI prediction has gained momentum, with the goal of improving the allocation of maintenance and rehabilitation resources by enabling accurate assessments of pavement conditions. Most prior reviews, however, have concentrated on flexible pavements, leaving a notable gap regarding rigid and composite pavements. To address this gap, the present study conducts a systematic review of Artificial Intelligence (AI) methods applied to IRI prediction for rigid and composite pavements. Literature published between 2004 and 2025 is synthesized to highlight prevailing trends, methodological contributions, and directions for future research. Particular attention is given to the types of models employed, the datasets used for training and validation, and the role of input variables and data-processing strategies. Across the included studies, ensemble learning methods (especially gradient boosting variants such as XGBoost), artificial neural networks, and hybrid architectures frequently achieved high predictive skill, with several models reporting test-set coefficients of determination approaching 0.9–0.96, indicating strong potential for capturing the influence of traffic, pavement structure, and climatic factors. Since these results are obtained from heterogeneous datasets and evaluation protocols, they are interpreted qualitatively rather than as strict cross-study rankings. Analysis of input variables revealed that pavement age and initial IRI were included in 91% (21 of 23) and 78% (18 of 23) of studies, respectively. Climatic variables such as the freezing index appeared in 57% (13 of 23), while traffic-related factors were considered in 65% (15 of 23). The findings underscore the importance of standardized, high-quality datasets, such as those from the Long-Term Pavement Performance (LTPP) program, along with data consistency, model interpretability, computational efficiency, and replicability in enhancing IRI prediction. Future research should focus on incorporating input variable selection techniques to identify the most influential predictors, thereby improving accuracy and robustness. Integrating these approaches with advanced non-linear data-driven models, coupled with robust hyperparameter optimization, holds considerable promise for strengthening the reliability of IRI prediction and supporting resilient pavement management strategies.

42 ENGINEERING↗

NASA Downscaling Project: Final Report

A team of researchers from NASA Ames Research Center, Goddard Space Flight Center, the Jet Propulsion Laboratory, and Marshall Space Flight Center, along with university partners at UCLA, conducted an investigation to explore whether downscaling coarse resolution global climate model (GCM) predictions might provide valid insights into the regional impacts sought by decision makers. Since the computational cost of running global models at high spatial resolution for any useful climate scale period is prohibitive, the hope for downscaling is that a coarse resolution GCM provides sufficiently accurate synoptic scale information for a regional climate model (RCM) to accurately develop fine scale features that represent the regional impacts of a changing climate. As a proxy for a prognostic climate forecast model, and so that ground truth in the form of satellite and in-situ observations could be used for evaluation, the MERRA and MERRA - 2 reanalyses were used to drive the NU - WRF regional climate model and a GEOS - 5 replay. This was performed at various resolutions that were at factors of 2 to 10 higher than the reanalysis forcing. A number of experiments were conducted that varied resolution, model parameterizations, and intermediate scale nudging, for simulations over the continental US during the period from 2000 - 2010. The results of these experiments were compared to observational datasets to evaluate the output.

dynamical downsizing↗

NASA Downscaling Project

A team of researchers from NASA Ames Research Center, Goddard Space Flight Center, the Jet Propulsion Laboratory, and Marshall Space Flight Center, along with university partners at UCLA, conducted an investigation to explore whether downscaling coarse resolution global climate model (GCM) predictions might provide valid insights into the regional impacts sought by decision makers. Since the computational cost of running global models at high spatial resolution for any useful climate scale period is prohibitive, the hope for downscaling is that a coarse resolution GCM provides sufficiently accurate synoptic scale information for a regional climate model (RCM) to accurately develop fine scale features that represent the regional impacts of a changing climate. As a proxy for a prognostic climate forecast model, and so that ground truth in the form of satellite and in-situ observations could be used for evaluation, the MERRA and MERRA-2 reanalyses were used to drive the NU-WRF regional climate model and a GEOS-5 replay. This was performed at various resolutions that were at factors of 2 to 10 higher than the reanalysis forcing. A number of experiments were conducted that varied resolution, model parameterizations, and intermediate scale nudging, for simulations over the continental US during the period from 2000-2010. The results of these experiments were compared to observational datasets to evaluate the output.

Ferraro, Robert↗

Self-supervised and multi-fidelity learning for extended predictive soil spectroscopy

Infrared spectroscopy is a cost-effective, non-destructive, and environmentally benign technology that is increasingly recognized as an important solution for meeting the global demand for soil data. While both near-infrared (NIR) and mid-infrared (MIR) diffuse reflectance spectroscopy enable rapid estimation of soil properties, they present a significant trade-off: NIR offers superior scalability and lower operational costs, whereas MIR provides higher analytical fidelity by capturing fundamental molecular vibrations. In this study, we propose a self-supervised, multi-fidelity learning framework designed to bridge this gap. Our approach leverages large-scale MIR spectral libraries to learn a compact, transferable latent representation, into which NIR spectra are subsequently aligned for downstream prediction. The workflow consists of pretraining a latent model on a large MIR library, adapting the representation using a smaller paired NIR–MIR dataset, and evaluating generalization on an independent external test set. Across a range of chemical and physical soil properties, we found that MIR-derived embeddings improved prediction accuracy relative to baseline models that used raw MIR inputs. Predictions derived from the spectrum conversion (NIR to MIR) task did not match the performance of the original MIR spectra but were similar or superior to predictive performance of NIR-only models, suggesting the unified spectral latent space can effectively leverage the larger and more diverse MIR dataset for prediction of soil properties not well represented in current NIR libraries.

54 ENVIRONMENTAL SCIENCES↗

Evaluation of Correction Methods for NASA GeneLab Transcriptomic Datasets

Conducting space biology experiments aboard the International Space Station, particularly those utilizing complex model organisms like mice, is expensive and difficult due to limited crew availability, hardware, and space. As a result, sample numbers from these studies are low, reducing the statistical power of any one experiment. Aggregating spaceflight datasets serves as a method to increase sample numbers, allowing for novel insights through bioinformatic analysis of ‘omics data from merged datasets. However, aggregating datasets can introduce unwanted variation including 1) differences in sample handling, processing, and sequencing platforms between datasets (technical variation) as well as 2) differences in experimental design between datasets such as sex or age of the model organism used. In the present study, NASA GeneLab-hosted RNAseq datasets from rodent liver tissues were used to evaluate several statistical methods to correct for this unwanted variation through two approaches, reference-based and standard. The following correction algorithms were applied with (reference-based) and/or without (standard) considering Universal Mouse RNA Reference samples: ComBat and ComBat_seq from the SVA package, median polish, empirical Bayes, and ANOVA-based algorithms from the MBatch package, and negative binomial regression normalization in the DESeq2 package. For each approach, after the correction algorithm was applied, differential gene expression (DGE) analysis of flight and ground control samples was performed with the combined data. The robustness of each tool was evaluated using BatchQC, to determine statistical differences between datasets before and after correction, Principal Component Analysis, to evaluate global gene expression in samples before and after correction, and by comparing DGE analysis of individual datasets and combined datasets before and after correction. The results showed that the reference-based approach introduced several additional (and likely artificial) DEGs when compared with the standard approach. Thus, the most robust standard correction will be implemented in the GeneLab Visualization 2.0 platform when datasets are combined.

GeneLab, RNA-seq, Batch Correction↗

Evaluation of Correction Methods for NASA GeneLab Transcriptomic Datasets

Conducting space biology experiments aboard the International Space Station, particularly those utilizing complex model organisms like mice, is expensive and difficult due to limited crew availability, hardware, and space. As a result, sample numbers from these studies are low, reducing the statistical power of any one experiment. Aggregating spaceflight datasets serves as a method to increase sample numbers, allowing for novel insights through bioinformatic analysis of ‘omics data from merged datasets. However, aggregating datasets can introduce unwanted variation including 1) differences in sample handling, processing, and sequencing platforms between datasets (technical variation) as well as 2) differences in experimental design between datasets. In the present study, NASA GeneLab-hosted RNAseq datasets from mouse liver tissues were used to evaluate several statistical methods to correct for this unwanted variation through two approaches, reference-based and standard. The following correction algorithms were applied with (reference-based) and/or without (standard) considering Universal Mouse RNA Reference samples: ComBat and ComBat_seq from the SVA package, median polish, empirical Bayes, and ANOVA-based algorithms from the MBatch package, and negative binomial regression normalization in the DESeq2 package. For each approach, after the correction algorithm was applied, differential gene expression (DGE) analysis of flight and ground control samples was performed with the combined data. The robustness of each tool was evaluated using BatchQC to determine statistical differences between datasets before and after correction, Principal Component Analysis to evaluate global gene expression in samples before and after correction, and by comparing DGE analysis of individual datasets and combined datasets before and after correction. The results showed that the reference-based approach introduced several additional (and likely artificial) DEGs when compared with the respective standard approach. Of the methods tested, standard ComBat and DESeq2 were identified as the most robust correction methods for combining spaceflight mouse liver RNAseq datasets hosted on GeneLab.

GeneLab↗

Hidden Features: How Subsurface and Landscape Heterogeneity Govern Hydrologic Connectivity and Stream Chemistry in a Montane Watershed

ABSTRACT Hydrologic connectivity is defined as the connection among stores of water within a watershed and controls the flux of water and solutes from the subsurface to the stream. Hydrologic connectivity is difficult to quantify because it is goverened by heterogeniety in subsurface storage and permeability and responds to seasonal changes in precipitation inputs and subsurface moisture conditions. How interannual climate variability impacts hydrologic connectivity, and thus stream flow generation and chemistry, remains unclear. Using a rare, four‐year synoptic stream chemistry dataset, we evaluated shifts in stream chemistry and stream flow source of Coal Creek, a montane, headwater tributary of the Upper Colorado River. We leveraged compositional principal component analysis and end‐member mixing to evaluate how seasonal and interannual variation in subsurface moisture conditions impacts stream chemistry. Overall, three main findings emerged from this work. First, three geochemically distinct end members were identified that constrained stream flow chemistry: reach inflows, and quick and slow flow groundwater contributions. Reach inflows were impacted by historic base and precious metal mine inputs. Bedrock fractures facilitated much of the transport of quick flow groundwater and higher‐storage subsurface features (e.g., alluvial fans) facilitated the transport of slow flow groundwater. Second, the contributions of different end members to the stream changed over the summer. In early summer, stream flow was composed of all three end members, while in late summer, it was composed predominantly of reach inflows and slow flow groundwater. Finally, we observed minimal differences in proportional composition in stream chemistry across all four years, indicating seasonal variability in subsurface moisture and spatial heterogeneity in landscape and geologic features had a greater influence than interannual climate fluctuation on hydrologic connectivity and stream water chemistry. These findings indicate that mechanisms controlling solute transport (e.g., hydrologic connectivity and flow path activation) may be resilient (i.e., able to rebound after perturbations) to predicted increases in climate variability. By establishing a framework for assessing compositional stream chemistry across variable hydrologic and subsurface moisture conditions, our study offers a method to evaluate watershed biogeochemical resilience to variations in hydrometeorological conditions.

Johnson, Keira [College of Earth, Ocean, and Atmos↗

Projections of future climate for U.S. national assessments: past, present, future

Climate assessments consolidate our understanding of possible future climate conditions as represented by climate projections, which are largely based on the output of global climate models. Over the past 30 years, the scientific insights gained from climate projections have been refined through model structural improvements, emerging constraints on climate feedbacks, and increased computational efficiency. Within the same period, the process of assessing and evaluating information from climate projections has become more defined and targeted to inform users. As the size and audience of climate assessments has expanded, the framing, relevancy, and accessibility of projections has become increasingly important. This paper reviews the use of climate projections in national climate assessments (NCA) while highlighting challenges and opportunities that have emerged over time. Reflections and lessons learned address the continuous process to understand the broadening assessment audience and evolving user needs. Insights for future NCA development include (1) identifying benchmarks and standards for evaluating downscaled datasets, (2) expanding efforts to gather research gaps and user needs to inform how climate projections are presented in the assessment (3) providing practitioner guidance on the use, interpretation, and reporting of climate projections and uncertainty to better inform decision-making.

Assessment↗

Energy-efficient cooperative resource allocation and task scheduling for Internet of Things environments

Offloading Internet of Things (IoT) tasks to the cloud for further processing might not always lead to an optimal execution time, particularly in situations such as resource contention, under-provisioning, over-provisioning, and fragmentation. In addition, dynamically optimizing the number of Virtual Machines (VMs) for resource scheduling in order to meet application requirements remains a major research challenge. Further, existing resource scheduling algorithms focus primarily on minimizing operational costs while maximizing resource sharing and utilization. Considering energy utilization as part of the resource allocation and scheduling process as an optimization objective for maintaining load balancing has often been neglected. To address these challenges and more, we propose a cooperative energy-aware resource allocation and scheduling strategy based on a Technique for Order of Preference by Similarity to Ideal Solution (TOPSIS) multi-criteria decision-making method. Here we used the Grid Workloads Archive dataset to evaluate our proposed approach named TOPREAL. Experimental results with respect to the allocation of VM resources when considering processing a large segment of tasks indicate that TOPREAL outperforms existing algorithms in terms of energy savings, with an average improvement of 40.25%, while maintaining an average improvement of 16.21% when it comes to execution time. Results also demonstrate that our method can save an average of 78.06 processing hours and 63,215kJ of energy when compared to existing scheduling algorithms. These results demonstrate the effectiveness of our proposed model and the viability of using multi-criteria decision-making techniques such as TOPSIS to solve the resource allocation and scheduling problem in edge environments.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

FL‐ADS: Federated learning anomaly detection system for distributed energy resource networks

Abstract With the ongoing development of Distributed Energy Resources (DER) communication networks, the imperative for strong cybersecurity and data privacy safeguards is increasingly evident. DER networks, which rely on protocols such as Distributed Network Protocol 3 and Modbus, are susceptible to cyberattacks such as data integrity breaches and denial of service due to their inherent security vulnerabilities. This paper introduces an innovative Federated Learning (FL)‐based anomaly detection system designed to enhance the security of DER networks while preserving data privacy. Our models leverage Vertical and Horizontal Federated Learning to enable collaborative learning while preserving data privacy, exchanging only non‐sensitive information, such as model parameters, and maintaining the privacy of DER clients' raw data. The effectiveness of the models is demonstrated through its evaluation on datasets representative of real‐world DER scenarios, showcasing significant improvements in accuracy and F1‐score across all clients compared to the traditional baseline model. Additionally, this work demonstrates a consistent reduction in loss function over multiple FL rounds, further validating its efficacy and offering a robust solution that balances effective anomaly detection with stringent data privacy needs.

Purohit, Shaurya [Iowa State University Ames Iowa ↗

Fast correlation function calculator: A high-performance pair-counting toolkit

A novel high-performance exact pair-counting toolkit called fast correlation function calculator (FCFC) is presented. With the rapid growth of modern cosmological datasets, the evaluation of correlation functions with observational and simulation catalogues has become a challenge. High-efficiency pair-counting codes are thus in great demand. We introduce different data structures and algorithms that can be used for pair-counting problems, and perform comprehensive benchmarks to identify the most efficient algorithms for real-world cosmological applications. We then describe the three levels of parallelisms used by FCFC, SIMD, OpenMP, and MPI, and run extensive tests to investigate the scalabilities. Finally, we compare the efficiency of FCFC with alternative pair-counting codes. The data structures and histogram update algorithms implemented in FCFC are shown to outperform alternative methods. FCFC does not benefit greatly from SIMD because the bottleneck of our histogram update algorithm is mainly cache latency. Nevertheless, the efficiency of FCFC scales well with the numbers of OpenMP threads and MPI processes, even though speedups may be degraded with over a few thousand threads in total. FCFC is found to be faster than most (if not all) other public pair-counting codes for modern cosmological pair-counting applications.

79 ASTRONOMY AND ASTROPHYSICS↗