Engineering PapersSearch

SEARCH · Engineering Papers

Results for “validation data sets”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Personalized, disease-stage specific, rapid identification of immunosuppression in sepsis

Introduction Data overlapping of different biological conditions prevents personalized medical decision-making. For example, when the neutrophil percentages of surviving septic patients overlap with those of non-survivors, no individualized assessment is possible. To ameliorate this problem, an immunological method was explored in the context of sepsis. Methods Blood leukocyte counts and relative percentages as well as the serum concentration of several proteins were investigated with 4072 longitudinal samples collected from 331 hospitalized patients classified as septic (n=286), non-septic (n=43), or not assigned (n=2). Two methodological approaches were evaluated: (i) a reductionist alternative, which analyzed variables in isolation; and (ii) a non-reductionist version, which examined interactions among six (leukocyte-, bacterial-, temporal-, personalized-, population-, and outcome-related) dimensions. Results The reductionist approach did not distinguish outcomes: the leukocyte and serum protein data of survivors and non-survivors overlapped. In contrast, the non-reductionist alternative differentiated several data groups, of which at least one was only composed of survivors (a finding observable since hospitalization day 1). Hence, the non-reductionist approach promoted personalized medical practices: every patient classified within a subset associated with 100% survival subset was likely to survive. The non-reductionist method also revealed five inflammatory or disease-related stages (provisionally named ‘early inflammation, early immunocompetence, intermediary immuno-suppression, late immuno-suppression, or other’). Mortality data validated these labels: both ‘suppression’ subsets revealed 100% mortality, the ‘immunocompetence’ group exhibited 100% survival, while the remaining sets reported two-digit mortality percentages. While the ‘intermediary’ suppression expressed an impaired monocyte-related function, the ‘late’ suppression displayed renal-related dysfunctions, as indicated by high concentrations of urea and creatinine. Discussion The data-driven differentiation of five data groups may foster early and non-overlapping biomedical decision-making, both upon admission and throughout their hospitalization. This approach could evaluate therapies, at personalized level, earlier. To ascertain repeatability and investigate the dynamics of the ‘other’ group, additional studies are recommended.

Immunology

Snow Distribution Patterns Revisited: A Physics-Based and Machine Learning Hybrid Approach to Snow Distribution Mapping in the Sub-Arctic

Snowpack distribution in Arctic and alpine landscapes often occurs in repeating, year-to-year patterns due to local topographic, weather, and vegetation characteristics. Previous studies have suggested that with years of observational data, these snow distribution patterns can be statistically integrated into a snow process modeling workflow. Recent advances in snow hydrology and machine learning (ML) have increased our ability to predict snowpack distribution using in-situ observations, remote sensing data sets, and simple landscape characteristics that can be easily obtained for most environments. Here, we propose a hybrid approach to couple a ML snow distribution pattern (MLSDP) map with a physics-based, snow process model. We trained a random forest ML algorithm on tens of thousands of snow survey observations from a subarctic study area on the Seward Peninsula, Alaska, collected during peak snow water equivalent (SWE). We validated hybrid model outputs using in-situ snow depth and SWE observations, as well as a light detection and ranging data set and a distributed temperature profiling sensor data set. When the hybrid results were compared with the physics-based method, the hybrid method more accurately depicted the spatial patterns of the snowpack, areas of drifting snow, and years when no in-situ observations were used in the random forest ML training data set. The hybrid method also showed improvements in root mean squared error at 61% of locations where time-series estimations of snow depth were observed. These results can be applied to any physics-based model to improve the snow distribution patterning to reflect observed conditions in high latitude and high elevation cold region environments.

54 ENVIRONMENTAL SCIENCES

SAXS Assistant: Automated SAXS analysis for structural discovery in biologics and polymeric nanoparticles

Small-angle x-ray scattering (SAXS) is a powerful technique for assessing macromolecular structure. High-throughput SAXS is limited by the time-consuming and, at times, subjective nature of SAXS data interpretation. Here, we present SAXS Assistant, a Python-based script that streamlines SAXS data analysis to extract features for machine learning (ML) and key structural parameters, including the Guinier radius of gyration (R g ), pair distance distribution function (PDDF)-derived R g , maximum particle dimension (D max ), and Kratky plots. The script builds upon BioXTAS RAW and validates reliability via Guinier/PDDF R g agreement, an important indicator of well-measured data sets. For assistance in D max estimation, a multilayer perceptron regressor was trained with 1940 data files from the Small Angle Scattering Biological Data Bank. The model achieved a test set performance R 2 = 0.90 and mean absolute error = 11.7 Å. Training exclusively with experimental data translates analyses from researchers, including experts in the field, to the ML model, which helps assess D max estimations from PDDF. Gaussian mixture model clustering was implemented to classify profiles into structural classes based on entries in the Small Angle Scattering Biological Data Bank. Users may therefore assess the similarity between experimental samples and known biomolecular shapes within the mapped repository entries. This probabilistic clustering aids in quantifying information from Kratky and generating shape-descriptive features. SAXS Assistant accelerates SAXS data analysis through enforced quality control, ML-ready outputs, and flags for low-confidence results. In addition to providing the ability to analyze large data sets at high throughput, this tool is versatile and may serve researchers in both biological and synthetic polymer research fields.

36 MATERIALS SCIENCE

The dark energy survey supernova program: investigating beyond-ΛCDM

We report constraints on a variety of non-standard cosmological models using the full 5-yr photometrically classified type Ia supernova sample from the Dark Energy Survey (DES-SN5YR). Both Akaike Information Criterion (AIC) and Suspiciousness calculations find no strong evidence for or against any of the non-standard models we explore. When combined with external probes, the AIC and Suspiciousness agree that 11 of the 15 models are moderately preferred over Flat-|$\Lambda$|CDM suggesting additional flexibility in our cosmological models may be required beyond the cosmological constant. We also provide a detailed discussion of all cosmological assumptions that appear in the DES supernova cosmology analyses, evaluate their impact, and provide guidance on using the DES Hubble diagram to test non-standard models. An approximate cosmological model, used to perform bias corrections to the data holds the biggest potential for harbouring cosmological assumptions. We show that even if the approximate cosmological model is constructed with a matter density shifted by |$\Delta \Omega _{\rm m}\sim 0.2$| from the true matter density of a simulated data set the bias that arises is subdominant to statistical uncertainties. Nevertheless, we present and validate a methodology to reduce this bias.

79 ASTRONOMY AND ASTROPHYSICS

Machine learning approaches for crystallographic classification from synthetic 2D X-ray diffraction data

Crystallographic structure identification is crucial for understanding material properties; however, current methodologies often depend on labor-intensive and time-consuming analyses of 2D X-ray diffraction (XRD) patterns. To address these limitations, this study employs synthetic 2D XRD patterns combined with deep learning (DL) techniques to enable automated and high-throughput classification of the seven crystal systems and 230 space groups. We introduce the novel Auto Diffraction Pipeline, designed to generate synthetic 2D XRD spot patterns from crystallographic information files under diverse conditions, including varying zone axes, atomic substitution, atomic depletion and mechanical loading. These conditions enhance the realism of synthetic data, mitigating the scarcity of experimental datasets and enabling the creation of large representative training sets. Convolutional neural networks were trained and validated on these synthetic datasets to classify crystallographic structures across multiple scenarios. Our results demonstrate that integrating synthetic 2D XRD patterns with DL facilitates rapid, accurate and automated crystallographic classification, promoting the wider adoption of data-driven approaches in materials science.

Shahnazari, Ayoub [Univ. of Rochester, NY (United

Machine Learning‐Assisted Microearthquake Location Workflow for Monitoring the Newberry Enhanced Geothermal System

Abstract Enhanced geothermal systems (EGS) offer a sustainable energy source but face challenges in accurately locating microearthquakes induced during reservoir stimulation. Locating these microearthquakes provides reliable feedback on the stimulation progress. Current deep learning methods for locating earthquakes require extensive data sets for training, which is problematic as detected microearthquakes are often limited. To address the scarcity of training data, we propose a practical workflow using probabilistic multilayer perceptron (PMLP) which predicts microearthquake locations from cross‐correlation time lags in waveforms. Utilizing a 3D velocity model of Newberry site derived from ambient noise interferometry, we generate numerous synthetic microearthquakes and 3D acoustic waveforms for PMLP training. Accurate synthetic tests prompt us to apply the trained network to the 2012 and 2014 stimulation field waveforms. To enhance the accuracy of source localization, we carefully handpick the P‐arrival times. Predictions on the 2012 stimulation data set show major microseismic activity at depths of 0.5–1.2 km, correlating with a known casing leakage scenario. In the 2014 data set, the majority of predictions concentrate at 2.0–2.9 km depths, consistent with results obtained from conventional physics‐based inversion, and align with the presence of natural fractures from 2.0 to 2.7 km. We validate our findings by comparing the synthetic and field picks, demonstrating a satisfactory match for the first arrivals. By combining the benefits of quick inference speeds and accurate location predictions, we demonstrate the feasibility of using realistic synthetic data set to locate microseismicity for EGS monitoring.

15 GEOTHERMAL ENERGY

Systematic exploration of the thermochemistry for a set of peroxy hydroperoxy-alkyl radicals

Here, the thermochemistry of peroxy hydroperoxy-alkyl ($\overline{O}$OQOOH) radicals has a significant influence on the reactivity of fuels and on the formation of highly oxygenated molecules (HOMs) in the atmosphere. Theoretical characterization of these radicals can be arduous due to their molecular size and complex fundamental interactions, such as hydrogen-bonding and torsional anharmonicity, and difficult to validate in the absence of any direct experimental thermochemical data. In this work, we systematically explore the thermochemistry of a set of $\overline{O}$OQOOH radicals with five increasingly affordable approaches with considerations for these interactions. In doing so, we present a novel conformer selection approach suited to predict properties at combustion temperatures. As a corollary, we also highlight the shortcomings in the standard choice of the ground conformer as the reference. The set of molecules is comprised of 149 C 2 -C 8 $\overline{O}$OQOOH radicals, selected to encompass a wide variety of branching and substitution patterns. Comparisons amongst the approaches help quantify the errors arising from various simplifying assumptions. For the largest of these radicals, and with the most affordable approach, final predictions of Gibbs energies are assigned a 2 σ uncertainty of 4 kcal mol -1 in the negative temperature coefficient region.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH

Cleaned 5-Minute Resolution Air Quality and Meteorological Data from Nine TCEQ CAMS Sites in Houston, Texas (Nov 2021 – Oct 2022)

These data encompass 5-minute air monitoring and meteorological observations collected in the greater Houston, Texas metropolitan region, at nine (9) Continuous Ambient Monitoring Stations (CAMS) operated by the Texas Commission on Environmental Quality (TCEQ) between November 1, 2021 and October 31, 2022. The CAMS sites (CAMS 1, 8, 35, 45, 148, 403, 405, 410, and 1052) were chosen because their instrumentation includes measurements of PM2.5. These sites also provide continuous multi-parameter air-quality and meteorological measurements. Particulate matter (PM2.5, PM10) was sampled along with several trace gases, including ozone (O3), nitrogen oxides (NO, NO2, NOx), sulfur dioxide (SO2), and carbon monoxide (CO). The data set also contains standard surface meteorological parameters (temperature, humidity, pressure, wind speed, and wind direction). Several sites also include AutoGC-based measurements of volatile organic compounds (VOCs). Air monitoring instruments deployed at the selected sites comprise the following systems: BAM-1020 or TEOM (PM2.5), Thermo Scientific TEI 49i (O3), TEI 42i (NOx), and AutoGCs (VOCs). This data set is similar to the data included within the houairq5mX1.00 datastream, except for a few additional quality control steps. A systematic data cleaning and verification process was performed on the data set to ensure its quality and preparation for analysis. Removal of non-numeric status flags (e.g., [LIM], [QAS], [SPZ], [CAL], [PMA], [AQI], [SPN], [MAL]) was accomplished by employing rule-based string parsing to extract valid numerical values. Missing entries were set to -9999; however, invalid or anomalous values (e.g., 99999) were retained as originally reported by the TCEQ to preserve data provenance. The time sequence was verified for completeness, removal of duplicates, and uniformity at 5-minute intervals. Column labeling was standardized, and corresponding values were assessed for physical plausibility. All timestamps in the data set were reported in Coordinated Universal Time (UTC) as provided by the TCEQ. Further, the latitude and longitude coordinates were added for each CAMS site. A subset of the data (June 1–September 30, 2022) has been used in the following publication: Subba et al. 2025. “Implications of sea breeze circulations on boundary layer aerosols in the southern coastal Texas region.” EGUsphere 2025: 1–49, https://doi.org/10.5194/egusphere-2025-2659.

latitude

Enclosed Sydney Swirl Burner Experimental Data Set - Release 1.0

This submission includes a comprehensive data set for the NETL Enclosed Sydney Swirl Burner, generated between 2012 and 2014. The data set includes global parametric data and advanced diagnostic data for (3) flames of interest to CFD model validation. The submission includes a companion summary document describing the data and associated references.

Diagnostics

ComStock Measure Documentation: High-Efficiency Rooftop Unit

Building on the 3-year End-Use Load Profiles project to calibrate and validate the U.S. Department of Energy's ResStock and ComStock models, this work produces national data sets that enable cities, states, utilities, and other stakeholders to answer a broad range of questions regarding their commercial building stock. ComStock is a highly granular, bottom-up model that uses various data sources, statistical sampling methods, and advanced building energy simulations to estimate the annual subhourly energy consumption of the commercial building stock across the United States. The "baseline" model intends to represent the U.S. commercial building stock as it existed in 2018. The methodology of the baseline model is discussed in the ComStock Reference Documentation. The goal of this work is to develop energy efficiency and demand flexibility measures that cover market-ready technologies and study their mass adoption impact on the baseline building stock. "Measures" refers to various "what-if" scenarios that can be applied to buildings. The results for the baseline and measure scenario simulations are published in public data sets that provide insights into building stock characteristics, operational behaviors, utility bill impacts, and annual and sub-hourly energy usage by fuel type and end use. This report describes the modeling methodology for a single ComStock measure scenario - high-efficiency rooftop unit (RTU) - and briefly introduces key results. The full public data set can be accessed on the Comstock data lake or via the Data Viewer at comstock.nlr.gov. The public data set enables users to create custom aggregations of results for their use case (e.g., filter to a specific county or building type).

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI

ComStock Measure Documentation: Variable-Speed Pumps

Building on the 3-year End-Use Load Profiles project to calibrate and validate the U.S. Department of Energy's ResStock and ComStock models, this work produces national data sets that enable cities, states, utilities, and other stakeholders to answer a broad range of questions regarding their commercial building stock. ComStock is a highly granular, bottom-up model that uses various data sources, statistical sampling methods, and advanced building energy simulations to estimate the annual subhourly energy consumption of the commercial building stock across the United States. The "baseline" model intends to represent the U.S. commercial building stock as it existed in 2018. The methodology of the baseline model is discussed in the ComStock Reference Documentation. The goal of this work is to develop energy efficiency and demand flexibility measures that cover market-ready technologies and study their mass adoption impact on the baseline building stock. "Measures" refers to various "what-if" scenarios that can be applied to buildings. The results for the baseline and measure scenario simulations are published in public data sets that provide insights into building stock characteristics, operational behaviors, utility bill impacts, and annual and sub-hourly energy usage by fuel type and end use. This report describes the modeling methodology for a single ComStock measure scenario - variable speed pumps - and briefly introduces key results. The full public data set can be accessed on the ComStock data lake or via the Data Viewer at comstock.nlr.gov. The public data set enables users to create custom aggregations of results for their use case (e.g., filter to a specific county or building type).

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI

Numerical simulations of liquid jetting with solid inclusions

The dynamics of finite-sized particles in fluids, and their influence on the overall flow, are of great interest across several industrial, environmental, and medical fields. In the context of inkjet printing, the presence of solid inclusions can be either intentional, as in additive manufacturing, or unintentional, as in standard printing processes. These inclusions can strongly impact the jetting process, causing effects such as jet asymmetry, bubble entrapment, and the formation of satellite droplets. Understanding and controlling particle behavior is therefore essential, particularly to predict how and when particles are ejected over multiple jetting cycles. It is therefore critical to develop reliable models that allow for a deeper understanding of the complex interplay between particle and fluid during the whole printing process. To address this, we present a tailored implementation of the Color-Gradient multicomponent Lattice Boltzmann Method for fully resolved three-dimensional (3D) simulations of multicycle liquid jetting with particles. Our method supports realistic parameter settings aligned with industrial inkjet systems, and we provide both qualitative and quantitative validation against experimental data. Additionally, we introduce a simplified model based on the Stokes drag law, in which solid particles are represented as point particles and do not influence the fluid flow. Despite this limitation, the model offers a computationally efficient means to explore the vast parameter space typically encountered in industrial applications, allowing, e.g., identifying critical ejection regions and estimating the number of cycles required for particle release. These qualitative insights are valuable for guiding and complement fully two-way coupled simulations.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC

Validation Data for Benchmarking Wire Arc Additive Manufacturing Process Simulations

Residual stresses cause geometric distortion and affect mechanical performance of additively manufactured structures, yet they are notoriously difficult to assess and predict. Distortion (warpage) can drive parts outside dimensional tolerance limits, leading to part rejection or rework. For parts that meet tolerance, locked-in residual stress fields can affect structural integrity during operation, particularly subcritical cracking by fatigue, creep, or corrosion. This work develops benchmark data for a common additive manufacturing process (Wire Arc Additive Manufacturing) that can be applied for calibration and validation of physical process models that predict residual stress fields. The work includes design of two different samples of differing geometry, detailed manufacturing records for a set of physical samples, and an extensive set of residual stress measurement data developed using two diverse techniques (the contour method and neutron diffraction). An initial application of the work is also reported, where a modeling challenge was issued to secure residual stress model predictions from two independent laboratories that were blind to residual stress measurement data. These initial blind residual stress predictions show significant discrepancies relative to the measurement data, illustrating the potential value of the underlying validation data. An open repository for this work, including the sample designs, manufacturing process records, and the residual stress data, is also provided for future application in non-blind validation efforts.

36 MATERIALS SCIENCE

Importance of Spatially Continuous Urban Surface Properties in Urban‐Resolving Earth System Modeling

Accurate representation of urban properties and processes at higher resolutions in global modeling systems is essential for advancing our ability to capture the complexities of urban systems and informing effective resilience strategies. However, the prescription of coarse global-scale urban properties in most state-of-the-art Earth system models (ESMs) is limiting their potential for capturing urban signals as they advance toward kilometer-scale simulation capabilities. To bridge this gap in inadequate urban property representation and to advance urban-resolving Earth system modeling, this work integrates the newly-developed global 1 km-resolution facet-level urban surface property data set, U-Surf, into the land component of Community Earth System Model (CESM)—Community Terrestrial System Model (CTSM). The land-only CTSM simulations are validated against satellite measurements, ground-based urban weather stations, flux tower observations, and reanalysis data. Results demonstrate that the enhanced urban properties allow improved simulations of urban meteorology and surface energy fluxes compared to the default coarse-resolution categorical urban canopy parameters. Spatial scaling analysis reveals regime-dependent information loss during resolution aggregation, as well as substantial scale-dependent variations in urban surface energy flux representation. Furthermore, these findings have critical implications for coupled Earth system modeling when including the effect of land-atmosphere interaction. This work establishes a foundation for future urban-resolving kilometer-scale ESM development, which will enable systematic intra- and inter-city comparisons that inform urban adaptation strategies across diverse global urban environments.

Cheng, Yifan [University of Illinois Urbana-Champa

Validation of HyRAM+ Version 5.1 Physics Models

The Hydrogen Plus Other Alternative Fuels Risk Assessment Models (HyRAM+) software has seen various improvements and additional physics capabilities since validation against experimental data was last published for HyRAM v3.1. Notably, HyRAM+ now includes four models allowing for the calculation of overpressure resulting from vapor cloud explosions from unconfined jet releases. As with the previous HyRAM validation report, validation data was gathered from available published literature and tested against HyRAM+ capabilities. The validation comparisons include tank blowdown, unignited dispersion jet plume, ignited jet flame, and enclosed accumulation and overpressure. The unconfined overpressure calculations in HyRAM+ v5.1.1 generally show good agreement with many of the experimental data sets for all four unconfined overpressure models, though HyRAM+ overpredicts the experimental data for small and cryogenic hydrogen releases. The comparisons for the other HyRAM+ physics models are largely unchanged from the previously published validation report.

08 HYDROGEN

Data for Discovery, Characterization, and Application of Chromosomal Integration Sites for Stable Heterologous Gene Expression in Rhodotorula toruloides

Rhodotorula toruloides is a non-model, oleaginous yeast uniquely suited to produce acetyl-CoA-derived chemicals. However, the lack of well-characterized genomic integration sites has impeded the metabolic engineering of this organism. Here we report a set of computationally predicted and experimentally validated chromosomal integration sites in R. toruloides . We first implemented an in silico platform by integrating essential gene information and transcriptomic data to identify candidate sites that meet stringent criteria. We then conducted a full experimental characterization of these sites, assessing integration efficiency, gene expression levels, impact on cell growth, and long-term expression stability. Among the identified sites, 12 exhibited integration efficiencies of 50% or higher, making them sufficient for most metabolic engineering applications. Using selected high-efficiency sites, we achieved simultaneous double and triple integrations and efficiently integrated long functional pathways (up to 14.7 kb). Additionally, we developed a new inducible marker recycling system that allows multiple rounds of integration at our characterized sites. We validated this system by performing five sequential rounds of GFP integration and three sequential rounds of MaFAR integration for fatty alcohol production, demonstrating, for the first time, precise gene copy number tuning in R. toruloides . These characterized integration sites should significantly advance metabolic engineering efforts and future genetic tool development in R. toruloides .

Conversion

Update on Radiochemical Assessment of High Burnup Commercially Irradiated Fuel

This work documents an effort to collect burnup measurements on a high burnup rod, designated 6XV, and first cycle accident tolerant fuel (ATF) rod, designated 47I, to enable benchmarking of fuel performance codes and neutronics codes. In addition to measurements, Virtual Environment for Reactor Applications (VERA) full-core-depletion analysis was also performed for the rods that were experimentally analyzed to provide an opportunity for code validation. This effort focuses on collecting data from rods irradiated at Byron Generating Station and shipped to the Oak Ridge National Laboratory (ORNL) hot-cells. This data will also anchor non-destructive examination evaluations of burnup of the various fuel rods undergoing postirradiation examination (PIE) at ORNL. Previous PIE of these fuel rods provides some guidance on the burnup trend across the fuel. Axial gamma spectroscopy scans provide a measure of relative changes in burnup across a fuel pin. Mass spectrometry based burnup measurements performed for this work at specific axial locations in the fuel are fully quantitative. By combining the mass spectrometry data with the gamma scans it is possible to more quantitatively evaluate axial variations in burnup across the entire fuel pin [1]. The combined set of burnup evaluations will be made available to other organizations that have an interest in high burnup radiochemistry data for validation of neutronic simulations and source term evaluation such as the Nuclear Regulatory Commission (NRC).

Harp, Jason [Oak Ridge National Laboratory (ORNL),

Quantifying Epistemic Uncertainty in Binary Classification via Accuracy Gain

ABSTRACT Recently, a surge of interest has been given to quantifying epistemic uncertainty (EU), the reducible portion of uncertainty due to lack of data. We propose a novel EU estimator in the binary classification setting, as the posterior expected value of the empirical gain in accuracy between the current prediction and the optimal prediction. In order to validate the performance of our EU estimator, we introduce an experimental procedure where we take an existing dataset, remove a set of points, and compare the estimated EU with the observed change in accuracy. Through real and simulated data experiments, we demonstrate the effectiveness of our proposed EU estimator.

97 MATHEMATICS AND COMPUTING