Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “evaluation datasets”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

Dynamical Downscaling of Earth System Model Data for Energy System Analysis

Assessing energy resources (e.g., solar, wind, and hydro) under future scenarios requires datasets with sufficient spatial and temporal detail to capture variability and extreme events. While global-scale Earth System Model (ESM) projections are widely used, their coarse resolution limits direct application to regional energy system analyses. Dynamical downscaling offers a robust approach to generate physically consistent, fine-scale datasets that better represent local atmospheric processes impacting energy resources. In this work, we present a two-stage approach for producing high-resolution historical and future projections over the contiguous United States (CONUS). First, we optimize the Weather Research and Forecasting (WRF) model configuration for energy-relevant variables - solar irradiance, wind speed, and precipitation - by conducting ERA5-driven simulations at 8-km and 28-km resolution. Multiple physics schemes and model configurations within the WRF are evaluated against observational datasets including the National Solar Radiation Database (NSRDB), the Parameter-elevation Regressions on Independent Slopes Model (PRISM), and the Stage IV multi-radar/multi-sensor precipitation product for the CONUS domain. Using the best-performing configuration, we dynamically downscale MPI-ESM1-2-HR simulations for 2000-2060 under SSP2-4.5 and SSP5-8.5 scenarios at 4-km spatial and hourly temporal resolution. This presentation will provide a comprehensive analysis of the results from multiple numerical experiments and high-resolution ESM projections. In addition, we will discuss potential applications of our high-resolution datasets within the energy sector and outline future research avenues dedicated to evaluating how extreme weather events influence system performance and resilience.

24 POWER TRANSMISSION AND DISTRIBUTION↗

An improved dataset for predicting mammal infecting viruses from genetic sequence information

There have been several attempts to develop machine learning (ML) models to identify human infecting viruses from their genomic sequences, with varying degrees of success. Direct comparison between models is problematic, because these models are typically trained and evaluated on different datasets with alternative data splitting schemes, features, and model performance metrics. In this paper we present a standardized dataset of mammal infecting and non-infecting viral pathogens, refined from the previous work of Mollentze et al. to include the latest literature evidence, roughly doubling the number of curated host-virus records available to the community, and new host target labels, primate and mammal. The new host labels were included for several reasons, including previous reports that classification performance is better at broader taxonomic ranks and the idea that there may be more data for primate infection that might serve as a suitable proxy for zoonotic potential and avoidance of false positives for human infection due to absence of evidence. On this dataset, we report the performance of eight machine learning models for predicting mammal-infecting viruses from their genomic sequences. We find that randomly assigning cases in our improved dataset to training/testing sets, when compared to the original assignments into training/testing in Mollentze et al., increases the overall average ROC AUC of prediction of human infection from 0.663 ± 0.070 to 0.784 ± 0.013, consistent with the reduction in phylogenetic distance between train and test sets (relative entropy change from 3.00 to 0.08). The broadest host category of mammal infection can be predicted most reliably at 0.850 ± 0.020. We share our improved dataset and code to enable standardized comparisons of machine learning methods to predict human host infections. Overall, we have presented preliminary evidence that classification of virus host infection is more tractable at higher taxonomic ranks, that unsurprisingly reducing the phylogenetic distance between training and test sets can improve predictive performance, that peptide kmer features appear to be harmful to out of sample model performance, and we are left with the question of whether models for virus host prediction can reasonably be expected to perform well in out of sample scenarios given the likelihood that viruses do not share a common ancestor. Consistent with this concern, when the data is resampled such that there is no overlap between viral families in training and test sets (relative entropy > 24), models perform no better than random chance at prediction of human infection regardless of whether kmers are included (ROC AUC 0.50 ± 0.08) or not (ROC AUC 0.50 ± 0.04).

59 BASIC BIOLOGICAL SCIENCES↗

SetGo: Metadata Readiness for Scientific AI Datasets

Scientific datasets intended for AI use require both computational readiness for model training and metadata readiness for discovery, sharing, and reuse. The Readiness Engine for Data Integration (REDI) addresses computational readiness, but no corresponding tool evaluates whether a dataset’s metadata are sufficiently complete, governed, and standards-compliant for publication and agent-based consumption. Existing FAIR assessors operate only on published repository records, and no single system covers FAIR compliance, licensing, provenance, governance, reproducibility, and catalog readiness together. We present SetGo, an open-source Python toolkit that assesses and repairs metadata readiness across these six dimensions before a dataset is published or archived. Applied to four scientific corpora, SetGo surfaces deficiencies that general-purpose tools do not detect: ERA5 climate metadata scores 4% on ACDD 1.3 compliance; materials datasets fail OPTIMADE species-definition requirements; and PDB-derived proteomics data carries licensing terms incompatible with standard SPDX identifiers. Guided enrichment raises overall FAIR scores from 52–57% to 81–91%, and a single setgo publish command pushes to Hugging Face Hub, CKAN, or OpenMetadata with ML Commons Croissant 1.0 metadata sidecars. To support interactive and automated workflows, SetGo integrates with coding agents powered by large language models (LLMs) through a /setgo skill that enables natural-language execution of the full assess–enrich–publish loop, with user involvement limited to supplying missing metadata values.

Wilkinson, Sean [ORNL] (ORCID:0000000214437479)↗

Evaluation of precipitation indices in suites of dynamically and statistically downscaled regional climate models over Florida

Abstract The present work evaluates historical precipitation and its indices defined by the Expert Team on Climate Change Detection and Indices (ETCCDI) in suites of dynamically and statistically downscaled regional climate models (RCMs) against NOAA’s Global Historical Climatology Network Daily (GHCN-Daily) dataset over Florida. The models examined here are: (1) nested RCMs involved in the North American CORDEX (NA-CORDEX) program, (2) variable resolution Community Earth System Models (VR-CESM), (3) Coupled Model Intercomparison Project phase 5 (CMIP5) models statistically downscaled using localized constructed analogs (LOCA) technique. To quantify observational uncertainty, three in situ-based (PRISM, Livneh, CPC) and three reanalysis (ERA5, MERRA2, NARR) datasets are also evaluated against the station data. The reanalyses and dynamically downscaled RCMs generally underestimate the magnitude of the monthly precipitation and the frequency of the extreme rainfall in summer. The models forced with CanESM2 miss the phase of the seasonality of extreme precipitation. All models and reanalyses severely underestimate both the mean and interannual variability of mean wet-day precipitation (SDII), consecutive dry days (CDD), and overestimate consecutive wet days (CWD). Metric analysis suggests large uncertainty across NA-CORDEX models. Both the LOCA and VR-CESM models perform better than the majority of models. Overall, RegCM4 and WRF models perform poorer than the median model performance. The performance uncertainty across models is comparable to that in the reanalyses. Specifically, NARR performs poorer than the median model performance in simulating the mean indices and MERRA2 performs worse than the majority of models in capturing the interannual variability of the indices.

54 ENVIRONMENTAL SCIENCES↗

Utah FORGE Well 16A(78)-32 Stimulation DFN Fracture Plane Evaluation and Data

This dataset includes files used to fit planar fractures through the preliminary earthquake catalogs of the three stages of the April 2022 well 16A(78)-32 stimulation which is linked bellow. These planar features have been used to update the FORGE reference Discrete Fracture Network (DFN) model. The files are provided to encourage other modelers to use additional workflows to find additional/alternative features. To this end, the dataset includes the cleaned earthquake catalog data translated to the FORGE reference model global reference frame, the well trajectory of 16A(78)-32 in those same coordinates, the fit 15 planar features in csv format, and a pdf file with slides illustrating the process used to fit the features. A recorded presentation of this material is available from the October 2022 FORGE Modeling and Simulation Forum which is also linked below.

15 GEOTHERMAL ENERGY↗

Predicting High Energy Arcing Fault Zones of Influence for Aluminum Using an Arc Flash Modeling Approach: Evaluation of a model bias, uncertainty, parameter sensitivity and zone of influence estimation

This report documents the development of an arc flash hazard model to calculate the incident energy and zone of influence from high energy arcing faults involving aluminum. The NRC has identified the potential for (HEAFs) involving aluminum to increase the damage zone beyond what is currently postulated in fire probabilistic risk assessment (PRA) methodologies. To estimate the hazard from HEAFs involving aluminum an arc flash model was developed. Differences between the initial model and nuclear power plant (NPP) fire PRA scenarios were identified. Modification of the initial model established from existing literature and test data was used to minimize these differences. The developed model was evaluated against NRC datasets to understand the model prediction and relative uncertainties. Finally, a range of fire PRA zone of influences (ZOI) were developed based on the developed model, target fragility estimates and update HEAF PRA methodology. The results were developed to support an NRC LIC-504 evaluation in tandem with other modeling efforts. The report documents the effort and provides a reference for any future advancements in arc flash modeling.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Evaluation of Rainfall-Snowfall Separation Performance in Remote Sensing Datasets

The first step to accurately measure global snowfall is to separate rainfall from snowfall correctly (i.e., precipitation phase discrimination). This study first evaluates the phase discrimination performance in four remote sensing datasets, including observations from ground radar, spaceborne radars, and spaceborne radiometer, relative to ground observations. Results show that the snowfall discrimination accuracy varies greatly among these datasets ranging from 42% to 96%, dependent on whether and how the temperature information are considered. For example, over half of the snowfall from the Global Precipitation Measurement Mission (GPM) spaceborne radar is actually rainfall at the surface since it detects snowfall in the air without considering the temperature information close to the surface. Second, we evaluate the discrimination performance using the temperature information from four reanalysis datasets. It is found that MERRA2 temperature close to the surface is colder than the other three datasets, leading to more rainfall being misclassified as snowfall.

Yalei You↗

Comparative study of machine learning techniques for post-combustion carbon capture systems

Computational analysis of countercurrent flows in packed absorption columns, often used in solvent-based post-combustion carbon capture systems (CCSs), is challenging. Typically, computational fluid dynamics (CFD) approaches are used to simulate the interactions between a solvent, gas, and column's packing geometry while accounting for the thermodynamics, kinetics, heat, and mass transfer effects of the absorption process. These simulations can then be used explain a column's hydrodynamic characteristics and evaluate its CO 2 -capture efficiency. However, these approaches are computationally expensive, making it difficult to evaluate numerous designs and operating conditions to improve efficiency at industrial scales. In this work, we comprehensively explore the application of statistical ML methods, convolutional neural networks (CNNs), and graph neural networks (GNNs) to aid and accelerate the scale-up and design optimization of solvent-based post-combustion CCSs. We apply these methods to CFD datasets of countercurrent flows in absorption columns with structured packings characterized by several geometric parameters. We train models to use these parameters, inlet velocity conditions, and other model-specific representations of the column to estimate key determinants of CO 2 -capture efficiency without having to simulate additional CFD datasets. We also evaluate the impact of different input types on the accuracy and generalizability of each model. We discuss the strengths and limitations of each approach to further elucidate the role of CNNs, GNNs, and other machine learning approaches for CO 2 -capture property prediction and design optimization.

97 MATHEMATICS AND COMPUTING↗

Evaluating probabilistic deep learning methods for uncertainty quantification of temperature downscaling

Deep learning (DL) has emerged as a promising tool for downscaling coarse-resolution climate data to high-resolution outputs, enabling improved regional climate predictions. A critical aspect of DL-based downscaling is the incorporation of uncertainty quantification (UQ), which enhances the interpretability and reliability of predictions—key factors for climate risk assessment and decision-making. This study develops a DL model to downscale 2 m temperature across the contiguous United States using reanalysis datasets. We systematically evaluate three epistemic UQ methods—deep ensembles (DEns), Monte Carlo dropout (MCD), and Flipout—based on their probabilistic accuracy, downscaling performance, sensitivity to geographical features, and computational efficiency. Results indicate that MCD generally outperforms Flipout and DEns in terms of calibration and downscaling accuracy. However, DEns demonstrate lower calibration errors in coastal regions, indicating its higher confidence within these areas. Flipout, in contrast, is more sensitive to elevation gradients and exhibits higher calibration errors in mountainous regions. Hence, the choice of UQ method for this task depends on the specific requirements of the application. For applications that prioritize overall calibration, downscaling accuracy, and computational efficiency, MCD is a strong candidate. These findings highlight the importance of selecting UQ methods based on application-specific requirements, such as geographical context and computational constraints. By addressing the trade-offs between UQ methods, this study provides actionable insights for improving the reliability, scalability, and utility of DL-based downscaling in climate science.

Environmental sciences↗

A 1 km soil moisture dataset over eastern CONUS generated by assimilating SMAP data into the Noah-MP land surface model

An improved fine-scale soil moisture (SM) dataset at 1 km grid spacing, covering much of the eastern continental US, was generated by assimilating 9 km Soil Moisture Active Passive (SMAP) SM data into the v4.0.1 Noah-MP land surface model. With 12 ensemble members, the assimilation was carried out using the ensemble Kalman filter algorithm within NASA's Land Information System. The SM analysis for 2016 was fully validated against in situ observations from four different networks and compared with four other existing datasets. Results indicate that this SM analysis surpasses other datasets in top-layer SM distribution, including a machine-learning-based product, despite all SM estimates being less heterogeneous than observed. The analysis of anomalous errors suggests that large similarity in intrinsic errors is likely due to overlapping data sources among the selected SM datasets. More detailed evaluations were performed over two geographic areas. The observations collected by the Atmospheric Radiation Measurement facility in Oklahoma suggest that soil temperature and surface heat fluxes are concurrently simulated with good accuracy. Investigation into the 2016 southeastern US drought response further indicates drier conditions and higher evapotranspiration estimates compared to GLEAMv4.1. Notably, large errors are associated with grids having clay soil textures, underscoring the need for refined model treatments for specific soil types to further improve SM estimates. The dataset is publicly available on Zenodo at https://doi.org/10.5281/zenodo.14370563 (Tai et al., 2024).

Tai, Sheng-Lun [Pacific Northwest National Laborat↗

Illustrating the Spatiotemporal Complexity of No2 Columns Using A Multi-Perspective Observing System: Moving Toward Geostationary Product Validation and Applications

As a precursor to secondary pollutants like ozone and PM2.5, nitrogen dioxide (NO2) is crucial to understand when addressing air quality issues. However, due to NO2’s short lifetime during the daytime and complexity of emission sources in urbanized regions, interpreting datasets from ground or satellite perspectives alone are challenged by variance in spatial and temporal resolutions. High resolution airborne mapping (< 1 km) of NO2 column densities across morning, midday, and afternoon add a unique perspective toward interpreting satellite data with respect to ground-measurements. This presentation focuses on the interpretation of spatiotemporal complexity of NO2 columns from the Synergistic TEMPO Air Quality Science Study (STAQS). The mission’s goal is to integrate geostationary observations from Tropospheric Emissions: Monitoring of Pollution (TEMPO) with traditional and enhanced air quality monitoring to improve the understanding of air quality science for increased societal benefit. We will demonstrate the interweaved perspective of NO2 columns from ground-based Pandora spectrometers and satellite-based observations (e.g., TROPOMI) as compared to high spatial resolution airborne observations from the GEOstationary Coastal and Air Pollution Events (GEO-CAPE) Airborne Simulator (GCAS). This includes the evaluation of each dataset through comparison to each other to identify potential biases in data products and the impact of heterogeneity on these comparisons. Airborne data will also be used as a proxy for geostationary observations with morning, midday, and afternoon raster maps collected over four cities (Los Angeles, Chicago, Toronto, and New York City). Finally, recent research outcomes will be presented to demonstrate how airborne and geostationary observations can be used to evaluate emission inventories and air quality models.

Laura Judd↗

Shortwave Array Spectroradiometer-Hemispheric (SAS-He): design and evaluation

A novel ground-based radiometer, referred to as the Shortwave Array Spectroradiometer-Hemispheric (SAS-He), is introduced. This radiometer uses the shadow-band technique to report total irradiance and its direct and diffuse components frequently (every 30 s) with continuous spectral coverage (350–1700 nm) and moderate spectral (~2.5 nm ultraviolet–visible and ~6 nm shortwave-infrared) resolution. The SAS-He's performance is evaluated using integrated datasets collected over coastal regions during three field campaigns supported by the US Department of Energy's Atmospheric Radiation Measurement (ARM) program, namely the (1) Two-Column Aerosol Project (TCAP; Cape Cod, Massachusetts), (2) Tracking Aerosol Convection Interactions Experiment (TRACER; in and around Houston, Texas), and (3) Eastern Pacific Cloud Aerosol Precipitation Experiment (EPCAPE; La Jolla, California). We compare (i) aerosol optical depth (AOD) and total optical depth (TOD) derived from the direct irradiance, as well as (ii) the diffuse irradiance and direct-to-diffuse ratio (DDR) calculated from two components of the total irradiance. As part of the evaluation, both AOD and TOD derived from the SAS-He direct irradiance are compared to those provided by a collocated Cimel sunphotometer (CSPHOT) at five (380, 440, 500, 675, 870 nm) and two (1020, 1640 nm) wavelengths, respectively. Additionally, the SAS-He diffuse irradiance and DDR are contrasted with their counterparts offered by a collocated multifilter rotating shadowband radiometer (MFRSR) at six (415, 500, 615, 675, 870, 1625 nm) wavelengths. Overall, reasonable agreement is demonstrated between the compared products despite the challenging observational conditions associated with varying aerosol loadings and diverse types of aerosols and clouds. For example, the AOD- and TOD-related values of root mean square error remain within 0.021 at 380, 440, 500, 675, 870, 1020, and 1640 nm wavelengths during the three field campaigns.

47 OTHER INSTRUMENTATION↗

a priori uncertainty quantification of reacting turbulence closure models using Bayesian neural networks

While many physics-based closure model forms have been posited for the sub-filter scale (SFS) in large eddy simulation (LES), vast amounts of data available from direct numerical simulations (DNS) create opportunities to leverage data-driven modeling techniques. Albeit flexible, data-driven models still depend on the dataset and the functional form of the model chosen. Increased adoption of such models requires reliable uncertainty estimates both in the data-informed and out-of-distribution regimes. Here, in this work, we employ Bayesian neural networks (BNNs) to capture both epistemic and aleatoric uncertainties in a reacting flow model. In particular, we model the filtered progress variable scalar dissipation rate which plays a key role in the dynamics of turbulent premixed flames. We demonstrate that BNN models can provide unique insights about the structure of uncertainty of the data-driven closure models. We also propose a method for the incorporation of out-of-distribution information in a BNN, which can be used for out-of-distribution query detection. The efficacy of the model is demonstrated by a priori evaluation on a dataset consisting of a variety of flame conditions and fuels.

97 MATHEMATICS AND COMPUTING↗

Foundational Dataset for Developing Large-Sample Stream Temperature Models in the Conterminous United States

This dataset provides inputs, evaluation results, and trained weights from a large-sample Long Short-Term Memory (LSTM) model designed to predict daily stream temperatures across unregulated river reaches in the conterminous United States (CONUS). It includes dynamic meteorological and hydrologic forcings, static physiographic attributes, and model outputs from cross-validation experiments spanning 300 basins. It supports reproducible modeling, direct application for new basins, and provides data suitable for integration with reservoir and river simulations under current and future climates. It contains two .zip files described below · RQ-AI_runs.zip: Model outputs from 10-fold cross-validation experiments, including observed and predicted daily stream temperatures, along with test performance metrics for water years 2017–2019. Two versions are included: 1. Model trained and validated using subbasin-area weighted dynamic features. 2. Model trained and validated using whole-basin area weighted dynamic features. · RQ-AI_inputs.zip: Collection of all formatted dynamic and static predictor datasets (meteorological, hydrologic, and physiographic features) used in model training and analysis. Detailed instructions and data structure is held at the following GitLab repository: https://code.ornl.gov/tempwise/training.

Gomez-Velez, Jesus [Oak Ridge National Laboratory ↗

JGI-Trichoderma v1.0

There is a series of Python and bash scripts to parse genomics datasets used to evaluate the coevolution of gene families and the feature importance of gene families using an SVM classifier. - Cover analysis: takes a list of single-copy genes in a set of genomes, aligns and builds the gene trees to determine if two gene families have a signature of covariation with one another. It parses the files to run phykit cover script described here: https://jlsteenwyk.com/PhyKIT/usage/index.html - SVM-classifier: This Python script is an SVM-based genomic classifier designed for biological data analysis. It combines machine learning with feature selection to identify important genomic markers and classify biological samples. Core Functionality: The script uses Support Vector Machines from scikit-learn to classify genomic data, incorporating SelectKBest for automated feature selection and leave-one-out cross-validation for performance assessment. It operates in multiple modes: feature ranking, optimal combination discovery, and sample prediction. Primary Applications: Genomic sample classification and biomarker discovery Feature importance analysis in high-dimensional biological datasets Prediction of sample categories based on genomic profiles Research applications requiring robust classification of biological data Key Advantages: High-dimensional handling: SVMs excel with genomic data's typical high feature-to-sample ratios Integrated feature selection: Reduces noise and computational overhead while identifying key markers Probability estimation: Provides confidence scores essential for biological interpretation Validation robustness: Leave-one-out cross-validation ensures reliable performance metrics Operational flexibility: Multiple analysis modes support different research phases from exploration to prediction

Stecca Steindorff, Andrei [Lawrence Berkeley Natio↗

Virtual Flight Demonstration of the Stratospheric Dual-Aircraft Platform

A baseline configuration for the dual-aircraft platform (DAP) concept is described and evaluated in a physics-based flight dynamics simulations for two month-long missions as a communications relay in the lower stratosphere above central Florida, within 150-miles of downtown Orlando.The DAP configuration features two large glider-like (130 ft wing span) unmanned aerial vehicles connected via a long adjustable cable (total extendible length of 3000 ft) which effectively sail without propulsion using available wind shear. Use of onboard LiDAR wind profilers to forecast wind distributions are found to be necessary to enable the platform to efficiently adjust flight conditions to remain sailing by finding sufficient wind shear across the platform. The aircraft derive power from solar cells, like a conventional solar aircraft, but also extract wind power using the propeller as a turbine when there is an excess of wind shear available.Month-long atmospheric profiles (at 3-5 min intervals) in the vicinity of 60,000-ft are derived from archived data measured by the 50-Mhz Doppler Radar Wind Profiler at Cape Canaveral and used in the DAP flight simulations. A cursory evaluation of these datasets show that sufficient wind shear for DAP sailing is persistent, suggesting that DAP could potentially sail over 90% of the month-long durations even when limited by modest ascent/descent rates.DAP's novel guidance software uses a non-linear constrained optimization technique to define waypoints such that sailing mode of flight is maintained where possible, and minimal thrust is required where sailing is not practical. A set of constraints are identified which result in waypoints that enable efficient flight (i.e., minimal use of propulsion) over the two month-long flight simulations. Waypoint solutions may need to be tabulated for a wide range of potential atmospheric conditions and stored onboard for quick retrieval on a real DAP.DAP's flight control software uses an unconventional mixture of spacecraft and aircraft control techniques. Flight simulations confirms that this controls approach enables the platform to consistently reach successive waypoints over the month-long flight simulations.The ability of DAP to transition between the sailing mode (i.e., cable tension is high) and standard formation flight (i.e., cable tension is low) is a vital capability (e.g., to enable intermittent turns while stationkeeping). A new method to perform these transitions has been identified and characterized with flight simulation which requires special aircraft modifications.The energy-usage of the DAP configuration during two month-long stationkeeping missions over central Florida (i.e., stationkeeping over Orlando) is evaluated and compared to that of a pure solar aircraft of the same weight and aerodynamic performance. DAP is shown to consistently reduce net propulsion usage while simultaneously increasing solar energy capture.A baseline 700 GHz communications system is described and its performance evaluated for the proposed mission over central Florida. It is found that the variable roll orientation of the aircraft would increase the power required to maintain coverage over the stationkeeping radius of 150 miles (e.g., by as much as 100% when DAP is 150 miles from Orlando), compared to level flight. This effect can be mitigated via additional antenna design complexity or a more restricted stationkeeping radius.

Demonstrations↗

Seasonal variations in composition and sources of atmospheric ultrafine particles in urban Beijing based on near-continuous measurements

Abstract. Understanding the composition and sources of atmospheric ultrafine particles (UFPs) is essential in evaluating their exposure risks. It requires long-term measurements with high time resolution, which are scarce to date. We performed near-continuous measurements of UFP composition during four seasons in urban Beijing using a thermal desorption chemical ionization mass spectrometer, accompanied by real-time size distribution measurements. We found that UFPs in urban Beijing are dominated by organic components, varying seasonally from 68 % to 81 %. CHO organics (i.e., molecules containing carbon, hydrogen, and oxygen) are the most abundant in summer, while sulfur-containing organics, some nitrogen-containing organics, nitrate, and chloride are the most abundant in winter. With the increase of particle diameter, the contribution of CHO organics decreases, while that of sulfur-containing and nitrogen-containing organics, nitrate, and chloride increases. Source apportionment analysis of the UFP organics indicates contributions from cooking and vehicle sources, photooxidation sources enriched in CHO organics, and aqueous/heterogeneous sources enriched in nitrogen- and sulfur-containing organics. The increased contributions of cooking, vehicle, and photooxidation components are usually accompanied by simultaneous increases in UFP number concentrations related to cooking emission, vehicle emission, and new particle formation, respectively, while the increased contribution of the aqueous/heterogeneous composition is usually accompanied by the growth of UFP mode diameters. The highest UFP number concentrations in winter are due to the strongest new particle formation, the strongest local primary particle number emissions, and the slowest condensational growth of UFPs to larger sizes. This study provides a comprehensive understanding of urban UFP composition and sources and offers valuable datasets for the evaluation of UFP exposure risks.

54 ENVIRONMENTAL SCIENCES↗

The Use of Thermal Cameras for Pedestrian Detection

Visible-range camera sensors have been widely used for pedestrian detection. However, most of the methods, which employ visible-range color cameras, do not perform well under low-light and no-light conditions, e.g. during night time. Since the working principle of thermal camera sensors is mainly based on temperature and not light, they have been employed for person detection to overcome the drawbacks of visible-range sensors under these conditions. Every object gives off thermal energy, which is captured by a thermal camera sensor. When an object becomes hotter, it emits more thermal energy, and is therefore captured as much brighter or vice versa. Yet, compared to visible-range cameras, there are many additional challenges that need to be addressed when detecting pedestrians from thermal camera images. These challenges include bright hot objects close to humans, similar pixel values in an image due to weather conditions, or objects that block thermal cameras such as concrete or glass. Glass acts like a mirror for infrared radiation and reflects whatever is in front of the camera. Thus, novel methods are still required to accomplish pedestrian detection task from thermal camera images. To contribute to these efforts, we propose a new method and a modified object detection network incorporating saliency maps of thermal camera images. The features obtained from thermal images and their corresponding saliency maps are combined to obtain richer representations of pedestrian regions, and better detection performance. We perform extensive evaluations on five different datasets to compare the performance of the proposed approach with two baselines. Moreover, we evaluate and compare the transferability of these approaches by doing leave-one-out cross validation across different datasets. Furthermore, the results show that the proposed approach outperforms the baselines, and has better transferability properties across different thermal image datasets.

47 OTHER INSTRUMENTATION↗