Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “evaluation datasets”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

A susceptibility gene signature for ERBB2-driven mammary tumour development and metastasis in collaborative cross mice

Background: Deeper insights into ERBB2-driven cancers are essential to develop new treatment approaches for ERBB2+ breast cancers (BCs). We employed the Collaborative Cross (CC) mouse model to unearth genetic factors underpinning Erbb2-driven mammary tumour development and metastasis. Methods: 732 F1 hybrid female mice between FVB/N MMTV-Erbb2 and 30 CC strains were monitored for mammary tumour phenotypes. GWAS pinpointed SNPs that influence various tumour phenotypes. Multivariate analyses and models were used to construct the polygenic score and to develop a mouse tumour susceptibility gene signature (mTSGS), where the corresponding human ortholog was identified and designated as hTSGS. The importance and clinical value of hTSGS in human BC was evaluated using public datasets, encompassing TCGA, METABRIC, GSE96058, and I-SPY2 cohorts. The predictive power of mTSGS for response to chemotherapy was validated in vivo using genetically diverse MMTV-Erbb2 mice. Findings: Distinct variances in tumour onset, multiplicity, and metastatic patterns were observed in F1-hybrid female mice between FVB/N MMTV-Erbb2 and 30 CC strains. Besides lung metastasis, liver and kidney metastases emerged in specific CC strains. GWAS identified specific SNPs significantly associated with tumour onset, multiplicity, lung metastasis, and liver metastasis. Multivariate analyses flagged SNPs in 20 genes (Stx6, Ramp1, Traf3ip1, Nckap5, Pfkfb2, Trmt1l, Rprd1b, Rer1, Sepsecs, Rhobtb1, Tsen15, Abcc3, Arid5b, Tnr, Dock2, Tti1, Fam81a, Oxr1, Plxna2, and Tbc1d31) independently tied to various tumour characteristics, designated as a mTSGS. hTSGS scores (hTSGSS) based on their transcriptional level showed prognostic values, superseding clinical factors and PAM50 subtype across multiple human BC cohorts, and predicted pathological complete response independent of and superior to MammaPrint score in I-SPY2 study. The power of mTSGS score for predicting chemotherapy response was further validated in an in vivo mouse MMTV-Erbb2 model, showing that, like findings in human patients, mouse tumours with low mTSGS scores were most likely to respond to treatment. Interpretation: Our investigation has unveiled many new genes predisposing individuals to ERBB2-driven cancer. Translational findings indicate that hTSGS holds promise as a biomarker for refining treatment strategies for patients with BC.

60 APPLIED LIFE SCIENCES↗

Data analytics for leak detection in a subcritical boiler

For decades, boiler leaks have been the leading cause of forced outages in the coal-fired unit. The leak occurrences are currently escalating since the existing plants must satisfy faster-ramping rates to support grid operation. Data analytics including Principal Component Analysis, Canonical Variate, and Fisher Discriminant Analysis were combined for detecting and characterizing the leak in a commercial 650 MW subcritical coal-fired power plant. The combined approach was shown to be highly effective in the fault investigation that would not have been easily achieved by an individual technique. The variability in both training and validation datasets was first evaluated using PCA. Then, the CV-FDA was employed to discriminate among faults, and to categorize the processed data into two main groups: no-leak (0) and leak (1), providing the timeframe and location of the leak occurrence. Furthermore, about 8,014 observations from 81 process variables were initially included in the calculation, while the variable count was reduced to 4 with less than 1% misclassification rate in total observations. Finally, the leak was isolated in the waterwall section. Thus, the outcome of this research may provide early detection and isolation of faulty operations in the coal-fired power plant that involves a considerable number of process variables.

20 FOSSIL-FUELED POWER PLANTS↗

Tightly-coupled camera/LiDAR integration for point cloud generation from GNSS/INS-assisted UAV mapping systems

Unmanned aerial vehicles (UAVs) equipped with integrated global navigation satellite systems/inertial navigation systems (GNSS/INS) together with cameras and/or LiDAR sensors are being widely used for topographic mapping in a variety of applications such as precision agriculture, coastal monitoring, and archaeological documentation. Integration of image-based and LiDAR point clouds can provide a comprehensive 3D model of the area of interest. For such integration, ensuring a good alignment between data from the different sources is critical. Although many works have been conducted on this topic, there is still a need for a rigorous integration approach that minimizes the discrepancy between camera and LiDAR data caused by inaccurate system calibration parameters and/or trajectory artifacts. This study proposes an automated tightly-coupled camera/LiDAR integration workflow for GNSS/INS-assisted UAV systems. The proposed strategy is conducted in three main steps. First, an image-based point cloud is generated using a LiDAR/GNSS/INS-assisted structure from motion (SfM) strategy. Then, feature correspondences between image-based and LiDAR point clouds are automatically identified. Finally, an integrated-bundle adjustment procedure including image points, LiDAR raw measurements, and GNSS/INS information is conducted to minimize the discrepancy between point clouds from different sensors while estimating system calibration parameters and refining the trajectory information. The proposed SfM strategy and integration framework are evaluated using five datasets. The SfM results show that using LiDAR data can facilitate feature matching and further increase the number of reconstructed 3D points. The experimental results also illustrate that the developed automated camera/LiDAR integration strategy is capable of accurately estimating system calibration parameters to achieve good alignment among camera/LiDAR data from single/multiple systems. Finally, an absolute accuracy in the range of 3–5 cm is achieved for the image/LiDAR point clouds after the integration process.

42 ENGINEERING↗

Inverse mapping of properties to composition through generative modeling for designing molten salts

Generative modeling (GM) has been increasingly used for the inverse design and optimization of materials, yet its application to molten salt mixtures remains unexplored despite how a successful approach to the inverse design of molten salts would contribute to efficiently exploiting their customizability and unlocking their advantages in applications, such as energy production and energy storage. This work presents a workflow for the inverse design of molten salts with targeted density values, addressing the challenge of representing these complex mixtures in GM. A dataset of critically evaluated molten salt densities is used to train a variational autoencoder coupled with a predictive deep neural network, which then can be used to generate new molten salt compositions with desired density values. The effectiveness of the approach is demonstrated by designing mixtures with distinct densities and validating the predicted values using ab initio molecular dynamics simulations.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Intelligent Sampling of Extreme-Scale Turbulence Datasets for Accurate and Efficient Spatiotemporal Model Training

With the end of Moore’s law and Dennard scaling, efficient training increasingly requires rethinking data volume. Can we train better models with significantly less data via intelligent subsampling? To explore this, we develop SICKLE, a sparse intelligent curation framework for efficient learning, featuring a novel maximum entropy (MaxEnt) sampling approach, scalable training, and energy benchmarking. We compare MaxEnt with random and phase-space sampling on large direct numerical simulation (DNS) datasets of turbulence. Evaluating SICKLE at scale on Frontier, we show that subsampling as a preprocessing step can, in many cases, improve model accuracy and substantially lower energy consumption, with observed reductions of up to 38×.

Brewer, Wes [ORNL] (ORCID:0000000236393956)↗

CYPminer: an automated cytochrome P450 identification, classification, and data analysis tool for genome data sets across kingdoms

Background: Cytochrome P450 monooxygenases (termed CYPs or P450s) are hemoproteins ubiquitously found across all kingdoms, playing a central role in intracellular metabolism, especially in metabolism of drugs and xenobiotics. The explosive growth of genome sequencing brings a new set of challenges and issues for researchers, such as a systematic investigation of CYPs across all kingdoms in terms of identification, classification, and pan-CYPome analyses. Such investigation requires an automated tool that can handle an enormous amount of sequencing data in a timely manner. Results: CYPminer was developed in the Python language to facilitate rapid, comprehensive analysis of CYPs from genomes of all kingdoms. CYPminer consists of two procedures i) to generate the Genome-CYP Matrix (GCM) that lists all occurrences of CYPs across the genomes, and ii) to perform analyses and visualization of the GCM, including pan-CYPomes (pan- and core-CYPome), CYP co-occurrence networks, CYP clouds, and genome clustering data. The performance of CYPminer was evaluated with three datasets from fungal and bacterial genome sequences. Conclusions: CYPminer completes CYP analyses for large-scale genomes from all kingdoms, which allows systematic genome annotation and comparative insights for CYPs. CYPminer also can be extended and adapted easily for broader usage.

59 BASIC BIOLOGICAL SCIENCES↗

2D reactive transport model of shale chemical weathering and biogeochemical fluxes along a mountainous hillslope, East River Watershed, Colorado: Input files and simulation results

This data package contains input files and simulation results for a two-dimensional (2D) reactive transport model used to quantitatively analyze the coupled hydrological and biogeochemical processes governing shale weathering and associated biogeochemical fluxes under realistic environmental conditions in the high-elevation East River Watershed. These data support the conclusions presented in Stolze et al. (Water Resources Research, under review), "Model-based interpretation of solute exports and carbon partitioning during shale weathering in a mountainous hillslope". The model simulates atmospheric-subsurface gas exchange, subsurface water flow, and shale weathering processes under dynamic, year-scale conditions along a shale-underlain hillslope located in the East River watershed. The simulations were performed using the PFLOTRAN flow and reactive transport code and executed on the Perlmutter supercomputer to leverage its large-scale parallel computing capabilities. The data package contains two zipped folders, "model_input_files" and "simulation_results", and one readme.txt file. "model_input_files" contains the necessary input files to run the calibrated base-base model presented in Stolze et al. (Water Resources Research, under review). "simulation_results" contains a single hdf5 file ("Output_2D_hillslope_model.h5") which includes the results of simulation performed using the base-case model. This file can be opened with HDFView 3.1.4, Python, or MATLAB. "readme.txt" contains relevant information about the base-case model and provides guidelines on how to run the associated input files provided in the folder "model_input_files". Furthermore, readme.txt provides information regarding the model results provided in "Output_2D_hillslope_model.h5" such as matrix dimensionality and output units. Field datasets used to evaluate model performance were collected at three monitoring wells located along a hillslope transect (PLM1, PLM2, and PLM3). Dissolved ion concentration data were collected from November 2016 to October 2021 for Ca, Mg, DIC, Na, K, SO4 (Dong et al., 2025 - dic_npoc_data_2014_2024.zip - DOI:10.15485/1660459; Williams et al., 2025 - anion_data_2014_2024.zip - DOI:10.15485/1668054; Dong et al., 2025 - cation_data_2014_2024.zip - DOI:10.15485/1668055). Note that we used the files named er_PLM1_xx_yy, er_PLM2_xx_yy, and er_PLM3_xx_yy where xx stands for the name of the aqueous species and yy stands for the depth where the measurements were performed. Soil water content ([0 - 1] m) and water table depth were collected from November 2016 to October 2021 (Wan et al., 2024 - Dynamic_water_table__depthsFig2b.csv and Soil_water_content_Fig4e.csv - DOI:10.15485/2322567). Gaseous CO2 concentration were collected from October 2020 to December 2021(Wan et al., 2024 - Soil_CO2_concentrations_Fig4h.csv - DOI:10.15485/2322567) Gaseous CO2 flux from the subsurface to the atmosphere were collected in the vicinity of PLM2 from October 2019 to May 2022 (Wu et al., 2025). Soil microbial biomass concentration was measured from August 2016 to June 2017 (Sorensen et al., 2019 - 2017_East_River_Pumphouse_Microbial_Biomass__1_.csv - DOI:10.15485/1577267) All field data are published as CSV files compatible with Microsoft Excel, MATLAB, and Python, or as text files. The coordinates of the monitoring wells and the CO2(g) flux sensor in the coordinate system WGS84 are: -PLM1: [38.9197710 ; -106.9492750] -PLM2: [38.9201580 ; -106.9487170] -PLM3: [38.9207843 ; -106.9483668] -PLM4: 38.9210060 ; -106.9479528] -CO2(g) flux sensor: [38.9199180 ; -106.9489906] ------------------------------------------------------------------------------------------- This work was supported by the Watershed Function Science Focus Area at Lawrence Berkeley National Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-05CH11231. This research used resources of the National Energy Research Scientific Computing Center (NERSC), a Department of Energy User Facility using NERSC award BER-ERCAP 23980, BER-ERCAP 28550, and BER-ERCAP 33789.

54 ENVIRONMENTAL SCIENCES↗

Gradient-Based Novelty Detection Boosted by Self-Supervised Binary Classification

Novelty detection aims to automatically identify out-of-distribution (OOD) data, without any prior knowledge of them. It is a critical step in data monitoring, behavior analysis and other applications, helping enable continual learning in the field. Conventional methods of OOD detection perform multi-variate analysis on an ensemble of data or features, and usually resort to the supervision with OOD data to improve the accuracy. In reality, such supervision is impractical as one cannot anticipate the anomalous data. In this paper, we propose a novel, self-supervised approach that does not rely on any pre-defined OOD data: (1) The new method evaluates the Mahalanobis distance of the gradients between the in-distribution and OOD data. (2) It is assisted by a self-supervised binary classifier to guide the label selection to generate the gradients, and maximize the Mahalanobis distance. In the evaluation with multiple datasets, such as CIFAR-10, CIFAR-100, SVHN and TinyImageNet, the proposed approach consistently outperforms state-of-the-art supervised and unsupervised methods in the area under the receiver operating characteristic (AUROC) and area under the precision-recall curve (AUPR) metrics. We further demonstrate that this detector is able to accurately learn one OOD class in continual learning.

Sun, Jingbo↗

An Exact Algorithm for the Linear Tape Scheduling Problem

Magnetic tapes are often considered as an outdated storage technology, yet they are still used to store huge amounts of data. Their main interests are a large capacity and a low price per gigabyte, which come at the cost of a much larger file access time than on disks. With tapes, finding the right ordering of multiple file accesses is thus key to performance. Moving the reading head back and forth along a kilometer long tape has a non-negligible cost and unnecessary movements thus have to be avoided. However, the optimization of tape request ordering has rarely been studied in the scheduling literature, much less than I/O scheduling on disks. For instance, minimizing the average service time for several read requests on a linear tape remains an open question. Therefore, in this paper, we aim at improving the quality of service experienced by users of tape storage systems, and not only the peak performance of such systems. To this end, we propose a reasonable polynomial-time exact algorithm while this problem and simpler variants have been conjectured NP-hard. We also refine the proposed model by considering U-turn penalty costs accounting for inherent mechanical accelerations. Then, we propose a low-cost variant of our optimal algorithm by restricting the solution space, yet still yielding an accurate suboptimal solution. Finally, we compare our algorithms to existing solutions from the literature on logs of the mass storage management system of a major datacenter. This allows us to assess the quality of previous solutions and the improvement achieved by our low-cost algorithm. Aiming for reproducibility, we make available the complete implementation of the algorithms used in our evaluation, alongside the dataset of tape requests that is, to the best of our knowledge, the first of its kind to be publicly released.

Honoré, Valentin↗

National Virtual Biotechnology Laboratory: Report on Rapid R&D Solutions to the COVID-19 Crisis

With funding from the CARES Act, the U.S Department of Energy (DOE) established the National Virtual Biotechnology Laboratory (NVBL) in March 2020 to address key challenges associated with the COVID-19 crisis. NVBL brought together the broad scientific and technical expertise and resources of DOE’s 17 national laboratories to help tackle medical supply short ages, discover potential drugs to fight the virus, develop and validate COVID-19 testing methods, model disease spread and impact across the nation, and understand virus transport in buildings and the environment. National laboratory resources leveraged for this effort include a suite of world-leading user facilities broadly available to the research community, such as light and neutron sources, nanoscale science research centers, sequencing and biocharacterization facilities, and high-performance computing facilities. Within months, NVBL teams produced innovations in materials and advanced manufacturing that mitigated shortages in test kits and personal protective equipment (PPE), creating nearly 1,000 new jobs. They used DOE’s high-performance computers and light and neutron sources to identify promising candidates for antibodies and antivirals that universities and drug companies are now evaluating. NVBL researchers also developed new diagnostic targets and sample collection approaches, and supported U.S. Food and Drug Administration (FDA), Centers for Disease Control and Prevention (CDC), and U.S. Department of Defense (DoD) efforts to establish national guidelines used in administering millions of tests. Researchers used artificial intelligence and high-performance computing to produce near-real-time data analysis to forecast disease transmission, stress on public health infrastructure, and economic impact, which supported decision-makers at the local, state, and national levels. NVBL teams also studied how to control indoor virus movement to minimize uptake and protect human health. NVBL’s accomplishments demonstrate not only the powerful resource represented by DOE’s national laboratories working together to meet national needs, but also the effectiveness of the integrated NVBL framework for rapidly responding to emergencies with research and development (R&D) solutions. As the fight against COVID continues, sustained efforts are needed to confront this pandemic as well as future threats. Examples include: 1) Establishing “supply chains on demand” to meet emergency production needs by leveraging the materials and manufacturing expertise of DOE national laboratories and developing advances in electronics, sensing, robotics, and automation capabilities; 2) Improving the speed and robustness of drug discovery by integrating experimental platforms with DOE’s computational and experimental user facilities, which provide unique resources to support the discovery of high-potential therapeutic agents; 3) Protecting public, environmental, and animal health by developing new testing protocols and instrumentation adaptable to diverse sample types (both physiological and environmental) to quickly detect a wide range of pathogens and monitor other biorisks; 4) Supporting near-real-time data needs of decision-makers at the local, regional, state, and national levels by advancing data curation, analysis, and modeling using artificial intelligence and new data science tools for managing and evaluating large diverse datasets; 5) Harnessing DOE’s expertise in environmental modeling to design rooms and air handling for offices, classrooms, restaurants, and other structures to minimize biorisk transmissions. Going forward, NVBL is poised to apply the unique capabilities and expertise of the national laboratory complex to future national and international emergencies, both natural and engineered. Through this framework, the Office of Science will continue to be an integral component of agency wide efforts to prepare for and respond to biorisks and other crises.

42 ENGINEERING↗

Rapid identification of enteric bacteria from whole genome sequences using average nucleotide identity metrics

Identification of enteric bacteria species by whole genome sequence (WGS) analysis requires a rapid and an easily standardized approach. We leveraged the principles of average nucleotide identity using MUMmer (ANIm) software, which calculates the percent bases aligned between two bacterial genomes and their corresponding ANI values, to set threshold values for determining species consistent with the conventional identification methods of known species. The performance of species identification was evaluated using two datasets: the Reference Genome Dataset v2 (RGDv2), consisting of 43 enteric genome assemblies representing 32 species, and the Test Genome Dataset (TGDv1), comprising 454 genome assemblies which is designed to represent all species needed to query for identification, as well as rare and closely related species. The RGDv2 contains six Campylobacter spp., three Escherichia/Shigella spp., one Grimontia hollisae, six Listeria spp., one Photobacterium damselae, two Salmonella spp., and thirteen Vibrio spp., while the TGDv1 contains 454 enteric bacterial genomes representing 42 different species. The analysis showed that, when a standard minimum of 70% genome bases alignment existed, the ANI threshold values determined for these species were ≥95 for Escherichia/Shigella and Vibrio species, ≥93% for Salmonella species, and ≥92% for Campylobacter and Listeria species. Using these metrics, the RGDv2 accurately classified all validation strains in TGDv1 at the species level, which is consistent with the classification based on previous gold standard methods.

59 BASIC BIOLOGICAL SCIENCES↗

Understanding the compound flood risk along the coast of the contiguous United States

Abstract. Compound flooding is a type of flood event caused by multiple flood drivers. The associated risk has usually been assessed using statistics-based analyses or hydrodynamics-based numerical models. This study proposes a compound flood (CF) risk assessment (CFRA) framework for coastal regions in the contiguous United States (CONUS). In this framework, a large-scale river model is coupled with a global ocean reanalysis dataset to (a) evaluate the CF exposure related to the coastal backwater effects on river basins, and (b) generate spatially distributed data for analyzing the CF hazard using a bivariate statistical model of river discharge and storm surge. The two kinds of risk are also combined to achieve a holistic understanding of the continental-scale CF risk. The estimated CF risk shows remarkable inter- and intra-basin variabilities along the CONUS coast with more variabilities in the CF hazard over the US west and Gulf coastal basins. Different risk assessment methods present significantly different patterns in a few key regions such as the San Francisco Bay area, the lower Mississippi River, and Puget Sound. Our results highlight the need to weigh different CF risk measures and avoid using single statistics-based or hydrodynamics-based CFRAs. Uncertainty sources in these CFRAs include the use of gauge observations, which cannot account for the flow physics or resolve the spatial variability of risks, and underestimations of the flood extremes and the dependence of CF drivers in large-scale models, highlighting the importance of understanding the CF risks for developing a more robust CFRA.

54 ENVIRONMENTAL SCIENCES↗

Venus gravity field - Pioneer Venus Orbiter navigation results

The gravity field of Venus has been modeled by a spherical harmonic expansion of the potential to degree and order seven. The estimates of these coeficients were obtained by combining information from 43 short arcs (4 hr) of line-of-sight Doppler data centered at periapsis. The data arcs were distributed in longitude and time over more than two circulations of Venus by the Pioneer Venus Orbiter subperiapsis point which was confined to the band of latitudes from 14 deg N to 17 deg N. Convergence of the solution has been assured by iterating upon the initial estimate. All estimates were performed with zero a priori information on the gravity coefficients. Since the altitude of periapsis for most of the orbits was within the sensible Venusian atmosphere, drag effects on the estimated harmonics have been removed using an exponential atmosphere density model. Estimates of the mass parameter (GM) of Venus using this dataset are also evaluated.

Williams, B. G.↗

Comparative solar EUV flux for the San Marco ASSI

The Airglow and Solar Spectrometer Instrument (ASSI) on the San Marco D/L satellite has measured solar extreme ultraviolet irradiances. The data are currently being released for analysis. As a preliminary step in evaluating this important dataset, modeled solar irradiances from 4 to 105 nm are presented for comparison to the San Marco data. The comparable flux for March-December 1988 is obtained from a revised and extended empirical solar EUV model derived from OSO 1, OSO 3, OSO 4, OSO 6, AEROS A, and AE-E satellite and six rocket flight datasets. Solar rotational features are prominent on several occasions in the model time series. A useful example is the modeled integrated flux between 30-31 nm which includes the Si XI (30.3-nm) and He II (30.4-nm) irradiance. The modeled flux in this 1-nm range shows both an absolute 22 percent increase from beginning to end of mission and a solar rotational variability with a typical peak-to-valley ratio of 14 percent.

Tobiska, W. K.↗

Towards A Representation of Vertically Resolved Ozone Changes in Reanalyses

The Solar Backscatter Ultraviolet Radiometer (SBUV) instruments on NASA and NOAA spacecraft provide a long-term record of total-column ozone and deep-layer partial columns since about 1980. These data have been carefully processed to extract long-term trends and offer a valuable resource for ozone monitoring. Studies assimilating limb-sounding observations in the Goddard Earth Observing System (GEOS) data assimilation system (DAS) demonstrate that vertical ozone gradients in the upper troposphere and lower stratosphere (UTLS) are much better represented than with the deep-layer SBUV observations. This is exemplified by the use of retrieved ozone from the EOS Microwave Limb Sounder (EOS-MLS) instrument in the MERRA-2 reanalysis, for the period after 2004. This study examines the potential for extending the use of limb-sounding observations at earlier times and into the future, so that future reanalyses may be more applicable to the study of long-term ozone changes.Historical data are available from NASA instruments: the Limb Infrared Monitor of the Stratosphere (LIMS: 1978-1979); the Upper Atmospheric Research Satellite (UARS: 1991-1995); Sounding of the Atmosphere using Broadband Emission Radiometry (SABER: 2000-onwards). For the post EOS-MLS period, the joint NASA-NOAA Ozone Monitoring and Profiling Suite Limb Profiler (OMPS-LP) instrument was launched on the Suomi-NPP platform in 201x and is planned for future platforms. This study will examine two aspects of these data pertaining to future reanalyses. First, the feasibility of merging the EOS-MLS and OMPS-LP instruments to provide a long-term record that extends beyond the potential lifetime of EOS-MLS. If feasible, this would allow for long-term monitoring of ozone recovery in a three-dimensional reanalysis context. Second, the skill of the GEOS DAS in ingesting historical data types will be investigated. Because these do not overlap with EOS-MLS, use will be made of system statistics and evaluation using independent datasets. Impacts of using a complete ozone chemistry module will also be considered.

ML↗

Assessment of NO2 Observations During DISCOVER-AQ and KORUS-AQ Field Campaigns

NASA’s Deriving Information on Surface Conditions from Column and Vertically Resolved Observations Relevant to Air Quality (DISCOVER-AQ, conducted in 2011–2014) campaign in the United States and the joint NASA and National Institute of Environmental Research (NIER) Korea–United States Air Quality Study (KORUS-AQ, conducted in 2016) in South Korea were two field study programs that provided comprehensive, integrated data sets of airborne and surface observations of atmospheric constituents, including nitrogen dioxide (NO2), with the goal of improving the interpretation of spaceborne remote sensing data. Various types of NO2 measurements were made, including in situ concentrations and column amounts of NO2 using ground- and aircraft-based instruments, while NO2 column amounts were being derived from the Ozone Monitoring Instrument (OMI) on the Aura satellite. This study takes advantage of these unique datasets by first evaluating in situ data taken from two different instruments on the same aircraft platform, comparing coincidently sampled profile-integrated columns from aircraft spirals with remotely sensed column observations from ground-based Pandora spectrometers, intercomparing column observations from the ground (Pandora), aircraft (in situ vertical spirals), and space (OMI), and evaluating NO2 simulations from coarse Global Modeling Initiative (GMI) and high-resolution regional models. We then use these data to interpret observed discrepancies due to differences in sampling and deficiencies in the data reduction process. Finally, we assess satellite retrieval sensitivity to observed and modeled a priori NO2 profiles. Contemporaneous measurements from two aircraft instruments that likely sample similar air masses generally agree very well but are also found to differ in integrated columns by up to 31.9 %. These show even larger differences with Pandora, reaching up to 53.9 %, potentially due to a combination of strong gradients in NO2 fields that could be missed by aircraft spirals and errors in the Pandora retrievals. OMI NO2 values are about a actor of 2 lower in these highly polluted environments due in part to inaccurate retrieval assumptions (e.g., a priori pro-files) but mostly to OMI’s large footprint (>312 km2).

Nitrogen dioxide↗

Fine particulate concentrations over East Asia derived from aerosols measured by the Advanced Himawari Imager using machine learning

Fine particulate matter with a diameter below 2.5 μm (PM 2.5 ) is deleterious to the cardiovascular and respiratory systems. It is often difficult to assess the effects of PM 2.5 on human health over regions with limited ground monitoring sites, especially in East Asia. As an alternative, we estimated near-surface PM 2.5 concentrations by analyzing Advanced Himawari Imager (AHI) Yonsei Aerosol Retrieval (YAER) products. This study incorporates daytime data for East Asia covering the Korean Peninsula, China, Japan, Southeast Asia, and southern Mongolia. We collocated AHI YAER product pixels with meteorological, land-cover, and other ancillary data for the period from March 2018 to February 2019. To estimate PM 2.5 concentrations over wide areas spanning many countries displaying various relationships between aerosol optical depth and PM 2.5 , monthly models were developed by considering both the spatial and temporal characteristics of ground-based PM 2.5 measurements. Random forest machine learning model estimated ground-level mass concentrations of PM 2.5 ; subsequent 10-fold cross validation (CV) yielded a CV R 2 value of 0.81 and a CV root mean squared error (RMSE) of 12.3 μg m -3 . We investigated the spatial pattern of PM 2.5 concentrations over multiple countries and seasonal variation in PM 2.5 concentrations. Diurnal variation of a severe PM 2.5 event in the Korean Peninsula was investigated as a case study. The model captured the extremely heterogeneous spatial distribution of PM 2.5 concentrations peaked around local noon. To measure the capability of the developed model to estimate PM 2.5 concentrations in areas with few in-situ data, its predictive performance was evaluated using a dataset independent of the training process with an R 2 of 0.60 and RMSE of 8.18 μg m −3 . This study demonstrates the potential for satellite-based PM 2.5 estimation for areas with insufficient measuring stations.

Pm2.5↗

Using Intelligent Targeting to increase the science return of a Smart Ice Storm Hunting Radar

Smart Ice Cloud Sensing (SMICES) is a small-sat concept in which a radar intelligently targets ice storms based on information collected by a lookahead radiometer. Often space observations are performed by continuously collecting data from an instrument aimed at nadir (e.g. directly below the space platform). However, if the platform has the ability to assess science utility of features being overflown, an intelligent measurement scheme can improve science return. This can be achieved by controlling the on/off state of the instrument if it is not able to continuously operate (e.g. due to energy or thermal constraints), and by allowing the instrument to view off nadir if it has pointing capabilities.In the case of SMICES, power constraints and the rarity of storms means that with blind nadir targeting SMICES would collect a limited amount of ice storm radar data. The algorithms proposed acquire measurements to maximize acquired high interest storms while concurrently collecting a background sampling of all features. We use a cloud classification system to identify five different cloud types. Six algorithms ranging from “blind” to more selective are described and results from evaluation on a dataset of 13 ground swaths covering 72,399,600 km2 of data are presented. This data is from high quality science simulations that contain all five cloud types and multiple storms. When utilizing the radiometer’s lookahead and the full range of the radar the results show a 23.7x and 1.9x increase over the base algorithm in the most and second most important cloud types respectively.

Cooke, Caitlyn↗