Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Principal component analysis”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12

Enhancing physicochemical, bioactive, and nutritional properties of sweet potatoes: Ultrasonic contact drying with slot jet nozzles compared to hot-air drying and freeze drying

Sweet potatoes are a rich source of nutrients and bioactive compounds, but their quality can be impacted by the drying process. This study investigates the impact of slot jet reattachment (SJR) nozzle and ultrasound (US) combined drying (SJR + US) on sweet potato quality, compared to freeze-drying (FD), SJR drying, and hot air drying (HAD). SJR + US drying at 50 °C closely resembled FD in enhancing quality attributes and outperformed HAD and SJR in key areas such as rehydration, shrinkage ratios, and nutritional composition. Notably, SJR + US at 50 °C produced the highest total starch (36.84 g/100 g), total dietary fiber (8.48 g/100 g), total phenolic content (158.19 mg GAE/100 g), total flavonoid content (119.08 mg QE/g), DPPH antioxidant activity (6.44 μmol TE/g), β-carotene (31.98 mg/100 g), and vitamin C (5.27 mg/100 g). It also exhibited higher glass transition temperatures (Tg: 14.49 °C), indicating better stability at room temperature. The hardness values for SJR + US samples were similar to FD, while HAD samples had the highest hardness. SJR + US at 50 °C resulted in the lowest total color changes (ΔE), indicating minimal impact on appearance. Additionally, FTIR analysis revealed that peaks in specific spectral regions indicated superior preservation of bioactive compounds in SJR + US samples compared to other methods, which was also confirmed by principal component analysis (PCA) and heatmap visualization. Overall, these findings suggest that SJR + US is an effective alternative to conventional drying techniques, significantly improving the quality of dried sweet potatoes.

Color↗

Anthropometric Accommodation in Space Suit Design

Design requirements for next generation hardware are in process at NASA. Anthropometry requirements are given in terms of minimum and maximum sizes for critical dimensions that hardware must accommodate. These dimensions drive vehicle design and suit design, and implicitly have an effect on crew selection and participation. At this stage in the process, stakeholders such as cockpit and suit designers were asked to provide lists of dimensions that will be critical for their design. In addition, they were asked to provide technically feasible minimum and maximum ranges for these dimensions. Using an adjusted 1988 Anthropometric Survey of U.S. Army (ANSUR) database to represent a future astronaut population, the accommodation ranges provided by the suit critical dimensions were calculated. This project involved participation from the Anthropometry and Biomechanics facility (ABF) as well as suit designers, with suit designers providing expertise about feasible hardware dimensions and the ABF providing accommodation analysis. The initial analysis provided the suit design team with the accommodation levels associated with the critical dimensions provided early in the study. Additional outcomes will include a comparison of principal components analysis as an alternate method for anthropometric analysis.

Rajulu, Sudhakar↗

Evaluation of Meterorite Amono Acid Analysis Data Using Multivariate Techniques

The amino acid distributions in the Murchison carbonaceous chondrite, Mars meteorite ALH84001, and ice from the Allan Hills region of Antarctica are shown, using a multivariate technique known as Principal Component Analysis (PCA), to be statistically distinct from the average amino acid compostion of 101 terrestrial protein superfamilies.

meteorites Amino acids Multivariate Analysis↗

Lipid droplet-associated proteins in alcohol-associated fatty liver disease: A proteomic approach

The earliest manifestation of alcohol-associated liver disease (ALD) is steatosis characterized by deposition of fat in specialized organelles called lipid droplets (LDs). While alcohol administration causes a rise in LD numbers in the hepatocytes, little is known regarding their characteristics that allow their accumulation and size to increase. The aim of the present study is to gain insights into underlying pathophysiological mechanisms by investigating the ethanol-induced changes in hepatic LD proteome as a function of LD size. Adult male Wistar rats (180–200 g BW) were fed with ethanol liquid diet for 6 weeks. At sacrifice, large-, medium-, and small-sized hepatic LD subpopulations (LD1, LD2, and LD3, respectively) were isolated and subjected to morphological and proteomic analyses. Morphological analysis of LD1-LD3 fractions of ethanol-fed rats clearly demonstrated that LD1 contained larger LDs compared with LD2 and LD3 fractions. Our preliminary results from principal component analysis showed that the proteome of different-sized hepatic LD fractions was distinctly different. Proteomic data analysis identified over 2000 proteins in each LD fraction with significant alterations in protein abundance among the three LD fractions. Among the altered proteins, several were related to fat metabolism, including synthesis, incorporation of fatty acid, and lipolysis. Ingenuity pathway analysis revealed increased fatty acid synthesis, fatty acid incorporation, LD fusion, and reduced lipolysis in LD1 compared to LD3. Overall, the proteomic findings indicate that the increased level of protein that facilitates fusion of LDs combined with an increased association of negative regulators of lipolysis dictates the generation of large-sized LDs during the development of alcohol-associated hepatic steatosis. Several significantly altered proteins were identified in different-sized LDs isolated from livers of ethanol-fed rats. Ethanol-induced increases in specific proteins that hinder LD lipid metabolism led to the accumulation and persistence of large-sized LDs in the liver.

60 APPLIED LIFE SCIENCES↗

Correlations between Texture Profile Analysis and Sensory Evaluation of Cured Largemouth Bass Meat (Micropterus salmoides)

Texture is an important factor in evaluating the quality of aquatic products. To evaluate the texture properties of cured large mouth bass, edible sodium chloride (0, 1.0, 2.0, and 3.0%) was smeared to the bass meat. Texture profile analysis (TPA) and sensory evaluation were performed to evaluate the quality of the cured samples, and the correlations of the indexes in the two methods were analyzed. Two principal components were obtained from the TPA indexes and sensory evaluation indexes, and the cumulative variance contribution rates were 73.87% and 72.99%, respectively. Results from the principal component analysis showed that the main indicators that affected the TPA were gumminess and springiness, while those that affected sensory evaluation were chewiness and adhesiveness. The TPA index and sensory evaluation could be effectively improved when the sodium chloride added to the bass meat was 1%. In the correlation analysis, sensory springiness was negatively correlated with TPA hardness ( P < 0.05 , r = −0.553) but positively correlated with TPA chewiness ( P < 0.05 , r = 0.596). After stepwise regression analysis, the prediction equation between the sensory springiness and TPA hardness was obtained as SSp=5.770−0.002Ha. These results provide a basis for predicting the quality of large mouth bass cured products.

Li, Meijin↗

Tropical intraseasonal oscillation and its prediction by the NMC operational model

Results are presented of an investigation of the tropical intraseasonal oscillation (ISO) and its impact on the extended-range forecast in the NMC operational model during Phase II (14 December 1986-31 March 1987) of the Dynamical Extended Range Forecast. Based on principal component analysis of the velocity potential and streamfunction, evidence was found of tropical-extratropical interaction associated with the ISO. The NMC model possess significant forecast skills for the principal streamfunction and velocity potential modes up to the first ten days. Results of the error growth analysis suggest that the principal modes of velocity potential have large errors comparable to the model random errors. By comparison, the initial errors in the streamfunction are much smaller. The error growth for both tropical and extratropical modes are found to be significantly suppressed during periods of strong ISO relative to periods of weak ISO. The increase in extratropical forecast skill is likely due to (1) the model's ability to better capture ISO signals in the tropics and (2) the increased coupling between the tropics and extratropics during periods of strong ISO.

Lau, K.-M.↗

Spectral Comparison and Stability of Red Regions on Jupiter

A study of absolute color on Jupiter from Hubble Space Telescope imaging data shows that the Great Red Spot (GRS) is not the reddest region of the planet. Rather, a transient red cyclone visible in 1995 and the North Equatorial Belt both show redder spectra than the GRS (i.e., more absorption at blue and green wavelengths). This cyclone is unique among vortices in that it is intensely colored yet low altitude, unlike the GRS. Temporal analysis shows that the darkest regions of the NEB are relative constant in color from 1995 to 2008, while the slope of the GRS core may vary slightly. Principal component analysis shows several spectral components are needed, in agreement with past work, and further highlights the differences between regions. These color differences may be indicative of the same chromophore(s) under different conditions, such as mixing with white clouds, longer UV irradiation at higher altitude, and thermal processing, or may indicate abundance variations in colored compounds. A single compound does not fit the spectrum of any region well and mixes of multiple compounds including NH4SH, photolyzed NH3, hydrocarbons, and possibly P4, are likely needed to fully match each spectrum.

cyclonic↗

Quantifying Drivers of Methane Hydrobiogeochemistry in a Tidal River Floodplain System

The influence of coastal ecosystems on global greenhouse gas (GHG) budgets and their response to increasing inundation and salinization remains poorly constrained. In this study, we have integrated an uncertainty quantification (UQ) and ensemble machine learning (ML) framework to identify and rank the most influential processes, properties, and conditions controlling methane behavior in a freshwater floodplain responding to recently restored seawater inundation. Our unique multivariate, multiyear, and multi-site dataset comprises tidal creek and floodplain porewater observations encompassing water level, salinity, pH, temperature, dissolved oxygen (DO), dissolved organic carbon (DOC), total dissolved nitrogen (TDN), partial pressure of carbon dioxide (pCO 2 ), nitrous oxide (pN 2 O), methane (pCH 4 ), and the stable isotopic composition of methane (δ 13 CH 4 ). Additionally, we incorporated topographical data, soil porosity, hydraulic conductivity, and water retention parameters for UQ analysis using a previously developed 3D variably saturated flow and transport floodplain model for a physical mechanistic understanding of factors influencing groundwater levels and salinity and, therefore, CH 4 . Principal component analysis revealed that groundwater level and salinity are the most significant predictors of overall biogeochemical variability. The ensemble ML models and UQ analyses identified DO, water level, salinity, and temperature as the most influential factors for porewater methane levels and indicated that approximately 80% of the total variability in hourly water levels and around 60% of the total variability in hourly salinity can be explained by permeability, creek water level, and two van Genuchten water retention function parameters: the air-entry suction parameter α and the pore size distribution parameter m. These findings provide insights on the physicochemical factors in methane behavior in coastal ecosystems and their representation in local- to global-scale Earth system models.

54 ENVIRONMENTAL SCIENCES↗

Principal Component and Machine Learning Approach to Gap Fill Hyperspectral Ocean Color Satellite Retrievals

Retrievals of ocean color properties from space are important for monitoring the health of the ocean ecosystem but such retrievals can be limited spatially due to conditions such as clouds, aerosols, and sun glint. Gap filling of ocean color retrievals is typically performed by combining retrievals from multiple satellites or temporally averaging multiple days of retrievals. Despite these techniques large gaps still exist posing challenges for near real time monitoring of events like harmful algae blooms. To address these limitations, we developed a spatial gap filling approach applying machine learning approach to hyperspectral instruments to learn how to perform an atmospheric correction under challenging retrieval conditions. In this approach a principal component analysis is used to decompose the hyperspectral measurements into spectral components that describe the scattering and absorption of the atmosphere mixed with the surface spectral signatures. The coefficients of the principal components are used to train a neural network to predict ocean color properties derived from a standard MODIS ocean color algorithm. We apply the approach to two hyperspectral UV/VIS sensors, the Ozone Monitoring Instrument (OMI) and TROPOspheric Monitoring Instrument (TROPOMI) to show that it can be used to estimate ocean color properties such as chlorophyll, remote sensing reflectance, and fluorescence line height. This method could be used as a gap-filling technique for the future Ocean Color Instrument (OCI) onboard upcoming NASA's Plankton, Aerosol Cloud, ocean Ecosystem (PACE) satellite to provide additional information for monitoring the health of our global oceans. Additionally, it could be applied to the first NASA and Smithsonian geostationary Tropospheric Emissions: Monitoring of Pollution (TEMPO) spectrometer to better understand diurnal variability in inland and coastal ocean ecology.

Zachary Fasnacht↗

Applications of array processors in the analysis of remote sensing images

The architectures, programming characteristics, and ranges of application of past, present, and planned array processors for the digital processing of remote-sensing images are compared. Such functions as radiometric and geometric corrections, principal-components analysis, cluster coding, histogram generation, grey-level mapping, convolution, classification, and mensuration and modeling operations are considered, and both pipeline-type and single-instruction/multiple-data-stream (SIMD) arrays are evaluated. Numerical results are presented in a table, and it is found that the pipeline-type arrays normally used with minicomputers increase their speed significantly at low cost, while even further gains are provided by the more expensive SIMD arrays. Most image-processing operations become I/O-limited when SIMD arrays are used with current I/O devices.

Ramapriyan, H. K.↗

Evaluation of Correction Methods for NASA GeneLab Transcriptomic Datasets

Conducting space biology experiments aboard the International Space Station, particularly those utilizing complex model organisms like mice, is expensive and difficult due to limited crew availability, hardware, and space. As a result, sample numbers from these studies are low, reducing the statistical power of any one experiment. Aggregating spaceflight datasets serves as a method to increase sample numbers, allowing for novel insights through bioinformatic analysis of ‘omics data from merged datasets. However, aggregating datasets can introduce unwanted variation including 1) differences in sample handling, processing, and sequencing platforms between datasets (technical variation) as well as 2) differences in experimental design between datasets such as sex or age of the model organism used. In the present study, NASA GeneLab-hosted RNAseq datasets from rodent liver tissues were used to evaluate several statistical methods to correct for this unwanted variation through two approaches, reference-based and standard. The following correction algorithms were applied with (reference-based) and/or without (standard) considering Universal Mouse RNA Reference samples: ComBat and ComBat_seq from the SVA package, median polish, empirical Bayes, and ANOVA-based algorithms from the MBatch package, and negative binomial regression normalization in the DESeq2 package. For each approach, after the correction algorithm was applied, differential gene expression (DGE) analysis of flight and ground control samples was performed with the combined data. The robustness of each tool was evaluated using BatchQC, to determine statistical differences between datasets before and after correction, Principal Component Analysis, to evaluate global gene expression in samples before and after correction, and by comparing DGE analysis of individual datasets and combined datasets before and after correction. The results showed that the reference-based approach introduced several additional (and likely artificial) DEGs when compared with the standard approach. Thus, the most robust standard correction will be implemented in the GeneLab Visualization 2.0 platform when datasets are combined.

GeneLab, RNA-seq, Batch Correction↗

Evaluation of Correction Methods for NASA GeneLab Transcriptomic Datasets

Conducting space biology experiments aboard the International Space Station, particularly those utilizing complex model organisms like mice, is expensive and difficult due to limited crew availability, hardware, and space. As a result, sample numbers from these studies are low, reducing the statistical power of any one experiment. Aggregating spaceflight datasets serves as a method to increase sample numbers, allowing for novel insights through bioinformatic analysis of ‘omics data from merged datasets. However, aggregating datasets can introduce unwanted variation including 1) differences in sample handling, processing, and sequencing platforms between datasets (technical variation) as well as 2) differences in experimental design between datasets. In the present study, NASA GeneLab-hosted RNAseq datasets from mouse liver tissues were used to evaluate several statistical methods to correct for this unwanted variation through two approaches, reference-based and standard. The following correction algorithms were applied with (reference-based) and/or without (standard) considering Universal Mouse RNA Reference samples: ComBat and ComBat_seq from the SVA package, median polish, empirical Bayes, and ANOVA-based algorithms from the MBatch package, and negative binomial regression normalization in the DESeq2 package. For each approach, after the correction algorithm was applied, differential gene expression (DGE) analysis of flight and ground control samples was performed with the combined data. The robustness of each tool was evaluated using BatchQC to determine statistical differences between datasets before and after correction, Principal Component Analysis to evaluate global gene expression in samples before and after correction, and by comparing DGE analysis of individual datasets and combined datasets before and after correction. The results showed that the reference-based approach introduced several additional (and likely artificial) DEGs when compared with the respective standard approach. Of the methods tested, standard ComBat and DESeq2 were identified as the most robust correction methods for combining spaceflight mouse liver RNAseq datasets hosted on GeneLab.

GeneLab↗

Combining RNA-SEQ Datasets from NASA GENELAB: An Evaluation of Correction Methods

Background: Conducting space biology experiments aboard the International Space Station, particularly those utilizing complex model organisms like mice, is expensive and difficult due to limited crew availability, hardware, and space. As a result, sample numbers from these studies are low, reducing the statistical power of any one experiment. Aggregating spaceflight datasets serves as a method to increase sample numbers, allowing for novel insights through bioinformatic analysis of ‘omics data from merged datasets. However, aggregating datasets can introduce unwanted variation including 1) differences in sample handling, processing, and sequencing platforms between datasets (technical variation) as well as 2) differences in experimental design between datasets. Methods: In the present study, NASA GeneLab-hosted RNAseq datasets from mouse liver tissues were used to evaluate several statistical methods to correct for this unwanted variation through two approaches, reference-based and standard. The following correction algorithms were applied with (reference-based) and/or without (standard) considering Universal Mouse RNA Reference samples: ComBat and ComBat_seq from the SVA package, the median polish, empirical Bayes, and ANOVA-based algorithms from the MBatch package, and negative binomial regression normalization in the DESeq2 package. For each approach, after the correction algorithm was applied, differential gene expression (DGE) analysis of flight and ground control samples was performed with the combined data. The robustness of each tool was evaluated using BatchQC to determine statistical differences between datasets before and after correction, Principal Component Analysis to evaluate global gene expression in samples before and after correction, and by comparing DGE analysis of individual datasets and combined datasets before and after correction. Results: The results showed that the reference-based approach introduced several additional (and likely artificial) differentially expressed genes when compared with the respective standard approach. Conclusions: Of the methods tested, standard ComBat_seq and DESeq2 were identified as the most robust correction methods for combining spaceflight mouse liver RNAseq datasets hosted on GeneLab.

Finsam Samson↗

A review on recent machine learning applications for imaging mass spectrometry studies

Imaging mass spectrometry (IMS) is a powerful analytical technique widely used in biology, chemistry, and materials science fields that continue to expand. IMS provides a qualitative compositional analysis and spatial mapping with high chemical specificity. The spatial mapping information can be 2D or 3D depending on the analysis technique employed. Due to the combination of complex mass spectra coupled with spatial information, large high-dimensional datasets (hyperspectral) are often produced. Therefore, the use of automated computational methods for an exploratory analysis is highly beneficial. The fast-paced development of artificial intelligence (AI) and machine learning (ML) tools has received significant attention in recent years. These tools, in principle, can enable the unification of data collection and analysis into a single pipeline to make sampling and analysis decisions on the go. There are various ML approaches that have been applied to IMS data over the last decade. Here, in this review, we discuss recent examples of the common unsupervised (principal component analysis, non-negative matrix factorization, k-means clustering, uniform manifold approximation and projection), supervised (random forest, logistic regression, XGboost, support vector machine), and other methods applied to various IMS datasets in the past five years. The information from this review will be useful for specialists from both IMS and ML fields since it summarizes current and representative studies of computational ML-based exploratory methods for IMS.

47 OTHER INSTRUMENTATION↗

Genetic variation in nitrogen‐use efficiency and its associated traits in dryland winter wheat ( Triticum aestivum L.) cultivars released from the 1940s to the 2010s in Shaanxi Province, China

Abstract BACKGROUND Improving the nitrogen‐use efficiency (NUE) of wheat can help mitigate the problems of poor soil fertility under dryland conditions. We conducted field experiments using three nitrogen (N) fertilization levels (0, 120, and 180 kg ha −1 ) applied to eight dryland wheat cultivars to assess NUE and its associated traits. RESULTS The grain yield significantly increased with the improvement in variety, mainly as a result of a substantial increase in 1000‐grain weight and harvest index. Modern wheat varieties have stabilized at an optimal plant height and exhibited improved performance in terms of NUE, partial N productivity, N harvest index, and grain protein content compared to older varieties. The NUE of wheat gradually increased with variety replacement. The net photosynthesis rate of the flag leaves in the filling stage improved with the year of cultivar release; Increasing soil–plant analysis development (SPAD) values of flag leaves in the flowering and filling stages were observed over time, with the flag leaves of modern varieties showing a high chlorophyll content in the filling stage. Additionally, the principal component analysis showed that the SPAD value, grain number per unit area, transpiration rate, leaf area, and grain protein content positively contributed to the clustering of the N180 and modern cultivars (from the 2000s to 2010s). CONCLUSION Overall, high levels of N application did not significantly improve the NUE of wheat. However, modern wheat varieties can optimize N distribution, increase flag leaf photosynthetic capacity, and improve photosynthesis ability, thus enhancing NUE to achieve high yields under a suitable level of N supply. © 2022 Society of Chemical Industry.

Lian, Huida↗

Finding Hidden Patterns in High Resolution Wind Flow Model Simulations

Wind flow data is critical in terms of investment decisions and policy making. High resolution data from wind flow model simulations serve as a supplement to the limited resource of original wind flow data collection. Given the large size of data, finding hidden patterns in wind flow model simulations are critical for reducing the dimensionality of the analysis. In this work, we first perform dimension reduction with two autoencoder models: the CNN-based autoencoder (CNN-AE) [1], and hierarchical autoencoder (HIER-AE) [2], and compare their performance with the Principal Component Analysis (PCA). We then investigate the super-resolution of the wind flow data. By training a Generative Adversarial Network (GAN) with 300 epochs, we obtained a trained model with 2× resolution enhancement. We compare the results of GAN with Convolutional Neural Network (CNN), and GAN results show finer structure as expected in the data field images. Also, the kinetic energy spectra comparisons show that GAN outperforms CNN in terms of reproducing the physical properties for high wavenumbers and is critical for analysis where high-wavenumber kinetics play an important role.

97 MATHEMATICS AND COMPUTING↗

Investigating the Impacts of Land Use Change on Urban Heat and Vulnerability in Cali, Columbia

The urban heat island effect (UHI) is an environmental phenomenon where cities experience higher temperatures than rural areas due to increased pavement and decreased cooling from vegetation. Approximately 76% of people in Colombia live in urban areas, and the city of Santiago de Cali is facing UHI challenges exacerbated by land use change. Wetlands and forests formerly surrounded the city but were replaced by development and agriculture. The Colombian municipal government agency Departamento Administrativo de Gestión del Medio Ambiente and the community organization Fundacion Dinamizadores Ambientales partnered with NASA DEVELOP to evaluate communities in Cali most vulnerable to urban heat. This project illustrated the utility of using NASA Earth observations to evaluate the relationship between land use, temperature, and social factors in Cali, Colombia between 2013 and 2023. The team used Landsat 7 Enhanced Thematic Mapper Plus (ETM+), Landsat 8 Operational Land Imager (OLI) and Thermal Infrared Sensor (TIRS), and Landsat 9 OLI-2/TIRS-2 to generate land surface temperature (LST), normalized difference vegetation index (NDVI), and albedo maps in Google Earth Engine. Heavy cloud cover limited the accuracy of the LST but incorporating up to three satellites for a median image reduced potential errors. Through further analysis in ArcGIS Pro, the team classified land use change using a deep learning model and found that LST was significantly higher in urban areas than in wetlands or forests. Using R studio, the team ran a principal component analysis to determine which social factors had the strongest correlation with LST. The team found that health care and green space access were negatively correlated, and Afro-Colombian ethnicity was positively correlated with LST. With awareness of the most impacted and vulnerable regions, the partner organizations can work to prioritize green space establishment in those areas to reduce the impacts of urban heat. Addressing the urban heat island effect will reduce environmental justice concerns within the city and improve overall health, air, and water quality for those who live there.

Brenna Bruffey↗

A New Machine Learning Based Analysis for Improving Satellite Retrieved Atmospheric Composition Data: OMI SO2 as an Example

Despite recent progress, satellite retrievals of anthropogenic SO2 still suffer from relatively low signal-tonoise ratios. In this study, we demonstrate a new machine learning data analysis method to improve the quality of satellite SO2 products. In the absence of large ground-truth datasets for SO2, we start from SO2 slant column densities (SCDs) retrieved from the Ozone Monitoring Instrument (OMI) using a data-driven, physically based algorithm and calculate the ratio between the SCD and the root mean square (rms) of the fitting residuals for each pixel. To build the training data, we select presumably clean pixels with small SCD / rms ratios (SRRs) and set their target SCDs to zero. For polluted pixels with relatively large SRRs, we set the target to the original retrieved SCDs. We then train neural networks (NNs) to reproduce the target SCDs using predictors including SRRs for individual pixels, solar zenith, viewing zenith and phase angles, scene reflectivity, and O3 column amounts, as well as the monthly mean SRRs. For data analysis, we employ two NNs: (1) one trained daily to produce analyzed SO2 SCDs for polluted pixels each day and (2) the other trained once every month to produce analyzed SCDs for less polluted pixels for the entire month. Test results for 2005 show that our method can significantly reduce noise and artifacts over background regions. Over polluted areas, the monthly mean NN-analyzed and original SCDs generally agree to within ±15 %, indicating that our method can retain SO2 signals in the original retrievals except for large volcanic eruptions. This is further confirmed by running both the NN-analyzed and original SCDs through a topdown emission algorithm to estimate the annual SO2 emissions for ∼ 500 anthropogenic sources, with the two datasets yielding similar results. We also explore two alternative approaches to the NN-based analysis method. In one, we employ a simple linear interpolation model to analyze the original SCD retrievals. In the other, we develop a PCA–NN algorithm that uses OMI measured radiances, transformed and dimension-reduced with a principal component analysis (PCA) technique, as inputs to NNs for SO2 SCD retrievals. While the linear model and the PCA–NN algorithm can reduce retrieval noise, they both underestimate SO2 over polluted areas. Overall, the results presented here demonstrate that our new data analysis method can significantly improve the quality of existing OMI SO2 retrievals. The method can potentially be adapted for other sensors and/or species and enhance the value of satellite data in air quality research and applications.

Can Li↗