Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “principal component analysis”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Derivation of Integrated Load Distributions from Resampled Computational Data

Principal component analysis (PCA) has been the center of many surrogate models used to characterize fluid flows in recent years. However, little work has been done to character- ize the uncertainty in the PCA transform itself and its effect on derived surrogate models. To explore the uncertainty, a typical interpolated surrogate model is constructed for a represen- tative aerodynamic body from computational data. The computational data is then resampled to explore the robustness of the PCA transformation. The variations of the model predictions during this resampling are analyzed to get a measure of confidence in the PCA transformation, which is then applied to the interpolated surrogate model to get uncertainty on integrated force predictions. An initial test case has been explored with promising results.

Computational Fluid Dynamics↗

Coronado Ecological Conservation: Assessing Vegetation Change Due to Border Wall Construction and Shifting Social Trails

Species monitoring is essential for mitigating the impacts of plant invasion, such as radical changes in an area’s ecosystem, degraded soil health, increased wildfire severity, landslides, and increased flooding. For this project, NASA DEVELOP partnered with the National Park Service (NPS) to investigate invasive species in disturbed lands: specifically, areas affected by off-trail travel and U.S.-Mexico border construction activities. The team assessed how construction has impacted the distribution of Lehmann’s lovegrass and Russian thistle invasives throughout Coronado National Memorial, AZ from 1986-2022. Using data from Landsat 5 and 8, Sentinel-2, NAIP, and PlanetScope, the team computed NDVI, NDMI, MSAVI2, EVI, and Tasseled Cap Wetness, Brightness, and Greenness transformations as vegetation health indicators to input into various machine learning algorithms. To minimize noise, the team conducted Principal Component Analysis on vegetation indices and spectral bands before running k-means clustering and random forest classification algorithms. Between all datasets, the team found that the median area fully overtaken by invasive plants was 5.37% of the park’s total area in 2022. The NPS will use end products to help increase restoration efforts in disturbed areas with high concentrations of invasive plants, and this project can serve as a jumping off point for future invasive species monitoring. The NPS’s collection of ground data for 2022-2023, in conjunction with future data collection, will notably improve the accuracy of classification models, leading to more precise monitoring of invasive species spread over time.

Coronado National Memorial↗

Moisture and Temperature Influences on Nonlinear Vegetation Trends in Serengeti National Park

While long-term vegetation greening trends have appeared across large land areas over the late 20th century, uncertainty remains in identifying and attributing finer-scale vegetation changes and trends, particularly across protected areas. Serengeti National Park (SNP) is a critical East African protected area, where seasonal vegetation cycles support vast populations of grazing herbivores and a host of ecosystem dynamics. Previous work has shown how non-climate drivers (e.g. land use) shape the SNP ecosystem, but it is still unclear to what extent changing climate conditions influence SNP vegetation, particularly at finer spatial and temporal scales. We fill this research gap by evaluating long-term (1982–2016) changes in SNP leaf area index (LAI) in relation to both temperature and moisture availability using Ensemble Empirical Mode Decomposition and Principal Component Analysis with regression techniques. We find that SNP LAI trends are nonlinear, display high sub-seasonal variation, and are influenced by lagged changes in both moisture and temperature variables and their interactions. LAI during the long rains (e.g. March) exhibits a greening-to-browning trend reversal starting in the early 2000s, partly due to antecedent precipitation declines. In contrast, LAI during the short rains (e.g. November, December) displays browning-to-greening alongside increasing moisture availability. Rising temperature trends also have important, secondary interactions with moisture variables to shape these SNP vegetation trends. Our findings show complex vegetation-climate interactions occurring at important temporal and spatial scales of the SNP, and our rigorous statistical approaches detect these complex climate-vegetation trends and interactions, while guarding against spurious vegetation signals.

Moisture↗

Development of a Data Fusion Methodology for Lineload Aerodynamic Databases for a Launch Vehicle during Liftoff and Transition

The need for databases for the distributed loading on launch vehicles during the early portion of flight necessitates the use of expensive computational flows in regimes where wake effects dominate. While also being expensive, this is a regime that computational tools tend to historically have problems simulating accurately. To help tackle this problem, a method of data fusion to combine computational results to wind tunnel derived force and moment data is developed. Using this method, significant reduction in computational costs and increases in confidence of the final product is possible and has been used to generate several databases for the Space Launch System (SLS) at NASA. While the full details of database generation are not part of this work, the crucial method at its core is developed here. Two SLS geometries are used throughout the work to demonstrate the techniques. These are two of the larger geometries and represent both planned crewed missions to the Moon as well as potential cargo missions to deep space. The method uses principal component analysis (PCA) to generate a reduced ordered model (ROM) to help fill in the full parameter space. Other similar techniques are explored, but were not found to have a significant result on the predictions of the ROM. Because the full number of components are kept to generate the model, this lack of difference is expected. This method is then extended to ensure that predicted surfaces match trusted force and moment data derived from wind tunnel testing. This extension is done by setting up a constrained optimization problem in order to minimize the deviation from the surface resolved computational data while still integrating to the desired values. When generating the constrained optimization problem, a weighting factor to balance these competing needs is introduced. The work compares previously introduced weighting terms from similar work to the proposed terms and shows that the previously used terms do not have as desirable behavior in this flow regime. This method is then expanded by developing a technique to incorporate uncertainty quantification into the developed data fusion methodology. This expansion takes a two pronged approach. One examines transferring the uncertainties in the force and moment database and characterizes how those adjustments change the predicted lineloads. The second looks at model form error and looks how rebuilding the model using slightly different data changes the predictions. These two terms are then combined in order to create an uncertainty model that takes both effects into account. The limitations of the proposed methods is then discussed as well as possible techniques to address these shortcomings.

Launch Vehicles↗

Transcriptomics-based Machine Learning Analysis Predicts Space-Exposed Murine Livers

Limited sample sizes, high data dimensionality, and sensitivity to technical and biological variability of next generation sequencing (NGS), typically limits machine learning (ML) approaches in spaceflight studies that include radiation effects. However, pooling smaller studies while addressing intra- and inter-study variabilities allows for ML predictive modeling. Here, integration methods were applied to whole transcriptome shotgun sequencing (RNA-seq) data from six mouse liver GeneLab datasets (GLDS) (n ranging from 6 to 39 samples) from with a total of 81 spaceflight and ground-control samples to determine top features (i.e. genes) relevant to spaceflight including the effect of radiation exposure. RNASeq counts were normalized for each study, then merged and scaled across all datasets. Data dimensionality was reduced using a minimum redundancy maximum relevance (MRMR) methodology. Redundancy and relevance were computed using the Pearson correlation and F-statistic, respectively. The top 100 MRMR features were used to predict spaceflight vs. ground-control samples using Random Forest (RF), Support Vector Machine (SVM), and Linear Discriminant Analysis (LDA) classifiers with 5-fold cross validation (CV). Principal component analysis (PCA) on the complete feature set versus the MRMR features shows separation between spaceflight samples and ground controls (Figure 1A). The ML-based gene sets were compared against differential gene expression results obtained with DESeq2 from individual GLDS. Using all features or randomly sampled subsets at matching set sizes with MRMR, a maximum classifier accuracy of 69% was shown on the test set over 5 folds. For all classifiers, CV training using at least the top 30 MRMR genes show minimum 89% accuracy and 0.95 AUC value on the test set over 5 folds (Figure 1B). Baseline set analysis on differentially expressed genes (DEGs) identified using padj ≤ 0.05 show 295 DEGs that overlap at least two studies and 13 DEGs that overlap three studies (Figure 1C). Set analysis between the top 100 MRMR features and the DEGs showed 47 genes that overlap at least one study and 24 genes that overlap two studies. Over-representation analysis showed overlapping biological processes related to fatty acid and lipid metabolism which may indicate these processes in the response to spaceflight stressors. MRMR feature selection for the selected ML methods improve performance relative to a classifier built on all features or randomly sampled subsets. Permutation feature importance within the decorrelated MRMR features showed concordance in feature ranking between ML methods. A challenge of applying ML methods across heterogeneous NGS data is accounting for signal:noise. Here, signal validation across studies was shown by intersecting sets between top MRMR genes and DEGs from DESeq2 analysis. Non-intersecting sets introduce opportunity to explore genes relevant to differentiating space flight exposed groups and implementing ML methods across existing NGS datasets may overcome sample size limitations.

Machine Learning↗

Transcriptomics-based Machine Learning (ML) Analysis Predicts Space-Exposed Murine Livers

Limited sample sizes, high data dimensionality, and sensitivity to technical and biological variability of next generation sequencing (NGS), typically limits machine learning (ML) approaches in spaceflight studies that include radiation effects. However, pooling smaller studies while addressing intra- and inter-study variabilities allows for ML predictive modeling. Here, integration methods were applied to whole transcriptome shotgun sequencing (RNA-seq) data from six mouse liver GeneLab datasets (GLDS) (n ranging from 6 to 39 samples) from with a total of 81 spaceflight and ground-control samples to determine top features (i.e. genes) relevant to spaceflight including the effect of radiation exposure. RNASeq counts were normalized for each study, then merged and scaled across all datasets. Data dimensionality was reduced using a minimum redundancy maximum relevance (MRMR) methodology. Redundancy and relevance were computed using the Pearson correlation and F-statistic, respectively. The top 100 MRMR features were used to predict spaceflight vs. ground-control samples using Random Forest (RF), Support Vector Machine (SVM), and Linear Discriminant Analysis (LDA) classifiers with 5-fold cross validation (CV). Principal component analysis (PCA) on the complete feature set versus the MRMR features shows separation between spaceflight samples and ground controls (Figure 1A). The ML-based gene sets were compared against differential gene expression results obtained with DESeq2 from individual GLDS. Using all features or randomly sampled subsets at matching set sizes with MRMR, a maximum classifier accuracy of 69% on the test set over 5 folds. For all classifiers, CV training using at least the top 30 MRMR genes show minimum 89% accuracy and 0.95 AUC value on the test set over 5 folds (Figure 1B). Baseline set analysis on differentially expressed genes (DEGs) identified using padj ≤ 0.05 show 295 DEGs that overlap at least two studies and 13 DEGs that overlap three studies (Figure 1C). Set analysis between the top 100 MRMR features and the DEGs showed 47 genes that overlap at least one study and 24 genes that overlap two studies. Over-representation analysis showed overlapping biological processes related to fatty acid and lipid metabolism which may indicate these processes in the response to spaceflight stressors. MRMR feature selection for the selected ML methods improve performance relative to a classifier built on all features or randomly sampled subsets. Permutation feature importance within the decorrelated MRMR features showed concordance in feature ranking between ML methods. A challenge of applying ML methods across heterogeneous NGS data is accounting for signal:noise. Here, signal validation across studies was shown by intersecting sets between top MRMR genes and DEGs from DESeq2 analysis. Non-intersecting sets introduce opportunity to explore genes relevant to differentiating space flight exposed groups and implementing ML methods across existing NGS datasets may overcome sample size limitations.

Machine Learning↗

Transcriptomics-based Machine Learning Analysis Predicts Space-Exposed Murine Livers

Limited sample sizes, high data dimensionality, and sensitivity to technical and biological variability of next generation sequencing (NGS), typically limits machine learning (ML) approaches in spaceflight studies that include radiation effects. However, pooling smaller studies while addressing intra- and inter-study variabilities allows for ML predictive modeling. Here, integration methods were applied to whole transcriptome shotgun sequencing (RNA-seq) data from six mouse liver GeneLab datasets (GLDS) (n ranging from 6 to 39 samples) from with a total of 81 spaceflight and ground-control samples to determine top features (i.e. genes) relevant to spaceflight including the effect of radiation exposure. RNASeq counts were normalized for each study, then merged and scaled across all datasets. Data dimensionality was reduced using a minimum redundancy maximum relevance (MRMR) methodology. Redundancy and relevance were computed using the Pearson correlation and F-statistic, respectively. The top 100 MRMR features were used to predict spaceflight vs. ground-control samples using Random Forest (RF), Support Vector Machine (SVM), and Linear Discriminant Analysis (LDA) classifiers with 5-fold cross validation (CV). Principal component analysis (PCA) on the complete feature set versus the MRMR features shows separation between spaceflight samples and ground controls (Figure 1A). The ML-based gene sets were compared against differential gene expression results obtained with DESeq2 from individual GLDS. Using all features or randomly sampled subsets at matching set sizes with MRMR, a maximum classifier accuracy of 69% was shown on the test set over 5 folds. For all classifiers, CV training using at least the top 30 MRMR genes show minimum 89% accuracy and 0.95 AUC value on the test set over 5 folds (Figure 1B). Baseline set analysis on differentially expressed genes (DEGs) identified using padj ≤ 0.05 show 295 DEGs that overlap at least two studies and 13 DEGs that overlap three studies (Figure 1C). Set analysis between the top 100 MRMR features and the DEGs showed 47 genes that overlap at least one study and 24 genes that overlap two studies. Over-representation analysis showed overlapping biological processes related to fatty acid and lipid metabolism which may indicate these processes in the response to spaceflight stressors. MRMR feature selection for the selected ML methods improve performance relative to a classifier built on all features or randomly sampled subsets. Permutation feature importance within the decorrelated MRMR features showed concordance in feature ranking between ML methods. A challenge of applying ML methods across heterogeneous NGS data is accounting for signal:noise. Here, signal validation across studies was shown by intersecting sets between top MRMR genes and DEGs from DESeq2 analysis. Non-intersecting sets introduce opportunity to explore genes relevant to differentiating space flight exposed groups and implementing ML methods across existing NGS datasets may overcome sample size limitations.

Machine Learning↗

Next Generation Aura-OMI SO2 Retrieval Algorithm: Introduction and Implementation Status

We introduce our next generation algorithm to retrieve SO2 using radiance measurements from the Aura Ozone Monitoring Instrument (OMI). We employ a principal component analysis technique to analyze OMI radiance spectral in 310.5-340 nm acquired over regions with no significant SO2. The resulting principal components (PCs) capture radiance variability caused by both physical processes (e.g., Rayleigh and Raman scattering, and ozone absorption) and measurement artifacts, enabling us to account for these various interferences in SO2 retrievals. By fitting these PCs along with SO2 Jacobians calculated with a radiative transfer model to OMI-measured radiance spectra, we directly estimate SO2 vertical column density in one step. As compared with the previous generation operational OMSO2 PBL (Planetary Boundary Layer) SO2 product, our new algorithm greatly reduces unphysical biases and decreases the noise by a factor of two, providing greater sensitivity to anthropogenic emissions. The new algorithm is fast, eliminates the need for instrument-specific radiance correction schemes, and can be easily adapted to other sensors. These attributes make it a promising technique for producing long-term, consistent SO2 records for air quality and climate research. We have operationally implemented this new algorithm on OMI SIPS for producing the new generation standard OMI SO2 products.

Sulfur dioxide↗

Principle Component Analysis of the Evolution of the Saharan Air Layer and Dust Transport: Comparisons between a Model Simulation and MODIS Retrievals

The onset and evolution of Saharan Air Layer (SAL) episodes during June-September 2002 are diagnosed by applying principal component analysis to the NCEP reanalysis temperature anomalies at 850 hPa, where the largest SAL-induced temperature anomalies are located. The first principal component (PC) represents the onset of SAL episodes, which are associated with large warm anomalies located at the west coast of Africa. The second PC represents two opposite phases of the evolution of the SAL. The positive phase of the second PC corresponds to the southwestward extension of the warm anomalies into the tropical-subtropical North Atlantic Ocean, and the negative phase corresponds to the northwestward extension into the subtropical to mid-latitude North Atlantic Ocean and the southwest Europe. A dust transport model (CARMA) and the MODIS retrievals are used to study the associated effects on dust distribution and deposition. The positive (negative) phase of the second PC corresponds to a strengthening (weakening) of the offshore flows in the lower troposphere around 10deg - 20degN, causing more (less) dust being transported along the tropical to subtropical North Atlantic Ocean. The variation of the offshore flow indicates that the subseasonal variation of African Easterly Jet is associated with the evolution of the SAL. Significant correlation is found between the second PC time series and the daily West African monsoon index, implying a dynamical linkage between West African monsoon and the evolution of the SAL and Saharan dust transport.

Wong, S.↗

Baseline Observations of Hemispheric Sea Ice with the Nimbus 7 Scanning Multichannel Microwave Radiometer

The Scanning Multichannel Microwave Radiometer (SMMR) on board the NASA Nimbus 7 satellite was designed to obtain data for sea surface temperatures (SSTs), near-surface wind speeds, sea ice coverage and type, rainfall rates over the oceans, cloud water content, snow water equivalent, and soil moisture. In this paper, I shall emphasize the sea ice observations and mention briefly some important SST observations. A prime factor contributing to the importance of SMMR sea ice observations lies in their successful integration into a long-term time series, presently being extended by observations from the series of Special Sensor Microwave/Imager (SSMI) on board the DOD/DMSP F8, Fl1, and F12 satellites. This currently constitutes a 19-year data set. Almost half of this was provided by the SMMR. Unfortunately, the 4-year data set produced earlier by the single-channel Electrically Scanned Microwave Radiometer (ESMR) was not successfully integrated into the SMMR/SSMI data set. This resulted primarily from the lack of an overlap period to provide intersensor adjustment, but also because of the large difference between the algorithms to produce ice concentrations and large temporal gaps in the ESMR data. The lack of overlap between the SeaSat and Nimbus 7 SMMR data sets was an important consideration for also excluding the SeatSat one, but the spatial gaps especially in the Southern Hemisphere daily SeaSat observations was another. The sea ice observations will continue into the future by means of the Advanced Microwave Scanning Radiometer (AMSR) on board the ADEOS II and EOS satellites due to be launched in mid- and late-2000, respectively. Analysis of the sea ice data has been carried out by a number of different techniques. Long-term trends have been examined by means of ordinary least squares and band-limited regression. Oscillations in the data have been examined by band-limited Fourier analysis. Here, I shall present results from a novel combination of Principal Component analysis and the recently-developed Empirical Mode Decomposition (EMD). In this method, the data are first separated into spatial and temporal parts, and then the temporal parts of the first few PCs are broken into intrinsic modes by the EMD method.

Gloersen, Per↗

Kernel PLS-SVC for Linear and Nonlinear Discrimination

A new methodology for discrimination is proposed. This is based on kernel orthonormalized partial least squares (PLS) dimensionality reduction of the original data space followed by support vector machines for classification. Close connection of orthonormalized PLS and Fisher's approach to linear discrimination or equivalently with canonical correlation analysis is described. This gives preference to use orthonormalized PLS over principal component analysis. Good behavior of the proposed method is demonstrated on 13 different benchmark data sets and on the real world problem of the classification finger movement periods versus non-movement periods based on electroencephalogram.

Rosipal, Roman↗

Bridgeport Urban Development: Leveraging NASA Earth Observations and Sociodemographic Data to Assess Urban Heat Vulnerability and Inform Cool Corridors in Bridgeport, Connecticut

Urban environments face hotter temperatures than suburban and rural areas due to higher concentrations of impervious surfaces, heat-retaining buildings, and lack of green space. Bridgeport, Connecticut, which was formerly a national manufacturing hub, is now the densest and most populous city in the state. Bridgeport experiences hotter temperatures, exposing its residents to more extreme temperatures than the surrounding affluent suburbs. Extreme heat affects the health of those exposed to it and intensifies energy demands. Understanding temperature differences is the first step in effectively directing mitigation efforts. Our partner, Groundwork Bridgeport, along with the Yale Urban Design Workshop, are planning a “cool corridors” project, implementing cooling infrastructure to combat urban heat. We used Landsat 8 Thermal Infrared Sensor and Landsat 9 Thermal Infrared Sensor-2 data to conduct a Land Surface Temperature analysis in Google Earth Engine for the county of Fairfield. A Principal Component Analysis was performed to identify indicators of social vulnerability in Bridgeport. We used the SOlar and LongWave Environmental Irradiance Geometry model to identify felt heat on the block level to inform where the partner should locate their cooling interventions to ensure they are most effective and equitable. We focused on the East Side of Bridgeport, which we found was 10 degrees hotter than other areas of Bridgeport and the neighboring town of Fairfield. We integrated our findings using Earth observations and additional sociodemographic and climate data into final communication products for our partners which will facilitate their selection of candidate locations for their Cool Corridors project.

Silas Kirsch↗

The luminosity structure and objective classification of galaxies

The luminosity structure of spiral galaxies is studied using the technique of principal component analysis. It is found that approximately 94% of the variation in the luminosity distribution of galaxies can be accounted for by just two principal components. The principal luminosity components may contain valuable information about star formation history or whatever luminosity-regulating process occurs in galaxies. Practically, these principal components provide a new approach for the investigation of the luminosity structures of galaxies and their dependence on other properties. They also serve as an excellent objective classification system for galaxies. We introduce in this paper such a classification scheme and explore its various properties. The new system shows a number of very impressive characteristics. Most important, it can well segregate virtually all the important galactic properties we tested and does so much better than the conventional morphological classification systems. Of particular interest is that some distance-dependent parameters can also be determined to a surprisingly good accuracy; for example, absolute magnitude may be determined to an accuracy of approximately 0.6 mag (yet further improvement is believed to be highly possible). Second, the system is objective, and the classification procedure can be automated to a large degree; also the new system can apply to much smaller and fainter images than do eye-based clasification systems. These properties make the new system suitable for practical application, especially on very large (and deeper) digital image catalogs. Third, the classification is expressed in dimensionless numbers, yet the simple notation bears significant and easily understandable meaning, making it easy and convenient to use. Finally, the new system has another extremely useful feature: it provides a very powerful and convenient platform not only for classification, but also for easily recording, examining, and studying the variations and correlations of galaxy properties-all these may be carried out graphically by using the C-vectors and the C-diagrams introduced in the paper. We wil also give an example to demonstrate the use of the classification system for the study of the internal extinction problem in spiral galaxies.

Han, Mingshen↗

The intermediate line region of QSOs

Our recent statistical investigations of broad UV lines of luminous quasars (QSOs) suggest that the traditional broadline region (BLR) consists of two components-one of width approximately 2000 km/s FWHM with the peak within a few hundred kilometers per second of the systemic redshift, and another very broad component of width greater than or equal to 7000 km/s FWHM and blueshifted by greater than or equal to 1000 km/s. Differences in the relative strengths of these components account for much of the diversity of broad-line profiles, as well as relations among line strength, line width, asymmetry, and peak blueshift. We have suggested that the narrower component arises in a distinct intermediate-line region (ILR) that is an inner extension of the narrow-line region (NLR). We form spectra of the ILR and very broad line region (VBLR) in two complementary ways. First, using a small sample of high-quality spectra, we difference two composite spectra, one with FWHM (sub C IV) approximately 3000 km/s, the other with FWHM (sub C IV) approximately 7000 km/s (essentially a VBLR spectrum)-revealing a narrower line spectrum with FWHM approximately 2000 km/s. Second, we use a principal component analysis of 200 low signal-to-noise ratio spectra from the Large Bright Quasar Survey. The ILR is identified with the first principal component-that component accounting for most of the spectrum-to-spectrum variation. The VBLR spectrum is derived by subtracting this ILR from the mean spectrum. The two methods yield similar results, and the spectra of the ILR and VBLR are very different. Additional support for the existence of two component is the lack of a correlation between the equivalent widths of the ILR and VBLR. We discuss the relationships between the VBLR, the ILR, and the traditional NLR. The ILR line intensity ratios are distinctly different from those of the VBLR, with (relative to C IV lambda 1549) stronger Lyman alpha and weaker N v lambda 1240, lambda 1400 feature, He II lambda 1640, O III lambda 1663, and Al III lambda 1860. Comparison with other active galactic nuclei (AGN) emission-line regions shows that the ILR spectrum tends to be intermediate between that of the VBLR and that of gas more distant from the ionizing continuum, such as the NLR and extended Lyman alpha nebulosity. Comparison of the spectra with results from photoionization models suggests that, compared with the VBLR, the ILR is about 10 times more distant from the ionizing continuum (approximately 1 pc), 100 to 1000 times less dense (10(exp 10)/cu cm), and has a smaller covering factor (less than or approximately 3%, compared with approximately 24% for the VBLR).

Brotherton, M. S.↗

Principal Components as a Data Reduction and Noise Reduction Technique

The potential of principal components as a pipeline data reduction technique for thematic mapper data was assessed and principal components analysis and its transformation as a noise reduction technique was examined. Two primary factors were considered: (1) how might data reduction and noise reduction using the principal components transformation affect the extraction of accurate spectral classifications; and (2) what are the real savings in terms of computer processing and storage costs of using reduced data over the full 7-band TM complement. An area in central Pennsylvania was chosen for a study area. The image data for the project were collected using the Earth Resources Laboratory's thematic mapper simulator (TMS) instrument.

Imhoff, M. L.↗

Global variability of precipitation according to the Tropical Rainfall Measuring Mission

Numerous studies have documented the effect of El Nino-Southern Oscillation (ENSO) on rainfall in many regions of the globe. The question of whether ENSO is the single most important factor in interannual rainfall variability has received less attention, mostly because the kind of data that would be required to make such an assessment were simply not available. Until 1979 the evidence linking El Nino with changes in rainfall around the world came from rain gauges measuring precipitation over land masses and a handful of islands. From 1980 until the launch of the Tropical Rainfall Measuring Mission (TRMM) in November 1997 the remote sensing evidence was confined to ocean rainfall because of the very poor sensitivity of the instruments over land. In this paper we summarize the results of a principal component analysis of TRMM's 60-month (January 1998 to December 2002) global land and ocean remote-sensing record of monthly rainfall accumulations. Contrary to the first principal component of the rainfall itself, the first three indices of the anomaly are most sensitive to precipitation over the ocean rather than over the land. With the help of archived surface station data the first TRMM rain anomaly index is extended back several decades. Comparison of the extended index with the Southern Oscillation Index confirms that the first principal component of the rainfall anomaly is strongly correlated with the ENSO indices.

rainfall↗

Representation of Probability Density Functions from Orbit Determination using the Particle Filter

Statistical orbit determination enables us to obtain estimates of the state and the statistical information of its region of uncertainty. In order to obtain an accurate representation of the probability density function (PDF) that incorporates higher order statistical information, we propose the use of nonlinear estimation methods such as the Particle Filter. The Particle Filter (PF) is capable of providing a PDF representation of the state estimates whose accuracy is dependent on the number of particles or samples used. For this method to be applicable to real case scenarios, we need a way of accurately representing the PDF in a compressed manner with little information loss. Hence we propose using the Independent Component Analysis (ICA) as a non-Gaussian dimensional reduction method that is capable of maintaining higher order statistical information obtained using the PF. Methods such as the Principal Component Analysis (PCA) are based on utilizing up to second order statistics, hence will not suffice in maintaining maximum information content. Both the PCA and the ICA are applied to two scenarios that involve a highly eccentric orbit with a lower apriori uncertainty covariance and a less eccentric orbit with a higher a priori uncertainty covariance, to illustrate the capability of the ICA in relation to the PCA.

Mashiku, Alinda K.↗

Using Machine Learning for Timely Estimates of Ocean Color Information From Hyperspectral Satellite Measurements in the Presence of Clouds, Aerosols, and Sunglint

Retrievals of ocean color from space are important for better understanding of the ocean ecosystem but can be limited under conditions such as clouds, aerosols, and sunglint. Many ocean color algorithms use a few selected spectral bands to perform an atmospheric correction and then derive the upwelling radiance from the ocean. The limitations in the atmospheric correction under certain conditions lead to many gaps in daily spatial coverage of ocean color retrievals. To address these limitations, we introduce a new approach that uses machine learning to estimate ocean color from top of atmosphere radiances or reflectance measurements. In this approach, a principal component analysis is used to decompose the hyperspectral measurements into spectral features that describe the scattering and absorption of the atmosphere and the underlying surface. The coefficients of the principal components are then used to train a neural network to predict ocean color properties derived from the MODIS atmospheric correction algorithm. This machine learning approach is independent of a priori information and does not rely on any radiative transfer modeling. We apply the approach to two hyperspectral UV/VIS instruments, the ozone monitoring instrument (OMI) and the TROPOspheric Monitoring Instrument (TROPOMI), using measurements from 320–500 nm to show that it can be used to reproduce ocean color properties in less-than-ideal conditions. This machine learning approach complements the current atmospheric correction ocean color retrievals by filling in the gaps resulting from cloud, aerosol, and sunglint contamination. This method can be applied to the future hyperspectral Ocean Color Instrument (OCI), which will be onboard NASA’s Plankton, Aerosol Cloud, ocean Ecosystem (PACE) ocean color satellite set to launch in 2024.

Ocean color↗