SEARCH · Engineering Papers
Results for “Principal Component Analysis (PCA)”
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Using Principal Component Analysis (PCA) to Speed up Radiative Transfer (RT) Computations
Multiple scattering RT calculations time-consuming. Need a speed improvement of about 1000 (for OCO)! Solution: Make use of redundancies in spectra. Correlated-k (Lacis and Wang, Lacis and Oinas, Goody et al, Fu and Liou) Problem: Assume that spectral variation of atmospheric optical properties spatially correlated at all points along optical path. High accuracy (HI) and 2-stream (2S) calculations have high correlation. Single scattering (SS) computations highly scenario-dependent, but not time consuming. Perform SS and 2S calculations at every wavelength. Perform small number of HI computations. Need to compute correction factor B at every wavelength.
Discerning the Impact of Powder Feedstock Variability on Structure, Property, and Performance of Selective Laser Melted Alloy 718: A Principal Component Analysis (PCA) of Feedstock Variability
Extensive mechanical, chemical and microstructural analyses were conducted on additively manufactured Alloy 718 to characterize powders from multiple vendors to determine the effects of variations observed in the powders had on the consolidated material. With over 190 variables examined, it was necessary to reduce the number of variables and identify the variables and classes of variables that had the greatest effect. Principle Component Analysis (PCA) was used to reduce the number of variable to effectively 12 while identifying several classes of variables as most important.
Variational encoder geostatistical analysis (VEGAS) with an application to large scale riverine bathymetry
Estimation of riverbed profiles, also known as bathymetry, plays a vital role in many applications, such as safe and efficient inland navigation, prediction of bank erosion, land subsidence, and flood risk management. The high cost and complex logistics of direct bathymetry surveys, i.e, depth imaging, have encouraged the use of indirect measurements such as surface flow velocities. However, estimating high-resolution bathymetry from indirect measurements is an inverse problem that can be computationally challenging. Here, we propose a reduced-order model (ROM) based approach that utilizes a variational autoencoder (VAE), a type of deep neural network with a narrow layer in the middle, to compress bathymetry and flow velocity information and accelerate bathymetry inverse problems from flow velocity measurements. In our application, the shallow-water equations (SWE) with appropriate boundary conditions (BCs), e.g., the discharge and/or the free surface elevation, constitute the forward problem, to predict flow velocity. Then, ROMs of the SWEs are constructed on a nonlinear manifold of low dimensionality through a variational encoder and the bathymetry inversion problem is derived on the low-dimensional latent space in a Hierarchical Bayesian setting. Further, the reformulation allows variational inference with a small number (e.g., $\mathscr{O}$ (100) of ROM runs and efficient uncertainty quantification. We have tested our inversion approach on a one-mile reach of the Savannah River, GA, USA. Once the neural network is trained (offline stage), the proposed technique can perform the inversion operation orders of magnitude faster than traditional inversion methods that are commonly based on linear projections, such as principal component analysis (PCA), or the principal component geostatistical approach (PCGA). Furthermore, tests show that the algorithm can estimate the bathymetry with good accuracy even with sparse flow velocity measurements.
Data structure characterization of miltispectral data using principal component and principal factor analysis
Both principal component analysis (PCA) and principal factor analysis (PFA) were used to analyze an experimental multispectral data structure in terms of common and unique variance. Only the common variance of the multispectral data was associated with the principal factor, while higher-order principal components were associated with both common and unique variance. The unique variance was found to represent small spectral variations within each cover type as well as noise vectors, and was most abundant in the lower-order principal components. The lower-order principal components can be useful in research designed to discriminate minor physical variations within features, and to highlight localized change when using multitemporal-multispectral data. Conversely, PFA of the multispectral data provided an insight into a great potential for discriminating basic land-cover types by excluding the unique variance which was related to the noise and minor spectral variations.
3-D Geologic Controls of Hydrothermal Fluid Flow at Brady Geothermal Field, Nevada using PCA
In many hydrothermal systems, fracture permeability along faults provides pathways for groundwater to transport heat from depth. Faulting generates a range of deformation styles that cross-cut heterogeneous geology, resulting in complex patterns of permeability, porosity, and hydraulic conductivity. Vertical connectivity (a through going network of permeable areas that allows advection of heat from depth to the shallow subsurface) is rare and is confined to relatively small volumes that have highly variable spatial distribution. This local compartmentalization of connectivity represents a significant challenge to understanding hydrothermal circulation and for exploring, developing, and managing hydrothermal resources. Here, we present an evaluation of the geologic characteristics that control this compartmentalization in hydrothermal systems through 3-D analysis of the Brady geothermal field in western Nevada. A published 3-D geologic map of the Brady area is used as a basis to develop structural and geological variables that are hypothesized to control or effect permeability or connectivity. The 3-D distribution of these variables is compared to the distribution of productive and non-productive fluid flow intervals along production wells and non-productive wells via principal component analysis (PCA). This comparison elucidates which geologic and structural variables are most closely associated with productive fluid flow intervals. Results indicate that production intervals at Brady are located: (1) within or near to known and stress-loaded macro-scale faults, and (2) in areas of high fault and fracture density. This submission includes the published journal article detailing this work, the published 3-D geologic map of the Brady Geothermal Area used as a basis to develop structural and geological variables that are hypothesized to control or effect permeability or connectivity, 3-D well data, along which geologic data were sampled for PCA analyses, and associated metadata file. This work was done using existing R programs.
SO(3)-invariant PCA with application to molecular data
Principal component analysis (PCA) is a fundamental technique for dimensionality reduction and denoising; however, its application to three-dimensional data with arbitrary orientations -- common in structural biology -- presents significant challenges. A naive approach requires augmenting the dataset with many rotated copies of each sample, incurring prohibitive computational costs. In this paper, we extend PCA to 3D volumetric datasets with unknown orientations by developing an efficient and principled framework for SO(3)-invariant PCA that implicitly accounts for all rotations without explicit data augmentation. By exploiting underlying algebraic structure, we demonstrate that the computation involves only the square root of the total number of covariance entries, resulting in a substantial reduction in complexity. We validate the method on real-world molecular datasets, demonstrating its effectiveness and opening up new possibilities for large-scale, high-dimensional reconstruction problems.
Continuum Fitting HST QSO Spectra
The Principal Component Analysis (PCA) method which we are using to fit and describe QSO spectra relies upon the fact that QSO continuum are generally very smooth and simple except for emission and absorption lines. To see this we need high signal-to-noise (S/N) spectra of QSOs at low redshift which have relatively few absorption lines in the Lyman-a forest. We need a large number of such spectra to use as the basis set for the PCA analysis which will find the set of principal component spectra which describe the QSO family as a whole. We have found that too few HST spectra have the required S/N and hence we need to supplement them with ground based spectra of QSOs at higher redshift. We have many such spectra and we have been working to make them suitable for this analysis. We have concentrated on this topic since 12/15/01.
What drives the variance of galaxy spectra?
We present a study aimed at understanding the physical phenomena underlying the formation and evolution of galaxies following a data-driven analysis of spectroscopic data based on the variance in a carefully selected sample. We apply principal component analysis (PCA) independently to three subsets of continuum-subtracted optical spectra, segregated into their nebular emission activity as quiescent, star-forming, and active galactic nuclei (AGNs). We emphasize that the variance of the input data in this work only relates to the absorption lines in the photospheres of the stellar populations. The sample is taken from the Sloan Digital Sky Survey (SDSS) in the stellar velocity dispersion range 100–150 km s −1 , to minimize the ‘blurring’ effect of the stellar motion. We restrict the analysis to the first three principal components (PCs) and find that PCA segregates the three types with the highest variance mapping SSP-equivalent age, along with an inextricable degeneracy with metallicity, even when all three PCs are included. Spectral fitting shows that stellar age dominates PC1, whereas PC2 and PC3 have a mixed dependence of age and metallicity. The trends support – independently of any model fitting – the hypothesis of an evolutionary sequence from star formation to AGN to quiescence. As a further test of the consistency of the analysis, we apply the same methodology in different spectral windows, finding similar trends, but the variance is maximal in the blue wavelength range, roughly around the 4000 Å break.
Short-term nodal load forecasting based on machine learning techniques
This paper introduces an advanced Short-term Nodal Load Forecasting (STNLF) method that forecasts nodal load profiles for the next day in power systems, based on the combined use of three machine learning techniques. Least Absolute Shrinkage and Selection Operator (LASSO) is employed to reduce the number of features for a single nodal load forecasting. Principal Component Analysis (PCA) is used to capture the features of historical loads in low-dimensional space compared to the original high-dimensional load space where features are barely possible to depict. Additionally, Bayesian Ridge Regression (BRR) is utilized to decide the parameters of the prediction model from a statistics perspective. Tests based on modified PJM load data demonstrate the effectiveness of the proposed STNLF method compared to the state-of-the-art General Regression Neural Network (GRNN) method. Moreover, the reliability of the day-ahead Unit Commitment (UC) solution is shown to have been improved, based on the forecasted load data using the proposed STNLF method.
Mapping Rare Earths and Toxics in E-Waste via Hyperspectral Imaging and Machine Learning
Electronic waste (e-waste) presents a mounting challenge to environmental sustainability due to its complex composition, which includes high-value rare earth elements, hazardous organic compounds, and non-recyclable plastics. Accurate and scalable material classification is essential for enabling efficient resource recovery and safe recycling practices. This study introduces a confidence-aware classification pipeline that combines mid-infrared hyperspectral imaging (HSI), spectral angle mapping (SAM), and iterative machine learning to perform pixel-level material identification across e-waste devices. A curated spectral library encompassing artificial materials (e.g., plastic iron oxide, galvanized metals), minerals (e.g., allanite, hematite), and organic compounds (e.g., benzanthracene, toluene) was used to generate pseudo-labels, each assigned a confidence score based on SAM-derived spectral similarity. High-confidence samples from seven consumer electronics—digital cameras, keyboards, laptop fans, modems, motherboards, TV remotes, and speakers—were iteratively expanded and classified using models such as Support Vector Machine (SVM), Random Forest, Gradient Boosting Classifier, Partial Least Squares Discriminant Analysis (PLSDA) and Logistic Regression. The best-performing classifiers achieved macro F1 scores approaching 1.0. Results revealed widespread plastic content (dominated by plastic iron oxide), the presence of rare earth-bearing minerals like cerium-containing allanite, and pervasive detection of hazardous organics such as benzanthracene. Principal Component Analysis (PCA) visualizations and confusion matrices confirmed high separability and robust classification performance. This methodology enables precise, non-destructive, and scalable classification of heterogeneous e-waste streams. It supports automated, hazard-aware sorting in recycling workflows, facilitating selective recovery of critical materials and compliance with circular economy goals. The confidence-aware framework provides a foundation for real-time deployment in industrial settings, offering significant implications for smart e-recycling infrastructure and policy-driven material stewardship.
Structuring Nutrient Yields throughout Mississippi/Atchafalaya River Basin Using Machine Learning Approaches
To minimize the eutrophication pressure along the Gulf of Mexico or reduce the size of the hypoxic zone in the Gulf of Mexico, it is important to understand the underlying temporal and spatial variations and correlations in excess nutrient loads, which are strongly associated with the formation of hypoxia. This study’s objective was to reveal and visualize structures in high-dimensional datasets of nutrient yield distributions throughout the Mississippi/Atchafalaya River Basin (MARB). For this purpose, the annual mean nutrient concentrations were collected from thirty-three US Geological Survey (USGS) water stations scattered in the upper and lower MARB from 1996 to 2020. Eight surface water quality indicators were selected to make comparisons among water stations along the MARB over the past two decades. Principal component analysis (PCA) was used to comprehensively evaluate the nutrient yields across thirty-three USGS monitoring stations and identify the major contributing nutrient loads. The results showed that all samples could be analyzed using two main components, which accounted for 81.6% of the total variance. The PCA results showed that yields of orthophosphate (OP), silica (SI), nitrate–nitrites (NO 3 -NO 2 ), and total suspended sediment (TSS) are major contributors to nutrient yields. It also showed that land-planted crops, density of population, domestic and industrial discharges, and precipitation are fundamental causes of excess nutrient loads in MARB. These factors are of great significance for the excess nutrient load management and pollution control of the Mississippi River. It was found that the average nutrient yields were stable within the sub-MARB area, but the large nitrogen yields in the upper MARB and the large phosphorus yields in the lower MARB were of great concern. t-distributed stochastic neighbor embedding (t-SNE) revealed interesting nonlinear and local structures in nutrient yield distributions. Clustering analysis (CA) showed the detailed development of similarities in the nutrient yield distribution. Moreover, PCA, t-SNE, and CA showed consistent clustering results. This study demonstrated that the integration of dimension reduction techniques, PCA, and t-SNE with CA techniques in machine learning are effective tools for the visualization of the structures of the correlations in high-dimensional datasets of nutrient yields and provide a comprehensive understanding of the correlations in the distributions of nutrient loads across the MARB.
Spatiotemporal Filtering Using Principal Component Analysis and Karhunen-Loeve Expansion Approaches for Regional GPS Network Analysis
Spatial filtering is an effective way to improve the precision of coordinate time series for regional GPS networks by reducing so-called common mode errors, thereby providing better resolution for detecting weak or transient deformation signals. The commonly used approach to regional filtering assumes that the common mode error is spatially uniform, which is a good approximation for networks of hundreds of kilometers extent, but breaks down as the spatial extent increases. A more rigorous approach should remove the assumption of spatially uniform distribution and let the data themselves reveal the spatial distribution of the common mode error. The principal component analysis (PCA) and the Karhunen-Loeve expansion (KLE) both decompose network time series into a set of temporally varying modes and their spatial responses. Therefore they provide a mathematical framework to perform spatiotemporal filtering.We apply the combination of PCA and KLE to daily station coordinate time series of the Southern California Integrated GPS Network (SCIGN) for the period 2000 to 2004. We demonstrate that spatially and temporally correlated common mode errors are the dominant error source in daily GPS solutions. The spatial characteristics of the common mode errors are close to uniform for all east, north, and vertical components, which implies a very long wavelength source for the common mode errors, compared to the spatial extent of the GPS network in southern California. Furthermore, the common mode errors exhibit temporally nonrandom patterns.
Portland Urban Development: Quantifying and Visualizing Urban Heat with Compounding Vulnerabilities to Support Community Depaving Initiatives
Urban heat is a pressing concern in Portland, Oregon as climate change induced heat waves increase. Cities experience higher temperatures due to the urban heat island effect (UHI), and environmental injustice and disenfranchisement in minority communities expose low-income and Black, Indigenous, and People of Color (BIPOC) residents to more extreme and debilitating heat events. Our team identified Portland’s communities on the frontlines of urban heat impacts by overlapping environmental and social vulnerabilities using NASA Earth observations. We partnered with Depave, a Portland-based nonprofit that works alongside communities to replace pavement with greenspace in historically disenfranchised areas. Using Landsat 8 Thermal Infrared Sensor (TIRS) imagery, we mapped Land Surface Temperature (LST) and developed a heat-specific Social Vulnerability Index (SVI) through a Principal Component Analysis (PCA) to identify Portland’s communities with the highest potential heat vulnerability. Then, we calculated the temperature change of depaving in six case studies to quantify Depave's efforts in heat mitigation and environmental justice. Our analysis demonstrated that, throughout Portland, there are frontline communities experiencing high potential social vulnerability to extreme temperatures due to environmental injustices and over-pavement. Finally, Depave’s impact on urban heat is observable and quantifiable using remote-sensing data and tools, with an average of 1ºF LST decrease across the six case studies. We illustrated the significance of local urban heat mitigation efforts and propose next steps for conducting inclusive and intentional research that highlights the lived experiences and resilience of frontline communities.
Reducing the matrix effect in mass spectral imaging of biofilms using flow-cell culture
The interactions between soil microorganisms and soil minerals play a crucial role in the formation and evolution of minerals and the stability of soil aggregates. Due to the heterogeneity and diversity of the soil environment, the under-standing of the functions of bacterial biofilms in soil minerals at the microscale is limited. A soil mineral-bacterial biofilm system was used as a model in this study, and it was analyzed by time-of-flight secondary ion mass spectrometry (ToF-SIMS) to acquire molecular level information. Static culture in multi-wells and dynamic flow-cell culture in microfluidics of biofilms were investigated. Our results show that more characteristic molecules of biofilms can be observed in SIMS spectra of the flow-cell culture. In contrast, biofilm signature peaks are buried under the mineral components in SIMS spectra in the static culture case. Spectral overlay was used in peak selection prior to performing Principal component analysis (PCA). Comparisons of the PCA results between the static and flow-cell culture show more pronounced molecular features and higher loadings of organic peaks of the dynamic cultured specimens. For example, fatty acids secreted from bacterial biofilm extracellular polymeric substance are likely to be responsible for biofilm dispersal due to mineral treatment up to 48 h. Such findings suggest that the use of microfluidic cells to dynamically culture biofilms be a more suitable method for reducing the matrix effect arisen from the growth medium and minerals as a perturbation fac-tor for improved spectral and multivariate analysis of complex mass spectral data in ToF-SIMS. These results show that the interaction mechanism between biofilms and soil minerals at the molecular level can be better studied using the flow-cell culture and advanced mass spectral imaging techniques like ToF-SIMS.
Predicting non-linear stress–strain response of mesostructured cellular materials using supervised autoencoder
Recent breakthroughs in advanced manufacturing capabilities have made it possible to design and print sophisticated topologies of cellular structures using diverse engineering materials such as metals, polymers, and ceramics. In these architectured materials, it is often desirable to tailor the mechanical properties by altering the unit cell topology. This necessitates an in-depth understanding of how the topology of the unit cell structure affects the macroscopic behavior of the material in both the linear and the non-linear regimes encountered under large compression. Here, we have developed a machine learning (ML) approach capable of accelerating the prediction of the stress–strain response of a polymer-based cellular structure under uniaxial confined compression. As part of generating the training data for ML, 60,000 mesostructures were generated using a relatively novel approach based on cellular automata, and their corresponding stress–strain responses were obtained from the finite element simulations. Principal component analysis (PCA) was used to reduce the dimensionality of the stress–strain curves. With only 20 principal components, PCA captured 99.89% of the variance in the stress–strain curves while reducing the dimensionality by 5X. ML using supervised autoencoder was able to successfully speed up the prediction of the non-linear stress–strain response of a unit cell by up to 4600X. The proposed method can serve as an efficient data generation tool and a rapid means for predicting the structure–property relationship through accelerated forward modeling of cellular materials under compaction, in cases where the macroscopic stress–strain response is governed by the unit-cell topology.
Principal Component Analysis of azimuthal flow in intermediate-energy heavy-ion reactions
Principal Component Analysis (PCA) via Singular Value Decomposition (SVD) of large datasets is an adaptive exploratory method to uncover natural patterns underlying the data. Several recent applications of the PCA-SVD to event-by-event single-particle azimuthal angle distribution matrices in ultra-relativistic heavy-ion collisions at RHIC-LHC energies indicate that the sine and cosine functions chosen a priori in the traditional Fourier analysis are naturally the most optimal basis for azimuthal flow studies according to the data itself. We perform PCA-SVD analyses of mid-central Au+Au collisions at $E$ $beam$ / $A$ =1.23 GeV simulated using an isospin-dependent Boltzmann-Uehling-Uhlenbeck (IBUU) transport model to address the following two questions: (1) if the principal components of the covariance matrix of nucleon azimuthal angle distributions in heavy-ion reactions around 1 GeV/nucleon are naturally sine and/or cosine functions and (2) what if any advantages the PCA-SVD may have over the traditional flow analysis using the Fourier expansion for studying the EOS of dense nuclear matter. In conclusion, we find that (1) in none of our analyses the principal components come out naturally as sine and/or cosine functions, (2) while both the eigenvectors and eigenvalues of the covariance matrix are appreciably EOS dependent, the PCA-SVD has no apparent advantage over the traditional Fourier analysis for studying the EOS of dense nuclear matter using the azimuthal collective flow in heavy-ion collisions.
Forward and Inverse Models for Satellite Remote Sensors using Principal Component Analysis
Satellite remote sensors such as AIRS on Aqua, CrIS on S-NPP, NOAA20 and JPSS-2, IASI on Metop A, B, and C make millions of observations each day with thousands of spectral channels for each observation; this poses challenges for efficiently inversion of the inherently large dataset as needed to retrieve atmospheric and surface properties. This presentation will illustrate the use of Principal Component Analysis (PCA) to speed up radiative transfer forward model calculations and to stabilize the inversion algorithms. A Principal Component-based radiative transfer model (PCRTM) developed at NASA Langley Research Center can simulate top of atmosphere (TOA) radiance or reflectance spectra from 50 cm-1 to 50000 cm-1 (200 m to 0.20 m quickly and accurately. PCRTM demonstrated very high accuracy relative to reference line-by-line radiative transfer models and it saves orders of magnitude computational time. Examples of the PCRTM model developed for hyperspectral sensors such as AIRS, CrIS, IASI, NAST-I, SHIS, CPF, TEMPO, SBG, OMI, and SCIAMACHY will be presented. In addition to using the PCRTM as forward model, the NASA Langley developed inversion algorithm also uses PCA to compress the state vector into a compressed dimension to speed up and stabilize the inversion process. Examples of retrieved atmospheric temperature, water vapor, CO2, CO, CH4, N2O, and O3 profiles, cloud properties (optical depth, size, phase, and height), and surface properties (surface emissivity spectra and skin temperatures) will be presented. This algorithm is being transitioned to the NASA Sounder SIPS and NASA's Goddard Earth Sciences Data and Information Services Center (GES DISC).