SEARCH · Engineering Papers
Results for “Principal component analysis”
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Towards a data-driven paradigm for characterizing plastic anisotropy using principal components analysis and manifold learning
Not Available
Covariance Matrix Preparation for Quantum Principal Component Analysis
Not Available
Principal Component Analysis for Modal Optimization in Resonant Ultrasound Spectroscopy
Explore the source record for details and available documents.
An evaluation of air quality in major urban areas of India
Rapid economic growth and burgeoning population have contributed to enhanced levels of PM 2.5 concentrations in urban regions of India. Evaluation of ambient air quality facilitates the assessment of effectiveness of emission control measures and early identification of new sources. This study provides a comprehensive statistical analysis of PM 2.5 concentrations in key urban areas across India, including Delhi, Kolkata, Mumbai, Chennai, Hyderabad, and several regional centers. Data from 2017 to 2023 was analyzed using trend analysis, cluster analysis, principal component analysis, and geostatistical interpolation to understand spatiotemporal variations and sources. The analysis reveals significant differences in spatial distribution of PM 2.5 concentrations with high annual averages in urban regions in Indo-Gangetic plain (82–123 μg m −3 ) and relatively lower concentrations (29–46 μg m −3 ) in southern urban areas of Kerala, Tamil Nadu and Andhra Pradesh. Delhi state had the highest 24-averaged PM 2.5 concentrations (112 μg m −3 ) followed by urban regions in Uttar Pradesh, Bihar and West Bengal (94 μg m −3 ). Trend analysis from 2017 to 2023 revealed an overall 2.5% decline in site-wide PM2.5 concentrations, with the exception of Ludhiana, which exhibited a consistent annual increase of 10%. Principal component analysis (PCA) attributes 30% of the variance to wintertime emissions, 13% to biomass burning, and 18% to the regional haze in the northern Indo-Gangetic Plain. Different analyses clearly demonstrates the contribution of biomass burning to pollution in Delhi and surrounding cities. Transboundary pollution to Kolkata is likely from the highly polluted region in Indo-Gangetic Plain. Coastal cities of Mumbai and Chennai has relatively lower pollution attributed to the influence of sea breeze dilution, with mostly local contribution and some potential transport from upwind industry clusters. Hyderabad also has local contribution due to high density of vehicular traffic and local small industries. This study shows that mitigation efforts targeting clusters of regions should be undertaken to curb the high PM2.5 pollution. Policy measures should be implemented both at local and the intra-state level to address shared sources and transport of pollution.
Uncertainty Quantification for Smooth Functional Data with Application to Material Properties
This document outlines a method for processing functional output (i.e., curves) for the ultimate purpose of sampling curves under specified input conditions for use in modeling and simulation uncertainty quantification (UQ) studies. A set of benchmark curves sufficiently representative of the relevant scenario(s) being simulated are provided to the process and formatted as described in Section 1. Principal Component Analysis (PCA) is utilized to discover the components of uncertainty in the benchmark curves and is outlined in Section 2. Section 3 describes the application of uncertainty quantification to the PCA results for the purpose of sampling curves to be used in UQ analysis. Section 4 applies these techniques to an example benchmark dataset. Concluding remarks are provided in the final section.
Combining ToF‐SIMS and Multivariate Analysis to Resolve Active Sites on Ni‐Based HER Catalysts
Unambiguous identification of active sites in heterogeneous catalysis remains a major challenge, particularly for materials with ultrathin, chemically mixed surface layers. Here, we demonstrate a generalizable approach that combines time-of-flight secondary ion mass spectrometry (ToF-SIMS) with multivariate statistical analysis (principal component analysis [PCA] and multivariate curve resolution [MCR]) to resolve catalytically relevant motifs at the nanoscale. Using Ni electrodes as a model system, PCA distinguished hydroxide-enriched domains from oxide- and metal-rich regions, while MCR decomposed depth profiles and 3D images into hydroxide, oxide, and metallic layers with nanometer resolution. A unique secondary-ion fragment, NiO 3 H 3 − (m/z 108.94), emerged as a marker of hydroxide-rich environments and correlated with hydrogen evolution reaction (HER) activity across a series of Ni electrodes. Complementary density functional theory (DFT) calculations revealed that Ni(OH) 2 clusters adjacent to metallic Ni offer the most favorable water dissociation energetics, establishing the structural origin of the marker. While illustrated here for Ni-based HER, this workflow provides a broadly applicable framework to isolate and rank near-surface patterns that govern catalytic activity, thereby extending ToF-SIMS from a qualitative probe to a predictive tool for active site identification.
Toward wide-spectrum antivirals against coronaviruses: Molecular characterization of SARS-CoV-2 NSP13 helicase inhibitors
To date, effective therapeutic treatments that confer strong attenuation against coronaviruses (CoVs) remain elusive. Among potential drug targets, the helicase of CoVs is attractive due to its sequence conservation and indispensability. We rely on atomistic molecular dynamics simulations to explore the structural coordination and dynamics associated with the SARS-CoV-2 Nsp13 apo enzyme, as well as their complexes with natural ligands. A complex communication network is revealed among the five domains of Nsp13, which is differentially activated because of the presence of the ligands, as shown by shear strain analysis, principal components analysis, dynamical cross-correlation matrix analysis, and water transport analysis. The binding free energy and the corresponding mechanism of action are presented for three small molecules that were shown to be efficient inhibitors of the previous SARS-CoV Nsp13 enzyme. Together, our findings provide critical fresh insights for rational design of broad-spectrum antivirals against CoVs.
Two (or more) for one: Identifying classes of household energy- and water-saving measures to understand the potential for positive spillover
A key component of behavior-based energy conservation programs is the identification of target behaviors. A common approach is to target behaviors with the greatest energy-saving potential. The concept of behavioral spillover introduces further considerations, namely that adoption of one energy-saving behavior may increase (or decrease) the likelihood of other energy-saving behaviors. This research aimed to identify and describe household energy- and water-saving measure classes within which positive spillover is likely to occur (e.g., adoption of energy-efficient appliances may correlate with adoption of water-efficient appliances), and explore demographic and psychographic predictors of each. Nearly 1,000 households in a California city were surveyed and asked to report whether they had adopted 75 different energy- and/or water-saving measures. Principal Component Analysis and Network Analysis based on correlations between adoption of these diverse measures revealed and characterized eight water-energy-saving measure classes: Water Conservation, Energy Conservation, Maintenance and Management, Efficient Appliance, Advanced Efficiency, Efficient Irrigation, Green Gardening, and Green Landscaping. Understanding these measure classes can help guide behavior-based energy program developers in selecting target behaviors and designing interventions.
Elastic functional changepoint detection of climate impacts from localized sources
Detecting changepoints in functional data has become an important problem as interest in monitoring of climate phenomenon has increased, where the data is functional in nature. Here, the observed data often contains both amplitude (y-axis) and phase (x-axis) variability. If not accounted for properly, true changepoints may be undetected, and the estimated underlying mean change functions will be incorrect. In this article, an elastic functional changepoint method is developed which properly accounts for these types of variability. The method can detect amplitude and phase changepoints which current methods in the literature do not, as they focus solely on the amplitude changepoint. This method can easily be implemented using the functions directly or can be computed via functional principal component analysis to ease the computational burden. We apply the method and its nonelastic competitors to both simulated data and observed data to show its efficiency in handling data with phase variation with both amplitude and phase changepoints. We use the method to evaluate potential changes in stratospheric temperature due to the eruption of Mt. Pinatubo in the Philippines in June 1991. Using an epidemic changepoint model, we find evidence of a increase in stratospheric temperature during a period that contains the immediate aftermath of Mt. Pinatubo, with most detected changepoints occurring in the tropics as expected.
A comparative study of the physical properties for a representative sample of Narrow and Broad-line Seyfert galaxies
ABSTRACT We present a comparative study of the physical properties of a homogeneous sample of 144 Narrow line Seyfert 1 (NLSy1) and 117 Broad-line Seyfert 1 (BLSy1) galaxies. These two samples are in a similar luminosity and redshift range and have optical spectra available in the 16th data release of Sloan Digital Sky Survey (SDSS-DR16) and X-ray spectra in either XMM-NEWTON or ROSAT. Direct correlation analysis and a principal component analysis (PCA) have been performed using ten observational and physical parameters obtained by fitting the optical spectra and the soft X-ray photon indices as another parameter. We confirm that the established correlations for the general quasar population hold for both types of galaxies in this sample despite significant differences in the physical properties. We characterize the sample also using the line shape parameters, namely the asymmetry and kurtosis indices. We find that the fraction of NLSy1 galaxies showing outflow signatures, characterized by blue asymmetries, is higher by a factor of about 3 compared to the corresponding fraction in BLSy1 galaxies. The presence of high iron content in the broad-line region of NLSy1 galaxies in conjunction with higher Eddington ratios can be the possible reason behind this phenomenon. We also explore the possibility of using asymmetry in the emission lines as a tracer of outflows in the inner regions of Active Galactic Nuclei. The PCA results point to the NLSy1 and BLSy1 galaxies occupying different parameter spaces, which challenges the notion that NLSy1 galaxies are a subclass of BLSy1 galaxies.
Time series methods for the analysis of soundscapes and other cyclical ecological data
Biodiversity monitoring has entered an era of ‘big data’, exemplified by a near-continuous collection of sounds, images, chemical and other signals from organisms in diverse ecosystems. Such data streams have the potential to help identify new threats, assess the effectiveness of conservation interventions, as well as generate new ecological insights. However, appropriate analytical methods are often still missing, particularly with respect to characterizing cyclical temporal patterns. Here, we present a framework for characterizing and analysing ecological responses that represent nonstationary, complex temporal patterns and demonstrate the value of using Fourier transforms to decorrelate continuous data points. In our example, we use a framework based on three approaches (spectral analysis, magnitude squared coherence, and principal component analysis) to characterize differences in tropical forest soundscapes within and across sites and seasons in Gabon. By reconstructing the underlying, cyclic behaviour of the soundscape for each site, we show how one can identify circadian patterns in acoustic activity. Soundscapes in the dry season had a complex diel cycle, requiring multiple harmonics to represent daily variation, while in the wet season there was less variance attributable to the daily cyclic patterns. Our framework can be applied to most continuous, or near-continuous ecological data collected at a fine temporal resolution, allowing ecologists to explore patterns of temporal autocorrelation at multiple levels for biologically meaningful trends. Such methods will become indispensable as biological big data are used to understand the impact of anthropogenic pressures on biodiversity and to inform efforts to mitigate them.
Cover crops and poultry litter impact on soil structural stability in dryland soybean production in southeastern United States
Abstract This study explored the efficacy of soil aggregate indices in quantifying soil structural development, utilizing 5‐year field experiment data from the Southeastern United States. The experiment utilized a split‐plot design with cover crops (native vegetation as control, cereal rye (Secale cerealeL.), winter wheat (Triticum aestivum), hairy vetch (Vicia villosa), and mustard (Brassica rapa) plus cereal rye as the main factor and fertilizer source (no fertilizer as control, inorganic fertilizer with phosphorus, potassium, and elemental sulfur, and poultry litter) as the secondary factor. Aggregate size fractions were determined using the wet‐sieving method, and aggregate stability index (ASI), mean weight diameter (MWD), geometric mean diameter (GMD), and fractal dimension (FD) were calculated to assess soil structural stability. Main effects results indicated that cereal rye (55.11%) and poultry litter (50.97%) exhibited the highest ASI values. The highest MWD, GMD, and FD were observed under mustard plus cereal rye (1.187 mm), cereal rye (0.462 mm), and hairy vetch (2.573), respectively. Principal component analysis revealed that cover crops significantly improved soil aggregate structure and stability, overcoming limitations of sole fertilization practices. Regression analysis suggested that ASI, MWD, and GWD positively correlated with soil organic carbon, whereas FD negatively correlated with MWD, GMD, and ASI. Principal component analysis exhibited that FD decreased with increasing soil organic carbon, ASI, MWD, and GMD, demonstrating that lower FD values indicate enhanced soil aggregation and structure. Assessed indices, FD included, effectively gauged soil structural stability. These metrics should be prioritized in managerial decisions to support soil productivity and health in agricultural systems.
Variational encoder geostatistical analysis (VEGAS) with an application to large scale riverine bathymetry
Estimation of riverbed profiles, also known as bathymetry, plays a vital role in many applications, such as safe and efficient inland navigation, prediction of bank erosion, land subsidence, and flood risk management. The high cost and complex logistics of direct bathymetry surveys, i.e, depth imaging, have encouraged the use of indirect measurements such as surface flow velocities. However, estimating high-resolution bathymetry from indirect measurements is an inverse problem that can be computationally challenging. Here, we propose a reduced-order model (ROM) based approach that utilizes a variational autoencoder (VAE), a type of deep neural network with a narrow layer in the middle, to compress bathymetry and flow velocity information and accelerate bathymetry inverse problems from flow velocity measurements. In our application, the shallow-water equations (SWE) with appropriate boundary conditions (BCs), e.g., the discharge and/or the free surface elevation, constitute the forward problem, to predict flow velocity. Then, ROMs of the SWEs are constructed on a nonlinear manifold of low dimensionality through a variational encoder and the bathymetry inversion problem is derived on the low-dimensional latent space in a Hierarchical Bayesian setting. Further, the reformulation allows variational inference with a small number (e.g., $\mathscr{O}$ (100) of ROM runs and efficient uncertainty quantification. We have tested our inversion approach on a one-mile reach of the Savannah River, GA, USA. Once the neural network is trained (offline stage), the proposed technique can perform the inversion operation orders of magnitude faster than traditional inversion methods that are commonly based on linear projections, such as principal component analysis (PCA), or the principal component geostatistical approach (PCGA). Furthermore, tests show that the algorithm can estimate the bathymetry with good accuracy even with sparse flow velocity measurements.
Using the optimal combined index weight ratio to improve the probability of anomaly detection in big area additive manufacturing
Big Area Additive Manufacturing (BAAM) of composites requires significant time, energy, and material, so it is critical to reduce production inefficiencies to make functional parts without multiple iterations. Statistical process control coupled with Principal Component Analysis (PCA) is a powerful technique that provides a quick, computationally inexpensive, and intuitive way for operators to detect defects that form in a manufacturing process without massive datasets. Recently, a combined index that is a weighted sum of the Hotelling's T 2 and squared residual error statistics has been proposed that can be monitored in one chart, improving interpretation accuracy and simplicity. However, the literature does not offer a formal method to optimise the weights. Here, we introduce two new approaches to the traditional weight selection approach using simulated and BAAM image data. Approach 1 uses a theoretically motivated optimum inspired by probabilistic principal component analysis. Approach 2 systematically varies the ratio of the weights to find the optimum. We show that approach 1 delivers optimal anomaly detection performance in select cases while approach 2 fares better in practice. Surprisingly, we also show that choosing a more complex PCA model has a minimal negative impact on anomaly detection performance compared to a more simplistic model.
Structuring Nutrient Yields throughout Mississippi/Atchafalaya River Basin Using Machine Learning Approaches
To minimize the eutrophication pressure along the Gulf of Mexico or reduce the size of the hypoxic zone in the Gulf of Mexico, it is important to understand the underlying temporal and spatial variations and correlations in excess nutrient loads, which are strongly associated with the formation of hypoxia. This study’s objective was to reveal and visualize structures in high-dimensional datasets of nutrient yield distributions throughout the Mississippi/Atchafalaya River Basin (MARB). For this purpose, the annual mean nutrient concentrations were collected from thirty-three US Geological Survey (USGS) water stations scattered in the upper and lower MARB from 1996 to 2020. Eight surface water quality indicators were selected to make comparisons among water stations along the MARB over the past two decades. Principal component analysis (PCA) was used to comprehensively evaluate the nutrient yields across thirty-three USGS monitoring stations and identify the major contributing nutrient loads. The results showed that all samples could be analyzed using two main components, which accounted for 81.6% of the total variance. The PCA results showed that yields of orthophosphate (OP), silica (SI), nitrate–nitrites (NO 3 -NO 2 ), and total suspended sediment (TSS) are major contributors to nutrient yields. It also showed that land-planted crops, density of population, domestic and industrial discharges, and precipitation are fundamental causes of excess nutrient loads in MARB. These factors are of great significance for the excess nutrient load management and pollution control of the Mississippi River. It was found that the average nutrient yields were stable within the sub-MARB area, but the large nitrogen yields in the upper MARB and the large phosphorus yields in the lower MARB were of great concern. t-distributed stochastic neighbor embedding (t-SNE) revealed interesting nonlinear and local structures in nutrient yield distributions. Clustering analysis (CA) showed the detailed development of similarities in the nutrient yield distribution. Moreover, PCA, t-SNE, and CA showed consistent clustering results. This study demonstrated that the integration of dimension reduction techniques, PCA, and t-SNE with CA techniques in machine learning are effective tools for the visualization of the structures of the correlations in high-dimensional datasets of nutrient yields and provide a comprehensive understanding of the correlations in the distributions of nutrient loads across the MARB.
Functional Data Analysis for Extracting the Intrinsic Dimensionality of Spectra: Application to Chemical Homogeneity in the Open Cluster M67
High-resolution spectroscopic surveys of the Milky Way have entered the Big Data regime and have opened avenues for solving outstanding questions in Galactic archeology. However, exploiting their full potential is limited by complex systematics, whose characterization has not received much attention in modern spectroscopic analyses. In this work, we present a novel method to disentangle the component of spectral data space intrinsic to the stars from that due to systematics. Using functional principal component analysis on a sample of 18,933 giant spectra from APOGEE, we find that the intrinsic structure above the level of observational uncertainties requires ≈10 functional principal components (FPCs). Our FPCs can reduce the dimensionality of spectra, remove systematics, and impute masked wavelengths, thereby enabling accurate studies of stellar populations. To demonstrate the applicability of our FPCs, we use them to infer stellar parameters and abundances of 28 giants in the open cluster M67. We employ Sequential Neural Likelihood, a simulation-based Bayesian inference method that learns likelihood functions using neural density estimators, to incorporate non-Gaussian effects in spectral likelihoods. By hierarchically combining the inferred abundances, we limit the spread of the following elements in M67: Fe ≲ 0.02 dex; C ≲ 0.03 dex; O, Mg, Si, Ni ≲ 0.04 dex; Ca ≲ 0.05 dex; N, Al ≲ 0.07 dex (at 68% confidence). Our constraints suggest a lack of self-pollution by core-collapse supernovae in M67, which has promising implications for the future of chemical tagging to understand the star formation history and dynamical evolution of the Milky Way.
DESI Emission-line Galaxies: Unveiling the Diversity of [O II ] Profiles and Its Links to Star Formation and Morphology
We study the [O II ] profiles of emission-line galaxies (ELGs) from the Early Data Release of the Dark Energy Spectroscopic Instrument (DESI). To this end, we decompose and classify the shape of [O II ] profiles with the first two eigenspectra derived from principal component analysis. Our results show that DESI ELGs have diverse line profiles, which can be categorized into three main types: (1) narrow lines with a median width of ∼50 km s −1 , (2) broad lines with a median width of ∼80 km s −1 , and (3) two redshift systems with a median velocity separation of ∼150 km s −1 , i.e., double-peak galaxies. To investigate the connections between the line profiles and galaxy properties, we utilize the information from the COSMOS data set and compare the properties of ELGs, including star formation rate (SFR) and galaxy morphology, with the average properties of reference star-forming galaxies with similar stellar mass, sizes, and redshifts. Our findings show that, on average, DESI ELGs have a higher SFR and more asymmetrical/disturbed morphology than the reference galaxies. Moreover, we uncover a relationship between the line profiles, the excess SFR, and the excess asymmetry parameter, showing that DESI ELGs with broader [O II ] line profiles have more disturbed morphology and higher SFR than the reference star-forming galaxies. Finally, we discuss possible physical mechanisms giving rise to the observed relationship and the implications of our findings on the galaxy clustering measurements, including the halo occupation distribution modeling of DESI ELGs and the observed excess velocity dispersion of the satellite ELGs.