Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Principal component analysis”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 523 records · Page 29

NMR Spectroscopy for Protein Higher Order Structure Similarity Assessment in Formulated Drug Products

Peptide and protein drug molecules fold into higher order structures (HOS) in formulation and these folded structures are often critical for drug efficacy and safety. Generic or biosimilar drug products (DPs) need to show similar HOS to the reference product. The solution NMR spectroscopy is a non-invasive, chemically and structurally specific analytical method that is ideal for characterizing protein therapeutics in formulation. However, only limited NMR studies have been performed directly on marketed DPs and questions remain on how to quantitively define similarity. Here, NMR spectra were collected on marketed peptide and protein DPs, including calcitonin-salmon, liraglutide, teriparatide, exenatide, insulin glargine and rituximab. The 1D 1 H spectral pattern readily revealed protein HOS heterogeneity, exchange and oligomerization in the different formulations. Principal component analysis (PCA) applied to two rituximab DPs showed consistent results with the previously demonstrated similarity metrics of Mahalanobis distance (D M ) of 3.3. The 2D 1 H- 13 C HSQC spectral comparison of insulin glargine DPs provided similarity metrics for chemical shift difference (Δδ) and methyl peak profile, i.e., 4 ppb for 1 H, 15 ppb for 13 C and 98% peaks with equivalent peak height. Finally, 2D 1 H- 15 N sofast HMQC was demonstrated as a sensitive method for comparison of small protein HOS. The application of NMR procedures and chemometric analysis on therapeutic proteins offer quantitative similarity assessments of DPs with practically achievable similarity metrics.

59 BASIC BIOLOGICAL SCIENCES↗

Predicting Short-Term Deformation in the Central Valley Using Machine Learning

Land subsidence caused by excessive groundwater pumping in Central Valley, California, is a major issue that has several negative impacts such as reduced aquifer storage and damaged infrastructures which, in turn, produce an economic loss due to the high reliance on crop production. This is why it is of utmost importance to routinely monitor and assess the surface deformation occurring. Two main goals that this paper attempts to accomplish are deformation characterization and deformation prediction. The first goal is realized through the use of Principal Component Analysis (PCA) applied to a series of Interferomtric Synthetic Aperture Radar (InSAR) images that produces eigenimages displaying the key characteristics of the subsidence. Water storage changes are also directly analyzed by the use of data from the Gravity Recovery and Climate Experiment (GRACE) twin satellites and the Global Land Data Assimilation System (GLDAS). The second goal is accomplished by building a Long Short-Term Memory (LSTM) model to predict short-term deformation after developing an InSAR time series using LiCSBAS, an open-source InSAR time series package. The model is applied to the city of Madera and produces better results than a baseline averaging model and a one dimensional convolutional neural network (CNN) based on a mean squared error metric showing the effectiveness of machine learning in deformation prediction as well as the potential for incorporation in hazard mitigation models. The model results can directly aid policy makers in determining the appropriate rate of groundwater withdrawal while maintaining the safety and well-being of the population as well as the aquifers’ integrity.

58 GEOSCIENCES↗

High-Throughput Field Plant Phenotyping: A Self-Supervised Sequential CNN Method to Segment Overlapping Plants

High-throughput plant phenotyping—the use of imaging and remote sensing to record plant growth dynamics—is becoming more widely used. The first step in this process is typically plant segmentation, which requires a well-labeled training dataset to enable accurate segmentation of overlapping plants. However, preparing such training data is both time and labor intensive. To solve this problem, we propose a plant image processing pipeline using a self-supervised sequential convolutional neural network method for in-field phenotyping systems. This first step uses plant pixels from greenhouse images to segment nonoverlapping in-field plants in an early growth stage and then applies the segmentation results from those early-stage images as training data for the separation of plants at later growth stages. The proposed pipeline is efficient and self-supervising in the sense that no human-labeled data are needed. We then combine this approach with functional principal components analysis to reveal the relationship between the growth dynamics of plants and genotypes. We show that the proposed pipeline can accurately separate the pixels of foreground plants and estimate their heights when foreground and background plants overlap and can thus be used to efficiently assess the impact of treatments and genotypes on plant growth in a field environment by computer vision techniques. This approach should be useful for answering important scientific questions in the area of high-throughput phenotyping.

59 BASIC BIOLOGICAL SCIENCES↗

Dimensionality Reduction of SDSS Spectra with Variational Autoencoders

High-resolution galaxy spectra contain much information about galactic physics, but the high dimensionality of these spectra makes it difficult to fully utilize the information they contain. We apply variational autoencoders (VAEs), a nonlinear dimensionality reduction technique, to a sample of spectra from the Sloan Digital Sky Survey (SDSS). In contrast to principal component analysis (PCA), a widely used technique, VAEs can capture nonlinear relationships between latent parameters and the data. We find that a VAE can reconstruct the SDSS spectra well with only six latent parameters, outperforming PCA with the same number of components. Different galaxy classes are naturally separated in this latent space, without class labels having been given to the VAE. The VAE latent space is interpretable because the VAE can be used to make synthetic spectra at any point in latent space. For example, making synthetic spectra along tracks in latent space yields sequences of realistic spectra that interpolate between two different types of galaxies. Using the latent space to find outliers may yield interesting spectra: in our small sample, we immediately find unusual data artifacts and stars misclassified as galaxies. In this exploratory work, we show that VAEs create compact, interpretable latent spaces that capture nonlinear features of the data. While a VAE takes substantial time to train (≈1 day for 48,000 spectra), once trained, VAEs can enable the fast exploration of large astronomical data sets.

79 ASTRONOMY AND ASTROPHYSICS↗

A Novel Machine Learning Approach to Disentangle Multitemperature Regions in Galaxy Clusters

The hot intracluster medium (ICM) surrounding the heart of galaxy clusters is a complex medium that comprises various emitting components. Although previous studies of nearby galaxy clusters, such as the Perseus, the Coma, or the Virgo cluster, have demonstrated the need for multiple thermal components when spectroscopically fitting the ICM’s X-ray emission, no systematic methodology for calculating the number of underlying components currently exists. In turn, underestimating or overestimating the number of components can cause systematic errors in the emission parameter estimations. In this paper, we present a novel approach to determining the number of components using an amalgam of machine learning techniques. Synthetic spectra containing a various number of underlying thermal components were created using well-established tools available from the Chandra X-ray Observatory. The dimensions of the training set was initially reduced using principal component analysis and then categorized based on the number of underlying components using a random forest classifier. Our trained and tested algorithm was subsequently applied to Chandra X-ray observations of the Perseus cluster. Our results demonstrate that machine learning techniques can efficiently and reliably estimate the number of underlying thermal components in the spectra of galaxy clusters, regardless of the thermal model (MEKAL versus APEC). We also confirm that the core of the Perseus cluster contains a mix of differing underlying thermal components. We emphasize that although this methodology was trained and applied on Chandra X-ray observations, it is readily portable to other current (e.g., XMM-Newton, eROSITA) and upcoming (e.g., Athena, Lynx, XRISM) X-ray telescopes. The code is publicly available at https://github.com/XtraAstronomy/Pumpkin.

79 ASTRONOMY AND ASTROPHYSICS↗

Star–Galaxy Image Separation with Computationally Efficient Gaussian Process Classification

Abstract We introduce a novel method for discerning optical telescope images of stars from those of galaxies using Gaussian processes (GPs). Although applications of GPs often struggle in high-dimensional data modalities such as optical image classification, we show that a low-dimensional embedding of images into a metric space defined by the principal components of the data suffices to produce high-quality predictions from real large-scale survey data. We develop a novel method of GP classification hyperparameter training that scales approximately linearly in the number of image observations, which allows for application of GP models to large-size Hyper Suprime-Cam Subaru Strategic Program data. In our experiments, we evaluate the performance of a principal component analysis embedded GP predictive model against other machine-learning algorithms, including a convolutional neural network and an image photometric morphology discriminator. Our analysis shows that our methods compare favorably with current methods in optical image classification while producing posterior distributions from the GP regression that can be used to quantify object classification uncertainty. We further describe how classification uncertainty can be used to efficiently parse large-scale survey imaging data to produce high-confidence object catalogs.

79 ASTRONOMY AND ASTROPHYSICS↗

Archetype-based Redshift Estimation for the Dark Energy Spectroscopic Instrument Survey

We present a computationally efficient galaxy archetype-based redshift estimation and spectral classification method for the Dark Energy Survey Instrument (DESI) survey. The DESI survey currently relies on a redshift fitter and spectral classifier using a linear combination of principal component analysis–derived templates, which is very efficient in processing large volumes of DESI spectra within a short time frame. However, this method occasionally yields unphysical model fits for galaxies and fails to adequately absorb calibration errors that may still be occasionally visible in the reduced spectra. Our proposed approach improves upon this existing method by refitting the spectra with carefully generated physical galaxy archetypes combined with additional terms designed to absorb data reduction defects and provide more physical models to the DESI spectra. We test our method on an extensive data set derived from the survey validation (SV) and Year 1 (Y1) data of DESI. Our findings indicate that the new method delivers marginally better redshift success for SV tiles while reducing catastrophic redshift failure by 10%–30%. At the same time, results from millions of targets from the main survey show that our model has relatively higher redshift success and purity rates (0.5%–0.8% higher) for galaxy targets while having similar success for QSOs. These improvements also demonstrate that the main DESI redshift pipeline is generally robust. Additionally, it reduces the false-positive redshift estimation by 5%–40% for sky fibers. We also discuss the generic nature of our method and how it can be extended to other large spectroscopic surveys, along with possible future improvements.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

SpecDis: Value Added Distance Catalog for 4 Million Stars from DESI Year-1 Data

We present the SpecDis value-added stellar distance catalog accompanying DESI Data Release 1. SpecDis trains a feed-forward neural network (NN) with Gaia parallaxes and gets the distance estimates. To build up an unbiased training sample, we do not apply selections on parallax error or signal-to-noise (S/N) of the stellar spectra, and instead, we incorporate parallax error into the loss function. Moreover, we employ principal component analysis to reduce the noise and dimensionality of stellar spectra. Validated by independent external samples of member stars with precise distances from globular clusters, dwarf galaxies, stellar streams, combined with blue horizontal branch stars, we demonstrate that our distance measurements show no significant bias up to 100 kpc, and are much more precise than Gaia parallax beyond 7 kpc. The median distance uncertainties are 23%, 19%, 11%, and 7% for S/N < 20, 20 ≤ S/N < 60, 60 ≤ S/N < 100, and S/N ≥ 100. Selecting stars with ${\mathrm{log}}\,g\lt 3.8$ and distance uncertainties smaller than 25%, we have more than 74,000 giant candidates within 50 kpc of the Galactic center and 1500 candidates beyond this distance. Additionally, we develop a Gaussian mixture model to identify unresolvable equal-mass binaries by modeling the discrepancy between the NN-predicted and the geometric absolute magnitudes from Gaia parallaxes and identify 120,000 equal-mass binary candidates. Our final catalog provides distances and distance uncertainties for >4 million stars, offering a valuable resource for Galactic astronomy.

astronomy data analysis↗

A Generative Model for Realistic Galaxy Cluster X-Ray Morphologies

Abstract The X-ray morphologies of clusters of galaxies display significant variations, reflecting their dynamical histories and the nonlinear dependence of X-ray emissivity on the density of the intracluster gas. Qualitative and quantitative assessments of X-ray morphology have long been considered a proxy for determining whether clusters are dynamically active or “relaxed.” Conversely, the use of circularly or elliptically symmetric models for cluster emission can be complicated by the variety of complex features realized in nature, spanning scales from megaparsecs down to the resolution limit of current X-ray observatories. In this work, we use mock X-ray images from simulated clusters from The Three Hundred project to define a basis set of cluster image features. We take advantage of the clusters’ approximate self-similarity to minimize the differences between images before encoding the remaining diversity through a distribution of high-order polynomial coefficients. Principal component analysis then provides an orthogonal basis for this distribution, corresponding to natural perturbations from an average model. This representation allows novel, realistically complex X-ray cluster images to be easily generated, and we provide code to do so. The approach provides a simple way to generate training data for cluster image analysis algorithms and could be straightforwardly adapted to generate clusters displaying specific types of features or selected by physical characteristics available in the original simulations.

79 ASTRONOMY AND ASTROPHYSICS↗

Dimensional Reduction for Sampled Priors and Application to Photometric Redshift Distributions

A typical Bayesian inference on the values of some parameters of interest q from some data D involves running a Markov Chain (MC) to sample from the posterior $p$($q$,$n$|$D$) $\propto$ $\mathcal{L}$($D$|$q$,$n$)$p$(q)$p$($n$), where n are some nuisance parameters with a separable prior. In some cases, the nuisance parameters are high-dimensional, and their prior p(n) is itself defined only by a set of samples that have been drawn from some other MC. The MC for the posterior will typically require evaluation of p(n) at arbitrary values of n, i.e., one needs to provide a density estimator over the full n space from the provided samples. But the high dimensionality of n hinders both the density estimation and the efficiency of the MC for the posterior. We describe a solution to this problem: a linear compression of the n space into a much lower-dimensional space u, which projects away directions in n space that cannot appreciably alter $\mathcal{L}$. The algorithm for doing so is a slight modification to principal components analysis, and is less restrictive on p(n) than other proposed solutions to this issue. We demonstrate this “mode projection” technique using the analysis of 2-point correlation functions of weak lensing fields and galaxy density in the Dark Energy Survey, where n is a binned representation of the redshift distribution n(z) of the galaxies.

79 ASTRONOMY AND ASTROPHYSICS↗

The Sloan Digital Sky Survey Quasar Catalog: Sixteenth Data Release

We present the final Sloan Digital Sky Survey IV (SDSS-IV) quasar catalog from Data Release 16 of the extended Baryon Oscillation Spectroscopic Survey (eBOSS). This catalog comprises the largest selection of spectroscopically confirmed quasars to date. The full catalog includes two subcatalogs (the current versions are DR16Q_v4 and DR16Q_Superset_v3 at https://data.sdss.org/sas/dr16/eboss/qso/DR16Q/): a "superset" of all SDSS-IV/eBOSS objects targeted as quasars containing 1,440,615 observations and a quasar-only catalog containing 750,414 quasars, including 225,082 new quasars appearing in an SDSS data release for the first time, as well as known quasars from SDSS-I/II/III. We present automated identification and redshift information for these quasars alongside data from visual inspections for 320,161 spectra. Here, the quasar-only catalog is estimated to be 99.8% complete with 0.3%-1.3% contamination. Automated and visual inspection redshifts are supplemented by redshifts derived via principal component analysis and emission lines. We include emission-line redshifts for Hα, Hβ, Mg II, C III], C IV, and Lyα. Identification and key characteristics generated by automated algorithms are presented for 99,856 broad absorption-line quasars and 35,686 damped Lyman alpha quasars. In addition to SDSS photometric data, we also present multiwavelength data for quasars from the Galaxy Evolution Explorer, UKIDSS, the Wide-field Infrared Survey Explorer, FIRST, ROSAT/2RXS, XMM-Newton, and Gaia. Calibrated digital optical spectra for these quasars can be obtained from the SDSS Science Archive Server.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Advanced Intra-Cycle Detection of Pre-Ignition Events through Phase-Space Transforms of Cylinder Pressure Data

The widespread adoption of boosted, downsized SI engines has brought pre-ignition phenomena into greater focus, as the knock events resulting from pre-ignitions can cause significant hardware damage. Much attention has been given to understanding the causes of pre-ignition and identify lubricant or fuel properties and engine design and calibration considerations that impact its frequency. This helps to shift the pre-ignition limit to higher specific loads and allow further downsizing but does not fundamentally eliminate the problem. Real-time detection and mitigation of pre-ignition would thus be desirable to allow safe engine operation in pre-ignition-prone conditions. This study focuses on advancing the time of detection of pre-ignition in an engine cycle where it occurs. Furthermore, phase space transforms through time-delay embedding of cylinder pressure and principal component analysis were applied to same-cycle detection of pre-ignition and shown to enable detection on the order of a crank degree earlier than deviation in cylinder pressure can be identified through direct statistical observation of the pressure data. Additionally, it appears that the deviation of the trajectory in phase space may offer the opportunity to extend this method to further extend the detection window and allow more time for mitigation actions to occur.

42 ENGINEERING↗

Environmental effects on aerosol–cloud interaction in non-precipitating marine boundary layer (MBL) clouds over the eastern North Atlantic

Abstract. Over the eastern North Atlantic (ENA) ocean, a total of 20 non-precipitating single-layer marine boundary layer (MBL) stratus and stratocumulus cloud cases are selected to investigate the impacts of the environmental variables on the aerosol–cloud interaction (ACIr) using the ground-based measurements from the Department of Energy Atmospheric Radiation Measurement (ARM) facility at the ENA site during 2016–2018. The ACIr represents the relative change in cloud droplet effective radius re with respect to the relative change in cloud condensation nuclei (CCN) number concentration at 0.2 % supersaturation (NCCN,0.2 %) in the stratified water vapor environment. The ACIr values vary from −0.01 to 0.22 with increasing sub-cloud boundary layer precipitable water vapor (PWVBL) conditions, indicating that re is more sensitive to the CCN loading under sufficient water vapor supply, owing to the combined effect of enhanced condensational growth and coalescence processes associated with higher Nc and PWVBL. The principal component analysis shows that the most pronounced pattern during the selected cases is the co-variations in the MBL conditions characterized by the vertical component of turbulence kinetic energy (TKEw), the decoupling index (Di), and PWVBL. The environmental effects on ACIr emerge after the data are stratified into different TKEw regimes. The ACIr values, under both lower and higher PWVBL conditions, more than double from the low-TKEw to high-TKEw regime. This can be explained by the fact that stronger boundary layer turbulence maintains a well-mixed MBL, strengthening the connection between cloud microphysical properties and the below-cloud CCN and moisture sources. With sufficient water vapor and low CCN loading, the active coalescence process broadens the cloud droplet size spectra and consequently results in an enlargement of re. The enhanced activation of CCN and the cloud droplet condensational growth induced by the higher below-cloud CCN loading can effectively decrease re, which jointly presents as the increased ACIr. This study examines the importance of environmental effects on the ACIr assessments and provides observational constraints to future model evaluations of aerosol–cloud interactions.

54 ENVIRONMENTAL SCIENCES↗

The development of rainfall retrievals from radar at Darwin

Abstract. The U.S. Department of Energy Atmospheric Radiation Measurement program Tropical Western Pacific site hosted a C-band polarization (CPOL) radar in Darwin, Australia. It provides 2 decades of tropical rainfall characteristics useful for validating global circulation models. Rainfall retrievals from radar assume characteristics about the droplet size distribution (DSD) that vary significantly. To minimize the uncertainty associated with DSD variability, new radar rainfall techniques use dual polarization and specific attenuation estimates. This study challenges the applicability of several specific attenuation and dual-polarization-based rainfall estimators in tropical settings using a 4-year archive of Darwin disdrometer datasets in conjunction with CPOL observations. This assessment is based on three metrics: statistical uncertainty estimates, principal component analysis (PCA), and comparisons of various retrievals from CPOL data. The PCA shows that the variability in R can be consistently attributed to reflectivity, but dependence on dual-polarization quantities was wavelength dependent for 1 10mmh-1. Rainfall estimates during these conditions primarily originate from deep convective clouds with median drop diameters greater than 1.5 mm. An uncertainty analysis and intercomparison with CPOL show that a Colorado State University blended technique for tropical oceans, with modified estimators developed from video disdrometer observations, is most appropriate for use in all cases, such as when 1 10mmh-1 (deeper convective rain).

54 ENVIRONMENTAL SCIENCES↗

Prediction of DIII-D Pedestal Structure from Externally Controllable Parameters

The sharp increase of pressure at the edge of a high confinement mode (H-mode) plasma, the pedestal, strongly impacts overall plasma performance. Predicting the pedestal is a necessity to control and optimize tokamak operations. An experimental data-driven machine learning (ML) approach is presented that predicts the pedestal heights and widths of electron density (ne) and electron temperature (Te) profiles as well as the separatrix ne from externally controllable parameters such as the plasma shape, heating method and power, and gas puff rate and integrated gas puff. The OMFIT framework was used with DIII-D data to efficiently, robustly, and automatically build a database of pedestal parameters to train machine learning models. Database creation was enabled by the search engine tool for DIII-D data, TokSearch, which parallelizes data fetching, enabling fast searches through basic signals of thousands of DIII-D shots and selection of relevant time intervals. Principal Component Analysis (PCA) separated the database into three clusters that represent classes of plasma shapes that are regularly used in DIII-D. The most important parameters for setting the pedestal structure were plasma current (Ip), toroidal magnetic field (Bφ), neutral beam heating power (PNBI) and shaping quantities. The Deep Jointly Informed Neural Networks (DJINN) algorithm was applied to identify suitable neural network (NN) architectures that appropriately capture the features of the pedestal database. Separate NNs were implemented for each pedestal parameter, and ensembling methods were used to improve the prediction accuracy and allowed estimation of the prediction uncertainty. The pedestal predictions of the test dataset lie within the measurement uncertainties of the pedestal parameters. The NN outperformed simple Linear Regression (LR) analysis, indicating non-linear dependencies in the pedestal structure. The presented achievements illustrate a promising path for future research, using feature extraction to infer experimental trends and thereby improve pedestal models as well as deploying NN for a fast pedestal prediction in DIII-D scenario development.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

A unified development of several techniques for the representation of random vectors and data sets

Linear vector space theory is used to develop a general representation of a set of data vectors or random vectors by linear combinations of orthonormal vectors such that the mean squared error of the representation is minimized. The orthonormal vectors are shown to be the eigenvectors of an operator. The general representation is applied to several specific problems involving the use of the Karhunen-Loeve expansion, principal component analysis, and empirical orthogonal functions; and the common properties of these representations are developed.

Bundick, W. T.↗

A statistical-chemical and thermodynamic approach to the study of lunar mineralogy

Principal components analysis is used to study the chemical compositions of pyroxenes of five Apollo 12 specimens. Important correlations are recognized in the variation of oxide weight per cent. These correlations indicating substitutional relationships can be interpreted as representative of stable and metastable trends of crystallization by using crystal-chemical and thermodynamic information. The per cent variance of pyroxene groups with characteristic trends in each specimen can be evaluated and interpreted in terms of history of crystallization. Distribution of Fe and Mg in certain pairs of olivine and pyroxene, which are found in contact in the rock and which may have crystallized simultaneously, is useful in recognizing the tendency towards chemical equilibrium in Fe-Mg distribution during a limited interval in the liquidus or subsolidus stages.

Saxena, S. K.↗

Application of remote sensing to reconnaissance geologic mapping and mineral exploration

A method of mapping geology at a reconnaissance scale and locating zones of possible hydrothermal alteration has been developed. This method is based on principal component analysis of Landsat digital data and is applied to the desert area of the Chagai Hills, Baluchistan, Pakistan. A method for airborne spectrometric detection of geobotanical anomalies associated with prophyry Cu-Mo mineralization at Heddleston, Montana has also been developed. This method is based on discriminants in the 0.67 micron and 0.79 micron region of the spectrum.

Birnie, R. W.↗