Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Principal component analysis”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10

The luminosity structure and objective classification of galaxies

The luminosity structure of spiral galaxies is studied using the technique of principal component analysis. It is found that approximately 94% of the variation in the luminosity distribution of galaxies can be accounted for by just two principal components. The principal luminosity components may contain valuable information about star formation history or whatever luminosity-regulating process occurs in galaxies. Practically, these principal components provide a new approach for the investigation of the luminosity structures of galaxies and their dependence on other properties. They also serve as an excellent objective classification system for galaxies. We introduce in this paper such a classification scheme and explore its various properties. The new system shows a number of very impressive characteristics. Most important, it can well segregate virtually all the important galactic properties we tested and does so much better than the conventional morphological classification systems. Of particular interest is that some distance-dependent parameters can also be determined to a surprisingly good accuracy; for example, absolute magnitude may be determined to an accuracy of approximately 0.6 mag (yet further improvement is believed to be highly possible). Second, the system is objective, and the classification procedure can be automated to a large degree; also the new system can apply to much smaller and fainter images than do eye-based clasification systems. These properties make the new system suitable for practical application, especially on very large (and deeper) digital image catalogs. Third, the classification is expressed in dimensionless numbers, yet the simple notation bears significant and easily understandable meaning, making it easy and convenient to use. Finally, the new system has another extremely useful feature: it provides a very powerful and convenient platform not only for classification, but also for easily recording, examining, and studying the variations and correlations of galaxy properties-all these may be carried out graphically by using the C-vectors and the C-diagrams introduced in the paper. We wil also give an example to demonstrate the use of the classification system for the study of the internal extinction problem in spiral galaxies.

Han, Mingshen↗

The intermediate line region of QSOs

Our recent statistical investigations of broad UV lines of luminous quasars (QSOs) suggest that the traditional broadline region (BLR) consists of two components-one of width approximately 2000 km/s FWHM with the peak within a few hundred kilometers per second of the systemic redshift, and another very broad component of width greater than or equal to 7000 km/s FWHM and blueshifted by greater than or equal to 1000 km/s. Differences in the relative strengths of these components account for much of the diversity of broad-line profiles, as well as relations among line strength, line width, asymmetry, and peak blueshift. We have suggested that the narrower component arises in a distinct intermediate-line region (ILR) that is an inner extension of the narrow-line region (NLR). We form spectra of the ILR and very broad line region (VBLR) in two complementary ways. First, using a small sample of high-quality spectra, we difference two composite spectra, one with FWHM (sub C IV) approximately 3000 km/s, the other with FWHM (sub C IV) approximately 7000 km/s (essentially a VBLR spectrum)-revealing a narrower line spectrum with FWHM approximately 2000 km/s. Second, we use a principal component analysis of 200 low signal-to-noise ratio spectra from the Large Bright Quasar Survey. The ILR is identified with the first principal component-that component accounting for most of the spectrum-to-spectrum variation. The VBLR spectrum is derived by subtracting this ILR from the mean spectrum. The two methods yield similar results, and the spectra of the ILR and VBLR are very different. Additional support for the existence of two component is the lack of a correlation between the equivalent widths of the ILR and VBLR. We discuss the relationships between the VBLR, the ILR, and the traditional NLR. The ILR line intensity ratios are distinctly different from those of the VBLR, with (relative to C IV lambda 1549) stronger Lyman alpha and weaker N v lambda 1240, lambda 1400 feature, He II lambda 1640, O III lambda 1663, and Al III lambda 1860. Comparison with other active galactic nuclei (AGN) emission-line regions shows that the ILR spectrum tends to be intermediate between that of the VBLR and that of gas more distant from the ionizing continuum, such as the NLR and extended Lyman alpha nebulosity. Comparison of the spectra with results from photoionization models suggests that, compared with the VBLR, the ILR is about 10 times more distant from the ionizing continuum (approximately 1 pc), 100 to 1000 times less dense (10(exp 10)/cu cm), and has a smaller covering factor (less than or approximately 3%, compared with approximately 24% for the VBLR).

Brotherton, M. S.↗

Principal Components as a Data Reduction and Noise Reduction Technique

The potential of principal components as a pipeline data reduction technique for thematic mapper data was assessed and principal components analysis and its transformation as a noise reduction technique was examined. Two primary factors were considered: (1) how might data reduction and noise reduction using the principal components transformation affect the extraction of accurate spectral classifications; and (2) what are the real savings in terms of computer processing and storage costs of using reduced data over the full 7-band TM complement. An area in central Pennsylvania was chosen for a study area. The image data for the project were collected using the Earth Resources Laboratory's thematic mapper simulator (TMS) instrument.

Imhoff, M. L.↗

Global variability of precipitation according to the Tropical Rainfall Measuring Mission

Numerous studies have documented the effect of El Nino-Southern Oscillation (ENSO) on rainfall in many regions of the globe. The question of whether ENSO is the single most important factor in interannual rainfall variability has received less attention, mostly because the kind of data that would be required to make such an assessment were simply not available. Until 1979 the evidence linking El Nino with changes in rainfall around the world came from rain gauges measuring precipitation over land masses and a handful of islands. From 1980 until the launch of the Tropical Rainfall Measuring Mission (TRMM) in November 1997 the remote sensing evidence was confined to ocean rainfall because of the very poor sensitivity of the instruments over land. In this paper we summarize the results of a principal component analysis of TRMM's 60-month (January 1998 to December 2002) global land and ocean remote-sensing record of monthly rainfall accumulations. Contrary to the first principal component of the rainfall itself, the first three indices of the anomaly are most sensitive to precipitation over the ocean rather than over the land. With the help of archived surface station data the first TRMM rain anomaly index is extended back several decades. Comparison of the extended index with the Southern Oscillation Index confirms that the first principal component of the rainfall anomaly is strongly correlated with the ENSO indices.

rainfall↗

Representation of Probability Density Functions from Orbit Determination using the Particle Filter

Statistical orbit determination enables us to obtain estimates of the state and the statistical information of its region of uncertainty. In order to obtain an accurate representation of the probability density function (PDF) that incorporates higher order statistical information, we propose the use of nonlinear estimation methods such as the Particle Filter. The Particle Filter (PF) is capable of providing a PDF representation of the state estimates whose accuracy is dependent on the number of particles or samples used. For this method to be applicable to real case scenarios, we need a way of accurately representing the PDF in a compressed manner with little information loss. Hence we propose using the Independent Component Analysis (ICA) as a non-Gaussian dimensional reduction method that is capable of maintaining higher order statistical information obtained using the PF. Methods such as the Principal Component Analysis (PCA) are based on utilizing up to second order statistics, hence will not suffice in maintaining maximum information content. Both the PCA and the ICA are applied to two scenarios that involve a highly eccentric orbit with a lower apriori uncertainty covariance and a less eccentric orbit with a higher a priori uncertainty covariance, to illustrate the capability of the ICA in relation to the PCA.

Mashiku, Alinda K.↗

Using Machine Learning for Timely Estimates of Ocean Color Information From Hyperspectral Satellite Measurements in the Presence of Clouds, Aerosols, and Sunglint

Retrievals of ocean color from space are important for better understanding of the ocean ecosystem but can be limited under conditions such as clouds, aerosols, and sunglint. Many ocean color algorithms use a few selected spectral bands to perform an atmospheric correction and then derive the upwelling radiance from the ocean. The limitations in the atmospheric correction under certain conditions lead to many gaps in daily spatial coverage of ocean color retrievals. To address these limitations, we introduce a new approach that uses machine learning to estimate ocean color from top of atmosphere radiances or reflectance measurements. In this approach, a principal component analysis is used to decompose the hyperspectral measurements into spectral features that describe the scattering and absorption of the atmosphere and the underlying surface. The coefficients of the principal components are then used to train a neural network to predict ocean color properties derived from the MODIS atmospheric correction algorithm. This machine learning approach is independent of a priori information and does not rely on any radiative transfer modeling. We apply the approach to two hyperspectral UV/VIS instruments, the ozone monitoring instrument (OMI) and the TROPOspheric Monitoring Instrument (TROPOMI), using measurements from 320–500 nm to show that it can be used to reproduce ocean color properties in less-than-ideal conditions. This machine learning approach complements the current atmospheric correction ocean color retrievals by filling in the gaps resulting from cloud, aerosol, and sunglint contamination. This method can be applied to the future hyperspectral Ocean Color Instrument (OCI), which will be onboard NASA’s Plankton, Aerosol Cloud, ocean Ecosystem (PACE) ocean color satellite set to launch in 2024.

Ocean color↗

Clustering High-dimensional Toxicogenomics Data with Rare Signals

Toxicogenomics studies the gene and protein activities to drug treatments or toxic exposures. As the drugs and genes are numerous, toxicogenomics data are naturally high dimensional, with dimension sizes up to millions. In addition, the distribution of toxicogenomics data is oftentimes skewed, and they contain rare but important signals representing a cell or organism’s response to toxicity. The combination of high dimension and extremely skewed distribution of toxicogenomics data makes clustering analysis extremely challenging.We present our study of clustering toxicogenomics data using classical approaches such as principal component analysis as well as deep learning approaches such as auto-encoders. Our experiments show that these approaches fail to preserve rare signals and produce high-quality clusters. We then explore augmenting matrix factorization with deep learning techniques such as attention mechanism to produce latent representations for clustering. Our technique is able to better preserve rare signals after dimensionality reduction than prior approaches. Furthermore, we combine our augmented matrix factorization with a mechanism similar to autoencoder to balance separable clusters and low regeneration errors. Our experiments demonstrate better clustering with our proposed approach.

Cong, Guojing↗

Preliminary Comparisons of the Information Content and Utility of TM Versus MSS Data

Some preliminary indications were provided as to the relative merits of actual TM data versus MSS data for land cover mapping related applications. Three analyses were designed which had sensitivity to the differences in spectral, spatial and radiometric parameters between the TM and MSS. In the water body analysis, a primarily spatially related test, the detectability of small uniform targets was examined. The principal components analysis, an examination of the inherent dimensionality of the data, was more spectrally and radiometrically related. The spectral clustering analysis, also heavily spectrally and radiometrically influenced, provided information on the types of targets separable on TM versus MSS data. These analyses were to be conducted with simultaneously collected LANDSAT-4 complete TM (7 band) and MSS (4 band) data. In actuality, 4-band TM data, and archived LANDSAT-2 MSS data of the same area were used.

Markham, B. L.↗

Pattern recognition of clouds and ice in polar regions

The study is based on AVHRR imagery and results from Landsat high-spatial-resolution scenes. Among the textual features investigated are the gray level difference vector (GLDV), and sum and difference histogram (SADH) approaches as well as gray level run length, spatial-coherence, and spectral-histogram measures. The traditional stepwise discriminant analysis and neural-network analysis are used for the identification of 20 Arctic surface and cloud classes. A principal-component analysis and hybrid architecture employing a modularized competitive learning layer are utilized. It is pointed out that the cloud-classification accuracy comparable to that of back-propagation could be achieved with a training time two orders of magnitude faster.

Welch, R. M.↗

A Principal Component and Machine Learning Approach to Spatially Gap Fill Hyperspectral Ocean Color Satellite Retrievals

Retrievals of ocean color properties from space are important for monitoring the health of the ocean ecosystem but such retrievals tend to be limited in spatial coverage due to conditions such as clouds, aerosols, and sun glint. Gap filling of ocean color retrievals is typically performed by combining retrievals from multiple satellites or temporally averaging multiple days of retrievals but despite these techniques large gaps still exist posing challenges for near real time monitoring of events like harmful algae blooms. To address these limitations, we propose a spatial gap filling approach using machine learning to learn how to perform an atmospheric correction under challenging retrieval conditions. In this approach a principal component analysis is used to decompose the hyperspectral measurements into spectral features that describe the scattering and absorption of the atmosphere as well as the underlying surface. The coefficients of the principal components are then used to train a neural network to predict ocean color properties derived from a standard ocean color algorithm such as the MODIS atmospheric correction algorithm. This machine learning approach is independent of a priori information and does not rely on any radiative transfer modeling. We apply the approach to two hyperspectral UV/VIS instruments, the Ozone Monitoring Instrument (OMI) and TROPOspheric Monitoring Instrument (TROPOMI) to show that it can be used to estimate ocean color properties such as chlorophyll, remote sensing reflectance, and fluorescence line height. This method could be used as a gap-filling technique for the future Ocean Color Instrument (OCI) which will be onboard NASA's Plankton, Aerosol Cloud, ocean Ecosystem (PACE) ocean color satellite to provide additional information for monitoring the health of our global oceans. Additionally, it could be applied to the geostationary satellite Tropospheric Emissions: Monitoring of Pollution (TEMPO) to better understand diurnal variability in ocean ecology.

MODIS atmospheric correction algorithm↗

A bi-level data-driven framework for fault-detection and diagnosis of HVAC systems

Long-term operation of heating, ventilation, and air conditioning (HVAC) systems will eventually lead to a range of HVAC system failures, resulting in excessive energy consumption and maintenance costs. Here, to avoid HVAC malfunctioning, fault detection diagnostic (FDD) is utilized as a common practice. Machine learning methods have lately received considerable interest for FDD analysis of HVAC systems due to their high detection accuracy. Meanwhile, HVAC malfunctions are regarded as rare occurrences, hence normal operating data samples are much more accessible than data samples in faulty and malfunctioning conditions. The dominating frequency of normal operation in HVAC datasets has also led to heavily biased classification algorithms within the literature. Moreover, the focus of previous literature has been on increasing the accuracy of the models which leads to a high number of false positives (misleading alarms) in the system. In order to enhance the performance of diagnostic procedures and fill the mentioned gaps, this study proposes a novel data-driven framework. A bi-level machine learning framework is developed for diagnosing faults in air handling units (AHUs) and rooftop units (RTUs) based on principal component analysis (PCA), time series anomaly detection, and random forest (RF). It is shown that PCA can reduce the dataset dimension with one principal component accounting for 95% of data variance. Also, the random forest could classify the faults with 89% precision for single-zone AHU, 85% precision for RTU, and 79% for multi-zone AHU. By proposing this framework, three persistent challenges are addressed: (I) minimizing false positives; (II) accounting for data imbalance; and (III) normal condition monitoring of equipment.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

DEEPEN Global Standardized Categorical Exploration Datasets for Magmatic Plays

DEEPEN stands for DE-risking Exploration of geothermal Plays in magmatic ENvironments. As part of the development of the DEEPEN 3D play fairway analysis (PFA) methodology for magmatic plays (conventional hydrothermal, superhot EGS, and supercritical), weights needed to be developed for use in the weighted sum of the different favorability index models produced from geoscientific exploration datasets. This was done using two different approaches: one based on expert opinions, and one based on statistical learning. This GDR submission includes the datasets used to produce the statistical learning-based weights. While expert opinions allow us to include more nuanced information in the weights, expert opinions are subject to human bias. Data-centric or statistical approaches help to overcome these potential human biases by focusing on and drawing conclusions from the data alone. The drawback is that, to apply these types of approaches, a dataset is needed. Therefore, we attempted to build comprehensive standardized datasets mapping anomalies in each exploration dataset to each component of each play. This data was gathered through a literature review focused on magmatic hydrothermal plays along with well-characterized areas where superhot or supercritical conditions are thought to exist. Datasets were assembled for all three play types, but the hydrothermal dataset is the least complete due to its relatively low priority. For each known or assumed resource, the dataset states what anomaly in each exploration dataset is associated with each component of the system. The data is only a semi-quantitative, where values are either high, medium, or low, relative to background levels. In addition, the dataset has significant gaps, as not every possible exploration dataset has been collected and analyzed at every known or suspected geothermal resource area, in the context of all possible play types. The following training sites were used to assemble this dataset: - Conventional magmatic hydrothermal: Akutan (from AK PFA), Oregon Cascades PFA, Glass Buttes OR, Mauna Kea (from HI PFA), Lanai (from HI PFA), Mt St Helens Shear Zone (from WA PFA), Wind River Valley (From WA PFA), Mount Baker (from WA PFA). - Superhot EGS: Newberry (EGS demonstration project), Coso (EGS demonstration project), Geysers (EGS demonstration project), Eastern Snake River Plain (EGS demonstration project), Utah FORGE, Larderello, Kakkonda, Taupo Volcanic Zone, Acoculco, Krafla. - Supercritical: Coso, Geysers, Salton Sea, Larderello, Los Humeros, Taupo Volcanic Zone, Krafla, Reyjanes, Hengill. **Disclaimer: Treat the supercritical fluid anomalies with skepticism. They are based on assumptions due to the general lack of confirmed supercritical fluid encounters and samples at the sites included in this dataset, at the time of assembling the dataset. The main assumption was that the supercritical fluid in a given geothermal system has shared properties with the hydrothermal fluid, which may not be the case in reality. Once the datasets were assembled, principal component analysis (PCA) was applied to each. PCA is an unsupervised statistical learning technique, meaning that labels are not required on the data, that summarized the directions of variance in the data. This approach was chosen because our labels are not certain, i.e., we do not know with 100% confidence that superhot resources exist at all the assumed positive areas. We also do not have data for any known non-geothermal areas, meaning that it would be challenging to apply a supervised learning technique. In order to generate weights from the PCA, an analysis of the PCA loading values was conducted. PCA loading values represent how much a feature is contributing to each principal component, and therefore the overall variance in the data.

15 GEOTHERMAL ENERGY↗

Automated Analysis, Classification, and Display of Waveforms

A computer program partly automates the analysis, classification, and display of waveforms represented by digital samples. In the original application for which the program was developed, the raw waveform data to be analyzed by the program are acquired from space-shuttle auxiliary power units (APUs) at a sampling rate of 100 Hz. The program could also be modified for application to other waveforms -- for example, electrocardiograms. The program begins by performing principal-component analysis (PCA) of 50 normal-mode APU waveforms. Each waveform is segmented. A covariance matrix is formed by use of the segmented waveforms. Three eigenvectors corresponding to three principal components are calculated. To generate features, each waveform is then projected onto the eigenvectors. These features are displayed on a three-dimensional diagram, facilitating the visualization of the trend of APU operations.

Kwan, Chiman↗

Towards multi-fidelity deep learning of wind turbine wakes

We report engineering wake models that accurately predict wake in a computationally efficient manner are very important for tasks such as layout optimization and control of wind farms. In this paper, we explore an application of deep learning (DL) to learn the wake model from hierarchies of physics-based approaches ranging from analytical models to an approximate form of the Reynolds-averaged Navier-Stokes equations. We first illustrate the application of principal component analysis to obtain a lower-dimensional representation that allows a computationally tractable training and deployment of DL models. Then, the DL model is trained to learn the mapping from input parameter space to the principal components, which are then used to reconstruct the three-dimensional flow field. Additionally, we investigate a composite framework consisting of two neural networks to learn the correlation between low- and high-fidelity data with Gauss and curl models treated as proxies for low- and high-fidelity models, respectively. The prediction from both DL models matches well with the high-fidelity data with a maximum relative percentage error for the kinetic energy flux of <1%. This work opens up possibilities for data-efficient construction of surrogate models for wake prediction that can be used to study the influence of wind speed and yaw angles on wind farm power production.

17 WIND ENERGY↗

Constraining the baryonic feedback with cosmic shear using the DES Year-3 small-scale measurements

ABSTRACT We use the small scales of the Dark Energy Survey (DES) Year-3 cosmic shear measurements, which are excluded from the DES Year-3 cosmological analysis, to constrain the baryonic feedback. To model the baryonic feedback, we adopt a baryonic correction model and use the numerical package baccoemu to accelerate the evaluation of the baryonic non-linear matter power spectrum. We design our analysis pipeline to focus on the constraints of the baryonic suppression effects, utilizing the implication given by a principal component analysis on the Fisher forecasts. Our constraint on the baryonic effects can then be used to better model and ameliorate the effects of baryons in producing cosmological constraints from the next-generation large-scale structure surveys. We detect the baryonic suppression on the cosmic shear measurements with a ∼2σ significance. The characteristic halo mass for which half of the gas is ejected by baryonic feedback is constrained to be $M_c \gt 10^{13.2} \, h^{-1} \, \mathrm{M}_{\odot }$ (95 per cent C.L.). The best-fitting baryonic suppression is $\sim 5{{\ \rm per\ cent}}$ at $k=1.0 \, {\rm Mpc}\ h^{-1}$ and $\sim 15{{\ \rm per\ cent}}$ at $k=5.0 \, {\rm Mpc} \ h^{-1}$. Our findings are robust with respect to the assumptions about the cosmological parameters, specifics of the baryonic model, and intrinsic alignments.

79 ASTRONOMY AND ASTROPHYSICS↗

Data Analysis for the SOLIS Vector Spectromagnetograph

The National Solar Observatory's SOLIS Vector Spectromagnetograph (VSM), which will produce three or more full-disk maps of the Sun's photospheric vector magnetic field every day for at least one solar magnetic cycle, is in the final stages of assembly. Initial observations, including cross-calibration with the current NASA/NSO spectromagnetograph (SPM) will soon be carried out at a test site in Tucson. This paper discusses data analysis techniques for reducing the raw data, calculation of line-of-sight magnetograms and both quick-look and high-precision inference of vector fields from Stokes spectral profiles. Existing SPM algorithms, suitably modified to accomodate the cameras, scanning pattern, and polarization calibration optics for the VSM, will be used to "clean" the raw data and to process line-of-sight, magnetograms. A recent. version of the High Altitude Observatory Milne-Eddington (HAO-ME) inversion code (Skumanich and Lites; 1987, 11)J 322, p. 473) will he used for high-precision vector fields since the algorithm has been extensively tested, is well understood, and is fast enough to complete data analysis within 24 hours of data acquisition. The simplified inversion algorithm of Auer, Heasley. arid House (1977, Sol. Phys. 55, p. 47) forms the initial guess for this version of the HAO-ME code and will be used for quick-look vector analysis of VSM data since its performance on simulated Stokes profiles is better than other candidate methods. Improvements (e.g., principal components analysis or neural networks) are under consideration and will be straightforward to implement. However, current resources are sufficient to store the original Stokes profiles only long enough for high-precision analysis. Retrospective reduction of Stokes data with improved methods will not be possible, and modifications will only be introduced when the advantages of doing so are compelling enough to justify discontinuity in the long-term data stream.

Jones, Harrison P.↗

Sensor selection and tool wear prediction with data‐driven models for precision machining

Abstract Estimation of tool wear in precision machining is vital in the traditional subtractive machining industry to reduce processing cost, improve manufacturing efficiency and product quality. In this vein, fusion of time and frequency‐domain features of commonly sensed signals can provide an early indication of tool wear and improve its prediction accuracy for prognostics and health management. This paper presents a data‐driven methodology and a complete tool chain for the inference of precision machining tool wear from fused machine measurements, such as cutting force, power, audio and vibration signals, and quantify the usefulness of each measurement. Indicators of tool wear are extracted from time‐domain signal statistics, frequency‐domain analysis, and time‐frequency domain analysis. Correlation coefficients between the extracted features (indicators) and the tool wear are used to select the most informative features. Principal Component Analysis and Partial Least‐Squares are used to reduce the dimensionality of the feature space. Regression models, including linear regression, support vector regression, Decision tree regression, neural network regression and Gaussian process regression, are used to predict the tool wear using data from a Haas milling machine performing spiral boss face milling. The performance of the regression models based on subsets of sensors validates the preliminary estimates about the saliency of the sensors. The experimental results show that the proposed methods can predict the machine tool wear precisely, with readily available sensor measurements. Neural network and Gaussian process regression were able to achieve good estimates of tool wear at different machine operating conditions. The most informative signal in predicting tool wear was shown to be the vibration signal. Time‐frequency domain features were the most informative features among the combination of features of three domains. In addition, using partial least squares components extracted from the original features of signals led to higher prediction accuracy.

Han, Seulki↗

Seeking regularity from irregularity: unveiling the synthesis–nanomorphology relationships of heterogeneous nanomaterials using unsupervised machine learning

Nanoscale morphology of functional materials determines their chemical and physical properties. However, despite increasing use of transmission electron microscopy (TEM) to directly image nanomorphology, it remains challenging to quantify the information embedded in TEM data sets, and to use nanomorphology to link synthesis and processing conditions to properties. We develop an automated, descriptor-free analysis workflow for TEM data that utilizes convolutional neural networks and unsupervised learning to quantify and classify nanomorphology, and thereby reveal synthesis–nanomorphology relationships in three different systems. While TEM records nanomorphology readily in two-dimensional (2D) images or three-dimensional (3D) tomograms, we advance the analysis of these images by identifying and applying a universal shape fingerprint function to characterize nanomorphology. After dimensionality reduction through principal component analysis, this function then serves as the input for morphology grouping through unsupervised learning. We demonstrate the wide applicability of our workflow to both 2D and 3D TEM data sets, and to both inorganic and organic nanomaterials, including tetrahedral gold nanoparticles mixed with irregularly shaped impurities, hybrid polymer-patched gold nanoprisms, and polyamide membranes with irregular and heterogeneous 3D crumple structures. In each of these systems, unsupervised nanomorphology grouping identifies both the diversity and the similarity of the nanomaterial across different synthesis conditions, revealing how synthetic parameters guide nanomorphology development. Our work opens possibilities for enhancing synthesis of nanomaterials through artificial intelligence and for understanding and controlling complex nanomorphology, both for 2D systems and in the far less explored case of 3D structures, such as those with embedded voids or hidden interfaces.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗