Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Traditional Machine Learning Models”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

Machine Learning Global Simulation of Nonlocal Gravity Wave Propagation

Global climate models typically operate at a grid resolution of hundreds of kilometers and fail to resolve atmospheric mesoscale processes, e.g., clouds, precipitation, and gravity waves (GWs).Model representation of these processes and their sources is essential to the global circulation and planetary energy budget, but subgrid scale contributions from these processes are often only approximately represented in models using parameterizations. These parameterizations are subject to approximations and idealizations, which limit their capability and accuracy. The most drastic of these approximations is the “single-column approximation” which completely neglects the horizontal evolution of these processes, resulting in key biases in current climate models. With a focus on atmospheric GWs, we present the first-ever global simulation of atmospheric GW fluxes using machine learning (ML) models trained on the WINDSET dataset to emulate global GW emulation in the atmosphere, as an alternative to traditional single-column parameterizations. Using an Attention U-Net-based architecture trained on globally resolved GW momentum fluxes, we illustrate the importance and effectiveness of global nonlocality, when simulating GWs using data-driven schemes.

Aman Gupta↗

Sentiment analysis of the United States public support of nuclear power on social media using large language models

This study utilized large language models (LLMs) to analyze public sentiment in the United States (US) regarding nuclear power on social media, focusing on X/Twitter, considering climate change challenges and advancements in nuclear power technology. Approximately, 1.26 million nuclear tweets from 2008–2023 were examined to fine-tune LLMs for sentiment classification. We found the crucial role of accurate data labeling for model performance, with potential implications for a 15% improvement, achieved through high-confidence labels. LLMs demonstrated better performance compared to traditional machine learning classifiers, with reduced susceptibility to overfitting and up to 96% classification accuracy. LLMs are used to segment the US public tweets into policy and energy-related categories, revealing that 68% are politically themed. Policy tweets tended to convey negative sentiment, often reflecting opposing political perspectives and focusing on nuclear deals and international relations. Energy-related tweets covered diverse topics with predominantly neutral to positive sentiment, indicating broad support for nuclear power in 48 out of 50 US states. The US public positive sentiments toward nuclear power stemmed from its high power density, reliability regardless of weather conditions, environmental benefits, application versatility, and recent innovations and advancements in both fission and fusion technologies. Negative sentiments primarily focused on waste management, high capital costs, and safety concerns. The neutral campaign highlighted global nuclear facts and advancements, with varying tones leaning towards positivity or negativity. An interesting neutral theme was the advocacy for the combined use of renewable and nuclear energy to attain net-zero goals.

Energy & Fuels↗

Utilizing IBM Spectrum LSF Simulator to Understand the Impacts of Adding AI Workloads to Capability Supercomputing

Machine Learning and Artificial Intelligence has been identified as an emerging priority science area within the Department of Energy. Large scale accelerator based supercomputers like Summit, while traditionally employed for modeling and simulation, provide architectures that are suitable for accelerating the ML/AI workloads at scale. With the release of Summit in 2018, there was an increase in the number of ML/AI based projects seeking time on the machine. It quickly became apparent that the allocations and job runtimes for this workload deviated from traditional large scale modeling and simulation. Accommodating this new workload requires understanding the impacts to traditional large scale modeling and simulation.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Implicit learning of convective organization explains precipitation stochasticity

Accurate prediction of precipitation intensity is crucial for both human and natural systems, especially in a warming climate more prone to extreme precipitation. Yet, climate models fail to accurately predict precipitation intensity, particularly extremes. One missing piece of information in traditional climate model parameterizations is subgrid-scale cloud structure and organization, which affects precipitation intensity and stochasticity at coarse resolution. Here, using global storm-resolving simulations and machine learning, we show that, by implicitly learning subgrid organization, we can accurately predict precipitation variability and stochasticity with a low-dimensional set of latent variables. Using a neural network to parameterize coarse-grained precipitation, we find that the overall behavior of precipitation is reasonably predictable using large-scale quantities only; however, the neural network cannot predict the variability of precipitation (R 2 ~ 0.45) and underestimates precipitation extremes. The performance is significantly improved when the network is informed by our organization metric, correctly predicting precipitation extremes and spatial variability (R 2 ~ 0.9). The organization metric is implicitly learned by training the algorithm on a high-resolution precipitable water field, encoding the degree of subgrid organization. The organization metric shows large hysteresis, emphasizing the role of memory created by subgrid-scale structures. We demonstrate that this organization metric can be predicted as a simple memory process from information available at the previous time steps. These findings stress the role of organization and memory in accurate prediction of precipitation intensity and extremes and the necessity of parameterizing subgrid-scale convective organization in climate models to better project future changes of water cycle and extremes.

54 ENVIRONMENTAL SCIENCES↗

GOOML Big Kahuna Forecast Modeling and Genetic Optimization Files

This submission includes example files associated with the Geothermal Operational Optimization using Machine Learning (GOOML) Big Kahuna fictional power plant, which uses synthetic data to model a fictional power plant. A forecast was produced using the GOOML data model framework and fictional input data, and a genetic optimization is included which determines optimal flash plant parameters. The inputs and outputs associated with the forecast and genetic optimization are included. The input and output files consist of data, configuration files, and plots. A link to the Physics-Guided Neural Networks (phygnn) GitHub repository is also included, which augments a traditional neural network loss function with a generic loss term that can be used to guide the neural network to learn physical or theoretical constraints. phygnn is used by the GOOML framework to help integrate its machine learning models into the relevant physics and engineering applications. Note that the data included in this submission are intended to provide a demonstration of GOOML's capabilities. Additional files that have not been released to the public are needed for users to run these models and reproduce these results. Units can be found in the readme data resource.

15 GEOTHERMAL ENERGY↗

Machine Learning Interatomic Potentials for Modeling Framework Flexibility and Water Uptake in NbOFFIVE-1-Ni Metal–Organic Framework

Metal–organic frameworks (MOFs), with their distinctive porous structures and tunable chemical properties, have shown immense promise in the separation and storage of gases. Currently, the accurate simulation of their adsorptive properties remains challenging, especially for systems where the molecules fit very tightly into the pores. Traditional simulation methods often approximate the frameworks as rigid and do not account for the framework flexibility seen in materials such as NbOFFIVE-1-Ni. First-principles molecular dynamics (FPMD) simulations offer the desired accuracy in modeling this flexibility but are limited by their extensive computational demands, rendering them impractical for long simulations. Conversely, classical force field-based simulations offer computational efficiency but lack the necessary accuracy. Here, to break this accuracy-efficiency trade-off, we have developed machine learning interatomic potentials trained on energies and forces from FPMD to model the framework flexibility of NbOFFIVE-1-Ni in the presence of water over nanosecond time scales. Furthermore, by integrating MLIP-driven molecular dynamics (MLIP-MD) with grand canonical Monte Carlo (GCMC) simulations, we further incorporated framework flexibility into adsorption predictions, yielding water adsorption isotherms that better align with experimental data compared to those of conventional GCMC simulations. These advances offer new opportunities for the design and optimization of MOFs in gas storage and separation applications.

adsorption↗

Super Resolving Unrolled Neural Networks for Remote Sensing

In remote sensing systems, the capabilities of the system are constrained by the complex interactions between size, weight, and power (SWAP) of potential designs. In electro-optical (EO) systems, examples of these critical parameters include the system’s sensitivity and resolution. Those parameters can be increased by ever larger optical apertures and focal planes but at the cost of more SWAP. Multi-image super resolution (MISR) techniques allow resolution to be enhanced via computation rather than more sophisticated optical hardware. These algorithms combine multiple images together into a single, higher resolution image, trading temporal resolution and computation for spatial resolution. Fielded MISR techniques, such as Drizzle, can require several hundred images to create a single super resolved image, implying reduced temporal resolution, increased data acquisition load, and limiting mission applications. Iterative techniques, such as model-based image reconstruction and compressive sensing, have been shown to create super resolved images using fewer images than Drizzle. They do this by posing an optimization problem that balances accuracy between a highly accurate physical model and an image model. In the case of super resolution, the physical model is defined by the relation between low resolution input images and the desired high resolution output image. The image model encodes some assumptions about the super resolved image. These assumptions are meant to suppress reconstruction artifacts that arise due to deterministic physical model error, stochastic measurement noise, and potential undersampling. In practice, the performance of iterative methods are limited by imaging models compatible with optimization. Deep learning-based methods can effectively learn image models of arbitrary complexity, but lack the theoretical explainability and robustness of iterative techniques. Consensus equilibrium (CE) generalizes the iterative techniques beyond optimization, enabling blackbox algorithms such as traditional and neural image denoisers to be used as the image model. CE-based approaches retain much of the explainability and robustness of iterative techniques while allowing the expressiveness of machine learning image models to be used. Additionally, by unrolling iterations of CE with an embedded image denoiser, the image denoiser can be further trained and specialized to the specific application with potentially higher quality reconstructions. Under this project, we demonstrated the feasibility of training an unrolled neural network based upon CE. While we didn’t train one, we showed that the CE process is differentiable and its gradient can be tractably computed. We also explored the usage of a variants of CE akin to generative neural works. Most importantly, we applied the CE framework to a number of problems including non-blind deconvolution, upsampling, single-image super resolution, MISR, event-based sensing, and saturated deconvolution. Our MISR prototype creates high quality reconstructions with an order of magnitude fewer images than previous approaches and, critically, produces these reconstructions fast enough for practical usage.

47 OTHER INSTRUMENTATION↗

Can artificial intelligence and data-driven machine learning models match or even replace process-driven hydrologic models for streamflow simulation?: A case study of four watersheds with different hydro-climatic regions across the CONUS

With recent developments in computational techniques, Data-driven Machine Learning Models (DMLs) have shown great potential in simulating streamflow and capturing the rainfall-runoff relationship in given watersheds, which are traditionally fulfilled by Process-based Hydrologic Models (PHMs). There are debates on whether the DMLs can outperform and possibly replace the classical PHMs for streamflow simulation and river forecasting, but no clear conclusions have been made. This study aims to investigate whether the newer DMLs have any potential in further improving the simulation accuracy of classical PHMs, and vice versa. To do this, we compared a few popular PHMs and DMLs over four watersheds across the Continental US (CONUS) that are associated with different input, climate, and regional conditions. A total of five hydrologic models were chosen, including (1) two classical lumped models, i.e., the Sacramento Soil Moisture Accounting (SAC-SMA) and Xinanjiang (XAJ); (2) one modern distributed model, termed Coupled Routing and Excess Storage (CREST); (3) and two DMLs including an Artificial Neural Networks (ANN) and a deep learning model, termed Long Short Term Memory (LSTM). Our results demonstrated that the DMLs still significantly biased when using the baseline input scenario with the PHMs. However, the DMLs fed with delayed input scenarios had great potential and can reach high simulation accuracy. The DMLs, especially the ANN, outperformed other employed models under the rainfall-runoff relationship in which rainfall dominantly drives. Furthermore, the DMLs also showed better performance in the high-flow regime, while the PHMs had a better performance for the low-flow regime, implying both PHMs and DMLs have their own merits and are worthy of joint development. In general, our study indicated a great potential of using DMLs to simulate streamflow, but further studies are still needed to verify the transferability and scalability of DMLs in large-scale experiments, such as the Distributed Model Intercomparison Projects 1&2 conducted by National Weather Services but to compare modern DMLs and PHMs.

58 GEOSCIENCES↗

A Perspective on Traditional and Data Driven Electrochemical Modeling and Analysis

To understand the behavior of electrochemical systems, we need to reduce the dimensionality of the measured current-voltage-time (I-V-t) data by fitting models, thus enabling us to analyze and compare the governing physics. Traditionally, the process for this is an 'expert first' approach: defining the model and its explicit assumptions based on inductive reasoning or empirical observation, fitting small portions of the I-V-t data where assumptions are most valid or carefully designing experiments to enforce key assumptions, and then interpreting the model parameters. However, modern data-driven methods enable a new paradigm: a 'data first' approach, where the latent behaviors governing the system's measured response are identified directly using machine-learning models that optimize both model structure and parameters from the I-V-t data, guaranteeing that the learned model explains as much of the observed system response as possible. After model identification, the model can then be interrogated by an expert to connect observed behaviors with underlying physics. This talk will review several different types of electrochemical analysis (electrochemical impedance, differential voltage-capacity, electrochemical kinetics) and compare the traditional and data-driven methods for analyzing the data.

42 ENGINEERING↗

Evaluating county-level lung cancer incidence from environmental radiation exposure, PM 2.5 , and other exposures with regression and machine learning models

Characterizing the interplay between exposures shaping the human exposome is vital for uncovering the etiology of complex diseases. For example, cancer risk is modified by a range of multifactorial external environmental exposures. Environmental, socioeconomic, and lifestyle factors all shape lung cancer risk. However, epidemiological studies of radon aimed at identifying populations at high risk for lung cancer often fail to consider multiple exposures simultaneously. For example, moderating factors, such as PM 2.5 , may affect the transport of radon progeny to lung tissue. This ecological analysis leveraged a population-level dataset from the National Cancer Institute’s Surveillance, Epidemiology, and End-Results data (2013–17) to simultaneously investigate the effect of multiple sources of low-dose radiation (gross γ activity and indoor radon) and PM 2.5 on lung cancer incidence rates in the USA. County-level factors (environmental, sociodemographic, lifestyle) were controlled for, and Poisson regression and random forest models were used to assess the association between radon exposure and lung and bronchus cancer incidence rates. Tree-based machine learning (ML) method perform better than traditional regression: Poisson regression: 6.29/7.13 (mean absolute percentage error, MAPE), 12.70/12.77 (root mean square error, RMSE); Poisson random forest regression: 1.22/1.16 (MAPE), 8.01/8.15 (RMSE). The effect of PM 2.5 increased with the concentration of environmental radon, thereby confirming findings from previous studies that investigated the possible synergistic effect of radon and PM 2.5 on health outcomes. In summary, the results demonstrated (1) a need to consider multiple environmental exposures when assessing radon exposure’s association with lung cancer risk, thereby highlighting (1) the importance of an exposomics framework and (2) that employing ML models may capture the complex interplay between environmental exposures and health, as in the case of indoor radon exposure and lung cancer incidence.

63 RADIATION, THERMAL, AND OTHER ENVIRON. POLLUTAN↗

Deep learning of dynamically responsive chemical Hamiltonians with semiempirical quantum mechanics

Conventional machine-learning (ML) models in computational chemistry learn to directly predict molecular properties using quantum chemistry only for reference data. While these heuristic ML methods show quantum-level accuracy with speeds several orders of magnitude faster than traditional quantum chemistry methods, they suffer from poor extensibility and transferability; i.e., their accuracy degrades on large or new chemical systems. Incorporating quantum chemistry frameworks into the ML models directly solves this problem. Here we take the structure of semiempirical quantum mechanics (SEQM) methods to construct dynamically responsive Hamiltonians. SEQM methods use empirical parameters fitted to experimental properties to construct reduced-order Hamiltonians, facilitating much faster calculations than ab initio methods but with compromised accuracy. By replacing these static parameters with machine-learned dynamic values inferred from the local environment, we greatly improve the accuracy of the SEQM methods. Trained on molecular energies and atomic forces, these dynamically generated Hamiltonian parameters show a strong correlation with atomic hybridization and bonding. Trained with only about 60,000 small organic molecular conformers, the resulting model retains interpretability, extensibility, and transferability when testing on much larger chemical systems and predicting various molecular properties. Overall, this work demonstrates the virtues of incorporating physics-based descriptions with ML to develop models that are simultaneously accurate, transferable, and interpretable.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Data for: A hybrid biophysical-machine learning framework for diurnal surface energy flux estimation using proximal sensing

Thermal-based remote sensing of surface energy fluxes has traditionally relied on high spatial resolution satellite data with revisit frequencies on the order of weeks. In this study, we evaluate a biophysics-based analytical surface energy balance model for predicting latent energy (LE) and sensible heat (H) fluxes using proximal sensing observations. The Surface Temperature Initiated Closure (STIC1.2) model has been extensively validated across a wide range of spatial and temporal scales using various satellite-derived thermal datasets. Here we extend this validation by applying STIC at sub-hourly temporal resolution over multiple growing seasons for four distinct agricultural systems. We further develop and evaluate novel STIC variants that incorporate machine learning (ML) techniques to eliminate the need for specific surface energy balance observations, specifically net radiation and soil heat flux, thereby enhancing model applicability in data-sparse settings. The integration of an ML component to estimate surface available energy is shown to have strong predictive performance for both LE (R2 = 0.81-0.94) and H (R2 = 0.46-0.72) across all agricultural systems examined here, demonstrating the potential of hybrid biophysical – machine learning approaches for surface energy balance modeling with minimal data requirements. This study concludes with a novel application of explainable machine learning (exML) to diagnose sources of model error. This exML framework attributes residual prediction errors to both model input variables and environmental drivers not explicitly included in the simulation experiments. This approach provides a new pathway for improving model design and integrating previously overlooked yet influential variables into future model iterations.

Agricultural Sciences↗

Machine learned potential for high-throughput phonon calculations of metal—organic frameworks

Metal–organic frameworks (MOFs) are highly porous and versatile materials studied extensively for applications such as carbon capture and water harvesting. However, computing phonon-mediated properties in MOFs, like thermal expansion and mechanical stability, remains challenging due to the large number of atoms per unit cell, making traditional Density Functional Theory (DFT) methods impractical for high-throughput screening. Recent advances in machine learning potentials have led to foundation atomistic models, such as MACE-MP-0, that accurately predict equilibrium structures but struggle with phonon properties of MOFs. In this work, we developed a workflow for computing phonons in MOFs within the quasi-harmonic approximation with a fine-tuned MACE model, MACE-MP-MOF0. The model was trained on a curated dataset of 127 representative and diverse MOFs. The fine-tuned MACE-MP-MOF0 improves the accuracy of phonon density of states and corrects the imaginary phonon modes of MACE-MP-0, enabling high-throughput phonon calculations with state-of-the-art precision. The model successfully predicts thermal expansion and bulk moduli in agreement with DFT and experimental data for several well-known MOFs. These results highlight the potential of MACE-MP-MOF0 in guiding MOF design for applications in energy storage and thermoelectrics.

Elena, Alin Marin↗

A Hybrid Biophysical‐Machine Learning Framework for Diurnal Surface Energy Flux Estimation Using Proximal Sensing

Thermal infrared-based remote sensing of surface energy fluxes has traditionally relied on high spatial resolution satellite data with revisit frequencies on the order of weeks. In this study, we evaluate a biophysics-based analytical surface energy balance model for predicting latent energy ( LE ) and sensible heat ( H ) fluxes using proximal sensing observations. The Surface Temperature Initiated Closure (STIC1.2) model has been extensively validated across a wide range of spatial and temporal scales using various satellite-derived thermal infrared data sets. Here we extend this validation by applying STIC at sub-hourly temporal resolution over multiple growing seasons for four distinct agricultural systems. We further develop and evaluate novel STIC variants that incorporate machine learning (ML) techniques to eliminate the need for surface energy balance observations, specifically net radiation and soil heat flux, thereby enhancing model applicability in data-sparse settings. The integration of a ML component to estimate surface available energy is shown to have strong predictive performance for both LE (R 2 = 0.81–0.94) and H (R 2 = 0.46–0.72) across all agricultural systems examined here, demonstrating the potential of hybrid biophysical-machine learning approaches for surface energy balance modeling with minimal data requirements. This study concludes with a novel application of explainable machine learning (exML) to diagnose sources of model error. This exML framework attributes residual prediction errors to both model input variables and environmental drivers not explicitly included in the simulation experiments. This approach provides a new pathway for improving model design and integrating previously overlooked yet influential variables into future model iterations.

evapotranspiration↗

Multifidelity Ensemble Kalman Filtering Using Surrogate Models Defined by Theory-Guided Autoencoders

Data assimilation is a Bayesian inference process that obtains an enhanced understanding of a physical system of interest by fusing information from an inexact physics-based model, and from noisy sparse observations of reality. The multifidelity ensemble Kalman filter (MFEnKF) recently developed by the authors combines a full-order physical model and a hierarchy of reduced order surrogate models in order to increase the computational efficiency of data assimilation. The standard MFEnKF uses linear couplings between models, and is statistically optimal in case of Gaussian probability densities. This work extends the MFEnKF into to make use of a broader class of surrogate model such as those based on machine learning methods such as autoencoders non-linear couplings in between the model hierarchies. We identify the right-invertibility property for autoencoders as being a key predictor of success in the forecasting power of autoencoder-based reduced order models. We propose a methodology that allows us to construct reduced order surrogate models that are more accurate than the ones obtained via conventional linear methods. Numerical experiments with the canonical Lorenz'96 model illustrate that nonlinear surrogates perform better than linear projection-based ones in the context of multifidelity ensemble Kalman filtering. We additionality show a large-scale proof-of-concept result with the quasi-geostrophic equations, showing the competitiveness of the method with a traditional reduced order model-based MFEnKF.

97 MATHEMATICS AND COMPUTING↗

Machine learning-enhanced model-based scenario optimization for DIII-D

Abstract Scenario development in tokamaks is an open area of investigation that can be approached in a variety of different ways. Experimental trial and error has been the traditional method, but this required a massive amount of experimental time and resources. As high fidelity predictive models have become available, offline development and testing of proposed scenarios has become an option to reduce the required experimental resources. The use of predictive models also offers the possibility of using a numerical optimization process to find the controllable inputs that most closely achieve the desired plasma state. However, this type of optimization can require as many as hundreds or thousands of predictive simulation cases to converge to a solution; many of the commonly used high fidelity models have high computational burdens, so it is only reasonable to run a handful of predictive simulations. In order to make use of numerical optimization approaches, a compromise needs to be found between model fidelity and computational burden. This compromise can be achieved using neural networks surrogates of high fidelity models that retain nearly the same level of accuracy as the models they are trained to replicate while reducing the computation time by orders of magnitude. In this work, a model-based numerical optimization tool for scenario development is described. The predictive model used by the optimizer includes neural network surrogate models integrated into the fast Control-Oriented Transport simulation framework. This optimization scheme is able to converge to the optimal values of the controllable inputs that produce the target plasma scenario by running thousands of predictive simulations in under an hour without sacrificing too much prediction accuracy.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Group structure selection with random forests

Choosing an appropriate group structure for multigroup transport is far from an exact science. For some applications, one blindly uses a group structure developed years ago by forgotten methods. Furthermore, one sometimes uses the same group structure for a variety of problems, even if the group structure was originally developed with a certain application in mind. In this work, we create optimized group structures with simulated annealing for critical assembly test problems and apply a random forest regressor with bagging to choose the best group structure based on parameters of the different test problems. The optimized group structures were generated using a simulated annealing optimizer for several simple, spherical, and unreflected problems. The optimization was performed to minimize a cost function that included fission rate, absorption rate, leakage, and k{sub eff}. A random forest regressor was then trained on a set of International Criticality Safety Benchmark Evaluation Project inputs and used to select one of these six group structures. The trained machine learning model chose the best group structure 65% of the time, and one of the three best 89% of the time. Furthermore, it decreased the L2 error over all test problems by a factor of 25 when compared to the standard Los Alamos 70-group structure. In other words, the model chose group structure that were far more appropriate for the test problems than the traditional LANL group structure. (authors)

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Creating ground truth for nanocrystal morphology: a fully automated pipeline for unbiased transmission electron microscopy analysis

Control over colloidal nanocrystal morphology (size, size distribution, and shape) is important for tailoring the functionality of individual nanocrystals and their ensemble behavior. Despite this, traditional methods to quantify nanocrystal morphology are laborious. New developments in automated morphology classification will accelerate these analyses but the assessment of machine learning models is limited by human accuracy for ground truth, causing even unsupervised machine learning models to have inherent bias. Herein, we introduce synthetic image rendering to solve the ground truth problem of nanocrystal morphology classification. By simulating 2D images of nanocrystal shapes via a function of high-dimensional parameter space, we trained a convolutional neural network to link unique morphologies to their simulated parameters, defining nanocrystal morphology quantitatively rather than qualitatively. An automated pipeline then processes, quantitatively defines, and classifies nanocrystal morphology from experimental transmission electron microscopy (TEM) images. Using improved computer vision techniques, 42,650 nanocrystals were identified, assessed, and labeled with quantitative parameters, offering a 600-fold improvement in efficiency over best-practice manual measurements. Further, a classification algorithm was trained with a prediction accuracy of 99.5%, which can successfully analyze a range of concave, convex, and irregular nanocrystal shapes. The resulting pipeline was applied to differentiating two syntheses of nominally cuboidal CsPbBr 3 nanocrystals and uniquely classifying binary nickel sulfide nanocrystal phase based on morphology. This pipeline provides a simple, efficient, and unbiased method to quantify nanocrystal morphology and represents a practical route to construct large datasets with an absolute ground truth for training unbiased morphology-based machine learning algorithms.

77 NANOSCIENCE AND NANOTECHNOLOGY↗