Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “machine learning, Random Forest”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

X-ray Spectra and Multiwavelength Machine Learning Classification for Likely Counterparts toFermi3FGL Unassociated Sources

We conduct X-ray spectral fits on 184 likely counterparts to Fermi-LAT 3FGL unassociated sources. Characterization and classification of these sources allows for more complete population studies of the high-energy sky. Most of these X-ray spectra are well fit by an absorbed power law model, as expected for a population dominated by blazars and pulsars. A small subset of 7 X-ray sources ave spectra unlike the power law expected from a blazar or pulsar and may be linked to coincident stars or background emission. We develop a multiwavelength machine learning classifier to categorize unassociated sources into pulsars and blazars using gamma- and X-ray observations. Training a random forest procedure with known pulsars and blazars, we achieve a cross-validated classification accuracy of 98.6%. Applying the random forest routine to the unassociated sources returned 126 likely blazar candidates (defined as P(bzr) ≥ 90%) and 5 likely pulsar candidates (P(bzr) ≤ 10%). Our new X-ray spectral analysis does not drastically alter the random forest classifications of these sources compared to previous works, but it builds a more robust classification scheme and highlights the importance of X-ray spectral fitting. Our procedure can be further expanded with UV, visual, or radio spectral parameters or by measuring flux variability.

Stephen Kerby↗

Machine-z: Rapid Machine-Learned Redshift Indicator for Swift Gamma-Ray Bursts

Studies of high-redshift gamma-ray bursts (GRBs) provide important information about the early Universe such as the rates of stellar collapsars and mergers, the metallicity content, constraints on the re-ionization period, and probes of the Hubble expansion. Rapid selection of high-z candidates from GRB samples reported in real time by dedicated space missions such as Swift is the key to identifying the most distant bursts before the optical afterglow becomes too dim to warrant a good spectrum. Here, we introduce 'machine-z', a redshift prediction algorithm and a 'high-z' classifier for Swift GRBs based on machine learning. Our method relies exclusively on canonical data commonly available within the first few hours after the GRB trigger. Using a sample of 284 bursts with measured redshifts, we trained a randomized ensemble of decision trees (random forest) to perform both regression and classification. Cross-validated performance studies show that the correlation coefficient between machine-z predictions and the true redshift is nearly 0.6. At the same time, our high-z classifier can achieve 80 per cent recall of true high-redshift bursts, while incurring a false positive rate of 20 per cent. With 40 per cent false positive rate the classifier can achieve approximately 100 per cent recall. The most reliable selection of high-redshift GRBs is obtained by combining predictions from both the high-z classifier and the machine-z regressor.

gamma-ray burst: general↗

Landslide Likelihood Prediction using Machine Learning Algorithms

The supply of electricity via power plants is criticalto the operation of many critical infrastructure systems in mod-ern society. Natural hazards can disrupt the power supply, causepower outages that can halt economic growth, and impede emer-gency response until power is restored. The proposed work aimsto predict the landslides likelihood in these critical infrastructurelocations in the Northeastern USA using integrated databases ofexplanatory variables and machine learning algorithms. First,data related to landslides are obtained and merged, includingtopographic, soil moisture, and precipitation-related data. Fiveregression algorithms, namely: Random Forest, Extreme Gradi-ent Boosting (XGBoost), K-Nearest Neighbor regression (KNN),Linear Support Vector Regressor (SVR), and Linear regression,are utilized to predict the landslide probability and evaluatedon the dataset. The accuracy of the models is assessed by usingstatistical metrics such as mean absolute error (MAE), meansquared error (MSE), and root mean squared error (RMSE).The study results show that Random Forest outperformed othermodels with the mutual information feature selection method.It achieved an MSE of 0.0011 with mutual information-basedfeature selection and an MSE of 0.00157 without feature selection.KNN regressor outperformed the other models with an MSEof 0.00139 with correlation-based information selection. Theproposed landslide identification model with Random Forestalgorithm shows outstanding robustness and great potential intackling the landslide likelihood prediction by employing MLalgorithms.

Vasundhara Acharya↗

Imbalanced Multi-layer Cloud Classification with Advanced Baseline Imager (ABI) and CloudSat/CALIPSO Data

Clouds at different altitudes play different roles in Earth’s climate. Comprehensive understanding of overlapping clouds is important for climate and weather prediction. The East Pacific region is where El Ni˜no and La Ni˜na originate and where multi-layer clouds frequently occur. The overlap of clouds at different altitudes in this region increases the classification complexity for cloud-based climatological studies. Unlike prior work in cloud layer classification that assumes single layer or two-layer of clouds, in this work, we consider multi-layer cloud classification with 8 cloud-level classes (clear-sky, high, middle, low, high+middle, high+low, middle+low, high+middle+low). We develop and analyze machine learning models on features extracted from satellite images from the East Pacific regions collected by GOES Advanced Baseline Imager (ABI). These are used to classify CloudSat/CALIPSO observed multi-layer clouds. Due to the imbalanced nature of the data, we investigate the adoption of conventional resampling methods, as well as deep learning methods with data augmentation. In our experiments, we utilize the random forest classifier and Multilayer perceptron classifier with data augmentation methods to reduce the class imbalance during training. With these approaches, we achieve a classification accuracy of 83.6% without exploiting any ancillary information.

machine learning↗

Coronado Ecological Conservation: Assessing Vegetation Change Due to Border Wall Construction and Shifting Social Trails

Species monitoring is essential for mitigating the impacts of plant invasion, such as radical changes in an area’s ecosystem, degraded soil health, increased wildfire severity, landslides, and increased flooding. For this project, NASA DEVELOP partnered with the National Park Service (NPS) to investigate invasive species in disturbed lands: specifically, areas affected by off-trail travel and U.S.-Mexico border construction activities. The team assessed how construction has impacted the distribution of Lehmann’s lovegrass and Russian thistle invasives throughout Coronado National Memorial, AZ from 1986-2022. Using data from Landsat 5 and 8, Sentinel-2, NAIP, and PlanetScope, the team computed NDVI, NDMI, MSAVI2, EVI, and Tasseled Cap Wetness, Brightness, and Greenness transformations as vegetation health indicators to input into various machine learning algorithms. To minimize noise, the team conducted Principal Component Analysis on vegetation indices and spectral bands before running k-means clustering and random forest classification algorithms. Between all datasets, the team found that the median area fully overtaken by invasive plants was 5.37% of the park’s total area in 2022. The NPS will use end products to help increase restoration efforts in disturbed areas with high concentrations of invasive plants, and this project can serve as a jumping off point for future invasive species monitoring. The NPS’s collection of ground data for 2022-2023, in conjunction with future data collection, will notably improve the accuracy of classification models, leading to more precise monitoring of invasive species spread over time.

Coronado National Memorial↗

Coronado Ecological Conservation: Assessing Vegetation Change Due to Border Wall Construction and Shifting Social Trails

Species monitoring is essential in mitigating the impacts of plant invasion, such as radical changes in an area’s ecosystem, degraded soil health, increased wildfire severity, landslides, and increased flooding. NASA DEVELOP partnered with the National Park Service (NPS) to investigate invasive species in disturbed lands: specifically, areas affected by off-trail walking and US-Mexico border construction activities. The team assessed how construction has impacted the distribution of Lehmann’s lovegrass and Russian thistle invasives throughout Coronado National Memorial, AZ from 1986 to 2022. Using data from Landsat 5 and 8, Sentinel-2, the National Agriculture Imagery Program, and PlanetScope, the team computed vegetation indices including the Normalized Difference Vegetation Index, Normalized Difference Moisture Index, Modified Soil Adjusted Vegetation Index 2, Enhanced Vegetation Index, and Tasseled Cap Wetness, Brightness, and Greenness transformations as vegetation health indicators to input into various machine learning algorithms. To minimize noise, the team conducted Principal Component Analysis on the vegetation indices and spectral bands before running k-means++ clustering and random forest classification algorithms. Between all datasets, we found the median area fully overtaken by invasive plants was 5.37% of the park’s total area in 2022. The NPS will use the end products to help increase restoration efforts in disturbed areas with high concentrations of invasive plants. The NPS’s collection of ground data for 2022–2023, in conjunction with future data collection, will notably improve the accuracy of classification models, leading to more precise monitoring of invasive spread over time.

Carson Schuetze↗

Diagnosis of Antarctic Blowing Snow Properties Using MERRA-2 Reanalysis with a Machine Learning Model

This paper presents the work on using a machine learning model to diagnose Antarctic blowing snow (BLSN) properties with the Modern Era Retrospective analysis for Research and Applications v2 (MERRA-2) data. We adopt the random forest classifier for BLSN identification and the random forest regressor for BLSN optical depth and height diagnosis. BLSN properties observed from the Cloud-Aerosol Lidar and Infrared Pathfinder Satellite Observation (CALIPSO) are used as the truth for training the model. Using MERRA-2 fields such as snow age, surface elevation and pressure, temperature, specific humidity, and temperature gradient at the 2m level, and wind speed at the 10m level as input, reasonable results are achieved. Hourly blowing snow property diagnostics are generated with the trained model. Using the year 2010 as an example, it is shown that the Antarctic BLSN frequency is much higher over East than West Antarctica. High frequency months are from April to September, during which BLSN frequency exceeds 20% over East Antarctica. For May 2010, the BLSN snow frequency in the region is as high as 37%. Due to the suppression by strong surface-based inversions, larger values of BLSN height and optical depth are usually limited to the coastal regions, wherein the strength of surface-based inversions is weaker.

Antarctic↗

Remote monitoring of agricultural systems using NDVI time series and machine learning methods: a tool for an adaptive agricultural policy

This study aims to provide accurate information about changes in agricultural systems (AS) using phenological metrics derived from the NDVI time series. Use of such information could help land managers optimize land use choices and monitor the status of agricultural lands, under a variety of environmental and socioeconomic conditions. For this purpose, the Moderate Resolution Imaging Spectroradiometer (MODIS) NDVI data were used to derive phenological metrics over the Oum Er-Rbia basin (central Morocco). Random forest (RF), support vector machine (SVM), and K-nearest neighbor (KNN) classifiers were explored and compared on their ability to classify AS classes over the study area. Four main AS classes have been considered: (1) irrigated annual crop (IAC), (2) irrigated perennial crop (IPC), (3) rainfed area (RA), and (4) fallow (FA). By comparing the accuracy of the three classifiers, the RF method showed the best performance with an overall accuracy of 0.97 and kappa coefficient of 0.96.The RF method was then chosen to examine time variations in AS over a 16-year period (2000–2016). The AS main variations were detected and evaluated for the four AS classes. These variations have been found to be linked well with other indicators of local agricultural land management, as well as the historical agricultural drought changes over the study area. Overall, the results present a tool for decision makers to improve agricultural management and provide a different perspective in understanding the spatiotemporal dynamics of agricultural systems.

Youssef Lebrini↗

Design of Materials with Alchemite

Machine learning models that establish the relationships between materials processing and properties can enable inverse design of materials through active learning. Alchemite is a commercial software that can perform inverse materials design on sparse data. Here we evaluate Alchemite’s performance on a dataset of shape memory alloys and a dataset of heat exchangers compared to baseline random forest models. Alchemite had higher accuracy when making predictions on sparse data and was more accurate or nearly as accurate as random forests on complete datasets while also quantifying uncertainty. The software was also used to suggest processing steps and design parameters to optimize properties and performance; however, physical validation of the suggested design parameters was beyond the scope of this work. Several useful design insights were gained about the impact of the design parameters on properties and performance including the importance of dopant choice and amount for shape memory alloys and the importance of height and weight on the thermal resistance of heat exchangers.

Machine learning↗

Using Machine-Learning to Dynamically Generate Operationally Acceptable Strategic Reroute Options

The newly developed Trajectory Option Set (TOS), a preference-weighted set of alternative routes submitted by flight operators, is a capability in the U.S. traffic flow management system that enables automated trajectory negotiation between flight operators and Air Navigation Service Providers. The objective of this paper is to describe and demonstrate an approach for automatically generating pre-departure and airborne TOSs that have a high probability of operational acceptance. The approach uses hierarchical clustering of historical route data to identify route candidates. The probability of operational acceptance is then estimated using predictors trained on historical flight plan amendment data using supervised machine learning algorithms, allowing the routes with highest probability of operational acceptance to be selected for the TOS. Features used describe historical route usage, difference in flight time and downstream demand to capacity imbalance. A random forest was found to be the best performing algorithm for learning operational acceptability, with a model accuracy of 0.96. The approach is demonstrated for an historical pre-departure flight from Dallas/Fort Worth International Airport to Newark Liberty International Airport.

Evans, Antony↗

Using Machine-Learning to Dynamically Generate Operationally Acceptable Strategic Reroute Options

The newly developed Trajectory Option Set (TOS), a preference-weighted set of alternative routes submitted by flight operators, is a capability in the U.S. traffic flow management system that enables automated trajectory negotiation between flight operators and Air Navigation Service Providers. The objective of this paper is to describe and demonstrate an approach for automatically generating pre-departure and airborne TOSs that have a high probability of operational acceptance. The approach uses hierarchical clustering of historical route data to identify route candidates. The probability of operational acceptance is then estimated using predictors trained on historical flight plan amendment data using supervised machine learning algorithms, allowing the routes with highest probability of operational acceptance to be selected for the TOS. Features used describe historical route usage, difference in flight time and downstream demand to capacity imbalance. A random forest was found to be the best performing algorithm for learning operational acceptability, with a model accuracy of 0.96. The approach is demonstrated for an historical pre-departure flight from Dallas/Fort Worth International Airport to Newark Liberty International Airport.

Evans, Antony↗

Advancements in Blowing Dust Detection at Night via Machine Learning

This presentation introduces operational users to a machine-learning based Dust Probability product developed by the NASA SPoRT program for the application of detecting and monitoring blowing dust plumes at night. Advances in earth observing satellites has improved monitoring and detection of dust both day and night through derived imagery such as the Dust RGB. However, limitations of the RGB at night result in less contrast between dust and land surface features, as seen by the user. A Machine Learning (ML) model has been developed and applied to GOES-16 ABI to overcome this limitation and improve nighttime dust detection. The ML capability is a subset of Artificial Intelligence methods. In this case the Dust ML model was developed using a simple Random Forest (RF) model, typically used to solve classification challenges (or to provide regression type output). The goal was to leverage the strengths of the RF model to learn how to identify blowing dust, and hence, overcome the limitation of a user trying to detect blowing dust within the satellite imagery by eye alone. A brief description of the ML model development will be provided. However, the focus of the presentation will be on the initial user feedback from the assessment of this tool for the 2022 blowing dust events of March through April. During this time several users across the U.S. Southwest collaborated to apply this Dust ML product at night as a complement to the existing Dust RGB in order to determine if it provided greater operational efficiency and value.

Machine Learning↗

Nearest-Neighbor Machine Learning Feature Selection for Interpretation of Microbial Molecular Signatures from Isotope Ratio Mass Spectrometry Data

Mass spectrometry (MS) promises to be a powerful tool for potential biosignature detection during astrobiological missions on ocean worlds in our solar system. Accurate and generalizable machine learning methods could enhance science return on investment by predicting seawater chemistry and classifying isotopic biosignatures, either as a signature consistent with microbial life (biotic) or as a novelty (unclassified/unique). However, machine learning models are likely to be complex and involve interactions between MS features, making biosignatures difficult to interpret. Feature selection methods provide biological and chemical context that help interpret the mechanisms of machine learning models, but these methods also need the ability to detect complex interactions. Previously, we developed a machine learning feature selection algorithm called nearest-neighbor projected distance regression (NPDR) that has the ability to identify important model features that involve complex interactions and automatically reduce correlation and the dimensionality in a high-dimensional variable space. The standard distance metrics used in NPDR – Manhattan and Euclidean – assume the multivariate data are isotropic, which is often violated in real data due to differences in the covariance between variables. Thus, we extend NPDR to include a random forest distance, and other anisotropic distance metrics, for computing nearest neighbors. We also augment the isotope-ratio MS data with time-series features from the raw MS signal to improve biotic classification. We test NPDR on our novel experimental ocean world seawater analog MS data. We measure isotope fractionations of volatile CO 2 that could be measured in exospheres or plumes. Samples include baseline abiotic conditions using a range of possible seawater chemistry consistent with Europa and Enceladus, and biotic samples that include microbes in these seawaters. We use penalized NPDR with random forest proximity to identify interpretable microbial molecular signatures. We compare features with random forest importance, and we train a classifier that discriminates between biotic and abiotic samples with high accuracy. These ML-trained ocean-world analog MS data could be used to assist in identifying biosignatures during future missions.

geochemistry↗

An Ensemble of Bayesian Neural Networks for Exoplanetary Atmospheric Retrieval

Machine learning (ML) is now used in many areas of astrophysics, from detecting exoplanets in Kepler transit signals to removing telescope systematics. Recent work demonstrated the potential of using ML algorithms for atmospheric retrieval by implementing a random forest (RF) to perform retrievals in seconds that are consistent with the traditional, computationally expensive nested-sampling retrieval method. We expand upon their approach by presenting a new ML model, plan-net, based on an ensemble of Bayesian neural networks (BNNs) that yields more accurate inferences than the RF for the same data set of synthetic transmission spectra. We demonstrate that an ensemble provides greater accuracy and more robust uncertainties than a single model. In addition to being the first to use BNNs for atmospheric retrieval, we also introduce a new loss function for BNNs that learns correlations between the model outputs. Importantly, we show that designing ML models to explicitly incorporate domain-specific knowledge both improves performance and provides additional insight by inferring the covariance of the retrieved atmospheric parameters. We apply plan-net to the Hubble Space Telescope Wide Field Camera 3 transmission spectrum for WASP-12b and retrieve an isothermal temperature and water abundance consistent with the literature. We highlight that our method is flexible and can be expanded to higher resolution spectra and a larger number of atmospheric parameters.

Adam D. Cobb↗

Document Classification Techniques for Aviation Letters of Agreement

Often when working with technical documents, it is helpful to classify them into specific categories. In this paper, we conduct a thorough review of natural language processing techniques to perform this classification task on Letters of Agreement (LOAs), technical aviation documents outlining rules for utilizing US airspace. We evaluate multiple techniques, including Transfer Learning, for representing the text in the documents as embeddings: unigram and bigram Term Frequency Inverse Document Frequency (TFIDF), Word2Vec, Doc2Vec, GloVe and RoBERTa. We investigate a wide range of classification models: K-Nearest Neighbors, Random Forest, Support Vector Machines (SVM), Logistic Regression, Naive Bayes, Feed-Forward Neural Network, Convolutional Neural Networks (CNNs) and Long-Short Term Memory (LSTM). By comparing the different methods, we found the best overall approach for our task was to use unigram TFIDF representations with SVM while also gaining insight into how the other methodologies performed on a small technical datasets.

Aayushi Batra↗

Machine Learning for Predicting Team Functioning in HERA Missions

Team functioning is integral to success in future long term space exploration missions. Proactively detecting declines in team functioning can mitigate conflict and ensure mission success. This project developed a speech-based artificial intelligence (AI) system that unobtrusively predicts degradation in team functioning, including performance and cohesion, in the Human Exploration Research Analog (HERA) Campaigns 4 and 5. The AI system conducted automated analysis of the prosodic (tone of voice) and linguistic (language content) components of speech, modeling interpersonal dynamics at both the turn-taking and day-wide levels. We investigated team functioning via observing structured interactions (i.e., multi-mission space exploration vehicle-extra vehicular activity [MMSEV-EVA], team interaction battery [TIB]) and unstructured interactions before the MMSEV-EVA task. We developed machine learning models to predict team functioning (objective task accuracy, self reported team efficacy and self reported team cohesion) by analyzing OpenSmile acoustic features, linguistic descriptors extracted via the linguistic inquiry and word count (LIWC) dictionary, and semantic embeddings. In the TIB, static models using logistic regression and random forests were not able to predict task accuracy, but predicted team efficacy and cohesion during both the decision making and relational tasks to a moderate level (60-70%). Majority voting on the individual turns to predict day long team efficacy further increased accuracies (70-80%). Finally, long short-term memory (LSTM) models showed the best performance across all variables (80-91%), including task performance. In the MMSEV-EVA, static models achieved an accuracy of 60% with majority voting, which increased to 80% through the incorporation of mission day as a variable, accounting for the learning effect. A key finding across both tasks was the "team-dependent" nature of these interactions; models achieved much higher accuracy when trained on prior days of the same team's data rather than attempting to generalize across entirely different teams, with even 1-2 days of prior data per team achieving 5-15% improvement over team-independent models. In addition, the incorporation of pre-task data from the same team also improves model performance, e.g., incorporating data from the decision-making task of the TIB, which preceded the relational task, improved the prediction of team efficacy and cohesion during the latter. We compared model performance when trained on machine-generated data compared to data that had been further corrected by human annotators. Overall, models trained on human-corrected data exhibited a modest improvement in performance, particularly when acoustic features were used. We found no significant correlation between word error rate (WER) and model accuracy (r(55) = -0.08, p = 0.51), but model’s accuracy was significantly higher for medium/high quality transcription (0.74 (SD = 0.48)) compared to the low-quality group (0.64 (SD = 0.36)) (t(63)=2.82, p = 0.006). Based on these, several design recommendation emerge, that could inform Standards at NASA. Models predicting team functioning should incorporate at least one to two days of historical interaction data, include brief pre-task discussions, and explicitly model temporal learning effects, especially for longer operational tasks. Minimum quality standards for automated speech-processing pipelines are needed, given the performance gains observed with manually corrected acoustic data. Finally, systems should leverage both acoustic features and language embeddings in complementary ways, with modality choices and fusion strategies tailored to mission context, task demands, and data quality requirements.

Shrivatsa Mishra↗

Algorithmic Classification of Raman Spectra Biosignatures: Improving Life Detection Confidence

“Agnostic” biosignatures – indicators of life (or the absence of life), independent of a particular biochemistry – are increasingly considered a high standard for life detection. The Ladder of Life Detection (2018) called for investigating how combinations of independent and different potential biosignatures affect confidence. To address this gap, statistical classification of elemental abundances, isotopic fractionation, and reflectance spectroscopy (VNIR) has been implemented. Raman spectroscopy, highly desirable due to its wide availability, has the potential to improve this predictive power. This work implemented biosignature classification algorithms on Raman data alone, in preparation for combination with the other data types. Raman spectroscopy data was collected from published databases and papers as part of a manually curated dataset of “indicative” and “non-indicative of life” samples. These currently include 61 non-indicative samples (meteorites, magnetite); 3 indicative living samples (bacteria); 20 indicative non-living samples (chalk, bone); and 12 indicative mixed (with non-indicative material) samples (soil, microbial mats). Laboratory work is ongoing to characterize additional samples, particularly a greater breadth of mixed systems. Spectra were interpolated, filtered with the Savitzsky-Golay filter, and de-noised. For a preliminary examination, agnostic features were manually extracted including mean intensity, number of peaks, and mean peak width. Different peak prominences and filtering polynomials were used to refine features. Classification algorithms were implemented: k-nearest neighbors (KNN), logistic regression (LR), linear support vector machines (SVM), random forest (RF), Gaussian naïve bayes (GNB). Lastly, Monte Carlo simulations on 1,000 50%-train-test-splits were used to validate classification performance and feature significance. The preliminary feature set achieved its highest AUC of 0.52 with LR, with no strongly discriminatory features. Work to improve feature extraction, such as through deep learning with back propagation, is planned. In future work, the Raman data will be combined with the other data types, and potentially new data types such as enantiomeric excess. This project was partially supported through the NASA Ames Project EXcellence (APEX) incubator program.

Astrobiology↗

A Comprehensive Machine Learning Study to Classify Precipitation Type over Land from Global Precipitation Measurement Microwave Imager (GPM-GMI) Measurements

Precipitation type is a key parameter used for better retrieval of precipitation characteristics as well as to understand the cloud–convection–precipitation coupling processes. Ice crystals and water droplets inherently exhibit different characteristics in different precipitation regimes (e.g., convection, stratiform), which reflect on satellite remote sensing measurements that help us distinguish them. The Global Precipitation Measurement (GPM) Core Observatory’s microwave imager (GMI) and dual-frequency precipitation radar (DPR) together provide ample information on global precipitation characteristics. As an active sensor, the DPR provides an accurate precipitation type assignment, while passive sensors such as the GMI are traditionally only used for empirical understanding of precipitation regimes. Using collocated precipitation type flags from the DPR as the “truth”, this paper employs machine learning (ML) models to train and test the predictability and accuracy of using passive GMI-only observations together with ancillary information from a reanalysis and GMI surface emissivity retrieval products. Out of six ML models, four simple ones (support vector machine, neural network, random forest, and gradient boosting) and the 1-D convolutional neural network (CNN) model are identified to produce 90–94% prediction accuracy globally for five types of precipitation (convective, stratiform, mixture, no precipitation, and other precipitation), which is much more robust than previous similar effort. One novelty of this work is to introduce data augmentation (subsampling and bootstrapping) to handle extremely unbalanced samples in each category. A careful evaluation of the impact matrices demonstrates that the polarization difference (PD), brightness temperature (Tc) and surface emissivity at high-frequency channels dominate the decision process, which is consistent with the physical understanding of polarized microwave radiative transfer over different surface types, as well as in snow and liquid clouds with different microphysical properties. Furthermore, the view-angle dependency artifact that the DPR’s precipitation flag bears with does not propagate into the conical-viewing GMI retrievals. This work provides a new and promising way for future physics-based ML retrieval algorithm development.

machine learning/artificial intelligence↗