Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “random forest regression”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

Next-Level Energy Management in Manufacturing: Facility-Level Energy Digital Twin Framework Based on Machine Learning and Automated Data Collection

This research introduces an energy prediction framework at the facility level supported by automated data collection and machine learning models. It investigates whether reducing the prediction time scale allows for applying more complex machine learning techniques and if those techniques improve the prediction accuracy. The primary advantages of this framework lie in its automation of the energy prediction process and its provision of real-time energy data suitable for use in energy dashboards or digital twins. A sitewide dataset was created by combining 15 min energy and daily production data of five shops—assembly, battery, body (electric), body (gas), and paint—from a globally recognized electric vehicle manufacturer. Various machine learning models were evaluated on daily, weekly, and monthly datasets, including, in increasingly complex order: naïve, simple linear regression, net regularized generalized linear regression, principal component regression, k-nearest neighbor, random forest, and Bayesian regularized neural network. Compared to the current state-of-the-art energy consumption prediction for the industrial facility level, this research investigates more complex models and smaller time intervals for higher accuracy. The findings revealed that the more complex monthly models require a minimum of a year and a half of data to operate, while weekly models demand a year of data to achieve improved accuracy. Daily models can operate with only six months of data but exhibit poor performance due to reduced prediction accuracy of production. Key challenges identified include access to reliable, high-quality energy and production data and the initial demand for human labor.

digital twin↗

Graph-based featurization methods for classifying small molecule compounds

For over a decade, drug-induced liver injury (DILI) has posed significant drawbacks in the synthesis and development of drugs and remains a consequential concern. With finite success within the existing preclinical models, DILI is one of the main causes of drug withdrawal or termination from the market. Particularly, this withdrawal occurs during the late stages of drug development (Kullak-Ublick, 2017). Since DILI is difficult to diagnose and treat, it has become an obstacle in the drug production market that in turn affects clinicians, pharmaceutical companies, and consumers. We propose a method for learning features of DILI-positive drugs based on the graphical relationships and patterns they possess within a network of biological databases. We also train various statistical and machine learning models on these learned features in order to classify the drugs as DILI-positive or negative. Our methods include Random Forest, Neural networks, and logistic regression classification. We utilize labeled DILI-positive and DILI-negative datasets, which were developed by the FDA and the National center for toxicological research, as well as additional literature datasets (Thakkar, 2020) in order to validate our results and assess our featurization and model accuracy.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Prediction of Weather Impacts on Airport Arrival Meter Fix Capacity

This paper introduces a data driven model for predicting airport arrival capacity with a look-ahead time 2-8 hour forecast. The model is suitable for air traffic flow management by explicitly investigating the impact of convective weather on airport arrival meter fix throughput. Estimation of the arrival airport capacity under arrival meter fix flow constraints due to severe weather is an important part of Air Traffic Management (ATM). Airport arrival capacity can be reduced if one or more airport arrival meter fixes are partially or completely blocked by convective weather. When the predicted airport arrival demands exceed the predicted available airport's arrival capacity for a sustained period, Ground Delay Program (GDP) operations will be triggered by ATM system. Serious imbalances between demand and capacity occur most frequently when the airport capacity is severely degraded due to either bad airport terminal surface weather or inclement convective weather around airport arrival fixes. A model that predicts the weather-impacted airport arrival meter fix throughput may help ATM personnel to plan GDP operations more efficiently. This paper identifies the characteristics of air traffic flow across arrival meter fixes at Newark Liberty International Airport (EWR). The proposed approach, based on machine-learning methods, is developed to predict the weather impacted EWR arrival Meter Fix (MF) throughput. Sector forecast coverage is used to envision the weather impact on airport arrival MF flow, and the validation is accomplished by using Convective Weather Avoidance Model (CWAM) 0.5 to 2-hour and Collaborative Convective Forecast Product (CCFP) 4 to 8-hour look-ahead forecast data for the period of April-September in 2014. Furthermore, the regression tree ensemble learning of random forests approach for translating a sector forecast coverage model to an EWR arrival meter fix throughput model is examined. The results suggest that ATM decision makers in charge of MF flow control and GDP planning may benefit from adopting the airport arrival meter capacity prediction models to estimate the inclement weather impacts.

Wang, Yao X.↗

Evaluating Combinations of Sentinel-2 Data and Machine-Learning Algorithms for Mangrove Mapping in West Africa

Creating a national baseline for natural resources, such as mangrove forests, and monitoring them regularly often requires a consistent and robust methodology. With freely available satellite data archives and cloud computing resources, it is now more accessible to conduct such large-scale monitoring and assessment. Yet, few studies examine the reproducibility of such mangrove monitoring frameworks, especially in terms of generating consistent spatial extent. Our objective was to evaluate a combination of image processing approaches to classify mangrove forests along the coast of Senegal and The Gambia. We used freely available global satellite data (Sentinel-2), and cloud computing platform (Google Earth Engine) to run two machine learning algorithms, random forest (RF), and classification and regression trees (CART). We calibrated and validated the algorithms using 800 reference points collected using high-resolution images. We further re-ran 10 iterations for each algorithm, utilizing unique subsets of the initial training data. While all iterations resulted in thematic mangrove maps with over 90% accuracy, the mangrove extent ranges between 827-2807 km2 for Senegal and 245-1271 km2 for The Gambia with one outlier for each country. We further report "Places of Agreement" (PoA) to identify areas where all iterations for both methods agree (506.6 km2 and 129.6 km2 for Senegal and The Gambia, respectively), thus have a high confidence in predicting mangrove extent. While we acknowledge the time- and cost-effectiveness of such methods for the landscape managers, we recommend utilizing them with utmost caution, as well as post-classification on-the-ground checks, especially for decision making.

Mondal, Pinki↗

Predicting Air Traffic Management Initiatives Using Supervised Learning

Terminal Traffic Management Initiatives (TMIs) such as Ground Stops (GS) and Ground Delay Programs (GDP) are implemented to manage excess demand or lowered capacity at an airport. Air Traffic Flow Management (TFM) specialists identify situations such as aviation constraints, current and forecasted weather conditions, airport demand and capacity, and initiate TMIs for safe and orderly movement of air traffic. In this paper, we outline supervised learning techniques that can be used to predict and recommend TMIs at an airport based on current weather and airport conditions. Our research involves building classic Machine Learning (ML) models such as Logistic Regression, K-Nearest Neighbor, Random Forest and XGBoost, as well as Long short-term memory (LSTM) networks. We trained the models on 3-year historical data (weather, airport demand, capacity and TMIs) from Newark (EWR) airport which was selected based on its higher TMI implementation rates and varied weather conditions. Although Random Forest and XGBoost algorithms are able to predict if a TMI is needed or not, they have difficulty in predicting specific program type. For this purpose, we found that LSTM time-series forecasting models performed better as they also learn from past TMI program type sequences. This study also lays down the foundation for advanced modeling techniques and architectures to predict TMIs in advance for future periods. The ability to predict TMIs in advance will be highly beneficial to the traffic controllers and managers as this will help them to prepare for and manage TMIs more efficiently.

Manoj Agrawal↗

Stand validation of lidar forest inventory modeling for a managed southern pine forest

We evaluated area-based approaches (ABAs) to light detection and ranging (lidar) predictions of plot- and stand-level forest attributes (tree count, height, basal area, volume, aboveground biomass, broadleaf/conifer, and diameter at breast height — “diameter”). ABA methods included post-stratification (PS), ordinary least squares (OLSs) regression, k nearest neighbors ( kNN), and random forest (RF). This study was conducted on the Savannah River Site in South Carolina, USA. Plot- and stand-level predictions were validated against fixed-radius 0.04 ha (0.1 acre) plots in 49 ≈2.0 ha (5 acre) stands. Our findings demonstrate that lidar can be incorporated operationally into forest inventory systems to provide stand-level inferences for a wide range of forest attributes. Volume predictions for specific diameter classes, however, often fared poorly (root mean squared error (RMSE) > 100%) for the methods we explored, especially for larger (less common) diameter trees. Stand-level results were consistently better than pixel-level results (10–200+ percentage points). kNN and RF performed similarly and better than OLS and PS, but RF was the most robust to model configurations, while kNN has practical advantages such as simultaneous predictions of many attributes.

Forestry↗

Predicting Elastic Constants of Refractory Complex Concentrated Alloys Using Machine Learning Approach

Refractory complex concentrated alloys (RCCAs) have drawn increasing attention recently owing to their balanced mechanical properties, including excellent creep resistance, ductility, and oxidation resistance. The mechanical and thermal properties of RCCAs are directly linked with the elastic constants. However, it is time consuming and expensive to obtain the elastic constants of RCCAs with conventional trial-and-error experiments. The elastic constants of RCCAs are predicted using a combination of density functional theory simulation data and machine learning (ML) algorithms in this study. The elastic constants of several RCCAs are predicted using the random forest regressor, gradient boosting regressor (GBR), and XGBoost regression models. Based on performance metrics R-squared, mean average error and root mean square error, the GBR model was found to be most promising in predicting the elastic constant of RCCAs among the three ML models. Additionally, GBR model accuracy was verified using the other four RHEAs dataset which was never seen by the GBR model, and reasonable agreements between ML prediction and available results were found. The present findings show that the GBR model can be used to predict the elastic constant of new RHEAs more accurately without performing any expensive computational and experimental work.

36 MATERIALS SCIENCE↗

Applying NIR and MIR spectroscopy for C and soil property prediction in northern cold-region ecosystems. Which approach works better?

Here, developing reliable predictions of soil attributes is necessary to understand northern cold-region climate-soil feedback. Calibration models using near-infrared (NIR) and mid-infrared (MIR) spectroscopy were developed to predict eight commonly measured soil properties for 119 soil samples representing a range of vegetation types, parent materials, and soil types spanning >23° of latitude from southeast Alaska to the Canadian high Arctic. In order to obtain a more accurate prediction, this study compared the performance of linear and non-linear calibration techniques, including lasso regression (Lasso), support vector machine (SVM), random forest (RF) and classic partial least squares (PLS) to predict different soil properties of these soils. Comparing the four models, we noticed that their performance was quite similar for MIR overall, while NIR achieved better results with a PLS model for our dataset. PLS coupled with MIR showed a better performance for soil parameters, such as total organic carbon (TOC), total nitrogen (TN), cation exchange capacity (CEC) and clay (R-squared of 0.9, 0.81, 0.80, and 0.84) when compared with NIR (R-squared of 0.85, 0.72, 0.81 and 0.68). However, using either MIR or NIR spectroscopy, PLS predictions for bulk density (BD) and sand content were not accurate. The variable importance analysis based on the PLS model successfully estimated the relative contribution of wavelengths influencing soil property predictions most. Overall, TOC, TN, CEC and clay mineral predictions are closely related to the occurrence of specific spectral bands in the MIR region. For example, wavelengths at 2978 and 1761 cm -1 for TOC and TN, as well as at 3064 cm -1 for CEC, were selected as the most influential predictor variables. We demonstrated that MIR spectroscopy is a powerful tool for more extensive monitoring in soils of the northern cold climate region; however, NIR could be utilized for rapid estimates when the highest accuracy is not essential.

54 ENVIRONMENTAL SCIENCES↗

The Global LAnd Surface Satellite (GLASS) evapotranspiration product Version 5.0: Algorithm development and preliminary validation

An accurate estimation of spatially and temporally continuous global terrestrial evapotranspiration (ET) is essential in the assessment of surface energy, water and carbon cycles. The Global LAnd Surface Satellite (GLASS) ET product Version 4.0 (v4.0) based on the Bayesian model averaging (BMA) method was generated to estimate global terrestrial ET. However, certain uncertainty for the GLASS ET product v4.0 limits its application. In this study, we introduced the deep neural networks (DNN) merging framework to improve terrestrial ET estimation for GLASS ET product Version 5.0 (v5.0) generation by integrating five satellite-derived ET products [Moderate Resolution Imaging Spectroradiometer (MODIS) ET product (MOD16), Shuttleworth–Wallace dual-source ET product (SW), Priestley–Taylor-based ET product (PT-JPL), modified satellite-based Priestley–Taylor ET product (MS-PT) and simple hybrid ET product (SIM)]. We compared the performance of DNN method against other merging methods, including GLASS ET algorithm v4.0 (BMA), the gradient boosting regression tree (GBRT) method and the random forest (RF) method, based on 195 global eddy covariance (EC) flux towers covering observations from 2000 through 2015. Validations indicated that the DNN had the highest accuracy among four merging methods across different land cover types, yielding the highest average determination coefficients (R 2 , 0.62), root-mean-squared-error (RMSE, 24.1 W/m 2 ) and Kling–Gupta efficiency (KGE, 0.77) with a of 99% confidence interval. Compared with GLASS ET algorithm v4.0, the DNN improved on the R 2 by approximately 7% (p < 0.01) and the KGE by 10%. Based on the DNN, we then generated 8-day GLASS ET product v5.0 globally with a 1 km spatial resolution from 2001 to 2015 driven by GLASS vegetation and surface net radiation (R n ) datasets and Modern-Era Retrospective Analysis for Research and Applications, Version 2 (MERRA2) datasets. Finally, this global terrestrial ET product provides a valuable dataset for monitoring regional and global water resources and environmental changes.

54 ENVIRONMENTAL SCIENCES↗

Use of Satellite, Surface Observations and Numerical Weather Prediction Model Data to Improve Cloud Base Height and Cloud Base Vertical Velocity Estimation

Cloud base height (CBH) and cloud base vertical velocity (CBVV) are important variables that impact the overall climate in a region as they influence the formulation, longevity, and evolution of clouds. Retrieval of both parameters have long used ground instrumentation (e.g., Doppler lidar (DL), ground base radar); however, retrieving CBH from satellites is particularly challenging given that space-based instruments only observe cloud tops. In this manuscript, CBH is retrieved using a multi-linear regression equation, while CBVV used a random forests model. Both retrievals combine satellite and numerical weather prediction data. The satellite data used are the Visible Infrared Imaging Radiometer Suite imagery, while measurements of CBH and CBVV include DL and radiosonde data at the Southern Great Plains (SGP) Atmospheric Radiation Measurement observatory. Data from 83 summer days (May-August) in 2018–2021 featuring cumulus clouds forced by solar heating were examined and used to train the models, with years 2022–2023 used for validation. Various spatial domains were defined with one large (2.4° longitude by 2.0° latitude) SGP domain being split into smaller sections (smallest being 0.99° and 0.61° longitude and latitude respectably). CBH and CBVV values obtained from the DL as compared to the models show root mean square errors between 150 and 200 m, with CBVV values between 0.45 and 1 ms -1 . Finally, it was found that the CBH formulation performs well over all domains, while the CBVV retrievals become less accurate due to more turbulence being introduced into the observations as the number of DL stations decreases in the smaller domains.

54 ENVIRONMENTAL SCIENCES↗

A Comparison of Machine Learning Methods to Forecast Tropospheric Ozone Levels in Delhi

Ground-level ozone is a pollutant that is harmful to urban populations, particularly in developing countries where it is present in significant quantities. It greatly increases the risk of heart and lung diseases and harms agricultural crops. This study hypothesized that, as a secondary pollutant, ground-level ozone is amenable to 24 h forecasting based on measurements of weather conditions and primary pollutants such as nitrogen oxides and volatile organic compounds. We developed software to analyze hourly records of 12 air pollutants and 5 weather variables over the course of one year in Delhi, India. To determine the best predictive model, eight machine learning algorithms were tuned, trained, tested, and compared using cross-validation with hourly data for a full year. The algorithms, ranked by R2 values, were XGBoost (0.61), Random Forest (0.61), K-Nearest Neighbor Regression (0.55), Support Vector Regression (0.48), Decision Trees (0.43), AdaBoost (0.39), and linear regression (0.39). When trained by separate seasons across five years, the predictive capabilities of all models increased, with a maximum R 2 of 0.75 during winter. Bidirectional Long Short-Term Memory was the least accurate model for annual training, but had some of the best predictions for seasonal training. Out of five air quality index categories, the XGBoost model was able to predict the correct category 24 h in advance 90% of the time when trained with full-year data. Separated by season, winter is considerably more predictable (97.3%), followed by post-monsoon (92.8%), monsoon (90.3%), and summer (88.9%). These results show the importance of training machine learning methods with season-specific data sets and comparing a large number of methods for specific applications.

54 ENVIRONMENTAL SCIENCES↗

Document Classification Techniques for Aviation Letters of Agreement

Often when working with technical documents, it is helpful to classify them into specific categories. In this paper, we conduct a thorough review of natural language processing techniques to perform this classification task on Letters of Agreement (LOAs), technical aviation documents outlining rules for utilizing US airspace. We evaluate multiple techniques, including Transfer Learning, for representing the text in the documents as embeddings: unigram and bigram Term Frequency Inverse Document Frequency (TFIDF), Word2Vec, Doc2Vec, GloVe and RoBERTa. We investigate a wide range of classification models: K-Nearest Neighbors, Random Forest, Support Vector Machines (SVM), Logistic Regression, Naive Bayes, Feed-Forward Neural Network, Convolutional Neural Networks (CNNs) and Long-Short Term Memory (LSTM). By comparing the different methods, we found the best overall approach for our task was to use unigram TFIDF representations with SVM while also gaining insight into how the other methodologies performed on a small technical datasets.

Aayushi Batra↗

Document Classification Techniques for Aviation Letters of Agreement

Often when working with historic air traffic management (ATM) documents, it is helpful to classify them into specific categories. In this paper, we conduct a thorough review of natural language processing techniques to perform this classification task on Letters of Agreement (LOAs), technical aviation documents outlining rules for utilizing US airspace. We evaluate multiple techniques for representing the text in the documents as embeddings: unigram and bigram Term Frequency Inverse Document Frequency (TFIDF), Word2Vec, Doc2Vec, GloVe and RoBERTa. We investigate a wide range of classification models: K-Nearest Neighbors, Random Forest, Support Vector Machines (SVM), Logistic Regression, Naive Bayes, Feed-Forward Neural Network, Convolutional Neural Networks (CNNs) and Long-Short Term Memory (LSTM). By comparing the different methods, we found the best overall approach for our task was to use unigram TFIDF representations with SVM while also gaining insight into how the other methodologies performed on a small technical datasets.

ATM↗

Document Classification Techniques for Aviation Letters of Agreement

Often when working with historic air traffic management (ATM) documents, it is helpful to classify them into specific categories. In this paper, we conduct a thorough review of natural language processing techniques to perform this classification task on Letters of Agreement (LOAs), technical aviation documents outlining rules for utilizing US airspace. We evaluate multiple techniques for representing the text in the documents as embeddings: unigram and bigram Term Frequency Inverse Document Frequency (TFIDF), Word2Vec, Doc2Vec, GloVe and RoBERTa. We investigate a wide range of classification models: K-Nearest Neighbors, Random Forest, Support Vector Machines (SVM), Logistic Regression, Naive Bayes, Feed-Forward Neural Network, Convolutional Neural Networks (CNNs) and Long-Short Term Memory (LSTM). By comparing the different methods, we found the best overall approach for our task was to use unigram TFIDF representations with SVM while also gaining insight into how the other methodologies performed on a small technical datasets.

ATM↗

Machine learning models for estimating contamination across different curbside collection strategies

Contaminated recyclables, which are frequently discarded as waste, pose a significant challenge to the implementation of a circular economy. These contaminated recyclables impede the circulation of resources, resulting in higher processing costs at material recovery facilities (MRFs). Over the past few decades, machine learning (ML) models such as linear regression (LR), support vector machine (SVM), and random forest (RF) have evolved to provide new methods for predicting inbound contamination rates in addition to traditional statistical models. In this study, we applied ML models to predict inbound contamination rates using demographic features from 15 counties in the U.S. with different curbside collection strategies. In general, we found that ML models outperformed linear mixed models. Specifically, SVM models had the highest performance (R 2 = 0.75; mean absolute error (MAE) = 0.06), which may be due to their ability to model nonlinear relationships between features and inbound contamination rates. Further, the key predictor was population, with poverty rate being positively correlated and median age negatively correlated with inbound contamination rates. To improve the management of contamination and enhance the implementation of a circular economy, better models are needed to understand and estimate inbound contamination rates as well as identify critical factors in the present and future.

54 ENVIRONMENTAL SCIENCES↗

Predicting measures of soil health using the microbiome and supervised machine learning

Soil health encompasses a range of biological, chemical, and physical soil properties that sustain the commercial and ecological value of agroecosystems. Monitoring soil health requires a comprehensive set of diagnostics that can be cost-prohibitive for routine analyses. The soil microbiome provides a rich source of information about soil properties, which can be assayed in a high-throughput, cost-effective way. We evaluated the accuracy of random forest (RF) and support vector machine (SVM) regression and classification models in predicting 12 measures of soil health, tillage status, and soil texture from 16S rRNA gene amplicon data with an operationally relevant sample set. We validated the efficacy of the best performing models against independent datasets and also tested best practices for processing microbiome data for use in machine learning. Soil health metrics could be predicted from microbiome data with the best models achieving a Kappa value of ~0.65, for categorical assessments, and a R2 value of ~0.8, for numerical scores. Biological health ratings were better predicted than chemical or physical ratings. Validation with independent datasets revealed that models had general predictive value for soil properties, including yield. The ecological profiles of several taxa important for model accuracy matched the observed relationships with soil health, including Pyrinomonadaceae, Nitrososphaeraceae, and Candidatus Udeaobacter. Models trained at the highest taxonomic resolution proved most accurate, with losses in accuracy resulting from rarefying, sparsity filtering, and aggregating at higher taxonomic ranks. Furthermore, our study provides the groundwork for developing scalable technology to use microbiome-based diagnostics for the assessment of soil health.

16S rRNA gene↗

Estimating Compressional Velocity and Bulk Density Logs in Marine Gas Hydrates Using Machine Learning

Compressional velocity (Vp) and bulk density (ρb) logs are essential for characterizing gas hydrates and near-seafloor sediments; however, it is sometimes difficult to acquire these logs due to poor borehole conditions, safety concerns, or cost-related issues. We present a machine learning approach to predict either compressional Vp or ρb logs with high accuracy and low error in near-seafloor sediments within water-saturated intervals, in intervals where hydrate fills fractures, and intervals where hydrate occupies the primary pore space. We use scientific-quality logging-while-drilling well logs, gamma ray, ρb, Vp, and resistivity to train the machine learning model to predict Vp or ρb logs. Of the six machine learning algorithms tested (multilinear regression, polynomial regression, polynomial regression with ridge regularization, K nearest neighbors, random forest, and multilayer perceptron), we find that the random forest and K nearest neighbors algorithms are best suited to predicting Vp and ρb logs based on coefficients of determination (R2) greater than 70% and mean absolute percentage errors less than 4%. Given the high accuracy and low error results for Vp and ρb prediction in both hydrate and water-saturated sediments, we argue that our model can be applied in most LWD wells to predict Vp or ρb logs in near-seafloor siliciclastic sediments on continental slopes irrespective of the presence or absence of gas hydrate.

Naim, Fawz↗

Algorithmic Classification of Raman Spectra Biosignatures: Improving Life Detection Confidence

“Agnostic” biosignatures – indicators of life (or the absence of life), independent of a particular biochemistry – are increasingly considered a high standard for life detection. The Ladder of Life Detection (2018) called for investigating how combinations of independent and different potential biosignatures affect confidence. To address this gap, statistical classification of elemental abundances, isotopic fractionation, and reflectance spectroscopy (VNIR) has been implemented. Raman spectroscopy, highly desirable due to its wide availability, has the potential to improve this predictive power. This work implemented biosignature classification algorithms on Raman data alone, in preparation for combination with the other data types. Raman spectroscopy data was collected from published databases and papers as part of a manually curated dataset of “indicative” and “non-indicative of life” samples. These currently include 61 non-indicative samples (meteorites, magnetite); 3 indicative living samples (bacteria); 20 indicative non-living samples (chalk, bone); and 12 indicative mixed (with non-indicative material) samples (soil, microbial mats). Laboratory work is ongoing to characterize additional samples, particularly a greater breadth of mixed systems. Spectra were interpolated, filtered with the Savitzsky-Golay filter, and de-noised. For a preliminary examination, agnostic features were manually extracted including mean intensity, number of peaks, and mean peak width. Different peak prominences and filtering polynomials were used to refine features. Classification algorithms were implemented: k-nearest neighbors (KNN), logistic regression (LR), linear support vector machines (SVM), random forest (RF), Gaussian naïve bayes (GNB). Lastly, Monte Carlo simulations on 1,000 50%-train-test-splits were used to validate classification performance and feature significance. The preliminary feature set achieved its highest AUC of 0.52 with LR, with no strongly discriminatory features. Work to improve feature extraction, such as through deep learning with back propagation, is planned. In future work, the Raman data will be combined with the other data types, and potentially new data types such as enantiomeric excess. This project was partially supported through the NASA Ames Project EXcellence (APEX) incubator program.

Astrobiology↗