Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “support vector regression”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Machine learning models for estimating contamination across different curbside collection strategies

Contaminated recyclables, which are frequently discarded as waste, pose a significant challenge to the implementation of a circular economy. These contaminated recyclables impede the circulation of resources, resulting in higher processing costs at material recovery facilities (MRFs). Over the past few decades, machine learning (ML) models such as linear regression (LR), support vector machine (SVM), and random forest (RF) have evolved to provide new methods for predicting inbound contamination rates in addition to traditional statistical models. In this study, we applied ML models to predict inbound contamination rates using demographic features from 15 counties in the U.S. with different curbside collection strategies. In general, we found that ML models outperformed linear mixed models. Specifically, SVM models had the highest performance (R 2 = 0.75; mean absolute error (MAE) = 0.06), which may be due to their ability to model nonlinear relationships between features and inbound contamination rates. Further, the key predictor was population, with poverty rate being positively correlated and median age negatively correlated with inbound contamination rates. To improve the management of contamination and enhance the implementation of a circular economy, better models are needed to understand and estimate inbound contamination rates as well as identify critical factors in the present and future.

54 ENVIRONMENTAL SCIENCES↗

BigNeuron: a resource to benchmark and predict performance of algorithms for automated tracing of neurons in light microscopy datasets

BigNeuron is an open community bench-testing platform with the goal of setting open standards for accurate and fast automatic neuron tracing. We gathered a diverse set of image volumes across several species that is representative of the data obtained in many neuroscience laboratories interested in neuron tracing. Here, we report generated gold standard manual annotations for a subset of the available imaging datasets and quantified tracing quality for 35 automatic tracing algorithms. The goal of generating such a hand-curated diverse dataset is to advance the development of tracing algorithms and enable generalizable benchmarking. Together with image quality features, we pooled the data in an interactive web application that enables users and developers to perform principal component analysis, t-distributed stochastic neighbor embedding, correlation and clustering, visualization of imaging and tracing data, and benchmarking of automatic tracing algorithms in user-defined data subsets. The image quality metrics explain most of the variance in the data, followed by neuromorphological features related to neuron size. Furthermore, we observed that diverse algorithms can provide complementary information to obtain accurate results and developed a method to iteratively combine methods and generate consensus reconstructions. The consensus trees obtained provide estimates of the neuron structure ground truth that typically outperform single algorithms in noisy datasets. However, specific algorithms may outperform the consensus tree strategy in specific imaging conditions. Finally, to aid users in predicting the most accurate automatic tracing results without manual annotations for comparison, we used support vector machine regression to predict reconstruction quality given an image volume and a set of automatic tracings.

97 MATHEMATICS AND COMPUTING↗

Machine learning prediction of electron density and temperature from He I line ratios

We propose to utilize machine learning to predict the electron density, ne, and temperature, T e , from He I line intensity ratios. In this approach, training data consist of measured He I line ratios as input and ne and T e measured using other diagnostic(s) as desired output, which is a Langmuir probe in our study. Support vector machine regression analysis is, then, performed with the training data to develop a predictive model for n e and T e , separately. It is confirmed that n e and T e predicted using the developed models agree well with those from the Langmuir probe in the ranges of 0.28 × 10 18 ≤ n e (m -3 ) ≤ 3.8 × 10 18 and 3.2 ≤ T e (eV) ≤ 7.5. The developed models are, further, examined with an evaluation data, which are not included in the training data, and are found to well reproduce absolute values and radial profiles of probe-measured n e and T e .

47 OTHER INSTRUMENTATION↗

A review on recent machine learning applications for imaging mass spectrometry studies

Imaging mass spectrometry (IMS) is a powerful analytical technique widely used in biology, chemistry, and materials science fields that continue to expand. IMS provides a qualitative compositional analysis and spatial mapping with high chemical specificity. The spatial mapping information can be 2D or 3D depending on the analysis technique employed. Due to the combination of complex mass spectra coupled with spatial information, large high-dimensional datasets (hyperspectral) are often produced. Therefore, the use of automated computational methods for an exploratory analysis is highly beneficial. The fast-paced development of artificial intelligence (AI) and machine learning (ML) tools has received significant attention in recent years. These tools, in principle, can enable the unification of data collection and analysis into a single pipeline to make sampling and analysis decisions on the go. There are various ML approaches that have been applied to IMS data over the last decade. Here, in this review, we discuss recent examples of the common unsupervised (principal component analysis, non-negative matrix factorization, k-means clustering, uniform manifold approximation and projection), supervised (random forest, logistic regression, XGboost, support vector machine), and other methods applied to various IMS datasets in the past five years. The information from this review will be useful for specialists from both IMS and ML fields since it summarizes current and representative studies of computational ML-based exploratory methods for IMS.

47 OTHER INSTRUMENTATION↗

IoT Intrusion Detection Taxonomy, Reference Architecture, and Analyses

This paper surveys the deep learning (DL) approaches for intrusion-detection systems (IDSs) in Internet of Things (IoT) and the associated datasets toward identifying gaps, weaknesses, and a neutral reference architecture. A comparative study of IDSs is provided, with a review of anomaly-based IDSs on DL approaches, which include supervised, unsupervised, and hybrid methods. All techniques in these three categories have essentially been used in IoT environments. To date, only a few have been used in the anomaly-based IDS for IoT. For each of these anomaly-based IDSs, the implementation of the four categories of feature(s) extraction, classification, prediction, and regression were evaluated. We studied important performance metrics and benchmark detection rates, including the requisite efficiency of the various methods. Four machine learning algorithms were evaluated for classification purposes: Logistic Regression (LR), Support Vector Machine (SVM), Decision Tree (DT), and an Artificial Neural Network (ANN). Therefore, we compared each via the Receiver Operating Characteristic (ROC) curve. The study model exhibits promising outcomes for all classes of attacks. The scope of our analysis examines attacks targeting the IoT ecosystem using empirically based, simulation-generated datasets (namely the Bot-IoT and the IoTID20 datasets).

97 MATHEMATICS AND COMPUTING↗

Vehicle Position Detection Based on Machine Learning Algorithms in Dynamic Wireless Charging

Dynamic wireless charging (DWC) has emerged as a viable approach to mitigate range anxiety by ensuring continuous and uninterrupted charging for electric vehicles in motion. DWC systems rely on the length of the transmitter, which can be categorized into long-track transmitters and segmented coil arrays. The segmented coil array, favored for its heightened efficiency and reduced electromagnetic interference, stands out as the preferred option. However, in such DWC systems, the need arises to detect the vehicle’s position, specifically to activate the transmitter coils aligned with the receiver pad and de-energize uncoupled transmitter coils. This paper introduces various machine learning algorithms for precise vehicle position determination, accommodating diverse ground clearances of electric vehicles and various speeds. Through testing eight different machine learning algorithms and comparing the results, the random forest algorithm emerged as superior, displaying the lowest error in predicting the actual position.

47 OTHER INSTRUMENTATION↗

Influence of Global Climate on Freshwater Changes in Africa’s Largest Endorheic Basin Using Multi-Scaled Indicators

The poor investments in gauge measurements for hydro-climatic research in Africa has necessitated the need to investigate how decision makers can leverage on sophisticated spaceborne measurements to improve knowledge on surface water hydrology that can feed directly into water accounting processes and risk assessment from extreme droughts and its impacts. To demonstrate such potential, a suite of satellite earth observations (Sentinel-2, altimetry, Landsat, GRACE, and TRMM) and model data are combined with the standardized precipitation evapotranspiration index to assess the impacts of global climate on freshwater dynamics over the LCB (Lake Chad basin), Africa’s largest endorheic basin. As shown in the results of this study, the significant relationship of climate modes (AMO; r = 0.68 and 0.59; and AMM; r = 0.2 and 0.47) with drought patterns in the LCB highlights the evidence of global climate influence in the region. The significant declines in drought extents and their intensities (2004 - 2015) over LCB coincide with the rise in surface water extent of the Lake Chad during the same period. Change detection analysis of open water features in the southern pool of Lake Chad during the 2015 - 2019 period shows that on the average, only 28.4% of inundated areas within the vicinity of the Lake persisted during the period. While the association of terrestrial water storage (TWS) with model-derived surface water storage (SWS) is strongest (r = 0.89) in the catchments that provide the most nourishment to the Lake Chad, the relationship of rainfall (2002 - 2017) with TWS (r = 0.85), model TWS (r = 0.87) and SWS (r = 0.88) confirm that the LCB’s hydrology is predominantly climate-driven. This notion is further reinforced as the predicted SWS over the LCB using a support vector machine regression scheme was found to be strongly correlated (r = 0.95 at = 0.05) with observed SWS.

Sentinel-2↗

The Added Value of SMAP Soil Moisture in Crop Yield Forecasting Over Argentina

Argentina is one of the major producers and exporter of soybeans, corn, and wheat to the world market; therefore, the accurate and timely forecasting of those crops yield is crucial to national crop management and global food security. Previous studies have mainly focused on developing forecasting models for a specific crop type and location using a single source of data (e.g., vegetation indices), thus providing little insight into the forecasting models' performance on different crop types and regions. Besides, these models are based on traditional statistical regression algorithms, while more advanced machine learning approaches have not been explored. This study investigated the estimation of crop yields of three major crops (corn, soybean, and winter wheat) using Multiple Linear Regression (MLR) and Support Vector Machine (SVM), over major growing provinces in Argentina. Our models were trained and evaluated on data from 2015 to 2020, where three remote sensing products (Normalized difference vegetation index (NDVI), SMAP soil moisture, and MODIS evapotranspiration) were used as predictors. Our results indicated that accurate crop yield forecasts using the developed regression models could be made one to two months before harvest. The MLR and SVM model performance varied among different crop types, where soybean and corn exhibited better predictability compare to the wheat. In most cases, the SVM outperformed the multiple linear regression model due to its ability to capture the nonlinear and complex features of the crop-production process. The forecasted model that combines data from multiple sources outperformed single-source satellite data. The highest accuracy was obtained when the three data sources were all considered in the model development. Results also indicated that the inclusion of SMAP soil moisture improved crop yield forecasting in most provinces, and the most significant improvements occurred in the drier region.

Nazmus Shams Sazib↗

Predicting measures of soil health using the microbiome and supervised machine learning

Soil health encompasses a range of biological, chemical, and physical soil properties that sustain the commercial and ecological value of agroecosystems. Monitoring soil health requires a comprehensive set of diagnostics that can be cost-prohibitive for routine analyses. The soil microbiome provides a rich source of information about soil properties, which can be assayed in a high-throughput, cost-effective way. We evaluated the accuracy of random forest (RF) and support vector machine (SVM) regression and classification models in predicting 12 measures of soil health, tillage status, and soil texture from 16S rRNA gene amplicon data with an operationally relevant sample set. We validated the efficacy of the best performing models against independent datasets and also tested best practices for processing microbiome data for use in machine learning. Soil health metrics could be predicted from microbiome data with the best models achieving a Kappa value of ~0.65, for categorical assessments, and a R2 value of ~0.8, for numerical scores. Biological health ratings were better predicted than chemical or physical ratings. Validation with independent datasets revealed that models had general predictive value for soil properties, including yield. The ecological profiles of several taxa important for model accuracy matched the observed relationships with soil health, including Pyrinomonadaceae, Nitrososphaeraceae, and Candidatus Udeaobacter. Models trained at the highest taxonomic resolution proved most accurate, with losses in accuracy resulting from rarefying, sparsity filtering, and aggregating at higher taxonomic ranks. Furthermore, our study provides the groundwork for developing scalable technology to use microbiome-based diagnostics for the assessment of soil health.

16S rRNA gene↗

Estimating mass-absorption cross-section of ambient black carbon aerosols: theoretical, empirical, and machine learning models

The mass-absorption cross-section of black carbon (MAC BC ) is an essential parameter to link the atmospheric concentration of black carbon (BC) with its radiative forcing. When a direct calculation of MAC BC based on observations of aerosol light absorption and BC mass concentration is impossible, we rely on modeling and simulations to estimate MAC BC , but currently, there is no consensus model that can be relied on for accurate predictions across all atmospheric environments when BC particles have different coating thicknesses. Here, we applied five MAC BC prediction models (including three light scattering theories, an empirical model based on observations of particle mass concentrations, and a machine learning model developed in our previous work) to aerosols from three Department of Energy (DOE) Atmospheric Radiation Measurement (ARM) field campaigns. While many studies have found that increasing the complexity of the models helps to constrain biases of the estimated MAC BC , our effort is to evaluate the models based on the criteria of simplicity and accuracy. We find that our machine learning model (support vector machine for regression, SVM) generally performs well across all DOE ARM field campaign data, while the accuracy of core-shell Mie theory depends on the bias correction algorithm applied to filter-based light absorption data. Generally, the empirical model for internally-mixed particles that we considered tends to over-predict MAC BC , while Mie theory for externally-mixed particles tends to under-predict MAC BC . An examination of the influence of coating material on BC cores suggests that the performance of our current SVM model is degraded when the BC is thickly-coated (e.g., it has undergone aging and mixing with other materials in the atmosphere).

54 ENVIRONMENTAL SCIENCES↗

Algorithmic Detection of Elemental Biosignatures

Machine learning models that classify a sample as indicative or non-indicative of life could play an important role in life-detection missions. Their predictions result from agnostic algorithms and thereby add redundancy to judgements resulting from human expertise. Additionally, their important features can reveal the most informative measurements within the operational constraints of a life-detection mission. The Ladder of Life Detection (Neveu 2018) identifies the need for an understanding of how combinations of multiple biosignatures affect overall confidence. The present work provides a starting point to answer this need, and future work will expand the data types to obtain even more predictive combinations of features. Elemental abundance was chosen as a starting set of features due to its availability in diverse sample types, which are needed to train a generalizable model. A standardized dataset was collected, including 35 non-indicative, e.g., lunar rock, basalt; 19 indicative mixed, e.g., seawater, agricultural soil; 46 indicative non-alive, e.g., coal, chalk; and 10 indicative alive, e.g., biofilm, bacteria. This dataset could be valuable for complementary biosignature research. The samples were standardized to the same limit of detection of a simulated mission scenario. Four classification models were used: k-nearest neighbors (KNN), logistic regression (LR), linear support vector machines (SVM), and Gaussian naïve Bayes (GNB). To obtain feature importances, KNN was run on three principal components of the training data and LR and SVM were run with L1 and L2 regularization. The performances and feature importances of the six model variants on 40:60 train to validation ratios were assessed with Monte Carlo simulations. ROC AUC and mean accuracy scores ranged between 82% - 94%, with sensitivity greater than specificity. For indicative of life predictors, all models had C and Ca as strong and Cl as medium; a majority of models had N, K, and P as medium. For non-indicative of life predictors, all models had Si as strong, and a majority of models had Mg, Al, and Ti as medium. Varied elements were Fe (slightly non-indicative), H (slightly indicative), O (widely varied), Na, Mn, and S. These results serve as a proof of concept and suggest important elemental signals beyond merely the CHNOPS of Earth-based life.

Algorithmic↗

Algorithmic Classification of Raman Spectra Biosignatures: Improving Life Detection Confidence

“Agnostic” biosignatures – indicators of life (or the absence of life), independent of a particular biochemistry – are increasingly considered a high standard for life detection. The Ladder of Life Detection (2018) called for investigating how combinations of independent and different potential biosignatures affect confidence. To address this gap, statistical classification of elemental abundances, isotopic fractionation, and reflectance spectroscopy (VNIR) has been implemented. Raman spectroscopy, highly desirable due to its wide availability, has the potential to improve this predictive power. This work implemented biosignature classification algorithms on Raman data alone, in preparation for combination with the other data types. Raman spectroscopy data was collected from published databases and papers as part of a manually curated dataset of “indicative” and “non-indicative of life” samples. These currently include 61 non-indicative samples (meteorites, magnetite); 3 indicative living samples (bacteria); 20 indicative non-living samples (chalk, bone); and 12 indicative mixed (with non-indicative material) samples (soil, microbial mats). Laboratory work is ongoing to characterize additional samples, particularly a greater breadth of mixed systems. Spectra were interpolated, filtered with the Savitzsky-Golay filter, and de-noised. For a preliminary examination, agnostic features were manually extracted including mean intensity, number of peaks, and mean peak width. Different peak prominences and filtering polynomials were used to refine features. Classification algorithms were implemented: k-nearest neighbors (KNN), logistic regression (LR), linear support vector machines (SVM), random forest (RF), Gaussian naïve bayes (GNB). Lastly, Monte Carlo simulations on 1,000 50%-train-test-splits were used to validate classification performance and feature significance. The preliminary feature set achieved its highest AUC of 0.52 with LR, with no strongly discriminatory features. Work to improve feature extraction, such as through deep learning with back propagation, is planned. In future work, the Raman data will be combined with the other data types, and potentially new data types such as enantiomeric excess. This project was partially supported through the NASA Ames Project EXcellence (APEX) incubator program.

Astrobiology↗

Estimation of Snow Mass Information via Assimilation of C-Band Synthetic Aperture Radar Backscatter Observations Into an Advanced and Surface Model

This study assimilated Sentinel-1 C-band backscatter observations over snow-covered terrain into the Noah-Multiparameterization land surface model using support vector machine (SVM) regression and an ensemble Kalman filter to improve the modeled terrestrial snow mass estimates. The data assimilation (DA) experiment was conducted across Western Colorado from September 2016 to August 2017. As part of the DA experiments, the impact of a rule-based update was evaluated by comparing snow water equivalent (SWE) estimates via DA (with [ DAv1 ] and without [ DAv2 ] the rule-based update) against SNOTEL SWE measurements. Results confirmed that rule-based update helped minimize SVM controllability issues, and in turn, improved the accuracy of SWE estimates relative to both open loop (OL) and DAv2 . Comparison of SWE estimates from Sentinel-1 DAv1 against SNOTEL SWE revealed that 75% of stations showed improvements in bias and correlation coefficient relative to the OL. Assimilated SWE estimates also showed statistical improvements during both the snow accumulation and snow ablation periods. However, unbiased root mean square error showed a slight increase during the snow ablation period due to the large variability in the electromagnetic response of C-band backscatter over deep and/or wet snow. Improvement of the SWE estimates also resulted in improving river discharge estimates compared to in situ measurements. River discharge using Sentinel-1 DAv1 improved the Nash–Sutcliffe efficiency at all available stations. These results suggest that physically constrained SVM can serve as an efficient observation operator for snow mass DA through explicit consideration of the first-order C-band scattering mechanisms over different terrestrial snow conditions.

Jongmin Park↗

Statistical Classification of Biosignature Information using Multiple Instrument Observations

The accurate identification of biosignatures (indications of life) from data taken from remote or in situ planetary exploration is one of the most important challenges in astrobiology, the interdisciplinary field examining habitability and the potential for extraterrestrial life. This study employs machine learning algorithms to optimize the identification of biosignatures, with an emphasis on those which are agnostic to a specific biochemical basis. We exploit the wealth of terrestrial data available from biogenic and abiogenic systems to enhance efficient feature prioritization. Our dataset, pulled from public databases and laboratory recorded measurements, includes elemental abundance, isotopic fractionation, and VNIR/Raman spectra The data curation process included standardization for detection limits and ranges. Subsequent feature extraction yielded detailed inputs for machine learning, including combinations of elemental content, isotopic ratios, and parameters of spectral peaks and troughs. Feature significance was evaluated across diverse machine learning methodologies, such as k-nearest neighbors, logistic regression, Random Forest, support vector machines, and Gaussian Naïve Bayes, along with a combined voting classifier. We utilized Receiver Operating Characteristic Area Under the Curve (ROC AUC) across 2,000 50% test-train splits as a robust metric of model performance. Results revealed a promising ROC AUC of 0.853 for the combined voting classifier. Removing elemental abundance data notably reduced model accuracy (13% decrease in AUC), highlighting its critical role in biosignature detection. Several other individual data features exhibited significance within their respective data types, offering additional granularity. This research fortifies the relevance of machine learning to astrobiology, potentially enhancing life detection missions by allowing algorithmic prioritization of high-interest samples for further investigation. Future work will refine data standardization, expand the dataset to include more terrestrial systems, and incorporate convolutional neural networks for spectral feature extraction. The potential for public data sharing is also under exploration, reinforcing our commitment to collective scientific advancement.

Statistical↗

Predicting nepheline precipitation in waste glasses using ternary submixture model and machine learning

Nepheline precipitation in nuclear waste glasses during vitrification can be detrimental due to its negative effect on chemical durability. Developing models to accurately predict nepheline precipitation from compositions is important to increase waste loading since existing models can be overly conservative. In this study, an expanded dataset containing 955 glasses was compiled from literature data, where 355 glasses are for high-level waste (HLW). Previously developed submixture models were refitted using the new dataset, where a misclassification rate of 7.8% was achieved. Nine machine learning (ML) algorithms (e.g., k-nearest neighbor, Gaussian process regression, artificial neural network, support vector machine, decision tree, etc.) were applied to evaluate their ability of predicting nepheline precipitation from compositions. Model accuracy, precision, recall/sensitivity, and F1 score were systemically compared between different ML algorithms and modeling protocols. Good model prediction with an accuracy ~0.9 (misclassification rate of ~10%) was observed with different algorithms under certain protocol. This study evaluated various ML models to predict nepheline precipitations in waste glasses, highlighting the importance of data preparation, modeling protocol, and their effect on model stability and reproducibility. The results provide insights into applying ML to predict glass properties and suggest areas for future research on modeling nepheline precipitations.

Lu, Xiaonan↗

Deep Learning Estimation of Daily Ground–Level NO 2 Concentrations from Remote Sensing Data

The limited number of nitrogen dioxide (NO 2 ) surface measurements calls for the development of highly accurate approaches to estimating surface NO 2 concentrations. In this study, we leverage a new satellite instrument, the TROPOspheric Monitoring Instrument (TROPOMI), along with other predictor variables, to estimate daily surface NO 2 concentrations over Texas in 2019. We use the deep convolutional neural network (Deep-CNN), an advanced deep learning algorithm, to obtain estimates and achieve a correlation coefficient (R) of 0.91, an index of agreement (IOA) of 0.95, and a mean absolute bias (MAB) of 1.75 ppb in surface NO 2 estimation. Additionally, we leverage a novel approach, SHapley Additive exPlanations (SHAP), to describe how Deep-CNN understands each predictor variable. The SHAP results show that the Deep-CNN model has an advanced understanding of the dataset, revealing that TROPOMI closely captures levels of NO 2 . In addition, we show the superiority of our Deep-CNN model at estimating surface NO 2 over other well-known machine learning and regression models in the field, including the support vector machines (SVM), random forest (RF), and multiple linear regression (MLR). Although SVM and RF show strong capabilities at estimating surface NO 2 concentrations, their accuracy is inferior to that of the Deep-CNN model, ranking second and third in model accuracy in this study. The MLR, however, shows a poor ability at NO 2 estimation and ranks last among all models. Furthermore, testing the impact of sample size on model performance, we also show that, compared to other models, Deep-CNN needs more samples to trigger its strength at surface NO 2 estimation.

54 ENVIRONMENTAL SCIENCES↗

Predicting melt pool depth and grain length using multiple signatures from in-situ single camera two-wavelength imaging pyrometry for laser powder bed fusion

In laser powder bed fusion (LPBF), the in-situ process signatures are known to have a direct correlation with the microstructural properties of the solidified melt pool (MP). It is known that the MP cooling and heating rates, and laser processing parameters can critically determine the grain structure and thereby affect the part properties. The objective of this work is to study the feasibility of using in-process, high-speed imaging pyrometry for evaluating the solidified MP properties “below” the surface, such as depth and microstructural properties. To accomplish this, we employ an in-house single camera-based two-wavelength imaging pyrometry (STWIP) system for monitoring the printing of single-scan tracks with Inconel 718 on a commercial LPBF printer (EOS M290). Further, the lab designed STWIP system is a coaxial high-speed (>10,000 fps) imaging system capable of monitoring MP temperature, morphology, and intensity profiles. The temperature measurements from STWIP are emissivity independent. The STWIP measured MP signatures of the printed tracks are correlated with the ex-situ microscopy characterized MP depth and the average grain lengths. From the data analysis, using support vector machine (SVM)-based regression models, we found that the MP temperature signatures are crucial for an accurate prediction of MP depth and the grain length, thus validating the novelty and necessity of the developed in-situ monitoring methods and analysis.

36 MATERIALS SCIENCE↗

Application of machine learning approaches in the analysis of mass absorption cross-section of black carbon aerosols: Aerosol composition dependencies and sensitivity analyses

Physics-based models typically require an in-depth understanding of a phenomenon and assumptions of the underlying process(es), which are often hard to obtain in practice, whereas data-driven machine learning models learn the structure and patterns in the training data without any prior theoretical assumptions and then use inference to develop useful predictions. A novel machine learning-based algorithm has been previously developed for the prediction of black carbon mass absorption cross-section (MAC BC ) and applied to a variety of different atmospheric environments. In contrast to light scattering theories which require assumptions about the underlying physics, this algorithm uses time-series data of aerosol properties to estimate the temporally-varying MAC BC at 870 nm. Here, we analyze our algorithm and discuss the influence of aerosol optical properties (such as Ångström exponents and single scattering albedo) and chemical composition on the model outputs and the associated accuracy. Additionally, we conduct sensitivity analyses on our models to understand how the predictions change in response to different sets of input variables. Our support vector machine (SVM) for regression model is the least sensitive to variations in the input variables, although all models tend to exhibit a degradation to their accuracy when scattering Ångström exponents are less than one.

54 ENVIRONMENTAL SCIENCES↗