Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “machine learning, Random Forest”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

A Machine Learning Approach to Quantitative Analysis of Enamel Microstructure from Scanning Electron Microscopy Images

Dental enamel, the outermost tissue of mammalian teeth, must withstand a lifetime of wear and cyclic contact. To meet this demand, enamel possesses a combination of high hardness and resistance to fracture, properties that are typically mutually exclusive. The impressive damage tolerance has been attributed largely to decussation of the enamel rods, the principal unit of its microstructure. As such, enamel is inspiring the design of next‐generation structural materials. However, quantitative descriptions of the decussated enamel rod microstructure remain limited due to challenges encountered in applying computed tomography and in acquiring quality images appropriate for traditional digital processing methods. Here, a machine learning segmentation method is applied to images of the enamel obtained using scanning electron microscopy to support quantitative analysis of the microstructure. A pretrained convolutional neural network is used to expand the input training image dataset to allow the training of a random forest classifier, which ultimately segments the image with a very small training set ( n = 3 images). A validation of this segmentation method is presented, in addition to its application to calculate relevant microstructural parameters for images of tooth enamel from selected mammalian species. The methodology applied here is equally applicable to other hard tissues.

36 MATERIALS SCIENCE↗

A comparative study of machine learning models for predicting the state of reactive mixing

Mixing phenomena are important mechanisms controlling flow, species transport, and reaction processes in fluids and porous media. Accurate predictions of reactive mixing are critical for many Earth and environmental science problems such as contaminant fate and remediation, macroalgae growth, and plankton biomass evolution. Here, to investigate the evolution of mixing dynamics under different scenarios (e.g., anisotropy, fluctuating velocity fields), a finite-element-based numerical model was built to solve the fast, irreversible bimolecular reaction-diffusion equations to simulate a range of reactive-mixing scenarios. A total of 2,315 simulations were performed using different sets of model input parameters comprising various spatial scales of vortex structures in the velocity field, time-scales associated with velocity oscillations, the perturbation parameter for the vortex-based velocity, anisotropic dispersion contrast (i.e., ratio of longitudinal-to-transverse dispersion), and molecular diffusion. The outputs comprised concentration profiles of reactants and products. The inputs to and outputs from these simulations were concatenated into feature and label matrices, respectively, to train 20 different machine learning (ML) models intended to emulate system behavior. These 20 ML emulators, based on linear methods, Bayesian methods, ensemble learning methods, and multilayer perceptrons (MLPs), were trained to classify the state of mixing and predict three quantities of interest (QoIs) characterizing species production, decay (i.e., average concentration, square of average concentration), and degree of mixing (i.e., variances of species concentration). Unsurprisingly, linear classifiers and regressors failed to reproduce the QoIs; however, ensemble methods (classifiers and regressors) and the MLP model accurately classified the state of reactive mixing and the QoIs. Among ensemble methods, random forest and decision-tree-based AdaBoost faithfully predicted the QoIs. At run time, trained ML emulators produced results times faster than the finite-element simulations. Due to their low computational expense and high accuracy, ensemble and MLP models are excellent emulators for these numerical simulations and great utilities in uncertainty quantification exercises, which can require 1,000s of forward model runs.

97 MATHEMATICS AND COMPUTING↗

Towards fast, accurate predictions of RF simulations via data-driven modeling: Forward and lateral models

Three machine learning techniques (multilayer perceptron, random forest, and Gaussian process) provide fast surrogate models for lower hybrid current drive (LHCD) simulations. A single GENRAY/CQL3D simulation without radial diffusion of fast electrons requires several minutes of wall-clock time to complete, which is acceptable for many purposes, but too slow for integrated modeling and real-time control applications. More accurate simulations with fast electron diffusion are even slower, requiring multiple hours of run time with parallel processing. The machine learning models use a database of 16,000+ GEN-RAY/CQL3D simulations for training, validation, and testing. Latin hypercube sampling methods implemented in πScope ensure that the database covers the range of 9 input parameters (n e0 , T e0 , I p , B t , R 0 , n ∥︀ , Z e f f , V loop , P LHCD ) with sufficient density in all regions of parameter space. The surrogate models reduce the computation time from minutes-hours to ms with high accuracy across the input parameter space. Data-driven surrogate models also allow for solving inverse and “lateral” problems. A surrogate model for the inverse problem maps from a desired current drive or power deposition profile to a set of input parameters that would result in such a profile, while a surrogate model for the lateral problem maps from a measured experimental quantity such as hard x-ray emission to a current drive or power deposition profile. In conclusion, the πScope database creation workflow is flexible and applicable to other RF simulation codes such as TORIC.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Random Forests as a Viable Method to Select and Discover High-redshift Quasars

We present a method of selecting quasars up to redshift ≈6 with random forests, a supervised machine-learning method, applied to Pan-STARRS1 and WISE data. We find that, thanks to the increasing set of known quasars, we can assemble a training set that enables supervised machine-learning algorithms to become a competitive alternative to other methods up to this redshift. We present a candidate set for the redshift range 4.8–6.3, which includes the region around z = 5.5 where selecting quasars is difficult due to their photometric similarity to red and brown dwarfs. We demonstrate that, under our survey restrictions, we can reach a high completeness (66% ± 7% below redshift 5.6/83{sub -9}{sup +6}% above redshift 5.6) while maintaining a high selection efficiency (78{sub -8}{sup +10}%/94{sub -8}{sup +5}%). Our selection efficiency is estimated via a novel method based on the different distributions of quasars and contaminants on the sky. The final catalog of 515 candidates includes 225 known quasars. We predict the candidate catalog to contain additional 148{sub -33}{sup +41} new quasars below redshift 5.6 and 45{sub -8}{sup +5} above, and we make the catalog publicly available. Spectroscopic follow-up observations of 37 candidates led us to discover 20 new high redshift quasars (18 at 4.6 ≤ z ≤ 5.5, 2 z ~ 5.7). These observations are consistent with our predictions on efficiency. We argue that random forests can lead to higher completeness because our candidate set contains a number of objects that would be rejected by common color cuts, including one of the newly discovered redshift 5.7 quasars.

79 ASTRONOMY AND ASTROPHYSICS↗

A general spatial-temporal framework for short-term building temperature forecasting at arbitrary locations with crowdsourcing weather data

Weather forecasting has been a critical component to predict and control building energy consumption for better building energy management. Without accessibility to other data sources, the onsite observed temperatures or the airport temperatures are used in forecast models. In this paper, we present a novel approach by utilizing the crowdsourcing weather data from neighboring personal weather stations (PWS) to improve the weather forecast accuracy around buildings using a general spatial-temporal modeling framework. The final forecast is based on the ensemble of local forecasts for the target location using neighboring PWSs. Our approach is distinguished from existing literature in various aspects. First, we leverage the crowdsourcing weather data from PWS in addition to public data sources. In this way, the data is at much finer time resolution (e.g., at 5-minute frequency) and spatial resolution (e.g., arbitrary location vs grid). Second, our proposed model incorporates spatial-temporal correlation information of weather variables between the target building and a set of neighboring PWSs so that underlying correlations can be effectively captured to improve forecasting performance. Here, we demonstrate the performance of the proposed framework by comparing to the benchmark models on temperature forecasting for a building located at an arbitrary location at San Antonio, Texas, USA. In general, the proposed model framework equipped with machine learning technique such as Random Forest can improve forecasting by 50% compares with persistent model and has 90% chance to outperform airport forecast in short-term forecasting. In a real-time setting, the proposed model framework can provide more accurate temperature forecasting results compared with using airport temperature forecast for most forecast horizon. Moreover, we analyze the sensitivity of model parameters to gain insights on how crowdsourcing data from the neighboring personal weather stations impacts forecasting performance. Finally, we implement our model in other cities such as Syracuse and Chicago to test the model's performance in different landforms and climate types.

54 ENVIRONMENTAL SCIENCES↗

Information Content of JWST NIRSpec Transmission Spectra of Warm Neptunes

Warm Neptunes offer a rich opportunity for understanding exo-atmospheric chemistry. With the upcoming James Webb Space Telescope (JWST), there is a need to elucidate the balance between investments in telescope time versus scientific yield. We use the supervised machine-learning method of the random forest to perform an information content (IC) analysis on a 11-parameter model of transmission spectra from the various NIRSpec modes. The three bluest medium-resolution NIRSpec modes (0.7–1.27 μm, 0.97–1.84 μm, 1.66–3.07 μm) are insensitive to the presence of CO. The reddest medium-resolution mode (2.87–5.10 μm) is sensitive to all of the molecules assumed in our model: CO, CO{sub 2}, CH{sub 4}, C{sub 2}H{sub 2}, H{sub 2}O, HCN, and NH{sub 3}. It competes effectively with the three bluest modes on the information encoded on cloud abundance and particle size. It is also competitive with the low-resolution prism mode (0.6–5.3 μm) on the inference of every parameter except for the temperature and ammonia abundance. We recommend astronomers to use the reddest medium-resolution NIRSpec mode for studying the atmospheric chemistry of 800–1200 K warm Neptunes; its corresponding high-resolution counterpart offers diminishing returns. We compare our findings to previous JWST IC analyses that favor the blue orders and suggest that the reliance on chemical equilibrium could lead to biased outcomes if this assumption does not apply. A simple, pressure-independent diagnostic for identifying chemical disequilibrium is proposed based on measuring the abundances of H{sub 2}O, CO, and CO{sub 2}.

79 ASTRONOMY AND ASTROPHYSICS↗

Classification Analysis of Southwest Pacific Tropical Cyclone Intensity Changes Prior to Landfall

This study evaluates the ability of a random forest classifier to identify tropical cyclone (TC) intensification or weakening prior to landfall over the western region of the Southwest Pacific Ocean (SWPO) basin. For both Australia mainland and SWPO island cases, when a TC first crosses land after spending ≥24 h over the ocean, the closest hour prior to the intersection is considered as the landfall hour. If the maximum wind speed (V max ) at the landfall hour increased or remained the same from the 24-h mark prior to landfall, the TC is labeled as intensifying and if the V max at the landfall hour decreases, the TC is labeled as weakening. Geophysical and aerosol variables closest to the 24 h before landfall hour were collected for each sample. The random forest model with leave-one-out cross validation and the random oversampling example technique was identified as the best-performing classifier for both mainland and island cases. The model identified longitude, initial intensity, and sea skin temperature as the most important variables for the mainland and island landfall classification decisions. Incorrectly classified cases from the test data were analyzed by sorting the cases by their initial intensity hour, landfall hour, monthly distribution, and 24-h intensity changes. TC intensity changes near land strongly impact coastal preparations such as wind damage and flood damage mitigations; hence, this study will contribute to improve identifying and prioritizing prediction of important variables contributing to TC intensity change before landfall.

54 ENVIRONMENTAL SCIENCES↗

Estimating soybean yields from high-temporal-resolution multi-source data using deep learning

Accurate and timely crop yield prediction is crucial for ensuring food security and maintaining stable agricultural markets. In recent years, there has been a surge in interest in leveraging high-temporal-resolution, multi-source data for effective crop growth monitoring and yield estimation. A notable challenge arises from the difficulty in capturing the intricate interactions between variables across different time steps within these high-temporal-resolution time series datasets. This complexity hinders the reliable extraction of yield information from voluminous and often noisy datasets, especially during periods of extreme weather events. Here, in this study, we propose an Attention and Graph Isomorphism Network-enhanced Bi-directional Long Short-Term Memory network (AGB-LSTM) for estimating county-level soybean yield in the United States. This model integrates a diverse set of remote sensing data, including Near-Infrared Reflectance of Vegetation (NIRv), Sun-Induced chlorophyll Fluorescence (SIF), and Gross Primary Productivity (GPP), along with environmental covariates. The AGB-LSTM effectively leverages information related to crop yield from high-temporal-resolution time series data (5-days), achieving an accuracy of R²= 0.67 and rRMSE = 14.46%. This approach significantly outperforms traditional machine learning methods such as Random Forest (RF) (R²= 0.52, rRMSE = 17.36%) and Bi-LSTM (R²= 0.58, rRMSE = 16.17%). Sensitivity experiments with different time steps and ranges demonstrated that our model could accurately and stably predict yields 1 to 2 months before harvest. Moreover, data with a finer temporal resolution consistently improved prediction performance, resulting in an approximately 20% increase in and an approximately 20% decrease in rRMSE compared to using monthly composites. We also evaluated the robustness of the model under extreme climate events and observed strong performance (R²= 0.50, rRMSE = 21.32%). Finally, yield mapping for major soybean-producing regions in North America in 2023 revealed spatial patterns that closely matched USDA yield reports. Our findings suggest that the AGB-LSTM model is a promising and effective method for estimating yield and has notable potential for global crop yield forecasting.

Deep learning↗

Machine Learning of Key Variables Impacting Extreme Precipitation in Various Regions of the Contiguous United States

Abstract Amplification in extreme precipitation intensity and frequency can cause severe flooding and impose significant social and economic consequences. Variations in extreme precipitation intensity, frequencies, and return periods can be attributed to many physical variables across spatial and temporal scales. Here we employ ensemble machine learning (ML) methods, namely random forest (RF), eXtreme Gradient Boosting (XGB), and artificial neural networks (ANN), to explore key contributing variables to monthly extreme precipitation intensity and frequency in six regions over the United States. We further establish emulators for return periods. Results show that the ML models for intensity perform better in regions with obvious seasonality (i.e., Northern Great Plains, Southern Great Plains, and West Coast) than the other three regions (Northeast, Southwest, and Rocky Mountains), while for frequency the models perform well for most regions. The Shapley additive explanation is used to help explain the relationships between extreme precipitation characteristics and identify top variables for RF and XGB. We find that latent heat flux, relative humidity, soil moisture, and large‐scale subsidence are key common variables across the regions for both monthly intensity and frequency, and their compound effects are non‐negligible. The developed ML models capture the probability and return period of extreme precipitation well for all regions and may be used for decision making (e.g., infrastructure planning and design).

54 ENVIRONMENTAL SCIENCES↗

Autonomous fabrication of tailored defect structures in 2D materials using machine learning-enabled scanning transmission electron microscopy

Materials with tailored quantum properties can be engineered from atomic-scale assembly techniques, but existing methods often lack the agility and accuracy to precisely and intelligently control the manufacturing process. Here, we demonstrate a fully autonomous approach for fabricating atomic-level defects using electron beams in scanning transmission electron microscopy (STEM) that combines advanced machine learning and automated beam control. As a proof of concept, we achieved controlled fabrication of MoS-nanowire (MoS-NW) edge structures by iterative and targeted exposure of MoS 2 monolayer to a focused electron beam to selectively eject sulfur atoms, utilizing high-angle annular dark-field (HAADF) imaging for feedback-controlled monitoring of structural evolution of defects. A machine learning framework combining a random forest model and a convolutional neural network (CNN) was developed to decode the HAADF image and accurately identify atomic positions and species. This atomic-level information was then integrated into an autonomous decision-making platform, which applied predefined fabrication strategies to instruct beam control about atomic sites to be ejected. The selected sites were subsequently exposed to a localized electron beam using an FPGA-controlled scan routine with precise control over beam positioning and duration. While the MoS-NW edge structures produced exhibit promising mechanical and electronic properties, the proposed methods to build the autonomous fabrication framework is material-agnostic and can be extended to other 2D materials for the creation of diverse defect structures and heterostructures beyond Mo S2 .

Engineering↗

The quenching of galaxies, bulges, and disks since cosmic noon

Here, we present an analysis of the quenching of star formation in galaxies, bulges, and disks throughout the bulk of cosmic history, from z = 2 – 0. We utilise observations from the Sloan Digital Sky Survey and the Mapping Nearby Galaxies at Apache Point Observatory survey at low redshifts. We complement these data with observations from the Cosmic Assembly Near-Infrared Deep Extragalactic Legacy Survey at high redshifts. Additionally, we compare the observations to detailed predictions from the LGalaxies semi-analytic model. To analyse the data, we developed a machine learning approach utilising a Random Forest classifier. We first demonstrate that this technique is extremely effective at extracting causal insight from highly complex and inter-correlated model data, before applying it to various observational surveys. Our primary observational results are as follows: at all redshifts studied in this work, we find bulge mass to be the most predictive parameter of quenching, out of the photometric parameter set (incorporating bulge mass, disk mass, total stellar mass, and B/T structure). Moreover, we also find bulge mass to be the most predictive parameter of quenching in both bulge and disk structures, treated separately. Hence, intrinsic galaxy quenching must be due to a stable mechanism operating over cosmic time, and the same quenching mechanism must be effective in both bulge and disk regions. Despite the success of bulge mass in predicting quenching, we find that central velocity dispersion is even more predictive (when available in spectroscopic data sets). In comparison to the LGalaxies model, we find that all of these observational results may be consistently explained through quenching via preventative ‘radio-mode’ active galactic nucleus feedback. Furthermore, many alternative quenching mechanisms (including virial shocks, supernova feedback, and morphological stabilisation) are found to be inconsistent with our observational results and those from the literature.

79 ASTRONOMY AND ASTROPHYSICS↗

Improved sensitivity of the DRIFT-IId directional dark matter experiment using machine learning

We demonstrate a new type of analysis for the DRIFT-IId directional dark matter detector using a machine learning algorithm called a Random Forest Classifier. The analysis labels events as signal or background based on a series of selection parameters, rather than solely applying hard cuts. The analysis efficiency is shown to be comparable to our previous result at high energy but with increased efficiency at lower energies. This leads to a projected sensitivity enhancement of one order of magnitude below a WIMP mass of 15 GeV c -2 and a projected sensitivity limit that reaches down to a WIMP mass of 9 GeV c -2 , which is a first for a directionally sensitive dark matter detector.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

The drivers and predictability of wildfire re-burns in the western United States (US)

Evidence is mounting that the effectiveness of using prescribed burns as a management tactic may be diminishing due to the higher incidence of wildfire re-burns. The development of predictive models of re-burns is thus essential to better understand their primary drivers so that forest management practices can be updated to account for these events. First, we assess the potential for human activity as a driver of re-burns by evaluating re-burn trends both within and outside of the wildland–urban interface (WUI) of the western US. Next, we investigate the predictability of re-burns through the application of both random forest and the explanatory machine learning non-negative matrix factorization using k-means clustering (NMFk) algorithms to predict re-burn occurrence over California based on a number of climate factors. Our findings indicate that while most states showed increasing trends within the WUI when trends were conducted over longer moving windows (e.g. 20 years), California was the only state where the rate of increase was consistently higher in the WUI, indicating a stronger potential for human activity as a driver in that location. Furthermore, we find model performance was found to be robust over most of California (Testing F1 scores = 0.688), although results were highly variable based on EPA level III Ecoregion (F1 scores = 0.0–0.778). Insights provided from this study will lead to a better understanding of climate and human activity drivers of re-burns and how these vary at broad spatial scales so that improvements in forest management practices can be tuned according to the level of change that is expected for a given region.

54 ENVIRONMENTAL SCIENCES↗

Fusion RF Modeling Machine Learning (FusionML_RF) v1.0

FusionML_RF consists of multiple codes and trained machine learning (ML) models that perform low-cost output modeling from the Genray-CQL3D. Three machine learning techniques (multilayer perceptron, random forest, and Gaussian process) provide fast surrogate models for lower hybrid current drive (LHCD) simulations. For example, completing a single GENRAY/CQL3D simulation without radial diffusion of fast electrons requires several minutes of wall-clock time. On the other hand, these ML models achieve ~ms of inference time with high accuracy across the input parameter space. This software collection consists of multiple components. (1) codes that use ML methods and precomputed Genray-CQL3D simulation output to build regression models that enable approximate computations of Genray-CLQ3D outputs from arbitrary but physically meaningful input parameters (surrogate modeling); (2) three trained models created by the team, using a database of 16,000+ GENRAY/CQL3D simulations, to study the performance of ML models for surrogate modeling; (3) codes that load the trained models and simulation data, and then compute mean squared error between the models' predictions and the ground truth of simulation output data. This collection is being made available in conjunction with a scientific publication about the work to promote reusability and provide an artifact of the scientific work.

Bai, Zhe↗

Machine learning models for maintenance cost estimation in delivery trucks using diesel and natural gas fuels

The maintenance costs can represent about 15%–60% of the cost of produced goods depending on the type of goods transported. To comply with stringent emissions regulations, diesel engines are incorporated with complex after-treatment systems that demand increased maintenance. The availability of alternative fuels such as natural gas and propane has fostered the natural gas and propane powertrain systems as well as electrification options for heavy- and medium-duty vehicles. A critical barrier to adopting alternative fuel vehicles has been the lack of knowledge on comparative vehicle maintenance/repair costs with conventional diesel. Moreover, the region of operation, the type of vehicle operation, and seasonal temperature changes also affect the duty cycle which impacts the maintenance and repair costs. This study focuses on estimating the cost-per-mile for heavy-duty vehicles using machine learning models such as random forest, xgboost, neural networks, and a super-learner model. The super-learner model achieved an error as low as 0.0068 $/mile for mean absolute error and 0.0086 $/mile for root mean square error with a coefficient of determination/R-Squared of 97.28%. Specifically, the paper investigates the data collected from the maintenance and repair costs associated with delivery trucks using diesel and natural gas fuels. Since the availability of data is the major constraint, we leveraged the data collected by West Virginia University and the partnership with fleet companies. This allows for additional information related to maintenance costs and fleet-specific maintenance practices of alternative fuel vehicles. This study promotes clean fuel technologies and enables fleet management companies to adopt alternative fuel vehicles in case of similar or lower cost of maintenance compared to diesel vehicles resulting in reduced emissions and total cost of ownership.

Katreddi, Sasanka↗

Remaining Useful Strength (RUS) Prediction of SiCf-SiCm Composite Materials Using Deep Learning and Acoustic Emission

Prognosis techniques for prediction of remaining useful life (RUL) are of crucial importance to the management of complex systems for they can lead to appropriate maintenance interventions and improvements in reliability. While various data-driven methods have been introduced to predict the remaining useful life (RUL) of machinery systems or batteries, no research has been reported on the remaining useful strength (RUS) prediction of silicon carbide fiber reinforced silicon carbide matrix (SiCf-SiCm) materials with pivotal role in its potential usage as a structural material in nuclear reactors and turbine engines. Knowledge of its degradation process is of the utmost importance to the manufacturers. For this purpose, two approaches based on the machine-learning techniques of random-forest (RF) and convolutional neural network (CNN) are proposed to predict the RUS of SiCf-SiCm using only acoustic emission (AE) signals generated during the material’s stress applying process. Experimental results show that the CNN models achieved better predictive performance than the RF models but the latter with expert-engineered features achieves better prediction for AE signals in the early stage of degradation. Additionally, our results demonstrate that both models can correctly predict the SiCf-SiCm RUS as evaluated by our robust testing method from which the best average root mean square error (RMSE) and Pearson correlation coefficient of 3.55 ksi units and 0.85 were obtained.

36 MATERIALS SCIENCE↗

Machine Learning-Based Classification of Lignocellulosic Biomass from Pyrolysis-Molecular Beam Mass Spectrometry Data

High-throughput analysis of biomass is necessary to ensure consistent and uniform feedstocks for agricultural and bioenergy applications and is needed to inform genomics and systems biology models. Pyrolysis followed by mass spectrometry such as molecular beam mass spectrometry (py-MBMS) analyses are becoming increasingly popular for the rapid analysis of biomass cell wall composition and typically require the use of different data analysis tools depending on the need and application. Here, the authors report the py-MBMS analysis of several types of lignocellulosic biomass to gain an understanding of spectral patterns and variation with associated biomass composition and use machine learning approaches to classify, differentiate, and predict biomass types on the basis of py-MBMS spectra. Py-MBMS spectra were also corrected for instrumental variance using generalized linear modeling (GLM) based on the use of select ions relative abundances as spike-in controls. Machine learning classification algorithms e.g., random forest, k-nearest neighbor, decision tree, Gaussian Naïve Bayes, gradient boosting, and multilayer perceptron classifiers were used. The k-nearest neighbors (k-NN) classifier generally performed the best for classifications using raw spectral data, and the decision tree classifier performed the worst. After normalization of spectra to account for instrumental variance, all the classifiers had comparable and generally acceptable performance for predicting the biomass types, although the k-NN and decision tree classifiers were not as accurate for prediction of specific sample types. Gaussian Naïve Bayes (GNB) and extreme gradient boosting (XGB) classifiers performed better than the k-NN and the decision tree classifiers for the prediction of biomass mixtures. The data analysis workflow reported here could be applied and extended for comparison of biomass samples of varying types, species, phenotypes, and/or genotypes or subjected to different treatments, environments, etc. to further elucidate the sources of spectral variance, patterns, and to infer compositional information based on spectral analysis, particularly for analysis of data without a priori knowledge of the feedstock composition or identity.

59 BASIC BIOLOGICAL SCIENCES↗

Evaluation of Classifier Complexity for Delay Tolerant Network Routing

The growing popularity of small cost effective satellites (SmallSats, CubeSats, etc.) creates the potential for a variety of new science applications involving multiple nodes functioning together or independently to achieve a task, such as swarms and constellations. As this technology develops and is deployed for missions in Low Earth Orbit and beyond, the use of delay tolerant networking (DTN) techniques may improve communication capabilities within the network. In this paper, a network hierarchy is developed from heterogeneous networks of SmallSats, surface vehicles, relay satellites and ground stations which form an integrated network. There is a tradeoff between complexity, flexibility, and scalability of user defined schedules versus autonomous routing as the number of nodes in the network increases. To address these issues, this work proposes a machine learning classifier based on DTN routing metrics. A framework is developed which will allow for the use of several categories of machine learning algorithms (decision tree, random forest and deep learning) to be applied to a dataset of historical network statistics, which allows for the evaluation of algorithm complexity versus performance to be explored. We develop the emulation of a hierarchical network, consisting of tens of nodes which form a cognitive network architecture. CORE (Common Open Research Emulator) is used to emulate the network using bundle protocol and DTN IP neighbor discovery.

Dudukovich, Rachel↗