Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “feature selection”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

How does ion temperature gradient turbulence depend on magnetic geometry? Insights from data and machine learning

Magnetic geometry has a significant effect on the level of turbulent transport in fusion plasmas. Here, we model and analyse this dependence using multiple machine learning methods and a dataset of >200 000 nonlinear gyrokinetic simulations of ion-temperature-gradient turbulence in diverse non-axisymmetric geometries. The dataset is generated using a large collection of both optimised and randomly generated stellarator equilibria. At fixed gradients and other input parameters, the turbulent heat flux varies between geometries by several orders of magnitude. Trends are apparent among the configurations with particularly high or particularly low heat flux. Regression and classification techniques from machine learning are then applied to extract patterns in the dataset. Due to a symmetry of the gyrokinetic equation, the heat flux and regressions thereof should be invariant to translations of the raw features in the parallel coordinate, similar to translation invariance in computer vision applications. Multiple regression models including convolutional neural networks (CNNs) and decision trees can achieve reasonable predictive power for the heat flux in held-out test configurations, with highest accuracy for the CNNs. Using Spearman correlation, sequential feature selection and Shapley values to measure feature importance, it is consistently found that the most important geometric lever on the heat flux is the flux surface compression in regions of bad curvature. The second most important geometric feature relates to the magnitude of geodesic curvature. These two features align remarkably with surrogates that have been proposed based on theory, while the methods here allow a natural extension to more features for increased accuracy. The dataset, released with this publication, may also be used to test other proposed surrogates, and we find that many previously published proxies do correlate well with both the heat flux and stability boundary.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Factors Controlling a Synthetic Aperture Radar (SAR) Derived Root-Zone Soil Moisture Product over The Seward Peninsula of Alaska

Root-zone soil moisture exerts a fundamental control on vegetation, energy balance, and the carbon cycle in Arctic ecosystems, but it is still not well understood in vast, remote, and understudied regions of discontinuous permafrost. The root-zone soil moisture product (30 m resolution) used in this analysis was retrieved from a time-series P-Band (420–440 MHz) synthetic aperture radar (SAR) backscatter observations (August 2017 & October 2017). While similar approaches have been taken to retrieve surface (0 cm to 5 cm) soil moisture from L-Band (1.2 GHz) SAR backscatter, this is one of the first known attempts at reaching the root-zone in permafrost regions. Here, we analyze secondary factors (excluding primary factors, such as precipitation) controlling summer (August) soil moisture at depths of 6 cm, 12 cm, and 20 cm over a 4500 km2 area on the Seward Peninsula of Alaska. Using a random forest model, we quantify the impact of topography, vegetation, and meteorological factors on soil moisture distributions. In developing the random forest model, we explore a variety of feature scales (30 m, 60 m, 90 m, 120 m, 180 m, and 240 m), tune hyperparameters (the structure of individual decision trees making up the ensemble including the number and depth of trees), and perform the final feature selection using cross-validated recursive feature elimination. Results suggest that root-zone soil moisture on the Seward Peninsula is primarily controlled by vegetation at 6 cm, but deeper in the soil column topography and meteorological factors, such as predominant winter wind direction and summer insolation, play a larger role. The random forest model accounts for 40% to 60% of the variation observed (R2 = 0.44 at 6 cm, R2 = 0.52 at 12 cm, R2 = 0.58 at 20 cm). These results indicate that vegetation is the dominant control on soil moisture shallow in the soil column, but the impact of vegetation does not extend to deeper layers retrieved from P-Band SAR backscatter.

Dann, Julian↗

Analysis and Benchmarking of feature reduction for classification under computational constraints

Abstract Machine learning is most often expensive in terms of computational and memory costs due to training with large volumes of data. Current computational limitations of many computing systems motivate us to investigate practical approaches, such as feature selection and reduction, to reduce the time and memory costs while not sacrificing the accuracy of classification algorithms. In this work, we carefully review, analyze, and identify the feature reduction methods that have low costs/overheads in terms of time and memory. Then, we evaluate the identified reduction methods in terms of their impact on the accuracy, precision, time, and memory costs of traditional classification algorithms. Specifically, we focus on the least resource intensive feature reduction methods that are available in Scikit-Learn library. Since our goal is to identify the best performing low-cost reduction methods, we do not consider complex expensive reduction algorithms in this study. In our evaluation, we find that at quadratic-scale feature reduction, the classification algorithms achieve the best trade-off among competitive performance metrics. Results show that the overall training times are reduced 61%, the model sizes are reduced 6×, and accuracy scores increase 25% compared to the baselines on average with quadratic scale reduction.

97 MATHEMATICS AND COMPUTING↗

Landslide Likelihood Prediction using Machine Learning Algorithms

The supply of electricity via power plants is criticalto the operation of many critical infrastructure systems in mod-ern society. Natural hazards can disrupt the power supply, causepower outages that can halt economic growth, and impede emer-gency response until power is restored. The proposed work aimsto predict the landslides likelihood in these critical infrastructurelocations in the Northeastern USA using integrated databases ofexplanatory variables and machine learning algorithms. First,data related to landslides are obtained and merged, includingtopographic, soil moisture, and precipitation-related data. Fiveregression algorithms, namely: Random Forest, Extreme Gradi-ent Boosting (XGBoost), K-Nearest Neighbor regression (KNN),Linear Support Vector Regressor (SVR), and Linear regression,are utilized to predict the landslide probability and evaluatedon the dataset. The accuracy of the models is assessed by usingstatistical metrics such as mean absolute error (MAE), meansquared error (MSE), and root mean squared error (RMSE).The study results show that Random Forest outperformed othermodels with the mutual information feature selection method.It achieved an MSE of 0.0011 with mutual information-basedfeature selection and an MSE of 0.00157 without feature selection.KNN regressor outperformed the other models with an MSEof 0.00139 with correlation-based information selection. Theproposed landslide identification model with Random Forestalgorithm shows outstanding robustness and great potential intackling the landslide likelihood prediction by employing MLalgorithms.

Vasundhara Acharya↗

Hierarchical, Self-Assembled Metasurfaces via Exposure-Controlled Reflow of Block Copolymer-Derived Nanopatterns

Here, nanopatterning for the fabrication of optical metasurfaces entails a need for high-resolution approaches like electron beam lithography that cannot be readily scaled beyond prototyping demonstrations. Block copolymer thin film self-assembly offers an attractive alternative for producing periodic nanopatterns across large areas, yet the pattern feature sizes are fixed by the polymer molecular weight and composition. Here, a general strategy is reported that overcomes the limitation of fixed feature size by treating the copolymer thin film as a hierarchical resist, in which the nanoscale pattern motif is defined by self-assembly. Feature sizes can then be tuned by thermal reflow controlled locally by irradiative crosslinking or chemical alteration using lithographic ultraviolet light or electron beam exposure. Using blends of polystyrene-block-poly(methylmethacrylate) (PS-b-PMMA) with PS and PMMA homopolymer, we demonstrate both self-assembled PS grating and hexagonal hole patterns; exposure-controlled reflow is then used to reduce hole diameter by as much as 50% or increase PS grating linewidth by more than 180%. Transferring these nanopatterns, or their inverse obtained by a lift-off approach, into silicon yields structural colors that may be prescriptively controlled based on nanoscale feature size. Furthermore, patterned exposure enables area-selective feature size control, yielding uniform structural color patterns across centimeter square areas. Electron beam lithography is also used to show that the lithographic resolution of this selective-area control can be extended to the nanoscale dimensions of the self-assembled features. The exposure-controlled reflow approach demonstrated here takes a pivotal step towards fabricating complex, hierarchical optical metasurfaces using scalable self-assembly methods.

77 NANOSCIENCE AND NANOTECHNOLOGY↗

Dimensionality Reduction Through Classifier Ensembles

In data mining, one often needs to analyze datasets with a very large number of attributes. Performing machine learning directly on such data sets is often impractical because of extensive run times, excessive complexity of the fitted model (often leading to overfitting), and the well-known "curse of dimensionality." In practice, to avoid such problems, feature selection and/or extraction are often used to reduce data dimensionality prior to the learning step. However, existing feature selection/extraction algorithms either evaluate features by their effectiveness across the entire data set or simply disregard class information altogether (e.g., principal component analysis). Furthermore, feature extraction algorithms such as principal components analysis create new features that are often meaningless to human users. In this article, we present input decimation, a method that provides "feature subsets" that are selected for their ability to discriminate among the classes. These features are subsequently used in ensembles of classifiers, yielding results superior to single classifiers, ensembles that use the full set of features, and ensembles based on principal component analysis on both real and synthetic datasets.

Oza, Nikunj C.↗

Machine Learning Prediction of the Experimental Transition Temperature of Fe(II) Spin-Crossover Complexes

Spin-crossover (SCO) complexes are materials that exhibit changes in the spin state in response to external stimuli, with potential applications in molecular electronics. It is challenging to know a priori how to design ligands to achieve the delicate balance of entropic and enthalpic contributions needed to tailor a transition temperature close to room temperature. Here, we leverage the SCO complexes from the previously curated SCO-95 data set [Vennelakanti et al. J. Chem. Phys. 159, 024120 (2023)] to train three machine learning (ML) models for transition temperature (T 1/2 ) prediction using graph-based revised autocorrelations as features. We perform feature selection using random forest-ranked recursive feature addition (RF-RFA) to identify the features essential to model transferability. Of the ML models considered, the full feature set RF and recursive feature addition RF models perform best, achieving moderate correlation to experimental T 1/2 values. We then compare ML T 1/2 predictions to those from three previously identified best-performing density functional approximations (DFAs) which accurately predict SCO behavior across SCO-95, finding that the ML models predict T 1/2 more accurately than the best-performing DFAs. In addition, we study ML model predictions for a set of 18 SCO complexes for which only estimated T 1/2 values are available. Upon excluding outliers from this set, the RF-RFA RF model shows a strong correlation to estimated T 1/2 values with a Pearson’s r of 0.82. In contrast, DFA-predicted T 1/2 values have large errors and show no correlation to estimated T 1/2 values over the same set of complexes. Overall, our study demonstrates slightly superior performance of ML models in comparison with some of the best-performing DFAs, and we expect ML models to improve further as larger data sets of SCO complexes are curated and become available for model training.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Machine Learning Self-Diffusion Prediction for Lennard-Jones Fluids in Pores

Predicting the diffusion coefficient of fluids under nanoconfinement is important for many applications including the extraction of shale gas from kerogen and product turnover in porous catalysts. Due to the large number of important variables, including pore shape and size, fluid temperature and density, and the fluid–wall interaction strength, simulating diffusion coefficients using molecular dynamics (MD) in a systematic study could prove to be prohibitively expensive. Here, we use machine learning models trained on a subset of MD data to predict the self-diffusion coefficients of Lennard-Jones fluids in pores. Our MD data set contains 2280 simulations of ideal slit pore, cylindrical pore, and hexagonal pore geometries. We use the forward feature selection method to determine the most useful features (i.e., descriptors) for developing an artificial neutral network (ANN) model with an emphasis on easily acquired features. Our model shows good predictive ability with a coefficient of determination (i.e., R 2 ) of ~0.99 and a mean squared error of ~2.9 × 10 –5 . Finally, we propose an alteration to our feature set that will allow the ANN model to be applied to nonideal pore geometries.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Quantile regression-enriched event modeling framework for dropout analysis in high-temperature superconductor manufacturing

High-temperature superconductor (HTS) tapes have shown promising characteristics of high critical current, which are prerequisites for applications in high-field magnets. Due to the unstable growth conditions in the HTS manufacturing process, however, the frequent occurrences of dropouts in the critical current impede the consistent performance of HTS tapes. To manufacture HTS tapes with large scale, high yield, and uniform performance, it is essential to develop novel data analysis approaches for modeling the dropouts and identifying the related important process parameters. Conventional methods for modeling recurrent events, such as the point process, require the extraction of events from quality measurements. As the critical current is a continuous process, it may not comprehensively represent the drop patterns by transforming the time-series measurements into a set of events. Here, to solve this issue, we develop a novel quantile regression-enriched event modeling (QREM) framework that integrates the non-homogeneous Poisson process for modeling the occurrence of dropouts and the quantile regression for capturing the drop patterns. By incorporating the feature selection and regularization, the proposed framework identifies a set of significant process parameters that can potentially cause the dropouts of HTS tapes. The proposed method is tested on real HTS tapes produced using an advanced manufacturing process, successfully identifying important parameters that influence dropout events including the substrate temperature and voltage. The results demonstrate that the proposed QREM method outperforms the standard point process in predicting the occurrence of dropouts.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

Localized keyhole pore prediction during laser powder bed fusion via multimodal process monitoring and X-ray radiography

Systematic fault detection and control during laser powder bed fusion (L-PBF) has been a long-standing objective for system manufacturers and researchers in the additive manufacturing (AM) industry. This manuscript investigates a data fusion approach for detection of keyhole porosity formation during laser irradiation of Ti-6Al-4V substrates by concurrent recording of thermally induced optical emission measured using both off-axis and coaxial photodiode sensors, and acoustic emission. Subsurface defect formation was monitored via high-speed synchrotron X-ray imaging at 20,000 frames per second, enabling temporal registration of keyhole pore formation events to the monitoring signals at a resolution of 50 µs. We developed data fusion machine learning (ML) models for localized prediction of keyhole pore formation at various time scales ranging from 0.5 ms to 2 ms. The signal segments were featurized using two independent approaches: (1) power spectral density (PSD) and (2) highly comparative time series analysis (HCTSA) framework. The extracted features from different sensor modalities were fused together to construct a multimodal feature space and sequential feature selection was used to determine the most informative features for training the ML models. The predictive performance was evaluated for three classifying algorithms: Support Vector Machine (SVM), K-Nearest Neighbor (KNN), and Gaussian Naive Bayes (GNB). As a result, pore formation events were predicted with up to 0.95 F1-score, 1.0 recall and 0.94 accuracy. The most heavily weighted features indicate that model performance is chiefly governed by the acoustic monitoring signal, with a secondary contribution from the optical emission sensors.

36 MATERIALS SCIENCE↗

Constructing a Database from Multiple 2D Images for Camera Pose Estimation and Robot Localization

The LMDB (Landmark Database) Builder software identifies persistent image features (landmarks) in a scene viewed multiple times and precisely estimates the landmarks 3D world positions. The software receives as input multiple 2D images of approximately the same scene, along with an initial guess of the camera poses for each image, and a table of features matched pair-wise in each frame. LMDB Builder aggregates landmarks across an arbitrarily large collection of frames with matched features. Range data from stereo vision processing can also be passed to improve the initial guess of the 3D point estimates. The LMDB Builder aggregates feature lists across all frames, manages the process to promote selected features to landmarks, and iteratively calculates the 3D landmark positions using the current camera pose estimations (via an optimal ray projection method), and then improves the camera pose estimates using the 3D landmark positions. Finally, it extracts image patches for each landmark from auto-selected key frames and constructs the landmark database. The landmark database can then be used to estimate future camera poses (and therefore localize a robotic vehicle that may be carrying the cameras) by matching current imagery to landmark database image patches and using the known 3D landmark positions to estimate the current pose.

Wolf, Michael↗

MODSIM World 2007 Conference and Expo: Select Papers and Presentations from the Education and Training Track

This NASA Conference Publication features select papers and PowerPoint presentations from the Education and Training Track of MODSIM World 2007 Conference and Expo. Invited speakers and panelists of national and international renown, representing academia, industry and government, discussed how modeling and simulation (M&S) technology can be used to accelerate learning in the K-16 classroom, especially when using M&S technology as a tool for integrating science, technology, engineering and mathematics (STEM) classes. The presenters also addressed the application ofM&S technology to learning and training outside of the classroom. Specific sub-topics of the presentations included: learning theory; curriculum development; professional development; tools/user applications; implementation/infrastructure/issues; and workforce development. There was a session devoted to student M&S competitions in Virginia too, as well as a poster session.

Pinelli, Thomas E.↗

A Confidence-Guided Technique for Tracking Time-Varying Features

Application scientists often employ feature tracking algorithms to capture the temporal evolution of various features in their simulation data. However, as the complexity of the scientific features is increasing with the advanced simulation modeling techniques, quantification of reliability of the feature tracking algorithms is becoming important. One of the desired requirements for any robust feature tracking algorithm is to estimate its confidence during each tracking step so that the results obtained can be interpreted without any ambiguity. To address this, we develop a confidence-guided feature tracking algorithm that allows reliable tracking of user-selected features and presents the tracking dynamics using a graph-based visualization along with the spatial visualization of the tracked feature. Here, the efficacy of the proposed method is demonstrated by applying it to two scientific datasets containing different types of time-varying features.

97 MATHEMATICS AND COMPUTING↗

Automatic variable selection in ecological niche modeling: A case study using Cassin’s Sparrow (Peucaea cassinii)

MERRA/Max provides a feature selection approach to dimensionality reduction that enables direct use of global climate model outputs in ecological niche modeling. The system accomplishes this reduction through a Monte Carlo optimization in which many independent MaxEnt runs, operating on a species occurrence file and a small set of randomly selected variables in a large collection of variables, converge on an estimate of the top contributing predictors in the larger collection. These top predictors can be viewed as potential candidates in the variable selection step of the ecological niche modeling process. MERRA/Max’s Monte Carlo algorithm operates on files stored in the underlying filesystem, making it scalable to large data sets. Its software components can run as parallel processes in a high-performance cloud computing environment to yield near real-time performance. In tests using Cassin’s Sparrow (Peucaea cassinii) as the target species, MERRA/Max selected a set of predictors from Worldclim’s Bioclim collection of 19 environmental variables that have been shown to be important determinants of the species’ bioclimatic niche. It also selected biologically and ecologically plausible predictors from a more diverse set of 86 environmental variables derived from NASA’s Modern-Era Retrospective Analysis for Research and Applications Version 2 (MERRA-2) reanalysis, an output product of the Goddard Earth Observing System Version 5 (GEOS-5) modeling system. We believe these results point to a technological approach that could expand the use global climate model outputs in ecological niche modeling, foster exploratory experimentation with otherwise difficult-to-use climate data sets, streamline the modeling process, and, eventually, enable automated bioclimatic modeling as a practical, readily accessible, low-cost, commercial cloud service.

John L. Schnase↗

Better understanding and prediction of antiviral peptides through primary and secondary structure feature importance

The emergence of viral epidemics throughout the world is of concern due to the scarcity of available effective antiviral therapeutics. The discovery of new antiviral therapies is imperative to address this challenge, and antiviral peptides (AVPs) represent a valuable resource for the development of novel therapies to combat viral infection. We present a new machine learning model to distinguish AVPs from non-AVPs using the most informative features derived from the physicochemical and structural properties of their amino acid sequences. To focus on those features that are most likely to contribute to antiviral performance, we filter potential features based on their importance for classification. These feature selection analyses suggest that secondary structure is the most important peptide sequence feature for predicting AVPs. Our Feature-Informed Reduced Machine Learning for Antiviral Peptide Prediction (FIRM-AVP) approach achieves a higher accuracy than either the model with all features or current state-of-the-art single classifiers. Understanding the features that are associated with AVP activity is a core need to identify and design new AVPs in novel systems. The FIRM-AVP code and standalone software package are available at https://github.com/pmartR/FIRM-AVP with an accompanying web application at https://msc-viz.emsl.pnnl.gov/AVPR.

59 BASIC BIOLOGICAL SCIENCES↗

Transcriptomics-based Machine Learning Analysis Predicts Space-Exposed Murine Livers

Limited sample sizes, high data dimensionality, and sensitivity to technical and biological variability of next generation sequencing (NGS), typically limits machine learning (ML) approaches in spaceflight studies that include radiation effects. However, pooling smaller studies while addressing intra- and inter-study variabilities allows for ML predictive modeling. Here, integration methods were applied to whole transcriptome shotgun sequencing (RNA-seq) data from six mouse liver GeneLab datasets (GLDS) (n ranging from 6 to 39 samples) from with a total of 81 spaceflight and ground-control samples to determine top features (i.e. genes) relevant to spaceflight including the effect of radiation exposure. RNASeq counts were normalized for each study, then merged and scaled across all datasets. Data dimensionality was reduced using a minimum redundancy maximum relevance (MRMR) methodology. Redundancy and relevance were computed using the Pearson correlation and F-statistic, respectively. The top 100 MRMR features were used to predict spaceflight vs. ground-control samples using Random Forest (RF), Support Vector Machine (SVM), and Linear Discriminant Analysis (LDA) classifiers with 5-fold cross validation (CV). Principal component analysis (PCA) on the complete feature set versus the MRMR features shows separation between spaceflight samples and ground controls (Figure 1A). The ML-based gene sets were compared against differential gene expression results obtained with DESeq2 from individual GLDS. Using all features or randomly sampled subsets at matching set sizes with MRMR, a maximum classifier accuracy of 69% was shown on the test set over 5 folds. For all classifiers, CV training using at least the top 30 MRMR genes show minimum 89% accuracy and 0.95 AUC value on the test set over 5 folds (Figure 1B). Baseline set analysis on differentially expressed genes (DEGs) identified using padj ≤ 0.05 show 295 DEGs that overlap at least two studies and 13 DEGs that overlap three studies (Figure 1C). Set analysis between the top 100 MRMR features and the DEGs showed 47 genes that overlap at least one study and 24 genes that overlap two studies. Over-representation analysis showed overlapping biological processes related to fatty acid and lipid metabolism which may indicate these processes in the response to spaceflight stressors. MRMR feature selection for the selected ML methods improve performance relative to a classifier built on all features or randomly sampled subsets. Permutation feature importance within the decorrelated MRMR features showed concordance in feature ranking between ML methods. A challenge of applying ML methods across heterogeneous NGS data is accounting for signal:noise. Here, signal validation across studies was shown by intersecting sets between top MRMR genes and DEGs from DESeq2 analysis. Non-intersecting sets introduce opportunity to explore genes relevant to differentiating space flight exposed groups and implementing ML methods across existing NGS datasets may overcome sample size limitations.

Machine Learning↗

Transcriptomics-based Machine Learning (ML) Analysis Predicts Space-Exposed Murine Livers

Limited sample sizes, high data dimensionality, and sensitivity to technical and biological variability of next generation sequencing (NGS), typically limits machine learning (ML) approaches in spaceflight studies that include radiation effects. However, pooling smaller studies while addressing intra- and inter-study variabilities allows for ML predictive modeling. Here, integration methods were applied to whole transcriptome shotgun sequencing (RNA-seq) data from six mouse liver GeneLab datasets (GLDS) (n ranging from 6 to 39 samples) from with a total of 81 spaceflight and ground-control samples to determine top features (i.e. genes) relevant to spaceflight including the effect of radiation exposure. RNASeq counts were normalized for each study, then merged and scaled across all datasets. Data dimensionality was reduced using a minimum redundancy maximum relevance (MRMR) methodology. Redundancy and relevance were computed using the Pearson correlation and F-statistic, respectively. The top 100 MRMR features were used to predict spaceflight vs. ground-control samples using Random Forest (RF), Support Vector Machine (SVM), and Linear Discriminant Analysis (LDA) classifiers with 5-fold cross validation (CV). Principal component analysis (PCA) on the complete feature set versus the MRMR features shows separation between spaceflight samples and ground controls (Figure 1A). The ML-based gene sets were compared against differential gene expression results obtained with DESeq2 from individual GLDS. Using all features or randomly sampled subsets at matching set sizes with MRMR, a maximum classifier accuracy of 69% on the test set over 5 folds. For all classifiers, CV training using at least the top 30 MRMR genes show minimum 89% accuracy and 0.95 AUC value on the test set over 5 folds (Figure 1B). Baseline set analysis on differentially expressed genes (DEGs) identified using padj ≤ 0.05 show 295 DEGs that overlap at least two studies and 13 DEGs that overlap three studies (Figure 1C). Set analysis between the top 100 MRMR features and the DEGs showed 47 genes that overlap at least one study and 24 genes that overlap two studies. Over-representation analysis showed overlapping biological processes related to fatty acid and lipid metabolism which may indicate these processes in the response to spaceflight stressors. MRMR feature selection for the selected ML methods improve performance relative to a classifier built on all features or randomly sampled subsets. Permutation feature importance within the decorrelated MRMR features showed concordance in feature ranking between ML methods. A challenge of applying ML methods across heterogeneous NGS data is accounting for signal:noise. Here, signal validation across studies was shown by intersecting sets between top MRMR genes and DEGs from DESeq2 analysis. Non-intersecting sets introduce opportunity to explore genes relevant to differentiating space flight exposed groups and implementing ML methods across existing NGS datasets may overcome sample size limitations.

Machine Learning↗

Transcriptomics-based Machine Learning Analysis Predicts Space-Exposed Murine Livers

Limited sample sizes, high data dimensionality, and sensitivity to technical and biological variability of next generation sequencing (NGS), typically limits machine learning (ML) approaches in spaceflight studies that include radiation effects. However, pooling smaller studies while addressing intra- and inter-study variabilities allows for ML predictive modeling. Here, integration methods were applied to whole transcriptome shotgun sequencing (RNA-seq) data from six mouse liver GeneLab datasets (GLDS) (n ranging from 6 to 39 samples) from with a total of 81 spaceflight and ground-control samples to determine top features (i.e. genes) relevant to spaceflight including the effect of radiation exposure. RNASeq counts were normalized for each study, then merged and scaled across all datasets. Data dimensionality was reduced using a minimum redundancy maximum relevance (MRMR) methodology. Redundancy and relevance were computed using the Pearson correlation and F-statistic, respectively. The top 100 MRMR features were used to predict spaceflight vs. ground-control samples using Random Forest (RF), Support Vector Machine (SVM), and Linear Discriminant Analysis (LDA) classifiers with 5-fold cross validation (CV). Principal component analysis (PCA) on the complete feature set versus the MRMR features shows separation between spaceflight samples and ground controls (Figure 1A). The ML-based gene sets were compared against differential gene expression results obtained with DESeq2 from individual GLDS. Using all features or randomly sampled subsets at matching set sizes with MRMR, a maximum classifier accuracy of 69% was shown on the test set over 5 folds. For all classifiers, CV training using at least the top 30 MRMR genes show minimum 89% accuracy and 0.95 AUC value on the test set over 5 folds (Figure 1B). Baseline set analysis on differentially expressed genes (DEGs) identified using padj ≤ 0.05 show 295 DEGs that overlap at least two studies and 13 DEGs that overlap three studies (Figure 1C). Set analysis between the top 100 MRMR features and the DEGs showed 47 genes that overlap at least one study and 24 genes that overlap two studies. Over-representation analysis showed overlapping biological processes related to fatty acid and lipid metabolism which may indicate these processes in the response to spaceflight stressors. MRMR feature selection for the selected ML methods improve performance relative to a classifier built on all features or randomly sampled subsets. Permutation feature importance within the decorrelated MRMR features showed concordance in feature ranking between ML methods. A challenge of applying ML methods across heterogeneous NGS data is accounting for signal:noise. Here, signal validation across studies was shown by intersecting sets between top MRMR genes and DEGs from DESeq2 analysis. Non-intersecting sets introduce opportunity to explore genes relevant to differentiating space flight exposed groups and implementing ML methods across existing NGS datasets may overcome sample size limitations.

Machine Learning↗