Engineering PapersSearch

SEARCH · Engineering Papers

Results for “Random forest”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Using Random Forest Models to Predict Organizational Violence

We present a methodology to access the proclivity of an organization to commit violence against nongovernment personnel. We fitted a Random Forest model using the Minority at Risk Organizational Behavior (MAROS) dataset. The MAROS data is longitudinal; so, individual observations are not independent. We propose a modification to the standard Random Forest methodology to account for the violation of the independence assumption. We present the results of the model fit, an example of predicting violence for an organization; and finally, we present a summary of the forest in a "meta-tree,"

Levine, Burton

Mapping tree canopy cover and canopy height with L-band SAR using LiDAR data and Random Forests

Light detection and ranging (LiDAR) data can provide direct measurements of vegetation structures but are limited by the sparse spatial coverage. Polarimetric synthetic aperture radar (SAR) can perform large-scale high-resolution mapping without weather constraints but the information about vegetation and ground subsurface are mixed in the backscatter data. In this paper, we adopted the Random Forests algorithm to train an upscaling function using tree canopy cover (TCC) and canopy height model (CHM) derived from Goddard’s LiDAR, Hyperspectral and Thermal Imager (G-LiHT) data. The regression model is then applied to the L-band Uninhabited Aerial Vehicle Synthetic Aperture Radar (UAVSAR) data acquired during the 2017 Arctic-Boreal Vulnerability Experiment (ABoVE) airborne campaign to map the TCC and CHM over the Delta Junction area in interior Alaska.

Moghaddam, Mahta

Classifying Microwave Radiometer Observations Over the Netherlands Into Dry, Shallow-, and Non-Shallow Precipitation Using A Random Forest Model

Spaceborne microwave radiometers represent an important component of the Global Precipitation Measurement (GPM) mission due to their frequent sampling of rain systems. Microwave radiometers measure microwave radiation (brightness temperatures, Tb), which can be converted into precipitation estimates with appropriate assumptions. However, detecting shallow precipitation systems using space-borne radiometers is challenging, especially over land, as their weak signals are hard to differentiate from those associated with dry conditions. This study uses a random forest model (RF) to classify microwave radiometer observations as dry, shallow, or nonshallow over the Netherlands - a region with varying surface conditions and frequent occurrence of shallow precipitation. The RF is trained on five years of data (2016-2020) and tested with two independent years (2015, 2021). The observations are classified using ground-based weather radar echo top heights. Various RF models are assessed, such as using only GPM’s Microwave Imager (GMI) Tb values as input features or including spatially aligned ERA-5 2-meter temperature and freezing level reanalysis and/or Dual Precipitation Radar (DPR) observations. Independent of the input features, the model performs best in summer and worst in winter. The model classifies observations from high-frequency channels (≥85 GHz) with lower Tb-values as non-shallow, higher values as dry, and those in between as shallow. Misclassified footprints exhibit radiometric characteristics corresponding to their assigned class. Case studies reveal dry observations misclassified as shallow are associated with lower Tb-values, likely resulting from the presence of ice particles in non-precipitating clouds. Shallow footprints misclassified as dry are likely related to the absence of ice particles.

Linda Bogerd

Mapping Inundation from Hurricane Florence (2018) with L-Band Synthetic Aperture Radar, Commercial Imagery, and Ancillary Data via Random Forest Classification

Mapping the extent of floodwaters following extreme rainfall aids in the distribution of resources, recovery efforts, and damage assessment practices. Development of a land cover classification system focused on mapping inundation after major hurricane events using synthetic aperture radar (SAR) data could allow for the production of near-real-time inundation mapping, enabling government and emergency response entities to get a preliminary idea of a developing situation. In response to Hurricane Florence of 2018, NASA JPL collected numerous swaths of quad-pol L-band SAR data with the Uninhabited Aerial Vehicle Synthetic Aperture Radar (UAVSAR) instrument observing the record-setting river stages across North and South Carolina. The resulting fully-polarized SAR images allow for mapping of inundation extent at a high spatial resolution with a unique advantage over optical imaging stemming from the sensor’s ability to penetrate cloud cover and dense vegetation. This study seeks to determine how accurately maps of inundation can be generated from L-band SAR imagery through Random Forest classification. Once the extent of water and inundated vegetation is classified, cleanup operations are performed using fuzzy logic to reduce false detections. Estimates of water extent are then combined with datasets describing the distribution of population, buildings, and roads throughout the domain to evaluate societal impacts. Results from the Hurricane Florence case study will be discussed along with the limitations of available validation data for assessment of the classifier’s accuracy.

Alexander Melancon

Microstructure Quantification and Random Forest Regression Models for Li4Ti5O12–Ni Property Prediction

All-solid-state structural lithium-ion batteries are sought to enable all-electric propulsion in next generation aerospace concepts through improved safety and systems level weight savings. In this work, the influence of processing conditions on microstructural evolution was evaluated for anode composites of strain-free Li4Ti5O12 and metallic nickel current collector. Beyond size distributions, this study explored methods of quantifying microstructural features that describe changes in the spatial distribution and coalescence of nickel particles as a function of sample composition and sintering conditions. Processing-microstructure-property relationships were described by microstructure quantifiers including nickel particle count per area, nearest neighbor distance distribution, and edge-to-edge distance distribution. Machine learning methods were applied to compare the relative influence of processing conditions and microstructural features on electrical conductivity and mechanical strength to optimize for simultaneous energy storage and load bearing performance. Insights gained from this work inform future evaluation of alternative energy storage materials and microstructures for multifunctional performance, and generation of microstructural descriptors strengthens modeling across length scales.

anode

Bhutan Agriculture: Developing a Crop Mask for Rice and Creating a Data Collection Protocol Utilizing Remotely Sensed Data in Bhutan

Rice cultivation in Bhutan has been increasingly threatened by deteriorating soil health and outbreaks of diseases and pests associated with the global change in climate patterns. Field surveys, which the national government of Bhutan has relied on to monitor remote agricultural lands, are becoming increasingly overwhelmed by growing threats to agricultural health. To address these concerns, NASA DEVELOP partnered with the Department of Agriculture of Bhutan, the Bhutan Foundation, and the Ugyen Wangchuck Institute of Conservation and Environmental Research (UWICER) and worked to increase the government of Bhutan’s agricultural monitoring capacity. Utilizing Earth observations including Landsat 8 Operational Land Imager (OLI), Sentinel-1 C-band Synthetic Aperture Radar (C-SAR), Shuttle Radar Topography Mission (SRTM), and Planet imagery, the DEVELOP team worked with NASA SERVIR and created a sampling protocol to identify rice plantations and supplement field surveys for more efficient agriculture monitoring. The analysis focused on districts Paro, Punakha, Samtse, Sarpang, Trongsa, Zhemgang, Wangdue Phodrang, and Samdrup Jongkhar in the year 2020 during the period of transplantation (June) to harvesting of rice (November). The team provided the partners with a sampling protocol for integrating NASA Earth observations into their crop monitoring methods, as well as a crop mask for rice identification and to aid crop management. The crop mask for rice was developed using the Random Forest (RF) classifier for the eight districts of Bhutan. Visually, the random forest model has proved to be more accurate and precise than the classification and Regression Tree model. Statistically, the Random Forest model was 91.8% accurate in identifying rice in Bhutan.

Yeshey Seldon

Benchmarking Bayesian Optimization Frameworks and Acquisition Strategies for Materials Discovery and Autonomous Laboratories

Bayesian optimization (BO) can accelerate materials discovery by guiding expensive experiments toward the most promising processing conditions. We systematically compare five BO surrogate and framework combinations (Gaussian processes in Ax, Gaussian processes and Monte-Carlo neural networks in BayBE, random forests in Lolopy, and tree-structured Parzen (TPE) estimators in Hyperopt) on three benchmarks that mimic common materials design tasks (a discrete solid-electrolyte composition space, a hybrid discrete/continuous laminate-composite design problem solved with micromechanics modeling, and the continuous Ishigami analytic function which is a standard optimization benchmark). Each BO surrogate is paired with posterior mean, probability of improvement, and expected improvement acquisition functions and run for 100 trials from randomized initial samples with uniform random search providing a control. Across five random seeds per setting, BayBE’s Gaussian-process surrogate with expected improvement consistently reached ≥95 % of the known optimum in the fewest evaluations, while Lolopy’s random forest matched or exceeded GP performance on purely categorical or mixed spaces at a higher computational cost. Posterior mean alone often stagnated at local optima, underscoring the need for exploration, whereas probability and expected improvement balanced exploration and exploitation leading to better optimization in fewer trials. Execution times ranged from milliseconds for TPE to minutes for neural-network and random-forest surrogates. These results establish baseline expectations for BO in automated materials laboratories and highlight expected improvement with Gaussian processes as a reliable first choice, with random forests offering a strong alternative when categorical variables dominate. The benchmark suite and code are released to facilitate future surrogate, acquisition, and constraint-handling research in data-driven materials optimization.

Bayesian optimization

X-ray Spectra and Multiwavelength Machine Learning Classification for Likely Counterparts toFermi3FGL Unassociated Sources

We conduct X-ray spectral fits on 184 likely counterparts to Fermi-LAT 3FGL unassociated sources. Characterization and classification of these sources allows for more complete population studies of the high-energy sky. Most of these X-ray spectra are well fit by an absorbed power law model, as expected for a population dominated by blazars and pulsars. A small subset of 7 X-ray sources ave spectra unlike the power law expected from a blazar or pulsar and may be linked to coincident stars or background emission. We develop a multiwavelength machine learning classifier to categorize unassociated sources into pulsars and blazars using gamma- and X-ray observations. Training a random forest procedure with known pulsars and blazars, we achieve a cross-validated classification accuracy of 98.6%. Applying the random forest routine to the unassociated sources returned 126 likely blazar candidates (defined as P(bzr) ≥ 90%) and 5 likely pulsar candidates (P(bzr) ≤ 10%). Our new X-ray spectral analysis does not drastically alter the random forest classifications of these sources compared to previous works, but it builds a more robust classification scheme and highlights the importance of X-ray spectral fitting. Our procedure can be further expanded with UV, visual, or radio spectral parameters or by measuring flux variability.

Stephen Kerby

Machine Learning Application to Atmospheric Chemistry Modeling

Atmospheric chemistry is a high-dimensionality, large-data problem and thus may be suited to machine-learning algorithms. We show here the potential of a random forest regression algorithm to replace the gas-phase chemistry solver in the GEOS-Chem chemistry model. In this proof-of-concept study, we used one month of model output to train random forest regression models to predict the concentrations of each long-lived chemical species after integration based upon the physical and chemical conditions before the chemical integration. The choice of prediction type has a strong impact on the skill of the regression model. We find best results from predicting the change in concentration for very long-lived species and the absolute concentration for shorter lived species. The skill of the machine learning algorithm is further improved by using a family approach for NO and NO2 rather than treating them independently.By replacing the numerical integrator with the random forest algorithm and running this model for one month, we find that the model is able to reproduce many of the features of the reference chemistry simulation. Replacing the integration methodology with a machine learning algorithm has the potential to be substantially faster. There are a wide range of applications for such an approach, e.g. to generate boundary conditions, for use in air quality forecasts or chemical data assimilation systems, etc.

Keller, Christoph A.

Diagnosis of Antarctic Blowing Snow Properties Using MERRA-2 Reanalysis with a Machine Learning Model

This paper presents the work on using a machine learning model to diagnose Antarctic blowing snow (BLSN) properties with the Modern Era Retrospective analysis for Research and Applications v2 (MERRA-2) data. We adopt the random forest classifier for BLSN identification and the random forest regressor for BLSN optical depth and height diagnosis. BLSN properties observed from the Cloud-Aerosol Lidar and Infrared Pathfinder Satellite Observation (CALIPSO) are used as the truth for training the model. Using MERRA-2 fields such as snow age, surface elevation and pressure, temperature, specific humidity, and temperature gradient at the 2m level, and wind speed at the 10m level as input, reasonable results are achieved. Hourly blowing snow property diagnostics are generated with the trained model. Using the year 2010 as an example, it is shown that the Antarctic BLSN frequency is much higher over East than West Antarctica. High frequency months are from April to September, during which BLSN frequency exceeds 20% over East Antarctica. For May 2010, the BLSN snow frequency in the region is as high as 37%. Due to the suppression by strong surface-based inversions, larger values of BLSN height and optical depth are usually limited to the coastal regions, wherein the strength of surface-based inversions is weaker.

Antarctic

Nearest-Neighbor Machine Learning Feature Selection for Interpretation of Microbial Molecular Signatures from Isotope Ratio Mass Spectrometry Data

Mass spectrometry (MS) promises to be a powerful tool for potential biosignature detection during astrobiological missions on ocean worlds in our solar system. Accurate and generalizable machine learning methods could enhance science return on investment by predicting seawater chemistry and classifying isotopic biosignatures, either as a signature consistent with microbial life (biotic) or as a novelty (unclassified/unique). However, machine learning models are likely to be complex and involve interactions between MS features, making biosignatures difficult to interpret. Feature selection methods provide biological and chemical context that help interpret the mechanisms of machine learning models, but these methods also need the ability to detect complex interactions. Previously, we developed a machine learning feature selection algorithm called nearest-neighbor projected distance regression (NPDR) that has the ability to identify important model features that involve complex interactions and automatically reduce correlation and the dimensionality in a high-dimensional variable space. The standard distance metrics used in NPDR – Manhattan and Euclidean – assume the multivariate data are isotropic, which is often violated in real data due to differences in the covariance between variables. Thus, we extend NPDR to include a random forest distance, and other anisotropic distance metrics, for computing nearest neighbors. We also augment the isotope-ratio MS data with time-series features from the raw MS signal to improve biotic classification. We test NPDR on our novel experimental ocean world seawater analog MS data. We measure isotope fractionations of volatile CO 2 that could be measured in exospheres or plumes. Samples include baseline abiotic conditions using a range of possible seawater chemistry consistent with Europa and Enceladus, and biotic samples that include microbes in these seawaters. We use penalized NPDR with random forest proximity to identify interpretable microbial molecular signatures. We compare features with random forest importance, and we train a classifier that discriminates between biotic and abiotic samples with high accuracy. These ML-trained ocean-world analog MS data could be used to assist in identifying biosignatures during future missions.

geochemistry

Classifying Forest Type in the National Forest Inventory Context with Airborne Hyperspectral and Lidar Data

Forest structure and composition regulate a range of ecosystem services, including biodiversity, water and nutrient cycling, and wood volume for resource extraction. Forest type is an important metric measured in the US Forest Service Forest Inventory and Analysis (FIA) program, the national forest inventory of the USA. Forest type information can be used to quantify carbon and other forest resources within specific domains to support ecological analysis and forest management decisions, such as managing for disease and pests. In this study, we developed a methodology that uses a combination of airborne hyperspectral and lidar data to map FIA-defined forest type between sparsely sampled FIA plot data collected in interior Alaska. To determine the best classification algorithm and remote sensing data for this task, five classification algorithms were tested with six different combinations of raw hyperspectral data, hyperspectral vegetation indices, and lidar-derived canopy and topography metrics. Models were trained using forest type information from 632 FIA subplots collected in interior Alaska. Of the thirty model and input combinations tested, the random forest classification algorithm with hyperspectral vegetation indices and lidar-derived topography and canopy height metrics had the highest accuracy (78% overall accuracy). This study supports random forest as a powerful classifier for natural resource data. It also demonstrates the benefits from combining both structural (lidar) and spectral (imagery) data for forest type classification.

random forest

Bhutan Agriculture III: Monitoring Cropland Changes in Bhutan using Remote Sensing to Bolster Food Security and Support Crop Monitoring

The Bhutan Agriculture III team aimed to improve agricultural efficiency in Bhutan. Bhutan is a nation heavily reliant on agriculture, but it faces challenges such as geophysical limitations and lack of scientific agricultural practice. The team partnered with a primary end user, Bhutan’s Department of Agriculture (DoA), and with collaborators; the Bhutan Foundation, National Plant Protection Centre (NPPC), Agricultural Research Department Centre (ARDC), National Statistics Bureau (NSB), and the Ugyen Wangchuck Institute for Conservation and Environment Research (UWICER). Advised by NASA SERVIR, the team developed crop masks and monitored rice distribution from 2015 to 2022 utilizing Earth observations such as Landsat 8 Operational Land Imager (OLI), Landsat 9 OLI-2, Sentinel-1 C-Band Synthetic Aperture Radar (C-SAR), Sentinel-2 MultiSpectral Instrument (MSI) and Shuttle Radar Topography Mission (SRTM). The team gathered 5,000 points from the five dzongkhags that yield the most rice in Bhutan (Paro, Punakha, Samtse, Sarpang and Wangue Phodrang) using Collect Earth Online (CEO). With the data collected, the team split the data into training and validation data on Google Earth Engine (GEE) for a random forest (RF) classifier for rice and non-rice classification. After running the data on the Random Forest (RF) model, the team got an accuracy score of 81.48%, a kappa score of 55.75% and an F1 score of 86.11%. This data supports better agricultural decision-making for the governing body of Bhutan, helps enhance farming efficiency and foster sustainable practices, assists in overcoming data inaccuracy and bolsters food security in the country.

Sonam Seldon Tshering

Design of Materials with Alchemite

Machine learning models that establish the relationships between materials processing and properties can enable inverse design of materials through active learning. Alchemite is a commercial software that can perform inverse materials design on sparse data. Here we evaluate Alchemite’s performance on a dataset of shape memory alloys and a dataset of heat exchangers compared to baseline random forest models. Alchemite had higher accuracy when making predictions on sparse data and was more accurate or nearly as accurate as random forests on complete datasets while also quantifying uncertainty. The software was also used to suggest processing steps and design parameters to optimize properties and performance; however, physical validation of the suggested design parameters was beyond the scope of this work. Several useful design insights were gained about the impact of the design parameters on properties and performance including the importance of dopant choice and amount for shape memory alloys and the importance of height and weight on the thermal resistance of heat exchangers.

Machine learning