Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Tree Classifiers”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

A new clustering algorithm applicable to multispectral and polarimetric SAR images

We describe an application of a scale-space clustering algorithm to the classification of a multispectral and polarimetric SAR image of an agricultural site. After the initial polarimetric and radiometric calibration and noise cancellation, we extracted a 12-dimensional feature vector for each pixel from the scattering matrix. The clustering algorithm was able to partition a set of unlabeled feature vectors from 13 selected sites, each site corresponding to a distinct crop, into 13 clusters without any supervision. The cluster parameters were then used to classify the whole image. The classification map is much less noisy and more accurate than those obtained by hierarchical rules. Starting with every point as a cluster, the algorithm works by melting the system to produce a tree of clusters in the scale space. It can cluster data in any multidimensional space and is insensitive to variability in cluster densities, sizes and ellipsoidal shapes. This algorithm, more powerful than existing ones, may be useful for remote sensing for land use.

Wong, Yiu-Fai↗

Feature Acquisition with Imbalanced Training Data

This work considers cost-sensitive feature acquisition that attempts to classify a candidate datapoint from incomplete information. In this task, an agent acquires features of the datapoint using one or more costly diagnostic tests, and eventually ascribes a classification label. A cost function describes both the penalties for feature acquisition, as well as misclassification errors. A common solution is a Cost Sensitive Decision Tree (CSDT), a branching sequence of tests with features acquired at interior decision points and class assignment at the leaves. CSDT's can incorporate a wide range of diagnostic tests and can reflect arbitrary cost structures. They are particularly useful for online applications due to their low computational overhead. In this innovation, CSDT's are applied to cost-sensitive feature acquisition where the goal is to recognize very rare or unique phenomena in real time. Example applications from this domain include four areas. In stream processing, one seeks unique events in a real time data stream that is too large to store. In fault protection, a system must adapt quickly to react to anticipated errors by triggering repair activities or follow- up diagnostics. With real-time sensor networks, one seeks to classify unique, new events as they occur. With observational sciences, a new generation of instrumentation seeks unique events through online analysis of large observational datasets. This work presents a solution based on transfer learning principles that permits principled CSDT learning while exploiting any prior knowledge of the designer to correct both between-class and withinclass imbalance. Training examples are adaptively reweighted based on a decomposition of the data attributes. The result is a new, nonparametric representation that matches the anticipated attribute distribution for the target events.

Thompson, David R.↗

Entire four-graviton EFT from the duality between color and kinematics

The Bern-Carrasco-Johansson (BCJ) double-copy construction reveals a fundamental structural connection between gauge and gravity theories. At its core, the BCJ double copy is directly due to a duality between the algebraic relations of a color root and those of a kinematic root. We generalize this principle beyond the conventional Lie algebra structure of tree-level Yang-Mills theory. By demanding color-kinematics duality for the complete basis of four-point color structures—including those involving the symmetric 𝑑 𝑎⁢𝑏⁢𝑐 constants—we define the universal double copy. We systematically classify the bases of all such parity-even generalized gauge-theory numerators and, independently, the space of all parity-even four-graviton higher-derivative operators. We demonstrate that our universal double-copy construction precisely spans the entire tower of parity-even four-graviton amplitudes in any dimension, except for the Lovelock 𝑅 3 contribution in 𝐷 > 6 which we can express in terms of a particularly simple universal triple-copy involving gauge theories coupled to scalars. Explicit machine-readable expressions for the complete basis of gauge-theory numerators and fundamental gravitational building blocks are provided in the Supplemental Material. This establishes that all possible four-point gravitational interactions can be factorized into products of gauge-theory building blocks governed by this universal notion of color-kinematics duality.

Carrasco, John Joseph M. [Northwestern Univ., Evan↗

Black Hills Wildfires: Mapping Post-fire Conifer Regeneration using Snow-on Imagery

The 2000 Jasper Fire in the Black Hills of South Dakota was the largest wildfire to date in the region, burning over 83,000 acres of ponderosa pine forest. In collaboration with partners from the United States Forest Service (USFS) Black Hills Experimental Forest, USFS Rocky Mountain Research Station, and United States Geological Survey Geosciences and Environmental Change Science Center, we characterized post-fire forest regeneration within high-severity burn patches. We accomplished this by implementing novel conifer detection techniques using a snow index mask to create a winter, snow-on image composite from Landsat 8 Operational Land Imager (OLI) and Sentinel-2 Multispectral Instrument (MSI) data. We utilized 2015 USFS stem maps of field-observed regeneration plots and ocularly sampled additional reforestation sites planted in 2001–2013. In Google Earth Engine (GEE), the field data and imagery were used to train a Random Forest (RF) model. The RF model classified 2021 conifer regeneration density as low, medium, or high across the high-severity burn area with an overall accuracy of 81.3%. Approximately 45.9% of the high-severity burn had low or no regeneration (0-40 trees per acre) 20 years post-fire. Given our partners' desire to find easily accessible low conifer regeneration zones, we identified 4,079 acres of priority planting sites that were within 1,500 feet of roads, had not been planted previously, and were larger than 50 acres. This method supports the use of snow-on imagery as a successful technique to identify conifer regeneration.

Casey Menick​↗

Black Hills Wildfires: Mapping Post-Fire Conifer Regeneration using Snow-On Imagery

The 2000 Jasper Fire in the Black Hills of South Dakota was the largest wildfire to date in the region, burning over 83,000 acres of ponderosa pine forest. In collaboration with partners from the United States Forest Service (USFS) Black Hills Experimental Forest, USFS Rocky Mountain Research Station, and United States Geological Survey Geosciences and Environmental Change Science Center, we characterized post-fire forest regeneration within high-severity burn patches. We accomplished this by implementing novel conifer detection techniques using a snow index mask to create a winter, snow-on image composite from Landsat 8 Operational Land Imager (OLI) and Sentinel-2 Multispectral Instrument (MSI) data. We utilized 2015 USFS stem maps of field-observed regeneration plots and ocularly sampled additional reforestation sites planted in 2001–2013. In Google Earth Engine (GEE), the field data and imagery were used to train a Random Forest (RF) model. The RF model classified 2021 conifer regeneration density as low, medium, or high across the high-severity burn area with an overall accuracy of 81.3%. Approximately 45.9% of the high-severity burn had low or no regeneration (0-40 trees per acre) 20 years post-fire. Given our partners' desire to find easily accessible low conifer regeneration zones, we identified 4,079 acres of priority planting sites that were within 1,500 feet of roads, had not been planted previously, and were larger than 50 acres. This method supports the use of snow-on imagery as a successful technique to identify conifer regeneration.

Casey Menick↗

Evaluating Combinations of Sentinel-2 Data and Machine-Learning Algorithms for Mangrove Mapping in West Africa

Creating a national baseline for natural resources, such as mangrove forests, and monitoring them regularly often requires a consistent and robust methodology. With freely available satellite data archives and cloud computing resources, it is now more accessible to conduct such large-scale monitoring and assessment. Yet, few studies examine the reproducibility of such mangrove monitoring frameworks, especially in terms of generating consistent spatial extent. Our objective was to evaluate a combination of image processing approaches to classify mangrove forests along the coast of Senegal and The Gambia. We used freely available global satellite data (Sentinel-2), and cloud computing platform (Google Earth Engine) to run two machine learning algorithms, random forest (RF), and classification and regression trees (CART). We calibrated and validated the algorithms using 800 reference points collected using high-resolution images. We further re-ran 10 iterations for each algorithm, utilizing unique subsets of the initial training data. While all iterations resulted in thematic mangrove maps with over 90% accuracy, the mangrove extent ranges between 827-2807 km2 for Senegal and 245-1271 km2 for The Gambia with one outlier for each country. We further report "Places of Agreement" (PoA) to identify areas where all iterations for both methods agree (506.6 km2 and 129.6 km2 for Senegal and The Gambia, respectively), thus have a high confidence in predicting mangrove extent. While we acknowledge the time- and cost-effectiveness of such methods for the landscape managers, we recommend utilizing them with utmost caution, as well as post-classification on-the-ground checks, especially for decision making.

Mondal, Pinki↗

Mapping Post-fire Conifer Regeneration using Snow-on Imagery

The 2000 Jasper Fire in the Black Hills of South Dakota was the largest wildfire to date in the region, burning over 83,000 acres of ponderosa pine forest. In collaboration with partners from the United States Forest Service (USFS) Black Hills Experimental Forest, USFS Rocky Mountain Research Station, and United States Geological Survey Geosciences and Environmental Change Science Center, we characterized post-fire forest regeneration within high severity burn patches. We accomplished this by implementing novel conifer detection techniques using a snow index mask to create a winter, snow-on image composite from Landsat 8 Operational Land Imager (OLI) and Sentinel-2 Multispectral Instrument (MSI) data. We utilized 2015 USFS stem maps of field-observed regeneration plots and ocularly sampled reforestation sites planted from 2001–2013. The field data and imagery were used to train a Random Forest (RF) model in Google Earth Engine. The RF model classified conifer regeneration density as low, medium, or high across the high-severity burn area with an overall accuracy of 81.3% for 2021. Approximately 45.9% of the high-severity burn area had low or no regeneration (0-40 trees per acre) 20 years post-fire. Given our partners' desire to find easily accessible low conifer regeneration zones, we identified 4,079 acres of priority planting sites that were within 1,500 feet of roads, had not been planted previously, and were larger than 50 acres. This method supports the use of snow-on imagery as a successful technique to identify conifer regeneration in a post-wildfire landscape.

Yeshey Seldon↗

The application of LANDSAT remote sensing technology to natural resources management. Section 1: Introduction to VICAR - Image classification module. Section 2: Forest resource assessment of Humboldt County.

A teaching module on image classification procedures using the VICAR computer software package was developed to optimize the training benefits for users of the VICAR programs. The field test of the module is discussed. An intensive forest land inventory strategy was developed for Humboldt County. The results indicate that LANDSAT data can be computer classified to yield site specific forest resource information with high accuracy (82%). The "Douglas-fir 80%" category was found to cover approximately 21% of the county and "Mixed Conifer 80%" covering about 13%. The "Redwood 80%" resource category, which represented dense old growth trees as well as large second growth, comprised 4.0% of the total vegetation mosaic. Furthermore, the "Brush" and "Brush-Regeneration" categories were found to be a significant part of the vegetative community, with area estimates of 9.4 and 10.0%.

Fox, L., III↗

Channel fading for mobile satellite communications using spread spectrum signaling and TDRSS

This paper will present some preliminary results from a propagation experiment which employed NASA's TDRSS and an 8 MHz chip rate spread spectrum signal. Channel fade statistics were measured and analyzed in 21 representative geographical locations covering urban/suburban, open plain, and forested areas. Cumulative distribution Functions (CDF's) of 12 individual locations are presented and classified based on location. Representative CDF's from each of these three types of terrain are summarized. These results are discussed, and the fade depths exceeded 10 percent of the time in three types of environments are tabulated. The spread spectrum fade statistics for tree-lined roads are compared with the Empirical Roadside Shadowing Model.

Jenkins, Jeffrey D.↗

Multiple Spectral-Spatial Classification Approach for Hyperspectral Data

A .new multiple classifier approach for spectral-spatial classification of hyperspectral images is proposed. Several classifiers are used independently to classify an image. For every pixel, if all the classifiers have assigned this pixel to the same class, the pixel is kept as a marker, i.e., a seed of the spatial region, with the corresponding class label. We propose to use spectral-spatial classifiers at the preliminary step of the marker selection procedure, each of them combining the results of a pixel-wise classification and a segmentation map. Different segmentation methods based on dissimilar principles lead to different classification results. Furthermore, a minimum spanning forest is built, where each tree is rooted on a classification -driven marker and forms a region in the spectral -spatial classification: map. Experimental results are presented for two hyperspectral airborne images. The proposed method significantly improves classification accuracies, when compared to previously proposed classification techniques.

Tarabalka, Yuliya↗

Utilizing NASA Earth Observing System (EOS) Data to Determine Ideal Planting Locations for Wetland Tree Species in St. Bernard Parish, Louisiana

St. Bernard Parish, in southeast Louisiana, is rapidly losing coastal forests and wetlands due to a combination of natural and anthropogenic disturbances (e.g. subsidence, saltwater intrusion, low sedimentation, nutrient deficiency, herbivory, canal dredging, levee construction, spread of invasive species, etc.). After Hurricane Katrina severely impacted the area in 2005, multiple Non-Governmental Organizations (NGOs) have worked not only on rebuilding destroyed dwellings, but on rebuilding the ecosystems that once protected the citizens of St. Bernard Parish. Volunteer groups, NGOs, and government entities often work separately and independently of each other and use different sets of information to choose the best planting sites for coastal forests. Using NASA EOS, NRCS soil surveys, and ancillary road and canal data in conjunction with ground truthing, the team created maps of optimal planting sites for several species of wetland trees to aid in unifying these organizations, who share a common goal, under one plan. The methodology for this project created a comprehensive Geographic Information System (GIS) to help identify suitable planting sites in St. Bernard Parish. This included supplementing existing elevation data using LIDAR data and classifying existing land cover in the study area from ASTER multispectral satellite data. Low altitude AVIRIS hyperspectral imagery was used to assess the health of vegetation over an area near the intersection of the Mississippi River Gulf Outlet Canal (MRGO) and Bayou la Loutre. Historic extent of coastal forests was mapped using aerial photos from USGS collected between 1952 and 1956. The final products demonstrated the utility of combining NASA EOS with other geospatial data in assessing, monitoring, and restoring of coastal ecosystems in Louisiana. This methodology also provides a useful template for other ecological forecasting and coastal restoration applications.

Reahard, Ross↗

Analysis of data acquired by Shuttle Imaging Radar SIR-A and Landsat Thematic Mapper over Baldwin County, Alabama

Seasonally compatible data collected by SIR-A and by Landsat 4 TM over the lower coastal plain in Alabama were coregistered, forming a SIR-A/TM multichannel data set with 30 m x 30 m pixel size. Spectral signature plots and histogram analysis of the data were used to observe data characteristics. Radar returns from pine forest classes correlated highly with the tree ages, suggesting the potential utility of microwave remote sensing for forest biomass estimation. As compared with the TM-only data set, the use of SIR-A/TM data set improved classification accuracy of the seven land cover types studied. In addition, the SIR-A/TM classified data support previous finding by Engheta and Elachi (1982) that microwave data appear to be correlated with differing bottomland hardwood forest vegetation as associated with varying water regimens (i.e., wet versus dry).

Wu, S.-T.↗

NETRA: A parallel architecture for integrated vision systems. 1: Architecture and organization

Computer vision is regarded as one of the most complex and computationally intensive problems. An integrated vision system (IVS) is considered to be a system that uses vision algorithms from all levels of processing for a high level application (such as object recognition). A model of computation is presented for parallel processing for an IVS. Using the model, desired features and capabilities of a parallel architecture suitable for IVSs are derived. Then a multiprocessor architecture (called NETRA) is presented. This architecture is highly flexible without the use of complex interconnection schemes. The topology of NETRA is recursively defined and hence is easily scalable from small to large systems. Homogeneity of NETRA permits fault tolerance and graceful degradation under faults. It is a recursively defined tree-type hierarchical architecture where each of the leaf nodes consists of a cluster of processors connected with a programmable crossbar with selective broadcast capability to provide for desired flexibility. A qualitative evaluation of NETRA is presented. Then general schemes are described to map parallel algorithms onto NETRA. Algorithms are classified according to their communication requirements for parallel processing. An extensive analysis of inter-cluster communication strategies in NETRA is presented, and parameters affecting performance of parallel algorithms when mapped on NETRA are discussed. Finally, a methodology to evaluate performance of algorithms on NETRA is described.

Choudhary, Alok N.↗

Training toward significance with the decorrelated event classifier transformer neural network

Experimental particle physics uses machine learning for many tasks, where one application is to classify signal and background events. This classification can be used to bin an analysis region to enhance the expected significance for a mass resonance search. In natural language processing, one of the leading neural network architectures is the transformer. In this work, an event classifier transformer is proposed to bin an analysis region, in which the network is trained with special techniques. The techniques developed here can enhance the significance and reduce the correlation between the network’s output and the reconstructed mass. It is found that this trained network can perform better than boosted decision trees and feed-forward networks. Published by the American Physical Society 2024

Astronomy & Astrophysics↗

Using machine learning techniques to automate sky survey catalog generation

We describe the application of machine classification techniques to the development of an automated tool for the reduction of a large scientific data set. The 2nd Palomar Observatory Sky Survey provides comprehensive photographic coverage of the northern celestial hemisphere. The photographic plates are being digitized into images containing on the order of 10(exp 7) galaxies and 10(exp 8) stars. Since the size of this data set precludes manual analysis and classification of objects, our approach is to develop a software system which integrates independently developed techniques for image processing and data classification. Image processing routines are applied to identify and measure features of sky objects. Selected features are used to determine the classification of each object. GID3* and O-BTree, two inductive learning techniques, are used to automatically learn classification decision trees from examples. We describe the techniques used, the details of our specific application, and the initial encouraging results which indicate that our approach is well-suited to the problem. The benefits of the approach are increased data reduction throughput, consistency of classification, and the automated derivation of classification rules that will form an objective, examinable basis for classifying sky objects. Furthermore, astronomers will be freed from the tedium of an intensely visual task to pursue more challenging analysis and interpretation problems given automatically cataloged data.

Fayyad, Usama M.↗

Uncertainty-Aware Machine Learning for Small-Angle X-ray Scattering Analysis in Autonomous Experimentation

Small-angle X-ray scattering (SAXS) is a powerful high-throughput characterization tool for probing nanoscale structure in native sample environments, providing real-time morphological information such as nanoparticle size and shape during synthesis. However, automated SAXS data analysis for extracting meaningful structural parameters is non-trivial and remains a bottleneck in closed-loop experimentation towards autonomous materials discovery, which demands fast, reliable, and uncertainty-aware data analysis. Here, we develop a machine-learning approach for automated SAXS analysis tailored to closed-loop nanoparticle synthesis. A Random Forest (RF) regression model is trained on 100,000 synthetic SAXS curves generated from polydisperse spherical nanoparticles with realistic background contributions. Using normalized one-dimensional SAXS intensity profiles as input, the RF model directly predicts nanoparticle radius, size polydispersity, and background parameters, while the ensemble standard deviation across trees provides built-in uncertainty quantification (UQ). On synthetic data, we show that combining fit-quality metrics (R 2 , MAE) with thresholds on prediction uncertainty reliably identifies accurate parameter estimates without access to ground truth. We then apply the trained model to 365 experimental SAXS profiles of citrate-reduced gold nanoparticles synthesized using an automated droplet-flow microreactor with in situ SAXS at a synchrotron beamline, classifying the results into high- and low-confidence subsets based on UQ metrics. Finally, we integrate RF-based SAXS analysis into a simulated closed-loop optimization campaign using Gaussian process Bayesian optimization to minimize nanoparticle polydispersity, benchmarking against conventional automated Levenberg–Marquardt fitting. The RF-guided campaign exhibits substantially faster convergence and lower relative opportunity cost (∼0.07 vs ∼0.3), demonstrating that uncertainty-aware machine-learning SAXS analysis significantly enhances the efficiency and robustness of autonomous nanomaterials synthesis workflows.

Bayesian optimization↗

Kingdoms in turmoil

How should the world's living organisms be classified? Into how many kingdoms should they be grouped? Scientists have been grappling with these questions since the time of Aristotle, drawing on a broad base of biological characteristics for clues. The fossil record, visible traits of living organisms and, more recently, results from cell biology have all shaped theories of biological classification. But last year a new and controversial concept emerged: a classification of life based solely on molecular traits. The focal point of the controversy is a tree of life, or "phylogeny", devised by Carl Woese of the University of Illinois, Otto Kandler of the University of Munich and Mark Wheelis of the University of California. The tree is unusual because, unlike all previous schemes, it is constructed solely from biochemical data such as DNA sequences rather than a range of different organism characteristics. But that is not all. The scheme also challenges the idea that life on Earth is best divided into five kingdoms, with the main split being between bacteria and all other organisms. Woese and his colleagues create three main groupings by dividing the bacteria in two and unifying all other organisms.

NASA Program Exobiology↗

Prediction of Weather Impacted Airport Capacity using Ensemble Learning

Ensemble learning with the Bagging Decision Tree (BDT) model was used to assess the impact of weather on airport capacities at selected high-demand airports in the United States. The ensemble bagging decision tree models were developed and validated using the Federal Aviation Administration (FAA) Aviation System Performance Metrics (ASPM) data and weather forecast at these airports. The study examines the performance of BDT, along with traditional single Support Vector Machines (SVM), for airport runway configuration selection and airport arrival rates (AAR) prediction during weather impacts. Testing of these models was accomplished using observed weather, weather forecast, and airport operation information at the chosen airports. The experimental results show that ensemble methods are more accurate than a single SVM classifier. The airport capacity ensemble method presented here can be used as a decision support model that supports air traffic flow management to meet the weather impacted airport capacity in order to reduce costs and increase safety.

Weather impact↗