Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “classification models”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Land Covering Classifications of Boreas Modeling Grid Using AIRSAR Images

Mapping forest types in the boreal ecosystem in an integrated part of any modeling excercise of biogeophysical processes characterizing the interaction of forest with the atmosphere. In this paper, we report the results of the land cover classification of the SAR data acquired during the BOREAS (BOReal Ecosystem Atmospheric Study) intensive field campaigns over the modeling sub-grid of the southern study area in Saskatchewan , Canada. A Bayesian-maximum-a-posteriori classifier has been applied on the NASA/JPL AIRSAR images covering the region during the peak of the growing season in July, 1994.

Atmospheric Study↗

Analytical models and system topologies for remote multispectral data acquisition and classification

Simple analytical models are presented of the radiometric and statistical processes that are involved in multispectral data acquisition and classification. Also presented are basic system topologies which combine remote sensing with data classification. These models and topologies offer a preliminary but systematic step towards the use of computer simulations to analyze remote multispectral data acquisition and classification systems.

Huck, F. O.↗

Apollo 16 rocks - Classification and petrogenetic model

The Apollo 16 rocks include cataclastic anorthosites, two varieties of unequilibrated breccia, two varieties of partly to fully equilibrated breccia, and a sequence of partially melted breccias. The latter, which dominate the Apollo 16 collection, include glass, divitrified glass, mesostasis-olivine-plagioclase rock, mesostasis-rich basalt, basalt, and poikilitic rocks. All sequence members contain vesicles and relics of plagioclase, olivine, pink spinel, and lithic fragments. Their equilibrated matrices define a series from glass, to a plagioclase-olivine-mesostasis assemblage displaying spherulitic and skeletal shaped crystals, through a plagioclase-pyroxene-olivine assemblage displaying euhedral shaped crystals. Such data suggest that the sequence lithologies were derived from breccias or soils that were partially melted in an impact event.

Warner, J. L.↗

Linear dimensionality of Landsat agricultural data with implications for classification

A model for the Landsat multispectral scanner data, representing a generalization of the commonly used Gaussian model, has been formulated and analyzed. The model hypothesizes that the data for different crop types essentially lie on distinct hyperplanes in the feature space. Tests of this model reveal that: (1) the agricultural data from any single acquisition (i.e., four-channel) of Landsat are essentially two dimensional, regardless of the crop type; and (2) the data from different sites and different stages of crop development all lie on planes which are parallel. These findings have significant implications for data display, classification, feature extraction, and signature extension.

Wheeler, S. G.↗

A Machine Learning Approach to Improve Air Traffic Management Initiatives

Collaborating closely with commercial air carriers and related organizations, the Federal Aviation Administration(FAA) regulates air traffic and ensures the safety and efficiency of air operations. Air traffic controllers make strategic decisions, such as delaying, rerouting, or canceling flights, partly based on guidance provided by the FAA’s Air TrafficControl System Command Center (ATCSCC). The guidance includes, among other things, control measures known asTraffic Management Initiatives (TMIs) designed to enhance safety and improve operational efficiency. TMIs play a crucial role in managing the demand and capacity within the U.S. National Airspace System (NAS). Two major TMIs that are routinely used (primarily to mitigate the adverse effects of bad weather) are Ground Delay Programs (GDPs) andGround Stops (GSs). In a GDP, flights destined for airports facing thunderstorm activity experience delays at their origin airports. This proactive approach minimizes the risk of routing aircraft through hazardous weather conditions and also replaces (fuel burning) airborne delays with ground delays. In a GS, a temporary restriction is imposed on the departure or arrival of aircraft at a specific airport or within a designated airspace. Although other TMIs (e.g., miles-in-trail) are also implemented as part of (air) traffic flow management in the NAS, the focus of this work is on GDPs and GSs. Since TMIs, by design, lead to flight delays or cancellations, it is crucial to put in place the right set of parameters(e.g., scope and duration of the GDP). For example, when the end time of a GDP extends beyond what is necessary, it imposes unnecessary delays on departing flights. This situation could occur as a result of inaccurate prediction of the(required) duration of the GDP based on the weather forecast. On the other hand, if a GDP ends prematurely before the underlying capacity constraints are resolved at the destination airport, it may result in airborne holding. The delicate balance lies in matching the termination of the GDP precisely with the resolution of capacity constraints, avoiding both the imposition of unnecessary ground delays and the need for airborne holding due to premature program termination.Failing to specify the right parameters for TMIs also leads to flight delays, creating a significant obstacle in managing the increasing traffic volumes causing increased work load for the controllers. To address this issue, we propose the integration of Machine Learning (ML) models in the traffic flow management(TFM) pipeline. In current operations, decisions are made by human experts based on extensive training, historical patterns, available traffic and weather data. Since we have an abundance of data from past events that tell us the likely impact of various TMIs, by ingesting historical data, properly trained ML models can offer valuable insights and aid human decision-making. With the FAA increasingly exploring advanced analytics, ML emerges as a focal point for enhancing TFM within the National Airspace System (NAS). As a first step, this study aims to provide traffic controllers with decision-making support for the issuance and adjustment of TMIs. Data analytics and machine learning have been previously employed to address some of the challenges associated with TMIs. Numerous studies have concentrated on various facets of TMI issuance, exploring factors influencing TMI parameters, including arrival rate, airport capacity, and delay prediction. For example, using weather forecasts, several statistical methods were used to produce probabilistic capacity profiles which in conjunction with deterministic models provided insights into the GDP planning process [1–4]. The downside of using deterministic models is that they rely on fixed inputs and predetermined rules, which lack the ability to account for the inherent uncertainty and variability present in real-world scenarios. In a separate series of studies, researchers aimed to predict the occurrences of GDPs and GSs. The majority of these studies utilized various supervised learning methods, including Decision Trees, Naive Bayes, Support VectorMachines, and Random Forests to analyze the influence of weather conditions and arrival demand on TMI incidents[5–8]. However, these studies primarily focused on predicting the incidence of TMIs without explicitly addressing the scope of TMIs, including their duration and their geographical coverage. Furthermore, the emphasis of these studies was largely on GDPs, given their higher frequency and longer duration when compared to GSs. A limited number of studies focused on predicting the parameters of TMIs, specifically addressing their duration and extent. In one such study focusing on optimizing the TMI parameters at San Francisco International Airport (SFO),the authors utilized a probabilistic forecast of fog [9]. They simulated various capacity scenarios based on the (fog)burn-off forecasts, selecting GDP parameters that minimized airborne and overall ground delays. However, this approach exclusively emphasizes stratus (fog) burn-off as the primary determinant of GDP and GS, neglecting other influential factors like severe weather events, runway closures, lower capacity than traffic demand, and other important variables. Given the complexity of predicting the TMI and determining its scope, we seek a more holistic approach. We aim to consider all significant factors that could impact TMIs and their parameters. What sets this research apart is the fusion of all data sources relevant to the issuance and adjustment of TMIs and it represents the first comprehensive attempt to optimize TMIs in this manner. Since this comprehensive solution involves various aspects, we break down the problem into smaller components and input all parameters into a unified model called the “TMI Adjuster”. Figure 1 shows the overall framework and the list of datasets used in each model. The objective of the TMI Adjuster module is to deliver reliable, consistent and expedited recommendations for the progression, adjustment, and termination of TMIs. The ML solution entails developing a pipeline capable of predicting the necessity of a TMI (e.g., GS or GDP) along with its various parameters. For example, in the case of a GS, this includes the scope of the GS either in terms of distance from the destination airport or based on pre-defined airspace sectors. Here, scope refers to those regions and departing airports that are subject to the GS. In this paper, we concentrate on the issuance of GSs in the three major airports in the New York area — LaGuardia(LGA), John F. Kennedy International (JFK), and Newark Liberty International (EWR). We fuse traffic, weather and other relevant aviation data from years 2017 to 2019 to train and validate the ML models. In particular, we use the following datasets: •Terminal Aerodrome Forecast (TAF): meteorological forecasts specific to each airport, issued four times a day, covering predefined time periods. •TMI data: includes all GSs and GDPs along with their respective parameters. •Aviation System Performance Metrics (ASPM): includes traffic related data such as aircraft delays, arrival, and departure rates. •Notices to Airmen (NOTAMs): utilized to extract runway closure data and manage interdependencies between terminals in close proximity. •Flight cancellation data •Airspace Flow Programs (AFP): includes information on flight airborne holdings caused by TMIs. The data preprocessing entails transforming ASPM, TMI, AFP, NOTAMs, and weather data into an hourly format and consolidating all datasets by merging them based on date and time as the primary key. The TMI Adjuster framework comprises two parallel models: one dedicated to GS and a second model focused on GDP. As previously mentioned, our specific focus is on the GS model as a multi-classification problem. In this framework, each data point of the GS model input summarizes ten hours of data. Specifically, the data loader for the GS model generates the input and output of the model as follows: at a given time step, the input includes the actual traffic, weather, and TMI data from the two-hour window before the time step, alongside the weather forecast and scheduled traffic for the next 8 hours starting from the time step. Based on this information, the output of the GS model for each time interval consists of three dimensions. The first dimension represents a binary decision on whether there should be a GS in place for the next hour or not. The second dimension is related to the scope of the GS in the United States, and the third dimension is related to the scope of the GS in Canada (i.e., to determine if the GS impacts airports in Canada).One of the challenges with TMI modeling is the sparsity of TMI events, particularly regarding its scope. To address this challenge in the scope of the GS model output, we implement grouping. The GS scope for the US region is defined based on a list of centers that should be included when the GS is in place. With 20 centers in the US, we utilized historical data to group them into 4 categories. In particular, we summarized our historical data in a graph format where nodes represent centers, and link weights are defined based on the co-occurrence of centers in the scope parameter ofTMIs. By identified strongly connected components in this graph, we were able to partition the centers into four groups. We consider two model structures for the GS Model. Firstly, a hierarchical classification model [10], where the human decision-making for a GS is of hierarchical nature. The decision-maker first decides whether there is a need fora GS, and if the answer is yes, determines the scope. A hierarchical classification model organizes the problem into a class hierarchy, typically a tree or a Directed Acyclic Graph (DAG) structure, and considers the dependency of the decision in the previous step to the next component [10]. Here, we employ the local classifier per level approach, which involves training one multi-class classifier for each level of the class hierarchy. The second structure is the independent structure. In this setting, as the name suggests, we do not consider the dependency of the decisions in the different dimensions of the output of the model. Instead, for each dimension, we train a multi-class classifier independently. Table 1 summarizes GS model statistics for training, validation and testing. The table documents the effect of limiting data to the time steps when there was actually a TMI in place or when a TMI had just terminated. This resulted in a more balanced distribution of the GS class(GS positive class)versus “No GS”(GS negative class), which might help the training process. While JFK and LGA follow very similar distributions, with 40% and 42% GS positive class respectively, EWR has proportionally fewer GS incidents at 28%. Our subsequent phase involves evaluating the performance of both hierarchical structure and independent structure using different state-of-the-art multi-class classifier models such as Random Forest, Decision Trees, K-nearest Neighbors, and Logistic Regression and forecast the duration and scope of the GSs.

Farzan Masrour Shalmani↗

New Neighbours: Modelling the Growing Population of gamma-ray Millisecond Pulsars

The Fermi Large Area Telescope, in collaboration with several groups from the radio community. have had marvelous success at uncovering new gamma-ray millisecond pulsars (MSPs). In fact, MSPs now make up a sizable fraction of the total number of known gamma-ray pulsars. The MSP population is characterized by a variety of pulse profile shapes, peak separations, and radio-to-gamma phase lags, with some members exhibiting nearly phase-aligned radio and gamma-ray light curves (LCs). The MSPs' short spin periods underline the importance of including special relativistic effects in LC calculations, even for emission originating from near the stellar surface. We present results on modelling and classification of MSP LCs using standard pulsar model geometries.

Venter, C.↗

Fitting a Two-Component Scattering Model to Polarimetric SAR Data

Classification, decomposition and modeling of polarimetric SAR data has received a great deal of attention in the recent literature. The objective behind these efforts is to better understand the scattering mechanisms which give rise to the polarimetric signatures seen in SAR image data. In this Paper an approach is described, which involves the fit of a combination of two simple scattering mechanisms to polarimetric SAR observations. The mechanisms am canopy scatter from a cloud of randomly oriented oblate spheroids, and a ground scatter term, which can represent double-bounce scatter from a pair of orthogonal surfaces with different dielectric constants or Bragg scatter from a moderately rough surface, seen through a layer of vertically oriented scatterers. An advantage of this model fit approach is that the scattering contributions from the two basic scattering mechanisms can be estimated for clusters of pixels in polarimetric SAR images. The solution involves the estimation of four parameters from four separate equations. The model fit can be applied to polarimetric AIRSAR data at C-, L- and P-Band.

Freeman, A.↗

A Three-component Scattering Model for Polarimetric SAR Data

Classification, decomposition and modeling of polarimetric SAR data has received a great deal of attention in the recent literature. The objective behind these efforts is to better understand the scattering mechanisms which give rise to the polarimetric signatures seen in SAR image data.

SAR data AIRSAR backscatter scattering↗

Bounding Species Distribution Models

Species distribution models are increasing in popularity for mapping suitable habitat for species of management concern. Many investigators now recognize that extrapolations of these models with geographic information systems (GIS) might be sensitive to the environmental bounds of the data used in their development, yet there is no recommended best practice for "clamping" model extrapolations. We relied on two commonly used modeling approaches: classification and regression tree (CART) and maximum entropy (Maxent) models, and we tested a simple alteration of the model extrapolations, bounding extrapolations to the maximum and minimum values of primary environmental predictors, to provide a more realistic map of suitable habitat of hybridized Africanized honey bees in the southwestern United States. Findings suggest that multiple models of bounding, and the most conservative bounding of species distribution models, like those presented here, should probably replace the unbounded or loosely bounded techniques currently used [Current Zoology 57 (5): 642-647, 2011].

Stohlgren, Thomas J.↗

Natural Language Processing Analysis of Notices to Airmen for Air Traffic Management Optimization

With new emerging technologies in the field of NLP, we explore their applications to digitize and analyze heritage Air Traffic Management (ATM) documents for planning and optimizing airspace operations. Specifically, this research focuses on harvesting semi-structured or un-structured information contained in Notices to Airmen (NOTAMs). Using NLP and other advanced data analytics, we will construct a data-driven framework which facilitates finding language patterns and the use of pretrained language models for classification and extraction of useful airspace constraints and restrictions. These may lead to tools that assist airspace users in understanding the constraints more efficiently, contributing to better route planning and safer execution. This paper explores three workflows entailing different NLP tasks. First, unsupervised techniques like word embedding and topic modeling are used for pattern finding and document classification. Second, a dataset is created by extracting information from the semi-structured NOTAM format as metadata for categorizing, visualizing, and extracting key entities driving NOTAM content. Third, modern pre-built deep learning based transformer models such as BERT, RoBERTa, and XLNet are evaluated on the question answering task, an even more robust approach to information extraction, as well as their respective fine-tuning tasks. In this work we include various performance metrics for the trained models to evaluate both accuracy and precision and we show that the models can be generalized for their respective tasks. The research work developed shows promise in uncovering trends in digital NOTAMs in the NAS and also offers a new framework for digitizing and inferring insights from free-form legacy NOTAMs, that are yet to be digitized.

Natural Language Processing↗

Natural Language Processing (NLP) Analysis of NOTAMs for Air Traffic Management Optimization

With new emerging technologies in the field of NLP, we explore their applications to digitize and analyze heritage Air Traffic Management (ATM) documents for planning and optimizing airspace operations. Specifically, this research focuses on harvesting semi-structured or un-structured information contained in Notices to Airmen (NOTAMs). Using NLP and other advanced data analytics, we will construct a data-driven framework which facilitates finding language patterns and the use of pretrained language models for classification and extraction of useful airspace constraints and restrictions. These may lead to tools that assist airspace users in understanding the constraints more efficiently, contributing to better route planning and safer execution. This paper explores three workflows entailing different NLP tasks. First, unsupervised techniques like word embedding and topic modeling are used for pattern finding and document classification. Second, a dataset is created by extracting information from the semi-structured NOTAM format as metadata for categorizing, visualizing, and extracting key entities driving NOTAM content. Third, modern pre-built deep learning based transformer models such as BERT, RoBERTa, and XLNet are evaluated on the question answering task, an even more robust approach to information extraction, as well as their respective fine-tuning tasks. In this work we include various performance metrics for the trained models to evaluate both accuracy and precision and we show that the models can be generalized for their respective tasks. The research work developed shows promise in uncovering trends in digital NOTAMs in the NAS and also offers a new framework for digitizing and inferring insights from free-form legacy NOTAMs, that are yet to be digitized. Video is an mp4 download, with a play time of 9 min 35 secs.

Natural Language Processing↗

A Rorschach Test for Visual Classification Strategies

Contemporary models of pattern, detection and discrimination often employ template matching, but there have been few direct tests of this proposition. Adopting a method developed by Ahumada, we have analyzed how human observers discriminate between two letters of the alphabet ('c' and 'x'). The stimulus consisted of a one degree tall letter plus a four degree field of static white noise, both displayed for 16 frames at a 67 Hz frame rate. Our font and display dimensions approximated those of Solomon and Pelli. The observer identified the letter presented. A QUEST staircase varied letter contrast to maintain a 75% correct rate. For each trial, we preserved the information required to reconstruct the noise field. Possible trial categories based on (signal, response) pairs are: (c,c), (c,x), (x,c), (x,x). Noise fields were averaged separately for each category, and a final classification image was obtained by averaging the four mean images after inverting the sign of categories in which x was the response. If the observer employs a template, it should be revealed in the classification image. The lowpass-filtered classification image derived from 2048 responses of one observer is shown here, along with the corresponding ideal template. An approximation to the ideal template can be seen appropriately located within the classification image. We have also simulated and will discuss the classification images expected from various discrimination models in this experimental context. The construction of classification images appears to be a powerful tool for studying classification strategies used by human observers. Like a Rorschach test, it surreptitiously discovers the inner desires of the visual system.

Watson, Andrew B.↗

Basalt generation at the Apollo 12 site. Part 1: New data, classification, and re-evaluation

New data are reported from five previously unanalyzed Apollo 12 mare basalts that are incorporated into an evaluation of previous petrogenetic models and classification schemes for these basalts. This paper proposes a classification for Apollo 12 mare basalts on the basis of whole-rock Mg# (molar 100*(Mg/(Mg+Fe))) and Rb/Sr ratio (analyzed by isotope dilution), whereby the ilmenite, olivine, and pigeonite basalt groups are readily distinguished from each other. Scrutiny of the Apollo 12 feldspathic 'suite' demonstrates that two of the three basalts previously assigned to this group (12031, 12038, 12072) can be reclassified: 12031 is a plagioclase-rich pigeonite basalt; and 12072 is an olivine basalt. Only basalt 12038 stands out as a unique sample to the Apollo 12 site, but whether this represents a single sample from another flow at the Apollo 12 site or is exotic to this site is equivocal. The question of whether the olivine and pigeonite basalt suites are co-magmatic is addressed by incompatible trace-element chemistry: the trends defined by these two suites when Co/Sm and Sm/Eu ratios are plotted against Rb/Sr ratio demonstrate that these two basaltic types cannot be co-magmatic. Crystal fractionation/accumulation paths have been calculated and show that neither the pigeonite, olivine, or ilmenite basalts are related by this process. Each suite requires a distinct and separate source region. This study also examines sample heterogeneity and the degree to which whole-rock analyses are representative, which is critical when petrogenetic interpretation is undertaken. Sample heterogeneity has been investigated petrographically (inhomogeneous mineral distribution) with consideration of duplicate analyses, and whether a specific sample (using average data) plots consistently upon a fractionation trend when a number of different compostional parameters are considered. Using these criteria, four basalts have been identified where reported analyses are not representative of the whole-rock composition: 12005, an ilmenite basalt; 12006 and 12036, olivine basalts; and 12031 previously classified as a feldspathic basalt, but reclassified as part of the pigeonite suite.

Neal, Clive R.↗

Satellite image analysis using neural networks

The tremendous backlog of unanalyzed satellite data necessitates the development of improved methods for data cataloging and analysis. Ford Aerospace has developed an image analysis system, SIANN (Satellite Image Analysis using Neural Networks) that integrates the technologies necessary to satisfy NASA's science data analysis requirements for the next generation of satellites. SIANN will enable scientists to train a neural network to recognize image data containing scenes of interest and then rapidly search data archives for all such images. The approach combines conventional image processing technology with recent advances in neural networks to provide improved classification capabilities. SIANN allows users to proceed through a four step process of image classification: filtering and enhancement, creation of neural network training data via application of feature extraction algorithms, configuring and training a neural network model, and classification of images by application of the trained neural network. A prototype experimentation testbed was completed and applied to climatological data.

Sheldon, Roger A.↗

Protein Kinase Classification with 2866 Hidden Markov Models and One Support Vector Machine

The main application considered in this paper is predicting true kinases from randomly permuted kinases that share the same length and amino acid distributions as the true kinases. Numerous methods already exist for this classification task, such as HMMs, motif-matchers, and sequence comparison algorithms. We build on some of these efforts by creating a vector from the output of thousands of structurally based HMMs, created offline with Pfam-A seed alignments using SAM-T99, which then must be combined into an overall classification for the protein. Then we use a Support Vector Machine for classifying this large ensemble Pfam-Vector, with a polynomial and chisquared kernel. In particular, the chi-squared kernel SVM performs better than the HMMs and better than the BLAST pairwise comparisons, when predicting true from false kinases in some respects, but no one algorithm is best for all purposes or in all instances so we consider the particular strengths and weaknesses of each.

Weber, Ryan↗

An Evaluation of Clouds and Radiation in a Large-Scale Atmospheric Model Using a Cloud Vertical Structure Classification

We revisit the concept of the cloud vertical structure (CVS) classes we have previously employed to classify the planet's cloudiness (Oreopoulos et al., 2017). The CVS classification reflects simple combinations of simultaneous cloud occurrence in the three standard layers traditionally used to separate low, middle, and high clouds and was applied to a dataset derived from active lidar and cloud radar observations. This classification is now introduced in an atmospheric global climate model, specifically a version of NASA's GEOS-5, in order to evaluate the realism of its cloudiness and of the radiative effects associated with the various CVS classes. Such classes can be defined in GEOS-5 thanks to a sub column cloud generator paired with the model's radiative transfer algorithm, and their associated radiative effects can be evaluated against observations. We find that the model produces 50% more clear skies than observations in relative terms and produces isolated high clouds that are slightly less frequent than in observations, but optically thicker, yielding excessive planetary and surface cooling. Low clouds are also brighter than in observations, but underestimates of the frequency of occurrence (by ~20% in relative terms) help restore radiative agreement with observations. Overall the model better reproduces the longwave radiative effects of the various CVS classes because cloud vertical location is substantially constrained in the CVS framework.

Lee, Dongmin↗

A simulation of remote sensor systems and data processing algorithms for spectral feature classification

A computational model of the deterministic and stochastic processes involved in multispectral remote sensing was designed to evaluate the performance of sensor systems and data processing algorithms for spectral feature classification. Accuracy in distinguishing between categories of surfaces or between specific types is developed as a means to compare sensor systems and data processing algorithms. The model allows studies to be made of the effects of variability of the atmosphere and of surface reflectance, as well as the effects of channel selection and sensor noise. Examples of these effects are shown.

Arduini, R. F.↗