Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “automatic relevance determination”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

Visualizing Multi-process CPU Utilization using CUSP

The CPU Utilization Statistics Plotter (CUSP) tool automates the interpretation of detailed CPU Utilization trace data and statistics. It puts you on the cusp of understanding how CPU resources are split among the many parallel components of a software system.CUSP combines time-sampled CPU utilization numbers and Event Log annotations to generate human-readable plots and tables. It automatically splits up large CPU usage log files around interesting events, determines and highlights just the tasks of primary relevance by evaluating their changing contribution to each plot's total CPU usage, automatically eliminates irrelevant tasks, provides context by labeling plots with names and durations of all active commands, and uses consistent color-coding to enable quick visual comparison across multiple plots.CUSP has been used to process CPU Utilization trace logs on the Mars Science Laboratory and the Mars 2020 Rover missions during flight software development and Flight Operations on the Martian surface since December 2013.

Maimone, Mark W↗

Evaluation of Technology Concepts for Traffic Data Management and Relevant Audio for Datalink in Commercial Airline Flight Decks

Datalink is currently operational for departure clearances and in oceanic environments and is currently being tested in high altitude domestic enroute airspace. Interaction with even simple datalink clearances may create more workload for flight crews than the voice system they replace if not carefully designed. Datalink may also introduce additional complexity for flight crews with hundreds of uplink messages now defined for use. Finally, flight crews may lose airspace awareness and operationally relevant information that they normally pickup from Air Traffic Control (ATC) voice communications with other aircraft (i.e., “party-line” transmissions). Once again, automation may be poised to increase workload on the flight deck for incremental benefit. Datalink implementation to support future air traffic management concepts needs to be carefully considered, understanding human communication norms and especially, the change from voice- to text-based communications modality and its effect on pilot workload and situation awareness. Increasingly autonomous systems, where autonomy is designed to support human-autonomy teaming, may be suited to solve these issues. NASA is conducting research and development of increasingly autonomous systems, utilizing machine-learning algorithms seamlessly integrated with humans whereby task performance of the combined system is significantly greater than the individual components. Increasingly autonomous systems offer the potential for significantly improved levels of performance and safety that are superior to either human or automation alone. Two increasingly autonomous systems concepts - a traffic data manager and a conversational co-pilot - were developed to intelligently address the datalink issues in a complex, future state environment with significant levels of traffic. The system was tested for suitability of datalink usage for terminal airspace. The traffic data manager allowed for automated declutter of the Automatic Dependent Surveillance-Broadcast (ADS-B) display. The system determined relevant traffic for display based on machine learning algorithms trained by experienced human pilot behaviors. The conversational co-pilot provided relevant audio air traffic control messages based on context and proximity to ownship. Both systems made use of the connected aircraft concepts to provide intelligent context to determine relevancy above and beyond proximity to ownship. A human-in-the-loop test was conducted in NASA Langley Research Center’s Integration Flight Deck B-737-800 simulator to evaluate the traffic data manager and the conversational co-pilot. Twelve airline crews flew various normal and non-normal procedures and their actions and performance were recorded in response to the procedural events. This paper details the flight crew performance and evaluation during the events.

Etherington, Timothy↗

Using the International Directory Network and connected information systems for research in the Earth and space sciences

Many researchers are becoming aware of the International Directory Network (IDN), an interconnected federation of international directories to Earth and space science data. Are you aware, however, of the many Earth-science-relevant information systems which can be accessed automatically from the directories? After determining potentially useful data sets in various disciplines through directories such as the Global Change Master Directory, it is becoming increasingly possible to get detailed information about the correlative possibilities of these data sets through the connected guide/catalog and inventory systems. Such capabilities as data set browse, subsetting, analysis, etc. are available now and will be improving in the future.

Thieman, J. R.↗

Machine Learning Approaches to Increasing Value of Spaceflight Omics Databases

The number of spaceflight bioscience mission opportunities is too small to allow all relevant biological and environmental parameters to be experimentally identified. Simulated spaceflight experiments in ground-based facilities (GBFs), such as clinostats, are each suitable only for particular investigations -- a rotating-wall vessel may be 'simulated microgravity' for cell differentiation (hours), but not DNA repair (seconds) -- and introduce confounding stimuli, such as motor vibration and fluid shear effects. This uncertainty over which biological mechanisms respond to a given form of simulated space radiation or gravity, as well as its side effects, limits our ability to baseline spaceflight data and validate mission science. Machine learning techniques autonomously identify relevant and interdependent factors in a data set given the set of desired metrics to be evaluated: to automatically identify related studies, compare data from related studies, or determine linkages between types of data in the same study. System-of-systems (SoS) machine learning models have the ability to deal with both sparse and heterogeneous data, such as that provided by the small and diverse number of space biosciences flight missions; however, they require appropriate user-defined metrics for any given data set. Although machine learning in bioinformatics is rapidly expanding, the need to combine spaceflight/GBF mission parameters with omics data is unique. This work characterizes the basic requirements for implementing the SoS approach through the System Map (SM) technique, a composite of a dynamic Bayesian network and Gaussian mixture model, in real-world repositories such as the GeneLab Data System and Life Sciences Data Archive. The three primary steps are metadata management for experimental description using open-source ontologies, defining similarity and consistency metrics, and generating testing and validation data sets. Such approaches to spaceflight and GBF omics data may soon enable unique insight into which measured phenomena correlate to biological mechanisms that are truly affected by spaceflight conditions; which are most likely to be confounded by other variables; and which are insufficiently characterized, significantly increasing existing and future science return from ISS and spaceflight missions.

Gentry, Diana↗

Patterns in Crew-Initiated Photography of Earth from ISS - Is Earth Observation a Salutogenic Experience?

To provide for the well-being of crewmembers on future exploration missions, understanding how space station crewmembers handle the inherently stressful isolation and confinement during long-duration missions is important. A recent retrospective survey of previously flown astronauts found that the most commonly reported psychologically enriching aspects of spaceflight had to do with their Perceptions of Earth. Crewmembers onboard the International Space Station (ISS) photograph Earth through the station windows. Some of these photographs are in response to requests from scientists on the ground through the Crew Earth Observations (CEO) payload. Other photographs taken by crewmembers have not been in response to these formal requests. The automatically recorded data from the camera provides a dataset that can be used to test hypotheses about factors correlated with self-initiated crewmember photography. The present study used objective in-flight data to corroborate the previous questionnaire finding and to further investigate the nature of voluntary Earth-Observation activity. We examined the distribution of photographs with respect to time, crew, and subject matter. We also determined whether the frequency fluctuated in conjunction with major mission events such as vehicle dockings, and extra-vehicular activities (EVAs, or spacewalks), relative to the norm for the relevant crew. We also examined the influence of geographic and temporal patterns on frequency of Earth photography activities. We tested the hypotheses that there would be peak photography intensity over locations of personal interest, and on weekends. From December 2001 through October 2005 (Expeditions 4-11) crewmembers took 144,180 photographs of Earth with time and date automatically recorded by the camera. Of the time-stamped photographs, 84.5% were crew-initiated, and not in response to CEO requests. Preliminary analysis indicated some phasing in patterns of photography during the course of a mission (significant quadratic and trimodal models). There was also a small but significant increase in photo activity on the weekends. In contrast, fewer photos were taken during major station events and for a period of time immediately preceding those events. Data on photography patterns presented here represent a relatively objective group-level measure of Earth observing activities on ISS. Crew Earth Observations offers a self-initiated positive activity that may be important in salutogenesis (maintenance of well-being) of astronauts on long-duration missions. Consideration should be given to developing substitute activities for crewmembers in future exploration missions where there will not be the opportunity to look at Earth, such as on long-duration transits to Mars.

Robinson, Julie A.↗

Automatic data partitioning on distributed memory multicomputers

Distributed-memory parallel computers are increasingly being used to provide high levels of performance for scientific applications. Unfortunately, such machines are not very easy to program. A number of research efforts seek to alleviate this problem by developing compilers that take over the task of generating communication. The communication overheads and the extent of parallelism exploited in the resulting target program are determined largely by the manner in which data is partitioned across different processors of the machine. Most of the compilers provide no assistance to the programmer in the crucial task of determining a good data partitioning scheme. A novel approach is presented, the constraints-based approach, to the problem of automatic data partitioning for numeric programs. In this approach, the compiler identifies some desirable requirements on the distribution of various arrays being referenced in each statement, based on performance considerations. These desirable requirements are referred to as constraints. For each constraint, the compiler determines a quality measure that captures its importance with respect to the performance of the program. The quality measure is obtained through static performance estimation, without actually generating the target data-parallel program with explicit communication. Each data distribution decision is taken by combining all the relevant constraints. The compiler attempts to resolve any conflicts between constraints such that the overall execution time of the parallel program is minimized. This approach has been implemented as part of a compiler called Paradigm, that accepts Fortran 77 programs, and specifies the partitioning scheme to be used for each array in the program. We have obtained results on some programs taken from the Linpack and Eispack libraries, and the Perfect Benchmarks. These results are quite promising, and demonstrate the feasibility of automatic data partitioning for a significant class of scientific application programs with regular computations.

Gupta, Manish↗

A dynamic localization model with stochastic backscatter

The modeling of subgrid scales in large-eddy simulation (LES) has been rationalized by the introduction of the dynamic localization procedure. This method allows one to compute rather than prescribe the unknown coefficients in the subgrid-scale model. Formally, the LES equations are supposed to be obtained by applying to the Navier-Stokes equations a 'grid filter' operation. Though the subgrid stress itself is unknown, an identity between subgrid stresses generated by different filters has been derived. Although preliminary tests of the Dynamic Localization Model (DLM) with k-equation have been satisfactory, the use of a negative eddy viscosity to describe backscatter is probably a crude representation of the physics of reverse transfer of energy. Indeed, the model is fully deterministic. Knowing the filtered velocity field and the subgrid-scale energy, the subgrid stress is automatically determined. We know that the LES equations cannot be fully deterministic since the small scales are not resolved. This stems from an important distinction between equilibrium hydrodynamics and turbulence. In equilibrium hydrodynamics, the molecular motions are also not resolved. However, there is a clear separation of scale between these unresolved motions and the relevant hydrodynamic scales. The result of molecular motions can then be separated into an average effect (the molecular viscosity) and some fluctuations. Due to the large number of molecules present in a box with size of the order of the hydrodynamic scale, the ratio between fluctuations and the average effect should be very small (as a result of the 'law of large numbers'). For that reason, the hydrodynamic balance equations are usually purely deterministic. In turbulence, however, there is no clear separation of scale between small and large eddies. In that case, the fluctuations around a deterministic eddy viscosity term could be significant. An eddy noise would then appear through a stochastic term in the subgrid-scale model and could be the source of backscatter.

Carati, Daniele↗

Intelligent Weather Agent

Method and system for automatically displaying, visually and/or audibly and/or by an audible alarm signal, relevant weather data for an identified aircraft pilot, when each of a selected subset of measured or estimated aviation situation parameters, corresponding to a given aviation situation, has a value lying in a selected range. Each range for a particular pilot may be a default range, may be entered by the pilot and/or may be automatically determined from experience and may be subsequently edited by the pilot to change a range and to add or delete parameters describing a situation for which a display should be provided. The pilot can also verbally activate an audible display or visual display of selected information by verbal entry of a first command or a second command, respectively, that specifies the information required.

Spirkovska, Liljana↗

Determining Surface Roughness in Urban Areas Using Lidar Data

An automated procedure has been developed to derive relevant factors, which can increase the ability to produce objective, repeatable methods for determining aerodynamic surface roughness. Aerodynamic surface roughness is used for many applications, like atmospheric dispersive models and wind-damage models. For this technique, existing lidar data was used that was originally collected for terrain analysis, and demonstrated that surface roughness values can be automatically derived, and then subsequently utilized in disaster-management and homeland security models. The developed lidar-processing algorithm effectively distinguishes buildings from trees and characterizes their size, density, orientation, and spacing (see figure); all of these variables are parameters that are required to calculate the estimated surface roughness for a specified area. By using this algorithm, aerodynamic surface roughness values in urban areas can then be extracted automatically. The user can also adjust the algorithm for local conditions and lidar characteristics, like summer/winter vegetation and dense/sparse lidar point spacing. Additionally, the user can also survey variations in surface roughness that occurs due to wind direction; for example, during a hurricane, when wind direction can change dramatically, this variable can be extremely significant. In its current state, the algorithm calculates an estimated surface roughness for a square kilometer area; techniques using the lidar data to calculate the surface roughness for a point, whereby only roughness elements that are upstream from the point of interest are used and the wind direction is a vital concern, are being investigated. This technological advancement will improve the reliability and accuracy of models that use and incorporate surface roughness.

Holland, Donald↗

The SAX Italian scientific satellite. The on-board implemented automation as a support to the ground control capability

This paper presents the capabilities implemented in the SAX system for an efficient operations management during its in-flight mission. SAX is an Italian scientific satellite for x-ray astronomy whose major mission objectives impose quite tight constraints on the implementation of both the space and ground segment. The most relevant mission characteristics require an operative lifetime of two years, performing scientific observations both in contact and in noncontact periods, with a low equatorial orbit supported by one ground station, so that only a few minutes of communications are available each orbit. This operational scenario determines the need to have a satellite capable of performing the scheduled mission automatically and reacting autonomously to contingency situations. The implementation approach of the on-board operations management, through which the necessary automation and autonomy are achieved, follows a hierarchical structure. This has been achieved adopting a distributed avionic architecture. Nine different on-board computers, in fact, constitute the on-board data management system. Each of them performs the local control and monitors its own functions while the system level control is performed at a higher level by the data handling applications software. The SAX on-board architecture provides the ground operators with different options of intervention by three classes of telecommands. The management of the scientific operations will be scheduled by the operation control center via dedicated operating plans. The SAX satellite flight mode is presently being integrated at Alenia Spazio premises in Turin for a launch scheduled for the end of 1995. Once in orbit, the SAX satellite will be subject to intensive check-out activities in order to verify the required mission performances. An overview of the envisaged procedure and of the necessary on-ground activities is therefore depicted as well.

Martelli, Andrea↗

A Scheme for finding the Front Boundary of an Interplanetary Magnetic Cloud

We developed a scheme for finding the front boundary of an interplanetary magnetic cloud (MC) based on criteria that depend on the possible existence of any one or all of six specific solar wind features. The features that the program looks for, within +/- 2 hours of a preliminarily determined time for the front boundary, estimated either by visual inspection or by an automatic MC identification scheme, are: (1) a sufficiently large directional discontinuity in the interplanetary magnetic field (IMF), (2) existence of a magnetic hole, (3) a significant proton plasma beta drop, (4) a significant proton temperature drop, (5) a marked increase in the IMF's intensity, and (6) a significant decrease in a normalized root-mean-square deviation (RMS)of the magnetic field - where the scheme was tested using 5, 10, 15, and 20 minute averages of the relevant physical quantities, in order to find the optimum average (and RMS) to use. Other criteria, besides these six, were examined and dismissed as not reliable, e.g., plasma speed. The scheme was developed specifically for aiding in forecasting the strength and timing of a geomagnetic storm due to the passage of an interplanetary MC in real-time, but can be used in post ground-data collection for imposition of consistency in choosing a MC's front boundary. The scheme has been extensively tested, first using 80 bona fide MCs over about 9 years of WIND data, and also for 121 MC-like structures as defined by a program that automatically identifies such structures over the same period. Optimum limits for various parameters in the scheme were found by statistical studies of the WIND MCs. The resulting limits can be user-adjusted for other data sets, if desired. Final testing of the 80 MCs showed that for 50 percent of the events the boundary estimates occurred within +/-10 minutes of visually determined times, 80 percent occurred within +/-30 minutes, and 91 percent occur within +/-60 minutes, and three or more individual boundary tests were passed for 88 percent of the total MCs. The scheme and its testing will be described.

Lepping, Ronald P.↗

Semi-Supervised Data Summarization: Using Spectral Libraries to Improve Hyperspectral Clustering

Hyperspectral imagers produce very large images, with each pixel recorded at hundreds or thousands of different wavelengths. The ability to automatically generate summaries of these data sets enables several important applications, such as quickly browsing through a large image repository or determining the best use of a limited bandwidth link (e.g., determining which images are most critical for full transmission). Clustering algorithms can be used to generate these summaries, but traditional clustering methods make decisions based only on the information contained in the data set. In contrast, we present a new method that additionally leverages existing spectral libraries to identify materials that are likely to be present in the image target area. We find that this approach simultaneously reduces runtime and produces summaries that are more relevant to science goals.

Wagstaff, K. L.↗

Automatic discovery of optimal classes

A criterion, based on Bayes' theorem, is described that defines the optimal set of classes (a classification) for a given set of examples. This criterion is transformed into an equivalent minimum message length criterion with an intuitive information interpretation. This criterion does not require that the number of classes be specified in advance, this is determined by the data. The minimum message length criterion includes the message length required to describe the classes, so there is a built in bias against adding new classes unless they lead to a reduction in the message length required to describe the data. Unfortunately, the search space of possible classifications is too large to search exhaustively, so heuristic search methods, such as simulated annealing, are applied. Tutored learning and probabilistic prediction in particular cases are an important indirect result of optimal class discovery. Extensions to the basic class induction program include the ability to combine category and real value data, hierarchical classes, independent classifications and deciding for each class which attributes are relevant.

Cheeseman, Peter↗

Automatic determination of fault effects on aircraft functionality

The problem of determining the behavior of physical systems subsequent to the occurrence of malfunctions is discussed. It is established that while it was reasonable to assume that the most important fault behavior modes of primitive components and simple subsystems could be known and predicted, interactions within composite systems reached levels of complexity that precluded the use of traditional rule-based expert system techniques. Reasoning from first principles, i.e., on the basis of causal models of the physical system, was required. The first question that arises is, of course, how the causal information required for such reasoning should be represented. The bond graphs presented here occupy a position intermediate between qualitative and quantitative models, allowing the automatic derivation of Kuipers-like qualitative constraint models as well as state equations. Their most salient feature, however, is that entities corresponding to components and interactions in the physical system are explicitly represented in the bond graph model, thus permitting systematic model updates to reflect malfunctions. Researchers show how this is done, as well as presenting a number of techniques for obtaining qualitative information from the state equations derivable from bond graph models. One insight is the fact that one of the most important advantages of the bond graph ontology is the highly systematic approach to model construction it imposes on the modeler, who is forced to classify the relevant physical entities into a small number of categories, and to look for two highly specific types of interactions among them. The systematic nature of bond graph model construction facilitates the process to the point where the guidelines are sufficiently specific to be followed by modelers who are not domain experts. As a result, models of a given system constructed by different modelers will have extensive similarities. Researchers conclude by pointing out that the ease of updating bond graph models to reflect malfunctions is a manifestation of the systematic nature of bond graph construction, and the regularity of the relationship between bond graph models and physical reality.

Feyock, Stefan↗

Mode Transitions in Glass Cockpit Aircraft: Results of a Field Study

One consequence of increased levels of automation in complex control systems is the presence of modes. A mode is a particular configuration of a control system that defines how human command inputs are interpreted. In complex systems, modes also often determine a specific allocation of control authority between the human and automated systems. Even in simple static devices (e.g., electronic watches, word processors), the presence of modes has been found to cause problems in either-the acquisition or production of skilled performance. Many of these problems arise due to the fact that the selection of a mode causes device behavior to be mediated by hidden internal state information. For these simple systems, many of these interaction problems can be solved by the design of appropriate feedback to communicate internal state information to the human operator. In complex dynamic systems, however, the design issues associated with modes seem to trancend the problem of merely communicating internal state information via displayed feedback. In complex supervisory control systems (e.g., aircraft, spacecraft, military command and control), a key function of modes is the selection of a particular configuration of control authority between the human operator and automated control systems. One mode may result in full manual control, another may result in a mix of manual and automatic control, while a third may result in full automatic control over the entire system. The human operator selects an appropriate mode as a function of current goals, operating conditions, and operating procedures. Thus, the operator is put in a position of essentially trying to control two coupled dynamic systems: the target system itself, and also a highly complex suite of automation controlling the target system. From a historical perspective, it should probably not come as a surprise that very little information is available to guide the design of mode-oriented control systems. The topic of function allocation (i.e., the proper division of control authority among human and computer) has a long history in human-machine systems research. Although this research has produced some relevant guidelines, a design approach capable of defining appropriate allocations of control function between the human and automation is not yet available. As a result, the function allocation decision itself has been allocated to the operator, to be performed in real-time, in the operation of mode-oriented control systems. A variety of documented aircraft accidents and incidents suggest that the real-time selection and monitoring of control modes is a weak link in the effective operation of complex supervisory control systems. Research in human-machine systems and human-computer interaction has barely scraped the surface of the problem of understanding how operators manage this task.The purpose of this paper is to present the results of a field study which examined how operators manage mode selection in a complex supervisory control system. Data on mode engagements using the Boeing B757/767 auto-flight system were collected during approach and descent into four major airports in the East Coast of the United States. Protocols documenting mode selection, automatic mode changes, pilot actions, quantitative records of flight-path variables, and verbal reports during and after mode engagements were collected by an observer from the jumpseat. Observations were conducted on two typical trips between three airports. Each trip was be replicated 11 times, which yielded a total of 22 trips and 66 legs on which data were collected. All data collected concerned the same flight numbers, and therefore, the same time of day, same type of aircraft, and identical operational environments (e.g., ATC facilities, weather patterns, traffic flow etc.)

Degani, Asaf↗

Water Across Synthetic Aperture Radar Data (WASARD): SAR Water Body Classification for the Open Data Cube

The detection of inland water bodies from Synthetic Aperture Radar (SAR) data provides a great advantage over water detection with optical data, since SAR imaging is not impeded by cloud cover. Traditional methods of detecting water from SAR data involves using thresholding methods that can be labor intensive and imprecise. This paper describes Water Across Synthetic Aperture Radar Data (WASARD): a method of water detection from SAR data which automates and simplifies the thresholding process using machine learning on training data created from Geoscience Australia’s WOFS algorithm. Of the machine learning models tested, the Linear Support Vector Machine was determined to be optimal, with the option of training using solely the VH polarization or a combination of the VH and VV polarizations. WASARD was able to identify water in the target area with a correlation of 97% with WOFS. Sentinel-1, Open Data Cube, Earth Observations, Machine Learning, Water Detection 1. INTRODUCTION Water classification is an important function of Earth imaging satellites, as accurate remote classification of land and water can assist in land use analysis, flood prediction, climate change research, as well as a variety of agricultural applications [2]. The ability to identify bodies of water remotely via satellite is immensely cheaper than contracting surveys of the areas in question, meaning that an application that can accurately use satellite data towards this function can make valuable information available to nations which would not be able to afford it otherwise. Highly reliable applications for the remote detection of water currently exist for use with optical satellite data such as that provided by LANDSAT. One such application, Geoscience Australia’s Water Observations from Space (WOFS) has already been ported for use with the Open Data Cube [6]. However, water detection using optical data from Landsat is constrained by its relatively long revisit cycle of 16 days [5], and water detection using any optical data is constrained in that it lacks the ability to make accurate classifications through cloud cover [2]. The alternative solution which solves these problems is water detection using SAR data, which images the Earth using cloud-penetrating microwaves. Because of its advantages over optical data, much research has been done into water detection using SAR data. Traditionally, this has been done using the thresholding method, which involves picking a polarization band and labeling all pixels for which this band’s value is below a certain threshold as containing water. The thresholding method works since water tends to return a much lower backscatter value to the satellite than land [1]. However, this method can be flawed since estimating the proper threshold is often imprecise, complicated, and labor intensive for the end user. Thresholding also tends to use data from only one SAR polarization, when a combination of polarizations can provide insight into whether water is present. [2] In order to alleviate these problems, this paper presents an application for the Open Data Cube to detect water from SAR data using support vector machine (SVM) classification. 2. PLATFORM WASARD is an application for the Open Data Cube, a mechanism which provides a simple yet efficient means of ingesting, storing, and retrieving remote sensing data. Data can be ingested and made analysis ready according to whatever specifications the researcher chooses, and easily resampled to artificially alter a scene’s resolution. Currently WASARD supports water detection on scenes from ESA’s Sentinel-1 and JAXA’s ALOS. When testing WASARD, Sentinel-1 was most commonly used due to its relatively high spatial resolution and its rapid 6 day revisit cycle [5]. With minor alterations to the application's code, however, it could support data from other satellites. 3. METHODOLOGY Using supervised classification, WASARD compares SAR data to a dataset pre-classified by WOFS in order to train an SVM classifier. This classifier is then used to detect water in other SAR scenes outside the training set. Accuracy was measured according to the following metrics:  Precision: a measure of what percentage of the points WASARD labels as water are truly water  Recall: a measure of what percentage of the total water cover WASARD was able to identify.  F1 Score: a harmonic average of the precision and recall scores Both precision and recall are calculated at the end of the training phase, when the trained classifier is compared to a testing dataset. Because the WOFS algorithm’s classifications are used as the truth values when training a WASARD classifier, when precision and recall are mentioned in this paper, they are always with respect to the values produced by WOFS on a similar scene of Landsat data, which themselves have a classification accuracy of 97% [6]. Visual representations of water identified by WASARD in this paper were produced using the function wasard_plot(), which is included in WASARD. 3.1 Algorithm Selection The machine learning model used by WASARD is the Linear Support Vector Machine (SVM). This model uses a supervised learning algorithm to develop a classifier, meaning it creates a vector which can be multiplied by the vector formed by the relevant data bands to determine whether a pixel in a SAR scene contains water. This classifier is trained by comparing data points from selected bands in a SAR scene to their respective labels, which in this case are “water” or “not water” as given by the WOFS algorithm. The SVM was selected over the Random Forest model, which outperformed the SVM in training speed, but had a greater classification time and lower accuracy, and the Multilayer Perceptron Artificial Neural Network, which had a slightly higher average accuracy than the SVM, but much greater training and classification times. Figure 1: Visual representation of the SVM Classifier. Each white point represents a pixel in a SAR scene. In Figure 1, the diagonal line separating pixels determined to be water from those determined not to be water represents the actual classification vector produced by the SVM. It is worth noting that once the model has been trained, classification of pixels is done in a similar manner as in the thresholding method. This is especially true if only one band was used to train the model. 3.1 Feature Selection Sentinel-1 collects data from two bands: the Vertical/Vertical polarization (VV) and the Vertical/Horizontal polarization (VH). When 100 SVM classifiers were created for each polarization individually, and for the combination of the two, the following results were achieved: Figure 2: Accuracy of classifiers trained using different polarization bands. Precision and Recall were measured with respect to the values produced by WOFS. Figure 2 demonstrates that using both the VV and VH bands trades slightly lower recall for significantly greater precision when compared with the VH band alone, and that using the VV band alone is inferior in both metrics. WASARD therefore defaults to using both the VV and VH bands, and includes the option to use solely the VH band. The VV polarization’s lower precision compared to the VH polarization is in contrast to results from previous research and may merit further analysis [4]. 3.2 Training a Classifier The steps in training a classifier with WASARD are 1. Selecting two scenes (one SAR, one optical) with the same spatial extents, and acquired close to each other in time, with a preference that the scenes are taken on the same day. 2. Using the WOFS algorithm to produce an array of the detected water in the scene of optical data, to be used as the labels during supervised learning 3. Data points from the selected bands from the SAR acquisition are bundled together into an array with the corresponding labels gathered from WOFS. A random sample with an equal number of points labeled “Water” and “Not Water” is selected to be partitioned into a training and a testing dataset 4. Using Scikit-Learn’s LinearSVC object, the training dataset is used to produce a classifier, which is then tested against the testing dataset to determine its precision and recall The result is a wasard_classifier object, which has the following attributes: 1. f1, recall, and precision: 3 metrics used to determine the classifier’s accuracy 2. Coefficient: Vector which the SVM uses to make its predictions. The classifier detects water when the dot product of the coefficient and the vector formed by the SAR bands is positive 3. Save(): allows a user to save a classifier to the disk in order to use it without retraining 4. wasard_classify(): Classifies an entire xarray of SAR data using the SVM classifier All of the above steps are performed automatically when the user creates a wasard_classifier object. 3.3 Classifying a Dataset Once the classifier has been created, it can be used to detect water in an xarray of SAR data using wasard_classify(). By taking the dot product of the classifier’s coefficients and the vector formed by the selected bands of SAR data, an array of predictions is constructed. A classifier can effectively be used on the same spatial extents as the ones where it was trained, or on any area with a similar landscape. While

Kreiser, Zachary↗