Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “automatic relevance determination”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

33 records · Page 2

Automated phase segmentation and quantification of high-resolution TEM image for alloy design

In the alloy design and development process, a wealth of atomically resolved structural high-resolution transmission electron microscopy (HRTEM) images are produced. Identifying the different nano-precipitate phases and tracking their evolution under various compositions and during manufacturing or post-processing requires hundreds of HRTEM images and thousands of precipitates. The nanoscopic phase information labeling and analysis purely relies on humans are prohibitively costly and time-consuming, sometimes not reliable because of the lack of authoritative knowledge. Here, in this work, we develop a novel unsupervised machine learning approach coupled with adaptive computer vision techniques with features in the Fourier space to automatically determine the number of phases and segment/quantify the phases with nanoscale resolution, allowing for quantitative correlation between nanostructure formation, processing and functional properties. To automate the phase extraction/quantification and ascertain its applicability, we have applied the developed framework to the HRTEM images from several alloy systems, processing conditions, image magnifications, and phase types and morphologies (precipitates, nano-twins, stacking faults, crystalline matrix, and amorphous structures) for verification. This study paves the road for compression, visualization, and translation of raw image structural data into physically relevant information in real-time with minimal human supervision. It shows the promise of enabling high-throughput materials characterization for the acceleration of alloy manufacturing and design.

36 MATERIALS SCIENCE↗

Automated Defect Identification for Tri-structural Isotropic Fuels (AUDIT)

During the manufacture of tri-structural isotropic (TRISO)-coated nuclear fuel particles, the potential exists for the formation of internal fissure defects in the uranium oxycarbide (UCO) kernels. These fissures result in a defective fuel particle that can fracture during subsequent fuel processing. Therefore, it is necessary to detect the presence of fissured kernels in a batch to determine if the batch meets specification prior to blending with other batches and upgrading processes. Previous attempts at identifying fissures involved manual inspection of micrographs of UCO fuel kernel cross-sections. This process is tedious, time-consuming and may introduce counting errors making it a good candidate for automation. This work presents a method for the automated detection of fissures in UCO kernels. Image segmentation is used for the extraction of relevant features in the micrographs which then serve as the input to a convolutional neural network used to automatically distinguish between fissured and non-fissured kernels.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

AI-Batt (Autonomous Identification of Battery Life Models) [SWR 21-36]

Autonomous Identification of Battery Life Models (AI-Batt) AI-Batt is a MATLAB code base for developing lifetime models for batteries from accelerated aging data. The code base provides many functions for processing, visualizing, and modeling battery aging data, making the data processing, exploration, and modeling workflow substantially faster. These tools are tailored for working with battery aging data sets, which usually consist of many separate time-series for each cell, with many test conditions and possible replicates at each condition, which makes it difficult to simply process or visualize the data set. Complex modeling tasks, such as cross-validation, sensitivity analysis, and uncertainty quantification have been implemented to enable thorough statistical investigation of model predictions. Additionally, several machine-learning algorithms are implemented to autonomously identify suitable models via symbolic regression. Data processing functions automatically cast data from the struct data type, which is commonly used to store experimental data, but is not an acceptable input for most algorithms, to the table data type, which can be easily used as input to any optimization algorithm. Also, the data can be separated into time-invariant and time-variant data tables, which is helpful for exploring the data set as well as developing separate models for time-variant and time-invariant aging mechanisms. For example, in aging tests with constant temperature, temperature is a time-invariant experimental condition. Visualization tools enable plotting of data, model fits, and model simulations possible with single-line function calls, empowering data exploration of complex data sets with both time-varying and time-invariant trends. Plots can be automatically generated for the whole data set, or separated by data group (groups of test replicates) or individual data series. Data points or data series can be automatically colored by the value of a variable with a variety of color maps, and model predictions can also be colored by the value of a fit statistic. Comparisons between data sets and the predictions/simulations of different models on the same data set can be easily plotted as well. Distributions of parameter values from bootstrap resampling can be plotted to visualize the reliability of parameter estimation, or determine any correlations between parameters. Modeling tools handle the complex task of creating and parsing symbolic equations for modeling battery lifetime. Equations are parsed to grab relevant data variables, parameter values, or specified sub-models for input into optimization, evaluation, or simulation functions. Models can be optimized locally (one set of parameters for each data series), bi-level (some parameters shared across the data set), or globally (single set of parameters for all data). Functions implementing symbolic regression algorithms help users to discover effective model equations, even in poorly sampled, high-dimensional data.

Smith, Kandler↗

Semi-Supervised Data Summarization: Using Spectral Libraries to Improve Hyperspectral Clustering

Hyperspectral imagers produce very large images, with each pixel recorded at hundreds or thousands of different wavelengths. The ability to automatically generate summaries of these data sets enables several important applications, such as quickly browsing through a large image repository or determining the best use of a limited bandwidth link (e.g., determining which images are most critical for full transmission). Clustering algorithms can be used to generate these summaries, but traditional clustering methods make decisions based only on the information contained in the data set. In contrast, we present a new method that additionally leverages existing spectral libraries to identify materials that are likely to be present in the image target area. We find that this approach simultaneously reduces runtime and produces summaries that are more relevant to science goals.

Wagstaff, K. L.↗

Optical emissivity dataset of multi-material heterogeneous designs generated with automated figure extraction

Optical device design is typically an iterative optimization process based on a good initial guess from prior reports. Optical properties databases are useful in this process but difficult to compile because their parsing requires finding relevant papers and manually converting graphical emissivity curves to data tables. Here, we present two contributions: one is a dataset of thermal emissivity records with design-related parameters, and the other is a software tool for automated colored curve data extraction from scientific plots. We manually collected 64 papers with 176 figures reporting thermal emissivity and automatically retrieved 153 colored curve data records. The automated figure analysis software pipeline uses Faster R-CNN for axes and legend object detection, EasyOCR for axes numbering recognition, and k-means clustering for colored curve retrieval. Additionally, we manually extracted geometry, materials, and method information from the text to add necessary metadata to each emissivity curve. Finally, we analyzed the dataset to determine the dominant classes of emissivity curves and determine the underlying design parameters leading to a type of emissivity profile.

47 OTHER INSTRUMENTATION↗

Automatic discovery of optimal classes

A criterion, based on Bayes' theorem, is described that defines the optimal set of classes (a classification) for a given set of examples. This criterion is transformed into an equivalent minimum message length criterion with an intuitive information interpretation. This criterion does not require that the number of classes be specified in advance, this is determined by the data. The minimum message length criterion includes the message length required to describe the classes, so there is a built in bias against adding new classes unless they lead to a reduction in the message length required to describe the data. Unfortunately, the search space of possible classifications is too large to search exhaustively, so heuristic search methods, such as simulated annealing, are applied. Tutored learning and probabilistic prediction in particular cases are an important indirect result of optimal class discovery. Extensions to the basic class induction program include the ability to combine category and real value data, hierarchical classes, independent classifications and deciding for each class which attributes are relevant.

Cheeseman, Peter↗

Assessment of Flow-Enhanced Electrochemical Sensor Testing and Deployments for MSRs

This report serves as the deliverable for Milestone M3RS-23AN0401061 that is part of Work Package RS-23AN040106 (Flow Enhanced Sensors for MSRs – ANL). The goal of this milestone was to determine performance of the flow enhanced electrochemical sensor (FEES) and modular flow instrumentation testbed (MFIT) in safeguards relevant scenarios. Flow enhanced electrochemical sensors are a type of electroanalytical sensor that has been developed at Argonne National Laboratory to be installed directly into MSR flow conduits to make measurements of the salt composition. These sensors represent a significant improvement in capabilities compared to earlier electroanalytical sensors that instead can only be operated in quiescent conditions. Previous work has focused on testing of the FEES in flowing conditions provided by the MFIT to assess the accuracy and precision of the sensor measurements. To further improve this capability, in FY23 we undertook a campaign of safeguards relevant scenarios in molten salt containing a range of uranium chloride concentrations (0 to 3 wt%). All the testing carried out in FY23 was aided by a control system designed to automatically actuate flow conditions and collect data. This new automation system is estimated to have increased experimental throughput by a factor of four and enabled testing in a variety of complex conditions. The advancements in throughput and repeatability led to improved quantification of actinide concentrations using the in-flow sensors, with a reduction of the mean absolute relative error from 5.6% in FY22 to 3.1% in FY23. In addition to safeguards scenarios run in the MFIT, FY23 work included deployment of a FEES at a partner institution where it will be tested in a flowing salt loop. The FEES was successfully integrated into that loop and is being tested prior to loop startup. In FY23, work also continued on the smaller flow system that we have named the mini-MFIT. This smaller system is capable of rapid prototyping of new sensor designs prior to installation in the larger MFIT radiological flow system. Work was carried out to test this new system in non-radiological molten salts in a separate glovebox. This work is helping us to enhance the accuracy of our salt monitoring capabilities through the integration of multiple types of sensors. The high degree of accuracy required by 10 CFR 74 represents a significant challenge, and further design evolution and integration of the sensors into multimodal sensing frameworks will be needed to further push the measurement accuracy to the needed level.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

BFSVBF (BatFinder Smart Video BioFilter) [SWR-22-87] and Multi-class BatFinder Smart Video BioFilter Keras

Bats are notoriously difficult to study, therefore, identifying specific behavioral trends and the precise environmental conditions at the time of collision requires a monitoring solution that can reliably collect relevant data. To date, thermal infrared video surveillance has been extensively applied to study bats and has proven to be a powerful yet cumbersome tool. Current analytical approaches are time consuming because data processing data has not been fully automated. In the past, steps have been taken to record avian and bat activity in conjunction with complicated image processing techniques that separate species from other moving objects within the field of view (i.e. clouds and portions of the wind turbine). Once the videos are collected, the post-processing does not allow real time monitoring and identification, leading to a delay in both studying the behavior of these species and determining the effectiveness of any impact reduction strategy being studied. Moreover, object identification capability is lacking, thus limiting the usefulness of video data. To resolve these issues, we are using open source computer vision and machine learning techniques allowing for automatic detection of objects in real-time with the ability to correlate these objects with environmental variables and recording the flight paths of each object. The code has gone through five rounds of development with images used to train the models. This advancement allows for automated real-time data collection, identification, and tracking, thereby eliminating the need for long and tedious post-analysis processing of the videos. We will discuss the two open source and publicly available machine learning models developed within this scope of this work: 1) a binary model with a 97.5% accuracy in identifying the difference between an object and an empty scene, including wind turbine and clouds; and 2) a multiple classification model with the capability of identifying the type of object detected: bats (90% accuracy), birds (83% accuracy), insects (69% accuracy) and non-biological (99% accuracy).

Yarbrough, John↗

3DBFSVBF (3D BatFinder Smart Video BioFilter and Multi-class BatFinder Smart Video BioFilter) [SWR-22-88]

Bats are notoriously difficult to study, therefore, identifying specific behavioral trends and the precise environmental conditions at the time of collision requires a monitoring solution that can reliably collect relevant data. To date, thermal infrared video surveillance has been extensively applied to study bats and has proven to be a powerful yet cumbersome tool. Current analytical approaches are time consuming because data processing data has not been fully automated. In the past, steps have been taken to record avian and bat activity in conjunction with complicated image processing techniques that separate species from other moving objects within the field of view (i.e. clouds and portions of the wind turbine). Once the videos are collected, the post-processing does not allow real time monitoring and identification, leading to a delay in both studying the behavior of these species and determining the effectiveness of any impact reduction strategy being studied. Moreover, object identification capability is lacking, thus limiting the usefulness of video data. To resolve these issues, we are using open source 3D computer vision and machine learning techniques allowing for automatic detection of objects in real-time with the ability to correlate these objects with environmental variables and recording the flight paths of each object. The machine learning has been trained on 3D data and allows for automated real-time data collection, identification and tracking, thereby eliminating the need for long and tedious post-analysis processing of the videos. This machine learning model is an added feature to the previous BatFinder Smart Video BioFilter and increases the accuracy of that systems classification by increasing the accuracy of identifying bats (90% accuracy) and insects (69% accuracy) to a 97% accuracy. There are two object classifier machine learning models, Binary and multi-classification. Binary object classifier labeled BatFinder_Smart_Video_BioFilter.h5 distinguishes between biological objects and non-biological objects. The main goal of this object classifier is to ignore the turbine blades while detecting biological object flying withing the rotor swept area of the turbine. Non-biological objects have a probability of 0 and biological objects have a probability of 1. Multi-classifier labeled Multiclass_BatFinder_Smart_Video_BioFilter.h5 distinguishes between bats, birds, insects and non-biological.

Yarbrough, John↗

Reliable and Efficient Machine Learning (Final Technical Report)

Modern scientific experiments generate massive amounts of data at a pace much faster than humans can manually analyze. While machine learning has revolutionized commercial data analysis (such as recommending movies or recognizing faces), applying these tools to complex scientific discovery is challenging because scientific answers must be precise, interpretable, and adhere to physical laws. The research under this project aims to develop new mathematical tools and computer algorithms specifically designed for scientific applications. Major progress has been made in automatically cleaning and deconstructing messy experimental data, analyzing the visual information of physical phenomena, determining the underlying physical variables, and providing rig orous mathematical analysis of interesting algorithms and concepts widely used in machine learning. This project addressed the critical gap between our ability to generate massive scientific data and our ability to extract interpretable information from it. We established mathematical foundations for Scientific Machine Learning (SciML) aimed at effective data analytics and automated discovery. Our work focused on three core objectives: (1) developing reliable feature extraction methods for dynamic high-dimensional data, (2) establishing mathematical foundations for discovering dynamics via neural networks, and (3) creating rigorous optimization techniques for these models. Key outcomes come from two fronts. On the practical side, they include the development of algorithms that significantly enhance the extraction of signals from field data, as well as the capability to handle situations that exhibit smooth variations or physical stretching due to temperature changes. They also include the creation of an automated framework for discovering fundamental state variables from raw experimental data, demonstrating the ability to identify intrinsic physical dimensions without prior knowledge of the governing laws. On the theoretical front, the research results in theoretical advances in Optimal Transport, a widely used notion in SciML, specifically regarding functions with fixed-size nodal sets, provide sharp bounds relevant to uncertainty quantification. Meanwhile, the outcomes also include the establishment of convergence theories for nonlocal gradient descent methods, enabling robust optimization with noisy data in high-dimensional settings commonly encountered in scientific modeling. The project also helps creating opportunities to train the next generation of researchers, equipping them with the necessary technical skills for today’s workplace and preparing them for future advances.

97 MATHEMATICS AND COMPUTING↗

Automatic determination of fault effects on aircraft functionality

The problem of determining the behavior of physical systems subsequent to the occurrence of malfunctions is discussed. It is established that while it was reasonable to assume that the most important fault behavior modes of primitive components and simple subsystems could be known and predicted, interactions within composite systems reached levels of complexity that precluded the use of traditional rule-based expert system techniques. Reasoning from first principles, i.e., on the basis of causal models of the physical system, was required. The first question that arises is, of course, how the causal information required for such reasoning should be represented. The bond graphs presented here occupy a position intermediate between qualitative and quantitative models, allowing the automatic derivation of Kuipers-like qualitative constraint models as well as state equations. Their most salient feature, however, is that entities corresponding to components and interactions in the physical system are explicitly represented in the bond graph model, thus permitting systematic model updates to reflect malfunctions. Researchers show how this is done, as well as presenting a number of techniques for obtaining qualitative information from the state equations derivable from bond graph models. One insight is the fact that one of the most important advantages of the bond graph ontology is the highly systematic approach to model construction it imposes on the modeler, who is forced to classify the relevant physical entities into a small number of categories, and to look for two highly specific types of interactions among them. The systematic nature of bond graph model construction facilitates the process to the point where the guidelines are sufficiently specific to be followed by modelers who are not domain experts. As a result, models of a given system constructed by different modelers will have extensive similarities. Researchers conclude by pointing out that the ease of updating bond graph models to reflect malfunctions is a manifestation of the systematic nature of bond graph construction, and the regularity of the relationship between bond graph models and physical reality.

Feyock, Stefan↗

Mode Transitions in Glass Cockpit Aircraft: Results of a Field Study

One consequence of increased levels of automation in complex control systems is the presence of modes. A mode is a particular configuration of a control system that defines how human command inputs are interpreted. In complex systems, modes also often determine a specific allocation of control authority between the human and automated systems. Even in simple static devices (e.g., electronic watches, word processors), the presence of modes has been found to cause problems in either-the acquisition or production of skilled performance. Many of these problems arise due to the fact that the selection of a mode causes device behavior to be mediated by hidden internal state information. For these simple systems, many of these interaction problems can be solved by the design of appropriate feedback to communicate internal state information to the human operator. In complex dynamic systems, however, the design issues associated with modes seem to trancend the problem of merely communicating internal state information via displayed feedback. In complex supervisory control systems (e.g., aircraft, spacecraft, military command and control), a key function of modes is the selection of a particular configuration of control authority between the human operator and automated control systems. One mode may result in full manual control, another may result in a mix of manual and automatic control, while a third may result in full automatic control over the entire system. The human operator selects an appropriate mode as a function of current goals, operating conditions, and operating procedures. Thus, the operator is put in a position of essentially trying to control two coupled dynamic systems: the target system itself, and also a highly complex suite of automation controlling the target system. From a historical perspective, it should probably not come as a surprise that very little information is available to guide the design of mode-oriented control systems. The topic of function allocation (i.e., the proper division of control authority among human and computer) has a long history in human-machine systems research. Although this research has produced some relevant guidelines, a design approach capable of defining appropriate allocations of control function between the human and automation is not yet available. As a result, the function allocation decision itself has been allocated to the operator, to be performed in real-time, in the operation of mode-oriented control systems. A variety of documented aircraft accidents and incidents suggest that the real-time selection and monitoring of control modes is a weak link in the effective operation of complex supervisory control systems. Research in human-machine systems and human-computer interaction has barely scraped the surface of the problem of understanding how operators manage this task.The purpose of this paper is to present the results of a field study which examined how operators manage mode selection in a complex supervisory control system. Data on mode engagements using the Boeing B757/767 auto-flight system were collected during approach and descent into four major airports in the East Coast of the United States. Protocols documenting mode selection, automatic mode changes, pilot actions, quantitative records of flight-path variables, and verbal reports during and after mode engagements were collected by an observer from the jumpseat. Observations were conducted on two typical trips between three airports. Each trip was be replicated 11 times, which yielded a total of 22 trips and 66 legs on which data were collected. All data collected concerned the same flight numbers, and therefore, the same time of day, same type of aircraft, and identical operational environments (e.g., ATC facilities, weather patterns, traffic flow etc.)

Degani, Asaf↗

Energy Intensity Baselining and Tracking Guidance

Each company joining the U.S. Department of Energy’s (DOE’s) Better Buildings, Better Plants Program (Better Plants) commits to establishing an energy consumption and energy intensity (EI) baseline and to tracking its energy performance over a 10-year period against that baseline. The baseline must reflect a company’s energy consumption over a 12-month period, covering all its U.S.-based operations. Energy consumption is calculated by fuel type in terms of primary energy (also known as source energy). EI is broadly defined as the amount of energy consumed per unit of output produced. For this guidance document and for the program, the term energy performance represents an evaluation of a facility’s capacity to use energy efficiently. Metrics used to assess a facility’s energy performance can include EI, energy consumption, improvements in EI, etc. Establishing an energy baseline and tracking system is a critical first step in effectively managing energy use. Developing a baseline can help a company understand energy use within the corporation and give it a point of comparison to evaluate future efforts to improve energy performance. It can also support efforts to validate a company’s energy management activities, improve comparative analyses when using benchmarks, and help in predicting future energy needs. In addition, a company that normalizes its performance data can determine highly defensible measures of energy savings generated through implemented energy efficiency projects. Establishing a baseline and tracking energy performance is also a requirement for ISO 50001 certification. Although basic energy data can be collected through utility bills, most manufacturers will have to perform additional analyses to develop accurate and robust energy baselines and tracking systems. Energy is consumed in many ways within the manufacturing sector and can come from multiple sources. Energy is sometimes generated and sold to other parties or captured and reused on-site. External events can exert a significant impact on a facility or company’s energy use independent of any purposeful efforts to improve energy efficiency. Operational changes, such as production shifts—which may be inevitable for some companies over the 10-year period covered by the program—can also make a big difference in energy use. Since Better Plants asks companies to account for all their U.S.-based operations, mergers, acquisitions, and divestitures can also have significant implications for a company’s energy metrics. This document aims to demystify the sometimes complex baselining process. It devotes special attention to the task of normalizing and adjusting energy consumption to account for external factors, such as weather and production changes. A key recommendation is that companies use regression analysis to normalize their energy consumption data whenever possible. Regression analysis is a statistical technique that estimates the dependence of a variable (i.e., energy use in the context of Better Plants) on one or more independent variables such as ambient temperature, while controlling for the influence of other variables at the same time. A properly developed regression analysis can provide a reliable estimate of energy savings resulting from energy improvement strategies and projects by accounting for the effects of variables such as annual production levels and weather. DOE has developed a companion Energy Performance Indicator software tool (EnPI) to simplify the baselining process. This tool can run regression models, calculate changes in EI at the facility level, and automatically compile facility-level data into a corporate-wide metric. Note that although the relevant equations used to calculate EI are provided in this document, the EnPI tool will automatically perform most calculations for the user. Additionally, Better Plants Partners (Partners) can call on their Technical Account Manager (TAM) to help them establish a baseline and assist with the necessary calculations to track progress.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Application and Certification of Comparative Vacuum Monitoring Sensors for Structural Health Monitoring of 737 Wing Box Fittings

Multi-site fatigue damage, hidden cracks in hard-to-reach locations, disbonded joints, erosion, impact, and corrosion are among the major flaws encountered in today's extensive fleet of aging aircraft and space vehicles. The use of in-situ sensors for real-time health monitoring of aircraft structures are a viable option to overcome inspection impediments stemming from accessibility limitations, complex geometries, and the location and depth of hidden damage. Reliable, structural health monitoring systems can automatically process data, assess structural condition, and signal the need for human intervention. Prevention of unexpected flaw growth and structural failure can be improved if on-board health monitoring systems are used to continuously assess structural integrity. Such systems are able to detect incipient damage before catastrophic failures occurs. Condition-based maintenance practices could be substituted for the current time-based maintenance approach. Other advantages of on-board distributed sensor systems are that they can eliminate costly, and potentially damaging, disassembly, improve sensitivity by producing optimum placement of sensors and decrease maintenance costs by eliminating more time- consuming manual inspections. This report presents a Sandia Labs-aviation industry effort to move SHM into routine use for aircraft maintenance. This program addressed formal SHM technology validation and certification issues so that the full spectrum of concerns, including design, deployment, performance and certification were appropriately considered. The Airworthiness Assurance NDI Validation Center (AANC) at Sandia Labs, in conjunction with Boeing, Delta Air Lines, Structural Monitoring Systems Ltd., Anodyne Electronics Manufacturing Corp. and the Federal Aviation Administration (FAA) carried out a certification program to formally introduce Comparative Vacuum Monitoring (CVM) as a structural health monitoring solution to a specific aircraft wing box application. Validation tasks were designed to address the SHM equipment, the health monitoring task, the resolution required, the sensor interrogation procedures, the conditions under which the monitoring will occur, the potential inspector population, adoption of CVM into an airline maintenance program and the document revisions necessary to allow for routine use of CVM as an alternate means of performing periodic structural inspects. To carry out the validation process, knowledge of aircraft maintenance practices was coupled with an unbiased, independent evaluation. Sandia Labs designed, implemented, and analyzed the results from a focused and statistically-relevant experimental effort to quantify the reliability of the CVM system applied to the Boeing 737 Wing Box fitting application. All factors that affect SHM sensitivity were included in this program: flaw size, shape, orientation and location relative to the sensors, as well as operational and environmental variables. Statistical methods were applied to performance data to derive Probability of Detection (POD) values for CVM sensors in a manner that agrees with current nondestructive inspection (NDI) validation requirements and also is acceptable to both the aviation industry and regulatory bodies. This report presents the use of several different statistical methods, some of them adapted from NDI performance assessments and some proposed to address the unique nature of damage detection via SHM systems, and discusses how they can converge to produce a confident quantification of SHM performance An important element in developing SHM validation processes is a clear understanding of the regulatory measures needed to adopt SHM solutions along with the knowledge of the structural and maintenance characteristics that may impact the operational performance of an SHM system. This report describes the major elements of an SHM validation approach and differentiates the SHM elements from those found in NDI validation. The activities conducted in this program demonstrated the feasibility of routine SHM usage in general and CVM in particular for the application selected. They also helped establish an optimum OEM-airline-regulator process and determined how to safely adopt SHM solutions. This formal SHM validation will allow aircraft manufacturers and airlines to confidently make informed decisions about the proper utilization of CVM technology. It will also streamline the regulatory actions and formal certification measures needed to assure the safe application of SHM solutions.

42 ENGINEERING↗

Water Across Synthetic Aperture Radar Data (WASARD): SAR Water Body Classification for the Open Data Cube

The detection of inland water bodies from Synthetic Aperture Radar (SAR) data provides a great advantage over water detection with optical data, since SAR imaging is not impeded by cloud cover. Traditional methods of detecting water from SAR data involves using thresholding methods that can be labor intensive and imprecise. This paper describes Water Across Synthetic Aperture Radar Data (WASARD): a method of water detection from SAR data which automates and simplifies the thresholding process using machine learning on training data created from Geoscience Australia’s WOFS algorithm. Of the machine learning models tested, the Linear Support Vector Machine was determined to be optimal, with the option of training using solely the VH polarization or a combination of the VH and VV polarizations. WASARD was able to identify water in the target area with a correlation of 97% with WOFS. Sentinel-1, Open Data Cube, Earth Observations, Machine Learning, Water Detection 1. INTRODUCTION Water classification is an important function of Earth imaging satellites, as accurate remote classification of land and water can assist in land use analysis, flood prediction, climate change research, as well as a variety of agricultural applications [2]. The ability to identify bodies of water remotely via satellite is immensely cheaper than contracting surveys of the areas in question, meaning that an application that can accurately use satellite data towards this function can make valuable information available to nations which would not be able to afford it otherwise. Highly reliable applications for the remote detection of water currently exist for use with optical satellite data such as that provided by LANDSAT. One such application, Geoscience Australia’s Water Observations from Space (WOFS) has already been ported for use with the Open Data Cube [6]. However, water detection using optical data from Landsat is constrained by its relatively long revisit cycle of 16 days [5], and water detection using any optical data is constrained in that it lacks the ability to make accurate classifications through cloud cover [2]. The alternative solution which solves these problems is water detection using SAR data, which images the Earth using cloud-penetrating microwaves. Because of its advantages over optical data, much research has been done into water detection using SAR data. Traditionally, this has been done using the thresholding method, which involves picking a polarization band and labeling all pixels for which this band’s value is below a certain threshold as containing water. The thresholding method works since water tends to return a much lower backscatter value to the satellite than land [1]. However, this method can be flawed since estimating the proper threshold is often imprecise, complicated, and labor intensive for the end user. Thresholding also tends to use data from only one SAR polarization, when a combination of polarizations can provide insight into whether water is present. [2] In order to alleviate these problems, this paper presents an application for the Open Data Cube to detect water from SAR data using support vector machine (SVM) classification. 2. PLATFORM WASARD is an application for the Open Data Cube, a mechanism which provides a simple yet efficient means of ingesting, storing, and retrieving remote sensing data. Data can be ingested and made analysis ready according to whatever specifications the researcher chooses, and easily resampled to artificially alter a scene’s resolution. Currently WASARD supports water detection on scenes from ESA’s Sentinel-1 and JAXA’s ALOS. When testing WASARD, Sentinel-1 was most commonly used due to its relatively high spatial resolution and its rapid 6 day revisit cycle [5]. With minor alterations to the application's code, however, it could support data from other satellites. 3. METHODOLOGY Using supervised classification, WASARD compares SAR data to a dataset pre-classified by WOFS in order to train an SVM classifier. This classifier is then used to detect water in other SAR scenes outside the training set. Accuracy was measured according to the following metrics:  Precision: a measure of what percentage of the points WASARD labels as water are truly water  Recall: a measure of what percentage of the total water cover WASARD was able to identify.  F1 Score: a harmonic average of the precision and recall scores Both precision and recall are calculated at the end of the training phase, when the trained classifier is compared to a testing dataset. Because the WOFS algorithm’s classifications are used as the truth values when training a WASARD classifier, when precision and recall are mentioned in this paper, they are always with respect to the values produced by WOFS on a similar scene of Landsat data, which themselves have a classification accuracy of 97% [6]. Visual representations of water identified by WASARD in this paper were produced using the function wasard_plot(), which is included in WASARD. 3.1 Algorithm Selection The machine learning model used by WASARD is the Linear Support Vector Machine (SVM). This model uses a supervised learning algorithm to develop a classifier, meaning it creates a vector which can be multiplied by the vector formed by the relevant data bands to determine whether a pixel in a SAR scene contains water. This classifier is trained by comparing data points from selected bands in a SAR scene to their respective labels, which in this case are “water” or “not water” as given by the WOFS algorithm. The SVM was selected over the Random Forest model, which outperformed the SVM in training speed, but had a greater classification time and lower accuracy, and the Multilayer Perceptron Artificial Neural Network, which had a slightly higher average accuracy than the SVM, but much greater training and classification times. Figure 1: Visual representation of the SVM Classifier. Each white point represents a pixel in a SAR scene. In Figure 1, the diagonal line separating pixels determined to be water from those determined not to be water represents the actual classification vector produced by the SVM. It is worth noting that once the model has been trained, classification of pixels is done in a similar manner as in the thresholding method. This is especially true if only one band was used to train the model. 3.1 Feature Selection Sentinel-1 collects data from two bands: the Vertical/Vertical polarization (VV) and the Vertical/Horizontal polarization (VH). When 100 SVM classifiers were created for each polarization individually, and for the combination of the two, the following results were achieved: Figure 2: Accuracy of classifiers trained using different polarization bands. Precision and Recall were measured with respect to the values produced by WOFS. Figure 2 demonstrates that using both the VV and VH bands trades slightly lower recall for significantly greater precision when compared with the VH band alone, and that using the VV band alone is inferior in both metrics. WASARD therefore defaults to using both the VV and VH bands, and includes the option to use solely the VH band. The VV polarization’s lower precision compared to the VH polarization is in contrast to results from previous research and may merit further analysis [4]. 3.2 Training a Classifier The steps in training a classifier with WASARD are 1. Selecting two scenes (one SAR, one optical) with the same spatial extents, and acquired close to each other in time, with a preference that the scenes are taken on the same day. 2. Using the WOFS algorithm to produce an array of the detected water in the scene of optical data, to be used as the labels during supervised learning 3. Data points from the selected bands from the SAR acquisition are bundled together into an array with the corresponding labels gathered from WOFS. A random sample with an equal number of points labeled “Water” and “Not Water” is selected to be partitioned into a training and a testing dataset 4. Using Scikit-Learn’s LinearSVC object, the training dataset is used to produce a classifier, which is then tested against the testing dataset to determine its precision and recall The result is a wasard_classifier object, which has the following attributes: 1. f1, recall, and precision: 3 metrics used to determine the classifier’s accuracy 2. Coefficient: Vector which the SVM uses to make its predictions. The classifier detects water when the dot product of the coefficient and the vector formed by the SAR bands is positive 3. Save(): allows a user to save a classifier to the disk in order to use it without retraining 4. wasard_classify(): Classifies an entire xarray of SAR data using the SVM classifier All of the above steps are performed automatically when the user creates a wasard_classifier object. 3.3 Classifying a Dataset Once the classifier has been created, it can be used to detect water in an xarray of SAR data using wasard_classify(). By taking the dot product of the classifier’s coefficients and the vector formed by the selected bands of SAR data, an array of predictions is constructed. A classifier can effectively be used on the same spatial extents as the ones where it was trained, or on any area with a similar landscape. While

Kreiser, Zachary↗