Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “clustering algorithm”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 271 records · Page 15

Rule groupings in expert systems

Currently, expert system shells do not address software engineering issues for developing or maintaining expert systems. As a result, large expert systems tend to be incomprehensible, difficult to debug or modify, and almost impossible to verify or validate Partitioning rule-based systems into rule groups which reflect the underlying subdomains of the problem should enhance the comprehensibility, maintainability, and reliability of expert-system software. In this paper, we investigate methods to semi-automatically structure a CLIPS rule base e into groups of rules that carry related information. We discuss three different distance metrics for measuring the relatedness of rules and describe two clustering algorithms based on these distance metrics. The results of our experiment with three sample rule bases are also presented.

Mehrotra, Mala↗

Fault Diagnosis of Power Systems Using Intelligent Systems

The power system operator's need for a reliable power delivery system calls for a real-time or near-real-time Al-based fault diagnosis tool. Such a tool will allow NASA ground controllers to re-establish a normal or near-normal degraded operating state of the EPS (a DC power system) for Space Station Alpha by isolating the faulted branches and loads of the system. And after isolation, re-energizing those branches and loads that have been found not to have any faults in them. A proposed solution involves using the Fault Diagnosis Intelligent System (FDIS) to perform near-real time fault diagnosis of Alpha's EPS by downloading power transient telemetry at fault-time from onboard data loggers. The FDIS uses an ANN clustering algorithm augmented with a wavelet transform feature extractor. This combination enables this system to perform pattern recognition of the power transient signatures to diagnose the fault type and its location down to the orbital replaceable unit. FDIS has been tested using a simulation of the LeRC Testbed Space Station Freedom configuration including the topology from the DDCU's to the electrical loads attached to the TPDU's. FDIS will work in conjunction with the Power Management Load Scheduler to determine what the state of the system was at the time of the fault condition. This information is used to activate the appropriate diagnostic section, and to refine if necessary the solution obtained. In the latter case, if the FDIS reports back that it is equally likely that the faulty device as 'start tracker #1' and 'time generation unit,' then based on a priori knowledge of the system's state, the refined solution would be 'star tracker #1' located in cabinet ITAS2. It is concluded from the present studies that artificial intelligence diagnostic abilities are improved with the addition of the wavelet transform, and that when such a system such as FDIS is coupled to the Power Management Load Scheduler, a faulty device can be located and isolated from the rest of the system. The benefit of these studies provides NASA with the ability to quickly restore the operating status of a space station from a critical state to a safe degraded mode, thereby saving costs in experimentation rescheduling, fault diagnostics, and prevention of loss-of-life.

Momoh, James A.↗

Approximation Of Multi-Valued Inverse Functions Using Clustering And Sugeno Fuzzy Inference

Finding the inverse of a continuous function can be challenging and computationally expensive when the inverse function is multi-valued. Difficulties may be compounded when the function itself is difficult to evaluate. We show that we can use fuzzy-logic approximators such as Sugeno inference systems to compute the inverse on-line. To do so, a fuzzy clustering algorithm can be used in conjunction with a discriminating function to split the function data into branches for the different values of the forward function. These data sets are then fed into a recursive least-squares learning algorithm that finds the proper coefficients of the Sugeno approximators; each Sugeno approximator finds one value of the inverse function. Discussions about the accuracy of the approximation will be included.

Walden, Maria A.↗

Genetic Network Inference: From Co-Expression Clustering to Reverse Engineering

Advances in molecular biological, analytical, and computational technologies are enabling us to systematically investigate the complex molecular processes underlying biological systems. In particular, using high-throughput gene expression assays, we are able to measure the output of the gene regulatory network. We aim here to review datamining and modeling approaches for conceptualizing and unraveling the functional relationships implicit in these datasets. Clustering of co-expression profiles allows us to infer shared regulatory inputs and functional pathways. We discuss various aspects of clustering, ranging from distance measures to clustering algorithms and multiple-duster memberships. More advanced analysis aims to infer causal connections between genes directly, i.e., who is regulating whom and how. We discuss several approaches to the problem of reverse engineering of genetic networks, from discrete Boolean networks, to continuous linear and non-linear models. We conclude that the combination of predictive modeling with systematic experimental verification will be required to gain a deeper insight into living organisms, therapeutic targeting, and bioengineering.

Dhaeseleer, Patrik↗

Observed and Simulated Radiative and Microphysical Properties of Tropical Convective Storms

Increases in the ice content, albedo and cloud cover of tropical convective storms in a warmer climate produce a large negative contribution to cloud feedback in the GISS GCM. Unfortunately, the physics of convective upward water transport, detrainment, and ice sedimentation, and the relationship of microphysical to radiative properties, are all quite uncertain. We apply a clustering algorithm to TRMM satellite microwave rainfall retrievals to identify contiguous deep precipitating storms throughout the tropics. Each storm is characterized according to its size, albedo, OLR, rain rate, microphysical structure, and presence/absence of lightning. A similar analysis is applied to ISCCP data during the TOGA/COARE experiment to identify optically thick deep cloud systems and relate them to large-scale environmental conditions just before storm onset. We examine the statistics of these storms to understand the relative climatic roles of small and large storms and the factors that regulate convective storm size and albedo. The results are compared to GISS GCM simulated statistics of tropical convective storms to identify areas of agreement and disagreement.

DelGenio, Anthony D.↗

Fatigue Crack Measurement in Composite Materials by Ultrasonic Methods

The nondestructive detection of intra-ply microcracking in unlined pressure vessels fabricated from composite materials is critical to ensuring mission success. Microcracking in composite structures due to combined fatigue and cryogenic thermal loading can be very troublesome to detect in-service and when it begins to link through the thickness can cause leakage and failure of the structure. These leaks may lead to loss of pressure/propellant, increased risk of explosion and possible cryo-pumping. The work presented herein develops a method and an instrument to locate and measure intraply fatigue cracking through the thickness of laminated composite material by means of correlation with ultrasonic resonance. Resonant ultrasound spectroscopy provides measurements which are, sensitive to both the microscopic and macroscopic properties of an object. Elastic moduli, acoustic attenuation, and geometry can all be probed. The approach is based on the premise of half-wavelength resonance. The method injects a broadband ultrasonic wave into the test structure using a swept frequency technique. This method provides dramatically increased energy input into the test article, as compared to conventional spike pulsed ultrasonics. This relative energy increase improves the ability to measure finer details in the materials character, such as micro-cracking and porosity. As the micro-crack density increases, more interactions occur with the higher frequency (small wavelength) components of the signal train causing the spectrum to shift toward lower frequencies. Preliminary experiments have verified a measurable effect on the resonance spectrum of the ultrasonic data to detect microcracking. Methods involving self organizing neural networks and other clustering algorithms show that the resonance ultrasound signatures from composites vary with the degree of microcracking and can be separated and identified.

Walker, James L.↗

Ultrasonic Characterization of Fatigue Cracks in Composite Materials

Microcracking in composite structures due to combined fatigue and cryogenic loading can cause leakage and failure of the structure and can be difficult to detect in-service. In aerospace systems, these leaks may lead to loss of pressure/propellant, increased risk of explosion and possible cryo-pumping. The success of nondestructive evaluation to detect intra-ply microcracking in unlined pressure vessels fabricated from composite materials is critical to the use of composite structures in future space systems. The work presented herein characterizes measurements of intraply fatigue cracking through the thickness of laminated composite material by means of correlation with ultrasonic resonance. Resonant ultrasound spectroscopy provides measurements which are sensitive to both the microscopic and macroscopic properties of the test article. Elastic moduli, acoustic attenuation, and geometry can all be probed. The approach is based on the premise of half-wavelength resonance. The method injects a broadband ultrasonic wave into the test structure using a swept frequency technique. This method provides dramatically increased energy input into the test article, as compared to conventional pulsed ultrasonics. This relative energy increase improves the ability to measure finer details in the materials characterization, such as microcracking and porosity. As the microcrack density increases, more interactions occur with the higher frequency (small wavelength) components of the signal train causing the spectrum to shift toward lower frequencies. Several methods are under investigation to correlate the degree of microcracking from resonance ultrasound measurements on composite test articles including self organizing neural networks, chemometric techniques used in optical spectroscopy and other clustering algorithms.

Workman, Gary L.↗

Fatigue Crack and Porosity Measurement in Composite Materials by Thermographic and Ultrasonic Methods

Many nondestructive methods exist for the detection of localized material anomalies in an otherwise good composite structure. The problem arises when the material system as a whole has degraded during service or was improperly manufactured. Porosity and intra-ply microcracking are two such conditions that in unlined composite pressure vessels can be very troublesome to detect and when linked through the thickness can be critical to mission success. These leak paths may lead to loss of pressure/propellant, increased risk of explosion and possible cryo-pumping. Research sought nondestructive methods for quantifying porosity and microcracking in composite tankage. Both thermographic and resonance ultrasound methods have been utilized with artificial neural network and statistical approaches to analyze the data. Resonant ultrasound spectroscopy provides measurements, which are sensitive to fine details in the materials character, such as micro-cracking and porosity. Here, the higher frequency (shorter wavelength) components of the signal train provide more significant interaction with the defects causing the spectral characteristics to shift toward lower amplitudes at the higher frequencies. As the density of the defects increases more interactions occur and more drastic amplitude changes are observed. From a thermal perspective, the higher the defect density the lower the through thickness thermal diffusivity will be. Utilizing a point heat source, and thermographically recording the heat profile with time, diffusivity calculations can be made which in turn can be related to the relative quality of the material. Preliminary experiments to verify the measurable effect on the resonance spectrum of the ultrasonic data to detect microcracking and for porosity detection thermographically are presented. Methods involving supervised and unsupervised artificial neural networks as well as other clustering algorithms are developed for signal identification.

Walker, James L.↗

Machine Learning for Biological Trajectory Classification Applications

Machine-learning techniques, including clustering algorithms, support vector machines and hidden Markov models, are applied to the task of classifying trajectories of moving keratocyte cells. The different algorithms axe compared to each other as well as to expert and non-expert test persons, using concepts from signal-detection theory. The algorithms performed very well as compared to humans, suggesting a robust tool for trajectory classification in biological applications.

Sbalzarini, Ivo F.↗

INDUCTIVE SYSTEM HEALTH MONITORING WITH STATISTICAL METRICS

Model-based reasoning is a powerful method for performing system monitoring and diagnosis. Building models for model-based reasoning is often a difficult and time consuming process. The Inductive Monitoring System (IMS) software was developed to provide a technique to automatically produce health monitoring knowledge bases for systems that are either difficult to model (simulate) with a computer or which require computer models that are too complex to use for real time monitoring. IMS processes nominal data sets collected either directly from the system or from simulations to build a knowledge base that can be used to detect anomalous behavior in the system. Machine learning and data mining techniques are used to characterize typical system behavior by extracting general classes of nominal data from archived data sets. In particular, a clustering algorithm forms groups of nominal values for sets of related parameters. This establishes constraints on those parameter values that should hold during nominal operation. During monitoring, IMS provides a statistically weighted measure of the deviation of current system behavior from the established normal baseline. If the deviation increases beyond the expected level, an anomaly is suspected, prompting further investigation by an operator or automated system. IMS has shown potential to be an effective, low cost technique to produce system monitoring capability for a variety of applications. We describe the training and system health monitoring techniques of IMS. We also present the application of IMS to a data set from the Space Shuttle Columbia STS-107 flight. IMS was able to detect an anomaly in the launch telemetry shortly after a foam impact damaged Columbia's thermal protection system.

Iverson, David L.↗

How Much Global Burned Area Can Be Forecast on Seasonal Time Scales Using Sea Surface Temperatures?

Large-scale sea surface temperature (SST) patterns influence the interannual variability of burned area in many regions by means of climate controls on fuel continuity, amount, and moisture content. Some of the variability in burned area is predictable on seasonal timescales because fuel characteristics respond to the cumulative effects of climate prior to the onset of the fire season. Here we systematically evaluated the degree to which annual burned area from the Global Fire Emissions Database version 4 with small fires (GFED4s) can be predicted using SSTs from 14 different ocean regions. We found that about 48 of global burned area can be forecast with a correlation coefficient that is significant at a p < 0.01 level using a single ocean climate index (OCI) 3 or more months prior to the month of peak burning. Continental regions where burned area had a higher degree of predictability included equatorial Asia, where 92% of the burned area exceeded the correlation threshold, and Central America, where 86% of the burned area exceeded this threshold. Pacific Ocean indices describing the El Nino-Southern Oscillation were more important than indices from other ocean basins, accounting for about 1/3 of the total predictable global burned area. A model that combined two indices from different oceans considerably improved model performance, suggesting that fires in many regions respond to forcing from more than one ocean basin. Using OCI-burned area relationships and a clustering algorithm, we identified 12 hotspot regions in which fires had a consistent response to SST patterns. Annual burned area in these regions can be predicted with moderate confidence levels, suggesting operational forecasts may be possible with the aim of improving ecosystem management.

Chen, Yang↗

FloodPlanet: High-Resolution Commercial Imagery for Training and Validation of Deep Learning-Based Models of Inundation Extent

Flooding events are becoming increasingly frequent worldwide and are known to cause extensive damage. Public optical and radar satellite imagery can be used to detect large areas of inundation in rural areas, however, long revisit times and coarse spatial resolution limit applications for short-lived events and urban areas. Commercial constellations such as those operated by Planet offer increased spatial and temporal resolution and can supplement mapping efforts to provide more information to disaster response, relief, and mitigation efforts. Deep learning requires high quality labeled data for training across coincident sensors. The FloodPlanet dataset presented here contains labeled surface water for 18 events across the world based on Planetscope imagery with coincident Harmonized Landsat Sentinel-2 ( HLS) or Sentinel-1 and builds upon the previously existing Sen1Floods11, xBD, and NASA Sentinel-1 datasets. Sen1Floods11 includes 4,831 512x512 pixel overlapping tiles of coincident Sentinel-1 and Sentinel-2 data observing 11 flood events across the world from 2017-2019. The dataset contains a combination of automated and hand-labeled surface water for use in training and validation of inundation modeling efforts. The xBD dataset identifies flood-damaged buildings and indicates the scale of damage to each (none, minor, moderate, and major) from four flood events which occurred in the United States, India, Nepal, and Bangladesh from the same time period. The NASA dataset contains hand-labeled water bodies observed in Sentinel-1 imagery during five flood events within the 2017-2019 period. The effort presented here utilizes observations from these previously investigated flood events to generate labels of surface water at the 3-5m spatial resolution provided by Planetscope and facilitate the comparison between public and commercial data. A data pipeline was built which uses clustering algorithms to pick the most suitable overlapping chips between the public data and PlanetScope data for manual labeling. Labels were created manually using NASA’s ImageLabeler tool and include areas of high- and low-confidence water. The high confidence designation is reserved for areas of open, unobstructed water while low confidence is used for areas of suspected water beneath vegetation, clouds, or cloud shadows. Expected to be released in late 2022, the FloodPlanet dataset will include tiled imagery with a unique ID for each 1024x1024 pixel tile, 7 bands of HLS data, and high- and low-confidence flood labels in both shapefile and tiff formats. The authors will follow Spatial Temporal Access Catalog (STAC) guidelines to release FloodPlanet on the Radiant Earth ML hub, which hosts public datasets for machine learning.

Alexander Melancon↗

Supporting Responsible Machine Learning in Heliophysics

Over the last decade, Heliophysics researchers have increasingly adopted a variety of machine learning methods such as artificial neural networks, decision trees, and clustering algorithms into their workflow. Adoption of these advanced data science methods had quickly outpaced institutional response, but many professional organizations such as the European Commission, the National Aeronautics and Space Administration (NASA), and the American Geophysical Union have now issued (or will soon issue) standards for artificial intelligence and machine learning that will impact scientific research. These standards add further (necessary) burdens on the individual researcher who must now prepare the public release of data and code in addition to traditional paper writing. Support for these is not reflected in the current state of institutional support, community practices, or governance systems. We examine here some of these principles and how our institutions and community can promote their successful adoption within the Heliophysics discipline.

Machine learning↗

How Have Hydrological Extremes Changed Over the Past 20 Years?

Severe floods and droughts, including their back-to-back occurrences (weather whiplash), have been increasing in frequency and severity around the world. Improved understanding of systematic changes in hydrological extremes is essential for preparation and adaptation. In this study, we identified and quantified extreme wet and dry events globally by applying a clustering algorithm to terrestrial water storage (TWS) data from the Gravity Recovery and Climate Experiment (GRACE) and GRACE Follow-On (FO). The most intense events, ranked using an intensity metric, often reflect impacts of large-scale oceanic oscillations such as El Niño–Southern Oscillation and consequences of climate change. The severity of both wet and dry events, represented by standardized TWS anomalies, increased significantly in most cases, likely associated with intensification of wet and dry weather regimes in a warmer world, and consequently, exhibited strongest correlation with global temperature. In the Dry climate, the number of wet events decreased while the number of dry events increased significantly, suggesting a drying trend that may be attributed to climate variability and possible increases in irrigation and reliance on groundwater. In the Continental climate where temperature has risen faster than global average, dry events increased significantly. Characteristics of extreme events often showed strong correlations with global temperature, especially when averaged over all climates. These results suggest changes in hydrological extremes and underscore the importance of quantifying total water storage changes when studying hydrological extremes. Extending the GRACE/FO record, which spans 2002 to the present, is essential to continuously tracking changes in TWS and hydrological extremes.

Bailing Li↗

Multiyear Dry Periods in Southern Africa

Characteristics and physical features related to low precipitation across many years in Southern Africa that lead to societal disruptions are diagnosed using observed analyses and an ensemble of historical coupled climate model simulations during 1921 to 2014. Four regions are evaluated, as identified through a hierarchical clustering algorithm applied to the Standardized Precipitation Index (SPI) during the October–April precipitation season. Although dryness spanning many October–April occurs periodically in each region, they seldom occur simultaneously, consistent with largely insignificant SPI cross-correlations between them. However, characteristics relevant to low precipitation across many years are generalizable between the four regions, including the serial persistence of October–April precipitation, the likelihood of consecutive dry October–April, and the likelihood of dry October–April in temporal extents of up to 10 consecutive such 7-month seasons. Systematic precipitation persistence is not a feature in any of the four Southern Africa regions, as serial correlations of October–April SPI are not statistically significant at any time lags. It follows that there is an exponential-folding decay in the likelihood of consecutive October–April for various SPI thresholds and that there is a large spread in the likelihood of low October–April SPI across many years. In terms of physical features, low October–April SPI in each Southern Africa region is closely related to local atmospheric circulations; however, they are not as closely related to sea surface temperatures (SSTs). These results suggest that dryness spanning many years is determined primarily by persistent local circulations related to atmospheric variability and to a lesser extent variability related to SST anomalies, including the El Niño–Southern Oscillation.

subtropical Indian Ocean dipole↗

Characterizing the California Current System through Sea Surface Temperature and Salinity

Characterizing temperature and salinity (T-S) conditions is a standard framework in oceanography to identify and describe deep water masses and their dynamics. At the surface, this practice is hindered by multiple air–sea–land processes impacting T-S properties at shorter time scales than can easily be monitored. Now, however, the unsurpassed spatial and temporal coverage and resolution achieved with satellite sea surface temperature (SST) and salinity (SSS) allow us to use these variables to investigate the variability of surface processes at climate-relevant scales. In this work, we use SSS and SST data, aggregated into domains using a cluster algorithm over a T-S diagram, to describe the surface characteristics of the California Current System (CCS), validating them with in situ data from uncrewed Saildrone vessels. Despite biases and uncertainties in SSS and SST values in highly dynamic coastal areas, this T-S framework has proven useful in describing CCS regional surface properties and their variability in the past and in real time, at novel scales. This analysis also shows the capacity of remote sensing data for investigating variability in land–air–sea interactions not previously possible due to limited in situ data.

Marisol García-Reyes↗

Discriminative Dimensionality Reduction using Deep Neural Networks for Clustering of LIGO Data

In this paper, leveraging the capabilities of neural networks for modeling the non-linearities that exist in the data, we propose several models that can project data into a low dimensional, discriminative, and smooth manifold. The proposed models can transfer knowledge from the domain of known classes to a new domain where the classes are unknown. A clustering algorithm is further applied in the new domain to find potentially new classes from the pool of unlabeled data. The research problem and data for this paper originated from the Gravity Spy project which is a side project of Advanced Laser Interferometer Gravitational-wave Observatory (LIGO). The LIGO project aims at detecting cosmic gravitational waves using huge detectors. However non-cosmic, non-Gaussian disturbances known as "glitches", show up in gravitational-wave data of LIGO. This is undesirable as it creates problems for the gravitational wave detection process. Gravity Spy aids in glitch identification with the purpose of understanding their origin. Since new types of glitches appear over time, one of the objective of Gravity Spy is to create new glitch classes. Towards this task, we offer a methodology in this paper to accomplish this.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Scaling Building Energy Audits through Machine Learning Methods on Novel Drone Image Data

Building energy audits are time-consuming and labor-intensive. This paper describes a new method using machine learning (ML) techniques on novel data sources (drone images) to improve the identification of building characteristics and retrofit opportunities, and thereby reduce the effort for audits. The new ML method includes: (1) Building footprint extraction using line extraction, polygonization, and polygon-merging, (2) Building envelope extraction using PIX4d modeling software to reconstruct a building 3D model, (3) Visualization tool for viewing images from the 3D model, (4) Window-to-wall ratio (WWR) using state-of-art deep neural network semantic segmentation, (5) Envelope thermal anomaly detection using an unsupervised machine learning clustering algorithm, and (6) Rooftop energy equipment detection based on an object detection algorithm. The testing of this method involved a comparison of additional ML-generated information overlaid on current ‘state-of-practice’ audit and remote assessment baselines using evaluation metrics: labor time and associated cost, marginal benefits of using ML-generated information in workflows for audits and remote assessments, integration potential with existing processes and tools, and replicability/scalability of the method. In two test buildings in California that had comprehensive drawings and meter data available, the ML method effectively generated a building footprint, envelope, rooftop equipment, WWR, and locations of envelope thermal anomalies. Projected target segments of the ML method are sites with minimal drawings and energy data, and underserved sectors such as multistoried housing, disadvantaged communities, and schools for which the ML method can enable identification of building asset characteristics and prioritization of envelope retrofits and decentralized energy equipment retrofits.

Singh, Reshma↗