Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Statistical Algorithms”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10

Systematic methods for knowledge acquisition and expert system development

Nine cooperating rule-based systems, collectively called AUTOCREW which were designed to automate functions and decisions associated with a combat aircraft's subsystems, are discussed. The organization of tasks within each system is described; performance metrics were developed to evaluate the workload of each rule base and to assess the cooperation between the rule bases. Simulation and comparative workload results for two mission scenarios are given. The scenarios are inbound surface-to-air-missile attack on the aircraft and pilot incapacitation. The methodology used to develop the AUTOCREW knowledge bases is summarized. Issues involved in designing the navigation sensor selection expert in AUTOCREW's NAVIGATOR knowledge base are discussed in detail. The performance of seven navigation systems aiding a medium-accuracy INS was investigated using Kalman filter covariance analyses. A navigation sensor management (NSM) expert system was formulated from covariance simulation data using the analysis of variance (ANOVA) method and the ID3 algorithm. ANOVA results show that statistically different position accuracies are obtained when different navaids are used, the number of navaids aiding the INS is varied, the aircraft's trajectory is varied, and the performance history is varied. The ID3 algorithm determines the NSM expert's classification rules in the form of decision trees. The performance of these decision trees was assessed on two arbitrary trajectories, and the results demonstrate that the NSM expert adapts to new situations and provides reasonable estimates of the expected hybrid performance.

Belkin, Brenda L.↗

Knowledge acquisition for expert systems using statistical methods

A common problem in the design of expert systems is the definition of rules from data obtained in system operation or simulation. A statistical method for generating rule bases from numerical data, motivated by an example based on aircraft navigation with multiple sensors is presented. The specific objective is to design an expert system that selects a satisfactory suite of measurements from a dissimilar, redundant set, given an arbitrary navigation geometry and possible sensor failures. The systematic development of a Navigation Sensor Management (NSM) Expert System from Kalman Filter covariance data is described. The development method invokes two statistical techniques: Analysis-of-Variance (ANOVA) and the ID3 algorithm. The ANOVA technique indicates whether variations of problem parameters give statistically different covariance results, and the ID3 algorithm identifies the relationships between the problem parameters using probabilistic knowledge extracted from a simulation example set.

Belkin, Brenda L.↗

DeepSAT: A Deep Learning Approach to Tree-Cover Delineation in 1-m NAIP Imagery for the Continental United States

High resolution tree cover classification maps are needed to increase the accuracy of current land ecosystem and climate model outputs. Limited studies are in place that demonstrates the state-of-the-art in deriving very high resolution (VHR) tree cover products. In addition, most methods heavily rely on commercial softwares that are difficult to scale given the region of study (e.g. continents to globe). Complexities in present approaches relate to (a) scalability of the algorithm, (b) large image data processing (compute and memory intensive), (c) computational cost, (d) massively parallel architecture, and (e) machine learning automation. In addition, VHR satellite datasets are of the order of terabytes and features extracted from these datasets are of the order of petabytes. In our present study, we have acquired the National Agriculture Imagery Program (NAIP) dataset for the Continental United States at a spatial resolution of 1-m. This data comes as image tiles (a total of quarter million image scenes with ~60 million pixels) and has a total size of ~65 terabytes for a single acquisition. Features extracted from the entire dataset would amount to ~8-10 petabytes. In our proposed approach, we have implemented a novel semi-automated machine learning algorithm rooted on the principles of "deep learning" to delineate the percentage of tree cover. Using the NASA Earth Exchange (NEX) initiative, we have developed an end-to-end architecture by integrating a segmentation module based on Statistical Region Merging, a classification algorithm using Deep Belief Network and a structured prediction algorithm using Conditional Random Fields to integrate the results from the segmentation and classification modules to create per-pixel class labels. The training process is scaled up using the power of GPUs and the prediction is scaled to quarter million NAIP tiles spanning the whole of Continental United States using the NEX HPC supercomputing cluster. An initial pilot over the state of California spanning a total of 11,095 NAIP tiles covering a total geographical area of 163,696 sq. miles has produced true positive rates of around 88 percent for fragmented forests and 74 percent for urban tree cover areas, with false positive rates lower than 2 percent for both landscapes.

Imagery↗

LACIE/ERIPS software system summary

The Earth resources interactive processing system (ERIPS) supports LACIE by classifying LANDSAT sensed data on the basis of the statistical similarity to those portions which were identified by analysts. The development and capabilities of the ERIPS software system are described with emphasis on (1) system requirements; (2) LACIE/ERIPS hardware; (3) system functions; (4) pattern recognition concept; and (5) LACIE/ERIPS data bases. Algorithms used in LACIE/ERIPS for statistics, divergence, feature selection, classification, registration, adaptive clustering, iterative clustering, clustering report functions, Sun angle correction, mean level adjustment, and bias correction are appended.

Johnson, C. L.↗

An Overview of the Total Lightning Jump Algorithm: Past, Present and Future Work

Rapid increases in total lightning prior to the onset of severe and hazardous weather have been observed for several decades. These rapid increases are known as lightning jumps and can precede the occurrence of severe weather by tens of minutes. Over the past decade, a significant effort has been made to quantify lightning jump behavior in relation to its utility as a predictor of severe and hazardous weather. Based on a study of 34 thunderstorms that occurred in the Tennessee Valley, early work conducted in our group at Huntsville determined that it was indeed possible to create a reasonable operational lightning jump algorithm (LJA) based on a statistical framework relying on the variance behavior of the lightning trending signal. We the expanded this framework and tested several variance-related LJA configurations on a much larger sample of 87 severe and non severe thunderstorms. This study determined that a configuration named the "2(sigma)" algorithm had the most promise in development of the operational LJA with a probability of detection (POD) of 87%, a false alarm rate (FAR) of 33%, a Heidke Skill Score (HSS) of 0.75. The 2(sigma) algorithm was then tested on an even larger sample of 711 thunderstorms of all types from four regions of the country where total lightning measurement capability existed. The result was very encouraging.Despite the larger number of storms and the inclusion of different regions of the country, the POD remained high (79%), the FAR was low (36%) and HSS was solid (0.71). Average lead time from jump to severe weather occurrence was 20.65 minutes, with a standard deviation of +/- 15 minutes. Also, trends in total lightning were compared to cloud to ground (CG) lightning trends, and it was determined that total lightning trends had a higher POD (79% vs 66%), lower FAR (36% vs 54 %) and a better HSS (0.71 vs 0.55). From the 711-storm case study it was determined that a majority of missed events were due to severe weather producing thunderstorms in low flashing environments. The latest efforts have been geared toward examining these low flashing storms in order to adjust the algorithm for such storms, thus enhancing the capability of the LJA. Future work will test the algorithm in real time using current satellite and radar based cell tracking methods, as well as, comparing total lightning jump occurrence to both satellite based and ground base observations of thunderstorms to create correlations between lightning jumps and the observed structures within thunderstorms. Finally this algorithm will need to be tested using Geostationary Lightning Mapper proxy data to transition the algorithm from VHF ground based lightning measurements to lower frequency space-based lightning measurements.

Schultz, Christopher J.↗

Algorithm for Identifying Erroneous Rain-Gauge Readings

An algorithm analyzes rain-gauge data to identify statistical outliers that could be deemed to be erroneous readings. Heretofore, analyses of this type have been performed in burdensome manual procedures that have involved subjective judgements. Sometimes, the analyses have included computational assistance for detecting values falling outside of arbitrary limits. The analyses have been performed without statistically valid knowledge of the spatial and temporal variations of precipitation within rain events. In contrast, the present algorithm makes it possible to automate such an analysis, makes the analysis objective, takes account of the spatial distribution of rain gauges in conjunction with the statistical nature of spatial variations in rainfall readings, and minimizes the use of arbitrary criteria. The algorithm implements an iterative process that involves nonparametric statistics.

Rickman, Doug↗

Astronomical data analysis software and systems I; Proceedings of the 1st Annual Conference, Tucson, AZ, Nov. 6-8, 1991

Consideration is given to a definition of a distribution format for X-ray data, the Einstein on-line system, the NASA/IPAC extragalactic database, COBE astronomical databases, Cosmic Background Explorer astronomical databases, the ADAM software environment, the Groningen Image Processing System, search for a common data model for astronomical data analysis systems, deconvolution for real and synthetic apertures, pitfalls in image reconstruction, a direct method for spectral and image restoration, and a discription of a Poisson imagery super resolution algorithm. Also discussed are multivariate statistics on HI and IRAS images, a faint object classification using neural networks, a matched filter for improving SNR of radio maps, automated aperture photometry of CCD images, interactive graphics interpreter, the ROSAT extreme ultra-violet sky survey, a quantitative study of optimal extraction, an automated analysis of spectra, applications of synthetic photometry, an algorithm for extra-solar planet system detection and data reduction facilities for the William Herschel telescope.

Worrall, Diana M.↗

Interpretable Machine Learning for Molecular Biosignatures: a Novel Single-Sample Feature Importance Method That Is Sensitive To Statistical Interactions

Isotope ratio mass spectrometry (IRMS) of volatiles (e.g., CO 2 ) promises to be a powerful tool for potential biosignature detection for future missions to ocean worlds (OW) such as Europa and Enceladus. Machine learning (ML) methods for IRMS data could enable science autonomy by onboard prediction of seawater chemistry and biosignature presence. However, ML models are likely to be complex and involve statistical interactions between features (variables), which can make predictions seem opaque and enigmatic. For ML predictions as significant as extraterrestrial biosignatures, we must place extraordinary confidence in models. It is therefore essential that these models make interpretable predictions (i.e., human-understandable) and include false-prediction diagnostics. We achieve high accuracy and interpretability in ML biosignature and seawater chemistry models for OW through a nearest-neighbors feature selection tool that detects statistical interactions between predictors, constructs interaction networks for visualization of selected features working together to make a prediction, and reports single-sample feature importance scores for false-detection diagnostics. Here we develop a novel single-sample nearest-neighbors projected distance regression(ssNPDR) feature selection method that improves upon existing single-sample algorithms through the inclusion of statistical interactions while providing false-prediction diagnostics for ML models.

geochemistry↗

Application of Artificial Neural Networks to the Development of Improved Multi-Sensor Retrievals of Near-Surface Air Temperature and Humidity Over Ocean

Improved estimates of near-surface air temperature and air humidity are critical to the development of more accurate turbulent surface heat fluxes over the ocean. Recent progress in retrieving these parameters has been made through the application of artificial neural networks (ANN) and the use of multi-sensor passive microwave observations. Details are provided on the development of an improved retrieval algorithm that applies the nonlinear statistical ANN methodology to a set of observations from the Advanced Microwave Scanning Radiometer (AMSR-E) and the Advanced Microwave Sounding Unit (AMSU-A) that are currently available from the NASA AQUA satellite platform. Statistical inversion techniques require an adequate training dataset to properly capture embedded physical relationships. The development of multiple training datasets containing only in-situ observations, only synthetic observations produced using the Community Radiative Transfer Model (CRTM), or a mixture of each is discussed. An intercomparison of results using each training dataset is provided to highlight the relative advantages and disadvantages of each methodology. Particular emphasis will be placed on the development of retrievals in cloudy versus clear-sky conditions. Near-surface air temperature and humidity retrievals using the multi-sensor ANN algorithms are compared to previous linear and non-linear retrieval schemes.

Roberts, J. Brent↗

Satellite Sampling and Retrieval Errors in Regional Monthly Rain Estimates from TMI AMSR-E, SSM/I, AMSU-B and the TRMM PR

Passive and active microwave rain sensors onboard earth-orbiting satellites estimate monthly rainfall from the instantaneous rain statistics collected during satellite overpasses. It is well known that climate-scale rain estimates from meteorological satellites incur sampling errors resulting from the process of discrete temporal sampling and statistical averaging. Sampling and retrieval errors ultimately become entangled in the estimation of the mean monthly rain rate. The sampling component of the error budget effectively introduces statistical noise into climate-scale rain estimates that obscure the error component associated with the instantaneous rain retrieval. Estimating the accuracy of the retrievals on monthly scales therefore necessitates a decomposition of the total error budget into sampling and retrieval error quantities. This paper presents results from a statistical evaluation of the sampling and retrieval errors for five different space-borne rain sensors on board nine orbiting satellites. Using an error decomposition methodology developed by one of the authors, sampling and retrieval errors were estimated at 0.25 resolution within 150 km of ground-based weather radars located at Kwajalein, Marshall Islands and Melbourne, Florida. Error and bias statistics were calculated according to the land, ocean and coast classifications of the surface terrain mask developed for the Goddard Profiling (GPROF) rain algorithm. Variations in the comparative error statistics are attributed to various factors related to differences in the swath geometry of each rain sensor, the orbital and instrument characteristics of the satellite and the regional climatology. The most significant result from this study found that each of the satellites incurred negative longterm oceanic retrieval biases of 10 to 30%.

Fisher, Brad↗

Mathematical algorithms for approximate reasoning

Most state of the art expert system environments contain a single and often ad hoc strategy for approximate reasoning. Some environments provide facilities to program the approximate reasoning algorithms. However, the next generation of expert systems should have an environment which contain a choice of several mathematical algorithms for approximate reasoning. To meet the need for validatable and verifiable coding, the expert system environment must no longer depend upon ad hoc reasoning techniques but instead must include mathematically rigorous techniques for approximate reasoning. Popular approximate reasoning techniques are reviewed, including: certainty factors, belief measures, Bayesian probabilities, fuzzy logic, and Shafer-Dempster techniques for reasoning. A group of mathematically rigorous algorithms for approximate reasoning are focused on that could form the basis of a next generation expert system environment. These algorithms are based upon the axioms of set theory and probability theory. To separate these algorithms for approximate reasoning various conditions of mutual exclusivity and independence are imposed upon the assertions. Approximate reasoning algorithms presented include: reasoning with statistically independent assertions, reasoning with mutually exclusive assertions, reasoning with assertions that exhibit minimum overlay within the state space, reasoning with assertions that exhibit maximum overlay within the state space (i.e. fuzzy logic), pessimistic reasoning (i.e. worst case analysis), optimistic reasoning (i.e. best case analysis), and reasoning with assertions with absolutely no knowledge of the possible dependency among the assertions. A robust environment for expert system construction should include the two modes of inference: modus ponens and modus tollens. Modus ponens inference is based upon reasoning towards the conclusion in a statement of logical implication, whereas modus tollens inference is based upon reasoning away from the conclusion. These algorithms allow one to reason accurately with uncertain data. The above environment can replicate state-f-the-art expert system environments which provides a continuity between the current expert systems which cannot be validated or verified and future expert systems which should be both validated and verified

Murphy, John H.↗

Learning classification trees

Algorithms for learning classification trees have had successes in artificial intelligence and statistics over many years. How a tree learning algorithm can be derived from Bayesian decision theory is outlined. This introduces Bayesian techniques for splitting, smoothing, and tree averaging. The splitting rule turns out to be similar to Quinlan's information gain splitting rule, while smoothing and averaging replace pruning. Comparative experiments with reimplementations of a minimum encoding approach, Quinlan's C4 and Breiman et al. Cart show the full Bayesian algorithm is consistently as good, or more accurate than these other approaches though at a computational price.

Buntine, Wray↗

Transiting Planet Search in the Kepler Pipeline

The Kepler Mission simultaneously measures the brightness of more than 160,000 stars every 29.4 minutes over a 3.5-year mission to search for transiting planets. Detecting transits is a signal-detection problem where the signal of interest is a periodic pulse train and the predominant noise source is non-white, non-stationary (1/f) type process of stellar variability. Many stars also exhibit coherent or quasi-coherent oscillations. The detection algorithm first identifies and removes strong oscillations followed by an adaptive, wavelet-based matched filter. We discuss how we obtain super-resolution detection statistics and the effectiveness of the algorithm for Kepler flight data.

Jenkins, Jon M.↗

Effects of preprocessing Landsat MSS data on derived features

Important to the use of multitemporal Landsat MSS data for earth resources monitoring, such as agricultural inventories, is the ability to minimize the effects of varying atmospheric and satellite viewing conditions, while extracting physically meaningful features from the data. In general, the approaches to the preprocessing problem have been derived from either physical or statistical models. This paper compares three proposed algorithms; XSTAR haze correction, Color Normalization, and Multiple Acquisition Mean Level Adjustment. These techniques represent physical, statistical, and hybrid physical-statistical models, respectively. The comparisons are made in the context of three feature extraction techniques; the Tasseled Cap, the Cate Color Cube. and Normalized Difference.

Parris, T. M.↗

ICAP: An Interactive Cluster Analysis Procedure for analyzing remotely sensed data

An Interactive Cluster Analysis Procedure (ICAP) was developed to derive classifier training statistics from remotely sensed data. The algorithm interfaces the rapid numerical processing capacity of a computer with the human ability to integrate qualitative information. Control of the clustering process alternates between the algorithm, which creates new centroids and forms clusters and the analyst, who evaluate and elect to modify the cluster structure. Clusters can be deleted or lumped pairwise, or new centroids can be added. A summary of the cluster statistics can be requested to facilitate cluster manipulation. The ICAP was implemented in APL (A Programming Language), an interactive computer language. The flexibility of the algorithm was evaluated using data from different LANDSAT scenes to simulate two situations: one in which the analyst is assumed to have no prior knowledge about the data and wishes to have the clusters formed more or less automatically; and the other in which the analyst is assumed to have some knowledge about the data structure and wishes to use that information to closely supervise the clustering process. For comparison, an existing clustering method was also applied to the two data sets.

Wharton, S. W.↗

A Fast-Time Study of Aircraft Reordering in Arrival Sequencing and Scheduling

In order to ensure that the safe capacity of the terminal area is not exceeded, Air Traffic Management ATM often places restrictions on arriving flights transitioning from en route airspace to terminal airspace. This restriction of arrival traffic is commonly referred to as arrival flow management, and includes techniques such as metering, vectoring, fix-load balancing, and the imposition of miles-in-trail separations. These restrictions are enacted without regard for the relative priority which airlines may be placing on individual flights based on factors such as crew criticality, passenger connectivity, critical turn times, gate availability, on-time performance, fuel status, or runway preference. The development of new arrival flow management techniques which take into consideration priorities expressed by air carriers will likely reduce the economic impact of ATM restrictions on the airlines and lead to increased airline economic efficiency by allowing airlines to have greater control over their individual arrival banks of aircraft. NASA and the Federal Aviation Administration (FAA) have designed and developed a suite of software decision support tools (DSTs) collectively known as the Center TRACON Automation System (CTAS). One of these tools, the Traffic Management Advisor (TMA) is currently being used at the Fort Worth Air Route Traffic Control Center to perform arrival flow management of traffic into the Dallas/Fort Worth airport (DFW). The TMA is a time-based strategic planning tool that assists Traffic Management Coordinators (TMCs) and En Route Air Traffic Controllers in efficiently balancing arrival demand with airport capacity. The primary algorithm in the TMA is a real-time scheduler which generates efficient landing sequences and landing times for arrivals within about 200 no a. from touchdown. This scheduler will sequence aircraft so that they arrive in a first- come - first-served (FCFS) order. While FCFS sequencing establishes a fair order based on estimated times of arrival, it does not take into account individual airline priorities among incoming flights. NASA is exploring the possibility of allowing airlines to express relative arrival priorities to air traffic management through the development of new CTAS scheduling algorithms which take into consideration airline arrival preferences. The accommodation of airline priorities in arrival sequencing and scheduling would under most circumstances result in a deviation from a "natural" or FCFS arrival order. As a First step toward developing airline influenced sequencing algorithms, an investigation was conducted to determine the feasibility of reordering arrival traffic from a strict FCFS sequence. A fast-time simulation has been developed which allows statistical evaluation of sequencing and scheduling algorithms for arrival traffic at the Dallas/Fort Worth Airport. In contrast to real-time simulation or field tests, which would require on the order of ninety minutes to examine a single traffic rush period, the fast-time simulation allows examination of multiple rush periods in a matter of seconds.

Carr, Greg↗

Scattering Properties and Brightness Temperatures Associated with Solid Precipitation

In the past few years, early solid precipitation detection and retrieval algorithms have been developed and shown to be applicable for snowing clouds and blizzards. NOAA has an operational snow versus rain classifier based on AMSU-B observations. Solid precipitation retrieval algorithms reported in the literature over the past two years include those that rely on neural nets, statistics, or physical relationships. All of the algorithms require the use of millimeter-wave radiometer observations. The millimeter-wave frequencies are especially sensitive to the scattering and emission properties of frozen particles due to the ice particle refractive index. Passive radiometric channels respond to both the integrated particle mass throughout the volume and field of view, and to the amount, location, and size distribution of the frozen (and liquid) particles with the sensitivity varying for different frequencies and hydrometeor types. This investigation probes the sensitivity of scattering and absorption coefficients, and hence computed brightness temperatures, resulting from variations in solid precipitation cloud profiles. The first study compares the single scattering, absorption, and asymmetry parameters associated with snow particles in clouds. Several methodologies are used to convert the physical characteristics (e.g., shape, size distributions, ice-air-water ratios) of ice particles to electromagnetic properties (e.g., absorption, scattering, and asymmetry factors). These methodologies include: conversion to solid ice particles, homogeneous dielectric mixing, or discrete dipole approximation. Changes in the conversion methodology can produce computed brightness temperature differences greater than 50 Kelvin.

Skofronick-Jackson, Gail M.↗

Evaluation of Classifier Complexity for Delay Tolerant Network Routing

The growing popularity of small cost effective satellites (SmallSats, CubeSats, etc.) creates the potential for a variety of new science applications involving multiple nodes functioning together or independently to achieve a task, such as swarms and constellations. As this technology develops and is deployed for missions in Low Earth Orbit and beyond, the use of delay tolerant networking (DTN) techniques may improve communication capabilities within the network. In this paper, a network hierarchy is developed from heterogeneous networks of SmallSats, surface vehicles, relay satellites and ground stations which form an integrated network. There is a tradeoff between complexity, flexibility, and scalability of user defined schedules versus autonomous routing as the number of nodes in the network increases. To address these issues, this work proposes a machine learning classifier based on DTN routing metrics. A framework is developed which will allow for the use of several categories of machine learning algorithms (decision tree, random forest and deep learning) to be applied to a dataset of historical network statistics, which allows for the evaluation of algorithm complexity versus performance to be explored. We develop the emulation of a hierarchical network, consisting of tens of nodes which form a cognitive network architecture. CORE (Common Open Research Emulator) is used to emulate the network using bundle protocol and DTN IP neighbor discovery.

Dudukovich, Rachel↗