Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “preprocessed data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13

Markov Chain Monte Carlo Bayesian Learning for Neural Networks

Conventional training methods for neural networks involve starting al a random location in the solution space of the network weights, navigating an error hyper surface to reach a minimum, and sometime stochastic based techniques (e.g., genetic algorithms) to avoid entrapment in a local minimum. It is further typically necessary to preprocess the data (e.g., normalization) to keep the training algorithm on course. Conversely, Bayesian based learning is an epistemological approach concerned with formally updating the plausibility of competing candidate hypotheses thereby obtaining a posterior distribution for the network weights conditioned on the available data and a prior distribution. In this paper, we developed a powerful methodology for estimating the full residual uncertainty in network weights and therefore network predictions by using a modified Jeffery's prior combined with a Metropolis Markov Chain Monte Carlo method.

Goodrich, Michael S.↗

Experiences with Text Mining Large Collections of Unstructured Systems Development Artifacts at JPL

Often repositories of systems engineering artifacts at NASA's Jet Propulsion Laboratory (JPL) are so large and poorly structured that they have outgrown our capability to effectively manually process their contents to extract useful information. Sophisticated text mining methods and tools seem a quick, low-effort approach to automating our limited manual efforts. Our experiences of exploring such methods mainly in three areas including historical risk analysis, defect identification based on requirements analysis, and over-time analysis of system anomalies at JPL, have shown that obtaining useful results requires substantial unanticipated efforts - from preprocessing the data to transforming the output for practical applications. We have not observed any quick 'wins' or realized benefit from short-term effort avoidance through automation in this area. Surprisingly we have realized a number of unexpected long-term benefits from the process of applying text mining to our repositories. This paper elaborates some of these benefits and our important lessons learned from the process of preparing and applying text mining to large unstructured system artifacts at JPL aiming to benefit future TM applications in similar problem domains and also in hope for being extended to broader areas of applications.

text mining↗

Mars Entry Atmospheric Data System Trajectory Reconstruction Algorithms and Flight Results

The Mars Entry Atmospheric Data System is a part of the Mars Science Laboratory, Entry, Descent, and Landing Instrumentation project. These sensors are a system of seven pressure transducers linked to ports on the entry vehicle forebody to record the pressure distribution during atmospheric entry. These measured surface pressures are used to generate estimates of atmospheric quantities based on modeled surface pressure distributions. Specifically, angle of attack, angle of sideslip, dynamic pressure, Mach number, and freestream atmospheric properties are reconstructed from the measured pressures. Such data allows for the aerodynamics to become decoupled from the assumed atmospheric properties, allowing for enhanced trajectory reconstruction and performance analysis as well as an aerodynamic reconstruction, which has not been possible in past Mars entry reconstructions. This paper provides details of the data processing algorithms that are utilized for this purpose. The data processing algorithms include two approaches that have commonly been utilized in past planetary entry trajectory reconstruction, and a new approach for this application that makes use of the pressure measurements. The paper describes assessments of data quality and preprocessing, and results of the flight data reduction from atmospheric entry, which occurred on August 5th, 2012.

Karlgaard, Christopher D.↗

Optimization of a Multi-Stage ATR System for Small Target Identification

An Automated Target Recognition system (ATR) was developed to locate and target small object in images and videos. The data is preprocessed and sent to a grayscale optical correlator (GOC) filter to identify possible regionsof- interest (ROIs). Next, features are extracted from ROIs based on Principal Component Analysis (PCA) and sent to neural network (NN) to be classified. The features are analyzed by the NN classifier indicating if each ROI contains the desired target or not. The ATR system was found useful in identifying small boats in open sea. However, due to "noisy background," such as weather conditions, background buildings, or water wakes, some false targets are mis-classified. Feedforward backpropagation and Radial Basis neural networks are optimized for generalization of representative features to reduce false-alarm rate. The neural networks are compared for their performance in classification accuracy, classifying time, and training time.

false alarm rate↗

Prediction of Aircraft Estimated Time of Arrival Using A Supervised Learning Approach

We present a novel data-driven approach for prediction of the estimated time of arrival (ETA) of aircraft in the terminal area via the implementation of a Random Forest regression model. The model uses data fused from a number of sources (flight track, weather, flight plan information, etc.) and provides predictions for the remaining flight time for aircraft landing at Dallas/Fort Worth (DFW) International Airport. The predictions are made when the aircraft is at a distance of 200-miles from the airport. The results show that the model is able to predict estimated time of arrival to within ± 5 min for 90% of the flights in the test data with the mean absolute error being lower at 145 seconds. This paper covers the entire pipeline of data collection, preprocessing, setup and training of the ML model, and the results obtained for DFW.

Machine learning↗

Mechanical separations of corn stover anatomical fractions in an integrated feedstock preprocessing system: An experimental and data-driven modeling study

High variabilities of material attributes in lignocellulosic biomass present risks for biofuel and biochemical productions and must be mitigated via preprocessing. Since almost no mechanical device is originally designed for processing biomass, how to operate existing apparatuses with efficient performance has not been investigated extensively. This work presents a study on an integrated screening and air classification to separate cobs and stalks from husks and leaves in corn stover. Prototype machine learning models were developed to assess the feasibility of predicting the process outcome based on the measurable parameters. The models trained upon limited experimental data rendered decent predictive accuracy of yield and purity. The experimental data and modeling results collectively suggest decreasing throughput leads to a higher purity. To the contrary, if throughput increases, a lower purity is likely. A possible trade-off between yield and purity of the separated streams indicates the need for optimal combinations of feedstock size, moisture, and throughput to achieve optimized separations. The results of this study also suggest the need to further improve model predictability by developing more accurate formulations for physics governing the integrated unit operations. To accomplish this, additional experimental data needs to be generated for model training.

09 - BIOMASS FUELS↗

Land Boundary Conditions for the Goddard Earth Observing System Model Version 5 (GEOS-5) Climate Modeling System: Recent Updates and Data File Descriptions

The Earths land surface boundary conditions in the Goddard Earth Observing System version 5 (GEOS-5) modeling system were updated using recent high spatial and temporal resolution global data products. The updates include: (i) construction of a global 10-arcsec land-ocean lakes-ice mask; (ii) incorporation of a 10-arcsec Globcover 2009 land cover dataset; (iii) implementation of Level 12 Pfafstetter hydrologic catchments; (iv) use of hybridized SRTM global topography data; (v) construction of the HWSDv1.21-STATSGO2 merged global 30 arc second soil mineral and carbon data in conjunction with a highly-refined soil classification system; (vi) production of diffuse visible and near-infrared 8-day MODIS albedo climatologies at 30-arcsec from the period 2001-2011; and (vii) production of the GEOLAND2 and MODIS merged 8-day LAI climatology at 30-arcsec for GEOS-5. The global data sets were preprocessed and used to construct global raster data files for the software (mkCatchParam) that computes parameters on catchment-tiles for various atmospheric grids. The updates also include a few bug fixes in mkCatchParam, as well as changes (improvements in algorithms, etc.) to mkCatchParam that allow it to produce tile-space parameters efficiently for high resolution AGCM grids. The update process also includes the construction of data files describing the vegetation type fractions, soil background albedo, nitrogen deposition and mean annual 2m air temperature to be used with the future Catchment CN model and the global stream channel network to be used with the future global runoff routing model. This report provides detailed descriptions of the data production process and data file format of each updated data set.

GEOS-5↗

Remote sensing program in earth resources

The basic features of the NASA remote sensing program are briefly outlined. Consideration is given to physical data acquisition and preprocessing, archiving for bulk retrieval, availability of Landsat data, and the role of foreign ground stations.

Billingsley, F. C.↗

Climatespark: an In-Memory Distributed Computing Framework for Big Climate Data Analytics

The unprecedented growth of climate data creates new opportunities for climate studies, and yet big climate data pose a grand challenge to climatologists to efficiently manage and analyze big data. The complexity of climate data content and analytical algorithms increases the difficulty of implementing algorithms on high performance computing systems. This paper proposes an in-memory, distributed computing framework, ClimateSpark, to facilitate complex big data analytics and time-consuming computational tasks. Chunking data structure improves parallel I/O efficiency, while a spatiotemporal index is built for the chunks to avoid unnecessary data reading and preprocessing. An integrated, multi-dimensional, array-based data model (ClimateRDD) and ETL operations are developed to address big climate data variety by integrating the processing components of the climate data lifecycle. ClimateSpark utilizes Spark SQL and Apache Zeppelin to develop a web portal to facilitate the interaction among climatologists, climate data, analytic operations and computing resources (e.g., using SQL query and Scala/Python notebook). Experimental results show that ClimateSpark conducts different spatiotemporal data queries/analytics with high efficiency and data locality. ClimateSpark is easily adaptable to other big multiple- dimensional, array-based datasets in various geoscience domains.

Hu, Fei↗

Advances in automatic extraction of information from multispectral scanner data

The state-of-the-art of automatic multispectral scanner data analysis and interpretation is reviewed. Sources of system variability which tend to obscure the spectral characteristics of the classes under consideration are discussed, and examples of the application of spatial and temporal discrimination bases are given. Automatic processing functions, techniques and methods, and equipment are described with particular attention to those that are applicable to large land surveys using satellite data. The development and characteristics of the Multivariate Interactive Digital Analysis System (MIDAS) for processing aircraft or satellite multispectral scanning data are discussed in detail. The MIDAS system combines the parallel digital implementation capabilities of a low-cost processor with a general purpose PDP-11/45 minicomputer to provide near-real-time data processing. The preprocessing functions are user-selectable. The input subsystem accepts data stored on high density digital tape, computer compatible tape, and analog tape.

Erickson, J. D.↗

Deep near-infrared survey of the Southern Sky (DENIS)

DENIS (Deep Near-Infrared Survey of the Southern Sky) will be the first complete census of astronomical sources in the near-infrared spectral range. The challenges of this novel survey are both scientific and technical. Phenomena radiating in the near-infrared range from brown dwarfs to galaxies in the early stages of cosmological evolution, the scientific exploitation of data relevant over such a wide range requires pooling expertise from several of the leading European astronomical centers. The technical challenges of a project which will provide an order of magnitude more sources than given by the IRAS space mission, and which will involve advanced data-handling and image-processing techniques, likewise require pooling of hardware and software resources, as well as of human expertise. The DENIS project team is composed of some 40 scientists, computer specialists, and engineers located in 5 European Community countries (France, Germany, Italy, The Netherlands, and Spain), with important contributions from specialists in Australia, Brazil, Chile, and Hungary. DENIS will survey the entire southern sky in 3 colors, namely in the I band at a wavelength of 0.8 micron, in the 1.25 micron J band, and in the 2.15 micron K' band. The sensitivity limits will be 18th magnitude in the I band, 16th in the J band, and 14.5th in the K' band. The angular resolution achieved will be 1 arcsecond in the I band, and 3.0 arcseconds in the J and K' bands. The European Southern Observatory 1 m telescope on La Silla will be dedicated to survey use during operations expected to last four years, commencing in late 1993. DENIS aims to provide the astronomical community with complete digitized infrared images of the full southern sky and a catalogue of extracted objects, both of the best quality and in readily accessible form. This will be achieved through dedicated software packages and specialized catalogues, and with assistance from the Leiden and Paris Data Analysis Centers. The data will be mailed on DAT tapes from La Silla to the two Data Analysis Centers for further processing. Two centers are necessary because of the shear quantity of data and because of the complementary roles the Centers will develop, each exploiting its own particular expertise. The Leiden Data Analysis Center (LDAC) will extract objects, establish their parameters, and archive them into a source catalogue. The LDAC will collaborate with the Groningen Space Research group that has gained experience in infrared image handling from the IRAS satellite. The Paris Data Analysis Center (PDAC) will be responsible for archiving and preprocessing the raw data to provide a homogeneous set of data suitable for further reduction in both the Leiden and Paris data analysis streams. The PDAC will also extract and archive images for the sources flagged by the LDAC as extended, and create a catalogue of galaxies. In exploiting the DENIS data we foresee the collaboration with other data analysis centers, such as the Observatoire de Lyon where the relevant DENIS catalogue of galaxies can be incorporated into their extragalactic database. The Point Sources and the Small Extended Sources catalogues could be incorporated in the Late Type Star database at Montpellier, and in the SIMBAD database as CDS. At Groningen the IRAS Point Source catalogue and/or image data can be merged with the DENIS catalogues. At Meudon algorithms and software will be developed with main goal assessing the limits reachable for the homogeneity and intrinsic consistency between the ensemble of the images in the data base (flat-fielding, relative positioning of the fields, bootstrapped flux calibration) but also for the data analysis.

Deul, E.↗

Estimation of transport airplane aerodynamics using multiple stepwise regression

This paper presents an application of multiple stepwise regression to the flight test data of a typical transport airplane. The flight test data was carefully preprocessed to eliminate aliasing, time skews and high frequency noise. The data consisted both of basic certification maneuvers, such as wind-up-turns and maneuvers suitable for parameter estimation, such as responses to elevator pulses and doublets. It is shown that the results of multiple stepwise regression techniques compare favorably with the results obtained from maximum likelihood estimation. Finally, it is concluded that multiple stepwise regression could be a fast economical way to estimate transport airplane aerodynamics.

Keskar, D. A.↗

Contour Error Map Algorithm

The contour error map (CEM) algorithm and the software that implements the algorithm are means of quantifying correlations between sets of time-varying data that are binarized and registered on spatial grids. The present version of the software is intended for use in evaluating numerical weather forecasts against observational sea-breeze data. In cases in which observational data come from off-grid stations, it is necessary to preprocess the observational data to transform them into gridded data. First, the wind direction is gridded and binarized so that D(i,j;n) is the input to CEM based on forecast data and d(i,j;n) is the input to CEM based on gridded observational data. Here, i and j are spatial indices representing 1.25-km intervals along the west-to-east and south-to-north directions, respectively; and n is a time index representing 5-minute intervals. A binary value of D or d = 0 corresponds to an offshore wind, whereas a value of D or d = 1 corresponds to an onshore wind. CEM includes two notable subalgorithms: One identifies and verifies sea-breeze boundaries; the other, which can be invoked optionally, performs an image-erosion function for the purpose of attempting to eliminate river-breeze contributions in the wind fields.

Merceret, Francis↗

Investigating tectonic and bathymetric features of the Indian Ocean using MAGSAT magnetic anomaly data

MAGSAT Investigator-B tapes were preprocessed by (1) removing all data points with obvious erroneous values and location errors; (2) removing smaller spikes (typically 15 nT or more), and deleting data tracks with fewer than 20 points; and (3) removing a linear trend from each track. The remaining data were recorded on tape for use by the equivalent source mapping (ESMAP) program which uses a least squares algorithm to fit the magnetization parameter of the grid of equivalent source dipoles in the crust to satellite data acquired at different times and locations. ESMAP was implemented on the TASC computing system and modified to read preprocessed MAGSAT tapes and interface with TASC plotting software. Some verification of the software was accomplished. Gridded 1-degree mean values of gravity anomaly and sea surface undulation computed from SEASAT radar altimeter were obtained and brought on line.

Lazarewicz, A. R.↗

Crop identification technology assessment for remote sensing (CITARS). Volume 10: Interpretation of results

The CITARS was an experiment designed to quantitatively evaluate crop identification performance for corn and soybeans in various environments using a well-defined set of automatic data processing (ADP) techniques. Each technique was applied to data acquired to recognize and estimate proportions of corn and soybeans. The CITARS documentation summarizes, interprets, and discusses the crop identification performances obtained using (1) different ADP procedures; (2) a linear versus a quadratic classifier; (3) prior probability information derived from historic data; (4) local versus nonlocal recognition training statistics and the associated use of preprocessing; (5) multitemporal data; (6) classification bias and mixed pixels in proportion estimation; and (7) data with differnt site characteristics, including crop, soil, atmospheric effects, and stages of crop maturity.

Bizzell, R. M.↗

Effect of shaping sensor data on pilot response

The pilot of a modern jet aircraft is subjected to varying workloads while being responsible for multiple, ongoing tasks. The ability to associate the pilot's responses with the task/situation, by modifying the way information is presented relative to the task, could provide a means of reducing workload. To examine the feasibility of this concept, a real time simulation study was undertaken to determine whether preprocessing of sensor data would affect pilot response. Results indicated that preprocessing could be an effective way to tailor the pilot's response to displayed data. The effects of three transformations or shaping functions were evaluated with respect to the pilot's ability to predict and detect out-of-tolerance conditions while monitoring an electronic engine display. Two nonlinear transformations, on being the inverse of the other, were compared to a linear transformation. Results indicate that a nonlinear transformation that increases the rate-or-change of output relative to input tends to advance the prediction response and improve the detection response, while a nonlinear transformation that decreases the rate-of-change of output relative to input tends to lengthen the prediction response and make detection more difficult.

Bailey, Roger M.↗

Self-Organizing-Map Program for Analyzing Multivariate Data

SOM_VIS is a computer program for analysis and display of multidimensional sets of Earth-image data typified by the data acquired by the Multi-angle Imaging Spectro-Radiometer [MISR (a spaceborne instrument)]. In SOM_VIS, an enhanced self-organizing-map (SOM) algorithm is first used to project a multidimensional set of data into a nonuniform three-dimensional lattice structure. The lattice structure is mapped to a color space to obtain a color map for an image. The Voronoi cell-refinement algorithm is used to map the SOM lattice structure to various levels of color resolution. The final result is a false-color image in which similar colors represent similar characteristics across all its data dimensions. SOM_VIS provides a control panel for selection of a subset of suitably preprocessed MISR radiance data, and a control panel for choosing parameters to run SOM training. SOM_VIS also includes a component for displaying the false-color SOM image, a color map for the trained SOM lattice, a plot showing an original input vector in 36 dimensions of a selected pixel from the SOM image, the SOM vector that represents the input vector, and the Euclidean distance between the two vectors.

Li, P. Peggy↗

Radiometric correction and equalization of satellite digital data

Satellite digital data from Landsat and NOAA satellites is often marred by striping or streaking errors due to variations in the response of the radiometric sensors. In this paper, we discuss the equalization of the digital data as a preprocessing step, prior to image enhancement or automatic classification. The methods described make use of statistics of the data itself to generate nonlinear or linear memory-less equalization algorithms. These algorithms, by contrast to multidimensional filtering, do not result in a loss of spatial resolution. Examples of applications to Landsat and NOAA-3 thermal infrared data are given and illustrated.

Algazi, V. R.↗