Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “sparse data representation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Representation-Independent Iteration of Sparse Data Arrays

An approach is defined that describes a method of iterating over massively large arrays containing sparse data using an approach that is implementation independent of how the contents of the sparse arrays are laid out in memory. What is unique and important here is the decoupling of the iteration over the sparse set of array elements from how they are internally represented in memory. This enables this approach to be backward compatible with existing schemes for representing sparse arrays as well as new approaches. What is novel here is a new approach for efficiently iterating over sparse arrays that is independent of the underlying memory layout representation of the array. A functional interface is defined for implementing sparse arrays in any modern programming language with a particular focus for the Chapel programming language. Examples are provided that show the translation of a loop that computes a matrix vector product into this representation for both the distributed and not-distributed cases. This work is directly applicable to NASA and its High Productivity Computing Systems (HPCS) program that JPL and our current program are engaged in. The goal of this program is to create powerful, scalable, and economically viable high-powered computer systems suitable for use in national security and industry by 2010. This is important to NASA for its computationally intensive requirements for analyzing and understanding the volumes of science data from our returned missions.

James, Mark↗

NASA Tech Briefs, January 2007

Topics covered include: Flexible Skins Containing Integrated Sensors and Circuitry; Artificial Hair Cells for Sensing Flows; Video Guidance Sensor and Time-of-Flight Rangefinder; Optical Beam-Shear Sensors; Multiple-Agent Air/Ground Autonomous Exploration Systems; A 640 512-Pixel Portable Long-Wavelength Infrared Camera; An Array of Optical Receivers for Deep-Space Communications; Microstrip Antenna Arrays on Multilayer LCP Substrates; Applications for Subvocal Speech; Multiloop Rapid-Rise/Rapid Fall High-Voltage Power Supply; The PICWidget; Fusing Symbolic and Numerical Diagnostic Computations; Probabilistic Reasoning for Robustness in Automated Planning; Short-Term Forecasting of Radiation Belt and Ring Current; JMS Proxy and C/C++ Client SDK; XML Flight/Ground Data Dictionary Management; Cross-Compiler for Modeling Space-Flight Systems; Composite Elastic Skins for Shape-Changing Structures; Glass/Ceramic Composites for Sealing Solid Oxide Fuel Cells; Aligning Optical Fibers by Means of Actuated MEMS Wedges; Manufacturing Large Membrane Mirrors at Low Cost; Double-Vacuum-Bag Process for Making Resin- Matrix Composites; Surface Bacterial-Spore Assay Using Tb3+/DPA Luminescence; Simplified Microarray Technique for Identifying mRNA in Rare Samples; High-Resolution, Wide-Field-of-View Scanning Telescope; Multispectral Imager With Improved Filter Wheel and Optics; Integral Radiator and Storage Tank; Compensation for Phase Anisotropy of a Metal Reflector; Optical Characterization of Molecular Contaminant Films; Integrated Hardware and Software for No-Loss Computing; Decision-Tree Formulation With Order-1 Lateral Execution; GIS Methodology for Planning Planetary-Rover Operations; Optimal Calibration of the Spitzer Space Telescope; Automated Detection of Events of Scientific Interest; Representation-Independent Iteration of Sparse Data Arrays; Mission Operations of the Mars Exploration Rovers; and More About Software for No-Loss Computing.

Source record↗

A path-oriented matrix-based knowledge representation system

Experience has shown that designing a good representation is often the key to turning hard problems into simple ones. Most AI (Artificial Intelligence) search/representation techniques are oriented toward an infinite domain of objects and arbitrary relations among them. In reality much of what needs to be represented in AI can be expressed using a finite domain and unary or binary predicates. Well-known vector- and matrix-based representations can efficiently represent finite domains and unary/binary predicates, and allow effective extraction of path information by generalized transitive closure/path matrix computations. In order to avoid space limitations a set of abstract sparse matrix data types was developed along with a set of operations on them. This representation forms the basis of an intelligent information system for representing and manipulating relational data.

Feyock, Stefan↗

Antarctic Stratospheric Ozone from the Assimilation of Occultation Data

Ozone data from the solar occultation Polar Ozone and Aerosol Measurement (POAM) III instrument are included in the ozone assimilation system at NASA's Global Modeling and Assimilation Office, which uses Solar Backscatter UItraViolet/2 (SBUV/2) instrument data. Even though POAM data are available at only one latitude in the southern hemisphere on each day, their assimilation leads to more realistic ozone distribution throughout the Antarctic region, especially inside the polar vortex. Impacts of POAM data were evaluated by comparisons of assimilated ozone profiles with independent ozone sondes. Major improvements in ozone representation are seen in the Antarctic lower stratosphere during austral Winter and spring in 1998. Limitations of assimilation of sparse occultation data are illustrated by an example.

Stajner, Ivanka↗

Enhancements of Bayesian Blocks; Application to Large Light Curve Databases

Bayesian Blocks are optimal piecewise linear representations (step function fits) of light-curves. The simple algorithm implementing this idea, using dynamic programming, has been extended to include more data modes and fitness metrics, multivariate analysis, and data on the circle (Studies in Astronomical Time Series Analysis. VI. Bayesian Block Representations, Scargle, Norris, Jackson and Chiang 2013, ApJ, 764, 167), as well as new results on background subtraction and refinement of the procedure for precise timing of transient events in sparse data. Example demonstrations will include exploratory analysis of the Kepler light curve archive in a search for "star-tickling" signals from extraterrestrial civilizations. (The Cepheid Galactic Internet, Learned, Kudritzki, Pakvasa1, and Zee, 2008, arXiv: 0809.0339; Walkowicz et al., in progress).

Kepler light curve archive↗

Multiresolution representation and numerical algorithms: A brief review

In this paper we review recent developments in techniques to represent data in terms of its local scale components. These techniques enable us to obtain data compression by eliminating scale-coefficients which are sufficiently small. This capability for data compression can be used to reduce the cost of many numerical solution algorithms by either applying it to the numerical solution operator in order to get an approximate sparse representation, or by applying it to the numerical solution itself in order to reduce the number of quantities that need to be computed.

Harten, Amiram↗

Predictive Modeling for Differential Diagnosis and Mortality Risk Assessment

The prevalence of electronic health record (EHR) systems has brought prodigious biomedical informatics opportunity. Automated machine learning methods can effectively utilize such data and have become common tools for healthcare predictive modeling. Researches in medical informatics have explored the potential of deep learning and classical models in emergent care scenarios. In particular, predicting differential diagnoses for admissions have proven useful in decreasing unnecessary lab tests and improving inpatient triage decision-making. Moreover, identification of high-risk patients for in-hospital mortality is vitally important to maximize allocation of medical resources.The Medical Information Mart for Intensive Care (MIMIC-III) database, containing de-identified critical care inpatient was used in our study. This data set captures hospital patient laboratory measurements, pharmacologic prescriptions, diagnostic data and procedure event recordings. When considering adult patients and discounting admissions with ICU length of stay less than 24 hours, there were 37,787 unique admissions and 30,414 total patients. We examined the top 25 most prevalent ICD-9 group-level disease specificities in MIMIC-III using a multi-label classification model. In-hospital mortality was modeled as binary classification with 4,155 (13%) adult patients that expired, of which 3,138 (75.5%) were in the ICU setting. The metrics AUC, F1 score, sensitivity and specificity values calculated for each disease label measured prediction performance.The usage of ICD-9 group codes reduced feature dimension from 14,567 to 942 and greatly improved distribution of patient diagnostic categories. Disease temporal patterns were captured by considering the most frequently sampled 6 vital signs and 13 laboratory values. Missing data were imputed at each time-stamp. Time-series raw hourly average values were converted into 5 summary features (mean, standard deviation, number of observations, min & max values). Patient demographic variables such as age, gender, marital status and ethnicity were also factored into the modeling. Choi et al showed that contextual embedding of medical data, diagnostic and procedural codes alone can predict future diagnoses with sensitivity as high as 0.79. We utilized an embedding technique called word2vec which allowed sparse representations of medical history to be transformed into dense word vectors. The mappings captured contextual information by treating each admission as a sentence and learning the most likely neighboring words in a sliding window fashion. Binary and multi-label classification was achieved via collapse models, which do not consider temporal information, as well as recurrent neural networks with regularization, Softmax output layer activation together with categorical cross-entropy as the loss function.

US Army collaboration↗

Arctic sea ice albedo from AVHRR

The seasonal cycle of surface albedo of sea ice in the Arctic is estimated from measurements made with the Advanced Very High Resolution Radiometer (AVHRR) on the polar-orbiting satellites NOAA-10 and NOAA-11. The albedos of 145 200-km-square cells are analyzed. The cells are from March through September 1989 and include only those for which the sun is more than 10 deg above the horizon. Cloud masking is performed manually. Corrections are applied for instrument calibration, nonisotropic reflection, atmospheric interference, narrowband to broadband conversion, and normalization to a common solar zenith angle. The estimated albedos are relative, with the instrument gain set to give an albedo of 0.80 for ice floes in March and April. The mean values for the cloud-free portions of individual cells range from 0.18 to 0.91. Monthly averages of cells in the central Arctic range from 0.76 in April to 0.47 in August. The monthly averages of the within-cell standard deviations in the central Arctic are 0.04 in April and 0.06 in September. The surface albedo and surface temperature are correlated most strongly in March (R = -0.77) with little correlation in the summer. The monthly average lead fraction is determined from the mean potential open water, a scaled representation of the temperature or albedo between 0.0 (for ice) and 1.0 (for water); in the central Arctic it rises from an average 0.025 in the spring to 0.06 in September. Sparse data on aerosols, ozone, and water vapor in the atmospheric column contribute uncertainties to instantaneous, area-average albedos of 0.13, 0.04, and 0.08. Uncertainties in monthly average albedos are not this large. Contemporaneous estimation of these variables could reduce the uncertainty in the estimated albedo considerably. The poor calibration of AVHRR channels 1 and 2 is another large impediment to making accurate albedo estimates.

Lindsay, R. W.↗

Numerical Reanalyses as a Gateway to Arctic Synthesis

Reanalyses are regularly gridded, retrospective depictions of the physical earth system, which are produced through the correction of a short-term forecast to available observations. In the Arctic, reanalyses are particularly well suited to marshal the sparse observing network to provide a plausible, multivariate representation of conditions. Atmospheric reanalyses such as MERRA-2 (NASA Modern-Era Retrospective analysis for Research and Applications, version 2) and ocean reanalyses such as SODA3 (Univ. Maryland Simple Ocean Data Assimilation version 3) are widely used in Arctic research for diagnostic studies of circulation, model evaluation, and as boundary conditions for a variety of process models. Here, we provide examples that illustrate the utility of reanalyses for providing information on the spatial and temporal scales of recent, rapid changes in the Arctic. Recent trends in Arctic surface temperatures, surface melt over Greenland and Arctic glaciers, and evolving freshwater conditions in the Arctic Ocean are examples where reanalyses can provide information that cannot easily be obtained via other means. These examples provide information on the scale, magnitude, and the uncertainty of recent Arctic change and provide a context for future scenarios. We further quantify uncertainties in key reanalyses variables and approaches for addressing these issues.

Arctic↗

Enhanced Reporting of Mars Exploration Rover Telemetry

Mars Exploration Rover Enhanced Telemetry Extraction and Reporting System (METERS) is software that generates a human-readable representation of the state of the mobility and arm-related systems of the Mars Exploration Rover (MER) vehicles on each Martian solar day (sol). Data are received from the MER spacecraft in multiple streams having various formats including text messages, sparsely-sampled engineering quantities, images, and individual motor-command histories.

Maimone, Mark W.↗

A Data Type for Efficient Representation of Other Data Types

A self-organizing, monomorphic data type denoted a sequence has been conceived to address certain concerns that arise in programming parallel computers. A sequence in the present sense can be regarded abstractly as a vector, set, bag, queue, or other construct. Heretofore, in programming a parallel computer, it has been necessary for the programmer to state explicitly, at the outset, what parts of the program and the underlying data structures must be represented in parallel form. Not only is this requirement not optimal from the perspective of implementation; it entails an additional requirement that the programmer have intimate understanding of the underlying parallel structure. The present sequence data type overcomes both the implementation and parallel structure obstacles. In so doing, the sequence data type provides unified means by which the programmer can represent a data structure for natural and automatic decomposition to a parallel computing architecture. Sequences exhibit the behavioral and structural characteristics of vectors, but the underlying representations are automatically synthesized from combinations of programmers advice and execution use metrics. Sequences can vary bidirectionally between sparseness and density, making them excellent choices for many kinds of algorithms. The novelty and benefit of this behavior lies in the fact that it can relieve programmers of the details of implementations. The creation of a sequence enables decoupling of a conceptual representation from an implementation. The underlying representation of a sequence is a hybrid of representations composed of vectors, linked lists, connected blocks, and hash tables. The internal structure of a sequence can automatically change from time to time on the basis of how it is being used. Those portions of a sequence where elements have not been added or removed can be as efficient as vectors. As elements are inserted and removed in a given portion, then different methods are utilized to provide both an access and memory strategy that is optimized for that portion and the use to which it is put.

James, Mark↗

The state of the atmosphere as inferred from the FGGE satellite observing systems during SOP-1

The statistical properties, and coverage, of satellite temperature sounding data are described. Tropical regions are observed every two days, extratropics from one to four times a day. Oceans are covered two to three times a day. Asynoptic coverage is comparable to the U.S. rawinsonde network twice daily coverage. Lack of ground truth for data sparse areas makes accuracy difficult to assess. The rms differences of layer mean temperatures obtained from collocating rawinsonde observations with satellite temperature profiles in space and time differ from rms differences of layer mean satellite temperature soundings. The FGGE satellite systems can infer the three dimensional motion field and improve the representation of the large scale state of the atmosphere.

Halem, M.↗

Efficient Implementation of an Optimal Interpolator for Large Spatial Data Sets

Scattered data interpolation is a problem of interest in numerous areas such as electronic imaging, smooth surface modeling, and computational geometry. Our motivation arises from applications in geology and mining, which often involve large scattered data sets and a demand for high accuracy. The method of choice is ordinary kriging. This is because it is a best unbiased estimator. Unfortunately, this interpolant is computationally very expensive to compute exactly. For n scattered data points, computing the value of a single interpolant involves solving a dense linear system of size roughly n x n. This is infeasible for large n. In practice, kriging is solved approximately by local approaches that are based on considering only a relatively small'number of points that lie close to the query point. There are many problems with this local approach, however. The first is that determining the proper neighborhood size is tricky, and is usually solved by ad hoc methods such as selecting a fixed number of nearest neighbors or all the points lying within a fixed radius. Such fixed neighborhood sizes may not work well for all query points, depending on local density of the point distribution. Local methods also suffer from the problem that the resulting interpolant is not continuous. Meyer showed that while kriging produces smooth continues surfaces, it has zero order continuity along its borders. Thus, at interface boundaries where the neighborhood changes, the interpolant behaves discontinuously. Therefore, it is important to consider and solve the global system for each interpolant. However, solving such large dense systems for each query point is impractical. Recently a more principled approach to approximating kriging has been proposed based on a technique called covariance tapering. The problems arise from the fact that the covariance functions that are used in kriging have global support. Our implementations combine, utilize, and enhance a number of different approaches that have been introduced in literature for solving large linear systems for interpolation of scattered data points. For very large systems, exact methods such as Gaussian elimination are impractical since they require 0(n(exp 3)) time and 0(n(exp 2)) storage. As Billings et al. suggested, we use an iterative approach. In particular, we use the SYMMLQ method, for solving the large but sparse ordinary kriging systems that result from tapering. The main technical issue that need to be overcome in our algorithmic solution is that the points' covariance matrix for kriging should be symmetric positive definite. The goal of tapering is to obtain a sparse approximate representation of the covariance matrix while maintaining its positive definiteness. Furrer et al. used tapering to obtain a sparse linear system of the form Ax = b, where A is the tapered symmetric positive definite covariance matrix. Thus, Cholesky factorization could be used to solve their linear systems. They implemented an efficient sparse Cholesky decomposition method. They also showed if these tapers are used for a limited class of covariance models, the solution of the system converges to the solution of the original system. Matrix A in the ordinary kriging system, while symmetric, is not positive definite. Thus, their approach is not applicable to the ordinary kriging system. Therefore, we use tapering only to obtain a sparse linear system. Then, we use SYMMLQ to solve the ordinary kriging system. We show that solving large kriging systems becomes practical via tapering and iterative methods, and results in lower estimation errors compared to traditional local approaches, and significant memory savings compared to the original global system. We also developed a more efficient variant of the sparse SYMMLQ method for large ordinary kriging systems. This approach adaptively finds the correct local neighborhood for each query point in the interpolation process.

Memarsadeghi, Nargess↗

Assimilation of DAWN Doppler Wind Lidar Data During the 2017 Convective Processes Experiment (CPEX): Impact on Precipitation and Flow Structure

An improved representation of 3-D air motion and precipitation structure through forecast models and assimilation of observations is vital for improvements in weather forecasting capabilities. However, there are few independent data to properly validate a model forecast of precipitation structure when the underlying dynamics are evolving on short convective timescales. Using data from the JPL Ku/Ka-band Airborne Precipitation Radar (APR-2) and the 2 μmDoppler Aerosol Wind (DAWN) lidar collected during the2017 Convective Processes Experiment (CPEX), the NASA Unified Weather Research and Forecasting (WRF) Ensemble Data Assimilation System (EDAS) modeling system was used to quantify the impact of high-resolution sparsely sampled DAWN measurements on the analyzed variables and on the forecast when the DAWN winds were assimilated. Over-all, the assimilation of the DAWN wind profiles had a discernible impact on the wind field as well as the evolution and timing of the 3-D precipitation structure. Analysis of individual variables revealed that the assimilation of the DAWN winds resulted in important and coherent modifications of the environment. It led to an increase in the near-surface convergence, temperature, and water vapor, creating more favorable conditions for the development of convection exactly where it was observed (but not present in the control run). Comparison to APR-2 and observations by the Global Precipitation Measurement (GPM) satellite shows a much-improved forecast after the assimilation of the DAWN winds – development of precipitation where there was none, more organized precipitation where there was some, and a much more intense and organized cold pool, similar to the analysis of the dropsonde data. The onset of the vertical evolution of the precipitation showed similar radar-derived cloud-top heights, but delayed in time. While this investigation was limited to a single CPEX flight date, the investigation design is appropriate for further investigation of the impact of airborne Doppler wind lidar observations upon short-term convective precipitation forecasts

DAWN↗

Recognition of simple visual images using a sparse distributed memory: Some implementations and experiments

Previously, a method was described of representing a class of simple visual images so that they could be used with a Sparse Distributed Memory (SDM). Herein, two possible implementations are described of a SDM, for which these images, suitably encoded, will serve both as addresses to the memory and as data to be stored in the memory. A key feature of both implementations is that a pattern that is represented as an unordered set with a variable number of members can be used as an address to the memory. In the 1st model, an image is encoded as a 9072 bit string to be used as a read or write address; the bit string may also be used as data to be stored in the memory. Another representation, in which an image is encoded as a 256 bit string, may be used with either model as data to be stored in the memory, but not as an address. In the 2nd model, an image is not represented as a vector of fixed length to be used as an address. Instead, a rule is given for determining which memory locations are to be activated in response to an encoded image. This activation rule treats the pieces of an image as an unordered set. With this model, the memory can be simulated, based on a method of computing the approximate result of a read operation.

Jaeckel, Louis A.↗

Data Mining and Optimization Tools for Developing Engine Parameters Tools

This project was awarded for understanding the problem and developing a plan for Data Mining tools for use in designing and implementing an Engine Condition Monitoring System. From the total budget of $5,000, Tricia and I studied the problem domain for developing ail Engine Condition Monitoring system using the sparse and non-standardized datasets to be available through a consortium at NASA Lewis Research Center. We visited NASA three times to discuss additional issues related to dataset which was not made available to us. We discussed and developed a general framework of data mining and optimization tools to extract useful information from sparse and non-standard datasets. These discussions lead to the training of Tricia Erhardt to develop Genetic Algorithm based search programs which were written in C++ and used to demonstrate the capability of GA algorithm in searching an optimal solution in noisy datasets. From the study and discussion with NASA LERC personnel, we then prepared a proposal, which is being submitted to NASA for future work for the development of data mining algorithms for engine conditional monitoring. The proposed set of algorithm uses wavelet processing for creating multi-resolution pyramid of the data for GA based multi-resolution optimal search. Wavelet processing is proposed to create a coarse resolution representation of data providing two advantages in GA based search: 1. We will have less data to begin with to make search sub-spaces. 2. It will have robustness against the noise because at every level of wavelet based decomposition, we will be decomposing the signal into low pass and high pass filters.

Dhawan, Atam P.↗

Global Assimilation of Loon Stratospheric Balloon Observations

Project Loon has an overall goal of providing worldwide internet coverage using a network of long-durationsuper-pressure balloons. Since 2013, Loon has launched over 1600 balloons from multiple tropical and middlelatitude locations. These GPS tracked balloon trajectories provide lower stratospheric wind information overthe oceans and remote land areas where traditional radiosonde soundings are sparse, thus providing uniquecoverage of lower stratospheric winds. To fully investigate these Loon winds we: 1) compare the Loon windsto winds produced by a global data assimilation system (DAS: NASA GEOS) and 2) assimilate the Loon windsinto the same comprehensive DAS. Results show that in middle latitudes the Loon winds and DAS winds agreewell, and the Loon wind assimilation has only a minor impact on the forecasts. However, in the Tropics, thereis often a substantial difference between the assimilated winds and the observed Loon winds, of 8 m/s or morein magnitude. In these cases, assimilating the Loon winds significantly improves the meteorological analysesand subsequently the forecasts of the Loon winds. By highlighting cases where the Loon and DAS winds differ,these results can lead to improved understanding of stratospheric winds, especially in the tropics, as well asaiding analyses of the representation of dynamical forcing mechanisms in the GEOS model.

Coy, Lawrence↗

Global Assimilation of EOS-Aura Data as a Means of Mapping Ozone Distribution in the Lower Stratosphere and Troposphere

Ozone in the lower stratosphere and the troposphere plays an important role in forcing the climate. However, the global ozone distribution in this region is not well known because of the sparse distribution of in-situ data and the poor sensitivity of satellite based observations to the lowermost of the atmosphere. The Ozone Monitoring Instrument (OMI) and Microwave Limb Sounder (MLS) instruments on EOS-Aura provide information on the total ozone column and the stratospheric ozone profile. This data has been assimilated into NASA s Global Earth Observing System, Version 5 (GEOS-5) data assimilation system (DAS). We will discuss the results of assimilating three years of OMI and MLS data into GEOS-5. This data was assimilated alongside meteorological observations from both conventional sources and satellite instruments. Previous studies have shown that combining observations from these instruments through the Trajectory Tropospheric Ozone Residual methodology (TTOR) or using data assimilation can yield useful, yet low biased, estimates of the tropospheric ozone budget. We show that the assimilated ozone fields in this updated version of GEOS-5 exhibit an excellent agreement with ozone sonde and High Resolution Dynamics Limb Sounder (HIRDLS) data in the lower stratosphere in terms of spatial and temporal variability as well as integrated ozone abundances. Good representation of small-scale vertical features follows from combining the MLS data with the assimilated meteorological fields. We then demonstrate how this information can be used to calculate the Stratosphere - Troposphere Exchange of ozone and its contribution to the tropospheric ozone column in GEOS-5. Evaluations of tropospheric ozone distributions from the assimilation will be made by comparisons with sonde and other in-situ observations.

Wargan, Krzysztof↗