Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “massive data sets”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Digital Technologies at NASA for Science and Engineering

While scientific and engineering advancements used to rely primarily on theoretical studies and physical experiments, today digital technology enabled by petaflops-scale supercomputers is an equal, if not a greater, contributor to such achievements. In addition, computational modeling and simulation serves as a predictive tool that is not otherwise available. As a result, the use of high performance computing is integral to NASA's work in all mission areas such as space exploration, aeronautics, and scientific discovery. But traditional supercomputing alone is not sufficient for all of the space agency's needs. The success of many NASA missions depends on solving complex computing challenges, some of which are NP-hard (decision theory) if using classical solution methods. Quantum computing promises an unprecedented ability to solve such intractable problems by harnessing quantum mechanical effects such as tunneling, superposition, and entanglement. Another disruptive digital technology is neuromorphic computing that uses brain-inspired lessons to generate new architectures that are much more energy efficient, and capable of massive parallel processing and learning in-situ. Finally, with large amounts of observational and computational data sets, the opportunities of big data and data analytics can be leveraged to enable deep learning and knowledge discovery - it's all a massive digital transformation. This talk will be an overview how NASA utilizes digital technologies for its science and engineering efforts.

Biswas, Rupak↗

Graphics Processing Unit Assisted Thermographic Compositing

Objective Develop a software application utilizing high performance computing techniques, including general purpose graphics processing units (GPGPUs), for the analysis and visualization of large thermographic data sets. Over the past several years, an increasing effort among scientists and engineers to utilize graphics processing units (GPUs) in a more general purpose fashion is allowing for previously unobtainable levels of computation by individual workstations. As data sets grow, the methods to work them grow at an equal, and often greater, pace. Certain common computations can take advantage of the massively parallel and optimized hardware constructs of the GPU which yield significant increases in performance. These common computations have high degrees of data parallelism, that is, they are the same computation applied to a large set of data where the result does not depend on other data elements. Image processing is one area were GPUs are being used to greatly increase the performance of certain analysis and visualization techniques.

Ragasa, Scott↗

Graphics Processing Unit Assisted Thermographic Compositing

Objective: To develop a software application utilizing general purpose graphics processing units (GPUs) for the analysis of large sets of thermographic data. Background: Over the past few years, an increasing effort among scientists and engineers to utilize the GPU in a more general purpose fashion is allowing for supercomputer level results at individual workstations. As data sets grow, the methods to work them grow at an equal, and often greater, pace. Certain common computations can take advantage of the massively parallel and optimized hardware constructs of the GPU to allow for throughput that was previously reserved for compute clusters. These common computations have high degrees of data parallelism, that is, they are the same computation applied to a large set of data where the result does not depend on other data elements. Signal (image) processing is one area were GPUs are being used to greatly increase the performance of certain algorithms and analysis techniques.

Ragasa, Scott↗

GeoDash: Assisting Visual Image Interpretation in Collect Earth Online by Leveraging Big Data on Google Earth Engine

Collect Earth Online (CEO) is a free and open online implementation of the FAO Collect Earth system for collaboratively collecting environmental data through the visual interpretation of Earth observation imagery. The primary collection mechanism in CEO is human interpretation of land surface characteristics in imagery served via Web Map Services (WMS). However, interpreters may not have enough contextual information to classify samples by only viewing the imagery served via WMS, be they high resolution or otherwise. To assist in the interpretation and collection processes in CEO, SERVIR, a joint NASA-USAID initiative that brings Earth observations to improve environmental decision making in developing countries, developed the GeoDash system, an embedded and critical component of CEO. GeoDash leverages Google Earth Engine (GEE) by allowing users to set up custom browser-based widgets that pull from GEE's massive public data catalog. These widgets can be quick looks of other satellite imagery, time series graphs of environmental variables, and statistics panels of the same. Users can customize widgets with any of GEE's image collections, such as the historical Landsat collection with data available since the 1970s, select date ranges, image stretch parameters, graph characteristics, and create custom layouts, all on-the-fly to support plot interpretation in CEO. This presentation focuses on the implementation and potential applications, including the back-end links to GEE and the user interface with custom widget building. GeoDash takes large data volumes and condenses them into meaningful, relevant information for interpreters. While designed initially with national and global forest resource assessments in mind, the system will complement disaster assessments, agriculture management, project monitoring and evaluation, and more.

SERVI↗

Distributed Parallel Processing and Dynamic Load Balancing Techniques for Multidisciplinary High Speed Aircraft Design

Multidisciplinary design optimization (MDO) for large-scale engineering problems poses many challenges (e.g., the design of an efficient concurrent paradigm for global optimization based on disciplinary analyses, expensive computations over vast data sets, etc.) This work focuses on the application of distributed schemes for massively parallel architectures to MDO problems, as a tool for reducing computation time and solving larger problems. The specific problem considered here is configuration optimization of a high speed civil transport (HSCT), and the efficient parallelization of the embedded paradigm for reasonable design space identification. Two distributed dynamic load balancing techniques (random polling and global round robin with message combining) and two necessary termination detection schemes (global task count and token passing) were implemented and evaluated in terms of effectiveness and scalability to large problem sizes and a thousand processors. The effect of certain parameters on execution time was also inspected. Empirical results demonstrated stable performance and effectiveness for all schemes, and the parametric study showed that the selected algorithmic parameters have a negligible effect on performance.

Krasteva, Denitza T.↗

Software Development to Assist in the Processing and Analysis of Data Obtained Using Fiber Bragg Grating Interrogation Systems

A fiber Bragg grating is a portion of a core of a fiber optic strand that has been treated to affect the way light travels through the strand. Light within a certain narrow range of wavelengths will be reflected along the fiber by the grating, while light outside that range will pass through the grating mostly undisturbed. Since the range of wavelengths that can penetrate the grating depends on the grating itself as well as temperature and mechanical strain, fiber Bragg gratings can be used as temperature and strain sensors. This capability, along with the light-weight nature of the fiber optic strands in which the gratings reside, make fiber optic sensors an ideal candidate for flight testing and monitoring in which temperature and wing strain are factors. The purpose of this project is to research the availability of software capable of processing massive amounts of data in both real-time and post-flight settings, and to produce software segments that can be integrated to assist in the task as well.

Hicks, Rebecca↗

The Contribution of Ionizing Stars to the Far-Infrared and Radio Emission in the Galaxy

This is the first report of a new contract. However, this project represents ongoing work, so there are completed tasks as well as newly started tasks to report. The project involves the completion of the IRAS Galaxy Atlas (IGA), a large image database produced using data from the Infrared Astronomical Satellite (IRAS). In this phase, the project switches from the production and characterization of the IGA to its use in astronomical research studies of massive star formation. The research utilizes the IGA as well as two other large data sets being produced by research partners.

Cao, Yu↗

An ALMA Study of the Massive Molecular Clump N159W-North in the Large Magellanic Cloud: A Possible Gas Flow Penetrating One of the Most Massive Protocluster Systems in the Local Group

Massive dense clumps in the Large Magellanic Cloud can be an important laboratory to explore the formation ofpopulous clusters. We report multiscale ALMA observations of the N159W-North clump, which is the most CO-intense region in the galaxy. High-resolution CO isotope and 1.3 mm continuum observations with an angularresolution of∼025(∼0.07 pc)revealed more thanfive protostellar sources with CO outflows within the mainridge clump. One of the thermal continuum sources, MMS-2, shows an especially massive/dense nature whosetotal H2mass and peak column density are∼104Meand∼1024cm−2, respectively, and harbors massive(∼100Me)starless core candidates identified as its internal substructures. The main ridge containing this sourcecan be categorized as one of the most massive protocluster systems in the Local Group. The CO high-resolutionobservations found several distinctfilamentary clouds extending southward from the star-forming spots. The CO(1–0)data set with a largerfield of view reveals a conical,∼30 pc long complex extending toward the northerndirection. These features indicate that a large-scale gas compression event may have produced the massive star-forming complex. Based on the striking similarity between the N159W-North complex and the other twopreviously reported high-mass star-forming clouds in the nearby regions, we propose a“teardrops inflow model”that explains the synchronized, extreme star formation across>50 pc, including one of the most massiveprotocluster clumps in the Local Group.

Kazuki Tokuda↗

Dark matter in the galactic disk

Observational data on the distributions of tracer stars, F dwarfs and K giants were used as input to obtain self-consistent solutions for the Poisson and Vlasov equations to set bounds on the amount of missing matter in the solar neighborhood. The numerical computations were carried out using Galaxy models which feature multiple disk components and an unseen massive halo. The star data included the mass components and velocity dispersions. Consideration of various possible distributions of the unseen matter leads to the conjecture that half of the disk material in the solar neighborhood has yet to be observed. Techniques for determining if brown dwarfs are a significant component of the missing mass are discussed, as are improved models which would use limited numbers of tracer stars to set further constraints on the amount and distribution of the missing mass.

Bahcall, John N.↗

Towards the Interoperability of Web, Database, and Mass Storage Technologies for Petabyte Archives

At the San Diego Supercomputer Center, a massive data analysis system (MDAS) is being developed to support data-intensive applications that manipulate terabyte sized data sets. The objective is to support scientific application access to data whether it is located at a Web site, stored as an object in a database, and/or storage in an archival storage system. We are developing a suite of demonstration programs which illustrate how Web, database (DBMS), and archival storage (mass storage) technologies can be integrated. An application presentation interface is being designed that integrates data access to all of these sources. We have developed a data movement interface between the Illustra object-relational database and the NSL UniTree archival storage system running in a production mode at the San Diego Supercomputer Center. With this interface, an Illustra client can transparently access data on UniTree under the control of the Illustr DBMS server. The current implementation is based on the creation of a new DBMS storage manager class, and a set of library functions that allow the manipulation and migration of data stored as Illustra 'large objects'. We have extended this interface to allow a Web client application to control data movement between its local disk, the Web server, the DBMS Illustra server, and the UniTree mass storage environment. This paper describes some of the current approaches successfully integrating these technologies. This framework is measured against a representative sample of environmental data extracted from the San Diego Ba Environmental Data Repository. Practical lessons are drawn and critical research areas are highlighted.

Moore, Reagan↗

Fast and Flexible Multivariate Time Series Subsequence Search

Multivariate Time-Series (MTS) are ubiquitous, and are generated in areas as disparate as sensor recordings in aerospace systems, music and video streams, medical monitoring, and financial systems. Domain experts are often interested in searching for interesting multivariate patterns from these MTS databases which often contain several gigabytes of data. Surprisingly, research on MTS search is very limited. Most of the existing work only supports queries with the same length of data, or queries on a fixed set of variables. In this paper, we propose an efficient and flexible subsequence search framework for massive MTS databases, that, for the first time, enables querying on any subset of variables with arbitrary time delays between them. We propose two algorithms to solve this problem (1) a List Based Search (LBS) algorithm which uses sorted lists for indexing, and (2) a R*-tree Based Search (RBS) which uses Minimum Bounding Rectangles (MBR) to organize the subsequences. Both algorithms guarantee that all matching patterns within the specified thresholds will be returned (no false dismissals). The very few false alarms can be removed by a post-processing step. Since our framework is also capable of Univariate Time-Series (UTS) subsequence search, we first demonstrate the efficiency of our algorithms on several UTS datasets previously used in the literature. We follow this up with experiments using two large MTS databases from the aviation domain, each containing several millions of observations. Both these tests show that our algorithms have very high prune rates (>99%) thus needing actual disk access for only less than 1% of the observations. To the best of our knowledge, MTS subsequence search has never been attempted on datasets of the size we have used in this paper.

Bhaduri, Kanishka↗

Fast Multivariate Search on Large Aviation Datasets

Multivariate Time-Series (MTS) are ubiquitous, and are generated in areas as disparate as sensor recordings in aerospace systems, music and video streams, medical monitoring, and financial systems. Domain experts are often interested in searching for interesting multivariate patterns from these MTS databases which can contain up to several gigabytes of data. Surprisingly, research on MTS search is very limited. Most existing work only supports queries with the same length of data, or queries on a fixed set of variables. In this paper, we propose an efficient and flexible subsequence search framework for massive MTS databases, that, for the first time, enables querying on any subset of variables with arbitrary time delays between them. We propose two provably correct algorithms to solve this problem (1) an R-tree Based Search (RBS) which uses Minimum Bounding Rectangles (MBR) to organize the subsequences, and (2) a List Based Search (LBS) algorithm which uses sorted lists for indexing. We demonstrate the performance of these algorithms using two large MTS databases from the aviation domain, each containing several millions of observations Both these tests show that our algorithms have very high prune rates (>95%) thus needing actual

Bhaduri, Kanishka↗

An Innovative Infrastructure with a Universal Geo-Spatiotemporal Data Representation Supporting Cost-Effective Integration of Diverse Earth Science Data

The SpatioTemporal Adaptive Resolution Encoding (STARE) is a unifying scheme encoding geospatial and temporal information for organizing data on scalable computing/storage resources, minimizing expensive data transfers. STARE provides a compact representation that turns set-logic functions into integer operations, e.g. conditional sub-setting, taking into account representative spatiotemporal resolutions of the data in the datasets. STARE geo-spatiotemporally aligns data placements of diverse data on massive parallel resources to maximize performance. Automating important scientific functions (e.g. regridding) and computational functions (e.g. data placement) allows scientists to focus on domain-specific questions instead of expending their efforts and expertise on data processing. With STARE-enabled automation, SciDB (Scientific Database) plus STARE provides a database interface, reducing costly data preparation, increasing the volume and variety of interoperable data, and easing result sharing. Using SciDB plus STARE as part of an integrated analysis infrastructure dramatically eases combining diametrically different datasets.

Rilee, Michael Lee↗

Managing Complexity in Multidisciplinary Visualization

As high performance computing technology progresses, computational simulations are becoming more advanced in their capabilities. In the computational aerosciences domain, single discipline steady-state simulations computed on a single grid are far from the state-of-the-art. In their place are complex, time-dependent multidisciplinary simulations that attempt to model a given geometry more realistically. The product of these multidisciplinary simulations is a massive amount of data stored in different formats, grid topologies, units of measure, etc., as a result of the differences in the simulated physical domains. In addition to the challenges posed by setting up and performing the simulation, additional challenges exist in analyzing computational results. Visualization plays an important role in the advancement of multidisciplinary simulations. To date, visualization has been used to aid in the interpretation of large amounts of simulation data. Because the human visual system is effective in digesting a large amount of information presented graphically, visualization has helped simulation scientists to understand complex simulation results. As these simulations become even more complex, integrating several different physical domains, visualization will be critical to digest the massive amount of information. Another important role for visualization is to provide a common communication medium from which the domain scientists can use to develop, debug, and analyze their work. Multidisciplinary analyses are the next step in simulation technology, not only in computational aerosciences, but in many other areas such as global climate modeling. Visualization researchers must understand and work towards the challenges posed by multidisciplinary simulation scenarios. This paper addresses some of these challenges, describing technologies that must be investigated to create a useful visualization analysis tool for domain scientists.

Miceli, Kristina D.↗

Programmable remapper with single flow architecture

The invention relates to image processing systems and methods and in particular to a machine which accepts a real time video image in the form of a matrix of picture elements (pixels) and remaps such image according to a selectable one of a plurality of mapping functions to create an output matrix of pixels. Such mapping functions, or transformations, may be any one of a number of different transformations depending on the objective of the user of the system. The system remaps input images from one coordinate system to another using a set of look-up tables for the data necessary for the transform. The transforms, which are operator selectable, are precomputed and loaded into massive look-up tables. Input pixels, via the look-up tables of any particular transform selected, are mapped into output pixels with the radiance information of the input pixels being appropriately weighted. An earlier embodiment of the system included two parallel processors: a collective processor which mapped multiple input pixels into a single output pixel and an interpolative processor. The interpolative processor performed an interpolation among pixels in the input image where a given input pixel may affect the value of many output pixels. Several advantages are provided over previous embodiments in that the two distinct processors are replaced by a single processor capable of performing both types of operations (collective and interpolative) with no more complexity. Previously, there has existed no image processor or 'remapper' that can operate with sufficient speed and flexibility to permit investigating different transformation patterns in real time.

Fisher, Timothy E.↗

Implementing Access to Data Distributed on Many Processors

A reference architecture is defined for an object-oriented implementation of domains, arrays, and distributions written in the programming language Chapel. This technology primarily addresses domains that contain arrays that have regular index sets with the low-level implementation details being beyond the scope of this discussion. What is defined is a complete set of object-oriented operators that allows one to perform data distributions for domain arrays involving regular arithmetic index sets. What is unique is that these operators allow for the arbitrary regions of the arrays to be fragmented and distributed across multiple processors with a single point of access giving the programmer the illusion that all the elements are collocated on a single processor. Today's massively parallel High Productivity Computing Systems (HPCS) are characterized by a modular structure, with a large number of processing and memory units connected by a high-speed network. Locality of access as well as load balancing are primary concerns in these systems that are typically used for high-performance scientific computation. Data distributions address these issues by providing a range of methods for spreading large data sets across the components of a system. Over the past two decades, many languages, systems, tools, and libraries have been developed for the support of distributions. Since the performance of data parallel applications is directly influenced by the distribution strategy, users often resort to low-level programming models that allow fine-tuning of the distribution aspects affecting performance, but, at the same time, are tedious and error-prone. This technology presents a reusable design of a data-distribution framework for data parallel high-performance applications. Distributions are a means to express locality in systems composed of large numbers of processor and memory components connected by a network. Since distributions have a great effect on the performance of applications, it is important that the distribution strategy is flexible, so its behavior can change depending on the needs of the application. At the same time, high productivity concerns require that the user be shielded from error-prone, tedious details such as communication and synchronization.

James, Mark↗

Hyperresolution Global Land Surface Modeling: Meeting a Grand Challenge for Monitoring Earth's Terrestrial Water

Monitoring Earth's terrestrial water conditions is critically important to many hydrological applications such as global food production; assessing water resources sustainability; and flood, drought, and climate change prediction. These needs have motivated the development of pilot monitoring and prediction systems for terrestrial hydrologic and vegetative states, but to date only at the rather coarse spatial resolutions (approx.10-100 km) over continental to global domains. Adequately addressing critical water cycle science questions and applications requires systems that are implemented globally at much higher resolutions, on the order of 1 km, resolutions referred to as hyperresolution in the context of global land surface models. This opinion paper sets forth the needs and benefits for a system that would monitor and predict the Earth's terrestrial water, energy, and biogeochemical cycles. We discuss six major challenges in developing a system: improved representation of surface-subsurface interactions due to fine-scale topography and vegetation; improved representation of land-atmospheric interactions and resulting spatial information on soil moisture and evapotranspiration; inclusion of water quality as part of the biogeochemical cycle; representation of human impacts from water management; utilizing massively parallel computer systems and recent computational advances in solving hyperresolution models that will have up to 10(exp 9) unknowns; and developing the required in situ and remote sensing global data sets. We deem the development of a global hyperresolution model for monitoring the terrestrial water, energy, and biogeochemical cycles a grand challenge to the community, and we call upon the international hydrologic community and the hydrological science support infrastructure to endorse the effort.

Wood, Eric F.↗

A Coordinated X-Ray and Optical Campaign of the Nearest Massive Eclipsing Binary, Delta Orionis Aa. II. X-Ray Variability

We present time-resolved and phase-resolved variability studies of an extensive X-ray high-resolution spectral data set of the delta Ori Aa binary system. The four observations, obtained with Chandra ACIS (Advanced CCD Imaging Spectrometer) HETGS (High Energy Transmission Grating), have a total exposure time approximately equal to 479 kiloseconds and provide nearly complete binary phase coverage. Variability of the total X-ray flux in the range of 5-25 angstroms is confirmed, with a maximum amplitude of about plus or minus15 percent within a single approximately equal to125 kiloseconds observation. Periods of 4.76 and 2.04 days are found in the total X-ray flux, as well as an apparent overall increase in the flux level throughout the nine-day observational campaign. Using 40 kiloseconds contiguous spectra derived from the original observations, we investigate the variability of emission line parameters and ratios. Several emission lines are shown to be variable, including S (sub XV), Si (sub XIII), and Ne (sub IX). For the first time, variations of the X-ray emission line widths as a function of the binary phase are found in a binary system, with the smallest widths at phi = 0.0 when the secondary delta Ori Aa2 is at the inferior conjunction. Using 3D hydrodynamic modeling of the interacting winds, we relate the emission line width variability to the presence of a wind cavity created by a wind-wind collision, which is effectively void of embedded wind shocks and is carved out of the X-ray-producing primary wind, thus producing phase-locked X-ray variability.

X-ray flux↗