Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “large data sets”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Integrating Large Scale Data Sets to Develop Predictive Hypotheses of Low-Dose Radiation-Induced Health Effects

Over one hundred years of radiation biology research has revealed much about the DNA damages induced by the deposition of energy from exposure to ionizing radiation and the subsequent cellular responses. However, there are still significant gaps in our understanding of how these might lead to detrimental health effects, particularly at low doses (100 mGy (milligray)). Recent advances in high throughput omics technologies enable interrogation of induced radiation effects at the genomic, proteomic and metabolomic levels. These include changes in gene expression, protein modifications, e.g., phosphorylation, acetylation, and methylation, and metabolic changes. We will discuss the integration of data obtained from multiple omics platforms to understand radiation dose, and dose rate effects in a complex human tissue model as a function of time. We will use as an example our results on the low dose responses in a 3D human skin model.

ionizing radiation↗

Encounter-Based Simulation Architecture for Detect-And-Avoid Modeling

This paper presents an encounter-based simulation architecture developed at NASA to facilitate flexible and efficient Detect and Avoid modeling in parametric or tradespace studies on large data sets. The basic premise of this tool is that large-scale input data can be reduced to a set of `canonical encounters' and that using the reduced data in simulations does not lead to loss of fidelity. A canonical encounter is specified as ownship and intruder flight portions potentially resulting in a loss of well clear along with a set of properties that characterize the encounter. The advantages of using canonical encounters include faster simulations, reduced memory footprint, ability to select encounters based on user-specified criteria, shared encounters across multiple teams, peer-reviewed encounters, and a better understanding of the input data set, to name a few.

MOPS↗

Encounter-Based Simulation Architecture for Detect and Avoid Modeling

This paper presents an encounter-based simulation architecture developed at NASA to facilitate flexible and efficient Detect and Avoid modeling in parametric or tradespace studies on large data sets. The basic premise of this tool is that large-scale input data can be reduced to a set of `canonical encounters' and that using the reduced data in simulations does not lead to loss of fidelity. A canonical encounter is specified as ownship and intruder flight portions potentially resulting in a loss of well clear along with a set of properties that characterize the encounter. The advantages of using canonical encounters include faster simulations, reduced memory footprint, ability to select encounters based on user-specified criteria, shared encounters across multiple teams, peer-reviewed encounters, and a better understanding of the input data set, to name a few.

MOPS↗

The Application of Principal Component Analysis Using Fixed Eigenvectors to the Infrared Thermographic Inspection of the Space Shuttle Thermal Protection System

The Nondestructive Evaluation Sciences Branch at NASA s Langley Research Center has been actively involved in the development of thermographic inspection techniques for more than 15 years. Since the Space Shuttle Columbia accident, NASA has focused on the improvement of advanced NDE techniques for the Reinforced Carbon-Carbon (RCC) panels that comprise the orbiter s wing leading edge. Various nondestructive inspection techniques have been used in the examination of the RCC, but thermography has emerged as an effective inspection alternative to more traditional methods. Thermography is a non-contact inspection method as compared to ultrasonic techniques which typically require the use of a coupling medium between the transducer and material. Like radiographic techniques, thermography can be used to inspect large areas, but has the advantage of minimal safety concerns and the ability for single-sided measurements. Principal Component Analysis (PCA) has been shown effective for reducing thermographic NDE data. A typical implementation of PCA is when the eigenvectors are generated from the data set being analyzed. Although it is a powerful tool for enhancing the visibility of defects in thermal data, PCA can be computationally intense and time consuming when applied to the large data sets typical in thermography. Additionally, PCA can experience problems when very large defects are present (defects that dominate the field-of-view), since the calculation of the eigenvectors is now governed by the presence of the defect, not the good material. To increase the processing speed and to minimize the negative effects of large defects, an alternative method of PCA is being pursued when a fixed set of eigenvectors is used to process the thermal data from the RCC materials. These eigen vectors can be generated either from an analytic model of the thermal response of the material under examination, or from a large cross section of experimental data. This paper will provide the details of the analytic model; an overview of the PCA process; as well as a quantitative signal-to-noise comparison of the results of performing both embodiments of PCA on thermographic data from various RCC specimens. Details of a system that has been developed to allow insitu inspection of a majority of shuttle RCC components will be presented along with the acceptance test results for this system. Additionally, the results of applying this technology to the Space Shuttle Discovery after its return from flight will be presented.

Cramer, K. Elliott↗

Analysis and representation of complex structures in separated flows

We discuss our recent work on extraction and visualization of topological information in separated fluid flow data sets. As with scene analysis, an abstract representation of a large data set can greatly facilitate the understanding of complex, high-level structures. When studying flow topology, such a representation can be produced by locating and characterizing critical points in the velocity field and generating the associated stream surfaces. In 3D flows, the surface topology serves as the starting point. The 2D tangential velocity field near the surface of the body is examined for critical points. The tangential velocity field is integrated out along the principal directions of certain classes of critical points to produce curves depicting the topology of the flow near the body. The points and curves are linked to form a skeleton representing the 2D vector field topology. This skeleton provides a basis for analyzing the 3D structures associated with the flow separation. The points along the separation curves in the skeleton are used to start tangent curve integrations. Integration origins are successively refined to produce stream surfaces. The map of the global topology is completed by generating those stream surfaces associated with 3D critical points.

Helman, James↗

Hybrid Storage Solution

With the rise of artificial intelligence and machine learning, data sets used to train models have become increasingly large. The availability, accessibility and integrity of large data sets has become important to the research conducted at Los Alamos National Laboratory. Ceph is a storage solution suitable for use with critical data because of its distributed nature and ability to keep multiple copies of a file in different locations. The amount of data means that bandwidth, latency, and cost are important factors and the reason most storage solutions are on-premises. However, there are distinct advantages to hosting services in the cloud, namely scalability and ease-of-use. In this paper, we explore the possibility of provisioning a hybrid Ceph cluster that leverages the benefits of both cloud architectures and on-premise performance.

97 MATHEMATICS AND COMPUTING↗

Ground State Energy Functional with Hartree–Fock Efficiency and Chemical Accuracy

We introduce the deep post Hartree–Fock (DeePHF) method, a machine learning-based scheme for constructing accurate and transferable models for the ground-state energy of electronic structure problems. DeePHF predicts the energy difference between results of highly accurate models such as the coupled cluster method and low accuracy models such as the Hartree–Fock (HF) method, using the ground-state electronic orbitals as the input. It preserves all the symmetries of the original high accuracy model. The added computational cost is less than that of the reference HF or DFT and scales linearly with respect to system size. We examine the performance of DeePHF on organic molecular systems using publicly available data sets and obtain the state-of-art performance, particularly on large data sets.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Arithmetic Data Cube as a Data Intensive Benchmark

Data movement across computational grids and across memory hierarchy of individual grid machines is known to be a limiting factor for application involving large data sets. In this paper we introduce the Data Cube Operator on an Arithmetic Data Set which we call Arithmetic Data Cube (ADC). We propose to use the ADC to benchmark grid capabilities to handle large distributed data sets. The ADC stresses all levels of grid memory by producing 2d views of an Arithmetic Data Set of d-tuples described by a small number of parameters. We control data intensity of the ADC by controlling the sizes of the views through choice of the tuple parameters.

Frumkin, Michael A.↗

The Application of Infrared Thermographic Inspection Techniques to the Space Shuttle Thermal Protection System

The Nondestructive Evaluation Sciences Branch at NASA s Langley Research Center has been actively involved in the development of thermographic inspection techniques for more than 15 years. Since the Space Shuttle Columbia accident, NASA has focused on the improvement of advanced NDE techniques for the Reinforced Carbon-Carbon (RCC) panels that comprise the orbiter s wing leading edge. Various nondestructive inspection techniques have been used in the examination of the RCC, but thermography has emerged as an effective inspection alternative to more traditional methods. Thermography is a non-contact inspection method as compared to ultrasonic techniques which typically require the use of a coupling medium between the transducer and material. Like radiographic techniques, thermography can be used to inspect large areas, but has the advantage of minimal safety concerns and the ability for single-sided measurements. Principal Component Analysis (PCA) has been shown effective for reducing thermographic NDE data. A typical implementation of PCA is when the eigenvectors are generated from the data set being analyzed. Although it is a powerful tool for enhancing the visibility of defects in thermal data, PCA can be computationally intense and time consuming when applied to the large data sets typical in thermography. Additionally, PCA can experience problems when very large defects are present (defects that dominate the field-of-view), since the calculation of the eigenvectors is now governed by the presence of the defect, not the "good" material. To increase the processing speed and to minimize the negative effects of large defects, an alternative method of PCA is being pursued where a fixed set of eigenvectors, generated from an analytic model of the thermal response of the material under examination, is used to process the thermal data from the RCC materials. Details of a one-dimensional analytic model and a two-dimensional finite-element model will be presented. An overview of the PCA process as well as a quantitative signal-to-noise comparison of the results of performing both embodiments of PCA on thermographic data from various RCC specimens will be shown. Finally, a number of different applications of this technology to various RCC components will be presented.

Cramer, K. E.↗

Field Model: An Object-Oriented Data Model for Fields

We present an extensible, object-oriented data model designed for field data entitled Field Model (FM). FM objects can represent a wide variety of fields, including fields of arbitrary dimension and node type. FM can also handle time-series data. FM achieves generality through carefully selected topological primitives and through an implementation that leverages the potential of templated C++. FM supports fields where the nodes values are paired with any cell type. Thus FM can represent data where the field nodes are paired with the vertices ("vertex-centered" data), fields where the nodes are paired with the D-dimensional cells in R(sup D) (often called "cell-centered" data), as well as fields where nodes are paired with edges or other cell types. FM is designed to effectively handle very large data sets; in particular FM employs a demand-driven evaluation strategy that works especially well with large field data. Finally, the interfaces developed for FM have the potential to effectively abstract field data based on adaptive meshes. We present initial results with a triangular adaptive grid in R(sup 2) and discuss how the same design abstractions would work equally well with other adaptive-grid variations, including meshes in R(sup 3).

Moran, Patrick J.↗

Social Networking Adapted for Distributed Scientific Collaboration

Share is a social networking site with novel, specially designed feature sets to enable simultaneous remote collaboration and sharing of large data sets among scientists. The site will include not only the standard features found on popular consumer-oriented social networking sites such as Facebook and Myspace, but also a number of powerful tools to extend its functionality to a science collaboration site. A Virtual Observatory is a promising technology for making data accessible from various missions and instruments through a Web browser. Sci-Share augments services provided by Virtual Observatories by enabling distributed collaboration and sharing of downloaded and/or processed data among scientists. This will, in turn, increase science returns from NASA missions. Sci-Share also enables better utilization of NASA s high-performance computing resources by providing an easy and central mechanism to access and share large files on users space or those saved on mass storage. The most common means of remote scientific collaboration today remains the trio of e-mail for electronic communication, FTP for file sharing, and personalized Web sites for dissemination of papers and research results. Each of these tools has well-known limitations. Sci-Share transforms the social networking paradigm into a scientific collaboration environment by offering powerful tools for cooperative discourse and digital content sharing. Sci-Share differentiates itself by serving as an online repository for users digital content with the following unique features: a) Sharing of any file type, any size, from anywhere; b) Creation of projects and groups for controlled sharing; c) Module for sharing files on HPC (High Performance Computing) sites; d) Universal accessibility of staged files as embedded links on other sites (e.g. Facebook) and tools (e.g. e-mail); e) Drag-and-drop transfer of large files, replacing awkward e-mail attachments (and file size limitations); f) Enterprise-level data and messaging encryption; and g) Easy-to-use intuitive workflow.

Karimabadi, Homa↗

Principal Component Analysis of Thermographic Data

Principal Component Analysis (PCA) has been shown effective for reducing thermographic NDE data. While a reliable technique for enhancing the visibility of defects in thermal data, PCA can be computationally intense and time consuming when applied to the large data sets typical in thermography. Additionally, PCA can experience problems when very large defects are present (defects that dominate the field-of-view), since the calculation of the eigenvectors is now governed by the presence of the defect, not the "good" material. To increase the processing speed and to minimize the negative effects of large defects, an alternative method of PCA is being pursued where a fixed set of eigenvectors, generated from an analytic model of the thermal response of the material under examination, is used to process the thermal data from composite materials. This method has been applied for characterization of flaws.

Winfree, William P.↗

Data transfer for STAR grid jobs

The Solenoidal Tracker at RHIC (STAR) is a multipurpose experiment at the Relativistic Heavy Ion Collider (RHIC) with the primary goal to study the formation and properties of the quark-gluon plasma. STAR is an international collaboration of member institutions and laboratories from around the world. Yearly data-taking period produces PBytes of raw data collected by the experiment. STAR primarily uses its dedicated facility at BNL to process this data, but has routinely leveraged distributed systems, both high throughput (HTC) and high performance (HPC) computing clusters, to significantly augment the processing capacity available to the experiment. The ability to automate the efficient transfer of large data sets on reliable, scalable, and secure infrastructure is critical for any large-scale distributed processing campaign. For more than a decade, STAR computing has relied upon GridFTP with its x509-based authentication to build such data transfer systems and integrate them into its larger production workflow. The end of support by the community for both GridFTP and the x509 standard requires STAR to investigate other approaches to meet its distributed processing needs. In this study we investigate two multi-purpose data distribution systems, Globus.org and XRootD, as alternatives to GridFTP. We compare both their performance and the ease by which each service is integrated into the type of secure and automated data transfer systems STAR has previously built using GridFTP. The presented approach and study may be applicable to other distributed data processing use cases beyond STAR.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

SLIP (Surrogate Launching and Integration Platform)

SLIP (Surrogate Launching and Integration Platform) is a software ecosystem for downloading and running ML/AI benchmarks. SLIP automatically downloads and sets up the code and data running a benchmark. The large data sets and model code are cached locally with a specific ID within a designated cache directory.

Brown, Cade↗

From Images to Dark Matter: End-to-end Inference of Substructure from Hundreds of Strong Gravitational Lenses

Abstract Constraining the distribution of small-scale structure in our universe allows us to probe alternatives to the cold dark matter paradigm. Strong gravitational lensing offers a unique window into small dark matter halos (<10 10 M ⊙ ) because these halos impart a gravitational lensing signal even if they do not host luminous galaxies. We create large data sets of strong lensing images with realistic low-mass halos, Hubble Space Telescope (HST) observational effects, and galaxy light from HST’s COSMOS field. Using a simulation-based inference pipeline, we train a neural posterior estimator of the subhalo mass function (SHMF) and place constraints on populations of lenses generated using a separate set of galaxy sources. We find that by combining our network with a hierarchical inference framework, we can both reliably infer the SHMF across a variety of configurations and scale efficiently to populations with hundreds of lenses. By conducting precise inference on large and complex simulated data sets, our method lays a foundation for extracting dark matter constraints from the next generation of wide-field optical imaging surveys.

79 ASTRONOMY AND ASTROPHYSICS↗

A distributed parallel storage architecture and its potential application within EOSDIS

We describe the architecture, implementation, use of a scalable, high performance, distributed-parallel data storage system developed in the ARPA funded MAGIC gigabit testbed. A collection of wide area distributed disk servers operate in parallel to provide logical block level access to large data sets. Operated primarily as a network-based cache, the architecture supports cooperation among independently owned resources to provide fast, large-scale, on-demand storage to support data handling, simulation, and computation.

Johnston, William E.↗