Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “data sciences”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 289 records · Page 16

The Field Guide to NASA’s Life Sciences Data Repositories

For over 30 years, NASA has invested in life sciences research both in space and on the ground. Data accessibility is an important tool for researchers, and NASA has committed to preserving this vital resource for ongoing use. The Life Sciences Data Archive’s multi-center collaboration between NASA’s Johnson Space Center, Ames Research Center, and Kennedy Space Center is geared toward preserving unique and high-value data from a wide variety of disciplines, data collection methods, and species within NASA’s Life Sciences Portal (NLSP). The data generated by the Human Research Program (HRP) require a systematic approach to data preservation that accounts for diverse data sources, formats, physical storage requirements, and security and privacy protections. This poster presentation will provide a guide to the repositories where the various types of human, non-human animal, plant, and microbial data NASA generates are archived and tips for navigating these data collections. Topics will include where different types of data, metadata, and biospecimens are archived or preserved, how the federated repositories work together as a data preservation ecosystem, and how researchers can access each repository’s collections.

Robert S Beaton↗

A Field Guide to NASA’s Life Sciences Data Repositories

For over 30 years, NASA has invested in life sciences research both in space and on the ground. Data accessibility is an important tool for researchers, and NASA has committed to preserving this vital resource for ongoing use. The Life Sciences Data Archive’s multi-center collaboration between NASA’s Johnson Space Center, Ames Research Center, and Kennedy Space Center is geared toward preserving unique and high-value data from a wide variety of disciplines, data collection methods, and species within NASA’s Life Sciences Portal (NLSP). The data generated by the Human Research Program (HRP) require a systematic approach to data preservation that accounts for diverse data sources, formats, physical storage requirements, and security and privacy protections. This poster presentation will provide a guide to the repositories where the various types of human, non-human animal, plant, and microbial data NASA generates are archived and tips for navigating these data collections. Topics will include where different types of data, metadata, and biospecimens are archived or preserved, how the federated repositories work together as a data preservation ecosystem, and how researchers can access each repository’s collections.

LSDA↗

The NPOESS Preparatory Project Science Data Segment: Brief Overview

The NPOESS Preparatory Project (NPP) provides remotely-sensed land, ocean, atmospheric, ozone, and sounder data that will serve the meteorological and global climate change scientific communities while also providing risk reduction for the National Polar-orbiting Operational Environmental Satellite System (NPOESS), the U.S. Government s future low-Earth orbiting satellite system monitoring global weather and environmental conditions. NPOESS and NPP are a new era, not only because the sensors will provide unprecedented quality and volume of data but also because it is a joint mission of three federal agencies, NASA, NOAA, and DoD. NASA's primary science role in NPP is to independently assess the quality of the NPP science and environmental data records. Such assessment is critical for making NPOESS products the best that they can be for operational use and ultimately for climate studies. The Science Data Segment (SDS) supports science assessment by assuring the timely provision of NPP data to NASA s science teams organized by climate measurement themes. The SDS breaks down into nine major elements, an input element that receives data from the operational agencies and acts as a buffer, a calibration analysis element, five elements devoted to measurement based quality assessment, an element used to test algorithmic improvements, and an element that provides overall science direction. This paper will describe how the NPP SDS will leverage on NASA experience to provide a mission-reliable research capability for science assessment of NPP derived measurements.

Schweiss, Robert J.↗

Framework for Integrating Science Data Processing Algorithms Into Process Control Systems

A software framework called PCS Task Wrapper is responsible for standardizing the setup, process initiation, execution, and file management tasks surrounding the execution of science data algorithms, which are referred to by NASA as Product Generation Executives (PGEs). PGEs codify a scientific algorithm, some step in the overall scientific process involved in a mission science workflow. The PCS Task Wrapper provides a stable operating environment to the underlying PGE during its execution lifecycle. If the PGE requires a file, or metadata regarding the file, the PCS Task Wrapper is responsible for delivering that information to the PGE in a manner that meets its requirements. If the PGE requires knowledge of upstream or downstream PGEs in a sequence of executions, that information is also made available. Finally, if information regarding disk space, or node information such as CPU availability, etc., is required, the PCS Task Wrapper provides this information to the underlying PGE. After this information is collected, the PGE is executed, and its output Product file and Metadata generation is managed via the PCS Task Wrapper framework. The innovation is responsible for marshalling output Products and Metadata back to a PCS File Management component for use in downstream data processing and pedigree. In support of this, the PCS Task Wrapper leverages the PCS Crawler Framework to ingest (during pipeline processing) the output Product files and Metadata produced by the PGE. The architectural components of the PCS Task Wrapper framework include PGE Task Instance, PGE Config File Builder, Config File Property Adder, Science PGE Config File Writer, and PCS Met file Writer. This innovative framework is really the unifying bridge between the execution of a step in the overall processing pipeline, and the available PCS component services as well as the information that they collectively manage.

Mattmann, Chris A.↗

Earth Science Data Fusion with Event Building Approach

Objectives of the NASA Information And Data System (NAIADS) project are to develop a prototype of a conceptually new middleware framework to modernize and significantly improve efficiency of the Earth Science data fusion, big data processing and analytics. The key components of the NAIADS include: Service Oriented Architecture (SOA) multi-lingual framework, multi-sensor coincident data Predictor, fast into-memory data Staging, multi-sensor data-Event Builder, complete data-Event streaming (a work flow with minimized IO), on-line data processing control and analytics services. The NAIADS project is leveraging CLARA framework, developed in Jefferson Lab, and integrated with the ZeroMQ messaging library. The science services are prototyped and incorporated into the system. Merging the SCIAMACHY Level-1 observations and MODIS/Terra Level-2 (Clouds and Aerosols) data products, and ECMWF re- analysis will be used for NAIADS demonstration and performance tests in compute Cloud and Cluster environments.

Lukashin, C.↗

Using NASA's Giovanni Web Portal to Access and Visualize Satellite-based Earth Science Data in the Classroom

One of the biggest obstacles for the average Earth science student today is locating and obtaining satellite-based remote sensing data sets in a format that is accessible and optimal for their data analysis needs. At the Goddard Earth Sciences Data and Information Services Center (GES-DISC) alone, on the order of hundreds of Terabytes of data are available for distribution to scientists, students and the general public. The single biggest and time-consuming hurdle for most students when they begin their study of the various datasets is how to slog through this mountain of data to arrive at a properly sub-setted and manageable data set to answer their science question(s). The GES DISC provides a number of tools for data access and visualization, including the Google-like Mirador search engine and the powerful GES-DISC Interactive Online Visualization ANd aNalysis Infrastructure (Giovanni) web interface.

Lloyd, Steven↗

The Earth Observing System - A multidisciplinary system for the long-term acquisition of earth science data from space

The development of the Earth Observing System (EOS) requirements is examined. An envisioned, EOS is to be NASA's multidisciplinary approach to acquisition of earth science data in the 1990s. The rationale and assumptions for the EOS are presented, together with a discussion of anticipated operations and use environments. Planned instrumentation concepts are discussed, involving a broad array of multiuse sensors operating across the electromagnetic spectrum from the ultraviolet to the microwave and including both active and passive techniques. Concepts for the near-polar orbiting observatory system, and essentially permanent man-serviced and -tended facility, are presented, along with concepts for information capture, processing and distribution and for common data-base referencing of acquired data. Milestones are indicated that are to be passed before a commitment to program execution can be made later in this decade.

Broome, D. R., Jr.↗

Earth Science Data Analytics: Bridging Tools and Techniques with the Co-Analysis of Large, Heterogeneous Datasets

The continuum of ever-evolving data management systems affords great opportunities to the enhancement of knowledge and facilitation of science research. To take advantage of these opportunities, it is essential to understand and develop methods that enable data relationships to be examined and the information to be manipulated. This presentation describes the efforts of the Earth Science Information Partners (ESIP) Federation Earth Science Data Analytics (ESDA) Cluster to understand, define, and facilitate the implementation of ESDA to advance science research. As a result of the void of Earth science data analytics publication material, the cluster has defined ESDA along with 10 goals to set the framework for a common understanding of tools and techniques that are available and still needed to support ESDA.

science data analysis↗

Square Kilometre Array Science Data Challenge 1: analysis and results

ABSTRACT As the largest radio telescope in the world, the Square Kilometre Array (SKA) will lead the next generation of radio astronomy. The feats of engineering required to construct the telescope array will be matched only by the techniques developed to exploit the rich scientific value of the data. To drive forward the development of efficient and accurate analysis methods, we are designing a series of data challenges that will provide the scientific community with high-quality data sets for testing and evaluating new techniques. In this paper, we present a description and results from the first such Science Data Challenge 1 (SDC1). Based on SKA MID continuum simulated observations and covering three frequencies (560, 1400, and 9200 MHz) at three depths (8, 100, and 1000 h), SDC1 asked participants to apply source detection, characterization, and classification methods to simulated data. The challenge opened in 2018 November, with nine teams submitting results by the deadline of 2019 April. In this work, we analyse the results for eight of those teams, showcasing the variety of approaches that can be successfully used to find, characterize, and classify sources in a deep, crowded field. The results also demonstrate the importance of building domain knowledge and expertise on this kind of analysis to obtain the best performance. As high-resolution observations begin revealing the true complexity of the sky, one of the outstanding challenges emerging from this analysis is the ability to deal with highly resolved and complex sources as effectively as the unresolved source population.

Bonaldi, A.↗

Enhancing NASA Earth Science Data Discovery from Scientific Publications

Earth observations from space borne instruments have evolved explosively in the past decades. Following closely are reanalysis systems assimilating model and observational data, yielding even longer records and larger number of variables. Thanks to advances in internet technology, it is now easier than ever to visualize and analyze these data using web interfaces. On the other hand, it also becomes an increasingly daunting task to build upon the existing knowledge published in various peer reviewed sources, and navigate toward the most relevant data, analysis, and visualization. We present an analysis of a subset of publications that utilized a popular visualization web interface at the NASA Goddard Earth Science Data and Information Services Center. Known as "Giovanni", it allows researchers from wide backgrounds to work with hundreds of variables from space observations and assimilation systems. Since coming online more than a decade ago, Giovanni has been credited in more than 100 papers per year, and the total count now is estimated to be nearly 1,500. Many of these papers contain valuable information about when, where and how Giovanni has been used, and hence forge an opportunity to learn and share the knowledge of which variables were used for what research projects. The purpose of our work is to retrieve the information from the papers and organize it as a knowledge repository which links together datasets, variables, places, dates and phenomena all of which reflect the essence of the published research. Since the publications are unstructured texts, we use natural language processing along with machine learning methods in the retrieval process. One of the challenges is deciphering the dataset names, because in many cases researchers refer to variables, rather than the datasets containing them. To constrain the number of terms, we deploy Earth Science ontologies as dictionaries for the term extraction. We demonstrate that storing these terms and underlying ontologies, along with datasets, variables and papers in the knowledge graph database, enables various linkages between all these entities facilitating the data discovery. Thus, we are setting a qualitatively new stage in improvements of web data interfaces, where machine learning techniques are used to establish and optimize usage-based discovery of data.

Irina V Gerasimov↗

Applications of wavelet-based compression to multidimensional Earth science data

A data compression algorithm involving vector quantization (VQ) and the discrete wavelet transform (DWT) is applied to two different types of multidimensional digital earth-science data. The algorithms (WVQ) is optimized for each particular application through an optimization procedure that assigns VQ parameters to the wavelet transform subbands subject to constraints on compression ratio and encoding complexity. Preliminary results of compressing global ocean model data generated on a Thinking Machines CM-200 supercomputer are presented. The WVQ scheme is used in both a predictive and nonpredictive mode. Parameters generated by the optimization algorithm are reported, as are signal-to-noise (SNR) measurements of actual quantized data. The problem of extrapolating hydrodynamic variables across the continental landmasses in order to compute the DWT on a rectangular grid is discussed. Results are also presented for compressing Landsat TM 7-band data using the WVQ scheme. The formulation of the optimization problem is presented along with SNR measurements of actual quantized data. Postprocessing applications are considered in which the seven spectral bands are clustered into 256 clusters using a k-means algorithm and analyzed using the Los Alamos multispectral data analysis program, SPECTRUM, both before and after being compressed using the WVQ program.

Bradley, Jonathan N.↗

Application of Data Science and Engineering

Metal additive manufacturing (AM) processes exhibit significant variability in the quality and properties of components that are produced. This variability has prevented the widespread adoption of AM in industry. The need for more advanced and descriptive process monitoring, part qualification, and process control has led to an increasing number of sensors on machines and subsequent data to analyze. Increasingly, data science principles are being leveraged in each of these domains in order to process this data and better understand the causes of variability and the corresponding quality inconsistencies that occur in additive manufacturing.

Halsey, William↗

In the Mix : A Workshop Merging Computational Chemistry and Electrochemistry Alongside Data Science

As chemistry expands to more complex and interdisciplinary areas, a new generation of diverse researchers must engage with science and learn effective cross-disciplinary collaboration and communication. To these ends, we designed and implemented In the Mix, a graduate student-led, two-day workshop for undergraduate students promoting collaborative science in the context of energy storage innovations. Here, the interactive workshop was designed for future and emerging researchers to gain hands-on experience with data science, computational chemistry, and electrochemistry techniques that are critical for developing materials for battery technologies. Participants also visited commercial renewable energy facilities to help them connect discovery-based research with industry and broader societal considerations. The workshop content and structure ensured that participants experienced the interrelatedness of the fields and understood the importance of collaborative research to yield scientific advances with real-world applications. An external team evaluated the workshop and participants’ perceptions of their experiences. While our research context was energy storage, the workshop goals and outcomes are applicable to other contexts. Interdisciplinary, experiential workshops are a key avenue to broadening participation in science and research, and the ideas presented here can be readily modified for other scientific contexts and/or incorporated as broader impact activities.

25 ENERGY STORAGE↗

Tracking Provenance of Earth Science Data

Tremendous volumes of data have been captured, archived and analyzed. Sensors, algorithms and processing systems for transforming and analyzing the data are evolving over time. Web Portals and Services can create transient data sets on-demand. Data are transferred from organization to organization with additional transformations at every stage. Provenance in this context refers to the source of data and a record of the process that led to its current state. It encompasses the documentation of a variety of artifacts related to particular data. Provenance is important for understanding and using scientific datasets, and critical for independent confirmation of scientific results. Managing provenance throughout scientific data processing has gained interest lately and there are a variety of approaches. Large scale scientific datasets consisting of thousands to millions of individual data files and processes offer particular challenges. This paper uses the analogy of art history provenance to explore some of the concerns of applying provenance tracking to earth science data. It also illustrates some of the provenance issues with examples drawn from the Ozone Monitoring Instrument (OMI) Data Processing System (OMIDAPS) run at NASA's Goddard Space Flight Center by the first author.

Tilmes, Curt↗

Data Science Enabled Enabled Discovery of Superconductors (Final Progress Report)

This Final Technical Report describes efforts by 4 PIs at the University of Florida (Peter Hirschfeld, Richard Hennig, Greg Stewart and James Hamlin), over the period September 2019-August 2023, to use data science and machine learning techniques to discover new conventional superconductors. The PIs constructed a discovery loop with two theorists and two experimentalists to: develop algorithms to machine learn descriptors correlating strongly with the critical temperature Tc (PI's Peter Hirschfeld, UF Physics and Richard Hennig, UF Materials Science and En), synthesize and measure properties of promising materials, and feed back the knowledge gained into the prediction algorithm. This work was motivated by the theoretical prediction and experimental discovery of high-pressure, high-pressure hydride superconductors, and to find ways to recreate the high critical temperatures in these systems at ambient pressure. Highlights from the grant include: 1) a new equation for Tc in terms of moments of the electron-phonon spectral function, improving on the so-called Allen-Dynes equation (1975); 2) study of the metastable A15 superconductor Nb3Si, formed under explosive compression at ~1000GPa to determine the kinetic barrier to the ground state structure; 3) the development of ultra-fast machine-learned atomic potentials for molecular dynamics, and 4) the discovery of superconductivity at 19K in WB2 arising from metastable defect structures in the crystal.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

NASA GIBS and Worldview: Leveraging FOSS for NASA Earth Science Data Visualizations

The NASA Global Imagery Browse Services (GIBS) and Worldview interactive mapping site leverage scientific and community best practices, open source software, and public standards to provide a scalable, compliant, and authoritative source for NASA Earth Observing System (EOS) Earth science data visualizations. GIBS and Worldview allow end users to easily and quickly interact with more than 800 full resolution pre-generated raster- and vector-based visualizations. This interactive discovery approach relies on visual observation and identification of phenomena that are not as simply identified otherwise. This eLightning presentation will exhibit the broad set of capabilities and visualization layers made possible through the GIBS and Worldview open source software. Specific dependencies on, and contributions to, open source software will be highlighted. Additionally, opportunities for future improvements for better interoperability and reuse through open source software will be discussed.

NASA Global Imagery Browse Services (GIBS)↗

Development of the Science Data System for the International Space Station Cold Atom Lab

Cold Atom Laboratory (CAL) is a facility that will enable scientists to study ultra-cold quantum gases in a microgravity environment on the International Space Station (ISS) beginning in 2016. The primary science data for each experiment consists of two images taken in quick succession. The first image is of the trapped cold atoms and the second image is of the background. The two images are subtracted to obtain optical density. These raw Level 0 atom and background images are processed into the Level 1 optical density data product, and then into the Level 2 data products: atom number, Magneto-Optical Trap (MOT) lifetime, magnetic chip-trap atom lifetime, and condensate fraction. These products can also be used as diagnostics of the instrument health. With experiments being conducted for 8 hours every day, the amount of data being generated poses many technical challenges, such as downlinking and managing the required data volume. A parallel processing design is described, implemented, and benchmarked. In addition to optimizing the data pipeline, accuracy and speed in producing the Level 1 and 2 data products is key. Algorithms for feature recognition are explored, facilitating image cropping and accurate atom number calculations.

bose einstein condensate↗

A Unified Level of Service Model for NASA Earth Science Data Stewardship

During the past year, the Interagency Implementation and Concepts Team (IMPACT) reviewed existing service models in use at various NASA data centers in an effort to produce a unified, cohesive, and comprehensive Level of Service Model for all of NASA Distributed Active Archive Centers (DAACs). NASA DAACs are responsible for ensuring NASA Earth Science data are accurately and securely ingested, distributed, supported, and preserved. The term “Service" as used here refers to the spectrum of data management activities and outputs provided by DAACs in support of the data cared fo by each data cente. The unified Level-of- Service (LoS) model described in this presentation utilizes both the NASA-defined data product category and the data processing level to easily identify an appropriate level-of-service to be applied to a data product throughout the full data life cycle. This LoS model is to be used by DAACs when appraising incoming data in order to determine the appropriate and required services to provide. The LoS model utilizes a 3-level system in which services build upon previous levels and thereby require greater commitment and effort both on the part of the DAAC personnel and the data producer at the highest level. The LoS model description also contains examples of ways to communicate with data producers and data users what services can be expected, thereby bringing more consistent user experiences across the enterprise. In this presentation, we will outline the features of the LoS model and describe how it relates to the FAIR data practices and the NOAA Maturity Matrix model.

Smith, Deborah↗