Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Scientific data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13

F*** workflows: when parts of FAIR are missing

The FAIR principles for scientific data (Findable, Accessible, Interoperable, Reusable) are also relevant to other digital objects such as research software and scientific workflows that operate on scientific data. The FAIR principles can be applied to the data being handled by a scientific workflow as well as the processes, software, and other infrastructure which are necessary to specify and execute a workflow. The FAIR principles were designed as guidelines, rather than rules, that would allow for differences in standards for different communities and for different degrees of compliance. There are many practical considerations which impact the level of FAIR-ness that can actually be achieved, including policies, traditions, and technologies. Because of these considerations, obstacles are often encountered during the workflow lifecycle that trace directly to shortcomings in the implementation of the FAIR principles. Here, we detail some cases, without naming names, in which data and workflows were Findable but otherwise lacking in areas commonly needed and expected by modern FAIR methods, tools, and users. We describe how some of these problems, all of which were overcome successfully, have motivated us to push on systems and approaches for fully FAIR workflows.

Wilkinson, Sean↗

Sim2Ls: FAIR simulation workflows and data

Just like the scientific data they generate, simulation workflows for research should be findable, accessible, interoperable, and reusable (FAIR). However, while significant progress has been made towards FAIR data, the majority of science and engineering workflows used in research remain poorly documented and often unavailable, involving ad hoc scripts and manual steps, hindering reproducibility and stifling progress. We introduce Sim2Ls (pronounced simtools) and the Sim2L Python library that allow developers to create and share end-to-end computational workflows with well-defined and verified inputs and outputs. The Sim2L library makes Sim2Ls , their requirements, and their services discoverable, verifies inputs and outputs, and automatically stores results in a globally-accessible simulation cache and results database. This simulation ecosystem is available in nanoHUB, an open platform that also provides publication services for Sim2Ls , a computational environment for developers and users, and the hardware to execute runs and store results at no cost. We exemplify the use of Sim2Ls using two applications and discuss best practices towards FAIR simulation workflows and associated data.

59 BASIC BIOLOGICAL SCIENCES↗

Array-Pattern-Match Compiler for Opportunistic Data Analysis

A computer program has been written to facilitate real-time sifting of scientific data as they are acquired to find data patterns deemed to warrant further analysis. The patterns in question are of a type denoted array patterns, which are specified by nested parenthetical expressions. [One example of an array pattern is ((>3) 0 (not=1)): this pattern matches a vector of at least three elements, the first of which exceeds 3, the second of which is 0, and the third of which does not equal 1.] This program accepts a high-level description of a static array pattern and compiles a highly optimal and compact other program to determine whether any given instance of any data array matches that pattern. The compiler implemented by this program is independent of the target language, so that as new languages are used to write code that processes scientific data, they can easily be adapted to this compiler. This program runs on a variety of different computing platforms. It must be run in conjunction with any one of a number of Lisp compilers that are available commercially or as shareware.

James, Mark↗

FunMC^2: A Filter for Uncertainty Visualization of Marching Cubes on Multi-Core Devices

Visualization is an important tool for scientists to extract understanding from complex scientific data. Scientists need to understand the uncertainty inherent in all scientific data in order to interpret the data correctly. Uncertainty visualization has been an active and growing area of research to address this challenge. Algorithms for uncertainty visualization can be expensive, and research efforts have been focused mainly on structured grid types. Further, support for uncertainty visualization in production tools is limited. In this paper, we adapt an algorithm for computing key metrics for visualizing uncertainty in Marching Cubes (MC) to multi-core devices and present the design, implementation, and evaluation for a Filter for uncertainty visualization of Marching Cubes on Multi-Core devices (FunMC2). FunMC2 accelerates the uncertainty visualization of MC significantly, and it is portable across multi-core CPUs and GPUs. Evaluation results show that FunMC2 based on OpenMP runs around 11× to 41× faster on multi-core CPUs than the corresponding serial version using one CPU core. FunMC2 based on a single GPU is around 5× to 9× faster than FunMC2 running by OpenMP. Moreover, FunMC2 is flexible enough to process ensemble data with both structured and unstructured mesh types. Furthermore, we demonstrate that FunMC2 can be seamlessly integrated as a plugin into ParaView, a production visualization tool for post-processing.

Wang, Jay↗

Methods and Experiences for Developing Abstractions for Data-intensive, Scientific Applications

Developing software for scientific applications that require the integration of diverse types of computing, instruments, and data present challenges that are distinct from commercial software. These applications require scale, and the need to integrate various programming and computational models with evolving and heterogeneous infrastructure. Pervasive and effective abstractions for distributed infrastructures are thus critical; however, the process of developing abstractions for scientific applications and infrastructures is not well understood. While theory-based approaches for system development are suited for well-defined, closed environments, they have severe limitations for designing abstractions for scientific systems and applications. The design science research (DSR) method provides the basis for designing practical systems that can handle real-world complexities at all levels. In contrast to theory-centric approaches, DSR emphasizes both practical relevance and knowledge creation by building and rigorously evaluating all artifacts. In this work, we show how DSR provides a well-defined framework for developing abstractions and middleware systems for distributed systems. Specifically, we address the critical problem of distributed resource management on heterogeneous infrastructure over a dynamic range of scales, a challenge that currently limits many scientific applications. We use the pilot-abstraction, a widely used resource management abstraction for high-performance, high throughput, big data, and streaming applications, as a case study for evaluating the DSR activities. For this purpose, we analyze the research process and artifacts produced during the design and evaluation of the pilot-abstraction. We find DSR provides a concise framework for iteratively designing and evaluating systems. Finally, we capture our experiences and formulate different lessons learned.

97 MATHEMATICS AND COMPUTING↗

Earth science and application

The University of Alabama in Huntsville (UAH) has completed the research proposed. The major tasks under this contract were: (1) research into visualization of scientific data sets (browse); (2) studies of standard data formatting procedures; and (3) investigations of approaches for submission of scientific data sets for archival. Summaries of each activity are presented along with travel reports and conclusions and recommendations.

Hardin, Danny↗

Snakes on a Spaceship - An Overview of Python in Heliophysics

Computational analysis has become ubiquitous within the heliophysics community. However, community standards for peer review of codes and analysis have lagged behind these developments. This absence has contributed to the reproducibility crisis, where inadequate analysis descriptions and loss of scientific data have made scientific studies difficult or impossible to replicate. The heliophysics community has responded to this challenge by expressing a desire for a more open, collaborative set of analysis tools. This article summarizes the current state of these efforts and presents an overview of many of the existing Python heliophysics tools. It also outlines the challenges facing community members who are working toward the goal of an open, collaborative, Python heliophysics toolkit and presents guidelines that can ease the transition from individualistic data analysis practices to an accountable, communalistic environment.

Burrell, A.G.↗

Structural Analysis Report for Sandia High Altitude Aerosol Research (SHAAR) Payloads

Sandia National Laboratories has interest in mounting enclosures to gather scientific data aboard NASA scientific balloons. This report documents the structural integrity of three separate payloads considered ‘piggybacks.’ To date, there are no design criteria for piggybacks, therefore each piggyback shall follow the same design requirements set by NASA per the Gondola Structural Design Requirements. This analysis report shall describe how each payload meets the design requirements laid out in the Gondola Design Requirements.

47 OTHER INSTRUMENTATION↗

Science Activity Planner for the MER Mission

The Maestro Science Activity Planner is a computer program that assists human users in planning operations of the Mars Explorer Rover (MER) mission and visualizing scientific data returned from the MER rovers. Relative to its predecessors, this program is more powerful and easier to use. This program is built on the Java Eclipse open-source platform around a Web-browser-based user-interface paradigm to provide an intuitive user interface to Mars rovers and landers. This program affords a combination of advanced display and simulation capabilities. For example, a map view of terrain can be generated from images acquired by the High Resolution Imaging Science Explorer instrument aboard the Mars Reconnaissance Orbiter spacecraft and overlaid with images from a navigation camera (more precisely, a stereoscopic pair of cameras) aboard a rover, and an interactive, annotated rover traverse path can be incorporated into the overlay. It is also possible to construct an overhead perspective mosaic image of terrain from navigation-camera images. This program can be adapted to similar use on other outer-space missions and is potentially adaptable to numerous terrestrial applications involving analysis of data, operations of robots, and planning of such operations for acquisition of scientific data.

Norris, Jeffrey S.↗

JPSS-3 / 4 VIIRS Response Versus Scan Angle Characterization and Performance

Scientific studies of the Earth’s climate increasingly rely on high-quality satellite observations. The Visible Infrared Imaging Radiometer Suite (VIIRS) is a key sensor onboard a series of satellites [Suomi National Polar-orbiting Partnership (SNPP) and Joint Polar-orbiting Satellite System 1–4 (JPSS-1–JPSS-4)] that generate scientific data from land, ocean, and atmosphere used in these climate models. Providing quality scientific data from space-borne sensors requires the instruments to be well-calibrated. While much of the calibration can be maintained on-orbit, some aspects of the calibration can best be measured prior to launch. One VIIRS parameter that needs to be measured pre-launch is the response versus scan angle (RVS). The RVS measures the relative change in the reflectance of the scanning optics as a function of the angle of incidence. With the RVS, the gain calibration measured on-orbit can be transferred to any scan angle. The JPSS-3 and JPSS-4 instruments have undergone ground testing including the RVS measurements, which is the subject of this work. Results indicate that the measurements are comparable to previous VIIRS builds and are expected to contribute to the generation of high-quality science data once JPSS-3 and JPSS-4 are on-orbit.

JPSS↗

ROSAT data analysis with EXSAS

For the x-ray observatory ROSAT, data from survey and pointed mission phases taken with different focal plane instruments and according to a complex mission timeline have to be handled. Data analysis therefore puts high demands on appropriate software tools. With EXSAS - the Extended Scientific Analysis System developed with an effort of 20 man years by the German ROSAT Scientific Data Center - a comfortable system for the reduction of data from the ROSAT x-ray and XUV instruments has been made available. EXSAS comprises a large collection of application modules as typically required in analyzing data of this wavelength regime and runs as a specific context in the wide-spread ESO-MIDAS environment. EXSAS, completely written in FORTRAN 77, takes full advantage of all the standards used in MIDAS and therefore, reflects the same portability (different UNIX installations and VMS). If required, the FORTRAN code also enables users to adapt the software in an easy way to their specific needs. To maintain independence from the specifics of different operating systems also on the data input side, all ROSAT data redistributed in the widely accepted FITS format. Although EXSAS has been developed specifically for data analysis of the ROSAT instruments, its structural design is sufficiently general to serve equally well also data from other X-ray and XUV instruments. EXSAS analysis modules are grouped into 4 application packages dealing with Data Preparation and Instrument Correction, Spatial Analysis, Spectral Analysis and Timing Analysis. A special EXSAS header, read and updated by each application, maintains the general information transfer on the origin, the history and the parameter space of the data stored in tables and images. About 100 genuine commands (most of which offer several additional options) allow to interactively explore the functionality of the system. Up to now 40 institutes all over the world have requested the EXSAS software. Maintenance and regular updates of the software and the comprehensive documentation are provided by the ROSAT Scientific Data Center at Garching.

Zimmermann, H. U.↗

A guide to NASA's Pilot Land Data System (PLDS)

NASA's Pilot Land Data System (PLDS) is a distributed information management system designed to support NASA's land science community. The PLDS provides a wide range of services including management of information about scientific data, access to a library of scientific data, a data ordering capability, communications, connection to data analysis facilities, and electronic mail. The PLDS provides these services by offering the scientist the capability to search for and order data, and to communicate electronically with other scientists and computers. Three functions enable scientists to find what data are available and where they reside. The first two, Find data summaries and Read detailed descriptions give summary and detailed descriptions about data sets or groups of related data sets, science, projects, and institutions which archive land data. The third, gives information about specific pieces of data. This last function has two components, Search systemwide inventory and Search local inventory. The first component enables the user to find data elements (images, geological samples, transects, maps, etc.) that exist anywhere in the PLDS while the second has only information about data at the local site. The first enables the user to find pieces of data from several different data sets with the same temporal and spatial coverage and other elements common to most data sets, while the second allows the user to select a data set based on these descriptors and on those that are unique to a data set. The PLDS provides capabilities that enable electronic file transfers, intercomputer connection, and electronic mail. Both TCP/IP and DECnet protocols are supported via the NASA Science Internet (NIS). Access is also available through Telenet.

Source record↗

NASA biological and physical sciences databases: who’s the FAIRest of them all?

Conceptual models are a key part of the foundation of scientific study. Scientific data discovery and retrieval are often inaccurate and incomplete because these models are not sufficiently well-incorporated into data retrieval systems. Systems often don’t provide the necessary tools to those producing scientific data to fully and unambiguously annotate them and the result is consumers of the data cannot find them efficiently. The capability of data archives to provide these tools to link data to underlying conceptual models is one of dimensions of the recently developed “FAIR” principles (https://www.go-fair.org/fair-principles/ ), and is key to many automated processes being able to operate on these data, particularly analytics involving artificial intelligence. We used an open-source web service to measure the FAIR compliance of the three data archives operated by NASA for the biological and physical sciences: the Life Sciences Data Archive, the Physical Sciences Informatics database, and GeneLab. The service ingests references to data sets in these archives, and then executes domain-non-specific examinations of these data and metadata that test compliance to the FAIR principles. Of the 22 metrics tested, GeneLab passed 11 (50%), and PSI and LSDA each passed 7 (32%). These data were gathered using only one representative data set from each archive and we anticipate variability in results as we continue to apply these metrics to other data. A preliminary study of the failure traces for each metric suggests there is a wide range of effort and complexity in the enhancements required for each system to elevate FAIR compliance, and this is the subject of continued investigation. This information has been and will likely continue to be important information in planning these enhancements, with the goal of increased readiness of the data for automated processes.

database↗

A Web of Data Analytics Services

Cloud Computing has become the ubiquitous approach to our Big Data challenge. However, one will quickly discover that moving (a.k.a. forklifting) existing on-premise data analytics solutions to the Cloud doesn’t always translate to costing saving and performance boost. The Cloud’s elasticity, its availability, and its wide selection of computing options and selections of costing models making Cloud an attractive environment to tackle our Big Data challenge. The fact is Cloud, on its own, is not the silver bullet to our daunting challenge need for analyze and derive scientific inferences through vast collections of multi-sensor measurements. We would like to have all scientific data in one easy to access environment, but getting the world of scientific data in one analytic system is immensely difficult to achieve. This paper describes the data analytics web architecture NASA is developing by infusing instances of Integrated Data Analytics systems next to the data. The goal is to minimize unnecessary data movement through collection of data access and analytics webservices for researchers to interact with and analyze measurements without have to download data to their local computer. These services are RESTful and provisioned by the data centers with the help from subject matter and science experts. These services encapsulate the physical computing infrastructure, which could local computing cluster, on-premise or public Cloud environment.

Huang, Thomas↗

Expanding Repository Data Available For Sharing and Knowledge Discovery

Some of the hardest space biology and space health challenges require data-intensive, bioinformatic, meta-analytical, and computer-assisted research approaches. These challenges include examining interdisciplinary space life science research across experiments and across interacting spaceflight hazards (radiation, altered gravity, confinement, hostile-closed environments, distance-duration from Earth). The approaches to confront these challenges involve mining multiple datasets simultaneously from various hierarchical organizations of biological complexity, all while concurrently evaluating how experimental design factors affect endpoints of standard assays. To enable this field, it is essential that principal investigators (PIs) submit data in a structure so it can be maximally re-used. The purpose of the NASA Ames Life Sciences Data Archive (ALSDA) is to collect, curate, and make publicly available all non-human space-relevant biological data. ALSDA must also ensure data are open-access, and maximally findable, accessible, interoperable, and reusable (FAIR). The scope of ALSDA data collected and submitted by PIs include subject and study design metadata, assay metadata parameters, raw and processed assay data, assay imagery/video, and subject-experienced mission data telemetry (radiation, temperature, humidity, acoustics, vibrations, etc.). ALSDA recently integrated into a collaborative group of Open Science projects to facilitate a suite of new tools and workflows that will improve data submission, accessibility, and reusability by implementing digital data submission agreements, and adopting the data management system originally developed by NASA GeneLab. ALSDA intends to bring current biological repository data and all future collected data into this new scientific data reuse reality. This new suite of tools will enable ALSDA to deploy a science curation system using scientific assay configurations for the data submission portal. It will capture essential assay parameters according to established standards in each sub-field within biology. The submission portal expedites data collection by enhancing ease of PI data submission, providing a user interface and specificity for which data is to be submitted. Data submissions can be brought into cutting-edge informatic analysis portals to enable mining of physiological, behavioral, biochemical, and imaging datasets in conjunction with ‘omics-level datasets. As ALSDA datasets are submitted, curated, and published (e.g., micro-computed tomography, histology, pulse oximetry, serum metabolites, magnetic resonance imaging, intraocular pressure, novel object recognition, etc.), the merging together of spaceflight data along this multi-hierarchical complexity of biology will enable informatics and data-intensive approaches resulting in knowledge discoveries across missions, space hazards, and biological disciplines.

Biology↗

Expanding Repository Data Available For Sharing And Knowledge Discovery

Some of the hardest space biology and space health challenges require data-intensive, bioinformatic, meta-analytical, and computer-assisted research approaches. These challenges include examining interdisciplinary space life science research across experiments and across interacting spaceflight hazards (radiation, altered gravity, confinement, hostile-closed environments, distance-duration from Earth). The approaches to confront these challenges involve mining multiple datasets simultaneously from various hierarchical organizations of biological complexity, all while concurrently evaluating how experimental design factors affect endpoints of standard assays. To enable this field, it is essential that principal investigators (PIs) submit data in a structure so it can be maximally re-used. The purpose of the NASA Ames Life Sciences Data Archive (ALSDA) is to collect, curate, and make publicly available all non-human space-relevant biological data. ALSDA must also ensure data are open-access, and maximally findable, accessible, interoperable, and reusable (FAIR). The scope of ALSDA data collected and submitted by PIs include subject and study design metadata, assay metadata parameters, raw and processed assay data, assay imagery/video, and subject-experienced mission data telemetry (radiation, temperature, humidity, acoustics, vibrations, etc.). ALSDA recently integrated into a collaborative group of Open Science projects to facilitate a suite of new tools and workflows that will improve data submission, accessibility, and reusability by implementing digital data submission agreements, and adopting the data management system originally developed by NASA GeneLab. ALSDA intends to bring current biological repository data and all future collected data into this new scientific data reuse reality. This new suite of tools will enable ALSDA to deploy a science curation system using scientific assay configurations for the data submission portal. It will capture essential assay parameters according to established standards in each sub-field within biology. The submission portal expedites data collection by enhancing ease of PI data submission, providing a user interface and specificity for which data is to be submitted. Data submissions can be brought into cutting-edge informatic analysis portals to enable mining of physiological, behavioral, biochemical, and imaging datasets in conjunction with ‘omics-level datasets. As ALSDA datasets are submitted, curated, and published (e.g., micro-computed tomography, histology, pulse oximetry, serum metabolites, magnetic resonance imaging, intraocular pressure, novel object recognition, etc.), the merging together of spaceflight data along this multi-hierarchical complexity of biology will enable informatics and data-intensive approaches resulting in knowledge discoveries across missions, space hazards, and biological disciplines.

life science↗