Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “scientific data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

Generic functional requirements for a NASA general-purpose data base management system

Generic functional requirements for a general-purpose, multi-mission data base management system (DBMS) for application to remotely sensed scientific data bases are detailed. The motivation for utilizing DBMS technology in this environment is explained. The major requirements include: (1) a DBMS for scientific observational data; (2) a multi-mission capability; (3) user-friendly; (4) extensive and integrated information about data; (5) robust languages for defining data structures and formats; (6) scientific data types and structures; (7) flexible physical access mechanisms; (8) ways of representing spatial relationships; (9) a high level nonprocedural interactive query and data manipulation language; (10) data base maintenance utilities; (11) high rate input/output and large data volume storage; and adaptability to a distributed data base and/or data base machine configuration. Detailed functions are specified in a top-down hierarchic fashion. Implementation, performance, and support requirements are also given.

Lohman, G. M.↗

Improve Data Mining and Knowledge Discovery Through the Use of MatLab

Data mining is widely used to mine business, engineering, and scientific data. Data mining uses pattern based queries, searches, or other analyses of one or more electronic databases/datasets in order to discover or locate a predictive pattern or anomaly indicative of system failure, criminal or terrorist activity, etc. There are various algorithms, techniques and methods used to mine data; including neural networks, genetic algorithms, decision trees, nearest neighbor method, rule induction association analysis, slice and dice, segmentation, and clustering. These algorithms, techniques and methods used to detect patterns in a dataset, have been used in the development of numerous open source and commercially available products and technology for data mining. Data mining is best realized when latent information in a large quantity of data stored is discovered. No one technique solves all data mining problems; challenges are to select algorithms or methods appropriate to strengthen data/text mining and trending within given datasets. In recent years, throughout industry, academia and government agencies, thousands of data systems have been designed and tailored to serve specific engineering and business needs. Many of these systems use databases with relational algebra and structured query language to categorize and retrieve data. In these systems, data analyses are limited and require prior explicit knowledge of metadata and database relations; lacking exploratory data mining and discoveries of latent information. This presentation introduces MatLab(R) (MATrix LABoratory), an engineering and scientific data analyses tool to perform data mining. MatLab was originally intended to perform purely numerical calculations (a glorified calculator). Now, in addition to having hundreds of mathematical functions, it is a programming language with hundreds built in standard functions and numerous available toolboxes. MatLab's ease of data processing, visualization and its enormous availability of built in functionalities and toolboxes make it suitable to perform numerical computations and simulations as well as a data mining tool. Engineers and scientists can take advantage of the readily available functions/toolboxes to gain wider insight in their perspective data mining experiments.

Shaykhian, Gholam Ali↗

Improve Data Mining and Knowledge Discovery through the use of MatLab

Data mining is widely used to mine business, engineering, and scientific data. Data mining uses pattern based queries, searches, or other analyses of one or more electronic databases/datasets in order to discover or locate a predictive pattern or anomaly indicative of system failure, criminal or terrorist activity, etc. There are various algorithms, techniques and methods used to mine data; including neural networks, genetic algorithms, decision trees, nearest neighbor method, rule induction association analysis, slice and dice, segmentation, and clustering. These algorithms, techniques and methods used to detect patterns in a dataset, have been used in the development of numerous open source and commercially available products and technology for data mining. Data mining is best realized when latent information in a large quantity of data stored is discovered. No one technique solves all data mining problems; challenges are to select algorithms or methods appropriate to strengthen data/text mining and trending within given datasets. In recent years, throughout industry, academia and government agencies, thousands of data systems have been designed and tailored to serve specific engineering and business needs. Many of these systems use databases with relational algebra and structured query language to categorize and retrieve data. In these systems, data analyses are limited and require prior explicit knowledge of metadata and database relations; lacking exploratory data mining and discoveries of latent information. This presentation introduces MatLab(TradeMark)(MATrix LABoratory), an engineering and scientific data analyses tool to perform data mining. MatLab was originally intended to perform purely numerical calculations (a glorified calculator). Now, in addition to having hundreds of mathematical functions, it is a programming language with hundreds built in standard functions and numerous available toolboxes. MatLab's ease of data processing, visualization and its enormous availability of built in functionalities and toolboxes make it suitable to perform numerical computations and simulations as well as a data mining tool. Engineers and scientists can take advantage of the readily available functions/toolboxes to gain wider insight in their perspective data mining experiments.

Shaykahian, Gholan Ali↗

Array-Pattern-Match Compiler for Opportunistic Data Analysis

A computer program has been written to facilitate real-time sifting of scientific data as they are acquired to find data patterns deemed to warrant further analysis. The patterns in question are of a type denoted array patterns, which are specified by nested parenthetical expressions. [One example of an array pattern is ((>3) 0 (not=1)): this pattern matches a vector of at least three elements, the first of which exceeds 3, the second of which is 0, and the third of which does not equal 1.] This program accepts a high-level description of a static array pattern and compiles a highly optimal and compact other program to determine whether any given instance of any data array matches that pattern. The compiler implemented by this program is independent of the target language, so that as new languages are used to write code that processes scientific data, they can easily be adapted to this compiler. This program runs on a variety of different computing platforms. It must be run in conjunction with any one of a number of Lisp compilers that are available commercially or as shareware.

James, Mark↗

Earth science and application

The University of Alabama in Huntsville (UAH) has completed the research proposed. The major tasks under this contract were: (1) research into visualization of scientific data sets (browse); (2) studies of standard data formatting procedures; and (3) investigations of approaches for submission of scientific data sets for archival. Summaries of each activity are presented along with travel reports and conclusions and recommendations.

Hardin, Danny↗

Snakes on a Spaceship - An Overview of Python in Heliophysics

Computational analysis has become ubiquitous within the heliophysics community. However, community standards for peer review of codes and analysis have lagged behind these developments. This absence has contributed to the reproducibility crisis, where inadequate analysis descriptions and loss of scientific data have made scientific studies difficult or impossible to replicate. The heliophysics community has responded to this challenge by expressing a desire for a more open, collaborative set of analysis tools. This article summarizes the current state of these efforts and presents an overview of many of the existing Python heliophysics tools. It also outlines the challenges facing community members who are working toward the goal of an open, collaborative, Python heliophysics toolkit and presents guidelines that can ease the transition from individualistic data analysis practices to an accountable, communalistic environment.

Burrell, A.G.↗

Science Activity Planner for the MER Mission

The Maestro Science Activity Planner is a computer program that assists human users in planning operations of the Mars Explorer Rover (MER) mission and visualizing scientific data returned from the MER rovers. Relative to its predecessors, this program is more powerful and easier to use. This program is built on the Java Eclipse open-source platform around a Web-browser-based user-interface paradigm to provide an intuitive user interface to Mars rovers and landers. This program affords a combination of advanced display and simulation capabilities. For example, a map view of terrain can be generated from images acquired by the High Resolution Imaging Science Explorer instrument aboard the Mars Reconnaissance Orbiter spacecraft and overlaid with images from a navigation camera (more precisely, a stereoscopic pair of cameras) aboard a rover, and an interactive, annotated rover traverse path can be incorporated into the overlay. It is also possible to construct an overhead perspective mosaic image of terrain from navigation-camera images. This program can be adapted to similar use on other outer-space missions and is potentially adaptable to numerous terrestrial applications involving analysis of data, operations of robots, and planning of such operations for acquisition of scientific data.

Norris, Jeffrey S.↗

JPSS-3 / 4 VIIRS Response Versus Scan Angle Characterization and Performance

Scientific studies of the Earth’s climate increasingly rely on high-quality satellite observations. The Visible Infrared Imaging Radiometer Suite (VIIRS) is a key sensor onboard a series of satellites [Suomi National Polar-orbiting Partnership (SNPP) and Joint Polar-orbiting Satellite System 1–4 (JPSS-1–JPSS-4)] that generate scientific data from land, ocean, and atmosphere used in these climate models. Providing quality scientific data from space-borne sensors requires the instruments to be well-calibrated. While much of the calibration can be maintained on-orbit, some aspects of the calibration can best be measured prior to launch. One VIIRS parameter that needs to be measured pre-launch is the response versus scan angle (RVS). The RVS measures the relative change in the reflectance of the scanning optics as a function of the angle of incidence. With the RVS, the gain calibration measured on-orbit can be transferred to any scan angle. The JPSS-3 and JPSS-4 instruments have undergone ground testing including the RVS measurements, which is the subject of this work. Results indicate that the measurements are comparable to previous VIIRS builds and are expected to contribute to the generation of high-quality science data once JPSS-3 and JPSS-4 are on-orbit.

JPSS↗

ROSAT data analysis with EXSAS

For the x-ray observatory ROSAT, data from survey and pointed mission phases taken with different focal plane instruments and according to a complex mission timeline have to be handled. Data analysis therefore puts high demands on appropriate software tools. With EXSAS - the Extended Scientific Analysis System developed with an effort of 20 man years by the German ROSAT Scientific Data Center - a comfortable system for the reduction of data from the ROSAT x-ray and XUV instruments has been made available. EXSAS comprises a large collection of application modules as typically required in analyzing data of this wavelength regime and runs as a specific context in the wide-spread ESO-MIDAS environment. EXSAS, completely written in FORTRAN 77, takes full advantage of all the standards used in MIDAS and therefore, reflects the same portability (different UNIX installations and VMS). If required, the FORTRAN code also enables users to adapt the software in an easy way to their specific needs. To maintain independence from the specifics of different operating systems also on the data input side, all ROSAT data redistributed in the widely accepted FITS format. Although EXSAS has been developed specifically for data analysis of the ROSAT instruments, its structural design is sufficiently general to serve equally well also data from other X-ray and XUV instruments. EXSAS analysis modules are grouped into 4 application packages dealing with Data Preparation and Instrument Correction, Spatial Analysis, Spectral Analysis and Timing Analysis. A special EXSAS header, read and updated by each application, maintains the general information transfer on the origin, the history and the parameter space of the data stored in tables and images. About 100 genuine commands (most of which offer several additional options) allow to interactively explore the functionality of the system. Up to now 40 institutes all over the world have requested the EXSAS software. Maintenance and regular updates of the software and the comprehensive documentation are provided by the ROSAT Scientific Data Center at Garching.

Zimmermann, H. U.↗

A guide to NASA's Pilot Land Data System (PLDS)

NASA's Pilot Land Data System (PLDS) is a distributed information management system designed to support NASA's land science community. The PLDS provides a wide range of services including management of information about scientific data, access to a library of scientific data, a data ordering capability, communications, connection to data analysis facilities, and electronic mail. The PLDS provides these services by offering the scientist the capability to search for and order data, and to communicate electronically with other scientists and computers. Three functions enable scientists to find what data are available and where they reside. The first two, Find data summaries and Read detailed descriptions give summary and detailed descriptions about data sets or groups of related data sets, science, projects, and institutions which archive land data. The third, gives information about specific pieces of data. This last function has two components, Search systemwide inventory and Search local inventory. The first component enables the user to find data elements (images, geological samples, transects, maps, etc.) that exist anywhere in the PLDS while the second has only information about data at the local site. The first enables the user to find pieces of data from several different data sets with the same temporal and spatial coverage and other elements common to most data sets, while the second allows the user to select a data set based on these descriptors and on those that are unique to a data set. The PLDS provides capabilities that enable electronic file transfers, intercomputer connection, and electronic mail. Both TCP/IP and DECnet protocols are supported via the NASA Science Internet (NIS). Access is also available through Telenet.

Source record↗

NASA biological and physical sciences databases: who’s the FAIRest of them all?

Conceptual models are a key part of the foundation of scientific study. Scientific data discovery and retrieval are often inaccurate and incomplete because these models are not sufficiently well-incorporated into data retrieval systems. Systems often don’t provide the necessary tools to those producing scientific data to fully and unambiguously annotate them and the result is consumers of the data cannot find them efficiently. The capability of data archives to provide these tools to link data to underlying conceptual models is one of dimensions of the recently developed “FAIR” principles (https://www.go-fair.org/fair-principles/ ), and is key to many automated processes being able to operate on these data, particularly analytics involving artificial intelligence. We used an open-source web service to measure the FAIR compliance of the three data archives operated by NASA for the biological and physical sciences: the Life Sciences Data Archive, the Physical Sciences Informatics database, and GeneLab. The service ingests references to data sets in these archives, and then executes domain-non-specific examinations of these data and metadata that test compliance to the FAIR principles. Of the 22 metrics tested, GeneLab passed 11 (50%), and PSI and LSDA each passed 7 (32%). These data were gathered using only one representative data set from each archive and we anticipate variability in results as we continue to apply these metrics to other data. A preliminary study of the failure traces for each metric suggests there is a wide range of effort and complexity in the enhancements required for each system to elevate FAIR compliance, and this is the subject of continued investigation. This information has been and will likely continue to be important information in planning these enhancements, with the goal of increased readiness of the data for automated processes.

database↗

A Web of Data Analytics Services

Cloud Computing has become the ubiquitous approach to our Big Data challenge. However, one will quickly discover that moving (a.k.a. forklifting) existing on-premise data analytics solutions to the Cloud doesn’t always translate to costing saving and performance boost. The Cloud’s elasticity, its availability, and its wide selection of computing options and selections of costing models making Cloud an attractive environment to tackle our Big Data challenge. The fact is Cloud, on its own, is not the silver bullet to our daunting challenge need for analyze and derive scientific inferences through vast collections of multi-sensor measurements. We would like to have all scientific data in one easy to access environment, but getting the world of scientific data in one analytic system is immensely difficult to achieve. This paper describes the data analytics web architecture NASA is developing by infusing instances of Integrated Data Analytics systems next to the data. The goal is to minimize unnecessary data movement through collection of data access and analytics webservices for researchers to interact with and analyze measurements without have to download data to their local computer. These services are RESTful and provisioned by the data centers with the help from subject matter and science experts. These services encapsulate the physical computing infrastructure, which could local computing cluster, on-premise or public Cloud environment.

Huang, Thomas↗

Expanding Repository Data Available For Sharing and Knowledge Discovery

Some of the hardest space biology and space health challenges require data-intensive, bioinformatic, meta-analytical, and computer-assisted research approaches. These challenges include examining interdisciplinary space life science research across experiments and across interacting spaceflight hazards (radiation, altered gravity, confinement, hostile-closed environments, distance-duration from Earth). The approaches to confront these challenges involve mining multiple datasets simultaneously from various hierarchical organizations of biological complexity, all while concurrently evaluating how experimental design factors affect endpoints of standard assays. To enable this field, it is essential that principal investigators (PIs) submit data in a structure so it can be maximally re-used. The purpose of the NASA Ames Life Sciences Data Archive (ALSDA) is to collect, curate, and make publicly available all non-human space-relevant biological data. ALSDA must also ensure data are open-access, and maximally findable, accessible, interoperable, and reusable (FAIR). The scope of ALSDA data collected and submitted by PIs include subject and study design metadata, assay metadata parameters, raw and processed assay data, assay imagery/video, and subject-experienced mission data telemetry (radiation, temperature, humidity, acoustics, vibrations, etc.). ALSDA recently integrated into a collaborative group of Open Science projects to facilitate a suite of new tools and workflows that will improve data submission, accessibility, and reusability by implementing digital data submission agreements, and adopting the data management system originally developed by NASA GeneLab. ALSDA intends to bring current biological repository data and all future collected data into this new scientific data reuse reality. This new suite of tools will enable ALSDA to deploy a science curation system using scientific assay configurations for the data submission portal. It will capture essential assay parameters according to established standards in each sub-field within biology. The submission portal expedites data collection by enhancing ease of PI data submission, providing a user interface and specificity for which data is to be submitted. Data submissions can be brought into cutting-edge informatic analysis portals to enable mining of physiological, behavioral, biochemical, and imaging datasets in conjunction with ‘omics-level datasets. As ALSDA datasets are submitted, curated, and published (e.g., micro-computed tomography, histology, pulse oximetry, serum metabolites, magnetic resonance imaging, intraocular pressure, novel object recognition, etc.), the merging together of spaceflight data along this multi-hierarchical complexity of biology will enable informatics and data-intensive approaches resulting in knowledge discoveries across missions, space hazards, and biological disciplines.

Biology↗

Expanding Repository Data Available For Sharing And Knowledge Discovery

Some of the hardest space biology and space health challenges require data-intensive, bioinformatic, meta-analytical, and computer-assisted research approaches. These challenges include examining interdisciplinary space life science research across experiments and across interacting spaceflight hazards (radiation, altered gravity, confinement, hostile-closed environments, distance-duration from Earth). The approaches to confront these challenges involve mining multiple datasets simultaneously from various hierarchical organizations of biological complexity, all while concurrently evaluating how experimental design factors affect endpoints of standard assays. To enable this field, it is essential that principal investigators (PIs) submit data in a structure so it can be maximally re-used. The purpose of the NASA Ames Life Sciences Data Archive (ALSDA) is to collect, curate, and make publicly available all non-human space-relevant biological data. ALSDA must also ensure data are open-access, and maximally findable, accessible, interoperable, and reusable (FAIR). The scope of ALSDA data collected and submitted by PIs include subject and study design metadata, assay metadata parameters, raw and processed assay data, assay imagery/video, and subject-experienced mission data telemetry (radiation, temperature, humidity, acoustics, vibrations, etc.). ALSDA recently integrated into a collaborative group of Open Science projects to facilitate a suite of new tools and workflows that will improve data submission, accessibility, and reusability by implementing digital data submission agreements, and adopting the data management system originally developed by NASA GeneLab. ALSDA intends to bring current biological repository data and all future collected data into this new scientific data reuse reality. This new suite of tools will enable ALSDA to deploy a science curation system using scientific assay configurations for the data submission portal. It will capture essential assay parameters according to established standards in each sub-field within biology. The submission portal expedites data collection by enhancing ease of PI data submission, providing a user interface and specificity for which data is to be submitted. Data submissions can be brought into cutting-edge informatic analysis portals to enable mining of physiological, behavioral, biochemical, and imaging datasets in conjunction with ‘omics-level datasets. As ALSDA datasets are submitted, curated, and published (e.g., micro-computed tomography, histology, pulse oximetry, serum metabolites, magnetic resonance imaging, intraocular pressure, novel object recognition, etc.), the merging together of spaceflight data along this multi-hierarchical complexity of biology will enable informatics and data-intensive approaches resulting in knowledge discoveries across missions, space hazards, and biological disciplines.

life science↗

Laboratory for Atmospheres 2002 Technical Highlights

How can we improve our ability to predict the weather-tomorrow, next week, and into the future? How is the Earth's climate changing? What causes such change? And what are its costs? What can the atmospheres of distant planets teach us about our own planet and its evolution? The Laboratory for Atmospheres is helping to answer these and other scientific questions about our planet and its neighbors. The Laboratory conducts a broad theoretical and experimental research program studying all aspects of the atmospheres of the Earth and other planets, including their structural, dynamical, radiative, and chemical properties, with the overarching goal to provide better understanding and to improve prediction of the Earth's climate. Vigorous research is central to NASA's exploration of the frontiers of knowledge. NASA scientists play a key role in conceiving new space missions, providing mission requirements, and carrying out research to explore the behavior of planetary systems, including, notably, the Earth's. Our Laboratory's scientists also supply outside scientists with technical assistance and scientific data to further investigations not immediately addressed by NASA itself. Laboratory scientists submit competitive research proposals with diverse scientific or technological approaches to NASA and other Federal agencies to acquire research support. The Laboratory management strives to provide a working environment that promotes creativity, competition, and openness. The Laboratory for Atmospheres is a vital participant in NASA's research program. Our Laboratory often has relatively large programs, sizable satellite missions, or observational campaigns that require the cooperative and collaborative efforts of many scientists. We ensure an appropriate balance between our scientists' responsibility for these large collaborative projects and their need for an active individual research agenda. This balance allows members of the Laboratory to continuously improve their scientific credentials. The Laboratory places high importance on promoting and measuring quality in its scientific research. We strive to assure high quality through peer-review funding processes that support approximately 90% of the work in the Laboratory. The overall quality of our scientific efforts is evaluated periodically by committees of advisors from the external scientific community, as detailed in Appendix 2 of this document. Members of the Laboratory interact with the general public to support a wide range of interests in the atmospheric sciences. Among other activities, the Laboratory raises the public's awareness of atmospheric science by presenting public lectures and demonstrations, by making scientific data available to wide audiences, by teaching, and by mentoring students and teachers. Section 6 presents details of the Laboratory's outreach activities during 2002. The Laboratory is also committed to addressing the demographic imbalances that exist today in the atmospheric and space sciences. We must address these imbalances for our field to enjoy the full benefit of all of the Nation's talent. The Laboratory makes substantial efforts to attract new scientists to the fields of atmospheric and space sciences. We strongly encourage the establishment of partnerships with Federal and state agencies that have operational responsibilities to promote the societal application of Earth sciences.

Steven E. Platnick↗

A Delay Tolerant Networking-Based Approach to a High Data Rate Architecture for Spacecraft

Historically, it has been the case that SWaP placed such severe constraints on radios that the links between spacecraft and the ground were relatively slow. This meant that the radio link was normally a significant bottleneck in returning scientific data. Over recent years, however, a combination of more efficient radio design, intelligent waveforms, and highly directed, high-frequency RF / optical systems have led to a rapid increase in the amount of data that can be pushed through radio and optical links. This has led to some cases where the radio links are capable of moving data much more quickly than the spacecraft and instruments are capable of actually generating it! In some instances, scientific data can therefore be lost not because the downlink is too slow to support the data rate, but instead because the spacecraft was not designed in a way that would let it fully utilize both the radio and the networking services available to it.The High Data Rate Architecture (HiDRA) project describes a packet-based approach to building modern, distributed spacecraft systems. It presents a means for spacecraft and other assets to participate in both present and future Delay Tolerant Networks (DTN), while simultaneously ensuring that the asset is able to fully utilize the new, high-speed links that have been seeing more widespread development and deployment in recent years. With this in mind, this paper begins with a discussion regarding HiDRA's evolution. Next, it discusses the capabilities and limitations of NASA's present DTN-enabled networks. Of particular note is the way in which principles of network design at the terrestrial level (e.g. use of programmable networks / software-defined networks, separation between data and control plane, infusion of COTS Ethernet switch chips, etc.) can all be translated into the space environment as well. After this, the paper discusses the design and implementation of a present prototype reference implementation of High-Rate DTN (HDTN), which is intended to demonstrate future high-rate networking concepts as part of a coherent demonstration on the International Space Station (ISS). The goal, of both the research and of this implementation, is to help develop a ready-made toolbox of ideas, approaches, and examples from which mission designers can draw when putting together new missions. Assuming all goes as planned, this should not only work to reduce the cost of individual mission design, but also improve the rate at which science data can be returned for mission participants to review.

Hylton, Alan↗