Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “scientific data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 289 records · Page 16

Data Sharing in Radiobiology; Towards FAIR

The value of scientific data depends on their findability, accessibility, integrability and reusability according to the FAIR principles. Together with the sustainability of data preservation and access, these principles underpin the long term benefits of scientific research. Within the domain of radiobiology we have a huge array of data types, themes and complexities which make standardisation of metadata, data structure and data integration very challenging. Moreover, it is clear that, for example, in the area of disaster preparedness, the ready discovery and availability of multiple types of data, for example on biological effects of exposure, climatology, ecology, human behavioural and attitudinal studies, is important for an integrated scientific approach. Because these data are spread over many databases, journal supplementary information resources and even the computers of the investigators, their discovery and reuse can be challenging. Despite exhortations from funding agencies and scientific institutions over the past two decades there is still a serious deficit in the willingness and in some cases the ability of investigators to share data, and although much may not be formally „Public domain“, information about the existence of the data, their metadata, and how to obtain them should always be available. We report the progress of work on three databases, the STORE and the NASA GeneLab and LSDA repositories to leverage the Radiation Biology Ontology (RBO), a structured terminology for metadata that can be used by all radiation biology-relevant databases to unite federated and automated data searches across multiple databases, for example using web services, and through semantic web technologies supporting data discovery. The initial primary use-cases for RBO were archiving data in the STORE database (https://www.storedb.org/), the repository used for the RadoNorm and Pianoforte Projects among others, and in the NASA Open Science Data Repository (https://osdr.nasa.gov/bio). The scope of radiobiology research ranges from basic physics to radiation oncology to sociolegal studies; no existing ontology had the necessary breadth or depth to fulfill this need. In addition, a formal ontology has the advantage of being usable for machine learning and, importantly, for tasks like data integration, knowledge extraction from the scientific literature and for query extension and data classification. Standardisation of metadata is one of the primary objectives of the FAIR principles for open data; RBO is an important landmark for FAIR-compliant radiation biology data sharing. The RBO is developed using the open-source tools of GitHub and the OBO Foundry-led Ontology Development Kit, and published through GitHub and the NIH/NCBI BioPortal website. This initial phase of concept modeling has yielded an ontology that has more than 300 declared concepts, with more than 3500 additional concepts imported from other OBO Foundry ontologies with relevance to radiation biology (for example, concepts from the ISO standard Basic Formal Ontology, the Environment Ontology and the Gene Ontology). We welcome input into the development of RBO and encourage its adoption.

ontologies↗

The Planetary Data System - A Case Study in the Development and Management of Meta-Data for a Scientific Digital Library

The Planetary Data System (PDS) is an active science data archive managed by scientists for NASA's planetary science community. With the advent of the World Wide Web the majority of the archive has been placed on-line as a science digital libraty for access by scientists, the educational community, and the general public.

Data System Meta-data scientific digital library↗

Data Sharing in Radiation Biology: Towards FAIR

The value of scientific data depends on their findability, accessibility, integrability and reusability according to the FAIR principles. Together with the sustainability of data preservation and access, these principles underpin the long term benefits of scientific research. Within the domain of radiobiology we have a huge array of data types, themes and complexities which make standardisation of metadata, data structure and data integration very challenging. Moreover, it is clear that, for example, in the area of disaster preparedness, the ready discovery and availability of multiple types of data, for example on biological effects of exposure, climatology, ecology, human behavioural and attitudinal studies, is important for an integrated scientific approach. Because these data are spread over many databases, journal supplementary information resources and even the computers of the investigators, their discovery and reuse can be challenging. Despite exhortations from funding agencies and scientific institutions over the past two decades there is still a serious deficit in the willingness and in some cases the ability of investigators to share data, and although much may not be formally "Public domain“, information about the existence of the data, their metadata, and how to obtain them should always be available. We report the progress of work on three databases, the STORE and the NASA GeneLab and LSDA repositories to leverage the Radiation Biology Ontology (RBO), a structured terminology for metadata that can be used by all radiation biology-relevant databases to unite federated and automated data searches across multiple databases, for example using web services, and through semantic web technologies supporting data discovery. The initial primary use-cases for RBO were archiving data in the STORE database (https://www.storedb.org/), the repository used for the RadoNorm and Pianoforte Projects among others, and in the NASA Open Science Data Repository (https://osdr.nasa.gov/bio). The scope of radiobiology research ranges from basic physics to radiation oncology to sociolegal studies; no existing ontology had the necessary breadth or depth to fulfill this need. In addition, a formal ontology has the advantage of being usable for machine learning and, importantly, for tasks like data integration, knowledge extraction from the scientific literature and for query extension and data classification. Standardisation of metadata is one of the primary objectives of the FAIR principles for open data; RBO is an important landmark for FAIR-compliant radiation biology data sharing. The RBO is developed using the open-source tools of GitHub and the OBO Foundry-led Ontology Development Kit, and published through GitHub and the NIH/NCBI BioPortal website. This initial phase of concept modeling has yielded an ontology that has more than 300 declared concepts, with more than 3500 additional concepts imported from other OBO Foundry ontologies with relevance to radiation biology (for example, concepts from the ISO standard Basic Formal Ontology, the Environment Ontology and the Gene Ontology). We welcome input into the development of RBO and encourage its adoption.

ontologies↗

Open Source Scalable Data Services and Data Fusion for Biological and Environmental Sciences (SBIR Phase I Final Scientific/ Technical Report)

The overarching goal of the project is to develop an integrated open-source scientific data management system (Apache V2 license), ResonantEco, that meets the need of biological and environmental researchers and developers for data management, curation, and data processing for analyses with a wide range of scale and complexity. ResonantEco will provide web enabled data services with features such as unified data interfaces and federated views of data and metadata for heterogeneous data sources with an interactive web client for data exploration. Our use of the term fusion is taken from geospatial (GIS) domain where data fusion is often synonymous with data integration. In particular, data integration in ResonantEco involves combining data residing in different sources and providing users with a unified view of them.

99 GENERAL AND MISCELLANEOUS↗

Visualization tools for the processing of airglow data from RAIDS

In anticipation of large data sets associated with a number of atmospheric imaging instruments being prepared for long term global coverage, NRL is developing graphical interfaces for all aspects of the program. For the first of these projects, RAIDS (the Remote Atmospheric and Ionospheric Detection System), a graphical approach to data handling, visualization, and analysis is envisioned and will set the stage for the satellites that follow. An overall system of hardware and a set of software 'tools,' that will allow for both the routine handling of all data and the analysis of large data sets assembled by scientists and instrument engineers, are currently being developed. The software for standard processing and visualization of instrument data is independent of computer platform and will allow for easy adaptation from one experiment to another. The processing will produce data sets that have similar characteristics, allowing for easy comparison of data obtained under similar circumstances. The visualization of both the engineering and scientific data is an important part of the system. By creating graphical environments for engineering evaluations and for scientific analysis data sets can be viewed and analyzed rapidly. This rapid analysis of data will contribute towards a greater portion of the RAIDS data being utilized.

Miller, Gordon J.↗

Science Goal Monitor: Science Goal Driven Automation for NASA Missions

Infusion of automation technologies into NASA s future missions will be essential because of the need to: (1) effectively handle an exponentially increasing volume of scientific data, (2) successfully meet dynamic, opportunistic scientific goals and objectives, and (3) substantially reduce mission operations staff and costs. While much effort has gone into automating routine spacecraft operations to reduce human workload and hence costs, applying intelligent automation to the science side, i.e., science data acquisition, data analysis and reactions to that data analysis in a timely and still scientifically valid manner, has been relatively under-emphasized. In order to introduce science driven automation in missions, we must be able to: capture and interpret the science goals of observing programs, represent those goals in machine interpretable language; and allow spacecrafts onboard systems to autonomously react to the scientist's goals. In short, we must teach our platforms to dynamically understand, recognize, and react to the scientists goals. The Science Goal Monitor (SGM) project at NASA Goddard Space Flight Center is a prototype software tool being developed to determine the best strategies for implementing science goal driven automation in missions. The tools being developed in SGM improve the ability to monitor and react to the changing status of scientific events. The SGM system enables scientists to specify what to look for and how to react in descriptive rather than technical terms. The system monitors streams of science data to identify occurrences of key events previously specified by the scientist. When an event occurs, the system autonomously coordinates the execution of the scientist s desired reactions. Through SGM, we will improve om understanding about the capabilities needed onboard for success, develop metrics to understand the potential increase in science returns, and develop an operational prototype so that the perceived risks associated with increased use of automation can be reduced.

Koratkar, Anuradha↗

Efficient Asynchronous I/O with Request Merging

With the advancement of exascale computing, the amount of scientific data is increasing day by day. Efficient data access is necessary for scientific discoveries. Unfortunately, the I/O performance is not improved, like the CPU and network speed. So, I/O operations take longer time than data generation or analysis. Asynchronous I/O has been proposed to extenuate the I/O bottleneck by overlapping I/O and computation time. However, multiple small write operations can diminish the benefits of asynchronous I/O, as the I/O time becomes significantly longer than the compute time, with little time to overlap with. To overcome these issues, we present an optimization technique to merge small contiguous write operations. We integrated our solution into the HDF5 asynchronous I/O VOL connector and demonstrated the effectiveness of merging HDF5 write operations automatically and transparently without requiring any code change from the application.

Chowdhury, Kamal Hossain↗

Smart Data Node in the Sky

A document discusses the physical and engineering principles affecting the design of the Smart Data Node in the Sky (SDNITS) -- a proposed Earth-orbiting satellite for relaying scientific data from other Earth-orbiting satellites to one or more ground station(s). The document characterizes the problem of designing the telecommunication architecture of the SDNITS as consisting of two main parts: (1) finding the most advantageous orbit for the SDNITS to gather data from the scientific satellites and relay the data to the ground, taking account of such factors as visibility and range; and (2) choosing a telecommunication architecture appropriate for the intended relay function.

Lansing, Faiza↗

Stakeholder analysis for designing an urban air quality data governance ecosystem in smart cities

Cities, the world over, are fuelling economic growth. At the same time, rapid urbanization is a root cause of serious environmental damage. Recent WHO global air pollution guidelines highlight air pollution as a critical environmental threat along with climate change. To address these threats, smart cities and clean air programs are on a rise. In smart cities, data and Information and Communication Technologies (ICT) are major drivers of city transformations. The 4th Industrial Revolution (4IR) technologies such as the Internet of Things (IoT), big data, artificial intelligence (AI), and cloud computing have the potential to accelerate these transformations toward urban resilience. However, the success of smart cities and clean air programs depends on cohesive multi-sector stakeholder contributions. This study conducted interdisciplinary participative stakeholder analysis to understand the data, and sectorial challenges, to outline the technological opportunities to facilitate clean air programs in Indian smart cities. The research highlights gaps due to siloed stakeholder operations, lack of data calibration, non-alignment of smart city and air quality management services, non-availability of health exposure data, and difficulty in translating scientific data into implementable actions. Stakeholders expressed potential ‘fit for the purpose’ use of IoT devices, satellites, smartphones, and mobility data augmented by AI methods in bridging these gaps. In conclusion, the analysis points toward a need to develop an easily accessible and ubiquitous urban data governance ecosystem enabling seamless cross-sector data exchanges to build trusting relationships among the stakeholders across the air quality management value chain.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

To Derive or Not to Derive: I/O Libraries Take Charge of Derived Quantities Computation

The ever-increasing volume of data produced by HPC simulations necessitates scalable methods for data exploration and knowledge extraction. Scientific data analysis often involves complex queries across distributed datasets, requiring manipulation of multiple primary variables and generating derived data that needs to be handled efficiently, creating challenges for applications that need to parse many large datasets. Relying on individual applications to handle all intermediate data generally leads to redundant computations across studies and unnecessary data transfers. In this paper, we investigate the performance of different approaches where applications define derived variables as quantities of interest (QoIs) and offload the computation and transfer of these QoIs to the I/O library. This significantly reduces redundancy and optimizes data movement across the distributed storage and processing infrastructure by allowing control over when and where derived variables are computed. We present a detailed analysis of the performance-storage trade-offs associated with different solutions and showcase results for our study on two large-scale datasets created from climate and combustion simulations.

Gainaru, Ana↗

The data array, a tool to interface the user to a large data base

Aspects of the processing of spacecraft data is considered. Use of the data array in a large address space as an intermediate form in data processing for a large scientific data base is advocated. Techniques for efficient indexing in data arrays are reviewed and the data array method for mapping an arbitrary structure onto linear address space is shown. A compromise between the two forms is given. The impact of the data array on the user interface are considered along with implementation.

Foster, G. H.↗

Pilot climate data system user's guide

Instructions for using the Pilot Climate Data System (PCDS), an interactive, scientific data management system for locating, obtaining, manipulating, and displaying climate-research data are presented. The PCDS currently provides this supoort for approximately twenty data sets. Figures that illustrate the terminal displays which a user sees when he/she runs the PCDS and some examples of the output from this system are included. The capabilities which are described in detail allow a user to perform the following: (1) obtain comprehensive descriptions of a number of climate parameter data sets and the associated sensor measurements from which they were derived; (2) obtain detailed information about the temporal coverage and data volume of data sets which are readily accessible via the PCDS; (3) extract portions of a data set using criteria such as time range and geographic location, and output the data to tape, user terminal, system printer, or online disk files in a special data-set-independent format; (4) access and manipulate the data in these data-set-independent files, performing such functions as combining the data, subsetting the data, and averaging the data; and (5) create various graphical representations of the data stored in the data-set-independent files.

Reph, M. G.↗

Measurement of the Earth-Observer-1 Satellite X-Band Phased Array

The recent launch and successful orbiting of the EO-1 Satellite has provided an opportunity to validate the performance of a newly developed X-Band transmit-only phased array aboard the satellite. This paper will compare results of planar near-field testing before and after spacecraft installation as well as on-orbit pattern characterization. The transmit-only array is used as a high data rate antenna for relaying scientific data from the satellite to earth stations. The antenna contains distributed solid-state amplifiers behind each antenna element that cannot be monitored except for radiation pattern measurements. A unique portable planar near-field scanner allows both radiation pattern measurements and also diagnostics of array aperture distribution before and after environmental testing over the ground-integration and prelaunch testing of the satellite. The antenna beam scanning software was confirmed from actual pattern measurements of the scanned beam positions during the spacecraft assembly testing. The scanned radiation patterns on-orbit were compared to the near-field patterns made before launch to confirm the antenna performance. The near-field measurement scanner has provided a versatile testing method for satellite high gain data-link antennas.

Perko, Kenneth↗

Flight Hardware Development and Research at MSFC for Optimizing Success on the International Space Station

To optimize biological crystallization success in microgravity in-house personnel at the MSFC are working on the development of innovative flight hardware such as Delta-L and the Iterative Biological Crystallization (IBC) apparatus as well as troubleshooting the performance of existing hardware. Delta-L will provide a diagnostic hardware to examine the relationship between crystal growth characteristics and crystal quality improvement in microgravity. IBC is a new hardware being designed to allow iteration of crystal growth experiments in microgravity using innovative lab on a chip technology. While being built to obtain scientific data of benefit to the scientific community, the design methods involved in the development of these hardware have directly benefited other groups within NASA and keep NASA at the forefront of innovation.

Source record↗

Data-Driven Art

In Fall 2023, Katie Baldwin (UAH) and Helen Parache (NASA) will follow up on their pilot activity from the spring that focused on collaboration between the arts and sciences at the UAH Art Department. Ms. Parache will present on open access data and artists that incorporate scientific data in their work, e.g. Tali Weinberg and Sarah Bryant (University of Alabama). Ms. Baldwin will demonstrate printmaking and mark making techniques. The students in Ms. Baldwin’s Book Arts class will participate in a series of generative activities and engage with data to develop content. The focus on the Art Department stems from the importance of Art as a cultural pillar. Tapping into the communication and social relevance of art could be an avenue to pursue Environmental Justice goals of interest to NASA. A creative perspective on data can bring about creative questions and solutions. The workshop incorporates changes based on feedback from the spring workshop.

data science↗

Performance and Reliability Assessment of the U.S. Department of Energy Atmospheric Radiation Measurement (ARM) Data Advisor (ADA)

The Atmospheric Radiation Measurement (ARM) User Facility provides one of the world's largest openly accessible repositories of atmospheric observations through the ARM Data Discovery platform. Although the repository contains more than three decades of measurements collected from permanent observatories, mobile facilities, aircraft campaigns, and field experiments, identifying appropriate datasets can be challenging, particularly for new users unfamiliar with ARM instrumentation and datastream organization. To improve data accessibility, the ARM Data Center developed the ARM Data Advisor (ADA), an artificial intelligence-powered assistant designed to facilitate scientific data discovery, dataset interpretation, and user guidance. This report evaluates ADA's performance as a domain-specific scientific assistant using realistic atmospheric science workflows. The evaluation examines five key capabilities: data retrieval and curation efficiency, hallucination resistance, scientific reasoning, response to ambiguous queries, and content retention and session continuity. Representative prompts were developed to simulate typical interactions between researchers and the ARM Data Discovery platform, and ADA's responses were assessed for retrieval completeness, scientific accuracy, consistency, and practical usefulness. In these representative tests, ADA reduced the complexity of discovering and accessing ARM datasets by recommending appropriate datastreams, explaining instrumentation, interpreting metadata, and assisting with data processing workflows. ADA also exhibits strong domain knowledge of atmospheric science terminology and generally resists hallucination by acknowledging unavailable datasets and requesting clarification when appropriate. Overall, the results indicate that ADA represents a promising advancement in scientific data discovery within the ARM User Facility and has considerable potential to improve researcher productivity, particularly for new users and interdisciplinary scientists seeking efficient access to ARM observations.

Salvador, Christian [ORNL] (ORCID:0000000283287777↗

Usable Data Abstractions for Next-Generation Scientific Workflows

Data- and computationally-intensive scientific research, such as numerical simulations and inversions or the training of large neural networks in machine learning applications, that are well suited for HPC environments also often require expert insight and evaluation throughout the computation which can be greatly facilitated with the use of interactive computing tools, such as those in the Jupyter ecosystem. HPC workflows and interactive workflows are typically treated as orthogonal, however, the next generation of research will require both. The first challenge we face in this project is thus designing the right level of abstractions to allow interactive capabilities in the JupyterLab environment to allow the working scientist to flexibly explore and query their data at multiple levels, with a minimal amount of customization required of the underlying optimized codes. In addition to these questions regarding the high-level representation of data for interactive use in HPC, we tackled two additional issues that are part of the entire lifecycle of research and that become particularly acute in HPC contexts: how to improve the experience of interfacing with the HPC system's scheduling environment for a scientist focused on exploratory questions, and how can that scientist then best share the results of their work with others in a self-contained, reproducible manner.

97 MATHEMATICS AND COMPUTING↗

Science Goal Driven Automation for NASA Missions: The Science Goal Monitor

Infusion of automation technologies into NASA s future missions will be essential not only to achieve substantial reduction in mission operations staff and costs, but also in order to both effectively handle an exponentially increasing volume of scientific data and to successfully meet dynamic, opportunistic scientific goals and objectives. Current spacecraft operations cannot respond to science driven events, such as intrinsically variable or short-lived phenomena in a timely manner. For such investigations, we must teach our platforms to dynamically understand, recognize, and react to the scientists goals. While much effort has gone into automating routine spacecraft operations to reduce human workload and hence costs, applying intelligent automation to the science side, i.e., science data acquisition, data analysis and reactions to that data analysis in a timely and still scientifically valid manner, has been relatively under-emphasized.

Korathkar, Anuradha↗