Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Modeling workflow”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

179 records · Page 10

TPSAS-NF1676L-33992-DND

The CERES Science Team integrates and fuses observations from 6 CERES instruments aboard the Terra, Aqua, S-NPP, and NOAA-20 missions with data from more than 20 other unique data sources. Following the November 2017 launch of CERES Flight Model 6 (FM6) onboard NOAA20, CERES has now amassed over 80 instrument-years of valuable Earth radiation budget data. The rapidly growing volume of CERES data coupled with the introduction of new data products alongside improvements to existing science algorithms fosters the requirement for faster, more flexible, and scalable data production and orchestration. New virtualized, cloud-centric compute hardware hosted by the NASA Langley Research Center’s (LaRC) Atmospheric Sciences Data Center (ASDC) provides an ideal environment for these ever-increasing data production demands for CERES. This poster discusses updates to the implementation of the CERES Data Management Team’s (DMT) CERES AuTomAted job Loading sYSTem (CATALYST), a custom data processing workflow engine for CERES, to use on-demand computing resources to perform automated CERES data production processing in a Linux-based container environment. Linux containers provide CERES the flexibility to build multiple production environments in containers tailored for specific workloads and allow effortless provisioning of resources based on the CERES Science Team’s data production requirements.

Thomas N. Hillyer↗

Harmonized Sentinel-1 SAR Global River Geometry and Inundation Database

Satellite-based observations on river geometries are sporadic in time, space, or both. Most satellite-based surface water maps, river widths, water surface elevations (WSE), slopes, and bathymetry are asynchronized in time and space. The current configuration of satellites such as Sentinel-6 measured the WSE but is missing the river width, slopes, and depths. To advance hydrological sciences research, there is a need to produce a harmonized time series of river geometry data of non-SWOT satellites in partnership with the upcoming SWOT mission. The SWOT satellite will measure river width, height, and slope but missing river depth measurements in space and time. Further, none of these current satellites measure the WSE, river width, and slopes synchronously. In this work, we use the Sentinel-1 SAR satellite data archive from 2015 to the present to create a global river width and surface water database at the reach scale. A modified version of the Sentinel SAR surface water classification algorithm from ASF is used to quantify the surface water extent on the stream approximately every six days (at the equator) at 10m spatial resolution globally. This 10m water mask is fed into a workflow to quantify the river widths, surface water inundations, slopes, and synthetic bathymetry in SWORD (SWOT River Database) stream networks. A Satellite HAND is used to address the cloud obscured surface water observations using a trained machine learning algorithm. We use WSE derived from the Global Water Monitor from NASA GSFC, Hydroweb from LEGOS, and ICESat-2 to harmonize the WSE observation. And Landsat-8/9 and Sentinel-2 water observations to fill the gaps in the Sentinel-1 SAR database. We use Congo River Basin as a test case where we have more than 500 radar altimetry-based WSE, continuous series of Sentinel-1, ICESat-2, Landsat-8/9, and Sentinel-2 observations. A Congo River hydrologic model is used to generate the streamflow discharge. The satellite observed river reaches are assimilated with the stream flows computed by the routing models. And the downstream reaches in the river network without satellite observations get optimized for discharge/river geometry at each observation cycle. Our final product is a harmonized river geometry dataset (reach's water extent, WSE, slope, synthetic bathymetry) for Congo Basin's SWORD reaches.

Chandana Gangodagamage↗

Extending the Reach of IGSN Beyond Earth: Implementing IGSN Registration to Link Nasa's Apollo Lunar Samples and Their Data

The rock and soil samples returned from the Apollo missions from 1969-72 have supported 46 years of research leading to advances in our understanding of the formation and evolution of the inner Solar System. NASA has been engaged in several initiatives that aim to restore, digitize, and make available to the public existing published and unpublished research data for the Apollo samples. One of these initiatives is a collaboration with IEDA (Interdisciplinary Earth Data Alliance) to develop MoonDB, a lunar geochemical database modeled after PetDB (Petrological Database of the Ocean Floor). In support of this initiative, NASA has adopted the use of IGSN (International Geo Sample Number) to generate persistent, unique identifiers for lunar samples that scientists can use when publishing research data. To facilitate the IGSN registration of the original 2,200 samples and over 120,000 subdivided samples, NASA has developed an application that retrieves sample metadata from the Lunar Curation Database and uses the SESAR API to automate the generation of IGSNs and registration of samples into SESAR (System for Earth Sample Registration). This presentation will describe the work done by NASA to map existing sample metadata to the IGSN metadata and integrate the IGSN registration process into the sample curation workflow, the lessons learned from this effort, and how this work can be extended in the future to help deal with the registration of large numbers of samples.

Todd, Nancy S.↗

Crew Health and Performance Integrated Data System Platform Project Updates

Future human exploration missions introduce a new paradigm as crews move further from the resupply and near real-time ground support typical of Low Earth Orbit missions today. Without immediate support from ground-based personnel, exploration crews will be more reliant on inflight data and technology to respond to emergencies and anomalies. Today, in-flight data is often siloed, unsynchronized, and largely inaccessible in real time. Many data sets require manual entry and/or data transfer between vehicles and the ground. These issues contribute to risks in supporting crew autonomy for future exploration missions. An integrated data system platform is needed to mitigate these risks by supporting a new generation of technologies and employing advanced analytical and predictive modeling techniques to enable crew autonomy for future exploration missions. The Crew Health and Performance Integrated Data System Platform (CHP-IDSP) project is laying a foundation for future in-flight informatics by providing a back-end architecture for collecting, storing, and integrating multiple sources of data generated by and around the crew. This cohesive integration point will streamline the management of CHP data (e.g., environmental, exercise, medical, sleep, performance, etc.) and facilitate situation awareness and decision support required by the crew and remote support of exploration missions. This presentation will describe the ongoing development effort of the path-to-flight CHP-IDSP software and the demonstration of its core capabilities. This includes a brief history of the project, the human-centered process used to identify data needs and workflows feeding the development of scenarios and requirements, and current subsystem development status. Current integrations, including the Chiron exploration electronic health record application, will be discussed. Future work includes collaboration with additional CHP domains and a flight technology demonstration.

Data integration↗

Crew Health and Performance Integrated Data Service Platform (CHP-IDSP): Project Updates

Future human exploration missions introduce a new paradigm as crews move further from the resupply and near real-time ground support typical of Low Earth Orbit missions today. Without immediate support from ground-based personnel, exploration crews will be more reliant on inflight data and technology to respond to emergencies and anomalies. Today, in-flight data is often siloed, unsynchronized, and largely inaccessible in real time. Many data sets require manual entry and/or data transfer between vehicles and the ground. These issues contribute to risks in supporting crew autonomy for future exploration missions. An integrated data services platform is needed to mitigate these risks by supporting a new generation of technologies and employing advanced analytical and predictive modeling techniques to enable crew autonomy for future exploration missions. The Crew Health and Performance Integrated Data System Platform (CHP-IDSP) project is laying a foundation for future in-flight informatics by providing a back-end architecture for collecting, storing, and integrating multiple sources of data generated by and around the crew. This cohesive integration point will streamline the management of CHP data (e.g., environmental, exercise, medical, sleep, performance, etc.) and facilitate situation awareness and decision support required by the crew and remote support of exploration missions. This presentation will describe the ongoing development effort of the path-to-flight CHP-IDSP software and the demonstration of its core capabilities. This includes a brief history of the project, the human-centered process used to identify data needs and workflows feeding the development of scenarios and requirements, and current subsystem development status. Current integrations, including the Chiron exploration electronic health record application, will be discussed. Future work includes collaboration with additional CHP domains and a flight technology demonstration.

Software↗

Integrating Cloud-Based Workflows in Continental-Scale Cropland Extent Classification

Accurate information on cropland spatial distribution is required for global-scale assessments and agricultural land use policies. Cloud computing platforms such as Google Earth Engine (GEE) provide unprecedented opportunities for large-scale classifications of Landsat data. We developed a novel method to fuse pixel-based random forest classification of continental-scale Landsat data on GEE and an object-based segmentation approach known as recursive hierarchical segmentation (RHSeg). Using our fusion method, we produced a continental-scale cropland extent map for North America at 30m spatial resolution for the nominal year 2010. The total cropland area for North America was estimated at 275.18 million hectares (Mha). The overall accuracies of the map are>90% across the continent. This map also compares well with the United States Department of Agriculture (USDA) cropland data layer (CDL), Agriculture and Agri-food Canada (AAFC) annual crop inventory (ACI), and the Mexican government agency Servicio de Informacion Agroalimentaria y Pesquera (SIAP)'s agricultural boundaries. Furthermore, our map compared well with sub-country statistics including state-wise and county-wise cropland statistics in regression models resulting in R2 > 0.84. This key contribution paves the way for more detailed products such as crop intensity, crop type, and crop irrigation, and provides a method for creating high-resolution cropland extent maps for other countries where spatial information about croplands are not as prevalent.

Massey, Richard↗

GeneLab: A Systems Biology Platform for Omics Analysis

NASA's GeneLab includes an open-access repository of some 200+ omics datasets generated by biological experiments relevant to spaceflight (including simulated cosmic radiation and microgravity). In order to maximize the intelligibility of these data, particularly for users with limited bioinformatics knowledge, GeneLab is now transforming the data in the repository into actual biological and physiological knowledge of the genetic and proteomic signatures found in these samples. This processed data is being derived by establishing standard data analysis workflows vetted by 114 scientists who are members of the four GeneLab Analysis Working Groups (Animal AWG, Plant AWG, Microbe AWG, Multi-Omics AWG). AWG members from institutes spanning the U.S. and four other countries participate on a voluntary basis. The AWGs meet monthly to discuss data mining, compare results and interpretations, and test forthcoming releases of the GeneLab Data Systems (GLDS). GLDS version 3.0 has been available to the general public since October 1st 2018, and has been providing a professional state-of-the-art bioinformatics platform for everyone in the space biology community to upload their data into a space biology omics data commons, to process their data with vetted standard workflows and to compare to existing analyses. The user interface for the platform is being designed to be accessible to a broad variety of users including those with limited bioinformatics experience, including high school and college students who can use it to learn about omics data analysis and space biology. As such, Genelab will constitute a powerful general public outreach capability of NASA and the Space Biology community at large. Data mining of the GeneLab database by the AWG has already started generating very interesting findings, including reports linking specific spaceflight conditions such as radiation, microgravity or carbon dioxide levels to molecular changes seen across various species. In this presentation, we will report on the current and future objectives for GeneLab, and review recent studies reported by the various AWGs relating molecular changes observed in various animal models and tissue with microgravity, radiation, circadian rhythm, hydration and carbon dioxide conditions.

Omics↗

GeneLab: A Systems Biology Platform for Omics Analysis: Disseminate and Reuse Data, Tools, and Samples Post-Project

NASA's GeneLab includes an open-access repository of some 200 plus omics datasets generated by biological experiments relevant to spaceflight (including simulated cosmic radiation and microgravity). In order to maximize the intelligibility of these data, particularly for users with limited bioinformatics knowledge, GeneLab is now transforming the data in the repository into actual biological and physiological knowledge of the genetic and proteomic signatures found in these samples. This processed data is being derived by establishing standard data analysis workflows vetted by 114 scientists who are members of the four GeneLab Analysis Working Groups (Animal AWG, Plant AWG, Microbe AWG, Multi-Omics AWG). AWG members from institutes spanning the U.S. and four other countries participate on a voluntary basis. The AWGs meet monthly to discuss data mining, compare results and interpretations, and test forthcoming releases of the GeneLab Data Systems (GLDS). GLDS version 3.0 has been available to the general public since October 1st 2018, and has been providing a professional state-of-the-art bioinformatics platform for everyone in the space biology community to upload their data into a space biology omics data commons, to process their data with vetted standard workflows and to compare to existing analyses. The user interface for the platform is being designed to be accessible to a broad variety of users including those with limited bioinformatics experience, including high school and college students who can use it to learn about omics data analysis and space biology. As such, Genelab will constitute a powerful general public outreach capability of NASA and the Space Biology community at large. Data mining of the GeneLab database by the AWG has already started generating very interesting findings, including reports linking specific spaceflight conditions such as radiation, microgravity or carbon dioxide levels to molecular changes seen across various species. In this presentation, we will report on the current and future objectives for GeneLab, and review recent studies reported by the various AWGs relating molecular changes observed in various animal models and tissue with microgravity, radiation, circadian rhythm, hydration and carbon dioxide conditions.

Omics↗

NASA GRC ICME Schema for Materials Data Management: An Executive Summary

Integrated Computational Materials Engineering (ICME) has received a growing emphasis in attention due its potential impact on rapid material design, reduction in cost and time to market for new applications, and the promise of ‘fit-for-purpose’ materials coupled with recent advances in high performance computing and material characterization tools. However, for an organization to implement ICME practices for material discovery and design, a series of both technical and cultural challenges must be overcome to foster an environment that enables efficient, traceable, and predictive multiscale simulations of material behavior to enable virtual design of materials. In 2016, NASA sponsored a 2040 Vision study to define the potential 25-year future state required for integrated multiscale modeling of materials and systems to improve both the associated time and cost for aerospace and aeronautical innovation. The study envisions a cyber-physical-social ecosystem of experimentally validated computational models, tools, and techniques, along with the associated digital tapestry, that can enable rapid, optimized, ‘fit-for-purpose’ design of materials, components, and systems. A key requirement for such an ecosystem is the development of a robust information management system for materials across their full lifecycle, including material pedigree, experimental (real) and virtual (simulation) data, developed material models, and the implementation of models in engineering applications, such that process-structure-property-performance relationships can be established, thereby enabling the virtual design and optimization of materials. Such an information management system must be able to effectively capture: i) material information at each length scale; ii) test data and analysis; iii) associated material models; and iv) material and model deployment in engineering applications. These systems must also provide traceability between experimental and virtual representations of the material to ensure, when appropriate, the material digital twin is maintained. Additionally, this robust material information management system must be able to seamlessly connect with both commercial and an organization’s in-house software tools, be they analysis tools, other material databases, product lifecycle management (PLM) or simulation data management (SDM) tools, etc., such that automation of the design and analysis of a material across multiple length scales is possible. In this paper, an executive summary of the NASA GRC ICME Schema for materials information management is presented. The database best practices and schema design philosophy specifically for ICME materials data management and an overview description of each element in the schema is given, along with its associated role in an ICME workflow. Additionally, auxiliary tools that interact with the database and provide judicious automation with regards to importing, exporting, and analyzing materials data are presented. Such tools are critical to an ICME ecosystem, not only for their role in enabling optimization, but also in relieving users of tedious manual tasks, thus helping to promote adoption and combat the cultural challenges organizations face in enabling ICME.

Materials↗

Grid Enabled Geospatial Catalogue Web Service

Geospatial Catalogue Web Service is a vital service for sharing and interoperating volumes of distributed heterogeneous geospatial resources, such as data, services, applications, and their replicas over the web. Based on the Grid technology and the Open Geospatial Consortium (0GC) s Catalogue Service - Web Information Model, this paper proposes a new information model for Geospatial Catalogue Web Service, named as GCWS which can securely provides Grid-based publishing, managing and querying geospatial data and services, and the transparent access to the replica data and related services under the Grid environment. This information model integrates the information model of the Grid Replica Location Service (RLS)/Monitoring & Discovery Service (MDS) with the information model of OGC Catalogue Service (CSW), and refers to the geospatial data metadata standards from IS0 19115, FGDC and NASA EOS Core System and service metadata standards from IS0 191 19 to extend itself for expressing geospatial resources. Using GCWS, any valid geospatial user, who belongs to an authorized Virtual Organization (VO), can securely publish and manage geospatial resources, especially query on-demand data in the virtual community and get back it through the data-related services which provide functions such as subsetting, reformatting, reprojection etc. This work facilitates the geospatial resources sharing and interoperating under the Grid environment, and implements geospatial resources Grid enabled and Grid technologies geospatial enabled. It 2!so makes researcher to focus on science, 2nd not cn issues with computing ability, data locztic~, processir,g and management. GCWS also is a key component for workflow-based virtual geospatial data producing.

Chen, Ai-Jun↗

Climate Analytics as a Service

Exascale computing, big data, and cloud computing are driving the evolution of large-scale information systems toward a model of data-proximal analysis. In response, we are developing a concept of climate analytics as a service (CAaaS) that represents a convergence of data analytics and archive management. With this approach, high-performance compute-storage implemented as an analytic system is part of a dynamic archive comprising both static and computationally realized objects. It is a system whose capabilities are framed as behaviors over a static data collection, but where queries cause results to be created, not found and retrieved. Those results can be the product of a complex analysis, but, importantly, they also can be tailored responses to the simplest of requests. NASA's MERRA Analytic Service and associated Climate Data Services API provide a real-world example of climate analytics delivered as a service in this way. Our experiences reveal several advantages to this approach, not the least of which is orders-of-magnitude time reduction in the data assembly task common to many scientific workflows.

big data↗

Computational Fluid Dynamics Methods Used in the Development of the Space Launch System Liftoff and Transition Lineloads Databases

The objective of this paper is to document the reasoning and trade studies that supported the selection of appropriate tools for constructing aerodynamic lineload databases for the Liftoff and Transition phases of flight for launch vehicles. These decisions were made amid the maturation of an evolving workflow for generating databases on variants of the Space Launch System launch vehicle, with most being based on results from brief developmental studies performed in response to specific, unforeseen challenges that were encountered in analyzing a given configuration. This report is intended to provide a summary of the results and the decision-making processes chronologically over the design cycles of various configurations, starting with isolated free-air bodies for the Block 1 Crew, then the Block 1B Crew and Cargo configurations, and most recently the Block 1B Crew configuration in proximity to the launch tower. The results from these analyses led to the selection of the CREATE-AV Kestrel flowsolver for simulating these problems. The need to accurately capture the expected leeward-wake flow field characteristics required the use of Delayed Detached Eddy Simulation (DDES) method, for which the vorticity magnitude was employed as the solution Adaptive Mesh Refinement (AMR) function over the off-body Cartesian grid region. In addition, the Spalart-Allmaras (SA) model is used to account for the flow turbulence effects.

Ratnayake, Nalin A.↗

ncompare: A Python Package for Comparing netCDF Structures

Earth science researchers and data engineers have a common problem: they often need to compare data files to see what is different between them. A lot of time is spent developing code to test differences. When it comes to comparing multidimensional data file formats like netCDFs (Network Common Data Form), this is particularly challenging and time-consuming, since there is frequently a need to evaluate the differences between dimension sizes, variable structures, and variable attributes, especially for regression testing. Since netCDFs are widely used in Earth science — with climate models, oceanographic or atmospheric reanalyses, and observational data — improved means of evaluating netCDF files can help enable a wide range of applications. We have developed a reusable open source approach through `ncompare`, which is a Python package for comparing netCDF structures [[https://github.com/nasa/ncompare]]. The `ncompare` tool compares the structure of two Network Common Data Form (NetCDF) files at the command line. It facilitates rapid comparisons by generating a formatted display of the matching and non-matching groups, variables, and associated metadata between two NetCDF datasets. The user has the option to colorize the terminal output for ease of viewing, and `ncompare` can optionally save comparison reports in text, comma-separated value (CSV), and/or Microsoft Excel formats. Despite the availability of tools (such as ncmpidiff or nccmp) that compare the values of variables, there was not previously a readily available, Python-based tool for rapid visual comparisons of group and variable structures, attributes, and chunking. `ncompare` was developed at NASA’s Atmospheric Science Data Center (ASDC) and is a collaboration with NASA Openscapes [[https://nasa-openscapes.github.io]] mentors across 11 of NASA’s data centers. Openscapes’ overarching vision is to support scientific researchers using NASA Earthdata as they migrate their workflows to the cloud. Relevant links: - https://github.com/nasa/ncompare - https://github.com/pyOpenSci/software-submission/issues/146 - https://nasa-openscapes.github.io

Daniel Kaufman↗

From the Knowledge-based Digital Platform (KbDP) Concept for Advanced Air Mobility Research to a Preliminary Prototype

Advanced Air Mobility (AAM) encompasses a range of innovative operational and technological changes to aviation (electric aircraft, increasingly automated aircraft, increasingly automated airspace operations, etc.) that are transforming aviation’s role in everyday movement of people and goods. There are multiple associated concepts and use cases for AAM, all interrelated, including small Unmanned Aircraft System (UAS) Traffic Management (UTM), Upper-Class E Traffic Management (ETM), Extensible Traffic Management (xTM), Regional Air Mobility (RAM), and Urban Air Mobility (UAM). These AAM operations must integrate with traditional Air Traffic Management (ATM) operations, as well as non-aviation modes of transportation and logistics. National Aeronautics and Space Administration (NASA) is spearheading an innovative digital engineering approach to integrate, communicate, and facilitate the research of multi-modal transportation systems. The Knowledge-based Digital Platform (KbDP) is a concept being developed that ties the workflows of Project Managers (PM), Principal Investigators (PI), and System Engineers together across organizational boundaries. It does so through the management of an information database defined by mathematical, data science, and system engineering principles. Machine Learning (ML) algorithms play a key role in this concept by extracting meaningful knowledge from the information database, which the human user leverages to greatly improve the efficiency and effectiveness of their research. Expected benefits of this concept include improved technology transfers from research to production, improved research portfolio investments, and research outcomes that are more integrated with all aspects of the multi-modal transportation problem. The preliminary KbDP prototype has been realized using UAM as a pathfinder use case and developed by a team of system engineer, software developer, data scientist, and interns.

Systems Engineering↗

Range Safety Flight Elevation Limit Calculation Software

This program was developed to fill a need within the Wallops Flight Facility workflow for automation of the development of vertical plan limit lines used by flight safety officers during the conduct of expendable launch vehicle missions. Vertical plane present-position-based destruct lines have been used by range safety organizations at numerous launch ranges to mitigate launch vehicle risks during the early phase of flight. Various ranges have implemented data submittal and processing workflows to develop these destruct lines. As such, there is significant prior art in this field. The ElLimits program was developed at NASA's Wallops Flight Facility to automate the process for developing vertical plane limit lines using current computing technologies. The ElLimits program is used to configure launch-phase range safety flight control lines for guided missiles. The name of the program derives itself from the fundamental quantity that is computed - flight elevation limits. The user specifies the extent and resolution of a grid in the vertical plane oriented along the launch azimuth. At each grid point, the program computes the maximum velocity vector flight elevation that can be permitted without endangering a specified back-range location. Vertical plane x-y limit lines that can be utilized on a present position display are derived from the flight elevation limit data by numerically propagating 'streamlines' through the grid. The failure turn and debris propagation simulation technique used by the application is common to all of its analysis options. A simulation is initialized at a vertical plane grid point chosen by the program. A powered flight failure turn is then propagated in the plane for the duration of the so-called RSO reaction time. At the end of the turn, a delta-velocity is imparted, and a ballistic trajectory is propagated to impact. While the program possesses capability for powered flight failure turn modeling, it does not require extensive user inputs of vehicle characteristics (e.g., thrust and aerodynamic data), nor does it require reams of turn data after the traditional fashion of the Air Force ranges. The program requires a nominal trajectory table (time, altitude, range, velocity, and flight elevation) and makes heavy use of it to initialize and model a failure turn.

Lanzi, Raymond J↗

Simplifying NASA Earth Science Data and Information Access Through Natural Language Processing Based Data Analysis and Visualization

NASA Earth science data collected from satellites, model assimilation, airborne missions, and field campaigns, are large, complex and evolving. Such characteristics pose great challenges for end users (e.g., Earth science and applied science users, students, citizen scientists), particularly for those who are unfamiliar with NASA's EOSDIS and thus unable to access and utilize datasets effectively. For example, a novice user may simply ask: what is the total rainfall for a flooding event in my county yesterday? For an experienced user (e.g., algorithm developer), a question can be: how did my rainfall product perform, compared to ground observations, during a flooding event? Nonetheless, with rapid information technology development such as natural language processing, it is possible to develop simplified Web interfaces and back-end processing components to handle such questions and deliver answers in terms of text, data, or graphic results directly to users.In this presentation, we describe the main challenges for end users with different levels of expertise in accessing and utilizing NASA Earth science data. Surveys reveal that most non-professional users normally do not want to download and handle raw data as well as conduct heavy-duty data processing tasks. Often they just want some simple graphics or data for various purposes. To them, simple and intuitive user interfaces are sufficient because complicated ones can be difficult and time-consuming to learn. Professionals also want such interfaces to answer many questions from datasets. One solution is to develop a natural language based search box like Google and the search results can be text, data, graphics and more. Now the challenge is, with natural language processing, can we design a system to process a scientific question typed in by a user? In this presentation, we describe our plan for such a prototype. The workflow is: 1) extract needed information (e.g., variables, spatial and temporal information, processing methods, etc.) from the input, 2) process the data in the backend, and 3) deliver the results (data or graphics) to the user.

Liu, Zhong↗

Expanding Repository Data Available For Sharing And Knowledge Discovery

Some of the hardest space biology and space health challenges require data-intensive, bioinformatic, meta-analytical, and computer-assisted research approaches. These challenges include examining interdisciplinary space life science research across experiments and across interacting spaceflight hazards (radiation, altered gravity, confinement, hostile-closed environments, distance-duration from Earth). The approaches to confront these challenges involve mining multiple datasets simultaneously from various hierarchical organizations of biological complexity, all while concurrently evaluating how experimental design factors affect endpoints of standard assays. To enable this field, it is essential that principal investigators (PIs) submit data in a structure so it can be maximally re-used. The purpose of the NASA Ames Life Sciences Data Archive (ALSDA) is to collect, curate, and make publicly available all non-human space-relevant biological data. ALSDA must also ensure data are open-access, and maximally findable, accessible, interoperable, and reusable (FAIR). The scope of ALSDA data collected and submitted by PIs include subject and study design metadata, assay metadata parameters, raw and processed assay data, assay imagery/video, and subject-experienced mission data telemetry (radiation, temperature, humidity, acoustics, vibrations, etc.). ALSDA recently integrated into a collaborative group of Open Science projects to facilitate a suite of new tools and workflows that will improve data submission, accessibility, and reusability by implementing digital data submission agreements, and adopting the data management system originally developed by NASA GeneLab. ALSDA intends to bring current biological repository data and all future collected data into this new scientific data reuse reality. This new suite of tools will enable ALSDA to deploy a science curation system using scientific assay configurations for the data submission portal. It will capture essential assay parameters according to established standards in each sub-field within biology. The submission portal expedites data collection by enhancing ease of PI data submission, providing a user interface and specificity for which data is to be submitted. Data submissions can be brought into cutting-edge informatic analysis portals to enable mining of physiological, behavioral, biochemical, and imaging datasets in conjunction with ‘omics-level datasets. As ALSDA datasets are submitted, curated, and published (e.g., micro-computed tomography, histology, pulse oximetry, serum metabolites, magnetic resonance imaging, intraocular pressure, novel object recognition, etc.), the merging together of spaceflight data along this multi-hierarchical complexity of biology will enable informatics and data-intensive approaches resulting in knowledge discoveries across missions, space hazards, and biological disciplines.

life science↗