Engineering PapersSearch

SEARCH · Engineering Papers

Results for “big data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

A 28 nm multiply-accumulate ASIC architecture for on-chip data compression in MHz frame rate X-ray and electron pixel detectors

Modern X-ray detector systems urgently require compact, efficient, and fast data compression schemes to handle the transmission of big data from pixel arrays, enabling frame rates in the MHz regime. Here, in this work, a data compression ASIC that implements a streaming fixed-length lossy compression scheme is introduced and analyzed, proving the feasibility and benefits of on-chip compression. The compression scheme utilizes a vector matrix product logic, which performs a number of floating-point multiplications, additions, and accumulations. The logic is verified, synthesized, and shown to fit in the area resource available for the X-ray detector under study, which comprises 192 × 168 pixels each of 12-bit width, and having a total area of 20 mm× 20 mm, about 2 mm× 20 mm of which are available for the digital logic. Several system architectures, precisions, and compression ratios ranging from 100 to 250 were analyzed to pave the way for on-chip fixed-length compression (e.g., principal component analysis, singular value decomposition) and data reduction (e.g., azimuthal integration) for X-ray and electron detectors.

Data compression

FAIRness and Usability for Open-Access Omics Data Systems

Omics data sharing is especially crucial to the biological research community, and the last decade or two has seen a huge rise in collaborative analysis systems, databases, and knowledge bases for omics and other systems biology data. We assessed the "FAIRness" of NASA's GeneLab Data Systems (GLDS) along with four similar kinds of systems in the research omics data domain, using 14 FAIRness metrics. 14 metrics. The range of Pass ratings was 29-79% of the 14 metrics, Partial Pass 0-21%, and Fail 7-50%. The range of overall FAIRness scores was 5-12 (out of 14). The systems we evaluated performed the best in the areas of data findability and accessibility, and worst in the area of data interoperability. We propose two new principles that Big Data systems, in particular, should consider for increasing data accessibility. We relate our experiences implementing semantic integration of omics data from several systems for the federated querying and retrieval functions of the GLDS, given the shortcomings in data interoperability of these systems.

Berrios, Daniel C.

Cloud Optimized Data Formats

Cloud computing offers the promise of being able to analyze Big Data earth Observations at scale, by allowing scientists to deploy many nodes at once to analyze the data. However, in order to take full advantage of cloud scalability, it is often necessary to reorganize and reformat the data to enable fine-grained, parallel access to the data in Web Object Storage. NASA recently conducted a study of several formats that are optimized for analysis in the cloud: Parquet, zarr, HDF (Hierarchical Data Format) in the Cloud, and Cloud-Optimized GeoTIFF (Tagged Image File Format). They were compared against non-cloud-optimized formats, netCDF (network Common Data Form) and GeoTIFF, with criteria based both on stewardship and analysis performance.

Christopher Lynnes

FAIRness and Usability for Open-access Omics Data Systems

Omics data sharing is crucial to the biological research community, and the last decade or two has seen a huge rise in collaborative analysis systems, databases, and knowledge bases for omics and other systems biology data. We assessed the "FAIRness" of NASA's GeneLab Data Systems (GLDS) along with four similar kinds of systems in the research omics data domain, using 14 FAIRness metrics. The range of overall FAIRness scores was 6-12 (out of 14), average 10.1, and standard deviation 2.4. The range of Pass ratings for the metrics was 29-79%, Partial Pass 0-21%, and Fail 7-50%. The systems we evaluated performed the best in the areas of data findability and accessibility, and worst in the area of data interoperability. Reusability of metadata, in particular, was frequently not well supported. We relate our experiences implementing semantic integration of omics data from some of the assessed systems for federated querying and retrieval functions, given their shortcomings in data interoperability. Finally, we propose two new principles that Big Data system developers, in particular, should consider for maximizing data accessibility.

Berrios, Daniel C.

Federated Cloud Challenges in NASA's Earth Science Data Systems (Why So Difficult?)

NASA is presented with a number of opportunities and challenges in federating its Earth Science Data Systems in the burgeoning world of cloud computing. Cloud hosting of Earth Science data provides a new way of bringing data together, at least from a virtual location sense, and is one of the main motives for NASA to host data there. However, NASA is also faced with a Big Data Variety challenge, brought on by the variety of the EO datasets in its archives. This diversity requires many diverse science archives to service the different science communities. As a result, nearly every major function in the Earth Observing System Data and Information System (EOSDIS) must also be federated across its data centers. This pattern is repeated with many of the outside agencies and organizations that EOSDIS federates with, such as the Committee for Earth Observing Satellites, leading to pioneering work on "deep federation" in a joint project with the European Space Agency to develop a Multi-Mission Algorithm and Analysis Platform.

Lynnes, Christopher

FAIRness and Usability for Open-access Omics Data Systems

Omics data sharing is crucial to the biological research community, and the last decade or two has seen a huge rise in collaborative analysis systems, databases, and knowledge bases for omics and other systems biology data. We assessed the “FAIRness” of NASA’s GeneLab Data Systems (GLDS) along with four similar kinds of systems in the research omics data domain, using 14 FAIRness metrics. The range of overall FAIRness scores was 6-12 (out of 14), average 10.1, and standard deviation 2.4. The range of Pass ratings for the metrics was 29-79%, Partial Pass 0-21%, and Fail 7-50%. The systems we evaluated performed the best in the areas of data findability and accessibility, and worst in the area of data interoperability. Reusability of metadata, in particular, was frequently not well supported. We relate our experiences implementing semantic integration of omics data from some of the assessed systems for federated querying and retrieval functions, given their shortcomings in data interoperability. Finally, we propose two new principles that Big Data system developers, in particular, should consider for maximizing data accessibility.

Berrios, Daniel C.

SNPP and N20 VIIRS Thermal Emissive Bands Calibration Comparison Using the GEO-LEO Double Difference Method

The VIIRS instruments onboard the SNPP and NOAA-20 satellites have identical spatial resolutions and the same spectral bands. Similar prelaunch tests and identical on-orbit calibration algorithms established the foundation for their consistent Earth measurements. Calibration assessment and consistency comparisons are useful to maintain their performance and measurement accuracy. Simultaneous nadir overpasses (SNO) between two satellites are commonly used for a direct calibration comparison between sensors. However, there are no SNO between SNPP and NOAA20. Hence, a reference sensor or Earth measurements are normally used to bridge the comparison. As a reference, we focus on the Advanced Baseline Imager (ABI) onboard the GOES-R series spacecraft and its application to the SNPP and NOAA-20 VIIRS comparison. GOES16 and GOES17are the first two satellites of the GOES-R series and were launched on November 19, 2016, and March 12, 2018, respectively. Their operational positions are on the equator with longitudes of 75.2° West over land and 137.2° West over ocean, respectively. The ABI is the primary imaging instrument of these spacecrafts for the Earth’s weather, oceans, and environment, with observations (every 10 minutes) that provide vast data for GEO-Low Earth orbit (LEO)and LEO-LEO comparisons utilizing it as an intermediate reference sensor. VIIRS and ABI have spectrally matched bands and can have simultaneous measurements over any selected site every day. The simultaneous measurements over the same site also have various scan angles. These features provide advantages for a VIIRS-to-ABI comparison. The spectral response function difference between instruments, sites selected, and view angles will have effects on the instrument measurements. Their impacts on the calibration comparison, including the use of double differences, will be discussed. By collecting VIIRS measurements over a large range of view angles, the view angle effect will also be investigated. The collection of an ample amount of data provides an advantage for statistical analyses and potential big data applications to sensor calibration assessments. This method can also be applied to other sensor calibration comparison and performance assessments, such as GOES16 and GOES17 ABI, and Terra and Aqua MODIS.

Tiejun Chang

Expanding Access to Science Participation: A FAIR Framework for Petascale Data Visualization and Analytics

The massive data generated by scientists daily serve as both a major catalyst for new discoveries and innovations, as well as a significant roadblock that restricts access to the data. Here, our paper introduces a new approach to removing Big Data barriers and democratizing access to petascale data for the broader scientific community. Our novel data fabric abstraction layer allows user-friendly querying of scientific information while hiding the complexities of dealing with file systems or cloud services. We enable FAIR (Findable, Accessible, Interoperable, and Reusable) access to datasets such as NASA’s petascale climate datasets. Our paper presents an approach to managing, visualizing, and analyzing petabytes of data within a browser on equipment ranging from the top NASA supercomputer to commodity hardware like a laptop. Our novel data fabric abstraction utilizes state-of-the art progressive compression algorithms and machine-learning insights to power scalable visualization dashboards for petascale data. The result provides users with the ability to identify extreme events or trends dynamically, expanding access to scientific data and further enabling discoveries. We validate our approach by improving the ability of climate scientists to visually explore their data via three fully interactive dashboards. We further validate our approach by deploying the dashboards and simplified training materials in the classroom at a minority-serving institution. These dashboards, released in simplified form to the general public, contribute significantly to a broader push to democratize the access and use of climate data.

Computer science

Detection of Hail Storms in Radar Imagery Using Deep Learning

In 2016, hail was responsible for 3.5 billion and 23 million dollars in damage to property and crops, respectively, making it the second costliest weather phenomenon in the United States. In an effort to improve hail-prediction techniques and reduce the societal impacts associated with hail storms, we propose a deep learning technique that leverages radar imagery for automatic detection of hail storms. The technique is applied to radar imagery from 2011 to 2016 for the contiguous United States and achieved a precision of 0.848. Hail storms are primarily detected through the visual interpretation of radar imagery (Mrozet al., 2017). With radars providing data every two minutes, the detection of hail storms has become a big data task. As a result, scientists have turned to neural networks that employ computer vision to identify hail-bearing storms (Marzbanet al., 2001). In this study, we propose a deep Convolutional Neural Network (ConvNet) to understand the spatial features and patterns of radar echoes for detecting hailstorms.

natural hazard

VEDA Visualization Exploration & Data Analysis

Why? - Interdisciplinary science depends on large amount of Earth science data and computational resources - Working with these datasets is non-trivial - Big data science requires advanced distributed computing knowledge What? VEDA is an open platform that brings key Earth science datasets next to open source tools for data processing, analysis, visualization, and exploration in a managed and more accessible computing environment.

Manil Maskey

Scalable Adaptive Graphics Environment (SAGE) Software for the Visualization of Large Data Sets on a Video Wall

The use of collaborative scientific visualization systems for the analysis, visualization, and sharing of "big data" available from new high resolution remote sensing satellite sensors or four‐dimensional numerical model simulations is propelling the wider adoption of ultra‐resolution tiled display walls interconnected by high speed networks. These systems require a globally connected and well‐integrated operating environment that provides persistent visualization and collaboration services. This abstract and subsequent presentation describes a new collaborative visualization system installed for NASA's Shortterm Prediction Research and Transition (SPoRT) program at Marshall Space Flight Center and its use for Earth science applications. The system consists of a 3 x 4 array of 1920 x 1080 pixel thin bezel video monitors mounted on a wall in a scientific collaboration lab. The monitors are physically and virtually integrated into a 14' x 7' for video display. The display of scientific data on the video wall is controlled by a single Alienware Aurora PC with a 2nd Generation Intel Core 4.1 GHz processor, 32 GB memory, and an AMD Fire Pro W600 video card with 6 mini display port connections. Six mini display‐to‐dual DVI cables are used to connect the 12 individual video monitors. The open source Scalable Adaptive Graphics Environment (SAGE) windowing and media control framework, running on top of the Ubuntu 12 Linux operating system, allows several users to simultaneously control the display and storage of high resolution still and moving graphics in a variety of formats, on tiled display walls of any size. The Ubuntu operating system supports the open source Scalable Adaptive Graphics Environment (SAGE) software which provides a common environment, or framework, enabling its users to access, display and share a variety of data‐intensive information. This information can be digital‐cinema animations, high‐resolution images, high‐definition video‐teleconferences, presentation slides, documents, spreadsheets or laptop screens. SAGE is cross‐platform, community‐driven, open‐source visualization and collaboration middleware that utilizes shared national and international cyberinfrastructure for the advancement of scientific research and education.

Jedlovec, Gary

An Integrated Data Analytics Platform

An Integrated Science Data Analytics Platform is an environment that enables the confluence of resources for scientific investigation. It harmonizes data, tools and computational resources which subsequently enable the research community to focus on the investigation rather than spending time on security, data preparation, management, etc. OceanWorks is a NASA technology integration project to establish a cloud-based Integrated Ocean Science Data Analytics Platform at NASA’s Physical Oceanography Distributed Active Archive Center (PO.DAAC) for big ocean science. It focuses on advancement and maturity by bringing together several NASA open-source, big data projects for parallel analytics, anomaly detection, in-situ to satellite data matchup, quality-screened data subsetting, search relevancy, and data discovery. Our communities are relying on data distributed through data centers such as the PO.DAAC, COAPS, NCAR, and many others to conduct their research. In typical investigations, scientists would engage in: search for data, evaluate the relevance of that data, download it, and then apply algorithms to identify trends. Such workflow cannot scale if the research involves a massive amount of data or multi-variate measurements. NASA’s Surface Water and Ocean Topography (SWOT) mission is expected to produce massive amount of observational data during its 3-year nominal mission. Collections like SWOT challenges all existing Earth Science data archival, distribution and analysis paradigms. In this paper, we will discuss how OceanWorks enhances the analysis of physical ocean data where the computation is done on an elastic cloud platform next to the archive to deliver fast, web-accessible services for working with oceanographic measurements.

Yang, Chaowei

Advancing Sustainability in Data Centers: Evaluation of Hybrid Air/Liquid Cooling Schemes for IT Payload Using Sea Water

Abstract-The growth in cloud computing, Big Data, AI and high-performance computing (HPC) necessitate the deployment of additional data centers (DC's) with high energy demands. The unprecedented increase in the Thermal Design Power (TDP) of the computing chips will require innovative cooling techniques. Furthermore, DC's are increasingly limited in their ability to add powerful GPU servers by power capacity constraints. As cooling energy use accounts for up to 40% of DC energy consumption, creative cooling solutions are urgently needed to allow deployment of additional servers, enhance sustainability and increase energy efficiency of DC's. The information in this study is provided from Start Campus' Sines facility supported by Alfa Laval for the heat exchanger and CO 2 emission calculations. The study evaluates the performance and sustainability impact of various data center cooling strategies including an air-only deployment and a subsequent hybrid air/water cooling solution all utilizing sea water as the cooling source. Here we evaluate scenarios from 3 MW to 15+1 MW of IT load in 3 MW increments which correspond to the size of heat exchangers used in the Start Campus' modular system design. This study also evaluates the CO 2 emissions compared to a conventional chiller system for all the presented scenarios. Results indicate that the effective use of the sea water cooled system combined with liquid cooled systems improve the efficiency of the DC, plays a role in decreasing the CO 2 emissions and supports in achieving sustainability goals.

97 MATHEMATICS AND COMPUTING

MPEX AI Digital Twins

All magnetically confined plasma fusion power plant concepts (Tokamak, Spherical Tokamak, Stellarator, Mirror, ...) must exhaust the heat and plasma from the core confinement region to the material walls. The primary channel for this exhaust is through a plasma divertor which directs plasma along open magnetic field lines to a material target. The Material Plasma Exposure eXperiment (MPEX) illustrated in Figure 1, is a high-power, steady-state linear plasma device designed to produce the plasma material interaction (PMI) conditions of the divertor of future magnetic confinement fusion power plants: energy flux 20MW/m 2 , ion fluence 1031/m 2 , pulse duration 106 sec. These goals of plasma exposure in MPEX are well beyond those achieved in magnetic fusion experimental devices. Successfully achieving these high power steady state conditions for long pulses requires operational control of the heating and particle sources and the plasma flux to the walls and target. The MPEX AI Hot Spot Controller, proposed in this project, will help achieve the operational milestones of MPEX. The MPEX device will begin commissioning at the end of FY26. A smaller proto-MPEX was operated for 14,666 plasma discharges and will resume operation in September of 2025 as proto-MPEX-lite, with reduced capability, to test a new window for the Helicon plasma source. The proto-MPEX data has undergone surrogate modeling with machine learning methods (R. Archibald, 2022 IEEE International Conference on Big Data). This proto-MPEX data will be used to begin development of the AI digital twins described in this white paper. The scientific mission of MPEX is to qualify materials of different composition for use in the high energy and plasma flux conditions of a fusion power plant. The materials exposed in MPEX will in some cases be exposed to high neutron fluxes at other ORNL facilities to measure the changes to their PMI properties. The targets exposed in MPEX will be transported under vacuum to a Surface Analysis Station (SAS). The SAS will be equipped with the following diagnostics: Focused Ion Beam (FIB) for trench milling, 100-400 angstrom resolution scanning electron microscope (SEM), surface mapping x-ray spectrometer, high resolution camera, and a future upgrade to a laser induced breakdown spectroscopy quadruple mass spectrometer (LIBS-QMS). The MPEX experiments will generate diverse pre- and post-exposure measurement data of detailed material properties down to the crystal grain level in 3D for post-exposure assessment of PMI damage (e.g. cracking, melting, erosion and redeposition of the material). Physics models for the PMI, and how the material composition and manufacturing impact its performance under high energy plasma exposure, need to be validated with MPEX data to guide the selection of new candidate materials. Our vision for the MPEX AI Digital Twins project is to supply experimental and physics model simulation data to train Artificial Intelligence (AI) models for data processing, analysis, operational control, PMI and materials simulation to maximize the scientific output of the MPEX device. Ultimately, an AI digital twin of MPEX material assessment metrics for tested and synthetic material types with simulated PMI will be trained by the AI Modeling Teams on the experimental and physics simulation data submitted to the American Science Cloud by this project. A purely empirical search for the best material is inefficient given the finite number of samples that can be tested on MPEX. In order to expand the material properties database for training the MPEX Material Assessment AI Digital Twin, and to gain physics understanding of the PMI processes, physics models of the material properties and PMI processes are required. The physics simulations provide detailed simulation data, like impact angles for plasma ions, sputtering yields, transport of the ionized sputtered target material in the plasma, and redeposition locations. This simulation data expands the measurement data for deeper physics understanding. The experimental data is essential to validate the PMI and material structure simulation models. The validated models can then be used to generate new simulation data of MPEX material assessments for synthetic material compositions that have not been exposed in MPEX. These predictive simulations, plus the whole experimental dataset, will be used to train the MPEX Material Assessment AI Digital Twin allowing a rapid generative AI search for new materials with reduced PMI damage by interpolating the domain of the training set. These new optimum materials can be simulated with the physics codes and/or tested in MPEX. The ability of AI neural networks to interpolate multi-dimensional parameter spaces and generate virtual data is exploited for a more efficient search for optimum materials. The advent of the Transformational AI Models Consortium (TAIMC) is an opportunity to engage with state of the art private and public AI developers to achieve the goals of the AI digital twins and AI accelerated physics models proposed in this project. Our partners at ORNL from the Advance Scientific Computing Research (ASCR) organization will collaborate in accelerating the integrated plasma material interaction simulation framework. This simulation framework will provide a platform for generating simulation data across a range of physical fidelities, including hybrid methods that produce multi-fidelity results. This data will be leveraged for AI model development, both for generation of surrogates and the automation of simulation campaigns. A part of the research below will include collaborative efforts with the TAIMC to (i) adapt data storage approaches to ensure AI-readiness, (ii) provide a protypical exemplar to inform and exercise constructed workflows, and (iii) generate and share data, using the TAIMC unified AI data standard, for foundational models that will be trained from multiple sources across the DOE complex. We will also collaborate with the TAIMC, as well as the planned AI modeling teams, to develop approaches for reducing the cost of data generation. These include tailored multi-fidelity approaches as well as fine-tuning strategies to augment general, large-scale foundational models.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY

Mining Twitter Data to Augment NASA GPM Validation

The Twitter data stream is an important new source of real-time and historical global information for potentially augmenting the validation program of NASA's Global Precipitation Measurement (GPM) mission. There have been other similar uses of Twitter, though mostly related to natural hazards monitoring and management. The validation of satellite precipitation estimates is challenging, because many regions lack data or access to data, especially outside of the U.S. and in remote and developing areas. The time-varying set of "precipitation" tweets can be thought of as an organic network of rain gauges, potentially providing a widespread view of precipitation occurrence. Twitter provides a large source of crowd for crowdsourcing. During a 24-hour period in the middle of the snow storm this past March in the U.S. Northeast, we collected more than 13,000 relevant precipitation tweets with exact geolocation. The overall objective of our project is to determine the extent to which processed tweets can provide additional information that improves the validation of GPM data. Though our current effort focuses on tweets and precipitation, our approach is general and applicable to other social media and other geophysical measurements. Specifically, we have developed an operational infrastructure for processing tweets, in a format suitable for analysis with GPM data; engaged with potential participants, both passive and active, to "enrich" the Twitter stream; and inter-compared "precipitation" tweet data, ground station data, and GPM retrievals. In this presentation, we detail the technical capabilities of our tweet processing infrastructure, including data abstraction, feature extraction, search engine, context-awareness, real-time processing, and high volume (big) data processing; various means for "enriching" the Twitter stream; and results of inter-comparisons. Our project should bring a new kind of visibility to Twitter and engender a new kind of appreciation of the value of Twitter by the science research communities.

validatio

Sherlock Data Warehouse

This slide deck provides an overview of the data and resources available in the Sherlock Data Warehouse. Sherlock was developed and is currently maintained by the Aviation Systems Division at NASA Ames Research Center. Sherlock contains a valuable collection of flight, air traffic management, and weather data. But Sherlock is not just a data archive. Sherlock also includes tools and resources to access, download, and visualize data, as well as resources to process the data. This overview summarizes Sherlock data sources, demonstrates data analytics and visualization with MicroStrategy, illustrates disparate data integration using the ATM Knowledge graph, and presents a machine learning use case using the Big Data system.

data warehouse

Achieving Fast Operational Intelligence in NASA's Deep Space Network Through Complex Event Processing

NASA’s Deep Space Network (DSN) is a complex, global project, in which the expertise of human operators remain crucial for its successful operation. To find ways to save costs in operations and to improve its services, a number of modernization efforts are underway in the DSN. One such effort is a research and technology development task at the Jet Propulsion Laboratory that is investigating the use of complex event processing (CEP) for intelligent assessment of situations, trend analysis, and advanced automation. The technology leverages the significant business intelligence (BI) and data science advancements made in the enterprise industries over the last several years. The open source big data processing engine Apache SparkTM and the high-throughput, distributed messaging system Apache Kafka form the core of the DSN Complex Event Processing (DCEP) framework. This paper discusses the system engineering perspective of why achieving efficient, lower-cost operations in the DSN is a challenging problem, how the DCEP system handles the use cases that help realize intelligent operations, and how this solution fits into the overall model of the planned DSN Follow-the- Sun Operations (FtSO).

Choi, Joshua S.

Open Science for Life in Space: Data Sharing and Tools for Knowledge Discovery

The next era in human space exploration is rapidly approaching and will require the use of countermeasures to deep space health hazards. The development of countermeasures (or, the re-purposing of existing agents) will be highly dependent on our understanding of basic biological responses to space stressors (e.g. ionizing radiation, altered gravitational fields, altered day-night cycles, confinement, isolation, hostile-closed environments, distance-duration from Earth, exposure to celestial regolith, etc.). The fast-growing array of space biological data, which in the past was simply archived after minimal analysis, holds great potential if it can be reorganized and formatted for Open Science. Organizing the data for such analysis is a challenge because of its diverse nature (molecular, cellular, tissue, imaging, whole organism and behavior). We will discuss here several strategies that NASA’s Biological and Physical Science Division has put in place to maximize the return on investment for spaceflight bioscience data. Open Science, as a scientific philosophy, is the concept that the more people who have access to the data, the more knowledge will be gained from it. This guiding principle led NASA to develop GeneLab in 2015. GeneLab houses spaceflight and relevant ground-based multi-omics data, and has grown to ~400 transcriptomic, proteomic, metabolomic and epigenomic datasets from plant, rodent, small animal, and microbial space experiments. GeneLab provides users with various tools for data analysis and a visualization portal that allows users to interact with gene expression data from space-related ‘omics experiments. Open Science is also about building scientific communities, and with this spirit in mind, GeneLab has spawned several Analysis Working Groups (AWGs), comprised of more than 200 volunteer scientists. The AWGs initially provided feedback on the processing pipeline and metadata ‘omics standards for GeneLab. Over the last few years, they have become a community-driven science enterprise, engaging in large meta-analysis of GeneLab datasets, resulting in 10 publications (beyond the originally submitted research). Overall, the Open Science nature of GeneLab has resulted in a high degree of data re-use, resulting in 38 additional publications derived from the original 67 publication over the past four years. The enormous success and knowledge gained from GeneLab has led to a collection of sister NASA “Open Science Data Repositories (OSDR)” and research support groups. These include the NASA Ames Life Sciences Data Archive (ALSDA), the NASA Biological Institutional Scientific Collection (NBISC), and the Biospecimen Sharing Program (BSP). All are adopting the GeneLab data architecture system to maximize open-access, find-ability, accessibility, interoperability, and reusability (FAIR). ALSDA collects and curates phenotypic-physiological bioimaging-behavioral data from space and space-relevant non-human experiments, oftentimes coming from the same omics-associated experimental datasets found in GeneLab. Since 2021, a community of ~100 researchers have rallied around ALSDA, to provide feedback in a new ALSDA AWG focused on phenotypic-physiological investigation-sample-assay metadata standards (e.g., Micro-Computed Tomography, Light/Fluorescence Microscopy, Western Blot, Flow Cytometry, Novel Object Recognition, Elevated Plus Maze, etc. of ~50 assays collected). These standards are part of a new single point-of-entry data submission portal for all non-human Space Biology and Human Research Program principal investigators, to submit, curate, and share their research data. With open-access space biological data now collected and curated together with rich metadata, and with the potential for linkage to “big data” from the international biological and medical communities (NIH, EBI, etc.), the artificial intelligence and machine learning (AI/ML) era has started for Space Biology. Several other talks will cover these topics in this conference.

life sciences