Engineering PapersSearch

SEARCH · Engineering Papers

Results for “DATA”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

The Environmental Data Application for Analysis of Space Telemetry Data

Sensors on the International Space Station (ISS) and multiple spacecraft elsewhere in Earth orbit and in deep space continuously monitor and collect environmental data, transmitting this information back to Earth. These data include ionizing radiation and, on the ISS and spacecrafts, CO2, relative humidity levels, and temperature, and are of great importance to space biology research. Looking ahead to future long duration crewed missions beyond low Earth orbit, the ability to study how factors including CO2 levels, light cycle, temperature modulate the response to ionizing radiation and microgravity is essential. To date, access to these data has been fragmented across space agencies, spacecraft, and databases. To address this issue, NASA’s Open Science Data Repository (OSDR) has developed a user interface for interrogation of telemetry data: the Environmental Data Application (EDA). The EDA provides the capability to visualize telemetry and radiation data collected on the International Space Station and corresponding ground platforms during the Rodent Research missions. Telemetry data includes temperature, relative humidity, and CO2 levels. Radiation data includes galactic cosmic rays, the contribution of the South Atlantic Anomaly, total radiation dose rate, and accumulated radiation dose. The application allows users to view single missions, compare multiple missions, and view and download summary or full data tables. In summary, the EDA provides GUIs for data visualization and exploration, as well as means for data export, making these data FAIR (Findable, Accessible, Interoperable, and Reusable), complementing the biological data contained in OSDR, and providing the space science community with a valuable resource for scientific analyses.

telemetry

Data and scripts from: “Denoising autoencoder for reconstructing sensor observation data and predicting evapotranspiration: noisy and missing values repair and uncertainty quantification”

This data package includes data and scripts from the manuscript “Denoising autoencoder for reconstructing sensor observation data and predicting evapotranspiration: noisy and missing values repair and uncertainty quantification”.The study addressed common challenges faced in environmental sensing and modeling, including uncertain input data, missing sensor observations, and high-dimensional datasets with interrelated but redundant variables. Point-scaled meteorological and soil sensor observations were perturbed with noises and missing values, and denoising autoencoder (DAE) neural networks were developed to reconstruct the perturbed data and further predict evapotranspiration. This study concluded that (1) the reconstruction quality of each variable depends on its cross-correlation and alignment to the underlying data structure, (2) uncertainties from the models were overall stronger than those from the data corruption, and (3) there was a tradeoff between reducing bias and reducing variance when evaluating the uncertainty of the machine learning models.This package includes:(1) Four ipython scripts (.ipynb): “DAE_train.ipynb” trains and evaluates DAE neural networks, “DAE_predict.ipynb” makes predictions from the trained DAE models, “ET_train.ipynb” trains and evaluates ET prediction neural networks, and “ET_predict.ipynb” makes predictions from trained ET models.(2) One python file (.py): “methods.py” includes all user-defined functions and python codes used in the ipython scripts.(3) A “sub_models” folder that includes five trained DAE neural networks (in pytorch format, .pt), which could be used to ingest input data before being fed to the downstream ET models in ‘ET_train.ipynb” or ‘ET_predict.ipynb’.(4) Two data files (.csv). Daily meteorological, vegetation, and soil data is in “df_data.csv”, where “df_meta.csv” contains the location and time information of “df_data.csv”. Each row (index) in “df_meta.csv” corresponds to each row in “df_data.csv”. These data files are formatted to follow the data structure requirements and be directly used in the ipython scripts, and they have been shuffled chronologically to train machine learning models. The meteorological and soil data was collected using point sensors between 2019-2023 at(4.a) Three shrub-dominated field sites in East River, Colorado (named “ph1”, “ph2” and “sg5” in “df_meta.csv”, where “ph1” and “ph2” were located at PumpHouse Hillslopes, and “sg5” was at Snodgrass Mountain meadow) and(4.b) One outdoor, mesoscale, and herbaceous-dominated experiment in Berkeley, California (named “tb” in “df_meta.csv”, short for Smartsoils Testbed at Lawrence Berkeley National Lab).- See "df_data_dd.csv" and "df_meta_dd.csv" for variable descriptions and the Methods section for additional data processing steps. See "flmd.csv" and "README.txt" for brief file descriptions.- All ipython scripts and python files are written in and require PYTHON language software.

54 ENVIRONMENTAL SCIENCES

RC-SFA Data Management Templates and Guidance for Standardized, Reusable AI-Ready Data Packages

This data package provides templates and supporting documentation developed by the River Corridor Science Focus Area (RC-SFA; https://www.pnnl.gov/projects/river-corridor) to communicate its approach to managing and publishing AI-ready data. The package is intended to help data users and data producers understand the structures, metadata practices, and quality-control approaches that support consistent, reusable, and machine-actionable data products across RC-SFA studies. Rather than focusing on a single experimental dataset, this package documents the data management framework used to make RC-SFA data easier to find, ingest, navigate, and interpret. The materials in this package reflect RC-SFA practices for standardized data package organization, including the use of a human- and machine-readable README, file-level metadata, data dictionaries, descriptive file naming, method identifiers, and automated and review-based quality assurance procedures. Together, these components illustrate how RC-SFA extends FAIR data principles toward AI-readiness by prioritizing deep metadata, consistency across data packages, and support for informed downstream reuse by both humans and computational tools. This dataset is comprised of (1) readme; (2) presentation slides with an overview of RC-SFA approach and guidance; (3) document of RC-SFA best practices; (4) data dictionary (dd); (5) file level metadata (flmd); and a subfolder containing templates for dd and flmd. All files are .csv and .pdf. For details on how to navigate data packages generated by this project, see https://data.ess-dive.lbl.gov/portals/PNNLRiverCorridorSFA/About.

AI-readiness

MST data exchange through the NCAR incoherent-scatter radar data base

One means of making MST (mesosphere stratosphere troposphere) radar data more easily accessible for scientific research by the general scientific community is through a centralized data base. Such a data base can be designed to readily provide information on data availability and quality, and to provide copies of data from any radar in a common format to the user. The ionospheric incoherent scatter community has established a centralized data base at NCAR that may serve not only as a model for a possible MST data base, but also as a catalyst for getting an MST data base started. (Some key elements of the NCAR data base are given.) The NCAR data base can include MST data in the same framework with relatively little extra effort. They are willing to handle MST data on a limited basis in order to permit assessment of community interest and in order to provide some experience with a centralized data base for MST data.

Richmond, A. D.

Marine Magnetic Data Holdings of World Data Center-a for Marine Geology and Geophysics

The World Data Center-A for Marine Geology and Geophysics is co-located with the Marine Geology & Geophysical Data Center, Boulder, CO. Fifteen million digital marine magnetic trackline measurements are managed within the GEOphysical DAta System (GEODAS). The bulk of these data were collected with proton precision magnetometers under Transit Satellite navigational control. Along-track sampling averages about 1 sample per kilometer, while spatial density, a function of ship's track and survey pattern, range from 4 to 0.02 data points/sq. km. In the near future, the entire geophysical data set will be available on CD-ROM. The Marine Geology and Geophysics Division (World Data Center-A for MGG), of the National Geophysical Data Center, handles a broad spectrum of marine geophysical data, including measurements of bathymetry, magnetics, gravity, seismic reflection subbottom profiles, and side-scan images acquired by ships throughout the world's oceans. Digital data encompass the first three, while the latter two are in analog form, recorded on 35mm microfilm. The marine geophysical digital trackline data are contained in the GEODAS data base which includes 11.6 million nautical miles of cruise trackline coverage contributed by more than 70 organizations worldwide. The inventory includes data from 3206 cruises with 33 million digital records and indexing to 5.3 million track miles of analog data on microfilm.

Sharman, George F.

Real Data and Rapid Results: Ocean Color Data Analysis with Giovanni (GES DISC Interactive Online Visualization and ANalysis Infrastructure)

The NASA Goddard Earth Sciences Data and Information Services Center (GES DISC) has taken a major step addressing the challenge of using archived Earth Observing System (EOS) data for regional or global studies by developing an infrastructure with a World Wide Web interface which allows online, interactive, data analysis: the GES DISC Interactive Online Visualization and ANalysis Infrastructure, or "Giovanni." Giovanni provides a data analysis environment that is largely independent of underlying data file format. The Ocean Color Time-Series Project has created an initial implementation of Giovanni using monthly Standard Mapped Image (SMI) data products from the Sea-viewing Wide Field-of-view Sensor (SeaWiFS) mission. Giovanni users select geophysical parameters, and the geographical region and time period of interest. The system rapidly generates a graphical or ASCII numerical data output. Currently available output options are: Area plot (averaged or accumulated over any available data period for any rectangular area); Time plot (time series averaged over any rectangular area); Hovmeller plots (image view of any longitude-time and latitude-time cross sections); ASCII output for all plot types; and area plot animations. Future plans include correlation plots, output formats compatible with Geographical Information Systems (GIs), and higher temporal resolution data. The Ocean Color Time-Series Project will produce sensor-independent ocean color data beginning with the Coastal Zone Color Scanner (CZCS) mission and extending through SeaWiFS and Moderate Resolution Imaging Spectroradiometer (MODIS) data sets, and will enable incorporation of Visible/lnfrared Imaging Radiometer Suite (VIIRS) data, which will be added to Giovanni. The first phase of Giovanni will also include tutorials demonstrating the use of Giovanni and collaborative assistance in the development of research projects using the SeaWiFS and Ocean Color Time-Series Project data in the online Laboratory for Ocean Color Users (LOCUS). The synergy of Giovanni with high-quality ocean color data provides users with the ability to investigate a variety of important oceanic phenomena, such as coastal primary productivity related to pelagic fisheries, seasonal patterns and interannual variability, interdependence of atmospheric dust aerosols and harmful algal blooms, and the potential effects of climate change on oceanic productivity.

Acker, J. G.

Combining Satellite and in Situ Data with Models to Support Climate Data Records in Ocean Biology

The satellite ocean color data record spans multiple decades and, like most long-term satellite observations of the Earth, comes from many sensors. Unfortunately, global and regional chlorophyll estimates from the overlapping missions show substantial biases, limiting their use in combination to construct consistent data records. SeaWiFS and MODIS-Aqua differed by 13% globally in overlapping time segments, 2003-2007. For perspective, the maximum change in annual means over the entire Sea WiFS mission era was about 3%, and this included an El NinoLa Nina transition. These discrepancies lead to different estimates of trends depending upon whether one uses SeaWiFS alone for the 1998-2007 (no significant change), or whether MODIS is substituted for the 2003-2007 period (18% decline, P less than 0.05). Understanding the effects of climate change on the global oceans is difficult if different satellite data sets cannot be brought into conformity. The differences arise from two causes: 1) different sensors see chlorophyll differently, and 2) different sensors see different chlorophyll. In the first case, differences in sensor band locations, bandwidths, sensitivity, and time of observation lead to different estimates of chlorophyll even from the same location and day. In the second, differences in orbit and sensitivities to aerosols lead to sampling differences. A new approach to ocean color using in situ data from the public archives forces different satellite data to agree to within interannual variability. The global difference between Sea WiFS and MODIS is 0.6% for 2003-2007 using this approach. It also produces a trend using the combination of SeaWiFS and MODIS that agrees with SeaWiFS alone for 1998-2007. This is a major step to reducing errors produced by the first cause, sensor-related discrepancies. For differences that arise from sampling, data assimilation is applied. The underlying geographically complete fields derived from a free-running model is unaffected by solar zenith angle requirements and obscuration from clouds and aerosols. Combined with in situ dataenhanced satellite data, the model is forced into consistency using data assimilation. This approach eliminates sampling discrepancies from satellites. Combining the reduced differences of satellite data sets using in situ data, and the removal of sampling biases using data assimilation, we generate consistent data records of ocean color. These data records can support investigations of long-term effects of climate change on ocean biology over multiple satellites, and can improve the consistency of future satellite data sets.

Gregg, Watson

Usage of Data-Encoded Web Maps with Client Side Color Rendering for Combined Data Access, Visualization and Modeling Purposes

Current approaches to satellite observation data storage and distribution implement separate visualization and data access methodologies which often leads to the need in time consuming data ordering and coding for applications requiring both visual representation as well as data handling and modeling capabilities. We describe an approach we implemented for a data-encoded web map service based on storing numerical data within server map tiles and subsequent client side data manipulation and map color rendering. The approach relies on storing data using the lossless compression Portable Network Graphics (PNG) image data format which is natively supported by web-browsers allowing on-the-fly browser rendering and modification of the map tiles. The method is easy to implement using existing software libraries and has the advantage of easy client side map color modifications, as well as spatial subsetting with physical parameter range filtering. This method is demonstrated for the ASTER-GDEM elevation model and selected MODIS data products and represents an alternative to the currently used storage and data access methods. One additional benefit includes providing multiple levels of averaging due to the need in generating map tiles at varying resolutions for various map magnification levels. We suggest that such merged data and mapping approach may be a viable alternative to existing static storage and data access methods for a wide array of combined simulation, data access and visualization purposes.

Pliutau, Denis

Constructing the 'Best' Reliability Data for the Job - Developing Generic Reliability Data from Alternative Sources Early in a Product's Development Phase

Reliability practitioners advocate getting reliability involved early in a product development process. However, when assigned to estimate or assess the (potential) reliability of a product or system early in the design and development phase, they are faced with lack of reasonable models or methods for useful reliability estimation. Developing specific data is costly and time consuming. Instead, analysts rely on available data to assess reliability. Finding data relevant to the specific use and environment for any project is difficult, if not impossible. Instead, analysts attempt to develop the "best" or composite analog data to support the assessments. Industries, consortia and vendors across many areas have spent decades collecting, analyzing and tabulating fielded item and component reliability performance in terms of observed failures and operational use. This data resource provides a huge compendium of information for potential use, but can also be compartmented by industry, difficult to find out about, access, or manipulate. One method used incorporates processes for reviewing these existing data sources and identifying the available information based on similar equipment, then using that generic data to derive an analog composite. Dissimilarities in equipment descriptions, environment of intended use, quality and even failure modes impact the "best" data incorporated in an analog composite. Once developed, this composite analog data provides a "better" representation of the reliability of the equipment or component. It can be used to support early risk or reliability trade studies, or analytical models to establish the predicted reliability data points. It also establishes a baseline prior that may updated based on test data or observed operational constraints and failures, i.e., using Bayesian techniques. This tutorial presents a descriptive compilation of historical data sources across numerous industries and disciplines, along with examples of contents and data characteristics. It then presents methods for combining failure information from different sources and mathematical use of this data in early reliability estimation and analyses.

Kleinhammer, Roger K.

Using Big Data Technologies with Earth Science Data in HDF5: HDF5 Scalable Solutions

HDF5 (Hierarchical Data Format 5) is open-source, high-performance software that consists of an abstract data model, library, and fileformat used for storing and managing extremely large and/or complex data collections. NASA Earth Observing System (EOS) Data and Information Systems use HDF5 as an archival format to store remote sensing data from EOS satellites. HDF5 is also used to store other types of Geoscience and Strophysical data, e.g., seismic data and data from Low-Frequency Array (LOFAR) radio telescopes. Data stored in HDF5 has reached tens of petabytes and is growing at an accelerated rate.With the growing amout of HDF5 Earth Science data to analyze and process, scientists need to adopt big data technologies including new storage paradigms such as cloud and object storage. To run models and perform data analysis they also need to utilizied efficient and diverse ways to access data, from high-performance computing's (HPC) Message Passing Interface (MPI) I/O and deep memory hierarchies (DMH) to non-HPC frameworks such as Apache Hadoop, Spark, and Drill. The HDF Group continually works to enable usage of big data technologies in HDF software.

Knox, Larry

Medical Data Architecture Platform and Recommended Requirements for a Medical Data System for Exploration Missions

The Medical Data Architecture (MDA) project supports the Exploration Medical Capability (ExMC) risk to minimize or reduce the risk of adverse health outcomes and decrements in performance due to in-flight medical capabilities on human exploration missions. To mitigate this risk, the ExMC MDA project addresses the technical limitations identified in ExMC Gap Med 07: We do not have the capability to comprehensively process medically- relevant information to support medical operations during exploration missions. This gap identifies that the current in-flight medical data management includes a combination of data collection and distribution methods that are minimally integrated with on-board medical devices and systems. Furthermore, there are a variety of data sources and methods of data collection. For an exploration mission, the seamless management of such data will enable a more medically autonomous crew than the current paradigm of medical data management on the International Space Station. ExMC has recognized that in order to make informed decisions about a medical data architecture framework, current methods for medical data management must not only be understood, but an architecture must also be identified that provides the crew with actionable insight to medical conditions. This medical data architecture will provide the necessary functionality to address the challenges of executing a self-contained medical system that approaches crew health care delivery without assistance from ground support. Hence, the products derived from the third MDA prototype development will directly inform exploration medical system requirements for Level of Care IV in Gateway missions. In fiscal year 2019, the MDA project developed Test Bed 3, the third iteration in a series of prototypes, that featured integrations with cognition tool data, ultrasound image analytics and core Flight Software (cFS). Maintaining a layered architecture design, the framework implemented a plug-in, modular approach in the integration of these external data sources. An early version of MDA Test Bed 3 software was deployed and operated in a simulated analog environment that was part of the Next Space Technologies for Exploration Partnerships (NextSTEP) Gateway tests of multiple habitat prototypes. In addition, the MDA team participated in the Gateway Test and Verification Demonstration, where the MDA cFS applications was integrated with Gateway-in-a-Box software to send and receive medically relevant data over a simulated vehicle network. This software demonstration was given to ExMC and Gateway Program stakeholders at the NASA Johnson Space Center Integrated Power, Avionics and Software (iPAS) facility. Also, the integrated prototypes served as a vehicle to provide Level 5 requirements for the Crew Health and Performance Habitat Data System for Gateway Missions (Medical Level of Care IV). In the upcoming fiscal year, the MDA project will continue to provide systems engineering and vertical prototypes to refine requirements for medical Level of Care IV and inform requirements for Level of Care V.

Krihak, M.

Enriching the physics program of the CMS experiment via data scouting and data parking

Specialized data-taking and data-processing techniques were introduced by the CMS experiment in Run 1 of the CERN LHC to enhance the sensitivity of searches for new physics and the precision of standard model measurements. These techniques, termed data scouting and data parking, extend the data-taking capabilities of CMS beyond the original design specifications. The novel data-scouting strategy trades complete event information for higher event rates, while keeping the data bandwidth within limits. Data parking involves storing a large amount of raw detector data collected by algorithms with low trigger thresholds to be processed when sufficient computational power is available to handle such data. The research program of the CMS Collaboration is greatly expanded with these techniques. The implementation, performance, and physics results obtained with data scouting and data parking in CMS over the last decade are discussed in this Report, along with new developments aimed at further improving low-mass physics sensitivity over the next years of data taking.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS

Super Resolution for Renewable Energy Resource Data With Wind From Reanalysis Data (Sup3rWind) and Application to Ukraine [Slides]

In this work we present a novel deep learning-based downscaling method, using generative adversarial networks (GANs), for generating high-resolution wind resource data from ECMWF Reanalysis v5 data (ERA5). We show that by training a GAN model on ERA5, as opposed to coarsened high-resolution data, we achieve results that are competitive with conventional dynamical downscaling. This GAN-based downscaling method additionally reduces computational costs over dynamical downscaling by two orders of magnitude. All GANs are trained on data sampled from CONUS, selected to provide a diverse sampling of terrain conditions, and validated on observational data along with data held out from training. This cross-validation shows low error and high correlations with observations and excellent agreement with hold out data across physical distributions. Our approach is finally used to downscale 30km hourly ERA5 to 2-km 5-minute wind data, for January 2000 through December 2023, at multiple hub heights, over Ukraine, Moldova, and part of Romania. Comparisons against observational data from Meteorological Assimilation Data Ingest System (MADIS) and multiple wind farms show the same level of performance as for CONUS validation. This 24 year data record is the first member of the "super resolution for renewable energy resource data with wind from reanalysis data" dataset (Sup3rWind).

17 WIND ENERGY

Remote Sensing Data from CLARET: A Prototype Cart Data Set

A data set containing radiation, meteorological, and cloud sensor observations is documented. It was prepared for use by the Department of Energy's Atmospheric Radiation Measurement (ARM) program and other interested scientists. These data are a precursor of the types of data that ARM Cloud And Radiation Testbed (CART) sites will provide. The data are from the Cloud Lidar And Radar Exploratory Test (CLARET) conducted by the Wave Propagation Laboratory during autumn 1989 in the Denver-Boulder area of Colorado primarily for the purpose of developing new cloud-sensing techniques on cirrus. After becoming aware of this experiment, ARM scientists requested archival of subsets or the data to assist in the developing ARM program. Five CLARET cases were selected: two with cirrus, one with stratus, one with mixed-phase clouds, and one with clear skies. The cases range from 2 to 9.5 h in length. A pyranometer, pyrgeometer, pyrheliometer, and an infrared radiometer constituted the ensemble of instruments that provided surface radiation data. A lidar, radar, and ceilometer observed the cloud geometrical structure, and visual reports and all-sky camera observations were assimilated to provide cloud cover data. Radiosondes, wind profiler, RASS (profiling virtual temperature), microwave radiometers (observing column integrated liquid water and water vapor), and standard surface measurements provided meteorological data. Satellite data from the stratus case and one cirrus case were analyzed for statistics on cloud cover and top height. The main body of the selected data are available on diskette from the Wave Propagation Laboratory or Los Alamos National Laboratory. In addition to documenting the data set, this report describes CLARET and gives a bibliography of publications associated with the project. Some preliminary results of CLARET' research are also summarized. Simultaneous CO 2 lidar and radar backscatter measurements were shown to provide estimates of the effective radius of ice particles. Simultaneous radar and infrared radiometer data appear useful for estimating column-integrated numbers and average sizes of ice cloud particles. Ice water content obtained with this method compared favorably with values from another empirical technique using radar data alone. Depolarization of the CO 2 lidar signal from ice clouds was surprisingly small, suggesting that calculation of backscatter from nonspherical particles for this lidar is a tractable problem. Examples are also cited of CO 2 lidar measurements of the effective radius of water cloud drop size distributions and of inference of the size of pristine ice crystals that assume a particular orientation in the air. These parameters are all important to radiative transfer through clouds.

Clouds (Meteorology)

Modelling above-ground biomass stock over Norway using national forest inventory data with ArcticDEM and Sentinel-2 data

Boreal forests constitute a large portion of the global forest area, yet they are undersampled through field surveys, and only a few remotely sensed data sources provide structural information wall-to-wall throughout the boreal domain. ArcticDEM is a collection of high-resolution (2 m) space-borne stereogrammetric digital surface models (DSM) covering the entire land area north of 60° of latitude. The free-availability of ArcticDEM data offers new possibilities for aboveground biomass mapping (AGB) across boreal forests, and thus it is necessary to evaluate the potential for these data to map AGB over alternative open-data sources (i.e., Sentinel-2). This study was performed over the entire land area of Norway north of 60° of latitude, and the Norwegian national forest inventory (NFI) was used as a source of field data composed of accurately geolocated field plots (n=7710) systematically distributed across the study area. Separate random forest models were fitted using NFI data, and corresponding remotely sensed data consisting of either: i) a canopy height model (ArcticCHM) obtained by subtracting a high-quality digital terrain model (DTM) from the ArcticDEM DSM height values, ii) Sentinel-2 (S2), or iii) a combination of the two (ArcticCHM+S2). Furthermore, we assessed the effect of the forest- and terrain-specific factors on the models’ predictive accuracy. The best model (,i.e., ArcticCHM+S2) explained nearly 60% of the variance of the training set, which translated in the largest accuracy in terms of root mean square error (RMSE=41.4 t/ha). This result highlights the synergy between 3D and multispectral data in AGB modelling. Furthermore, this study showed that despite the importance of ArcticCHM variables, the S2 model performed slightly better than ArcticCHM model. This finding highlights some of the limitations of ArcticDEM, which, despite the unprecedented spatial resolution, is highly heterogeneous due to the blending of multiple acquisitions across different years and seasons. We found that both forest- and terrain-specific characteristics affected the uncertainty of the ArcticCHM+S2 model and concluded that the combined use of ArcticCHM and Sentinel-2 represents a viable solution for AGB mapping across boreal forests. The synergy between the two data sources allowed for a reduction of the saturation effects typical of multispectral data while ensuring the spatial consistency in the output predictions due to the removal of artifacts and data voids present in ArcticCHM data. While the main contribution of this study is to provide the first evidence of the best-case-scenario (i.e., availability of accurate terrain models) that ArcticDEM data can provide for large-scale AGB modelling, it remains critically important for other studies to investigate how ArcticDEM may be used in areas where no DTMs are available as is the case for large portions of the boreal zone.

space-borne imagery

The NASA Open Science Data Repository: Biomedical Fair Data, Analysis Tools, User Communities, Publications, and Discoveries for Deep Space Missions

Increased biomedical risks and challenges associated with deep space missions require new knowledge discovery, new health countermeasures, and development of novel ecosystems, life support, crop production, and biomedical support capabilities. To meet NASA’s Moon to Mars strategic program goals for Human and Biological Sciences, findable, accessible, interoperable, reusable (FAIR), and maximally open-access data is going to be required to enable humanity to thrive in deep space. Indeed, this cornerstone perspective on FAIR and maximally open access data was also recommended in the recent 2023-2032 Decadal Survey from the National Academies of Sciences, Engineering, and Medicine. The NASA Open Science Data Repository (OSDR) is a maximally open access and FAIR database, and meets various scientific, technical, and operational spaceflight needs. It offers public users and submitters the ability to upload, download, search, share, analyze, and visualize data across ‘omics, physiological, phenotypic, behavioral, bioimaging, video, and environmental monitoring telemetry datasets. OSDR includes NASA GeneLab, NASA Ames Life Sciences Data Archive, and the NASA Biological Institutional Scientific Collection. OSDR has >455 studies with datasets from model organisms and non-NASA human astronauts. There are ~12 datasets from the Inspiration 4 (I4) mission, spanning metagenomics, comprehensive metabolic panels, clonal hematopoiesis, spatial transcriptomics, proteomics, and cytokine panels. In the interest of data privacy, two I4 datasets have raw FASTQ and FASTA files relating to the epitranscriptome, and a new request feature is live in OSDR (with a backend review process established) which was developed based on industry norms. OSDR also recently began a collaboration with the European Space Agency (ESA) to scientifically curate and make available >200 terabytes of human and model organism space-relevant data. The OSDR submission portal is designed to ingest and curate ~25 ‘omics assay data types, and ~50 physiological-phenotypic-imaging assay data types, spanning ultrasonography, micro-computed tomography, histology, morphometric photography, rebound tonometry, gait analysis, optical coherence tomography, novel object recognition, flow cytometry, and immunohistochemistry. A suite of analysis tools are available for OSDR users including: 1) an Environmental Data Application to compare radiation, CO2, relative humidity, temperature, and other telemetry across missions and subjects, 2) the RadLab database, a collaboration between NASA, ESA, the German and Italian Space Agencies, and the Bulgarian Academy of Sciences, which compiles radiation measurements relevant to human spaceflight and provides tools for accessing and manipulating the data, and 3) a Multi-study visualization tool which enables users to look across and combine GeneLab’s omics datasets across different experiments and missions. There are ~600 volunteer OSDR Analysis Working Group (AWG) members who: 1) provide feedback on scientific standards for reuse (subject and assay metadata; processing pipelines; dataset formats and uniformed structures for machine-readability), and 2) collaborate to mine-reuse OSDR data conducting scientific analysis. OSDR has enabled 60 publications as of September 2023, many directly from AWG collaborations most notably the Cell Press package in 2020. Lastly, there are at least 15 articles which mine OSDR data part of a package of ~50 articles across Nature Portfolio with research stemming from I4, the Japan Aerospace Exploration Agency, NASA Space Biology, and the NASA Human Research Program.

space biology

NASA Open Science Data Repository: Biomedical FAIR Data, Analysis Tools, User Communities, and Discoveries for Deep Space Missions

Increased biomedical risks and challenges associated with deep space missions require new knowledge discovery, new health countermeasures, and development of novel ecosystems, life support, crop production, and biomedical support capabilities. To meet NASA’s Moon to Mars strategic program goals for Human and Biological Sciences, findable, accessible, interoperable, reusable (FAIR), and maximally open-access data is going to be required to enable humanity to thrive in deep space. Indeed, this cornerstone perspective on FAIR and maximally open access data was also recommended in the recent 2023-2032 Decadal Survey from the National Academies of Sciences, Engineering, and Medicine. The NASA Open Science Data Repository (OSDR) is a maximally open access and FAIR database, and meets various scientific, technical, and operational spaceflight needs. It offers public users and submitters the ability to upload, download, search, share, analyze, and visualize data across ‘omics, physiological, phenotypic, behavioral, bioimaging, video, and environmental monitoring telemetry datasets. OSDR includes NASA GeneLab, NASA Ames Life Sciences Data Archive, and the NASA Biological Institutional Scientific Collection. OSDR has >455 studies with datasets from model organisms and non-NASA human astronauts. There are ~12 datasets from the Inspiration 4 (I4) mission, spanning metagenomics, comprehensive metabolic panels, clonal hematopoiesis, spatial transcriptomics, proteomics, and cytokine panels. In the interest of data privacy, two I4 datasets have raw FASTQ and FASTA files relating to the epitranscriptome, and a new request feature is live in OSDR (with a backend review process established) which was developed based on industry norms. OSDR also recently began a collaboration with the European Space Agency (ESA) to scientifically curate and make available >200 terabytes of human and model organism space-relevant data. The OSDR submission portal is designed to ingest and curate ~25 ‘omics assay data types, and ~50 physiological-phenotypic-imaging assay data types, spanning ultrasonography, micro-computed tomography, histology, morphometric photography, rebound tonometry, gait analysis, optical coherence tomography, novel object recognition, flow cytometry, and immunohistochemistry. A suite of analysis tools are available for OSDR users including: 1) an Environmental Data Application to compare radiation, CO2, relative humidity, temperature, and other telemetry across missions and subjects, 2) the RadLab database, a collaboration between NASA, ESA, the German and Italian Space Agencies, and the Bulgarian Academy of Sciences, which compiles radiation measurements relevant to human spaceflight and provides tools for accessing and manipulating the data, and 3) a Multi-study visualization tool which enables users to look across and combine GeneLab’s omics datasets across different experiments and missions. There are ~600 volunteer OSDR Analysis Working Group (AWG) members who: 1) provide feedback on scientific standards for reuse (subject and assay metadata; processing pipelines; dataset formats and uniformed structures for machine-readability), and 2) collaborate to mine-reuse OSDR data conducting scientific analysis. OSDR has enabled 60 publications as of September 2023, many directly from AWG collaborations most notably the Cell Press package in 2020. Lastly, there are at least 15 articles which mine OSDR data part of a package of ~50 articles across Nature Portfolio with research stemming from I4, the Japan Aerospace Exploration Agency, NASA Space Biology, and the NASA Human Research Program.

open access

Building a FAIR data ecosystem for incorporating single-cell transcriptomics data into agricultural genome to phenome research

Introduction The agriculture genomics community has numerous data submission standards available, but the standards for describing and storing single-cell (SC, e.g., scRNA- seq) data are comparatively underdeveloped. Methods To bridge this gap, we leveraged recent advancements in human genomics infrastructure, such as the integration of the Human Cell Atlas Data Portal with Terra, a secure, scalable, open-source platform for biomedical researchers to access data, run analysis tools, and collaborate. In parallel, the Single Cell Expression Atlas at EMBL-EBI offers a comprehensive data ingestion portal for high-throughput sequencing datasets, including plants, protists, and animals (including humans). Developing data tools connecting these resources would offer significant advantages to the agricultural genomics community. The FAANG data portal at EMBL-EBI emphasizes delivering rich metadata and highly accurate and reliable annotation of farmed animals but is not computationally linked to either of these resources. Results Herein, we describe a pilot-scale project that determines whether the current FAANG metadata standards for livestock can be used to ingest scRNA-seq datasets into Terra in a manner consistent with HCA Data Portal standards. Importantly, rich scRNA-seq metadata can now be brokered through the FAANG data portal using a semi-automated process, thereby avoiding the need for substantial expert curation. We have further extended the functionality of this tool so that validated and ingested SC files within the HCA Data Portal are transferred to Terra for further analysis. In addition, we verified data ingestion into Terra, hosted on Azure, and demonstrated the use of a workflow to analyze the first ingested porcine scRNA-seq dataset. Additionally, we have also developed prototype tools to visualize the output of scRNA-seq analyses on genome browsers to compare gene expression patterns across tissues and cell populations. This JBrowse tool now features distinct tracks, showcasing PBMC scRNA-seq alongside two bulk RNA-seq experiments. Discussion We intend to further build upon these existing tools to construct a scientist-friendly data resource and analytical ecosystem based on Findable, Accessible, Interoperable, and Reusable (FAIR) SC principles to facilitate SC-level genomic analysis through data ingestion, storage, retrieval, re-use, visualization, and comparative annotation across agricultural species.

Genetics & Heredity