Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “database for machine learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Observations on the Application of Machine Learning Techniques to Aviation Operations

There is an increasing interest in applying methods based on Machine Learning Techniques (MLT) to problems in aviation operations. The current interest is based on developments in Cloud Computing, the availability of open software and the success of MLT in automation, consumer behavior and finance involving large database. Historically aviation operations have been analyzed using physics-based models and provide information for making operational decisions. This talk describes issues to be addressed in applying either model-driven or data-driven methods. Aviation operations involving many decision makers, multiple objectives, poor or unavailable physics-based models and a rich historical database are prime candidates for analysis using data-driven methods. The issues are illustrated by a detailed example and summary of current research in the area. The application of MLT to aviation operations falls into two categories: (a) based on the lack of a physics-based model, MLT is the favored approach and (b) marginal difference between regression methods using physics-based models and MLT. Further research is needed in the selection of MLT to critical aviation operations. As always, the best approach depends on the task, the physical understanding of the problem and the quality and quantity of the available data.

Sridhar, Banavar↗

Application of Machine Learning Techniques to Aviation Operations: A Case Study

There is an increasing interest in applying methods based on Machine Learning Techniques (MLT) to problems in aviation operations. The current interest is based on developments in Cloud Computing, the availability of open software and the success of MLT in automation, consumer behavior and finance involving large database. Historically aviation operations have been analyzed using physics-based models and provide information for making operational decisions. This talk describes issues to be addressed in applying either model-driven or data-driven methods. Aviation operations involving many decision makers, multiple objectives, poor or unavailable physics-based models and a rich historical database are prime candidates for analysis using data-driven methods. The issues are illustrated by a detailed example and summary of current research in the area. The application of MLT to aviation operations falls into two categories 58; (a) based on the lack of a physics-based model, MLT is the favored approach and (b) marginal difference between regression methods using physics-based models and MLT. Further research is needed in the selection of MLT to critical aviation operations. As always, the best approach depends on the task, the physical understanding of the problem and the quality and quantity of the available data.

Sridhar, Banavar↗

Application of Machine Learning Techniques to Aviation Operations: NASA Case Studies

There is an increasing interest in applying methods based on Machine Learning Techniques(MLT) to problems in aviation operations. The current interest is based on developments in Cloud Computing, the availability of open software and the success of MLT in automation, consumer behavior and finance involving large database. Historically aviation operations have been analyzed using physics-based models and provide information for making operational decisions. This talk describes issues to be addressed in applying either model-driven or data-driven methods. Aviation operations involving many decision makers, multiple objectives, poor or unavailable physics-based models and a rich historical database are prime candidates for analysis using data-driven methods. The issues relating to data, feature selection and validation of the models are illustrated by examining case studies of the application of MLT to problems in air traffic management at NASA. Further research is needed in the application of MLT to critical aviation operations. As always, the best approach depends on the task, the physical understanding of the problem and the quality and quantity of the available data.

Sridhar, Banavar↗

The Radiation Biology Ontology: A New Tool Supporting FAIR Principles Across Radiation Biology Facilitating Data Discovery and Integration

Development of the Radiation Biology Ontology (RBO) was motivated by the need for a comprehensive, well-structured ontology for encoding radiation biology metadata. The primary use-cases were archiving data in the STORE database (https://www.storedb.org/), the repository for the RadoNorm Project, and in GeneLab (https://genelab.nasa.gov), NASA’s ‘omics database. The scope of radiobiology research ranges from physics to radiation oncology to socio-legal studies; no existing ontology has the necessary breadth or depth. In addition, a formal ontology has the advantage of being usable for machine learning and, importantly, for tasks like data integration, knowledge extraction from the scientific literature and for query extension and data classification. Standardisation of metadata is one of the primary objectives of the FAIR principles for open data; RBO is an important landmark for FAIR radiation biology data.

ontology↗

Datascope to Enable Earth Independent Medical Operations (EIMO)

BACKGROUND: NASA has amassed sixty years of knowledge and experience relevant to maintenance of crew health and performance in low earth orbit. The Apollo Program introduced the importance of ensuring progressively autonomous operational capability. Earth Independent Medical Operations (EIMO) will require a gradual shift in the balance of medical responsibility, management, and authority from terrestrial to space-based assets. Terrestrial assets will continue to be essential for pre-mission screening and planning in addition to maintenance of crew health and performance. However, new capabilities are needed to enable EIMO and the amount of data required to support these systems, and mitigate the impacts of data transmission delays and reduced bandwidth coupled with lack of cloud-like resources and on-board computing capacity that is currently unclear or operationally insufficient. OVERVIEW: The overall goal of EIMO is to develop artificial intelligence (AI)-based solutions to analyze crew health and performance data utilizing a clinical decision support system (CDSS) to provide crew medical officers (CMO) with the equivalent of real-time, on-board medical consults. The EIMO ecosystem is envisioned as a “system of systems” where embedded reference databases and real-time data streams from multiple input vectors continuously and seamlessly assess crew health and performance. EIMO will be designed to make recommendations to the CMO using multi-modal AI-based natural language processing and machine learning methods with interoperability to push/pull data within and between multiple vehicle and habitat architectures. DISCUSSION: Data flows and storage/retrieval capacity are severely constrained during space missions and the challenges will become even greater during exploration missions. Just as each past program from Mercury to the International Space Station (ISS) required rethinking the interaction between ground-based controllers and space-based crew, so too will future missions to the Moon and Mars. While the NASA High-Performance Spaceflight Computing Processor project aims to increase computational capacity by 100 times over current spaceflight computers, the projected deliverable still lags considerably behind what will be needed to enable an AI-driven CDSS. Restrictions in processing speed and data storage capacity, coupled with transmission bottlenecks and delays, necessitate definition and optimization of an integrated data architecture to enable a progressively autonomous medical capability.

Medical operations↗

Objectively Identifying Transverse Cirrus Bands in Tropical Cyclones using a Convolutional Neural Network

Transverse cirrus bands (TCBs) are bands of upper-level clouds that are regularly seen in mesoscale and synoptic-scale weather systems. In tropical cyclones, their appearance has been subjectively linked to intensification and the diurnal cycle, but these hypothesized relationships have not been rigorously tested due to the difficulty of objectively identifying TCBs in satellite images. This presentation describes a machine learning technique that successfully identifies TCBs in imagery from the GOES-16 Advanced Baseline Imager. The technique uses a convolutional neural network (CNN) that assigns a probability of each pixel in the image being associated with a TCB. Using the CNN, a database of TCBs from 2018 to 2022 was developed for the Atlantic basin. Statistics using this database will be presented, including the relationship between TCBs and deep-layer vertical wind shear, intensity change, and the TC diurnal cycle.

John Mark Mayhall↗

Automatic cataloguing and characterization of Earth science data using SE-trees

In the future, NASA's Earth Observing System (EOS) platforms will produce enormous amounts of remote sensing image data that will be stored in the EOS Data Information System. For the past several years, the Intelligent Data Management group at Goddard's Information Science and Technology Office has been researching techniques for automatically cataloguing and characterizing image data (ADCC) from EOS into a distributed database. At the core of the approach, scientists will be able to retrieve data based upon the contents of the imagery. The ability to automatically classify imagery is key to the success of contents-based search. We report results from experiments applying a novel machine learning framework, based on Set-Enumeration (SE) trees, to the ADCC domain. We experiment with two images: one taken from the Blackhills region in South Dakota; and the other from the Washington DC area. In a classical machine learning experimentation approach, an image's pixels are randomly partitioned into training (i.e. including ground truth or survey data) and testing sets. The prediction model is built using the pixels in the training set, and its performance is estimated using the testing set. With the first Blackhills image, we perform various experiments achieving an accuracy level of 83.2 percent, compared to 72.7 percent using a Back Propagation Neural Network (BPNN) and 65.3 percent using a Gaussain Maximum Likelihood Classifier (GMLC). However, with the Washington DC image, we were only able to achieve 71.4 percent, compared with 67.7 percent reported for the BPNN model and 62.3 percent for the GMLC.

Rymon, Ron↗

Virtual Assistant for First Responders Using Natural Language Understanding and Optical Character Recognition

Commercial deep learning capabilities are available for many applications such as computer vision processing and intelligent chat bots. The Google Cloud Platform product Google Dialogflow provides lifelike conversational artificial intelligence (AI) using machine learning (ML) to generate natural conversations between computers and humans. This ML utilizes natural language understanding (NLU) to recognize a user’s intent and extracts key information into a form of entities. We have developed a user-friendly application through understanding the hazardous material database, first aid safety guidelines and observing the process of first responders who access this information in the field. We created the Trusted and Explainable Artificial Intelligence for Saving Lives (TruePAL) virtual assistant using Dialogflow1 and TensorFlow2 paired with EasyOCR.3 The chatbot supports first responders by providing voice interaction which helps limit additional steps such as browsing through multiple categories when searching for information. Using feedback from our field interviews, the voice interface has been developed to enable the first responder to focus on the immediate emergency. With less distractions, the first responder is able to engage the incident more effectively. The partial hands-free TruePAL chatbot assistant improves the accessibility to the correct guidance by an average of 1.9 seconds compared to the widely used application, NIH WISER, which requires full attention to operate. We combined this intelligent chatbot with a separate visual processing capability to produce hazardous signage analysis and generate the proper guidance for first responders. With the evolving functionality of AI tools, the use of virtual assistants in first responder technology will be an advancement, benefiting the safety of both first responders and civilians.

Chow, Edward↗

A Chlorophyll-a Algorithm for Landsat-8 Based on Mixture Density Networks

Retrieval of aquatic biogeochemical variables, such as the near-surface concentration of chlorophyll-a (Chla) in inland and coastal waters via remote observations, has long been regarded as a challenging task. This manuscript applies Mixture Density Networks (MDN) that use the visible spectral bands available by the Operational Land Imager (OLI) aboard Landsat-8 to estimate Chla. We utilize a database of co-located in situ radiometric and Chla measurements (N = 4,354), referred to as Type A data, to train and test an MDN model (MDN(A)). This algorithm’s performance, having been proven for other satellite missions, is further evaluated against other widely used machine learning models (e.g., support vector machines), as well as other domain-specific solutions (OC3), and shown to offer significant advancements in the field. Our performance assessment using a held-out test data set suggests that a 49% (median) accuracy with near-zero bias can be achieved via the MDN(A) model, offering improvements of 20 to 100% in retrievals with respect to other models. The sensitivity of the MDN(A) model and benchmarking methods to uncertainties from atmospheric correction (AC) methods, is further quantified through a semi-global matchup dataset (N = 3,337), referred to as Type B data. To tackle the increased uncertainties, alternative MDN models (MDN(B)) are developed through various features of the Type B data (e.g., Rayleigh-corrected reflectance spectra ρ(s)). Using held-out data, along with spatial and temporal analyses, we demonstrate that these alternative models show promise in enhancing the retrieval accuracy adversely influenced by the AC process. Results lend support for the adoption of MDN(B) models for regional and potentially global processing of OLI imagery, until a more robust AC method is developed. Index Terms—Chlorophyll-a, coastal water, inland water, Landsat-8, machine learning, ocean color, aquatic remote sensing.

Brandon Smith↗

Simultaneous Retrieval of Selected Optical Water Quality Indicators From Landsat-8, Sentinel-2, and Sentinel-3

Constructing multi-source satellite-derived water quality (WQ) products in inland and nearshore coastal waters from the past, present, and future missions is a long-standing challenge. Despite inherent differences in sensors’ spectral capability, spatial sampling, and radiometric performance, research efforts focused on formulating, implementing, and validating universal WQ algorithms continue to evolve. This research extends a recently developed machine-learning (ML) model, i.e., Mixture Density Networks (MDNs) (Pahlevan et al., 2020; Smith et al., 2021), to the inverse problem of simultaneously retrieving WQ indicators, including chlorophyll-a (Chla), Total Suspended Solids (TSS), and the absorption by Colored Dissolved Organic Matter at 440 nm (a cdom (440)), across a wide array of aquatic ecosystems. We use a database of in situ measurements to train and optimize MDN models developed for the relevant spectral measurements (400–800 nm) of the Operational Land Imager (OLI), MultiSpectral Instrument (MSI), and Ocean and Land Color Instrument (OLCI) aboard the Landsat-8, Sentinel-2, and Sentinel-3 missions, respectively. Our two performance assessment approaches, namely hold-out and leave-one-out, suggest significant, albeit varying degrees of improvements with respect to second-best algorithms, depending on the sensor and WQ indicator (e.g., 68%, 75%, 117% improvements based on the hold-out method for Chla, TSS, and a cdom (440), respectively from MSI-like spectra). Using these two assessment methods, we provide theoretical upper and lower bounds on model performance when evaluating similar and/or out-of-sample datasets. To evaluate multi-mission product consistency across broad spatial scales, map products are demonstrated for three near-concurrent OLI, MSI, and OLCI acquisitions. Overall, estimated TSS and a cdom (440) from these three missions are consistent within the uncertainty of the model, but Chla maps from MSI and OLCI achieve greater accuracy than those from OLI. By applying two different atmospheric correction processors to OLI and MSI images, we also conduct matchup analyses to quantify the sensitivity of the MDN model and best-practice algorithms to uncertainties in reflectance products. Our model is less or equally sensitive to these uncertainties compared to other algorithms. Recognizing their uncertainties, MDN models can be applied as a global algorithm to enable harmonized retrievals of Chla, TSS, and a cdom (440) in various aquatic ecosystems from multi-source satellite imagery. Local and/or regional ML models tuned with an apt data distribution (e.g., a subset of our dataset) should nevertheless be expected to outperform our global model.

Machine learning↗

Managing the Digital Thread for Structural Applications With Fit for Purpose Materials

With the increased emphasis on reducing the cost and time to market of new materials, the need for analytical tools that enable the virtual design and optimization of materials throughout their processing - internal structure - property - performance envelope, along with the capturing and storing of the associated material and model information across its lifecycle, has become critical. This need is also fueled by the demands for higher efficiency in material testing; consistency, quality and traceability of data; product design; engineering analysis; as well as control of access to proprietary or sensitive information. Consequently, at NASA Glenn Research Center a robust information management system that manages the digital thread across the full material life (i.e., capture, analysis, maintenance, and dissemination of data) cycle directed at the design of ‘fit-for-purpose materials’ is under development. To this end the Application Table has been incorporated within NASA Glenn Research Center’s ICME Information Management framework within the ANSYS Granta MI tool. The Application Table provides a place where material and structural application information/requirements can be linked to marry the “design-the-material” (structural engineering) and the “design-with-material” (material science) paradigms and thereby enable application-driven design and optimization of materials and structures. In additional several associated toolsets, specifically: AIMAOS (Automated Information Management Across Organizations and Scales), Py MILab, and JARIMIS (Just A Rather Intelligent Material Interrogation System) are also under development to assist in the judicious automation of this process. AIMOAS offers users an interactive graphical user interface for connecting material information management systems with both commercial and in-house simulation tools at various length scales to enable such automation in the handoff across scales and maintenance of material digital twins and the digital thread. Py MILab, is an automatic framework for the capture, analysis, maintenance, and storage of material test data. Py MILab uses a modular approach for capturing raw data, analyzing the data, and storing the data in a database, interfaced by neutral file structures, to promote plug-and-play capabilities for various analysis types. Finally, JARIMIS is an expert system that integrates various materials informatics tools (e.g., MicroNet, Surrogate ML models, ANSYS Granta MI, etc.) to enable inverse design of materials and facilitate the application of machine learning (ML) and data science with human in the loop decision making to rapidly discover and optimize new materials.

Digital Transformation↗

Robust Algorithm for Estimating Total Suspended Solids (TSS) in Inland and Nearshore Coastal Waters

One of the challenging tasks in modern aquatic remote sensing is the retrieval of near-surface concentrations of Total Suspended Solids (TSS). This study aims to present a Statistical, inherent Optical property (IOP) -based, and muLti-conditional Inversion proceDure (SOLID) for enhanced retrievals of satellite-derived TSS under a wide range of in-water bio-optical conditions in rivers, lakes, estuaries, and coastal waters. In this study, using a large in situ database (N > 3500), the SOLID model is devised using a three-step procedure: (a) water-type classification of the input remote sensing reflectance (R(sub rs)), (b) retrieval of particulate backscattering (b(sub bp)) in the red or near-infrared (NIR) regions using semi-analytical, machine-learning, and empirical models, and (c) estimation of TSS from b(sub bp) via water-type-specific empirical models. Using an independent subset of our in situ data (N = 2729) with TSS ranging from 0.1 to 2626.8 [g/m (exp 3)], the SOLID model is thoroughly examined and compared against several state-of-the-art algorithms (Miller and McKee, 2004; Nechad et al., 2010; Novoa et al., 2017; Ondrusek et al., 2012; Petus et al., 2010). We show that SOLID outperforms all the other models to varying degrees, i.e., from 10 to > 100%, depending on the statistical attributes (e.g., global versus water-type-specific metrics). For demonstration purposes, the model is implemented for images acquired by the MultiSpectral Imager aboard Sentinel-2A/B over the Chesapeake Bay, San-Francisco-Bay-Delta Estuary, Lake Okeechobee, and Lake Taihu. To enable generating consistent, multimission TSS products, its performance is further extended to, and evaluated for, other missions, such as the Ocean and Land Color Instrument (OLCI), Moderate Resolution Imaging Spectroradiometer (MODIS), Visible Infrared Imaging Radiometer Suite (VIIRS), and Operational Land Imager (OLI). Sensitivity analyses on uncertainties induced by the atmospheric correction indicate that 10% uncertainty in Rrs leads to < 20% uncertainty in TSS retrievals from SOLID. While this study suggests that SOLID has a potential for producing TSS products in global coastal and inland waters, our statistical analysis certainly verifies that there is still a need for improving retrievals across a wide spectrum of particle loads.

Total suspended solids↗

ASCoT 3: Nonlinear Principal Components Analysis and Uncertainty Quantification in Early Concept Spacecraft Flight Software Cost Estimation

For mission planners and evaluators alike, value in cost models comes from a mean or median prediction, an understanding of the uncertainty on that prediction, and an understanding of model performance. Here we apply advanced statistical and machine learning methods to spacecraft flight software cost, effort, and SLOC estimation, and present the results in the latest version of the Analogy Software Cost Tool (ASCoT). We present in- and out-of-sample performance metrics for our models, each of which incorporate some amount of epistemic uncertainty. ASCoT, hosted on the One NASA Cost Engineering (ONCE) database via the Online NASA Space Estimation Tool (ONSET), was first showcased in 2016 as a number of analogy-based models and methods (kNN and Clustering) to support early project formulation. This ASCoT update improves upon the previous analogic methods by incorporating uncertainty in the data transformations. In particular, we use a Nonlinear Principal Components Analysis (NLPCA) to deal with ordinal data.

Robotic Spacecraft↗

Time series comparisons in Deep Space Network

The Deep Space Network (DSN) is NASA’s international array of antennas that support interplanetary spacecraft missions. DSN provides radar and radio astronomy observations that enhance our understanding of the solar system and the larger universe. A track is a block of continuous multi-dimensional time series from the beginning to end of DSN communication with the target spacecraft, containing 129 monitor data items lasting several hours at a frequency of 0.2-1Hz. Monitor data on each track reports on the performance of specific spacecraft operations and the DSN itself. DSN is receiving signals from 32 spacecraft across the solar system. DSN has pressure to reduce costs while maintaining the quality of support for DSN mission users. DSN operators need to simultaneously monitor multiple tracks and identify anomalies in real time. DSN has seen that as the number of missions increases, the data that needs to be processed increases over time. In this project, we look at the last 8 years of data for analysis. Any anomaly in the track indicates a problem with either the spacecraft, DSN equipment, or weather conditions. DSN operators typically write “discrepancy reports” for further analysis. It is recognized that it would be quite helpful to identify 10 similar historical tracks out of the huge database to quickly find/match anomalies. This tool has three functions: (1) identification of the top 10 similar historical tracks, (2) detection of anomalies compared to the reference normal track, and (3) comparison of statistical differences between two given tracks. The requirements for these features were confirmed by survey responses from 21 DSN operators and engineers. The preliminary machine learning model has shown promising performance (AUC=0.92). We plan to increase the number of data sets and perform additional testing to improve performance further before its planned integration into the Track Visualizer to assist DSN field operators and engineers.

Rebbapragada, Umaa↗

Remote Sensing of CDOM, CDOM Spectral Slope, and Dissolved Organic Carbon in the Global Ocean

A Global Ocean Carbon Algorithm Database (GOCAD) has been developed from over 500 oceanographic field campaigns conducted worldwide over the past 30 years including in situ reflectances and coincident satellite imagery, multi- and hyperspectral Chromophoric Dissolved Organic Matter (CDOM) absorption coefficients from 245–715 nm, CDOM spectral slopes in eight visible and ultraviolet wavebands, dissolved and particulate organic carbon (DOC and POC, respectively), and inherent optical, physical, and biogeochemical properties. From field optical and radiometric data and satellite measurements, several semi-analytical, empirical, and machine learning algorithms for retrieving global DOC, CDOM, and CDOM slope were developed, optimized for global retrieval, and validated. Global climatologies of satellite-retrieved CDOM absorption coefficient and spectral slope based on the most robust of these algorithms lag seasonal patterns of phytoplankton biomass belying Case 1 assumptions, and track terrestrial runoff on ocean basin scales. Variability in satellite retrievals of CDOM absorption and spectral slope anomalies are tightly coupled to changes in atmospheric and oceanographic conditions associated with El Niño Southern Oscillation (ENSO), strongly covary with the multivariate ENSO index in a large region of the tropical Pacific, and provide insights into the potential evolution and feedbacks related to sea surface dissolved carbon in a warming climate. Further validation of the DOC algorithm developed here is warranted to better characterize its limitations, particularly in mid-ocean gyres and the southern oceans.

Dissolved organic carbon↗

Machine Learning Based AFP Inspection: A Tool for Characterization and Integration

Automated Fiber Placement (AFP) has become a standard manufacturing technique in the creation of large scale composite structures due to its high production rates. However, the associated rapid layup that accompanies AFP manufacturing has a tendency to induce defects. We forward an inspection system that utilizes machine learning (ML) algorithms to locate and characterize defects from profilometry scans coupled with a data storage system and a user interface (UI) that allows for informed manufacturing. A Keyence LJ-7080 blue light profilometer is used for fast 2D height profiling. After scans are collected, they are process by ML algorithms, displayed to an operator through the UI, and stored in a database. The overall goal of the inspection system is to add an additional tool for AFP manufacturing. Traditional AFP inspection is done manually adding to manufacturing time and being subject to inspector errors or fatigue. For large parts, the inspection process can be cumbersome. The proposed inspection system has the capability of accelerating this process while still keeping a human inspector integrated and in control. This allows for the rapid capability of the automated inspection software and the robustness of a human checking for defects that the system either missed or misclassified.

Sacco, Christopher↗

Objectively Identifying Transverse Cirrus Bands in Tropical Cyclones using a Convolutional Neural Network

Transverse cirrus bands (TCBs) are bands of upper-level clouds regularly seen in mesoscale and synoptic-scale weather systems. In tropical cyclones, their appearance has been subjectively linked to intensification and the diurnal cycle. However, these hypothesized relationships have not been rigorously tested due to the subjective nature of TCBs in satellite images. A machine learning technique that successfully identifies TCBs objectively in imagery from the GOES-16 Advanced Baseline Imager (ABI) has been developed to solve this problem. The technique uses a U-Net convolutional neural network (CNN) that assigns a probability to each pixel in an image based on the likelihood of the pixel being associated with a TCB. Using the U-Net CNN, a database of TCBs from 2019 to 2022 was developed for the Atlantic tropical cyclone basin by defining an appropriate probability threshold that defines the difference between TCB and non-TCB pixels. This threshold is where the Jaccard score, calculated using manually identified TCBs and model identified TCBs, is maximized. Statistics for TCB occurrence will also be presented, including the relationships between TCBs and storm relative motion, shear relative direction, cardinal direction, tropical cyclone intensity, tropical cyclone intensification rates, and time of day.

John Mark Mayhall↗

Data Sharing in Radiobiology; Towards FAIR

The value of scientific data depends on their findability, accessibility, integrability and reusability according to the FAIR principles. Together with the sustainability of data preservation and access, these principles underpin the long term benefits of scientific research. Within the domain of radiobiology we have a huge array of data types, themes and complexities which make standardisation of metadata, data structure and data integration very challenging. Moreover, it is clear that, for example, in the area of disaster preparedness, the ready discovery and availability of multiple types of data, for example on biological effects of exposure, climatology, ecology, human behavioural and attitudinal studies, is important for an integrated scientific approach. Because these data are spread over many databases, journal supplementary information resources and even the computers of the investigators, their discovery and reuse can be challenging. Despite exhortations from funding agencies and scientific institutions over the past two decades there is still a serious deficit in the willingness and in some cases the ability of investigators to share data, and although much may not be formally „Public domain“, information about the existence of the data, their metadata, and how to obtain them should always be available. We report the progress of work on three databases, the STORE and the NASA GeneLab and LSDA repositories to leverage the Radiation Biology Ontology (RBO), a structured terminology for metadata that can be used by all radiation biology-relevant databases to unite federated and automated data searches across multiple databases, for example using web services, and through semantic web technologies supporting data discovery. The initial primary use-cases for RBO were archiving data in the STORE database (https://www.storedb.org/), the repository used for the RadoNorm and Pianoforte Projects among others, and in the NASA Open Science Data Repository (https://osdr.nasa.gov/bio). The scope of radiobiology research ranges from basic physics to radiation oncology to sociolegal studies; no existing ontology had the necessary breadth or depth to fulfill this need. In addition, a formal ontology has the advantage of being usable for machine learning and, importantly, for tasks like data integration, knowledge extraction from the scientific literature and for query extension and data classification. Standardisation of metadata is one of the primary objectives of the FAIR principles for open data; RBO is an important landmark for FAIR-compliant radiation biology data sharing. The RBO is developed using the open-source tools of GitHub and the OBO Foundry-led Ontology Development Kit, and published through GitHub and the NIH/NCBI BioPortal website. This initial phase of concept modeling has yielded an ontology that has more than 300 declared concepts, with more than 3500 additional concepts imported from other OBO Foundry ontologies with relevance to radiation biology (for example, concepts from the ISO standard Basic Formal Ontology, the Environment Ontology and the Gene Ontology). We welcome input into the development of RBO and encourage its adoption.

ontologies↗