Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “metadata extraction”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Recovery, Restoration and Archiving of Previously Lost Data and Metadata from the Apollo Lunar Surface Experiments Package (ALSEP)

The Apollo Lunar Surface Experiments Package (ALSEP) is the name used to collectively represent the geophysical instruments deployed on the lunar surface by the astronauts on Apollo 12, 14, 15, 16, and 17. These instruments were active from the times of their deployment (November 1969 – December 1972) to September 1977. During that time, fourteen types of experiments were conducted, and their data were transmitted to Earth. The experiment PIs processed them. At the conclusion of the experiments, some of these data were submitted to the NASA Space Science Data Coordinated Archive (NSSDCA) for archiving, while others were not. The raw instrument data received from the Moon prior to March 1976 were not archived, either. The unarchived data, resided on open-reel magnetic tapes, became lost in the decades since, along with much of the metadata (the information necessary/useful in properly processing/analyzing the data). This article retraces the history of the ALSEP data archiving efforts in the 1970s, the subsequent loss of the data tapes, and the search, recovery, and restoration of the lost data by contemporary researchers in the 21st century. In 2006, NSSDCA began reformatting some of the ALSEP data archived in the 1970s to conform with the current Planetary Data System (PDS). In 2010, 440 of the previously lost magnetic tapes containing the raw ALSEP data were recovered. From these tapes, the data were extracted, re-packaged for individual experiments, and, for those with sufficient metadata, processed into higher order data readily usable by researchers. All of these data products have been recently archived with either PDS or NSSDCA. These newly restored data fill a number of gaps in the previously existing archive of the ALSEP data. In addition, tens of thousands of pages of Apollo era documents have been optically scanned and compiled into an online searchable catalog. This article also describes the content, organization, and usage of the restored raw ALSEP data and metadata.

S Nagihara↗

The Radiation Biology Ontology: A New Tool Supporting FAIR Principles Across Radiation Biology Facilitating Data Discovery and Integration

Development of the Radiation Biology Ontology (RBO) was motivated by the need for a comprehensive, well-structured ontology for encoding radiation biology metadata. The primary use-cases were archiving data in the STORE database (https://www.storedb.org/), the repository for the RadoNorm Project, and in GeneLab (https://genelab.nasa.gov), NASA’s ‘omics database. The scope of radiobiology research ranges from physics to radiation oncology to socio-legal studies; no existing ontology has the necessary breadth or depth. In addition, a formal ontology has the advantage of being usable for machine learning and, importantly, for tasks like data integration, knowledge extraction from the scientific literature and for query extension and data classification. Standardisation of metadata is one of the primary objectives of the FAIR principles for open data; RBO is an important landmark for FAIR radiation biology data.

ontology↗

Data Sharing in Radiobiology; Towards FAIR

The value of scientific data depends on their findability, accessibility, integrability and reusability according to the FAIR principles. Together with the sustainability of data preservation and access, these principles underpin the long term benefits of scientific research. Within the domain of radiobiology we have a huge array of data types, themes and complexities which make standardisation of metadata, data structure and data integration very challenging. Moreover, it is clear that, for example, in the area of disaster preparedness, the ready discovery and availability of multiple types of data, for example on biological effects of exposure, climatology, ecology, human behavioural and attitudinal studies, is important for an integrated scientific approach. Because these data are spread over many databases, journal supplementary information resources and even the computers of the investigators, their discovery and reuse can be challenging. Despite exhortations from funding agencies and scientific institutions over the past two decades there is still a serious deficit in the willingness and in some cases the ability of investigators to share data, and although much may not be formally „Public domain“, information about the existence of the data, their metadata, and how to obtain them should always be available. We report the progress of work on three databases, the STORE and the NASA GeneLab and LSDA repositories to leverage the Radiation Biology Ontology (RBO), a structured terminology for metadata that can be used by all radiation biology-relevant databases to unite federated and automated data searches across multiple databases, for example using web services, and through semantic web technologies supporting data discovery. The initial primary use-cases for RBO were archiving data in the STORE database (https://www.storedb.org/), the repository used for the RadoNorm and Pianoforte Projects among others, and in the NASA Open Science Data Repository (https://osdr.nasa.gov/bio). The scope of radiobiology research ranges from basic physics to radiation oncology to sociolegal studies; no existing ontology had the necessary breadth or depth to fulfill this need. In addition, a formal ontology has the advantage of being usable for machine learning and, importantly, for tasks like data integration, knowledge extraction from the scientific literature and for query extension and data classification. Standardisation of metadata is one of the primary objectives of the FAIR principles for open data; RBO is an important landmark for FAIR-compliant radiation biology data sharing. The RBO is developed using the open-source tools of GitHub and the OBO Foundry-led Ontology Development Kit, and published through GitHub and the NIH/NCBI BioPortal website. This initial phase of concept modeling has yielded an ontology that has more than 300 declared concepts, with more than 3500 additional concepts imported from other OBO Foundry ontologies with relevance to radiation biology (for example, concepts from the ISO standard Basic Formal Ontology, the Environment Ontology and the Gene Ontology). We welcome input into the development of RBO and encourage its adoption.

ontologies↗

Data Sharing in Radiation Biology: Towards FAIR

The value of scientific data depends on their findability, accessibility, integrability and reusability according to the FAIR principles. Together with the sustainability of data preservation and access, these principles underpin the long term benefits of scientific research. Within the domain of radiobiology we have a huge array of data types, themes and complexities which make standardisation of metadata, data structure and data integration very challenging. Moreover, it is clear that, for example, in the area of disaster preparedness, the ready discovery and availability of multiple types of data, for example on biological effects of exposure, climatology, ecology, human behavioural and attitudinal studies, is important for an integrated scientific approach. Because these data are spread over many databases, journal supplementary information resources and even the computers of the investigators, their discovery and reuse can be challenging. Despite exhortations from funding agencies and scientific institutions over the past two decades there is still a serious deficit in the willingness and in some cases the ability of investigators to share data, and although much may not be formally "Public domain“, information about the existence of the data, their metadata, and how to obtain them should always be available. We report the progress of work on three databases, the STORE and the NASA GeneLab and LSDA repositories to leverage the Radiation Biology Ontology (RBO), a structured terminology for metadata that can be used by all radiation biology-relevant databases to unite federated and automated data searches across multiple databases, for example using web services, and through semantic web technologies supporting data discovery. The initial primary use-cases for RBO were archiving data in the STORE database (https://www.storedb.org/), the repository used for the RadoNorm and Pianoforte Projects among others, and in the NASA Open Science Data Repository (https://osdr.nasa.gov/bio). The scope of radiobiology research ranges from basic physics to radiation oncology to sociolegal studies; no existing ontology had the necessary breadth or depth to fulfill this need. In addition, a formal ontology has the advantage of being usable for machine learning and, importantly, for tasks like data integration, knowledge extraction from the scientific literature and for query extension and data classification. Standardisation of metadata is one of the primary objectives of the FAIR principles for open data; RBO is an important landmark for FAIR-compliant radiation biology data sharing. The RBO is developed using the open-source tools of GitHub and the OBO Foundry-led Ontology Development Kit, and published through GitHub and the NIH/NCBI BioPortal website. This initial phase of concept modeling has yielded an ontology that has more than 300 declared concepts, with more than 3500 additional concepts imported from other OBO Foundry ontologies with relevance to radiation biology (for example, concepts from the ISO standard Basic Formal Ontology, the Environment Ontology and the Gene Ontology). We welcome input into the development of RBO and encourage its adoption.

ontologies↗

SPASE, Metadata, and the Heliophysics Virtual Observatories

To provide data search and access capability in the field of Heliophysics (the study of the Sun and its effects on the Solar System, especially the Earth) a number of Virtual Observatories (VO) have been established both via direct funding from the U.S. National Aeronautics and Space Administration (NASA) and through other funding agencies in the U.S. and worldwide. At least 15 systems can be labeled as Virtual Observatories in the Heliophysics community, 9 of them funded by NASA. The problem is that different metadata and data search approaches are used by these VO's and a search for data relevant to a particular research question can involve consulting with multiple VO's - needing to learn a different approach for finding and acquiring data for each. The Space Physics Archive Search and Extract (SPASE) project is intended to provide a common data model for Heliophysics data and therefore a common set of metadata for searches of the VO's. The SPASE Data Model has been developed through the common efforts of the Heliophysics Data and Model Consortium (HDMC) representatives over a number of years. We currently have released Version 2.1 of the Data Model. The advantages and disadvantages of the Data Model will be discussed along with the plans for the future. Recent changes requested by new members of the SPASE community indicate some of the directions for further development.

Thieman, James↗

The SPASE Data Model: A Metadata Standard for Registering, Finding, Accessing, and Using Heliophysics Data Obtained from Observations and Modeling

The Space Physics Archive Search and Extract Consortium has developed and implemented the SPASE Data Model that provides a common language for registering a wide range of Heliophysics data and other products. The Data Model enables discovery and access tools such that any researcher can obtain data easily, thereby facilitating research, including on space weather. The Data Model includes descriptions of Simulation Models and Numerical Output, pioneered by the Integrated Medium for Planetary Exploration (IMPEx) group in Europe, and subsequently adopted by the Community Coordinated Modeling Center (CCMC). The SPASE group intends to register all relevant Heliophysics data resources, including space-, ground-, and model-based. Substantial progress has been made, especially for space-based observational data and associated observatories, instruments, and display data. Legacy product registrations and access go back more than 50 years. Real-time data will be included. The National Aeronautics and Space Administration (NASA) portion of the SPASE group has funding that assures continuity in the upkeep of the Data Model and aids with adding new products. Tools are being developed for making and editing data descriptions. Digital Object Identifiers (DOIs) for Data Products can now be included in the descriptions. The data access that SPASE facilitates is becoming more uniform, and work is progressing on Web Service access via a standard Application Programming Interface. The SPASE Data Model is stable; changes over the past 9 years were additions of terms and capabilities that are backward compatible. This paper provides a summary of the history, structure, use, and future of the SPASE Data Model.

Roberts, D. Aaron↗

Application of a Dataset-Publication Knowledge Graph for Improving Earth Science Data Search

Finding a dataset at a NASA data center that is the best fit for the researcher’s application presents a challenge, not only for a novice user but for an experienced one, due to the data complexity and a multitude of choices of the existing data. Users often search for the data based on the application they are interested in, their research domain, phenomena, research topic, etc. As existing dataset metadata may not cover these search terms, the user may not obtain the most relevant results for their purpose. This problem was addressed by leveraging the content of the titles and abstracts of the research papers that utilize NASA datasets. For this, features from the paper titles and abstracts were extracted, and then a knowledge graph (KG) was used to link these features to the datasets used in that paper. The search for the datasets was tested by querying this knowledge graph through various terms extracted from Earth Science ontologies such as Semantic Web for Earth and Environment Technology (SWEET), and it was shown that this KG search outperforms the existing search that exclusively queries the dataset metadata.

Kristina Stoyanova↗

The Heliophysics Data Environment, Virtual Observatories, NSSDC, and SPASE

Heliophysics (the study of the Sun and its effects on the Solar System, especially the Earth) has an interesting data environment in that the data are often to be found in relatively small data sets widely scattered in archives around the world. Within the last decade there have been more concentrated efforts to organize the data access methods and create a Heliophysics Data and Model Consortium (HDMC). To provide data search and access capability a number of Virtual Observatories (VO's) have been established both via funding from the U.S. National Aeronautics and Space Administration (NASA) and through other funding agencies in the U.S. and worldwide. At least 15 systems can be labeled as Heliophysics Virtual Observatories, 9 of them funded by NASA. Other parts of this data environment include Resident Archives, and the final, or "deep" archive at the National Space Science Data Center (NSSDC). The problem is that different data search and access approaches are used by all of these elements of the HDMC and a search for data relevant to a particular research question can involve consulting with multiple VO's - needing to learn a different approach for finding and acquiring data for each. The Space Physics Archive Search and Extract (SPASE) project is intended to provide a common data model for Heliophysics data and therefore a common set of metadata for searches of the VO's and other data environment elements. The SPASE Data Model has been developed through the common efforts of the HDMC representatives over a number of years. We currently have released Version 2.1. of the Data Model. The advantages and disadvantages of the Data Model will be discussed along with the plans for the future. Recent changes requested by new members of the SPASE community indicate some of the directions for further development.

Thieman, James↗

Metadata Entry Optimization for NASA's Biological Institutional Scientific Collection (NBISC)

The NASA Biological Institutional Sample Collection (NBISC) at NASA’s Ames Research Center is a critical resource housing non-human samples collected from spaceflight missions and ground analog studies, primarily consisting of specimens from rats, mice, and select microbes. The primary objective of NBISC is to systematically receive, document, preserve, and facilitate access to these samples for the global scientific community. NBISC promotes international collaboration and maximizes the return on investment for precious tissues from spaceflight and analog experiments. Researchers can request physical samples through an online request form and subsequent written proposal review process. This study addresses two core research objectives: streamlining the NBISC sample lifecycle processes and strategizing for managing an influx of 50,000 tissue samples from a series of cosmic radiation analog experiments carried out at the NASA Space Radiation Laboratory (NSRL) by Drs. Eleanor Chang (Lawrence Berkeley Laboratory) and Polly Blakely (SRI). The Chang/Blakely studies investigated Harderian gland (HG) tumorigenesis in mice exposed to low dose and LET radiation comprising 8 different exposure protocols in over 4000 mice. NBISC sample metadata is stored in a Laboratory Information Management System (SLIMS). To streamline sample data entry, we customize python scripts using information extracted from the individual experimental protocols. The scripts automate entry into multiple SLIMS data fields including protocol name, unique sample barcode, tissue and sub-tissue information, freezer location, sample preservation method, etc. The semi-automated procedure significantly decreases the time spent on data entry by several orders of magnitude. Automation and data organization are essential, as they free up time for curation and promotion of the collection which, in turn, increase the accessibility of samples to the broader research community. NBISC benefits from streamlined data ingestion, and the methodologies developed here are applicable to other projects which use SLIMS including the NASA Biospecimen Sharing Program and GeneLab. As of Fall 2023, plans include transferring sample data from SLIMS to public facing repositories (OSDR and NLSP), expanding the reach of the Chang/Blakely sample collection. The Human Research Program Space Radiation Element plans to transfer non-human tissues from many more investigations to NBISC in the coming year.

Sample Repository↗

Metadata Entry Optimization For NASA's Biological Institutional Scientific Collection (NBISC)

The NASA Biological Institutional Sample Collection (NBISC) at NASA’s Ames Research Center is a critical resource housing non-human samples collected from spaceflight missions and ground analog studies, primarily consisting of specimens from rats, mice, and select microbes. The primary objective of NBISC is to systematically receive, document, preserve, and facilitate access to these samples for the global scientific community. NBISC promotes international collaboration and maximizes the return on investment for precious tissues from spaceflight and analog experiments. Researchers can request physical samples through an online request form and subsequent written proposal review process. This study addresses two core research objectives: streamlining the NBISC sample lifecycle processes and strategizing for managing an influx of 50,000 tissue samples from a series of cosmic radiation analog experiments carried out at the NASA Space Radiation Laboratory (NSRL) by Drs. Eleanor Chang (Lawrence Berkeley Laboratory) and Polly Blakely (SRI). The Chang/Blakely studies investigated Harderian gland (HG) tumorigenesis in mice exposed to low dose and LET radiation comprising 8 different exposure protocols in over 4000 mice. NBISC sample metadata is stored in a Laboratory Information Management System (SLIMS). To streamline sample data entry, we customize python scripts using information extracted from the individual experimental protocols. The scripts automate entry into multiple SLIMS data fields including protocol name, unique sample barcode, tissue and sub-tissue information, freezer location, sample preservation method, etc. The semi-automated procedure significantly decreases the time spent on data entry by several orders of magnitude. Automation and data organization are essential, as they free up time for curation and promotion of the collection which, in turn, increase the accessibility of samples to the broader research community. NBISC benefits from streamlined data ingestion, and the methodologies developed here are applicable to other projects which use SLIMS including the NASA Biospecimen Sharing Program and GeneLab. As of Fall 2023, plans include transferring sample data from SLIMS to public facing repositories (OSDR and NLSP), expanding the reach of the Chang/Blakely sample collection. The Human Research Program Space Radiation Element plans to transfer non-human tissues from many more investigations to NBISC in the coming year.

Biospecimen↗

ImageLabler: Labeling and Managing Image Data for Machine Learning in the Earth Sciences

While machine learning techniques for image classification have been around for a long time, storing and managing the vast number of images required as training data is still a problem for scientists. This is especially true for the field of Earth science, where only recently have experts begun using machine learning techniques for image-based phenomena classification. Image Labeler, a fast and scalable cloud-based tagging platform for Earth science images, seeks to improve upon existing methods of managing images and associated metadata, such as maintaining categorized folders of images on a local machine, a process that can be cumbersome and difficult to scale. The platform facilitates rapid development of image-based Earth science phenomena training datasets by allowing scientists to upload their existing imagery as well as extract new samples from open satellite imagery services made available through NASA’s Global Imagery Browse Service (GIBS). Image Labeler also supports GeoTIFF data, with capabilities such as displaying GeoTIFFs on an interactive map, drawing shapefiles over them, and tagging them with additional metadata. This allows scientists to perform spatiotemporal subsetting with geographic information and develop training data more quickly. Built using modern web technologies, Image Labeler includes additional capabilities such as team collaboration for large-scale image tagging projects. Users can download their data in a machine-learning-ready format, allowing scientists to spend time on experimentation rather than on the collection of training data. In this presentation, we demonstrate how Image Labeler seeks to become a one-stop image data management solution for machine learning applications in Earth science.

Ashish Acharya↗

Availability of Previously Unprocessed ALSEP Raw Instrument Data, Derivative Data, and Metadata Products

In year 2010, 440 original data archival tapes for the Apollo Lunar Science Experiment Package (ALSEP) experiments were found at the Washington National Records Center. These tapes hold raw instrument data received from the Moon for all the ALSEP instruments for the period of April through June 1975. We have recently completed extraction of binary files from these tapes, and we have delivered them to the NASA Space Science Data Cordinated Archive (NSSDCA). We are currently processing the raw data into higher order data products in file formats more readily usable by contemporary researchers. These data products will fill a number of gaps in the current ALSEP data collection at NSSDCA. In addition, we have estabilished a digital, searcheable archive of ALSEP document and metadata as part of the web portal of the Lunar and Planetary Institute. It currently holds approx. 700 documents totaling approx. 40,000 pages

ALSEP↗

Capturing, Analyzing, Maintaining, and Disseminating Shape Memory Material Data Between Information Management Systems

With an increased demand on reducing the time, cost, and effort to develop new materials, Integrated Computational Materials Engineering (ICME) has received widespread attention in various engineering disciplines as a catalyst for significantly reducing experimental testing during the material design process. An ICME approach to design can enable ‘fit-for-purpose’ materials to be realized in engineering applications by incorporating well-understood process-property-performance relationships between the various length and time scales in a material’s structure, enabling material optimization. However, such an approach requires validated multiscale models at the various length scales for a material, which in turn requires a large amount of data, a robust means of storing the data, and the ability to link data to developed material models. The NASA Vision 2040 [1] has identified nine key elements to enabling ICME approaches in system level design, with one being “Data, Information, and Visualization”, thus outlining the importance of a robust information management system for ICME. As the relationship between microstructure, properties, and material performance become better understood and incorporated into multiscale models that can be leveraged in application design, the emergence of new materials with application-driven properties can be realized. One such new material class that has seen growing attention are shape memory materials (SMM), in which a material can transition between a deformed and undeformed state via a reversible phase transformation when subject to a thermal, mechanical, or magnetic load [2]. SMMs have been used widely in aerospace and biomedical industries, including applications such as actuators, low-shock mechanisms, medical staples, braces, and stents [3, 4]. These materials exhibit unique behavior due to their ability to transition between phases, and thus the mechanisms that enable this transition must be captured in a data information management system and incorporated into SMM material models. At NASA Glenn Research Center, the Shape Memory Materials Database (SMMD) Tool has been developed to capture the necessary information that governs SMM material behavior and provide users the ability to select and visualize various SMMs for a specific application [5]. The database contains point-wise data for published SMM materials, along with the pedigree metadata for traceability necessary for a robust information management system. The database is also capable of storing in-house test data performed at NASA GRC by interacting with the developed Shape Memory Alloy (SMA) Analytics tool to extract the necessary point-wise values and populate the database. Although the SMMD Tool offers its users a single, authoritative source for SMM material data that is critical for model development and material design, the full material pedigree of the in-house test data for SMMs is not currently captured and is out of the scope for the SMMD tool. In this work, the schema for capturing SMM test data within the larger NASA GRC ICME Schema [6, 7, 8, 9] will be developed and implemented for thermomechanical tests conducted at NASA GRC. The developed schema will not only store the relevant data needed for the SMMD tool, but also the material pedigree (i.e., production of the bulk material, bulk material analysis, sample cut-out diagrams, sample fabrication procedure, etc.), test pedigree (i.e., test equipment used, measurement systems used, raw test data), and analysis pedigree (i.e., how the data in the SMMD tool is calculated). Furthermore, a Python-based framework will be developed to seamlessly interact between the SMA Analytics and SMMD tools, which will write the full dataset and associated metadata to the GRC Information Management System before passing the required point-wise data to the SMMD tool. Data informatics is a key element of the NASA Vision 2040, which requires not only that data is stored and maintained throughout the material lifecycle, but that the data is also accessible and reusable such that material development efforts can be minimized. Therefore, for an ICME design approach to be realized, a centralized information management system that drives the ICME process must be able to communicate with other databases. The work that will be presented in this presentation will therefore not only demonstrate the ability of NASA GRC’s information management system to capture SMM data, but also its ability to interact with pre-existing tools specialized for such materials.

Data management↗

History and Status of ALSEP and the Apollo Lunar Data Project

A suite of automated scientific instruments (the Apollo Lunar Surface Experiment Package, or ALSEP) was installed at each of the landing sites of Apollo 12, 14, 15, 16, and 17 from 1969 to 1972. They operated from deployment until decommissioning on 30 September 1977. These data were continuously transmitted to Earth and saved on the Range Tapes, which were recorded at the Manned Space Flight Network stations. These data were also broken out by experiment and sent to the experiment Principal Investigators on what were called the P.I. Tapes. Starting in April 1973 the Range Tape data were stored in digital format on 7-track magnetic tapes, the ARCSAV Tapes. In February 1976, the handling of the Range Tapes was transferred to UT Galveston. They produced 9-track tapes referred to as the Work Tapes. Following the Apollo program the Range and ARCSAV tapes, which were never archived, were lost. The Work Tapes were archived at the National Space Science Data Center (NSSDC). Some investigators archived their individual experiment data with NSSDC as well, but much of the data had minimal documentation, were not in digital form, or were stored in difficult to translate formats. Data from many experiments were never delivered to the NSSDC. The Lunar Data Project was started to address the problem of both missing and not readily usable data. Our effort has resulted in recovery of some of the ARCSAV tapes, recovery and digitization of a large volume of Apollo scientific and technical documentation, and restoration of many ALSEP and other Apollo data collections. Restoration involves deciphering formats, assembling necessary ancillary data (metadata), and packaging data in digital format to be archived with the Planetary Data System (PDS). Recovery of the data from the ARCSAV tapes involved having the tapes read on special equipment and extracting the individual experiment data out of the integrated data stream. We will report on the history and status of the various recovery efforts.

Work Tapes↗

Availability of previously lost data and metadata from the Apollo Lunar Surface Experiments Package (ALSEP)

Fourteen types of geophysical instruments deployed at the Apollo 12, 14, 15, 16, and 17 sites by the astronauts for long-term observation were collectively called the Apollo Lunar Surface Experiments Package (ALSEP). These instruments were active from the times of their deployment (November 1969–December 1972) to September 1977. At the conclusion of the experiments, the raw instrument data received from the Moon prior to March 1976 were left unarchived. Portions of the data processed by the principal investigators (PIs) of these experiments had been archived at the NASA Space Science Data Coordinated Archive (NSSDCA) in various formats. The unarchived data, residing then on open-reel magnetic tapes, became lost in the decades since, along with much of the metadata (the supporting documents for these data). We have recently recovered 440 of the previously lost tapes, containing raw ALSEP instrument data from April through June of 1975. Here we describe the data extracted from these tapes and summarize the data products generated for archiving at the NASA Planetary Data System (PDS) and NSSDCA, along with their historical narrative. In addition, we have reformatted many of the datasets delivered to NSSDCA by the PIs in the 1970s for archiving at the PDS. Finally, we have compiled an online searchable repository of ALSEP-related documents by optically scanning tens of thousands of pages of them kept at the Lunar and Planetary Institute in Texas.

S. Nagihara↗

Machine Learning Enabled Quantitative Risk Assessment of Aerial Wildfire Response

Aerial wildfire operations are high risk and account for a large number of firefighter deaths. Increasing intensity of wildfires is driving a surge in aerial operations, while simultaneously there is growing interest in improving system safety and performance. In this work, wildfire aviation mishaps documented using the SAFECOM system are analyzed using a previously developed framework for hazard extraction and analysis of trends (HEAT). Hazards and specific failure modes are extracted from the narrative data in SAFECOM forms using natural language processing techniques. Metrics for each hazard are calculated, including frequency, rate, and severity. We examine whether these metrics change over time, and whether they are related to metadata, such as region and aircraft type. The results of the hazard analysis are presented in a risk matrix, identifying the highest and lowest risk hazards based on rate of occurrence and average severity. Results identify jumper operations hazards as high-risk, in addition to bucket drop failures, cargo let down failures, and severe weather as medium risk.

machine learning↗

Machine Learning Enabled Quantitative Risk Assessment of Aerial Wildfire Response

Aerial wildfire operations are high risk and account for a large number of firefighter deaths. Increasing intensity of wildfires is driving a surge in aerial operations, while simultaneously there is growing interest in improving system safety and performance. In this work, wildfire aviation mishaps documented using the SAFECOM system are analyzed using a previously developed framework for hazard extraction and analysis of trends (HEAT). Hazards and specific failure modes are extracted from the narrative data in SAFECOM forms using natural language processing techniques. Metrics for each hazard are calculated, including frequency, rate, and severity. We examine whether these metrics change over time, and whether they are related to metadata, such as region and aircraft type. The results of the hazard analysis are presented in a risk matrix, identifying the highest and lowest risk hazards based on rate of occurrence and average severity. Results identify jumper operations hazards as high-risk, in addition to bucket drop failures, cargo let down failures, and severe weather as medium risk.

machine learning↗

Textural-Contextual Labeling and Metadata Generation for Remote Sensing Applications

Despite the extensive research and the advent of several new information technologies in the last three decades, machine labeling of ground categories using remotely sensed data has not become a routine process. Considerable amount of human intervention is needed to achieve a level of acceptable labeling accuracy. A number of fundamental reasons may explain why machine labeling has not become automatic. In addition, there may be shortcomings in the methodology for labeling ground categories. The spatial information of a pixel, whether textural or contextual, relates a pixel to its surroundings. This information should be utilized to improve the performance of machine labeling of ground categories. Landsat-4 Thematic Mapper (TM) data taken in July 1982 over an area in the vicinity of Washington, D.C. are used in this study. On-line texture extraction by neural networks may not be the most efficient way to incorporate textural information into the labeling process. Texture features are pre-computed from cooccurrence matrices and then combined with a pixel's spectral and contextual information as the input to a neural network. The improvement in labeling accuracy with spatial information included is significant. The prospect of automatic generation of metadata consisting of ground categories, textural and contextual information is discussed.

Kiang, Richard K.↗