Engineering PapersSearch

SEARCH · Engineering Papers

Results for “Machine Learning for Data Analysis”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

A New Generation of Intelligent Trainable Tools for Analyzing Large Scientific Image Databases

In a variety of scientific disciplines two-dimensional digital image data is now relied on as a basic component of routine scientific investigation. The proliferation of image acquisition hardware such as multi-spectral remote-sensing platforms, medical imaging sensors, and high-resolution cameras have led to the widespread use of image data in fields such as atmospheric studies, planetary geology, ecology, agriculture, glacielogy, forestry, astronomy, diagnostic medicine, to name but a few.

machine

The Nasa Multiscale Analysis Tool: an Enabling Platform for Achieving Vision 2040

Vision 2040 is a community-driven consensus document, written in 2018, aimed at defining the potential 25-year future state required for performing integrated multiscale modeling of materials and systems for future aerospace and aeronautical applications. Nine Vision Key Elements (KEs) were defined along with associated technical gaps. This paper will address current NASA GRC research efforts utilizing the NASA Multiscale Analysis Tool (NASMAT). This paper will specifically focus on NASMAT’s ability to address gaps in three of the nine Vision 2040 KEs: 1) Models and Methods, 2) Multiscale Measurements and Characterization Tools and Methods, and 6) Data, Informatics, and Visualization. NASMAT is a versatile platform for performing computationally efficient multiscale analyses of heterogeneous materials. NASMAT offers the user flexibility to define an arbitrary number of length scales (levels) where a variety of micromechanics theories can be implemented at each level. Micromechanics theories can be selected to balance accuracy and computational efficiency and range from analytical (Mori-Tanaka) to several semi-analytical (method of cells) formulations. NASMAT can also be coupled with external software and used to perform multiscale analyses of more complex structures. The paper will include a recent application of NASMAT to model a complex, three-dimensional woven composite, with a particular emphasis placed on multiscale measurements utilized to enhance the quality of the multiscale analysis. Since typical NASMAT analyses can be completed in on the order of seconds to minutes, a second example will demonstrate NASMAT’s ability to generate large quantities of data useful for sensitivity analysis, uncertainty quantification, or machine learning applications. Current progress on developing multiscale data visualization tools will also be addressed along with the challenges associated with and proposed solutions for sifting through large amounts of data. These examples will demonstrate that NASMAT is an enabling platform for achieving the goals in Vision 2040.

Vision 2040

Machine Learning Framework for Hazard Extraction and Analysis of Trends (HEAT) in Wildfire Response

This research proposes a natural language processing enabled risk analysis framework, named Hazard Extraction andAnalysis of Trends (HEAT), and applies the framework to the ICS-209-PLUS data set of wildfire incident responseforms. The HEAT framework produces safety- and risk- relevant analyses, consisting of: (1) a set of hazards extractedfrom text data, (2) a primary analysis using hazard-relevant metrics, such as rate and severity, to form an FMEA-styletable and risk matrix, (3) a time series analysis of metric trends, and (4) a secondary analysis examining potentialpredictors for hazards. Results from HEAT provide quantitative risk-relevant information for high-level hazards doc-umented in existing-state operations. Because of the generalizability of the steps and limited data requirements, HEATcan be applied to any dataset containing narrative text, thus providing a framework for data-driven machine learning-enabled quantitative risk analysis across a variety of domains. To demonstrate HEAT in a case study, we apply theframework to the ICS-209-PLUS dataset of wildland fire incident response forms. Hazards identified in wildfire re-sponse arise from environmental conditions, the mission, and the wildland urban interface. The resulting risk matrixidentifies evacuations as high-risk hazards, while all other identified hazards are medium or serious risk.

natural language processing

Machine Learning in Heliophysics and space weather forecasting: a white paper of finding and recommendations

The authors of this white paper met on 16-17 January 2020 at the New Jersey Institute of Technology,Newark, NJ, for a 2-day workshop that brought together a group of heliophysicists, data providers,expert modelers, and computer/data scientists. Their objective was to discuss critical developments and prospects of the application of machine and/or deep learning techniques for data analysis, modeling and forecasting in Heliophysics, and to shape a strategy for further developments in the field. The workshop combined a set of plenary sessions featuring invited introductory talks interleaved with a set of open discussion sessions. The outcome of the discussion is encapsulated in this white paper that also features a top-level list of recommendations agreed by participants

HSR

Comparative Analysis of Empirical and Machine Learning Models for Chla Extraction Using Sentinel-2 and Landsat OLI Data: Opportunities, Limitations, and Challenges

Remote retrieval of near-surface chlorophyll-a (Chla) concentration in small inland waters is challenging due to substantial optical interferences of various water constituents and uncertainties in the atmospheric correction (AC) process. Although various algorithms have been developed to estimate Chla from moderate-resolution terrestrial missions (∼10–60 m), the production of both accurate distribution maps and time series of Chla has proven challenging, limiting the use of remote analyses for lake monitoring. Here, we develop a support vector regression (SVR) model, which uses satellite-derived remote-sensing reflectance spectra () from Sentinel-2 and Landsat-8 images as input for Chla retrieval in a representative eutrophic prairie lake, Buffalo Pound Lake (BPL), Saskatchewan, Canada. Validated against in situ Chla from seven ice-free seasons (N ∼ 200; 2014–2020), the SVR model outperformed both locally tuned, -fed empirical models (Normalized Difference Chlorophyll Index, 2- and 3-band, and OC3) and Mixture Density Networks (MDNs) by 15–65%, while exhibiting comparable performance to a locally trained MDN, with an error of ∼35%. Comparison of Chla retrieval models, AC processors (iCOR, ACOLITE), and radiometric products (Rayleigh-corrected, surface, and top-of-atmosphere reflectance) showed that the best Chla maps and optimal time series (up to 100 mg m−3) were produced using a coupled SVR-iCOR system.

algal blooms

Materials Informatics at NASA GRC: Machine Learning Surrogate Modeling, Data Management, and Integrated Toolsets for Establishing/Maintaining the Digital Thread

Integrated Computational Materials Engineering (ICME) has recently received widespread attention due to its promises in reducing dependence on physical testing for engineering design by relying on simulation, reducing both time and cost to market for various applications. ICME however requires validated multiscale material models, which heavily depend on available test data with full material and test pedigree, including material processing, test and measurement equipment, raw data collection, and analysis methodology and results that is findable and usable, along with integrated, efficient toolsets for effectively passing information across various length and time scales across such models. At the NASA Glenn Research Center under the Transformational Tools and Technologies Project, significant recent efforts have been directed towards establishing the required cyberinfrastructure to enable optimized ICME processes and the design of “fit-for-purpose” materials to achieve the goals outlined in the NASA Vision 2040 report. Such efforts include development of multiscale physics-based material models, which can be used to train highly efficient surrogate machine learning models, development of best practices and infrastructure for effective, traceable materials information management, and development of toolsets that integrate with physics-based codes, machine learning models, and an information management system to enable high throughput of materials data collection and analysis, establishment of digital twins and the digital thread, and automation of the ICME design process for material optimization.

Machine Learning

Automated Analysis of a Large-Scale Sky Survey: The SKICAT System

We describe the application of decision tree based classification techniques to the development of an automated tool for the reduction of a large scientific data set. The primary benefits of the SKICAT approach are increased data reduction throughput, repeatability, and consistency of classification.

data analysis image databases

TPSAS-NF1676L-32493-DND

The Committee on Earth Observation Satellites (CEOS) System Engineering Office (SEO) has supported the Open Data Cube (ODC) initiative to provide a data architecture solution that has value to its global users and increases the impact of EO satellite data. ODC is an open-source platform for processing satellite data. We have developed software products and tools around the core ODC that would help users perform machine learning on EO satellite data. The recent United Nations (UN) Sustainable Development Agenda provides a shared blueprint for peace and prosperity for people and for the planet, considering our current situation and helping to create a plan. The core of this agenda is a set of seventeen Sustainable Development Goals (SDGs), which represent an urgent call for action by all countries - both developed and developing - in a global partnership. The CEOS SEO team has recently developed and released a set of innovative Jupyter notebooks addressing UN SDGs 6.6.1 (spatial extents of water-related ecosystems), 11.3.1 (ratio of land consumption rate to population growth rate), and 15.3.1 (proportion of land that is degraded over total land area). These notebooks empower users by providing features that will assist with streamlining analysis ready data retrieval, processing, and visualization. We have recently incorporated several machine learning techniques in these notebooks. In this paper, we present the lessons learned from our experience on classifying land using supervised and unsupervised machine learning techniques using ODC framework for UN SDGs. We identify the current limitations of ODC to seamlessly support machine learning techniques. We propose features that would help machine learning, specifically within the ODC framework. We propose a thematic indexing/loading of data for both unsupervised learning as well as data annotation/labeling pipeline. Currently, ODC supports machine learning by separating data-management from the analysis process. It works as a mechanism to load cubes of data. ODC does not natively support features that are vital in machine learning such as validation splits, fair/balanced sampling, establishing load size constraints, etc. We believe that our proposed features will empower users by providing features that bring machine learning techniques closed to ODC. Enhancements to ODC to better accommodate machine learning techniques can assist in fulfilling UN SDGs such as 6.3.2, 6.4.2, 6.6.1, 11.3.1, 14.1.1, 15.1.1, 15.3.1, and 15.4.2.

Syed R Rizvi

A Gaussian Process Enhancement to Linear Parameter Varying Models

Simulation and analysis for modern engineering systems now routinely requires the merging of multiple disciplines, physical-domains, time-scales, and data sets — all at ever increasing levels. These capabilities are especially needed in the domain of Advanced Air Mobility, where rapidly emerging vehicle designs are significantly more complex, while having to be both cost-effective and safe. To meet these engineering challenges, machine learning methods are an attractive option for merging models and data across multiple areas while providing uncertainty quantification and maintaining computational efficiency. This paper examines the use of Gaussian process machine learning to generalize and enhance the commonly used class of quasi-Linear Parameter Varying models for fast full-envelope simulation while also supporting control system design and analysis with model uncertainty. Gaussian process machine learning is selected because it: can fuse multiple data sets, enables an easy trade-off between data fitting and smoothing, provides model uncertainty quantification, scales well with increasing complexity, and does not generally require starting from a large training data set. To demonstrate the benefits of the approach, a robust stability analysis with Gaussian process uncertainty is shown for a NASA reference design of an electric quad-rotor air-taxi concept vehicle with motor parameter uncertainty.

Gaussian Process

Space Flown Rodent Liver RNA Sequencing Data for Machine Learning in Space Biology Research

High-throughput nucleic acid sequencing (DNA-seq, RNA-seq) has become widespread in biomedical research due to the growing availability and affordability of these assays. Data analysis has been accelerated in recent years by the adoption of artificial intelligence (AI) and machine learning (ML) techniques by biomedical researchers. In space biology research, RNAseq datasets from space-flown experimental samples are critical for characterizing the gene expression aberrations associated with exposure to spaceflight stressors. However, space biological experiments tend to be very low sample size, so identifying proper AI/ML algorithms for sequencing data analysis is an ongoing challenge since these algorithms typically require large sample size. The NASA Science Mission Directorate (SMD) has started the “Benchmark Initiative for AI/ML”, focused on creating datasets meant for three main applications: 1) scientific benchmarking, which finds the best algorithm for a specific problem; 2) application benchmarking, which measures algorithm performance against a set of parameters; and 3) system benchmarking, which evaluates performance of hardware and software architecture. These scientific benchmarks consist of an AI-ready dataset and a reference implementation on a specific scientific question. In this work, we focused on generating standardized datasets to allow the scientific community to benchmark AI/ML algorithms in the domain of space biology. We present here a standardized, AI-ready, publicly available benchmark dataset for space biology RNA-seq data as a collaboration between the NASA AI4LS (Artificial Intelligence for Life Sciences) working group. and NASA’s SMD. This dataset consists of space-flown and ground control mouse liver found in the NASA GeneLab omics database. However, to amplify the small sample number (n=112 samples) for ML purposes, we employ Gaussian noise and a generative adversarial network to extend this dataset to 6,000 synthetic samples, matching the original gene expression characteristics.

James Casaletto

Biological Data for Deep Space Mission Support

Increased biomedical risks and challenges associated with deep space missions (cis-Lunar, Mars transit, Mars surface) require new knowledge discovery and development of novel ecosystem and biomedical support capabilities. This paradigm shift supporting distant and long-duration missions requires biological data to be findable, accessible, interoperable, reusable (FAIR), and maximally open-access (i.e., there is a data governance continuum from closed to mediated to embargoed to open). The NASA “Open Science Data Repositories” (OSDR) aims to meet scientific, technical, and operational spaceflight needs, and offers the ability to upload, download, search, share, analyze, and visualize data across physiological, behavioral, ‘omics, and environmental monitoring telemetry datasets. OSDR includes NASA GeneLab, NASA Ames Life Sciences Data Archive (ALSDA), and NASA Biological Institutional Scientific Collection (NBISC). In the past year, ALSDA has undergone a transformation in its data collection, curation, and architecture methods. Standardizing non-genomic (phenotypic) datasets was, and will continue to be, a challenge because of their diverse nature (e.g., molecular, cellular, tissue, whole organism, behavior; micro-computed tomography, intraocular pressure, fluorescence microscopy, western blot, ultrasonography; tabular, images, video). This year ALSDA, alongside GeneLab, introduced the Biological Data Management Environment (BDME) with the purpose to accept submission of data from space relevant experiments including spaceflight, radiation, simulated gravity, gravitropism, isolation and confinement, hostile closed environments and/or distance from Earth. In addition to bringing together omics, phenotypic, physiological, bioimaging, and behavioral data into one repository. By integrating with GeneLab a multi-project submission portal aims to reduce the burden on PIs submitting data and enabling the discovery of both omics and phenotypic data. The purpose of ALSDA is to collect, curate, and make all non-human space-relevant biological data maximally findable, accessible, interoperable, and reusable (FAIR). These scope of ALSDA data collected and submitted by PIs include study design metadata, subject metadata, assay metadata (parameters), raw and processed assay data, assay imagery/video, and subject-experienced mission data telemetry (radiation, temperature, humidity, acoustics, vibrations, etc.). In 2021, a community of researchers rallied to form the ALSDA Analysis Working Group (AWG) and provided scientific consensus on dataset sample and assay metadata. The community and excitement around the ALSDA/OSDR system has already led to several data reuse studies, demonstrating value using machine learning (ML), knowledge graphs, and meta-analysis approaches.

space biology

Science Autonomy for Ocean Worlds Astrobiology: A Perspective

Astrobiology missions to ocean worlds in our solar system must overcome both scientific and technological challenges due to extreme temperature and radiation conditions, long communication times, and limited bandwidth. While such tools could not replace ground-based analysis by science and engineering teams, machine learning algorithms could enhance the science return of these missions through development of autonomous science capabilities. Examples of science autonomy include onboard data analysis and subsequent instrument optimization, data prioritization (for transmission), and real-time decision-making based on data analysis. Similar advances could be made to develop streamlined data processing software for rapid ground-based analyses. Here we discuss several ways machine learning and autonomy could be used for astrobiology missions, including landing site selection, prioritization and targeting of samples, classification of “features” (e.g., proposed biosignatures) and novelties (uncharacterized, “new” features, which may be of most interest to agnostic astrobiological investigations), and data transmission.

ocean worlds

Biological Data for Deep Space Mission Support

Increased biomedical risks and challenges associated with deep space missions (cis-Lunar, Mars transit, Mars surface) require new knowledge discovery and development of novel ecosystem and biomedical support capabilities. This paradigm shift supporting distant and long-duration missions requires biological data to be findable, accessible, interoperable, reusable (FAIR), and maximally open-access (i.e., there is a data governance continuum from closed to mediated to embargoed to open). The NASA “Open Science Data Repositories” (OSDR) aims to meet scientific, technical, and operational spaceflight needs, and offers the ability to upload, download, search, share, analyze, and visualize data across physiological, behavioral, ‘omics, and environmental monitoring telemetry datasets. OSDR includes NASA GeneLab, NASA Ames Life Sciences Data Archive (ALSDA), and NASA Biological Institutional Scientific Collection (NBISC). In the past year, ALSDA has undergone a transformation in its data collection, curation, and architecture methods. Standardizing non-genomic (phenotypic) datasets was, and will continue to be, a challenge because of their diverse nature (e.g., molecular, cellular, tissue, whole organism behavior; micro-computed tomography, intraocular pressure, fluorescence microscopy, western blot, ultrasonography; tabular, images, video). This year ALSDA, alongside GeneLab, introduced the Biological Data Management Environment (BDME) with the purpose to accept submission of data from space relevant experiments including spaceflight, radiation, simulated gravity, gravitropism, isolation and confinement, hostile closed environments and/or distance from Earth. In addition to bringing together omics, phenotypic, physiological, bioimaging, and behavioral data into one repository. By integrating with GeneLab a multi-project submission portal aims to reduce the burden on PIs submitting data and enabling the discovery of both omics and phenotypic data. The purpose of ALSDA is to collect, curate, and make all non-human space-relevant biological data maximally findable, accessible, interoperable, and reusable (FAIR). These scope of ALSDA data collected and submitted by PIs include study design metadata, subject metadata, assay metadata (parameters), raw and processed assay data, assay imagery/video, and subject-experienced mission data telemetry (radiation, temperature, humidity, acoustics, vibrations, etc.). In 2021, a community of researchers rallied to form the ALSDA Analysis Working Group (AWG) and provided scientific consensus on dataset sample and assay metadata. The community and excitement around the ALSDA/OSDR system has already led to several data reuse studies, demonstrating value using machine learning (ML), knowledge graphs, and meta-analysis approaches.

space biology

An Enabling Platform for Achieving Multiscale Multiphysics Analysis of Multiphase Materials

This paper will address current NASA GRC research efforts utilizing the NASA Multiscale Analysis Tool (NASMAT) which address technical gaps in three of the Vision 2040 key discipline areas. NASMAT is a versatile platform for performing computationally efficient multiscale analyses of heterogeneous materials and is available free through the NASA software catalog. It offers the user flexibility to define an arbitrary number of length scales (levels) where a variety of micromechanics theories can be implemented at each level. Micromechanics theories can be selected to balance accuracy and computational efficiency and range from analytical (Mori-Tanaka) to several semi-analytical (method of cells) formulations. The resulting anisotropic, evolving nonlinear, thermomechanical constitutive model can also be coupled with external software and used to perform multiscale analyses of more complex structures. A recent application to model a complex, three-dimensional woven composite, with a particular emphasis placed on multiscale measurements utilized to enhance the quality of the multiscale analysis will be discussed. Since typical NASMAT analyses can be completed on the order of seconds to minutes, a second example will demonstrate the ability to generate large quantities of data useful for sensitivity analysis, uncertainty quantification, or machine learning applications. Current progress on implementing NASMAT within a multiscale digital thread/digital twin framework will also be addressed. These examples will demonstrate that NASMAT is an enabling platform for achieving the goals in Vision 2040.

Multiscale

Increasing accessibility to deep learning-based analytics for space biology: pretrained models, transfer learning, and analytics platform development

Biological systems react in complex ways to the stressors of spaceflight, and the data capturing these relationships is concomitantly high-dimensional and complex. Deep learning and machine learning approaches are increasingly popular as an analytical approach for space biosciences, due to their ability to model complex relationships in complex data. However, such approaches often require large datasets and extensive computational resources. New approaches that minimize data sizes and computational power needed to leverage machine learning, and resources that make these approaches accessible, are needed to increase accessibility and adoption of machine learning in the space biosciences. Transfer learning, in which a pretrained model of broad utility is trained on a large dataset, and subsequently reused on downstream applications for which data is more limited, is one approach to minimizing data and computational intensity of deep learning applications. This transfer learning approach results in more performant models in high-dimensional, low-sample-size settings such as space biology, as compared to training models on limited data from scratch. This presentation will outline efforts to generate pretrained models for the space biology community, and highlight transfer learning applications modeling microbial antibiotic resistance during spaceflight. Finally, in order to increase accessibility of these models and tools, as well as others, for the broader space biology community, we present a modeling and analysis platform facilitating machine learning applications in space biology. This platform streamlines machine learning training and analysis in a notebook format, facilitates download and use of space biology data from the NASA GeneLab database, and can be utilized on NASA-hosted servers or downloaded and hosted locally. This effort, as part of the AI4LS (Artificial Intelligence for Life in Space) working group, will increase accessibility, feasibility, and performance of machine learning approaches for the space biology community.

Adrienne Hoarfrost

Study of Antarctic Blowing Snow Storms Using MODIS and CALIOP Observations With a Machine Learning Model

As a common phenomenon over Antarctica, blowing snow (BLSN), especially the large BLSN storms, play an important role in the Antarctic surface mass balance, radiation budget, and planetary boundary layer processes. This study presents the work on BLSN storm identification and analysis with observations from the Moderate Resolution Imaging Spectroradiometer (MODIS) onboard the Aqua satellite. Spectral analysis shows that BLSN identification is feasible with MODIS daytime data. A random forest machine learning model is developed and observations from the Cloud‐Aerosol Lidar with Orthogonal Polarization are used for training. Model performance results show that machine‐learning based classification can achieve over 90% overall accuracy when classifying MODIS pixels into cloud, clear, and BLSN categories. The machine learning model is applied to MODIS observations during the month of October 2009 for BLSN storm analysis. Results show that the size of BLSN storms has a large spectrum and can reach hundreds of thousands km2. The MODIS based BLSN storm frequency map extends the Cloud‐Aerosol Lidar and Infrared Pathfinder Satellite Observations coverage limit from 82°S to the South Pole. A BLSN storm belt, which extends from the South Pole region to the coastal area between 130°E and 160°E along the Transantarctic Mountains, provides a potential pathway of snow transport. These results are important in improving the understanding of BLSN impact on Antarctic surface mass balance and boundary layer processes.

Antarctic

Machine Learning Prototype App For Recognition of Fruits

As the incidence of obesity and associated negative health consequences is rising, it becomes crucial to monitor the dietary choices of individuals. Unfortunately, traditional methods to collect this information involve collecting food frequency questionnaires from individuals using paper. Electronic food trackers have been developed to collect food data, but they require participants to manually label and describe the content of their meals, and which may be difficult for researchers to interpret in a standardized fashion. Machine learning, however, provides an easy and efficient method for both participants and researchers to label food items with standardized descriptions. This project aims to create a prototype phone application that can identify and label photos of apples. This is done by making a machine learning model through Turicreate, a python module, which is then implemented into an iOS app through Xcode and Swift. The modules used in Swift include CoreML and AVFoundation. This machine learning application will be incorporated with a MealLogger phone app that is also under development. The MealLogger app will be used to keep track of participants' calorie intake and other personal details throughout the sleep study. The machine learning model will present several potential identities of the foods found in the photo, and the user will only need to select the correct option. This will be a user-friendly method for participants to easily log their food consumption without the hard work of manually inputting each and every description. Some limitations to this project include the wide variety of food, including those within different cultures. To deal with this, the model will include the most generic food categories, which the participant may select, and produce a drop-down menu of more specific dishes under that specified category, with the option of self-input. Additional questionnaires may be implemented according to the food type selected This will allow the process to be quick and easy, but also specific for the purpose of analysis. The release of the application will require a much longer process, but the machine learning prototype presents a first step toward an application that may change data analysis for researchers interested in collecting food intake from individuals living in the real world.

Food tracker