Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “generalizability”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Bayesian Model Selection for Reducing Bloat and Overfitting in Genetic Programming for Symbolic Regression

When performing symbolic regression using genetic programming, overfitting and bloat can negatively impact generalizability and interpretability of the resulting equations as well as increase computation times. A Bayesian fitness metric is introduced and its impact on bloat and overfitting during population evolution is studied and compared to common alternatives in the literature. The proposed approach was found to be more robust to noise and data sparsity in numerical experiments, guiding evolution to a level of complexity appropriate to the dataset. Further evolution of the population resulted not in overfitting or bloat, but rather in slight simplifications in model form. The ability to identify an equation of complexity appropriate to the scale of noise in the training data was also demonstrated. In general, the Bayesian model selection algorithm was shown to be an effective means of regularization which resulted in less bloat and overfitting when any amount of noise was present in the training data.

Uncertainty quantification↗

Application of a Bayesian Framework for Plasticity Model Selection

Interpretable Machine Learning (IML) has performed well when tasked with deriving constitutive material models. However, IML has been shown to prefer models that overfit noise in data, which tends to lead to bloat and a decrease in interpretability. Due to these issues, the ability of IML to reliably derive models that fit the data and are both interpretable and generalizable is limited. A method developed recently has shown promise to improve upon traditional IML by using a Bayesian fitness definition for the evolution of free-form models with non-deterministic parameters. This framework was developed for genetic-programming-based symbolic regression(GPSR) and involves model parameter estimation using Sequential Monte Carlo sampling (SMC).The method has demonstrated a reduction in bloat when dealing with noisy data in comparison to conventional GPSR. The results of this framework applied to stress-strain data for copper show models that more effectively predict the experimental data better than was previously shown with GPSR.

plasticity↗

What Went Wrong: A Survey of Wildfire UAS Mishaps through Named Entity Recognition

Increasingly, unmanned aircraft systems (UAS) are being applied to wildfire incidents for tasks such as mapping, aerial ignition, and delivery. As a result, aviation incident reporting systems for wildfires are beginning to accumulate data related to UAS mishaps in wildfire response. In this research, we apply state-of-the-art natural language processing (NLP) techniques to develop a custom Named Entity Recognition (NER) model which extracts entities relevant to safety analysts. The custom NER model is built by fine-tuning an existing Bidirectional Encoder Representations from Transformers (BERT) model, resulting in a generalizable NER model that can extract engineering relevant entities including failure modes, causes, effects, control processes, and recommendations from failure-relevant text. This model performs passably, with a weighted average f1 score of 0.33 across entity types, indicating more labeled training data is needed. Extracted entities are used to form a Failure Modes and Effects Analysis (FMEA)-style survey of wildfire UAS mishaps reported using the SAFECOM system. Similar mishaps are manually clustered and reported as single rows within an FMEA. Foreach cluster, we compute frequency, severity, and overall riskin accordance with FAA standards. This methodology can beapplied as part of a broader safety management system totrack trends in mishaps (e.g., likelihood, severity) and discoverknowledge (e.g., causes, effects) that can be utilized to improvesafety outcomes and system performance.

Machine Learning↗

Nearest-Neighbor Machine Learning Feature Selection for Interpretation of Microbial Molecular Signatures from Isotope Ratio Mass Spectrometry Data

Mass spectrometry (MS) promises to be a powerful tool for potential biosignature detection during astrobiological missions on ocean worlds in our solar system. Accurate and generalizable machine learning methods could enhance science return on investment by predicting seawater chemistry and classifying isotopic biosignatures, either as a signature consistent with microbial life (biotic) or as a novelty (unclassified/unique). However, machine learning models are likely to be complex and involve interactions between MS features, making biosignatures difficult to interpret. Feature selection methods provide biological and chemical context that help interpret the mechanisms of machine learning models, but these methods also need the ability to detect complex interactions. Previously, we developed a machine learning feature selection algorithm called nearest-neighbor projected distance regression (NPDR) that has the ability to identify important model features that involve complex interactions and automatically reduce correlation and the dimensionality in a high-dimensional variable space. The standard distance metrics used in NPDR – Manhattan and Euclidean – assume the multivariate data are isotropic, which is often violated in real data due to differences in the covariance between variables. Thus, we extend NPDR to include a random forest distance, and other anisotropic distance metrics, for computing nearest neighbors. We also augment the isotope-ratio MS data with time-series features from the raw MS signal to improve biotic classification. We test NPDR on our novel experimental ocean world seawater analog MS data. We measure isotope fractionations of volatile CO 2 that could be measured in exospheres or plumes. Samples include baseline abiotic conditions using a range of possible seawater chemistry consistent with Europa and Enceladus, and biotic samples that include microbes in these seawaters. We use penalized NPDR with random forest proximity to identify interpretable microbial molecular signatures. We compare features with random forest importance, and we train a classifier that discriminates between biotic and abiotic samples with high accuracy. These ML-trained ocean-world analog MS data could be used to assist in identifying biosignatures during future missions.

geochemistry↗

Machine Learning Framework for Hazard Extraction and Analysis of Trends (HEAT) in Wildfire Response

This research proposes a natural language processing enabled risk analysis framework, named Hazard Extraction andAnalysis of Trends (HEAT), and applies the framework to the ICS-209-PLUS data set of wildfire incident responseforms. The HEAT framework produces safety- and risk- relevant analyses, consisting of: (1) a set of hazards extractedfrom text data, (2) a primary analysis using hazard-relevant metrics, such as rate and severity, to form an FMEA-styletable and risk matrix, (3) a time series analysis of metric trends, and (4) a secondary analysis examining potentialpredictors for hazards. Results from HEAT provide quantitative risk-relevant information for high-level hazards doc-umented in existing-state operations. Because of the generalizability of the steps and limited data requirements, HEATcan be applied to any dataset containing narrative text, thus providing a framework for data-driven machine learning-enabled quantitative risk analysis across a variety of domains. To demonstrate HEAT in a case study, we apply theframework to the ICS-209-PLUS dataset of wildland fire incident response forms. Hazards identified in wildfire re-sponse arise from environmental conditions, the mission, and the wildland urban interface. The resulting risk matrixidentifies evacuations as high-risk hazards, while all other identified hazards are medium or serious risk.

natural language processing↗

What Went Wrong: A Survey of Wildfire UAS Mishaps through Named Entity Recognition

Increasingly, unmanned aircraft systems (UAS) are being applied to wildfire incidents for tasks such as mapping, aerial ignition, and delivery. As a result, incident reporting systems for wildfires are beginning to accumulate data related to UAS mishaps in wildfire response. In this research, we apply state-of-the-art natural language processing (NLP) techniques to develop a custom Named Entity Recognition (NER) model which extracts a Failure Modes and Effects Analysis (FMEA)-style survey of wildfire UAS mishaps reported in SAFECOM. The custom NER model is built by fine-tuning an existing (BERT) model, resulting in a generalizable NER model that can extract engineering relevant entities including failure modes, causes, effects, control processes, and recommendations from any failure-relevant text. Similar mishaps are clustered and reported as single rows within the FMEA. For each cluster, frequency, severity, and overall risk are computed. The methodology can be applied as part of a broader safety management system to track trends in mishaps and discover knowledge that can be utilized to improve safety outcomes and system performance.

Machine Learning↗

Habitability Assessments And Lessons-learned From 3-day And 11-day Enriched Oxygen Hypobaric Chamber Tests At NASA Johnson Space Center

INTRODUCTION: Decompression sickness (DCS) is a risk to the health and performance of astronauts and high-altitude aircrew. Tolerance to flammability, hypoxia, prebreathe duration, and DCS risk varies across different organizations, vehicles, suits, and destinations, necessitating a variety of DCS risk mitigation approaches. Existing models of altitude DCS risk are often insufficient to enable accurate risk-informed decisions during hardware development, mission planning, and flight operations. METHODS: NASA completed outfitting of a dedicated facility at Johnson Space Center to support testing of up to eight human subjects for multiple days in hypobaric and enriched oxygen atmospheres. The primary purpose of the testing capability is validation of DCS risk mitigation protocols for Artemis missions to the Moon; however, it will also support development and validation of a generalizable altitude DCS risk estimation tool. A 3-day and an 11-day prebreathe validation test were completed in 2022, each with 8 human subjects living at 56.5 kPa (8.2 psia), 34% O2, 66% N2, with 5 simulated EVAs performed on masks at 29.6 kPa (4.3 psi), 85% O2, 15% N2. Facility and organizational lessons-learned and process improvements were recorded during and following the tests, and subjective habitability ratings were recorded daily during the 11-day test. Hypoxia and DCS-related physiological and cognitive outcome measures were recorded during both tests and are reported in companion presentations. RESULTS & DISCUSSION: All subjects completed each of the tests. Primary habitability issues related to mask discomfort during simulated EVAs and poor sleep quality due to thin mattresses. Polybenzimidazole (PBI) clothing was worn by all subjects due to the increased fire risk and may be required for Artemis missions; clothing was found to be acceptable overall with the worst ratings being due to poor fit and inelasticity. Chamber O2 and CO2 sensor inconsistency was observed that did not result in test termination but required post-test follow-up. Forward plans include additional hypobaric testing and integration of existing and future physiological outcome data into an open-source Aerospace Estimation Tool for Hypobaric Exposure Risk (AETHER). NASA is also working to make the testing capability available to commercial companies.

Andrew F J Abercromby↗

Predicting Airport Runway Configurations for Decision-Support Using Supervised Learning

One of the most challenging tasks for air traffic controllers is runway configuration management (RCM). It deals with the optimal selection of runways to operate on (for arrivals and departures) based on traffic, surface wind speed, wind direction, other environmental variables, noise constraints, and several other airport-specific factors. It affects the efficiency of the National Airspace System (NAS) and both surface and airspace operations can benefit from better understanding future runway configurations. In this paper, we present a comprehensive implementation of predictive models for runway configuration estimation from large volumes of historical data. Specifically, operational data from two full years (2018 and 2019) is collected, analyzed, and fused together to build the data product used in this work. The data set differs from prior work in the field in terms of its scope, resolution, and variety of factors collected and considered. Meteorological data is collected from two different sources – current weather conditions from METAR (Meteorological Terminal Aviation Routine Weather Report) and forecast weather conditions from Localized Aviation MOS Program (LAMP). Operational data from the Federal Aviation Administration (FAA) Aviation System Performance Metrics (ASPM) related to scheduled and actual number of arrivals and departures, average taxi times, etc. are collected. NASA’s Sherlock Data Warehouse is used to identify critical information such as go-arounds, and other events that might impact RCM decision-making. All data is collected and aggregated over 15-minute intervals throughout the two years. This provides a resolution like the timescales that might be necessary for runway configuration management decision-making. A variety of supervised learning algorithms are tested including Support Vector Machine, Random Forest, Gradient Boosting, etc. including tuning of the model hyperparameters. The modeling process is applied and presented on two representative U.S. airports – Charlotte Douglas International Airport (KCLT) and Denver International Airport (KDEN). The two airports present different levels of complexity in terms of the total number of configurations used and provide a balanced perspective on the generalizability of the developed approach to other airports in the NAS. Initial results are promising (F1 score of 0.91 at KCLT and 0.83 at KDEN) for data in the test set. The final paper will contain a comprehensive comparison between different models and model building strategies as well as further refined results. Most important predictors for each airport will be identified along with a discussion and recommendations on adapting the framework to other scenarios.

Tejas G Puranik↗

A Comparison of Rotor Disk Modeling and Blade-Resolved CFD Simulations for NASA's Tiltwing Air Taxi

A multi-fidelity computational fluid dynamics analysis is carried out for NASA’s tiltwing air taxi concept operating in airplane and helicopter mode. High-fidelity simulations are computationally expensive due to individual rotor blade modeling in a time-dependent computational domain with rotating grids. The mid-fidelity rotor disk option, in its source term implementation, is explored as a more affordable alternative. Computations are performed with NASA’s OVERFLOW flow solver loosely-coupled with the comprehensive code CAMRAD II for appropriate rotor trim. Detailed comparisons are shown for the trim solution, airloads, wake geometry, and rotor performance. While the rotor disk model is able to capture the flow field with satisfactory agreement in airplane mode, it faces difficulties in helicopter mode due to the three-dimensional effects of the wake. Although this study is limited to a specific vehicle geometry, it is expected that the results are somewhat generalizable to the analysis of multi-rotor configurations.

ARMD↗

SatNet: A Benchmark for Satellite Scheduling Optimization

Satellites provide essential services such as networking and weather tracking, and the number of near-earth and deep space satellites are expected to grow rapidly in the coming years. Communications with terrestrial ground stations is one of the critical functionalities of any space mission. Satellite scheduling is a problem that has been scientifically investigated since the 1970s. A central aspect of this problem is the need to consider resource contention and satellite visibility constraints as they require line of sight. Due to the combinatorial nature of the problem, prior solutions such as linear programs and evolutionary algorithms require extensive compute capabilities to output a feasible schedule for each scenario. Machine learning based scheduling can provide an alternative solution by training a model with historical data and generating a schedule quickly with model inference. We present SatNet, a benchmark for satellite scheduling optimization based on historical data from the NASA Deep Space Network. We propose formulation of the satellite scheduling problem as a Markov Decision Process and use reinforcement learning (RL) policies to generate schedules. The nature of constraints imposed by SatNet differ from other combinatorial optimization problems such as vehicle routing studied in prior literature. Our initial results indicate that RL is an alternative optimization approach that can generate candidate solutions of comparable quality to existing state-of-the-practice results. However, we also find that RL policies overfit to the training dataset and do not generalize well to new data, thereby necessitating continued research on reusable and generalizable agents.

Wilson, Brian↗

Habitability Assessments And Lessons-learned From 3-day And 11-day Enriched Oxygen Hypobaric Chamber Tests At NASA Johnson Space Center

INTRODUCTION: Decompression sickness (DCS) is a risk to the health and performance of astronauts and high-altitude aircrew. Tolerance to flammability, hypoxia, prebreathe duration, and DCS risk varies across different organizations, vehicles, suits, and destinations, necessitating a variety of DCS risk mitigation approaches. Existing models of altitude DCS risk are often insufficient to enable accurate risk-informed decisions during hardware development, mission planning, and flight operations. METHODS: NASA completed outfitting of a dedicated facility at Johnson Space Center to support testing of up to eight human subjects for multiple days in hypobaric and enriched oxygen atmospheres. The primary purpose of the testing capability is validation of DCS risk mitigation protocols for Artemis missions to the Moon; however, it will also support development and validation of a generalizable altitude DCS risk estimation tool. A 3-day and an 11-day prebreathe validation test were completed in 2022, each with 8 human subjects living at 56.5 kPa (8.2 psia), 34% O2, 66% N2, with 5 simulated EVAs performed on masks at 29.6 kPa (4.3 psi), 85% O2, 15% N2. Facility and organizational lessons-learned and process improvements were recorded during and following the tests, and subjective habitability ratings were recorded daily during the 11-day test. Hypoxia and DCS-related physiological and cognitive outcome measures were recorded during both tests and are reported in companion presentations. RESULTS & DISCUSSION: All subjects completed each of the tests. Primary habitability issues related to mask discomfort during simulated EVAs and poor sleep quality due to thin mattresses. Polybenzimidazole (PBI) clothing was worn by all subjects due to the increased fire risk and may be required for Artemis missions; clothing was found to be acceptable overall with the worst ratings being due to poor fit and inelasticity. Chamber O2 and CO2 sensor inconsistency was observed that did not result in test termination but required post-test follow-up. Forward plans include additional hypobaric testing and integration of existing and future physiological outcome data into an open-source Aerospace Estimation Tool for Hypobaric Exposure Risk (AETHER). NASA is also working to make the testing capability available to commercial companies.

Andrew Abercromby↗

Strategies for Identifying Resilient Behavior In Aviation

When we imagine a situation where people fly aircraft and nothing scary happens, we assume it is the system that affords this phenomenon. That is, the overall design is the cause of the success. However, this is not always true. There are many examples of how the presence of a human in the system is the reason for a successful outcome, despite flaws in the system’s design. This resilient behavior is often overlooked and challenging to characterize. In an attempt to identify this phenomenon in a generalizable way, we named and investigated two strategies, as well as identification methods, that exemplify resilient performance: 1) controls; and 2) modifications. First, controls are resilient actions in situations known to be problematic. To capture this, instead of looking at the examples of how the problem became a reality, we look at the examples of how the problem was successfully avoided or controlled. For example, a country road may have a hairpin turn where a higher-than-normal rate of accidents occur. Given that there is a likely system flaw identified, we would look at how the successful drivers navigated the turn. This can be accomplished using current safety reporting systems. Second, modifications are augmentations or changes that people create to fill in the gap between work-as-imagined and work-as-done. This type of resilient performance is directly related to poor design. In aviation, work-as-imagined is often scripted explicitly, so it can be compared to work-as-done through the use of examining system-generated data as well as narrative reports written by the system operators. These two approaches aim to identify resilient human actions to better understand how current systems function, as well as how people contribute to successes that were otherwise unknown.

human factors↗

Predicting Airport Runway Configurations for Decision-Support Using Supervised Learning

One of the most challenging tasks for air traffic controllers is runway configuration management (RCM). It deals with the optimal selection of runways to operate on (for arrivals and departures) based on current and forecast of traffic, surface wind speed, wind direction, other environmental variables, noise constraints, and several other airport-specific factors. In this paper, a methodology using supervised learning is developed to build a predictive model for RCM decision-support from large volumes of historical data. Data from two full years (2018 and 2019) related to current and forecast weather, demand/capacity, etc. is collected, analyzed, and fused together. A variety of supervised learning algorithms are tested for predicting runway configuration and hyperparameter tuning is carried out to select the best performing model. The validation process involves two airports of low (Charlotte Douglas International Airport, CLT) and high (Denver International Airport, DEN) complexity of configuration decision-making. The results show significant promise for the two airports with test accuracy of 93% (CLT) and 73% (DEN). The methodology is scalable and generalizable to other airports across the U.S. National Airspace System.

air traffic management↗

A Dynamic PCA and Machine Learning Tool for Automated Identification of Solar Wind Disturbances Impacting Earth’s Magnetosphere

Earth’s magnetosphere is continuously impacted by solar wind and interplanetary magnetic field (IMF) disturbances, such as shocks, discontinuities, magnetic clouds and more. Understanding how such disturbances propagate from the Sun and what is their impact on the different magnetospheric domains is key to understanding and forecasting energy transfer from the solar wind to Earth. The large number of overlapping solar wind and magnetospheric missions carrying magnetometers and the recent advances in communications and data storage technologies have enabled an unprecedented quantity of high-fidelity magnetic field data captured by in-situ spacecraft to be available at the click of a button. However, this massive quantity of available data can prove unwieldy for researchers, limiting the identification of interesting phenomena and disturbances to a relatively small percentage of the total dataset. Several techniques have been previously developed for automated identification of specific types of magnetic anomalies, but these methods are typically mission-specific and can be difficult to generalize. We present initial results for a generic method of automated anomaly detection in magnetic field measurements based on dimensionality reduction and unsupervised clustering via machine learning. The benefit of our technique is its high degree of generalizability and flexibility which make it a most useful data survey tool for a wide range of magnetic field datasets. This method can also be applied simultaneously to other observed time-series properties like plasma density, pressure, and velocity for more accurate event identification. Additionally, the application of this method to data captured by multiple spacecraft enables the simultaneous identification of disturbances and the determination of their propagation characteristics. Initial evaluation of this technique has been performed using data from Magnetospheric MultiScale (MMS) and THEMIS-ARTEMIS missions, providing a testbed scenario for the future Heliophysics Environmental and Radiation Measurement Experiment Suite (HERMES) platform instruments that will measure solar wind and IMF properties from lunar orbit onboard the Gateway station.

Miguel Martinez-Ledesma↗

Predicting Airport Runway Configuration for Decision-Support Using Supervised Learning

One of the most challenging tasks for air traffic controllers is runway configuration management (RCM). It deals with the optimal selection of runways to operate on (for arrivals and departures) based on current and forecast of traffic, surface wind speed, wind direction, other environmental variables, noise constraints, and several other airport-specific factors. In this paper, a methodology using supervised learning is developed to build a predictive model for RCM decision-support from large volumes of historical data. Data from two full years (2018 and 2019) related to current and forecast weather, demand/capacity, etc. is collected, analyzed, and fused together. A variety of supervised learning algorithms are tested for predicting runway configuration and hyperparameter tuning is carried out to select the best performing model. The validation process involves two airports of low (Charlotte Douglas International Airport, CLT) and high (Denver International Airport, DEN) complexity of configuration decision-making. The results show significant promise for the two airports with test accuracy of 93% (CLT) and 73% (DEN). The methodology is scalable and generalizable to other airports across the U.S. National Airspace System.

air traffic management↗

A Dynamic PCA and Machine Learning Tool for Automated Identification of Solar Wind Disturbances Impacting Earth’s Magnetosphere

Earth’s magnetosphere is continuously impacted by solar wind and interplanetary magnetic field (IMF) disturbances, such as shocks, discontinuities, magnetic clouds and more. Understanding how such disturbances propagate from the Sun and what is their impact on the different magnetospheric domains is key to understanding and forecasting energy transfer from the solar wind to Earth. The large number of overlapping solar wind and magnetospheric missions carrying magnetometers and the recent advances in communications and data storage technologies have enabled an unprecedented quantity of high-fidelity magnetic field data captured by in-situ spacecraft to be available at the click of a button. However, this massive quantity of available data can prove unwieldy for researchers, limiting the identification of interesting phenomena and disturbances to a relatively small percentage of the total dataset. Several techniques have been previously developed for automated identification of specific types of magnetic anomalies, but these methods are typically mission-specific and can be difficult to generalize. We present initial results for a generic method of automated anomaly detection in magnetic field measurements based on dimensionality reduction and unsupervised clustering via machine learning. The benefit of our technique is its high degree of generalizability and flexibility which make it a most useful data survey tool for a wide range of magnetic field datasets. This method can also be applied simultaneously to other observed time-series properties like plasma density, pressure, and velocity for more accurate event identification. Additionally, the application of this method to data captured by multiple spacecraft enables the simultaneous identification of disturbances and the determination of their propagation characteristics. Initial evaluation of this technique has been performed using data from Magnetospheric MultiScale (MMS) and THEMIS-ARTEMIS missions, providing a testbed scenario for the future Heliophysics Environmental and Radiation Measurement Experiment Suite (HERMES) platform instruments that will measure solar wind and IMF properties from lunar orbit onboard the Gateway station.

Miguel Martinez-Ledesma↗

Multiyear Dry Periods in Southern Africa

Characteristics and physical features related to low precipitation across many years in Southern Africa that lead to societal disruptions are diagnosed using observed analyses and an ensemble of historical coupled climate model simulations during 1921 to 2014. Four regions are evaluated, as identified through a hierarchical clustering algorithm applied to the Standardized Precipitation Index (SPI) during the October–April precipitation season. Although dryness spanning many October–April occurs periodically in each region, they seldom occur simultaneously, consistent with largely insignificant SPI cross-correlations between them. However, characteristics relevant to low precipitation across many years are generalizable between the four regions, including the serial persistence of October–April precipitation, the likelihood of consecutive dry October–April, and the likelihood of dry October–April in temporal extents of up to 10 consecutive such 7-month seasons. Systematic precipitation persistence is not a feature in any of the four Southern Africa regions, as serial correlations of October–April SPI are not statistically significant at any time lags. It follows that there is an exponential-folding decay in the likelihood of consecutive October–April for various SPI thresholds and that there is a large spread in the likelihood of low October–April SPI across many years. In terms of physical features, low October–April SPI in each Southern Africa region is closely related to local atmospheric circulations; however, they are not as closely related to sea surface temperatures (SSTs). These results suggest that dryness spanning many years is determined primarily by persistent local circulations related to atmospheric variability and to a lesser extent variability related to SST anomalies, including the El Niño–Southern Oscillation.

subtropical Indian Ocean dipole↗

Recommendations on Evidence and Process for Certification of Learning-enabled Components in Aerospace Systems

This report primarily identifies a collection of relevant and necessary evidence for assurance of machine learnt components (MLCs)—also known as learning-enabled components—integrated into aircraft systems, and gives preliminary suggestions on the elements of a certification process that invoke the identified evidence. The main focus is on feedforward neural networks that are static and trained offline through supervised learning. A brief background on the generic elements of the lifecycle of an MLC is given to contextualize the assurance considerations and, consequently, the evidence that is relevant and necessary to support certification. At the level of an MLC, those considerations relate to: (i) the consistency and correctness of MLC contributions to system functions in the context of a validated functional intent; and (ii) the absence of MLC contributions to aircraft-level failure conditions. At an ML model level, confidence in model and data properties contribute to assurance of the containing MLC, in particular: (a) generalizability and robustness of models, in the presence of inputs not previously seen during training, disturbances to inputs, and unexpected inputs; and (b) valid data, i.e., data that are at least representative, relevant, complete, and accurate. Evidence for the above span the elements of the ML lifecycle, and includes, at a minimum, lifecycle artifacts that pertain to: (1) properties of requirements capturing functional intent, safety constraints, and aspects of the intended use and operating environment; (2) model performance, model complexity and design, and algorithm choice; (3) achievement of required performance at the levels of a trained model during model development, a trained model after model development is complete, and a trained model that is transformed into an executable equivalent; (4) model implementation aspects necessary for transforming a trained model into the executable equivalent; (5) integration of the executable trained model into the containing MLC, and eventually the larger system; and, (6) lastly, the verification and validation (V&V) of each of the above. Such V&V lifecycle artifacts themselves include: aspects of coverage, e.g., of various levels of requirements by the input space of the model and the data; traceability (where applicable); application of formal methods for property specification, analysis, and checking. Examples of evidence generation methods and tools further ground the discussion on what constitutes evidence, and the contribution to assurance during certification. The identified assurance considerations and supporting evidence is not a comprehensive set. Additionally, neither what should be considered as sufficient evidence relative to the assigned criticality of an MLC, nor how criticality ought to be determined and adjusted, have been considered in this report. However, suggestions are made for potential activities of the ML lifecycle that are aimed at providing confidence that an MLC can be relied upon when integrated into its containing (aircraft) system. Those activities are proposed as candidate elements of a certification process for MLCs. The main purpose of this report to inform regulatory guidance and consensus standards that may be used to meet the safety intent of the applicable regulations.

Aviation safety↗