Engineering PapersSearch

SEARCH · Engineering Papers

Results for “Ensemble methods”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Multimodel Ensemble Methods for Prediction of Wake-Vortex Transport and Decay Originating NASA

Several multimodel ensemble methods are selected and further developed to improve the deterministic and probabilistic prediction skills of individual wake-vortex transport and decay models. The different multimodel ensemble methods are introduced, and their suitability for wake applications is demonstrated. The selected methods include direct ensemble averaging, Bayesian model averaging, and Monte Carlo simulation. The different methodologies are evaluated employing data from wake-vortex field measurement campaigns conducted in the United States and Germany.

Korner, Stephan

Infrared, microwave, and spaceborne radar simulations of a deep convective system using a 3-D cloud ensemble method

A 3D cloud model is used to simulate the storm structure, and the results are linked to microwave and infrared radiative transfer models for simulation of aircraft observations. Spaceborne radar data are also simulated along the aircraft flight track. The cloud and radiative model simulations are studied and compared with aircraft observations. The initial results indicate that the 3D cloud model is capable of simulating the major features of observed storm systems when given a representative atmospheric sounding to initialize the convective systems. The simulations of infrared and microwave radiances provide reasonably good comparisons with the observations.

Yeh, H.-Y. M.

Ensemble Data Mining Methods

Ensemble Data Mining Methods, also known as Committee Methods or Model Combiners, are machine learning methods that leverage the power of multiple models to achieve better prediction accuracy than any of the individual models could on their own. The basic goal when designing an ensemble is the same as when establishing a committee of people: each member of the committee should be as competent as possible, but the members should be complementary to one another. If the members are not complementary, Le., if they always agree, then the committee is unnecessary---any one member is sufficient. If the members are complementary, then when one or a few members make an error, the probability is high that the remaining members can correct this error. Research in ensemble methods has largely revolved around designing ensembles consisting of competent yet complementary models.

Oza, Nikunj C.

Multi-Model Ensemble Wake Vortex Prediction

Several multi-model ensemble methods are investigated for predicting wake vortex transport and decay. This study is a joint effort between National Aeronautics and Space Administration and Deutsches Zentrum fuer Luft- und Raumfahrt to develop a multi-model ensemble capability using their wake models. An overview of different multi-model ensemble methods and their feasibility for wake applications is presented. The methods include Reliability Ensemble Averaging, Bayesian Model Averaging, and Monte Carlo Simulations. The methodologies are evaluated using data from wake vortex field experiments.

Koerner, Stephan

Multi-Model Ensemble Wake Vortex Prediction

Several multi-model ensemble methods are investigated for predicting wake vortex transport and decay. This study is a joint effort between National Aeronautics and Space Administration and Deutsches Zentrum fuer Luft- und Raumfahrt to develop a multi-model ensemble capability using their wake models. An overview of different multi-model ensemble methods and their feasibility for wake applications is presented. The methods include Reliability Ensemble Averaging, Bayesian Model Averaging, and Monte Carlo Simulations. The methodologies are evaluated using data from wake vortex field experiments.

Koerner, Stephan

Prediction of Weather Impacted Airport Capacity using Ensemble Learning

Ensemble learning with the Bagging Decision Tree (BDT) model was used to assess the impact of weather on airport capacities at selected high-demand airports in the United States. The ensemble bagging decision tree models were developed and validated using the Federal Aviation Administration (FAA) Aviation System Performance Metrics (ASPM) data and weather forecast at these airports. The study examines the performance of BDT, along with traditional single Support Vector Machines (SVM), for airport runway configuration selection and airport arrival rates (AAR) prediction during weather impacts. Testing of these models was accomplished using observed weather, weather forecast, and airport operation information at the chosen airports. The experimental results show that ensemble methods are more accurate than a single SVM classifier. The airport capacity ensemble method presented here can be used as a decision support model that supports air traffic flow management to meet the weather impacted airport capacity in order to reduce costs and increase safety.

Weather impact

Virtual Sensors: Using Data Mining to Efficiently Estimate Spectra

Detecting clouds within a satellite image is essential for retrieving surface geophysical parameters, such as albedo and temperature, from optical and thermal imagery because the retrieval methods tend to be valid for clear skies only. Thus, routine satellite data processing requires reliable automated cloud detection algorithms that are applicable to many surface types. Unfortunately, cloud detection over snow and ice is difficult due to the lack of spectral contrast between clouds and snow. Snow and clouds are both highly reflective in the visible wavelen,ats and often show little contrast in the thermal Infrared. However, at 1.6 microns, the spectral signatures of snow and clouds differ enough to allow improved snow/ice/cloud discrimination. The recent Terra and Aqua Moderate Resolution Imaging Spectro-Radiometer (MODIS) sensors have a channel (channel 6) at 1.6 microns. Presently the most comprehensive, long-term information on surface albedo and temperature over snow- and ice-covered surfaces comes from the Advanced Very High Resolution Radiometer ( AVHRR) sensor that has been providing imagery since July 1981. The earlier AVHRR sensors (e.g. AVHRR/2) did not however have a channel designed for discriminating clouds from snow, such as the 1.6 micron channel available on the more recent AVHRR/3 or the MODIS sensors. In the absence of the 1.6 micron channel, the AVHRR Polar Pathfinder (APP) product performs cloud detection using a combination of time-series analysis and multispectral threshold tests based on the satellite's measuring channels to produce a cloud mask. The method has been found to work reasonably well over sea ice, but not so well over the ice sheets. Thus, improving the cloud mask in the APP dataset would be extremely helpful toward increasing the accuracy of the albedo and temperature retrievals, as well as extending the time-series of albedo and temperature retrievals from the more recent sensors to the historical ones. In this work, we use data mining methods to construct a model of MODIS channel 6 as a function of other channels that are common to both MODIS and AVHRR. The idea is to use the model to generate the equivalent of MODIS channel 6 for AVHRR as a function of the AVHRR equivalents to MODIS channels. We call this a Virtual Sensor because it predicts unmeasured spectra. The goal is to use this virtual channel 6. to yield a cloud mask superior to what is currently used in APP . Our results show that several data mining methods such as multilayer perceptrons (MLPs), ensemble methods (e.g., bagging), and kernel methods (e.g., support vector machines) generate channel 6 for unseen MODIS images with high accuracy. Because the true channel 6 is not available for AVHRR images, we qualitatively assess the virtual channel 6 for several AVHRR images.

Srivastava, Ashok

Similar Estimates of Temperature Impacts on Global Wheat Yield by Three Independent Methods

The potential impact of global temperature change on global crop yield has recently been assessed with different methods. Here we show that grid-based and point-based simulations and statistical regressions (from historic records), without deliberate adaptation or CO2 fertilization effects, produce similar estimates of temperature impact on wheat yields at global and national scales. With a 1 C global temperature increase, global wheat yield is projected to decline between 4.1% and 6.4%. Projected relative temperature impacts from different methods were similar for major wheat-producing countries China, India, USA and France, but less so for Russia. Point-based and grid-based simulations, and to some extent the statistical regressions, were consistent in projecting that warmer regions are likely to suffer more yield loss with increasing temperature than cooler regions. By forming a multi-method ensemble, it was possible to quantify 'method uncertainty' in addition to model uncertainty. This significantly improves confidence in estimates of climate impacts on global food security.

Climate impacts

Exploring uncertainties in global crop yield projections in a large ensemble of crop models and CMIP5 and CMIP6 climate scenarios

Concerns over climate change are motivated in large part because of their impact on human society. Assessing the effect of that uncertainty on specific potential impacts is demanding, since it requires a systematic survey over both climate and impacts models. We provide a comprehensive evaluation of uncertainty in projected crop yields for maize, spring and winter wheat, rice, and soybean, using a suite of nine crop models and up to 45 CMIP5 and 34 CMIP6 climate projections for three different forcing scenarios. To make this task computationally tractable, we use a new set of statistical crop model emulators. We find that climate and crop models contribute about equally to overall uncertainty. While the ranges of yield uncertainties under CMIP5 and CMIP6 projections are similar, median impact in aggregate total caloric production is typically more negative for the CMIP6 projections (+1% to −19%) than for CMIP5 (+5% to −13%). In the first half of the 21st century and for individual crops is the spread across crop models typically wider than that across climate models, but we find distinct differences between crops: globally, wheat and maize uncertainties are dominated by the crop models, but soybean and rice are more sensitive to the climate projections. Climate models with very similar global mean warming can lead to very different aggregate impacts so that climate model uncertainties remain a significant contributor to agricultural impacts uncertainty. These results show the utility of large-ensemble methods that allow comprehensively evaluating factors affecting crop yields or other impacts under climate change. The crop model ensemble used here is unbalanced and pulls the assumption that all projections are equally plausible into question. Better methods for consistent model testing, also at the level of individual processes, will have to be developed and applied by the crop modeling community.

crop yield projections

Statistical properties of ideal three-dimensional magnetohydrodynamics

Classical Gibbs ensemble methods are used to study the spectral structure of three-dimensional ideal MHD in periodic geometry. In this paper the equilibrium ensemble incorporates constraints of total energy, magnetic helicity, and cross helicity. Several new results are proven for ensemble averages, including the constraint that magnetic energy equal or exceed kinetic energy, and that cross helicity represents a constant fraction of magnetic energy across the spectral domain, for arbitrary size systems. Two zero-temperature limits are considered in detail, emphasizing the role of complete and partial condensaiton of spectral quantities to the longest wavelength states. The ensemble predictions are compared to direct numerical solution using a low-order truncation Galerkin spectral code. Implications for spectral transfer of nonequilibrium, dissipative turbulent MHD systems are discussed.

Stribling, T.

Improving Climate Projections Using "Intelligent" Ensembles

Recent changes in the climate system have led to growing concern, especially in communities which are highly vulnerable to resource shortages and weather extremes. There is an urgent need for better climate information to develop solutions and strategies for adapting to a changing climate. Climate models provide excellent tools for studying the current state of climate and making future projections. However, these models are subject to biases created by structural uncertainties. Performance metrics-or the systematic determination of model biases-succinctly quantify aspects of climate model behavior. Efforts to standardize climate model experiments and collect simulation data-such as the Coupled Model Intercomparison Project (CMIP)-provide the means to directly compare and assess model performance. Performance metrics have been used to show that some models reproduce present-day climate better than others. Simulation data from multiple models are often used to add value to projections by creating a consensus projection from the model ensemble, in which each model is given an equal weight. It has been shown that the ensemble mean generally outperforms any single model. It is possible to use unequal weights to produce ensemble means, in which models are weighted based on performance (called "intelligent" ensembles). Can performance metrics be used to improve climate projections? Previous work introduced a framework for comparing the utility of model performance metrics, showing that the best metrics are related to the variance of top-of-atmosphere outgoing longwave radiation. These metrics improve present-day climate simulations of Earth's energy budget using the "intelligent" ensemble method. The current project identifies several approaches for testing whether performance metrics can be applied to future simulations to create "intelligent" ensemble-mean climate projections. It is shown that certain performance metrics test key climate processes in the models, and that these metrics can be used to evaluate model quality in both current and future climate states. This information will be used to produce new consensus projections and provide communities with improved climate projections for urgent decision-making.

Baker, Noel C.

Improving Climate Projections Using "Intelligent" Ensembles

Recent changes in the climate system have led to growing concern, especially in communities which are highly vulnerable to resource shortages and weather extremes. There is an urgent need for better climate information to develop solutions and strategies for adapting to a changing climate. Climate models provide excellent tools for studying the current state of climate and making future projections. However, these models are subject to biases created by structural uncertainties. Performance metrics-or the systematic determination of model biases-succinctly quantify aspects of climate model behavior. Efforts to standardize climate model experiments and collect simulation data-such as the Coupled Model Intercomparison Project (CMIP)-provide the means to directly compare and assess model performance. Performance metrics have been used to show that some models reproduce present-day climate better than others. Simulation data from multiple models are often used to add value to projections by creating a consensus projection from the model ensemble, in which each model is given an equal weight. It has been shown that the ensemble mean generally outperforms any single model. It is possible to use unequal weights to produce ensemble means, in which models are weighted based on performance (called "intelligent" ensembles). Can performance metrics be used to improve climate projections? Previous work introduced a framework for comparing the utility of model performance metrics, showing that the best metrics are related to the variance of top-of-atmosphere outgoing longwave radiation. These metrics improve present-day climate simulations of Earth's energy budget using the "intelligent" ensemble method. The current project identifies several approaches for testing whether performance metrics can be applied to future simulations to create "intelligent" ensemble-mean climate projections. It is shown that certain performance metrics test key climate processes in the models, and that these metrics can be used to evaluate model quality in both current and future climate states. This information will be used to produce new consensus projections and provide communities with improved climate projections for urgent decision-making.

Baker, Noel C.

Artificial Neural Network Modeling for Airline Disruption Management

Since the 1970s, most airlines have incorporated computerized support for managing disruptions during flight schedule execution. However, existing platforms for airline disruption management (ADM) employ monolithic system design methods that rely on the creation of specific rules and requirements through explicit optimization routines, before a system that meets the specifications is designed. Thus, current platforms for ADM are unable to readily accommodate additional system complexities resulting from the introduction of new capabilities, such as the introduction of unmanned aerial systems (UAS), operations and infrastructure, to the system. To this end, we use historical data on airline scheduling and operations recovery to develop a system of artificial neural networks (ANNs), which describe a predictive transfer function model (PTFM) for promptly estimating the recovery impact of disruption resolutions at separate phases of flight schedule execution during ADM. Furthermore, we provide a modular approach for assessing and executing the PTFM by employing a parallel ensemble method to develop generative routines that amalgamate the system of ANNs. Our modular approach ensures that current industry standards for tardiness in flight schedule execution during ADM are satisfied, while accurately estimating appropriate time-based performance metrics for the separate phases of flight schedule execution.

Kolawole Ogunsina

Basin-Scale Assessment of the Land Surface Energy Budget in the National Centers for Environmental Prediction Operational and Research NLDAS-2 Systems

This paper compares the annual and monthly components of the simulated energy budget from the North American Land Data Assimilation System phase 2 (NLDAS-2) with reference products over the domains of the 12 River Forecast Centers (RFCs) of the continental United States (CONUS). The simulations are calculated from both operational and research versions of NLDAS-2. The reference radiation components are obtained from the National Aeronautics and Space Administration Surface Radiation Budget product. The reference sensible and latent heat fluxes are obtained from a multitree ensemble method applied to gridded FLUXNET data from the Max Planck Institute, Germany. As these references are obtained from different data sources, they cannot fully close the energy budget, although the range of closure error is less than 15%formean annual results. The analysis here demonstrates the usefulness of basin-scale surface energy budget analysis for evaluating model skill and deficiencies. The operational (i.e., Noah, Mosaic, and VIC) and research (i.e., Noah-I and VIC4.0.5) NLDAS-2 land surface models exhibit similarities and differences in depicting basin-averaged energy components. For example, the energy components of the five models have similar seasonal cycles, but with different magnitudes. Generally, Noah and VIC overestimate (underestimate) sensible (latent) heat flux over several RFCs of the eastern CONUS. In contrast, Mosaic underestimates (overestimates) sensible (latent) heat flux over almost all 12 RFCs. The research Noah-I and VIC4.0.5 versions show moderate-to-large improvements (basin and model dependent) relative to their operational versions, which indicates likely pathways for future improvements in the operational NLDAS-2 system.

Energy

Assimilation of Lidar Planetary Boundary Layer Height Observations

Lidar backscatter and wind retrievals of the planetary boundary layer height (PBLH) are assimilated into 22 hourly forecasts from the NASA Unified - Weather and Research Forecast (NU-WRF) model during the Plains Elevated Convection Convection at Night (PECAN) campaign on July 11, 2015 in Greensburg, Kansas, using error statistics collected from the model profiles to compute the necessary covariance matrices. Two separate forecast runs using different PBL physics schemes were employed, and comparisons with 6 independent radiosonde profiles were made for each run. Both of the forecast runs accurately predicted the PBLH and the state variable profiles within the planetary boundary layer during the early morning, and the assimilation had a small impact during this time. In the late afternoon, the forecast runs showed decreased accuracy as the convective boundary layer developed. However, assimilation of the Doppler lidar PBLH observations were found to improve the temperature and V velocity profiles relative to independent radiosonde profiles. Water vapor was overcorrected, leading to increased differences with independent data. Errors in the U velocity were made slightly larger. The computed forecast error covariances between the PBLH and state variables were found to rise in the late afternoon, leading to the larger improvements in the afternoon. This work represents the first effort to assimilate PBLH into forecast states using ensemble methods.

Andrew Tangborn

Underlying Fundamentals of Kalman Filtering for River Network Modeling

The grand challenge of producing hydrometeorological estimates every time and everywhere has motivated the fusion of sparse observations with dense numerical models, with a particular interest on discharge in river modeling. Ensemble methods are largely preferred as they enable the estimation of error properties, but at the expense of computational load and generally with underestimations. These imperfect stochastic estimates motivate the use of correction methods, that is, error localization and inflation, although the physical justifications for their optimality are limited. The purpose of this study is to use one of the simplest forms of data assimilation when applied to river modeling and reveal the underlying mechanisms impacting its performance. Our framework based on assimilating daily averaged in situ discharge measurements to correct daily averaged runoff was tested over a 4-yr case study of two rivers in Texas. Results show that under optimal conditions of inflation and localization, discharge simulations are consistently improved such that the mean values of Nash–Sutcliffe efficiency are enhanced from211.32 to 0.55 at observed gauges and from212.24 to21.10 at validation gauges. Yet, parameters controlling the inflation and the localization have a large impact on the performance. Further investigations of these sensitivities showed that optimal inflation occurs when compensating exactly for discrepancies in the magnitude of errors while optimal localization matches the distance traveled during one assimilation window. These results may be applicable to more advanced data assimilation methods as well as for larger applications motivated by upcoming river-observing satellite missions, such as NASA’s Surface Water and Ocean Topography mission.

Streamflow

Machine Learning Models to Predict Cognitive Impairment of Rodents Subjected to Space Radiation

This research uses machine-learned computational analyses to predict the cognitive performance impairment of rats induced by irradiation. The experimental data in the analyses is from a rodent model exposed to ≤ 15 cGy of individual Galactic Cosmic Radiation (GCR) ions: 4He, 16O, 28Si, 48Ti, or 56Fe, expected for a Lunar or Mars mission. This work investigates rats at a subject-based level and uses performance scores taken before irradiation to predict impairment in Attentional Set-shifting (ATSET) data post-irradiation. Here, the worst performing rats of the control group define the impairment thresholds based on population analyses via cumulative distribution functions, leading to the labeling of impairment for each subject. A significant finding is the exhibition of a dose-dependent increasing probability of impairment for 1 to 10 cGy of 28Si or 56Fe in the Simple Discrimination (SD) stage of the ATSET, and for 1 to 10 cGy of 56Fe in the Compound Discrimination (CD) stage. On a subject-based level, implementing Machine Learning (ML) classifiers such as the Gaussian Naïve Bayes, Support Vector Machine, and Artificial Neural Networks identifies rats that have a higher tendency for impairment after GCR exposure. The algorithms employ the experimental prescreenperformance scores as multidimensional input features to predict each rodent’s susceptibility to cognitive impairment due to space radiation exposure. The receiver operating characteristic and the precision-recall curves of the ML models show a better prediction of impairment when 56Feis the ion in question in both SD and CD stages. They, however, do not depict impairment due to 4Hein SD and 28Siin CD, suggesting no dose-dependent impairment response in these cases. One key finding of our study is that prescreen performance scores can be used to predict the ATSET performance impairments. This result is significant to crewed space missions as it supports the potential of predicting an astronaut’s impairment in a specific task before spaceflight through the implementation of appropriately trained ML tools. Future research can focus on constructing ML ensemble methods to integrate the findings from the methodologies implemented in this study for morerobust predictionsof cognitive decrements due to space radiation exposure.

space radiation

Automated Data Accountability for Missions in Mars Rover Data

As the Mars Curiosity Rover transmits data to the JPL Ground Data System (GDS), it frequently observes data loss and corruption, requiring re-transmits from the rover and Ground Data System Analysts (GDSA) to monitor the downlink process. As new missions are launched, the GDSA team redistributes analysts to these new missions, causing shortages in previous missions. The GDSA team can significantly benefit from the automation and optimization of the downlink process of telemetry data. In fact, there is a need for a better understanding of why the data is corrupted, so that the GDSA team can best determine the root cause of the issues in the GDS. This paper presents machine learning and deep learning based approaches to automate and optimize the detection of data loss. We first created a pipeline to automatically accumulate data from the telemetry databases (MAROS, Telemetry Data Storage, and GDS Elastic Search Database) in the downlink process. With our newly created datasets, we perform feature selection to supplement the GDSA understanding of the downlink process and provide supplemental analysis on the importance of different features. We implement various machine learning and deep learning based models, including support vector machines, ensemble methods, and deep neural networks and evaluate their accuracies in identifying whether a downlink process is complete or incomplete. We utilize fast hyperparameter optimization methods that allow our models to quickly be re-trained, allowing them to quickly be tuned and optimized on daily incoming data in real time. This hyperparameter optimization also allows our methods to be quickly integrated into other JPL missions. Our results show that our best-performing machine learning and deep learning based models outperform the existing GDSA detection software by 6 accuracy points and can aid analysts by providing insights into the data accountability problem. Since these various machine learning and deep learning approaches vary significantly in interpretability, we provide a discussion on the tradeoffs between their performance and trustworthiness in helping detect issues in data transmission.

Divsalar, Dariush