Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “ensemble learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

An Inverse Chance-constrained Approach to the Calibration of Robust Models

This paper proposes a strategy to calibrate computational models according to uncertain input-output data. To this end, uncertainty in the data is first quantified by creating adversarial data sets. Samples drawn from such sets are then mapped from the input-output space to the parameter space using an inverse mapping. This mapping minimizes the collective output spread of an ensemble of point predictions while satisfying a set of individual data-matching requirements. The distribution of the resulting parameter points, which often exhibits strong parameter dependencies, is then modeled using sliced-normals. The chance-constrained formulation used to learn this distribution enables the analyst to trade-off a greater likelihood for most of the data against a lower likelihood for some of the data thereby relaxing the conservatism of the calibrated model. This formulation not only neglects the worst-performing quantiles of each adversarial distribution but also eliminates the potentially serious effects that outliers might have on the resulting model. This calibration approach not only has a considerably lower computational cost than the standard forward approach but it also allows for the identification of suitable distribution classes, which in turn yield better calibrated models.

Calibration↗

Landslide Hazard is Projected to Increase Across High Mountain Asia

High Mountain Asia has long been known as a hotspot for landslide risk, and studies have suggested that landslide hazard is likely to increase in this region over the coming decades. Extreme precipitation may become more frequent, with a nonlinear response relative to increasing global temperatures. However, these changes are geographically varied. This article maps probable changes to landslide hazard, as shown by a landslide hazard indicator (LHI) derived from downscaled precipitation and temperature. In order to capture the nonlinear response of slopes to extreme precipitation, a simple machine-learning model was trained on a database of landslides across High Mountain Asia to develop a regional LHI. This model was applied to statistically downscaled data from the 30 members of the Seamless System for Prediction and Earth System Research large ensembles to produce a range of possible outcomes under the Shared Socioeconomic Pathways 2-4.5 and 5-8.5. The LHI reveals that landslide hazard will increase in most parts of High Mountain Asia. Absolute increases will be highest in already hazardous areas such as the Central Himalaya, but relative change is greatest on the Tibetan Plateau. Even in regions where landslide hazard declines by year 2100, it will increase prior to the mid-century mark. However, the seasonal cycle of landslide occurrence will not change greatly across High Mountain Asia. Although substantial uncertainty remains in these projections, the overall direction of change seems reliable. These findings highlight the importance of continued analysis to inform disaster risk reduction strategies for stakeholders across High Mountain Asia.

Thomas A Stanley↗

An Ensemble Approach to Building Mercer Kernels with Prior Information

This paper presents a new methodology for automatic knowledge driven data mining based on the theory of Mercer Kernels, which are highly nonlinear symmetric positive definite mappings from the original image space to a very high, possibly dimensional feature space. we describe a new method called Mixture Density Mercer Kernels to learn kernel function directly from data, rather than using pre-defined kernels. These data adaptive kernels can encode prior knowledge in the kernel using a Bayesian formulation, thus allowing for physical information to be encoded in the model. Specifically, we demonstrate the use of the algorithm in situations with extremely small samples of data. We compare the results with existing algorithms on data from the Sloan Digital Sky Survey (SDSS) and demonstrate the method's superior performance against standard methods. The code for these experiments has been generated with the AUTOBAYES tool, which automatically generates efficient and documented C/C++ code from abstract statistical model specifications. The core of the system is a schema library which contains templates for learning and knowledge discovery algorithms like different versions of EM, or numeric optimization methods like conjugate gradient methods. The template instantiation is supported by symbolic-algebraic computations, which allows AUTOBAYES to find closed-form solutions and, where possible, to integrate them into the code.

Srivastava, Ashok N.↗

Development and Implementation of a Comprehensive Radiometric Validation Protocol for the CERES Earth Radiation Budget Climate Record Sensors

The CERES Flight Models 1 through 4 instruments were launched aboard NASA's Earth Observing System (EOS) Terra and Aqua Spacecraft into 705 Km sun-synchronous orbits with 10:30 a.m. and 1:30 p.m. equatorial crossing times. These instruments supplement measurements made by the CERES Proto Flight Model (PFM) instrument launched aboard NASA's Tropical Rainfall Measuring Mission (TRMM) into a 350 Km, 38-degree mid-inclined orbit. CERES Climate Data Records consist of geolocated and calibrated instantaneous filtered and unfiltered radiances through temporally and spatially averaged TOA, Surface and Atmospheric fluxes. CERES filtered radiance measurements cover three spectral bands including shortwave (0.3 to 5 microns), total (0.3 to 100 microns) and an atmospheric window channel (8 to 12 microns). The CERES Earth Radiation Budget measurements represent a new era in radiation climate data, realizing a factor of 2 to 4 improvement in calibration accuracy and stability over the previous ERBE climate records, while striving for the next goal of 0.3-percent per decade absolute stability. The current improvement is derived from two sources: the incorporation of lessons learned from the ERBE mission in the design of the CERES instruments and the development of a rigorous and comprehensive radiometric validation protocol consisting of individual studies covering different spatial, spectral and temporal time scales on data collected both pre and post launch. Once this ensemble of individual perspectives is collected and organized, a cohesive and highly rigorous picture of the overall end-to-end performance of the CERES instrument's and data processing algorithms may be clearly established. This approach has resulted in unprecedented levels of accuracy for radiation budget instruments and data products with calibration stability of better than 0.2-percent and calibration traceability from ground to flight of 0.25-percent. The current work summarizes the development, philosophy and implementation of the protocol designed to rigorously quantify the quality of the data products as well as the level of agreement between the CERES TRMM, Terra and Aqua climate data records.

Priestley, K. J.↗

Comparisons of some large scientific computers

In 1975, the National Aeronautics and Space Administration (NASA) began studies to assess the technical and economic feasibility of developing a computer having sustained computational speed of one billion floating point operations per second and a working memory of at least 240 million words. Such a powerful computer would allow computational aerodynamics to play a major role in aeronautical design and advanced fluid dynamics research. Based on favorable results from these studies, NASA proceeded with developmental plans. The computer was named the Numerical Aerodynamic Simulator (NAS). To help insure that the estimated cost, schedule, and technical scope were realistic, a brief study was made of past large scientific computers. Large discrepancies between inception and operation in scope, cost, or schedule were studied so that they could be minimized with NASA's proposed new compter. The main computers studied were the ILLIAC IV, STAR 100, Parallel Element Processor Ensemble (PEPE), and Shuttle Mission Simulator (SMS) computer. Comparison data on memory and speed were also obtained on the IBM 650, 704, 7090, 360-50, 360-67, 360-91, and 370-195; the CDC 6400, 6600, 7600, CYBER 203, and CYBER 205; CRAY 1; and the Advanced Scientific Computer (ASC). A few lessons learned conclude the report.

Credeur, K. R.↗

Machine Learning Models to Predict Cognitive Impairment of Rodents Subjected to Space Radiation

This research uses machine-learned computational analyses to predict the cognitive performance impairment of rats induced by irradiation. The experimental data in the analyses is from a rodent model exposed to ≤ 15 cGy of individual Galactic Cosmic Radiation (GCR) ions: 4He, 16O, 28Si, 48Ti, or 56Fe, expected for a Lunar or Mars mission. This work investigates rats at a subject-based level and uses performance scores taken before irradiation to predict impairment in Attentional Set-shifting (ATSET) data post-irradiation. Here, the worst performing rats of the control group define the impairment thresholds based on population analyses via cumulative distribution functions, leading to the labeling of impairment for each subject. A significant finding is the exhibition of a dose-dependent increasing probability of impairment for 1 to 10 cGy of 28Si or 56Fe in the Simple Discrimination (SD) stage of the ATSET, and for 1 to 10 cGy of 56Fe in the Compound Discrimination (CD) stage. On a subject-based level, implementing Machine Learning (ML) classifiers such as the Gaussian Naïve Bayes, Support Vector Machine, and Artificial Neural Networks identifies rats that have a higher tendency for impairment after GCR exposure. The algorithms employ the experimental prescreenperformance scores as multidimensional input features to predict each rodent’s susceptibility to cognitive impairment due to space radiation exposure. The receiver operating characteristic and the precision-recall curves of the ML models show a better prediction of impairment when 56Feis the ion in question in both SD and CD stages. They, however, do not depict impairment due to 4Hein SD and 28Siin CD, suggesting no dose-dependent impairment response in these cases. One key finding of our study is that prescreen performance scores can be used to predict the ATSET performance impairments. This result is significant to crewed space missions as it supports the potential of predicting an astronaut’s impairment in a specific task before spaceflight through the implementation of appropriately trained ML tools. Future research can focus on constructing ML ensemble methods to integrate the findings from the methodologies implemented in this study for morerobust predictionsof cognitive decrements due to space radiation exposure.

space radiation↗

Machine Learning Applications to Metal-Silicate Equilibria and their Insights into Core Formation

An extensive number of studies have experimentally investigated how elements distribute between metal and silicate phases, to better constrain core-mantle chemical equilibrium. Here, we present a new database compiling all (to our knowledge) experimental data on liquid metal-silicate partitioning from 118 peer-reviewed publications. We applied various machine learning techniques to gain further insights into these partitioning equilibria and their dependencies. We performed a network analysis to investigate the relationship between experiments and partition coefficients, which enables visualizing gaps in the experimental dataset and biases related to varying experimental conditions and analytical setup. In addition, semi-empirical thermodynamic models are commonly used to extrapolate these chemical reactions to the wide range of pressure, temperature and compositional conditions of planetary differentiation. These models are based on linear regressions that assume continuous relationship between partition coefficients and experimental variables. Here, we considered random forest regressions, which are algorithms based on ensembles of decision trees and does not consider continuous effects of each variable. The application of this regression significantly improves the prediction of metal-silicate partitioning for several elements including Ni, Si and Cr. We will show how this new approach improves our understanding of elemental exchange between metal and silicate and their implications for the Earth’s core formation.

siderophile element↗

The Ensemble Canon

Ensemble is an open architecture for the development, integration, and deployment of mission operations software. Fundamentally, it is an adaptation of the Eclipse Rich Client Platform (RCP), a widespread, stable, and supported framework for component-based application development. By capitalizing on the maturity and availability of the Eclipse RCP, Ensemble offers a low-risk, politically neutral path towards a tighter integration of operations tools. The Ensemble project is a highly successful, ongoing collaboration among NASA Centers. Since 2004, the Ensemble project has supported the development of mission operations software for NASA's Exploration Systems, Science, and Space Operations Directorates.

applications programs (computers),↗

Exploring Anomalous PM 2.5 from Wildfires and Dust Storms using Data and Services at NASA GES DISC

The presence of fine particles in the atmosphere with a diameter of less than 2.5 µm, called particulate matter 2.5 (PM 2.5 ), poses a significant threat to human health as a criteria air pollutant. Fortunately, NASA's Goddard Earth Sciences Data and Information Services Center (GES DISC) provides easy access to several PM 2.5 concentration products. These datasets include the reanalysis of global hourly and monthly aerosol components including PM 2.5 data from the Modern-Era Retrospective analysis for Research and Applications, version 2 (MERRA-2), as well as 3-hourly real-time ensemble forecasts of PM 2.5 from the Hazardous Air Quality Ensemble System (HAQES). The HAQES products are developed by the George Mason University Air Quality Laboratory as part of NASA's Health Air Quality Applied Science Team (HAQAST). The GES DISC is actively collaborating with scientists in the HAQAST program to further expand air quality data collections. Two new datasets are currently being archived: one is the machine learning-based global hourly PM 2.5 derived from MERRA-2; the other is the localized data (NO 2 , O 3 , and PM 2.5 ) time series derived from NASA's GEOS Composition Forecasting (GEOS-CF) system. In this presentation, we will explore the spatial patterns and long-distance transport characteristics of elevated PM 2.5 during extreme pollution events, such as the June 2023 Canadian wildfires, which are still active at the time of writing; and severe spring dust storms in 2023 over Asia. To gain comprehensive insights, we will utilize various PM 2.5 data in conjunction with satellite-observed aerosol data from TROPOspheric Monitoring Instrument (TROPOMI) on Sentinel-5P. The primary focus of this presentation will be to demonstrate effective use of data tools and services to visualize and explore extreme air pollution phenomena. Additionally, we will provide guidance on how users can download specific data of interest, facilitating further analysis and research in this critical area.

air quality↗

State Predictor of Classification Cognitive Engine Applied to Channel Fading

This study presents the application of machine learning (ML) to a space-to-ground communication link, showing how ML can be used to detect the presence of detrimental channel fading. Using this channel state information, the communication link can be used more efficiently by reducing the amount of lost data during fading. The motivation for this work is based on channel fading observed during on-orbit operations with NASA's Space Communication and Navigation (SCaN) testbed on the International Space Station (ISS). This paper presents the process to extract a target concept (fading and not-fading) from the raw data. The pre-processing and data exploration effort is explained in detail, with a list of assumptions made for parsing and labelling the dataset. The model selection process is explained, specifically emphasizing the benefits of using an ensemble of algorithms with majority voting for binary classification of the channel state. Experimental results are shown, highlighting how an end-to-end communication system can utilize knowledge of the channel fading status to identity fading and take appropriate action. With a laboratory testbed to emulate channel fading, the overall performance is compared to standard adaptive methods without fading knowledge, such as adaptive coding and modulation.

Fading↗

Exploring New Pathways in Precipitation Assimilation

Precipitation assimilation poses a special challenge in that the forward model for rain in a global forecast system is based on parameterized physics, which can have large systematic errors that must be rectified to use precipitation data effectively within a standard statistical analysis framework. We examine some key issues in precipitation assimilation and describe several exploratory studies in assimilating rainfall and latent heating information in NASA's global data assimilation systems using the forecast model as a weak constraint. We present results from two research activities. The first is the assimilation of surface rainfall data using a time-continuous variational assimilation based on a column model of the full moist physics. The second is the assimilation of convective and stratiform latent heating retrievals from microwave sensors using a variational technique with physical parameters in the moist physics schemes as a control variable. We will show the impact of assimilating these data on analyses and forecasts. Among the lessons learned are (1) that the time-continuous application of moisture/temperature tendency corrections to mitigate model deficiencies offers an effective strategy for assimilating precipitation information, and (2) that the model prognostic variables must be allowed to directly respond to an improved rain and latent heating field within an analysis cycle to reap the full benefit of assimilating precipitation information. of microwave radiances versus retrieval information in raining areas, and initial efforts in developing ensemble techniques such as Kalman filter/smoother for precipitation assimilation. Looking to the future, we discuss new research directions including the assimilation

Hou, Arthur↗

Feature Selection in High-Dimensional Space with Applications to Gene Expression Data

Recent years have seen rapid growth in high-dimensional datasets. Most existing machine learning (ML) algorithms fail in high-dimensional settings where many features could be redundant. A critical process of feature selection is thus applied in such a setting that helps in identifying the most relevant features while removing redundant ones. With the increase in high dimensionality, one is also faced with problems of efficiency and interpretation in performing such selection methods. Therefore, this paper proposes a “novel” feature selection framework that uses an ensemble of interpretable ML algorithms to perform feature selection and the ranking of final features. Finally, this framework is applied to a gene expression dataset obtained through collaboration with the National Aeronautics and Space Administration (NASA)’s Biological and Physical Sciences (BPS) team and helps identify important and relevant genes contributing to specific target attributes through classification tasks.

Nishan Pantha↗

CATIA V5 Virtual Environment Support for Constellation Ground Operations

This summer internship primarily involved using CATIA V5 modeling software to design and model parts to support ground operations for the Constellation program. I learned several new CATIA features, including the Imagine and Shape workbench and the Tubing Design workbench, and presented brief workbench lessons to my co-workers. Most modeling tasks involved visualizing design options for Launch Pad 39B operations, including Mobile Launcher Platform (MLP) access and internal access to the Ares I rocket. Other ground support equipment, including a hydrazine servicing cart, a mobile fuel vapor scrubber, a hypergolic propellant tank cart, and a SCAPE (Self Contained Atmospheric Protective Ensemble) suit, was created to aid in the visualization of pad operations.

Kelley, Andrew↗

Vegetation Greening Mitigates the Impacts of Increasing Extreme Rainfall on Runoff Events

Future flood risk assessment has primarily focused on heavy rainfall as the main driver, with the assumption that projected increases in extreme rain events will lead to subsequent flooding. However, the presence of and changes in vegetation have long been known to influence the relationship between rainfall and runoff. Here, we extract historical (1850–1880) and projected (2070–2100) daily extreme rainfall events, the corresponding runoff, and antecedent conditions simulated in a prominent large Earth system model ensemble to examine the shifting extreme rainfall and runoff relationship. Even with widespread projected increases in the magnitude (78% of the land surface) and number (72%) of extreme rainfall events, we find projected declines in event-based runoff ratio (runoff/rainfall) for a majority (57%) of the Earth surface. Runoff ratio declines are linked with decreases in antecedent soil water driven by greater transpiration and canopy evaporation (both linked to vegetation greening) compared to areas with runoff ratio increases. Using a machine learning regression tree approach, we find that changes in canopy evaporation is the most important variable related to changes in antecedent soil water content in areas of decreased runoff ratios (with minimal changes in antecedent rainfall) while antecedent ground evaporation is the most important variable in areas of increased runoff ratios. Our results suggest that simulated interactions between vegetation greening, increasing evaporative demand, and antecedent soil drying are projected to diminish runoff associated with extreme rainfall events, with important implications for society.

soil moisture↗

Simulating Activities: Relating Motives, Deliberation and Attentive Coordination

Activities are located behaviors, taking time, conceived as socially meaningful, and usually involving interaction with tools and the environment. In modeling human cognition as a form of problem solving (goal-directed search and operator sequencing), cognitive science researchers have not adequately studied "off-task" activities (e.g., waiting), non-intellectual motives (e.g., hunger), sustaining a goal state (e.g., playful interaction), and coupled perceptual-motor dynamics (e.g., following someone). These aspects of human behavior have been considered in bits and pieces in past research, identified as scripts, human factors, behavior settings, ensemble, flow experience, and situated action. More broadly, activity theory provides a comprehensive framework relating motives, goals, and operations. This paper ties these ideas together, using examples from work life in a Canadian High Arctic research station. The emphasis is on simulating human behavior as it naturally occurs, such that "working" is understood as an aspect of living. The result is a synthesis of previously unrelated analytic perspectives and a broader appreciation of the nature of human cognition. Simulating activities in this comprehensive way is useful for understanding work practice, promoting learning, and designing better tools, including human-robot systems.

Clancey, William J.↗

Incremental Learning for Passive Microwave Precipitation Retrievals using Advanced Technology Microwave Sounder

Spaceborne passive microwave (PMW) radiometry is central to global precipitation monitoring, yet retrieval uncertainties remain substantial, particularly for cross-track sounders whose variable footprints and channel configurations are optimized for atmospheric temperature and moisture profiling rather than precipitation. Consequently, existing operational products often exhibit angular-dependent biases, limited effective swath utilization, unrealistic rainfall probability distributions, and systematic misclassification of precipitation phase. These limitations are further compounded by the scarcity of globally accurate and representative precipitation observations, as training data from the Dual-frequency Precipitation Radar (DPR) and the Cloud Profiling Radar (CPR) are spatially sparse, lack uniform global coverage, and exhibit heterogeneous error characteristics across precipitation regimes. To address these challenges, this study presents a supervised retrieval algorithm that incrementally trains an ensemble of extreme gradient-boosted decision trees by augmenting base learners with pre-training on reanalysis data and post-training on coincident DPR and CPR observations matched with the Advanced Technology Microwave Sounder (ATMS). By transferring prior information from reanalysis to posterior constraints from radar observations and adopting a sequential detection–estimation strategy for precipitation phase and rate retrieval, the proposed approach yields retrievals across the full ATMS swath that are largely free from persistent deficiencies in current Global Precipitation Measurement (GPM) passive microwave operational products. In particular, the method resolves bimodal artifacts in rainfall retrievals and mitigates systematic high-latitude snowfall biases, including overestimation across the Arctic and underestimation across the Antarctic. Validation against independent Multi-Radar Multi-Sensor (MRMS) data over the Contiguous United States (CONUS) further demonstrates improved performance in precipitation phase detection and rate estimation relative to both reanalysis and current GPM PMW products.

Mahyar Garshasbi↗

Transcriptomics-based Machine Learning Analysis Predicts Space-Exposed Murine Livers

Limited sample sizes, high data dimensionality, and sensitivity to technical and biological variability of next generation sequencing (NGS), has typically limited machine learning (ML) in space studies and further study of radiation effects. However, pooling smaller studies while addressing intra- and inter-study variabilities allows for ML predictive modeling. Here, integration methods were applied to whole transcriptome shotgun sequencing (RNAseq) data from 6 mouse liver GeneLab datasets (GLDS) with a total of 113 spaceflight and ground-control samples to determine top features relevant to spaceflight including the effect of radiation exposure. Data was normalized within each study, then merged and scaled across all datasets. Data dimensionality was reduced using a minimum redundancy maximum relevance (MRMR) methodology. The top MRMR features were used to predict spaceflight vs. ground-control samples using a Random Forest (RF) classifier with 5-fold cross validation (CV). The ML-based gene sets were further compared against differential gene expression results from individual GLDS. CV training using the top 100 MRMR genes show averages of 86% accuracy and 0.95 AUC value on the validation set over 5 folds (Figure 1A). Baseline set analysis on differentially expressed genes (DEGs) identified using padj ≤ 0.05 show 811 or 68 DEGs overlapping between at least 2 or 3 studies, respectively (Figure 1B). Over-representation analysis showed overlapping biological processes related to fatty acid and lipid metabolism. Set analysis between the MRMR features and the DEGs showed 60 or 8 genes overlapping with at least 1 or 2 studies, respectively. MRMR feature selection and ensemble ML methods (e.g. RF) improve performance relative to a Naïve Bayes classifier when NGS data sets are analyzed. A challenge of applying ML methods across heterogeneous NGS data is accounting for signal:noise ratio. Here, signal validation across studies was shown by intersecting sets between top MRMR genes and DEGs from RNASeq analysis. Non-intersecting sets introduce opportunity to explore spaceflight relevant genes and implementing ML methods across existing NGS datasets may overcome sample size limitations. ML coupled with existing analytical methods enhances understanding of disease by revealing common underlying pathways across datasets.

Machine Learning↗