Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “ensemble learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Machine Learning of Key Variables Impacting Extreme Precipitation in Various Regions of the Contiguous United States

Abstract Amplification in extreme precipitation intensity and frequency can cause severe flooding and impose significant social and economic consequences. Variations in extreme precipitation intensity, frequencies, and return periods can be attributed to many physical variables across spatial and temporal scales. Here we employ ensemble machine learning (ML) methods, namely random forest (RF), eXtreme Gradient Boosting (XGB), and artificial neural networks (ANN), to explore key contributing variables to monthly extreme precipitation intensity and frequency in six regions over the United States. We further establish emulators for return periods. Results show that the ML models for intensity perform better in regions with obvious seasonality (i.e., Northern Great Plains, Southern Great Plains, and West Coast) than the other three regions (Northeast, Southwest, and Rocky Mountains), while for frequency the models perform well for most regions. The Shapley additive explanation is used to help explain the relationships between extreme precipitation characteristics and identify top variables for RF and XGB. We find that latent heat flux, relative humidity, soil moisture, and large‐scale subsidence are key common variables across the regions for both monthly intensity and frequency, and their compound effects are non‐negligible. The developed ML models capture the probability and return period of extreme precipitation well for all regions and may be used for decision making (e.g., infrastructure planning and design).

54 ENVIRONMENTAL SCIENCES↗

A meta-learning based distribution system load forecasting model selection framework

This paper presents a meta-learning based, automatic distribution system load forecasting model selection framework. Furthermore, the framework includes the following processes: feature extraction, candidate model preparation and labeling, offline training, and online model recommendation. Using load forecasting needs and data characteristics as input features, multiple metalearners are used to rank the candidate load forecast models based on their forecasting accuracy. Then, a scoring-voting mechanism is proposed to weights recommendations from each meta-leaner and make the final recommendations. Heterogeneous load forecasting tasks with different temporal and technical requirements at different load aggregation levels are set up to train, validate, and test the performance of the proposed framework. Simulation results demonstrate that the performance of the meta-learning based approach is satisfactory in both seen and unseen forecasting tasks.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Earth's record-high greenness and its attributions in 2020

Terrestrial vegetation is a crucial component of Earth's biosphere, regulating global carbon and water cycles and contributing to human welfare. Despite an overall greening trend, terrestrial vegetation exhibits a significant inter-annual variability. The mechanisms driving this variability, particularly those related to climatic and anthropogenic factors, remain poorly understood, which hampers our ability to project the long-term sustainability of ecosystem services. Here, in this work, by leveraging diverse remote sensing measurements, we pinpointed 2020 as a historic landmark, registering as the greenest year in modern satellite records from 2001 to 2020. Using ensemble machine learning and Earth system models, we found this exceptional greening primarily stemmed from consistent growth in boreal and temperate vegetation, attributed to rising CO 2 levels, climate warming, and reforestation efforts, alongside a transient tropical green-up linked to the enhanced rainfall. Contrary to expectations, the COVID-19 pandemic lockdowns had a limited impact on this global greening anomaly. Our findings highlight the resilience and dynamic nature of global vegetation in response to diverse climatic and anthropogenic influences, offering valuable insights for optimizing ecosystem management and informing climate mitigation strategies.

54 ENVIRONMENTAL SCIENCES↗

General-Purpose Unsupervised Cyber Anomaly Detection via Non-Negative Tensor Factorization

Distinguishing malicious anomalous activities from unusual but benign activities is a fundamental challenge for cyber defenders. Prior studies have shown that statistical user behavior analysis yields accurate detections by learning behavior profiles from observed user activity. These unsupervised models are able to generalize to unseen types of attacks by detecting deviations from normal behavior, without knowledge of specific attack signatures. However, approaches proposed to date based on probabilistic matrix factorization are limited by the information conveyed in a two-dimensional space. Non-negative tensor factorization, on the other hand, is a powerful unsupervised machine learning method that naturally models multi-dimensional data, capturing complex and multi-faceted details of behavior profiles. Herein, our new unsupervised statistical anomaly detection methodology matches or surpasses state-of-the-art supervised learning baselines across several challenging and diverse cyber application areas, including detection of compromised user credentials, botnets, spam e-mails, and fraudulent credit card transactions.

97 MATHEMATICS AND COMPUTING↗

High-fidelity retrieval from instantaneous line-of-sight returns of nacelle-mounted lidar including supervised machine learning

Abstract. Wind turbine applications that leverage nacelle-mounted Doppler lidar are hampered by several sources of uncertainty in the lidar measurement, affecting both bias and random errors. Two problems encountered especially for nacelle-mounted lidar are solid interference due to intersection of the line of sight with solid objects behind, within, or in front of the measurement volume and spectral noise due primarily to limited photon capture. These two uncertainties, especially that due to solid interference, can be reduced with high-fidelity retrieval techniques (i.e., including both quality assurance/quality control and subsequent parameter estimation). Our work compares three such techniques, including conventional thresholding, advanced filtering, and a novel application of supervised machine learning with ensemble neural networks, based on their ability to reduce uncertainty introduced by the two observed nonideal spectral features while keeping data availability high. The approach leverages data from a field experiment involving a continuous-wave (CW) SpinnerLidar from the Technical University of Denmark (DTU) that provided scans of a wide range of flows both unwaked and waked by a field turbine. Independent measurements from an adjacent meteorological tower within the sampling volume permit experimental validation of the instantaneous velocity uncertainty remaining after retrieval that stems from solid interference and strong spectral noise, which is a validation that has not been performed previously. All three methods perform similarly for non-interfered returns, but the advanced filtering and machine learning techniques perform better when solid interference is present, which allows them to produce overall standard deviations of error between 0.2 and 0.3 m s−1, or a 1 %–22 % improvement versus the conventional thresholding technique, over the rotor height for the unwaked cases. Between the two improved techniques, the advanced filtering produces 3.5 % higher overall data availability, while the machine learning offers a faster runtime (i.e., ∼ 1 s to evaluate) that is therefore more commensurate with the requirements of real-time turbine control. The retrieval techniques are described in terms of application to CW lidar, though they are also relevant to pulsed lidar. Previous work by the authors (Brown and Herges, 2020) explored a novel attempt to quantify uncertainty in the output of a high-fidelity lidar retrieval technique using simulated lidar returns; this article provides true uncertainty quantification versus independent measurement and does so for three techniques rather than one.

47 OTHER INSTRUMENTATION↗

Quantifying Drivers of Methane Hydrobiogeochemistry in a Tidal River Floodplain System

The influence of coastal ecosystems on global greenhouse gas (GHG) budgets and their response to increasing inundation and salinization remains poorly constrained. In this study, we have integrated an uncertainty quantification (UQ) and ensemble machine learning (ML) framework to identify and rank the most influential processes, properties, and conditions controlling methane behavior in a freshwater floodplain responding to recently restored seawater inundation. Our unique multivariate, multiyear, and multi-site dataset comprises tidal creek and floodplain porewater observations encompassing water level, salinity, pH, temperature, dissolved oxygen (DO), dissolved organic carbon (DOC), total dissolved nitrogen (TDN), partial pressure of carbon dioxide (pCO 2 ), nitrous oxide (pN 2 O), methane (pCH 4 ), and the stable isotopic composition of methane (δ 13 CH 4 ). Additionally, we incorporated topographical data, soil porosity, hydraulic conductivity, and water retention parameters for UQ analysis using a previously developed 3D variably saturated flow and transport floodplain model for a physical mechanistic understanding of factors influencing groundwater levels and salinity and, therefore, CH 4 . Principal component analysis revealed that groundwater level and salinity are the most significant predictors of overall biogeochemical variability. The ensemble ML models and UQ analyses identified DO, water level, salinity, and temperature as the most influential factors for porewater methane levels and indicated that approximately 80% of the total variability in hourly water levels and around 60% of the total variability in hourly salinity can be explained by permeability, creek water level, and two van Genuchten water retention function parameters: the air-entry suction parameter α and the pore size distribution parameter m. These findings provide insights on the physicochemical factors in methane behavior in coastal ecosystems and their representation in local- to global-scale Earth system models.

54 ENVIRONMENTAL SCIENCES↗

Peatland fires in Alaska will double by the end of the century

During recent summers, warm and dry conditions have increased the occurrence of wildfires and potentially peat-fires across Alaska. Limitations in resolving the fine-scale distribution of peatlands and climate observations have constrained our ability to accurately predict peat-fire dynamics. Using a new high-resolution peatland map of Alaska, we evaluated the climate and environmental controls of past and future peat-fire activity. Ensemble machine learning models identified reduced soil moisture, higher temperatures, and evapotranspiration as key predictors of annual total burned peatland area (tenfold CV R 2 = 0.62, RMSE = 221.1 km 2 ). By the end of the twenty-first century, models forced with climate datasets from representative concentration pathways (RCPs) 4.5, 6.0, and 8.5 emission scenarios project a statewide doubling of burned peatlands (increasing 61–121%), with regional increases ranging from 25–165% in polar, 61–95% in boreal, and 102–106% in maritime ecoregions. These projections indicate that wildfires will progressively encroach further into organic-rich moist and wet peaty soils, potentially amplifying soil carbon release across Alaska.

climate-change ecology↗

Identifications of RR Lyrae Stars and Quasars from the Simulated Data of Mephisto-W Survey

We have investigated the feasibilities and accuracies of the identifications of RR Lyrae stars and quasars from the simulated data of the Multi-channel Photometric Survey Telescope (Mephisto) W Survey. Based on the variable sources light curve libraries from the Sloan Digital Sky Survey (SDSS) Stripe 82 data and the observation history simulation from the Mephisto-W Survey Scheduler, we have simulated the uvgriz multi-band light curves of RR Lyrae stars, quasars and other variable sources for the first-year observation of Mephisto W Survey. We have applied the ensemble machine learning algorithm Random Forest Classifier (RFC) to identify RR Lyrae stars and quasars, respectively. We build training and test samples and extract ~150 features from the simulated light curves and train two RFCs respectively for the RR Lyrae star and quasar classification. We find that, our RFCs are able to select the RR Lyrae stars and quasars with remarkably high precision and completeness, with purity = 95.4% and completeness = 96.9% for the RR Lyrae RFC and purity = 91.4% and completeness = 90.2% for the quasar RFC. In conclusion, we have also derived relative importances of the extracted features utilized to classify RR Lyrae stars and quasars.

(galaxies:) quasars: general↗

Navigating Uncertainty: Challenges in Visualizing Ensemble Data and Surrogate Models for Decision Systems

Uncertainty visualization plays a critical role in transforming ensemble simulation data into actionable insights by effectively communicating various dimensions of uncertainty within a system. The emergence of artificial intelligence-driven surrogate models trained on multirun ensemble data offers a transformative opportunity to replace computationally intensive simulations with fast estimates, enabling users to explore data spaces with unprecedented depth and interactivity. However, integrating ensemble data and surrogate models into decision-making workflows and tools introduces novel challenges for uncertainty visualization. These include reconciling and clearly communicating the unique uncertainties associated with ensembles and their surrogate model estimates, and leveraging these approximations to inform actionable decisions. This work explores these challenges in the context of high-dimensional data visualization, bridging discrete datasets with their continuous representations and addressing the complexities of systems that support iterative navigation between input and output spaces. We evaluate the role of uncertainty visualization in fostering intuitive, actionable interactions and identify critical hurdles in advancing this frontier of computational simulation.

97 MATHEMATICS AND COMPUTING↗

Enhancing Solar Power Forecasting with Regularized Constrained Quantile Regression Averaging and Bootstrapping Techniques

Probabilistic solar power forecasting (SPF) plays an essential role in optimizing power-grid operations by quantifying the forecast uncertainty. To improve the accuracy and robustness of probabilistic SPF, this paper introduces the regularized constrained quantile regression averaging (rCQRA) method to combine outputs from multiple PSPF models. In addition, a bootstrapping method was used to quantify model uncertainty, providing insights into the reliability and significance of each ensemble component. To evaluate its efficacy, the proposed rCQRA method is used to integrate four PSPF methods. The resulting SPF models are trained and validated using a real-world six-year dataset from a rooftop solar plant in the USA. The performance of the proposed rCQRA method is evaluated and compared with two benchmark methods under three categories of weather conditions. It is shown that the rCQRA method has superior performance in its forecast reliability, sharpness, and accuracy.

Ensemble learning, probabilistic solar power forec↗

Estimating Subhourly Inverter Clipping Loss From Satellite-Derived Irradiance Data

Photovoltaic system production simulations are conventionally run using hourly weather datasets. Hourly simulations are sufficiently accurate to predict the majority of long-term system behavior but cannot resolve high-frequency effects like inverter clipping caused by short-duration irradiance variability. Direct modeling of this subhourly clipping error is only possible for the few locations with high-resolution irradiance datasets. This paper describes a method of predicting the magnitude of this error using a machine learning regressor ensemble model, comprised of a random forest and an XGBoost model, and 30-minute satellite irradiance data. The method predicts a correction for each 30-minute interval with the potential to roll up into 60-minute corrections to match an hourly energy model. The model is trained and validated at locations where the error can be directly simulated from 1-minute ground data. The validation shows low bias at most ground station locations. The model is also applied to gridded satellite irradiance to produce a heatmap of the estimated clipping error across the United States. Finally, the relative importance of each predictor satellite variable is retrieved from the model and discussed.

41 EE - Solar Energy Technologies Office (EE-4S)↗

Continental United States may lose 1.8 petagrams of soil organic carbon under climate change by 2100

Abstract Aims High‐resolution information on soils’ vulnerability to climate‐induced soil organic carbon (SOC) loss can enable environmental scientists, land managers, and policy makers to develop targeted mitigation strategies. This study aims to estimate baseline and decadal changes in continental US surface SOC stocks under future emission scenarios. Location Continental United States. Time period 2014–2100. Methods We used recent SOC field observations ( n = 6,213 sites), environmental factors ( n = 32), and an ensemble machine learning (ML) approach to estimate baseline SOC stocks in surface soils across the continental United States at 100‐m spatial resolution, and decadal changes under the projected climate scenarios of Coupled Model Intercomparison Project Phase Six (CMIP6) earth system models (ESMs). Results Baseline SOC projections from ML approaches captured more than 50% of variability in SOC observations, whereas ESMs represented only 6–16% of observed SOC variability. ML estimates showed a mean total loss of 1.8 Pg C from US surface soils under the high‐emission scenario by 2100, whereas ESMs showed no significant change in SOC stocks with wide variation among ESMs. Both ML and ESM predictions agree on the direction of SOC change (net emissions or sequestration) across 46–51% of continental US land area. These differences are attributable to the high‐resolution site‐specific data used in the ML models compared to the relatively coarse grid represented in CMIP6 ESMs. Main conclusions Our high‐resolution estimates of baseline SOC stocks, identification of key environmental controllers, and projection of SOC changes from US land cover types under future climate scenarios suggest the need for high‐resolution simulations of SOC in ESMs to represent the heterogeneity of SOC. We found that the SOC change is sensitive to key soil related factors (e.g. soil drainage and soil order) that have not been historically considered as input parameters in ESMs, because currently more than 95% variability in the SOC of CMIP6 ESMs is controlled by net primary productivity, temperature, and precipitation. Using additional environmental factors to estimate the baseline SOC stocks and predict the future trajectory of SOC change can provide more accurate results.

54 ENVIRONMENTAL SCIENCES↗

Identifying precursors of daily to seasonal hydrological extremes over the USA using deep learning techniques and climate model ensembles

Focal Area(s): We focus on two areas of crosscutting interest for DOE: 1) predictability of extreme precipitation and drought in the USA and 2) the integration of climate models with new AI tools, such as convolutional neural networks (CNN) and methods to understand their output (e.g. layer-wise relevance propagation; LRP). This project fits into focus area 3 of this call for white paper using AI to gain insight from complex data, including explainable AI tools. Science Challenge: Predicting hydrological extremes is important due to their impacts on people, agriculture and infrastructure. This prediction is difficult due to the infrequent occurrence of extremes and their complexity. However, extreme events can be related to more predictable conditions in the ocean, such as El Nino, long-term soil moisture or large scale modes of climate variability, such as the North Atlantic Oscillation (NAO).

54 ENVIRONMENTAL SCIENCES↗

Constraining microphysical processes of warm rain formulation using advanced spectral separations, an ensemble retrieval framework and machine learning techniques

Drizzle, a common feature of marine boundary layer clouds formed through collision coalescence, plays a key role in cloud microphysics and evolution. Yet, simultaneously retrieving cloud and drizzle properties from remote-sensing observations remains challenging because drizzle droplets often dominate radar signals, masking cloud contributions. The goal of the proposed research is to provide constraints for the process of autoconversion and accretion using ARM cloud measurements. Specifically, we provide concurrent retrievals of cloud and drizzle that allows users to derive corresponding autoconversion and accretion rates.

54 ENVIRONMENTAL SCIENCES↗

Robust Explanations using Diverse Adversarially Trained Ensembles, Multi-Modal Contrastive Learning, and Attribution-based Confidence Metrics

The primary objective of this project is to strengthen the trustworthiness of AI systems by designing algorithms that make their internal decision-making processes more understandable to human users. This involves creating clear, interpretable explanations for AI decisions and developing metrics to assess these explanations' validity and reliability. Significant progress has been achieved through (i) developing symbolic explanations, (ii) generating meaningful interpretive insights, (iii) establishing accuracy and confidence metrics, and (iv) devising methods to evaluate the knowledge boundaries of AI models. To date, the research findings have been shared in peer-reviewed publications, with accompanying scientific and technical information (STI) detailed below.

97 MATHEMATICS AND COMPUTING↗