Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “ensemble learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Nature of molybdenum carbide surfaces for catalytic hydrogen dissociation using machine-learned potentials: an ensemble-averaged perspective

Molybdenum carbides with an electronic structure similar to noble metals have gained attention as a promising low-cost catalyst for biomass valorization and the hydrogen evolution reaction. However, our fundamental understanding of the catalyst surface and how different phases of these catalysts behave at varying reaction conditions is limited to ground state density functional theory calculations as ab initio molecular dynamics (AIMD) is computationally prohibitive at relevant length and time scales. Here, in this work, we train a multi-atomic cluster expansion (MACE) machine-learned interatomic potentials (MLIP) to study hydrogen dissociation and dynamics over Mo, δ-MoC, α-Mo 2 C, and β-Mo 2 C surfaces at varying temperatures and hydrogen partial pressures. Our simulations identify unique and different molecular and atomic hydrogen adsorption sites on different surfaces that do not depend on the temperature. At low hydrogen pressures, the surface coverage is monolayer, which transitions to two-layer adsorption at higher pressures. We find that atomic hydrogen diffusion and recombinations are preferred over molybdenum atom hollow sites, while the diffusion over carbon-terminated facets was negligible, signifying particularly strong C–H interactions. In contrast, molecular hydrogen adsorption occurs mostly atop Mo or the bridging sites. At a comparable hydrogen loading, β-Mo 2 C (001) is the most active surface for hydrogen dissociation reaction. This work provides insights into the dynamic nature of the hydrogen dissociation chemistry and the diversity of hydrogen adsorption sites on molybdenum carbides.

08 HYDROGEN↗

Machine Learning Approach for Spatiotemporal Multivariate Optimization of Environmental Monitoring Sensor Locations

Abstract Long-term environmental monitoring is critical for managing the soil and groundwater at contaminated sites. Recent improvements in state-of-the-art sensor technology, communication networks, and artificial intelligence have created opportunities to modernize this monitoring activity for automated, fast, robust, and predictive monitoring. In such modernization, it is required that sensor locations be optimized to capture the spatiotemporal dynamics of all monitoring variables as well as to make it cost-effective. The legacy monitoring datasets of the target area are important to perform this optimization. In this study, we have developed a machine-learning approach to optimize sensor locations for soil and groundwater monitoring based on ensemble supervised learning and majority voting. For spatial optimization, Gaussian process regression (GPR) is used for spatial interpolation, while the majority voting is applied to accommodate the multivariate temporal dimension. Results show that the algorithms significantly outperform the random selection of the sensor locations for predictive spatiotemporal interpolation. While the method has been applied to a four-dimensional dataset (with two-dimensional space, time, and multiple contaminants), we anticipate that it can be generalizable to higher-dimensional datasets for environmental monitoring sensor location optimization.

Siddiquee, Masudur R.↗

Using Ultrasound Image Augmentation and Ensemble Predictions to Prevent Machine-Learning Model Overfitting

Deep learning predictive models have the potential to simplify and automate medical imaging diagnostics by lowering the skill threshold for image interpretation. However, this requires predictive models that are generalized to handle subject variability as seen clinically. Here, we highlight methods to improve test accuracy of an image classifier model for shrapnel identification using tissue phantom image sets. Using a previously developed image classifier neural network—termed ShrapML—blind test accuracy was less than 70% and was variable depending on the training/test data setup, as determined by a leave one subject out (LOSO) holdout methodology. Introduction of affine transformations for image augmentation or MixUp methodologies to generate additional training sets improved model performance and overall accuracy improved to 75%. Further improvements were made by aggregating predictions across five LOSO holdouts. This was done by bagging confidences or predictions from all LOSOs or the top-3 LOSO confidence models for each image prediction. Top-3 LOSO confidence bagging performed best, with test accuracy improved to greater than 85% accuracy for two different blind tissue phantoms. This was confirmed by gradient-weighted class activation mapping to highlight that the image classifier was tracking shrapnel in the image sets. Overall, data augmentation and ensemble prediction approaches were suitable for creating more generalized predictive models for ultrasound image analysis, a critical step for real-time diagnostic deployment.

60 APPLIED LIFE SCIENCES↗

Multimodal sensor fusion framework for residential building occupancy detection

For several years now, smart building energy systems have been a research area of intensive activity. In light of the increasing need for sustainable buildings and energy systems, this trend motivates an increasing need for a solution to reduce carbon dioxide emissions and improve energy efficiency. This work proposes a high-performing and transferable occupancy detection framework that combines sensor data from different data modalities, including time series environmental data (temperature, humidity, and illuminance), image data, and acoustic energy data using ensemble method. To draw out the best prediction performance in each modality, the proposed framework was developed, including various models that were designed to learn the occupancy patterns reflected in the physical data streams. To tackle the time series environmental data, we designed two variants of an occupancy detection spatiotemporal pattern network (Occ-STPN) that performs both feature level and decision level fusion, respectively. We also propose a new metric; the fading memory mean square error (FMMSE), that provides a fair evaluation and penalization of delayed occupancy predictions. Multiple open-sourced datasets, including the Electricity Consumption and Occupancy and the University of California, Irvine's (UCI) building occupancy detection dataset, along with our own real data collected from six different houses, were used to validate the algorithms' performance. The experimental results presented herein break down the performance for each sensing modality, and a detailed analysis of the performance is also discussed.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Applications of flow models to the generation of correlated lattice QCD ensembles

Machine-learned normalizing flows can be used in the context of lattice quantum field theory to generate statistically correlated ensembles of lattice gauge fields at different action parameters. This work demonstrates how these correlations can be exploited for variance reduction in the computation of observables. Three different proof-of-concept applications are demonstrated using a novel residual flow architecture: continuum limits of gauge theories, the mass dependence of QCD observables, and hadronic matrix elements based on the Feynman–Hellmann approach. In all three cases, it is shown that statistical uncertainties are significantly reduced when machine-learned flows are incorporated as compared with the same calculations performed with uncorrelated ensembles or direct reweighting. Published by the American Physical Society 2024

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Atomistic learning in the electronically grand-canonical ensemble

Abstract A strategy is presented for the machine-learning emulation of electronic structure calculations carried out in the electronically grand-canonical ensemble. The approach relies upon a dual-learning scheme, where both the system charge and the system energy are predicted for each image. The scheme is shown to be capable of emulating basic electrochemical reactions at a range of potentials, and coupling it with a bootstrap-ensemble approach gives reasonable estimates of the prediction uncertainty. The method is also demonstrated to accelerate saddle-point searches, and to extrapolate to systems with one to five water layers. We anticipate that this method will allow for larger length- and time-scale simulations necessary for electrochemical simulations.

36 MATERIALS SCIENCE↗

Statistical upscaling of ecosystem CO 2 fluxes across the terrestrial tundra and boreal domain: Regional patterns and uncertainties

Abstract The regional variability in tundra and boreal carbon dioxide (CO 2 ) fluxes can be high, complicating efforts to quantify sink‐source patterns across the entire region. Statistical models are increasingly used to predict (i.e., upscale) CO 2 fluxes across large spatial domains, but the reliability of different modeling techniques, each with different specifications and assumptions, has not been assessed in detail. Here, we compile eddy covariance and chamber measurements of annual and growing season CO 2 fluxes of gross primary productivity (GPP), ecosystem respiration (ER), and net ecosystem exchange (NEE) during 1990–2015 from 148 terrestrial high‐latitude (i.e., tundra and boreal) sites to analyze the spatial patterns and drivers of CO 2 fluxes and test the accuracy and uncertainty of different statistical models. CO 2 fluxes were upscaled at relatively high spatial resolution (1 km 2 ) across the high‐latitude region using five commonly used statistical models and their ensemble, that is, the median of all five models, using climatic, vegetation, and soil predictors. We found the performance of machine learning and ensemble predictions to outperform traditional regression methods. We also found the predictive performance of NEE‐focused models to be low, relative to models predicting GPP and ER. Our data compilation and ensemble predictions showed that CO 2 sink strength was larger in the boreal biome (observed and predicted average annual NEE −46 and −29 g C m −2 yr −1 , respectively) compared to tundra (average annual NEE +10 and −2 g C m −2 yr −1 ). This pattern was associated with large spatial variability, reflecting local heterogeneity in soil organic carbon stocks, climate, and vegetation productivity. The terrestrial ecosystem CO 2 budget, estimated using the annual NEE ensemble prediction, suggests the high‐latitude region was on average an annual CO 2 sink during 1990–2015, although uncertainty remains high.

Virkkala, Anna‐Maria↗

AL4GAP: Active learning workflow for generating DFT-SCAN accurate machine-learning potentials for combinatorial molten salt mixtures

Machine learning interatomic potentials have emerged as a powerful tool for bypassing the spatiotemporal limitations of ab initio simulations, but major challenges remain in their efficient parameterization. We present AL4GAP, an ensemble active learning software workflow for generating multicomposition Gaussian approximation potentials (GAP) for arbitrary molten salt mixtures. The workflow capabilities include: (1) setting up user-defined combinatorial chemical spaces of charge neutral mixtures of arbitrary molten mixtures spanning 11 cations (Li, Na, K, Rb, Cs, Mg, Ca, Sr, Ba and two heavy species, Nd, and Th) and 4 anions (F, Cl, Br, and I), (2) configurational sampling using low-cost empirical parameterizations, (3) active learning for down-selecting configurational samples for single point density functional theory calculations at the level of Strongly Constrained and Appropriately Normed (SCAN) exchange-correlation functional, and (4) Bayesian optimization for hyperparameter tuning of two-body and many-body GAP models. Here, we apply the AL4GAP workflow to showcase high throughput generation of five independent GAP models for multicomposition binary-mixture melts, each of increasing complexity with respect to charge valency and electronic structure, namely: LiCl–KCl, NaCl–CaCl 2 , KCl–NdCl 3 , CaCl 2 –NdCl 3 , and KCl–ThCl 4 . Our results indicate that GAP models can accurately predict structure for diverse molten salt mixture with density functional theory (DFT)-SCAN accuracy, capturing the intermediate range ordering characteristic of the multivalent cationic melts.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Machine Learning of Key Variables Impacting Extreme Precipitation in Various Regions of the Contiguous United States

Abstract Amplification in extreme precipitation intensity and frequency can cause severe flooding and impose significant social and economic consequences. Variations in extreme precipitation intensity, frequencies, and return periods can be attributed to many physical variables across spatial and temporal scales. Here we employ ensemble machine learning (ML) methods, namely random forest (RF), eXtreme Gradient Boosting (XGB), and artificial neural networks (ANN), to explore key contributing variables to monthly extreme precipitation intensity and frequency in six regions over the United States. We further establish emulators for return periods. Results show that the ML models for intensity perform better in regions with obvious seasonality (i.e., Northern Great Plains, Southern Great Plains, and West Coast) than the other three regions (Northeast, Southwest, and Rocky Mountains), while for frequency the models perform well for most regions. The Shapley additive explanation is used to help explain the relationships between extreme precipitation characteristics and identify top variables for RF and XGB. We find that latent heat flux, relative humidity, soil moisture, and large‐scale subsidence are key common variables across the regions for both monthly intensity and frequency, and their compound effects are non‐negligible. The developed ML models capture the probability and return period of extreme precipitation well for all regions and may be used for decision making (e.g., infrastructure planning and design).

54 ENVIRONMENTAL SCIENCES↗

DeepONet-Assisted Optimization of Surface Topography for Transition Delay in a Mach 4.5 Boundary Layer

We use deep learning, an ensemble variational technique (EnVar), and direct numerical simulations(DNS) to design an optimal topography for a two-dimensional roughness element that delays the on-set of laminar-turbulent transition in a Mach 4.5 flat-plate boundary layer. Deep operator networks (DeepONets), which have the known ability to learn complex nonlinear operators within dynamical systems, are used for machine learning. For the baseline configuration of a smooth flat plate, the second-mode waves at the DNS inflow cause a quick nonlinear breakdown of the high-speed boundary layer within the computational domain. Results reported in the present study validate the ability of DeepONets to model the transition delay via a given topography of the roughness element. The computing cost to optimize the rough-ness element for minimal skin-friction drag is substantially lowered by the DeepONets-based reduced-order model. In comparison to the baseline method of EnVar optimization based on DNS alone, the DeepONets-based EnVar optimizer is able to delay transition past the outflow boundary of the computational domain while utilizing almost 5–6 times fewer DNS.

Machine Learning↗

DeepONet-Assisted Optimization of Surface Topography for Transition Delay in A Mach 4.5 Boundary Layer

We use deep learning, an ensemble variationaltechnique (EnVar), and direct numerical simulations(DNS) to design an optimal topography for a two-dimensional roughness element that delays the on-set of laminar-turbulent transition in a Mach 4.5 flat-plate boundary layer. Deep operator networks (Deep-ONets), which have the known ability to learn com-plex nonlinear operators within dynamical systems,are used for machine learning. For the baseline config-uration of a smooth flat plate, the second-mode wavesat the DNS inflow cause a quick nonlinear breakdownof the high-speed boundary layer within the computa-tional domain. Results reported in the present studyvalidate the ability of DeepONets to model the tran-sition delay via a given topography of the roughnesselement. The computing cost to optimize the rough-ness element for minimal skin-friction drag is substan-tially lowered by the DeepONets-based reduced-ordermodel. In comparison to the baseline method of EnVaroptimization based on DNS alone, the DeepONets-based EnVar optimizer is able to delay transition pastthe outflow boundary of the computational domainwhile utilizing almost 5–6 times fewer DNS.

Machine Learning↗

A meta-learning based distribution system load forecasting model selection framework

This paper presents a meta-learning based, automatic distribution system load forecasting model selection framework. Furthermore, the framework includes the following processes: feature extraction, candidate model preparation and labeling, offline training, and online model recommendation. Using load forecasting needs and data characteristics as input features, multiple metalearners are used to rank the candidate load forecast models based on their forecasting accuracy. Then, a scoring-voting mechanism is proposed to weights recommendations from each meta-leaner and make the final recommendations. Heterogeneous load forecasting tasks with different temporal and technical requirements at different load aggregation levels are set up to train, validate, and test the performance of the proposed framework. Simulation results demonstrate that the performance of the meta-learning based approach is satisfactory in both seen and unseen forecasting tasks.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Earth's record-high greenness and its attributions in 2020

Terrestrial vegetation is a crucial component of Earth's biosphere, regulating global carbon and water cycles and contributing to human welfare. Despite an overall greening trend, terrestrial vegetation exhibits a significant inter-annual variability. The mechanisms driving this variability, particularly those related to climatic and anthropogenic factors, remain poorly understood, which hampers our ability to project the long-term sustainability of ecosystem services. Here, in this work, by leveraging diverse remote sensing measurements, we pinpointed 2020 as a historic landmark, registering as the greenest year in modern satellite records from 2001 to 2020. Using ensemble machine learning and Earth system models, we found this exceptional greening primarily stemmed from consistent growth in boreal and temperate vegetation, attributed to rising CO 2 levels, climate warming, and reforestation efforts, alongside a transient tropical green-up linked to the enhanced rainfall. Contrary to expectations, the COVID-19 pandemic lockdowns had a limited impact on this global greening anomaly. Our findings highlight the resilience and dynamic nature of global vegetation in response to diverse climatic and anthropogenic influences, offering valuable insights for optimizing ecosystem management and informing climate mitigation strategies.

54 ENVIRONMENTAL SCIENCES↗

General-Purpose Unsupervised Cyber Anomaly Detection via Non-Negative Tensor Factorization

Distinguishing malicious anomalous activities from unusual but benign activities is a fundamental challenge for cyber defenders. Prior studies have shown that statistical user behavior analysis yields accurate detections by learning behavior profiles from observed user activity. These unsupervised models are able to generalize to unseen types of attacks by detecting deviations from normal behavior, without knowledge of specific attack signatures. However, approaches proposed to date based on probabilistic matrix factorization are limited by the information conveyed in a two-dimensional space. Non-negative tensor factorization, on the other hand, is a powerful unsupervised machine learning method that naturally models multi-dimensional data, capturing complex and multi-faceted details of behavior profiles. Herein, our new unsupervised statistical anomaly detection methodology matches or surpasses state-of-the-art supervised learning baselines across several challenging and diverse cyber application areas, including detection of compromised user credentials, botnets, spam e-mails, and fraudulent credit card transactions.

97 MATHEMATICS AND COMPUTING↗

High-fidelity retrieval from instantaneous line-of-sight returns of nacelle-mounted lidar including supervised machine learning

Abstract. Wind turbine applications that leverage nacelle-mounted Doppler lidar are hampered by several sources of uncertainty in the lidar measurement, affecting both bias and random errors. Two problems encountered especially for nacelle-mounted lidar are solid interference due to intersection of the line of sight with solid objects behind, within, or in front of the measurement volume and spectral noise due primarily to limited photon capture. These two uncertainties, especially that due to solid interference, can be reduced with high-fidelity retrieval techniques (i.e., including both quality assurance/quality control and subsequent parameter estimation). Our work compares three such techniques, including conventional thresholding, advanced filtering, and a novel application of supervised machine learning with ensemble neural networks, based on their ability to reduce uncertainty introduced by the two observed nonideal spectral features while keeping data availability high. The approach leverages data from a field experiment involving a continuous-wave (CW) SpinnerLidar from the Technical University of Denmark (DTU) that provided scans of a wide range of flows both unwaked and waked by a field turbine. Independent measurements from an adjacent meteorological tower within the sampling volume permit experimental validation of the instantaneous velocity uncertainty remaining after retrieval that stems from solid interference and strong spectral noise, which is a validation that has not been performed previously. All three methods perform similarly for non-interfered returns, but the advanced filtering and machine learning techniques perform better when solid interference is present, which allows them to produce overall standard deviations of error between 0.2 and 0.3 m s−1, or a 1 %–22 % improvement versus the conventional thresholding technique, over the rotor height for the unwaked cases. Between the two improved techniques, the advanced filtering produces 3.5 % higher overall data availability, while the machine learning offers a faster runtime (i.e., ∼ 1 s to evaluate) that is therefore more commensurate with the requirements of real-time turbine control. The retrieval techniques are described in terms of application to CW lidar, though they are also relevant to pulsed lidar. Previous work by the authors (Brown and Herges, 2020) explored a novel attempt to quantify uncertainty in the output of a high-fidelity lidar retrieval technique using simulated lidar returns; this article provides true uncertainty quantification versus independent measurement and does so for three techniques rather than one.

47 OTHER INSTRUMENTATION↗

Quantifying Drivers of Methane Hydrobiogeochemistry in a Tidal River Floodplain System

The influence of coastal ecosystems on global greenhouse gas (GHG) budgets and their response to increasing inundation and salinization remains poorly constrained. In this study, we have integrated an uncertainty quantification (UQ) and ensemble machine learning (ML) framework to identify and rank the most influential processes, properties, and conditions controlling methane behavior in a freshwater floodplain responding to recently restored seawater inundation. Our unique multivariate, multiyear, and multi-site dataset comprises tidal creek and floodplain porewater observations encompassing water level, salinity, pH, temperature, dissolved oxygen (DO), dissolved organic carbon (DOC), total dissolved nitrogen (TDN), partial pressure of carbon dioxide (pCO 2 ), nitrous oxide (pN 2 O), methane (pCH 4 ), and the stable isotopic composition of methane (δ 13 CH 4 ). Additionally, we incorporated topographical data, soil porosity, hydraulic conductivity, and water retention parameters for UQ analysis using a previously developed 3D variably saturated flow and transport floodplain model for a physical mechanistic understanding of factors influencing groundwater levels and salinity and, therefore, CH 4 . Principal component analysis revealed that groundwater level and salinity are the most significant predictors of overall biogeochemical variability. The ensemble ML models and UQ analyses identified DO, water level, salinity, and temperature as the most influential factors for porewater methane levels and indicated that approximately 80% of the total variability in hourly water levels and around 60% of the total variability in hourly salinity can be explained by permeability, creek water level, and two van Genuchten water retention function parameters: the air-entry suction parameter α and the pore size distribution parameter m. These findings provide insights on the physicochemical factors in methane behavior in coastal ecosystems and their representation in local- to global-scale Earth system models.

54 ENVIRONMENTAL SCIENCES↗

Peatland fires in Alaska will double by the end of the century

During recent summers, warm and dry conditions have increased the occurrence of wildfires and potentially peat-fires across Alaska. Limitations in resolving the fine-scale distribution of peatlands and climate observations have constrained our ability to accurately predict peat-fire dynamics. Using a new high-resolution peatland map of Alaska, we evaluated the climate and environmental controls of past and future peat-fire activity. Ensemble machine learning models identified reduced soil moisture, higher temperatures, and evapotranspiration as key predictors of annual total burned peatland area (tenfold CV R 2 = 0.62, RMSE = 221.1 km 2 ). By the end of the twenty-first century, models forced with climate datasets from representative concentration pathways (RCPs) 4.5, 6.0, and 8.5 emission scenarios project a statewide doubling of burned peatlands (increasing 61–121%), with regional increases ranging from 25–165% in polar, 61–95% in boreal, and 102–106% in maritime ecoregions. These projections indicate that wildfires will progressively encroach further into organic-rich moist and wet peaty soils, potentially amplifying soil carbon release across Alaska.

climate-change ecology↗