Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “extreme gradient boost”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

86 records · Page 5

PV Generation and Load Forecasting for Adjuntas PR Community Microgrids

Existing frameworks to forecast time-series photovoltaic (PV) output power and consumer load for microgrid operations and controls assume a near-continuous availability of real-time input features from the field assets such as PV inverters, energy meters, and weather station. These incoming data points are used to periodically retrain models and update forecast snapshots over a moving horizon window, be it one hour-ahead, one-day ahead, or one-week ahead. However, such frameworks are not resilient to disruptions in data availability caused by losses in communications between the field sensors and data loggers. Hence, there is a need for programs that assume no availability of real-time microgrid asset data and still make reliable forecasts that can be used for decision-making. Such programs would be apt to function in extreme weather events such as hurricanes and would use lightweight recursive time-series models to independently forecast solar irradiance and ambient temperature, then compute PV power from those forecasts, as well as independently forecast consumer load. The codebase performs forecasting for the scenario of when the microgrid does not have a reliable access to forecasts or real-time observations of solar irradiance (I) and ambient temperature (AT) and load (Load) to be able to adequately forecast, in real-time, the PV power production or a business' load. In this case, using historical values of PV power and load, a univariate forecasting of generation and consumption are respectively made. The use-case in particular has two sub-scenarios: one, a normal 7-day ahead forecast where the unavailability of real-time data is assumed due to infrastructure issues such as loss of communication or sensor maintenance or service downtimes. Whereas a hurricane-caused unavailability of real-time data requires a second model trained specifically on historical hurricane days to be able to capture the extreme day behavior of generation in particular, and load if applicable. A gradient boosted regression tree comprises an ensemble of additive models that map between the input of historical values (be it irradiance, temperature, or load) and their corresponding output forecasts of a given horizon such that the individual learner predictions are summed up over the total number of such learners in the ensemble to produce an aggregate forecast. A weighting mechanism is applied to the training data in each iteration, where actual and forecast values are compared to penalize incorrect forecasts by increasing the weight and reducing it to reward correct forecasts. The code's benefits are that it: (a) accounts for a contingency where communication loss renders newly measured real-time data unavailable for model tuning and snapshot updates; (b) presents blind forecasting that recursively determines the next time-step value in a horizon using the forecast of the same attribute from a prior step; and (c) employs lightweight models that, once trained, can reliably generalize for different horizons, which make them suitable for enhancing the resilience of field microgrids prone to extreme events that encounter disruptions to data availability.

Sundararajan, Aditya [Oak Ridge National Laborator↗

Evaluating Recursive Blind Forecast Against API and Baseline: A Puerto Rican Case Study on Solar Irradiance for Normal and Extreme Weather

This paper leverages ongoing work in a community microgrid in Adjuntas, Puerto Rico to forecast global horizontal irradiance (GHI) and compare performance in normal and extreme weather. Given a positive correlation of 0.98 between GHI and PV power, forecasting GHI can be an effective, indirect forecast of photovoltaic (PV) power, especially in microgrids where the end-users, owners, operators, or other stakeholders are reluctant to share data for training or validation due to privacy and security concerns. A recursive one-shot (termed as "blind") forecast is, hence, formulated, wherein a gradient-boosted regression tree (GBR) is built to forecast GHI for a 7-day horizon in normal weather, and a 2-day horizon in extreme weather. To demonstrate its resilience, the architecture is trained on normal and hurricane weather GHI from 2002-2022. It is generalized on February 9-16, 2023, and on the landfall of Hurricane Nicole (Nov 4-5, 2022), respectively. Forecasts from GBR are compared against that from a satellite-based API resource and three baselines: persistence, averaging, and exponential smoothing. Results show GBR and persistence outperform sophisticated API in both types of weather for this case study.

Sundararajan, Aditya↗

Recent increases in annual, seasonal, and extreme methane fluxes driven by changes in climate and vegetation in boreal and temperate wetland ecosystems

Climate warming is expected to increase global methane (CH 4 ) emissions from wetland ecosystems. Although in situ eddy covariance (EC) measurements at ecosystem scales can potentially detect CH 4 flux changes, most EC systems have only a few years of data collected, so temporal trends in CH 4 remain uncertain. Here, we use established drivers to hindcast changes in CH 4 fluxes (FCH 4 ) since the early 1980s. We trained a machine learning (ML) model on CH 4 flux measurements from 22 [methane-producing sites] in wetland, upland, and lake sites of the FLUXNET-CH 4 database with at least two full years of measurements across temperate and boreal biomes. The gradient boosting decision tree ML model then hindcasted daily FCH 4 over 1981-2018 using meteorological reanalysis data. We found that, mainly driven by rising temperature, half of the sites (n = 11) showed significant increases in annual, seasonal, and extreme FCH 4 , with increases in FCH 4 of ca. 10% or higher found in the fall from 1981–1989 to 2010–2018. The annual trends were driven by increases during summer and fall, particularly at high-CH 4 -emitting fen sites dominated by aerenchymatous plants. We also found that the distribution of days of extremely high FCH 4 (defined according to the 95th percentile of the daily FCH 4 values over a reference period) have become more frequent during the last four decades and currently account for 10–40% of the total seasonal fluxes. The share of extreme FCH 4 days in the total seasonal fluxes was greatest in winter for boreal/taiga sites and in spring for temperate sites, which highlights the increasing importance of the non-growing seasons in annual budgets. Our results shed light on the effects of climate warming on wetlands, which appears to be extending the CH 4 emission seasons and boosting extreme emissions.

54 ENVIRONMENTAL SCIENCES↗

Stream Temperature Predictions for River Basin Management in the Pacific Northwest and Mid-Atlantic Regions Using Machine Learning

Stream temperature (Ts) is an important water quality parameter that affects ecosystem health and human water use for beneficial purposes. Accurate Ts predictions at different spatial and temporal scales can inform water management decisions that account for the effects of changing climate and extreme events. In particular, widespread predictions of Ts in unmonitored stream reaches can enable decision makers to be responsive to changes caused by unforeseen disturbances. In this study, we demonstrate the use of classical machine learning (ML) models, support vector regression and gradient boosted trees (XGBoost), for monthly Ts predictions in 78 pristine and human-impacted catchments of the Mid-Atlantic and Pacific Northwest hydrologic regions spanning different geologies, climate, and land use. The ML models were trained using long-term monitoring data from 1980–2020 for three scenarios: (1) temporal predictions at a single site, (2) temporal predictions for multiple sites within a region, and (3) spatiotemporal predictions in unmonitored basins (PUB). In the first two scenarios, the ML models predicted Ts with median root mean squared errors (RMSE) of 0.69–0.84 °C and 0.92–1.02 °C across different model types for the temporal predictions at single and multiple sites respectively. For the PUB scenario, we used a bootstrap aggregation approach using models trained with different subsets of data, for which an ensemble XGBoost implementation outperformed all other modeling configurations (median RMSE 0.62 °C).The ML models improved median monthly Ts estimates compared to baseline statistical multi-linear regression models by 15–48% depending on the site and scenario. Air temperature was found to be the primary driver of monthly Ts for all sites, with secondary influence of month of the year (seasonality) and solar radiation, while discharge was a significant predictor at only 10 sites. The predictive performance of the ML models was robust to configuration changes in model setup and inputs, but was influenced by the distance to the nearest dam with RMSE <1 °C at sites situated greater than 16 and 44 km from a dam for the temporal single site and regional scenarios, and over 1.4 km from a dam for the PUB scenario. Our results show that classical ML models with solely meteorological inputs can be used for spatial and temporal predictions of monthly Ts in pristine and managed basins with reasonable (<1 °C) accuracy for most locations.

54 ENVIRONMENTAL SCIENCES↗

Recursive Blind Forecasting of Photovoltaic Generation and Consumer Load for Microgrids

Existing forecasting frameworks that predict time-series photovoltaic (PV) generation and consumer load for micro-grids' operation and control assume near-continuous availability of real-time predictors from the field. The incoming data are used to periodically re-train the models and update forecast snapshots over a moving horizon window. However, such frameworks are not resilient to disruptions in data availability caused by losses in communications between the field sensors and data loggers. This paper bridges the shortcoming by leveraging a previously proposed forecasting framework that is resilient to abrupt changes in data quality caused by communication losses. Assuming no availability of real-time field system data, which is typical in extreme weather events such as hurricanes, the framework uses lightweight recursive time-series models to independently forecast solar irradiance, ambient temperature, PV power, and consumer load for three horizon windows: 24 hours, 12 hours, and 1 hour. Four types of ensemble-based regression trees-simple gradient boosted trees (GBR), GBR with an adaptive component (A-GBR), random forests (RF), and extra trees (ExTR)-are leveraged and their performances are compared against a simple historical weekly mean. Numerical results show that A-GBR performs better on average by 32% for 24-hour horizon and 39% for 12-hour horizon, whereas ExTR outdoes the other models on average by 10% for 1-hour horizon.

Sundararajan, Aditya↗

A Comprehensive Machine Learning Study to Classify Precipitation Type over Land from Global Precipitation Measurement Microwave Imager (GPM-GMI) Measurements

Precipitation type is a key parameter used for better retrieval of precipitation characteristics as well as to understand the cloud–convection–precipitation coupling processes. Ice crystals and water droplets inherently exhibit different characteristics in different precipitation regimes (e.g., convection, stratiform), which reflect on satellite remote sensing measurements that help us distinguish them. The Global Precipitation Measurement (GPM) Core Observatory’s microwave imager (GMI) and dual-frequency precipitation radar (DPR) together provide ample information on global precipitation characteristics. As an active sensor, the DPR provides an accurate precipitation type assignment, while passive sensors such as the GMI are traditionally only used for empirical understanding of precipitation regimes. Using collocated precipitation type flags from the DPR as the “truth”, this paper employs machine learning (ML) models to train and test the predictability and accuracy of using passive GMI-only observations together with ancillary information from a reanalysis and GMI surface emissivity retrieval products. Out of six ML models, four simple ones (support vector machine, neural network, random forest, and gradient boosting) and the 1-D convolutional neural network (CNN) model are identified to produce 90–94% prediction accuracy globally for five types of precipitation (convective, stratiform, mixture, no precipitation, and other precipitation), which is much more robust than previous similar effort. One novelty of this work is to introduce data augmentation (subsampling and bootstrapping) to handle extremely unbalanced samples in each category. A careful evaluation of the impact matrices demonstrates that the polarization difference (PD), brightness temperature (Tc) and surface emissivity at high-frequency channels dominate the decision process, which is consistent with the physical understanding of polarized microwave radiative transfer over different surface types, as well as in snow and liquid clouds with different microphysical properties. Furthermore, the view-angle dependency artifact that the DPR’s precipitation flag bears with does not propagate into the conical-viewing GMI retrievals. This work provides a new and promising way for future physics-based ML retrieval algorithm development.

machine learning/artificial intelligence↗

A Near-Real-Time Model for Predicting Electricity Disruptions in Texas During Winter Storms

There has been an increase in extreme weather events, posing a threat to power grid systems, potentially influenced by factors such as population growth, changes in ecosystems, land cover, and land use in the service area, as well as the growth of certain vegetation types. This research seeks to develop a predictive model to mitigate potential damages caused by future winter storms. This research utilizes the Light Gradient Boosting Machine (LightGBM), incorporating the number of power outages experienced at the county level, geographic details, weather information, and lagged outage and lagged weather data. The developed models were broadly divided into two groups, with six models in each group - one group without optimization and another with optimization, totaling 12 trained models. For model optimization, Bayesian optimization was employed using Root Mean Squared Error (RMSE) as the objective function. In results, when comparing Group 2 (the optimized group) with Group 1 (the non-optimized group), it was found that optimization did not always lead to a reduction in RMSE and Mean Absolute Error (MAE). However, in terms of Mean Directional Accuracy (MDA), while all results in Group 1 were below the baseline accuracy of 0.33, all results in Group 2 exceeded 0.33, with some cases showing an increase of more than three times the baseline. The results indicated that, in the optimized model group, Population and Pressure were the most influential factors when using current weather data and geographical information. When using lagged data, lagged recorded outages and lagged Pressure emerged as the most significant factors. Among the 12 developed models, the L-1-2-O model showed the lowest RMSE and MAE, as well as the highest accuracy, with values of 390.62 households and 168.13 households, respectively. To normalize the RMSE and MAE values, each metric was divided by the average number of households among the counties in Texas. For the L-1-2-O model, the scaled RMSE was 0.88% and the scaled MAE was 0.38%. In terms of MDA, which indicates the accuracy of the prediction direction, the L-1-O model achieved the highest score of 0.41. Although this study focused on Texas, which suffered the greatest impact from the winter storms in 2021, with additional validation, the methodology used in this research could be applied to other regions.

Lee, Jangjae [Texas A & M Univ., College Station, ↗

Vickers hardness prediction from machine learning methods

Abstract The search for new superhard materials is of great interest for extreme industrial applications. However, the theoretical prediction of hardness is still a challenge for the scientific community, given the difficulty of modeling plastic behavior of solids. Different hardness models have been proposed over the years. Still, they are either too complicated to use, inaccurate when extrapolating to a wide variety of solids or require coding knowledge. In this investigation, we built a successful machine learning model that implements Gradient Boosting Regressor (GBR) to predict hardness and uses the mechanical properties of a solid (bulk modulus, shear modulus, Young’s modulus, and Poisson’s ratio) as input variables. The model was trained with an experimental Vickers hardness database of 143 materials, assuring various kinds of compounds. The input properties were calculated from the theoretical elastic tensor. The Materials Project’s database was explored to search for new superhard materials, and our results are in good agreement with the experimental data available. Other alternative models to compute hardness from mechanical properties are also discussed in this work. Our results are available in a free-access easy to use online application to be further used in future studies of new materials at www.hardnesscalculator.com .

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

Carbon nanotube (CNT) metal composites exhibit greatly reduced radiation damage

Radiation damage of structural materials leads to mechanical property degradation, eventually inducing failure. Secondary-phase dispersoids or other radiation defect sinks are often added to materials to boost their radiation resistance. We demonstrate that a metal composite made by adding 1D carbon nanotubes (CNTs) to aluminum (Al) exhibits superior radiation resistance. In situ ion irradiation with transmission electron microscopy (TEM) and atomistic simulations together reveal the mechanisms of rapid defect migration to CNTs, facilitating defect recombination and enhancing radiation tolerance. The origin of this effect is an evolving stress gradient in the Al matrix resulting from CNT transformation under irradiation, and the stability of resulting carbides. Extreme value statistics of large defect behavior in our simulations highlight the role of CNTs in reducing accumulated damage. Furthermore, this approach to controlling defect migration represents a promising opportunity to enhance the radiation resistance of nuclear materials without detrimental effects.

36 MATERIALS SCIENCE↗

Saturn Neutron Exosphere as Source for Inner and Innermost Radiation Belts

Energetic proton and electron measurements by the ongoing Cassini orbiter mission are expanding our knowledge of the highest energy components of the Saturn magnetosphere in the inner radiation belt region after the initial discoveries of these belts by the Pioneer 11 and Voyager 2 missions. Saturn has a neutron exosphere that extends throughout the magnetosphere from the cosmic ray albedo neutron source at the planetary main rings and atmosphere. The neutrons emitted from these sources at energies respectively above 4 and 8 eV escape the Saturn system, while those at lower energies are gravitationally bound. The neutrons undergo beta decay in average times of about 1000 seconds to provide distributed sources of protons and electrons throughout Saturn's magnetosphere with highest injection rates close to the Saturn and ring sources. The competing radiation belt source for energetic electrons is rapid inward diffusion and acceleration of electrons from the middle magnetosphere and beyond. Minimal losses during diffusive transport across the moon orbits, e.g. of Mimas and Enceladus, and local time asymmetries in electron intensity, suggest that drift resonance effects preferentially boost the diffusion rates of electrons from both sources. Energy dependences of longitudinal gradient-curvature drift speeds relative to the icy moons are likely responsible for hemispheric differences (e.g., Mimas, Tethys) in composition and thermal properties as at least partly produced by radiolytic processes. A continuing mystery is the similar radial profiles of lower energy (<10 MeV) protons in the inner belt region. Either the source of these lower energy protons is also neutron decay, but perhaps alternatively from atmospheric albedo, or else all protons from diverse distributed sources are similarly affected by losses at the moon' orbits, e.g. because the proton diffusion rates are extremely low. Enceladus cryovolcanism, and radiolytic processing elsewhere on the icy moon and ring surfaces, are additional sources of protons via ionization and charge exchange from breakup of water molecules. But one must then account somehow for local acceleration to the observed keV-MeV energies, since moon sweeping and E-ring absorption would remove protons diffusing inward from the middle magnetosphere. Although the main rings block further inward diffusion from the inner radiation belts, the exospheric neutron-decay source, combined with much slower diffusion of protons relative to electrons, may produce an innermost radiation belt in the gap between the upper atmosphere and the D-ring. This innermost belt will first be explored in-situ during the final proximal orbits of the Cassini mission.

Cooper, John↗

Photon acceleration of high-intensity vector vortex beams into the extreme ultraviolet

Extreme ultraviolet (XUV) light sources allow for the probing of bound electron dynamics on attosecond scales, interrogation of high-energy-density matter, and access to novel regimes of strong-field quantum electrodynamics. Despite the importance of these applications, coherent XUV sources remain relatively rare, and those that do exist are limited in their peak intensity and spatio-polarization structure. Here, we demonstrate that photon acceleration of an optical vector vortex pulse in the moving density gradient of an electron beam–driven plasma wave can produce a high-intensity, tunable-wavelength XUV pulse with the same vector vortex structure as the original pulse. Quasi-3D, boosted-frame particlein- cell simulations show the transition of optical vector vortex pulses with 800-nm wavelengths and intensities below 10 18 W/cm 2 to XUV vector vortex pulses with 36-nm wavelengths and intensities exceeding 10 20 W/cm 2 over a distance of 1.2 cm. The XUV pulses have sub-femtosecond durations and nearly flat phase fronts. The production of such high-quality, high-intensity XUV vector vortex pulses could expand the utility of XUV light as a diagnostic and driver of novel light–matter interactions.

43 PARTICLE ACCELERATORS↗

High-throughput injection-acceleration of electron bunches from a linear accelerator to a laser wakefield accelerator

Plasma-based accelerators (PBAs) driven by either intense lasers (laser wakefield accelerators, LWFAs) or particle beams (plasma wakefield accelerators, PWFAs), can accelerate charged particles at extremely high gradients compared to conventional radio-frequency (RF) accelerators. In the past two decades, great strides have been made in this field, making PBA a candidate for next-generation light sources and colliders. However, these challenging applications necessarily require beams with good stability, high quality, controllable polarization and excellent reproducibility. To date, such beams are generated only by conventional RF accelerators. As such, it is important to demonstrate the injection and acceleration of beams first produced using a conventional RF accelerator, by a PBA. In some recent studies on LWFA staging and external injection-acceleration in PWFA only a very small fraction (from below 0.1% to few percent) of the injected charge (the coupling efficiency) was accelerated. For future colliders where beam energy will need to be boosted using multiple stages, the coupling efficiency per stage must approach 100%. Here we report the first demonstration of external injection from a photocathode-RF-gun-based conventional linear accelerator (LINAC) into a LWFA and subsequent acceleration without any significant loss of charge or degradation of quality, which is achieved by properly shaping and matching the beam into the plasma structure. Furthermore, this is an important step towards realizing a high-throughput, multi-stage, high-energy, hybrid conventional-plasma accelerator.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Reliable photometric membership (RPM) of galaxies in clusters – I. A machine learning method and its performance in the local universe

ABSTRACT We introduce a new method to determine galaxy cluster membership based solely on photometric properties. We adopt a machine learning approach to recover a cluster membership probability from galaxy photometric parameters and finally derive a membership classification. After testing several machine learning techniques (such as stochastic gradient boosting, model averaged neural network and k-nearest neighbours), we found the support vector machine algorithm to perform better when applied to our data. Our training and validation data are from the Sloan Digital Sky Survey main sample. Hence, to be complete to $M_r^* + 3$, we limit our work to 30 clusters with $z$phot-cl ≤ 0.045. Masses (M200) are larger than $\sim 0.6\times 10^{14} \, \mathrm{M}_{\odot }$ (most above $3\times 10^{14} \, \mathrm{M}_{\odot }$). Our results are derived taking in account all galaxies in the line of sight of each cluster, with no photometric redshift cuts or background corrections. Our method is non-parametric, making no assumptions on the number density or luminosity profiles of galaxies in clusters. Our approach delivers extremely accurate results (completeness, C $\sim 92{\rm{ per\ cent}}$ and purity, P $\sim 87{\rm{ per\ cent}}$) within R200, so that we named our code reliable photometric membership. We discuss possible dependencies on magnitude, colour, and cluster mass. Finally, we present some applications of our method, stressing its impact to galaxy evolution and cosmological studies based on future large-scale surveys, such as eROSITA, EUCLID, and LSST.

Lopes, Paulo A. A.↗

EQC: Ensembled Quantum Computing for Variational Quantum Algorithms

Variational quantum algorithms (VQA), which are comprised of a classical optimizer and a parameterized quantum circuit, emerges as one of the most promising approaches of harvesting quantum power in the noisy-intermediate-scale-quantum (NISQ) era. However, the deployment of VQAs on today's NISQ devices often faces considerable system noise and prohibitively slow training speeds. On the other hand, the expensive supporting sources and infrastructure make quantum computers extremely keen on high utilization. In this paper, we propose a novel way of thinking about a quantum backend: rather than relying on one physical device which tends to introduce platform-specific noise and bias, a quantum ensemble, which distributes quantum tasks across parallel devices, can serve as a virtualized quantum computer for offering reduced noise levels through an adaptive mixture and also provide significantly improved training speeds through parallelization. With this idea, we build a distributive VQA optimization framework called DVQA, serving as the first effort in adopting parallel quantum devices for cooperative VQA training. To further constraint noise and speed-up convergence, we design a model for individual NISQ devices concerning their properties and running conditions, and propose a weighting mechanism for regularizing the returned gradients. Extensive evaluations on 10 IBM-Q quantum devices using the VQE example show that the distributive VQA training framework can substantially boost the training speed by 10.5x on average (up to 86x and at least 5.2x) with improved training accuracy.

Stein, Samuel A.↗