Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “random forest regression”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11

Machine learning models for rat multigeneration reproductive toxicity prediction

Reproductive toxicity is one of the prominent endpoints in the risk assessment of environmental and industrial chemicals. Due to the complexity of the reproductive system, traditional reproductive toxicity testing in animals, especially guideline multigeneration reproductive toxicity studies, take a long time and are expensive. Therefore, machine learning, as a promising alternative approach, should be considered when evaluating the reproductive toxicity of chemicals. We curated rat multigeneration reproductive toxicity testing data of 275 chemicals from ToxRefDB (Toxicity Reference Database) and developed predictive models using seven machine learning algorithms (decision tree, decision forest, random forest, k-nearest neighbors, support vector machine, linear discriminant analysis, and logistic regression). A consensus model was built based on the seven individual models. An external validation set was curated from the COSMOS database and the literature. The performances of individual and consensus models were evaluated using 500 iterations of 5-fold cross-validations and the external validation data set. The balanced accuracy of the models ranged from 58% to 65% in the 5-fold cross-validations and 45%–61% in the external validations. Prediction confidence analysis was conducted to provide additional information for more appropriate applications of the developed models. The impact of our findings is in increasing confidence in machine learning models. We demonstrate the importance of using consensus models for harnessing the benefits of multiple machine learning models (i.e., using redundant systems to check validity of outcomes). While we continue to build upon the models to better characterize weak toxicants, there is current utility in saving resources by being able to screen out strong reproductive toxicants before investing in vivo testing. The modeling approach (machine learning models) is offered for assessing the rat multigeneration reproductive toxicity of chemicals. Our results suggest that machine learning may be a promising alternative approach to evaluate the potential reproductive toxicity of chemicals.

consensus model↗

Interpretable Machine Learning Models for Autonomous Characterization of Analogue Ocean World Seawater Chemistry and Biosignature Potential Using Isotope Ratio Data

Background: Future missions to ocean worlds, such as Enceladus and Europa, will attempt to characterize the subsurface seawater chemistry and assess the potential for life. Such missions will be equipped with capabilities to precisely measure volatile isotopes in plumes, atmospheres, and exospheres. Motivation: While large isotopic fractionations can indicate a biological source, there are signatures resulting from abiotic geochemical processes that mimic isotopic biosignatures. While machine learning (ML) has the potential to disentangle competing effects and biotic mimicry, high-dimensional isotope ratio mass spectrometry (IRMS) data is likely to contain noise/irrelevant features and involve complex statistical interactions that make human inference and interpretation difficult. Further, ML predictions with as far-reaching implications as an extraterrestrial biosignature on an ocean world requires the use of interpretable models (i.e., not “black box” models) with physically and mathematically meaningful feature spaces along with false positive diagnostics. Methods: We use volatile CO2 IRMS data of analogue ocean world seawaters to validate an ML approach to provide biogeochemical context for biosignature detection. We employ a feature selection method called nearest-neighbor projected distance regression (NPDR) that detects statistical interactions and helps elucidate the mechanisms of the Random Forest classification models. Results: We train and validate predictive ML models on volatile CO2 IRMS data of analogue ocean world seawaters to predict major salt components (e.g., MgSO4, NaHCO3), pH, ionic strength, and the presence of biosignatures. Features derived from IRMS measurements are augmented with extracted time-series features. Our results show high test accuracy and interpretability, which is increased by interaction network visualization, sample-wise variable importance scores, and single-sample class probability estimates. We demonstrate an ML mission software solution that triggers autonomous data transmission and biogeochemical sample prediction.

geochemistry↗

A Machine Learning Initializer for Newton-Raphson AC Power Flow Convergence

Power flow computations are fundamental to many power system studies. Obtaining a converged power flow case is not a trivial task especially in large power grids due to the non-linear nature of the power flow equations. One key challenge is that the widely used Newton based power flow methods are sensitive to the initial voltage magnitude and angle estimates, and a bad initial estimate would lead to non-convergence. This paper addresses this challenge by developing a random-forest (RF) machine learning model to provide better initial voltage magnitude and angle estimates towards achieving power flow convergence. This method was implemented on a real ERCOT 6102 bus system under various operating conditions. By providing better Newton-Raphson initialization, the RF model precipitated the solution of 2,106 cases out of 3,899 non-converging dispatches. These cases could not be solved from flat start or by initialization with the voltage solution of a reference case. Finally, results obtained from the RF initializer performed better when compared with DC power flow initialization, Linear regression, and Decision Trees.

random forest↗

Characterization and prediction of the electromechanical wear of contact tips during wire arc additive manufacturing of 316L stainless steel

Here, this study seeks to better understand the degradation of the contact tip with respect to WAAM for a 316L wire electrode as well as explore methods of monitoring the contact tip state from process data. The contact tip, a consumable component, positions the wire and serves as the electrical contact surface between the wire electrode and the welding power supply. The wear of the contact tip was characterized in terms of material loss and material contamination for a set of tips worn to discrete levels as measured by the amount of wire fed or arc time. Geometrical characterization found a 49% increase in the bore exit area at 180 meters of wire fed. Machine learning models were developed to predict the relative bore exit area of the contact tip from arc-based process data and a random forest classifier exhibited favorable performance with a cross-validated f1-score of 0.84. The regression architecture implemented a multi-layer perceptron with the ability to predict the relative exit area with an $R^2$ score of 0.75. Key features used in the prediction include the standard deviation of the voltage and the time between shorts.

Contact tip wear↗

Recursive Blind Forecasting of Photovoltaic Generation and Consumer Load for Microgrids

Existing forecasting frameworks that predict time-series photovoltaic (PV) generation and consumer load for micro-grids' operation and control assume near-continuous availability of real-time predictors from the field. The incoming data are used to periodically re-train the models and update forecast snapshots over a moving horizon window. However, such frameworks are not resilient to disruptions in data availability caused by losses in communications between the field sensors and data loggers. This paper bridges the shortcoming by leveraging a previously proposed forecasting framework that is resilient to abrupt changes in data quality caused by communication losses. Assuming no availability of real-time field system data, which is typical in extreme weather events such as hurricanes, the framework uses lightweight recursive time-series models to independently forecast solar irradiance, ambient temperature, PV power, and consumer load for three horizon windows: 24 hours, 12 hours, and 1 hour. Four types of ensemble-based regression trees-simple gradient boosted trees (GBR), GBR with an adaptive component (A-GBR), random forests (RF), and extra trees (ExTR)-are leveraged and their performances are compared against a simple historical weekly mean. Numerical results show that A-GBR performs better on average by 32% for 24-hour horizon and 39% for 12-hour horizon, whereas ExTR outdoes the other models on average by 10% for 1-hour horizon.

Sundararajan, Aditya↗

Machine Learning Classification of Molten Salt Heat Exchanger Channel Plugging using Synthetic Data

This report addresses the requirements of Milestone M3.4 AI capability to identify and predict maintenance events. Development of digital twins (DT) for molten salt reactor (MSR) components is crucial for reducing operating and maintenance costs (O&M) and ensuring commercial viability of these reactors. Our focus is on development of DT for MSR primary system heat exchanger (HX), a critical component, the fault in which can reduce operating efficiency and force reactor shutdown. We are investigating the feasibility of a conceptual DT of HX consisting of internal distributed temperature sensing with fiber optics and machine learning (ML) algorithms to detect and localize faults. To determine the optimal approach to detection and localization of channel plugging, we benchmark seven different ML models: Logistic Regression, K-Nearest Neighbors (KNN), Gaussian Naïve Bayes, Support Vector Machines (SVM), Decision Tree Classifier, Random Forest Tree Classifier, and Feed-Forward Neural Network. ML algorithms are benchmarked using synthetic HX plugging data generated with computational fluid dynamics COMSOL software, with added brown noise to represent experimental noise. We show that the best performance is obtained with the Decision Tree classifier.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Increased salinity decreases annual gross primary productivity at a Northern California brackish tidal marsh

Tidal marshes sequester 11.4–87.0 Tg C yr –1 globally, but climate change impacts can threaten the carbon capture potential of these ecosystems. Tidal marshes occur across a wide range of salinity, with brackish marshes (0.5–18 ppt (parts per thousand)) dominating global tidal marsh extents. A diverse mix of freshwater- and saltwater-tolerant plant and microbial communities has led researchers to predict that carbon cycling in brackish wetlands may be less sensitive to changes in salinity than fresh- or saltwater wetlands. Rush Ranch, a well-monitored brackish tidal wetland of the San Francisco Bay National Estuarine Research Reserve, experiences highly variable annual salinity regimes. Within a five-year period (2014–2018), Rush Ranch experienced particularly extreme drought-induced salinization during the 2014 and 2015 growing seasons. During drought years, tidal channel salinity rose from a 15 year baseline of 4.7 ppt to growing season peaks of 10.3 ppt and 12.5 ppt. Continuous eddy covariance data from 2014 to 2018 demonstrate that during drought summers, gross primary productivity (GPP) decreased by 24%, whereas ecosystem respiration remained similar among all five years. Stepwise linear regression revealed that salinity, not air temperature or tidal height, was the dominant driver of annual GPP. A random forest model trained to predict GPP based on environmental data from low salinity years (i.e. naive to salinization) significantly over predicted GPP in drought years. When growing season salinities were doubled, annual estimates of net ecosystem exchange of CO 2 decreased by up to 30%. These results provide ecosystem-scale evidence that increased salinity influences CO 2 fluxes dominantly through reductions in GPP. This relationship provides a starting point for incorporating the effect of changes in salinity in wetland carbon models, which could improve wetland carbon forecasting and management for climate resilience.

54 ENVIRONMENTAL SCIENCES↗

Integrating very-high-resolution UAS data and airborne imaging spectroscopy to map the fractional composition of Arctic plant functional types in Western Alaska

Widespread changes in vegetation cover and composition are driving strong impacts on Arctic ecosystem functioning and global climate feedbacks. An accurate characterization of tundra vegetation composition is required to understand how the Arctic will respond to future climate change. However, quantifying tundra vegetation composition over large areas is challenging as commonly-used satellite observations are too coarse, spatially and spectrally, to differentiate low-lying tundra vegetation types. Recent airborne and spaceborne imaging spectroscopy platforms provide better data to characterize vegetation composition. Yet, our ability to characterize vegetation composition with imaging spectroscopy remains largely unexplored in the Arctic, particularly due to a lack of ground observations needed to train and test classification models. To address this problem, we collected very-high-resolution (VHR, ~5 cm) unoccupied aerial system (UAS) imagery at three low-Arctic tundra sites located on the Seward Peninsula, western Alaska. In this paper, we examine the feasibility of integrating imagery from the UAS and the hyperspectral Airborne Visible/Infrared Imaging Spectrometer, Next Generation (AVIRIS-NG) airborne instrument to map the fractional composition of 12 key Arctic plant functional types (PFTs). To this end, we first mapped the 12 PFTs from our VHR UAS imagery using random forest classification. We then used these UAS-derived PFT maps as ground truth to develop partial least squares regression (PLSR) models to predict the fractional cover (FCover) of each PFT from AVIRIS-NG imagery. Further, we evaluated the performance of our PLSR models using reserved UAS samples, as well as by mapping PFT FCover and dominant PFT for large tundra landscapes. Our results show that 1) Arctic PFTs can be effectively mapped using VHR UAS imagery, with overall accuracy between 86% and 92%, 2) when the UAS mapped PFTs were used to inform PLSR scaling models, the FCover of the 12 PFTs could be effectively estimated from AVIRIS-NG imagery with a mean absolute error (MAE) <0.13, and 3) our PLSR models outperformed traditional, fully constrained least-squares (FCLS) linear mixture analysis and produced high-quality, spatially contiguous PFT FCover and PFT maps that captured vegetation spatial patterns with similar accuracy to those developed from UAS imagery. The developed PLSR models have the potential to be broadly applied for quantifying vegetation composition with AVIRIS-NG images to help monitor tundra vegetation dynamics and improve process-based modeling of tundra ecosystems.

54 ENVIRONMENTAL SCIENCES↗

A framework to evaluate machine learning crystal stability predictions

The rapid adoption of machine learning in various scientific domains calls for the development of best practices and community agreed-upon benchmarking tasks and metrics. We present Matbench Discovery as an example evaluation framework for machine learning energy models, here applied as pre-filters to first-principles computed data in a high-throughput search for stable inorganic crystals. We address the disconnect between (1) thermodynamic stability and formation energy and (2) retrospective and prospective benchmarking for materials discovery. Alongside this paper, we publish a Python package to aid with future model submissions and a growing online leaderboard with adaptive user-defined weighting of various performance metrics allowing researchers to prioritize the metrics they value most. To answer the question of which machine learning methodology performs best at materials discovery, our initial release includes random forests, graph neural networks, one-shot predictors, iterative Bayesian optimizers and universal interatomic potentials. We highlight a misalignment between commonly used regression metrics and more task-relevant classification metrics for materials discovery. Accurate regressors are susceptible to unexpectedly high false-positive rates if those accurate predictions lie close to the decision boundary at 0 eV per atom above the convex hull. The benchmark results demonstrate that universal interatomic potentials have advanced sufficiently to effectively and cheaply pre-screen thermodynamic stable hypothetical materials in future expansions of high-throughput materials databases.

Riebesell, Janosh↗

Machine learning enhanced predictions of ICRF heating: Overcoming numerical limitations via data curation

In this work, we present the development of robust surrogate models for Ion Cyclotron Range of Frequencies (ICRF) and High-Harmonic Fast Wave (HHFW) heating predictions in fusion plasmas. Building upon our previous efforts to achieve real-time capable models, we identify the cause of the outliers found using TORIC in certain HHFW heating scenarios. The outliers are observed to be spurious ion Bernstein wave (IBW)-like modes caused by a wavelength control algorithm designed to address challenging scenarios with high perpendicular wavenumbers. The effect arises from the modulation in the perpendicular susceptibility, which can induce sign reversal and IBW-like propagation for scenarios featuring normalized ion Larmor radius λ i ≫ 1. We use TORIC with this algorithm disabled to generate a novel HHFW-NSTX database that is free of outliers. Surrogate models trained on this database, including Random Forest Regressor (RFR), Multi-Layer Perceptrons, and Gaussian Process Regressors (GPR), demonstrate the ability to accurately predict HHFW heating profiles, with regression scores of R 2 ∈[0.93−0.99]. Additionally we demonstrate that it is possible to generalize predictions beyond training data by the use of both RFR and GPR models, enabling the prediction of scenarios previously limited to the original model. GPR models also provide uncertainty quantification, offering insights into model confidence. This work introduces a comprehensive Verification, Validation, and Uncertainty Quantification methodology for surrogate modeling, applicable not only to ICRF heating but also to other RF heating challenges and fusion physics problems. Beyond accelerated inference, these models show effective extrapolation capabilities, providing an alternative for addressing numerical challenges.

Artificial neural networks↗

Predicting Biomass Yields of Advanced Switchgrass Cultivars for Bioenergy and Ecosystem Services Using Machine Learning

The production of advanced perennial bioenergy crops within marginal areas of the agricultural landscape is gaining interest due to its potential to sustainably produce feedstocks for biofuels and bioproducts while also improving the sustainability and resilience of commodity crop production. However, predicting the biomass yields of this production system is challenging because marginal areas are often relatively small and spread around agricultural fields and are typically associated with various abiotic conditions that limit crop production. Machine learning (ML) offers a viable solution as a biomass yield prediction tool because it is suited to predicting relationships with complex functional associations. The objectives of this study were to (1) evaluate the accuracy of commonly applied ML algorithms in agricultural applications for predicting the biomass yields of advanced switchgrass cultivars for bioenergy and ecosystem services and (2) determine the most important biomass yield predictors. Datasets on biomass yield, weather, land marginality, soil properties, and agronomic management were generated from three field study sites in two U.S. Midwest states (Illinois and Iowa) over three growing seasons. The ML algorithms evaluated in the study included random forests (RFs), gradient boosting machines (GBMs), artificial neural networks (ANNs), K-neighbors regressor (KNR), AdaBoost regressor (ABR), and partial least squares regression (PLSR). Coefficient of determination (R 2 ) and mean absolute error (MAE) were used to evaluate the predictive accuracy of the tested algorithms. Results showed that the ensemble methods, RF (R 2 = 0.86, MAE = 0.62 Mg/ha), GBM (R 2 = 0.88, MAE = 0.57 Mg/ha), and GBM (R 2 = 0.78, MAE = 0.66 Mg/ha), were the most accurate in predicting biomass yields of the Independence, Liberty, and Shawnee switchgrass cultivars, respectively. This is in agreement with similar studies that apply ML to multi-feature problems where traditional statistical methods are less applicable and datasets used were considered to be relatively small for ANNs. Consistent with previous studies on switchgrass, the most important predictors of biomass yield included average annual temperature, average growing season temperature, sum of the growing season precipitation, field slope, and elevation. This study helps pave the way for applying ML as a management tool for alternative bioenergy landscapes where understanding agronomic and environmental performance of a multifunctional cropping system seasonally and interannually at the sub-field scale is critical.

09 BIOMASS FUELS↗

Comparing Calibration Algorithms for the Rapid Characterization of Pretreated Corn Stover Using Near-Infrared Spectroscopy

Rapid characterization of biomass composition is a key enabling technology for biorefineries—the ability to measure the chemical composition of biomass materials entering the biorefinery as well as the composition of key process intermediate streams would allow real-time process control and the development of robust models to predict process performance. The utility of near-infrared (NIR) spectroscopy for rapid characterization requires multivariate algorithms for building calibration models. The most prevalent algorithm used for building calibration models using NIR spectra is the linear modeling algorithm Partial Least Squares Regression (PLS). Nonlinear regression algorithms (which are typically more computationally intensive than linear modeling approaches) have gained popularity in recent years due to their ability to solve a wide variety of classification and regression problems and the dramatic increase in available computational resources. In this work, we demonstrate that a calibration model can predict the composition of corn stover process intermediate samples pretreated with three different treatments—hot water (HW), dilute acid (DA), and deacetylation followed by dilute acid (DDA). We quantitatively compare three different algorithms for building prediction models based on near-infrared spectroscopy—partial least squares (PLS), support vector machines (SVM), and random forests (RF). We demonstrate the utility of improving model performance by accounting for instrument performance variability using repeated measurements of standard materials (e.g., the “repeatability file” strategy) and investigate its performance with nonlinear regression techniques, and we discuss methods for quantifying the uncertainties of specific predictions among the three methods.

09 BIOMASS FUELS↗

Risk-Aware Framework Development for Disruption Prediction: Alcator C-Mod and DIII-D Survival Analysis

Abstract Survival regression models can achieve longer warning times at similar receiver operating characteristic performance than previously investigated models. Survival regression models are also shown to predict the time until a disruption will occur with lower error than other predictors. Time-to-event predictions from time-series data can be obtained with a survival analysis statistical framework, and there have been many tools developed for this task which we aim to apply to disruption prediction. Using the open-source Auton-Survival package we have implemented disruption predictors with the survival regression models Cox Proportional Hazards, Deep Cox Proportional Hazards, and Deep Survival Machines. To compare with previous work, we also include predictors using a Random Forest binary classifier, and a conditional Kaplan-Meier formalism. We benchmarked the performance of these five predictors using experimental data from the Alcator C-Mod and DIII-D tokamaks by simulating alarms on each individual shot. We find that developing machine-relevant metrics to evaluate models is an important area for future work. While this study finds cases where disruptive conditions are not predicted, there are instances where the desired outcome is produced. Giving the plasma control system the expected time-to-disruption will allow it to determine the optimal actuator response in real time to minimize risk of damage to the device.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Verbal Learning and Memory Deficits across Neurological and Neuropsychiatric Disorders: Insights from an ENIGMA Mega Analysis

Deficits in memory performance have been linked to a wide range of neurological and neuropsychiatric conditions. While many studies have assessed the memory impacts of individual conditions, this study considers a broader perspective by evaluating how memory recall is differentially associated with nine common neuropsychiatric conditions using data drawn from 55 international studies, aggregating 15,883 unique participants aged 15–90. The effects of dementia, mild cognitive impairment, Parkinson’s disease, traumatic brain injury, stroke, depression, attention-deficit/hyperactivity disorder (ADHD), schizophrenia, and bipolar disorder on immediate, short-, and long-delay verbal learning and memory (VLM) scores were estimated relative to matched healthy individuals. Random forest models identified age, years of education, and site as important VLM covariates. A Bayesian harmonization approach was used to isolate and remove site effects. Regression estimated the adjusted association of each clinical group with VLM scores. Memory deficits were strongly associated with dementia and schizophrenia (p < 0.001), while neither depression nor ADHD showed consistent associations with VLM scores (p > 0.05). Differences associated with clinical conditions were larger for longer delayed recall duration items. By comparing VLM across clinical conditions, this study provides a foundation for enhanced diagnostic precision and offers new insights into disease management of comorbid disorders.

Neurosciences & Neurology↗

Southern Rockies Western Slope Agriculture: Identifying Drivers of Rangeland Production for Drought Planning on the Western Slope of the Southern Rockies

Over the last decade, the southern Rocky Mountains of the United States experienced severe and variable drought. Local ranchers and landowners have reported strain on their operations, citing decreasing forage for their cattle and a need to adjust their business models. This study identified Major Land Resource Area-48 (MLRA-48) and northwestern Colorado as the key region for analysis. NASA DEVELOP partnered with the BLM Colorado River Field Office, Colorado State University Extension, USDA Forest Service, and the National Drought Mitigation Center to address concerns regarding the efficacy of remotely sensed rangeland production platforms and identify early warning climatic indicators of drought. The study identified two key platforms, The Rangeland Productivity Monitoring Service (RPMS) and Rangeland Analysis Platform (RAP), which use NASA Landsat 5 TM, Landsat 7 ETM+, Landsat 8 OLI, and Landsat 9 OLI-2 to estimate rangeland biomass. We regressed these with in situ biomass data to validate their efficacy and found that RAP was more effective than RPMS in estimating rangeland biomass, though it presents a tendency to overestimate. Our study performed a random forest analysis, comparing monthly RAP biomass estimates to a variety of climate variables, including mean precipitation, temperature, Palmer Drought Severity Index, snow water equivalent, snow persistence from Terra MODIS, wind speed and direction, and vapor pressure deficit. We determined that vapor pressure deficit and precipitation are key indicators in predicting forage production in MLRA-48. Our climate analysis provided our partners with greater understanding of the influence of various climate variables in determining rangeland production and allows them to assist land managers in drought mitigation.

remote sensing↗

Improving Prediction of Surface Solar Irradiance Variability by Integrating Observed Cloud Characteristics and Machine Learning

A 5-year, 1-minute resolution observational dataset of clouds and solar radiation was produced that includes two metrics of the variability in surface solar irradiance due to cloud type and fractional sky cover. Multiple regression models were trained to fit observations of surface solar irradiance variability from those two cloud property predictors. We found that ensemble tree-based methods, Random Forest and Gradient Boosting Machine, have the least overfitting issues and showed the best performance with an R2 of 0.42. While the observational data trained in this study was only from one site, the U.S. Department of Energy (DOE) Atmospheric Radiation Measurement (ARM) Southern Great Plains (SGP) site in Oklahoma, initial comparisons of the seasonality of the statistics suggest that these results are relatively weather regime independent; the generality of such a finding across sites will be tested in future work. The observational data and developed machine learning model are being used to create a numerical weather prediction model parameterization to enable day-ahead solar variability prediction in a computationally efficient way. This is a first step towards creating a new paradigm of predicting day-ahead variability with the potential to provide a new tool to improve grid operation, planning, and resilience.

Riihimaki, Laura↗

What drives the scatter of local star-forming galaxies in the BPT diagrams? A Machine Learning based analysis

ABSTRACT We investigate which physical properties are most predictive of the position of local star forming galaxies on the BPT diagrams, by means of different Machine Learning (ML) algorithms. Exploiting the large statistics from the Sloan Digital Sky Survey (SDSS), we define a framework in which the deviation of star-forming galaxies from their median sequence can be described in terms of the relative variations in a variety of observational parameters. We train artificial neural networks (ANN) and random forest (RF) trees to predict whether galaxies are offset above or below the sequence (via classification), and to estimate the exact magnitude of the offset itself (via regression). We find, with high significance, that parameters primarily associated to variations in the nitrogen-over-oxygen abundance ratio (N/O) are the most predictive for the [N ii]-BPT diagram, whereas properties related to star formation (like variations in SFR or EW(H α)) perform better in the [S ii]-BPT diagram. We interpret the former as a reflection of the N/O–O/H relationship for local galaxies, while the latter as primarily tracing the variation in the effective size of the S+ emitting region, which directly impacts the [S ii] emission lines. This analysis paves the way to assess to what extent the physics shaping local BPT diagrams is also responsible for the offsets seen in high redshift galaxies or, instead, whether a different framework or even different mechanisms need to be invoked.

79 ASTRONOMY AND ASTROPHYSICS↗

A Machine Learning-based Reliability Evaluation Model for Integrated Power-Gas Systems

This article proposes a hybrid machine learning method for the reliability evaluation of integrated power-gas systems (IPGS) under the uncertain component failure probability distributions. The Random Forest (RF) method is designed to select important features to solve the insufficient quantity of data and the curse of dimensionality problems. The Extreme Gradient Boosting (XGBoost) regression algorithm is developed to quantify the relationship between the uncertain parameters and reliability metrics. Moreover, a ten-fold cross-validation method is employed to further improve the accuracy of the regression model. Simulation results on three test systems show that the proposed method can achieve high accuracy for the reliability evaluation.

24 POWER TRANSMISSION AND DISTRIBUTION↗