Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “machine learning, Random Forest”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

Machine Learning-Based Classification of Lignocellulosic Biomass from Pyrolysis-Molecular Beam Mass Spectrometry Data

High-throughput analysis of biomass is necessary to ensure consistent and uniform feedstocks for agricultural and bioenergy applications and is needed to inform genomics and systems biology models. Pyrolysis followed by mass spectrometry such as molecular beam mass spectrometry (py-MBMS) analyses are becoming increasingly popular for the rapid analysis of biomass cell wall composition and typically require the use of different data analysis tools depending on the need and application. Here, the authors report the py-MBMS analysis of several types of lignocellulosic biomass to gain an understanding of spectral patterns and variation with associated biomass composition and use machine learning approaches to classify, differentiate, and predict biomass types on the basis of py-MBMS spectra. Py-MBMS spectra were also corrected for instrumental variance using generalized linear modeling (GLM) based on the use of select ions relative abundances as spike-in controls. Machine learning classification algorithms e.g., random forest, k-nearest neighbor, decision tree, Gaussian Naïve Bayes, gradient boosting, and multilayer perceptron classifiers were used. The k-nearest neighbors (k-NN) classifier generally performed the best for classifications using raw spectral data, and the decision tree classifier performed the worst. After normalization of spectra to account for instrumental variance, all the classifiers had comparable and generally acceptable performance for predicting the biomass types, although the k-NN and decision tree classifiers were not as accurate for prediction of specific sample types. Gaussian Naïve Bayes (GNB) and extreme gradient boosting (XGB) classifiers performed better than the k-NN and the decision tree classifiers for the prediction of biomass mixtures. The data analysis workflow reported here could be applied and extended for comparison of biomass samples of varying types, species, phenotypes, and/or genotypes or subjected to different treatments, environments, etc. to further elucidate the sources of spectral variance, patterns, and to infer compositional information based on spectral analysis, particularly for analysis of data without a priori knowledge of the feedstock composition or identity.

59 BASIC BIOLOGICAL SCIENCES↗

Evaluation of Classifier Complexity for Delay Tolerant Network Routing

The growing popularity of small cost effective satellites (SmallSats, CubeSats, etc.) creates the potential for a variety of new science applications involving multiple nodes functioning together or independently to achieve a task, such as swarms and constellations. As this technology develops and is deployed for missions in Low Earth Orbit and beyond, the use of delay tolerant networking (DTN) techniques may improve communication capabilities within the network. In this paper, a network hierarchy is developed from heterogeneous networks of SmallSats, surface vehicles, relay satellites and ground stations which form an integrated network. There is a tradeoff between complexity, flexibility, and scalability of user defined schedules versus autonomous routing as the number of nodes in the network increases. To address these issues, this work proposes a machine learning classifier based on DTN routing metrics. A framework is developed which will allow for the use of several categories of machine learning algorithms (decision tree, random forest and deep learning) to be applied to a dataset of historical network statistics, which allows for the evaluation of algorithm complexity versus performance to be explored. We develop the emulation of a hierarchical network, consisting of tens of nodes which form a cognitive network architecture. CORE (Common Open Research Emulator) is used to emulate the network using bundle protocol and DTN IP neighbor discovery.

Dudukovich, Rachel↗

Tree-Based Ensemble Learning Models for Wall Temperature Predictions in Post-Critical Heat Flux Flow Regimes at Subcooled and Low-Quality Conditions

Accurately predicting post-critical heat flux (CHF) heat transfer is an important but challenging task in water-cooled reactor design and safety analysis. Although numerous heat transfer correlations have been developed to predict post-CHF heat transfer, these correlations are only applicable to relatively narrow ranges of flow conditions due to the complex physical nature of the post-CHF heat transfer regimes. In this paper, a large quantity of experimental data is collected and summarized from the literature for steady-state subcooled and low-quality film boiling regimes with water as the working fluid in vertical tubular test sections. In addition, a low-quality water film boiling (LWFB) database is consolidated with a total of 22,813 experimental data points, which cover a wide flow range of the system pressure from 0.1 to 9.0 MPa, mass flux from 25 to 2750 kg/m 2 s, and inlet subcooling from 1 to 70 °C. Two machine learning (ML) models, based on random forest (RF) and gradient boosted decision tree (GBDT), are trained and validated to predict wall temperatures in post-CHF flow regimes. The trained ML models demonstrate significantly improved accuracies compared to conventional empirical correlations. To further evaluate the performance of these two ML models from a statistical perspective, three criteria are investigated and three metrics are calculated to quantitatively assess the accuracy of these two ML models. For the full LWFB database, the root-mean-square errors between the measured and predicted wall temperatures by the GBDT and RF models are 5.7% and 6.2%, respectively, confirming the accuracy of the two ML models.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Machine Learning for Improving Surface-Layer-Flux Estimates

Abstract Flows in the atmospheric boundary layer are turbulent, characterized by a large Reynolds number, the existence of a roughness sublayer and the absence of a well-defined viscous layer. Exchanges with the surface are therefore dominated by turbulent fluxes. In numerical models for atmospheric flows, turbulent fluxes must be specified at the surface; however, surface fluxes are not known a priori and therefore must be parametrized. Atmospheric flow models, including global circulation, limited area models, and large-eddy simulation, employ Monin–Obukhov similarity theory (MOST) to parametrize surface fluxes. The MOST approach is a semi-empirical formulation that accounts for atmospheric stability effects through universal stability functions. The stability functions are determined based on limited observations using simple regression as a function of the non-dimensional stability parameter representing a ratio of distance from the surface and the Obukhov length scale (Obukhov in Trudy Inst Theor Geofiz AN SSSR 1:95–115, 1946), $$z/L$$ z / L . However, simple regression cannot capture the relationship between governing parameters and surface-layer structure under the wide range of conditions to which MOST is commonly applied. We therefore develop, train, and test two machine-learning models, an artificial neural network (ANN) and random forest (RF), to estimate surface fluxes of momentum, sensible heat, and moisture based on surface and near-surface observations. To train and test these machine-learning algorithms, we use several years of observations from the Cabauw mast in the Netherlands and from the National Oceanic and Atmospheric Administration’s Field Research Division tower in Idaho. The RF and ANN models outperform MOST. Even when we train the RF and ANN on one set of data and apply them to the second set, they provide more accurate estimates of all of the fluxes compared to MOST. Estimates of sensible heat and moisture fluxes are significantly improved, and model interpretability techniques highlight the logical physical relationships we expect in surface-layer processes.

Meteorology & Atmospheric Sciences↗

Large-Scale High-Resolution Coastal Mangrove Forests Mapping Across West Africa With Machine Learning Ensemble and Satellite Big Data

Coastal mangrove forests provide important ecosystem goods and services, including carbon sequestration, biodiversity conservation, and hazard mitigation. However, they are being destroyed at an alarming rate by human activities. To characterize mangrove forest changes, evaluate their impacts, and support relevant protection and restoration decision making, accurate and up-to-date mangrove extent mapping at large spatial scales is essential. Available large-scale mangrove extent data products use a single machine learning method commonly with 30 m Landsat imagery, and significant inconsistencies remain among these data products. With huge amounts of satellite data involved and the heterogeneity of land surface characteristics across large geographic areas, finding the most suitable method for large-scale high-resolution mangrove mapping is a challenge. The objective of this study is to evaluate the performance of a machine learning ensemble for mangrove forest mapping at 20 m spatial resolution across West Africa using Sentinel-2 (optical) and Sentinel-1 (radar) imagery. The machine learning ensemble integrates three commonly used machine learning methods in land cover and land use mapping, including Random Forest (RF), Gradient Boosting Machine (GBM), and Neural Network (NN). The cloud-based big geospatial data processing platform Google Earth Engine (GEE) was used for pre-processing Sentinel-2 and Sentinel-1 data. Extensive validation has demonstrated that the machine learning ensemble can generate mangrove extent maps at high accuracies for all study regions in West Africa (92%–99% Producer’s Accuracy, 98%–100% User’s Accuracy, 95%–99% Overall Accuracy). This is the first-time that mangrove extent has been mapped at a 20 m spatial resolution across West Africa. The machine learning ensemble has the potential to be applied to other regions of the world and is therefore capable of producing high-resolution mangrove extent maps at global scales periodically.

coastal environment↗

How accurate is a machine learning-based wind speed extrapolation under a round-robin approach?

As the size of commercial wind turbines keeps increasing, having accurate ways to vertically extrapolate wind speed is essential to obtain a precise characterization of the wind resource for wind energy production. Recently, machine learning has been proposed and applied to extrapolate wind speed to hub heights. However, previous studies trained and tested the machine learning methods at the same site, giving them an unfair advantage over the conventional extrapolation techniques, which are instead more universal. Here, we use data from four sites in Oklahoma to test a round-robin validation approach for machine learning, under which we train a random forest at a site, and test it at a different site, where the model has no prior knowledge of the wind resource. We quantify how the accuracy of this technique varies with distance from the training site, and we find that it outperforms conventional techniques for wind extrapolation at all the considered spatial separations. We then assess how the accuracy of the machine-learning based approach varies when it is used to predict wind speed in a wind farm far wake. Finally, we explore as case study the performance of the random forest in extrapolating winds during a low-level jet event.

17 WIND ENERGY↗

Machine learning study of magnetism in uranium-based compounds

Actinide and lanthanide-based materials display exotic properties that originate from the presence of itinerant or localized f electrons and include unconventional superconductivity and magnetism, hidden order, and heavy-fermion behavior. Due to the strongly correlated nature of the 5f electrons, magnetic properties of these compounds depend sensitively on applied magnetic field and pressure, as well as on chemical doping. However, precise connection between the structure and magnetism in actinide-based materials is currently unclear. In this investigation, we established such structure-property links by assembling and mining two datasets that aggregate, respectively, the results of high-throughput density functional theory simulations and experimental measurements for the families of uranium- and neptunium-based binary compounds. Various regression algorithms were utilized to identify correlations among accessible attributes (features or descriptors) of the material systems and predict their cation magnetic moments and general forms of magnetic ordering. Descriptors representing compound structural parameters and cation f-subshell occupation numbers were identified as most important for accurate predictions. The best machine learning model developed employs the random forest regression algorithm. It can predict both spin and orbit moment size with root-mean-square error of 0.17 μ B and 0.19 μ B , respectively. Lastly, the random forest classification algorithm is used to predict the ordering (paramagnetic, ferromagnetic, and antiferromagnetic) of such systems with 76% accuracy.

36 MATERIALS SCIENCE↗

Vehicle Position Detection Based on Machine Learning Algorithms in Dynamic Wireless Charging

Dynamic wireless charging (DWC) has emerged as a viable approach to mitigate range anxiety by ensuring continuous and uninterrupted charging for electric vehicles in motion. DWC systems rely on the length of the transmitter, which can be categorized into long-track transmitters and segmented coil arrays. The segmented coil array, favored for its heightened efficiency and reduced electromagnetic interference, stands out as the preferred option. However, in such DWC systems, the need arises to detect the vehicle’s position, specifically to activate the transmitter coils aligned with the receiver pad and de-energize uncoupled transmitter coils. This paper introduces various machine learning algorithms for precise vehicle position determination, accommodating diverse ground clearances of electric vehicles and various speeds. Through testing eight different machine learning algorithms and comparing the results, the random forest algorithm emerged as superior, displaying the lowest error in predicting the actual position.

47 OTHER INSTRUMENTATION↗

Uncertainty-Aware Machine Learning for Small-Angle X-ray Scattering Analysis in Autonomous Experimentation

Small-angle X-ray scattering (SAXS) is a powerful high-throughput characterization tool for probing nanoscale structure in native sample environments, providing real-time morphological information such as nanoparticle size and shape during synthesis. However, automated SAXS data analysis for extracting meaningful structural parameters is non-trivial and remains a bottleneck in closed-loop experimentation towards autonomous materials discovery, which demands fast, reliable, and uncertainty-aware data analysis. Here, we develop a machine-learning approach for automated SAXS analysis tailored to closed-loop nanoparticle synthesis. A Random Forest (RF) regression model is trained on 100,000 synthetic SAXS curves generated from polydisperse spherical nanoparticles with realistic background contributions. Using normalized one-dimensional SAXS intensity profiles as input, the RF model directly predicts nanoparticle radius, size polydispersity, and background parameters, while the ensemble standard deviation across trees provides built-in uncertainty quantification (UQ). On synthetic data, we show that combining fit-quality metrics (R 2 , MAE) with thresholds on prediction uncertainty reliably identifies accurate parameter estimates without access to ground truth. We then apply the trained model to 365 experimental SAXS profiles of citrate-reduced gold nanoparticles synthesized using an automated droplet-flow microreactor with in situ SAXS at a synchrotron beamline, classifying the results into high- and low-confidence subsets based on UQ metrics. Finally, we integrate RF-based SAXS analysis into a simulated closed-loop optimization campaign using Gaussian process Bayesian optimization to minimize nanoparticle polydispersity, benchmarking against conventional automated Levenberg–Marquardt fitting. The RF-guided campaign exhibits substantially faster convergence and lower relative opportunity cost (∼0.07 vs ∼0.3), demonstrating that uncertainty-aware machine-learning SAXS analysis significantly enhances the efficiency and robustness of autonomous nanomaterials synthesis workflows.

Bayesian optimization↗

Low-Cost Sensor Performance Intercomparison, Correction Factor Development, and 2+ Years of Ambient PM2.5 Monitoring in Accra, Ghana

Particulate matter air pollution is a leading cause of global mortality, particularly in Asia and Africa. Addressing the high and wide-ranging air pollution levels requires ambient monitoring, but many low- and middle-income countries (LMICs) remain scarcely monitored. To address these data gaps, recent studies have utilized low-cost sensors. These sensors have varied performance, and little literature exists about sensor intercomparison in Africa. By colocating 2 QuantAQ Modulair-PM, 2 PurpleAir PA-II SD, and 16 Clarity Node-S Generation II monitors with a reference-grade Teledyne monitor in Accra, Ghana, we present the first intercomparisons of different brands of low-cost sensors in Africa, demonstrating that each type of low-cost sensor PM2.5 is strongly correlated with reference PM2.5, but biased high for ambient mixture of sources found in Accra. When compared to a reference monitor, the QuantAQ Modulair-PM has the lowest mean absolute error at 3.04 μg/m3, followed by PurpleAir PA-II (4.54 μg/m3) and Clarity Node-S (13.68 μg/m3). We also compare the usage of 4 statistical or machine learning models (Multiple Linear Regression, Random Forest, Gaussian Mixture Regression, and XGBoost) to correct low-cost sensors data, and find that XGBoost performs the best in testing (R2: 0.97, 0.94, 0.96; mean absolute error: 0.56, 0.80, and 0.68 μg/m3 for PurpleAir PA-II, Clarity Node-S, and Modulair-PM, respectively), but tree-based models do not perform well when correcting data outside the range of the colocation training. Therefore, we used Gaussian Mixture Regression to correct data from the network of 17 Clarity Node-S monitors deployed around Accra, Ghana, from 2018 to 2021. We find that the network daily average PM2.5 concentration in Accra is 23.4 μg/m3, which is 1.6 times the World Health Organization Daily PM2.5 guideline of 15 μg/m3. While this level is lower than those seen in some larger African cities (such as Kinshasa, Democratic Republic of the Congo), mitigation strategies should be developed soon to prevent further impairment to air quality as Accra, and Ghana as a whole, rapidly grow.

Humidity↗

Mapping wall-to-wall fractional cover of Arctic tundra plant functional types in Alaska using 20-m spatial resolution satellite imagery and harmonized plot observations

Estimates of fractional cover (fCover) across given land surfaces are used to assess, and often model, vegetation composition and diversity, which are crucial for understanding the health and functioning of terrestrial ecosystems. Remote sensing provides a useful means for scaling local, plot-measured fCover estimates to regional scales. Leveraging a recently synthesized and harmonized plot database, this study generated wall-to-wall maps of fCover for six Alaskan-Arctic plant functional types (PFT), including non-vascular plants, forbs, graminoids, and deciduous and evergreen shrubs, using 20-m satellite data (Sentinel-1, Sentinel-2, ArcticDEM) using a machine learning regression approach, specifically the random forest (RF) algorithm, which is well-suited for handling nonlinear relationships and high-dimensional satellite datasets. This study additionally addressed the spatio-temporal inconsistencies e.g., sampling scale, plot size, and collection year in plot measured fCover by adopting a multivariate outlier detection approach—Cook’s distance—to identify high-quality plots for model training and validation. Our approach achieves high accuracy (R 2 = 0.59–0.93, root mean squared errors = 0.02–0.10 for all PFTs) between plot-observed and satellite-derived fCover when using high-quality plot samples. The mapped fCover characterizes the spatial patterns of different PFTs across the tundra biome at a 20-m resolution, providing key information needed for improved representation of Arctic tundra vegetation in terrestrial biosphere models to better understand climate-vegetation feedback across the Arctic tundra.

Arctic tundra↗

High-throughput screening of tribological properties of monolayer films using molecular dynamics and machine learning

Monolayer films have shown promise as a lubricating layer to reduce friction and wear of mechanical devices with separations on the nanoscale. These films have a vast design space with many tunable properties that can affect their tribological effectiveness. For example, terminal group chemistry, film composition, and backbone chemistry can all lead to films with significantly different tribological properties. This design space, however, is very difficult to explore without a combinatorial approach and an automatable, reproducible, and extensible workflow to screen for promising candidate films. Here, using the Molecular Simulation Design Framework (MoSDeF), a combinatorial screening study was performed to explore 9747 unique monolayer films (116 964 total simulations) and a machine learning (ML) model using a random forest regressor, an ensemble learning technique, to explore the role of terminal group chemistry and its effect on tribological effectiveness. The most promising films were found to contain small terminal groups such as cyano and ethylene. The ML model was subsequently applied to screen terminal group candidates identified from the ChEMBL small molecule library. Approximately 193 131 unique film candidates were screened with approximately a five order of magnitude speed-up in analysis compared to simulation alone. The ML model was thus able to be used as a predictive tool to greatly speed up the initial screening of promising candidate films for future simulation studies, suggesting that computational screening in combination with ML can greatly increase the throughput in combinatorial approaches to generate in silico data and then train ML models in a controlled, self-consistent fashion.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Data-driven acoustic measurement of moisture content in flowing biomass

Measuring the moisture content in flowing biomass is critical to processes such as liquid biofuel conversion, such as biogasoline, biodiesel, bio jet kerosene, etc. However, biomass tends to flow in aggregates, which results in significant inhomogeneities in the amount of biomass flowing in front of a sensor at a given time, and there can be significant overlap in the material properties of dry vs wet biomass, leading to poor signal-to-noise ratio. We present a technique for identifying biomass moisture content using a series of acoustic pitch-catch measurements to quantify the sound speed and acoustic amplitude through the biomass, in conjunction with classical machine learning techniques, including Naive Bayes, Random Forest, and K-Nearest Neighbors classification. We amplify the differences between the acoustic measurements in different moisture levels by collecting a series of pulse-echo measurements, which we sort in order of ascending sound speed. We test the accuracy of the technique on experimentally-prepared batches of corn stover biomass with specified moisture levels and measure the average error in the estimated moisture level as a function of the number of pitch-catch measurements used. We observe average estimation errors as low as 6.7% by increasing the number of measurements and optimizing the hyperparameters. This work presents a novel method determining moisture content in flowing biomass with inhomogeneous flow. Additionally, this technique has application in optimizing biomass conversion processes, as well as other fields including, paper production, natural fiber processing, and mineral extraction.

09 BIOMASS FUELS↗

Deciphering the source of primary biological aerosol particles: a pollen case study

Primary biological aerosol particles (PBAPs) are microscopic solids suspended in the atmosphere emitted by biological systems and play critical roles in the atmosphere and at the atmosphere-biosphere interface, impacting human health, climate, and the ecosystem function. Understanding the sources of PBAPs is necessary to decipher the mechanistic interactions between aerosols, climate, and distinct ecosystem components. However, the detection of specific PBAPs in complex ambient aerosol samples is challenging. We performed metabolomics analyses of pollen from three pollinating tree species and ambient samples collected during the peak pollination period of each species. Random Forest and sPLS-DA machine learning methods were employed to evaluate whether metabolic signatures of ambient samples can reveal the source of the main pollen particles present in the atmosphere. Our results suggest that atmospheric eco-metabolomics techniques combined with sophisticated statistical methods can decipher the origin of abundant PBAPs from complex ambient samples. Developing complete libraries containing high-resolution metabolomic fingerprints of the major PBAPs present in the atmosphere would significantly advance future research to accurately understand the role of PBAPs in the atmosphere, ecosystems and human health.

Rivas-Ubach, Albert↗

Barium stars as tracers of s -process nucleosynthesis in AGB stars

Barium (Ba) stars help to verify asymptotic giant branch (AGB) star nucleosynthesis models since they experienced pollution from an AGB binary companion and thus their spectra carry the signatures of the slow neutron capture process (s process). For a large number (180) of Ba stars, we searched for AGB stellar models that match the observed abundance patterns. We aim to uncover any systematic deviations of the sample abundances from the predictions of the nucleosynthesis models. We employed three machine learning algorithms as classifiers: a Random Forest method, developed for this work, and the two classifiers used in our previous study. Compared to that work, we also expanded our observational sample with 11 Ba stars available in the supersolar metallicity range. We studied the statistical behaviour of the different s-process elements in the observational sample to investigate if the AGB models systematically under- or overpredict the abundances observed in the Ba stars and show the results in the form of violin plots of the residuals between spectroscopic abundances and model predictions. We inspected the correlations between the observed [Fe/H], the s-process elemental abundances, and the residuals. We employed the [Zr/Fe] and [Nb/Fe] abundances as a thermometer to constrain the operational temperature that rules the production of these elements in the sample stars, assuming a steady-state s process. We also investigated the mass distribution of the identified polluter AGB stars and the behaviour of the δ parameter, which describes the fraction of accreted AGB material relative to the Ba star envelope. We find a significant trend in the residuals that implies an underproduction of the elements just after the first s-process peak (Nb, Mo, and Ru) in the models relative to the observations. This may originate from a neutron-capture process (e.g. the intermediate neutron-capture process, i process) not yet included in the AGB models of metallicity from solar to roughly 1/5 solar, corresponding to the range of the Ba stars. Correlations are found between the residuals of these peculiar elements, suggesting a common origin for the deviations from the models. In addition, there is a weak metallicity dependence of the residuals of these elements. The s-process temperatures derived with the [Zr/Fe] – [Nb/Fe] thermometer have an unrealistic value for the majority of our stars. The most likely explanation is that at least a fraction of these elements are not produced in a steady-state s process, and instead may be due to processes not included in the AGB models. The mass distribution of the identified models confirms that our sample of Ba stars was polluted by low-mass AGB stars (< 4 M ⊙ ). Most of the matching AGB models require low accreted mass, but a few systems with high accreted mass are needed to explain the observations.

79 ASTRONOMY AND ASTROPHYSICS↗

How can a diverse set of integral and semi-integral measurements inform identification of discrepant nuclear data?

Nuclear data are used for a variety of applications, including criticality safety, reactor performance, and material safeguards. Despite the breadth of use-cases, the effective neutron multiplication factor, keff, of ICSBEP critical assemblies are primarily used for nuclear data validation; these are sensitive to specific energy regions and nuclides and are unable to uniquely constrain nuclear data. As a consequence, general-purpose nuclear data libraries, such as ENDF/B-VIII.0, may have deficiencies that, while not apparent in criticality applications, negatively impact other applications, such as non-destructive analysis of special nuclear material and neutron diagnosed subcritical experiments. Recent work by the Experiments Underpinned by Computational Learning for Improvements in Nuclear Data (EUCLID) project developed a machine learning tool, RAFIEKI, which uses random forests and the SHAP metric to determine which nuclear data contribute most to predicted bias between measured and simulated responses (e.g. keff). This paper contrasts RAFIEKI analysis applied to keff only against RAFIEKI analysis with keff paired with either LLNL pulsed sphere measurements or subcritical benchmarks. Two examples show that a) including pulsed sphere measurements substantially increases 9Be nuclear data importance to bias between 2 and 15 MeV, and b) including subcritical benchmarks has the potential for disentangling compensating errors between 240Pu (n,el) and (n,il) cross-sections between 0.1 and 10 MeV. These results show that RAFIEKI analysis applied to response sets that include, but go beyond, keff can aid nuclear data evaluators in identifying issues in nuclear data.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Using machine learning to identify extragalactic globular cluster candidates from ground-based photometric surveys of M87

Globular clusters (GCs) have been at the heart of many longstanding questions in many sub-fields of astronomy and, as such, systematic identification of GCs in external galaxies has immense impacts. In this study, we take advantage of M87’s well-studied GC system to implement supervised machine learning (ML) classification algorithms – specifically random forest and neural networks – to identify GCs from foreground stars and background galaxies, using ground-based photometry from the Canada–France–Hawaii Telescope (CFHT). We compare these two ML classification methods to studies of ‘human-selected’ GCs and find that the best-performing random forest model can reselect 61.2 per cent ± 8.0 per cent of GCs selected from HST data (ACSVCS) and the best-performing neural network model reselects 95.0 per cent ± 3.4 per cent. When compared to human-classified GCs and contaminants selected from CFHT data – independent of our training data – the best-performing random forest model can correctly classify 91.0 per cent ± 1.2 per cent and the best-performing neural network model can correctly classify 57.3 per cent ± 1.1 per cent. ML methods in astronomy have been receiving much interest as Vera C. Rubin Observatory prepares for first light. The observables in this study are selected to be directly comparable to early Rubin Observatory data and the prospects for running ML algorithms on the upcoming data set yields promising results.

79 ASTRONOMY AND ASTROPHYSICS↗

A Machine Learning-Based Vulnerability Analysis for Cascading Failures of Integrated Power-Gas Systems

This article proposes a cascading failure simulation (CFS) method and a hybrid machine learning method for vulnerability analysis of integrated power-gas systems (IPGSs). The CFS method is designed to study the propagating process of cascading failures between the two systems, generating data for machine learning with initial states randomly sampled. The proposed method considers generator and gas well ramping, transmission line and gas pipeline tripping, island issue handling and load shedding strategies. Then, a hybrid machine learning model with a combined random forest (RF) classification and regression algorithms is proposed to investigate the impact of random initial states on the vulnerability metrics of IPGSs. Extensive case studies are carried out on three test IPGSs to verify the proposed models and algorithms. Simulation results show that the proposed models and algorithms can achieve high accuracy for the vulnerability analysis of IPGSs.

24 POWER TRANSMISSION AND DISTRIBUTION↗