Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Random forests”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13

Artificial intelligence driven laser parameter search: Inverse design of photonic surfaces using greedy surrogate-based optimization

Photonic surfaces designed with specific optical characteristics are becoming increasingly crucial for novel energy harvesting and storage systems. The design of these surfaces can be achieved by texturing materials using lasers. The optimal adjustment of laser fabrication parameters to achieve target surface optical properties is an open challenge. Thus, we develop a surrogate-based optimization approach. Our framework employs the Random Forest algorithm to model the forward relationship between the laser fabrication parameters and the resulting optical characteristics. During the optimization process, we use a greedy, prediction-based exploration strategy that iteratively selects batches of laser parameters to be used in experimentation by minimizing the predicted discrepancy between the surrogate model’s outputs and the user-defined target optical characteristics. This strategy allows for efficient identification of optimal fabrication parameters without the need to model the error landscape directly. We demonstrate the efficiency and effectiveness of our approach on two synthetic benchmarks and two specific experimental applications of photonic surface inverse design targets. By calculating the average performance of our algorithm compared to other state of the art optimization methods, we show that our algorithm performs, on average, twice as well across all benchmarks. Additionally, a warm starting inverse design technique for changed target optical characteristics enhances the performance of the introduced approach.

97 MATHEMATICS AND COMPUTING↗

Machine learning enables identification of an alternative yeast galactose utilization pathway

How genomic differences contribute to phenotypic differences is a major question in biology. The recently characterized genomes, isolation environments, and qualitative patterns of growth on 122 sources and conditions of 1,154 strains from 1,049 fungal species (nearly all known) in the yeast subphylum Saccharomycotina provide a powerful, yet complex, dataset for addressing this question. We used a random forest algorithm trained on these genomic, metabolic, and environmental data to predict growth on several carbon sources with high accuracy. Known structural genes involved in assimilation of these sources and presence/absence patterns of growth in other sources were important features contributing to prediction accuracy. By further examining growth on galactose, we found that it can be predicted with high accuracy from either genomic (92.2%) or growth data (82.6%) but not from isolation environment data (65.6%). Prediction accuracy was even higher (93.3%) when we combined genomic and growth data. After the GALactose utilization genes, the most important feature for predicting growth on galactose was growth on galactitol, raising the hypothesis that several species in two orders, Serinales and Pichiales (containing the emerging pathogen Candida auris and the genus Ogataea, respectively), have an alternative galactose utilization pathway because they lack the GAL genes. Growth and biochemical assays confirmed that several of these species utilize galactose through an alternative oxidoreductive D-galactose pathway, rather than the canonical GAL pathway. Machine learning approaches are powerful for investigating the evolution of the yeast genotype–phenotype map, and their application will uncover novel biology, even in well-studied traits.

59 BASIC BIOLOGICAL SCIENCES↗

A Machine Learning Approach to Quantitative Analysis of Enamel Microstructure from Scanning Electron Microscopy Images

Dental enamel, the outermost tissue of mammalian teeth, must withstand a lifetime of wear and cyclic contact. To meet this demand, enamel possesses a combination of high hardness and resistance to fracture, properties that are typically mutually exclusive. The impressive damage tolerance has been attributed largely to decussation of the enamel rods, the principal unit of its microstructure. As such, enamel is inspiring the design of next‐generation structural materials. However, quantitative descriptions of the decussated enamel rod microstructure remain limited due to challenges encountered in applying computed tomography and in acquiring quality images appropriate for traditional digital processing methods. Here, a machine learning segmentation method is applied to images of the enamel obtained using scanning electron microscopy to support quantitative analysis of the microstructure. A pretrained convolutional neural network is used to expand the input training image dataset to allow the training of a random forest classifier, which ultimately segments the image with a very small training set ( n = 3 images). A validation of this segmentation method is presented, in addition to its application to calculate relevant microstructural parameters for images of tooth enamel from selected mammalian species. The methodology applied here is equally applicable to other hard tissues.

36 MATERIALS SCIENCE↗

Model Inputs, Outputs, and Scripts associated with: “Combined effects of stream hydrology and land use on basin-scale hyporheic zone denitrification in the Columbia River Basin”

This data package is associated with the publication “Combined effects of stream hydrology and land use on basin‐scale hyporheic zone denitrification in the Columbia River Basin”, published in Water Resource Research (Son et al.2022) available at https://doi.org/10.1029/2021WR031131. This data package includes the key model inputs/outputs of the river corridor model for the Columbia River Basin (CRB) and the model source codes used in the manuscript. The model is a carbon-nitrogen-coupled river corridor model (RCM), and the model is used to quantify hyporheic zone (HZ) denitrification at the NHDPLUS stream reach scales. The RCM used in this study combines empirical substrate models derived from observations and three microbially driven reactions, including two-step denitrification and aerobic respiration, are considered within the HZ. The key input data of the model are exchange flux, residence time, and stream solute (dissolved organic carbon (DOC), dissolved oxygen (DO), and nitrate concentrations). These inputs are constant over time and represent long-term averaged values. This study uses the RCM to explore the spatial patterns of HZ denitrification across reaches with different sizes and land use in the CRB. Our main objective is to use the RCM as a virtual reality model, and the machine-learning models as surrogates that encapsulate the complexities of the physics-based model while identifying the importance of different variables that are not evident in the model conceptualization. We do not include a direct comparison of the modeled HZ denitrification and measurements; however, the RCM can capture the overall spatial patterns of the HZ denitrification because the model inputs and its reaction networks are based on well-established theory and a physical-based model. The combination of the model-based predictions and a machine-learning approach (e.g., random forest) is used to improve our understanding of what variables of the model are associated with spatial patterns of the modeled denitrification across reaches with different sizes and land uses, and to develop a proxy model using measurable variables to reproduce the simulated patterns.This dataset contains five folders: (1) model_inputs, (2) model_outputs, (3) Rscripts, (4) figures, and (5) model_codes. It also contains a readme, file level metadata (FLMD), and data dictionary (dd). Please see the FLMD for a list of all the files contained in this data package and descriptions for each. The model_inputs folder contains the model inputs used to drive the model simulations. The model_outputs folder contains key model output files from the river corridor model. The Rscripts folder contains the Rscripts for pre- and post- processing model results. The figures folder contains the raw figures associated with the manuscript. The model_codes folder includes key model source codes/input files. All files are .jpg, .jpeg, .out, .e, .od, .dat, .sub, .F90, .0, .R, .sbx, .cpg, .sbn, .shx, .shp, .dbf, .prj, .tfw, .tif, .xml, .pdf, or .csv.

54 ENVIRONMENTAL SCIENCES↗

Employing Machine Learning for New Particle Formation Identification and Mechanistic Analysis: Insights From a Six‐Year Observational Study in the Southern Great Plains

We present a supervised machine learning (ML) framework to automatically identify new particle formation (NPF) events and analyze key atmospheric factors associated with their occurrence and growth. We applied ML to detect NPF events using start time and particle concentrations across size ranges, while identifying atmospheric variables including ambient temperature, relative humidity, solar radiation intensity (SRI), wind speed, wind direction, boundary layer height, total organics, sulfate, nitrate, total surface area concentration, sulfur dioxide, and turbulent kinetic energy (TKE). We analyzed a 6-year data set from the Atmospheric Radiation Measurement at the Southern Great Plains (SGP) site in Oklahoma, USA. Using long-term ground-based measurements, we identified NPF events and applied Random Forest Classifiers, which achieved 90%–95% prediction accuracy. Feature importance analysis highlighted SRI, relative humidity, and ambient temperature as the most influential variables, contributing normalized importances of 28%, 17%, and 10%. Partial Dependence Plots (PDPs) indicated that higher SRI and lower relative humidity were critical in promoting NPF formation at SGP. Seasonally, NPF events were more frequent in winter (42.1%) and spring (35.5%), and least in summer (4.0%). Particle growth rates also exhibited a seasonal variation, with the lowest in winter (below 2 nm hr −1 ) and highest in late spring and early summer (exceeding 5 nm hr −1 ). Temperature, turbulent kinetic energy, and aerosol properties were the primary factors of growth rate variability. This study advances predictive modeling of NPF, offers insights for future campaign deployments, and demonstrates the effectiveness of ML in understanding the formation and growth of atmospheric aerosols.

54 ENVIRONMENTAL SCIENCES↗

Machine Learning Downscaling of SoilMERGE in the United States Southern Great Plains

SoilMERGE (SMERGE) is a root-zone soil moisture (RZSM) product that covers the entire continental United States and spans 1978 to 2019. Machine learning techniques, Random Forest (RF), eXtreme Gradient Boosting (XGBoost), and Gradient Boost (GBoost) downscaled SMERGE to spatial resolutions straddling the field scale domain (100 to 3000 m). Study area was northern Oklahoma and southern Kansas. The coarse resolution of SMERGE (0.125 degree) limits this product’s utility. To validate downscaled results in situ data from four sources were used that included: United States Department of Energy Atmospheric Radiation Measurement (ARM) observatory, United States Climate Reference Network (USCRN), Soil Climate Analysis Network (SCAN), and Soil moisture Sensing Controller and oPtimal Estimator (SoilSCAPE). In addition, RZSM retrievals from NASA’s Airborne Microwave Observatory of Subcanopy and Surface (AirMOSS) campaign provided a nearly spatially continuous comparison. Three periods were examined: era 1 (2016 to 2019), era 2 (2012 to 2015), and era 3 (2003 to 2007). During eras 1 and 2, RF outperformed XGBoost and GBoost, whereas during era 3 no model dominated. Performance was better during eras 1 and 2 as opposed to the pre-L band era 3. Improvements across all eras, regions, and models realized from downscaling included an increase in correlation from 0.03 to 0.42 and a decrease in ub RMSE from -0.0005 to -0.0118 m 3 /m 3 . This study demonstrates the feasibility of SMERGE downscaling opening the prospect for the development of a long-term RZSM dataset at a more desirable field-scale resolution with the potential to support diverse hydrometeorological and agricultural applications.

54 ENVIRONMENTAL SCIENCES↗

Uncertainty-Aware Machine Learning for Small-Angle X-ray Scattering Analysis in Autonomous Experimentation

Small-angle X-ray scattering (SAXS) is a powerful high-throughput characterization tool for probing nanoscale structure in native sample environments, providing real-time morphological information such as nanoparticle size and shape during synthesis. However, automated SAXS data analysis for extracting meaningful structural parameters is non-trivial and remains a bottleneck in closed-loop experimentation towards autonomous materials discovery, which demands fast, reliable, and uncertainty-aware data analysis. Here, we develop a machine-learning approach for automated SAXS analysis tailored to closed-loop nanoparticle synthesis. A Random Forest (RF) regression model is trained on 100,000 synthetic SAXS curves generated from polydisperse spherical nanoparticles with realistic background contributions. Using normalized one-dimensional SAXS intensity profiles as input, the RF model directly predicts nanoparticle radius, size polydispersity, and background parameters, while the ensemble standard deviation across trees provides built-in uncertainty quantification (UQ). On synthetic data, we show that combining fit-quality metrics (R 2 , MAE) with thresholds on prediction uncertainty reliably identifies accurate parameter estimates without access to ground truth. We then apply the trained model to 365 experimental SAXS profiles of citrate-reduced gold nanoparticles synthesized using an automated droplet-flow microreactor with in situ SAXS at a synchrotron beamline, classifying the results into high- and low-confidence subsets based on UQ metrics. Finally, we integrate RF-based SAXS analysis into a simulated closed-loop optimization campaign using Gaussian process Bayesian optimization to minimize nanoparticle polydispersity, benchmarking against conventional automated Levenberg–Marquardt fitting. The RF-guided campaign exhibits substantially faster convergence and lower relative opportunity cost (∼0.07 vs ∼0.3), demonstrating that uncertainty-aware machine-learning SAXS analysis significantly enhances the efficiency and robustness of autonomous nanomaterials synthesis workflows.

Bayesian optimization↗

Modeling Spatial Distribution of Snow Water Equivalent by Combining Meteorological and Satellite Data with Lidar Maps

Abstract An accurate characterization of the water content of snowpack, or snow water equivalent (SWE), is necessary to quantify water availability and constrain hydrologic and land surface models. Recently, airborne observations (e.g., lidar) have emerged as a promising method to accurately quantify SWE at high resolutions (scales of ∼100 m and finer). However, the frequency of these observations is very low, typically once or twice per season in the Rocky Mountains of Colorado. Here, we present a machine learning framework that is based on random forests to model temporally sparse lidar-derived SWE, enabling estimation of SWE at unmapped time points. We approximated the physical processes governing snow accumulation and melt as well as snow characteristics by obtaining 15 different variables from gridded estimates of precipitation, temperature, surface reflectance, elevation, and canopy. Results showed that, in the Rocky Mountains of Colorado, our framework is capable of modeling SWE with a higher accuracy when compared with estimates generated by the Snow Data Assimilation System (SNODAS). The mean value of the coefficient of determination R 2 using our approach was 0.57, and the root-mean-square error (RMSE) was 13 cm, which was a significant improvement over SNODAS (mean R 2 = 0.13; RMSE = 20 cm). We explored the relative importance of the input variables and observed that, at the spatial resolution of 800 m, meteorological variables are more important drivers of predictive accuracy than surface variables that characterize the properties of snow on the ground. This research provides a framework to expand the applicability of lidar-derived SWE to unmapped time points. Significance Statement Snowpack is the main source of freshwater for close to 2 billion people globally and needs to be estimated accurately. Mountainous snowpack is highly variable and is challenging to quantify. Recently, lidar technology has been employed to observe snow in great detail, but it is costly and can only be used sparingly. To counter that, we use machine learning to estimate snowpack when lidar data are not available. We approximate the processes that govern snowpack by incorporating meteorological and satellite data. We found that variables associated with precipitation and temperature have more predictive power than variables that characterize snowpack properties. Our work helps to improve snowpack estimation, which is critical for sustainable management of water resources.

54 ENVIRONMENTAL SCIENCES↗

Data and Scripts associated with a manuscript on ecosystem responses to wildfires in the Columbia River Basin

This data package is associated with the publication “Ecosystem leaf area, gross primary production, and evapotranspiration responses to wildfire in the Columbia River Basin” submitted to Biogeosciences (Shi et al., 2024; doi: 10.22541/au.171053013.30286044/v1). In this research, data products, leaf area index (LAI), gross primary production (GPP), and evapotranspiration (ET), from the Moderate Resolution Imaging Spectroradiometer (MODIS) are used to quantify the resistance and resilience of different ecosystem types in the Columbia River Basin (CRB). A machine learning algorithm, random forest (RF), was used to examine the impacts of precipitation, vapor pressure deficit (VPD), and burn severity from Monitoring Trends in Burn Severity (MTBS) on ecosystem resilience. The data package includes the processed MODIS data products, precipitation, VPD, and burn severity in 138 fire regions in CRB and the input files for RF model training. This data package includes six folders. The MODIS products are included in three MODIS_* folders with shell scripts for data clipping and *ncl files for data processing: (1) “/MODIS_LAI_CRB”; (2) “/MODIS_GPP_CRB”; and (3) “/MODIS_ET_CRB”. All the processed data for each fire event are NetCDF formatted. The MTBS burn severity data and the shell and *ncl scripts used for data processing are in the folder named (4) “MTBS_fire”. The ERA meteorological fields and the data processing scritps are in (5) “ERA_Var_CR”. All the scripts for figure development are in the format of *ncl and in the folder (6) “paper_scripts”. See the file ending in “flmd.csv” for a list of all files contained in this data package and descriptions for each. Tabular column headers and units are described in the data dictionary file ending in “dd.csv”.

54 ENVIRONMENTAL SCIENCES↗

Upscaling Wetland Methane Emissions From the FLUXNET–CH4 Eddy Covariance Network (UpCH4 v1.0): Model Development, Network Assessment, and Budget Comparison

Wetlands are responsible for 20%–31% of global methane (CH 4 ) emissions and account for a large source of uncertainty in the global CH 4 budget. Data-driven upscaling of CH 4 fluxes from eddy covariance measurements can provide new and independent bottom-up estimates of wetland CH 4 emissions. Here, we develop a six-predictor random forest upscaling model (UpCH4), trained on 119 site-years of eddy covariance CH 4 flux data from 43 freshwater wetland sites in the FLUXNET-CH4 Community Product. Network patterns in site-level annual means and mean seasonal cycles of CH 4 fluxes were reproduced accurately in tundra, boreal, and temperate regions (Nash-Sutcliffe Efficiency ~0.52–0.63 and 0.53). UpCH4 estimated annual global wetland CH 4 emissions of 146 ± 43 TgCH 4 y –1 for 2001–2018 which agrees closely with current bottom-up land surface models (102–181 TgCH 4 y –1 ) and overlaps with top-down atmospheric inversion models (155–200 TgCH 4 y –1 ). However, UpCH4 diverged from both types of models in the spatial pattern and seasonal dynamics of tropical wetland emissions. We conclude that upscaling of eddy covariance CH 4 fluxes has the potential to produce realistic extra-tropical wetland CH 4 emissions estimates which will improve with more flux data. To reduce uncertainty in upscaled estimates, researchers could prioritize new wetland flux sites along humid-to-arid tropical climate gradients, from major rainforest basins (Congo, Amazon, and SE Asia), into monsoon (Bangladesh and India) and savannah regions (African Sahel) and be paired with improved knowledge of wetland extent seasonal dynamics in these regions.

54 ENVIRONMENTAL SCIENCES↗

Identifying dominant environmental predictors of freshwater wetland methane fluxes across diurnal to seasonal time scales

While wetlands are the largest natural source of methane (CH 4 ) to the atmosphere, they represent a large source of uncertainty in the global CH 4 budget due to the complex biogeochemical controls on CH 4 dynamics. Here we present, to our knowledge, the first multi-site synthesis of how predictors of CH 4 fluxes (FCH4) in freshwater wetlands vary across wetland types at diel, multiday (synoptic), and seasonal time scales. In this work we used several statistical approaches (correlation analysis, generalized additive modeling, mutual information, and random forests) in a wavelet-based multi-resolution framework to assess the importance of environmental predictors, nonlinearities and lags on FCH4 across 23 eddy covariance sites. Seasonally, soil and air temperature were dominant predictors of FCH4 at sites with smaller seasonal variation in water table depth (WTD). In contrast, WTD was the dominant predictor for wetlands with smaller variations in temperature (e.g., seasonal tropical/subtropical wetlands). Changes in seasonal FCH4 lagged fluctuations in WTD by ~17 ± 11 days, and lagged air and soil temperature by median values of 8 ± 16 and 5 ± 15 days, respectively. Temperature and WTD were also dominant predictors at the multiday scale. Atmospheric pressure (PA) was another important multiday scale predictor for peat-dominated sites, with drops in PA coinciding with synchronous releases of CH 4 . At the diel scale, synchronous relationships with latent heat flux and vapor pressure deficit suggest that physical processes controlling evaporation and boundary layer mixing exert similar controls on CH 4 volatilization, and suggest the influence of pressurized ventilation in aerenchymatous vegetation. In addition, 1- to 4-h lagged relationships with ecosystem photosynthesis indicate recent carbon substrates, such as root exudates, may also control FCH4. By addressing issues of scale, asynchrony, and nonlinearity, this work improves understanding of the predictors and timing of wetland FCH4 that can inform future studies and models, and help constrain wetland CH 4 emissions.

59 BASIC BIOLOGICAL SCIENCES↗

Mapping Vegetation at Species Level with High-Resolution Multispectral and Lidar Data Over a Large Spatial Area: A Case Study with Kudzu

Mapping vegetation species is critical to facilitate related quantitative assessment, and mapping invasive plants is important to enhance monitoring and management activities. Integrating high-resolution multispectral remote-sensing (RS) images and lidar (light detection and ranging) point clouds can provide robust features for vegetation mapping. However, using multiple sources of high-resolution RS data for vegetation mapping on a large spatial scale can be both computationally and sampling intensive. Here, we designed a two-step classification workflow to potentially decrease computational cost and sampling effort and to increase classification accuracy by integrating multispectral and lidar data in order to derive spectral, textural, and structural features for mapping target vegetation species. We used this workflow to classify kudzu, an aggressive invasive vine, in the entire Knox County (1362 km2) of Tennessee (U.S.). Object-based image analysis was conducted in the workflow. The first-step classification used 320 kudzu samples and extensive, coarsely labeled samples (based on national land cover) to generate an overprediction map of kudzu using random forest (RF). For the second step, 350 samples were randomly extracted from the overpredicted kudzu and labeled manually for the final prediction using RF and support vector machine (SVM). Computationally intensive features were only used for the second-step classification. SVM had constantly better accuracy than RF, and the producer’s accuracy, user’s accuracy, and Kappa for the SVM model on kudzu were 0.94, 0.96, and 0.90, respectively. SVM predicted 1010 kudzu patches covering 1.29 km2 in Knox County. We found the sample size of kudzu used for algorithm training impacted the accuracy and number of kudzu predicted. The proposed workflow could also improve sampling efficiency and specificity. Our workflow had much higher accuracy than the traditional method conducted in this research, and could be easily implemented to map kudzu in other regions as well as map other vegetation species.

59 BASIC BIOLOGICAL SCIENCES↗

A Machine Learning-Based Vulnerability Analysis for Cascading Failures of Integrated Power-Gas Systems

This article proposes a cascading failure simulation (CFS) method and a hybrid machine learning method for vulnerability analysis of integrated power-gas systems (IPGSs). The CFS method is designed to study the propagating process of cascading failures between the two systems, generating data for machine learning with initial states randomly sampled. The proposed method considers generator and gas well ramping, transmission line and gas pipeline tripping, island issue handling and load shedding strategies. Then, a hybrid machine learning model with a combined random forest (RF) classification and regression algorithms is proposed to investigate the impact of random initial states on the vulnerability metrics of IPGSs. Extensive case studies are carried out on three test IPGSs to verify the proposed models and algorithms. Simulation results show that the proposed models and algorithms can achieve high accuracy for the vulnerability analysis of IPGSs.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Functionally Assembled Terrestrial Ecosystem Simulator (FATES) for Hurricane Disturbance and Recovery

Tropical cyclones are an important cause of forest disturbance, and major storms caused severe structural damage and elevated tree mortality in coastal tropical forests. Model capabilities that can be used to understand post-hurricane forest recovery are still limited. We use a vegetation demography model, the Functionally Assembled Terrestrial Ecosystem Simulator, coupled with the Energy Exascale Earth System Model Land Model (ELM-FATES) to study the processes and the key factors regulating post-hurricane forest recovery. We implemented hurricane-induced forest damage, including defoliation, structural biomass reduction, and tree mortality, performed ensemble model simulations, and used random forest feature importance. For the simulation in the Luquillo Experimental Forest, Puerto Rico, we identified factors controlling the post-hurricane forest recovery, and quantified the sensitivity of key model parameters to the post-hurricane forest recovery. The results indicate a tendency for the Bisley forests to shift toward the light demanding plant functional type (PFT) when the pre-hurricane biomass between the light demanding and shade tolerant PFTs is nearly equal and forests experience hurricane disturbance with mortality >60% for both the two PFTs. Under more realistic conditions where the shade tolerant PFT is initially dominant, mortality >80% is required for a shift toward dominance of the light demanding PFT at Bisley. Hurricane mortality and background mortality are the two major factors regulating post-hurricane forest recovery in simulations. This research improves understanding of the ELM-FATES model behavior associated with hurricane disturbance and provides guidance for dynamic vegetation model development in representing hurricane induced forest damage with varied intensities.

54 ENVIRONMENTAL SCIENCES↗

Structure–activity relationship-based chemical classification of highly imbalanced Tox21 datasets

Abstract The specificity of toxicant-target biomolecule interactions lends to the very imbalanced nature of many toxicity datasets, causing poor performance in Structure–Activity Relationship (SAR)-based chemical classification. Undersampling and oversampling are representative techniques for handling such an imbalance challenge. However, removing inactive chemical compound instances from the majority class using an undersampling technique can result in information loss, whereas increasing active toxicant instances in the minority class by interpolation tends to introduce artificial minority instances that often cross into the majority class space, giving rise to class overlapping and a higher false prediction rate. In this study, in order to improve the prediction accuracy of imbalanced learning, we employed SMOTEENN, a combination of Synthetic Minority Over-sampling Technique (SMOTE) and Edited Nearest Neighbor (ENN) algorithms, to oversample the minority class by creating synthetic samples, followed by cleaning the mislabeled instances. We chose the highly imbalanced Tox21 dataset, which consisted of 12 in vitro bioassays for > 10,000 chemicals that were distributed unevenly between binary classes. With Random Forest (RF) as the base classifier and bagging as the ensemble strategy, we applied four hybrid learning methods, i.e., RF without imbalance handling (RF), RF with Random Undersampling (RUS), RF with SMOTE (SMO), and RF with SMOTEENN (SMN). The performance of the four learning methods was compared using nine evaluation metrics, among which F 1 score, Matthews correlation coefficient and Brier score provided a more consistent assessment of the overall performance across the 12 datasets. The Friedman’s aligned ranks test and the subsequent Bergmann-Hommel post hoc test showed that SMN significantly outperformed the other three methods. We also found that a strong negative correlation existed between the prediction accuracy and the imbalance ratio (IR), which is defined as the number of inactive compounds divided by the number of active compounds. SMN became less effective when IR exceeded a certain threshold (e.g., > 28). The ability to separate the few active compounds from the vast amounts of inactive ones is of great importance in computational toxicology. This work demonstrates that the performance of SAR-based, imbalanced chemical toxicity classification can be significantly improved through the use of data rebalancing.

Idakwo, Gabriel↗

Stand validation of lidar forest inventory modeling for a managed southern pine forest

We evaluated area-based approaches (ABAs) to light detection and ranging (lidar) predictions of plot- and stand-level forest attributes (tree count, height, basal area, volume, aboveground biomass, broadleaf/conifer, and diameter at breast height — “diameter”). ABA methods included post-stratification (PS), ordinary least squares (OLSs) regression, k nearest neighbors ( kNN), and random forest (RF). This study was conducted on the Savannah River Site in South Carolina, USA. Plot- and stand-level predictions were validated against fixed-radius 0.04 ha (0.1 acre) plots in 49 ≈2.0 ha (5 acre) stands. Our findings demonstrate that lidar can be incorporated operationally into forest inventory systems to provide stand-level inferences for a wide range of forest attributes. Volume predictions for specific diameter classes, however, often fared poorly (root mean squared error (RMSE) > 100%) for the methods we explored, especially for larger (less common) diameter trees. Stand-level results were consistently better than pixel-level results (10–200+ percentage points). kNN and RF performed similarly and better than OLS and PS, but RF was the most robust to model configurations, while kNN has practical advantages such as simultaneous predictions of many attributes.

Forestry↗

Forest Carbon Storage in the Western United States: Distribution, Drivers, and Trends

Abstract Forests are a large carbon sink and could serve as natural climate solutions that help moderate future warming. Thus, establishing forest carbon baselines is essential for tracking climate‐mitigation targets. Western US forests are natural climate solution hotspots but are profoundly threatened by drought and altered disturbance regimes. How these factors shape spatial patterns of carbon storage and carbon change over time is poorly resolved. Here, we estimate live and dead forest carbon density in 19 forested western US ecoregions with national inventory data (2005–2019) to determine: (a) current carbon distributions, (b) underpinning drivers, and (c) recent trends. Potential drivers of current carbon included harvest, wildfire, insect and disease, topography, and climate. Using random forests, we evaluated driver importance and relationships with current live and dead carbon within ecoregions. We assessed trends using linear models. Pacific Northwest (PNW) and Southwest (SW) ecoregions were most and least carbon dense, respectively. Climate was an important carbon driver in the SW and Lower Rockies. Fire reduced live and increased dead carbon, and was most important in the Upper Rockies and California. No ecoregion was unaffected by fire. Harvest and private ownership reduced carbon, particularly in the PNW. Since 2005, live carbon declined across much of the western US, likely from drought and fire. Carbon has increased in PNW ecoregions, likely recovering from past harvest, but recent record fire years may alter trajectories. Our results provide insight into western US forest carbon function and future vulnerabilities, which is vital for effective climate change mitigation strategies.

Environmental Sciences & Ecology↗

A decreasing carbon allocation to belowground autotrophic respiration in global forest ecosystems

Belowground autotrophic respiration (RAsoil) depends on carbohydrates from photosynthesis flowing to roots and rhizospheres, and is one of the most important but least understood components in forest carbon cycling. Carbon allocation plays an important role in forest carbon cycling and reflects forest adaptation to changing environmental conditions. However, carbon allocation to RAsoil has not been fully examined at the global scale. To fill this knowledge gap, first, the spatio-temporal patterns of RAsoil from 1981 to 2017 were predicted by a Random Forest (RF) algorithm using the most updated Global Soil Respiration Database (v5) with global environmental variables; second, carbon allocation from photosynthesis to RAsoil (CAB), was calculated as the ratio of RAsoil to gross primary production; and its temporal and spatial patterns were assessed in global forest ecosystems. . Globally, mean RAsoil from forests was 8.9 ± 0.08 Pg C yr-1 (mean ± standard deviation) from 1981 to 2017 with strong spatial variabilities. Temporally, RAsoil increased at a rate of 0.0059 Pg C yr-2, paralleling broader soil respiration changes and indicating increasing carbon respired by roots. Mean CAB was 0.243 ± 0.016 and decreased over time. The temporal trend of CAB varied greatly in space, reflecting uneven responses of CAB to environmental changes. This study is the first attempt to predict global CAB and analyze its temporal and spatial patterns. With the linkage of carbon use efficiency, the developed CAB offers an completely independent approach to quantify global aboveground autotropic respiration spatially and temporally, which could provide crucial insights into carbon flux partition and global carbon cycling under climate change.

Tang, Xiaolu↗