Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “random forest regression”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Upscaling Soil Organic Carbon Measurements at the Continental Scale Using Multivariate Clustering Analysis and Machine Learning

Abstract Estimates of soil organic carbon (SOC) stocks are essential for many environmental applications. However, significant inconsistencies exist in SOC stock estimates for the U.S. across current SOC maps. We propose a framework that combines unsupervised multivariate geographic clustering (MGC) and supervised Random Forests regression, improving SOC maps by capturing heterogeneous relationships with SOC drivers. We first used MGC to divide the U.S. into 20 SOC regions based on the similarity of covariates (soil biogeochemical, bioclimatic, biological, and physiographic variables). Subsequently, separate Random Forests models were trained for each SOC region, utilizing environmental covariates and SOC observations. Our estimated SOC stocks for the U.S. (52.6 ± 3.2 Pg for 0–30 cm and 108.3 ± 8.2 Pg for 0–100 cm depth) were within the range estimated by existing products like Harmonized World Soil Database, HWSD (46.7 Pg for 0–30 cm and 90.7 Pg for 0–100 cm depth) and SoilGrids 2.0 (45.7 Pg for 0–30 cm and 133.0 Pg for 0–100 cm depth). However, independent validation with soil profile data from the National Ecological Observatory Network showed that our approach ( R 2 = 0.51) outperformed the estimates obtained from Harmonized World Soil Database ( R 2 = 0.23) and SoilGrids 2.0 ( R 2 = 0.39) for the topsoil (0–30 cm). Uncertainty analysis (e.g., low representativeness and high coefficients of variation) identified regions requiring more measurements, such as Alaska and the deserts of the U.S. Southwest. Our approach effectively captures the heterogeneous relationships between widely available predictors and the current SOC baseline across regions, offering reliable SOC estimates at 1 km resolution for benchmarking Earth system models.

58 GEOSCIENCES↗

Application of machine learning to discover new intermetallic catalysts for the hydrogen evolution and the oxygen reduction reactions

The adsorption energies for hydrogen, oxygen, and hydroxyl were calculated by means of density functional theory on the lowest energy surface of 24 pure metals and 332 binary intermetallic compounds with stoichiometries AB, A 2 B, and A 3 B taking into account the effect of biaxial elastic strains. This information was used to train two random forest regression models, one for the hydrogen adsorption and another for the oxygen and hydroxyl adsorption, based on 9 descriptors that characterized the geometrical and chemical features of the adsorption site as well as the applied strain. All the descriptors for each compound in the models could be obtained from physico-chemical databases. The random forest models were used to predict the adsorption energy for hydrogen, oxygen, and hydroxyl of ≈2700 binary intermetallic compounds with stoichiometries AB, A 2 B, and A 3 B made of metallic elements, excluding those that were environmentally hazardous, radioactive, or toxic. This information was used to search for potential good catalysts for the HER and ORR from the criteria that their adsorption energy for H and O/OH, respectively, should be close to that of Pt. Further, this investigation shows that the suitably trained machine learning models can predict adsorption energies with an accuracy not far away from density functional theory calculations with minimum computational cost from descriptors that are readily available in physico-chemical databases for any compound. Moreover, the strategy presented in this paper can be easily extended to other compounds and catalytic reactions, and is expected to foster the use of ML methods in catalysis.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Sub-pilot-scale Production of High-Value Products from U.S. Coals

Investigators from the University of Utah, University of Wyoming and Marshall University pursued a program to study the conversion of raw coal to high-value products of carbon fiber and silicon carbide. Team members also developed an initial framework for a data portal that can incorporate laboratory data on coal processing and product quality, and also work with tools for machine learning for data analysis, data visualization and economic assessment. Experimental R&D efforts focused on the conversion of raw coal to coal tar and other byproducts, and the resulting tar intermediates were upgraded to form anisotropic and isotropic pitch materials. These pitch materials were produced from coal using both thermal (pyrolysis) and chemical (mild solvolysis liquefaction) decomposition of raw coal. Four different coals were studied: Utah bituminous coal (Sufco), Wyoming PRB coal (Black Thunder), Illinois bituminous coal (Illinois #6), and West Virginia bituminous coal (Flying Eagle). Both metallurgical-grade coking coals and lower-grade steam coals were investigated, and controlled secondary gas-phase reactions were used during a two-stage pyrolysis process to induce cracking and condensation reactions among the pyrolytic tar species. This approach successfully improved the performance of the lower grade coals for yielding pitch materials, with properties more consistent with a commercial-grade pitch that had previously demonstrated success for quality carbon fiber production. The use of waste plastic materials was also studied, to help improve physical and chemical characteristics of the intermediate tars and final pitch product; in particular, for lowering the pitch softening point to an acceptable level for melt spinning carbon fiber. Mild solvolysis liquefaction was also used as a method for producing pitch for carbon fiber production. As expected, significantly higher pitch yields were obtained using this approach, and waste plastic materials were also successfully used to reduce pitch softening point to an acceptable level. The plastic materials were also utilized to create a solvent for the mild solvolysis process, and this plastic-derived solvent was shown to provide results consistent with more expensive commercial chemical solvents, and could thus avoid the need for costly recovery and recycle of a liquefaction solvent. Additional experimental R&D focused on the production of silicon carbide (β-SiC) from the residual char byproduct from pitch production, and also on the production of carbon fiber from the anisotropic pitch. SiC was successfully synthesized using a mixture of residual char and sandstone at a ratio of 1:1. Reaction temperature and residence time were optimized and yielded a product purity of 81%. For carbon fiber production, the most successful pitch samples were obtained from the mild solvolysis liquefaction approach, combined with the use of a plastic (HDPE)-derived solvent. Fiber properties improved over time as laboratory fiber production methodologies improved, and final yields of carbon fiber were obtained with a diameter of 12.14 ± 1.10 um, Modulus of 173.73 ± 15.25 GPa, and Tensile Strength of 1.04 ± 0.10 GPa. A proof-of-concept Modern Community Research Data Portal (MCRDP) was developed and deployed for coal and coal-derived pitch characterization, with the full support of (i) remote web-based access, (ii) distributed analysis, (iii) interactive visualization and exploration, (iv) shared and long-term data access, (v) advanced query capabilities and (vi) real-time collaboration. The Coal to Products Data Portal “coaltoproducts.org” provides researchers with space to store and share data within a project, tools for analyzing and understanding data for scientific investigation, and the ability to publish data to the broader community for reproducibility. The portal leverages the Material Commons 2.0 (MC) platform developed by the Center for PRedictive Integrated Structural Materials Science (PRISMS) of the University of Michigan, to achieve long-term longevity of data collections and, more importantly, collaborative science. A number of data visualization tools were also assessed and implemented for interrogating the experimental and modeling data. The machine learning portion of this project analyzed datasets from two different coal conversion processes performed on a diverse set of coal samples from both the coal pyrolysis experiments and the solvent liquefaction experiments. The work was initiated by exploring standard regression models on the pyrolysis data, aiming to understand the impact of sample characteristics and processing conditions on key product metrics. Over the course of the project, the focus expanded to include a variety of machine learning tools, delving into both supervised and unsupervised learning methods. Models tested on the pyrolysis data included linear, ridge, lasso, elastic-net, Gaussian process, random forest regression, and AutoSklearn, and the approach was continually refined to enhance predictive accuracy and model interpretability. Similar techniques were applied to the liquefaction data with an additional focus on feature engineering. Along with mesophase content, additional outputs of interest were the pitch yield, softening point, and QI content. Insights derived from these analyses are crucial in determining the factors influencing the quality and yield of coal-derived products. As the work progressed, the research evolved from foundational model comparisons to analyses of random forests, decision paths, and feature importance scores. A thorough market analysis was performed to examine the prospects of coal-based carbon fibers. The best opportunities for coal come from its lower and more stable price relative to petroleum, particularly for subbituminous coals, which is the primary advantage that a coal refinery may have over a petroleum refinery. Before a commercial CTP production facility can be modeled, however, several things need to be understood regarding the nature of the would-be coal refinery. These include the technology to be deployed, the size of facility, the volume(s) of co-product(s), and the waste and emissions profile of the plant. The volume of co-products and waste may be substantial and will require separate market analysis to ensure viability. In the near-term, the importance of coal tar pitch, in the form of carbon pitch, to the aluminum and steel industries is likely to overshadow the alternative use of this material as an input for carbon fiber. The importance of steel and aluminum in building materials, and the need for carbon materials in their manufacturing, will ensure that demand for these products remains for the long run. In addition, carbon fiber may also be the best substitute for steel and aluminum well into the future. While society will eventually be able to shift production of much of its electricity needs to renewables, it will not be able to shift away from fossil fuels for production of high-strength construction and vehicular materials. Demand for carbon fiber is expected to increase quickly, but the volume of carbon fiber and the amount of coal that would be needed to produce even a sizeable share of this market may still be relatively small compared to current coal production. Thus, other coal-based products like graphene, graphite, carbon foams, resins, and carbon-based building products will play important roles in sustaining coal production as coal-fired power generation continues to decline.

01 COAL, LIGNITE, AND PEAT↗

A Centralized AI Lakehouse Framework for Brain Tumor MRI Classification and Segmentation, University KPI Forecasting, and Water Potability Prediction

In many university and healthcare projects, models are built for very different data types such as tables, institutional time series, and medical images, but they are deployed as separate applications. In this work, that separation made testing and maintenance difficult because each module had its own pipeline and runtime requirements. This paper presents an integrated AI lakehouse-style implementation that runs three model pipelines inside one containerized backend. For medical imaging, we used MRI datasets from IEEE DataPort: a four-class classification set with 7012 images (5708 train/1304 test) and a segmentation set with 3063 image–mask pairs. The classification model (ResNet50 transfer learning) is evaluated using a proper train–validation–test protocol across multiple splits (80/10/10, 70/10/20, 60/10/30, and 10/30/60), achieving a test accuracy of 99.00% under the standard 80/10/10 split. Additionally, a patient-level evaluation is conducted using an external glioma dataset to provide a more realistic assessment without data leakage. The segmentation model (DeepLabV3-ResNet50) achieved 83.09% validation mIoU and 88.79% Dice score. For university KPI forecasting, we used annual IPEDS and NSF HERD data from 2010 to 2023 for three universities (BSU, EOU, and UAB). To examine the effect of preprocessing on forecasting performance, two case studies are conducted. In the first case, linear interpolation is applied to generate semester-level data. In the second case, the original annual data is used directly without interpolation. Random Forest regression and ARIMA models are evaluated using MAE, RMSE, MAPE, and R 2 . The results showed that interpolation improved apparent forecasting performance due to smoothing, while evaluation on the original annual data provided a more realistic assessment of model behavior. To further validate the framework on a larger dataset, an additional case study is conducted using a student dropout dataset. For water potability, we trained and compared multiple tabular classifiers on a large dataset (1,048,575 samples). A Random Forest model (100 trees, max depth 10) achieved 85.86% test accuracy and high recall for unsafe samples (0.8447). All modules are served via FastAPI and deployed together using Docker, with workflow automation routing requests to the correct endpoint. System-level benchmarking indicates that the backend maintains stable throughput and latency under concurrent requests.

97 MATHEMATICS AND COMPUTING↗

Lithium-Ion Battery Diagnostics Using Electrochemical Impedance via Machine-Learning

Diagnosing battery states such as health, state-of-charge, or temperature is crucial for ensuring the safety and reliability of electrochemical energy storage systems. While some states, such as temperature, may be measured using cheap sensors, accurate diagnosis of battery health metrics usually requires time-consuming performance measurements, making them infeasible for use in real-world operation. These health metrics can be measured during lab-testing and then estimated on-line using predictive life models or via state observer algorithms such as Kalman filters, but these predictive methods should be supplemented by actual measurement of battery health whenever possible to ensure reliability. Rapid measurement of battery health may be done by various types of fast diagnostic techniques such as electrochemical impedance spectroscopy (EIS), which can be performed in only a few minutes and require only a fraction of the energy and power needed for a full charge and discharge measurement. But there is a substantial challenge for estimating battery health using EIS data, as EIS is sensitive to cell temperature, state-of-charge, current, and resting time in addition to health. Thus, utilizing EIS data to predict battery capacity requires correcting for all these additional variables, a task that is extremely difficult to handle analytically. This talk utilizes machine-learning methods to estimate the effectiveness of battery capacity prediction from EIS data, leveraging a data set of hundreds of EIS measurements recorded at varying temperature and state-of-charge throughout a 500-day aging study of 32 commercial, large-format NMC-Graphite lithium-ion batteries. Using EIS as input to machine-learning models is complicated by the nonlinear response of impedance to battery health, temperature, and state-of-charge, as well as the collinearity between the impedance response at neighboring frequencies, which can easily lead to overfit models. To train robust models, features from EIS data need to be extracted from the data or some subset of critical frequencies selected. Many approaches for extracting and selecting features from EIS data from electrochemical analysis and machine-learning fields were identified for analysis: using the entire raw spectra; selection of one, two, or many frequencies from the entire spectra; selecting interesting points from the EIS measurement using domain knowledge; fitting EIS with an equivalent-circuit model; calculating statistics on the raw impedance values; and reducing the dimensionality of the data using unsupervised linear (principal component analysis) and non-linear (uniform manifold approximation and projection) methods. These approaches were rigorously compared using a machine-learning pipeline approach, training linear, Gaussian process, and random forest regression models and quantifying performance using cross-validation as well as a held-out test set. An artificial neural network model trained on the raw spectra was also tested. Promising pipelines were fine-tuned via Bayesian hyperparameter optimization using cross-validation loss and training with class-specific weights to counter data set imbalance. The most reliable method for utilizing impedance in this work was the selection of two optimal frequencies through an exhaustive search, resulting in about 2% mean absolute error on test data for both Gaussian process and random forest model architectures. Interrogation of a variety of models reveals critical frequencies of 100 Hz and 103 Hz for this data set, though the optimal set of frequencies is not necessarily intuitive, i.e., the best performing models are not simply those that use impedance at frequencies that have the highest correlation to the relative discharge capacity. The best performing model is an ensemble model, which is able to predict battery capacity with 1.9% mean absolute error for unseen cells using impedance recorded at a variety of temperatures and states-of-charge.

battery↗

Estimating Fine-Resolution Shortwave Broadband Albedo of Croplands from Harmonized Landsat and Sentinel-2 Data

Altered surface albedo due to land-cover conversions and management is a significant driver of global climate change. Albedo can be directly measured at ground stations, and remote sensing data can be used to scale-up albedo values to regional and global levels. Some previous studies have retrieved fine-resolution (10–30 m) instantaneous albedo and coarse-resolution (500–1000 m) daily mean albedo from remote sensing data, but they all required the input of Moderate Resolution Imaging Spectroradiometer (MODIS) albedo information at 500-m resolution, and none have assembled both instantaneous and daily albedo based exclusively on fine-resolution satellite data. Here, to address this issue, we compiled 387 instantaneous and 346 daily albedo records using field net radiometer measurements from the bioenergy croplands at the W. K. Kellogg Biological Station in southwest Michigan. We then connected these albedo records with a suite of variables derived from harmonized Landsat and Sentinel-2 data through two machine learning algorithms (random forest regression and extreme gradient boosting) to retrieve clear-sky instantaneous and daily shortwave broadband albedo. The performance statistics indicate reasonable accuracy of model results [root-mean-square error (RMSE)] around or below 0.03 except for snow-covered surfaces), suggesting that the retrieval of both instantaneous and daily albedo based exclusively on fine-resolution satellite data is promising. To facilitate the use of fine-resolution albedo products at the global level, future efforts need to include more albedo records of diverse surface cover types, as well as to accurately model daily albedo for cloudy days to address the “clear-sky bias.”

Harmonized Landsat and Sentinel-2↗

Effects of spatial variability in vegetation phenology, climate, landcover, biodiversity, topography, and soil property on soil respiration across a coastal ecosystem

Coastal terrestrial-aquatic interfaces (TAIs) are crucial contributors to global biogeochemical cycles and carbon exchange. A systematic evaluation of the interaction between coastal catchment properties and carbon dioxide (CO2) emission by soil respiration is significant for assessing carbon dynamics and predicting the future trajectory of atmospheric CO2 concentrations in coastal TAIs. The soil CO2 efflux in these transition zones is however poorly understood due to the high spatiotemporal dynamics of TAIs, as various sub-ecosystems in this region are compressed and expanded by complex influences of tides, changes in river levels, climate, and land use. We focus on the Chesapeake Bay region to (i) investigate the spatial heterogeneity of the coastal ecosystem and identify spatial zones with similar environmental characteristics based on the spatial data layers, including vegetation index (kNDVI), climate, landcover, diversity, topography, soil property, and relative tidal elevation; (ii) understand the primary driving factors affecting soil respiration within sub-ecosystems of the coastal ecosystem. Specifically, we employed hierarchical clustering analysis to identify spatial regions with distinct environmental characteristics, followed by the determination of main driving factors using Random Forest regression and SHapley Additive exPlanations. Maximum and minimum temperature are the main drivers common to all sub-ecosystems, while each region also has additional unique major drivers that differentiate them from one another. Precipitation exerts an influence on vegetated lands, while soil pH value holds importance specifically in forested lands. In croplands characterized by high clay content and low sand content, the significant role is attributed to bulk density. Wetlands demonstrate the importance of both elevation and sand content, with clay content being more relevant in non-inundated wetlands than in inundated wetlands. The topographic wetness index significantly contributes to the mixed vegetation areas, including shrub, grass, pasture, and forest. Additionally, our research reveals that dense vegetation land covers and urban/developed areas exhibit distinct soil property drivers. Overall, there is no one-size-fits-all approach to modeling carbon fluxes in coastal TAIs, and our study highlights the importance of further research and monitoring practices to improve our understanding of carbon dynamics and promote the sustainable management of coastal TAIs.

54 ENVIRONMENTAL SCIENCES↗

Predictive links between microbial communities and biological oxygen utilization in the Arctic Ocean

Microbial metabolism influences rates of net community production (NCP), exerting a direct biological control on marine oxygen and carbon fluxes. In the Arctic, it is increasingly important to understand and quantify this process, as ecological and oceanographic conditions shift due to changing climate. Here, we describe potential ecological links between pelagic microbial diversity and an NCP precursor, biological oxygen utilization, using machine learning and paired observations of community structure and metabolic activity from a seasonally and spatially variable transect of the Arctic Ocean (2019–2020 MOSAiC Expedition). Community structure was determined using 16S (prokaryotic) and 18S (eukaryotic) rRNA gene amplicon sequencing, and metabolic activity was derived from ΔO 2 /Ar. Using self-organizing maps, we identified clear successional patterns in observed microbial community structure that were seasonally driven in the upper ocean and vertically stratified with depth. Metabolic activity was also stratified, with a primarily net heterotrophic water column (median −1.5% biological oxygen saturation), excepting periodic oxygen supersaturation (maximum: 13.6%) within the mixed layer. Using DNA sequences as predictor variables, we then constructed a random forest regression model that reliably reconstructed biological oxygen concentrations (root mean squared error = 4.14 μmol kg −1 ). Top predictors from this model were from heterotrophic (bacteria) or potentially mixotrophic (dinoflagellate) taxa. These analyses highlight biologically driven diagnostic tools that can be used to expand biogeochemical datasets and improve the microbial perspectives and metabolisms represented in ecological models of net productivity and carbon flux in a changing Arctic Ocean.

Chamberlain, Emelia J. [Univ. of San Diego, San Di↗

Field evaluation of semi‐automated moisture estimation from geophysics using machine learning

Geophysical methods can provide three-dimensional (3D), spatially continuous estimates of soil moisture. However, point-to-point comparisons of geophysical properties to measure soil moisture data are frequently unsatisfactory, resulting in geophysics being used for qualitative purposes only. This is because (1) geophysics requires models that relate geophysical signals to soil moisture, (2) geophysical methods have potential uncertainties resulting from smoothing and artifacts introduced from processing and inversion, and (3) results from multiple geophysical methods are not easily combined within a single soil moisture estimation framework. To investigate these potential limitations, an irrigation experiment was performed wherein soil moisture was monitored through time, and several surface geophysical datasets indirectly sensitive to soil moisture were collected before and after irrigation: ground penetrating radar, electrical resistivity tomography (ERT), and frequency domain electromagnetics (FDEM). Data were exported in both raw and processed form, and then snapped to a common 3D grid to facilitate moisture prediction by standard calibration techniques, multivariate regression, and machine learning. A combination of inverted ERT data, raw FDEM, and inverted FDEM data was most informative for predicting soil moisture using a random regression forest model (one-thousand 60/40 training/test cross-validation folds produced root mean squared errors ranging from 0.025–0.046 cm 3 /cm 3 ). This cross-validated model was further supported by a separate evaluation using a test set from a physically separate portion of the study area. Machine learning was conducive to a semi-automated model-selection process that could be used for other sites and datasets to locally improve accuracy.

54 ENVIRONMENTAL SCIENCES↗

Effects of forest structural and compositional change on forest microclimates across a gradient of disturbance severity

Forest structural diversity and community composition are key in regulating forest microclimates. When disturbance affects structural diversity or composition, forest microclimates may be altered due to changes in soil temperature, soil water content, and light availability. It is unclear however which structural or compositional components, when changed or to what extent, result in microclimatic change. To address this question, we used data from a large scale, manipulative stem-girdling experiment in northern, lower Michigan—the Forest Resilience and Threshold Experiment (FoRTE). FoRTE follows a factorial design with multiple levels of disturbance severity (0, 45, 65, 85%) based on targeted reductions in gross leaf area index via stem-girdling induced mortality. These disturbance severity treatments are applied in two ways: either as top-down (largest trees are killed) or bottom-up (small to medium trees killed) treatments. We examined how multiple components of structural diversity and community composition changed as a product of disturbance severity and type, and then tested for resulting effects on forest microclimates (light availability, soil temperature, and soil water), using a multivariate, Random Forest framework. We found that measures of community composition (species richness, species evenness, and Shannon-Wiener Diversity Index) and stand structure (basal area, standard deviation of DBH, tree size diversity) declined more following disturbance than did measures of canopy cover, heterogeneity, arrangement, or height. However, when changes in each variable from pre- to post-disturbance, measured as log change, were employed in a multivariate, Random Forest regression framework, structural diversity measures of heterogeneity (rugosity, top rugosity), cover (canopy cover), and arrangement (porosity) were the most influential variables, but with differences among bottom-up and top-down treatments We found that the death of large trees from disturbance impacts soil temperature, water, and light environments more substantially and uniformly across disturbance gradients than does the death of smaller trees. Furthermore, our results have implications for both statistical and process-based modeling of forest disturbance.

54 ENVIRONMENTAL SCIENCES↗

Evaluating the impact of wildfire smoke on solar photovoltaic production

There are growing needs to understand how extreme weather events impact the electrical grid. Renewable energy sources such as solar photovoltaics are expanding in use to help sustainably meet electricity demands. Wildfires and, notably, the widespread smoke resulting from them, are one such extreme event that can impair the performance of solar photovoltaics. However, isolating the impact that smoke has on photovoltaic energy production, separate from ambient conditions, can be difficult. In this work, we seek to understand and quantify the impacts of wildfire smoke on solar photovoltaic production within the Western United States. Our analysis focuses on the construction of a random forest regression model to predict overall solar photovoltaic production. The model is used to separate and quantify the impacts of wildfire smoke in particular. To do so, we fuse historical weather, solar photovoltaic energy production, and PM2.5 particulate matter (primary smoke pollutant) data to train and test our model. The additional weather data allows us to capture interactions between wildfire smoke and other ambient conditions, as well as to create a more powerful predictive model capable of better quantifying the impacts of wildfire smoke on its own. We find that solar PV energy production decreases 8.3% on average during high smoke days at PV sites as compared to similar conditions without smoke present. Finally, this work allows us to improve our understanding of the potential impact on photovoltaic-based energy production estimates due to wildfire events and can help inform grid and operational planning as solar photovoltaic penetration levels continue to grow.

14 SOLAR ENERGY↗

Spatiotemporal features of traffic help reduce automatic accident detection time

Quick and reliable automatic detection of traffic accidents is of paramount importance to save human lives in transportation systems. However, automatically detecting when accidents occur has proven challenging, and minimizing the time to detect accidents (TTDA) by using traditional features in machine learning (ML) classifiers has plateaued. We hypothesize that accidents affect traffic farther from the accident location than previously reported. Therefore, leveraging traffic signatures from neighboring sensors that are adjacent to accidents should help improve their detection. We confirm this hypothesis by using verified ground-truth accident data, traffic data from radar detection system sensors, and light and weather conditions and show that we can minimize the TTDA while maximizing classification performance by considering spatiotemporal features of traffic. Specifically, we compare the performance of different ML classifiers (i.e, logistic regression, random forest, and XGBoost) when controlling for different numbers of neighboring sensors and TTDA horizons. We use data from interstates 75 and 24 in the metropolitan area that surrounds Chattanooga, TN. Our results show that the XGBoost classifier produces the best results by detecting accidents as quickly as 1.0 min after their occurrence with an area under the receiver operating characteristic curve of up to 83% and an average precision of up to 49%. We describe limitations, open challenges, and how the proposed framework can be used for quicker operational accident detection.

33 ADVANCED PROPULSION SYSTEMS↗

Fundamental Insights into Cathode Stability: Linking Compositional Tuning and Local Coordination in Complex Metal Oxides under Aqueous Transformations

Compositional tuning of complex metal oxides in Li-ion battery materials influences their performance as well as their end-of-life behavior, in particular, the tendency to release toxic metal cations in aqueous solution. We modeled ternary variants of a parent LiCoO 2 delafossite structure by varying the metal identity and relative amounts. This yielded ten model formulations of Li(A 4/6 B 1/6 C 1/6 )O 2 , where the material is enriched with the A metal and doped with B and C, with Ni, Mn, Co, Fe, Al, V, and Ti as constituent metals. To assess their stability in aqueous conditions, metal release energetics were calculated using a combination of Density Functional Theory calculations and thermodynamics. Metal release in ternary oxides is dictated by subtle variations in the coordination environment of the leaving group. To identify governing chemical features across diverse compositions with varying local coordination environments, we leverage random forest regression and descriptor importance analysis. A key result is that metal–oxygen orbital hybridization, quantified using a projected density-of-states-derived descriptor, H d/p , provides a physically grounded measure of interaction strength that governs metal release energetics. This refined perspective goes beyond conventional oxidation state considerations and offers more robust insights for materials science. Finally, we model defect surface-bound O 2 dimer formation as a proxy for reactive oxygen species (ROS) generation. The results show that Ni-rich compositions more readily stabilize spin-polarized O 2 dimers, corroborating experimental reports of an increased ROS-driven biological response. In conclusion, our results establish a compositional and electronic basis for metal release and surface oxygen reactivity that form a rationale for complex metal oxide design principles.

36 MATERIALS SCIENCE↗

Ensemble Estimation of Historical Evapotranspiration for the Conterminous U.S.

Abstract Evapotranspiration (ET) is the largest component of the water budget, accounting for the majority of the water available from precipitation. ET is challenging to quantify because of the uncertainties associated with the many ET equations currently in use, and because observations of ET are uncertain and sparse. In this study, we combine information provided by available ET data and equations to produce a new monthly data set for ET for the conterminous U.S. (CONUS). These maps are produced from 1895 to 2018 at an 800 m spatial scale, marking a finer resolution than currently available products over this time period. In our approach, the relative performance of a suite of ET equations is assessed using water balance, flux tower, and remotely sensed ET estimates. At the observation locations, we use error distributions to quantify relative weights for the equations and use these in a modified Bayesian model averaging weighted ensemble approach. The relative weights are spatially generalized using a random forest regression, which is applied to wall‐to‐wall explanatory variable maps to generate CONUS‐wide relative weight maps and ensemble estimates. We assess the performance of the ensemble using a reserved subset of the observations and compare this performance against other national‐scale map products for historical to modern ET. The ensemble ET maps are shown to provide an improved accuracy over the alternative comparison products. These ET maps could be useful for a variety of hydrologic modeling and assessment applications that benefit from a long record, such as the study of periods of water scarcity through time.

Environmental Sciences & Ecology↗

Strontium isoscape of sub-Saharan Africa allows tracing origins of victims of the transatlantic slave trade

Abstract Strontium isotope ( 87 Sr/ 86 Sr) analysis with reference to strontium isotope landscapes (Sr isoscapes) allows reconstructing mobility and migration in archaeology, ecology, and forensics. However, despite the vast potential of research involving 87 Sr/ 86 Sr analysis particularly in Africa, Sr isoscapes remain unavailable for the largest parts of the continent. Here, we measure the 87 Sr/ 86 Sr ratios in 778 environmental samples from 24 African countries and combine this data with published data to model a bioavailable Sr isoscape for sub-Saharan Africa using random forest regression. We demonstrate the efficacy of this Sr isoscape, in combination with other lines of evidence, to trace the African roots of individuals from historic slavery contexts, particularly those with highly radiogenic 87 Sr/ 86 Sr ratios uncommon in the African Diaspora. Our study provides an extensive African 87 Sr/ 86 Sr dataset which includes scientifically marginalized regions of Africa, with significant implications for the archaeology of the transatlantic slave trade, wildlife ecology, conservation, and forensics.

Science & Technology - Other Topics↗

A mapped dataset of surface ocean acidification indicators in large marine ecosystems of the United States

Mapped monthly data products of surface ocean acidification indicators from 1998 to 2022 on a 0.25° by 0.25° spatial grid have been developed for eleven U.S. large marine ecosystems (LMEs). The data products were constructed using observations from the Surface Ocean CO 2 Atlas, co-located surface ocean properties, and two types of machine learning algorithms: Gaussian mixture models to organize LMEs into clusters of similar environmental variability and random forest regressions (RFRs) that were trained and applied within each cluster to spatiotemporally interpolate the observational data. The data products, called RFR-LMEs, have been averaged into regional timeseries to summarize the status of ocean acidification in U.S. coastal waters, showing a domain-wide carbon dioxide partial pressure increase of 1.4 ± 0.4 μatm yr -1 and pH decrease of 0.0014 ± 0.0004 yr -1 . RFR-LMEs have been evaluated via comparisons to discrete shipboard data, fixed timeseries, and other mapped surface ocean carbon chemistry data products. Regionally averaged timeseries of RFR-LME indicators are provided online through the NOAA National Marine Ecosystem Status web portal.

54 ENVIRONMENTAL SCIENCES↗

Predictive machine learning approaches for the microstructural behavior of multiphase zirconium alloys

Abstract Zirconium alloys are widely used in harsh environments characterized by high temperatures, corrosivity, and radiation exposure. These alloys, which have a hexagonal closed packed (h.c.p.) structure thermo-mechanically degrade, when exposed to severe operating environments due to hydride formation. These hydrides have a different crystalline structure, than the matrix, which results in a multiphase alloy. To accurately model these materials at the relevant physical scale, it is necessary to fully characterize them based on a microstructural fingerprint, which is defined here as a combination of features that include hydride geometry, parent and hydride texture and crystalline structure of these multiphase alloys. Hence, this investigation will develop a reduced order modeling approach, where this microstructural fingerprint is used to predict critical fracture stress levels that are physically consistent with microstructural deformation and fracture modes. Machine Learning (ML) methodologies based on Gaussian Process Regression, random forests, and multilayer perceptrons (MLP) were used to predict material fracture critical stress states. MLPs, or neural networks, had the highest accuracy on held-out test sets across three predetermined strain levels of interest. Hydride orientation, grain orientation or texture, and hydride volume fraction had the greatest effect on critical fracture stress levels and had partial dependencies that were highly significant, and in comparison hydride length and hydride spacing have less effects on fracture stresses. Furthermore, these models were also used accurately predicted material response to nominal applied strains as a function of the microstructural fingerprint.

36 MATERIALS SCIENCE↗

Computational multiphysics modeling of radioactive aerosol deposition in diverse human respiratory tract geometries

The evaluation of aerosol exposure relies on generic mathematical models that assume uniform particle deposition profiles over the human respiratory tract and do not account for subject-specific characteristics. Here we introduce a hybrid-automated computational workflow that generates personalized particle deposition profiles in 3D reconstructed human airways from computed tomography scans using Computational Fluid and Particle Dynamics simulations. This is the first large-scale study to consider realistic airways variability, where 380 lower and 40 upper human respiratory tract 3D geometries are reconstructed and parameterized. The data is clustered into nine groups using random forest regression. Computational fluid and particle dynamics simulations are conducted on these representative geometries using a realistic heavy-breathing respiratory cycle and radioactive iodine-131 as a source term. Monte Carlo radiation transport simulations are performed to obtain detailed energy deposition maps. Our findings emphasize the importance of personalized studies, as minor respiratory tract variations notably influence deposition patterns rather than global parameters of the lower airways, observing more than 30% variance in the mass deposition fraction.

62 RADIOLOGY AND NUCLEAR MEDICINE↗