Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “random forest regression”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Identification of novel organic polar materials: A machine learning study with importance sampling

Recent advances in the synthesis of polar molecular materials have produced practical alternatives to ferroelectric ceramics, opening up exciting new avenues for their incorporation into modern electronic devices. However, in order to realize the full potential of polar polymer and molecular crystals for modern technological applications, it is paramount to assemble and evaluate all the available data for such compounds, identifying descriptors that could be associated with an emergence of ferroelectricity. In this paper, we utilized data-driven approaches to judiciously shortlist candidate materials from a wide chemical space that could possess ferroelectric functionalities. A machine learning study with importance sampling was employed to address the challenge of having a limited amount of available data on already-known organic ferroelectrics. Sets of molecular- and crystal-level descriptors were combined with a Random Forest Regression algorithm in order to predict the spontaneous polarization of the shortlisted compounds. First-principles simulations were performed to further validate the predictions obtained from the machine learning model.

36 MATERIALS SCIENCE↗

Data-driven electrolyte design for lithium metal anodes

Improving Coulombic efficiency (CE) is key to the adoption of high energy density lithium metal batteries. Liquid electrolyte engineering has emerged as a promising strategy for improving the CE of lithium metal batteries, but its complexity renders the performance prediction and design of electrolytes challenging. Here, we develop machine learning (ML) models that assist and accelerate the design of high-performance electrolytes. Using the elemental composition of electrolytes as the features of our models, we apply linear regression, random forest, and bagging models to identify the critical features for predicting CE. Our models reveal that a reduction in the solvent oxygen content is critical for superior CE. We use the ML models to design electrolyte formulations with fluorine-free solvents that achieve a high CE of 99.70%. This work highlights the promise of data-driven approaches that can accelerate the design of high-performance electrolytes for lithium metal batteries.

25 ENERGY STORAGE↗

The metallicity’s fundamental dependence on both local and global galactic quantities

ABSTRACT We study the scaling relations between gas-phase metallicity, stellar mass surface density (Σ*), star formation rate surface density (ΣSFR), and molecular gas surface density ($\Sigma _{{\rm H}_2}$) in local star-forming galaxies on scales of a kpc. We employ optical integral field spectroscopy from the Mapping Nearby Galaxies at Apache Point Observatory (MaNGA) survey, and ALMA data for a subset of MaNGA galaxies. We use partial correlation coefficients and Random Forest regression to determine the relative importance of local and global galactic properties in setting the gas-phase metallicity. We find that the local metallicity depends primarily on Σ* (the resolved mass–metallicity relation, rMZR), and has a secondary anticorrelation with ΣSFR (i.e. a spatially resolved version of the ‘Fundamental Metallicity Relation’, rFMR). We find that $\Sigma _{{\rm H}_2}$ is less important than ΣSFR in determining the local metallicity. This result indicates that gas accretion, resulting in local metallicity dilution and local boosting of star formation, is unlikely to be the primary origin of the rFMR. The local metallicity depends also on the global properties of galaxies. We find a strong dependence on the total stellar mass (M*) and a weaker (inverse) dependence on the total SFR. The global metallicity scaling relations, therefore, do not simply stem out of their resolved counterparts; global properties and processes, such as the global gravitational potential well, galaxy-scale winds and global redistribution/mixing of metals, likely contribute to the local metallicity, in addition to local production and retention.

79 ASTRONOMY AND ASTROPHYSICS↗

Plastics Environmental Risk Calculator

SF-24-074 The Plastics Environmental Risk Calculator (PERC) was developed in Microsoft Excel and estimates the environmental distribution and lifetime of new biobased and conventional plastics from commonly measured properties of the plastics. A Random Forest regression model embedded in the calculator calculates plastic degradation rates and lifetimes as a proxy for environmental risk. Model default assumptions may be overwritten by the user.

Beckman, Kevin [Argonne National Laboratory (ANL),↗

Calibration and Rapid-Adoption Forecasting Techniques

CRAFT (Calibration and Rapid-Adoption Forecasting Techniques) CRAFT is a Python-based project for processing, analyzing, and modeling atmospheric or environmental data. It uses machine learning techniques, specifically Random Forest Regression, to create emulators for various environmental variables such as gross primary production and soil water content. It then uses these emulators to robustly test the parameter space of mechanistic models to provide posterior estimations of the free parameters.

Robins, Zachary↗

Remote Sensing and GIS data at 1km-grid over Chesapeake Bay used in “He et al. 2024, Effects of spatial variability in vegetation phenology, climate, landcover, biodiversity, topography, and soil property on soil respiration across a coastal ecosystem”

The package contains the data layers used in “He et al. 2024, Effects of spatial variability in vegetation phenology, climate, landcover, biodiversity, topography, and soil property on soil respiration across a coastal ecosystem”. The study aims to use multi-source remote sensing and GIS datasets to investigate the spatial heterogeneity and identify spatial zones with similar environmental characteristics and understand the primary driving factors affecting soil respiration within sub-ecosystems of the coastal ecosystem. We employed unsupervised hierarchical clustering analysis to identify spatial regions with distinct environmental characteristics, then determined the main driving factors using Random Forest regression and SHapley Additive exPlanations (SHAP). Spatial data layers include soil respiration, kernel Normalized Difference Vegetation Index (kNDVI) computed from Harmonized Landsat 8 and Sentinel-2 time series, climate variables from the Daymet dataset, land cover, biodiversity, topographical metrics, soil property, and tidal elevation.

54 ENVIRONMENTAL SCIENCES↗

Using Deep Learning to Automate Inference of Meteoroid Pre-Entry Properties

Properly assessing the asteroid threat depends on the knowledge of asteroid pre-entry parameters, such as size, velocity, mass, density, and strength. Although a vast number of possible bodies to study exist, such characterization of asteroid populations is currently limited by substantial costs associated with space rendezvous missions and rare meteorite findings. As asteroids fragment, ablate, and decelerate in the atmosphere, they emit light detectable by ground-based and space-borne instruments. Earth’s atmosphere, thus, becomes an accessible laboratory that enables impactor risk assessments by facilitating inference of the pre-entry parameters. These asteroid pre-entry conditions are typically deduced by modeling the entry and breakup physics that best reproduce the observed light or energy deposition curve. However, this process requires extensive manual trial-and-error of uncertain modeling parameters. Automating meteor modeling and inference would improve property distributions used in risk assessments and enable population characterization as more light curves become more readily available through the presence of space assets and ground-based camera networks. We previously developed a genetic algorithm to automate meteor modeling by using the fragment-cloud model (FCM) to search for the values of the FCM input parameters (e.g., diameter) that generate energy deposition profiles that match the observed one. Now, we apply deep learning to infer asteroid diameter, velocity, and density from observed energy deposition curves. We trained and tested our neural network models with synthetic energy deposition curves modeled using the FCM rubble pile implementation. We present an application of a 1D convolutional neural network and compare its performance to other attempted regressors and machine learning techniques, such as a fully connected neural network and Random Forest regression, to demonstrate its capabilities. We validate our model weights and approach using the Chelyabinsk, Tagish Lake, Benešov, Košice, and Lost City meteors.

Tarano, Ana Maria↗

Utilizing Airborne and Space-Based Remote Sensing Imagery to Implement the Unvegetated-Vegetated Ratio to Assess Salt Marsh Vulnerability in South Carolina

Among the most productive ecosystems on earth, salt marshes provide crucial ecosystem services including water filtration, shoreline protection, storm surge buffering, and flood mitigation. Marshes are largely dependent on their sediment budget which can significantly vary across a region and can be used to determine the life span of the marsh. Upstream land use change near Charleston, South Carolina, along with rising sea levels, are expected to alter sediment budgets and threaten marsh stability and long-term health. The unvegetated-vegetated ratio (UVVR), developed by researchers at USGS, is a scalable and efficient method to assess vulnerability. The NASA DEVELOP National Program collaborated with the South Carolina Department of Natural Resources, the South Carolina Department of Health and Environmental Control, and the United States Geological Survey Woods Hole Coastal and Marine Science Center to apply the UVVR method within Google Earth Engine. Marsh vulnerability was analyzed using UVVR derived from clustering and manual interpretation of National Agriculture Imagery Program (NAIP) high-resolution aerial imagery. NAIP derived UVVR was aggregated to Landsat 8 Operational Land Imager (OLI) and Landsat 7 Enhanced Thematic Mapper (ETM+) resolution and projection. A Random Forest Regression between Landsat derived data and UVVR was modeled to estimate a potential relationship. The estimation of this relationship was used to produce temporal change analysis maps of salt marsh vulnerability back to 1984. The NAIP imagery processed through Google Earth Engine allowed us to make detailed UVVR maps for 2009, 2015, 2017, and 2019 for decision making within South Carolina. Google Earth Engine scripting provided a novel approach to UVVR methodology that will allow decision makers to input new marsh regions and easily calculate marsh vulnerability without external data downloading. These results were used to understand what areas of the marsh need most resource allocation in the future.

NASA DEVELOP↗

The benefit of brightness temperature assimilation for the SMAP Level-4 surface and root-zone soil moisture analysis

The Soil Moisture Active Passive (SMAP) Level-4 (L4) product provides global estimates of surface soil moisture (SSM) and root-zone soil moisture (RZSM) via the assimilation of SMAP brightness temperature (Tb) observations into the NASA Catchment Land Surface Model (CLSM). Here, using in situ measurements from 2474 sites in China, we evaluate the performance of soil moisture estimates from the L4 data assimilation (DA) system and from a baseline “open-loop” (OL) simulation of CLSM without Tb assimilation. Using random forest regression, the efficiency of the L4 DA system (i.e., the performance improvement in DA relative to OL) is attributed to eight control factors related to the CLSM as well as τ–ω radiative transfer model (RTM) components of the L4 system. Results show that the Spearman rank correlation (R) for L4 SSM with in situ measurements increases for 77 % of the in situ measurement locations (relative to that of OL), with an average R increase of approximately 14 % (ΔR=0.056). RZSM skill is improved for about 74 % of the in situ measurement locations, but the average R increase for RZSM is only 7 % (ΔR=0.034). Results further show that the SSM DA skill improvement is most strongly related to the difference between the RTM-simulated Tb and the SMAP Tb observation, followed by the error in precipitation forcing data and estimated microwave soil roughness parameter h. For the RZSM DA skill improvement, these three dominant control factors remain the same, although the importance of soil roughness exceeds that of the Tb simulation error, as the soil roughness strongly affects the ingestion of DA increments and further propagation to the subsurface. For the skill of the L4 and OL estimates themselves, the top two control factors are the precipitation error and the SSM–RZSM coupling strength error, both of which are related to the CLSM component of the L4 system. Finally, we find that the L4 system can effectively filter out errors in precipitation. Therefore, future development of the L4 system should focus on improving the characterization of the SSM–RZSM coupling strength.

Jianxiu Qiu↗

South Carolina Water Resources Project Summary - Implementing the Unvegetated-Vegetated Ratio to Assess Salt Marsh Vulnerability in South Carolina Using Airborne and Space-Based Remote Sensing Imagery

Among the most productive ecosystems on earth, salt marshes provide crucial ecosystem services including water filtration, shoreline protection, storm surge buffering, and flood mitigation. Marshes are largely dependent on their sediment budget which can significantly vary across a region. Upstream land use change near Charleston, South Carolina, along with rising sea levels, are expected to alter sediment budgets and threaten marsh stability and long-term health. The unvegetated-vegetated ratio (UVVR) is a scalable and efficient method to assess vulnerability. This NASA DEVELOP project collaborated with the South Carolina Department of Natural Resources, the South Carolina Department of Health and Environmental Control, and the United States Geological Survey Woods Hole Coastal and Marine Science Center. Marsh vulnerability was analyzed using UVVR derived from Landsat 8 Operational Land Imager (OLI) and Landsat 7 Enhanced Thematic Mapper (ETM+) in conjunction with National Agriculture Imagery Program (NAIP) high-resolution aerial imagery. A Landsat random forest regression showed a low correlation (r2 = 0.247) between Landsat 7 ETM+ bands and NAIP aggregated UVVR suggesting the need for a more complex model and higher resolution sensors. Google Earth Engine scripting provided a novel approach to UVVR methodology that will allow decision makers to input new marsh areas and easily calculate UVVR without external data downloading.

DEVELOP Project Summary↗

Gila Water Resources III - Modeling the Impacts of Post-fire Restoration Methods on Vegetation Recovery in the Gila National Forest

In recent years, wildfires in New Mexico’s Gila National Forest have become increasingly common and more severe. Wildfires can have powerful impacts on hydrology and soil stability, including erosion, flooding, and debris-flows that threaten lives and infrastructure downstream. Vegetation restoration treatments like seeding and mulching can mitigate these effects and facilitate ecosystem recovery. Understanding the effectiveness of various restoration methods is vital to planning a cost-effective and successful post-fire recovery strategy. The immediate response to a fire on US Forest Service land is coordinated by a Burned Area Emergency Response (BAER) team, a group responsible for mitigating immediate post-fire risks to human life, property, and critical natural and cultural resources. This study created a proof-of-concept methodology for a decision-support tool designed to help BAER teams identify the restoration treatments most likely to succeed in a given burned area. Leveraging random forest regression, Google Earth Engine, and Landsat 7 and 8 Earth observations, this study modeled vegetation recovery after the 2013 Silver Fire for seeded areas, seeded/mulched areas, and untreated areas. Treatment type and initial burn severity were the largest drivers of vegetation recovery across the landscape. Seeded/mulched areas showed higher recovery levels than untreated areas three months post-fire, but by four years post-fire, treated and untreated areas displayed similar recovery levels. To produce a robust predictive tool for the Gila National Forest, the model should be trained on many more fires and incorporate post-fire weather conditions into the process. Such a model will help partners ensure efficient resource use and plan effective post-fire restoration strategies.

DEVELOP Project Summary↗

Global benefits of non-continuous flooding to reduce greenhouse gases and irrigation water use without rice yield penalty

Non-continuous flooding is an effective practice for reducing greenhouse gas emissions (GHGs) and irrigation water use (IRR) in rice fields. However, advancing global implementation is hampered by the lack of comprehensive understanding of GHGs and IRR reduction benefits without compromising rice yield. Here, we present the largest observational data set for such effects as of yet. By using Random Forest regression models based on 636 field trials at 105 globally georeferenced sites, we identified the key drivers of effects of non-continuous flooding practices and mapped maximum GHGs or IRR reduction benefits under optimal non-continuous flooding strategies. The results show that variation in effects of non-continuous flooding practices are primarily explained by the UnFlooded days Ratio (UFR, that is the ratio of the number of days without standing water in the field to total days of the growing period). Non-continuous flooding practices could be feasible to be adopted in 76% of global rice harvested areas. This would reduce the global warming potential (GWP) of CH4 and N2O combined from rice production by 47% or the total GWP by 7% and alleviate irrigation water use by 25%, while maintaining yield levels. The identified UFR targets far exceed currently observed levels particularly in South and Southeast Asia, suggesting large opportunities for climate mitigation and water use conservation, associated with the rigorous implementation of non-continuous flooding practices in global rice cultivation.

climate change mitigation↗

Machine Learning based Aircraft Performance Model Estimation for Trajectory Prediction

The accurate prediction of aircraft trajectory by ground-based decision support tools is a critical component of air traffic management in the US National Airspace System (NAS). Accurate predictions of where the aircraft will be in the future or when they will arrive at specific locations (e.g., fixes) is a key enabler for sequencing and efficient arrival management of flights. Traditional physics based aircraft trajectory prediction relies on a simplified point-mass total energy model whose parameters are referred to as Aircraft Performance Model (APM) parameters. Even though the performance coefficients and weight of an aircraft are a vital part of the aircraft performance model’s predictions and accuracy, these coefficients are proprietary in nature and therefore, unavailable to decision-support tools. Current approaches freeze some coefficients to default base of aircraft data (BADA) values and optimize others. However, the APM parameters are highly coupled by the flight dynamics and prioritizing one parameter over others leads to bias and skewed predictions. To alleviate this problem, we provide a combined optimization framework to predict all the critical (thrust, drag and weight) APM parameters. This paper is focused on training Machine Learning (ML) models that map historical flights to optimized APM parameters that provide the best fit (in terms of prediction error). Our dataset obtained from NASA’s Sherlock data warehouse is comprised of thousands of historical flights and includes weather and track data collected from 2019. Using different subsets of relevant features (e.g., aircraft type), we trained several ML models to estimate the aircraft’s take off weight, drag polar coefficients (both parasitic and lift induced), and thrust settings (multiplier applied to the maximum engine thrust). The chosen flights are from three of the most common aircraft types (B738, B737, and A320) arriving at four airports (LAX, DEN, MSP, and DFW). Our ML approach is comprised of two different solutions: 1- using a subset of features that are known prior to the flight departure and do not change during flight (such as engine type, current temperature at departure & destination airports, aircraft type) and 2 - using a subset of temporal features of the flight trajectory (such as cruise altitude, Mach, airspeed, and rate of climb) in addition to the pre-departure features from the first solution. The labels or target variables are the APM parameters that were obtained by an optimized ordinary differential equations (ODE) fitting process (applied to individual flights). The ODE-fitting is very time intensive and is therefore performed offline. Thus, training an ML model to learn the relationship between the flight features and ODE-generated labels enables faster estimation of the APM parameters and is therefore amenable to real-time prediction. Various ML models including linear regression, random forest, XGBoost, and neural network were trained, and the results are compared. After model validation and hyperparameter-tuning, we observed that the Random Forest model outperformed the other three models by the overall mean square error (MSE) of 2% for the first solution and 1.5% for the second solution. Finally, the ML-derived parameters are compared against default BADA APM parameters using NASA’s Autonomy Development toolkit (ADK) simulation software. The simulation results for one of each aircraft type is shown and discussed.

Aida Sharif Rohani↗

Statistical Classification of Biosignature Information using Multiple Instrument Observations

The accurate identification of biosignatures (indications of life) from data taken from remote or in situ planetary exploration is one of the most important challenges in astrobiology, the interdisciplinary field examining habitability and the potential for extraterrestrial life. This study employs machine learning algorithms to optimize the identification of biosignatures, with an emphasis on those which are agnostic to a specific biochemical basis. We exploit the wealth of terrestrial data available from biogenic and abiogenic systems to enhance efficient feature prioritization. Our dataset, pulled from public databases and laboratory recorded measurements, includes elemental abundance, isotopic fractionation, and VNIR/Raman spectra The data curation process included standardization for detection limits and ranges. Subsequent feature extraction yielded detailed inputs for machine learning, including combinations of elemental content, isotopic ratios, and parameters of spectral peaks and troughs. Feature significance was evaluated across diverse machine learning methodologies, such as k-nearest neighbors, logistic regression, Random Forest, support vector machines, and Gaussian Naïve Bayes, along with a combined voting classifier. We utilized Receiver Operating Characteristic Area Under the Curve (ROC AUC) across 2,000 50% test-train splits as a robust metric of model performance. Results revealed a promising ROC AUC of 0.853 for the combined voting classifier. Removing elemental abundance data notably reduced model accuracy (13% decrease in AUC), highlighting its critical role in biosignature detection. Several other individual data features exhibited significance within their respective data types, offering additional granularity. This research fortifies the relevance of machine learning to astrobiology, potentially enhancing life detection missions by allowing algorithmic prioritization of high-interest samples for further investigation. Future work will refine data standardization, expand the dataset to include more terrestrial systems, and incorporate convolutional neural networks for spectral feature extraction. The potential for public data sharing is also under exploration, reinforcing our commitment to collective scientific advancement.

Statistical↗

Exploring Flooded Fraction Prediction through Machine Learning Models Focusing on Medical Infrastructure in the Southeast U.S. Coastal Areas

Rising sea levels due to climate change increasingly threaten medical infrastructure through flooding. This study develops machine learning models to predict flood exposure for 11,508 medical facilities in the southeastern coastal regions of the United States by integrating datasets including meteorological, hydrological, topographic, and geological data, the Natural Risk Index, and historical flood records from NASA, HIFLD, and FEMA. Six regression models, namely Linear Regression, Support Vector Regression, Random Forest, k-Nearest Neighbors, XGBoost, and Artificial Neural Networks, are trained using 16 explanatory variables identified through literature review and correlation analysis. Data preprocessing employs the SMOGN for class imbalance and Winsorization for outliers. Model performance is evaluated using MAE, MSE, and RMSE, with Random Forest and XGBoost models achieving the highest performance (MSE of 2.58e-5 and 3.69e-5, respectively). This multifactorial approach allows the models to capture complex flood-influencing relationships, enhancing adaptability and performance across geographic regions. Future work focuses on expanding across the U.S. and developing a near real-time flood monitoring system.

Jihoon Chung↗

Quantifying mean, variability, and uncertainty in indoor radon exposure in Pennsylvania using random forest and quantile regression forest models

Radon is a naturally occurring radioactive gas that poses a serious health risk as the primary cause of lung cancer in non-smokers. Despite the well-known adverse association with health outcomes, current radon exposure assessments are limited to county-level or average-level estimates, which fail to capture regional variability. This study uses Machine Learning models, including Random Forest (RF) and Quantile Regression Forest (QRF), to estimate the indoor radon concentrations at the ZCTA (Zip code tabulation area)-level and characterize uncertainties in model estimates. Incorporating geological, meteorological, and building-specific data, the models aim to improve radon risk assessment by capturing mean exposure, variability, and extreme concentration levels. Processed radon test data (n = 718,111) were analyzed using average, variability, and quantile prediction methods. Models that estimate the average radon exposure at the ZCTA-level can yield promising model-fit results, but they do not capture the underlying variability of indoor radon exposure within a ZCTA. We utilize volatility analyses to identify characteristics indicative of high variability of indoor radon exposure. We also show that a QRF model can be used to estimate upper quantiles of residential radon exposure, thereby uncovering localized areas of elevated exposure that were not apparent in mean estimates. The results highlighted the need for a deep characterization of exposure risk and show that regions with moderate average exposure levels could still harbor extreme outliers with implications for evaluating health risks. Utilizing multiple radon exposure models allows for a deeper characterization of radon risk within a geographic area and can better identify high-risk areas. The results from this study provide a foundation for developing mitigation strategies and examining associations between radon exposure and health outcomes at fine scales. Future research should extend the geographic scope and incorporate additional environmental risk factors to establish a comprehensive framework for risk assessment.

Lee, Heechan [ORNL]↗

Evaluating county-level lung cancer incidence from environmental radiation exposure, PM 2.5 , and other exposures with regression and machine learning models

Characterizing the interplay between exposures shaping the human exposome is vital for uncovering the etiology of complex diseases. For example, cancer risk is modified by a range of multifactorial external environmental exposures. Environmental, socioeconomic, and lifestyle factors all shape lung cancer risk. However, epidemiological studies of radon aimed at identifying populations at high risk for lung cancer often fail to consider multiple exposures simultaneously. For example, moderating factors, such as PM 2.5 , may affect the transport of radon progeny to lung tissue. This ecological analysis leveraged a population-level dataset from the National Cancer Institute’s Surveillance, Epidemiology, and End-Results data (2013–17) to simultaneously investigate the effect of multiple sources of low-dose radiation (gross γ activity and indoor radon) and PM 2.5 on lung cancer incidence rates in the USA. County-level factors (environmental, sociodemographic, lifestyle) were controlled for, and Poisson regression and random forest models were used to assess the association between radon exposure and lung and bronchus cancer incidence rates. Tree-based machine learning (ML) method perform better than traditional regression: Poisson regression: 6.29/7.13 (mean absolute percentage error, MAPE), 12.70/12.77 (root mean square error, RMSE); Poisson random forest regression: 1.22/1.16 (MAPE), 8.01/8.15 (RMSE). The effect of PM 2.5 increased with the concentration of environmental radon, thereby confirming findings from previous studies that investigated the possible synergistic effect of radon and PM 2.5 on health outcomes. In summary, the results demonstrated (1) a need to consider multiple environmental exposures when assessing radon exposure’s association with lung cancer risk, thereby highlighting (1) the importance of an exposomics framework and (2) that employing ML models may capture the complex interplay between environmental exposures and health, as in the case of indoor radon exposure and lung cancer incidence.

63 RADIATION, THERMAL, AND OTHER ENVIRON. POLLUTAN↗

Machine learning and deep learning for mineralogy interpretation and CO 2 saturation estimation in geological carbon Storage: A case study in the Illinois Basin

Carbon capture and storage (CCS) is a promising approach to simultaneously maintaining energy security and reducing carbon dioxide (CO 2 ) emissions under the current energy portfolio that is dominated by fossil fuel energy. Pre-injection formation characterization and post-injection CO 2 monitoring are two critical tasks to guarantee storage efficiency in CCS. The CCS projects in the Illinois Basin, the first large-scale CO 2 injection into saline aquifers in the United States, employed conventional and the latest pulsed neutron logging (PNL) tools for mineralogy interpretation and CO 2 saturation estimation, which provide valuable references for future CCS projects. Because of the inherent fuzziness of petrophysical measurements and complex subsurface heterogeneity, interpreting well-logging data is time-consuming, and its accuracy can be user-biased. In recent years, data-driven methods have been widely used to capture the non-linear patterns between input features and interpretation results. This work applied and evaluated four commonly used machine learning (ML) models, including ridge regression (RR), random forest (RF), gradient boosting regression (GBR), support vector regression (SVR), and one deep learning (DL) model, the artificial neural network (ANN). We optimized the hyperparameters of the four ML models and the DL model using the simulated annealing algorithm and the grid search strategy, respectively. The input features of the mineralogy interpretation models were eleven conventional well-logging parameters, and the label data (i.e., ground truth) were the porosity and volumetric fractions of six minerals, including quartz, feldspar, dolomite, calcite, clay, and iron minerals. The results demonstrated that the GBR and RF models were superior in predicting volumetric fractions of minerals and porosity; label data with low coefficient of variation (CV) values tended to yield better performance. For CO 2 saturation estimation, the RF was the best-performing model, followed by SVR, ANN, GBR, and RR. Furthermore, we conducted feature importance ranking using the permutation importance algorithm and found that the formation sigma and well pressure were the most important features in this study. In conclusion, the study of CCS projects in the Illinois Basin bridges the gap between the limited knowledge and understanding of geological carbon storage and the increasing demand for reliable, cost-effective, and sustainable energy solutions.

58 GEOSCIENCES↗