Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “iterative random forest”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Evaluating the performance of random forest and iterative random forest based methods when applied to gene expression data

Gene-to-gene networks, such as Gene Regulatory Networks (GRN) and Predictive Expression Networks (PEN) capture relationships between genes and are beneficial for use in downstream biological analyses. There exists multiple network inference tools to produce these gene-to-gene networks from matrices of gene expression data. Random Forest-Leave One Out Prediction (RF-LOOP) is a method that has been shown to be efficient at producing these gene-to-gene networks, frequently known as GEne Network Inference with Ensemble of trees (GENIE3). Random Forest can be replaced in this process by iterative Random Forest (iRF), which performs variable selection and boosting. Here we validate that iterative Random Forest-Leave One Out Prediction (iRF-LOOP) produces higher quality networks than GENIE3 (RF-LOOP). We use both synthetic and empirical networks from the Dialogue for Reverse Engineering Assessment and Methods (DREAM) Challenges by Sage Bionetworks, as well as two additional empirical networks created from Arabidopsis thaliana and Populus trichocarpa expression data.

59 BASIC BIOLOGICAL SCIENCES↗

Evaluating the Performance of Random Forest and Iterative Random Forest Based Methods when Applied to Gene Expression Data

Gene-to-gene networks, such as Gene Regulatory Networks (GRN) and Predictive Expression Networks (PEN) capture relationships between genes and are beneficial for use in downstream biological analyses. There exists multiple network inference tools to produce these gene-to-gene networks from matrices of gene expression data. Random Forest-Leave One Out Prediction (RF-LOOP) is a method that has been shown to be efficient at producing these gene-to-gene networks, frequently known as GEne Network Inference with Ensemble of trees (GENIE3). Here we validate that iterative Random Forest-Leave One Out Prediction (iRF-LOOP) produces higher quality networks than GENIE3. We use both synthetic and empirical networks from the Dialogue for Reverse Engineering Assessment and Methods (DREAM) Challenges by Sage Bionetworks, as well as two additional empirical networks created from Arabidopsis thaliana and Populus trichocarpa expression data.

iRF-Loop, expression network, Populus Trichocarpa↗

Using iterative random forest to find geospatial environmental and Sociodemographic predictors of suicide attempts

Despite a recent global decrease in suicide rates, death by suicide has increased in the United States. It is therefore imperative to identify the risk factors associated with suicide attempts to combat this growing epidemic. In this study, we aim to identify potential risk factors of suicide attempt using geospatial features in an Artificial intelligence framework. We use iterative Random Forest, an explainable artificial intelligence method, to predict suicide attempts using data from the Million Veteran Program. This cohort incorporated 405,540 patients with 391,409 controls and 14,131 attempts. Our predictive model incorporates multiple climatic features at ZIP-code-level geospatial resolution. We additionally consider demographic features from the American Community Survey as well as the number of firearms and alcohol vendors per 10,000 people to assess the contributions of proximal environment, access to means, and restraint decrease to suicide attempts. In total 1,784 features were included in the predictive model. Our results show that geographic areas with higher concentrations of married males living with spouses are predictive of lower rates of suicide attempts, whereas geographic areas where males are more likely to live alone and to rent housing are predictive of higher rates of suicide attempts. We also identified climatic features that were associated with suicide attempt risk by age group. Additionally, we observed that firearms and alcohol vendors were associated with increased risk for suicide attempts irrespective of the age group examined, but that their effects were small in comparison to the top features. Taken together, our findings highlight the importance of social determinants and environmental factors in understanding suicide risk among veterans.

60 APPLIED LIFE SCIENCES↗

Quantum biological insights into CRISPR-Cas9 sgRNA efficiency from explainable-AI driven feature engineering

Abstract CRISPR-Cas9 tools have transformed genetic manipulation capabilities in the laboratory. Empirical rules-of-thumb have been developed for only a narrow range of model organisms, and mechanistic underpinnings for sgRNA efficiency remain poorly understood. This work establishes a novel feature set and new public resource, produced with quantum chemical tensors, for interpreting and predicting sgRNA efficiency. Feature engineering for sgRNA efficiency is performed using an explainable-artificial intelligence model: iterative Random Forest (iRF). By encoding quantitative attributes of position-specific sequences for Escherichia coli sgRNAs, we identify important traits for sgRNA design in bacterial species. Additionally, we show that expanding positional encoding to quantum descriptors of base-pair, dimer, trimer, and tetramer sequences captures intricate interactions in local and neighboring nucleotides of the target DNA. These features highlight variation in CRISPR-Cas9 sgRNA dynamics between E. coli and H. sapiens genomes. These novel encodings of sgRNAs enhance our understanding of the elaborate quantum biological processes involved in CRISPR-Cas9 machinery.

59 BASIC BIOLOGICAL SCIENCES↗

NF-κB perturbation reveals unique immunomodulatory functions in Prx1 + fibroblasts that promote development of atopic dermatitis

Skin is composed of diverse cell populations that cooperatively maintain homeostasis. Up-regulation of the nuclear factor κB (NF-κB) pathway may lead to the development of chronic inflammatory disorders of the skin, but its role during the early events remains unclear. Here, through analysis of single-cell RNA sequencing data via iterative random forest leave one out prediction, an explainable artificial intelligence method, we identified an immunoregulatory role for a unique paired related homeobox-1 (Prx1) + fibroblast subpopulation. Disruption of Ikkb–NF-κB under homeostatic conditions in these fibroblasts paradoxically induced skin inflammation due to the overexpression of C-C motif chemokine ligand 11 (CCL11; or eotaxin-1) characterized by eosinophil infiltration and a subsequent T H 2 immune response. Because the inflammatory phenotype resembled that seen in human atopic dermatitis (AD), we examined human AD skin samples and found that human AD fibroblasts also overexpressed CCL11 and that perturbation of Ikkb–NF-κB in primary human dermal fibroblasts up-regulated CCL11. Monoclonal antibody treatment against CCL11 was effective in reducing the eosinophilia and T H 2 inflammation in a mouse model. Together, the murine model and human AD specimens point to dysregulated Prx1 + fibroblasts as a previously unrecognized etiologic factor that may contribute to the pathogenesis of AD and suggest that targeting CCL11 may be a way to treat AD-like skin lesions.

60 APPLIED LIFE SCIENCES↗

Learning epistatic polygenic phenotypes with Boolean interactions

Detecting epistatic drivers of human phenotypes is a considerable challenge. Traditional approaches use regression to sequentially test multiplicative interaction terms involving pairs of genetic variants. For higher-order interactions and genome-wide large-scale data, this strategy is computationally intractable. Moreover, multiplicative terms used in regression modeling may not capture the form of biological interactions. Building on the Predictability, Computability, Stability (PCS) framework, we introduce the epiTree pipeline to extract higher-order interactions from genomic data using tree-based models. The epiTree pipeline first selects a set of variants derived from tissue-specific estimates of gene expression. Next, it uses iterative random forests (iRF) to search training data for candidate Boolean interactions (pairwise and higher-order). We derive significance tests for interactions, based on a stabilized likelihood ratio test, by simulating Boolean tree-structured null (no epistasis) and alternative (epistasis) distributions on hold-out test data. Finally, our pipeline computes PCS epistasis p-values that probabilisticly quantify improvement in prediction accuracy via bootstrap sampling on the test set. We validate the epiTree pipeline in two case studies using data from the UK Biobank: predicting red hair and multiple sclerosis (MS). In the case of predicting red hair, epiTree recovers known epistatic interactions surrounding MC1R and novel interactions, representing non-linearities not captured by logistic regression models. In the case of predicting MS, a more complex phenotype than red hair, epiTree rankings prioritize novel interactions surrounding HLA-DRB1 , a variant previously associated with MS in several populations. Taken together, these results highlight the potential for epiTree rankings to help reduce the design space for follow up experiments.

59 BASIC BIOLOGICAL SCIENCES↗

Invasion in the Niger Delta: Remote Sensing of Mangrove Conversion to Invasive Nypa fruticans from 2015-2020

Invasive species are a leading threat to biodiversity worldwide. Nypa palm ( Nypa fruticans ) has emerged as the predominant invasive species in the Niger Delta region of Nigeria. While endemic mangroves have high rates of carbon sequestration, stabilize coastlines, and protect biodiversity, Nypa does not provide these services outside its native region of Southeast Asia. Oil exploration and urbanization in this region also exacerbates mangrove loss and Nypa spread. As Nypa is difficult to distinguish from endemic mangrove species in remotely sensed data, estimates of mangrove and ecosystem services losses in Nigeria are highly uncertain. Here, we analyze multisensor satellite data with machine learning to quantify the rapid expansion of Nypa from 2015-2020 in Nigeria. Using Landsat imagery and random forest classification, we quantify total potential Nypa extent in Nigeria in 2019. We then produced a Nypa extent map using iterative combinations of Sentinel-1 SAR, Sentinel-2 MSI, and ALOS PALSAR. Random forest classifications using SAR data from ALOS and Sentinel-1 were best suited for mapping Nypa extent with similar accuracies (78% and 75% respectively). Based on data availability and accuracy, we focused our change analysis on Sentinel-1 SAR. Our results show ~28,000 ha of mangroves were converted to Nypa in Nigeria by 2020 and covered a larger extent than endemic mangroves, compounding the effect of the existing degradation and deforestation in the region. We also compared forest height and complexity estimates from GEDI (Global Ecosystem Dynamics Investigation) LiDAR to further distinguish between endemic mangroves and Nypa in three dimensions. Nypa structural variability, measured by top-of-canopy height, vegetation cover, plant area index, and foliage height diversity, was lower than that of mangroves. At current rates of Nypa expansion, the entire area of study would be invaded by Nypa by 2028, with potentially detrimental consequences to the ecosystem services provided by mangroves.

GEE↗

Developing a SARS-CoV-2 main protease binding prediction random forest model for drug repurposing for COVID-19 treatment

The coronavirus disease 2019 (COVID-19) global pandemic resulted in millions of people becoming infected with the severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2) virus and close to seven million deaths worldwide. It is essential to further explore and design effective COVID-19 treatment drugs that target the main protease of SARS-CoV-2, a major target for COVID-19 drugs. In this study, machine learning was applied for predicting the SARS-CoV-2 main protease binding of Food and Drug Administration (FDA)-approved drugs to assist in the identification of potential repurposing candidates for COVID-19 treatment. Ligands bound to the SARS-CoV-2 main protease in the Protein Data Bank and compounds experimentally tested in SARS-CoV-2 main protease binding assays in the literature were curated. These chemicals were divided into training (516 chemicals) and testing (360 chemicals) data sets. To identify SARS-CoV-2 main protease binders as potential candidates for repurposing to treat COVID-19, 1188 FDA-approved drugs from the Liver Toxicity Knowledge Base were obtained. A random forest algorithm was used for constructing predictive models based on molecular descriptors calculated using Mold2 software. Model performance was evaluated using 100 iterations of fivefold cross-validations which resulted in 78.8% balanced accuracy. The random forest model that was constructed from the whole training dataset was used to predict SARS-CoV-2 main protease binding on the testing set and the FDA-approved drugs. Model applicability domain and prediction confidence on drugs predicted as the main protease binders discovered 10 FDA-approved drugs as potential candidates for repurposing to treat COVID-19. Our results demonstrate that machine learning is an efficient method for drug repurposing and, thus, may accelerate drug development targeting SARS-CoV-2.

Research & Experimental Medicine↗

A framework to evaluate machine learning crystal stability predictions

The rapid adoption of machine learning in various scientific domains calls for the development of best practices and community agreed-upon benchmarking tasks and metrics. We present Matbench Discovery as an example evaluation framework for machine learning energy models, here applied as pre-filters to first-principles computed data in a high-throughput search for stable inorganic crystals. We address the disconnect between (1) thermodynamic stability and formation energy and (2) retrospective and prospective benchmarking for materials discovery. Alongside this paper, we publish a Python package to aid with future model submissions and a growing online leaderboard with adaptive user-defined weighting of various performance metrics allowing researchers to prioritize the metrics they value most. To answer the question of which machine learning methodology performs best at materials discovery, our initial release includes random forests, graph neural networks, one-shot predictors, iterative Bayesian optimizers and universal interatomic potentials. We highlight a misalignment between commonly used regression metrics and more task-relevant classification metrics for materials discovery. Accurate regressors are susceptible to unexpectedly high false-positive rates if those accurate predictions lie close to the decision boundary at 0 eV per atom above the convex hull. The benchmark results demonstrate that universal interatomic potentials have advanced sufficiently to effectively and cheaply pre-screen thermodynamic stable hypothetical materials in future expansions of high-throughput materials databases.

Riebesell, Janosh↗

Mapping Rare Earths and Toxics in E-Waste via Hyperspectral Imaging and Machine Learning

Electronic waste (e-waste) presents a mounting challenge to environmental sustainability due to its complex composition, which includes high-value rare earth elements, hazardous organic compounds, and non-recyclable plastics. Accurate and scalable material classification is essential for enabling efficient resource recovery and safe recycling practices. This study introduces a confidence-aware classification pipeline that combines mid-infrared hyperspectral imaging (HSI), spectral angle mapping (SAM), and iterative machine learning to perform pixel-level material identification across e-waste devices. A curated spectral library encompassing artificial materials (e.g., plastic iron oxide, galvanized metals), minerals (e.g., allanite, hematite), and organic compounds (e.g., benzanthracene, toluene) was used to generate pseudo-labels, each assigned a confidence score based on SAM-derived spectral similarity. High-confidence samples from seven consumer electronics—digital cameras, keyboards, laptop fans, modems, motherboards, TV remotes, and speakers—were iteratively expanded and classified using models such as Support Vector Machine (SVM), Random Forest, Gradient Boosting Classifier, Partial Least Squares Discriminant Analysis (PLSDA) and Logistic Regression. The best-performing classifiers achieved macro F1 scores approaching 1.0. Results revealed widespread plastic content (dominated by plastic iron oxide), the presence of rare earth-bearing minerals like cerium-containing allanite, and pervasive detection of hazardous organics such as benzanthracene. Principal Component Analysis (PCA) visualizations and confusion matrices confirmed high separability and robust classification performance. This methodology enables precise, non-destructive, and scalable classification of heterogeneous e-waste streams. It supports automated, hazard-aware sorting in recycling workflows, facilitating selective recovery of critical materials and compliance with circular economy goals. The confidence-aware framework provides a foundation for real-time deployment in industrial settings, offering significant implications for smart e-recycling infrastructure and policy-driven material stewardship.

Circular economy↗

Navigating Team Dynamics: Automated Detection of Micro-Behaviors Between Team Members Through Longitudinal Interaction Data

The success in future long term space exploration missions will depend on the cooperation, coordination, and mutual understanding among the crew members. Micro-behaviors are momentary, subtle linguistic and paralinguistic indicators of thinking and feeling toward another member of the team (Cortina et al., 2001; Smith & Griffiths, 2022) that can significantly impact team dynamics and influence the overall team performance (Paromita & Chaspari, 2024). Due to their interactive nature, micro-behaviors have a sender (i.e., the team member expressing the micro-behavior) and a target (the team member impacted by the micro-behavior). Detection of these behaviors can assist in avoiding possible conflict among crew members and promoting the overall team success. Our prior research focused on an initial proof of concept of machine learning (ML) models and natural language processing (NLP) techniques that were used for automatically detect micro-behaviors among crew members of the US National Aeronautics and Space Administration’s (NASA) Human Exploration Research Analog (HERA) Campaigns 4 and 5 missions (Paromita et al., 2023). Results underscored the importance of incorporating contextual information in the ML models in the form of sentiment analysis, type of task, and dyadic interaction among team members. Here, we expand the scope of our prior work in two ways. First, we assess ML/NLP methods on new behavioral annotations coded using an adapted version of Smith & Griffins (2022) theoretical framework in terms of Violation (i.e., presence of valenced behavior, uplifting/positive or discouraging/negative), Intensity (i.e., force of behavior in terms of how uplifting or discouraging is the behavior), and Intent (i.e., motive of the behavior in terms of whether it was deliberate or unintentional). Second, we expand the design of the ML model to preserve information about the role of each team member within the occurrence of the micro-behavior (in contrast to the previous model that only considered the sender and the target without determining the team member role). This allows to consider all team members' contributions in the conversation and model long-term dependencies in the dialogue. Our experiments for this study are conducted on data from 5 teams of the NASA HERA C4 (NASA grant NNX16AQ48G (PI: Bell)). Conversations were extracted from the 1.5 hour Team Interaction Battery (TIB) task that occurred 5 times in-mission per crew. This resulted in a total of 13,058 conversational turns (i.e., 17.8% uplifting, 3.3% discouraging, 75.76% neutral, 3.14% nulls). Our findings with the revised behavioral coding and ML/NLP models indicate a 43.66% macro F1-score (i.e., 38.29% precision (P), 50.8% recall (R)) for a dialog state-tracking model that includes information from the sender only, and a 40.9% F1-score (i.e., 38.7% P, 43.36% R) for the same model that includes information from both the sender and the target of the micro-behavior. These are significantly higher compared to simple random forest models that classify behaviors strictly based on speech content and do not consider iterative team dynamics, achieving a 36.07% F1-score (i.e., 39.04% R, 33.53% P). Our findings demonstrate potential ways to leverage large conversational datasets to better capture complex team dynamics. We will discuss future directions including proposed models that can incorporate additional mission days and tasks beyond the TIB for objectively quantifying team behavior at high temporal resolution in space exploration missions.

Projna Paromita↗

Polarimetric signatures of a coniferous forest canopy based on vector radiative transfer theory

Complete polarization signatures of a coniferous forest canopy are studied by the iterative solution of the vector radiative transfer equations up to the second order. The forest canopy constituents (leaves, branches, stems, and trunk) are embedded in a multi-layered medium over a rough interface. The branches, stems and trunk scatterers are modeled as finite randomly oriented cylinders. The leaves are modeled as randomly oriented needles. For a plane wave exciting the canopy, the average Mueller matrix is formulated in terms of the iterative solution of the radiative transfer solution and used to determine the linearly polarized backscattering coefficients, the co-polarized and cross-polarized power returns, and the phase difference statistics. Numerical results are presented to investigate the effect of transmitting and receiving antenna configurations on the polarimetric signature of a pine forest. Comparison is made with measurements.

Karam, M. A.↗

Active Learning for Rapid Targeted Synthesis of Compositionally Complex Alloys

The next generation of advanced materials is tending toward increasingly complex compositions. Synthesizing precise composition is time-consuming and becomes exponentially demanding with increasing compositional complexity. An experienced human operator does significantly better than a novice but still struggles to consistently achieve precision when synthesis parameters are coupled. The time to optimize synthesis becomes a barrier to exploring scientifically and technologically exciting compositionally complex materials. This investigation demonstrates an active learning (AL) approach for optimizing physical vapor deposition synthesis of thin-film alloys with up to five principal elements. We compared AL-based on Gaussian process (GP) and random forest (RF) models. The best performing models were able to discover synthesis parameters for a target quinary alloy in 14 iterations. We also demonstrate the capability of these models to be used in transfer learning tasks. RF and GP models trained on lower dimensional systems (i.e., ternary, quarternary) show an immediate improvement in prediction accuracy compared to models trained only on quinary samples. Furthermore, samples that only share a few elements in common with the target composition can be used for model pre-training. We believe that such AL approaches can be widely adapted to significantly accelerate the exploration of compositionally complex materials.

Chemistry↗

Artificial intelligence driven laser parameter search: Inverse design of photonic surfaces using greedy surrogate-based optimization

Photonic surfaces designed with specific optical characteristics are becoming increasingly crucial for novel energy harvesting and storage systems. The design of these surfaces can be achieved by texturing materials using lasers. The optimal adjustment of laser fabrication parameters to achieve target surface optical properties is an open challenge. Thus, we develop a surrogate-based optimization approach. Our framework employs the Random Forest algorithm to model the forward relationship between the laser fabrication parameters and the resulting optical characteristics. During the optimization process, we use a greedy, prediction-based exploration strategy that iteratively selects batches of laser parameters to be used in experimentation by minimizing the predicted discrepancy between the surrogate model’s outputs and the user-defined target optical characteristics. This strategy allows for efficient identification of optimal fabrication parameters without the need to model the error landscape directly. We demonstrate the efficiency and effectiveness of our approach on two synthetic benchmarks and two specific experimental applications of photonic surface inverse design targets. By calculating the average performance of our algorithm compared to other state of the art optimization methods, we show that our algorithm performs, on average, twice as well across all benchmarks. Additionally, a warm starting inverse design technique for changed target optical characteristics enhances the performance of the introduced approach.

97 MATHEMATICS AND COMPUTING↗

Evaluating Combinations of Sentinel-2 Data and Machine-Learning Algorithms for Mangrove Mapping in West Africa

Creating a national baseline for natural resources, such as mangrove forests, and monitoring them regularly often requires a consistent and robust methodology. With freely available satellite data archives and cloud computing resources, it is now more accessible to conduct such large-scale monitoring and assessment. Yet, few studies examine the reproducibility of such mangrove monitoring frameworks, especially in terms of generating consistent spatial extent. Our objective was to evaluate a combination of image processing approaches to classify mangrove forests along the coast of Senegal and The Gambia. We used freely available global satellite data (Sentinel-2), and cloud computing platform (Google Earth Engine) to run two machine learning algorithms, random forest (RF), and classification and regression trees (CART). We calibrated and validated the algorithms using 800 reference points collected using high-resolution images. We further re-ran 10 iterations for each algorithm, utilizing unique subsets of the initial training data. While all iterations resulted in thematic mangrove maps with over 90% accuracy, the mangrove extent ranges between 827-2807 km2 for Senegal and 245-1271 km2 for The Gambia with one outlier for each country. We further report "Places of Agreement" (PoA) to identify areas where all iterations for both methods agree (506.6 km2 and 129.6 km2 for Senegal and The Gambia, respectively), thus have a high confidence in predicting mangrove extent. While we acknowledge the time- and cost-effectiveness of such methods for the landscape managers, we recommend utilizing them with utmost caution, as well as post-classification on-the-ground checks, especially for decision making.

Mondal, Pinki↗

Autonomous fabrication of tailored defect structures in 2D materials using machine learning-enabled scanning transmission electron microscopy

Materials with tailored quantum properties can be engineered from atomic-scale assembly techniques, but existing methods often lack the agility and accuracy to precisely and intelligently control the manufacturing process. Here, we demonstrate a fully autonomous approach for fabricating atomic-level defects using electron beams in scanning transmission electron microscopy (STEM) that combines advanced machine learning and automated beam control. As a proof of concept, we achieved controlled fabrication of MoS-nanowire (MoS-NW) edge structures by iterative and targeted exposure of MoS 2 monolayer to a focused electron beam to selectively eject sulfur atoms, utilizing high-angle annular dark-field (HAADF) imaging for feedback-controlled monitoring of structural evolution of defects. A machine learning framework combining a random forest model and a convolutional neural network (CNN) was developed to decode the HAADF image and accurately identify atomic positions and species. This atomic-level information was then integrated into an autonomous decision-making platform, which applied predefined fabrication strategies to instruct beam control about atomic sites to be ejected. The selected sites were subsequently exposed to a localized electron beam using an FPGA-controlled scan routine with precise control over beam positioning and duration. While the MoS-NW edge structures produced exhibit promising mechanical and electronic properties, the proposed methods to build the autonomous fabrication framework is material-agnostic and can be extended to other 2D materials for the creation of diverse defect structures and heterostructures beyond Mo S2 .

Engineering↗

Machine learning models for rat multigeneration reproductive toxicity prediction

Reproductive toxicity is one of the prominent endpoints in the risk assessment of environmental and industrial chemicals. Due to the complexity of the reproductive system, traditional reproductive toxicity testing in animals, especially guideline multigeneration reproductive toxicity studies, take a long time and are expensive. Therefore, machine learning, as a promising alternative approach, should be considered when evaluating the reproductive toxicity of chemicals. We curated rat multigeneration reproductive toxicity testing data of 275 chemicals from ToxRefDB (Toxicity Reference Database) and developed predictive models using seven machine learning algorithms (decision tree, decision forest, random forest, k-nearest neighbors, support vector machine, linear discriminant analysis, and logistic regression). A consensus model was built based on the seven individual models. An external validation set was curated from the COSMOS database and the literature. The performances of individual and consensus models were evaluated using 500 iterations of 5-fold cross-validations and the external validation data set. The balanced accuracy of the models ranged from 58% to 65% in the 5-fold cross-validations and 45%–61% in the external validations. Prediction confidence analysis was conducted to provide additional information for more appropriate applications of the developed models. The impact of our findings is in increasing confidence in machine learning models. We demonstrate the importance of using consensus models for harnessing the benefits of multiple machine learning models (i.e., using redundant systems to check validity of outcomes). While we continue to build upon the models to better characterize weak toxicants, there is current utility in saving resources by being able to screen out strong reproductive toxicants before investing in vivo testing. The modeling approach (machine learning models) is offered for assessing the rat multigeneration reproductive toxicity of chemicals. Our results suggest that machine learning may be a promising alternative approach to evaluate the potential reproductive toxicity of chemicals.

consensus model↗

Machine Learning Models for Mapping Groundwater Pollution Risk: Advancing Water Security and Sustainable Development Goals in Georgia, USA

The widespread use of pesticides, such as atrazine and malathion, in agricultural systems raises significant concerns regarding the contamination of groundwater, which serves as a critical resource for drinking water. This study applies machine learning techniques to predict the concentrations of atrazine and malathion in groundwater across Georgia, USA, using 2019 data. A Random Forest classifier was employed to integrate various environmental and demographic factors, including pesticide application rates, precipitation, lithology, and population density, to predict pesticide contamination in groundwater. The models demonstrated high training accuracies of 100% and moderate average testing accuracy of 55% for atrazine and 60% for malathion across five iterations. The low test accuracy of the model, ranging from 50% to 75%, is likely due to overfitting, which can be attributed to the small dataset size and the complex nature of pesticide-contamination patterns, making it challenging for the model to generalize to unseen data. Feature importance analysis revealed that average pesticide usage emerged as the most influential factor for atrazine, while aquifer lithology and precipitation played crucial roles in both models. These results provide valuable insights into the dynamics of pesticide contamination, highlighting areas at greater risk of contamination. The findings underscore the importance of integrating environmental, geological, and agricultural variables for more effective groundwater management and sustainable agricultural practices, contributing to the protection of water resources and public health.

54 ENVIRONMENTAL SCIENCES↗