Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “machine learning, Random Forest”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Machine learning snow depth predictions at sites in Alaska, Norway, Siberia, Colorado and New Mexico

Temporally continuous snow depth estimates are vital for understanding changing snow patterns in the Arctic and impacts on permafrost. We trained random forest machine learning models to predict snow depth from temperature data recorded at or just below the ground surface. Training data was collected at the Teller 27 Watershed and Kougarok 64 Hillslope during the 2021 - 2022 water year on the Seward Peninsula, Alaska using distributed temperature profiling (DTP) systems. We then applied this model to other sites where ground surface or shallow soil temperature data was available for at least one water year (see Related Datasets). Many of these temperature measurements were collocated with snow depth observations. Ground surface temperature (i.e. snow-ground interface temperature) is easy to measure using small, cheap and easy-to-deploy temperature sensors such as iButtons and TinyTags, and such measurements have previously been used to calculate a variety of snow metrics (e.g. snow onset date). However, this is the first study to estimate snow depth directly from ground surface temperature data. The present dataset contains one *.csv file which includes machine learning snow depth predictions at sites in Alaska, Norway, Siberia, Colorado, and New Mexico and one *.kml file including the locations of sites with snow depth predictions. No training data predictions are included in the *.csv file. The Next-Generation Ecosystem Experiments: Arctic (NGEE Arctic), was a research effort to reduce uncertainty in Earth System Models by developing a predictive understanding of carbon-rich Arctic ecosystems and feedbacks to climate. NGEE Arctic was supported by the Department of Energy’s Office of Biological and Environmental Research. The NGEE Arctic project had two field research sites: 1) located within the Arctic polygonal tundra coastal region on the Barrow Environmental Observatory (BEO) and the North Slope near Utqiagvik (Barrow), Alaska and 2) multiple areas on the discontinuous permafrost region of the Seward Peninsula north of Nome, Alaska. Through observations, experiments, and synthesis with existing datasets, NGEE Arctic provided an enhanced knowledge base for multi-scale modeling and contributed to improved process representation at global pan-Arctic scales within the Department of Energy’s Earth system Model (the Energy Exascale Earth System Model, or E3SM), and specifically within the E3SM Land Model component (ELM).

54 ENVIRONMENTAL SCIENCES↗

Cross-domain digital twin architecture for predictive maintenance via machine learning and Large Language Models

This research introduces a comprehensive framework for creating and deploying a digital twin platform for continuous monitoring and predictive maintenance within industrial settings. Through utilizing advanced technologies, including Unreal Engine 5, Unity 3D, the Message Queue Telemetry Transport protocol, Random Forest machine learning algorithms, and Large Language Models (LLMs), we establish a platform that digitally reproduces physical equipment and translates digital controls into real-world actions. This facilitates preventive maintenance approaches and improves operational effectiveness. The digital twin platform gathers sensor data from operational equipment, analyzes it using machine learning, and delivers practical insights to prevent potential malfunctions and enhance equipment performance. Furthermore, the incorporation of a web portal enables efficient monitoring and access to historical data, educational materials, and equipment status information. Preliminary findings indicate that digital twins can transform industrial equipment management and maintenance methodologies.

97 MATHEMATICS AND COMPUTING↗

Machine Learning Assisted Gap-Filled Discharge Data for the East River Community Watershed, Colorado, for Water Years 2014-2021

This dataset contains a collection of machine learning assisted gap-filled discharge data created for all discharge stations across the East River Watershed, Colorado. This data was generated by using raw discharge data collected by Rosemary Carroll, and conducting a random forest machine learning analysis to gap-fill discharge data across all years at the hourly time level. Discharge data with gaps creates problems for analysis of measured and modeled fluxes of carbon and nitrogen exported out of each sub-watershed. Gap-filled data is also required as an input to surface water models, which helps to address our main research question related to how snowmelt timing impacts the timing and magnitude of nitrogen exports. Data is provided in one csv file.

54 ENVIRONMENTAL SCIENCES↗

MArVD2: a machine learning enhanced tool to discriminate between archaeal and bacterial viruses in viral datasets

Abstract Our knowledge of viral sequence space has exploded with advancing sequencing technologies and large-scale sampling and analytical efforts. Though archaea are important and abundant prokaryotes in many systems, our knowledge of archaeal viruses outside of extreme environments is limited. This largely stems from the lack of a robust, high-throughput, and systematic way to distinguish between bacterial and archaeal viruses in datasets of curated viruses. Here we upgrade our prior text-based tool (MArVD) via training and testing a random forest machine learning algorithm against a newly curated dataset of archaeal viruses. After optimization, MArVD2 presented a significant improvement over its predecessor in terms of scalability, usability, and flexibility, and will allow user-defined custom training datasets as archaeal virus discovery progresses. Benchmarking showed that a model trained with viral sequences from the hypersaline, marine, and hot spring environments correctly classified 85% of the archaeal viruses with a false detection rate below 2% using a random forest prediction threshold of 80% in a separate benchmarking dataset from the same habitats.

Vik, Dean (ORCID:000000027546899X)↗

Study of Antarctic Blowing Snow Storms Using MODIS and CALIOP Observations With a Machine Learning Model

As a common phenomenon over Antarctica, blowing snow (BLSN), especially the large BLSN storms, play an important role in the Antarctic surface mass balance, radiation budget, and planetary boundary layer processes. This study presents the work on BLSN storm identification and analysis with observations from the Moderate Resolution Imaging Spectroradiometer (MODIS) onboard the Aqua satellite. Spectral analysis shows that BLSN identification is feasible with MODIS daytime data. A random forest machine learning model is developed and observations from the Cloud‐Aerosol Lidar with Orthogonal Polarization are used for training. Model performance results show that machine‐learning based classification can achieve over 90% overall accuracy when classifying MODIS pixels into cloud, clear, and BLSN categories. The machine learning model is applied to MODIS observations during the month of October 2009 for BLSN storm analysis. Results show that the size of BLSN storms has a large spectrum and can reach hundreds of thousands km2. The MODIS based BLSN storm frequency map extends the Cloud‐Aerosol Lidar and Infrared Pathfinder Satellite Observations coverage limit from 82°S to the South Pole. A BLSN storm belt, which extends from the South Pole region to the coastal area between 130°E and 160°E along the Transantarctic Mountains, provides a potential pathway of snow transport. These results are important in improving the understanding of BLSN impact on Antarctic surface mass balance and boundary layer processes.

Antarctic↗

Seasonal drivers of dissolved oxygen across a tidal creek–marsh interface revealed by machine learning

Abstract Dissolved oxygen (DO) is a key biogeochemical control in coastal systems, and its concentration and drivers vary markedly through time and space. This makes it difficult to accurately represent coastal DO and associated biogeochemical processes in models, limiting our ability to predict how these systems will respond to global change. We obtained high‐frequency (5‐min) in situ measurements of DO collected at three locations across the interface of a tidal creek and coastal marsh in the Pacific Northwest, USA. Random Forest machine learning models quantified the importance of three categories of environmental drivers (Aquatic, Climatic, and Terrestrial) of DO variability across the creek–marsh interface. We selected two 4‐month datasets representing Summer and Winter seasonal periods to test two hypotheses on the dominant drivers of DO at the coastal interface. We found that the Terrestrial driver—characterized by long periods of anaerobic conditions and episodic pulses in DO after floods—was most important during the Winter, whereas the Aquatic driver—characterized by variability over tidal, diel, and lunar cycles—was most important during the Summer. We explored how future climate change scenarios could alter the drivers of DO variability using a cumulative sums driver–response framework. Our results suggest that under climate change, Aquatic and Climatic drivers may increase in importance during the Summer, potentially linked to changing metabolic regimes and sea level, with Terrestrial driver importance potentially increasing during the Winter. Our approach highlights useful methods for understanding the spatiotemporal complexity of oxygen across coastal interfaces and quantifying the relative importance of distinct environmental drivers.

54 ENVIRONMENTAL SCIENCES↗

Predictive Modeling of NOx Emissions from Lean Direct Injection of Hydrogen and Hydrogen/Natural Gas Blends Using Flame Imaging and Machine Learning

This research paper explores the use of machine learning to relate images of flame structure and luminosity to measured NOx emissions. Images of reactions produced by 16 aero-engine derived injectors for a ground-based turbine operated on a range of fuel compositions, air pressure drops, preheat temperatures and adiabatic flame temperatures were captured and postprocessed. The experimental investigations were conducted under atmospheric conditions, capturing CO, NO and NOx emissions data and OH* chemiluminescence images from 27 test conditions. The injector geometry and test conditions were based on a statistically designed test plan. These results were first analyzed using the traditional analysis approach of analysis of variance (ANOVA). The statistically based test plan yielded 432 data points, leading to a correlation for NOx emissions as a function of injector geometry, test conditions and imaging responses, with 70.2% accuracy. As an alternative approach to predicting emissions using imaging diagnostics as well as injector geometry and test conditions, a random forest machine learning algorithm was also applied to the data and was able to achieve an accuracy of 82.6%. This study offers insights into the factors influencing emissions in ground-based turbines while emphasizing the potential of machine learning algorithms in constructing predictive models for complex systems.

08 HYDROGEN↗

Quantifying wildfire drivers and predictability in boreal peatlands using a two-step error-correcting machine learning framework in TeFire v1.0

Abstract. Wildfires are becoming an increasing challenge to the sustainability of boreal peatland (BP) ecosystems and can alter the stability of boreal carbon storage. However, predicting the occurrence of rare and extreme BP fires proves to be challenging, and gaining a quantitative understanding of the factors, both natural and anthropogenic, inducing BP fires remains elusive. Here, we quantified the predictability of BP fires and their primary controlling factors from 1997 to 2015 using a two-step correcting machine learning (ML) framework that combines multiple ML classifiers, regression models, and an error-correcting technique. We found that (1) the adopted oversampling algorithm effectively addressed the unbalanced data and improved the recall rate by 26.88 %–48.62 % when using multiple datasets, and the error-correcting technique tackled the overestimation of fire sizes during fire seasons; (2) nonparametric models outperformed parametric models in predicting fire occurrences, and the random forest machine learning model performed the best, with the area under the receiver operating characteristic curve ranging from 0.83 to 0.93 across multiple fire datasets; and (3) four sets of factor-control simulations consistently indicated the dominant role of temperature, air dryness, and climate extreme (i.e., frost) for boreal peatland fires, overriding the effects of precipitation, wind speed, and human activities. Our findings demonstrate the efficiency and accuracy of ML techniques in predicting rare and extreme fire events and disentangle the primary factors determining BP fires, which are critical for predicting future fire risks under climate change.

54 ENVIRONMENTAL SCIENCES↗

Fast and Accurate Machine Learning Strategy for Calculating Partial Atomic Charges in Metal–Organic Frameworks

Computational high-throughput screening using molecular simulations is a powerful tool for identifying top-performing metal–organic frameworks (MOFs) for gas storage and separation applications. Accurate partial atomic charges are often required to model the electrostatic interactions between the MOF and the adsorbate, especially when the adsorption involves molecules with dipole or quadrupole moments such as water and CO 2 . Although ab initio methods can be used to calculate accurate partial atomic charges, these methods are impractical for screening large material databases because of the high computational cost. We developed a random forest machine learning model to predict the partial atomic charges in MOFs using a small yet meaningful set of features that represent both the elemental properties and the local environment of each atom. The model was trained and tested on a collection of about 320 000 density-derived electrostatic and chemical (DDEC) atomic charges calculated on a subset of the Computation-Ready Experimental Metal–Organic Framework (CoRE MOF-2019) database and separately on charge model 5 (CM5) charges. The model predicts accurate atomic charges for MOFs at a fraction of the computational cost of periodic density functional theory (DFT) and is found to be transferable to other porous molecular crystals and zeolites. In conclusion, a strong correlation is observed between the partial atomic charge and the average electronegativity difference between the central atom and its bonded neighbors.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Use of Longitudinal Serum Analysis and Machine Learning to Develop a Classifier for Cancer Early Detection

Early detection of solid tumors through a simple screening process, such as the proteomic analysis of biofluids, has the potential to significantly alter the management and outcomes of cancers. The application of advanced targeted proteomics measurements and data analysis strategies to uniformly collected serum or plasma samples would enable longitudinal studies of cancer risk, progression, and response to therapy that have the potential to significantly reduce cancer burden in general. In this article, we describe a generalizable workflow combining robust, multiplexed targeted proteomics measurements applied to longitudinal samples from the Department of Defense Serum Repository with a Random Forest machine learning method for developing and initially evaluating the performance of candidate biomarker panels for early detection of cancers. The effectiveness of this approach was demonstrated in a cohort of 175 head and neck squamous cell carcinoma patients. The outlined protocols include methods for sample preparation, instrument analysis, and data analysis and interpretation using this workflow.

Longitudinal analysis, machine learning, cancer, e↗

Importance of Depth and Artificial Structure as Predictors of Female Red Snapper Reproductive Parameters

Abstract The Red Snapper Lutjanus campechanus is a structure‐associated species occurring across a wide depth range in the northern Gulf of Mexico. We used the random forest machine learning algorithm to understand which habitat and individual fish characteristics could predict reproductive parameters of female Red Snapper. We evaluated fish captured from 2016 to 2018 on three artificial structure types with various structure heights at depths of 100 m or less. Overall, we found that depth and month were important predictors for most reproductive parameters, but the type of structure (artificial reefs, oil platforms, and rigs‐to‐reefs structures) was not important. Maturity was correctly classified in 88.9% of the cases when using the random forest ensemble model, with important predictors including FL, depth, structure height, and month of collection. Spawning seasonality (measured as gonadosomatic index [GSI]) was correctly classified in 59.5% of the cases when using histology reproductive phase, FL, month, and depth variables. Reproductively active or inactive females were correctly classified in 89.3% of the cases using GSI, month, FL, and depth, while females in the developing versus spawning capable phases were correctly classified in 82.2% of the cases using GSI, FL, month, and depth. Histological indicators that show potential spawning within a 36‐h period were correctly classified 61.5% of the time, with the best predictors being depth, FL, GSI, and month. Stepwise regression indicated that month was the only factor that significantly predicted contrasts in relative batch fecundity, with significantly greater values in August compared to all other months. Our findings suggest that female Red Snapper reproductive effort is not consistently or well predicted by artificial structure type or height but that a combination of fish FL, month, and depth can predict reproductive characteristics of female Red Snapper.

Brown‐Peterson, Nancy J.↗

Code Description for "Brief Communication: Monitoring snow depth using small, cheap, and easy-to-deploy ground surface temperature sensors"

Temporally continuous snow depth estimates are vital for understanding changing snow patterns and impacts on permafrost in the Arctic. We train a random forest machine learning model to predict snow depth from variability in ground surface temperature. To our knowledge, this is the first time that small ground surface temperature sensors have been used to estimate snow depth. The model performs well at sites where the model was trained and at pan-arctic evaluation sites (RMSE <= 0.15 m). Small temperature sensors are cheap and easy-to-deploy, so this technique enables spatially distributed and temporally continuous snowpack monitoring to an extent previously infeasible. The model is flexible and can be applied to datasets retroactively to retrieve snow depth estimates at additional sites. This code package includes a *.joblib file of the trained random forest model and a *.ipynb file showing how to clean input data, train the random forest model, and apply the model.

Bachand, Claire↗

Fine particulate concentrations over East Asia derived from aerosols measured by the Advanced Himawari Imager using machine learning

Fine particulate matter with a diameter below 2.5 μm (PM 2.5 ) is deleterious to the cardiovascular and respiratory systems. It is often difficult to assess the effects of PM 2.5 on human health over regions with limited ground monitoring sites, especially in East Asia. As an alternative, we estimated near-surface PM 2.5 concentrations by analyzing Advanced Himawari Imager (AHI) Yonsei Aerosol Retrieval (YAER) products. This study incorporates daytime data for East Asia covering the Korean Peninsula, China, Japan, Southeast Asia, and southern Mongolia. We collocated AHI YAER product pixels with meteorological, land-cover, and other ancillary data for the period from March 2018 to February 2019. To estimate PM 2.5 concentrations over wide areas spanning many countries displaying various relationships between aerosol optical depth and PM 2.5 , monthly models were developed by considering both the spatial and temporal characteristics of ground-based PM 2.5 measurements. Random forest machine learning model estimated ground-level mass concentrations of PM 2.5 ; subsequent 10-fold cross validation (CV) yielded a CV R 2 value of 0.81 and a CV root mean squared error (RMSE) of 12.3 μg m -3 . We investigated the spatial pattern of PM 2.5 concentrations over multiple countries and seasonal variation in PM 2.5 concentrations. Diurnal variation of a severe PM 2.5 event in the Korean Peninsula was investigated as a case study. The model captured the extremely heterogeneous spatial distribution of PM 2.5 concentrations peaked around local noon. To measure the capability of the developed model to estimate PM 2.5 concentrations in areas with few in-situ data, its predictive performance was evaluated using a dataset independent of the training process with an R 2 of 0.60 and RMSE of 8.18 μg m −3 . This study demonstrates the potential for satellite-based PM 2.5 estimation for areas with insufficient measuring stations.

Pm2.5↗

Microglia are implicated in the development of paclitaxel chemotherapy-associated cognitive impairment in female mice

Chemotherapy remains a mainstay in the treatment of many types of cancer even though it is associated with debilitating behavioral side effects referred to as “chemobrain,” including difficulty concentrating and memory impairment. The predominant hypothesis in the field is that systemic inflammation drives these cognitive impairments, although the brain mechanisms by which this occurs remain poorly understood. Here, we hypothesized that microglia are activated by chemotherapy and drive chemotherapy-associated cognitive impairments. To test this hypothesis, we treated female C57BL/6 mice with a clinically-relevant regimen of a common chemotherapeutic, paclitaxel (6 i.p. doses at 30 mg/kg), which impairs memory of an aversive stimulus as assessed via a contextual fear conditioning (CFC) paradigm. In this work, paclitaxel increased the percent area of IBA1 staining in the dentate gyrus of the hippocampus. Moreover, using a machine learning random forest classifier we identified immunohistochemical features of reactive microglia in multiple hippocampal subregions that were distinct between vehicle- and paclitaxel-treated mice. Paclitaxel treatment also increased gene expression of inflammatory cytokines in a microglia-enriched population of cells from mice. Lastly, a selective inhibitor of colony stimulating factor 1 receptor, PLX5622, was employed to deplete microglia and then assess CFC performance following paclitaxel treatment. PLX5622 significantly reduced hippocampal gene expression of paclitaxel-induced proinflammatory cytokines and restored memory, suggesting that microglia play a critical role in the development of chemotherapy-associated neuroinflammation and cognitive impairments. This work provides critical evidence that microglia drive paclitaxel-associated cognitive impairments, a key mechanistic detail for determining preventative and intervention strategies for these burdensome side effects.

60 APPLIED LIFE SCIENCES↗

Unraveling Adsorbate-Induced Structural Evolution of Iron Carbide Nanoparticles

Iron carbide (Fe x C y ) nanoparticles (NPs) are promising candidates for replacing platinum group metals in industrial applications, such as high-temperature Fischer–Tropsch synthesis. However, due to their amorphous nature, characterization of the active sites has been challenging experimentally and computationally. Here, using a combined density functional theory (DFT), neural network interatomic potential-assisted global optimization, and ensemble learning study, we evaluate dynamic surface changes associated with syngas (H and CO) interactions. For this purpose, we have developed a general procedure that we use to model an experimentally relevant 270-atom Fe 182 C 88 NP using the neural network-assisted stochastic surface walk global optimization algorithm (SSW-NN). Once generated, the Fe 182 C 88 NP active sites and particle morphology are thoroughly characterized before the effects of syngas adsorbate interactions are explored by using DFT and molecular dynamics simulations. Lastly, we explore correlations between geometric and electronic features of the active sites and the adsorption of H (H ads ), using a regularized random forest machine learning algorithm. In doing so, we identified the Fe–C coordination number and p orbital occupancy as the most important descriptors affecting H ads . Furthermore, using a combined ML and quantum chemistry approach, our work demonstrates a general and efficient procedure for generating and probing complex surface phenomena on binary nanoparticles.

Adsorption↗

Adrastea: An Efficient FPGA Design Environment for Heterogeneous Scientific Computing and Machine Learning

We present Adrastea, an efficient FPGA design environment for developing scientific machine learning applications. FPGA development is challenging, from deployment, proper toolchain setup, programming methods, interfacing FPGA kernels, and more importantly, the need to explore design space choices to get the best performance and area usage from the FPGA kernel design. Adrastea provides an automated and scalable design flow to parameterize, implement, and optimize complex FPGA kernels and associated interfaces. We show how virtualization of the development environment via virtual machines is leveraged to simplify the setup of the FPGA toolchain while deploying the FPGA boards and while scaling up the automated design space exploration to leverage multiple machines concurrently. Adrastea provides an automated build and test environment of FPGA kernels. By exposing design space hyper-parameters, Adrastea can automatically search the design space in parallel to optimize the FPGA design for a given metric, usually performance or area. Adrastea simplifies the task of interfacing with the FPGA kernels with a simplified interface API. To demonstrate the capabilities of Adrastea, we implement a complex random forest machine learning kernel with 10,000 input features while achieving extremely low computing latency without loss of prediction accuracy, which is required by a scientific edge application at SNS. We also demonstrate Adrastea using an FFT kernel and show that for both applications Adrastea is able to systematically and efficiently evaluate different design options, which reduced the time and effort required to develop the kernel from months of manual work to days of automatic builds.

Young, Aaron↗

Energy Infrastructure Futures: A Multiscale Evaluation of Projected Power Plant Siting Across the Western Interconnection

Energy Infrastructure Futures: A Multiscale Evaluation of Projected Power Plant Siting Across the Western Interconnection Description This dataset contains input and output data for the manuscript Mongird, K. et al. (under review) titled "Energy Infrastructure Futures: A Multiscale Evaluation of Projected Power Plant Siting Across the Western Interconnection". Input data corresponds to gridded spatial siting attributes that are necessary to conduct a random forest machine learning analysis of siting feature importance. Output data includes SHAP feature analysis outputs, and classification report values. For data on power plant siting results referred to in the manuscript, please refer to the CERF: IM3 Projected Western US Power Plant Locations data download page. The downloadable data includes values for eight different future scenarios for the Western US. The scenarios include combinations of two Shared Socioeconomic Pathways (SSP3 and SSP5) with four high-resolution climate projections specific to the United States (see, https://tgw-data.msdlive.org/). These climate projections include "hotter" and "cooler" variants for two Representative Concentration Pathways (RCP4.5 and RCP8.5). The resulting eight simulations are: rcp45cooler_ssp3 rcp45cooler_ssp5 rcp45hotter_ssp3 rcp45hotter_ssp5 rcp85cooler_ssp3 rcp85cooler_ssp5 rcp85hotter_ssp3 rcp85hotter_ssp5 Technical Information The dataset includes two sets of data files: (1) CERF gridded siting parameters and (2) Feature analysis outputs and classification reports. All downloadable data is in csv file format. Files with x/y coordinate information use the Albers Equal Area Conic projection (ESRI:102003). 1. CERF Gridded Siting Parameters This directory provides a balanced sample of gridded CERF siting parameters data for eight different scenarios for the Western US through 2055, seven different technologies, and eight timesteps. This data serves as input to the feature analysis. It contains the following parameters. region_name - name of region (i.e., state) sited - binary value representing whether the grid cell received a siting of that technology type (1=True) rcp - binary value representing scenario resource concentration pathway (0 = RCP4.5, 1 = RCP8.5) ssp - binary value representing scenario shared socioeconomic pathway (0 = SSP3, 1 = SSP5) climate - binary value representing cooler (0) or hotter (1) GCM forcing tech_name - generation technology name sited_year - year that values correspond to transmission_cost - cost of transmission interconnection pipeline_cost - cost of natural gas pipeline interconnection interconnection_cost - total interconnection cost (sum of transmission cost and gas pipeline cost) lmp - associated locational marginal value ($/MWh) associated with the grid cell, timestep, scenario, and technology xcoord - x-coordinate of location ycoord - y-coordinate of location 2a. Feature Analysis Output The dataset includes the feature analysis shap output for locational marginal price and interconnection cost. It contains the following parameters. technology - generator technology name scenario - name of scenario feature - name of feature, either locational_marginal_price or interconnection_cost value - the mean of absolute value of SHAP values for given feature 2b. Feature Analysis Classification Report This download includes the classification report associated with each random forest model. The dataset contains the following parameters. technology - generation technology name scenario - name of scenario test - one of precision (the proportion of predicted positives that are actually correct), recall (the proportion of actual positives that were correctly identified), f1-score (the harmonic mean of precision and recall) 0.0 - value of test for classification of 0 (grid cell not chosen for siting) 1.0 - value of test for classification of 1 (grid cell chosen for siting) accuracy - accuracy of model (i.e., fraction of all predictions that were right) macro avg - Simple average of test values for all classes weighted avg - Weighted average of test values for all classes, weighted based on Acknowledgment IM3 is a multi-institutional effort led by Pacific Northwest National Laboratory and supported by the U.S. Department of Energy's Office of Science as part of research in MultiSector Dynamics, Earth and Environmental Systems Modeling Program. License This data is made available under a CCBY4 License Disclaimer This material was prepared as an account of work sponsored by an agency of the United States Government. Neither the United States Government nor the United States Department of Energy, nor the Contractor, nor any or their employees, nor any jurisdiction or organization that has cooperated in the development of these materials, makes any warranty, express or implied, or assumes any legal liability or responsibility for the accuracy, completeness, or usefulness or any information, apparatus, product, software, or process disclosed, or represents that its use would not infringe privately owned rights. Reference herein to any specific commercial product, process, or service by trade name, trademark, manufacturer, or otherwise does not necessarily constitute or imply its endorsement, recommendation, or favoring by the United States Government or any agency thereof, or Battelle Memorial Institute. The views and opinions of authors expressed herein do not necessarily state or reflect those of the United States Government or any agency thereof. PACIFIC NORTHWEST NATIONAL LABORATORYoperated byBATTELLEfor theUNITED STATES DEPARTMENT OF ENERGYunder Contract DE-AC05-76RL01830

Mongird, Kendall [Pacific Northwest National Labor↗