Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “machine learning algorithms”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 307 records · Page 17

RU-net for automatic characterization of TRISO fuel cross sections

During irradiation, phenomena such as kernel swelling and buffer densification may impact the performance of tristructural isotropic (TRISO) particle fuel. Post-irradiation microscopy is often used to identify these irradiation-induced morphologic changes. However, each fuel compact generally contains thousands of TRISO particles. Manually performing the work to get statistical information on these phenomena is cumbersome and subjective. Here, to reduce the subjectivity inherent in that process and to accelerate data analysis, we used convolutional neural networks (CNNs) to automatically segment cross-sectional images of microscopic TRISO layers. CNNs are a class of machine-learning algorithms specifically designed for processing structured grid data. They have gained popularity in recent years due to their remarkable performance in various computer vision tasks, including image classification, object detection, and image segmentation. In this research, we generated a large irradiated TRISO layer dataset with more than 2,000 microscopic images of cross-sectional TRISO particles and the corresponding annotated images. Based on these annotated images, we used different CNNs to automatically segment different TRISO layers. These CNNs include RU-Net (developed in this study), as well as three existing architectures: U-Net, Residual Network (ResNet), and Attention U-Net. The preliminary results show that the model based on RU-Net performs best in terms of Intersection over Union (IoU). Using CNN models, we can expedite the analysis of TRISO particle cross sections, significantly reducing the manual labor involved and improving the objectivity of the segmentation results.

11 - NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

The Pixel Anomaly Detection Tool : a user-friendly GUI for classifying detector frames using machine-learning approaches

Data collection at X-ray free electron lasers has particular experimental challenges, such as continuous sample delivery or the use of novel ultrafast high-dynamic-range gain-switching X-ray detectors. This can result in a multitude of data artefacts, which can be detrimental to accurately determining structure-factor amplitudes for serial crystallography or single-particle imaging experiments. Here, a new data-classification tool is reported that offers a variety of machine-learning algorithms to sort data trained either on manual data sorting by the user or by profile fitting the intensity distribution on the detector based on the experiment. This is integrated into an easy-to-use graphical user interface, specifically designed to support the detectors, file formats and software available at most X-ray free electron laser facilities. The highly modular design makes the tool easily expandable to comply with other X-ray sources and detectors, and the supervised learning approach enables even the novice user to sort data containing unwanted artefacts or perform routine data-analysis tasks such as hit finding during an experiment, without needing to write code.

47 OTHER INSTRUMENTATION↗

Data, scripts, and figures associated with a manuscript studying impact of climate and topography on post-fire vegetation recovery.

This data package is associated with the publication “Impact of Topography and Climate on Post-fire Vegetation Recovery Across Different Burn Severity and Land Cover Types through Machine Learning” submitted to Remote Sensing of Environment (Zahura et al. 2023). In this research, a machine learning algorithm, random forest (RF), was utilized to examine the impact of climate and topography on post-fire vegetation recovery. We used enhanced vegetation index (EVI) to examine varying burn severity and land cover types. The data package includes the input files for RF model training, outputs from model predictions and analysis, and python scripts to run the model, analyze the results to understand model performance and interpretability, and plot manuscript figures. This data package contains three folders (Data, Scripts, and Figures), a file-level metadata (FLMD) csv, and a data dictionary (dd) csv. Please see Postfire_recovery_flmd.csv for a list of all files contained in this data package and descriptions for each. The data dictionary (Postfire_recovery_dd.csv) describes the csv column headers. The “Data” folder provides all the inputs and outputs to train the RF model, evaluate performance, and interpret predictions. The “Scripts” folder contains python scripts and jupyter notebooks for model training and result analysis. The “Figures” folder includes the figures used in the manuscript in “.png” and “.jpg” format.

54 ENVIRONMENTAL SCIENCES↗

Southwest Pacific tropical cyclone development classification utilizing machine learning and synoptic composites

This study evaluates the ability of machine learning algorithms to classify tropical depressions (TDs) and tropical storms (TSs) in the western region of the southwest Pacific Ocean (SWPO). Decision rules are generated to predict the environment required for a depression to fully develop into a mature storm, and the most influential predictors in the classification decision are ranked. TD and TS are discriminated based on a maximum sustained wind speed threshold (≥17 ms -1 ). Various aerosol, thermodynamic, and dynamic parameters are extracted closest to the initiation point of each non-developing and developing sample. The covariates associated with each labelled sample are used to train a decision tree and random forest model. Results using a testing dataset suggest the random forest approach more accurately distinguishes between non-developing and developing samples. The classification accuracy of the decision tree and random forest are 72% and 91%, respectively. Random forest outperformed the decision tree by providing higher accuracy in test data. The most important variables for binary classification are sea salt aerosol optical depth (AOD), 1,000 mb relative humidity, and sea surface temperature. AOD is a quantitative estimate of the aerosols presents in the air through the extinction of a ray of light as it passes through the atmosphere. Mean composite maps constructed in an unsupervised manner have been created for the most important variables identified by the random forest classifier during TD and TS events to highlight the difference in geophysical and aerosol variables' climatology during the two different classifications. This work will advance the risk management strategies for northeastern Australia and other SWPO basin islands to control their tropical cyclone related losses through prioritizing forecasting variables that are the strongest predictors of the strengthening of tropical depressions into tropical cyclones.

54 ENVIRONMENTAL SCIENCES↗

Using soil library hyperspectral reflectance and machine learning to predict soil organic carbon: Assessing potential of airborne and spaceborne optical soil sensing

Soil organic carbon (SOC) is a key variable to determine soil functioning, ecosystem services, and global carbon cycles. Spectroscopy, particularly optical hyperspectral reflectance coupled with machine learning, can provide rapid, efficient, and cost-effective quantification of SOC. However, how to exploit soil hyperspectral reflectance to predict SOC concentration, and the potential performance of airborne and satellite data for predicting surface SOC at large scales remain relatively underknown. Here, this study utilized a continental-scale soil laboratory spectral library (37,540 full-pedon 350–2500 nm reflectance spectra with SOC concentration of 0–780 g·kg –1 across the US) to thoroughly evaluate seven machine learning algorithms including Partial-Least Squares Regression (PLSR), Random Forest (RF), K-Nearest Neighbors (KNN), Ridge, Artificial Neural Networks (ANN), Convolutional Neural Networks (CNN), and Long Short-Term Memory (LSTM) along with four preprocessed spectra, i.e. original, vector normalization, continuum removal, and first-order derivative, to quantify SOC concentration. Furthermore, by using the coupled soil-vegetation-atmosphere radiative transfer model, we simulated twelve airborne and spaceborne hyper/multi-spectral remote sensing data from surface bare soil laboratory spectra to evaluate their potential for estimating SOC concentration of surface bare soils. Results show that LSTM achieved best predictive performance of quantifying SOC concentration for the whole data sets (R 2 = 0.96, RMSE = 30.81 g·kg –1 ), mineral soils (SOC ≤ 120 g·kg –1 , R 2 = 0.71, RMSE = 10.60 g·kg –1 ), and organic soils (SOC > 120 g·kg –1 , R 2 = 0.78, RMSE = 62.31 g·kg –1 ). Spectral data preprocessing, particularly the first-order derivative, improved the performance of PLSR, RF, Ridge, KNN, and ANN, but not LSTM or CNN. We found that the SOC models of mineral and organic soils should be distinguished given their distinct spectral signatures. Finally, we identified that the shortwave infrared is vital for airborne and spaceborne hyperspectral sensors to monitor surface SOC. This study highlights the high accuracy of LSTM with hyperspectral/multispectral data to mitigate a certain level of noise (soil moisture <0.4 m 3 ·m –3 , green leaf area < 0.3 m 2 ·m –2 , plant residue <0.4 m 2 ·m –2 ) for quantifying surface SOC concentration. Forthcoming satellite hyperspectral missions like Surface Biology and Geology (SBG) have a high potential for future global soil carbon monitoring, while high-resolution satellite multispectral fusion data can be an alternative.

54 ENVIRONMENTAL SCIENCES↗

Fine-tuning machine-learned particle-flow reconstruction for new detector geometries in future colliders

We demonstrate transfer learning capabilities in a machine-learned algorithm trained for particle-flow reconstruction in high energy particle colliders. This paper presents a cross-detector fine-tuning study, where we initially pretrain the model on a large full simulation dataset from one detector design, and subsequently fine-tune the model on a sample with a different collider and detector design. Specifically, we use the Compact Linear Collider detector (CLICdet) model for the initial training set and demonstrate successful knowledge transfer to the CLIC-like detector (CLD) proposed for the Future Circular Collider in electron-positron mode. We show that with an order of magnitude less samples from the second dataset, we can achieve the same performance as a costly training from scratch, across particle-level and event-level performance metrics, including jet and missing transverse momentum resolution. Furthermore, we find that the fine-tuned model achieves comparable performance to the traditional rule-based particle-flow approach on event-level metrics after training on 100,000 CLD events, whereas a model trained from scratch requires at least 1 million CLD events to achieve similar reconstruction performance. To our knowledge, this represents the first full-simulation cross-detector transfer learning study for particle-flow reconstruction. These findings offer valuable insights towards building large foundation models that can be fine-tuned across different detector designs and geometries, helping to accelerate the development cycle for new detectors and opening the door to rapid detector design and optimization using machine learning.

43 PARTICLE ACCELERATORS↗

Automated characterization of spatial and dynamical heterogeneity in supercooled liquids via implementation of machine learning

Abstract A computational approach by an implementation of the principle component analysis (PCA) with K -means and Gaussian mixture (GM) clustering methods from machine learning algorithms to identify structural and dynamical heterogeneities of supercooled liquids is developed. In this method, a collection of the average weighted coordination numbers ( W C N s ‾ ) of particles calculated from particles’ positions are used as an order parameter to build a low-dimensional representation of feature (structural) space for K -means clustering to sort the particles in the system into few meso-states using PCA. Nano-domains or aggregated clusters are also formed in configurational (real) space from a direct mapping using associated meso-states’ particle identities with some misclassified interfacial particles. These classification uncertainties can be improved by a co-learning strategy which utilizes the probabilistic GM clustering and the information transfer between the structural space and configurational space iteratively until convergence. A final classification of meso-states in structural space and domains in configurational space are stable over long times and measured to have dynamical heterogeneities. Armed with such a classification protocol, various studies over the thermodynamic and dynamical properties of these domains indicate that the observed heterogeneity is the result of liquid–liquid phase separation after quenching to a supercooled state.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

Background subtraction in inelastic scattering measurements using machine learning

Identifying, isolating, and subtracting background from the signal of interest is vital for nuclear physics experiments. These backgrounds introduce unwanted uncertainties that must be accounted for properly to extract accurate results from the signals. In nuclear reaction measurements, the typical contaminants are carbon and oxygen, contributing to background signals, and complicating the measurement of the light ejectiles. For instance, in the inelastic scattering measurement of a 20.9-MeV proton beam on 96 Mo, the 96 Mo target was contaminated with carbon and oxygen. Here, we used random forest, a machine learning algorithm commonly used for classification and regression tasks, to separate the inelastic scattering on the carbon and oxygen contaminants from the data of interest resulting from 96 Mo(p, p').

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

A machine learning method of modern urban building energy modeling: A case study of Chicago

Urban-scale building energy modeling is vital for urban planning. However, it can be challenging to assimilate reliable non-geometry building data for urban-scale modeling without extensive investment. Here, this study introduces a novel approach to developing modern urban-scale building energy stock data using geographic information systems and machine learning algorithms without necessarily requiring pre-supplied non-geometric metadata. The proposed framework integrates building footprint and height data to estimate gross floor areas, and matches each building to a pool of candidate records from ComStock or ResStock—filtered to the same county and ranked by geometric similarity—demonstrate a proof-of-concept case study in Chicago for predicting energy use intensity (EUI) using scalable datasets. The model achieved a mean bias error (MBE) of 0.08 kWh/m² and root mean square error (RMSE) of 14.84 kWh/m² under full metadata input for EUI prediction. With only location inputs, the model captured 69.2 % of EUI within predicted ranges. These results demonstrate the model’s potential to support early-stage urban planning, identify candidates for energy-efficient retrofits. By removing the dependency on detailed pre-surveys or extensive building metadata, the approach overcomes a key barrier in traditional urban-scale building energy modeling, illustrating a pathway toward broader and more cost-effective application, though further multi-city validation and improved treatment of pre-1925 buildings are needed.

Energy Use Intensity↗

Hybrid Analytics Solution to Improve Coal Power Plant Operations

This project focused on developing advanced methods for thermal performance monitoring of a coal-fueled power plant. The specific goal was to develop and demonstrate a new thermal performance monitoring approach using a hybrid model that integrates a physics-based heat balance model with a machine learning-based pattern recognition model. The hybrid model enables increased accuracy and scope of the thermal analysis and an improved ability to monitor and detect changes in plant operation. This new approach takes full advantage of the individual model capabilities and creates an important new set of capabilities not previously possible using the two types of models separately. Using the heat balance model, a rich set of derived parameters (virtual sensors) are calculated from the measured plant operating data at each time point. The combined measured and derived data values are used by machine learning algorithms to create pattern recognition models over the range of normal unit operation. To create the monitoring models, historical data from normal operation of the plant is first processed by the heat balance model to compute the derived parameter data. The result is a greatly expanded set of normal operating data that can be used as input to create the pattern recognition model. Once the models are calibrated for normal operation, the hybrid model is suitable for use in continuous online monitoring. During online monitoring, new plant operating data is processed first by the heat balance model and then by the pattern recognition model. Results from the pattern recognition model quantify the deviation of each measured or derived parameter from its expected value in normal operation. The hybrid models can detect abnormal changes in plant operating data with very high accuracy and sensitivity. When abnormal behavior is detected, alerts are generated automatically for evaluation by the plant monitoring staff. The new hybrid solution product was developed and verified in the performance of the project. The hybrid solution was tested first in a simulation environment that mimicked the plant data systems and infrastructure used by U.S. power generating plants and utilities. The hybrid solution was then deployed for real-time, online monitoring of an operating coal-fueled power plant at a field test site. Field testing demonstrated that all hybrid solution development objectives were accomplished. The project work was based on combining the capabilities of two existing software products to create the new hybrid solution product. One of these was the existing MapEx® heat balance product and the other was the existing SureSense® advanced pattern recognition product. Each of these separate products was assessed to be at a Technology Readiness Level (TRL) of 9 at the start of the effort. The hybrid solution product was assessed to be at a TRL of 2 at the start of the project based on early feasibility work by the project team. At completion of the field testing performed in the project, the hybrid solution product was assessed to be at a TRL of 7. The project team expects that the hybrid solution product will be deployed commercially and will achieve a TRL of 9 within one year after completion of the project.

01 COAL, LIGNITE, AND PEAT↗

cTULIP: application of a human-based RNA-seq primary tumor classification tool for cross-species primary tumor classification in canine

The domestic dog, Canis familiaris, is quickly gaining traction as an advantageous model for use in the study of cancer, one of the leading causes of death worldwide. Naturally occurring canine cancers share clinical, histological, and molecular characteristics with the corresponding human diseases. In this study, we take a deep-learning approach to test how similar the gene expression profile of canine glioma and bladder cancer (BLCA) tumors are to the corresponding human tumors. We likewise develop a tool for identifying misclassified or outlier samples in large canine oncological datasets, analogous to that which was developed for human datasets. We test a number of machine learning algorithms and found that a convolutional neural network outperformed logistic regression and random forest approaches. We use a recently developed RNA-seq-based convolutional neural network, TULIP, to test the robustness of a human-data-trained primary tumor classification tool on cross-species primary tumor prediction. Our study ultimately highlights the molecular similarities between canine and human BLCA and glioma tumors, showing that protein-coding one-to-one homologs shared between humans and canines, are sufficient to distinguish between BLCA and gliomas. The results of this study indicate that using protein-coding one-to-one homologs as the features in the input layer of TULIP performs good primary tumor prediction in both humans and canines. Furthermore, our analysis shows that our selected features also contain the majority of features with known clinical relevance in BLCA and gliomas. Our success in using a human-data-trained model for cross-species primary tumor prediction also sheds light on the conservation of oncological pathways in humans and canines, further underscoring the importance of the canine model system in the study of human disease.

60 APPLIED LIFE SCIENCES↗

A Multi-Sensor Approach for Measuring Bird and Bat Collisions with Offshore Wind Turbines (Final Technical Report)

Collision of birds and bats with wind turbines is a conservation concern for both land-based and offshore wind projects. The fatality rates of birds and bats at land-based turbines are well documented. The measurement strategies on land focus on finding carcasses following collision, estimating the number of carcasses missed through searcher efficiency, carcass persistence trials and carcass fall distributions, and modeling statistically robust fatality rates. Few technologies have been developed to monitor offshore bird and bat collisions, and many that have been developed focused on detecting collisions with large birds. The few studies that have attempted to document collisions at offshore turbines do not account for smaller bodied animals or for collisions that might be missed, which prevents the calculation of statistically robust fatality rates. The overall goal of this report, A Multi-Sensor Approach for Measuring Bird and Bat Collisions with Offshore Wind Turbines (Project), was to develop an effective multi-sensor system for quantifying bird and bat collision rates, specifically for offshore wind facilities. The Project goal and resulting automated collision detection system was achieved through two major technological advancements: 1) refining The Netherlands Organisation for Applied Scientific Research’s (TNO’s) existing WT-Bird® vibration sensing system, that had successfully detected large bird collisions during daytime, to allow for improved detection of smaller birds and bats during both daytime and nighttime hours and 2) improving image processing systems and developing and integrating machine learning algorithms to automatically detect and classify small and large bird and bat collisions with offshore turbines. This final technical report (FTR) summarizes Methods , Results , Conclusions , and Lessons Learned during each of the five Tasks identified for this research and development effort. This FTR includes summaries of the following: Task 1. Initial Engineering Tests to Improve WT-Bird® Task 2. Installation of WT‐Bird® on a Utility-scale Turbine at the National Wind Technology Center – National Renewable Energy Laboratory Task 3. Field Tests and Refinement of the Object Detection System Task 4. Validation of WT-Bird® on a Land-based Turbine Task 5. Preparation for the Implementation of WT-Bird® on an Offshore Turbine. This research and development effort documented successful improvement of the WT Bird® collision detection system to detect small birds and bats, and WT-Bird® is the first collision detection system to validate results compared to land-based post-construction monitoring. The collision trials provide estimates of missed targets that can be used to estimate fatality rates, a significant improvement relative to other offshore collision monitoring systems. Advances were made in developing an edge-processing solution to reduce data storage requirements, which is important if the system is deployed for long periods of time at offshore turbines. The improved WT-Bird® system also provides an important option for wind operators on land or offshore who need to document specific details about when collisions occur, particularly efforts to further research on bat impact minimization, or when standard fatality searches are impractical (e.g. offshore) or inadequate (e.g. challenging locations on land).

17 WIND ENERGY↗

Adiabatic quantum support vector machines

Adiabatic quantum computers can solve difficult optimization problems (e.g., the quadratic unconstrained binary optimization problem), and they seem well suited to train machine learning models. In this paper, we describe an adiabatic quantum approach for training support vector machines. We show that the time complexity of our quantum approach is an order of magnitude better than the classical approach. Next, we compare the test accuracy of our quantum approach against a classical approach that uses the Scikit-learn library in Python across five benchmark datasets (Iris, Wisconsin Breast Cancer (WBC), Wine, Digits, and Lambeq). We show that our quantum approach obtains accuracies on par with the classical approach. Finally, we perform a scalability study in which we compute the total training times of the quantum approach and the classical approach with an increasing number of features and an increasing number of data points in the training dataset. In conclusion, our scalability results show that the quantum approach obtains a 3.5–4.5x speedup over the classical approach on datasets with many (millions of) features.

Computational Complexity↗

Phase Selection Rules of Multi‐Principal Element Alloys

Abstract Computational prediction of phase stability of multi‐principal element alloys (MPEAs) holds a lot of promise for rapid exploration of the enormous design space and autonomous discovery of superior structural and functional properties. Regardless of many plausible works that rely on phenomenological theory and machine learning, precise prediction is still limited by insufficient data and the lack of interpretability of some machine learning algorithms, e.g., convolutional neural network. In this work, a comprehensive approach is presented, encompassing the development of a complete dataset that contains 72 387 density functional theory calculations, as well as a predictive global phenomenological descriptor. The phase selection descriptor, based on atomic electronegativity and valence electron concentration, significantly outperforms the widely used valence electron concentration, excelling in both accuracy (with an f1 score of 63% compared to 47%) and its ability to predict the HCP phase (0.48 recall compared to 0). The comprehensive data mining on the global design space of 61 425 quaternary MPEAs made from 28 possible metals, together with the phenomenological theory and physical interpretation, will set up a solid computational science foundation for data‐driven exploration of MPEAs.

Chemistry↗

Contaminant Investigation and Pre‐Processing Opportunities for Textile‐To‐Textile Recycling

Millions of metric tons of textiles are landfilled or incinerated each year in the United States, with less than 1% of textiles recycled into new clothing or fabrics. To counter this trend, a growing number of companies and researchers are exploring how a circular economy can be applied to support textile‐to‐textile recycling. A significant barrier they face comes down to quickly and efficiently extracting pure feedstock material from post‐consumer garments that feature a mix of natural and synthetic fibers. Textile recyclers prefer pure feedstocks, as working with mixed sources typically means lower throughput, higher risk of equipment failure, and diminished business margins. To facilitate a circular economy for textiles, methods, and technologies are needed that can efficiently separate out materials and contaminants from end‐of‐life textiles to increase the flow of pure feedstocks to recyclers. This paper summarizes findings from interviews with a cross section of textile recyclers and from a review of literature to define basic feedstock requirements. In addition to our qualitative research, we deconstruct a bale of post‐consumer textiles and analyze them using computer‐vision imaging, Fourier transform infrared spectroscopy (FTIR), and machine learning. The resulting data are used to set system‐level design inputs for an automated contaminant removal system to process post‐consumer clothing into appropriate feedstocks for recycling. To set the system's levels for automated real‐time near‐infrared analysis, we identify the minimum percentage of primary material that any single garment in a load of used clothing must contain for the average of the full output stream to meet the target purity levels of recyclers. Here, the envisioned automated system can also address undesirable trace materials that might contaminate the processed stream by using imaging cameras coupled with artificial intelligence to identify sections of clothing for de‐trimming. Proof‐of‐concept machine learning algorithms are evaluated to locate and identify trims or garment areas with hidden contaminant materials. Integrating these methods into automated textile cutting systems can provide a cost‐effective means for increasing feedstock purity from used clothing, which can advance circularity for textiles by helping recyclers to reach production volumes and quality targets that were not possible solely with manual dismantling operations.

Parsons, Ryan [Rochester Institute of Technology, ↗

Synchrotron‐source micro‐x‐ray computed tomography for examining butterfly eyes

Comparative anatomy is an important tool for investigating evolutionary relationships among species, but the lack of scalable imaging tools and stains for rapidly mapping the microscale anatomies of related species poses a major impediment to using comparative anatomy approaches for identifying evolutionary adaptations. We describe a method using synchrotron source micro-x-ray computed tomography (syn-μXCT) combined with machine learning algorithms for high-throughput imaging of Lepidoptera (i.e., butterfly and moth) eyes. Our pipeline allows for imaging at rates of ~15 min/mm 3 at 600 nm 3 resolution. Image contrast is generated using standard electron microscopy labeling approaches (e.g., osmium tetroxide) that unbiasedly labels all cellular membranes in a species-independent manner thus removing any barrier to imaging any species of interest. To demonstrate the power of the method, we analyzed the 3D morphologies of butterfly crystalline cones, a part of the visual system associated with acuity and sensitivity and found significant variation within six butterfly individuals. Despite this variation, a classic measure of optimization, the ratio of interommatidial angle to resolving power of ommatidia, largely agrees with early work on eye geometry across species. We show that this method can successfully be used to determine compound eye organization and crystalline cone morphology. Our novel pipeline provides for fast, scalable visualization and analysis of eye anatomies that can be applied to any arthropod species, enabling new questions about evolutionary adaptations of compound eyes and beyond.

59 BASIC BIOLOGICAL SCIENCES↗

Modeling the distribution of the endangered Jemez Mountains salamander (Plethodon neomexicanus) in relation to geology, topography, and climate

The Jemez Mountains salamander (Plethodon neomexicanus; hereafter JMS) is an endangered salamander restricted to the Jemez Mountains in north-central New Mexico, United States. This strictly terrestrial and lungless species requires moist surface conditions for activities such as mating and foraging. Threats to its current habitat include fire suppression and ensuing severe fires, changes in forest composition, habitat fragmentation, and climate change. Forest composition changes resulting from reduced fire frequency and increased tree density suggest that its current aboveground habitat does not mirror its historically successful habitat regime. However, because of its limited habitat area and underground behavior, we hypothesized that geology and topography might play a significant role in the current distribution of the salamander. We modeled the distribution of the JMS using a machine learning algorithm to assess how geology, topography, and climate variables influence its distribution. The best habitat suitability model indicates that geology type and maximum winter temperature (November to March) were most important in predicting the distribution of the salamander (23.5% and 50.3% permutation importance, respectively). Minimum winter temperature was also an important variable (21.4%), suggesting this also plays a role in salamander habitat. Our habitat suitability map reveals low uncertainty in model predictions, and we found slight discrepancies between the designated critical habitat and the most suitable areas for the JMS. Because geological features are important to its distribution, we recommend that geological and topographical data are considered, both during survey design and in the description of localities of JMS records once detected.

59 BASIC BIOLOGICAL SCIENCES↗

BrazilClim : The overcoming of limitations of pre‐existing bioclimate data

Abstract Species distribution modelling has become instrumental in assessing the influence of environmental conditions on the occurrence or abundance of taxa. The set of environmental layers used for this purpose is a crucial aspect, for which different climate‐based (bioclimatic) datasets have been recently developed. These bioclimatic variables result from combinations of precipitation and temperatures surfaces. Here, we explored both the performance and possibility of improving some of the currently available bioclimatic databases, through an evaluation of the precipitation and temperatures surfaces used to generate them. For this purpose, we used a combination of statistic and graphic approaches. We focused on Brazil, not only due to its natural megadiversity, but also due to its continental size and orographic heterogeneity: an excellent ground for refining methods replicable elsewhere. We found a better match between the climatic data measured on‐field and Tropical Rainfall Measuring Mission (TRMM 3B43 v7) in the case of precipitation, and the surfaces provided by the National Oceanic and Atmospheric Administration (NOAA) in the case of temperatures, sources uncommonly used for species niche modelling. We gauge‐calibrated the best performing surfaces using machine‐learning algorithms and generated corrected surfaces that allowed us to create BrazilClim: a database of bioclimatic variables, based on improved primary surfaces, which will result in more assertive predicted distributions and more actual pictures of the species' ecological requirements for megadiverse Brazil, an approach replicable elsewhere. All primary and bioclimatic surfaces generated for this study may be freely downloaded.

Ramoni‐Perazzi, Paolo↗