Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “classification models”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 361 records · Page 20

Adaptive Discovery and Mixed-Variable Optimization of Next Generation Synthesizable Microelectronic Materials

Design of new microelectronic materials is characterized by several challenges such as high-dimensionality of the atomic structure-composition variable space, formidable cost of directly using high-fidelity simulations for design optimization, dispersity in literature-reported similar materials and synthesis methods, complex physical mechanisms, and mixed qualitative and quantitative design variables that lead to a disjointed design space. Even though machine learning (ML) techniques have been employed to expedite materials innovation, existing methods treat ML and design optimization as two separate processes, failing to resolve the fundamental challenges associated with high dimensionality and mixed-variable complexity. We have developed a ML enhanced mixed-variable material design optimization framework to efficiently extract useful information from existing data in literature and physics-based simulations to guide the autonomous search for optimal materials. Our proposed framework is composed of four computational modules: (1) a natural language processing (NLP) based virtual screening module, (2) classification based concept exploration module, (3) a density functional theory (DFT)-based high-fidelity evaluation model, and (4) a novel latent-variable Gaussian process (LVGP) ML model for mixed-variable problems with uncertainty quantification, which seamlessly integrates with Bayesian Optimization (BO) and achieves superb efficiency through embedded physics-based dimension reduction. Our approach is demonstrated and validated using the testbed of functional materials exhibiting metal-insulation transitions (MITs), with the targeted reversible resistivity changes (∼10^5) near room temperature. At the end of the 30-month project, we have developed a series of new ML techniques using NLP, conditional variational autoencoders, active learning, latent-variable Gaussian processes, integrated with Bayesian optimization. Our project has resulted in new predicted MITs compounds and improved understanding of MITs microscopic mechanisms, which in turn will revolutionize microelectronics science to provide energy-saving solutions. Our research has improved both creativity and efficiency in transforming rare-event discoveries of new functional materials to persistent innovations. In addition to open-sourcing the online MIT database and the classification model, the LVGP open source code has been downloaded more than 15,000 times within two years. More than 40 MIT compounds have been identified and many have been pursued experimentally via collaborators. The research results are published in close to 20 collaborative papers in high-impact journals, such as Chem. Mater., Appl. Phys. Rev., Sci. Rep., among others of design space.

36 MATERIALS SCIENCE↗

Path-BigBird: An AI-Driven Transformer Approach to Classification of Cancer Pathology Reports

PURPOSE Surgical pathology reports are critical for cancer diagnosis and management. To accurately extract information about tumor characteristics from pathology reports in near real time, we explore the impact of using domain-specific transformer models that understand cancer pathology reports. METHODS We built a pathology transformer model, Path-BigBird, by using 2.7 million pathology reports from six SEER cancer registries. We then compare different variations of Path-BigBird with two less computationally intensive methods: Hierarchical Self-Attention Network (HiSAN) classification model and an offthe-shelf clinical transformer model (Clinical BigBird). We use five pathology information extraction tasks for evaluation: site, subsite, laterality, histology, and behavior. Model performance is evaluated by using macro and micro F 1 scores. RESULTS We found that Path-BigBird and Clinical BigBird outperformed the HiSAN in all tasks. Clinical BigBird performed better on the site and laterality tasks. Versions of the Path-BigBird model performed best on the two most difficult tasks: subsite (micro F 1 score of 72.53, macro F 1 score of 35.76) and histology (micro F 1 score of 80.96, macro F 1 score of 37.94). The largest performance gains over the HiSAN model were for histology, for which a Path-BigBird model increased the micro F 1 score by 1.44 points and the macro F 1 score by 3.55 points. Overall, the results suggest that a Path-BigBird model with a vocabulary derived from wellcurated and deidentified data is the best-performing model. CONCLUSION The Path-BigBird pathology transformer model improves automated information extraction from pathology reports. Although Path-BigBird outperforms Clinical BigBird and HiSAN, these less computationally expensive models still have utility when resources are constrained.

60 APPLIED LIFE SCIENCES↗

Ensemble models for circuit topology estimation, fault detection and classification in distribution systems

This paper presents a methodology for simultaneous fault detection, classification, and topology estimation for adaptive protection of distribution systems. The methodology estimates the probability of the occurrence of each one of these events by using a hybrid structure that combines three sub-systems, a convolutional neural network for topology estimation, a fault detection based on predictive residual analysis, and a standard support vector machine with probabilistic output for fault classification. The input to all these sub-systems is the local voltage and current measurements. A convolutional neural network uses these local measurements in the form of sequential data to extract features and estimate the topology conditions. The fault detector is constructed with a Bayesian stage (a multitask Gaussian process) that computes a predictive distribution (assumed to be Gaussian) of the residuals using the input. Since the distribution is known, these residuals can be transformed into a Standard distribution, whose values are then introduced into a one-class support vector machine. The structure allows using a one-class support vector machine without parameter cross-validation, so the fault detector is fully unsupervised. Finally, a support vector machine uses the input to perform the classification of the fault types. All three sub-systems can work in a parallel setup for both performance and computation efficiency. In conclusion, we test all three sub-systems included in the structure on a modified IEEE123 bus system, and we compare and evaluate the results with standard approaches.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Masked Particle Modeling on Sets: Towards Self-Supervised High Energy Physics Foundation Models

Abstract We propose masked particle modeling (MPM) as a self-supervised method for learning generic, transferable, and reusable representations on unordered sets of inputs for use in high energy physics (HEP) scientific data. This work provides a novel scheme to perform masked modeling based pre-training to learn permutation invariant functions on sets. More generally, this work provides a step towards building large foundation models for HEP that can be generically pre-trained with self-supervised learning and later fine-tuned for a variety of down-stream tasks. In MPM, particles in a set are masked and the training objective is to recover their identity, as defined by a discretized token representation of a pre-trained vector quantized variational autoencoder. We study the efficacy of the method in samples of high energy jets at collider physics experiments, including studies on the impact of discretization, permutation invariance, and ordering. We also study the fine-tuning capability of the model, showing that it can be adapted to tasks such as supervised and weakly supervised jet classification, and that the model can transfer efficiently with small fine-tuning data sets to new classes and new data domains.

Heinrich, Lukas (ORCID:0000000240487584)↗

SNM Radiation Signature Classification Using Different Semi-Supervised Machine Learning Models

The timely detection of special nuclear material (SNM) transfers between nuclear facilities is an important monitoring objective in nuclear nonproliferation. Persistent monitoring enabled by successful detection and characterization of radiological material movements could greatly enhance the nuclear nonproliferation mission in a range of applications. Supervised machine learning can be used to signal detections when material is present if a model is trained on sufficient volumes of labeled measurements. However, the nuclear monitoring data needed to train robust machine learning models can be costly to label since radiation spectra may require strict scrutiny for characterization. Therefore, this work investigates the application of semi-supervised learning to utilize both labeled and unlabeled data. As a demonstration experiment, radiation measurements from sodium iodide (NaI) detectors are provided by the Multi-Informatics for Nuclear Operating Scenarios (MINOS) venture at Oak Ridge National Laboratory (ORNL) as sample data. Anomalous measurements are identified using a method of statistical hypothesis testing. After background estimation, an energy-dependent spectroscopic analysis is used to characterize an anomaly based on its radiation signatures. In the absence of ground-truth information, a labeling heuristic provides data necessary for training and testing machine learning models. Supervised logistic regression serves as a baseline to compare three semi-supervised machine learning models: co-training, label propagation, and a convolutional neural network (CNN). In each case, the semi-supervised models outperform logistic regression, suggesting that unlabeled data can be valuable when training and demonstrating value in semi-supervised nonproliferation implementations.

98 NUCLEAR DISARMAMENT, SAFEGUARDS, AND PHYSICAL P↗

Deep Learning for Fish Identification from Sonar Data: CRADA 481 [Abstract only]

To help solve the challenges of hydropower energy production related to the potential for eel injury and mortality from passage through hydropower turbines, we will develop a deep learning method for identifying migrating eels from imaging sonar. This project continues with a prior project conducted by the Pacific Northwest National Laboratory (PNNL) and the Electric Power Research Institute (EPRI) in FY2018-2019. The proposed method employs Convolution Neural Network (CNN), a powerful deep learning method for image classification, to distinguish between images of eels and non-eel moving objects. We propose to collect more laboratory data and add more existing field data to train a powerful deep learning model. In addition to eels and sticks as classified in previous studies, we will add images containing several non-eel fish species and macrophyte mats to the training data. A multi-class classification model will be developed to distinguish these objects. Object detection algorithm will be explored and developed to locate and identify multiple objects in each sonar frame. Motion analysis will be performed to track the movement of objects in sonar video clips. We will also improve the data conversion algorithm so that it can read in both DIDSON and ARIS (both are imaging sonars developed by Sound Metrics Corp) data files and convert them to images with comparably high resolution, regardless of the varying detection ranges in different environments. The developed algorithms will be packaged as a software with a graphic user interface. The software will be evaluated by external collaborators in the field. The developed framework can be generalized for automatic monitoring of fish passage and migration using other imaging sonars like ARIS and will benefit the design and operation of ecologically friendly hydroelectric projects. The developed wavelet and CNN model configuration parameters can potentially be transferred to lamprey detection in similar riverine environments.

13 HYDRO ENERGY↗

Characterizing Quantum Classifier Utility in Natural Language Processing Workflows

Quantum Natural Language Processing (QNLP) develops natural language processing (NLP) models for deployment on quantum computers. We explore feature and data prototype selection techniques to address challenges posed by encoding high dimensional features. Our study builds quantum circuit classifiers that includes classical feature pre-processing, quantum embedding and quantum model training. The quantum models are built on 4 or 6 qubits and the quantum neural network (QNN) uses the established bricklayer design. We compare the dependence of model performance (in terms of accuracy and F1 scores) on feature length, embedding gates and parameterized unitary design. We compare the performance of quantum machine learning models to classical convolution neural network model (CNN) on binary and multi-class classification tasks using two datasets of synthetic features and labels. The first is the ECP-CANDLE P3B3 dataset a corpus of synthetically generated cancer pathology reports. The second dataset is extracted from well-known benchmark dataset (MADELON) - features are generated with a combination of informative, repeated and uninformative features. Both datasets are used for binary classification and multi-class classification with 3 classes. We observe robust, accurate performance from all models on the binary classification tasks, but multiclass classification is a challenge for the quantum models-there is a notable decrease in accuracy when using 3 classes. Overall the performance is comparable in terms of recall and accuracy between QNNs and CNNs, even with large datasets. These results provide a point of comparison between quantum and classical models on real-world datasets.

Hamilton, Kathleen↗

Robust Machine Learning Inference from X-ray Absorption Near Edge Spectra through Featurization

X-ray absorption spectroscopy (XAS) is a commonly employed technique for characterizing functional materials. In particular, X-ray absorption near edge spectra (XANES) encode local coordination and electronic information, and machine learning approaches to extract this information are of significant interest. To date, most ML approaches for XANES have primarily focused on using the raw spectral intensities as input, overlooking the potential benefits of incorporating spectral transformations and dimensionality reduction techniques into ML predictions. Here, in this work, we focused on systematically comparing the impact of different featurization methods on the performance of ML models for XAS analysis. We evaluated the classification and regression capabilities of these models on computed data sets and validated their performance on previously unseen experimental data sets. Our analysis revealed an intriguing discovery: the cumulative distribution function feature achieves both high prediction accuracy and exceptional transferability. This remarkably robust performance can be attributed to its tolerance to horizontal shifts in the spectra, which is crucial when validating models using experimental data. While this work exclusively focuses on XANES analysis, we anticipate that the methodology presented here will hold promise as a versatile asset to the broader spectroscopy community.

36 MATERIALS SCIENCE↗

Finding predictive models for singlet fission by machine learning

Singlet fission (SF), the conversion of one singlet exciton into two triplet excitons, could significantly enhance solar cell efficiency. Molecular crystals that undergo SF are scarce. Computational exploration may accelerate the discovery of SF materials. However, many-body perturbation theory (MBPT) calculations of the excitonic properties of molecular crystals are impractical for large-scale materials screening. We use the sure-independence-screening-and-sparsifying-operator (SISSO) machine-learning algorithm to generate computationally efficient models that can predict the MBPT thermodynamic driving force for SF for a dataset of 101 polycyclic aromatic hydrocarbons (PAH101). SISSO generates models by iteratively combining physical primary features. The best models are selected by linear regression with cross-validation. The SISSO models successfully predict the SF driving force with errors below 0.2 eV. Based on the cost, accuracy, and classification performance of SISSO models, we propose a hierarchical materials screening workflow. Three potential SF candidates are found in the PAH101 set.

36 MATERIALS SCIENCE↗

Combinatorial Evaluation of Physical Feature Engineering, Classical Machine Learning, and Deep Learning Models for Synchrophasor Data at Scale

A major objective of the project was to train and evaluate the effectiveness of multiple event and anomaly detection, identification and classification deep temporal learning models for processing of real-time phasor measurement unit (PMU) data streams. A vast dataset, consisting of two years of phasor measurements from all three U.S. Interconnections, was curated and released by the Department of Energy (DOE) through Pacific Northwest National Laboratory (PNNL). The dataset also included an event log that provided event times and types (e.g. generator trips, line trips, planned service events, transformer operations, etc.). Our analysis of this dataset addressed six (6) of the eleven (11) research priorities identified in Funding Opportunity Announcement (FOA) DE-FOA-0001861 “Big Data Analysis of Synchrophasor Data” (FOA 1861). Rather than being limited to pre-determined specific algorithms, this project relied on the uniquely structured, highly performant underlying time series database capabilities of the PredictiveGrid platform to assess the vast dataset utilizing a wide variety of algorithms.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Language Model For Earth Science: Exploring Potential Downstream Applications As Well As Current Challenges

The use of deep learning techniques to build transformer language models such as SciBERT and GPT3 have transformed the natural language technology (NLT) landscape. These new NLTs are being used in speech to text and vice versa, auto-mated text classification, sentiment analysis, topic modeling, text summarization, and cognitive assistants. While Earth science has no shortage of unstructured data such as journal and conference papers, little efforts have focused on harnessing NLTs for knowledge extraction and supporting the scientific process. This paper surveys the use of language models in different science. BERT-E, a new Earth science-specific language model, is presented. BERT-E is generated using a transfer learning solution. A language model that has already been trained for general Science (SciBERT) is fine-tuned using abstracts and full text extracted from various Earth science-related articles. A downstream keywords classification application is used for evaluation, and the use of BERT-E shows improved performance. The need to develop a robust set of benchmarks in evaluating the language model such as BERT-E is discussed. Finally, example applications are presented to inspire additional ideas for applications using domain-specific language models.

R Ramachandran↗

Initial Verification of GEOS-4 Aerosols Using CALIPSO and MODIS: Scene Classification

A-train sensors such as MODIS and MISR provide column aerosol properties, and in the process a means of estimating aerosol type (e.g. smoke vs. dust). Correct classification of aerosol type is important because retrievals are often dependent upon selection of the right aerosol model. In addition, aerosol scene classification helps place the retrieved products in context for comparisons and analysis with aerosol transport models. The recent addition of CALIPSO to the A-train now provides a means of classifying aerosol distribution with altitude. CALIPSO level 1 products include profiles of attenuated backscatter at 532 and 1064 nm, and depolarization at 532 nm. Backscatter intensity, wavelength ratio, and depolarization provide information on the vertical profile of aerosol concentration, size, and shape. Thus similar estimates of aerosol type using MODIS or MISR are possible with CALIPSO, and the combination of data from all sensors provides a means of 3D aerosol scene classification. The NASA Goddard Earth Observing System general circulation model and data assimilation system (GEOS-4) provides global 3D aerosol mass for sulfate, sea salt, dust, and black and organic carbon. A GEOS-4 aerosol scene classification algorithm has been developed to provide estimates of aerosol mixtures along the flight track for NASA's Geoscience Laser Altimeter System (GLAS) satellite lidar. GLAS launched in 2003 and did not have the benefit of depolarization measurements or other sensors from the A-train. Aerosol typing from GLAS data alone was not possible, and the GEOS-4 aerosol classifier has been used to identify aerosol type and improve the retrieval of GLAS products. Here we compare 3D aerosol scene classification using CALIPSO and MODIS with the GEOS-4 aerosol classifier. Dust, smoke, and pollution examples will be discussed in the context of providing an initial verification of the 3D GEOS-4 aerosol products. Prior model verification has only been attempted with surface mass comparisons and column optical depth from AERONET and MODIS.

Welton, Ellsworth J.↗

The Dark Energy Survey Supernova Programme: Modelling Selection Efficiency and Observed Core-collapse Supernova Contamination

The analysis of current and future cosmological surveys of Type Ia supernovae (SNe Ia) at high redshift depends on the accuratephotometric classification of the SN events detected. Generating realistic simulations of photometric SN surveys constitutes anessential step for training and testing photometric classification algorithms, and for correcting biases introduced by selectioneffects and contamination arising from core-collapse SNe in the photometric SN Ia samples. We use published SN time-seriesspectrophotometric templates, rates, luminosity functions, and empirical relationships between SNe and their host galaxies toconstruct a framework for simulating photometric SN surveys. We present this framework in the context of the Dark EnergySurvey (DES) 5-yr photometric SN sample, comparing our simulations of DES with the observed DES transient populations.We demonstrate excellent agreement in many distributions, including Hubble residuals, between our simulations and data.We estimate the core collapse fraction expected in the DES SN sample after selection requirements are applied and beforephotometric classification. After testing different modelling choices and astrophysical assumptions underlying our simulation,we find that the predicted contamination varies from 7.2 to 11.7 per cent, with an average of 8.8 per cent and an r.m.s. of 1.1 percent. Our simulations are the first to reproduce the observed photometric SN and host galaxy properties in high-redshift surveyswithout fine-tuning the input parameters. The simulation methods presented here will be a critical component of the cosmologyanalysis of the DES photometric SN Ia sample: correcting for biases arising from contamination, and evaluating the associatedsystematic uncertainty.

M Vincenzi↗

Deep-learning-aided forward optical coherence tomography endoscope for percutaneous nephrostomy guidance

Percutaneous renal access is the critical initial step in many medical settings. In order to obtain the best surgical outcome with minimum patient morbidity, an improved method for access to the renal calyx is needed. In our study, we built a forward-view optical coherence tomography (OCT) endoscopic system for percutaneous nephrostomy (PCN) guidance. Porcine kidneys were imaged in our experiment to demonstrate the feasibility of the imaging system. Three tissue types of porcine kidneys (renal cortex, medulla, and calyx) can be clearly distinguished due to the morphological and tissue differences from the OCT endoscopic images. To further improve the guidance efficacy and reduce the learning burden of the clinical doctors, a deep-learning-based computer aided diagnosis platform was developed to automatically classify the OCT images by the renal tissue types. Convolutional neural networks (CNN) were developed with labeled OCT images based on the ResNet34, MobileNetv2 and ResNet50 architectures. Nested cross-validation and testing was used to benchmark the classification performance with uncertainty quantification over 10 kidneys, which demonstrated robust performance over substantial biological variability among kidneys. ResNet50-based CNN models achieved an average classification accuracy of 82.6%±3.0%. The classification precisions were 79%±4% for cortex, 85%±6% for medulla, and 91%±5% for calyx and the classification recalls were 68%±11% for cortex, 91%±4% for medulla, and 89%±3% for calyx. Interpretation of the CNN predictions showed the discriminative characteristics in the OCT images of the three renal tissue types. The results validated the technical feasibility of using this novel imaging platform to automatically recognize the images of renal tissue structures ahead of the PCN needle in PCN surgery.

Wang, Chen↗

The Dark Energy Survey supernova programme: modelling selection efficiency and observed core-collapse supernova contamination

ABSTRACT The analysis of current and future cosmological surveys of Type Ia supernovae (SNe Ia) at high redshift depends on the accurate photometric classification of the SN events detected. Generating realistic simulations of photometric SN surveys constitutes an essential step for training and testing photometric classification algorithms, and for correcting biases introduced by selection effects and contamination arising from core-collapse SNe in the photometric SN Ia samples. We use published SN time-series spectrophotometric templates, rates, luminosity functions, and empirical relationships between SNe and their host galaxies to construct a framework for simulating photometric SN surveys. We present this framework in the context of the Dark Energy Survey (DES) 5-yr photometric SN sample, comparing our simulations of DES with the observed DES transient populations. We demonstrate excellent agreement in many distributions, including Hubble residuals, between our simulations and data. We estimate the core collapse fraction expected in the DES SN sample after selection requirements are applied and before photometric classification. After testing different modelling choices and astrophysical assumptions underlying our simulation, we find that the predicted contamination varies from 7.2 to 11.7 per cent, with an average of 8.8 per cent and an r.m.s. of 1.1 per cent. Our simulations are the first to reproduce the observed photometric SN and host galaxy properties in high-redshift surveys without fine-tuning the input parameters. The simulation methods presented here will be a critical component of the cosmology analysis of the DES photometric SN Ia sample: correcting for biases arising from contamination, and evaluating the associated systematic uncertainty.

79 ASTRONOMY AND ASTROPHYSICS↗

A Framework for Identifying Building Energy Models of Localized Utility Service Areas Using Smart Meter Data

Bottom-up load modeling of buildings offers a versatile approach to simulating baseline demand and scenarios of future technology evolution and adoption at the individual building level. This capability is essential to understanding how future load shapes may change with the adoption of electric equipment and vehicles, particularly as it relates to grid planning and infrastructure investments. Traditionally, grid planning techniques have used historical load data to predict future load and infrastructure needs. However, with the anticipated rise in adoption of electrification technologies such as heat pumps and electric vehicles, historical data become less reliable predictors of the future. By employing ResStock, a high-fidelity building stock modeling tool, we can fine-tune electrification scenarios and aggregate models to represent varying geographic resolutions of the grid system, while considering the underlying features of homes. This may enable a more accurate and responsive approach to anticipate and plan for the evolving landscape of energy demands. We present a new framework that leverages building stock energy modeling to identify building models that align with the load shapes and housing attributes of buildings with AMI data. This approach applies two model layers: (1) a classification step that identifies the presence of air conditioning, electric heating, and electric water heating, and (2) an optimization routine that identifies building energy models aligning with load profile data from advanced metering infrastructure meters. This report demonstrates one approach to deploying this framework, and presents results for three test cases that use both modeled and AMI data to assess performance. For a test case using AMI data in Fort Collins, Colorado, we observed a median monthly electricity load CV-RMSE of 16.6%, and a top ten daily heating and cooling median absolute percent error of 7.7% and 8.3%, respectively. For each AMI meter, we identify a set of potential energy models so that downstream use-cases can account for uncertainty driven by variability of baseline technologies and occupant behavior, which impact the response to electrification and energy efficiency scenarios. Our results indicate that ResStock has potential as a scalable solution for modeling residential energy demand at local grid resolutions. Its performance depends on location-specific factors, underlying building characteristics, and the level of aggregation, offering a path towards more precise and adaptive distribution grid planning for the evolving energy landscape.

24 POWER TRANSMISSION AND DISTRIBUTION↗

An overview of the neuron ring model

The Neuron Ring model employs an avalanche structure with two important distinctions at the neuron level. Each neuron has two memory latches; one traps maximum neuronal activation during pattern presentation, and the other records the time of latch content change. The latches filter short term memory. In the process, they preserve length 1 snapshots of activation theory history. The model finds utility in pattern classification. Its synaptic weights are first conditioned with sample spectra. The model then receives a test or unknown signal. The objective is to identify the sample closest to the test signal. Class decision follows complete presentation of the test data. The decision maker relies exclusively on the latch contents. Presented here is an overview of the Neuron Ring at the seminar level.

Taber, Rod↗

Machine Learning for the Prediction of Local Asteroid Damages

Risk assessment studies of local asteroid hazards traditionally simulate the physics of meteors with engineering models tailored to analyze tens-of-millions of scenarios. However, these simplified approaches still need to solve time-dependent ODEs to model the entry process and the resulting ground damage. With a computational cost of O(0.01 CPU.s) per scenario, simulating these large numbers of potential entry conditions in risk assessment studies can take several days on local computers. To improve computational efficiency, we propose in this paper an orthogonal approach based on machine learning models to predict the size of damaged areas given a list of entry parameters. We train 5 machine learning methods and compare the predictions to the outputs of the PAIR model, first only with primitive entry condition variables, and then with more advanced features. We find that complex models like neural networks are well-suited to estimate blast hazards, while simpler linear models can accurately assess thermal damage. For both types of hazards, the radii of damaged areas can be predicted with around 10% average errors and a coefficient of determination (R2) of 0.99. The CPU time is decreased by a factor O(10 3 ) compared to the PAIR model, which enables the simulation of millions of scenarios in minutes, on a local computer. We then use the same machine learning approaches for a classification task where the models are trained to predict if an asteroid will produce a given level of damage. Results show that complex models like the gradient boosting classifier and the neural network can perform this task with 98% accuracy. Beyond surrogate models, we finally incorporate the machine learning algorithms to the state-of-the-art Shapley sensitivity analysis and present a ranking of the entry parameters based on their contributions to ground damages.

SMD↗