Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “feature selection”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

Machine Learning–Augmented Laser-Induced Breakdown Spectroscopy for Spectral Discrimination of Iron Oxalates

Enhanced characterization and phase identification of post-PUREX Pu Oxalates (PuOXA) are pivotal for nonproliferation and pre-detonation nuclear forensics. Despite significant advances in the characterization of PuO 2 samples, little is known about the impact of both the chemical structure and oxidation states of PuOXA (i.e., Pu(III) and Pu(IV)) have on optical emission signatures. Here, we demonstrate the analytical capabilities of laser-induced breakdown spectroscopy (LIBS) applied to Fe(II) and Fe(III) oxalate samples as surrogates for PuOXA, highlighting the discriminating features in the LIBS emission spectra arising from differences in the oxidation states within mixed FeOXA samples. We report the enhancement of spectral feature selection using Principal Component Analysis (PCA), which enables the analytical superiority of machine learning algorithms such as Linear Discriminant Analysis (LDA), Quadratic Discriminant Analysis (QDA), Partial Least Squares Regression (PLSR), Support Vector Regression (SVR), and Random Forest Regression (RFR) over conventional univariate techniques for phase discrimination and chemometric analysis. Cluster analysis revealed how both matrix effects and laser ablation influence cluster separability by introducing spectral artifacts that misdirect the maximization of variance. PCA-selected emission lines were used in the regression models, demonstrating that both univariate and multivariate linear regression models (i.e., PLSR and SVR) can achieve acceptable performance, with machine learning models outperforming conventional calibration regressions. Furthermore, the application of non-linearly activated PCA-selected emission lines illustrates how simplifying the data while retaining captured variance enables the use of less complex and more computationally efficient models. Furthermore, this is particularly evident in the underperformance of RFR, which suffers from increased computational costs and overfitting owing to its high complexity.

Oxalates↗

Enhancing dimensionality prediction in hybrid metal halides via feature engineering and class-imbalance mitigation

We present a machine learning (ML) framework for predicting the structural dimensionality of hybrid metal halides (HMHs), including organic-inorganic perovskites, using a combination of chemically-informed feature engineering and advanced class-imbalance handling techniques. This study is motivated by the small and highly imbalanced nature of experimentally available HMH datasets, which limits the applicability and reliability of conventional ML approaches. The dataset, consisting of 494 HMH structures, is highly imbalanced across dimensionality classes (0D, 1D, 2D, 3D), posing significant challenges to predictive modeling. To mitigate this limitation, the dataset was augmented to 1336 samples using the synthetic minority oversampling technique, enabling improved learning of underrepresented dimensionality classes while preserving chemically meaningful feature relationships. We developed interaction-based descriptors designed to capture coupled steric and polarity effects relevant to dimensionality prediction, which are not readily captured by standard single-parameter or composition-only descriptors. These descriptors are integrated into a multi-stage workflow combining feature selection, ensemble stacking, and performance optimization. Our approach significantly improves F1-scores for underrepresented classes, achieving robust cross-validation performance across all dimensionalities. This work demonstrates a generalizable strategy for extracting reliable and interpretable structure–dimensionality relationships from limited experimental data, enabling pre-synthesis screening of organic cations and providing a practical blueprint for small-data ML in hybrid materials systems.

36 MATERIALS SCIENCE↗

Deep graph representations embed network information for robust disease marker identification

We report that the accurate disease diagnosis and prognosis based on omics data rely on the effective identification of robust prognostic and diagnostic markers that reflect the states of the biological processes underlying the disease pathogenesis and progression. In this article, we present GCNCC, a Graph Convolutional Network-based approach for Clustering and Classification, that can identify highly effective and robust network-based disease markers. Based on a geometric deep learning framework, GCNCC learns deep network representations by integrating gene expression data with protein interaction data to identify highly reproducible markers with consistently accurate prediction performance across independent datasets possibly from different platforms. GCNCC identifies these markers by clustering the nodes in the protein interaction network based on latent similarity measures learned by the deep architecture of a graph convolutional network, followed by a supervised feature selection procedure that extracts clusters that are highly predictive of the disease state. By benchmarking GCNCC based on independent datasets from different diseases (psychiatric disorder and cancer) and different platforms (microarray and RNA-seq), we show that GCNCC outperforms other state-of-the-art methods in terms of accuracy and reproducibility.

59 BASIC BIOLOGICAL SCIENCES↗

Z-Sequence: photometric redshift predictions for galaxy clusters with sequential random k-nearest neighbours

ABSTRACT We introduce Z-Sequence, a novel empirical model that utilizes photometric measurements of observed galaxies within a specified search radius to estimate the photometric redshift of galaxy clusters. Z-Sequence itself is composed of a machine learning ensemble based on the k-nearest neighbours algorithm. We implement an automated feature selection strategy that iteratively determines appropriate combinations of filters and colours to minimize photometric redshift prediction error. We intend for Z-Sequence to be a standalone technique but it can be combined with cluster finders that do not intrinsically predict redshift, such as our own DEEP-CEE. In this proof-of-concept study, we train, fine-tune, and test Z-Sequence on publicly available cluster catalogues derived from the Sloan Digital Sky Survey. We determine the photometric redshift prediction error of Z-Sequence via the median value of |Δ$z$|/(1 + $z$) (across a photometric redshift range of 0.05 ≤ $z$ ≤ 0.6) to be ∼0.01 when applying a small search radius. The photometric redshift prediction error for test samples increases by 30–50 per cent when the search radius is enlarged, likely due to line-of-sight interloping galaxies. Eventually, we aim to apply Z-Sequence to upcoming imaging surveys such as the Legacy Survey of Space and Time to provide photometric redshift estimates for large samples of as yet undiscovered and distant clusters.

Chan, Matthew C.↗

Photometric redshift estimation of galaxies in the DESI Legacy Imaging Surveys

ABSTRACT The accurate estimation of photometric redshifts plays a crucial role in accomplishing science objectives of the large survey projects. Template-fitting and machine learning are the two main types of methods applied currently. Based on the training set obtained by cross-correlating the DESI Legacy Imaging Surveys DR9 galaxy catalogue and the SDSS DR16 galaxy catalogue, the two kinds of methods are used and optimized, such as eazy for template-fitting approach and catboost for machine learning. Then, the created models are tested by the cross-matched samples of the DESI Legacy Imaging Surveys DR9 galaxy catalogue with LAMOST DR7, GAMA DR3, and WiggleZ galaxy catalogues. Moreover, three machine learning methods (catboost, Multi-Layer Perceptron, and Random Forest) are compared; catboost shows its superiority for our case. By feature selection and optimization of model parameters, catboost can obtain higher accuracy with optical and infrared photometric information, the best performance ($\rm MSE=0.0032$, σNMAD = 0.0156, and $O=0.88{{\ \rm per\ cent}}$) with g ≤ 24.0, r ≤ 23.4, and z ≤ 22.5 is achieved. But eazy can provide more accurate photometric redshift estimation for high redshift galaxies, especially beyond the redshift range of training sample. Finally, we finish the redshift estimation of all DESI Legacy Imaging Surveys DR9 galaxies with catboost and eazy, which will contribute to the further study of galaxies and their properties.

Astronomy & Astrophysics↗

Predicting band gaps and band-edge positions of oxide perovskites using density functional theory and machine learning

Density functional theory (DFT) within the local or semilocal density approximations, i.e., the local density approximation (LDA) or generalized gradient approximation (GGA), has become a workhorse in the electronic structure theory of solids, being extremely fast and reliable for energetics and structural properties, yet remaining highly inaccurate for predicting band gaps of semiconductors and insulators. The accurate prediction of band gaps using first-principles methods is time consuming, requiring hybrid functionals, quasiparticle GW, or quantum Monte Carlo methods. Efficiently correcting DFT-LDA/GGA band gaps and unveiling the main chemical and structural factors involved in this correction is desirable for discovering novel materials in high-throughput calculations. In this direction, we, in this study, use DFT and machine learning techniques to correct band gaps and band-edge positions of a representative subset of ABO 3 perovskite oxides. Relying on the results of HSE06 hybrid functional calculations as target values of band gaps, we find a systematic band-gap correction of ~1.5 eV for this class of materials, where ~1 eV comes from downward shifting the valence band and ~0.5 eV from uplifting the conduction band. The main chemical and structural factors determining the band-gap correction are determined through a feature selection procedure.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

Explainable machine learning for hydrogen diffusion in metals and random binary alloys

Hydrogen diffusion in metals and alloys plays an important role in the discovery of new materials for fuel cell and energy storage technology. While analytic models use hand-selected features that have clear physical ties to hydrogen diffusion, they often lack accuracy when making quantitative predictions. Machine learning models are capable of making accurate predictions, but their inner workings are obscured, rendering it unclear which physical features are truly important. To develop interpretable machine learning models to predict the activation energies of hydrogen diffusion in metals and random binary alloys, we create a database for physical and chemical properties of the species and use it to fit six machine learning models. Our models achieve root-mean-squared errors between 98–119 meV on the testing data and accurately predict that elemental Ru has a large activation energy, while elemental Cr and Fe have small activation energies. By analyzing the feature importances of these fitted models, we identify relevant physical properties for predicting hydrogen diffusivity. While metrics for measuring the individual feature importances for machine learning models exist, correlations between the features lead to disagreement between models and limit the conclusions that can be drawn. Instead grouped feature importance, formed by combining the features via their correlations, agree across the six models and reveal that the two groups containing the packing factor and electronic specific heat are particularly significant for predicting hydrogen diffusion in metals and random binary alloys. In conclusion, this framework allows us to interpret machine learning models and enables rapid screening of new materials with the desired rates of hydrogen diffusion.

36 MATERIALS SCIENCE↗

SRBench++: Principled Benchmarking of Symbolic Regression With Domain-Expert Interpretation

Symbolic regression searches for analytic expressions that accurately describe studied phenomena. The main promise of this approach is that it may return an interpretable model that can be insightful to users, while maintaining high accuracy. The current standard for benchmarking these algorithms is SRBench, which evaluates methods on hundreds of datasets that are a mix of real-world and simulated processes spanning multiple domains. At present, the ability of SRBench to evaluate interpretability is limited to measuring the size of expressions on real-world data, and the exactness of model forms on synthetic data. In practice, model size is only one of many factors used by subject experts to determine how interpretable a model truly is. Furthermore, SRBench does not characterize algorithm performance on specific, challenging sub-tasks of regression such as feature selection and evasion of local minima. In this work, we propose and evaluate an approach to benchmarking SR algorithms that addresses these limitations of SRBench by 1) incorporating expert evaluations of interpretability on a domain-specific task, and 2) evaluating algorithms over distinct properties of data science tasks. We evaluate 12 modern symbolic regression algorithms on these benchmarks and present an in-depth analysis of the results, discuss current challenges of symbolic regression algorithms and highlight possible improvements for the benchmark itself.

97 MATHEMATICS AND COMPUTING↗

Pursuing sources of heterogeneity in modeling clustered population

Abstract Researchers often have to deal with heterogeneous population with mixed regression relationships, increasingly so in the era of data explosion. In such problems, when there are many candidate predictors, it is not only of interest to identify the predictors that are associated with the outcome, but also to distinguish the true sources of heterogeneity , that is, to identify the predictors that have different effects among the clusters and thus are the true contributors to the formation of the clusters. We clarify the concepts of the source of heterogeneity that account for potential scale differences of the clusters and propose a regularized finite mixture effects regression to achieve heterogeneity pursuit and feature selection simultaneously. We develop an efficient algorithm and show that our approach can achieve both estimation and selection consistency. Simulation studies further demonstrate the effectiveness of our method under various practical scenarios. Three applications are presented, namely, an imaging genetics study for linking genetic factors and brain neuroimaging traits in Alzheimer's disease, a public health study for exploring the association between suicide risk among adolescents and their school district characteristics, and a sport analytics study for understanding how the salary levels of baseball players are associated with their performance and contractual status.

Li, Yan↗

Switchgrass Steroidal Saponins Reduce Fungal Disease but Decrease Yeast Fermentation Yield

Increasing the production of bioproducts from lignocellulosic feedstocks requires improvement in both field production and biorefinery efficiency. When plant traits arise that improve field production but decrease biofuel yield, these trade-offs can represent challenges in the entire production process. To examine trade-offs between field and production traits, we examined factors underlying switchgrass resistance to fungal rust pathogens in field conditions and factors that impede yeast fermentation in the lab using repeated measurements on a switchgrass genetic diversity panel. We found that the same switchgrass genotypes that showed high fungal pathogen resistance also showed recalcitrance to yeast fermentation. These switchgrass genotypes were mostly from the Atlantic genetic group, which had high levels of specialized metabolites of the saponin class. Among 1589 metabolites identified through metabolomics, we found that saponins were among the most likely to explain variation in both rust infection and fermentation yield using random forest feature selection, and that only four of these were sufficient to explain 57.9% of the variation in rust susceptibility. Through follow-up testing in recalcitrant biomass, we found that the bacterium Zymomonas mobilis does not suffer the same inhibition as the yeast Saccharomyces cerevisiae, and that the addition of ergosterol (thought to be the fungal cellular target of saponin inhibition) rescues yeast fermentation. Several lines of evidence point to a central role for saponins as key metabolites protecting switchgrass from fungal pathogens and interfering with yeast fermentation, underscoring an ongoing need for collaboration between plant breeders and biofuel production scientists.

VanWallendael, Acer [North Carolina State Universi↗

Materials characterization: Can artificial intelligence be used to address reproducibility challenges?

Material characterization techniques are widely used to characterize the physical and chemical properties of materials at the nanoscale and, thus, play central roles in material scientific discoveries. However, the large and complex datasets generated by these techniques often require significant human effort to interpret and extract meaningful physicochemical insights. Artificial intelligence (AI) techniques such as machine learning (ML) have the potential to improve the efficiency and accuracy of surface analysis by automating data analysis and interpretation. In this perspective paper, we review the current role of AI in surface analysis and discuss its future potential to accelerate discoveries in surface science, materials science, and interface science. We highlight several applications where AI has already been used to analyze surface analysis data, including the identification of crystal structures from XRD data, analysis of XPS spectra for surface composition, and the interpretation of TEM and SEM images for particle morphology and size. We also discuss the challenges and opportunities associated with the integration of AI into surface analysis workflows. These include the need for large and diverse datasets for training ML models, the importance of feature selection and representation, and the potential for ML to enable new insights and discoveries by identifying patterns and relationships in complex datasets. Most importantly, AI analyzed data must not just find the best mathematical description of the data, but it must find the most physical and chemically meaningful results. In addition, the need for reproducibility in scientific research has become increasingly important in recent years. The advancement of AI, including both conventional and the increasing popular deep learning, is showing promise in addressing those challenges by enabling the execution and verification of scientific progress. By training models on large experimental datasets and providing automated analysis and data interpretation, AI can help to ensure that scientific results are reproducible and reliable. Although integration of knowledge and AI models must be considered for the transparency and interpretability of models, the incorporation of AI into the data collection and processing workflow will significantly enhance the efficiency and accuracy of various surface analysis techniques and deepen our understanding at an accelerated pace.

Materials Science↗

Codon2Vec v1.0

Background: Codon2Vec is an embedding neural network that predicts 'high' or 'low' gene expression directly from the protein-coding sequences. Embedding neural networks are commonly used for natural language processing (NLP) applications. Analogous to how an English sentence is a string of words, a gene can be thought of as a string of codons. Similar to how NLP neural networks model English sentences as a non-random sequence of words, we considered a coding sequence as a non-random non-overlapping array of codons (k-mers of length = 3). Value Proposition: - Codon2Vec achieved a high median AUC-ROC score of 83.8% when trained and applied to transcriptomic data from 300 fungal species - Unlike Codo2Vec, conventional methods predicting for expression based on codon usage rely on a priori knowledge of optimal codons or a set of reference genes. - Unlike Codon2vec, these methods do not account for the effect of codon order on gene expression. - Codon2Vec neural network bypasses the need for artisanal feature selection step that is necessary for traditional machine learning models.

Wint, Rhondene↗

MACAW v1.0

The ability to embed molecules in a numeric space is essential in order to build mathematical and machine-learning models describing molecular properties or other processes affected by molecules. MACAW is a cheminformatic tool that allows embedding small molecules into a multidimensional numeric space. In the embedding, each molecule is assigned a numeric vector that captures information of the molecule in relation to other molecules, and that vector can be used as input to mathematical models. Molecules that are more similar to each other are embedded closer in this numeric space, whereas molecules that are more different are embedded further away. One advantage of the MACAW embedding technology compared to established alterantives is that it is fast and does not require extensive computational resources or expertise. In particular, MACAW embeddings can be used as input to mathematical models without the need for variable cleaning or feature selection, saving time and simplifying their use. On the other hand, MACAW also contains methods to generate new molecules and to recommend new molecules satisfying a desired molecular property. The generation of new molecules can be biased based on an input set of molecules, effectively generating molecular diversity around it. In its turn, MACAW's molecular recommendation tool is a novel method for evolving molecules in silico towards a desired molecular specification. In this method, the biased molecular generator tool is applied iteratively in combination with a molecular selection step. As a result, in each iteration the molecules selected by the software are increasingly closer to the desired specification. Both the molecular generation and the molecular recommendation tools are very fast, efficient, and intuitive to use.

Roger, VincentBlay↗

Identification of integrated proteomics and transcriptomics signature of alcohol-associated liver disease using machine learning

Distinguishing between alcohol-associated hepatitis (AH) and alcohol-associated cirrhosis (AC) remains a diagnostic challenge. In this study, we used machine learning with transcriptomics and proteomics data from liver tissue and peripheral mononuclear blood cells (PBMCs) to classify patients with alcohol-associated liver disease. The conditions in the study were AH, AC, and healthy controls. We processed 98 PBMC RNAseq samples, 55 PBMC proteomic samples, 48 liver RNAseq samples, and 53 liver proteomic samples. First, we built separate classification and feature selection pipelines for transcriptomics and proteomics data. The liver tissue models were validated in independent liver tissue datasets. Next, we built integrated gene and protein expression models that allowed us to identify combined gene-protein biomarker panels. For liver tissue, we attained 90% nested-cross validation accuracy in our dataset and 82% accuracy in the independent validation dataset using transcriptomic data. We attained 100% nested-cross validation accuracy in our dataset and 61% accuracy in the independent validation dataset using proteomic data. For PBMCs, we attained 83% and 89% accuracy with transcriptomic and proteomic data, respectively. The integration of the two data types resulted in improved classification accuracy for PBMCs, but not liver tissue. We also identified the following gene-protein matches within the gene-protein biomarker panels: CLEC4M-CLC4M, GSTA1-GSTA2 for liver tissue and SELENBP1-SBP1 for PBMCs. In this study, machine learning models had high classification accuracy for both transcriptomics and proteomics data, across liver tissue and PBMCs. The integration of transcriptomics and proteomics into a multi-omics model yielded improvement in classification accuracy for the PBMC data. The set of integrated gene-protein biomarkers for PBMCs show promise toward developing a liquid biopsy for alcohol-associated liver disease.

60 APPLIED LIFE SCIENCES↗

NRAP-Open-IAM Analytical Reservoir Model: Development and Testing

Geological carbon sequestration (GCS) is a key technology for reducing global carbon dioxide (CO 2 ) emissions. Over the last decade, the U.S. Department of Energy has invested in understanding the science base, developing practical implementation methods, and demonstrating secure GCS technologies to mitigate the environmental impacts associated with the atmospheric release of CO 2 . As part of the National Risk Assessment Partnership, a systems-level risk assessment tool, called the NRAP-Open-IAM, has been developed to conduct risk assessment and enable safe operations at a GCS site. The current NRAP-Open-IAM contains a simple reservoir model component that calculates the evolution of CO 2 saturation and fluid pressure in a storage reservoir during CO 2 injection operations. This report presents the development and testing of a new analytical reservoir reduced-order model (ROM), which is extended from an existing semi-analytical model for estimation of CO 2 and brine leakage along legacy wells, and enhances the capability of the NRAP-Open-IAM to simulate more types of reservoir conditions. The developed model is validated against three reference studies, and the results indicate that the new ROM predicts the behavior of the two-phase fluids (brine and injected CO 2 ) well and is applicable to different reservoir simulation boundary conditions (i.e., constant pressure boundary and infinite-acting boundary) without a priori user specification of the boundary type. Sensitivity analysis for a set of model parameters is performed using 4,000 synthetic cases prepared via a fully automated process and using machine-learning-based feature selection. The stochastic analysis identifies gravitational number (i.e., ratio of gravitational forces to viscous force) and distance between the injection well and observation location as the most impactful parameters for matching the pressure and CO 2 saturation, respectively, between the numerical simulations and the ROM. This report details the possible ROM uncertainties and serves as a guide for users to understand the use and limitations of this ROM. The code implementation of the model will be released as a module within the NRAP-Open-IAM.

54 ENVIRONMENTAL SCIENCES↗

WISP: Watching grid Infrastructure Stealthily through Proxies (Final Technical Report)

The complex interdependencies of cyber systems (sensors and communications), physical grids and associated electricity market operations make protecting electric power grids a significant challenge. The energy sector is constantly under new, targeted, advanced and dangerous cyber-attacks that have the potential to result in the loss of human life. These threats are further exacerbated by our need to modernize the grid. One focus of cyber security research in smart grids is the securing of the SCADA system through advanced intrusion detection systems (IDS) and bad data detection algorithms in state estimation. These methods either require full knowledge of the system topology and parameters or fail to understand the physical behaviors under attack. WISP (Watching grid Infrastructure Stealthily through Proxies) is designed to provide additional protection to the power grid using only publicly available data. In particular, WISP exploits the spatio-temporal nature of the real time locational marginal prices (LMPs), in conjunction with other information such as bids, weather, outages and load data to analyze anomalous power pricing behaviors and then correlate those observations to localize regions of interest and identify potential cyber events. WISP is non-intrusive as the tool is deployed as a service in the Cloud or on premise and provides reliable information to system operators for enhanced situational awareness, without impeding energy delivery functions. The WISP technology comprises three modules: the data-driven anomaly detection core, the vulnerability and risk analysis and the root cause analysis. The data-driven anomaly detection core performs the tasks of feature selection, anomaly detection and attack region localization. The vulnerability and risk analysis module provides system level information of the vulnerable variables and times, assisting the operators in selecting monitoring and protection nodes. The root cause analysis module takes the detection results and identifies potential operational conditions that contribute to the detected anomalies. In Phase I, we have demonstrated the feasibility and effectiveness of WISP. We developed a realistic electricity market simulator capable of generating normal and attack market data under various operational conditions. We developed a series of cyber-attack detection and analysis algorithms and evaluated them under multiple data sources. Finally, we integrated all modules into an end-to-end software, providing functions for data management, data analytics and visualization. Specifically, we have achieved: (i) real-time data acceptance from external utility interfaces with >99% acceptance rate; (ii) high performance anomaly detection algorithms with >98% detection accuracy and <0.1% false alarm rate; and (iii) ultra-low computing delay <50 milliseconds. Additionally, our team developed algorithms to identify the vulnerable variables in electricity market operations and root cause analysis functions to identify major contributors to the price spikes. These ancillary modules are necessary when deploying WISP in real world industry environment. In Phase II, we have demonstrated the effectiveness of WISP software on realistic largescale power systems. We performed red team testing for the Phase I WISP software and identified software vulnerabilities and implemented corresponding mitigation solutions. We adapted the electricity market simulator for the Texas synthetic 2000-bus system and generated datasets for the false data injection attacks. We created database and visualization interfaces for the Texas system and the ISO New England system. We performed software optimization in terms of operation efficiency, computing speed and detection accuracy. Finally, we tested the software on the Texas system and the ISO New England system and evaluated the detection performance. Overall, we achieved above 89% detection rate, below 3% false alarm rate and below 37 seconds of end-to-end detection delay.

24 POWER TRANSMISSION AND DISTRIBUTION↗

AMPX Developments in FY2022 [Slides]

This talk centers on AMPX developments in fiscal year of 2022. This presentation provides links to the open-source subset of SCALE, including AMPX. The discussion additionally provides an overview of GNDS support in AMPX. The presentation includes selected features on several SCALE versions and features.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Assessment of Machine Learning for Ultrasonic Nondestructive Evaluation of Alkali–Silica Reaction in Concrete

Alkali–silica reaction (ASR) is a type of material degradation in concrete structures that leads to concrete cracking and rebar corrosion, thereby reducing the material’s structural integrity and the overall structure’s lifetime and raising safety concerns. Ultrasonic nondestructive evaluation (NDE) has been proven to be a valuable technique for assessing concrete properties and monitoring ASR progression in concrete. However, the deployment and analysis of ultrasonic NDE and its data requires specialized expertise, often relying on the engineer’s subjective interpretation. With the surge in computational power, artificial intelligence (AI) and machine learning (ML) algorithms have become popular in automating NDE data analysis. Various industrial sectors are increasingly adopting ML algorithms for NDE data analysis with a growing emphasis on AI–assisted automation. Regulatory agencies are also preparing for this technological shift, anticipating corresponding revisions in standards. Thus, there is an urgent need to identify the capabilities and limitations of current ML technologies for the evaluation of concrete material properties and damage status. Furthermore, the effects of various factors on ML model performance must be thoroughly investigated. The study summarized herein evaluated the effectiveness of two ML models (i.e., support vector regression (SVR) and deep neural network (DNN)) in predicting concrete material damage induced by ASR based on the long-term ultrasonic monitoring data. Four distinct concrete specimens were cast with artificially induced ASR, and over a period exceeding 500 days, ultrasonic signals and expansion data were continuously collected. For the SVR model, wave velocity and 12 other wave features were extracted from the ultrasonic signals, with 6 out of 13 features selected as input for the model. Different combinations of training and testing datasets were designed to explore factors influencing prediction performance, including the range of data within training and testing sets, in addition to various signal preprocessing methodologies. These findings suggest the importance of using a training dataset with a broader data range compared with testing datasets for improved model performance alongside consistent signal preprocessing across datasets.

36 MATERIALS SCIENCE↗