Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “supervised learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 343 records · Page 19

DONUT: physics-aware machine learning for real-time X-ray nanodiffraction analysis

Coherent X-ray scattering techniques are critical for investigating the fundamental structural properties of materials at the nanoscale. While advancements have made these experiments more accessible, real-time analysis remains a significant bottleneck, often hindered by artifacts and computational demands. In scanning X-ray nanodiffraction microscopy, which is widely used to spatially resolve structural heterogeneities, this challenge is compounded by the convolution of the divergent beam with the sample’s local structure. To address this, we introduce DONUT (Diffraction with Optics for Nanobeam by Unsupervised Training), a physics-aware neural network designed for the rapid and automated analysis of nanobeam diffraction data. By incorporating a differentiable geometric diffraction model directly into its architecture, DONUT learns to predict crystal lattice strain and orientation in real-time. Crucially, this is achieved without reliance on labeled datasets or pre-training, overcoming a fundamental limitation for supervised machine learning in X-ray science. We demonstrate experimentally that DONUT accurately extracts all features within the data over 200 times more efficiently than conventional fitting methods.

Materials science↗

DONUT: Physics-aware Machine Learning for Real-time X-ray Nanodiffraction Analysis

SF-25-088 Coherent X-ray scattering techniques are critical for investigating the fundamental structural properties of materials at the nanoscale. While advancements have made these experiments more accessible, real-time analysis remains a significant bottleneck, often hindered by artifacts and computational demands. In scanning X-ray nanodiffraction microscopy, which is widely used to spatially resolve structural heterogeneities, this challenge is compounded by the convolution of the divergent beam with the sample’s local structure. To address this, we introduce DONUT (Diffraction with Optics for Nanobeam by Unsupervised Training), a physics-aware neural network designed for the rapid and automated analysis of nanobeam diffraction data. By incorporating a differentiable geometric diffraction model directly into its architecture, DONUT learns to predict crystal lattice strain and orientation in real-time. Crucially, this is achieved without reliance on labeled datasets or pre-training, overcoming a fundamental limitation for supervised machine learning in X-ray science. We demonstrate experimentally that DONUT accurately extracts all features within the data over 200 times more efficiently than conventional fitting methods.

Zhou, Tao [Argonne National Laboratory (ANL), Argo↗

Sub-pilot-scale Production of High-Value Products from U.S. Coals

Investigators from the University of Utah, University of Wyoming and Marshall University pursued a program to study the conversion of raw coal to high-value products of carbon fiber and silicon carbide. Team members also developed an initial framework for a data portal that can incorporate laboratory data on coal processing and product quality, and also work with tools for machine learning for data analysis, data visualization and economic assessment. Experimental R&D efforts focused on the conversion of raw coal to coal tar and other byproducts, and the resulting tar intermediates were upgraded to form anisotropic and isotropic pitch materials. These pitch materials were produced from coal using both thermal (pyrolysis) and chemical (mild solvolysis liquefaction) decomposition of raw coal. Four different coals were studied: Utah bituminous coal (Sufco), Wyoming PRB coal (Black Thunder), Illinois bituminous coal (Illinois #6), and West Virginia bituminous coal (Flying Eagle). Both metallurgical-grade coking coals and lower-grade steam coals were investigated, and controlled secondary gas-phase reactions were used during a two-stage pyrolysis process to induce cracking and condensation reactions among the pyrolytic tar species. This approach successfully improved the performance of the lower grade coals for yielding pitch materials, with properties more consistent with a commercial-grade pitch that had previously demonstrated success for quality carbon fiber production. The use of waste plastic materials was also studied, to help improve physical and chemical characteristics of the intermediate tars and final pitch product; in particular, for lowering the pitch softening point to an acceptable level for melt spinning carbon fiber. Mild solvolysis liquefaction was also used as a method for producing pitch for carbon fiber production. As expected, significantly higher pitch yields were obtained using this approach, and waste plastic materials were also successfully used to reduce pitch softening point to an acceptable level. The plastic materials were also utilized to create a solvent for the mild solvolysis process, and this plastic-derived solvent was shown to provide results consistent with more expensive commercial chemical solvents, and could thus avoid the need for costly recovery and recycle of a liquefaction solvent. Additional experimental R&D focused on the production of silicon carbide (β-SiC) from the residual char byproduct from pitch production, and also on the production of carbon fiber from the anisotropic pitch. SiC was successfully synthesized using a mixture of residual char and sandstone at a ratio of 1:1. Reaction temperature and residence time were optimized and yielded a product purity of 81%. For carbon fiber production, the most successful pitch samples were obtained from the mild solvolysis liquefaction approach, combined with the use of a plastic (HDPE)-derived solvent. Fiber properties improved over time as laboratory fiber production methodologies improved, and final yields of carbon fiber were obtained with a diameter of 12.14 ± 1.10 um, Modulus of 173.73 ± 15.25 GPa, and Tensile Strength of 1.04 ± 0.10 GPa. A proof-of-concept Modern Community Research Data Portal (MCRDP) was developed and deployed for coal and coal-derived pitch characterization, with the full support of (i) remote web-based access, (ii) distributed analysis, (iii) interactive visualization and exploration, (iv) shared and long-term data access, (v) advanced query capabilities and (vi) real-time collaboration. The Coal to Products Data Portal “coaltoproducts.org” provides researchers with space to store and share data within a project, tools for analyzing and understanding data for scientific investigation, and the ability to publish data to the broader community for reproducibility. The portal leverages the Material Commons 2.0 (MC) platform developed by the Center for PRedictive Integrated Structural Materials Science (PRISMS) of the University of Michigan, to achieve long-term longevity of data collections and, more importantly, collaborative science. A number of data visualization tools were also assessed and implemented for interrogating the experimental and modeling data. The machine learning portion of this project analyzed datasets from two different coal conversion processes performed on a diverse set of coal samples from both the coal pyrolysis experiments and the solvent liquefaction experiments. The work was initiated by exploring standard regression models on the pyrolysis data, aiming to understand the impact of sample characteristics and processing conditions on key product metrics. Over the course of the project, the focus expanded to include a variety of machine learning tools, delving into both supervised and unsupervised learning methods. Models tested on the pyrolysis data included linear, ridge, lasso, elastic-net, Gaussian process, random forest regression, and AutoSklearn, and the approach was continually refined to enhance predictive accuracy and model interpretability. Similar techniques were applied to the liquefaction data with an additional focus on feature engineering. Along with mesophase content, additional outputs of interest were the pitch yield, softening point, and QI content. Insights derived from these analyses are crucial in determining the factors influencing the quality and yield of coal-derived products. As the work progressed, the research evolved from foundational model comparisons to analyses of random forests, decision paths, and feature importance scores. A thorough market analysis was performed to examine the prospects of coal-based carbon fibers. The best opportunities for coal come from its lower and more stable price relative to petroleum, particularly for subbituminous coals, which is the primary advantage that a coal refinery may have over a petroleum refinery. Before a commercial CTP production facility can be modeled, however, several things need to be understood regarding the nature of the would-be coal refinery. These include the technology to be deployed, the size of facility, the volume(s) of co-product(s), and the waste and emissions profile of the plant. The volume of co-products and waste may be substantial and will require separate market analysis to ensure viability. In the near-term, the importance of coal tar pitch, in the form of carbon pitch, to the aluminum and steel industries is likely to overshadow the alternative use of this material as an input for carbon fiber. The importance of steel and aluminum in building materials, and the need for carbon materials in their manufacturing, will ensure that demand for these products remains for the long run. In addition, carbon fiber may also be the best substitute for steel and aluminum well into the future. While society will eventually be able to shift production of much of its electricity needs to renewables, it will not be able to shift away from fossil fuels for production of high-strength construction and vehicular materials. Demand for carbon fiber is expected to increase quickly, but the volume of carbon fiber and the amount of coal that would be needed to produce even a sizeable share of this market may still be relatively small compared to current coal production. Thus, other coal-based products like graphene, graphite, carbon foams, resins, and carbon-based building products will play important roles in sustaining coal production as coal-fired power generation continues to decline.

01 COAL, LIGNITE, AND PEAT↗

Machine Learning for Anomaly Detection in Neural Network Security and SRF Cavities

This dissertation explores the development and deployment of machine learning approaches to address critical challenges in anomaly detection across two distinct domains: neural network security in federated learning settings and cavity behavior analysis in particle accelerator operations at Jefferson Lab in Newport News, Virginia. Anomaly detection identifies deviations from expected patterns, safeguarding systems in cybersecurity, industry, and research against malicious activities and failures. This dissertation demonstrates how our machine learning approaches enhance detection accuracy and efficiency in both neural network security and industrial applications. First, we investigate vulnerabilities in deep neural networks deployed in federated learning. Although federated learning preserves user privacy by training models locally, it remains vulnerable to backdoor attacks, in which malicious participants embed hidden triggers that induce targeted misbehavior. We propose a self-supervised contrastive learning framework to detect and mitigate such backdoor attacks. In our experiments, this method achieves higher detection accuracy and lower false positive rates than existing defenses, while operating without access to local model updates or original training data and thus preserving the privacy guarantees of the federated setting. Second, we address the operational reliability of superconducting radio-frequency (SRF) cavities at the Continuous Electron Beam Accelerator Facility (CEBAF). Our research leverages an unsupervised learning approach, combined with Principal Component Analysis (PCA) and k-means clustering, to identify anomalous behaviors in SRF cavities. Our method detects subtle anomalous behavior by analyzing SRF signal data. This knowledge allows for the early detection and resolution of potential faults, significantly improving the efficiency and reliability of operations. Third, we extend these insights to time-series anomaly detection more broadly. We design a contrastive-learning based model tailored to increasingly dynamic environments and academic research. This model improves detection accuracy in settings that require real-time monitoring and predictive maintenance. Our research underscores the broader applicability and impact of advanced machine learning techniques in anomaly detection. By extracting meaningful patterns from complex data, machine learning can significantly enhance security in distributed neural networks and improve the efficiency of particle accelerator operations. This dissertation serves as a stepping stone for future investigations into the vast possibilities of anomaly detection, inspiring further exploration and development of machine learning techniques in this field.

Ferguson, Hal [Old Dominion University]↗

Root identification in minirhizotron imagery with multiple instance learning

In this study, multiple instance learning (MIL) algorithms to automatically perform root detection and segmentation in minirhizotron imagery using only image-level labels are proposed. Root and soil characteristics vary from location to location, and thus, supervised machine learning approaches that are trained with local data provide the best ability to identify and segment roots in minirhizotron imagery. However, labeling roots for training data (or otherwise) is an extremely tedious and time-consuming task. This paper aims to address this problem by labeling data at the image level (rather than the individual root or root pixel level) and train algorithms to perform individual root pixel level segmentation using MIL strategies. Three MIL methods (multiple instance adaptive cosine coherence estimator, multiple instance support vector machine, multiple instance learning with randomized trees) were applied to root detection and compared to non-MIL approaches. The results show that MIL methods improve root segmentation in challenging minirhizotron imagery and reduce the labeling burden. In our results, multiple instance support vector machine outperformed other methods. The multiple instance adaptive cosine coherence estimator algorithm was a close second with an added advantage that it learned an interpretable root signature which identified the traits used to distinguish roots from soil and did not require parameter selection.

59 BASIC BIOLOGICAL SCIENCES↗

Enhanced Detection of Primary Biological Aerosol Particles Using Machine Learning and Single-Particle Measurement

Accurately identifying primary biological aerosol particles (PBAPs) using analytical techniques poses inherent challenges due to their resemblance to other atmospheric carbonaceous particles. Here, we present a study of an enhanced method for detecting PBAPs by combining single-particle measurement with advanced supervised machine learning (SML) techniques. We analyzed ambient particles from a variety of environments and lab-generated standards, focusing on chemical composition for traditional rule-based and clustering approaches and incorporating morphological features into the SML approaches, neural networks and XGBoost, for improved accuracy. This study demonstrates that SML methods outperform traditional methods in quantifying PBAPs, achieving significant improvements in precision, recall, F1-score, and accuracy, leading to an increased number of detected PBAPs by at least 19%. The adaptability of the proposed XGBoost-based SML model is showcased in comparison to traditional methods in categorizing PBAPs for blind data sets from different geographical locations. Two field case studies were investigated, over agricultural land and Amazonia rain forest, representing relatively low and high concentrations of PBAPs, respectively, where XGBoost consistently detected up to 3.5 times more PBAPs than traditional methods. Precise detection of PBAPs in the atmosphere could significantly improve the prediction of climatic impacts by them.

42 ENGINEERING↗

Physics constrained unsupervised deep learning for rapid, high resolution scanning coherent diffraction reconstruction

By circumventing the resolution limitations of optics, coherent diffractive imaging (CDI) and ptychography are making their way into scientific fields ranging from X-ray imaging to astronomy. Yet, the need for time consuming iterative phase recovery hampers real-time imaging. While supervised deep learning strategies have increased reconstruction speed, they sacrifice image quality. Furthermore, these methods’ demand for extensive labeled training data is experimentally burdensome. Here, we propose an unsupervised physics-informed neural network reconstruction method, PtychoPINN, that retains the factor of 100-to-1000 speedup of deep learning-based reconstruction while improving reconstruction quality by combining the diffraction forward map with real-space constraints from overlapping measurements. In particular, PtychoPINN gains a factor of 4 in linear resolution and an 8 dB improvement in PSNR while also accruing improvements in generalizability and robustness. This blend of performance and computational efficiency offers exciting prospects for high-resolution real-time imaging in high-throughput environments such as X-ray free electron lasers (XFELs) and diffraction-limited light sources.

97 MATHEMATICS AND COMPUTING↗

Simultaneous prediction of structural properties in epitaxially–grown GaN with quantum and conventional multi–output learning algorithms

Hundreds of GaN thin film crystal plasma–assisted molecular beam epitaxy synthesis experiment records spanning two decades were organized into a dataset correlating the growth experiment design parameters with discrete, binary determinations of crystallinity and surface morphology. Conventional data science techniques as well as both quantum and classical multi–output supervised machine learning algorithms were implemented to investigate the relationships between the operating parameter data and the structural figures of merit. Correlation coefficients, decision tree nodes, p–values, and SHAP values all support substrate temperature and gallium effusion cell conditions as being statistically significant for simultaneously influencing GaN crystallinity and surface morphology. Here, a conventional deep neural network learned best from the data, followed by a quantum–classical hybrid gradient boosting algorithm. When combined with calculations of uncertainty intervals based on VennAbers predictors, machine learning predictions of both structural properties show good agreement with results reported in published experimental literature.

36 MATERIALS SCIENCE↗

In-process monitoring and prediction of droplet quality in droplet-on-demand liquid metal jetting additive manufacturing using machine learning

Abstract In droplet-on-demand liquid metal jetting (DoD-LMJ) additive manufacturing, complex physical interactions govern the droplet characteristics, such as size, velocity, and shape. These droplet characteristics, in turn, determine the functional quality of the printed parts. Hence, to ensure repeatable and reliable part quality it is necessary to monitor and control the droplet characteristics. Existing approaches for in-situ monitoring of droplet behavior in DoD-LMJ rely on high-speed imaging sensors. The resulting high volume of droplet images acquired is computationally demanding to analyze and hinders real-time control of the process. To overcome this challenge, the objective of this work is to use time series data acquired from an in-process millimeter-wave sensor for predicting the size, velocity, and shape characteristics of droplets in DoD-LMJ process. As opposed to high-speed imaging, this sensor produces data-efficient time series signatures that allows rapid, real-time process monitoring. We devise machine learning models that use the millimeter-wave sensor data to predict the droplet characteristics. Specifically, we developed multilayer perceptron-based non-linear autoregressive models to predict the size and velocity of droplets. Likewise, a supervised machine learning model was trained to classify the droplet shape using the frequency spectrum information contained in the millimeter-wave sensor signatures. High-speed imaging data served as ground truth for model training and validation. These models captured the droplet characteristics with a statistical fidelity exceeding 90%, and vastly outperformed conventional statistical modeling approaches. Thus, this work achieves a practically viable sensing approach for real-time quality monitoring of the DoD-LMJ process, in lieu of the existing data-intensive image-based techniques.

Gaikwad, Aniruddha (ORCID:0000000285642621)↗

How machine learning can help select capping layers to suppress perovskite degradation

Environmental stability of perovskite solar cells (PSCs) has been improved by trial-and-error exploration of thin low-dimensional (LD) perovskite deposited on top of the perovskite absorber, called the capping layer. In this study, a machine-learning framework is presented to optimize this layer. We featurize 21 organic halide salts, apply them as capping layers onto methylammonium lead iodide (MAPbI 3 ) films, age them under accelerated conditions, and determine features governing stability using supervised machine learning and Shapley values. We find that organic molecules’ low number of hydrogen-bonding donors and small topological polar surface area correlate with increased MAPbI 3 film stability. The top performing organic halide, phenyltriethylammonium iodide (PTEAI), successfully extends the MAPbI3 stability lifetime by 4 ± 2 times over bare MAPbI 3 and 1.3 ± 0.3 times over state-of-the-art octylammonium bromide (OABr). Through characterization, we find that this capping layer stabilizes the photoactive layer by changing the surface chemistry and suppressing methylammonium loss.

36 MATERIALS SCIENCE↗

Detecting technological maturity from bibliometric patterns

We report the capability to identify emergent technologies based upon easily accessed open-source indicators, such as publications, is important for decision-makers in industry and government. The scientific contribution of this work is the proposition of a machine learning approach to the detection of the maturity of emerging technologies based on publication counts. Time-series of publication counts have universal features that distinguish emerging and growing technologies. We train an artificial neural network classifier, a supervised machine learning algorithm, upon these features to predict the maturity (emergent vs. growth) of an arbitrary technology. With a training set comprised of 22 technologies we obtain a classification accuracy ranging from 58.3% to 100% with an average accuracy of 84.6% for six test technologies. To enhance classifier performance, we augmented the training corpus with synthetic time-series technology life cycle curves, formed by calculating weighted averages of curves in the original training set. Training the classifier on the synthetic data set resulted in improved accuracy, ranging from 83.3% to 100% with an average accuracy of 90.4% for the test technologies. The performance of our classifier exceeds that of competing machine learning approaches in the literature, which report an average classification accuracy of only 85.7% at maximum. Moreover, in contrast to current methods our approach does not require subject matter expertise to generate training labels, and it can be automated and scaled.

97 MATHEMATICS AND COMPUTING↗

Explainable Machine Learning for Functional Data

Black-box machine learning models are recognized as useful tools for prediction applications, but the algorithmic complexity of some models causes interpretation challenges. Explainability methods have been proposed to provide insight into these models, but there is little research focused on supervised modeling with functional data inputs. We argue that, especially in applications of high consequence, it is important to explicitly model the functional dependence in a black-box analysis to not obscure or misrepresent patterns in explanations. As such, we propose the V ariable importance E xplainable E lastic S hape A nalysis (VEESA) pipeline for training supervised machine learning models with functional inputs. The pipeline is an analysis process that includes the data preprocessing, modeling, and post-hoc explanations. The preprocessing is done using elastic functional principal components analysis, which accounts for vertical and horizontal variability in functional data and, ultimately, allows for explanations in the original data space that identify the important functional variability without bias due to correlated variables. Here, we demonstrate the pipeline on two high-consequence applications: explosives classification for national security and inkjet printer identification in forensic science. The applications exhibit the VEESA pipeline’s ability to provide an understanding of the characteristics of the functional data useful for prediction. Code for implementing the pipeline is available in the veesa R package (and supplemental python code).

Elastic Shape Analysis↗

Predicting peak day and peak hour of electricity demand with ensemble machine learning

Battery energy storage systems can be used for peak demand reduction in power systems, leading to significant economic benefits. Two practical challenges are 1) accurately determining the peak load days and hours and 2) quantifying and reducing uncertainties associated with the forecast in probabilistic risk measures for dispatch decision-making. In this study, we develop a supervised machine learning approach to generate 1) the probability of the next operation day containing the peak hour of the month and 2) the probability of an hour to be the peak hour of the day. Guidance is provided on preparation and augmentation of data as well as selection of machine learning models and decision-making thresholds. The proposed approach is applied to the Duke Energy Progress system and successfully captures 69 peak days out of 72 testing months with a 3% exceedance probability threshold. On 90% of the peak days, the actual peak hour is among the 2 h with the highest probabilities.

25 ENERGY STORAGE↗

Modern Senicide in the Face of a Pandemic: An Examination of Public Discourse and Sentiment About Older Adults and COVID-19 Using Machine Learning

Objectives This study examined public discourse and sentiment regarding older adults and COVID-19 on social media and assessed the extent of ageism in public discourse. Methods Twitter data (N = 82,893) related to both older adults and COVID-19 and dated from January 23 to May 20, 2020, were analyzed. We used a combination of data science methods (including supervised machine learning, topic modeling, and sentiment analysis), qualitative thematic analysis, and conventional statistics. Results The most common category in the coded tweets was “personal opinions” (66.2%), followed by “informative” (24.7%), “jokes/ridicule” (4.8%), and “personal experiences” (4.3%). The daily average of ageist content was 18%, with the highest of 52.8% on March 11, 2020. Specifically, more than 1 in 10 (11.5%) tweets implied that the life of older adults is less valuable or downplayed the pandemic because it mostly harms older adults. A small proportion (4.6%) explicitly supported the idea of just isolating older adults. Almost three-quarters (72.9%) within “jokes/ridicule” targeted older adults, half of which were “death jokes.” Also, 14 themes were extracted, such as perceptions of lockdown and risk. A bivariate Granger causality test suggested that informative tweets regarding at-risk populations increased the prevalence of tweets that downplayed the pandemic. Discussion Ageist content in the context of COVID-19 was prevalent on Twitter. Information about COVID-19 on Twitter influenced public perceptions of risk and acceptable ways of controlling the pandemic. Finaly, public education on the risk of severe illness is needed to correct misperceptions.

60 APPLIED LIFE SCIENCES↗

Artificial neural network approach for multiphase segmentation of battery electrode nano-CT images

The segmentation of tomographic images of the battery electrode is a crucial processing step, which will have an additional impact on the results of material characterization and electrochemical simulation. However, manually labeling X-ray CT images (XCT) is time-consuming, and these XCT images are generally difficult to segment with histographical methods. We propose a deep learning approach with an asymmetrical depth encode-decoder convolutional neural network (CNN) for real-world battery material datasets. This network achieves high accuracy while requiring small amounts of labeled data and predicts a volume of billions voxel within few minutes. While applying supervised machine learning for segmenting real-world data, the ground truth is often absent. The results of segmentation are usually qualitatively justified by visual judgement. We try to unravel this fuzzy definition of segmentation quality by identifying the uncertainty due to the human bias diluted in the training data. Further CNN trainings using synthetic data show quantitative impact of such uncertainty on the determination of material’s properties. Nano-XCT datasets of various battery materials have been successfully segmented by training this neural network from scratch. We will also show that applying the transfer learning, which consists of reusing a well-trained network, can improve the accuracy of a similar dataset.

25 ENERGY STORAGE↗

Using machine learning to identify extragalactic globular cluster candidates from ground-based photometric surveys of M87

Globular clusters (GCs) have been at the heart of many longstanding questions in many sub-fields of astronomy and, as such, systematic identification of GCs in external galaxies has immense impacts. In this study, we take advantage of M87’s well-studied GC system to implement supervised machine learning (ML) classification algorithms – specifically random forest and neural networks – to identify GCs from foreground stars and background galaxies, using ground-based photometry from the Canada–France–Hawaii Telescope (CFHT). We compare these two ML classification methods to studies of ‘human-selected’ GCs and find that the best-performing random forest model can reselect 61.2 per cent ± 8.0 per cent of GCs selected from HST data (ACSVCS) and the best-performing neural network model reselects 95.0 per cent ± 3.4 per cent. When compared to human-classified GCs and contaminants selected from CFHT data – independent of our training data – the best-performing random forest model can correctly classify 91.0 per cent ± 1.2 per cent and the best-performing neural network model can correctly classify 57.3 per cent ± 1.1 per cent. ML methods in astronomy have been receiving much interest as Vera C. Rubin Observatory prepares for first light. The observables in this study are selected to be directly comparable to early Rubin Observatory data and the prospects for running ML algorithms on the upcoming data set yields promising results.

79 ASTRONOMY AND ASTROPHYSICS↗

Anomaly Detection in Seismic Data with Deep Learning: Application for Instrument Failure Detection and Forecasting

Seismic data quality assessment (QA) is the first and one of the most important steps before conducting any further data analysis. Traditional methods involve checking various metrics, such as spike detection and power spectral density, by setting strict thresholds or comparing data against synthetic benchmarks. However, these approaches often rely on pre-existing knowledge and assumptions about data anomalies, leading to potential misclassification of unusual cases. Here, in this study, we propose a deep autoencoder model, an unsupervised learning approach that evaluates data quality without making assumptions about normal and anomalous data, which can be used to identify deviations in recorded data that may indicate nascent instrument failure. We test the model with the U.S. International Monitoring System (IMS) seismic stations and demonstrate the capability of detecting anomalies on a monthly scale. This could prompt station operators to examine potential problems early, allowing sufficient time for instrument maintenance to prevent data outages. In addition, we use a new manually selected testing dataset to compare our model performance against two supervised machine learning (ML) approaches and a standard QA package, as baseline models. When applied to the dataset containing known data anomalies, performance of the supervised and unsupervised ML approaches is similar, with an accuracy of 88.1% for our model compared to ∼90% for the supervised ML approach and 78.2% for the standard QA package. Our model outperforms the baseline models when applied to new stations, where new types of data anomalies can be station-specific and not included in the training dataset. Finally, we show model transferability by training the model with data from the Global Seismograph Network only and applying it to the IMS network data. The results suggest that our model is generalizable and can be applied to new stations with good accuracy.

Lin, Jiun-Ting [Lawrence Livermore National Labora↗