Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Principal component analysis”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10

Analysis of Superconducting Magnet Quench Antenna Data

Quenching poses a serious problem for superconducting magnets operating at high currents. It occurs when the material transitions from the superconducting to the normal state, which leads to heating and potential damage to the magnet. To understand and mitigate quenching, the Magnet Department at Fermilab is developing and testing superconducting magnet quench antenna arrays. This study delves into the anomalous events preceding the quench during magnet training by analyzing the collected data. With the moving average and Fast Fourier Transform techniques, we investigate the trends and frequency patterns of the data. Moreover, we introduce an unsupervised anomaly detection algorithm based on Principal Component Analysis and DBSCAN clustering. It can autonomously identify events within background noise, without relying on any predefined event features. Our analysis reveals that the spatio-temporal distribution of these anomalous events has little connection to the quench location, indicating that a majority of them bear no relation to the quenching process.

43 PARTICLE ACCELERATORS↗

Homomorphic Encryption for Machine Learning and Artificial Intelligence Applications

Third-party and expert analysis is a cost-effective solution for solving specialized problems or processing large datasets related to reactor structural health monitoring and nondestructive evaluation. However, when handling proprietary information, third-party and expert analysts pose a privacy risk. To address this challenge, Homomorphic Encryption (HE) permits arithmetic operations on encrypted data without exposing the underlying data. Implementations of Machine Learning (ML) and Artificial Intelligence (AI) algorithms using HE greatly enhances the capabilities of third-party analysts while maintaining a low security risk. This paper details current success in applying Principal Component Analysis (PCA) and Fully Connected Neural Networks (NN) using the Microsoft SEAL implementation of the popular CKKS Fully Homomorphic Encryption (FHE) algorithm. The MNIST Handwritten Dataset is analyzed as a proof-of-concept demonstration of the implementations.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Detecting anomalous packets in network transfers: investigations using PCA, autoencoder and isolation forest in TCP

Large-scale scientific workflows rely heavily on high-performance file transfers. These transfers require strict quality parameters such as guaranteed bandwidth, no packet loss or data duplication. To have successful file transfers, methods such as predetermined thresholds and statistical analysis need to be done to determine abnormal patterns. Network administrators routinely monitor and analyze network data for diagnosing and alleviating these, making decisions based on their experience. However, as networks grow and become complex, monitoring large data files and quickly processing them, makes it improbable to identify errors and rectify these. Abnormal file transfers have been classified by simply setting alert thresholds, via tools such as PerfSonar and TCP statistics (Tstat). This paper investigates the feasibility of unsupervised feature extraction methods for identifying network anomaly patterns with three unsupervised classification methods—principal component analysis, autoencoder and isolation forest. Here, we collect file transfer statistics from two experiment sets—synthetic iPerf generated traffic and 1000 Genome workflow runs, with synthetically introduced anomalies. Our results show that while PCA and a simple autoencoder finds it difficult to detect clusters, the tree-variant isolation forest is able to identify anomalous packets by breaking down TCP traces into tree classes early.

97 MATHEMATICS AND COMPUTING↗

Modeled production, oxidation, and transport processes of wetland methane emissions in temperate, boreal, and Arctic regions

Abstract Wetlands are the largest natural source of methane (CH 4 ) to the atmosphere. The eddy covariance method provides robust measurements of net ecosystem exchange of CH 4 , but interpreting its spatiotemporal variations is challenging due to the co‐occurrence of CH 4 production, oxidation, and transport dynamics. Here, we estimate these three processes using a data‐model fusion approach across 25 wetlands in temperate, boreal, and Arctic regions. Our data‐constrained model—iPEACE—reasonably reproduced CH 4 emissions at 19 of the 25 sites with normalized root mean square error of 0.59, correlation coefficient of 0.82, and normalized standard deviation of 0.87. Among the three processes, CH 4 production appeared to be the most important process, followed by oxidation in explaining inter‐site variations in CH 4 emissions. Based on a sensitivity analysis, CH 4 emissions were generally more sensitive to decreased water table than to increased gross primary productivity or soil temperature. For periods with leaf area index (LAI) of ≥20% of its annual peak, plant‐mediated transport appeared to be the major pathway for CH 4 transport. Contributions from ebullition and diffusion were relatively high during low LAI (<20%) periods. The lag time between CH 4 production and CH 4 emissions tended to be short in fen sites (3 ± 2 days) and long in bog sites (13 ± 10 days). Based on a principal component analysis, we found that parameters for CH 4 production, plant‐mediated transport, and diffusion through water explained 77% of the variance in the parameters across the 19 sites, highlighting the importance of these parameters for predicting wetland CH 4 emissions across biomes. These processes and associated parameters for CH 4 emissions among and within the wetlands provide useful insights for interpreting observed net CH 4 fluxes, estimating sensitivities to biophysical variables, and modeling global CH 4 fluxes.

Ueyama, Masahito↗

Comparing gas composition from fast pyrolysis of live foliage measured in bench-scale and fire-scale experiments

Background: Fire models have used pyrolysis data from oxidising and non-oxidising environments for flaming combustion. In wildland fires pyrolysis, flaming and smouldering combustion typically occur in an oxidising environment (the atmosphere). Aims: Using compositional data analysis methods, determine if the composition of pyrolysis gases measured in non-oxidising and ambient (oxidising) atmospheric conditions were similar. Methods: Permanent gases and tars were measured in a fuel-rich (non-oxidising) environment in a flat flame burner (FFB). Permanent and light hydrocarbon gases were measured for the same fuels heated by a fire flame in ambient atmospheric conditions (oxidising environment). Log-ratio balances of the measured gases common to both environments (CO, CO 2 , CH 4 , H 2 , C 6 H 6 O (phenol), and other gases) were examined by principal components analysis (PCA), canonical discriminant analysis (CDA) and permutational multivariate analysis of variance (PERMANOVA). Key results: Mean composition changed between the non-oxidising and ambient atmosphere samples. PCA showed that flat flame burner (FFB) samples were tightly clustered and distinct from the ambient atmosphere samples. CDA found that the difference between environments was defined by the CO-CO 2 log-ratio balance. PERMANOVA and pairwise comparisons found FFB samples differed from the ambient atmosphere samples which did not differ from each other. Conclusion: Relative composition of these pyrolysis gases differed between the oxidising and non-oxidising environments. This comparison was one of the first comparisons made between bench-scale and field scale pyrolysis measurements using compositional data analysis. Implications: These results indicate the need for more fundamental research on the early time-dependent pyrolysis of vegetation in the presence of oxygen.

54 ENVIRONMENTAL SCIENCES↗

TOFHunter—unlocking rapid untargeted screening of inductively coupled plasma–time-of-flight–mass spectrometry data

This study provides an overview of a newly developed open source program written in Python, TOFHunter, which permits the rapid and untargeted screening of inductively coupled plasma (ICP)-time-of-flight (TOF)-mass spectrometry (MS) datasets. ICP-TOF-MS is an analytical tool capable of providing quasi simultaneous detection of all nuclides from Li to Pu. This capability has triggered an increase in studies investigating single-particle analysis in which the TOF-MS provides correlated elemental/isotopic signatures on a particle basis in time. Similarly, laser ablation mapping has seen rapid growth owing to ICP-TOF-MS's capacity to handle fast washout times (<10 ms) while providing a broad nuclide coverage. The caveat to this broad mass coverage and high time resolution comes in the form of large, overwhelming datasets. With datasets typically on the scale of gigabytes, it is easy for a user to only focus on very targeted analytes; however, this focus diminishes the opportunity offered by the TOF-MS detector. TOFHunter applies chemometric methods, principal component analysis (PCA), and interesting features finder (IFF) on ICP-TOF-MS data, allowing for investigation of correlations, major and minor variance sources, and sample screening. The unique spectra identified by the (IFF) are used to generate a list of mass peaks, which are then matched with both nuclides and potential interferences before being exported for the user to investigate. Several case studies are discussed herein, demonstrating TOFHunter's ability to screen aqueous injections, single-particle/single-cell analysis, and probe laser ablation mapping files for unique regions of interest.

47 OTHER INSTRUMENTATION↗

Precise relative magnitude measurement improves fracture characterization during hydraulic fracturing

SUMMARY Microseismic monitoring is an important technique to obtain detailed knowledge of in-situ fracture size and orientation during stimulation to maximize fluid flow throughout the rock volume and optimize production. Furthermore, considering that the frequency of earthquake magnitudes empirically follows a power law (i.e. Gutenberg–Richter), the accuracy of microseismic event magnitude distributions is potentially crucial for seismic risk management. In this study, we analyse microseismicity observed during four hydraulic fracture treatments of the legacy Cotton Valley experiment in 1997 at the Carthage gas field of East Texas, where fractures were activated at the base of the sand-shale Upper Cotton Valley formation. We perform waveform cross-correlation to detect similar event clusters, measure relative amplitude from aligned waveform pairs with a principal component analysis, then measure precise relative magnitudes. The new magnitudes significantly reduce the deviations between magnitude differences and relative amplitudes of event pairs. This subsequently reduces the magnitude differences between clusters located at different depths. Reduction in magnitude differences between clusters suggests that some attenuation-related biases could be effectively mitigated with relative magnitude measurements. The maximum likelihood method is applied to understand the magnitude frequency distributions and quantify the seismogenic index of the clusters. Statistical analyses with new magnitudes suggest that fractures that are more favourably oriented for shear failure have lower b-value and higher seismogenic index, suggesting higher potential for relatively larger earthquakes, rather than fractures subparallel to maximum horizontal principal stress orientation.

58 GEOSCIENCES↗

Baycal

Bayesian Model Calibration (BayCal) toolkit is a software plugin for Risk Analysis Virtual Environment (RAVEN) framework, arming at inversely quantifying the uncertainties associated with simulation model parameters based on available experiment data. BayCal seeks statistical inference of the uncertain input parameters that are consistent with the available measurement data or observed data. The unique feature of BayCal is the capability to be linked with RAVEN to build corresponding calibration workflows for complex multi-physics simulations. In addition to be able to use the machine learning capability of RAVEN to significantly reduce the computational cost of expensive simulation models, another distinctive feature of BayCal is the capability to deal with high-dimensional correlated model outputs, such as time series observations at multiple locations, via principal component analysis (PCA) technique.

Wang, Congjian↗

Revealing Phase Heterogeneity in Vertically Aligned Nanocomposites via Plan-View Electron Energy Loss Spectroscopy

Hydrogen utilization in clean energy technologies is challenged by limited storage and transport within materials, owing to the complex hydrogen kinetics at interfaces [1]. Understanding these interfacial mechanisms at the nanoscale is crucial for developing improved materials for hydrogen applications, particularly proton-conducting fuel cells (PCFCs). Vertically aligned nanocomposites (VANs) grown by pulsed laser deposition (PLD) offer a unique platform for investigating the interfacial effects on hydrogen transport due to their well-defined interfaces parallel to the direction of charge transport [2-4]. To investigate hydrogen transport, the two phases within the VANs were chosen as BaZr 0.9 Y 0.1 O 3-x (BZY), a known proton conductor, and Pr 0.1 Ce 0.9 O 2-x (PCO), a mixed ionic-electronic conductor [5]. This PCO-BZY VANs architecture allows the investigation of how the interface between a proton conductor and a mixed conductor influences hydrogen transport. However, because of the small size of hydrogen, it is difficult to discern the nature of its interactions with interfaces from bulk measurements at the macroscale, thus necessitating nanoscale measurements [6]. Electron energy loss spectroscopy (EELS) allows for nanometer-resolution probing of the local atomic structure and chemistry at the BZY/PCO interface. In this study, plan-view analysis of PCO-BZY VANs films was employed to characterize the structure and phase distribution of the VANs and investigate the interface between the nanostructures. The films were imaged using scanning electron microscopy (SEM) in the Hitachi S-4800 SEM, collecting secondary electron images using mixed upper and lower detectors. Then, plan-view transmission electron microscopy (TEM) and scanning transmission electron microscopy (STEM) EELS were employed using a JEOL ARM300 microscope operated at 300kV with a Gatan K3 GIF Continuum detector to study the distribution of the BZY and PCO phases through the film. As a result, spectrum images were acquired at a dispersion of 0.18eV per channel and denoised afterward by principal component analysis (PCA) method.

Griffin, Elizabeth [Northwestern University, Evans↗

raogroupuiuc/lipo11557_growth

Project: Systems analysis of Lipomyces starkeyi during growth on various plant-based sugars Authors: Anshu Deewan*, Jing-Jing Liu*, Sujit Sadashiv Jagtap, Eun Ju Yun, Hanna E Walukiewicz, Yong-Su Jin, Christopher V Rao (* - co-authors) Abstract : Oleaginous yeasts have received significant attention due to their substantial lipids storage capability. The accumulated lipids can be utilized directly or processed into various bioproducts and biofuels. Lipomyces starkeyi is an oleaginous yeast capable of using multiple plant-based sugars, such as glucose, xylose, and cellobiose. It is, however, a relatively unexplored yeast due to limited knowledge about its physiology. In this study, we have evaluated the growth of L. starkeyi on different sugars and performed transcriptomic and metabolomic analyses to understand the underlying mechanisms of sugar metabolism. Principal component analysis showed clear differences resulting from growth on different sugars. We have further reported various metabolic pathways activated during growth on these sugars. We also observed non-specific regulation in L. starkeyi and have updated the gene annotations for the NRRL Y-11557 strain. This analysis provides a foundation for understanding the metabolism of these plant-based sugars and potentially valuable information to guide the metabolic engineering of L. starkeyi to produce bioproducts and biofuels.

Deewan, Anshu↗

Machine Learning-Driven Reliability Estimation of PV Inverters Considering Alert-Ambient Variability

Weather-induced spatio-temporal degradation limits outdoor PV inverter lifetime and reliability, necessitating advanced data analysis. This study employs a top-down, data-driven approach utilizing multiple machine learning (ML) algorithms to estimate inverter reliability in a 1.4 MW PV power plant, considering factors such as irradiance, humidity, temperature, time of day, and weather conditions. An extensive alert dataset from 17 identical inverters, including alert types, propagation, and frequency, reveals significant correlations with environmental factors and inverter output power, enabling the construction of a performance reliability model. Dual-stage supervised-ML models are evaluated for accuracy, with the ‘classification-regression’ model by an artificial neural network (ANN) tested on the averaged “Alert-Ambient” dataset, which is outperformed by ‘clustering-regression’ models using random forest (RF) and K-Nearest Neighbors (KNN) on individual inverter datasets. K-means clustering applies principal component analysis to reduce dimensions, achieving improved accuracy beyond the 80% achieved by ANN on the averaged dataset. Second-stage regression estimates inverter reliability with a mean square error of 0.0195 on the averaged dataset and as low as 0.002 on individual inverter datasets using RF. Furthermore, these findings highlight the method's suitability for estimating PV inverter output reliability under ambient conditions, essential for digital twin development and related applications.

14 SOLAR ENERGY↗

New Measurements of the Lyα Forest Continuum and Effective Optical Depth with LyCAN and DESI Y1 Data

Abstract We present the Ly α Continuum Analysis Network (LyCAN), a convolutional neural network that predicts the unabsorbed quasar continuum within the rest-frame wavelength range of 1040–1600 Å based on the red side of the Ly α emission line (1216–1600 Å). We developed synthetic spectra based on a Gaussian mixture model representation of nonnegative matrix factorization (NMF) coefficients. These coefficients were derived from high-resolution, low-redshift ( z < 0.2) Hubble Space Telescope/Cosmic Origins Spectrograph (COS) quasar spectra. We supplemented this COS-based synthetic sample with an equal number of DESI Year 5 mock spectra. LyCAN performs extremely well on testing sets, achieving a median error in the forest region of 1.5% on the DESI mock sample, 2.0% on the COS-based synthetic sample, and 4.1% on the original COS spectra. LyCAN outperforms principal component analysis (PCA) and NMF-based prediction methods using the same training set by 40% or more. We predict the intrinsic continua of 83,635 DESI Year 1 spectra in the redshift range of 2.1 ≤ z ≤ 4.2 and perform an absolute measurement of the evolution of the effective optical depth. This is the largest sample employed to measure the optical depth evolution to date. We fit a power law of the form τ ( z ) = τ 0 ( 1 + z ) γ to our measurements and find τ 0 = (2.46 ± 0.14) × 10 −3 and γ = 3.62 ± 0.04. Our results show particular agreement with high-resolution, ground-based observations around z = 2, indicating that LyCAN is able to predict the quasar continuum in the forest region with only spectral information outside the forest.

79 ASTRONOMY AND ASTROPHYSICS↗

Exploring Geothermal Potential of Great Basin Sub-Regions: Preprint

The INnovative Geothermal Exploration through Novel Investigations Of Undiscovered Systems (INGENIOUS) project aims to discover new, economically viable hidden geothermal systems in the Great Basin region by building on previous work in play fairway analysis and machine learning. A key objective of this project is to develop an exploration workflow to reduce geothermal exploration risks for hidden geothermal systems. A single preliminary play fairway workflow was developed from the assessment of the regional INGENIOUS geological, geophysical, and geochemical datasets. This workflow provided new preliminary predictive geothermal fairway maps for the INGENIOUS study area, which encompasses most of Nevada, western Utah, southern Idaho, southeastern Oregon, and easternmost California. However, a recent study (incorporating machine learning techniques) of a portion of Nevada identified four geologic domains and determined that the relative importance of individual datasets or features as indicators of geothermal potential may differ across these domains. The INGENIOUS study area includes a much larger and more geologically diverse region; therefore, additional geologic domains or sub-regions are expected. To assess the sub-regions in the INGENIOUS study area, principal component analysis and k-means clustering were applied. Preliminary results indicate that the INGENIOUS regional data cluster into groups that relate to different geologic domains in the Great Basin region. These include domains such as the Walker Lane, extensional western Great Basin region, broad lower strain region in the eastern Great Basin of western Utah and eastern Nevada, Quaternary volcanic fields, and the area adjacent to the Snake River Plain. These clusters are assessed to determine the key geologic drivers of the identified clusters. Understanding this variability can provide key insights for the exploration and characterization of hidden geothermal systems in the Great Basin region and could indicate the need to develop multiple geothermal conceptual models and play fairway workflows for the INGENIOUS study area.

GEOTHERMAL ENERGY↗

Exploring Geothermal Potential of Great Basin Sub-Regions

The INnovative Geothermal Exploration through Novel Investigations Of Undiscovered Systems (INGENIOUS) project aims to discover new, economically viable hidden geothermal systems in the Great Basin region by building on previous work in play fairway analysis and machine learning. A key objective of this project is to develop an exploration workflow to reduce geothermal exploration risks for hidden geothermal systems. A single preliminary play fairway workflow was developed from the assessment of the regional INGENIOUS geological, geophysical, and geochemical datasets. This workflow provided new preliminary predictive geothermal fairway maps for the INGENIOUS study area, which encompasses most of Nevada, western Utah, southern Idaho, southeastern Oregon, and easternmost California. However, a recent study (incorporating machine learning techniques) of a portion of Nevada identified four geologic domains and determined that the relative importance of individual datasets or features as indicators of geothermal potential may differ across these domains. The INGENIOUS study area includes a much larger and more geologically diverse region; therefore, additional geologic domains or sub-regions are expected. To assess the sub-regions in the INGENIOUS study area, principal component analysis and k-means clustering were applied. Preliminary results indicate that the INGENIOUS regional data cluster into groups that relate to different geologic domains in the Great Basin region. These include domains such as the Walker Lane, extensional western Great Basin region, broad lower strain region in the eastern Great Basin of western Utah and eastern Nevada, Quaternary volcanic fields, and the area adjacent to the Snake River Plain. These clusters are assessed to determine the key geologic drivers of the identified clusters. Understanding this variability can provide key insights for the exploration and characterization of hidden geothermal systems in the Great Basin region and could indicate the need to develop multiple geothermal conceptual models and play fairway workflows for the INGENIOUS study area.

exploration↗

Improved precision in As speciation analysis with HERFD-XANES at the As K -edge: the case of As speciation in mine waste

High-energy-resolution fluorescence-detected (HERFD) X-ray absorption near-edge spectroscopy (XANES) is a spectroscopic method that allows for increased spectral feature resolution, and greater selectivity to decrease complex matrix effects compared with conventional XANES. XANES is an ideal tool for speciation of elements in solid-phase environmental samples. Accurate speciation of As in mine waste materials is important for understanding the mobility and toxicity of As in near-surface environments. In this study, linear combination fitting (LCF) was performed on synthetic spectra generated from mixtures of eight measured reference compounds for both HERFD-XANES and transmission-detected XANES to evaluate the improvement in quantitative speciation with HERFD-XANES spectra. The reference compounds arsenolite (As 2 O 3 ), orpiment (As 2 S 3 ), getchellite (AsSbS 3 ), arsenopyrite (FeAsS), kaňkite (FeAsO 4 ·3.5H 2 O), scorodite (FeAsO 4 ·2H 2 O), sodium arsenate (Na 3 AsO 4 ), and realgar (As 4 S 4 ) were selected for their importance in mine waste systems. Statistical methods of principal component analysis and target transformation were employed to determine whether HERFD improves identification of the components in a dataset of mixtures of reference compounds. LCF was performed on HERFD- and total fluorescence yield (TFY)-XANES spectra collected from mine waste samples. Arsenopyrite, arsenolite, orpiment, and sodium arsenate were more accurately identified in the synthetic HERFD-XANES spectra compared with the transmission-XANES spectra. In mine waste samples containing arsenopyrite and either scorodite or kaňkite, LCF with HERFD-XANES measurements resulted in fits with smaller R -factors than concurrently collected TFY measurements. The improved accuracy of HERFD-XANES analysis may provide enhanced delineation of As phases controlling biogeochemical reactions in mine wastes, contaminated soils, and remediation systems.

58 GEOSCIENCES↗

Temporal Analysis and Scene Change Detection in Multispectral Overhead Imagery

Scene change detection can be a tedious and time consuming process especially when concerning large geographical areas, and the process can be even more cumbersome when analyzing changes in an area over large spans of time. Developing a useful way to help analysts recognize at what points in time significant changes to a scene have occurred can allow them to better focus their efforts in characterizing events. Applications include: Facility monitoring, Construction chronology, Monitoring of vehicle/aircraft activity, Characterization of larger sequences of events. In large areas exceeding hundreds to thousands of square kilometers in size, it can be difficult localizing when scene changes have occurred. Analysts can spend hours going through imagery to try to identify new construction, monitor facility activities, monitor vehicle movement, etc. where the object of interest may only be a few square meters. Our goal is to help cut down this time by giving analysts change maps with hot spots of change, allowing them to focus on regions that have experienced actual change in time frames they're interested in. Additionally, by combining these change maps into layers within a data cube, analysts can examine the change maps from a temporal perspective, allowing events to be characterized over spans of time. By opening the data cube in an imaging software capable of separating the layers, we can analyze the change maps sequentially, allowing us to examine scene changes occurring over time. As an example, we examined overhead imagery from Planet Labs of what appears to be a parking lot on Fort Irwin over the course of a year using ENVI, a geospatial satellite imagery analysis software. Using ENVI, we generate a graph of changes over time, and notice a particular segment near the end of our analysis window where no changes are detected. Examination of the actual satellite imagery reveals that during this time span, the parking lot was empty. This could be due to facility shutdown for maintenance or upgrades, or possibly even total workforce/vehicle fleet movement. Information like this could help analysts better characterize events, as well as to help create clearer timelines in larger sequences of events. Workflow steps: - Collect multiple maps of the same AOI (Area of Interest) during a time span of interest; - Generate change maps from AOI maps; - Generate data cube from change maps. An analyst can use the data cube to help inspect an AOI for activities within a time span of interest. If an event of interest is discovered, the analyst can then refer to the maps corresponding to the appropriate dates and times in the data cube to see precisely what is transpiring. The biggest objective being worked on is improving the change detection methodology employed. We currently use PCA-EM (Principal Component Analysis with Expectation Maximization), but we are currently focusing on implementing IR-MAD (Iteratively Reweighted Multivariate Alteration Detection) to be used in conjunction with PCA-EM in an effort to decrease false positivity and noise in the change maps we generate.

42 ENGINEERING↗

Machine Learning for Anomaly Detection in Neural Network Security and SRF Cavities

This dissertation explores the development and deployment of machine learning approaches to address critical challenges in anomaly detection across two distinct domains: neural network security in federated learning settings and cavity behavior analysis in particle accelerator operations at Jefferson Lab in Newport News, Virginia. Anomaly detection identifies deviations from expected patterns, safeguarding systems in cybersecurity, industry, and research against malicious activities and failures. This dissertation demonstrates how our machine learning approaches enhance detection accuracy and efficiency in both neural network security and industrial applications. First, we investigate vulnerabilities in deep neural networks deployed in federated learning. Although federated learning preserves user privacy by training models locally, it remains vulnerable to backdoor attacks, in which malicious participants embed hidden triggers that induce targeted misbehavior. We propose a self-supervised contrastive learning framework to detect and mitigate such backdoor attacks. In our experiments, this method achieves higher detection accuracy and lower false positive rates than existing defenses, while operating without access to local model updates or original training data and thus preserving the privacy guarantees of the federated setting. Second, we address the operational reliability of superconducting radio-frequency (SRF) cavities at the Continuous Electron Beam Accelerator Facility (CEBAF). Our research leverages an unsupervised learning approach, combined with Principal Component Analysis (PCA) and k-means clustering, to identify anomalous behaviors in SRF cavities. Our method detects subtle anomalous behavior by analyzing SRF signal data. This knowledge allows for the early detection and resolution of potential faults, significantly improving the efficiency and reliability of operations. Third, we extend these insights to time-series anomaly detection more broadly. We design a contrastive-learning based model tailored to increasingly dynamic environments and academic research. This model improves detection accuracy in settings that require real-time monitoring and predictive maintenance. Our research underscores the broader applicability and impact of advanced machine learning techniques in anomaly detection. By extracting meaningful patterns from complex data, machine learning can significantly enhance security in distributed neural networks and improve the efficiency of particle accelerator operations. This dissertation serves as a stepping stone for future investigations into the vast possibilities of anomaly detection, inspiring further exploration and development of machine learning techniques in this field.

Ferguson, Hal [Old Dominion University]↗

Surface analysis insight note: Differentiation methods applicable to noisy data for determination of sp2‐ versus sp3‐hybridization of carbon allotropes and AES signal strengths

The derivatives of the spectra are commonly used for quantification in Auger Electron Spectroscopy (AES) spectra, while the derivative of the KLL C Auger line has proven to be valuable in obtaining a measure of the relative proportions of sp 2 ‐ and sp 3 ‐hybridization using the D‐parameter in both AES and X‐ray Photoelectron Spectroscopy (XPS). Differentiation of X‐ray Photoelectron Spectroscopy (XPS) and Auger Electron Spectroscopy (AES) spectra by numerical means is presented and illustrated for polymeric, such as PEEK and Nylon, as well as for graphitic materials including highly ordered pyrolytic graphite and graphene oxide. The most commonly available Savitzky–Golay method is explained mathematically and developed through the case of constructing a 5‐point quadratic polynomial convolution kernel suitable for differentiating spectra of adequate signal to noise. The concept of differentiation of spectra where signal to noise is less than adequate is also developed. Two alternative strategies to Savitzky–Golay differentiation are presented, which fit curves to data that allow derivatives to be obtained where Savitzky–Golay would otherwise fail. These alternative methods involve constructing a parametric curve that fits data over the entire energy interval of interest. Derivatives of spectra are then obtained by differentiating these parametric curves directly. A comparison of results for different materials for which specific sp 2 ‐ vs sp 3 ‐hybridized carbon proportions are of interest is used to emphasize the importance of characterizing methods used to differentiate spectra and understanding the characteristics of instrumentation used to measure spectra. The case for using Principal Component Analysis noise reduction with C KLL spectra is made for spectra collected from a heterogeneous graphene oxide sample.

Fairley, Neal↗