Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Principal component analysis”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 451 records · Page 25

Towards revealing intrinsic vortex-core states in Fe-based superconductors through statistical discovery

Abstract In type-II superconductors, electronic states within magnetic vortices hold crucial information about the paring mechanism and can reveal non-trivial topology. While scanning tunneling microscopy/spectroscopy (STM/S) is a powerful tool for imaging superconducting vortices, it is challenging to isolate the intrinsic electronic properties from extrinsic effects like subsurface defects and disorders. Here we combine STM/STS with basic machine learning to develop a method for screening out the vortices pinned by embedded disorder in iron-based superconductors. Through a principal component analysis of large STS data within vortices, we find that the vortex-core states in Ba(Fe 0.96 Ni 0.04 ) 2 As 2 start to split into two categories at certain magnetic field strengths, reflecting vortices with and without pinning by subsurface defects or disorders. Our machine-learning analysis provides an unbiased approach to reveal intrinsic vortex-core states in novel superconductors and shed light on ongoing puzzles in the possible emergence of a Majorana zero mode.

Guo, Yueming↗

An end-to-end trainable hybrid classical-quantum classifier

Abstract We introduce a hybrid model combining a quantum-inspired tensor network and a variational quantum circuit to perform supervised learning tasks. This architecture allows for the classical and quantum parts of the model to be trained simultaneously, providing an end-to-end training framework. We show that compared to the principal component analysis, a tensor network based on the matrix product state with low bond dimensions performs better as a feature extractor for the input data of the variational quantum circuit in the binary and ternary classification of MNIST and Fashion-MNIST datasets. The architecture is highly adaptable and the classical-quantum boundary can be adjusted according to the availability of the quantum resource by exploiting the correspondence between tensor networks and quantum circuits.

97 MATHEMATICS AND COMPUTING↗

Whole-genome sequencing distinguishes the two most common giant kelp ecomorphs

Abstract Giant kelp, Macrocystis pyrifera, exists as distinct morphological variants—or “ecomorphs”—in different populations, yet the mechanism for this variation is uncertain, and environmental drivers for either adaptive or plastic phenotypes have not been identified. The ecomorphs Macrocystis “pyrifera” and M. “integrifolia” are distributed throughout temperate waters of North and South America with almost no geographic overlap and exhibit an incongruous, non-mirrored, distribution across the equator. This study evaluates the degree of genetic divergence between M. “pyrifera” and M. “integrifolia” across 18 populations in Chile and California using whole-genome sequencing and single-nucleotide polymorphism markers. Our results based on a principal component analysis, admixture clustering by genetic similarity, and phylogenetic inference demonstrate that M. “pyrifera” and M. “integrifolia” are genetically distinguishable. Analyses reveal separation by Northern and Southern Hemispheres and between morphs within hemispheres, suggesting that the convergent “integrifolia” morphology arose separately in each hemisphere. This is the first study to use whole-genome sequencing to understand genetic divergence in giant kelp ecomorphs, identifying 83 potential genes under selection and providing novel insights about Macrocystis evolution that were not evident with previous genetic techniques. Future studies are needed to uncover the environmental forces driving local adaptation and presumed convergent evolution of these morphs.

Environmental Sciences & Ecology↗

The influence of physical and algorithmic factors on simulated far-field waveforms and source–time functions of underground explosions using unsupervised machine learning

SUMMARY Characterizing explosion sources and differentiating between earthquake and underground explosions using distributed seismic networks becomes non-trivial when explosions are detonated in cavities or heterogeneous ground material. Moreover, there is little understanding of how changes in subsurface physical properties affect the far-field waveforms we record and use to infer information about the source. Simulations of underground explosions and the resultant ground motions can be a powerful tool to systematically explore how different subsurface properties affect far-field waveform features, but there are added variables that arise from how we choose to model the explosions that can confound interpretation. To assess how both subsurface properties and algorithmic choices affect the seismic wavefield and the estimated source functions, we ran a series of 2-D axisymmetric non-linear numerical explosion experiments and wave propagation simulations that explore a wide array of parameters. We then inverted the synthetic far-field waveform data using a linear inversion scheme to estimate source–time functions (STFs) for each simulation case. We applied principal component analysis (PCA), an unsupervised machine learning method, to both the far-field waveforms and STFs to identify the most important factors that control variance in the waveform data and differences between cases. For the far-field waveforms, the largest variance occurs in the shallower radial receiver channels in the 0–50 Hz frequency band. For the STFs, both peak amplitude and rise times across different frequencies contribute to the variance. We find that the ground equation of state (i.e. lithology and rheology) and the explosion emplacement conditions (i.e. tamped versus cavity) have the greatest effect on the variance of the far-field waveforms and STFs, with the ground yield strength and fracture pressure being secondary factors. Differences in the PCA results between the far-field waveforms and STFs could possibly be due to near-field non-linearities of the source that are not accounted for in the estimation of STFs and could be associated with yield strength, fracture pressure, cavity radius and cavity shape parameters. Other algorithmic parameters are found to be less important and cause less variance in both the far-field waveforms and STFs, meaning algorithmic choices in how we model explosions are less important, which is encouraging for the further use of explosion simulations to study how physical Earth properties affect seismic waveform features and estimated STFs.

58 GEOSCIENCES↗

Surrogate modelling the Baryonic Universe – I. The colour of star formation

The spectral energy distribution of a galaxy emerges from the complex interplay of many physical ingredients, including its star formation history (SFH), metallicity evolution, and dust properties. Using GALAXPY , a new galaxy spectral prediction tool, and SFHs predicted by the empirical model UNIVERSEMACHINE and the cosmological hydrodynamical simulation IllustrisTNG, we isolate the influence of SFH on optical and near-infrared colours from 320 to 1080 Å at z = 0. By carrying out a principal component analysis, we demonstrate that physically motivated SFH variations modify galaxy colours along a single direction in colour space: the SFH-direction. We find that the projection of a galaxy’s present-day colours on to the SFH-direction is almost completely regulated by the fraction of stellar mass that the galaxy formed over the last billion years. Together with cosmic downsizing, this results in galaxies becoming redder as their host halo mass increases. We additionally study the change in galaxy colours due to variations in metallicity, dust attenuation, and nebular emission lines, finding that these properties vary broad-band colours along distinct directions in colour space relative to the SFH-direction. Finally, we show that the colours of low-redshift Sloan Digital Sky Survey galaxies span an ellipsoid with significant extent along two independent dimensions, and that the SFH-direction is well-aligned with the major axis of this ellipsoid. Our analysis supports the conclusion that variations in SFH are the dominant influence on present-day galaxy colours, and that the nature of this influence is strikingly simple.

79 ASTRONOMY AND ASTROPHYSICS↗

Non-Gaussianity in the weak lensing correlation function likelihood – implications for cosmological parameter biases

ABSTRACT We study the significance of non-Gaussianity in the likelihood of weak lensing shear two-point correlation functions, detecting significantly non-zero skewness and kurtosis in 1D marginal distributions of shear two-point correlation functions in simulated weak lensing data. We examine the implications in the context of future surveys, in particular LSST, with derivations of how the non-Gaussianity scales with survey area. We show that there is no significant bias in 1D posteriors of Ωm and σ8 due to the non-Gaussian likelihood distributions of shear correlations functions using the mock data (100 deg2). We also present a systematic approach to constructing approximate multivariate likelihoods with 1D parametric functions by assuming independence or more flexible non-parametric multivariate methods after decorrelating the data points using principal component analysis (PCA). While the use of PCA does not modify the non-Gaussianity of the multivariate likelihood, we find empirically that the 1D marginal sampling distributions of the PCA components exhibit less skewness and kurtosis than the original shear correlation functions. Modelling the likelihood with marginal parametric functions based on the assumption of independence between PCA components thus gives a lower limit for the biases. We further demonstrate that the difference in cosmological parameter constraints between the multivariate Gaussian likelihood model and more complex non-Gaussian likelihood models would be even smaller for an LSST-like survey. In addition, the PCA approach automatically serves as a data compression method, enabling the retention of the majority of the cosmological information while reducing the dimensionality of the data vector by a factor of ∼5.

79 ASTRONOMY AND ASTROPHYSICS↗

Reionization history constraints from neural network based predictions of high-redshift quasar continua

ABSTRACT Observations of the early Universe suggest that reionization was complete by z ∼ 6, however, the exact history of this process is still unknown. One method for measuring the evolution of the neutral fraction throughout this epoch is via observing the Lyα damping wings of high-redshift quasars. In order to constrain the neutral fraction from quasar observations, one needs an accurate model of the quasar spectrum around Lyα, after the spectrum has been processed by its host galaxy but before it is altered by absorption and damping in the intervening intergalactic medium (IGM). In this paper, we present a novel machine learning approach, using artificial neural networks, to reconstruct quasar continua around Lyα. Our Quasar Spectra from Artificial Neural Network based predictive Regression Algorithm(QSANNdRA) improves the error in this reconstruction compared to the state-of-the-art principal component analysis (PCA) based model in the literature by 14.2 per cent on average, and provides an improvement of 6.1 per cent on average when compared to an extension thereof. In comparison with the extended PCA model, QSANNdRA further achieves an improvement of 22.1 per cent and 16.8 per cent when evaluated on low-redshift quasars most similar to the two high-redshift quasars under consideration, ULAS J1120+0641 at z = 7.0851 and ULAS J1342+0928 at z = 7.5413, respectively. Using our more accurate reconstructions of these two z > 7 quasars, we estimate the neutral fraction of the IGM using a homogeneous reionization model and find $\bar{x}_\mathrm{H\, \small{I}} = 0.25^{+0.05}_{-0.05}$ at z = 7.0851 and $\bar{x}_\mathrm{H\, \small{I}} = 0.60^{+0.11}_{-0.11}$ at z = 7.5413. Our results are consistent with the literature and favour a rapid end to reionization.

Ďurovčíková, Dominika↗

Surrogate modelling the Baryonic Universe II: On forward modelling the colours of individual and populations of galaxies

ABSTRACT Among the properties shaping the light of a galaxy, the star formation history (SFH) is one of the most challenging to model due to the variety of correlated physical processes regulating star formation. In this work, we leverage the stellar population synthesis model fsps, together with SFHs predicted by the hydrodynamical simulation IllustrisTNG and the empirical model universemachine, to study the impact of star formation variability on galaxy colours. We start by introducing a model-independent metric to quantify the burstiness of a galaxy formation model, and we use this metric to demonstrate that universemachine predicts SFHs with more burstiness relative to IllustrisTNG. Using this metric and principal component analysis, we construct families of SFH models with adjustable variability, and we show that the precision of broad-band optical and near-infrared colours degrades as the level of unresolved short-term variability increases. We use the same technique to demonstrate that variability in metallicity and dust attenuation presents a practically negligible impact on colours relative to star formation variability. We additionally provide a model-independent fitting function capturing how the level of unresolved star formation variability translates into imprecision in predictions for galaxy colours; our fitting function can be used to determine the minimal SFH model that reproduces colours with some target precision. Finally, we show that modelling the colours of individual galaxies with per cent-level precision demands resorting to complex SFH models, while producing precise colours for galaxy populations can be achieved using models with just a few degrees of freedom.

Chaves-Montero, Jonás (ORCID:0000000295534261)↗

Exploring the dependence of gas cooling and heating functions on the incident radiation field with machine learning

ABSTRACT Gas cooling and heating functions play a crucial role in galaxy formation. But, it is computationally expensive to exactly compute these functions in the presence of an incident radiation field. These computations can be greatly sped up by using interpolation tables of pre-computed values, at the expense of making significant and sometimes even unjustified approximations. Here, we explore the capacity of machine learning to approximate cooling and heating functions with a generalized radiation field. Specifically, we use the machine learning algorithm XGBoost to predict cooling and heating functions calculated with the photoionization code cloudy at fixed metallicity, using different combinations of photoionization rates as features. We perform a constrained quadratic fit in metallicity to enable a fair comparison with traditional interpolation methods at arbitrary metallicity. We consider the relative importance of various photoionization rates through both a principal component analysis (PCA) and calculation of SHapley Additive exPlanation (shap) values for our XGBoost models. We use feature importance information to select different subsets of rates to use in model training. Our XGBoost models outperform a traditional interpolation approach at each fixed metallicity, regardless of feature selection. At arbitrary metallicity, we are able to reduce the frequency of the largest cooling and heating function errors compared to an interpolation table. We find that the primary bottleneck to increasing accuracy lies in accurately capturing the metallicity dependence. This study demonstrates the potential of machine learning methods such as XGBoost to capture the non-linear behaviour of cooling and heating functions.

79 ASTRONOMY AND ASTROPHYSICS↗

Detecting anomalous SRF cavity behavior with unsupervised learning

We present an unsupervised learning framework for detecting anomalous superconducting radio-frequency (SRF) cavity behavior at the Continuous Electron Beam Accelerator Facility (CEBAF), emphasizing its initial performance and effectiveness. Key to the system’s success was the development of data acquisition systems (DAQs) that capture fast-sampled, information-rich signals, essential for detecting transient effects. The approach involves creating daily cavity-specific models using principal component analysis to handle variations in rf signal behavior and mitigate performance degradation from data drift. This unsupervised method eliminates the need for expensive labeling by continuously updating models with recent data. Deployed and operational for 3 months before a scheduled shutdown, the system successfully identified several issues with DAQ signals, confirming its effectiveness. Despite access to only a fraction of CEBAF’s SRF cavity signals, the framework efficiently detected several instances requiring intervention, demonstrating a significant improvement over traditional, labor-intensive methods of manual plot inspection. Published by the American Physical Society 2025

43 PARTICLE ACCELERATORS↗

Statistical tools for a better optical model

Background: Modern statistical tools provide the ability to compare the information content of observables and provide a path to explore which experiments would be most useful to give insight into and constrain theoretical models. Purpose: Here we study three such tools in the context of nuclear reactions with the goal of constraining the optical potential. Method: The three statistical tools examined are (i) the principal component analysis, (ii) the sensitivity analysis based on derivatives, and (iii) the Bayesian evidence. We first apply these tools to a toy-model case, comparing the form of the imaginary part of the optical potential. Then we consider two different reaction observables, elastic angular distributions and polarization data for reactions on 48 Ca and 208 Pb at two different beam energies. Results: For the toy-model case, we find significant discrimination power in the sensitivities and the Bayesian evidence, showing clearly that the volume imaginary term is more useful to describe scattering at higher energies. When comparing between elastic cross sections and polarization data using realistic optical models, sensitivity studies indicate that both observables are roughly equally sensitive but the variability of the optical model parameters is strongly angle dependent. The Bayesian evidence shows some variability between the two observables, but the Bayes factor obtained is not sufficient to discriminate between angular distributions and polarization. Conclusions: From the cases considered, we conclude that, in general, elastic scattering angular distributions have similar impact in constraining the optical potential parameters compared with the polarization data. The angular ranges for the optimum experimental constraints can vary significantly with the observable considered.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Training and projecting: A reduced basis method emulator for many-body physics

Here, we present the reduced basis method as a tool for developing emulators for equations with tun able parameters within the context of the nuclear many-body problem. The method uses a basis expansion informed by a set of solutions for a few values of the model parameters and then projects the equations over a well-chosen low-dimensional subspace. We connect some of the results in the eigenvector continuation literature to the formalism of reduced basis methods and show how these methods can be applied to a broad set of problems. As we illustrate, the possible success of the formalism on such problems can be diagnosed beforehand by a principal component analysis. We apply the reduced basis method to the one-dimensional Gross-Pitaevskii equation with a harmonic trap ping potential and to nuclear density functional theory for 48 Ca, achieving speed-ups of more than x150 and x250, respectively, when compared to traditional solvers. The outstanding performance of the approach, together with its straightforward implementation, show promise for its application to the emulation of computationally demanding calculations, including uncertainty quantification.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

New method for fitting coefficients in standard model effective theory

We present an alternative method for carrying out a principal-component analysis of Wilson coefficients in standard model effective field theory (SMEFT). The method is based on singular-value decomposition (SVD). The SVD method provides information about the sensitivity of experimental observables to physics beyond the standard model that is not accessible in the Fisher-information method. In principle, the SVD method can also have computational advantages over diagonalization of the Fisher information matrix. We demonstrate the SVD method by applying it to the dimension-6 coefficients for the process of top-quark decay to a b quark and a W boson and use this example to illustrate some pitfalls in widely used fitting procedures. We also outline an iterative procedure for applying the SVD method to dimension-8 SMEFT coefficients.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Machine learning BPS spectra and the gap conjecture

We explore statistical properties of Bogomol’nyi-Prasad-Sommerfield q-series for strongly coupled supersymmetric theories that correspond to a particular family of three-manifolds. We discover that gaps between exponents in the -series are statistically more significant at the beginning of the -series compared to gaps that appear in higher powers of. Our observations are obtained by calculating saliencies of -series features used as input data for principal component analysis, which is a standard example of an explainable machine learning technique that allows for a direct calculation and a better analysis of feature saliencies.

97 MATHEMATICS AND COMPUTING↗

Symbolic pregression: Discovering physical laws from distorted video

In this work, we present a method for unsupervised learning of equations of motion for objects in raw and optionally distorted unlabeled synthetic video (or, more generally, for discovering and modeling predictable features in time-series data). We first train an autoencoder that maps each video frame into a low-dimensional latent space where the laws of motion are as simple as possible, by minimizing a combination of nonlinearity, acceleration, and prediction error. Differential equations describing the motion are then discovered using Pareto-optimal symbolic regression. We find that our pre-regression (“pregression”) step is able to rediscover Cartesian coordinates of unlabeled moving objects even when the video is distorted by a generalized lens. Using intuition from multidimensional knot theory, we find that the pregression step is facilitated by first adding extra latent space dimensions to avoid topological problems during training and then removing these extra dimensions via principal component analysis. An inertial frame is autodiscovered by minimizing the combined equation complexity for multiple experiments.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Random projection using random quantum circuits

The random sampling task performed by Google's Sycamore processor gave us a glimpse of the “quantum supremacy era.” This has definitely shed some light on the power of random quantum circuits in this abstract task of sampling outputs from the (pseudo)random circuits. In this paper, we explore a practical near-term use of local random quantum circuits in dimensional reduction of large low-rank data sets. We make use of the well-studied dimensionality reduction technique called the random projection method. This method has been extensively used in various applications such as image processing, logistic regression, entropy computation of low-rank matrices, etc. We prove that the matrix representations of local random quantum circuits with sufficiently shorter depths [ ∼ O ( n ) ] serve as good candidates for random projection. We demonstrate numerically that their projection abilities are not far off from the computationally expensive classical principal components analysis on MNIST and CIFAR-100 image datasets. We also benchmark the performance of quantum random projection against the commonly used classical random projection in the tasks of dimensionality reduction of image data sets and computing von Neumann entropies of large low-rank density matrices. And finally, using variational quantum singular value decomposition, we demonstrate a near-term implementation of extracting the singular vectors with dominant singular values after quantum random projecting a large low-rank matrix to lower dimensions. All such numerical experiments unequivocally demonstrate the ability of local random circuits to randomize a large Hilbert space at sufficiently shorter depths with robust retention of properties of large data sets in reduced dimensions. Published by the American Physical Society 2024

Kumaran, Keerthi (ORCID:0009000949125721)↗

Active learning approach to simulations of strongly correlated matter with the ghost Gutzwiller approximation

Quantum embedding (QE) methods such as the ghost Gutzwiller approximation (gGA) offer a powerful approach to simulating strongly correlated systems, but come with the computational bottleneck of computing the ground state of an auxiliary embedding Hamiltonian (EH) iteratively. In this work, we introduce an active learning (AL) framework integrated within the gGA to address this challenge. The methodology is applied to the single-band Hubbard model and results in a significant reduction in the number of instances where the EH must be solved. Through a principal component analysis (PCA), we find that the EH parameters form a low-dimensional structure that is largely independent of the geometric specifics of the systems, especially in the strongly correlated regime. Our AL strategy enables us to discover this low-dimensionality structure on the fly, while leveraging it for reducing the computational cost of gGA, laying the groundwork for more efficient simulations of complex strongly correlated materials. Published by the American Physical Society 2024

36 MATERIALS SCIENCE↗

Validation of non-negative matrix factorization for rapid assessment of large sets of atomic pair distribution function data

The use of the non-negative matrix factorization (NMF) technique is validated for automatically extracting physically relevant components from atomic pair distribution function (PDF) data from time-series data such as in situ experiments. The use of two matrix-factorization techniques, principal component analysis and NMF, on PDF data is compared in the context of a chemical synthesis reaction taking place in a synchrotron beam, applying the approach to synthetic data where the correct composition is known and on measured PDFs from previously published experimental data. The NMF approach yields mathematical components that are very close to the PDFs of the chemical components of the system and a time evolution of the weights that closely follows the ground truth. Lastly, it is discussed how this would appear in a streaming context if the analysis were being carried out at the beamline as the experiment progressed.

36 MATERIALS SCIENCE↗