Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Principal component analysis”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 469 records · Page 26

Active learning approach to simulations of strongly correlated matter with the ghost Gutzwiller approximation

Quantum embedding (QE) methods such as the ghost Gutzwiller approximation (gGA) offer a powerful approach to simulating strongly correlated systems, but come with the computational bottleneck of computing the ground state of an auxiliary embedding Hamiltonian (EH) iteratively. In this work, we introduce an active learning (AL) framework integrated within the gGA to address this challenge. The methodology is applied to the single-band Hubbard model and results in a significant reduction in the number of instances where the EH must be solved. Through a principal component analysis (PCA), we find that the EH parameters form a low-dimensional structure that is largely independent of the geometric specifics of the systems, especially in the strongly correlated regime. Our AL strategy enables us to discover this low-dimensionality structure on the fly, while leveraging it for reducing the computational cost of gGA, laying the groundwork for more efficient simulations of complex strongly correlated materials. Published by the American Physical Society 2024

36 MATERIALS SCIENCE↗

Validation of non-negative matrix factorization for rapid assessment of large sets of atomic pair distribution function data

The use of the non-negative matrix factorization (NMF) technique is validated for automatically extracting physically relevant components from atomic pair distribution function (PDF) data from time-series data such as in situ experiments. The use of two matrix-factorization techniques, principal component analysis and NMF, on PDF data is compared in the context of a chemical synthesis reaction taking place in a synchrotron beam, applying the approach to synthetic data where the correct composition is known and on measured PDFs from previously published experimental data. The NMF approach yields mathematical components that are very close to the PDFs of the chemical components of the system and a time evolution of the weights that closely follows the ground truth. Lastly, it is discussed how this would appear in a streaming context if the analysis were being carried out at the beamline as the experiment progressed.

36 MATERIALS SCIENCE↗

Machine learning inversion from scattering for mechanically driven polymers

A machine learning inversion method is developed for analyzing scattering functions of mechanically driven polymers and extracting the corresponding feature parameters, which include energy parameters and conformation variables. The polymer is modeled as a chain of fixed-length bonds constrained by bending energy, and it is subject to external forces such as stretching and shear. We generate a data set consisting of random combinations of energy parameters, including bending modulus, stretching and shear force, along with Monte Carlo-calculated scattering functions and conformation variables such as end-to-end distance, radius of gyration and off-diagonal component of the gyration tensor. The effects of the energy parameters on the polymer are captured by the scattering function, and principal component analysis ensures the feasibility of the machine learning inversion. Finally, we train a Gaussian process regressor using part of the data set as a training set and validate the trained regressor for inversion using the rest of the data. The regressor successfully extracts the feature parameters.

Gaussian process regressors↗

Unsupervised learning approaches to characterizing heterogeneous samples using X-ray single-particle imaging

One of the outstanding analytical problems in X-ray single-particle imaging (SPI) is the classification of structural heterogeneity, which is especially difficult given the low signal-to-noise ratios of individual patterns and the fact that even identical objects can yield patterns that vary greatly when orientation is taken into consideration. Proposed here are two methods which explicitly account for this orientation-induced variation and can robustly determine the structural landscape of a sample ensemble. The first, termed common-line principal component analysis (PCA), provides a rough classification which is essentially parameter free and can be run automatically on any SPI dataset. The second method, utilizing variation auto-encoders (VAEs), can generate 3D structures of the objects at any point in the structural landscape. Both these methods are implemented in combination with the noise-tolerant expand–maximize–compress (EMC) algorithm and its utility is demonstrated by applying it to an experimental dataset from gold nanoparticles with only a few thousand photons per pattern. Both discrete structural classes and continuous deformations are recovered. These developments diverge from previous approaches of extracting reproducible subsets of patterns from a dataset and open up the possibility of moving beyond the study of homogeneous sample sets to addressing open questions on topics such as nanocrystal growth and dynamics, as well as phase transitions which have not been externally triggered.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

SO(3)-invariant PCA with application to molecular data

Principal component analysis (PCA) is a fundamental technique for dimensionality reduction and denoising; however, its application to three-dimensional data with arbitrary orientations -- common in structural biology -- presents significant challenges. A naive approach requires augmenting the dataset with many rotated copies of each sample, incurring prohibitive computational costs. In this paper, we extend PCA to 3D volumetric datasets with unknown orientations by developing an efficient and principled framework for SO(3)-invariant PCA that implicitly accounts for all rotations without explicit data augmentation. By exploiting underlying algebraic structure, we demonstrate that the computation involves only the square root of the total number of covariance entries, resulting in a substantial reduction in complexity. We validate the method on real-world molecular datasets, demonstrating its effectiveness and opening up new possibilities for large-scale, high-dimensional reconstruction problems.

Fraiman, Michael [Tel Aviv Univ., Tel Aviv (Israel↗

W2VPCA: A Machine Learning Method for Measuring Attitudes With Natural Language

Company strategy influences many decisions in freight transportation. Behavioral models of company decision-making therefore could benefit from including strategy variables. However, strategy is difficult to observe and quantify. Attitudinal surveys of company executives can be used to collect measurements of latent strategy to use in quantitative models. However, surveys are costly and burdensome. Text mining methods to collect measurements overcome these issues somewhat, but typically require manual intervention and ignore the context of words, which can be problematic. This study introduces a new machine learning method to generate strategy measurement data from existing big text data. The new method, called W2VPCA, combines Natural Language Processing and Principal Components Analysis. W2VPCA produces measurement data that serve as quantitative indicators of latent strategy in behavioral models. W2VPCA is unsupervised, data-driven, and uses information on word context. We apply W2VPCA to generate measurements of latent strategies using readily available, large-scale text data: annual company reports. The empirical measurements are used successfully to associate two latent strategies, one focusing on distribution and the other on products, with truck fleet and distribution center outsourcing decisions. The main empirical outcome is that the W2VPCA measurements outperform Bag-of-Words measurements in a psychometric analysis of latent firm strategies. While this study focuses on freight behavioral models, W2VPCA may also have applications in behavioral modeling in other domains.

97 MATHEMATICS AND COMPUTING↗

Identifying Entangled Physics Relationships Through Sparse Matrix Decomposition to Inform Plasma Fusion Design

We report a sustainable burn platform through inertial confinement fusion (ICF) has been an ongoing challenge for over 50 years. Mitigating engineering limitations and improving the current design involves an understanding of the complex coupling of physical processes. While sophisticated simulation codes are used to model ICF implosions, these tools contain necessary numerical approximation but miss physical processes that limit predictive capability. Identification of relationships between controllable design inputs to ICF experiments and measurable outcomes (e.g., neutron yield, neutron velocity, areal density) from performed experiments can help guide the future design of experiments and development of simulation codes, to potentially improve the accuracy of the computational models used to simulate ICF experiments. We use sparse matrix decomposition methods to identify clusters of a few related design variables. Sparse principal component analysis (SPCA) identifies groupings that are related to the physical origin of the variables (laser, hohlraum, and capsule). A variable importance analysis finds that in addition to variables highly correlated with neutron yield, such as picket power and laser energy, variables that represent a dramatic change of the ICF design, such as number of pulse steps, are also very important. The obtained sparse components are then used to train a random forest (RF) regression surrogate for predicting total yield. The RF performance on the training and testing data compares with the performance of the RF trained using all the design variables considered. This work is intended to inform design changes in future ICF experiments by augmenting the expert intuition and simulation results.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Prediction of DIII-D Pedestal Structure From Externally Controllable Parameters

The sharp increase of pressure at the edge of a high confinement mode (H-mode) plasma, the pedestal, strongly impacts overall plasma performance. Predicting the pedestal is a necessity to control and optimize tokamak operations. Here, an experimental data-driven machine learning (ML) approach is presented that predicts the pedestal heights and widths of electron density (n e ) and electron temperature (T e ) profiles as well as the separatrix ne from externally controllable parameters such as the plasma shape, heating method and power, and gas puff rate and integrated gas puff. The OMFIT framework was used with DIII-D data to efficiently, robustly, and automatically build a database of pedestal parameters to train machine learning models. Database creation was enabled by the search engine tool for DIII-D data, TokSearch, which parallelizes data fetching, enabling fast searches through basic signals of thousands of DIII-D shots and selection of relevant time intervals. Principal Component Analysis (PCA) separated the database into three clusters that represent classes of plasma shapes that are regularly used in DIII-D. The most important parameters for setting the pedestal structure were plasma current (I p ), toroidal magnetic field (B Φ ), neutral beam heating power (P NBI ) and shaping quantities. The Deep Jointly Informed Neural Networks (DJINN) algorithm was applied to identify suitable neural network (NN) architectures that appropriately capture the features of the pedestal database. Separate NNs were implemented for each pedestal parameter, and ensembling methods were used to improve the prediction accuracy and allowed estimation of the prediction uncertainty. The pedestal predictions of the test dataset lie within the measurement uncertainties of the pedestal parameters. The NN outperformed simple Linear Regression (LR) analysis, indicating non-linear dependencies in the pedestal structure. The presented achievements illustrate a promising path for future research, using feature extraction to infer experimental trends and thereby improve pedestal models as well as deploying NN for a fast pedestal prediction in DIII-D scenario development.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Predicting soil carbon changes in switchgrass grown on marginal lands under climate change and adaptation strategies

Abstract The United States Great Lakes Region (USGLR) is a critical geographic area for future bioenergy production. Switchgrass ( Panicum virgatum ) is widely considered a carbon (C)‐neutral or C‐negative bioenergy production system, but projected increases in air temperature and precipitation due to climate change might substantially alter soil organic C (SOC) dynamics and storage in soils. This study examined long‐term SOC changes in switchgrass grown on marginal land in the USGLR under current and projected climate, predicted using a process‐based model (Systems Approach to Land‐Use Sustainability) extensively calibrated with a wealth of plant and soil measurements at nine experimental sites. Simulations indicate that these soils are likely a net C sink under switchgrass (average gain 0.87 Mg C ha −1 year −1 ), although substantial variation in the rate of SOC accumulation was predicted (range: 0.2–1.3 Mg C ha −1 year −1 ). Principal component analysis revealed that the predicted intersite variability in SOC sequestration was related in part to differences in climatic characteristics, and to a lesser extent, to heterogeneous soils. Although climate change impacts on switchgrass plant growth were predicted to be small (4%–6% decrease on average), the increased soil respiration was predicted to partially negate SOC accumulations down to 70% below historical rates in the most extreme scenarios. Increasing N fertilizer rate and decreasing harvest intensity both had modest SOC sequestration benefits under projected climate, whereas introducing genotypes better adapted to the longer growing seasons was a much more effective strategy. Best‐performing adaptation scenarios were able to offset >60% of the climate change impacts, leading to SOC sequestration 0.7 Mg C ha −1 year −1 under projected climate. On average, this was 0.3 Mg C ha −1 year −1 more C sequestered than the no adaptation baseline. These findings provide crucial knowledge needed to guide policy and operational management for maximizing SOC sequestration of future bioenergy production on marginal lands in the USGLR.

09 BIOMASS FUELS↗

Assisted migration in a warmer and drier climate: less climate buffering capacity, less facilitation and more fires at temperate latitudes?

Assisted tree migration has been proposed as a conceptual solution to mitigate lags in biotic responses to anthropogenic climate change. The rationale behind this concept is that tree species currently growing under warmer and drier climates will be more resistant and resilient to the new climatic conditions than tree species naturally growing in currently wetter and colder climates. However, we hypothesize that, by being more stress‐tolerant to warmer and drier conditions, translocated species should exhibit different functional attributes, which could induce important ecological and societal costs and overcome the desired benefits of maintaining wood production and other ecosystem services. We used principal component analysis (PCA) to analyze variation in seven traits of 106 tree and tall shrub species from contrasting latitudinal distributions in western North America and Europe to predict the potential functional changes of forest ecosystems due to the translocation of tree species from low to high latitudes. We show that species from both continents differed primarily by their position on the leaf economy spectrum (LES) and their size traits. Even though, in Europe, differences in LES were significantly correlated to species southern latitudinal positions, in both continents differences in size traits were significantly correlated to latitude. These results suggest that assisted migration by translocating more conservative species of shorter stature in currently cooler climates should decrease the buffering capacity of forest canopies, decrease facilitation for understory species, and increase wildfire risks, whose effects have the potential to accelerate climate warming through negative atmospheric feedback processes. As an alternative solution to assisted migration that may accelerate rather than mitigate climate change, we recommend that foresters gradually diversify the vertical structure and layering of the existing forest canopy to maintain a sustainable water cycle and energy balance between the soil, the tree and the atmosphere without increasing the wildfire risk.

Environmental Sciences & Ecology↗

Mass spectral imaging showing the plant growth-promoting rhizobacteria's effect on the Brachypodium awn

The plant growth-promoting rhizobacteria (PGPR) on the host plant surface play a key role in biological control and pathogenic response in plant functions and growth. However, it is difficult to elucidate the PGPR effect on plants. Such information is important in biomass production and conversion. Brachypodium distachyon (Brachypodium), a genomics model for bioenergy and native grasses, was selected as a C3 plant model; and the Gram-negative Pseudomonas fluorescens SBW25 ( P.) and Gram-positive Arthrobacter chlorophenolicus A6 ( A.) were chosen as representative PGPR strains. The PGPRs were introduced to the Brachypodium seed's awn prior to germination, and their possible effects on the seeding and growth were studied using different modes of time-of-flight secondary ion mass spectrometry (ToF-SIMS) measurements, including a high mass-resolution spectral collection and delayed image extraction. We observed key plant metabolic products and biomarkers, such as flavonoids, phenolic compounds, fatty acids, and auxin indole-3-acetic acid in the Brachypodium awns. Furthermore, principal component analysis and two-dimensional imaging analysis reveal that the Brachypodium awns are sensitive to the PGPR, leading to chemical composition and morphology changes on the awn surface. Our results show that ToF-SIMS can be an effective tool to probe cell-to-cell interactions at the biointerface. This work provides a new approach to studying the PGPR effects on awn and shows its potential for the research of plant growth in the future.

metabolite↗

Quantum advantage in learning from experiments

Quantum technology promises to revolutionize how we learn about the physical world. An experiment that processes quantum data with a quantum computer could have substantial advantages over conventional experiments in which quantum states are measured and outcomes are processed with a classical computer. We proved that quantum machines could learn from exponentially fewer experiments than the number required by conventional experiments. This exponential advantage is shown for predicting properties of physical systems, performing quantum principal component analysis, and learning about physical dynamics. Furthermore, the quantum resources needed for achieving an exponential advantage are quite modest in some cases. Conducting experiments with 40 superconducting qubits and 1300 quantum gates, we demonstrated that a substantial quantum advantage is possible with today’s quantum processors.

Science & Technology - Other Topics↗

VAIM-CFF: a variational autoencoder inverse mapper solution to Compton form factor extraction from deeply virtual exclusive reactions

We develop a new methodology for extracting Compton form factors (CFFs) from deeply virtual exclusive reactions such as the unpolarized DVCS cross section using a specialized inverse problem solver, a variational autoencoder inverse mapper (VAIM). The VAIM-CFF framework not only allows us access to a fitted solution set possibly containing multiple solutions in the extraction of all 8 CFFs from a single cross section measurement, but also accesses the lost information contained in the forward mapping from CFFs to cross section. We investigate various assumptions and their effects on the predicted CFFs such as cross section organization, number of extracted CFFs, use of uncertainty quantification technique, and inclusion of prior physics information. We then use dimensionality reduction techniques such as principal component analysis to visualize the missing physics information tracked in the latent space of the VAIM framework. Through re-framing the extraction of CFFs as an inverse problem, we gain access to fundamental properties of the problem not comprehensible in standard fitting methodologies: exploring the limits of the information encoded in deeply virtual exclusive experiments.

Accelerator Physics↗

Sampling-based Sublinear Low-rank Matrix Arithmetic Framework for Dequantizing Quantum Machine Learning

We present an algorithmic framework for quantum-inspired classical algorithms on close-to-low-rank matrices, generalizing the series of results started by Tang’s breakthrough quantum-inspired algorithm for recommendation systems [STOC’19]. Motivated by quantum linear algebra algorithms and the quantum singular value transformation (SVT) framework of Gilyén et al. [STOC’19], we develop classical algorithms for SVT that run in time independent of input dimension, under suitable quantum-inspired sampling assumptions. Our results give compelling evidence that in the corresponding QRAM data structure input model, quantum SVT does not yield exponential quantum speedups. Since the quantum SVT framework generalizes essentially all known techniques for quantum linear algebra, our results, combined with sampling lemmas from previous work, suffice to generalize all prior results about dequantizing quantum machine learning algorithms. In particular, our classical SVT framework recovers and often improves the dequantization results on recommendation systems, principal component analysis, supervised clustering, support vector machines, low-rank regression, and semidefinite program solving. We also give additional dequantization results on low-rank Hamiltonian simulation and discriminant analysis. Our improvements come from identifying the key feature of the quantum-inspired input model that is at the core of all prior quantum-inspired results: ℓ 2 -norm sampling can approximate matrix products in time independent of their dimension. We reduce all our main results to this fact, making our exposition concise, self-contained, and intuitive.

Computer Science↗

Visual Instance-aware Prompt Tuning

Visual Prompt Tuning (VPT) has emerged as a parameter-efficient fine-tuning paradigm for vision transformers, with conventional approaches utilizing dataset-level prompts that remain the same across all input instances. We observe that this strategy results in sub-optimal performance due to high variance in downstream datasets. To address this challenge, we propose Visual Instance-aware Prompt Tuning (ViaPT), which generates instance-aware prompts based on each individual input and fuses them with dataset-level prompts, leveraging Principal Component Analysis (PCA) to retain important prompting information. Moreover, we reveal that VPT-Deep and VPT-Shallow represent two corner cases based on a conceptual understanding, in which they fail to effectively capture instance-specific information, while random dimension reduction on prompts only yields performance between the two extremes. Instead, ViaPT overcomes these limitations by balancing dataset-level and instance-level knowledge, while reducing the amount of learnable parameters compared to VPT-Deep. Extensive experiments across 34 diverse datasets demonstrate that our method consistently outperforms state-of-the-art baselines, establishing a new paradigm for analyzing and optimizing visual prompts for vision transformers.

Xiao, Xi [ORNL] (ORCID:0009000009316982)↗

Automated Stellar Spectra Classification with Ensemble Convolutional Neural Network

Large sky survey telescopes have produced a tremendous amount of astronomical data, including spectra. Machine learning methods must be employed to automatically process the spectral data obtained by these telescopes. Classification of stellar spectra by applying deep learning is an important research direction for the automatic classification of high-dimensional celestial spectra. In this paper, a robust ensemble convolutional neural network (ECNN) was designed and applied to improve the classification accuracy of massive stellar spectra from the Sloan digital sky survey. We designed six classifiers which consist six different convolutional neural networks (CNN), respectively, to recognize the spectra in DR16. Then, according the cross-entropy testing error of the spectra at different signal-to-noise ratios, we integrate the results of different classifiers in an ensemble learning way to improve the effect of classification. The experimental result proved that our one-dimensional ECNN strategy could achieve 95.0% accuracy in the classification task of the stellar spectra, a level of accuracy that exceeds that of the classical principal component analysis and support vector machine model.

79 ASTRONOMY AND ASTROPHYSICS↗

clif

SAND2022-12903 O clif contains code to perform climate fingerprinting. It includes capabilities for principal component analysis of climate data, Fourier analysis, a collection of pre-processing scripts, and time series statistical tools. It uses a templated object-oriented programming approach to integrate all classes with the operator in Scikit Learn. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Nichol, Jeffrey↗

GP-BayesOpInf

SAND2025-01851O GP-BayesOpInf is a software tool that uses algorithms to combine Gaussian process regression, principal component analysis, and linear Bayesian inference to produce a probabilistic reduced-order model for time-dependent systems. Numerical examples include the compressible Euler equations for an ideal gas, a heat diffusion process with a nonlinear reaction term, and a set of ordinary differential equations describing a compartmental model in epidemiology. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

SciDAC↗