Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Principal component analysis”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

An a priori evaluation of a principal component and artificial neural network based combustion model in diesel engine conditions

A principal component analysis (PCA) and artificial neural network (ANN) based chemistry tabulation approach is presented. ANNs are used to map the thermochemical state onto a low-dimensional manifold consisting of five control variables that have been identified using PCA. Three canonical configurations are considered to train the PCA-ANN model: a series of homogeneous reactors, a nonpremixed flamelet, and a two-dimensional lifted flame. The performance of the model in predicting the thermochemical manifold of a spatially-developing turbulent jet flame in diesel engine thermochemical conditions is a priori evaluated using direct numerical simulation (DNS) data. The PCA-ANN approach is compared with a conventional tabulation approach (tabulation using ad hoc defined control variables and linear interpolation). The PCA-ANN model provides higher accuracy and requires several orders of magnitude less memory. Here, these observations indicate that the PCA-ANN model is superior for chemistry tabulation, especially for modelling complex chemistries that present multiple combustion modes as observed in diesel combustion. The performance of the PCA-ANN model is then compared to the optimal estimator, i.e. the conditional mean from the DNS. The results indicate that the PCA-ANN model gives high prediction accuracy, comparable to the optimal estimator, especially for major species and the thermophysical properties. Higher errors are observed for the minor species and reaction rate predictions when compared to the optimal estimator. It is shown that the prediction of minor species and reaction rates can be improved by using training data that exhibits a variation of parameters as observed in the turbulent flame. The output of the ANN is analysed to assess mass conservation. It is observed that the ANN incurs a mean absolute error of 0.05% in mass conservation. Furthermore, it is demonstrated that this error can be reduced by modifying the cost function of the ANN to penalise for deviation from mass conservation.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

The Fe 4+ / 3+ Redox Mechanism in NaFeO 2 : A Simultaneous Operando Nuclear Resonance and X-ray Scattering Study

Simultaneous operando Nuclear Forward Scattering and transmission X-ray diffraction and 57 Fe Mössbauer spectroscopy measurements were carried out in order to investigate the electrochemical mechanism of NaFeO 2 vs. Na metal using a specifically designed in situ cell. The obtained data were analysed using an alternative and innovative data analysis approach based on chemometric tools such as Principal Component Analysis (PCA) and Multivariate Curve Resolution - Alternating Least Squares (MCR-ALS). This approach, which allows the unbiased extraction of all possible information from the operando data, enabled the stepwise reconstruction of the independent “real” components permitting the description of the desodiation mechanism of NaFeO 2 . This wealth of information allows a clear description of the electrochemical reaction at the redox-active iron centres, and thus an improved comprehension of the cycling mechanisms of this material vs. sodium.

25 ENERGY STORAGE↗

Predicting non-linear stress–strain response of mesostructured cellular materials using supervised autoencoder

Recent breakthroughs in advanced manufacturing capabilities have made it possible to design and print sophisticated topologies of cellular structures using diverse engineering materials such as metals, polymers, and ceramics. In these architectured materials, it is often desirable to tailor the mechanical properties by altering the unit cell topology. This necessitates an in-depth understanding of how the topology of the unit cell structure affects the macroscopic behavior of the material in both the linear and the non-linear regimes encountered under large compression. Here, we have developed a machine learning (ML) approach capable of accelerating the prediction of the stress–strain response of a polymer-based cellular structure under uniaxial confined compression. As part of generating the training data for ML, 60,000 mesostructures were generated using a relatively novel approach based on cellular automata, and their corresponding stress–strain responses were obtained from the finite element simulations. Principal component analysis (PCA) was used to reduce the dimensionality of the stress–strain curves. With only 20 principal components, PCA captured 99.89% of the variance in the stress–strain curves while reducing the dimensionality by 5X. ML using supervised autoencoder was able to successfully speed up the prediction of the non-linear stress–strain response of a unit cell by up to 4600X. The proposed method can serve as an efficient data generation tool and a rapid means for predicting the structure–property relationship through accelerated forward modeling of cellular materials under compaction, in cases where the macroscopic stress–strain response is governed by the unit-cell topology.

36 MATERIALS SCIENCE↗

VEESA R package

SAND2024-04584O R package for applying the VEESA pipeline method is a technique used for explainable machine learning with functional data. The VEESA pipeline makes use of the elastic-shape analysis framework for functional data. It also implements functional principal component analysis and permutation feature importance. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy's National Nuclear Security Administration under contract DE-NA0003525.

Tucker, James↗

A flow cytometric assay to detect viability and persistence of Salmonella enterica subsp. enterica serotypes in nuclease-free water at 4 and 25°C

Salmonella spp. is one of the most isolated microorganisms reported to be responsible for human foodborne diseases and death. Water constitutes a major reservoir where the Salmonella spp. can persist and go undetected when present in low numbers. In this study, we assessed the viability of 12 serotypes of Salmonella enterica subsp. enterica for 160 days in nuclease-free water at 4 and 25°C using flow cytometry and Tryptic Soy Agar (TSA) plate counts. The results show that all 12 serotypes remain viable after 160 days in distilled water using flow cytometry, whereas traditional plate counts failed to detect ten serotypes incubated at 25°C. Moreover, the findings demonstrate that 4°C constitutes a more favorable environment where Salmonella can remain viable for prolonged periods without nutrients. Under such conditions, however, Salmonella exhibits a higher susceptibility to all tested antibiotics and benzalkonium chloride (BZK). The pre-enrichment with Universal Pre-enrichment Broth (UP) and 1/10 × Tryptic Soy broth (1/10 × TSB) resuscitated all tested serotypes on TSA plates, nevertheless cell size decreased after 160 days. Furthermore, phenotype microarray (PM) analysis of S. Inverness and S. Enteritidis combined with principal component analysis (PCA) revealed an inter-individual variability in serotypes with their phenotype characteristics, and the impact of long-term storage at 4 and 25°C for 160 days in nuclease-free water. This study provides an insight to Salmonella spp. long-term survivability at different temperatures and highlights the need for powerful tools to detect this microorganism to reduce the risk of disease transmission of foodborne pathogens via nuclease-free water.

59 BASIC BIOLOGICAL SCIENCES↗

Multivariate Analysis as a Tool for Validating Tester Matching

A method of applying Principal Component Analysis, Soft Independent Modeling of Class Analysis, and statistical analysis is described that can be applied to many types of testers to ascertain how well matched the performance of the testers in the analysis are to one another or how well matched a tester is to itself at a later time. This method is most useful for situations for which the same units have not been run across the testers being analyzed for matched performance.

Multari, Rosalie A [Sandia National Laboratories (↗

Explainable Machine Learning for Functional Data

Black-box machine learning models are recognized as useful tools for prediction applications, but the algorithmic complexity of some models causes interpretation challenges. Explainability methods have been proposed to provide insight into these models, but there is little research focused on supervised modeling with functional data inputs. We argue that, especially in applications of high consequence, it is important to explicitly model the functional dependence in a black-box analysis to not obscure or misrepresent patterns in explanations. As such, we propose the V ariable importance E xplainable E lastic S hape A nalysis (VEESA) pipeline for training supervised machine learning models with functional inputs. The pipeline is an analysis process that includes the data preprocessing, modeling, and post-hoc explanations. The preprocessing is done using elastic functional principal components analysis, which accounts for vertical and horizontal variability in functional data and, ultimately, allows for explanations in the original data space that identify the important functional variability without bias due to correlated variables. Here, we demonstrate the pipeline on two high-consequence applications: explosives classification for national security and inkjet printer identification in forensic science. The applications exhibit the VEESA pipeline’s ability to provide an understanding of the characteristics of the functional data useful for prediction. Code for implementing the pipeline is available in the veesa R package (and supplemental python code).

Elastic Shape Analysis↗

Comparison of machine learning techniques to optimize the analysis of plutonium surrogate material via a portable LIBS device

The utilization of machine learning techniques has become commonplace in the analysis of optical emission spectra. These methods are often limited to variants of principal components analysis (PCA), partial-least squares (PLS), and artificial neural networks (ANNs). A plethora of other techniques exist and are well established in the world of data science, yet are seldom investigated for their use in spectroscopic problems. In this study, machine learning techniques were used to analyze optical emission spectra of laser-induced plasma from ceria pellets doped with silicon in order to predict silicon content. Additionally, a boosted regression ensemble model was created, and its predictive accuracy was compared to that of traditional PCA, PLS, and ANN regression models. Boosted regression tree ensembles yielded fits with R-squared (R2) values as high as 0.964 and mean-squared errors of prediction (MSEPs) as low as 0.074, providing the most accurate predictive model. Neural networks performed with slightly lower R2 values and higher MSEPs compared to the ensemble methods, thus indicating susceptibility to overfitting.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Identification and correction of temporal and spatial distortions in scanning transmission electron microscopy

Scanning transmission electron microscopy (STEM) has become the technique of choice for quantitative characterization of atomic structure of materials, where the minute displacements of atomic columns from high-symmetry positions can be used to map strain, polarization, octahedra tilts, and other physical and chemical order parameter fields. The latter can be used as inputs into mesoscopic and atomistic models, providing insight into the correlative relationships and generative physics of materials on the atomic level. However, these quantitative applications of STEM necessitate understanding the microscope induced image distortions and developing the pathways to compensate them both as part of a rapid calibration procedure for in situ imaging, and the post-experimental data analysis stage. Here, we explore the spatiotemporal structure of the microscopic distortions in STEM using multivariate analysis of the atomic trajectories in the image stacks. Based on the behavior of principal component analysis (PCA), we develop the Gaussian process (GP)-based regression method for quantification of the distortion function. The limitations of such an approach and possible strategies for implementation as a part of in-line data acquisition in STEM are discussed. Here, the analysis workflow is summarized in a Jupyter notebook that can be used to retrace the analysis and analyze the reader's data.

36 MATERIALS SCIENCE↗

Clustering High-dimensional Toxicogenomics Data with Rare Signals

Toxicogenomics studies the gene and protein activities to drug treatments or toxic exposures. As the drugs and genes are numerous, toxicogenomics data are naturally high dimensional, with dimension sizes up to millions. In addition, the distribution of toxicogenomics data is oftentimes skewed, and they contain rare but important signals representing a cell or organism’s response to toxicity. The combination of high dimension and extremely skewed distribution of toxicogenomics data makes clustering analysis extremely challenging.We present our study of clustering toxicogenomics data using classical approaches such as principal component analysis as well as deep learning approaches such as auto-encoders. Our experiments show that these approaches fail to preserve rare signals and produce high-quality clusters. We then explore augmenting matrix factorization with deep learning techniques such as attention mechanism to produce latent representations for clustering. Our technique is able to better preserve rare signals after dimensionality reduction than prior approaches. Furthermore, we combine our augmented matrix factorization with a mechanism similar to autoencoder to balance separable clusters and low regeneration errors. Our experiments demonstrate better clustering with our proposed approach.

Cong, Guojing↗

A bi-level data-driven framework for fault-detection and diagnosis of HVAC systems

Long-term operation of heating, ventilation, and air conditioning (HVAC) systems will eventually lead to a range of HVAC system failures, resulting in excessive energy consumption and maintenance costs. Here, to avoid HVAC malfunctioning, fault detection diagnostic (FDD) is utilized as a common practice. Machine learning methods have lately received considerable interest for FDD analysis of HVAC systems due to their high detection accuracy. Meanwhile, HVAC malfunctions are regarded as rare occurrences, hence normal operating data samples are much more accessible than data samples in faulty and malfunctioning conditions. The dominating frequency of normal operation in HVAC datasets has also led to heavily biased classification algorithms within the literature. Moreover, the focus of previous literature has been on increasing the accuracy of the models which leads to a high number of false positives (misleading alarms) in the system. In order to enhance the performance of diagnostic procedures and fill the mentioned gaps, this study proposes a novel data-driven framework. A bi-level machine learning framework is developed for diagnosing faults in air handling units (AHUs) and rooftop units (RTUs) based on principal component analysis (PCA), time series anomaly detection, and random forest (RF). It is shown that PCA can reduce the dataset dimension with one principal component accounting for 95% of data variance. Also, the random forest could classify the faults with 89% precision for single-zone AHU, 85% precision for RTU, and 79% for multi-zone AHU. By proposing this framework, three persistent challenges are addressed: (I) minimizing false positives; (II) accounting for data imbalance; and (III) normal condition monitoring of equipment.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

DEEPEN Global Standardized Categorical Exploration Datasets for Magmatic Plays

DEEPEN stands for DE-risking Exploration of geothermal Plays in magmatic ENvironments. As part of the development of the DEEPEN 3D play fairway analysis (PFA) methodology for magmatic plays (conventional hydrothermal, superhot EGS, and supercritical), weights needed to be developed for use in the weighted sum of the different favorability index models produced from geoscientific exploration datasets. This was done using two different approaches: one based on expert opinions, and one based on statistical learning. This GDR submission includes the datasets used to produce the statistical learning-based weights. While expert opinions allow us to include more nuanced information in the weights, expert opinions are subject to human bias. Data-centric or statistical approaches help to overcome these potential human biases by focusing on and drawing conclusions from the data alone. The drawback is that, to apply these types of approaches, a dataset is needed. Therefore, we attempted to build comprehensive standardized datasets mapping anomalies in each exploration dataset to each component of each play. This data was gathered through a literature review focused on magmatic hydrothermal plays along with well-characterized areas where superhot or supercritical conditions are thought to exist. Datasets were assembled for all three play types, but the hydrothermal dataset is the least complete due to its relatively low priority. For each known or assumed resource, the dataset states what anomaly in each exploration dataset is associated with each component of the system. The data is only a semi-quantitative, where values are either high, medium, or low, relative to background levels. In addition, the dataset has significant gaps, as not every possible exploration dataset has been collected and analyzed at every known or suspected geothermal resource area, in the context of all possible play types. The following training sites were used to assemble this dataset: - Conventional magmatic hydrothermal: Akutan (from AK PFA), Oregon Cascades PFA, Glass Buttes OR, Mauna Kea (from HI PFA), Lanai (from HI PFA), Mt St Helens Shear Zone (from WA PFA), Wind River Valley (From WA PFA), Mount Baker (from WA PFA). - Superhot EGS: Newberry (EGS demonstration project), Coso (EGS demonstration project), Geysers (EGS demonstration project), Eastern Snake River Plain (EGS demonstration project), Utah FORGE, Larderello, Kakkonda, Taupo Volcanic Zone, Acoculco, Krafla. - Supercritical: Coso, Geysers, Salton Sea, Larderello, Los Humeros, Taupo Volcanic Zone, Krafla, Reyjanes, Hengill. **Disclaimer: Treat the supercritical fluid anomalies with skepticism. They are based on assumptions due to the general lack of confirmed supercritical fluid encounters and samples at the sites included in this dataset, at the time of assembling the dataset. The main assumption was that the supercritical fluid in a given geothermal system has shared properties with the hydrothermal fluid, which may not be the case in reality. Once the datasets were assembled, principal component analysis (PCA) was applied to each. PCA is an unsupervised statistical learning technique, meaning that labels are not required on the data, that summarized the directions of variance in the data. This approach was chosen because our labels are not certain, i.e., we do not know with 100% confidence that superhot resources exist at all the assumed positive areas. We also do not have data for any known non-geothermal areas, meaning that it would be challenging to apply a supervised learning technique. In order to generate weights from the PCA, an analysis of the PCA loading values was conducted. PCA loading values represent how much a feature is contributing to each principal component, and therefore the overall variance in the data.

15 GEOTHERMAL ENERGY↗

Towards multi-fidelity deep learning of wind turbine wakes

We report engineering wake models that accurately predict wake in a computationally efficient manner are very important for tasks such as layout optimization and control of wind farms. In this paper, we explore an application of deep learning (DL) to learn the wake model from hierarchies of physics-based approaches ranging from analytical models to an approximate form of the Reynolds-averaged Navier-Stokes equations. We first illustrate the application of principal component analysis to obtain a lower-dimensional representation that allows a computationally tractable training and deployment of DL models. Then, the DL model is trained to learn the mapping from input parameter space to the principal components, which are then used to reconstruct the three-dimensional flow field. Additionally, we investigate a composite framework consisting of two neural networks to learn the correlation between low- and high-fidelity data with Gauss and curl models treated as proxies for low- and high-fidelity models, respectively. The prediction from both DL models matches well with the high-fidelity data with a maximum relative percentage error for the kinetic energy flux of <1%. This work opens up possibilities for data-efficient construction of surrogate models for wake prediction that can be used to study the influence of wind speed and yaw angles on wind farm power production.

17 WIND ENERGY↗

Constraining the baryonic feedback with cosmic shear using the DES Year-3 small-scale measurements

ABSTRACT We use the small scales of the Dark Energy Survey (DES) Year-3 cosmic shear measurements, which are excluded from the DES Year-3 cosmological analysis, to constrain the baryonic feedback. To model the baryonic feedback, we adopt a baryonic correction model and use the numerical package baccoemu to accelerate the evaluation of the baryonic non-linear matter power spectrum. We design our analysis pipeline to focus on the constraints of the baryonic suppression effects, utilizing the implication given by a principal component analysis on the Fisher forecasts. Our constraint on the baryonic effects can then be used to better model and ameliorate the effects of baryons in producing cosmological constraints from the next-generation large-scale structure surveys. We detect the baryonic suppression on the cosmic shear measurements with a ∼2σ significance. The characteristic halo mass for which half of the gas is ejected by baryonic feedback is constrained to be $M_c \gt 10^{13.2} \, h^{-1} \, \mathrm{M}_{\odot }$ (95 per cent C.L.). The best-fitting baryonic suppression is $\sim 5{{\ \rm per\ cent}}$ at $k=1.0 \, {\rm Mpc}\ h^{-1}$ and $\sim 15{{\ \rm per\ cent}}$ at $k=5.0 \, {\rm Mpc} \ h^{-1}$. Our findings are robust with respect to the assumptions about the cosmological parameters, specifics of the baryonic model, and intrinsic alignments.

79 ASTRONOMY AND ASTROPHYSICS↗

Sensor selection and tool wear prediction with data‐driven models for precision machining

Abstract Estimation of tool wear in precision machining is vital in the traditional subtractive machining industry to reduce processing cost, improve manufacturing efficiency and product quality. In this vein, fusion of time and frequency‐domain features of commonly sensed signals can provide an early indication of tool wear and improve its prediction accuracy for prognostics and health management. This paper presents a data‐driven methodology and a complete tool chain for the inference of precision machining tool wear from fused machine measurements, such as cutting force, power, audio and vibration signals, and quantify the usefulness of each measurement. Indicators of tool wear are extracted from time‐domain signal statistics, frequency‐domain analysis, and time‐frequency domain analysis. Correlation coefficients between the extracted features (indicators) and the tool wear are used to select the most informative features. Principal Component Analysis and Partial Least‐Squares are used to reduce the dimensionality of the feature space. Regression models, including linear regression, support vector regression, Decision tree regression, neural network regression and Gaussian process regression, are used to predict the tool wear using data from a Haas milling machine performing spiral boss face milling. The performance of the regression models based on subsets of sensors validates the preliminary estimates about the saliency of the sensors. The experimental results show that the proposed methods can predict the machine tool wear precisely, with readily available sensor measurements. Neural network and Gaussian process regression were able to achieve good estimates of tool wear at different machine operating conditions. The most informative signal in predicting tool wear was shown to be the vibration signal. Time‐frequency domain features were the most informative features among the combination of features of three domains. In addition, using partial least squares components extracted from the original features of signals led to higher prediction accuracy.

Han, Seulki↗

Seeking regularity from irregularity: unveiling the synthesis–nanomorphology relationships of heterogeneous nanomaterials using unsupervised machine learning

Nanoscale morphology of functional materials determines their chemical and physical properties. However, despite increasing use of transmission electron microscopy (TEM) to directly image nanomorphology, it remains challenging to quantify the information embedded in TEM data sets, and to use nanomorphology to link synthesis and processing conditions to properties. We develop an automated, descriptor-free analysis workflow for TEM data that utilizes convolutional neural networks and unsupervised learning to quantify and classify nanomorphology, and thereby reveal synthesis–nanomorphology relationships in three different systems. While TEM records nanomorphology readily in two-dimensional (2D) images or three-dimensional (3D) tomograms, we advance the analysis of these images by identifying and applying a universal shape fingerprint function to characterize nanomorphology. After dimensionality reduction through principal component analysis, this function then serves as the input for morphology grouping through unsupervised learning. We demonstrate the wide applicability of our workflow to both 2D and 3D TEM data sets, and to both inorganic and organic nanomaterials, including tetrahedral gold nanoparticles mixed with irregularly shaped impurities, hybrid polymer-patched gold nanoprisms, and polyamide membranes with irregular and heterogeneous 3D crumple structures. In each of these systems, unsupervised nanomorphology grouping identifies both the diversity and the similarity of the nanomaterial across different synthesis conditions, revealing how synthetic parameters guide nanomorphology development. Our work opens possibilities for enhancing synthesis of nanomaterials through artificial intelligence and for understanding and controlling complex nanomorphology, both for 2D systems and in the far less explored case of 3D structures, such as those with embedded voids or hidden interfaces.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Image Distinguishability Analysis Testing Through Principal Components and Its Application to Hot Spot Scale Invariance

Hot spots are spatial regions of intense energy localization that govern initiation of secondary high explosives. Studies that characterize or compare simulated hot spots are frequently either qualitatively descriptive or resort to quantitative distribution functions that neglect stochastic variations and spatial correlations—effects that are also neglected in common comparison tests like the Kolmogorov–Smirnov test. To this end, we develop an image distinguishability analysis (IDA) test based on principal component (PC) analysis that makes pixel-by-pixel comparisons between small, for example, O(<10), image data sets. The IDA test makes comparisons through a generalized distance metric in the PC space and a test statistic that is derived to calculate mathematical equation-values. Here, we derive a statistical distribution and criticality criterion to determine whether images are distinguishable from established baselines. We apply the IDA test on images generated from molecular dynamics simulations of hot spots from pore collapse in TATB to assess scale invariance in the complex patterns of hot spots that form in a representative high explosive crystal. The IDA test shows that TATB hot spot spatial temperature fields and their derived temperature histograms exhibit scale-invariant features over specific intervals of shock orientation, strength, and initial pore diameter. However, the IDA test also shows that qualitatively different conclusions regarding invariance can be reached depending on whether the hot spot is treated as a spatially correlated field as opposed to a distribution function that lacks spatial information.

organic↗

Impact of duration and missing data on the long-term photovoltaic degradation rate estimation

Accurate quantification of photovoltaic (PV) system degradation rate (R D ) is essential for lifetime yield predictions. Although R D is a critical parameter, its estimation lacks a standardized methodology that can be applied on outdoor field data. The purpose of this paper is to investigate the impact of time period duration and missing data on R D by analyzing the performance of different techniques applied to synthetic PV system data at different linear R D patterns and known noise conditions. The analysis includes the application of different techniques to a 10-year synthetic dataset of a crystalline Silicon PV system, with emulated degradation levels and imputed missing data. Here, the analysis demonstrated that the accuracy of ordinary least squares (OLS), year-on-year (YOY), autoregressive integrated moving average (ARIMA) and robust principal component analysis (RPCA) techniques is affected by the evaluation duration with all techniques converging to lower R D deviations over the 10-year evaluation, apart from RPCA at high degradation levels. Moreover, the estimated R D is strongly affected by the amount of missing data. Filtering out the corrupted data yielded more accurate R D results for all techniques. It is proven that the application of a change-point detection stage is necessary and guidelines for accurate R D estimation are provided.

14 SOLAR ENERGY↗