Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Principal component analysis”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 271 records · Page 15

Data for System Analysis of Lipomyces starkeyi During Growth on Various Plant‑Based Sugars

Oleaginous yeasts have received significant attention due to their substantial lipid storage capability. The accumulated lipids can be utilized directly or processed into various bioproducts and biofuels. Lipomyces starkeyi is an oleaginous yeast capable of using multiple plant-based sugars, such as glucose, xylose, and cellobiose. It is, however, a relatively unexplored yeast due to limited knowledge about its physiology. In this study, we have evaluated the growth of L. starkeyi on different sugars and performed transcriptomic and metabolomic analyses to understand the underlying mechanisms of sugar metabolism. Principal component analysis showed clear differences resulting from growth on different sugars. We have further reported various metabolic pathways activated during growth on these sugars. We also observed non-specific regulation in L. starkeyi and have updated the gene annotations for the NRRL Y-11557 strain. This analysis provides a foundation for understanding the metabolism of these plant-based sugars and potentially valuable information to guide the metabolic engineering of L. starkeyi to produce bioproducts and biofuels.

Conversion↗

Systems analysis of Lipomyces starkeyi during growth on various plant-based sugars

Oleaginous yeasts have received significant attention due to their substantial lipid storage capability. The accumulated lipids can be utilized directly or processed into various bioproducts and biofuels. Lipomyces starkeyi is an oleaginous yeast capable of using multiple plant-based sugars, such as glucose, xylose, and cellobiose. It is, however, a relatively unexplored yeast due to limited knowledge about its physiology. In this study, we have evaluated the growth of L. starkeyi on different sugars and performed transcriptomic and metabolomic analyses to understand the underlying mechanisms of sugar metabolism. Principal component analysis showed clear differences resulting from growth on different sugars. We have further reported various metabolic pathways activated during growth on these sugars. We also observed non-specific regulation in L. starkeyi and have updated the gene annotations for the NRRL Y-11557 strain. Furthermore, this analysis provides a foundation for understanding the metabolism of these plant-based sugars and potentially valuable information to guide the metabolic engineering of L. starkeyi to produce bioproducts and biofuels.

59 BASIC BIOLOGICAL SCIENCES↗

Molecular identification of wines using in situ liquid SIMS and PCA analysis

Composition analysis in wine is gaining increasing attention because it can provide information about the wine quality, source, and nutrition. In this work, in situ liquid secondary ion mass spectrometry (SIMS) was applied to 14 representative wines, including six wines manufactured by a manufacturer in Washington State, United States, four Cabernet Sauvignon wines, and four Chardonnay wines from other different manufacturers and locations. In situ liquid SIMS has the unique advantage of simultaneously examining both organic and inorganic compositions from liquid samples. Principal component analysis (PCA) of SIMS spectra showed that red and white wines can be clearly differentiated according to their aromatic and oxygen-contained organic species. Furthermore, the identities of different wines, especially the same variety of wines, can be enforced with a combination of both organic and inorganic species. Meanwhile, in situ liquid SIMS is sample-friendly, so liquid samples can be directly analyzed without any prior sample dilution or separation. Taken together, we demonstrate the great potential of in situ liquid SIMS in applications related to the molecular investigation of various liquid samples in food science.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Solving the structure of “single-atom” catalysts using machine learning – assisted XANES analysis

We show that "single-atom” catalysts (SACs) have demonstrated excellent activity and selectivity in challenging chemical transformations such as photocatalytic CO 2 reduction. For heterogeneous photocatalytic SAC systems, it is essential to obtain sufficient information of their structure at the atomic level in order to understand reaction mechanisms. In this work, a SAC was prepared by grafting a molecular cobalt catalyst on a light-absorbing carbon nitride surface. Due to the sensitivity of the X-ray absorption near edge structure (XANES) spectra to subtle variances in the Co SAC structure in reaction conditions, different machine learning (ML) methods, including principal component analysis, K-means clustering, and neural network (NN), were utilized for in situ Co XANES data analysis. As a result, we obtained quantitative structural information of the SAC nearest atomic environment thereby extending the NN-XANES approach previously demonstrated for nanoparticles and size-selective clusters.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Machine Learning-Based Anomaly Detection for PMT Data Quality Monitoring in the SBN and DUNE

Maintaining high-quality detector data is essential for achieving the scientific objectives of the Short-Baseline Neutrino (SBN) Program at Fermilab. Current data quality monitoring (DQM) procedures rely primarily on threshold-based metrics and manual inspection of detector monitoring plots, making the detection of subtle or gradually developing anomalies both time-consuming and dependent on expert interpretation. This project developed and evaluated a machine-learning workflow for automatically identifying anomalous photomultiplier tube (PMT) channels in the Short-Baseline Near Detector (SBND) using optical-hit amplitude data. A Python-based analysis program was developed to process ROOT files, extract statistical features describing individual PMT amplitude distributions, and generate feature vectors for anomaly detection. These features were used to train an Isolation Forest model using data representing normal detector operation. The trained model was subsequently applied to independent detector runs to identify channels exhibiting statistically unusual behavior relative to the learned reference response. To support expert interpretation, the workflow generated complementary diagnostic products, including anomaly score distributions, normalized amplitude comparisons, decision-tree visualizations, and principal component analysis (PCA) projections. This project demonstrated the feasibility of integrating unsupervised machine learning into detector data-quality monitoring and developed a complete workflow for automated PMT performance assessment to aid expert-driven review. Beyond its technical contributions, the VFP appointment fostered a research collaboration between Aurora University and Fermilab and provided direct workforce development benefits by training the visiting faculty member in detector-scale machine-learning methods that are now being incorporated into undergraduate coursework and research. The methodology developed here provides a foundation for future applications to ProtoDUNE and other liquid argon time projection chamber (LArTPC) detectors, contributing to ongoing efforts to improve detector reliability, reduce manual monitoring requirements, and enable scalable data quality monitoring for future large-scale neutrino experiments, including the Deep Underground Neutrino Experiment (DUNE).

Colón Santana, Juan A. [Unlisted, US, IL]↗

Novel Cell-Type-Specific Drought-Responsive Proteins in Root Tips of Field-Grown Perennial Switchgrass

The root-tip region of plants, including the root cap, forms the most basal terminal of the root and exhibits a high degree of cellular complexity in terms of morphology, cytological function, and interaction with environmental cues in the soil. Cells in this region follow a developmental trajectory, transitioning from stem cells to meristematic cells, and ultimately to fully differentiated cell types. However, our understanding of root-tip cell-type specific proteomic responses to abiotic stresses, such as drought, particularly under field conditions, remains limited. This study aimed to identify spatially resolved, cell type-specific proteomes in switchgrass (Panicum virgatum) root tips under drought stress. Root tips were collected from seven-year-old, field-grown switchgrass ‘Alamo’ plants excavated under both well-watered and long-term drought conditions. Cell type-specific proteins were identified using laser capture microdissection (LCM) coupled with nanoPOTS (Nanodroplet Processing in One Pot for Trace Samples) and nano-LC-MS proteomics analysis. Five distinct cell types were targeted: (1) cells in the quiescent center and stem cell niche (QuC), (2) protodermal epidermal cells (PEC) in the meristematic zone, (3) epidermal cells in the transition and elongation zones above the root cap (Epi), (4) peripheral root cap cells (PRC), forming 2–3 layers below the PEC and 1–2 layers above the root border cells, and (5) columella root cap cells (Col) comprising of the columella initials and a single underlying layer of cells undergoing active growth. Principal component analysis (PCA) revealed clear separation among the five targeted cell types, confirming distinct proteomic profiles. Proteins predominantly enriched in each cell type were linked to distinct cellular functions, with QuC cells showing involvement in chromosomal behavior, DNA replication, and mitosis—key processes for stem cell niche regulation. Drought stress resulted in alterations of proteostasis, as evidenced by significant decreases in ribosomal proteins and increases in protein synthesis inhibitors. Moreover, drought stress induced unique cell-type–specific proteins involved in phytohormone biosynthesis and signaling pathways, including auxin, cytokinin, and jasmonic acid. In particular, QuC cells were more highly enriched in proteins associated with DNA repair and mitotic processes. Metabolic pathways related to amino acids, carbohydrates, and lipids were differentially affected in a cell-type–dependent manner, whereas general stress-responsive proteins exhibited consistent changes across all five cell types. Overall, this study provides unique spatially resolved, cell-type-specific proteomic profiles in root tips, representing a significant advancement in our understanding of the cellular mechanisms underlying plant responses to drought stress in natural field conditions.

perennial grass↗

Anomaly Detection for the Roman Space Telescope Wide Field Instrument’s Science Data Processing Pipeline

The Roman Space Telescope (RST) Wide Field Instrument (WFI) will be utilizing a preliminary Science Data Processing (SDP) pipeline during its Integration and Test, and to some extent during Operations, to track basic statistics and identify known features such as cosmic rays, snowballs as well as possible anomalies in raw detector data. In our detectors, these anomalies appear as jumps in the ramp of a readout and are classified as cosmic rays if they appear as a streak or snowballs if they’re more circular. The WFI employs an array of 18 H4RG-10 detectors that collect image samples. Each set of raw frames within a non-destructive exposure is packaged by the SDP pipeline into image cubes for each detector. Each cube is a time series of 4096 × 4096 accumulating pixel frames. The preliminary analysis pipeline is used to locate anomalies in these time-series accumulation frames and identify the type of anomaly, either natural phenomena or detector characteristic. To compare different methods, we’ve implemented both heuristic-based and data-driven methods to identify anomalies. For the heuristic-based approach, we identify snowballs and cosmic rays by the size and shape of outlier pixel clusters between consecutive frames. For data driven methods, we evaluated a Convolutional Neural Network (CNN) model, and more traditional methods like Principal Component Analysis (PCA). CNN is a supervised learning/classification method. Thus, we used a labeled dataset of anomalies to perform segmentation of the image and identify anomalies. We used previously identified cosmic rays and snowballs to measure the accuracy and efficiency of the mentioned approaches. In evaluating these methods, we aim to pick the best fit for the SDP pipeline’s anomaly detection in terms of both performance and runtime.

Paul Horton↗

Analysis of Superconducting Magnet Quench Antenna Data

Quenching poses a serious problem for superconducting magnets operating at high currents. It occurs when the material transitions from the superconducting to the normal state, which leads to heating and potential damage to the magnet. To understand and mitigate quenching, the Magnet Department at Fermilab is developing and testing superconducting magnet quench antenna arrays. This study delves into the anomalous events preceding the quench during magnet training by analyzing the collected data. With the moving average and Fast Fourier Transform techniques, we investigate the trends and frequency patterns of the data. Moreover, we introduce an unsupervised anomaly detection algorithm based on Principal Component Analysis and DBSCAN clustering. It can autonomously identify events within background noise, without relying on any predefined event features. Our analysis reveals that the spatio-temporal distribution of these anomalous events has little connection to the quench location, indicating that a majority of them bear no relation to the quenching process.

43 PARTICLE ACCELERATORS↗

Homomorphic Encryption for Machine Learning and Artificial Intelligence Applications

Third-party and expert analysis is a cost-effective solution for solving specialized problems or processing large datasets related to reactor structural health monitoring and nondestructive evaluation. However, when handling proprietary information, third-party and expert analysts pose a privacy risk. To address this challenge, Homomorphic Encryption (HE) permits arithmetic operations on encrypted data without exposing the underlying data. Implementations of Machine Learning (ML) and Artificial Intelligence (AI) algorithms using HE greatly enhances the capabilities of third-party analysts while maintaining a low security risk. This paper details current success in applying Principal Component Analysis (PCA) and Fully Connected Neural Networks (NN) using the Microsoft SEAL implementation of the popular CKKS Fully Homomorphic Encryption (FHE) algorithm. The MNIST Handwritten Dataset is analyzed as a proof-of-concept demonstration of the implementations.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Modeled production, oxidation, and transport processes of wetland methane emissions in temperate, boreal, and Arctic regions

Abstract Wetlands are the largest natural source of methane (CH 4 ) to the atmosphere. The eddy covariance method provides robust measurements of net ecosystem exchange of CH 4 , but interpreting its spatiotemporal variations is challenging due to the co‐occurrence of CH 4 production, oxidation, and transport dynamics. Here, we estimate these three processes using a data‐model fusion approach across 25 wetlands in temperate, boreal, and Arctic regions. Our data‐constrained model—iPEACE—reasonably reproduced CH 4 emissions at 19 of the 25 sites with normalized root mean square error of 0.59, correlation coefficient of 0.82, and normalized standard deviation of 0.87. Among the three processes, CH 4 production appeared to be the most important process, followed by oxidation in explaining inter‐site variations in CH 4 emissions. Based on a sensitivity analysis, CH 4 emissions were generally more sensitive to decreased water table than to increased gross primary productivity or soil temperature. For periods with leaf area index (LAI) of ≥20% of its annual peak, plant‐mediated transport appeared to be the major pathway for CH 4 transport. Contributions from ebullition and diffusion were relatively high during low LAI (<20%) periods. The lag time between CH 4 production and CH 4 emissions tended to be short in fen sites (3 ± 2 days) and long in bog sites (13 ± 10 days). Based on a principal component analysis, we found that parameters for CH 4 production, plant‐mediated transport, and diffusion through water explained 77% of the variance in the parameters across the 19 sites, highlighting the importance of these parameters for predicting wetland CH 4 emissions across biomes. These processes and associated parameters for CH 4 emissions among and within the wetlands provide useful insights for interpreting observed net CH 4 fluxes, estimating sensitivities to biophysical variables, and modeling global CH 4 fluxes.

Ueyama, Masahito↗

Statistical analysis of astronomical data containing upper bounds - General methods and examples drawn from X-ray astronomy

Statistical procedures taken from the field of survival analysis have been adapted to astronomical usage and have been applied to a sample of stars in the B-V color range between 0.1 and 0.5 with measured soft X-ray luminosities and projected equatorial velocities. The two-sample problem and linear regression problem with arbitrarily censored data were studied. A new method for determining the linear regression between two random variables in the presence of arbitrary censoring has been developed which can also be used for a likelihood-ratio test for the independence of two random variables and for principal-component analysis in the presence of arbitrary censoring. The required numerical computations can be carried out straightforwardly and rapidly.

Schmitt, J. H. M. M.↗

Comparing gas composition from fast pyrolysis of live foliage measured in bench-scale and fire-scale experiments

Background: Fire models have used pyrolysis data from oxidising and non-oxidising environments for flaming combustion. In wildland fires pyrolysis, flaming and smouldering combustion typically occur in an oxidising environment (the atmosphere). Aims: Using compositional data analysis methods, determine if the composition of pyrolysis gases measured in non-oxidising and ambient (oxidising) atmospheric conditions were similar. Methods: Permanent gases and tars were measured in a fuel-rich (non-oxidising) environment in a flat flame burner (FFB). Permanent and light hydrocarbon gases were measured for the same fuels heated by a fire flame in ambient atmospheric conditions (oxidising environment). Log-ratio balances of the measured gases common to both environments (CO, CO 2 , CH 4 , H 2 , C 6 H 6 O (phenol), and other gases) were examined by principal components analysis (PCA), canonical discriminant analysis (CDA) and permutational multivariate analysis of variance (PERMANOVA). Key results: Mean composition changed between the non-oxidising and ambient atmosphere samples. PCA showed that flat flame burner (FFB) samples were tightly clustered and distinct from the ambient atmosphere samples. CDA found that the difference between environments was defined by the CO-CO 2 log-ratio balance. PERMANOVA and pairwise comparisons found FFB samples differed from the ambient atmosphere samples which did not differ from each other. Conclusion: Relative composition of these pyrolysis gases differed between the oxidising and non-oxidising environments. This comparison was one of the first comparisons made between bench-scale and field scale pyrolysis measurements using compositional data analysis. Implications: These results indicate the need for more fundamental research on the early time-dependent pyrolysis of vegetation in the presence of oxygen.

54 ENVIRONMENTAL SCIENCES↗

Visualization of Global Sensitivity Analysis Results Based on a Combination of Linearly Dependent and Independent Directions

A useful technique for the validation and verification of complex flight systems is Monte Carlo Filtering -- a global sensitivity analysis that tries to find the inputs and ranges that are most likely to lead to a subset of the outputs. A thorough exploration of the parameter space for complex integrated systems may require thousands of experiments and hundreds of controlled and measured variables. Tools for analyzing this space often have limitations caused by the numerical problems associated with high dimensionality and caused by the assumption of independence of all of the dimensions. To combat both of these limitations, we propose a technique that uses a combination of the original variables with the derived variables obtained during a principal component analysis.

Davies, Misty D.↗

IMERG and GPCP Seasonality and Response to Climate and Weather Variability

The Integrated Multi-satellitE Retrievals for GPM (IMERG) and the The Global Precipitation Climatology Project (GPCP) are two of the most popular precipitation products. IMERG is a relatively new dataset that targets the needs primarily of the hydrological community by resolving the hydrological cycle of precipitation at fine temporal (30-minutes) and spatial (10-km) scales. IMERG only recently exceeded the user base of the highly successful, but discontinued in 2019, TRMM Multi-satellite Precipitation Analysis (TMPA). GPCP, on the other hand, has traditionally been strong in the climate research community, and recently has been revised under the framework of NASA's Making Earth System Data records for Use in Research Environments (MEaSUREs) program. Both IMERG and GPCP are similar in the underlying approaches to achieve global coverage, in particular using satellite microwave and infrared observations, and adjusting the precipitation retrieval with rain gauge information. While there is a tendency to use both datasets interchangeably, differences remain and some of them limit the usage of IMERG as a climate data record at this point. By applying Principal Component Analysis, we identify the differences and similarities mode-by-mode, and further the guidance on suggested usage of IMERG as a climate data record. As an example in the attached figure, IMERG vs GPSP differences in the explained variance by the two leading seasonal modes can be identified, mainly in the extreme southern latitudes, and over western boundary currents (Gulfstream and Kuroshio). At the preparation time for this presentation, the new version "07"of IMERG was in the works that may resolve the issues presented here. Nevertheless, our analysis can help to gauge the uncertainties of the studies already done using the currently existing IMERG version "06", and evaluate the improvements in the upcoming version "07".

Andrey Savtchenko↗

TOFHunter—unlocking rapid untargeted screening of inductively coupled plasma–time-of-flight–mass spectrometry data

This study provides an overview of a newly developed open source program written in Python, TOFHunter, which permits the rapid and untargeted screening of inductively coupled plasma (ICP)-time-of-flight (TOF)-mass spectrometry (MS) datasets. ICP-TOF-MS is an analytical tool capable of providing quasi simultaneous detection of all nuclides from Li to Pu. This capability has triggered an increase in studies investigating single-particle analysis in which the TOF-MS provides correlated elemental/isotopic signatures on a particle basis in time. Similarly, laser ablation mapping has seen rapid growth owing to ICP-TOF-MS's capacity to handle fast washout times (<10 ms) while providing a broad nuclide coverage. The caveat to this broad mass coverage and high time resolution comes in the form of large, overwhelming datasets. With datasets typically on the scale of gigabytes, it is easy for a user to only focus on very targeted analytes; however, this focus diminishes the opportunity offered by the TOF-MS detector. TOFHunter applies chemometric methods, principal component analysis (PCA), and interesting features finder (IFF) on ICP-TOF-MS data, allowing for investigation of correlations, major and minor variance sources, and sample screening. The unique spectra identified by the (IFF) are used to generate a list of mass peaks, which are then matched with both nuclides and potential interferences before being exported for the user to investigate. Several case studies are discussed herein, demonstrating TOFHunter's ability to screen aqueous injections, single-particle/single-cell analysis, and probe laser ablation mapping files for unique regions of interest.

47 OTHER INSTRUMENTATION↗

Incorporating Endmember Variability into Spectral Mixture Analysis Through Endmember Bundles

Variation in canopy structure and biochemistry induces a concomitant variation in the top-of-canopy spectral reflectance of a vegetation type. Hence, the use of a single endmember spectrum to track the fractional abundance of a given vegetation cover in a hyperspectral image may result in fractions with considerable error. One solution to the problem of endmember variability is to increase the number of endmembers used in a spectral mixture analysis of the image. For example, there could be several tree endmembers in the analysis because of differences in leaf area index (LAI) and multiple scatterings between leaves and stems. However, it is often difficult in terms of computer or human interaction time to select more than six or seven endmembers and any non-removable noise, as well as the number of uncorrelated bands in the image, limits the number of endmembers that can be discriminated. Moreover, as endmembers proliferate, their interpretation becomes increasingly difficult and often applications simply need the aerial fractions of a few land cover components which comprise most of the scene. In order to incorporate endmember variability into spectral mixture analysis, we propose representing a landscape component type not with one endmember spectrum but with a set or bundle of spectra, each of which is feasible as the spectrum of an instance of the component (e.g., in the case of a tree component, each spectrum could reasonably be the spectral reflectance of a tree canopy). These endmember bundles can be used with nonlinear optimization algorithms to find upper and lower bounds on endmember fractions. This approach to endmember variability naturally evolved from previous work in deriving endmembers from the data itself by fitting a triangle, tetrahedron or, more generally, a simplex to the data cloud reduced in dimension by a principal component analysis. Conceptually, endmember variability could make it difficult to find a simplex that both surrounds the data cloud and has vertices that are realistic endmember spectra with reflectances between 0 and 1. In this paper, we create endmember bundles and bounding fraction images for an AVIRIS subscene simulated with a plant canopy radiative transfer model. The simulated subscene is spatially patterned after a subscene from the AVIRIS image acquired August, 1993 over La Copita, Texas. In addition, for comparison, we performed a traditional unmixing with image endmembers.

Bateson, C. Ann↗

An Initial Analysis of LANDSAT-4 Thematic Mapper Data for the Discrimination of Agricultural, Forested Wetlands, and Urban Land Cover

The capabilities of TM data for discriminating land covers within three particular cultural and ecological realms was assessed. The agricultural investigation in Poinsett County, Arkansas illustrates that TM data can successfully be used to discriminate a variety of crop cover types within the study area. The single-date TM classification produced results that were significantly better than those developed from multitemporal MSS data. For the Reelfoot Lake area of Tennessee TM data, processed using unsupervised signature development techniques, produced a detailed classification of forested wetlands with excellent accuracy. Even in a small city of approximately 15,000 people (Union City, Tennessee). TM data can successfully be used to spectrally distinguish specific urban classes. Furthermore, the principal components analysis evaluation of the data shows that through photointerpretation, it is possible to distinguish individual buildings and roof responses with the TM.

Quattrochi, D. A.↗

Precise relative magnitude measurement improves fracture characterization during hydraulic fracturing

SUMMARY Microseismic monitoring is an important technique to obtain detailed knowledge of in-situ fracture size and orientation during stimulation to maximize fluid flow throughout the rock volume and optimize production. Furthermore, considering that the frequency of earthquake magnitudes empirically follows a power law (i.e. Gutenberg–Richter), the accuracy of microseismic event magnitude distributions is potentially crucial for seismic risk management. In this study, we analyse microseismicity observed during four hydraulic fracture treatments of the legacy Cotton Valley experiment in 1997 at the Carthage gas field of East Texas, where fractures were activated at the base of the sand-shale Upper Cotton Valley formation. We perform waveform cross-correlation to detect similar event clusters, measure relative amplitude from aligned waveform pairs with a principal component analysis, then measure precise relative magnitudes. The new magnitudes significantly reduce the deviations between magnitude differences and relative amplitudes of event pairs. This subsequently reduces the magnitude differences between clusters located at different depths. Reduction in magnitude differences between clusters suggests that some attenuation-related biases could be effectively mitigated with relative magnitude measurements. The maximum likelihood method is applied to understand the magnitude frequency distributions and quantify the seismogenic index of the clusters. Statistical analyses with new magnitudes suggest that fractures that are more favourably oriented for shear failure have lower b-value and higher seismogenic index, suggesting higher potential for relatively larger earthquakes, rather than fractures subparallel to maximum horizontal principal stress orientation.

58 GEOSCIENCES↗