Engineering PapersSearch

SEARCH · Engineering Papers

Results for “cluster analysis”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Clustering analysis of medium-band selected high-redshift galaxies

Next-generation large-scale structure spectroscopic surveys will probe cosmology at high redshifts (2.3 < z < 3.5), relying on abundant galaxy tracers such as Lyα emitters (LAEs) and Lyman break galaxies (LBGs). Medium-band photometry has emerged as a potential technique for efficiently selecting these high-redshift galaxies. In this work, we present clustering analysis of medium-band selected galaxies at high redshift, utilizing photometric data from the Intermediate Band Imaging Survey (IBIS) and spectroscopic data from the Dark Energy Spectroscopic Instrument (DESI). We interpret the clustering of such samples using both Halo Occupation Distribution (HOD) modeling and a perturbation theory description of large-scale structure. Our modeling indicates that the current target sample is composed from an overlapping mixture of LAEs and LBGs with emission lines. Despite differences in target selection, we find that the clustering properties are consistent with previous studies, with correlation lengths r 0 ≃ 3-4 h -1 Mpc and a linear bias of b ∼ 1.8-2.5. Finally, we discuss the simulation requirements implied by these measurements and demonstrate that the properties of the samples would make them excellent targets to enhance our understanding of the high-z universe.

79 ASTRONOMY AND ASTROPHYSICS

Exploring HOD-dependent systematics for the DESI 2024 Full-Shape galaxy clustering analysis

We analyze the robustness of the DESI 2024 cosmological inference from the full shape of the galaxy power spectrum to uncertainties in the Halo Occupation Distribution (HOD) model of the galaxy-halo connection and the choice of priors on nuisance parameters. We assess variations in the recovered cosmological parameters across a range of mocks populated with different HOD models and find that shifts are often greater than 20% of the expected statistical uncertainties from the DESI data. We encapsulate the effect of such shifts in terms of a systematic covariance term, C HOD , and an additional diagonal contribution quantifying the impact of our choice of nuisance parameter priors on the ability of the effective field theory (EFT) model to correctly recover the cosmological parameters of the simulations. These two covariance contributions are designed to be added to the usual covariance term, C stat , describing the statistical uncertainty in the power spectrum measurement, in order to fairly represent these sources of systematic uncertainty. This novel approach should be more general and robust to the choice of model or additional external datasets used in cosmological fits than the alternative approach of adding systematic uncertainties to the recovered marginalised parameter posteriors. We compare the approaches within the context of a fixed ΛCDM model and demonstrate that our method gives conservative estimates of the systematic uncertainty that nevertheless have little impact on the final posteriors obtained from DESI data.

79 ASTRONOMY AND ASTROPHYSICS

Machine Learning–Augmented Laser-Induced Breakdown Spectroscopy for Spectral Discrimination of Iron Oxalates

Enhanced characterization and phase identification of post-PUREX Pu Oxalates (PuOXA) are pivotal for nonproliferation and pre-detonation nuclear forensics. Despite significant advances in the characterization of PuO 2 samples, little is known about the impact of both the chemical structure and oxidation states of PuOXA (i.e., Pu(III) and Pu(IV)) have on optical emission signatures. Here, we demonstrate the analytical capabilities of laser-induced breakdown spectroscopy (LIBS) applied to Fe(II) and Fe(III) oxalate samples as surrogates for PuOXA, highlighting the discriminating features in the LIBS emission spectra arising from differences in the oxidation states within mixed FeOXA samples. We report the enhancement of spectral feature selection using Principal Component Analysis (PCA), which enables the analytical superiority of machine learning algorithms such as Linear Discriminant Analysis (LDA), Quadratic Discriminant Analysis (QDA), Partial Least Squares Regression (PLSR), Support Vector Regression (SVR), and Random Forest Regression (RFR) over conventional univariate techniques for phase discrimination and chemometric analysis. Cluster analysis revealed how both matrix effects and laser ablation influence cluster separability by introducing spectral artifacts that misdirect the maximization of variance. PCA-selected emission lines were used in the regression models, demonstrating that both univariate and multivariate linear regression models (i.e., PLSR and SVR) can achieve acceptable performance, with machine learning models outperforming conventional calibration regressions. Furthermore, the application of non-linearly activated PCA-selected emission lines illustrates how simplifying the data while retaining captured variance enables the use of less complex and more computationally efficient models. Furthermore, this is particularly evident in the underperformance of RFR, which suffers from increased computational costs and overfitting owing to its high complexity.

Oxalates

An evaluation of air quality in major urban areas of India

Rapid economic growth and burgeoning population have contributed to enhanced levels of PM 2.5 concentrations in urban regions of India. Evaluation of ambient air quality facilitates the assessment of effectiveness of emission control measures and early identification of new sources. This study provides a comprehensive statistical analysis of PM 2.5 concentrations in key urban areas across India, including Delhi, Kolkata, Mumbai, Chennai, Hyderabad, and several regional centers. Data from 2017 to 2023 was analyzed using trend analysis, cluster analysis, principal component analysis, and geostatistical interpolation to understand spatiotemporal variations and sources. The analysis reveals significant differences in spatial distribution of PM 2.5 concentrations with high annual averages in urban regions in Indo-Gangetic plain (82–123 μg m −3 ) and relatively lower concentrations (29–46 μg m −3 ) in southern urban areas of Kerala, Tamil Nadu and Andhra Pradesh. Delhi state had the highest 24-averaged PM 2.5 concentrations (112 μg m −3 ) followed by urban regions in Uttar Pradesh, Bihar and West Bengal (94 μg m −3 ). Trend analysis from 2017 to 2023 revealed an overall 2.5% decline in site-wide PM2.5 concentrations, with the exception of Ludhiana, which exhibited a consistent annual increase of 10%. Principal component analysis (PCA) attributes 30% of the variance to wintertime emissions, 13% to biomass burning, and 18% to the regional haze in the northern Indo-Gangetic Plain. Different analyses clearly demonstrates the contribution of biomass burning to pollution in Delhi and surrounding cities. Transboundary pollution to Kolkata is likely from the highly polluted region in Indo-Gangetic Plain. Coastal cities of Mumbai and Chennai has relatively lower pollution attributed to the influence of sea breeze dilution, with mostly local contribution and some potential transport from upwind industry clusters. Hyderabad also has local contribution due to high density of vehicular traffic and local small industries. This study shows that mitigation efforts targeting clusters of regions should be undertaken to curb the high PM2.5 pollution. Policy measures should be implemented both at local and the intra-state level to address shared sources and transport of pollution.

Hysplitbacktrajectories

Single-nuclei transcriptome analysis of IgM+ cells isolated from channel catfish (Ictalurus punctatus) spleen

Catfish production is the primary aquaculture sector in the United States, and the key cultured species is channel catfish (Ictalurus punctatus). The major causes of production losses are pathogenic diseases, and the spleen, an important site of adaptive immunity, is implicated in these diseases. To examine the channel catfish immune system, single-nuclei transcriptomes of sorted and captured IgM + cells were produced from adult channel catfish. Three channel catfish (~1 kg) were euthanized, the spleen dissected, and the tissue dissociated. The lymphocytes were isolated using a Ficoll gradient and IgM + cells were then sorted with flow cytometry. The IgM + cells were lysed and single-nuclei libraries generated using a Chromium Next GEM Single Cell 3’ GEM Kit and the Chromium X Instrument (10x Genomics) and sequenced with the Illumina NovaSeq X Plus sequencer. The reads were aligned to theI. punctatusreference assembly (Coco_2.0) using Cell Ranger, and normalization, cluster analysis, and differential gene expression analysis were carried out with Seurat. Across the three samples, approximately 753.5 million reads were generated for 18,686 cells. After filtering, 10,637 cells remained for the cluster analysis. The cluster analysis identified 16 clusters which were classified as B cells (10,276), natural killer-like (NK-like) cells (178), T cells or natural killer cells (45), hematopoietic stem and progenitor cells (HSPC)/megakaryocytes (MK) (66), myeloid/epithelial cells (40), and plasma cells (32). The B cell clusters were further defined as different populations of mature B cells, cycling B cells, and plasma cells. The plasma cells highly expressedighmand we demonstrated that the secreted form of the transcript was largely being expressed by these cells. This atlas provides insight into the gene expression of IgM + immune cells in channel catfish. The atlas is publicly available and could be used garner more important information regarding the gene expression of splenic immune cells.

Immunology

Dark Energy Survey Year 6 Results: Weak Lensing and Galaxy Clustering Cosmological Analysis Framework

We present the methodology for the weak lensing and galaxy clustering analyses of the Dark Energy Survey (DES) Year 6 data set. In this work, we design and validate the analysis pipeline for the cosmic shear, galaxy clustering plus galaxy$-$galaxy lensing ($2 \times 2$pt), and the joint analysis in the $3 \times 2$pt. Our framework accounts for key theoretical uncertainties, such as baryonic feedback and galaxy bias, incorporating both linear and non-linear models. We apply scale cuts in regimes where theoretical modeling becomes unreliable. The robustness of the pipeline is validated using mock data and simulations, confirming unbiased cosmological constraints and highlighting the importance of posterior projection effects in the validation process. As a result, we deliver robust and validated analysis pipelines for cosmic shear, $2 \times 2$pt, and $3 \times 2$pt in $Λ$CDM and $w$CDM scenarios, including a well-defined set of scales suitable for real data analysis, a robust prescription for theoretical systematics, and the theoretical covariance of the signal. This comprehensive methodology also lays the groundwork for future galaxy surveys such as the Vera C. Rubin Observatory Legacy Survey of Space and Time.

Sanchez-Cid, D. [Zurich U.; Madrid, CIEMAT; Madrid

Spatio-temporal multivariate cluster evolution analysis for detecting and tracking climate impacts

Recent years have seen a growing concern about climate change and its impacts. While Earth System Models (ESMs) can be invaluable tools for studying the impacts of climate change, the complex coupling processes encoded in ESMs and the large amounts of data produced by these models, together with the high internal variability of the Earth system, can obscure important source-to-impact relationships. Here, this paper presents a novel and efficient unsupervised data-driven approach for detecting statistically-significant impacts and tracing spatio-temporal source-impact pathways in the climate through a unique combination of ideas from anomaly detection, clustering and Natural Language Processing (NLP). Using as an exemplar the 1991 eruption of Mount Pinatubo in the Philippines, we demonstrate that the proposed approach is capable of detecting known post-eruption impacts/events. We additionally describe a methodology for extracting meaningful sequences of post-eruption impacts/events by using NLP to efficiently mine frequent multivariate cluster evolutions, which can be used to confirm or discover the chain of physical processes between a climate source and its impact(s).

Anomaly detection

Detecting Living-off-the-land Attacks Using K-means And Graph Convolutional Networks

The code ingests Zeek logs derived from network packet captures and goes through data preprocessing before it gets passed into a K-Means model that labels each device as either a client or server. Graph Convolutional Network (GCN) model is used to obtain the embeddings to represent the features in lower dimension. Last, K-means cluster analysis is used to cluster the embeddings for each class.

Quach, Anna [Idaho National Laboratory (INL), Idah

Oleaginous Yeast Biology Elucidated With Comparative Transcriptomics

ABSTRACT Extremophilic yeasts have favorable metabolic and tolerance traits for biomanufacturing‐ like lipid biosynthesis, flavinogenesis, and halotolerance – yet the connection between these favorable phenotypes and strain genotype is not well understood. To this end, this study compares the phenotypes and gene expression patterns of biotechnologically relevant yeasts Yarrowia lipolytica , Debaryomyces hansenii , and Debaryomyces subglobosus grown under nitrogen starvation, iron starvation, and salt stress. To analyze the large data set across species and conditions, two approaches were used: a “network‐first” approach where a generalized metabolic network serves as a scaffold for mapping genes and a “cluster‐first” approach where unsupervised machine learning co‐expression analysis clusters genes. Both approaches provide insight into strain behavior. The network‐first approach corroborates that Yarrowia upregulates lipid biosynthesis during nitrogen starvation and provides new evidence that riboflavin overproduction in Debaryomyces yeasts is overflow metabolism that is routed to flavin cofactor production under salt stress. The cluster‐first approach does not rely on annotation; therefore, the coexpression analysis can identify known and novel genes involved in stress responses, mainly transcription factors and transporters. Therefore, this work links the genotype to the phenotype of biotechnologically relevant yeasts and demonstrates the utility of complementary computational approaches to gain insight from transcriptomics data across species and conditions.

Weintraub, Sarah J. [Department of Bioinformatics

Enhancing Cluster Identification in Atom Probe Tomography Data Using Transfer Learning

Atom Probe Tomography (APT) is a powerful technique for visualizing the atomic-scale distribution of solutes in materials, but quantitative cluster analysis of APT datasets remains a challenge due to the need for subjective parameter selection in clustering algorithms. While distance-based and density-based methods such as HDBSCAN are widely used, their performance is highly sensitive to user-defined parameters, which undermines reproducibility and accuracy. This study proposes an image-based, deep learning-aided workflow for automating parameter selection and cluster detection in APT data analysis. By projecting 3D APT point clouds onto 2D planes, we leverage pretrained convolutional neural networks (ConvNeXt-Tiny and ResNet-50) through transfer learning to predict the number of clusters present in synthetic datasets. The output is used to guide K-means clustering and estimate HDBSCAN parameters, specifically minimum cluster size and minimum sample points. This approach reduces reliance on manual parameter tuning, improving consistency and scalability. The methodology demonstrates the feasibility of using image-based deep learning for interpreting complex spatial patterns in APT data, enabling faster and more objective analysis. The complete workflow and code are made publicly available to support reproducibility and future research.

Density-based clustering

The 3D clustering of Lyman Alpha Emitters measured with DESI

We present a clustering analysis of Lyman-$α$ emitters (LAEs) using spectroscopic observations from the Dark Energy Spectroscopic Instrument (DESI) of candidates selected from the Blanco/DECam Intermediate-Band Imaging Survey (IBIS). We measure the two-point correlation function and the power spectrum, including cross-correlations with DESI quasars. Using both analytical and halo occupation distribution (HOD) simulation-based modeling, we find a linear bias of $b \sim 2.31$--$2.62$ for LAEs over the redshift range $2.26 < z < 3.41$. The analytical modeling also provides constraints on the strength of radiative transfer effects, while the HOD analysis characterizes the LAE-halo connection across multiple models. Finally, we quantify the magnitude of non-perturbative clustering effects such as Fingers of God in the LAE population, providing essential input for the accurate modeling of LAE-based cosmological analyses in forthcoming high-redshift surveys such as DESI-II.

Ebina, H. [UC, Berkeley; LBL, Berkeley] (ORCID:000

A generalizable machine learning-assisted fast Fourier transform algorithm to simulate the large strain phenomena in polycrystalline materials

Machine learning methods have shown initial promise in constitutive modeling for single crystals or homogenized polycrystals, delivering notable computational efficiency. However, existing machine learning-based constitutive models often lack generalizability, limiting their application across diverse boundary value problems. This study introduces a thermodynamics-informed artificial neural network model to accelerate rate-tangent crystal plasticity fast Fourier transform simulations for cross-scale deformation behaviors of polycrystals under complex loading. Our model integrates microstructural variability and local interactions effectively. To address local effects in each grain, we employ K-means clustering to group Gauss points within the microstructure into clusters assumed to be in similar mechanical states. This approach, based on self-clustering analysis, extends model scope from macroscopic stress response to the granular level, capturing mechanical responses and orientation evolution across grains. This reduces the number of nonlinear problems to solve, with cluster responses propagated throughout each group. The thermodynamics-based artificial neural network-extracted features are further processed using local material state clusters to account for history-dependent deformation and evolving microstructures. Additionally, representative volume element simulations with rate-tangent crystal plasticity fast Fourier transform provide reliable datasets for model training. The proposed model demonstrates high efficiency, accuracy, self-consistency, and enhanced generalizability in predicting strain–stress responses and orientation evolution at both individual grain and aggregate scales under complex loading conditions, such as biaxial tension and arbitrary loading scenarios.

36 MATERIALS SCIENCE

Feasibility Analyses for the Microgrid of the Mountain in the Cordillera Central Region of Puerto Rico

Distributed energy resource (DER) development can benefit communities through improvement of access to reliable electricity and a resilient energy system. Many place-specific considerations must be accounted for when designing electrical systems with co-located generation and loads. This paper presents an overview of four studies focusing on different aspects for regional DER development in the Cordillera Central region of Puerto Rico, including a grid and load stability study with the introduction of solar energy generation, microgrid design based on the solar resource in the area, resilience considerations during power outage scenarios, and qualitative risk aspects for components in the microgrid. Additionally, a conceptual substation microgrid analysis is conducted using the Microgrid Design Toolkit (MDT) to evaluate trade-offs between cost and energy availability, providing insights into optimal design strategies. A microgrid resilience assessment is performed using the Resilient Nodal Cluster Analysis Tool (ReNCAT) to identify critical loads and evaluates their accessibility during outages, emphasizing the importance of community needs. Finally, a qualitative risk assessment utilizing Hazard and Operability Analysis (HAZOP) methods highlights potential hazards and mitigations for microgrid components, ensuring robust system design and operation.

24 POWER TRANSMISSION AND DISTRIBUTION

Enhancing Cluster Identification in Atom Probe Tomography Data Using Transfer Learning

Atom probe tomography (APT) has enabled the direct visualization of solute clusters, providing valuable insights into material structures. This clustering is crucial for understanding the nanoscale composition and behavior of materials, which can significantly influence their mechanical and physical properties. However, the widely used clustering methods in the APT community face challenges such as subjective parametric selection and limited applicability, particularly in dealing with overlapping clusters, nested clusters, and artifacts across different scales, such as precipitates and dislocations. To address these challenges, we present a framework based on density-based cluster analysis that aims to be less dependent on user input, reproducible, and robust.

Density-based clustering

The Cocytos Stream: A Disrupted Globular Cluster from our Last Major Merger?

The census of stellar streams and dwarf galaxies in the Milky Way provides direct constraints on galaxy formation models and the nature of dark matter. The DESI Milky Way survey (with a footprint of 14,000$~deg{^2}$ and a depth of $r<19$ mag) delivers the largest sample of distant metal-poor stars compared to previous optical fiber-fed spectroscopic surveys. This makes DESI an ideal survey to search for previously undetected streams and dwarf galaxies. We present a detailed characterization of the Cocytos stream, which was re-discovered using a clustering analysis with a catalog of giants in the DESI year 3 data, supplemented with Magellan/MagE spectroscopy. Our analysis reveals a relatively metal-rich ([Fe/H]$=-1.3$) and thick stream (width$=1.5^\circ$) at a heliocentric distance of $\approx 25$ kpc, with an internal velocity dispersion of 6.5-9 km s$^{-1}$. The stream's metallicity, radial orbit, and proximity to the Virgo stellar overdensities suggest that it is most likely a disrupted globular cluster that came in with the Gaia-Enceladus merger. We also confirm its association with the Pyxis globular cluster. Our result showcases the ability of wide-field spectroscopic surveys to kinematically discover faint disrupted dwarfs and clusters, enabling constraints on the dark matter distribution in the Milky Way.

79 ASTRONOMY AND ASTROPHYSICS

Emerging Per- and Polyfluoroalkyl Substances in Tap Water from the American Healthy Homes Survey II

Humans experience widespread exposure to anthropogenic per- and polyfluoroalkyl substances (PFAS) through various media, which can lead to a wide range of negative health impacts. Tap water is an important source of exposure in communities with any degree of contamination but routine or large-scale PFAS monitoring often depends on targeted analytical methods limited to measuring specific PFAS. We analyzed 680 tap water samples from the American Healthy Homes Survey II for PFAS using non-targeted analysis (NTA) to expand the range of detectable PFAS. Based on detection frequency and relative abundance, about half of the identified PFAS were found only by NTA. We identified (with varying degrees of confidence) 75 distinct PFAS, including 57 exclusively detected by NTA. The identified PFAS are members of seven structural subclasses differentiated by their head groups and degree of fluorination. Clustering analysis categorized the PFAS into four coabundance groups dominated by specific PFAS subclasses. One group uniquely identified by NTA contains zwitterionic PFAS and other PFAS transformation products which are likely associated with aqueous firefighting foam contaminants in a small number of spatially correlated samples. These results help further characterize the scope of exposure to emerging PFAS experienced by the U.S. population via tap water and augment nationwide targeted-PFAS monitoring programs.

63 RADIATION, THERMAL, AND OTHER ENVIRON. POLLUTAN

Assessing Dynamic Time Warping Techniques for Discriminating Seismic Sources at Local and Regional Distances

Effective monitoring of seismic explosions and hazard assessment relies heavily on the accurate discrimination of underground seismic sources. This study investigates the application of novel nonlinear alignment techniques, specifically Dynamic Time Warping (DTW), for event-type discrimination at regional and local distances. Building on prior research that used DTW and Elastic Shape Analysis (ESA) in discrimination at regional distances, we evaluate the performance of recently developed variants of DTW, including a method that employs Pearson cross-correlation as a measure of warping distance and a time distortion coefficient that quantifies the type and degree of time distortion between signals. By analyzing observational datasets that include different source types, we assess the performance of these approaches for realistic monitoring scenarios. Specifically, we consider a dataset recorded at regional distances in the Korean Peninsula and a local-distance subset from the Unconstrained Utah Event Bulletin catalog to evaluate DTW-based discrimination across multiple distance scales. Additionally, we introduce the maximum cross-correlations of warped waveforms as a similarity metric for event classification. Through hierarchical cluster analysis and dendrogram interpretation, we present our findings, highlighting the strengths and limitations of these techniques in seismic event classification.

45 MILITARY TECHNOLOGY, WEAPONRY, AND NATIONAL DEF

Exploring the exact limits of the real-time equation-of-motion coupled cluster cumulant Green’s functions

In this paper, we analyze the properties of the recently proposed real-time equation-of-motion coupled-cluster (RT-EOM-CC) cumulant Green’s function approach [Rehr et al., J. Chem. Phys. 152, 174113 (2020)]. We specifically focus on identifying the limitations of the original time-dependent coupled cluster (TDCC) ansatz and propose an enhanced double TDCC ansatz, ensuring the exactness in the expansion limit. In addition, we introduce a practical cluster-analysis-based approach for characterizing the peaks in the computed spectral function from the RT-EOM-CC cumulant Green’s function approach, which is particularly useful for the assignments of satellite peaks when many-body effects dominate the spectra. Our preliminary numerical tests focus on reproducing, approximating, and characterizing the exact impurity Green’s function of the three-site and four-site single impurity Anderson models using the RT-EOM-CC cumulant Green’s function approach. The numerical tests allow us to have a direct comparison between the RT-EOM-CC cumulant Green’s function approach and other Green’s function approaches in the numerical exact limit.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH