Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “clustering algorithm”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 343 records · Page 19

Netostat: analyzing dynamic flow patterns in high-speed networks

Understanding flow traffic patterns in networks, such as the Internet or service provider networks, is crucial to improving their design and building them robustly. However, as networks grow and become more complex, it is increasingly cumbersome and challenging to study how the many flow patterns, sizes and the continually changing source-destination pairs in the network evolve with time. Here, we present Netostat, a visualization-based network analysis tool that uses visual representation and a mathematics framework to study and capture flow patterns, using graph theoretical methods such as clustering, similarity and difference measures. Netostat generates an interactive graph of all traffic patterns in the network, to isolate key elements that can provide insights for traffic engineering. We present results for U.S. and European research networks, ESnet and GEANT, demonstrating network state changes, to identify major flow trends, potential points of failure, and bottlenecks.

97 MATHEMATICS AND COMPUTING↗

CASM — A software package for first-principles based study of multicomponent crystalline solids

CASM is a software package that enables first-principles based studies of crystalline materials. It has been designed to treat coupled chemical, mechanical, vibrational, and magnetic degrees of freedom to determine ground state and finite temperature properties of crystals. The symmetry of the underlying parent crystal structure is used to enumerate perturbations of the parent crystal structure and generate derivative structures which can be input to first-principles calculations in order to explore the ground state energy landscape. CASM algorithmically constructs cluster expansions that fully couple discrete and continuous degrees of freedom, and generates highly efficient code to evaluate the cluster expansion basis functions. Widely used machine learning methods are integrated for fitting expansion coefficients to first-principle calculations. The fully parameterized cluster expansions can be combined with (kinetic) Monte Carlo methods to calculate finite temperature thermodynamic and kinetic properties. CASM Alloy Manager identifies distinct parent crystal structures in an alloy system, creating individual projects for each, and integrating the results. The integrated infrastructure facilitates the linkage between first-principles statistical mechanics predictions with higher length scale computational methods, such as phase field simulations. In conclusion, CASM projects are designed to be easy to share in repositories, to re-use, and to extend to include additional chemical species or types of degrees of freedom.

36 MATERIALS SCIENCE↗

Thermochemical Data Fusion Using Graph Representation Learning

Large databases are required for “Big Data” applications in catalysis and materials science. Thermochemical databases can be created by combining data from various sources and by correcting low-fidelity datasets to higher accuracy with minimal computation. To achieve this “data fusion”, thermochemical quantities of interest, calculated at various levels of density functional theory (DFT), need to be mapped to the same, high levels of theory. In this work, a graph theoretical, statistical framework is proposed for such tasks. Subgraph frequencies are shown to provide a natural representation for learning these fusion maps. The maps are linear and are learnt with automated descriptor selection. Using a dataset of as few as ~1% from the QM9 database of 133,885 molecules, these models can predict multiple thermochemical quantities at a higher level of theory with an accuracy of 1 kcal/mol. Here, the method is explainable, generalizable, and provides a diagnostic tool for outlier identification

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Thermal Weight Determination and Interstate Coupling in State-Averaged ADAPT-VQE

Characterizing electronic thermal states at low temperatures is an important but challenging task in quantum chemistry and condensed matter physics, making it a prime candidate for a useful application in quantum computing. One of the most successful methods for state preparation on quantum computers is the Adaptive, Problem-Tailored (ADAPT) Variational Quantum Eigensolver (VQE), which has recently been generalized to treat excited states within a state-averaged framework as well as Gibbs states. In this work, we introduce Helmholtz-Optimized Thermal (HOT) ADAPT-VQE, an ancilla-free strategy for preparing Gibbs states that directly minimizes the Helmholtz free energy by targeting the dominant eigenstates of the thermal ensemble. We demonstrate the usefulness of HOT-ADAPT-VQE by predicting the free energy of two model systems with strongly correlated ground states: (1) the Fe 2+ cation in a magnetic field and (2) a [Cu 2 O 7 ] 10– fragment of the Mott insulator La 2 CuO 4 . Our results demonstrate that HOT-ADAPT-VQE significantly improves upon Gibbs-state estimates from multistate variants of ADAPT-VQE, often with substantially shallower quantum circuits, making it a promising candidate for thermal-state calculations.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Cancer genomics predicts disease relapse and therapeutic response to neoadjuvant chemotherapy of hormone sensitive breast cancers

Several studies provide insight into the landscape of breast cancer genomics with the genomic characterization of tumors offering exceptional opportunities in defining therapies tailored to the patient’s specific need. However, translating genomic data into personalized treatment regimens has been hampered partly due to uncertainties in deviating from guideline based clinical protocols. Here we report a genomic approach to predict favorable outcome to treatment responses thus enabling personalized medicine in the selection of specific treatment regimens. The genomic data were divided into a training set of N = 835 cases and a validation set consisting of 1315 hormone sensitive, 634 triple negative breast cancer (TNBC) and 1365 breast cancer patients with information on neoadjuvant chemotherapy responses. Patients were selected by the following criteria: estrogen receptor (ER) status, lymph node invasion, recurrence free survival. The k-means classification algorithm delineated clusters with low- and high- expression of genes related to recurrence of disease; a multivariate Cox’s proportional hazard model defined recurrence risk for disease. Classifier genes were validated by Immunohistochemistry (IHC) using tissue microarray sections containing both normal and cancerous tissues and by evaluating findings deposited in the human protein atlas repository. Based on the leave-on-out cross validation procedure of 4 independent data sets we identified 51-genes associated with disease relapse and selected 10, i.e. TOP2A, AURKA, CKS2, CCNB2, CDK1 SLC19A1, E2F8, E2F1, PRC1, KIF11 for in depth validation. Expression of the mechanistically linked disease regulated genes significantly correlated with recurrence free survival among ER-positive and triple negative breast cancer patients and was independent of age, tumor size, histological grade and node status. Importantly, the classifier genes predicted pathological complete responses to neoadjuvant chemotherapy (P < 0.001) with high expression of these genes being associated with an improved therapeutic response toward two different anthracycline-taxane regimens; thus, highlighting the prospective for precision medicine. Our study demonstrates the potential of classifier genes to predict risk for disease relapse and treatment response to chemotherapies. The classifier genes enable rational selection of patients who benefit best from a given chemotherapy thus providing the best possible care. The findings encourage independent clinical validation.

59 BASIC BIOLOGICAL SCIENCES↗

Dynamic PRA-Based Estimation of PWR Coping Time Using a Surrogate Model for Accident Tolerant Fuel

In this study, we propose an interpolation-based response surface surrogate methodology to manage a large number of scenarios in dynamic probabilistic risk assessment. It adopts the shape Dynamic Time Warping algorithm to cluster the interpolation neighborhood from time series sample data. The interpolation method was adapted from Taylor Kriging to allow a reduced-order model of the Taylor series. In order to demonstrate its applicability to complex issues in risk assessment for nuclear engineering, an example risk response surface to estimate emergency core cooling system (ECCS) criteria for triplex silicon carbide (SiC) accident-tolerant fuel was constructed. The response surface was exploited to estimate the cumulative failure probability of the fuel cladding structure due to the uncertainties in operator actions and safety systems. The functional failures were assessed based on a combination of individual layer failures computed by coupling Risk Analysis Virtual Environment software with a pressurized water reactor 1000-MW(electric) RELAP5 model and the in-house fuel performance assessment module. Results showed that SiC cladding failure probability spiked less than 1 min after a large-break loss-of- coolant accident whenever the current ECCS criteria for Zircaloy-4 (Zr-4) cladding was used. However, it still provides an increased safety margin of three orders of magnitude compared to Zr-4. This positive margin could be utilized to relax active ECCS requirements by allowing deviations of up to 450 s in its actuation time. The proposed surrogate methodology generated a response surface of SiC cladding failure probability reasonably well, with a significant savings of computation time. This methodology is expected to be useful in the analysis of system response with complex uncertainty sources.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

The drivers and predictability of wildfire re-burns in the western United States (US)

Evidence is mounting that the effectiveness of using prescribed burns as a management tactic may be diminishing due to the higher incidence of wildfire re-burns. The development of predictive models of re-burns is thus essential to better understand their primary drivers so that forest management practices can be updated to account for these events. First, we assess the potential for human activity as a driver of re-burns by evaluating re-burn trends both within and outside of the wildland–urban interface (WUI) of the western US. Next, we investigate the predictability of re-burns through the application of both random forest and the explanatory machine learning non-negative matrix factorization using k-means clustering (NMFk) algorithms to predict re-burn occurrence over California based on a number of climate factors. Our findings indicate that while most states showed increasing trends within the WUI when trends were conducted over longer moving windows (e.g. 20 years), California was the only state where the rate of increase was consistently higher in the WUI, indicating a stronger potential for human activity as a driver in that location. Furthermore, we find model performance was found to be robust over most of California (Testing F1 scores = 0.688), although results were highly variable based on EPA level III Ecoregion (F1 scores = 0.0–0.778). Insights provided from this study will lead to a better understanding of climate and human activity drivers of re-burns and how these vary at broad spatial scales so that improvements in forest management practices can be tuned according to the level of change that is expected for a given region.

54 ENVIRONMENTAL SCIENCES↗

I-GCN: A Graph Convolutional Network Accelerator with Runtime Locality Enhancement through Islandization

In this paper, we propose a novel hardware accelerator for GCN inference called I-GCN that significantly improves data locality and reduces unnecessary computation through a new online graph restructuring algorithm we refer to as islandization. The proposed algorithm finds clusters of nodes with strong internal but weak external connections. The islandization process yields two major benefits. First, by processing islands rather than individual nodes, there is better on-chip data reuse and fewer off-chip memory accesses. Second, there is less redundant computation as aggregation for common/shared neighbors in an island can be reused. The parallel search, identification, and leverage of graph islands are all handled purely in hardware at runtime working in an incremental pipelined manner. This is done without any preprocessing of the graph data or adjustment of the GCN model structure.

Geng, Tong↗

GenomeFace v1.0

GenomeFace is meta-genome binning software. Metagenomic binning, the process of grouping DNA sequences into taxonomic units, is critical for understanding the functions, interactions, and evolutionary dynamics of microbial communities. We propose a deep learning approach to binning using two neural networks, one based on composition and another on environmental abundance, dynamically weighting the contribution of each based on characteristics of the input data. Trained on over 43,000 prokaryotic genomes, our network for composition-based binning is inspired by metric learning techniques used for facial recognition. Using a task-specific, multi-GPU accelerated algorithm to cluster the embeddings produced by our network, our binner leverages marker genes observed to be universally present in nearly all taxa to grade and select optimal clusters of sequences from a hierarchy of candidates. We evaluate our approach on four simulated datasets with known ground truth. Our linear time integration of marker genes recovers more near complete genomes than state of the art but computationally infeasible solutions using them, while being over an order of magnitude faster. Finally, we demonstrate the scalability and acuity of our approach by testing it on three of the largest metagenome assemblies ever performed. Compared to other binners, we produced 47%-183% more near complete genomes. From these datasets, we find over the genomes of over 3000 new candidate species which have never been previously cataloged, representing a potential 4% expansion of the known bacterial tree of life.

Lettich, Richard [Lawrence Berkeley National Labor↗

Methods of extending crop signatures from one area to another

Efforts to develop a technology for signature extension during LACIE phases 1 and 2 are described. A number of haze and Sun angle correction procedures were developed and tested. These included the ROOSTER and OSCAR cluster-matching algorithms and their modifications, the MLEST and UHMLE maximum likelihood estimation procedures, and the ATCOR procedure. All these algorithms were tested on simulated data and consecutive-day LANDSAT imagery. The ATCOR, OSCAR, and MLEST algorithms were also tested for their capability to geographically extend signatures using LANDSAT imagery.

Minter, T. C.↗

Onboard Algorithms for Data Prioritization and Summarization of Aerial Imagery

Many current and future NASA missions are capable of collecting enormous amounts of data, of which only a small portion can be transmitted to Earth. Communications are limited due to distance, visibility constraints, and competing mission downlinks. Long missions and high-resolution, multispectral imaging devices easily produce data exceeding the available bandwidth. To address this situation computationally efficient algorithms were developed for analyzing science imagery onboard the spacecraft. These algorithms autonomously cluster the data into classes of similar imagery, enabling selective downlink of representatives of each class, and a map classifying the terrain imaged rather than the full dataset, reducing the volume of the downlinked data. A range of approaches was examined, including k-means clustering using image features based on color, texture, temporal, and spatial arrangement

Chien, Steve A.↗

Electromagnetic Shower Reconstruction and Identification in FASER's Emulsion Detector for LHC Forward Neutrino Measurements

We present methods for electromagnetic shower reconstruction and identification in the FASERnu emulsion detector using 100 GeV and 200 GeV electron test-beam data from the CERN SPS H4 beamline. The reconstruction employs a clustering-based algorithm without energy-dependent tuning to determine shower axes. A multi-level identification chain comprising track pre-selection, a cut-based selection, and a BDT classifier achieves combined background rejection rates of 99.99% (100 GeV) and 99.94% (200 GeV). The method reaches total reconstruction and identification efficiencies of 58.9% (100 GeV) and 70.8% (200 GeV) evaluated from simulated samples. Energy reconstruction using the total number of reconstructed segments as the calorimetric estimator yields relative biases of +0.6% (100 GeV) and -0.8% (200 GeV), with resolutions of 25.4% and 22.6%, respectively. Systematic uncertainties on the energy reconstruction are dominated by variations in emulsion film detection efficiency, contributing (+10.9%/-8.2%) at 100 GeV and (+10.3%/-6.9%) at 200 GeV. The methodology provides a validated framework for electron neutrino identification with the FASERnu detector at the LHC.

Mammen Abraham, Roshan [UC, Irvine]↗

Unsupervised, Robust Estimation-based Clustering for Multispectral Images

To prepare for the challenge of handling the archiving and querying of terabyte-sized scientific spatial databases, the NASA Goddard Space Flight Center's Applied Information Sciences Branch (AISB, Code 935) developed a number of characterization algorithms that rely on supervised clustering techniques. The research reported upon here has been aimed at continuing the evolution of some of these supervised techniques, namely the neural network and decision tree-based classifiers, plus extending the approach to incorporating unsupervised clustering algorithms, such as those based on robust estimation (RE) techniques. The algorithms developed under this task should be suited for use by the Intelligent Information Fusion System (IIFS) metadata extraction modules, and as such these algorithms must be fast, robust, and anytime in nature. Finally, so that the planner/schedule module of the IlFS can oversee the use and execution of these algorithms, all information required by the planner/scheduler must be provided to the IIFS development team to ensure the timely integration of these algorithms into the overall system.

Netanyahu, Nathan S.↗

State Preparation in Quantum Algorithms for Fragment-Based Quantum Chemistry

State preparation for quantum algorithms is crucial for achieving high accuracy in quantum chemistry and competing with classical algorithms. The localized active space–unitary coupled cluster (LAS–UCC) algorithm iteratively loads a fragment-based multireference wave function onto a quantum computer. Here, in this study, we compare two state preparation methods, quantum phase estimation (QPE) and direct initialization (DI), for each fragment. We test the two state preparation methods on three systems, ranging from a model system, a set of interacting hydrogen molecules, to more realistic chemical problems, like the C–C double bond breaking in transbutadiene and the spin ladder in a bimetallic system. We analyze the impact of QPE parameters, such as the number of ancilla qubits and Trotter steps, on the prepared state. We find a trade-off between the methods, where DI requires fewer resources for smaller fragments, while QPE is more efficient for larger fragments. Our resource estimates highlight the benefits of system fragmentation in state preparation for subsequent quantum chemical calculations. These findings have broad applications for preparing multireference quantum chemical wave functions on quantum circuits that can be used for realistic chemical applications.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Discovery of correlated electron molecular orbital materials using graph representations

Correlated electron molecular orbital (CEMO) materials host emergent electronic states built from molecular orbitals localized over clusters of transition metal ions yet have historically been discovered sporadically and generally been treated as isolated case studies. Here we establish CEMO materials as a systematically discoverable class and introduce a graph-based framework to identify, classify, and organize transition-metal cluster motifs in inorganic solids. Starting from crystal structures in the Materials Project, we construct transition metal connectivity graphs, extract cluster motifs using a bond-cutting algorithm, and determine cluster point groups, effective cluster sublattice dimensionality, and translational symmetry. Applying this approach in a high-throughput screen of 34,548 compounds yields 5,306 cluster-containing materials, including 2,627 stable or metastable compounds with isolated clusters and 984 materials featuring mixed-metal clusters. The resulting dataset reveals symmetry and element dependent trends in cluster formation. By integrating cluster classification with flat band lattice topology and battery-relevant information, we provide further relevant information to multiple scientific communities. The accompanying open dataset, Cluster Finder software, and interactive web platform enable systematic exploration of cluster driven electronic phenomena and establish a general pathway for discovering correlated quantum materials and functional materials with cluster-based or extended metal-metal bonding in inorganic solids.

Akhond, Md. Rajbanul [Department of Chemistry, 800↗

Exploring the Landscape of Distributed Graph Clustering on Leadership Supercomputers

The rapid growth of large-scale datasets in fields like biology and social networks has driven the need for advanced graph analytics techniques. Community detection, a fundamental task in graph analytics, identifies closely connected groups of nodes within a network, providing valuable insights across various disciplines. This study focuses on two classic community detection methods, the Louvain algorithm and Markov Clustering (MCL), and evaluates the performance of two prominent distributed community detection algorithms: HiPDPL-GPU, our prior implementation, and HipMCL. We conduct experiments on GPU-accelerated heterogeneous HPC systems, Summit and Frontier, to assess their performance under varying conditions. Our objective is to identify the strengths and weaknesses of these algorithms in terms of scalability, and quality of solutions. We evaluate these algorithms on a diverse set of 70+ networks spanning 13 domains, with sizes ranging up to 4.2 billion edges. Our results demonstrate that HiPDPL-GPU consistently outperforms HipMCL, especially for large-scale networks. HiPDPL-GPU achieves significantly faster runtimes (47x to 1439x), higher modularity scores, and improved scalability. These findings highlight HiPDPL-GPU as a promising solution for efficient and effective large-scale graph analytics in diverse application domains, and provide insights into the feasibility of using MCL-based approaches for certain application domains.

Community detection, graph algorithms↗

A Generative Model for Realistic Galaxy Cluster X-Ray Morphologies

Abstract The X-ray morphologies of clusters of galaxies display significant variations, reflecting their dynamical histories and the nonlinear dependence of X-ray emissivity on the density of the intracluster gas. Qualitative and quantitative assessments of X-ray morphology have long been considered a proxy for determining whether clusters are dynamically active or “relaxed.” Conversely, the use of circularly or elliptically symmetric models for cluster emission can be complicated by the variety of complex features realized in nature, spanning scales from megaparsecs down to the resolution limit of current X-ray observatories. In this work, we use mock X-ray images from simulated clusters from The Three Hundred project to define a basis set of cluster image features. We take advantage of the clusters’ approximate self-similarity to minimize the differences between images before encoding the remaining diversity through a distribution of high-order polynomial coefficients. Principal component analysis then provides an orthogonal basis for this distribution, corresponding to natural perturbations from an average model. This representation allows novel, realistically complex X-ray cluster images to be easily generated, and we provide code to do so. The approach provides a simple way to generate training data for cluster image analysis algorithms and could be straightforwardly adapted to generate clusters displaying specific types of features or selected by physical characteristics available in the original simulations.

79 ASTRONOMY AND ASTROPHYSICS↗