Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “clustering algorithm”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

Clustering Acoustic Background Noise in the Stratosphere Using Machine Learning

Infrasound, characterized by low-frequency sound inaudible to humans (<20 Hz), emanates from natural and anthropogenic sources. Its efficacy for monitoring phenomena necessitates robust sensing networks. Traditional ground-based infrasound sensors have limitations due to atmospheric dynamics and noise interference. Balloon-bore sensors have emerged as an alternative, offering reduced noise and improved capabilities. This study bridges clustering algorithms with balloon borne infrasound data, a domain yet to be explored. Employing K-Means, DBSCAN, and GMM algorithms on normalized and reshaped data and only normalized data from a New Zealand-based NASA balloon flight, insights into background noise at stratospheric altitudes were revealed. Despite challenges arising from distinguishing signals amid unique background noise, this research provides vital reference material for noise analysis and calibration. Beyond infrasound event capture, the dataset enriches comprehension of background noise characteristics in the southern hemisphere.

47 OTHER INSTRUMENTATION↗

Prong Segmentation using Point Set Transformers in Multiple View Neutrino Detectors

NOvA is a long-baseline neutrino experiment studying neutrino oscillations by detecting neutrinos from the NuMI beam at Fermilab. Its physics analysis relies on accurate prong segmentation, which involves matching each hit to its source particle and identifying the particle type. This task has commonly been addressed using a combination of traditional clustering algorithms and convolutional neural networks (CNNs). However, NOvA’s detector design presents data as two sparse and decoupled 2D images (XZ and YZ views) rather than a native 3D representation, posing a significant challenge for traditional CNN-based models. In this talk, we propose a novel neural network based on the Point Set Transformer. By treating detector hits as sparse point clouds and implementing a cross-view attention mechanism, our model enables efficient information mixing between both views. Evaluated on NOvA simulated data, our model achieves superior accuracy while requiring significantly fewer computational resources compared to other models. Furthermore, the model demonstrates great performance when applied to Liquid Argon Time Projection Chamber (LArTPC) data, which shows its potential as a universal prong segmentation algorithm for multiple view neutrino detectors.

Liu, Jiaxi [UC, Irvine]↗

Uncertainty Quantification in CO2 Trapping Mechanisms: A Case Study of PUNQ-S3 Reservoir Model Using Representative Geological Realizations and Unsupervised Machine Learning

Evaluating uncertainty in CO2 injection projections often requires numerous high-resolution geological realizations (GRs) which, although effective, are computationally demanding. This study proposes the use of representative geological realizations (RGRs) as an efficient approach to capture the uncertainty range of the full set while reducing computational costs. A predetermined number of RGRs is selected using an integrated unsupervised machine learning (UML) framework, which includes Euclidean distance measurement, multidimensional scaling (MDS), and a deterministic K-means (DK-means) clustering algorithm. In the context of the intricate 3D aquifer CO2 storage model, PUNQ-S3, these algorithms are utilized. The UML methodology selects five RGRs from a pool of 25 possibilities (20% of the total), taking into account the reservoir quality index (RQI) as a static parameter of the reservoir. To determine the credibility of these RGRs, their simulation results are scrutinized through the application of the Kolmogorov–Smirnov (KS) test, which analyzes the distribution of the output. In this assessment, 40 CO2 injection wells cover the entire reservoir alongside the full set. The end-point simulation results indicate that the CO2 structural, residual, and solubility trapping within the RGRs and full set follow the same distribution. Simulating five RGRs alongside the full set of 25 GRs over 200 years, involving 10 years of CO2 injection, reveals consistently similar trapping distribution patterns, with an average value of Dmax of 0.21 remaining lower than Dcritical (0.66). Using this methodology, computational expenses related to scenario testing and development planning for CO2 storage reservoirs in the presence of geological uncertainties can be substantially reduced.

Mahjour, Seyed Kourosh↗

Chasing Accreted Structures within Gaia DR2 Using Deep Learning

In previous work, we developed a deep neural network classifier that only relies on phase-space information to obtain a catalog of accreted stars based on the second data release of Gaia (DR2). In this paper, we apply two clustering algorithms to identify velocity substructure within this catalog. We focus on the subset of stars with line-of-sight velocity measurements that fall in the range of Galactocentric radii $r\in [6.5,9.5]\,{\rm{kpc}}$ and vertical distances $| z| \lt 3\,{\rm{kpc}}$. Known structures such as Gaia Enceladus and the Helmi stream are identified. The largest previously unknown structure, Nyx, is a vast stream consisting of at least 200 stars in the region of interest. This study displays the power of the machine-learning approach by not only successfully identifying known features but also discovering new kinematic structures that may shed light on the merger history of the Milky Way.

Astronomy & Astrophysics↗

Unsupervised Machine Learning for Exploratory Data Analysis of Exoplanet Transmission Spectra

Abstract Transit spectroscopy is a powerful tool for decoding the chemical compositions of the atmospheres of extrasolar planets. In this paper, we focus on unsupervised techniques for analyzing spectral data from transiting exoplanets. After cleaning and validating the data, we demonstrate methods for: (i) initial exploratory data analysis, based on summary statistics (estimates of location and variability); (ii) exploring and quantifying the existing correlations in the data; (iii) preprocessing and linearly transforming the data to its principal components; (iv) dimensionality reduction and manifold learning; (v) clustering and anomaly detection; and (vi) visualization and interpretation of the data. To illustrate the proposed unsupervised methodology, we use a well-known public benchmark data set of synthetic transit spectra. We show that there is a high degree of correlation in the spectral data, which calls for appropriate low-dimensional representations. We explore a number of different techniques for such dimensionality reduction and identify several suitable options in terms of summary statistics, principal components, etc. We uncover interesting structures in the principal component basis, namely well-defined branches corresponding to different chemical regimes of the underlying atmospheres. We demonstrate that those branches can be successfully recovered with a K-means clustering algorithm in a fully unsupervised fashion. We advocate for lower-dimensional representations of the spectroscopic data in terms of the main principal components, in order to reveal the existing structure in the data and quickly characterize the chemical class of a planet.

Matchev, Konstantin T. (ORCID:0000000341829096)↗

rustpix

rustpix is a high-performance, open-source Rust library with first-class Python bindings (via PyO3) for processing pixel-detector data in neutron imaging. It targets time-stamping detectors such as Timepix3 (TPX3) at ORNL's Spallation Neutron Source (VENUS beamline), where each detected neutron deposits charge across a cluster of pixels within a very high-rate event stream (96M+ hits/sec). rustpix parses TPX3 event data in parallel using memory-mapped I/O, offers four interchangeable clustering algorithms (ABS adjacency-based search, DBSCAN, graph/union-find connected components, and a parallel grid method), and extracts weighted, super-resolved centroids to produce neutron-event lists. A streaming architecture lets it process files larger than available memory. rustpix is distributed as a pip-installable Python package (with NumPy integration), Rust crates, a command-line tool, and an interactive GUI; it writes HDF5, Apache Arrow, and CSV; and it is designed to extend to TPX4 and other detector types. Released as open-source under the MIT License.

Zhang, Chen [Oak Ridge National Laboratory (ORNL),↗

Data augmentation for disruption prediction via robust surrogate models

The goal of this work is to generate large statistically representative datasets to train machine learning models for disruption prediction provided by data from few existing discharges. Such a comprehensive training database is important to achieve satisfying and reliable prediction results in artificial neural network classifiers. Here, we aim for a robust augmentation of the training database for multivariate time series data using Student-t process regression. We apply Student-t process regression in a state space formulation via Bayesian filtering to tackle challenges imposed by outliers and noise in the training data set and to reduce the computational complexity. Thus, the method can also be used if the time resolution is high. We use an uncorrelated model for each dimension and impose correlations afterwards via coloring transformations. We demonstrate the efficacy of our approach on plasma diagnostics data of three different disruption classes from the DIII-D tokamak. To evaluate if the distribution of the generated data is similar to the training data, we additionally perform statistical analyses using methods from time series analysis, descriptive statistics, and classic machine learning clustering algorithms.

97 MATHEMATICS AND COMPUTING↗

Discriminative Dimensionality Reduction using Deep Neural Networks for Clustering of LIGO Data

In this paper, leveraging the capabilities of neural networks for modeling the non-linearities that exist in the data, we propose several models that can project data into a low dimensional, discriminative, and smooth manifold. The proposed models can transfer knowledge from the domain of known classes to a new domain where the classes are unknown. A clustering algorithm is further applied in the new domain to find potentially new classes from the pool of unlabeled data. The research problem and data for this paper originated from the Gravity Spy project which is a side project of Advanced Laser Interferometer Gravitational-wave Observatory (LIGO). The LIGO project aims at detecting cosmic gravitational waves using huge detectors. However non-cosmic, non-Gaussian disturbances known as "glitches", show up in gravitational-wave data of LIGO. This is undesirable as it creates problems for the gravitational wave detection process. Gravity Spy aids in glitch identification with the purpose of understanding their origin. Since new types of glitches appear over time, one of the objective of Gravity Spy is to create new glitch classes. Towards this task, we offer a methodology in this paper to accomplish this.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Scaling Building Energy Audits through Machine Learning Methods on Novel Drone Image Data

Building energy audits are time-consuming and labor-intensive. This paper describes a new method using machine learning (ML) techniques on novel data sources (drone images) to improve the identification of building characteristics and retrofit opportunities, and thereby reduce the effort for audits. The new ML method includes: (1) Building footprint extraction using line extraction, polygonization, and polygon-merging, (2) Building envelope extraction using PIX4d modeling software to reconstruct a building 3D model, (3) Visualization tool for viewing images from the 3D model, (4) Window-to-wall ratio (WWR) using state-of-art deep neural network semantic segmentation, (5) Envelope thermal anomaly detection using an unsupervised machine learning clustering algorithm, and (6) Rooftop energy equipment detection based on an object detection algorithm. The testing of this method involved a comparison of additional ML-generated information overlaid on current ‘state-of-practice’ audit and remote assessment baselines using evaluation metrics: labor time and associated cost, marginal benefits of using ML-generated information in workflows for audits and remote assessments, integration potential with existing processes and tools, and replicability/scalability of the method. In two test buildings in California that had comprehensive drawings and meter data available, the ML method effectively generated a building footprint, envelope, rooftop equipment, WWR, and locations of envelope thermal anomalies. Projected target segments of the ML method are sites with minimal drawings and energy data, and underserved sectors such as multistoried housing, disadvantaged communities, and schools for which the ML method can enable identification of building asset characteristics and prioritization of envelope retrofits and decentralized energy equipment retrofits.

Singh, Reshma↗

Modeling of H2 Dispersion at ARIES

Hydrogen is a versatile and clean energy carrier that can be produced from various renewable sources such as wind, solar, and hydropower and help decarbonize electricity grids, industry, and transportation. Using the Hydrogen Research Facility under Advanced Research on Integrated Energy Systems (ARIES) at the National Renewable Energy Laboratory's (NREL) Flatirons campus as a test bench, the study examines the feasibility, useability, and value of using computational fluid dynamics (CFD) techniques to model hydrogen dispersion. The ARIES facility was chosen because controlled hydrogen releases can be performed at a rate of 27 kg-H2/hr. Site-specific atmospheric and weather condition data such as wind speed and temperature were used as inputs to the model. The results show statistical distributions and ranges of hydrogen concentrations at locations throughout the domain. Wind conditions are found to significantly impact the release behavior, including the hydrogen cloud's direction and concentrations. At low wind speeds (below 1 mph), hydrogen forms a cloud and at higher wind speeds (> 2-4 mph) hydrogen plume stretches in the direction of wind momentum. From >100 simulations for ARIES site-specific conditions, statistical quantities combined with a clustering algorithm were used to propose sensor location at various elevations from ground.

dispersion↗

NWPEsSe: an Adaptive-Learning Global Optimization Algorithm for Nanosized Cluster Systems

Global optimization constitutes an important and fundamental problem in theoretical studies in many chemical fields, such as catalysis, materials or separations problems. In this paper, a novel algorithm has been developed for the global optimization of large systems including neat and ligated clusters in gas phase, and supported clusters in periodic boundary conditions. The method is based on an updated artificial bee colony (ABC) algorithm method, that allows for adaptive-learning during the search process. The new algorithm is tested against four classes of systems of diverse chemical nature: gas phase Au_55, ligated Au_8^(2+), Au_8 supported on graphene oxide and defected rutile, and a large cluster assembly ?[Co?_6 Te_8 (PEt_3 )_6][C_60 ]_n, with sizes ranging between 1 to 3 nm and containing up to 1300 atoms. Reliable global minima (GMs) are obtained for all cases, either confirming published data or reporting new lower energy structures. The algorithm and interface to other codes in the form of an independent program, Northwest Potential Energy Search Engine (NWPEsSe), is freely available and it provides a powerful and efficient approach for global optimization of nanosized cluster systems. The work described in this publication was performed at Pacific Northwest National Laboratory (PNNL), which is operated by Battelle for the United States Department of Energy (DOE) under Contract DE-AC05-76RL0180. J. Z. and V.-A. G. acknowledge support from DOE, Office of Science, Office of Basic Energy Sci-ences, Chemical, Geological and Biological Sciences Division and computing resources from PNNL’s Research Computing Facility and the National Energy Research Scientific Computing Center.

Zhang, Jun↗

ParChain: a framework for parallel hierarchical agglomerative clustering using nearest-neighbor chain

This paper studies the hierarchical clustering problem, where the goal is to produce a dendrogram that represents clusters at varying scales of a data set. We propose the ParChain framework for designing parallel hierarchical agglomerative clustering (HAC) algorithms, and using the framework we obtain novel parallel algorithms for the complete linkage, average linkage, and Ward's linkage criteria. Compared to most previous parallel HAC algorithms, which require quadratic memory, our new algorithms require only linear memory, and are scalable to large data sets. ParChain is based on our parallelization of the nearest-neighbor chain algorithm, and enables multiple clusters to be merged on every round. We introduce two key optimizations that are critical for efficiency: a range query optimization that reduces the number of distance computations required when finding nearest neighbors of clusters, and a caching optimization that stores a subset of previously computed distances, which are likely to be reused. Experimentally, we show that our highly-optimized implementations using 48 cores with two-way hyper-threading achieve 5.8--110.1x speedup over state-of-the-art parallel HAC algorithms and achieve 13.75--54.23x self-relative speedup. Compared to state-of-the-art algorithms, our algorithms require up to 237.3x less space. Our algorithms are able to scale to data set sizes with tens of millions of points, which existing algorithms are not able to handle.

Computer Science↗

Miscentring of optical galaxy clusters based on Sunyaev–Zeldovich counterparts

ABSTRACT The ‘miscentring effect’, i.e. the offset between a galaxy cluster’s optically defined centre and the centre of its gravitational potential, is a significant systematic effect on brightest cluster galaxy (BCG) studies and cluster lensing analyses. We perform a cross-match between the optical cluster catalogue from the Hyper Suprime-Cam (HSC) Survey S19A Data Release and the Sunyaev–Zeldovich cluster catalogue from Data Release 5 of the Atacama Cosmology Telescope (ACT). We obtain a sample of 186 clusters in common in the redshift range $0.1 \le z \le 1.4$ over an area of 469 deg$^2$. By modelling the distribution of centring offsets in this fiducial sample, we find a miscentred fraction (corresponding to clusters offset by more than 330 kpc) of ∼25 per cent, a value consistent with previous miscentring studies. We examine the image of each miscentred cluster in our sample and identify one of several reasons to explain the miscentring. Some clusters show significant miscentring for astrophysical reasons, i.e. ongoing cluster mergers. Others are miscentred due to non-astrophysical, systematic effects in the HSC data or the cluster-finding algorithm. After removing all clusters with clear, non-astrophysical causes of miscentring from the sample, we find a considerably smaller miscentred fraction, $\sim 10~\,\rm per\,cent$. We show that the gravitational lensing signal within 1 Mpc of miscentred clusters is considerably smaller than that of well-centred clusters, and we suggest that the ACT SZ centres are a better estimate of the true cluster potential centroid.

Ding, Jupiter (ORCID:0000000296119799)↗

A Perspective on Quantum Computing Applications in Quantum Chemistry Using 25-100 Logical Qubits

The intersection of quantum computing and quantum chemistry represents a promising frontier for achieving quantum utility in domains of both scientific and societal relevance. Owing to the exponential growth of classical resource requirements for simulating quantum systems, quantum chemistry has long been recognized as a natural candidate for quantum computation. This perspective focuses on identifying scientifically meaningful use cases where early fault-tolerant quantum computers, which are considered to be equipped with approximately 25-100 logical qubits, could deliver tangible impact. While recent advances in classical computing have pushed the boundaries of tractable simulations to unprecedented scales, this logical-qubit regime represents the first window where quantum devices can pursue qualitatively distinct strategies, such as polynomial-scaling phase estimation, direct simulation of quantum dynamics, and active-space embedding, that remain challenging for classical solvers, such as multireference charge-transfer and conical-intersection states central to photochemistry and materials design. We highlight near-term opportunities in algorithm and software design, discuss representative chemical problems suited for quantum acceleration, and propose strategic roadmaps and collaborative pathways for advancing practical quantum utility in quantum chemistry.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Reducing Communication in Graph Neural Network Training

Graph Neural Networks (GNNs) are powerful and flexible neural networks that use the naturally sparse connectivity information of the data. GNNs represent this connectivity as sparse matrices, which have lower arithmetic intensity and thus higher communication costs compared to dense matrices, making GNNs harder to scale to high concurrencies than convolutional or fully-connected neural networks. Here, we introduce a family of parallel algorithms for training GNNs and show that they can asymptotically reduce communication compared to previous parallel GNN training methods. We implement these algorithms, which are based on 1D, 1. 5D, 2D, and 3D sparse-dense matrix multiplication, using torch.distributed on GPU-equipped clusters. Our algorithms optimize communication across the full GNN training pipeline. We train GNNs on over a hundred GPUs on multiple datasets, including a protein network with over a billion edges.

97 MATHEMATICS AND COMPUTING↗