Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “clustering algorithm”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

Evaluation of the procedure 1A component of the 1980 US/Canada wheat and barley exploratory experiment

Several techniques which use clusters generated by a new clustering algorithm, CLASSY, are proposed as alternatives to random sampling to obtain greater precision in crop proportion estimation: (1) Proportional Allocation/relative count estimator (PA/RCE) uses proportional allocation of dots to clusters on the basis of cluster size and a relative count cluster level estimate; (2) Proportional Allocation/Bayes Estimator (PA/BE) uses proportional allocation of dots to clusters and a Bayesian cluster-level estimate; and (3) Bayes Sequential Allocation/Bayesian Estimator (BSA/BE) uses sequential allocation of dots to clusters and a Bayesian cluster level estimate. Clustering in an effective method in making proportion estimates. It is estimated that, to obtain the same precision with random sampling as obtained by the proportional sampling of 50 dots with an unbiased estimator, samples of 85 or 166 would need to be taken if dot sets with AI labels (integrated procedure) or ground truth labels, respectively were input. Dot reallocation provides dot sets that are unbiased. It is recommended that these proportion estimation techniques are maintained, particularly the PA/BE because it provides the greatest precision.

Chapman, G. M.↗

Clustering and Recurring Anomaly Identification: Recurring Anomaly Detection System (ReADS)

This viewgraph presentation reviews the Recurring Anomaly Detection System (ReADS). The Recurring Anomaly Detection System is a tool to analyze text reports, such as aviation reports and maintenance records: (1) Text clustering algorithms group large quantities of reports and documents; Reduces human error and fatigue (2) Identifies interconnected reports; Automates the discovery of possible recurring anomalies; (3) Provides a visualization of the clusters and recurring anomalies We have illustrated our techniques on data from Shuttle and ISS discrepancy reports, as well as ASRS data. ReADS has been integrated with a secure online search

McIntosh, Dawn↗

Data-driven evaluation of HVAC operation and savings in commercial buildings

Commercial buildings consumed 36% of electricity, or 1.35 trillion kWh, in the United States in 2017, and almost 30% of this energy was wasted. Much of this loss can be attributed to inefficient heating ventilation and air con­ditioning (HVAC) systems. By improving the operational conditions of HVAC, significant savings can be achieved. However, most buildings and building equipment do not use costly sub-meters to monitor and address performance issues, and on-site auditing can be expensive and insufficient. Alternatively in this study, we propose a data-driven method to identify savings opportunities using only whole building meter data and without setting foot in the building. For this purpose, we introduced two algorithms that virtually quantify the value of a thermostat setpoint setback and HVAC rescheduling. Additionally, we developed novel methods for detecting occupancy patterns and quantifying the baseload of the HVAC operation. Using a clustering algorithm, we identified those buildings for which HVAC savings was significant and further categorized the buildings based on their potential for savings. A population study of over 432 commercial buildings demonstrated a median percentage energy savings of 1.6% from a baseload reduction and 2.1% from HVAC rescheduling. Additionally, results indicate that retail buildings have the highest potential for savings among the building types studied.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Segmentation of multifrequency polarimetric radar images to facilitate the inference of geophysical parameters

An unsupervised clustering algorithm is used to segment multifrequency polarimetric radar data from the NASA/JPL airborne SAR (synthetic aperture radar). Twenty-two parameters are evaluated for their discriminatory capability for each pixel of an image. A clustering analysis is then performed using different subsets of these parameters. This analysis relies on data taken as part of an intensive field experiment during the summer of 1988 in the vicinity of the Pisgah lava flow in the Mojave Desert in southern California. As part of the experiment, extensive ground truth was acquired, including dielectric constant and topography measurements. Segmentation results show good agreement with these measurements.

Burnette, C F.↗

5S ribosomal ribonucleic acid sequences in Bacteroides and Fusobacterium: evolutionary relationships within these genera and among eubacteria in general

The 5S ribosomal ribonucleic acid (rRNA) sequences were determined for Bacteroides fragilis, Bacteroides thetaiotaomicron, Bacteroides capillosus, Bacteroides veroralis, Porphyromonas gingivalis, Anaerorhabdus furcosus, Fusobacterium nucleatum, Fusobacterium mortiferum, and Fusobacterium varium. A dendrogram constructed by a clustering algorithm from these sequences, which were aligned with all other hitherto known eubacterial 5S rRNA sequences, showed differences as well as similarities with respect to results derived from 16S rRNA analyses. In the 5S rRNA dendrogram, Bacteroides clustered together with Cytophaga and Fusobacterium, as in 16S rRNA analyses. Intraphylum relationships deduced from 5S rRNAs suggested that Bacteroides is specifically related to Cytophaga rather than to Fusobacterium, as was suggested by 16S rRNA analyses. Previous taxonomic considerations concerning the genus Bacteroides, based on biochemical and physiological data, were confirmed by the 5S rRNA sequence analysis.

NASA Discipline Exobiology↗

Partitioning of Large-Scale Power Electronics-Based Power Systems for Small-Signal Stability Analysis

The nodal admittance matrix (NAM)-based approach is suitable for analyzing the small-signal stability of large-scale power electronics-based power systems (PEPSs) as it preserves the system structure by utilizing the admittance matrix. Previously, NAM-based area partition has been proposed, which divides the system into various subareas and interconnections for easier analysis of the low-dimension matrix compared to the entire system-based high-dimension matrix. However, no partition algorithm has been presented for the NAM-based area partition method. This paper focuses on implementing the spectral partitioning algorithm for partitioning large-scale PEPSs into a low-dimension matrix to reduce the computation complexity of the analysis. These spectral components facilitate data transformation into a new space, enabling the application of traditional clustering methods like k-means. To evaluate the performance of the partitioning method, the subareas and interconnections obtained from the spectral clustering algorithm are incorporated into the NAM-based area partition method for a large system with 140 buses. The computational times of the original method, where the NAM-based criterion is directly applied to the entire system, are compared with those of the NAM-based partition method in MATLAB. PSCAD simulations of the whole system and the obtained subareas are conducted to validate the effectiveness of the proposed algorithm.

Nupur, Nupur↗

A new solution-adaptive grid generation method for transonic airfoil flow calculations

The clustering algorithm is controlled by a second-order, ordinary differential equation which uses the airfoil surface density gradient as a forcing function. The solution to this differential equation produces a surface grid distribution which is automatically clustered in regions with large gradients. The interior grid points are established from this surface distribution by using an interpolation scheme which is fast and retains the desirable properties of the original grid generated from the standard elliptic equation approach.

Nakamura, S.↗

Sub-10 nm Probing of Ferroelectricity in Heterogeneous Materials by Machine Learning Enabled Contact Kelvin Probe Force Microscopy

Reducing the dimensions of ferroelectric materials down to the nanoscale has strong implications on the ferroelectric polarization pattern and on the ability to switch the polarization. As the size of ferroelectric domains shrinks to the nanometer scale, the heterogeneity of the polarization pattern becomes increasingly pronounced, enabling a large variety of possible polar textures in nanocrystalline and nanocomposite materials. Critical to the understanding of fundamental physics of such materials and hence their applications in electronic nanodevices is the ability to investigate their ferroelectric polarization at the nanoscale in a nondestructive way. We show that contact Kelvin probe force microscopy (cKPFM) combined with a k-means response clustering algorithm enables to measure the ferroelectric response at a mapping resolution of 8 nm. In a BaTiO 3 thin film on silicon composed of tetragonal and hexagonal nanocrystals, we determine a nanoscale lateral distribution of discrete ferroelectric response clusters, fully consistent with the nanostructure determined by transmission electron microscopy. Moreover, we apply this data clustering method to the cKPFM responses measured at different temperatures, which allows us to follow the corresponding change in the polarization pattern as the Curie temperature is approached and across the phase transition. This work opens up perspectives for mapping complex ferroelectric polarization textures such as curled/swirled polar textures that can be stabilized in epitaxial heterostructures and more generally for mapping the polar domain distribution of any spatially highly heterogeneous ferroelectric materials.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Nearby stellar substructures in the Galactic halo from DESI Milky Way Survey Year 1 Data Release

We report five nearby ($d_{\mathrm{helio}} < 5$ kpc) stellar substructures in the Galactic halo from a subset of 138 661 stars in the Dark Energy Spectroscopic Instrument (DESI) Milky Way Survey Year 1 Data Release. With an unsupervised clustering algorithm, HDBSCAN*, these substructures are independently identified in Integrals of Motion ($E_{\rm tot}$, $L_{\rm z}$, $\log {J_r}$, $\log {J_z}$) space and Galactocentric cylindrical velocity space ($V_{R}$, $V_{\phi }$, $V_{z}$). We associate all identified clusters with known nearby substructures (Helmi streams, M18-Cand10/MMH-1, Sequoia, Antaeus, and ED-2) previously reported in various studies. With metallicities precisely measured by DESI, we confirm that the Helmi streams, M18-Cand10, and ED-2 are chemically distinct from local halo stars. We have characterized the chemodynamic properties of each dynamic group, including their metallicity dispersions, to associate them with their progenitor types (globular cluster or dwarf galaxy). Our approach for searching substructures with HDBSCAN* reliably detects real substructures in the Galactic halo, suggesting that applying the same method can lead to the discovery of new substructures in future DESI data. With more stars from future DESI data releases and improved astrometry from the upcoming Gaia Data Release 4, we will have a more detailed blueprint of the Galactic halo, offering a significant improvement in our understanding of the formation and evolutionary history of the Milky Way Galaxy.

dynamics↗

NASA Software Cost Estimation Model: An Analogy Based Estimation Model

The cost estimation of software development activities is increasingly critical for large scale integrated projects such as those at DOD and NASA especially as the software systems become larger and more complex. As an example MSL (Mars Scientific Laboratory) developed at the Jet Propulsion Laboratory launched with over 2 million lines of code making it the largest robotic spacecraft ever flown (Based on the size of the software). Software development activities are also notorious for their cost growth, with NASA flight software averaging over 50% cost growth. All across the agency, estimators and analysts are increasingly being tasked to develop reliable cost estimates in support of program planning and execution. While there has been extensive work on improving parametric methods there is very little focus on the use of models based on analogy and clustering algorithms. In this paper we summarize our findings on effort/cost model estimation and model development based on ten years of software effort estimation research using data mining and machine learning methods to develop estimation models based on analogy and clustering. The NASA Software Cost Model performance is evaluated by comparing it to COCOMO II, linear regression, and K-­ nearest neighbor prediction model performance on the same data set.

Hihn, Jairus↗

Strong chemical tagging with APOGEE: 21 candidate star clusters that have dissolved across the Milky Way disc

ABSTRACT Chemically tagging groups of stars born in the same birth cluster is a major goal of spectroscopic surveys. To investigate the feasibility of such strong chemical tagging, we perform a blind chemical tagging experiment on abundances measured from APOGEE survey spectra. We apply a density-based clustering algorithm to the 8D chemical space defined by [Mg/Fe], [Al/Fe], [Si/Fe], [K/Fe], [Ti/Fe], [Mn/Fe], [Fe/H], and [Ni/Fe], abundances ratios which together span multiple nucleosynthetic channels. In a high-quality sample of 182 538 giant stars, we detect 21 candidate clusters with more than 15 members. Our candidate clusters are more chemically homogeneous than a population of non-member stars with similar [Mg/Fe] and [Fe/H], even in abundances not used for tagging. Group members are consistent with having the same age and fall along a single stellar-population track in log g versus Teff space. Each group’s members are distributed over multiple kpc, and the spread in their radial and azimuthal actions increases with age. We qualitatively reproduce this increase using N-body simulations of cluster dissolution in Galactic potentials that include transient winding spiral arms. Observing our candidate birth clusters with high-resolution spectroscopy in other wavebands to investigate their chemical homogeneity in other nucleosynthetic groups will be essential to confirming the efficacy of strong chemical tagging. Our initially spatially compact but now widely dispersed candidate clusters will provide novel limits on chemical evolution and orbital diffusion in the Galactic disc, and constraints on star formation in loosely bound groups.

Price-Jones, Natalie↗

Framework for a space shuttle main engine health monitoring system

A framework developed for a health management system (HMS) which is directed at improving the safety of operation of the Space Shuttle Main Engine (SSME) is summarized. An emphasis was placed on near term technology through requirements to use existing SSME instrumentation and to demonstrate the HMS during SSME ground tests within five years. The HMS framework was developed through an analysis of SSME failure modes, fault detection algorithms, sensor technologies, and hardware architectures. A key feature of the HMS framework design is that a clear path from the ground test system to a flight HMS was maintained. Fault detection techniques based on time series, nonlinear regression, and clustering algorithms were developed and demonstrated on data from SSME ground test failures. The fault detection algorithms exhibited 100 percent detection of faults, had an extremely low false alarm rate, and were robust to sensor loss. These algorithms were incorporated into a hierarchical decision making strategy for overall assessment of SSME health. A preliminary design for a hardware architecture capable of supporting real time operation of the HMS functions was developed. Utilizing modular, commercial off-the-shelf components produced a reliable low cost design with the flexibility to incorporate advances in algorithm and sensor technology as they become available.

Hawman, Michael W.↗

The use of unsupervised clustering as a classifier for LACIE MSS data

The author has identified the following significant results. This classification method appears to give accurate field center results and to give practical, statistically consistent and accurate estimates of crop proportions. The accuracy of this method is attributable to certain qualities of the particular clustering algorithm. These qualities are freedom from assumptions about Gaussian data, and the continual updating of distribution estimates, including updating the number of modes. This method is relatively tolerant of errors in the determination of crop type, as crop identity is used only for identifying clusters, and not for computing signatures.

Pentland, A. P.↗

Development of advanced acreage estimation methods

The development of an accurate and efficient algorithm for analyzing the structure of MSS data, the application of the Akaiki information criterion to mixture models, and a research plan to delineate some of the technical issues and associated tasks in the area of rice scene radiation characterization are discussed. The AMOEBA clustering algorithm is refined and documented.

Guseman, L. F., Jr.↗

Examination of Radiation Belt Dynamics During Substorm Clusters: Activity Drivers and Dependencies of Trapped Flux Enhancements

Here, dynamical variations of radiation belt trapped electron fluxes are examined to better understand the variability of enhancements linked to substorm clusters. Analysis is undertaken using the Substorm Onsets and Phases from Indices of the Electrojet substorm cluster algorithm for event detection. Observations from low earth orbit are complemented by additional measurements from medium earth orbit to allow a major expansion in the energy range considered, from medium energy energetic electrons up to ultra-relativistic electrons. The number of substorms identified inside a cluster does not depend strongly on solar wind drivers or geomagnetic indices either before, during, or after the cluster start time. Clusters of substorms linked to moderate (100 nT < AE ≤ 300 nT) or strong AE (AE ≥ 300 nT) disturbances are associated with radiation belt flux enhancements, including up to ultra-relativistic energies by the strongest substorms (as measured by strong southward Bz and high AE). These clusters reliably occur during times of high speed solar winds streams with associated increased magnetospheric convection. However, substorm clusters associated with quiet AE disturbances (AE ≤ 100 nT) lead to no significant chorus whistler mode intensity enhancements, or increases in energetic, relativistic, or ultra-relativistic electron flux in the outer radiation belts. In these cases the solar wind speed is low, and the geomagnetic Kp index indicates a lack of magnetospheric convection. Our study clearly indicates that clusters of substorms occurring outside of high speed wind streams are not by themselves sufficient to drive acceleration, which may be due to the lack of pre-cluster convection.

79 ASTRONOMY AND ASTROPHYSICS↗

Online Voltage Event Detection Using Synchrophasor Data with Structured Sparsity-Inducing Norms

This paper develops an accurate and computationally efficient data-driven framework to detect voltage events from PMU data streams. It develops an innovative Proximal Bilateral Random Projection (PBRP) algorithm to quickly decompose the PMU data matrix into a low-rank matrix, a row-sparse event-pattern matrix and a noise matrix. Here, the row-sparse pattern matrix significantly distinguishes events from normal behavior. These matrices are then fed into a clustering algorithm to separate voltage events from normal operating conditions. Large-scale numerical study results on real-world PMU data show that the proposed algorithm is computationally more efficient and achieves higher F scores than state-of-the-art benchmarks.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Enhancing Cluster Identification in Atom Probe Tomography Data Using Transfer Learning

Atom Probe Tomography (APT) is a powerful technique for visualizing the atomic-scale distribution of solutes in materials, but quantitative cluster analysis of APT datasets remains a challenge due to the need for subjective parameter selection in clustering algorithms. While distance-based and density-based methods such as HDBSCAN are widely used, their performance is highly sensitive to user-defined parameters, which undermines reproducibility and accuracy. This study proposes an image-based, deep learning-aided workflow for automating parameter selection and cluster detection in APT data analysis. By projecting 3D APT point clouds onto 2D planes, we leverage pretrained convolutional neural networks (ConvNeXt-Tiny and ResNet-50) through transfer learning to predict the number of clusters present in synthetic datasets. The output is used to guide K-means clustering and estimate HDBSCAN parameters, specifically minimum cluster size and minimum sample points. This approach reduces reliance on manual parameter tuning, improving consistency and scalability. The methodology demonstrates the feasibility of using image-based deep learning for interpreting complex spatial patterns in APT data, enabling faster and more objective analysis. The complete workflow and code are made publicly available to support reproducibility and future research.

Density-based clustering↗

Building Stock Segmentation Cluster Development: Technical Reference Document

The building stock in the United States (U.S.) varies significantly as a function of several macro variables such as: climate, building type, vintage, and density. These variables change across the U.S. and can also significantly impact energy usage of the individual buildings and overall stock. For example, the square foot density and building type varies by several orders of magnitude from Manhattan to the eastern plains of Colorado. The diversity in energy use of the building stock of different areas of the U.S. is significant, and as a result, analyses that require localized results need to consider the relevant geography and the current makeup of the building stock. This document discusses the development and implementation of a stock clustering algorithm that produces a technically rigorous, consistent, and repeatable collection of geographies which are used as the basis for localized analysis. This framework considers the impact of built environment density, diversity, and climate in creating groupings of counties that create a far more nuanced analysis framework than national averages. Clustering of counties together represents a similarity of building characteristics and climate zone.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗