Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “K-means method”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Assessing Machine Learning as a Tool to Explain Variance in Deployed Photovoltaic (PV) System Degradation

Degradation remains a large uncertainty in forecasting production for PV plants, creating significant risk for developers and financiers. This study aims to quantify the distribution and drivers of degradation across 10,000 PV systems deployed for distributed or utility generation by training a machine learning model to predict year-over-year degradation rates from metadata characteristics. A combination of K-Means clustering and random forest regressor were found to associate multiple metadata features as potential drivers of degradation, including module characteristics, system design, and climate features. From this, it is inferred that if machine learning is able to find complex patterns between metadata features and system performance loss, such methods can be employed to help developers and financiers make data-informed decisions when estimating long-term energy production forecasts in financial models.

Dunn, Jimmy C.↗

Joint Management and Optimization of Residential Natural Gas and Electricity Distribution Networks Coupled via Fuel Cells

The interesting properties of natural gas as well as the growing electric power demand worldwide have led to increasing attention to natural-gas-based distributed generation applications in electric distribution systems. This paper goes over the interdependency between a residential natural gas network and an electric distribution network that are coupled via fuel cells. The modeling of the gas network is introduced first, and then the algorithm for gas flow study is presented. The optimal placement and sizing of fuel cell based distributed generation systems are formulated to minimize the losses in both the gas and electric distribution networks, subject to their model constraints. In addition to this, in order to capture the probabilistic nature of the optimization problem under study, the K-means clustering algorithm is applied to the gas and electricity demands to determine hourly load states and their corresponding probabilities. Furthermore, simulation studies are carried out on an integrated system consisting of the IEEE 69-bus distribution feeder and a radial 27-node natural gas network to verify the developed optimization model and the proposed method.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Polynomial chaos expansions on principal geodesic Grassmannian submanifolds for surrogate modeling and uncertainty quantification

In this work we introduce a manifold learning-based surrogate modeling framework for uncertainty quantification in high-dimensional stochastic systems. Our first goal is to perform data mining on the available simulation data to identify a set of low-dimensional (latent) descriptors that efficiently parameterize the response of the high-dimensional computational model. To this end, we employ Principal Geodesic Analysis on the Grassmann manifold of the response to identify a set of disjoint principal geodesic submanifolds, of possibly different dimension, that captures the variation in the data. Since operations on the Grassmann require the data to be concentrated, we propose an adaptive algorithm based on Riemannian K-means and the minimization of the sample Fréchet variance on the Grassmann manifold to identify “local” principal geodesic submanifolds that represent different system behavior across the parameter space. Polynomial chaos expansion is then used to construct a mapping between the random input parameters and the projection of the response on these local principal geodesic submanifolds. Here, the method is demonstrated on four test cases, a toy-example that involves points on a hypersphere, a Lotka-Volterra dynamical system, a continuous-flow stirred-tank chemical reactor system, and a two-dimensional Rayleigh-Bénard convection problem.

42 ENGINEERING↗

Sub-10 nm Probing of Ferroelectricity in Heterogeneous Materials by Machine Learning Enabled Contact Kelvin Probe Force Microscopy

Reducing the dimensions of ferroelectric materials down to the nanoscale has strong implications on the ferroelectric polarization pattern and on the ability to switch the polarization. As the size of ferroelectric domains shrinks to the nanometer scale, the heterogeneity of the polarization pattern becomes increasingly pronounced, enabling a large variety of possible polar textures in nanocrystalline and nanocomposite materials. Critical to the understanding of fundamental physics of such materials and hence their applications in electronic nanodevices is the ability to investigate their ferroelectric polarization at the nanoscale in a nondestructive way. We show that contact Kelvin probe force microscopy (cKPFM) combined with a k-means response clustering algorithm enables to measure the ferroelectric response at a mapping resolution of 8 nm. In a BaTiO 3 thin film on silicon composed of tetragonal and hexagonal nanocrystals, we determine a nanoscale lateral distribution of discrete ferroelectric response clusters, fully consistent with the nanostructure determined by transmission electron microscopy. Moreover, we apply this data clustering method to the cKPFM responses measured at different temperatures, which allows us to follow the corresponding change in the polarization pattern as the Curie temperature is approached and across the phase transition. This work opens up perspectives for mapping complex ferroelectric polarization textures such as curled/swirled polar textures that can be stabilized in epitaxial heterostructures and more generally for mapping the polar domain distribution of any spatially highly heterogeneous ferroelectric materials.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Unsupervised learning-enabled pulsed infrared thermographic microscopy of subsurface defects in stainless steel

Metallic structures produced with laser powder bed fusion (LPBF) additive manufacturing method (AM) frequently contain microscopic porosity defects, with typical approximate size distribution from one to 100 microns. Presence of such defects could lead to premature failure of the structure. In principle, structural integrity assessment of LPBF metals can be accomplished with nondestructive evaluation (NDE). Pulsed infrared thermography (PIT) is a non-contact, one-sided NDE method that allows for imaging of internal defects in arbitrary size and shape metallic structures using heat transfer. PIT imaging is performed using compact instrumentation consisting of a flash lamp for deposition of a heat pulse, and a fast frame infrared (IR) camera for measuring surface temperature transients. However, limitations of imaging resolution with PIT include blurring due to heat diffusion, sensitivity limit of the IR camera. We demonstrate enhancement of PIT imaging capability with unsupervised learning (UL), which enables PIT microscopy of subsurface defects in high strength corrosion resistant stainless steel 316 alloy. PIT images were processed with UL spatial–temporal separation-based clustering segmentation (STSCS) algorithm, refined by morphology image processing methods to enhance visibility of defects. The STSCS algorithm starts with wavelet decomposition to spatially de-noise thermograms, followed by UL principal component analysis (PCA), fine-tuning optimization, and neural learning-based independent component analysis (ICA) algorithms to temporally compress de-noised thermograms. The compressed thermograms were further processed with UL-based graph thresholding K-means clustering algorithm for defects segmentation. The STSCS algorithm also includes online learning feature for efficient re-training of the model with new data. For this study, metallic specimens with calibrated microscopic flat bottom hole defects, with diameters in the range from 203 to 76 µm, were produced using electro discharge machining (EDM) drilling. While the raw thermograms do not show any material defects, using STSCS algorithm to process PIT images reveals defects as small as 101 µm in diameter. To the best of our knowledge, this is the smallest reported size of a sub-surface defect in a metal imaged with PIT, which demonstrates the PIT capability of detecting defects in the size range relevant to quality control requirements of LPBF-printed high-strength metals.

36 MATERIALS SCIENCE↗

Exploratory analysis of machine learning techniques in the Nevada geothermal play fairway analysis

Play fairway analysis (PFA) is commonly used to generate geothermal potential maps and guide exploration studies, with a particular focus on locating and characterizing blind geothermal systems. This study evaluates the application of machine learning techniques to PFA in the Great Basin region of Nevada. Following the evaluation of various techniques, we identified two approaches to PFA that produced promising results, 1) supervised Bayesian probabilistic neural networks to generate geothermal potential maps with confidence intervals, and 2) unsupervised principal component analysis paired with k-means clustering to generate both cluster maps to help identify spatial patterns, as well as new combined feature inputs. We applied these techniques to perform a comparative analysis between two principal sets of geological and geophysical features related to permeability and heat and a set of positive (known geothermal resources) and negative training sites (known drill sites with unsuitable geothermal conditions). We found that these methods constrain previously unrecognized feature controls on geothermal favorability, many of which are spatially organized within the extent of cluster groups and the major structural-hydrologic domains of the study area. Furthermore, we utilized exploratory unsupervised modeling to highlight spatial relationships between input data and predictive output results of our supervised modeling. As a result, we demonstrate how our models compare to the previous Nevada PFA and how the rapid insights these machine learning techniques offer may support future assessments of both known and undiscovered blind geothermal systems in the Great Basin region of Nevada and beyond.

15 GEOTHERMAL ENERGY↗

A Satellite-Based Estimate of Convective Vertical Velocity and Convective Mass Flux: Global Survey and Comparison with Radar Wind Profiler Observations

Convective vertical velocity (w c ) and convective mass flux (M c ) lie at the heart of GCM cumulus parameterizations, but few observations of these critical parameters are available. In this paper, we develop and evaluate a novel, satellite-based method for estimating profiles of w c and M c . Here, comparisons with collocated ground-based radar wind profiler (RWP) observations show that satellite estimated median w c is slightly greater than the RWP estimates, but they show solid agreement when compared at the 95th percentiles (intense updrafts). RWP-derived and satellite estimated M c are broadly comparable in the lower and middle troposphere, with some differences in the upper troposphere due to differences in convective core sampling. A k-means cluster analysis of multiple years of w c data shows that convective characteristics are distinctly different among extratropical convection, tropical land convection, and tropical oceanic convection. Tropical land convection is significantly more intense and more variable than the oceanic counterpart.

54 ENVIRONMENTAL SCIENCES↗

Enhancing Cluster Identification in Atom Probe Tomography Data Using Transfer Learning

Atom Probe Tomography (APT) is a powerful technique for visualizing the atomic-scale distribution of solutes in materials, but quantitative cluster analysis of APT datasets remains a challenge due to the need for subjective parameter selection in clustering algorithms. While distance-based and density-based methods such as HDBSCAN are widely used, their performance is highly sensitive to user-defined parameters, which undermines reproducibility and accuracy. This study proposes an image-based, deep learning-aided workflow for automating parameter selection and cluster detection in APT data analysis. By projecting 3D APT point clouds onto 2D planes, we leverage pretrained convolutional neural networks (ConvNeXt-Tiny and ResNet-50) through transfer learning to predict the number of clusters present in synthetic datasets. The output is used to guide K-means clustering and estimate HDBSCAN parameters, specifically minimum cluster size and minimum sample points. This approach reduces reliance on manual parameter tuning, improving consistency and scalability. The methodology demonstrates the feasibility of using image-based deep learning for interpreting complex spatial patterns in APT data, enabling faster and more objective analysis. The complete workflow and code are made publicly available to support reproducibility and future research.

Density-based clustering↗

A generalizable machine learning-assisted fast Fourier transform algorithm to simulate the large strain phenomena in polycrystalline materials

Machine learning methods have shown initial promise in constitutive modeling for single crystals or homogenized polycrystals, delivering notable computational efficiency. However, existing machine learning-based constitutive models often lack generalizability, limiting their application across diverse boundary value problems. This study introduces a thermodynamics-informed artificial neural network model to accelerate rate-tangent crystal plasticity fast Fourier transform simulations for cross-scale deformation behaviors of polycrystals under complex loading. Our model integrates microstructural variability and local interactions effectively. To address local effects in each grain, we employ K-means clustering to group Gauss points within the microstructure into clusters assumed to be in similar mechanical states. This approach, based on self-clustering analysis, extends model scope from macroscopic stress response to the granular level, capturing mechanical responses and orientation evolution across grains. This reduces the number of nonlinear problems to solve, with cluster responses propagated throughout each group. The thermodynamics-based artificial neural network-extracted features are further processed using local material state clusters to account for history-dependent deformation and evolving microstructures. Additionally, representative volume element simulations with rate-tangent crystal plasticity fast Fourier transform provide reliable datasets for model training. The proposed model demonstrates high efficiency, accuracy, self-consistency, and enhanced generalizability in predicting strain–stress responses and orientation evolution at both individual grain and aggregate scales under complex loading conditions, such as biaxial tension and arbitrary loading scenarios.

36 MATERIALS SCIENCE↗

A Hydrogen Load Modeling Method for Integrated Hydrogen Energy System Planning

The integrated hydrogen energy system incorporates hydrogen energy into the power grid, which has been recognized as a promising option for reaching a 100% renewable electricity supply. It can make a profit because the hydrogen produced can be sold as fuel or used to generate electricity for grid services. In this paper, we develop a planning model for the integrated hydrogen energy system that considers the uncertainty of the load demand, the renewable energy generation, and the market prices. To calculate the hydrogen load, we simulate the refueling operations at a hydrogen fueling station over the course of one day and generate representative load profiles with K-means clustering. Moreover, the long-term profitability of the integrated system under both current and future conditions is validated in 10-year planning results.

grid service↗

Correlative piezoresponse and micro-Raman imaging of CuInP 2 S 6 –In 4/3 P 2 S 6 flakes unravels phase-specific phononic fingerprint via unsupervised learning

Characterizing the novel properties of layered van der Waals materials is key for their application in functional devices. A better understanding of this type of material requires correlative imaging of diverse nanoscale material properties. Within this class of materials, CuInP 2 S 6 (CIPS) has received a significant degree of interest due to its ionically mediated room temperature ferroelectricity. Moreover, it is possible to form stable self-assembled heterostructures of ferroelectric CuInP 2 S 6 (CIPS) and non-ferroelectric (i.e., lacking Cu) In 4/3 P 2 S 6 (IPS) phases, by controlling the targeted composition and kinetics of synthesis. In this work, we present a correlative nanometric imaging study of the phononic modes and piezoelectricity of the phase-separated thin heteroepitaxial CIPS/IPS flakes. Here, we show that it is possible to isolate the different phononic modes of the two phases by spatially correlating them with their distinct ferroelectric behavior. The coupling of our experimental data with unsupervised learning statistical methods enables unraveling specific Raman peaks that are characteristic of each chemical phase (CIPS and IPS) present in the composite sample, discarding the less significant ones.

correlative microscopy↗

Crash Risks Evaluation of Urban Expressways: A Case Study in Shanghai

We report that proactive traffic safety management systems can reduce crashes by identifying crash precursors, evaluating real-time crash risks, and implementing suitable interventions. The basic prerequisite for developing such a system is to propose a reliable crash risk evaluation model that takes real-time traffic flow data as input. Previous studies have primarily focused on real-time crash prediction using some statistical or machine-learning methods. However, further quantitative evaluation and classification of crash risks have been ignored. In this study, we conduct a systematic crash risk evaluation workflow, including crash risk prediction, crash risk quantification, and crash risk classification. Specifically, the crash risk prediction using an extended logit model is proposed, from which CAS, CSD, UAS, DAS, DTV are identified to be contributing factors of crash risks. Then a crash risk quantification model based on the parameter evaluation of the extended logit model is developed. The crash risks of urban expressways and their spatial-temporal evolution trends are quantified. Finally, the crash risks are classified into high crash risk level, moderate crash risk level, and low crash risk level by the k-means cluster algorithm. Then the threshold boundaries of different crash risk levels are determined. The research results provide a proactive guidance for traffic safety management of urban expressways.

33 ADVANCED PROPULSION SYSTEMS↗

Topological data analysis of task-based fMRI data from experiments on schizophrenia

We use methods from computational algebraic topology to study functional brain networks, in which nodes represent brain regions and weighted edges represent similarity of fMRI time series from each region. With these tools, which allow one to characterize topological invariants such as loops in high-dimensional data, we are able to gain understanding into low-dimensional structures in networks in a way that complements traditional approaches based on pairwise interactions. In the present paper, we analyze networks constructed from task-based fMRI data from schizophrenia patients, healthy controls, and healthy siblings of schizophrenia patients using persistent homology, which allows us to explore the persistence of topological structures such as loops at different scales in the networks. We use persistence landscapes, persistence images, and Betti curves to create output summaries from our persistent-homology calculations, and we study the persistence landscapes and images using k-means clustering and community detection. Based on our analysis of persistence landscapes, we find that the members of the sibling cohort have topological features (specifically, their 1-dimensional loops) that are distinct from the other two cohorts. From the persistence images, we are able to distinguish all three subject groups and to determine the brain regions in the loops (with four or more edges) that allow us to make these distinctions.

60 APPLIED LIFE SCIENCES↗

Unsupervised Machine Learning for Exploratory Data Analysis of Exoplanet Transmission Spectra

Abstract Transit spectroscopy is a powerful tool for decoding the chemical compositions of the atmospheres of extrasolar planets. In this paper, we focus on unsupervised techniques for analyzing spectral data from transiting exoplanets. After cleaning and validating the data, we demonstrate methods for: (i) initial exploratory data analysis, based on summary statistics (estimates of location and variability); (ii) exploring and quantifying the existing correlations in the data; (iii) preprocessing and linearly transforming the data to its principal components; (iv) dimensionality reduction and manifold learning; (v) clustering and anomaly detection; and (vi) visualization and interpretation of the data. To illustrate the proposed unsupervised methodology, we use a well-known public benchmark data set of synthetic transit spectra. We show that there is a high degree of correlation in the spectral data, which calls for appropriate low-dimensional representations. We explore a number of different techniques for such dimensionality reduction and identify several suitable options in terms of summary statistics, principal components, etc. We uncover interesting structures in the principal component basis, namely well-defined branches corresponding to different chemical regimes of the underlying atmospheres. We demonstrate that those branches can be successfully recovered with a K-means clustering algorithm in a fully unsupervised fashion. We advocate for lower-dimensional representations of the spectroscopic data in terms of the main principal components, in order to reveal the existing structure in the data and quickly characterize the chemical class of a planet.

Matchev, Konstantin T. (ORCID:0000000341829096)↗

Improving Medication Regimen Recommendation for Parkinson’s Disease Using Sensor Technology

Parkinson’s disease medication treatment planning is generally based on subjective data obtained through clinical, physician-patient interactions. The Personal KinetiGraph™ (PKG) and similar wearable sensors have shown promise in enabling objective, continuous remote health monitoring for Parkinson’s patients. In this proof-of-concept study, we propose to use objective sensor data from the PKG and apply machine learning to cluster patients based on levodopa regimens and response. The resulting clusters are then used to enhance treatment planning by providing improved initial treatment estimates to supplement a physician’s initial assessment. We apply k-means clustering to a dataset of within-subject Parkinson’s medication changes—clinically assessed by the MDS-Unified Parkinson’s Disease Rating Scale-III (MDS-UPDRS-III) and the PKG sensor for movement staging. A random forest classification model was then used to predict patients’ cluster allocation based on their respective demographic information, MDS-UPDRS-III scores, and PKG time-series data. Clinically relevant clusters were partitioned by levodopa dose, medication administration frequency, and total levodopa equivalent daily dose—with the PKG providing similar symptomatic assessments to physician MDS-UPDRS-III scores. A random forest classifier trained on demographic information, MDS-UPDRS-III scores, and PKG time-series data was able to accurately classify subjects of the two most demographically similar clusters with an accuracy of 86.9%, an F1 score of 90.7%, and an AUC of 0.871. A model that relied solely on demographic information and PKG time-series data provided the next best performance with an accuracy of 83.8%, an F1 score of 88.5%, and an AUC of 0.831, hence further enabling fully remote assessments. These computational methods demonstrate the feasibility of using sensor-based data to cluster patients based on their medication responses with further potential to assist with medication recommendations.

59 BASIC BIOLOGICAL SCIENCES↗

Web-based wide-area monitoring platform for ringdown and clustering analytics in power systems

This paper introduces an open-source research platform for monitoring the Mexican interconnected power grid, allowing real-time processing and information extraction of the grid’s dynamic condition. Moreover, the platform is a Python-based development that embeds different ringdown and clustering analytics tools. In the case of ringdown analysis, the modal information can be extracted using some of the most known algorithms, i.e., Prony analysis, eigensystem realization algorithm (ERA), and matrix pencil (MP). For clustering analysis, the coherent behaviour of generator and non-generator buses is provided by applying recent state-of-the-art techniques such as affinity propagation, K-means, hierarchical agglomerative clustering, and typicality data analysis. The results of up to 93 PMUs show that this open-source platform suits researchers’ and engineers’ power system dynamic analysis requirements.

Clustering↗

Machine Learning Based Resilience Testing of an Address Randomization Cyber Defense

Moving target defenses (MTDs) are widely used as an active defense strategy for thwarting cyberattacks on cyber-physical systems by increasing diversity of software and network paths. Recently, machine Learning (ML) and deep Learning (DL) models have been demonstrated to defeat some of the cyber defenses by learning attack detection patterns and defense strategies. It raises concerns about the susceptibility of MTD to ML and DL methods. Here, in this article, we analyze the effectiveness of ML and DL models when it comes to deciphering MTD methods and ultimately evade MTD-based protections in real-time systems. Specifically, we consider a MTD algorithm that periodically randomizes address assignments within the MIL-STD-1553 protocol—a military standard serial data bus. Two ML and DL-based tasks are performed on MIL-STD-1553 protocol to measure the effectiveness of the learning models in deciphering the MTD algorithm: 1) determining whether there is an address assignments change i.e., whether the given system employs a MTD protocol and if it does 2) predicting the future address assignments. The supervised learning models (random forest and k-nearest neighbors) effectively detected the address assignment changes and classified whether the given system is equipped with a specified MTD protocol. On the other hand, the unsupervised learning model (K-means) was significantly less effective. The DL model (long short-term memory) was able to predict the future addresses with varied effectiveness based on MTD algorithm's settings.

45 MILITARY TECHNOLOGY, WEAPONRY, AND NATIONAL DEF↗

Machine Learning-Driven Reliability Estimation of PV Inverters Considering Alert-Ambient Variability

Weather-induced spatio-temporal degradation limits outdoor PV inverter lifetime and reliability, necessitating advanced data analysis. This study employs a top-down, data-driven approach utilizing multiple machine learning (ML) algorithms to estimate inverter reliability in a 1.4 MW PV power plant, considering factors such as irradiance, humidity, temperature, time of day, and weather conditions. An extensive alert dataset from 17 identical inverters, including alert types, propagation, and frequency, reveals significant correlations with environmental factors and inverter output power, enabling the construction of a performance reliability model. Dual-stage supervised-ML models are evaluated for accuracy, with the ‘classification-regression’ model by an artificial neural network (ANN) tested on the averaged “Alert-Ambient” dataset, which is outperformed by ‘clustering-regression’ models using random forest (RF) and K-Nearest Neighbors (KNN) on individual inverter datasets. K-means clustering applies principal component analysis to reduce dimensions, achieving improved accuracy beyond the 80% achieved by ANN on the averaged dataset. Second-stage regression estimates inverter reliability with a mean square error of 0.0195 on the averaged dataset and as low as 0.002 on individual inverter datasets using RF. Furthermore, these findings highlight the method's suitability for estimating PV inverter output reliability under ambient conditions, essential for digital twin development and related applications.

14 SOLAR ENERGY↗