Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “k mean”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Spatio-Temporal Surrogates for Interaction of a Jet with High Explosives: Part II - Clustering Extremely High-Dimensional Grid-Based Data

Building an accurate surrogate model for the spatio-temporal outputs of a computer simulation is a challenging task. A simple approach to improve the accuracy of the surrogate is to cluster the outputs based on similarity and build a separate surrogate model for each cluster. This clustering is relatively straightforward when the output at each time step is of moderate size. However, when the spatial domain is represented by a large number of grid points, numbering in the millions, the clustering of the data becomes more challenging. In this report, we consider output data from simulations of a jet interacting with high explosives. These data are available on spatial domains of different sizes, at grid points that vary in their spatial coordinates, and in a format that distributes the output across multiple files at each time step of the simulation. We first describe how we bring these data into a consistent format prior to clustering. Borrowing the idea of random projections from data mining, we reduce the dimension of our data by a factor of thousand, making it possible to use the iterative k-means method for clustering. We show how we can use the randomness of both the random projections, and the choice of initial centroids in k-means clustering, to determine the number of clusters in our data set. Our approach makes clustering of extremely high dimensional data tractable, generating meaningful cluster assignments for our problem, despite the approximation introduced in the random projections.

97 MATHEMATICS AND COMPUTING↗

Machine-learning identification of the variability of mean velocity and turbulence intensity for wakes generated by onshore wind turbines: Cluster analysis of wind LiDAR measurements

Light detection and ranging (LiDAR) measurements of isolated wakes generated by wind turbines installed at an onshore wind farm are leveraged to characterize the variability of the wake mean velocity and turbulence intensity during typical operations, which encompass a breadth of atmospheric stability regimes and rotor thrust coefficients. The LiDAR measurements are clustered through the k-means algorithm, which enables identifying the most representative realizations of wind turbine wakes while avoiding the imposition of thresholds for the various wind and turbine parameters. Considering the large number of LiDAR samples collected to probe the wake velocity field, the dimensionality of the experimental dataset is reduced by projecting the LiDAR data on an intelligently truncated basis obtained with the proper orthogonal decomposition (POD). The coefficients of only five physics-informed POD modes are then injected in the k-means algorithm for clustering the LiDAR dataset. The analysis of the clustered LiDAR data and the associated supervisory control and data acquisition and meteorological data enables the study of the variability of the wake velocity deficit, wake extent, and wake-added turbulence intensity for different thrust coefficients of the turbine rotor and regimes of atmospheric stability. Furthermore, the cluster analysis of the LiDAR data allows for the identification of systematic off-design operations with a certain yaw misalignment of the turbine rotor with the mean wind direction.

17 WIND ENERGY↗

Discriminating Quantum States with Quantum Machine Learning

Quantum machine learning (QML) algorithms have obtained great relevance in the machine learning (ML) field due to the promise of quantum speedups when performing basic linear algebra subroutines (BLAS), a fundamental element in most ML algorithms. By making use of BLAS operations, we propose, implement and analyze a quantum k-means (qk-means) algorithm with a low time complexity of O(NKlog(D)I/C) to apply it to the fundamental problem of discriminating quantum states at readout. Discriminating quantum states allows the identification of quantum states |0⟩ and |1⟩ from low-level in-phase and quadrature signal (IQ) data, and can be done using custom ML models. In order to reduce dependency on a classical computer, we use the qk-means to perform state discrimination on the IBMQ Bogota device and managed to find assignment fidelities of up to 98.7% that were only marginally lower than that of the k-means algorithm. We also performed a cross-talk benchmark on the quantum device by applying both algorithms to perform state discrimination on a combination of quantum states and using Pearson Correlation coefficients and assignment fidelities of discrimination results to conclude on the presence of cross-talk on qubits. Evidence shows cross-talk in the (1, 2) and (2, 3) neighboring qubit couples for the analyzed device.

Quiroga, David↗

Fast Image Texture Classification Using Decision Trees

Texture analysis would permit improved autonomous, onboard science data interpretation for adaptive navigation, sampling, and downlink decisions. These analyses would assist with terrain analysis and instrument placement in both macroscopic and microscopic image data products. Unfortunately, most state-of-the-art texture analysis demands computationally expensive convolutions of filters involving many floating-point operations. This makes them infeasible for radiation- hardened computers and spaceflight hardware. A new method approximates traditional texture classification of each image pixel with a fast decision-tree classifier. The classifier uses image features derived from simple filtering operations involving integer arithmetic. The texture analysis method is therefore amenable to implementation on FPGA (field-programmable gate array) hardware. Image features based on the "integral image" transform produce descriptive and efficient texture descriptors. Training the decision tree on a set of training data yields a classification scheme that produces reasonable approximations of optimal "texton" analysis at a fraction of the computational cost. A decision-tree learning algorithm employing the traditional k-means criterion of inter-cluster variance is used to learn tree structure from training data. The result is an efficient and accurate summary of surface morphology in images. This work is an evolutionary advance that unites several previous algorithms (k-means clustering, integral images, decision trees) and applies them to a new problem domain (morphology analysis for autonomous science during remote exploration). Advantages include order-of-magnitude improvements in runtime, feasibility for FPGA hardware, and significant improvements in texture classification accuracy.

Thompson, David R.↗

QUBO formulations for training machine learning models

Abstract Training machine learning models on classical computers is usually a time and compute intensive process. With Moore’s law nearing its inevitable end and an ever-increasing demand for large-scale data analysis using machine learning, we must leverage non-conventional computing paradigms like quantum computing to train machine learning models efficiently. Adiabatic quantum computers can approximately solve NP-hard problems, such as the quadratic unconstrained binary optimization (QUBO), faster than classical computers. Since many machine learning problems are also NP-hard, we believe adiabatic quantum computers might be instrumental in training machine learning models efficiently in the post Moore’s law era. In order to solve problems on adiabatic quantum computers, they must be formulated as QUBO problems, which is very challenging. In this paper, we formulate the training problems of three machine learning models—linear regression, support vector machine (SVM) and balanced k-means clustering—as QUBO problems, making them conducive to be trained on adiabatic quantum computers. We also analyze the computational complexities of our formulations and compare them to corresponding state-of-the-art classical approaches. We show that the time and space complexities of our formulations are better (in case of SVM and balanced k-means clustering) or equivalent (in case of linear regression) to their classical counterparts.

97 MATHEMATICS AND COMPUTING↗

Predicting Dynamic-to-Static Correction Factor from Petrophysical Data and Chemostratigraphy using Unsupervised Machine Learning

Estimating static mechanical properties of stratigraphic layers is critical for optimizing subsurface engineering applications. To estimate dynamic-to-static correction factor F ds (static-to-dynamic Young’s modulus ratio) across the Caney shale interval in Oklahoma, USA, we integrated triaxial test measurements and petrophysical data, including well logs and X-ray fluorescence (XRF) using unsupervised machine learning (ML). We used a novel workflow that includes principal component analysis (PCA) to reduce data set dimensionality of well logs and XRF data sets—both separately and combined—creating three scenarios, and later applied inverse distance weighting (IDW) to derive F ds profiles for these scenarios. Furthermore, we applied K-means clustering on each scenario to predict depositional facies, and built a stiffness zonation profile through chemostratigraphic analysis of the terrigenous elements to validate the predicted F ds . The predicted F ds profile from each scenario using the PCA-IDW method was compared with the constant F ds approach from our previous study by calculating the root mean square error (RMSE). The combined data sets scenario yielded the lowest RMSE value of 0.113, while the RMSE values for the well logs and XRF scenarios were 0.131 and 0.129, respectively. In addition, the predicted F ds from the XRF scenario well-matched the stiffness zonation from the chemostratigraphic analysis that was built using the optimized K-means clustering of nine clusters for that scenario. These methods and findings offer a valuable tool for refining lithological classification and improving the F ds profile, potentially enhancing drilling and stimulation strategies for subsurface energy engineering applications.

clastic rock↗

Comparison of wheat classification accuracy using different classifiers of the image-100 system

Classification results using single-cell and multi-cell signature acquisition options, a point-by-point Gaussian maximum-likelihood classifier, and K-means clustering of the Image-100 system are presented. Conclusions reached are that: a better indication of correct classification can be provided by using a test area which contains various cover types of the study area; classification accuracy should be evaluated considering both the percentages of correct classification and error of commission; supervised classification approaches are better than K-means clustering; Gaussian distribution maximum likelihood classifier is better than Single-cell and Multi-cell Signature Acquisition Options of the Image-100 system; and in order to obtain a high classification accuracy in a large and heterogeneous crop area, using Gaussian maximum-likelihood classifier, homogeneous spectral subclasses of the study crop should be created to derive training statistics.

Dejesusparada, N.↗

Automated characterization of spatial and dynamical heterogeneity in supercooled liquids via implementation of machine learning

Abstract A computational approach by an implementation of the principle component analysis (PCA) with K -means and Gaussian mixture (GM) clustering methods from machine learning algorithms to identify structural and dynamical heterogeneities of supercooled liquids is developed. In this method, a collection of the average weighted coordination numbers ( W C N s ‾ ) of particles calculated from particles’ positions are used as an order parameter to build a low-dimensional representation of feature (structural) space for K -means clustering to sort the particles in the system into few meso-states using PCA. Nano-domains or aggregated clusters are also formed in configurational (real) space from a direct mapping using associated meso-states’ particle identities with some misclassified interfacial particles. These classification uncertainties can be improved by a co-learning strategy which utilizes the probabilistic GM clustering and the information transfer between the structural space and configurational space iteratively until convergence. A final classification of meso-states in structural space and domains in configurational space are stable over long times and measured to have dynamical heterogeneities. Armed with such a classification protocol, various studies over the thermodynamic and dynamical properties of these domains indicate that the observed heterogeneity is the result of liquid–liquid phase separation after quenching to a supercooled state.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

Benchmarking image processing techniques for porosity measurement in polymer additive manufacturing: Review and experimental analysis

An image processing workflow is proposed for porosity measurement in polymer additive manufacturing. Various techniques, including global and local thresholding, region growing, and K-means clustering, were applied to microscopic images of carbon fiber reinforced acrylonitrile butadiene styrene (CF-ABS) and benchmarked for their ability to accurately measure porosity. Global methods included Otsu, minimum error, iterative, and entropy-based thresholding, while local methods included Niblack, Bernsen, Sauvola, and Bradley-Roth algorithms. Artificial uneven illumination was introduced to test local adaptive thresholds. Results showed significant differences in porosity values across methods. Otsu, region growing, and K-means clustering excelled under uniform illumination, while Sauvola and Bradley-Roth performed better with uneven illumination. Comparison with X-ray computed tomography (XCT) revealed slightly lower porosity values (2.55 %) than optimized methods (2.73–2.79 %) due to XCT's lower resolution excluding smaller pores. While XCT offers finer pore detection, it limits sample volume and underestimates porosity due to spatial variation. Validation using artificial grayscale images with 5 % porosity confirmed that Otsu, Bradley-Roth, region growing, and Sauvola algorithms produced accurate results. Although tested on a single material system, these methods can be adapted to others with optimization. In conclusion, given XCT's high computational and time costs, this study highlights suitable image processing techniques as cost-effective alternatives for porosity analysis in polymer composites.

Additive manufacturing↗

Evaluating Grid Strength under Uncertain Renewable Generation

The increasing displacement of synchronous generators with renewable resources such as wind and solar via power electronic interfaces causes a reduction in short-circuit strength and weak grid issues. The variation and uncertainty of renewable energy increase challenges for identifying weak grid conditions. This paper proposes an efficient method to analyze the impact of uncertain renewable energy on grid strength. The proposed method uses the probabilistic collocation method (PCM) to approximate the results of grid strength assessment under uncertain renewable generation, in order to reduce computational burden without compromising result accuracy when compared with traditional Monte Carlo simulation (MCS). To improve the accuracy of the approximation results, the proposed method integrates the K-means clustering technique with PCM to select the approximation samples of input variables. The efficacy of the proposed method is demonstrated by comparison with MCS on the modified IEEE 9-bus system and modified IEEE 39-bus system with multiple renewable generators.

grid strength↗

The Empirical Effect of Fleet Optimization on Synchronization and Rebound Effects in Heat Pump Water Heaters

Demand response is a growing concept in light of the internet of things and an increasing need for grid flexibility. Water heaters are one of the preferred devices for providing demand response for grid services and peak management due to their capability to store energy. The efficient use of water heaters for demand response requires consideration of the associated load effects such as synchronization of device schedules and rebound effect. These effects present a significant challenge. Despite the importance of the mentioned effects for water heater queuing and scheduling, there has been no effort to quantify and empirically validate their impact. This study attempts to address this gap by offering two methods - Ward clustering and Euclidean K-means - to evaluate the extent of synchronization in a fleet of 42 water heaters in Atlanta, GA. Using the aforementioned methods on the measured data, we find evidence of convergence of water heater loads as a result of optimization compared to an idle period and analyzed their impact.

demand response↗

Identifying Climate Patterns Using Clustering Autoencoder Techniques

Abstract The complexity of growing spatiotemporal resolution of climate simulations produces a variety of climate patterns under different projection scenarios. This paper proposes a new data-driven climate classification workflow via an unsupervised deep learning technique that can dimensionally reduce the vast volume of spatiotemporal numerical climate projection data into a compact representation. We aim to identify distinct zones that capture multiple climate variables as well as their future changes under different climate change scenarios. Our approach leverages convolutional autoencoders combined with k -means clustering (standard autoencoder) and online clustering based on the Sinkhorn–Knopp algorithm (clustering autoencoder) across the conterminous United States (CONUS) to capture unique climate patterns in a data-driven fashion from the Geophysical Fluid Dynamics Laboratory Earth System Model with GOLD component (GFDL-ESM2G). The developed approach compresses 70 years of GFDL-ESM2G simulation at 0.125° spatial resolution across the CONUS under multiple warming scenarios to a lower-dimensional space by a factor of 660 000 and then tested on 150 years of GFDL-ESM2G simulation data. The results show that five climate clusters capture physically reasonable and spatially stable climatological patterns matched to known climate classes defined by human experts. Results also show that using a clustering autoencoder can reduce the computational time for clustering by up to 9.2 times when compared to using a standard autoencoder. Our five unique climate patterns resulting from the deep learning–based clustering of the lower-dimensional space thereby enable us to provide insights on hydrometeorology and its spatial heterogeneity across the conterminous United States immediately without downloading large climate datasets. Significance Statement This paper presents a data-driven climate classification approach using unsupervised deep learning to dimensionally reduce climate model outputs and to identify distinct climate regions for their future changes. Our approach compresses climate information for 70 years of Geophysical Fluid Dynamics Laboratory Earth System Model data across the conterminous United States (CONUS) at 0.125° spatial resolution. The results reveal that five climate clusters capture reasonable and stable climatological patterns matched to known climate patterns. The embedded clustering process in deep learning provides ×9.2 times faster execution than the k -means clustering technique. These results give us insight about climate spatial patterns and heterogeneity of hydrological patterns across the conterminous United States without downloading large climate datasets.

Kurihana, Takuya↗

Programs and Code for Geothermal Exploration Artificial Intelligence

The scripts below are used to run the Geothermal Exploration Artificial Intelligence developed within the "Detection of Potential Geothermal Exploration Sites from Hyperspectral Images via Deep Learning" project. It includes all scripts for pre-processing and processing, including: - Land Surface Temperature K-Means classifier - Labeling AI using Self Organizing Maps (SOM) - Post-processing for Permanent Scatterer InSAR (PSInSAR) analysis with SOM - Mineral marker summarizing - Artificial Intelligence (AI) Data splitting: creates data set from a single raster file - Artificial Intelligence Model: creates AI from a single data set, after splitting in Train, Validation and Test subsets - AI Mapper: creates a classification map based on a raster file

15 GEOTHERMAL ENERGY↗

Uncovering acoustic signatures of pore formation in laser powder bed fusion

Abstract We present a machine learning workflow to discover signatures in acoustic measurements that can be utilized to create a low-dimensional model to accurately predict the location of keyhole pores formed during additive manufacturing processes. Acoustic measurements were sampled at 100 kHz during single-layer laser powder bed fusion (LPBF) experiments, and spatio-temporal registration of pore locations was obtained from post-build radiography. Power spectral density (PSD) estimates of the acoustic data were then decomposed using non-negative matrix factorization with custom $$\varvec{k}$$ k -means clustering (NMF $$\varvec{k}$$ k ) to learn the underlying spectral patterns associated with pore formation. NMF $$\varvec{k}$$ k returned a library of basis signals and matching coefficients to blindly construct a feature space based on the PSD estimates in an optimized fashion. Moreover, the NMF $$\varvec{k}$$ k decomposition led to the development of computationally inexpensive machine learning models which are capable of quickly and accurately identifying pore formation with classification accuracy of supervised and unsupervised label learning greater than 95% and 90%, respectively. The intrinsic data compression of NMF k , the relatively light computational cost of the machine learning workflow, and the high classification accuracy makes the proposed workflow an attractive candidate for edge computing toward in-situ keyhole pore prediction in LPBF.

36 MATERIALS SCIENCE↗

Spatial Distribution and Clustering of Glycosaminoglycans in Electrospun Gelatin-Based Scaffolds

The extracellular matrix (ECM) is comprised of components like collagen, elastin, and glycosaminoglycans (GAGs). Electrospun fibrous scaffolds are designed to replicate the form and composition of the native ECM, often requiring blending of various ECM component mimics to enhance cellular responses. However, the spatial distribution of blended components within these fibers remains unclear. This study investigates the spatial distribution of chondroitin sulfate-C (CSC) in electrospun gelatin-based scaffolds. scanning electron microscopy (SEM), attenuated reflectance-Fourier transform infrared (ATR-FTIR) spectroscopy, X-ray photoelectron spectroscopy (XPS), and Time-of-flight Secondary Ion Mass Spectrometry (ToF-SIMS) were applied for surface and subsurface chemical characterization of the fibrous scaffolds. SEM confirmed a fibrous morphology, while ATR-FTIR and XPS analyses indicated the presence of CSC through the identification of sulfate groups. ToF-SIMS imaging, alongside K-means clustering and Ripley’s K function, revealed a nonuniform CSC distribution with higher concentrations at the top layer of the scaffold. This study demonstrates that CSC presentation at the fiber surface varies with depth and differs from bulk incorporation while reveals nanoscale clustering and spatial heterogeneity at both the surface and subsurface of electrospun gelatin fibers. These findings define an underexplored design consideration with potential to influence cell–scaffold interactions.

Animal derived food↗

Machine Learning to Identify Geologic Factors Associated with Production in Geothermal Fields: A Case-Study Using 3D Geologic Data from Brady Geothermal Field and NMFk

In this paper, we present an analysis using unsupervised machine learning (ML) to identify the key geologic factors that contribute to the geothermal production in Brady geothermal field. Brady is a hydrothermal system in northwestern Nevada that supports both electricity production and direct use of hydrothermal fluids. Transmissive fuid-fow pathways are relatively rare in the subsurface, but are critical components of hydrothermal systems like Brady and many other types of fuid-fow systems in fractured rock. Here, we analyze geologic data with ML methods to unravel the local geologic controls on these pathways. The ML method, non-negative matrix factorization with k-means clustering (NMFk), is applied to a library of 14 3D geologic characteristics hypothesized to control hydrothermal circulation in the Brady geothermal field. Our results indicate that macro-scale faults and a local step-over in the fault system preferentially occur along production wells when compared to injection wells and non-productive wells. We infer that these are the key geologic characteristics that control the through-going hydrothermal transmission pathways at Brady. Our results demonstrate: (1) the specific geologic controls on the Brady hydrothermal system and (2) the efficacy of pairing ML techniques with 3D geologic characterization to enhance the understanding of subsurface processes. This submission includes the published journal article detailing this work, the published 3D geologic map of the Brady Geothermal Area used as a basis to develop structural and geological variables that are hypothesized to control or effect permeability or connectivity, 3D well data, along which geologic data were sampled for PCA analyses, and associated metadata file. This work was done using the GeoThermalCloud framework, which is part of SmartTensors (both are linked below).

15 GEOTHERMAL ENERGY↗

Covariance Shaping Over Riemannian Manifolds for Massive MIMO Communication

Acquiring accurate instantaneous channel state information (CSI) is a challenging aspect of massive multi-input multi-output (MIMO) communication. Utilizing statistical information, such as channel covariance matrix, to design statistical beamforming vectors is robust when compared to instantaneous CSI. In this paper, we propose a novel MIMO covariance shaping scheme over Riemannian manifolds. It serves as an effective statistical beamforming solution to a number of close proximity user equipment (UE) that are undergoing substantial channel correlation. Proposed algorithm exploits the Hermitian positive definite nature of covariance matrices lying over Riemannian manifold. We introduce Wasserstein distance function as a Riemannian metric to measure distances between channel covariance matrices. Furthermore, K-means clustering technique is utilized to effectively identify the optimal shape of effective optimal covariance matrices. Our findings suggest that maximizing the geodesic distance between covariance matrices ultimately leads to a corresponding increase in the network throughput, as determined by the beamforming vector used to shape the covariance matrices. Simulation results validate that the proposed solution converges faster than Euclidean-based state-of-the-art, while maintaining the same computational complexity. Finally, the sum rate performance asymptotically achieves full capacity for two-UE case and more than 96% of the upper bound exhaustive search benchmark for multi-UE scenario.

42 ENGINEERING↗

Machine-learning-assisted analysis of transition metal dichalcogenide thin-film growth

In situ reflective high-energy electron diffraction (RHEED) is widely used to monitor the surface crystalline state during thin-film growth by molecular beam epitaxy (MBE) and pulsed laser deposition. With the recent development of machine learning (ML), ML-assisted analysis of RHEED videos aids in interpreting the complete RHEED data of oxide thin films. The quantitative analysis of RHEED data allows us to characterize and categorize the growth modes step by step, and extract hidden knowledge of the epitaxial film growth process. In this study, we employed the ML-assisted RHEED analysis method to investigate the growth of 2D thin films of transition metal dichalcogenides (ReSe2) on graphene substrates by MBE. Principal component analysis (PCA) and K-means clustering were used to separate statistically important patterns and visualize the trend of pattern evolution without any notable loss of information. Using the modified PCA, we could monitor the diffraction intensity of solely the ReSe2 layers by filtering out the substrate contribution. These findings demonstrate that ML analysis can be successfully employed to examine and understand the film-growth dynamics of 2D materials. Further, the ML-based method can pave the way for the development of advanced real-time monitoring and autonomous material synthesis techniques.

36 MATERIALS SCIENCE↗