Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “k mean”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Automated characterization of spatial and dynamical heterogeneity in supercooled liquids via implementation of machine learning

Abstract A computational approach by an implementation of the principle component analysis (PCA) with K -means and Gaussian mixture (GM) clustering methods from machine learning algorithms to identify structural and dynamical heterogeneities of supercooled liquids is developed. In this method, a collection of the average weighted coordination numbers ( W C N s ‾ ) of particles calculated from particles’ positions are used as an order parameter to build a low-dimensional representation of feature (structural) space for K -means clustering to sort the particles in the system into few meso-states using PCA. Nano-domains or aggregated clusters are also formed in configurational (real) space from a direct mapping using associated meso-states’ particle identities with some misclassified interfacial particles. These classification uncertainties can be improved by a co-learning strategy which utilizes the probabilistic GM clustering and the information transfer between the structural space and configurational space iteratively until convergence. A final classification of meso-states in structural space and domains in configurational space are stable over long times and measured to have dynamical heterogeneities. Armed with such a classification protocol, various studies over the thermodynamic and dynamical properties of these domains indicate that the observed heterogeneity is the result of liquid–liquid phase separation after quenching to a supercooled state.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

Benchmarking image processing techniques for porosity measurement in polymer additive manufacturing: Review and experimental analysis

An image processing workflow is proposed for porosity measurement in polymer additive manufacturing. Various techniques, including global and local thresholding, region growing, and K-means clustering, were applied to microscopic images of carbon fiber reinforced acrylonitrile butadiene styrene (CF-ABS) and benchmarked for their ability to accurately measure porosity. Global methods included Otsu, minimum error, iterative, and entropy-based thresholding, while local methods included Niblack, Bernsen, Sauvola, and Bradley-Roth algorithms. Artificial uneven illumination was introduced to test local adaptive thresholds. Results showed significant differences in porosity values across methods. Otsu, region growing, and K-means clustering excelled under uniform illumination, while Sauvola and Bradley-Roth performed better with uneven illumination. Comparison with X-ray computed tomography (XCT) revealed slightly lower porosity values (2.55 %) than optimized methods (2.73–2.79 %) due to XCT's lower resolution excluding smaller pores. While XCT offers finer pore detection, it limits sample volume and underestimates porosity due to spatial variation. Validation using artificial grayscale images with 5 % porosity confirmed that Otsu, Bradley-Roth, region growing, and Sauvola algorithms produced accurate results. Although tested on a single material system, these methods can be adapted to others with optimization. In conclusion, given XCT's high computational and time costs, this study highlights suitable image processing techniques as cost-effective alternatives for porosity analysis in polymer composites.

Additive manufacturing↗

Evaluating Grid Strength under Uncertain Renewable Generation

The increasing displacement of synchronous generators with renewable resources such as wind and solar via power electronic interfaces causes a reduction in short-circuit strength and weak grid issues. The variation and uncertainty of renewable energy increase challenges for identifying weak grid conditions. This paper proposes an efficient method to analyze the impact of uncertain renewable energy on grid strength. The proposed method uses the probabilistic collocation method (PCM) to approximate the results of grid strength assessment under uncertain renewable generation, in order to reduce computational burden without compromising result accuracy when compared with traditional Monte Carlo simulation (MCS). To improve the accuracy of the approximation results, the proposed method integrates the K-means clustering technique with PCM to select the approximation samples of input variables. The efficacy of the proposed method is demonstrated by comparison with MCS on the modified IEEE 9-bus system and modified IEEE 39-bus system with multiple renewable generators.

grid strength↗

The Empirical Effect of Fleet Optimization on Synchronization and Rebound Effects in Heat Pump Water Heaters

Demand response is a growing concept in light of the internet of things and an increasing need for grid flexibility. Water heaters are one of the preferred devices for providing demand response for grid services and peak management due to their capability to store energy. The efficient use of water heaters for demand response requires consideration of the associated load effects such as synchronization of device schedules and rebound effect. These effects present a significant challenge. Despite the importance of the mentioned effects for water heater queuing and scheduling, there has been no effort to quantify and empirically validate their impact. This study attempts to address this gap by offering two methods - Ward clustering and Euclidean K-means - to evaluate the extent of synchronization in a fleet of 42 water heaters in Atlanta, GA. Using the aforementioned methods on the measured data, we find evidence of convergence of water heater loads as a result of optimization compared to an idle period and analyzed their impact.

demand response↗

Identifying Climate Patterns Using Clustering Autoencoder Techniques

Abstract The complexity of growing spatiotemporal resolution of climate simulations produces a variety of climate patterns under different projection scenarios. This paper proposes a new data-driven climate classification workflow via an unsupervised deep learning technique that can dimensionally reduce the vast volume of spatiotemporal numerical climate projection data into a compact representation. We aim to identify distinct zones that capture multiple climate variables as well as their future changes under different climate change scenarios. Our approach leverages convolutional autoencoders combined with k -means clustering (standard autoencoder) and online clustering based on the Sinkhorn–Knopp algorithm (clustering autoencoder) across the conterminous United States (CONUS) to capture unique climate patterns in a data-driven fashion from the Geophysical Fluid Dynamics Laboratory Earth System Model with GOLD component (GFDL-ESM2G). The developed approach compresses 70 years of GFDL-ESM2G simulation at 0.125° spatial resolution across the CONUS under multiple warming scenarios to a lower-dimensional space by a factor of 660 000 and then tested on 150 years of GFDL-ESM2G simulation data. The results show that five climate clusters capture physically reasonable and spatially stable climatological patterns matched to known climate classes defined by human experts. Results also show that using a clustering autoencoder can reduce the computational time for clustering by up to 9.2 times when compared to using a standard autoencoder. Our five unique climate patterns resulting from the deep learning–based clustering of the lower-dimensional space thereby enable us to provide insights on hydrometeorology and its spatial heterogeneity across the conterminous United States immediately without downloading large climate datasets. Significance Statement This paper presents a data-driven climate classification approach using unsupervised deep learning to dimensionally reduce climate model outputs and to identify distinct climate regions for their future changes. Our approach compresses climate information for 70 years of Geophysical Fluid Dynamics Laboratory Earth System Model data across the conterminous United States (CONUS) at 0.125° spatial resolution. The results reveal that five climate clusters capture reasonable and stable climatological patterns matched to known climate patterns. The embedded clustering process in deep learning provides ×9.2 times faster execution than the k -means clustering technique. These results give us insight about climate spatial patterns and heterogeneity of hydrological patterns across the conterminous United States without downloading large climate datasets.

Kurihana, Takuya↗

Programs and Code for Geothermal Exploration Artificial Intelligence

The scripts below are used to run the Geothermal Exploration Artificial Intelligence developed within the "Detection of Potential Geothermal Exploration Sites from Hyperspectral Images via Deep Learning" project. It includes all scripts for pre-processing and processing, including: - Land Surface Temperature K-Means classifier - Labeling AI using Self Organizing Maps (SOM) - Post-processing for Permanent Scatterer InSAR (PSInSAR) analysis with SOM - Mineral marker summarizing - Artificial Intelligence (AI) Data splitting: creates data set from a single raster file - Artificial Intelligence Model: creates AI from a single data set, after splitting in Train, Validation and Test subsets - AI Mapper: creates a classification map based on a raster file

15 GEOTHERMAL ENERGY↗

Uncovering acoustic signatures of pore formation in laser powder bed fusion

Abstract We present a machine learning workflow to discover signatures in acoustic measurements that can be utilized to create a low-dimensional model to accurately predict the location of keyhole pores formed during additive manufacturing processes. Acoustic measurements were sampled at 100 kHz during single-layer laser powder bed fusion (LPBF) experiments, and spatio-temporal registration of pore locations was obtained from post-build radiography. Power spectral density (PSD) estimates of the acoustic data were then decomposed using non-negative matrix factorization with custom $$\varvec{k}$$ k -means clustering (NMF $$\varvec{k}$$ k ) to learn the underlying spectral patterns associated with pore formation. NMF $$\varvec{k}$$ k returned a library of basis signals and matching coefficients to blindly construct a feature space based on the PSD estimates in an optimized fashion. Moreover, the NMF $$\varvec{k}$$ k decomposition led to the development of computationally inexpensive machine learning models which are capable of quickly and accurately identifying pore formation with classification accuracy of supervised and unsupervised label learning greater than 95% and 90%, respectively. The intrinsic data compression of NMF k , the relatively light computational cost of the machine learning workflow, and the high classification accuracy makes the proposed workflow an attractive candidate for edge computing toward in-situ keyhole pore prediction in LPBF.

36 MATERIALS SCIENCE↗

Spatial Distribution and Clustering of Glycosaminoglycans in Electrospun Gelatin-Based Scaffolds

The extracellular matrix (ECM) is comprised of components like collagen, elastin, and glycosaminoglycans (GAGs). Electrospun fibrous scaffolds are designed to replicate the form and composition of the native ECM, often requiring blending of various ECM component mimics to enhance cellular responses. However, the spatial distribution of blended components within these fibers remains unclear. This study investigates the spatial distribution of chondroitin sulfate-C (CSC) in electrospun gelatin-based scaffolds. scanning electron microscopy (SEM), attenuated reflectance-Fourier transform infrared (ATR-FTIR) spectroscopy, X-ray photoelectron spectroscopy (XPS), and Time-of-flight Secondary Ion Mass Spectrometry (ToF-SIMS) were applied for surface and subsurface chemical characterization of the fibrous scaffolds. SEM confirmed a fibrous morphology, while ATR-FTIR and XPS analyses indicated the presence of CSC through the identification of sulfate groups. ToF-SIMS imaging, alongside K-means clustering and Ripley’s K function, revealed a nonuniform CSC distribution with higher concentrations at the top layer of the scaffold. This study demonstrates that CSC presentation at the fiber surface varies with depth and differs from bulk incorporation while reveals nanoscale clustering and spatial heterogeneity at both the surface and subsurface of electrospun gelatin fibers. These findings define an underexplored design consideration with potential to influence cell–scaffold interactions.

Animal derived food↗

Machine Learning to Identify Geologic Factors Associated with Production in Geothermal Fields: A Case-Study Using 3D Geologic Data from Brady Geothermal Field and NMFk

In this paper, we present an analysis using unsupervised machine learning (ML) to identify the key geologic factors that contribute to the geothermal production in Brady geothermal field. Brady is a hydrothermal system in northwestern Nevada that supports both electricity production and direct use of hydrothermal fluids. Transmissive fuid-fow pathways are relatively rare in the subsurface, but are critical components of hydrothermal systems like Brady and many other types of fuid-fow systems in fractured rock. Here, we analyze geologic data with ML methods to unravel the local geologic controls on these pathways. The ML method, non-negative matrix factorization with k-means clustering (NMFk), is applied to a library of 14 3D geologic characteristics hypothesized to control hydrothermal circulation in the Brady geothermal field. Our results indicate that macro-scale faults and a local step-over in the fault system preferentially occur along production wells when compared to injection wells and non-productive wells. We infer that these are the key geologic characteristics that control the through-going hydrothermal transmission pathways at Brady. Our results demonstrate: (1) the specific geologic controls on the Brady hydrothermal system and (2) the efficacy of pairing ML techniques with 3D geologic characterization to enhance the understanding of subsurface processes. This submission includes the published journal article detailing this work, the published 3D geologic map of the Brady Geothermal Area used as a basis to develop structural and geological variables that are hypothesized to control or effect permeability or connectivity, 3D well data, along which geologic data were sampled for PCA analyses, and associated metadata file. This work was done using the GeoThermalCloud framework, which is part of SmartTensors (both are linked below).

15 GEOTHERMAL ENERGY↗

Covariance Shaping Over Riemannian Manifolds for Massive MIMO Communication

Acquiring accurate instantaneous channel state information (CSI) is a challenging aspect of massive multi-input multi-output (MIMO) communication. Utilizing statistical information, such as channel covariance matrix, to design statistical beamforming vectors is robust when compared to instantaneous CSI. In this paper, we propose a novel MIMO covariance shaping scheme over Riemannian manifolds. It serves as an effective statistical beamforming solution to a number of close proximity user equipment (UE) that are undergoing substantial channel correlation. Proposed algorithm exploits the Hermitian positive definite nature of covariance matrices lying over Riemannian manifold. We introduce Wasserstein distance function as a Riemannian metric to measure distances between channel covariance matrices. Furthermore, K-means clustering technique is utilized to effectively identify the optimal shape of effective optimal covariance matrices. Our findings suggest that maximizing the geodesic distance between covariance matrices ultimately leads to a corresponding increase in the network throughput, as determined by the beamforming vector used to shape the covariance matrices. Simulation results validate that the proposed solution converges faster than Euclidean-based state-of-the-art, while maintaining the same computational complexity. Finally, the sum rate performance asymptotically achieves full capacity for two-UE case and more than 96% of the upper bound exhaustive search benchmark for multi-UE scenario.

42 ENGINEERING↗

Machine-learning-assisted analysis of transition metal dichalcogenide thin-film growth

In situ reflective high-energy electron diffraction (RHEED) is widely used to monitor the surface crystalline state during thin-film growth by molecular beam epitaxy (MBE) and pulsed laser deposition. With the recent development of machine learning (ML), ML-assisted analysis of RHEED videos aids in interpreting the complete RHEED data of oxide thin films. The quantitative analysis of RHEED data allows us to characterize and categorize the growth modes step by step, and extract hidden knowledge of the epitaxial film growth process. In this study, we employed the ML-assisted RHEED analysis method to investigate the growth of 2D thin films of transition metal dichalcogenides (ReSe2) on graphene substrates by MBE. Principal component analysis (PCA) and K-means clustering were used to separate statistically important patterns and visualize the trend of pattern evolution without any notable loss of information. Using the modified PCA, we could monitor the diffraction intensity of solely the ReSe2 layers by filtering out the substrate contribution. These findings demonstrate that ML analysis can be successfully employed to examine and understand the film-growth dynamics of 2D materials. Further, the ML-based method can pave the way for the development of advanced real-time monitoring and autonomous material synthesis techniques.

36 MATERIALS SCIENCE↗

Sub-10 nm Probing of Ferroelectricity in Heterogeneous Materials by Machine Learning Enabled Contact Kelvin Probe Force Microscopy

Reducing the dimensions of ferroelectric materials down to the nanoscale has strong implications on the ferroelectric polarization pattern and on the ability to switch the polarization. As the size of ferroelectric domains shrinks to the nanometer scale, the heterogeneity of the polarization pattern becomes increasingly pronounced, enabling a large variety of possible polar textures in nanocrystalline and nanocomposite materials. Critical to the understanding of fundamental physics of such materials and hence their applications in electronic nanodevices is the ability to investigate their ferroelectric polarization at the nanoscale in a nondestructive way. We show that contact Kelvin probe force microscopy (cKPFM) combined with a k-means response clustering algorithm enables to measure the ferroelectric response at a mapping resolution of 8 nm. In a BaTiO 3 thin film on silicon composed of tetragonal and hexagonal nanocrystals, we determine a nanoscale lateral distribution of discrete ferroelectric response clusters, fully consistent with the nanostructure determined by transmission electron microscopy. Moreover, we apply this data clustering method to the cKPFM responses measured at different temperatures, which allows us to follow the corresponding change in the polarization pattern as the Curie temperature is approached and across the phase transition. This work opens up perspectives for mapping complex ferroelectric polarization textures such as curled/swirled polar textures that can be stabilized in epitaxial heterostructures and more generally for mapping the polar domain distribution of any spatially highly heterogeneous ferroelectric materials.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Exploring Geothermal Potential of Great Basin Sub-Regions

The INnovative Geothermal Exploration through Novel Investigations Of Undiscovered Systems (INGENIOUS) project aims to discover new, economically viable hidden geothermal systems in the Great Basin region by building on previous work in play fairway analysis and machine learning. A key objective of this project is to develop an exploration workflow to reduce geothermal exploration risks for hidden geothermal systems. A single preliminary play fairway workflow was developed from the assessment of the regional INGENIOUS geological, geophysical, and geochemical datasets. This workflow provided new preliminary predictive geothermal fairway maps for the INGENIOUS study area, which encompasses most of Nevada, western Utah, southern Idaho, southeastern Oregon, and easternmost California. However, a recent study (incorporating machine learning techniques) of a portion of Nevada identified four geologic domains and determined that the relative importance of individual datasets or features as indicators of geothermal potential may differ across these domains. The INGENIOUS study area includes a much larger and more geologically diverse region; therefore, additional geologic domains or sub-regions are expected. To assess the sub-regions in the INGENIOUS study area, principal component analysis and k-means clustering were applied. Preliminary results indicate that the INGENIOUS regional data cluster into groups that relate to different geologic domains in the Great Basin region. These include domains such as the Walker Lane, extensional western Great Basin region, broad lower strain region in the eastern Great Basin of western Utah and eastern Nevada, Quaternary volcanic fields, and the area adjacent to the Snake River Plain. These clusters are assessed to determine the key geologic drivers of the identified clusters. Understanding this variability can provide key insights for the exploration and characterization of hidden geothermal systems in the Great Basin region and could indicate the need to develop multiple geothermal conceptual models and play fairway workflows for the INGENIOUS study area.

exploration↗

Crystallographic variant mapping using precession electron diffraction data

In this work, we developed three methods to map crystallographic variants of samples at the nanoscale by analyzing precession electron diffraction data using a high-temperature shape memory alloy and a VO2 thin film on sapphire as the model systems. The three methods are (I) a user-selecting-reference pattern approach, (II) an algorithm-selecting-reference-pattern approach, and (III) a k-means approach. In the first two approaches, Euclidean distance, Cosine, and Structural Similarity (SSIM) algorithms were assessed for the diffraction pattern similarity quantification. We demonstrated that the Euclidean distance and SSIM methods outperform the Cosine algorithm. We further revealed that the random noise in the diffraction data can dramatically affect similarity quantification. Denoising processes could improve the crystallographic mapping quality. With the three methods mentioned above, we were able to map the crystallographic variants in different materials systems, thus enabling fast variant number quantification and clear variant distribution visualization. The advantages and disadvantages of each approach are also discussed. We expect these methods to benefit researchers who work on martensitic materials, in which the variant information is critical to understand their properties and functionalities.

Crystallographic variant mapping↗

Accelerating the Structure Exploration of Diverse Bi–Pt Nanoclusters via Physics‐Informed Machine Learning Potential and Particle Swarm Optimization

Bimetallic Bi–Pt nanoclusters exhibit diverse structural motifs, including core-shell, Janus, and mixed alloy configurations, due to the unique bonding characteristics between Bi and Pt atoms. Using density functional theory refinements from ChIMES physically machine-learned potential and CALYPSO particle swarm optimization global searches, 34 Bi20-Pt20 nanoclusters are systematically classified. The results reveal that Bi atoms predominantly occupy surface sites, driven by charge transfer effects. Cohesive energy trends alone prove insufficient for structure differentiation, necessitating a data-driven approach employing principal component analysis and K-means clustering. Furthermore, vibrational, electronic, and infrared spectral analyses provide additional insights into structure-property relationships. The findings offer an original framework for the automated classification and analysis of bimetallic nanoclusters, enhancing the understanding of their stability and functional properties.

bimetallic nanoparticles↗

Shared Use Travel Behavior for Improving Rural Mobility: Insights from Greene County, Pennsylvania

Rural communities are considered disadvantaged communities as they suffer from a lack of transport options. Thus, rural regionsprovide less accessibility for commuters to reach their destination as opposed to urban regions. However, the issues of transport disadvantageand shared use mobility in rural areas within the United States (US) have not been well investigated. Furthermore, transport disadvantagediffers between communities and regions across the globe; thus, there is a need to study the behavioral choices of rural commuters within theUS context. This study contributes by analyzing the behavioral choices of rural communities within the US through a case study site ofWaynesburg, Pennsylvania, for adopting a shared use shuttle service. K-means clusters showed that trips from the survey data were a goodrepresentation of real trips from Ecolane. Furthermore, random parameter-based binary logit models were calibrated using data collected fromstudents, faculty, and residents in Waynesburg, Greene County, to study the behavioral choices of commuters. The findings for the faculty andstudents group revealed that prior experience with shared services increases the likelihood of using a shared shuttle. An important personalcharacteristic of inconvenience showed a higher propensity toward using existing modes as opposed to a shared shuttle. Such commutersvalue personal vehicles as more convenient as they have childcare responsibilities and varying schedules for work that require them to moveback and forth across locations, thus making a shared shuttle less attractive for them. The socioeconomic factors of age and gender show ahigher propensity for using shared shuttles. Furthermore, the findings from this study could be helpful for agencies in improving rural mobility andconsidering such shared mobility services for rural communities

42 ENGINEERING↗

Anomalously high elastic modulus of a poly(ethylene oxide)-based composite electrolyte

The practical use of lithium metal anodes in solid-state batteries requires a polymer membrane with high lithium-ion conductivity, thermal/electrochemical stability, and mechanical strength. The primary challenge is to effectively decouple the ionic conductivity and mechanical strength of the polymer electrolytes. We report a remarkably facile single step synthetic strategy based on in-situ crosslinking of poly(ethylene oxide) (xPEO) in the presence of a woven glass fiber (GF). Such a simple method yields composite polymer electrolytes (CPE) of anomalously high elastic modulus up to 2.5 GPa over a broad temperature range (20 °C – 245 °C) that has never been previously documented. An unsupervised machine learning algorithm, K-mean clustering analysis, was implemented on the hyperspectral Raman mapping at the xPEO/GF interface. Using such a unique means, we show for the first time that the promoted mechanical strength originates from xPEO and GF interactions through dynamic hydrogen and ionic bonding. High ionic conductivity is achieved by the addition plasticizer (e.g. tetraglyme), where trifluoromethanesulfonate anions are tethered to the xPEO matrix and Li + cations are favorably transported through coordination with the plasticizer. Further, stringent galvanostatic cycling tests indicates the CPE can be stably cycled for >3000 h in a Li-metal symmetric cell at a moderate temperature (nearly 1500 Coulombs/cm 2 Li equivalents), outperforming most of the PEO-based electrolytes. The GF reinforced CPE reported here has multifunctional uses, such as solid electrolytes for all solid-state batteries and membranes for redox-flow batteries. Although the focus of this study is on lithium-based batteries, the results are equally promising for other alkali metal based batteries such as sodium and potassium.

25 ENERGY STORAGE↗

Discovering Hidden Geothermal Signatures using Unsupervised Machine Learning

Discovering hidden geothermal resources is a very challenging task. It requires the mining of large datasets, including various diverse data attributes representing subsurface hydrogeological and geothermal conditions. The commonly used Play Fairway Analysis (PFA) typically relies on subject-matter expertise to analyze site or regional data to estimate geothermal conditions and prospectivity. Here, we demonstrate an alternative approach based on machine learning (ML) to process a geothermal dataset of Southwest New Mexico (SWNM). The study region includes low- and medium-temperature hydrothermal systems. However, most of these systems are poorly characterized because of insufficient existing data and limited past explorative studies. This study aims to discover hidden patterns and relationships in the SWNM geothermal dataset to better understand regional hydrothermal conditions. This is achieved by applying an unsupervised machine learning algorithm based on non-negative matrix factorization coupled with customized k-means clustering (NMFk). NMFk can automatically identify (1) hidden (latent) signatures characterizing datasets, (2) the optimal number of these signatures, (3) dominant data attributes associated with each signature, and (4) spatial distribution of the extracted signatures. Here, NMFk is applied to analyze 18 geological, geophysical, hydrogeological, geothermal attributes at 44 locations in SWNM. NMFk successfully finds data patterns and identifies the spatial associations of hydrothermal signatures with the four physiographic provinces in SWNM (Colorado Plateau, Volcanic Field, Basin and Range, and the Rio Grande rift). The algorithm identified up to 5 hydrothermal signatures in the SWNM datasets that differentiate between low- and medium-temperature hydrothermal systems in different provinces. Also, the algorithm identifies two medium-temperature hydrothermal systems in SWNM that require further exploration for geothermal resource development. Based on our analyses, 12 of the attributes are important to identify medium-temperature hydrothermal systems, and the remaining six attributes are critical to characterize low-temperature hydrothermal systems. Based on the obtained results, we identify potential physiographic provinces for further exploration to characterize them as geothermal resources. The resulting NMFk model can be applied to predict geothermal conditions and their uncertainties at new SWNM locations based on limited data from unexplored areas.

58 GEOSCIENCES↗