Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “k mean”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

Machine-learning predictions of the shale wells’ performance

The ultra-low permeability nature of shale reservoirs leads to an extended linear flow and necessitates horizontal wells with multi-stage engineered fractures to efficiently extract hydrocarbons resources. These artificially-generated and naturally-occurring fractures form complex networks that create complex flow regimes which control oil production. These fractures are neither identical nor equally-spaced, which leads to a production profile with a masked onset of the boundary-dominated flow. The combination of the extended linear flow with the indeterminate onset of the boundary-dominated flow challenges the current deterministic analytic approaches to forecast the estimated ultimate recovery (EUR). In this work, we propose a novel machine-learning approach which overcomes these challenges and provides reliable EUR estimates based on field-wide analyses. We implement a novel unsupervised machine learning (ML) methodology, which allows for automatic identification of the optimal number of features (signals) present in the data based on non-negative matrix/tensor factorization coupled with k-means clustering incorporating regularization and physics constraints. In the presented analyses, the input data to the ML algorithm is the available (public) production history from the field collected at existing unconventional reservoirs. We validate our approach through hindcasting of the production data, where we achieved an excellent agreement. In addition, our approach is able to identify the poorly-performing wells, which could benefit from early refracing. Our approach provides fast and accurate estimations of the well performance without presumptions about the state of the well or the flow regime.

03 NATURAL GAS↗

Distributed Tomographic Reconstruction with Quantization

Conventional tomographic reconstruction typically depends on centralized servers for both data storage and computation, leading to concerns about memory limitations and data privacy. Distributed reconstruction algorithms mitigate these issues by partitioning data across multiple nodes, reducing server load and enhancing privacy. However, these algorithms often encounter challenges related to memory constraints and communication overhead between nodes. In this paper, we introduce a decentralized Alternating Directions Method of Multipliers (ADMM) with configurable quantization. By distributing local objectives across nodes, our approach is highly scalable and can efficiently reconstruct images while adapting to available resources. To overcome communication bottlenecks, we propose two quantization techniques based on K-means clustering and JPEG compression. Numerical experiments with benchmark images illustrate the tradeoffs between communication efficiency, memory use, and reconstruction accuracy.

Miao, Runxuan↗

Automated identification of deformation twin systems in Mg WE43 from SEM DIC

In this study, the application of machine learning and computer vision approaches to microscale deformation data for the automated identification of deformation twinning systems, and their associated twin area, is introduced. Deformation data was obtained during in-situ SEM compression testing of a WE43 Mg alloy using digital image correlation (DIC) modified for use with electron microscopy. A 5.7 mm × 3.4 mm area of interest was analyzed, generating ~100 million data points of deformation. Experimental twin trace directions were determined by applying k-means clustering, morphological thinning, and Hough transforms to the deformation data. The identification of deformation twinning was achieved by consideration of the twin trace direction, its strain value, and its evolution. The deformation map was divided into areas corresponding to individual active twin systems, enabling the analysis of microstructure dependence of deformation behavior and twinning activity. The performance of the proposed twinning identification approach was evaluated by accuracy, precision, and sensitivity metrics. The effect of the tuning parameters used in the algorithm on the performance is also discussed.

36 MATERIALS SCIENCE↗

Unveiling the structural transformations in glassy solid electrolyte adapted to high-stress cycles

Brittleness of glass–ceramic and ceramic ion conductors is considered as a main roadblock for their implementation as electrolytes in solid-state batteries where the fractures often occur due to the pressure exerted by metallic lithium. In this regard, nano- and micro-scale ductility of the solid electrolyte allows reducing such pressure without formation of cracks. Among different types of solid state ion conductors, phosphate invert glasses seem to be promising in achieving such ductility. We report the mechanical behavior of lithium phosphorous oxynitride (LiPON) invert glass probed by static and cyclic nanoindentation. Repeated application of high intensity stress results in densification and shear deformation of material allowing LiPON to accommodate the ~ 22% nominal strain imposed by the nanoindenter without cracking. Ability to do this under repeated loading indicates robustness of LiPON when used in lithium metal batteries under cyclic charge and discharge. Using Raman spectroscopy with unsupervised K-means clustering we reveal pressure-induced formation of P 2 O 7 units which migrate to the periphery of the residual hardness impressions with cycling resulting in surface morphology of LiPON superficially similar to that of deformed bulk metallic glass.

36 MATERIALS SCIENCE↗

Machine learning and shallow groundwater chemistry to identify geothermal prospects in the Great Basin, USA

This study discovers various geothermal prospects in the Great Basin, USA based on shallow groundwater chemical (geochemical) data. The geochemical data are expected to include hidden (latent) information that is a proxy for geothermal prospectivity. We processed the sparse geochemical data in the Great Basin at 14,341 locations including 18 attributes. Next, a non-negative matrix factorization with customized k-means clustering is applied to the geochemical data matrix that automatically finds three hidden geothermal signatures representing modestly, moderately, and highly confident geothermal prospects. The algorithm also evaluated the probability of occurrence of these types of resources through the studied region. There is a consistency between regional geothermal prospectivity as estimated by our ML methodology and the traditional play fairway analysis conducted over a portion of the study area. We also identify the dominant data attributes associated with each signature. Finally, our ML analyses allow us to reconstruct attributes from sparse into continuous over the study domain. The predicted continuous attributes can be used for future detailed geothermal explorations in the Great Basin.

15 GEOTHERMAL ENERGY↗

Navigating Large Chemical Spaces Using Graph Theory and Integer Programming

Navigating and analyzing large chemical spaces are necessary to accelerate the design and discovery of new molecules and chemical processes. In this work, we introduce a computational framework that integrates graph theory and integer programming to enable the efficient navigation of large chemical spaces. Our framework represents the chemical space as a graph, wherein nodes represent molecules and edges represent the degree of similarity or connectivity based on domain-specific information. Using the graph representation, we identify representative molecules by computing the so-called minimum dominating set (MDS), which in our context is the minimum set of molecules that is connected to all other molecules. We present a suite of solution strategies for the MDS problem including heuristic and rigorous integer programming (IP) approaches. We show that these approaches allow us to capture physicochemical properties and domain-specific logic and constraints, facilitating the identification of molecules with the target properties. We demonstrate the effectiveness of the proposed approach by navigating the chemical space of per- and polyfluoroalkyl substances (PFAS); this comprises approximately 15,000 molecular structures. We compare our framework against traditional dimensionality reduction and clustering methods such as t-SNE and K-means clustering.

Chemical structure↗

Assessing Shifts in Regional Hydroclimatic Conditions of U.S. River Basins in Response to Climate Change over the 21st Century

Characterization of shifts in regional hydroclimatic conditions helps reduce negative consequences on agriculture, environment, economy, society, and ecosystem. This study assesses shifts in regional hydroclimatic conditions across the conterminous United States in response to climate change over the 21 st Century. The hydrological responses of five downscaled climate models from the Multivariate Adaptive Constructed Analogs (MACA) dataset ranging from the driest to wettest and least warm to hottest were simulated using the Variable Infiltration Capacity (VIC) model. Shifts in regional hydroclimatic conditions at 8-digit hydrologic unit scale (HUC8) were evaluated by the magnitude and direction of movements in the Budyko space. HUC8 river basins were then clustered into seven unique hydroclimatic behavior groups using the K-means method. A tree classification method was proposed to illustrate the relationships between hydroclimatic behavior groups and regional characteristics. The results indicate that hydroclimatic responses may vary from a river basin to another, but basins in the same neighborhood follow a similar movement in the Budyko space. The systematic hydroclimatic behavior of river basins is highly associated with their regional landform, climate, and ecosystem characteristics. Most HUC8s with Mountain, Plateau and Basin landform types will likely experience less arid conditions. However, most HUC8s with Plain landform type behave differently according to the regional ecosystem and climate. This study provides a potential roadmap of shifts in regional hydroclimatic conditions of U.S. river basins, which can be used to improve regional preparedness and ability of various sectors to mitigate or adapt to the impacts of future hydroclimate change.

54 ENVIRONMENTAL SCIENCES↗

A Satellite-Based Estimate of Convective Vertical Velocity and Convective Mass Flux: Global Survey and Comparison with Radar Wind Profiler Observations

Convective vertical velocity (w c ) and convective mass flux (M c ) lie at the heart of GCM cumulus parameterizations, but few observations of these critical parameters are available. In this paper, we develop and evaluate a novel, satellite-based method for estimating profiles of w c and M c . Here, comparisons with collocated ground-based radar wind profiler (RWP) observations show that satellite estimated median w c is slightly greater than the RWP estimates, but they show solid agreement when compared at the 95th percentiles (intense updrafts). RWP-derived and satellite estimated M c are broadly comparable in the lower and middle troposphere, with some differences in the upper troposphere due to differences in convective core sampling. A k-means cluster analysis of multiple years of w c data shows that convective characteristics are distinctly different among extratropical convection, tropical land convection, and tropical oceanic convection. Tropical land convection is significantly more intense and more variable than the oceanic counterpart.

54 ENVIRONMENTAL SCIENCES↗

Characterizing Drought Behavior in the Colorado River Basin Using Unsupervised Machine Learning

Drought is a pressing issue for the Colorado River Basin (CRB) due to the social and economic value of water resources in the region and the significant uncertainty of future drought under climate change. Here, we use climate simulations from various Earth System Models (ESMs) to force the Variable Infiltration Capacity hydrologic model and project multiple drought indicators for the sub-watersheds within the CRB. We apply an unsupervised machine learning (ML) based on Non-Negative Matrix Factorization using K-means clustering (NMFk) to synthesize the simulated historical, future, and change in drought indicators. The unsupervised ML approach can identify sub-watersheds where key changes to drought indicator behavior occur, including shifts in snowpack, snowmelt timing, precipitation, and evapotranspiration. While changes in future precipitation vary across ESMs, the results indicate that the Upper CRB will experience increasing evaporative demand and surface-water scarcity, with some locations experiencing a shift from a radiation-limited to a water-limited evaporation regime in the summer. Large shifts in peak runoff are observed in snowmelt-dominant sub-watersheds, with complete disappearance of the snowmelt signal for some sub-watersheds. The work demonstrates the utility of the NMFk algorithm to efficiently identify behavioral changes of drought indicators across space and time and to quickly analyze and interpret hydro climate model results.

54 ENVIRONMENTAL SCIENCES↗

Extension of Large Fire Emissions From Summer to Autumn and Its Drivers in the Western US

Abstract Burned areas in the western US have increased ten‐fold since 1980s, which are attributable to multiple factors, including increasing heat, changing precipitation patterns, and extended drought. To better understand how these factors contribute to large fire emissions (gridded monthly fire emissions >95th percentile of all the fire emissions in the western US; 0.009 Gg/month), we build a machine learning model to predict fire emissions (PM 2.5 ) over the western US at 0.25° resolution, interpreted using explainable artificial intelligence (XAI). From the predictor contributions derived from XAI, we conduct k‐means clustering analysis to identify four clusters of predictor variables representing different drivers of large fire emissions. The four clusters feature the contributions of fuel load (Cluster 1) and different levels of dryness (Cluster 2–4), controlled by fuel moisture, drought condition, and fire‐favorable large‐scale meteorological patterns featuring high temperature, high pressure, and low relative humidity. In the past two decades, large fire emissions peak in summer. However, large fire emissions increased significantly in September and October in 2010–2020 relative to 2000–2009, extending the peak large fire emissions from summer to autumn. The larger enhancements of large fire emissions during autumn compared to summer are contributed by decreased fuel moisture, along with more frequent concurrent fire‐favorable large‐scale meteorological patterns and drought. These results highlight fuel drying as a common driver supported by multiple drivers, such as warmer temperature and more frequent synoptic patterns favorable for fires, in increasing the autumn risk of large fire emissions across the western US.

54 ENVIRONMENTAL SCIENCES↗

Antecedent Hydrometeorological Conditions of Wildfire Occurrence in the Western U.S. in a Changing Climate

Abstract Wildfires have significant hydrological and ecological impacts in the western U.S. Using a high‐resolution regional climate simulation and wildfire observations for 1984–2018, this study investigates the antecedent hydrometeorological conditions (AHCs) of wildfires in the western U.S. During the warm season (April‐September), the wildfire AHCs feature diverse surface pressure (PS), soil moisture, and longwave/shortwave radiation (LW/SW) conditions. K‐means clustering classifies wildfires into four types with distinct AHCs: low‐PS‐type and high‐PS‐type with lower and higher PS anomalies, respectively, LW‐type featuring intense LW but weak SW anomalies, and wet‐soil‐type with wet soil anomalies. Each fire cluster represents 22%–27% of all the wildfires, featuring different combinations of climate and vegetation conditions and their diverse relations to regional hydrometeorological conditions, with wet‐soil‐type fires often exhibiting opposite correlations with AHCs compared to those of the other three types. In five major Köppen climate zones over the western U.S., clustering‐based predictions improve the seasonal wildfire prediction accuracy ( R 2 ) by 10% compared to prediction without classification. Such improvement comes from separating the opposite relationships between wet‐soil‐type fires and their seasonal AHCs from the other three types, along with separating LW‐type fires, which include most of the lightning‐ignited fires that occur more randomly. Increases in wildfire occurrence during 1984–2018 are dominated by the increases in the LW‐type fires, while the wet‐soil‐type fires have decreased, consistent with the long‐term drying in the western U.S.

54 ENVIRONMENTAL SCIENCES↗

Large‐Scale Statistically Meaningful Patterns (LSMPs) Associated With Precipitation Extremes Over Northern California

Abstract We analyze large‐scale statistically meaningful patterns (LSMPs) that precede extreme precipitation (PEx) events over Northern California (NorCal). We find LSMPs by applying k‐means clustering to the two leading principal components of daily 500 hPa geopotential height anomalies two days before the onset, from October to March during 1948–2015. Statistical significance testing based on Monte Carlo simulations suggests a minimum of four statistically distinguished LSMP clusters. The four LSMP clusters are characterized as Northwest continental negative height anomaly, Eastward positive “Pacific‐North American Pattern (PNA),” Westward negative “PNA,” and Prominent Alaskan ridge. These four clusters, shown in multiple variables, evolve very differently and have differing links to the Arctic and tropical Pacific regions. Using binary forecast skill measures and a new copula‐based framework for predicting PEx events, we find LSMP indices that are useful predictors of NorCal PEx events, with moisture‐based variables being the best predictors of PEx events at least 6 days before the onset, and the lower atmospheric variables being better than their upper atmospheric counterparts any day in advance tested. To ensure statistical rigor, the LSMPs analyzed here (with the modified acronym) include local tests of both significance and consistency, which are not always featured in the literature on large‐scale meteorological patterns.

54 ENVIRONMENTAL SCIENCES↗

Optical emissivity dataset of multi-material heterogeneous designs generated with automated figure extraction

Optical device design is typically an iterative optimization process based on a good initial guess from prior reports. Optical properties databases are useful in this process but difficult to compile because their parsing requires finding relevant papers and manually converting graphical emissivity curves to data tables. Here, we present two contributions: one is a dataset of thermal emissivity records with design-related parameters, and the other is a software tool for automated colored curve data extraction from scientific plots. We manually collected 64 papers with 176 figures reporting thermal emissivity and automatically retrieved 153 colored curve data records. The automated figure analysis software pipeline uses Faster R-CNN for axes and legend object detection, EasyOCR for axes numbering recognition, and k-means clustering for colored curve retrieval. Additionally, we manually extracted geometry, materials, and method information from the text to add necessary metadata to each emissivity curve. Finally, we analyzed the dataset to determine the dominant classes of emissivity curves and determine the underlying design parameters leading to a type of emissivity profile.

47 OTHER INSTRUMENTATION↗

Cancer genomics predicts disease relapse and therapeutic response to neoadjuvant chemotherapy of hormone sensitive breast cancers

Several studies provide insight into the landscape of breast cancer genomics with the genomic characterization of tumors offering exceptional opportunities in defining therapies tailored to the patient’s specific need. However, translating genomic data into personalized treatment regimens has been hampered partly due to uncertainties in deviating from guideline based clinical protocols. Here we report a genomic approach to predict favorable outcome to treatment responses thus enabling personalized medicine in the selection of specific treatment regimens. The genomic data were divided into a training set of N = 835 cases and a validation set consisting of 1315 hormone sensitive, 634 triple negative breast cancer (TNBC) and 1365 breast cancer patients with information on neoadjuvant chemotherapy responses. Patients were selected by the following criteria: estrogen receptor (ER) status, lymph node invasion, recurrence free survival. The k-means classification algorithm delineated clusters with low- and high- expression of genes related to recurrence of disease; a multivariate Cox’s proportional hazard model defined recurrence risk for disease. Classifier genes were validated by Immunohistochemistry (IHC) using tissue microarray sections containing both normal and cancerous tissues and by evaluating findings deposited in the human protein atlas repository. Based on the leave-on-out cross validation procedure of 4 independent data sets we identified 51-genes associated with disease relapse and selected 10, i.e. TOP2A, AURKA, CKS2, CCNB2, CDK1 SLC19A1, E2F8, E2F1, PRC1, KIF11 for in depth validation. Expression of the mechanistically linked disease regulated genes significantly correlated with recurrence free survival among ER-positive and triple negative breast cancer patients and was independent of age, tumor size, histological grade and node status. Importantly, the classifier genes predicted pathological complete responses to neoadjuvant chemotherapy (P < 0.001) with high expression of these genes being associated with an improved therapeutic response toward two different anthracycline-taxane regimens; thus, highlighting the prospective for precision medicine. Our study demonstrates the potential of classifier genes to predict risk for disease relapse and treatment response to chemotherapies. The classifier genes enable rational selection of patients who benefit best from a given chemotherapy thus providing the best possible care. The findings encourage independent clinical validation.

59 BASIC BIOLOGICAL SCIENCES↗

Unsupervised learning-enabled pulsed infrared thermographic microscopy of subsurface defects in stainless steel

Metallic structures produced with laser powder bed fusion (LPBF) additive manufacturing method (AM) frequently contain microscopic porosity defects, with typical approximate size distribution from one to 100 microns. Presence of such defects could lead to premature failure of the structure. In principle, structural integrity assessment of LPBF metals can be accomplished with nondestructive evaluation (NDE). Pulsed infrared thermography (PIT) is a non-contact, one-sided NDE method that allows for imaging of internal defects in arbitrary size and shape metallic structures using heat transfer. PIT imaging is performed using compact instrumentation consisting of a flash lamp for deposition of a heat pulse, and a fast frame infrared (IR) camera for measuring surface temperature transients. However, limitations of imaging resolution with PIT include blurring due to heat diffusion, sensitivity limit of the IR camera. We demonstrate enhancement of PIT imaging capability with unsupervised learning (UL), which enables PIT microscopy of subsurface defects in high strength corrosion resistant stainless steel 316 alloy. PIT images were processed with UL spatial–temporal separation-based clustering segmentation (STSCS) algorithm, refined by morphology image processing methods to enhance visibility of defects. The STSCS algorithm starts with wavelet decomposition to spatially de-noise thermograms, followed by UL principal component analysis (PCA), fine-tuning optimization, and neural learning-based independent component analysis (ICA) algorithms to temporally compress de-noised thermograms. The compressed thermograms were further processed with UL-based graph thresholding K-means clustering algorithm for defects segmentation. The STSCS algorithm also includes online learning feature for efficient re-training of the model with new data. For this study, metallic specimens with calibrated microscopic flat bottom hole defects, with diameters in the range from 203 to 76 µm, were produced using electro discharge machining (EDM) drilling. While the raw thermograms do not show any material defects, using STSCS algorithm to process PIT images reveals defects as small as 101 µm in diameter. To the best of our knowledge, this is the smallest reported size of a sub-surface defect in a metal imaged with PIT, which demonstrates the PIT capability of detecting defects in the size range relevant to quality control requirements of LPBF-printed high-strength metals.

36 MATERIALS SCIENCE↗

A review on recent machine learning applications for imaging mass spectrometry studies

Imaging mass spectrometry (IMS) is a powerful analytical technique widely used in biology, chemistry, and materials science fields that continue to expand. IMS provides a qualitative compositional analysis and spatial mapping with high chemical specificity. The spatial mapping information can be 2D or 3D depending on the analysis technique employed. Due to the combination of complex mass spectra coupled with spatial information, large high-dimensional datasets (hyperspectral) are often produced. Therefore, the use of automated computational methods for an exploratory analysis is highly beneficial. The fast-paced development of artificial intelligence (AI) and machine learning (ML) tools has received significant attention in recent years. These tools, in principle, can enable the unification of data collection and analysis into a single pipeline to make sampling and analysis decisions on the go. There are various ML approaches that have been applied to IMS data over the last decade. Here, in this review, we discuss recent examples of the common unsupervised (principal component analysis, non-negative matrix factorization, k-means clustering, uniform manifold approximation and projection), supervised (random forest, logistic regression, XGboost, support vector machine), and other methods applied to various IMS datasets in the past five years. The information from this review will be useful for specialists from both IMS and ML fields since it summarizes current and representative studies of computational ML-based exploratory methods for IMS.

47 OTHER INSTRUMENTATION↗

Atomic-scale electronic inhomogeneity in single-layer iron chalcogenide alloys revealed by machine learning of STM/S data

Chemical pressure from the isovalent substitution of Se by a larger Te atom in the epitaxial film of iron chalcogenide FeSe can effectively tune its superconducting, topological, and magnetic properties. However, such substitution during epitaxial growth inherently leads to defects and structural inhomogeneity, making the determination of alloy composition and atomic sites for the substitutional Te atoms challenging. Here, we utilize machine learning to distinguish between Se and Te atoms in scanning tunneling microscopy images of single-layer FeSe1−xTex on SrTiO3(001) substrates. Defect locations are first identified by analyzing spatial-dependent dI/dV tunneling spectra using the K-means clustering method. After excluding the defect regions, the remaining dI/dV spectra are further analyzed using the singular value decomposition method to determine the Se/Te ratio. Our findings demonstrate an effective and reliable approach for determining alloy composition and atomic-scale electronic inhomogeneity in superconducting single-layer iron chalcogenide films.

Materials Science↗

Graphic contrastive learning analyses of discontinuous molecular dynamics simulations: Study of protein folding upon adsorption

A comprehensive understanding of the interfacial behaviors of biomolecules holds great significance in the development of biomaterials and biosensing technologies. In this work, we used discontinuous molecular dynamics (DMD) simulations and graphic contrastive learning analysis to study the adsorption of ubiquitin protein on a graphene surface. Our high-throughput DMD simulations can explore the whole protein adsorption process including the protein structural evolution with sufficient accuracy. Contrastive learning was employed to train a protein contact map feature extractor aiming at generating contact map feature vectors. Subsequently, these features were grouped using the k-means clustering algorithm to identify the protein structural transition stages throughout the adsorption process. The machine learning analysis can illustrate the dynamics of protein structural changes, including the pathway and the rate-limiting step. Our study indicated that the protein–graphene surface hydrophobic interactions and the π–π stacking were crucial to the seven-stage adsorption process. Upon adsorption, the secondary structure and tertiary structure of ubiquitin disintegrated. The unfolding stages obtained by contrastive learning-based algorithm were not only consistent with the detailed analyses of protein structures but also provided more hidden information about the transition states and pathway of protein adsorption process and structural dynamics. Our combination of efficient DMD simulations and machine learning analysis could be a valuable approach to studying the interfacial behaviors of biomolecules.

97 MATHEMATICS AND COMPUTING↗