Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “clustering algorithm”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11

Machine Learning-Enabled Quantitative Analysis of Optically Obscure Scratches on Nickel-Plated Additively Manufactured (AM) Samples

Additively manufactured metal components often have rough and uneven surfaces, necessitating post-processing and surface polishing. Hardness is a critical characteristic that affects overall component properties, including wear. This study employed K-means unsupervised machine learning to explore the relationship between the relative surface hardness and scratch width of electroless nickel plating on additively manufactured composite components. The Taguchi design of experiment (TDOE) L9 orthogonal array facilitated experimentation with various factors and levels. Initially, a digital light microscope was used for 3D surface mapping and scratch width quantification. However, the microscope struggled with the reflections from the shiny Ni-plating and scatter from small scratches. To overcome this, a scanning electron microscope (SEM) generated grayscale images and 3D height maps of the scratched Ni-plating, thus enabling the precise characterization of scratch widths. Optical identification of the scratch regions and quantification were accomplished using Python code with a K-means machine-learning clustering algorithm. The TDOE yielded distinct Ni-plating hardness levels for the nine samples, while an increased scratch force showed a non-linear impact on scratch widths. The enhanced surface quality resulting from Ni coatings will have significant implications in various industrial applications, and it will play a pivotal role in future metal and alloy surface engineering.

36 MATERIALS SCIENCE↗

Multivariate Methods for Prediction of Geologic Sample Composition with Laser-Induced Breakdown Spectroscopy

Laser-induced breakdown spectroscopy (LIBS) uses pulses of laser light to ablate a material from the surface of a sample and produce an expanding plasma. The optical emission from the plasma produces a spectrum which can be used to classify target materials and estimate their composition. The ChemCam instrument on the Mars Science Laboratory (MSL) mission will use LIBS to rapidly analyze targets remotely, allowing more resource- and time-intensive in-situ analyses to be reserved for targets of particular interest. ChemCam will also be used to analyze samples that are not reachable by the rover's in-situ instruments. Due to these tactical and scientific roles, it is important that ChemCam-derived sample compositions are as accurate as possible. We have compared the results of partial least squares (PLS), multilayer perceptron (MLP) artificial neural networks (ANNs), and cascade correlation (CC) ANNs to determine which technique yields better estimates of quantitative element abundances in rock and mineral samples. The number of hidden nodes in the MLP ANNs was optimized using a genetic algorithm. The influence of two data preprocessing techniques were also investigated: genetic algorithm feature selection and averaging the spectra for each training sample prior to training the PLS and ANN algorithms. We used a ChemCam-like laboratory stand-off LIBS system to collect spectra of 30 pressed powder geostandards and a diverse suite of 196 geologic slab samples of known bulk composition. We tested the performance of PLS and ANNs on a subset of these samples, choosing to focus on silicate rocks and minerals with a loss on ignition of less than 2 percent. This resulted in a set of 22 pressed powder geostandards and 80 geologic samples. Four of the geostandards were used as a validation set and 18 were used as the training set for the algorithms. We found that PLS typically resulted in the lowest average absolute error in its predictions, but that the optimized MLP ANN and the CC ANN often gave results comparable to PLS. Averaging the spectra for each training sample and/or using feature selection to choose a small subset of wavelengths to use for predictions gave mixed results, with degraded performance in some cases and similar or slightly improved performance in other cases. However, training time was significantly reduced for both PLS and ANN methods by implementing feature selection, making this a potentially appealing method for initial, rapid-turn-around analyses necessary for Chemcam's tactical role on MSL. Choice of training samples has a strong influence on the accuracy of predictions. We are currently investigating the use of clustering algorithms (e.g. k-means, neural gas, etc.) to identify training sets that are spectrally similar to the unknown samples that are being predicted, and therefore result in improved predictions

Morris, Richard↗

Characterizing Signatures of Geothermal Exploration Data with Machine Learning Techniques: An Application to the Nevada Play Fairway Analysis

We are introducing machine learning methods to the play fairway analysis to generate geothermal potential maps to support the evaluation of geothermal resource potential and the exploration for undiscovered blind geothermal systems in the Nevada Great Basin region. Our project aims to identify new ways to combine the play fairway data and empirically organize relationships between feature weights and labels in an improved workflow. As a means of doing this, we introduce machine learning methods to evaluate the influence of certain geological and geophysical features/feature sets in predicting geothermal favorability. This report highlights promising approaches based on supervised and unsupervised learning methods. First, we demonstrate a filter method applied to supervised classification modeling. The supervised filter method is based on permutation analysis to evaluate every possible feature combination/drop out scenario and rank feature influence based on the performance variance of supervised classification models. Additionally, we present an unsupervised factor analysis based on principal component analysis coupled with a semi-supervised kmeans clustering algorithm. This analysis allows us to identify the optimal number of groups/clusters for training sites and structural settings to identify feature patterns including correlation, variance, and latent and dominant feature relationships. The results from these methods offer a promising avenue for identifying favorable sources of predictive information to identify the locations of blind geothermal systems and furthering our understanding of complex geothermal feature and label relationships in the Great Basin region and beyond.

15 GEOTHERMAL ENERGY↗

Photon Reconstruction in the Belle II Calorimeter Using Graph Neural Networks

We present the study of a fuzzy clustering algorithm for the Belle II electromagnetic calorimeter using Graph Neural Networks. We use a realistic detector simulation including simulated beam backgrounds and focus on the reconstruction of both isolated and overlapping photons. We find significant improvements of the energy resolution compared to the currently used reconstruction algorithm for both isolated and overlapping photons of more than 30% for photons with energies E γ < 0.5 GeV and high levels of beam backgrounds. Overall, the GNN reconstruction improves the resolution and reduces the tails of the reconstructed energy distribution and therefore is a promising option for the upcoming high luminosity running of Belle II.

Calorimeter↗

Unsupervised Image-Based Classification of Corrosion Severity in Automobile Engine Connecting Rods

Corrosion in engine connecting rods is a critical issue in the automotive industry, potentially leading to catastrophic engine failure, monetary losses, and safety hazards. The labor shortage in the industry further emphasizes the need for fast, accurate, and automated corrosion detection methods to ensure appropriate surface treatments can be applied to restore component integrity. We present an unsupervised image-based framework for classifying corrosion severity in automobile engine connecting rods using short-wave infrared (SWIR) and telecentric grayscale imaging. We employ the structural similarity index measure (SSIM) as a dissimilarity metric and the k-medians clustering algorithm for classification. Our algorithm achieves an overall accuracy of 80.64% for SWIR images, with 100% accuracy in classifying highly corroded samples. For grayscale images, the method attains an overall accuracy of 77.42%, with 90.91% accuracy for highly corroded samples. The method’s ability to work with different imaging modalities and its high accuracy in identifying severe corrosion cases make it a promising tool for automated corrosion assessment in the automotive industry, potentially improving efficiency and safety in engine component maintenance.

42 ENGINEERING↗

‘Flux+Mutability’: a conditional generative approach to one-class classification and anomaly detection

Abstract Anomaly Detection is becoming increasingly popular within the experimental physics community. At experiments such as the Large Hadron Collider, anomaly detection is growing in interest for finding new physics beyond the Standard Model. This paper details the implementation of a novel Machine Learning architecture, called Flux+Mutability, which combines cutting-edge conditional generative models with clustering algorithms. In the ‘flux’ stage we learn the distribution of a reference class. The ‘mutability’ stage at inference addresses if data significantly deviates from the reference class. We demonstrate the validity of our approach and its connection to multiple problems spanning from one-class classification to anomaly detection. In particular, we apply our method to the isolation of neutral showers in an electromagnetic calorimeter and show its performance in detecting anomalous dijets events from standard QCD background. This approach limits assumptions on the reference sample and remains agnostic to the complementary class of objects of a given problem. We describe the possibility of dynamically generating a reference population and defining selection criteria via quantile cuts. Remarkably this flexible architecture can be deployed for a wide range of problems, and applications like multi-class classification or data quality control are left for further exploration.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Renewable hydrogen and ammonia for combined heat and power systems in remote locations: Optimal design and scheduling

Abstract Using hydrogen (H ) and ammonia (NH ) for renewable energy storage has the potential to enable economical power and heat supply with high renewable penetrations, especially in remote locations which are characterized by high energy costs. In this work we assess the economic competitiveness of renewable combined heat and power (CHP) systems in Mahaka HI, Nantucket MA, and Northwest Arctic Borough (NWAB) AK by optimally designing these systems for scenarios in which power and heat can be purchased over a range of historical energy prices as well as when 100% renewable supply is required. We use a combined optimal design and scheduling model which minimizes annualized net present cost by determining optimal technology selection and size simultaneously with optimal schedules for each period of a system operating horizon aggregated from full year hourly resolution data via a consecutive temporal clustering algorithm. We find that renewable generation meets at least 85% of power demands and 75% of heat demands under the lowest energy prices investigated. Higher conventional energy prices lead to increased renewable penetration which is facilitated by renewable NH as a seasonal energy storage medium, as are 100% renewable CHP systems. NH is used for power generation with heat cogeneration in all three locations, as well as directly for heating in NWAB. On an annual cost basis, NH ‐enabled 100% renewable CHP is only 3% more expensive in Mahaka and NWAB than systems which can purchase energy at the lowest prices, while it is 15% more expensive in Nantucket.

Palys, Matthew J.↗

Advanced Image Reconstruction for MCP Detector in Event Mode

A two-step data reduction framework is proposed in this study to reconstruct a radiograph from the data collected with a micro-channel plate (MCP) detector operating under event mode. One clustering algorithm and three neutron event back-tracing models are proposed and evaluated using both example data and a full scan data. The reconstructed radiographs are analyzed, the results of which are used to suggest future development.

Zhang, Chen↗

Measurement and QCD analysis of double-differential inclusive jet cross sections in proton-proton collisions at $\sqrt{s}$ = 13 TeV

A measurement of the inclusive jet production in proton-proton collisions at the LHC at $ \sqrt{s} $ = 13 TeV is presented. The double-differential cross sections are measured as a function of the jet transverse momentum p$_{T}$ and the absolute jet rapidity |y|. The anti-k$_{T}$ clustering algorithm is used with distance parameter of 0.4 (0.7) in a phase space region with jet p$_{T}$ from 97 GeV up to 3.1 TeV and |y| < 2.0. Data collected with the CMS detector are used, corresponding to an integrated luminosity of 36.3 fb$^{-1}$ (33.5 fb$^{-1}$). The measurement is used in a comprehensive QCD analysis at next-to-next-to-leading order, which results in significant improvement in the accuracy of the parton distributions in the proton. Simultaneously, the value of the strong coupling constant at the Z boson mass is extracted as α$_{S}$(m$_{Z}$) = 0.1170±0.0019. For the first time, these data are used in a standard model effective field theory analysis at next-to-leading order, where parton distributions and the QCD parameters are extracted simultaneously with imposed constraints on the Wilson coefficient c$_{1}$ of 4-quark contact interactions.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

ThickBrick: optimal event selection and categorization in high energy physics. Part I. Signal discovery

We provide a prescription called ThickBrick to train optimal machine-learning-based event selectors and categorizers that maximize the statistical significance of a potential signal excess in high energy physics (HEP) experiments, as quantified by any of six different performance measures. For analyses where the signal search is performed in the distribution of some event variables, our prescription ensures that only the information complementary to those event variables is used in event selection and categorization. This eliminates a major misalignment with the physics goals of the analysis (maximizing the significance of an excess) that exists in the training of typical ML-based event selectors and categorizers. In addition, this decorrelation of event selectors from the relevant event variables prevents the background distribution from becoming peaked in the signal region as a result of event selection, thereby ameliorating the challenges imposed on signal searches by systematic uncertainties. Our event selectors (categorizers) use the output of machine-learning-based classifiers as input and apply optimal selection cutoffs (categorization thresholds) that are functions of the event variables being analyzed, as opposed to flat cutoffs (thresholds). These optimal cutoffs and thresholds are learned iteratively, using a novel approach with connections to Lloyd’s k-means clustering algorithm. We provide a public, Python implementation of our prescription, also called ThickBrick, along with usage examples.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Global teleconnections influencing large-scale drought in the United States using SVDI

Understanding recent large-scale drought patterns and the mechanisms producing extreme drought events is vital for future drought forecasts and understanding future drought risks. Increasingly, vapor pressure deficit (VPD) has been used as an important measure of evaporative demand and proxy for drought detection. In this study, VPD is used to calculate the new Standardized VPD Drought Index (SVDI) with NASA North American Land Data Assimilation System (NLDAS) data. Previous studies have shown that SVDI accurately identifies the timing and magnitude short-term droughts in the United States (U.S). In the present study, SVDI is now used to identify large-scale drought patterns between 1980 and 2021 and drought variability driven by selected global teleconnections originating in the Pacific and Atlantic Oceans. Spatial drought characteristics were extracted from SVDI using empirical orthogonal function (EOF) analysis. Then a k-means clustering algorithm was applied to both EOF principal components and primary teleconnections, including the El Nino-Southern Oscillation (ENSO) and Pacific Decadal Oscillation (PDO) to identify drought events driven by the Pacific Ocean. Results show that the SVDI is useful in evaluating large-scale drought variability in the U.S. related to global teleconnections, and that mechanisms influencing summer drought patterns in the Western and Southwestern U.S. are driven by a tropical-extratropical interactions originating in the equatorial Pacific Ocean related to ENSO dynamics with interdecadal variability modulated by PDO. The large-scale droughts in the Central and Southern U.S., like those in 2011 and 2012, on the other hand, are driven by the North Pacific Ocean warm pool during a strong negative PDO, which subsequently influenced variability in the Bermuda-Azores High in the Atlantic Ocean. In summer 2011, the Bermuda-Azores High weakened, reducing the onshore winds and moisture transport along the eastern Gulf of Mexico and contributing to ongoing drought in the region. The Northern Pacific and Atlantic Ocean sea surface temperatures (SSTs) have increased between 1980 and 2021. In conclusion, as SSTs continue to rise in the Northern Pacific Ocean, one consequence of the coupled North Pacific warm pool and atmospheric dynamics, is to increase summer drought variability over a large region in the southern and midwestern U.S. under global warming.

54 ENVIRONMENTAL SCIENCES↗

Haar-Like Wavelets on Hierarchical Trees

Here, discrete wavelet methods, originally formulated in the setting of regularly sampled signals, can be adapted to data defined on a point cloud if some multiresolution structure is imposed on the cloud. A wide variety of hierarchical clustering algorithms can be used for this purpose, and the multiresolution structure obtained can be encoded by a hierarchical tree of subsets of the cloud. Prior work introduced the use of Haar-like bases defined with respect to such trees for approximation and learning tasks on unstructured data. This paper builds on that work in two directions. First, we present an algorithm for constructing Haar-like bases on general discrete hierarchical trees. Second, with an eye towards data compression, we present thresholding techniques for data defined on a point cloud with error controlled in the $L$ $\infty$ norm and in a Hölder-type norm. In a concluding trio of numerical examples, we apply our methods to compress a point cloud dataset, study the tightness of the $L$ $\infty$ error bound, and use thresholding to identify MNIST classifiers with good generalizability.

97 MATHEMATICS AND COMPUTING↗

Using hydrogen and ammonia for renewable energy storage: A geographically comprehensive techno-economic study

Hydrogen and, more recently, ammonia have received worldwide attention as energy storage media. In this work we investigate the economics of using each of these chemicals as well as the two in combination for islanded renewable energy supply systems in 15 American cities representing different climate regions throughout the country. We use an optimal combined capacity planning and scheduling model which minimizes the levelized cost of energy (LCOE) by determining optimal unit selection and size along with unit commitments, production rates, and storage inventories for each period of system operation. These periods are aggregated from full year hourly resolution data via a consecutive temporal clustering algorithm. Ammonia is generally more economical than hydrogen as a single method of energy storage. Additionally, systems which use both hydrogen and ammonia outperform those which use only one storage option and have LCOE between $\$ 0.17$/kWh and $\$ 0.28$/kWh, including full investment in renewable generation infrastructure.

25 ENERGY STORAGE↗

An image-driven machine learning approach to kinetic modeling of a discontinuous precipitation reaction

Micrograph quantification is an essential component of several materials science studies. Machine learning methods, in particular convolutional neural networks, have previously demonstrated performance in image recognition tasks across several disciplines (e.g. materials science, medical imaging, facial recognition). Here, we apply these well-established methods to develop an approach to microstructure quantification for kinetic modeling of a discontinuous precipitation reaction in a case study on the uranium-molybdenum system. Prediction of material processing history based on image data (classification), calculation of area fraction of phases present in the micrographs (segmentation), and kinetic modeling from segmentation results were performed. Results indicate that convolutional neural networks represent microstructure image data well, and segmentation using the k-means clustering algorithm yields results that agree well with manually annotated images. Classification accuracies of original and segmented images are both 94% for a 5-class classification problem. Kinetic modeling results agree well with previously reported data using manual thresholding. The image quantification and kinetic modeling approach developed and presented here aims to reduce researcher bias introduced into the characterization process, and allows for leveraging information in limited image data sets.

36 MATERIALS SCIENCE↗

On the resolution of dual readout calorimeters

Dual readout calorimeters allow state-of-the-art resolutions for hadronic energy measurements. Their various incarnations are leading candidates for the calorimeter systems for future colliders. In this paper, we present a simple formula for the resolution of a dual readout calorimeter, which we verify with a toy simulation and with full simulation results. This formula can help those new to dual readout calorimetry understand its strengths and limitations. The paper also highlights that the dual readout correction works not just to compensate for binding energy loss, but also for energies escaping the calorimeter or clustering algorithm. Formulae are also presented for approximate resolutions and energy scales in terms of different sources of response.

Calorimeters↗

The impact of urban configuration types on urban heat islands, air pollution, CO 2 emissions, and mortality in Europe: a data science approach

The world is becoming increasingly urbanized. As cities around the world continue to grow, it is important for urban planners and policymakers to understand how different urban configuration patterns affect the environment and human health. We aimed at identifying European urban configuration types, based on the Local Climate Zones categories and street design variables from Open Street Map, and evaluating their association with motorized traffic flows, Surface Urban Heat Island (SUHI) intensities, tropospheric nitrogen dioxide (NO 2 ), CO 2 per capita emissions and age-standardized mortality. We considered 946 European cities from 31 countries for the analysis defined in the 2018 Urban Audit database, of which 919 European cities were analysed. Data were collected at a 250 m × 250 m grid cell resolution. We divided all cities into five concentric rings based on the Burgess concentric urban planning model and calculated the mean values of all variables for each ring. First, to identify distinct urban configuration types, we applied the Uniform Manifold Approximation and Projection for Dimension Reduction method, followed by the k-means clustering algorithm. Next, statistical differences in exposures (including SUHI) and mortality between the resulting urban configuration types were evaluated using a Kruskal–Wallis test followed by a post-hoc Dunn's test. We identified four distinct urban configuration types characterising European cities: compact high density (n=246), open low-rise medium density (n=245), open low-rise low density (n=261), and green low density (n=167). Compact high density cities were a small size, had high population densities, and a low availability of natural areas. In contrast, green low-density cities were a large size, had low population densities, and a high availability of natural areas and cycleways. The open low-rise medium and low-density cities were a small to medium size with medium to low population densities and low to moderate availability of green areas. Motorised traffic flows and NO 2 exposure were significantly higher in compact high density and open low rise medium density cities when compared with green low density and open low-rise low density cities. Additionally, green low-density cities had a significantly lower SUHI effect compared with all other urban configuration types. Per person CO 2 emissions were significantly lower in compact high density cities compared with green low density cities. Lastly, green low density cities had significantly lower mortality rates when compared with all other urban configuration types. Our findings indicate that, although the compact city model is more sustainable, European compact cities still face challenges related to poor environmental quality and health. Our results have notable implications for urban and transport planning policies in Europe and contribute to the ongoing discussion on which city models can bring the greatest benefits for the environment, climate, and health.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Data augmentation for disruption prediction via robust surrogate models

The goal of this work is to generate large statistically representative data sets to train machine learning models for disruption prediction provided by data from few existing discharges. Such a comprehensive training database is important to achieve satisfying and reliable prediction results in artificial neural network classifiers. Here, we aim for a robust augmentation of the training database for multivariate time series data using Student t process regression. We apply Student t process regression in a state space formulation via Bayesian filtering to tackle challenges imposed by outliers and noise in the training data set and to reduce the computational complexity. Thus, the method can also be used if the time resolution is high. We use an uncorrelated model for each dimension and impose correlations afterwards via colouring transformations. We demonstrate the efficacy of our approach on plasma diagnostics data of three different disruption classes from the DIII-D tokamak. To evaluate if the distribution of the generated data is similar to the training data, we additionally perform statistical analyses using methods from time series analysis, descriptive statistics and classic machine learning clustering algorithms.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗