Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “clustering algorithm”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Unsupervised Image-Based Classification of Corrosion Severity in Automobile Engine Connecting Rods

Corrosion in engine connecting rods is a critical issue in the automotive industry, potentially leading to catastrophic engine failure, monetary losses, and safety hazards. The labor shortage in the industry further emphasizes the need for fast, accurate, and automated corrosion detection methods to ensure appropriate surface treatments can be applied to restore component integrity. We present an unsupervised image-based framework for classifying corrosion severity in automobile engine connecting rods using short-wave infrared (SWIR) and telecentric grayscale imaging. We employ the structural similarity index measure (SSIM) as a dissimilarity metric and the k-medians clustering algorithm for classification. Our algorithm achieves an overall accuracy of 80.64% for SWIR images, with 100% accuracy in classifying highly corroded samples. For grayscale images, the method attains an overall accuracy of 77.42%, with 90.91% accuracy for highly corroded samples. The method’s ability to work with different imaging modalities and its high accuracy in identifying severe corrosion cases make it a promising tool for automated corrosion assessment in the automotive industry, potentially improving efficiency and safety in engine component maintenance.

42 ENGINEERING↗

‘Flux+Mutability’: a conditional generative approach to one-class classification and anomaly detection

Abstract Anomaly Detection is becoming increasingly popular within the experimental physics community. At experiments such as the Large Hadron Collider, anomaly detection is growing in interest for finding new physics beyond the Standard Model. This paper details the implementation of a novel Machine Learning architecture, called Flux+Mutability, which combines cutting-edge conditional generative models with clustering algorithms. In the ‘flux’ stage we learn the distribution of a reference class. The ‘mutability’ stage at inference addresses if data significantly deviates from the reference class. We demonstrate the validity of our approach and its connection to multiple problems spanning from one-class classification to anomaly detection. In particular, we apply our method to the isolation of neutral showers in an electromagnetic calorimeter and show its performance in detecting anomalous dijets events from standard QCD background. This approach limits assumptions on the reference sample and remains agnostic to the complementary class of objects of a given problem. We describe the possibility of dynamically generating a reference population and defining selection criteria via quantile cuts. Remarkably this flexible architecture can be deployed for a wide range of problems, and applications like multi-class classification or data quality control are left for further exploration.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Renewable hydrogen and ammonia for combined heat and power systems in remote locations: Optimal design and scheduling

Abstract Using hydrogen (H ) and ammonia (NH ) for renewable energy storage has the potential to enable economical power and heat supply with high renewable penetrations, especially in remote locations which are characterized by high energy costs. In this work we assess the economic competitiveness of renewable combined heat and power (CHP) systems in Mahaka HI, Nantucket MA, and Northwest Arctic Borough (NWAB) AK by optimally designing these systems for scenarios in which power and heat can be purchased over a range of historical energy prices as well as when 100% renewable supply is required. We use a combined optimal design and scheduling model which minimizes annualized net present cost by determining optimal technology selection and size simultaneously with optimal schedules for each period of a system operating horizon aggregated from full year hourly resolution data via a consecutive temporal clustering algorithm. We find that renewable generation meets at least 85% of power demands and 75% of heat demands under the lowest energy prices investigated. Higher conventional energy prices lead to increased renewable penetration which is facilitated by renewable NH as a seasonal energy storage medium, as are 100% renewable CHP systems. NH is used for power generation with heat cogeneration in all three locations, as well as directly for heating in NWAB. On an annual cost basis, NH ‐enabled 100% renewable CHP is only 3% more expensive in Mahaka and NWAB than systems which can purchase energy at the lowest prices, while it is 15% more expensive in Nantucket.

Palys, Matthew J.↗

Advanced Image Reconstruction for MCP Detector in Event Mode

A two-step data reduction framework is proposed in this study to reconstruct a radiograph from the data collected with a micro-channel plate (MCP) detector operating under event mode. One clustering algorithm and three neutron event back-tracing models are proposed and evaluated using both example data and a full scan data. The reconstructed radiographs are analyzed, the results of which are used to suggest future development.

Zhang, Chen↗

Measurement and QCD analysis of double-differential inclusive jet cross sections in proton-proton collisions at $\sqrt{s}$ = 13 TeV

A measurement of the inclusive jet production in proton-proton collisions at the LHC at $ \sqrt{s} $ = 13 TeV is presented. The double-differential cross sections are measured as a function of the jet transverse momentum p$_{T}$ and the absolute jet rapidity |y|. The anti-k$_{T}$ clustering algorithm is used with distance parameter of 0.4 (0.7) in a phase space region with jet p$_{T}$ from 97 GeV up to 3.1 TeV and |y| < 2.0. Data collected with the CMS detector are used, corresponding to an integrated luminosity of 36.3 fb$^{-1}$ (33.5 fb$^{-1}$). The measurement is used in a comprehensive QCD analysis at next-to-next-to-leading order, which results in significant improvement in the accuracy of the parton distributions in the proton. Simultaneously, the value of the strong coupling constant at the Z boson mass is extracted as α$_{S}$(m$_{Z}$) = 0.1170±0.0019. For the first time, these data are used in a standard model effective field theory analysis at next-to-leading order, where parton distributions and the QCD parameters are extracted simultaneously with imposed constraints on the Wilson coefficient c$_{1}$ of 4-quark contact interactions.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

ThickBrick: optimal event selection and categorization in high energy physics. Part I. Signal discovery

We provide a prescription called ThickBrick to train optimal machine-learning-based event selectors and categorizers that maximize the statistical significance of a potential signal excess in high energy physics (HEP) experiments, as quantified by any of six different performance measures. For analyses where the signal search is performed in the distribution of some event variables, our prescription ensures that only the information complementary to those event variables is used in event selection and categorization. This eliminates a major misalignment with the physics goals of the analysis (maximizing the significance of an excess) that exists in the training of typical ML-based event selectors and categorizers. In addition, this decorrelation of event selectors from the relevant event variables prevents the background distribution from becoming peaked in the signal region as a result of event selection, thereby ameliorating the challenges imposed on signal searches by systematic uncertainties. Our event selectors (categorizers) use the output of machine-learning-based classifiers as input and apply optimal selection cutoffs (categorization thresholds) that are functions of the event variables being analyzed, as opposed to flat cutoffs (thresholds). These optimal cutoffs and thresholds are learned iteratively, using a novel approach with connections to Lloyd’s k-means clustering algorithm. We provide a public, Python implementation of our prescription, also called ThickBrick, along with usage examples.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Global teleconnections influencing large-scale drought in the United States using SVDI

Understanding recent large-scale drought patterns and the mechanisms producing extreme drought events is vital for future drought forecasts and understanding future drought risks. Increasingly, vapor pressure deficit (VPD) has been used as an important measure of evaporative demand and proxy for drought detection. In this study, VPD is used to calculate the new Standardized VPD Drought Index (SVDI) with NASA North American Land Data Assimilation System (NLDAS) data. Previous studies have shown that SVDI accurately identifies the timing and magnitude short-term droughts in the United States (U.S). In the present study, SVDI is now used to identify large-scale drought patterns between 1980 and 2021 and drought variability driven by selected global teleconnections originating in the Pacific and Atlantic Oceans. Spatial drought characteristics were extracted from SVDI using empirical orthogonal function (EOF) analysis. Then a k-means clustering algorithm was applied to both EOF principal components and primary teleconnections, including the El Nino-Southern Oscillation (ENSO) and Pacific Decadal Oscillation (PDO) to identify drought events driven by the Pacific Ocean. Results show that the SVDI is useful in evaluating large-scale drought variability in the U.S. related to global teleconnections, and that mechanisms influencing summer drought patterns in the Western and Southwestern U.S. are driven by a tropical-extratropical interactions originating in the equatorial Pacific Ocean related to ENSO dynamics with interdecadal variability modulated by PDO. The large-scale droughts in the Central and Southern U.S., like those in 2011 and 2012, on the other hand, are driven by the North Pacific Ocean warm pool during a strong negative PDO, which subsequently influenced variability in the Bermuda-Azores High in the Atlantic Ocean. In summer 2011, the Bermuda-Azores High weakened, reducing the onshore winds and moisture transport along the eastern Gulf of Mexico and contributing to ongoing drought in the region. The Northern Pacific and Atlantic Ocean sea surface temperatures (SSTs) have increased between 1980 and 2021. In conclusion, as SSTs continue to rise in the Northern Pacific Ocean, one consequence of the coupled North Pacific warm pool and atmospheric dynamics, is to increase summer drought variability over a large region in the southern and midwestern U.S. under global warming.

54 ENVIRONMENTAL SCIENCES↗

Haar-Like Wavelets on Hierarchical Trees

Here, discrete wavelet methods, originally formulated in the setting of regularly sampled signals, can be adapted to data defined on a point cloud if some multiresolution structure is imposed on the cloud. A wide variety of hierarchical clustering algorithms can be used for this purpose, and the multiresolution structure obtained can be encoded by a hierarchical tree of subsets of the cloud. Prior work introduced the use of Haar-like bases defined with respect to such trees for approximation and learning tasks on unstructured data. This paper builds on that work in two directions. First, we present an algorithm for constructing Haar-like bases on general discrete hierarchical trees. Second, with an eye towards data compression, we present thresholding techniques for data defined on a point cloud with error controlled in the $L$ $\infty$ norm and in a Hölder-type norm. In a concluding trio of numerical examples, we apply our methods to compress a point cloud dataset, study the tightness of the $L$ $\infty$ error bound, and use thresholding to identify MNIST classifiers with good generalizability.

97 MATHEMATICS AND COMPUTING↗

Using hydrogen and ammonia for renewable energy storage: A geographically comprehensive techno-economic study

Hydrogen and, more recently, ammonia have received worldwide attention as energy storage media. In this work we investigate the economics of using each of these chemicals as well as the two in combination for islanded renewable energy supply systems in 15 American cities representing different climate regions throughout the country. We use an optimal combined capacity planning and scheduling model which minimizes the levelized cost of energy (LCOE) by determining optimal unit selection and size along with unit commitments, production rates, and storage inventories for each period of system operation. These periods are aggregated from full year hourly resolution data via a consecutive temporal clustering algorithm. Ammonia is generally more economical than hydrogen as a single method of energy storage. Additionally, systems which use both hydrogen and ammonia outperform those which use only one storage option and have LCOE between $\$ 0.17$/kWh and $\$ 0.28$/kWh, including full investment in renewable generation infrastructure.

25 ENERGY STORAGE↗

An image-driven machine learning approach to kinetic modeling of a discontinuous precipitation reaction

Micrograph quantification is an essential component of several materials science studies. Machine learning methods, in particular convolutional neural networks, have previously demonstrated performance in image recognition tasks across several disciplines (e.g. materials science, medical imaging, facial recognition). Here, we apply these well-established methods to develop an approach to microstructure quantification for kinetic modeling of a discontinuous precipitation reaction in a case study on the uranium-molybdenum system. Prediction of material processing history based on image data (classification), calculation of area fraction of phases present in the micrographs (segmentation), and kinetic modeling from segmentation results were performed. Results indicate that convolutional neural networks represent microstructure image data well, and segmentation using the k-means clustering algorithm yields results that agree well with manually annotated images. Classification accuracies of original and segmented images are both 94% for a 5-class classification problem. Kinetic modeling results agree well with previously reported data using manual thresholding. The image quantification and kinetic modeling approach developed and presented here aims to reduce researcher bias introduced into the characterization process, and allows for leveraging information in limited image data sets.

36 MATERIALS SCIENCE↗

On the resolution of dual readout calorimeters

Dual readout calorimeters allow state-of-the-art resolutions for hadronic energy measurements. Their various incarnations are leading candidates for the calorimeter systems for future colliders. In this paper, we present a simple formula for the resolution of a dual readout calorimeter, which we verify with a toy simulation and with full simulation results. This formula can help those new to dual readout calorimetry understand its strengths and limitations. The paper also highlights that the dual readout correction works not just to compensate for binding energy loss, but also for energies escaping the calorimeter or clustering algorithm. Formulae are also presented for approximate resolutions and energy scales in terms of different sources of response.

Calorimeters↗

The impact of urban configuration types on urban heat islands, air pollution, CO 2 emissions, and mortality in Europe: a data science approach

The world is becoming increasingly urbanized. As cities around the world continue to grow, it is important for urban planners and policymakers to understand how different urban configuration patterns affect the environment and human health. We aimed at identifying European urban configuration types, based on the Local Climate Zones categories and street design variables from Open Street Map, and evaluating their association with motorized traffic flows, Surface Urban Heat Island (SUHI) intensities, tropospheric nitrogen dioxide (NO 2 ), CO 2 per capita emissions and age-standardized mortality. We considered 946 European cities from 31 countries for the analysis defined in the 2018 Urban Audit database, of which 919 European cities were analysed. Data were collected at a 250 m × 250 m grid cell resolution. We divided all cities into five concentric rings based on the Burgess concentric urban planning model and calculated the mean values of all variables for each ring. First, to identify distinct urban configuration types, we applied the Uniform Manifold Approximation and Projection for Dimension Reduction method, followed by the k-means clustering algorithm. Next, statistical differences in exposures (including SUHI) and mortality between the resulting urban configuration types were evaluated using a Kruskal–Wallis test followed by a post-hoc Dunn's test. We identified four distinct urban configuration types characterising European cities: compact high density (n=246), open low-rise medium density (n=245), open low-rise low density (n=261), and green low density (n=167). Compact high density cities were a small size, had high population densities, and a low availability of natural areas. In contrast, green low-density cities were a large size, had low population densities, and a high availability of natural areas and cycleways. The open low-rise medium and low-density cities were a small to medium size with medium to low population densities and low to moderate availability of green areas. Motorised traffic flows and NO 2 exposure were significantly higher in compact high density and open low rise medium density cities when compared with green low density and open low-rise low density cities. Additionally, green low-density cities had a significantly lower SUHI effect compared with all other urban configuration types. Per person CO 2 emissions were significantly lower in compact high density cities compared with green low density cities. Lastly, green low density cities had significantly lower mortality rates when compared with all other urban configuration types. Our findings indicate that, although the compact city model is more sustainable, European compact cities still face challenges related to poor environmental quality and health. Our results have notable implications for urban and transport planning policies in Europe and contribute to the ongoing discussion on which city models can bring the greatest benefits for the environment, climate, and health.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Data augmentation for disruption prediction via robust surrogate models

The goal of this work is to generate large statistically representative data sets to train machine learning models for disruption prediction provided by data from few existing discharges. Such a comprehensive training database is important to achieve satisfying and reliable prediction results in artificial neural network classifiers. Here, we aim for a robust augmentation of the training database for multivariate time series data using Student t process regression. We apply Student t process regression in a state space formulation via Bayesian filtering to tackle challenges imposed by outliers and noise in the training data set and to reduce the computational complexity. Thus, the method can also be used if the time resolution is high. We use an uncorrelated model for each dimension and impose correlations afterwards via colouring transformations. We demonstrate the efficacy of our approach on plasma diagnostics data of three different disruption classes from the DIII-D tokamak. To evaluate if the distribution of the generated data is similar to the training data, we additionally perform statistical analyses using methods from time series analysis, descriptive statistics and classic machine learning clustering algorithms.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Synoptic Meteorology Explains Temperate Forest Carbon Uptake

Abstract While substantial attention has been paid to the effects of both global climate oscillations and local meteorological conditions on the interannual variability of ecosystem carbon exchange, the relationship between the interannual variability of synoptic meteorology and ecosystem carbon exchange has not been well studied. Here we use a clustering algorithm to identify a summertime cyclonic precipitation system northwest of the Great Lakes to determine (a) the association at a daily scale between the occurrence of this system and the local meteorology and net ecosystem exchange at three Great Lakes region forested eddy covariance sites and (b) the association between the seasonal prevalence of this system and the summertime net ecosystem exchange of these sites. We find that temperature, in addition to precipitation and cloud cover, is an important explanatory factor for the suppression of net ecosystem productivity that occurs during these cyclonic events in this region. In addition, the prevalence of this cyclonic system can explain a significant proportion of the interannual variability in summertime forest ecosystem exchange in this region. This explanatory power is not due to a simple accumulation of low‐productivity days that cooccur with this meteorological event, but rather a broader association between the frequency of these events and several aspects of prevailing seasonal conditions. This work demonstrates the usefulness of conceptualizing meteorology in terms of synoptic systems for explaining the interannual variability of regional carbon fluxes.

Randazzo, Nina A.↗

Coherent correlation imaging for resolving fluctuating states of matter

Fluctuations and stochastic transitions are ubiquitous in nanometre-scale systems, especially in the presence of disorder. However, their direct observation has so far been impeded by a seemingly fundamental, signal-limited compromise between spatial and temporal resolution. Here we develop coherent correlation imaging (CCI) to overcome this dilemma. Our method begins by classifying recorded camera frames in Fourier space. Contrast and spatial resolution emerge by averaging selectively over same-state frames. Temporal resolution down to the acquisition time of a single frame arises independently from an exceptionally low misclassification rate, which we achieve by combining a correlation-based similarity metric with a modified, iterative hierarchical clustering algorithm. We apply CCI to study previously inaccessible magnetic fluctuations in a highly degenerate magnetic stripe domain state with nanometre-scale resolution. We uncover an intricate network of transitions between more than 30 discrete states. Our spatiotemporal data enable us to reconstruct the pinning energy landscape and to thereby explain the dynamics observed on a microscopic level. CCI massively expands the potential of emerging high-coherence X-ray sources and paves the way for addressing large fundamental questions such as the contribution of pinning and topology in phase transitions and the role of spin and charge order fluctuations in high-temperature superconductivity.

36 MATERIALS SCIENCE↗

The Sagittarius stream in Gaia Early Data Release 3 and the origin of the bifurcations

The Sagittarius dwarf spheroidal (Sgr) is a dissolving galaxy being tidally disrupted by the Milky Way (MW). Its stellar stream still poses serious modelling challenges, which hinders our ability to use it effectively as a prospective probe of the MW gravitational potential at large radii. Our goal is to construct the largest and most stringent sample of stars in the stream with which we can advance our understanding of the Sgr-MW interaction, focusing on the characterization of the bifurcations. We improved on previous methods based on the use of the wavelet transform to systematically search for the kinematic signature of the Sgr stream throughout the whole sky in the Gaia data. We then refined our selection via the use of a clustering algorithm on the statistical properties of the colour-magnitude diagrams. Our final sample contains more than 700 000 candidate stars and is three times larger than previous Gaia samples. With it, we have been able to detect the bifurcation of the stream in both the northern and southern hemispheres, which requires four branches (two bright and two faint) to fully describe the system. We present the detailed proper motion distribution of the trailing arm as a function of the angular coordinate along the stream, showing, for the first time, the presence of a sharp edge (on the side of the small proper motions) beyond which there are no Sgr stars. We also characterize the correlation between kinematics and distance. Finally, the chemical analysis of our sample shows that the faint branch of the bifurcation is more metal poor than the bright. We provide analytical descriptions for the proper motion trends as well as for the sky distribution of the four branches of the stream. Based on our analysis, we interpret the bifurcations as a misaligned overlap of the material stripped at the antepenultimate pericentre (faint branches) with the stars ejected at the penultimate pericentre (bright branch), given that Sgr just went through its perigalacticon. The source of this misalignment is still unknown, but we argue that models with some internal rotation in the progenitor – at least during the time of stripping of the stars that are now in the faint branches – are worth exploring.

79 ASTRONOMY AND ASTROPHYSICS↗

StarHorse results for spectroscopic surveys and Gaia DR3: Chrono-chemical populations in the solar vicinity, the genuine thick disk, and young alpha-rich stars

The Gaia mission has provided an invaluable wealth of astrometric data for more than a billion stars in our Galaxy. The synergy between Gaia astrometry, photometry, and spectroscopic surveys gives us comprehensive information about the Milky Way. Using the Bayesian isochrone-fitting code StarHorse, we derive distances and extinctions for more than 10 million unique stars listed in both Gaia Data Release 3 and public spectroscopic surveys: 557 559 in GALAH+ DR3, 4 531 028 in LAMOST DR7 LRS, 347 535 in LAMOST DR7 MRS, 562 424 in APOGEE DR17, 471 490 in RAVE DR6, 249 991 in SDSS DR12 (optical spectra from BOSS and SEGUE), 67 562 in the Gaia-ESO DR5 survey, and 4 211 087 in the Gaia RVS part of the Gaia DR3 release. StarHorse can increase the precision of distance and extinction measurements where Gaia parallaxes alone would be uncertain. We used StarHorse for the first time to derive stellar ages for main-sequence turnoff and subgiant branch stars, around 2.5 million stars, with age uncertainties typically around 30%; the uncertainties drop to 15% for subgiant-branch-only stars, depending on the resolution of the survey. With the derived ages in hand, we investigated the chemical-age relations. In particular, the α and neutron-capture element ratios versus age in the solar neighbourhood show trends similar to previous works, validating our ages. We used the chemical abundances from local subgiant samples of GALAH DR3, APOGEE DR17, and LAMOST MRS DR7 to map groups with similar chemical compositions and StarHorse ages, using the dimensionality reduction technique t-SNE and the clustering algorithm HDBSCAN. We identify three distinct groups in all three samples, confirmed by their kinematic properties: the genuine chemical thick disk, the thin disk, and a considerable number of young alpha-rich stars (427) that are also a part of the delivered catalogues. We confirm that the genuine thick disk’s kinematics and age properties are radically different from those of the thin disk and compatible with high-redshift (z ≈ 2) star-forming disks with high dispersion velocities. We also find a few extra chemical populations in GALAH DR3 thanks to the availability of neutron-capture element information.

79 ASTRONOMY AND ASTROPHYSICS↗