Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “clustering algorithm”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12

Synoptic Meteorology Explains Temperate Forest Carbon Uptake

Abstract While substantial attention has been paid to the effects of both global climate oscillations and local meteorological conditions on the interannual variability of ecosystem carbon exchange, the relationship between the interannual variability of synoptic meteorology and ecosystem carbon exchange has not been well studied. Here we use a clustering algorithm to identify a summertime cyclonic precipitation system northwest of the Great Lakes to determine (a) the association at a daily scale between the occurrence of this system and the local meteorology and net ecosystem exchange at three Great Lakes region forested eddy covariance sites and (b) the association between the seasonal prevalence of this system and the summertime net ecosystem exchange of these sites. We find that temperature, in addition to precipitation and cloud cover, is an important explanatory factor for the suppression of net ecosystem productivity that occurs during these cyclonic events in this region. In addition, the prevalence of this cyclonic system can explain a significant proportion of the interannual variability in summertime forest ecosystem exchange in this region. This explanatory power is not due to a simple accumulation of low‐productivity days that cooccur with this meteorological event, but rather a broader association between the frequency of these events and several aspects of prevailing seasonal conditions. This work demonstrates the usefulness of conceptualizing meteorology in terms of synoptic systems for explaining the interannual variability of regional carbon fluxes.

Randazzo, Nina A.↗

Coherent correlation imaging for resolving fluctuating states of matter

Fluctuations and stochastic transitions are ubiquitous in nanometre-scale systems, especially in the presence of disorder. However, their direct observation has so far been impeded by a seemingly fundamental, signal-limited compromise between spatial and temporal resolution. Here we develop coherent correlation imaging (CCI) to overcome this dilemma. Our method begins by classifying recorded camera frames in Fourier space. Contrast and spatial resolution emerge by averaging selectively over same-state frames. Temporal resolution down to the acquisition time of a single frame arises independently from an exceptionally low misclassification rate, which we achieve by combining a correlation-based similarity metric with a modified, iterative hierarchical clustering algorithm. We apply CCI to study previously inaccessible magnetic fluctuations in a highly degenerate magnetic stripe domain state with nanometre-scale resolution. We uncover an intricate network of transitions between more than 30 discrete states. Our spatiotemporal data enable us to reconstruct the pinning energy landscape and to thereby explain the dynamics observed on a microscopic level. CCI massively expands the potential of emerging high-coherence X-ray sources and paves the way for addressing large fundamental questions such as the contribution of pinning and topology in phase transitions and the role of spin and charge order fluctuations in high-temperature superconductivity.

36 MATERIALS SCIENCE↗

The Sagittarius stream in Gaia Early Data Release 3 and the origin of the bifurcations

The Sagittarius dwarf spheroidal (Sgr) is a dissolving galaxy being tidally disrupted by the Milky Way (MW). Its stellar stream still poses serious modelling challenges, which hinders our ability to use it effectively as a prospective probe of the MW gravitational potential at large radii. Our goal is to construct the largest and most stringent sample of stars in the stream with which we can advance our understanding of the Sgr-MW interaction, focusing on the characterization of the bifurcations. We improved on previous methods based on the use of the wavelet transform to systematically search for the kinematic signature of the Sgr stream throughout the whole sky in the Gaia data. We then refined our selection via the use of a clustering algorithm on the statistical properties of the colour-magnitude diagrams. Our final sample contains more than 700 000 candidate stars and is three times larger than previous Gaia samples. With it, we have been able to detect the bifurcation of the stream in both the northern and southern hemispheres, which requires four branches (two bright and two faint) to fully describe the system. We present the detailed proper motion distribution of the trailing arm as a function of the angular coordinate along the stream, showing, for the first time, the presence of a sharp edge (on the side of the small proper motions) beyond which there are no Sgr stars. We also characterize the correlation between kinematics and distance. Finally, the chemical analysis of our sample shows that the faint branch of the bifurcation is more metal poor than the bright. We provide analytical descriptions for the proper motion trends as well as for the sky distribution of the four branches of the stream. Based on our analysis, we interpret the bifurcations as a misaligned overlap of the material stripped at the antepenultimate pericentre (faint branches) with the stars ejected at the penultimate pericentre (bright branch), given that Sgr just went through its perigalacticon. The source of this misalignment is still unknown, but we argue that models with some internal rotation in the progenitor – at least during the time of stripping of the stars that are now in the faint branches – are worth exploring.

79 ASTRONOMY AND ASTROPHYSICS↗

StarHorse results for spectroscopic surveys and Gaia DR3: Chrono-chemical populations in the solar vicinity, the genuine thick disk, and young alpha-rich stars

The Gaia mission has provided an invaluable wealth of astrometric data for more than a billion stars in our Galaxy. The synergy between Gaia astrometry, photometry, and spectroscopic surveys gives us comprehensive information about the Milky Way. Using the Bayesian isochrone-fitting code StarHorse, we derive distances and extinctions for more than 10 million unique stars listed in both Gaia Data Release 3 and public spectroscopic surveys: 557 559 in GALAH+ DR3, 4 531 028 in LAMOST DR7 LRS, 347 535 in LAMOST DR7 MRS, 562 424 in APOGEE DR17, 471 490 in RAVE DR6, 249 991 in SDSS DR12 (optical spectra from BOSS and SEGUE), 67 562 in the Gaia-ESO DR5 survey, and 4 211 087 in the Gaia RVS part of the Gaia DR3 release. StarHorse can increase the precision of distance and extinction measurements where Gaia parallaxes alone would be uncertain. We used StarHorse for the first time to derive stellar ages for main-sequence turnoff and subgiant branch stars, around 2.5 million stars, with age uncertainties typically around 30%; the uncertainties drop to 15% for subgiant-branch-only stars, depending on the resolution of the survey. With the derived ages in hand, we investigated the chemical-age relations. In particular, the α and neutron-capture element ratios versus age in the solar neighbourhood show trends similar to previous works, validating our ages. We used the chemical abundances from local subgiant samples of GALAH DR3, APOGEE DR17, and LAMOST MRS DR7 to map groups with similar chemical compositions and StarHorse ages, using the dimensionality reduction technique t-SNE and the clustering algorithm HDBSCAN. We identify three distinct groups in all three samples, confirmed by their kinematic properties: the genuine chemical thick disk, the thin disk, and a considerable number of young alpha-rich stars (427) that are also a part of the delivered catalogues. We confirm that the genuine thick disk’s kinematics and age properties are radically different from those of the thin disk and compatible with high-redshift (z ≈ 2) star-forming disks with high dispersion velocities. We also find a few extra chemical populations in GALAH DR3 thanks to the availability of neutron-capture element information.

79 ASTRONOMY AND ASTROPHYSICS↗

Mobility Gaps between Low-Income and Not Low-Income Households: A Case Study in New York State

Understanding the travel challenges faced by low-income residents has always been and continues to be one of the most important transportation equity topics. This study aims to explore the mobility gaps between low-income households (HHs) and not low-income HHs, and how the gaps vary within different socio-demographic population groups in New York State (NYS). The latest National Household Travel Survey data was used as the primary data source for the analysis. The study first employed the K-prototype clustering algorithm to categorize the HHs in NYS based on their socio-demographic attributes. Five population groups were identified based on nine different household (HH) features such as HH size, vehicle ownership, and elderly status of its members. Then, the mobility differences, measured by trip frequency, trip distance, travel time, and person miles traveled, were examined among the five population groups. Results suggest that the individuals in low-income HHs consistently took fewer trips and made shorter trips compared to their not low-income counterparts in NYS. The travel distance gaps were most obvious among white HHs with more vehicles than drivers. In addition, while the population from low-income HHs made shorter trips on average (2.7 mi shorter per trip), they experienced longer travel time than those from not low-income HHs (1.8 min longer per trip). These key findings provide a deeper understanding of the travel behavior disparities between low-income and not low-income households. The findings could also support policymakers and transportation planners in addressing the critical needs of residents in low-income households in NYS and provide inputs for designing a more equitable transportation system.

Liu, Yuandong↗

Graphic contrastive learning analyses of discontinuous molecular dynamics simulations: Study of protein folding upon adsorption

A comprehensive understanding of the interfacial behaviors of biomolecules holds great significance in the development of biomaterials and biosensing technologies. In this work, we used discontinuous molecular dynamics (DMD) simulations and graphic contrastive learning analysis to study the adsorption of ubiquitin protein on a graphene surface. Our high-throughput DMD simulations can explore the whole protein adsorption process including the protein structural evolution with sufficient accuracy. Contrastive learning was employed to train a protein contact map feature extractor aiming at generating contact map feature vectors. Subsequently, these features were grouped using the k-means clustering algorithm to identify the protein structural transition stages throughout the adsorption process. The machine learning analysis can illustrate the dynamics of protein structural changes, including the pathway and the rate-limiting step. Our study indicated that the protein–graphene surface hydrophobic interactions and the π–π stacking were crucial to the seven-stage adsorption process. Upon adsorption, the secondary structure and tertiary structure of ubiquitin disintegrated. The unfolding stages obtained by contrastive learning-based algorithm were not only consistent with the detailed analyses of protein structures but also provided more hidden information about the transition states and pathway of protein adsorption process and structural dynamics. Our combination of efficient DMD simulations and machine learning analysis could be a valuable approach to studying the interfacial behaviors of biomolecules.

97 MATHEMATICS AND COMPUTING↗

Reaction–drift–diffusion models from master equations: application to material defects

We present a general method to produce well-conditioned continuum reaction–drift–diffusion equations directly from master equations on a discrete, periodic state space. We assume the underlying data to be kinetic Monte Carlo models (i.e. continuous-time Markov chains) produced from atomic sampling of point defects in locally periodic environments, such as perfect lattices, ordered surface structures or dislocation cores, possibly under the influence of a slowly varying external field. Our approach also applies to any discrete, periodic Markov chain. Here, the analysis identifies a previously omitted non-equilibrium drift term, present even in the absence of external forces, which can compete in magnitude with the reaction rates, thus being essential to correctly capture the kinetics. To remove fast modes which hinder time integration, we use a generalized Bloch relation to efficiently calculate the eigenspectrum of the master equation. A well conditioned continuum equation then emerges by searching for spectral gaps in the long wavelength limit, using an established kinetic clustering algorithm to define a proper reduced, Markovian state space.

36 MATERIALS SCIENCE↗

Variability in terrestrial characteristics and erosion rates on the Alaskan Beaufort Sea coast

Abstract Arctic coastal environments are eroding and rapidly changing. A lack of pan-Arctic observations limits our ability to understand controls on coastal erosion rates across the entire Arctic region. Here, we capitalize on an abundance of geospatial and remotely sensed data, in addition to model output, from the North Slope of Alaska to identify relationships between historical erosion rates and landscape characteristics to guide future modeling and observational efforts across the Arctic. Using existing datasets from the Alaska Beaufort Sea coast and a hierarchical clustering algorithm, we developed a set of 16 coastal typologies that captures the defining characteristics of environments susceptible to coastal erosion. Relationships between landscape characteristics and historical erosion rates show that no single variable alone is a good predictor of erosion rates. Variability in erosion rate decreases with increasing coastal elevation, but erosion rate magnitudes are highest for intermediate elevations. Areas along the Alaskan Beaufort Sea coast (ABSC) protected by barrier islands showed a three times lower erosion rate on average, suggesting that barrier islands are critical to maintaining mainland shore position. Finally, typologies with the highest erosion rates are not broadly representative of the ABSC and are generally associated with low elevation, north- to northeast-facing shorelines, a peaty pebbly silty lithology, and glaciomarine deposits with high ice content. All else being equal, warmer permafrost is also associated with higher erosion rates, suggesting that warming permafrost temperatures may contribute to higher future erosion rates on permafrost coasts. The suite of typologies can be used to guide future modeling and observational efforts by quantifying the distribution of coastlines with specific landscape characteristics and erosion rates.

54 ENVIRONMENTAL SCIENCES↗

Learning nuclear cross sections across the chart of nuclides with graph neural networks

We explore the use of deep learning techniques to learn how nuclear cross sections change as we add or remove protons and neutrons. As a proof of principle, we focus on the neutron-induced reactions in the fast energy regime. Our approach follows a two-stage learning framework. First, we apply representation learning to encode cross section data into a latent space using either variational autoencoders (VAEs) or implicit neural representations (INRs). Then, we train graph neural networks (GNNs) on the resulting embeddings to predict missing values across the nuclear chart by leveraging the topological structure of neighboring isotopes. We demonstrate accurate cross section predictions within a 9 × 9 block of missing nuclei. We also find that the optimal GNN training strategy depends on the type of latent representation used, with VAE embeddings performing best under end-to-end optimization in the original space, while INR embeddings achieve better results when the GNN is trained only in the latent space. Furthermore, using clustering algorithms, we map groups of latent vectors into regions of the nuclear chart and show that VAEs and INRs can discover some of the neutron magic numbers. These findings suggest that deep-learning models based on the representation encoding of cross sections combined with graph neural networks hold significant potential in augmenting nuclear theory models, e.g., by providing reliable estimates of covariances of cross sections, including cross-material covariances.

Machine learning↗

From hard spheres to hard-core spins

A system of hard spheres exhibits physics that is controlled only by their density. This comes about because the interaction energy is either infinite or zero, so all allowed configurations have exactly the same energy. The low-density phase is liquid, while the high-density phase is crystalline, an example of “order by disorder” as it is driven purely by entropic considerations. Here we study a family of hard spin models, which we call hard-core spin models, where we replace the translational degrees of freedom of hard spheres with the orientational degrees of freedom of lattice spins. Their hard-core interaction serves analogously to divide configurations of the many spin system into allowed and disallowed sectors. We present detailed results on the square lattice in d = 2 for a set of models with $\mathbb{Z}_n$ symmetry, which generalize Potts models, and their U(1) limits, for ferromagnetic and antiferromagnetic senses of the interaction, which we refer to as exclusion and inclusion models. As the exclusion and inclusion angles are varied, we find a Kosterlitz-Thouless phase transition between a disordered phase and an ordered phase with quasi-long-ranged order, which is the form order by disorder takes in these systems. These results follow from a set of height representations, an ergodic cluster algorithm, and transfer matrix calculations.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

Representative Period Selection for Robust Capacity Expansion Planning in Low-carbon Grids

With the increasing urgency to decarbonize power systems, while mitigating extreme events, capacity expansion models can play a vital role in reliably planning the expansion of power systems and facilitating the integration of renewable energy sources. Optimizing capacity expansion generally involves selecting surrogate representative days from forecasts of load and the generation profiles of variable renewable energy resources. To properly select those representative days, we propose a novel input-based approach in combination with the k-means clustering algorithm that utilize three unique operational inputs: load shedding, renewable curtailment, and transmission congestion. The proposed method allows for more robust and cost-effective capacity planning. The method is validated using a capacity expansion model and a production cost model based on California Independent System Operator (CAISO)'s decarbonization goals, and results in reduced costs and drastically lower load shedding.

Anderson, Osten P.↗

Crash Risks Evaluation of Urban Expressways: A Case Study in Shanghai

We report that proactive traffic safety management systems can reduce crashes by identifying crash precursors, evaluating real-time crash risks, and implementing suitable interventions. The basic prerequisite for developing such a system is to propose a reliable crash risk evaluation model that takes real-time traffic flow data as input. Previous studies have primarily focused on real-time crash prediction using some statistical or machine-learning methods. However, further quantitative evaluation and classification of crash risks have been ignored. In this study, we conduct a systematic crash risk evaluation workflow, including crash risk prediction, crash risk quantification, and crash risk classification. Specifically, the crash risk prediction using an extended logit model is proposed, from which CAS, CSD, UAS, DAS, DTV are identified to be contributing factors of crash risks. Then a crash risk quantification model based on the parameter evaluation of the extended logit model is developed. The crash risks of urban expressways and their spatial-temporal evolution trends are quantified. Finally, the crash risks are classified into high crash risk level, moderate crash risk level, and low crash risk level by the k-means cluster algorithm. Then the threshold boundaries of different crash risk levels are determined. The research results provide a proactive guidance for traffic safety management of urban expressways.

33 ADVANCED PROPULSION SYSTEMS↗

Resilience-Oriented DG Siting and Sizing Considering Stochastic Scenario Reduction

In this paper, a fuel-based distributed generator (DG) allocation strategy is proposed to enhance the distribution system resilience against extreme weather. The long-term planning problem is formulated as a two-stage stochastic mixed-integer programming (SMIP). The first stage is to make decisions of DG siting and sizing under the given budget constraint. In the second stage, a post-extreme-event-restoration (PEER) is employed to minimize the operating cost in an uncertain fault scenario. In particular, this study proposes a method to select the most representative scenarios for the SMIP. First, a Monte Carlo Simulation (MCS) is introduced to generate sufficient scenarios considering random fault locations and load profiles. Then, the number of scenarios is reduced by the K-means clustering algorithm. The advantage of scenario reduction is to make a trade-off between accuracy and computational efficiency. Finally, the SMIP is solved by the progressive hedging algorithm. Here, the case studies of the IEEE 33-bus and 123-bus test systems demonstrate the effectiveness of the proposed algorithm in reducing the expected energy not served (EENS), which is a critical criterion of resilience.

42 ENGINEERING↗

Measurement of groomed event shape observables in deep-inelastic electron-proton scattering at HERA

The H1 Collaboration at HERA reports the first measurement of groomed event shape observables in deep inelastic electron-proton scattering (DIS) at $\sqrt{s} =319$ GeV, using data recorded between the years 2003 and 2007 with an integrated luminosity of 351 pb -1 . Event shapes provide incisive probes of perturbative and non-perturbative QCD. Grooming techniques have been used for jet measurements in hadronic collisions; this paper presents the first application of grooming to DIS data. The analysis is carried out in the Breit frame, utilizing the novel Centauro jet clustering algorithm that is designed for DIS event topologies. Events are required to have squared momentum-transfer $Q^2 > 150$ GeV 2 and inelasticity $0.2< y < 0.7$. We report measurements of the production cross section of groomed event 1-jettiness and groomed invariant mass for several choices of grooming parameter. Monte Carlo model calculations and analytic calculations based on Soft Collinear Effective Theory are compared to the measurements.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

MatLab Package for Whispering Gallery Mode Data

Dye-doped whispering gallery mode resonator (WGMR) microspheres yield highly structured emission spectra that are extremely sensitive to their environment and are of intense interest for use in a variety of sensing applications. Efforts to leverage the unique sensitivities of WGMRs have relied on stringent experimental requirements to correlate specific spectral shifts/changes to an analyte/stimulus such as 1) precise positional knowledge, 2) reference spectra for each microsphere, and 3) high mechanical stability. Consequently, these can hinder adequate mixing or incorporation of analytes and creates challenges for remote sensing. The MATLAB codes provided here are to be used in conjunction with a continuous flow technique for measuring WGM spectra of dye-doped microspheres suspended in solution. One MATLAB script, smooths the data, automatically baseline corrects it to isolate the whispering gallery modes (WGM) from the unwanted bulk emission, and assesses the similarity of each spectrum to aid in selecting a set of unique WGM spectra for further analysis. The next script is designed to analyze WGM spectra to determine the size of the resonator and the refractive index (RI) of its local environment without a priori knowledge of the individual microsphere. The final script allows the user to cluster spheres based on the product of their RI and radius and the contrast ratio of the RI of the sphere material and that of its environment by using a shared nearest neighbor spectral clustering algorithm.

Lilley, Laura↗

BiG-SLiCE 2 v1.0.0

BiG-SLiCE was originally an open source Python-based command line bioinformatics software that offers a highly scalable clustering analysis on biosynthetic gene clusters (BGC) data. It allows a simultaneous analysis of millions of BGCs, exceeding the capability of other existing tools (around one hundred thousands). As a tradeoff, the clustering accuracy is relatively lower and sometimes fall short in corner cases and specific BGC classes such as the RiPPs (Ribosomally-translated, Post-translationally modified Peptides). In BiG-SLiCE V2 (developed in LBNL), the clustering algorithm has been significantly improved to deliver a much accurate result even for RiPPs and other previous corner case classes. Moreover, the speed of the overall pipeline has been improved by 50-100%. Finally, additional features were implemented to support downstream analyses of BiG-SLiCE results, such as customized tabular (TSV/CSV) and columnar (Parquet) outputs.

Kautsar, Satria↗

Differentially Private DBSCAN

SAND2022-15185 O Differentially Private DBSCAN is a differential private clustering algorithm used for running experiments and conducting data analysis. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy's National Nuclear Security Administration under contract DE-NA0003525.

Smith, MichaelReed↗

Frosted Tracks

SAND2025-01893O Frosted Tracks is a software tool to group trajectories according to sequences of their behavior. The goal is to start with a very large number of trajectories and identify groups that exhibit similar behavior patterns. The application combines TICC and Metric DBSCAN clustering algorithms for behavioral segmentation and labeling of air/sea trajectory data. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Dalbey, Keith↗