Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “clustering algorithm”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

Mobility Gaps between Low-Income and Not Low-Income Households: A Case Study in New York State

Understanding the travel challenges faced by low-income residents has always been and continues to be one of the most important transportation equity topics. This study aims to explore the mobility gaps between low-income households (HHs) and not low-income HHs, and how the gaps vary within different socio-demographic population groups in New York State (NYS). The latest National Household Travel Survey data was used as the primary data source for the analysis. The study first employed the K-prototype clustering algorithm to categorize the HHs in NYS based on their socio-demographic attributes. Five population groups were identified based on nine different household (HH) features such as HH size, vehicle ownership, and elderly status of its members. Then, the mobility differences, measured by trip frequency, trip distance, travel time, and person miles traveled, were examined among the five population groups. Results suggest that the individuals in low-income HHs consistently took fewer trips and made shorter trips compared to their not low-income counterparts in NYS. The travel distance gaps were most obvious among white HHs with more vehicles than drivers. In addition, while the population from low-income HHs made shorter trips on average (2.7 mi shorter per trip), they experienced longer travel time than those from not low-income HHs (1.8 min longer per trip). These key findings provide a deeper understanding of the travel behavior disparities between low-income and not low-income households. The findings could also support policymakers and transportation planners in addressing the critical needs of residents in low-income households in NYS and provide inputs for designing a more equitable transportation system.

Liu, Yuandong↗

Graphic contrastive learning analyses of discontinuous molecular dynamics simulations: Study of protein folding upon adsorption

A comprehensive understanding of the interfacial behaviors of biomolecules holds great significance in the development of biomaterials and biosensing technologies. In this work, we used discontinuous molecular dynamics (DMD) simulations and graphic contrastive learning analysis to study the adsorption of ubiquitin protein on a graphene surface. Our high-throughput DMD simulations can explore the whole protein adsorption process including the protein structural evolution with sufficient accuracy. Contrastive learning was employed to train a protein contact map feature extractor aiming at generating contact map feature vectors. Subsequently, these features were grouped using the k-means clustering algorithm to identify the protein structural transition stages throughout the adsorption process. The machine learning analysis can illustrate the dynamics of protein structural changes, including the pathway and the rate-limiting step. Our study indicated that the protein–graphene surface hydrophobic interactions and the π–π stacking were crucial to the seven-stage adsorption process. Upon adsorption, the secondary structure and tertiary structure of ubiquitin disintegrated. The unfolding stages obtained by contrastive learning-based algorithm were not only consistent with the detailed analyses of protein structures but also provided more hidden information about the transition states and pathway of protein adsorption process and structural dynamics. Our combination of efficient DMD simulations and machine learning analysis could be a valuable approach to studying the interfacial behaviors of biomolecules.

97 MATHEMATICS AND COMPUTING↗

Reaction–drift–diffusion models from master equations: application to material defects

We present a general method to produce well-conditioned continuum reaction–drift–diffusion equations directly from master equations on a discrete, periodic state space. We assume the underlying data to be kinetic Monte Carlo models (i.e. continuous-time Markov chains) produced from atomic sampling of point defects in locally periodic environments, such as perfect lattices, ordered surface structures or dislocation cores, possibly under the influence of a slowly varying external field. Our approach also applies to any discrete, periodic Markov chain. Here, the analysis identifies a previously omitted non-equilibrium drift term, present even in the absence of external forces, which can compete in magnitude with the reaction rates, thus being essential to correctly capture the kinetics. To remove fast modes which hinder time integration, we use a generalized Bloch relation to efficiently calculate the eigenspectrum of the master equation. A well conditioned continuum equation then emerges by searching for spectral gaps in the long wavelength limit, using an established kinetic clustering algorithm to define a proper reduced, Markovian state space.

36 MATERIALS SCIENCE↗

Variability in terrestrial characteristics and erosion rates on the Alaskan Beaufort Sea coast

Abstract Arctic coastal environments are eroding and rapidly changing. A lack of pan-Arctic observations limits our ability to understand controls on coastal erosion rates across the entire Arctic region. Here, we capitalize on an abundance of geospatial and remotely sensed data, in addition to model output, from the North Slope of Alaska to identify relationships between historical erosion rates and landscape characteristics to guide future modeling and observational efforts across the Arctic. Using existing datasets from the Alaska Beaufort Sea coast and a hierarchical clustering algorithm, we developed a set of 16 coastal typologies that captures the defining characteristics of environments susceptible to coastal erosion. Relationships between landscape characteristics and historical erosion rates show that no single variable alone is a good predictor of erosion rates. Variability in erosion rate decreases with increasing coastal elevation, but erosion rate magnitudes are highest for intermediate elevations. Areas along the Alaskan Beaufort Sea coast (ABSC) protected by barrier islands showed a three times lower erosion rate on average, suggesting that barrier islands are critical to maintaining mainland shore position. Finally, typologies with the highest erosion rates are not broadly representative of the ABSC and are generally associated with low elevation, north- to northeast-facing shorelines, a peaty pebbly silty lithology, and glaciomarine deposits with high ice content. All else being equal, warmer permafrost is also associated with higher erosion rates, suggesting that warming permafrost temperatures may contribute to higher future erosion rates on permafrost coasts. The suite of typologies can be used to guide future modeling and observational efforts by quantifying the distribution of coastlines with specific landscape characteristics and erosion rates.

54 ENVIRONMENTAL SCIENCES↗

Learning nuclear cross sections across the chart of nuclides with graph neural networks

We explore the use of deep learning techniques to learn how nuclear cross sections change as we add or remove protons and neutrons. As a proof of principle, we focus on the neutron-induced reactions in the fast energy regime. Our approach follows a two-stage learning framework. First, we apply representation learning to encode cross section data into a latent space using either variational autoencoders (VAEs) or implicit neural representations (INRs). Then, we train graph neural networks (GNNs) on the resulting embeddings to predict missing values across the nuclear chart by leveraging the topological structure of neighboring isotopes. We demonstrate accurate cross section predictions within a 9 × 9 block of missing nuclei. We also find that the optimal GNN training strategy depends on the type of latent representation used, with VAE embeddings performing best under end-to-end optimization in the original space, while INR embeddings achieve better results when the GNN is trained only in the latent space. Furthermore, using clustering algorithms, we map groups of latent vectors into regions of the nuclear chart and show that VAEs and INRs can discover some of the neutron magic numbers. These findings suggest that deep-learning models based on the representation encoding of cross sections combined with graph neural networks hold significant potential in augmenting nuclear theory models, e.g., by providing reliable estimates of covariances of cross sections, including cross-material covariances.

Machine learning↗

From hard spheres to hard-core spins

A system of hard spheres exhibits physics that is controlled only by their density. This comes about because the interaction energy is either infinite or zero, so all allowed configurations have exactly the same energy. The low-density phase is liquid, while the high-density phase is crystalline, an example of “order by disorder” as it is driven purely by entropic considerations. Here we study a family of hard spin models, which we call hard-core spin models, where we replace the translational degrees of freedom of hard spheres with the orientational degrees of freedom of lattice spins. Their hard-core interaction serves analogously to divide configurations of the many spin system into allowed and disallowed sectors. We present detailed results on the square lattice in d = 2 for a set of models with $\mathbb{Z}_n$ symmetry, which generalize Potts models, and their U(1) limits, for ferromagnetic and antiferromagnetic senses of the interaction, which we refer to as exclusion and inclusion models. As the exclusion and inclusion angles are varied, we find a Kosterlitz-Thouless phase transition between a disordered phase and an ordered phase with quasi-long-ranged order, which is the form order by disorder takes in these systems. These results follow from a set of height representations, an ergodic cluster algorithm, and transfer matrix calculations.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

Representative Period Selection for Robust Capacity Expansion Planning in Low-carbon Grids

With the increasing urgency to decarbonize power systems, while mitigating extreme events, capacity expansion models can play a vital role in reliably planning the expansion of power systems and facilitating the integration of renewable energy sources. Optimizing capacity expansion generally involves selecting surrogate representative days from forecasts of load and the generation profiles of variable renewable energy resources. To properly select those representative days, we propose a novel input-based approach in combination with the k-means clustering algorithm that utilize three unique operational inputs: load shedding, renewable curtailment, and transmission congestion. The proposed method allows for more robust and cost-effective capacity planning. The method is validated using a capacity expansion model and a production cost model based on California Independent System Operator (CAISO)'s decarbonization goals, and results in reduced costs and drastically lower load shedding.

Anderson, Osten P.↗

Crash Risks Evaluation of Urban Expressways: A Case Study in Shanghai

We report that proactive traffic safety management systems can reduce crashes by identifying crash precursors, evaluating real-time crash risks, and implementing suitable interventions. The basic prerequisite for developing such a system is to propose a reliable crash risk evaluation model that takes real-time traffic flow data as input. Previous studies have primarily focused on real-time crash prediction using some statistical or machine-learning methods. However, further quantitative evaluation and classification of crash risks have been ignored. In this study, we conduct a systematic crash risk evaluation workflow, including crash risk prediction, crash risk quantification, and crash risk classification. Specifically, the crash risk prediction using an extended logit model is proposed, from which CAS, CSD, UAS, DAS, DTV are identified to be contributing factors of crash risks. Then a crash risk quantification model based on the parameter evaluation of the extended logit model is developed. The crash risks of urban expressways and their spatial-temporal evolution trends are quantified. Finally, the crash risks are classified into high crash risk level, moderate crash risk level, and low crash risk level by the k-means cluster algorithm. Then the threshold boundaries of different crash risk levels are determined. The research results provide a proactive guidance for traffic safety management of urban expressways.

33 ADVANCED PROPULSION SYSTEMS↗

Resilience-Oriented DG Siting and Sizing Considering Stochastic Scenario Reduction

In this paper, a fuel-based distributed generator (DG) allocation strategy is proposed to enhance the distribution system resilience against extreme weather. The long-term planning problem is formulated as a two-stage stochastic mixed-integer programming (SMIP). The first stage is to make decisions of DG siting and sizing under the given budget constraint. In the second stage, a post-extreme-event-restoration (PEER) is employed to minimize the operating cost in an uncertain fault scenario. In particular, this study proposes a method to select the most representative scenarios for the SMIP. First, a Monte Carlo Simulation (MCS) is introduced to generate sufficient scenarios considering random fault locations and load profiles. Then, the number of scenarios is reduced by the K-means clustering algorithm. The advantage of scenario reduction is to make a trade-off between accuracy and computational efficiency. Finally, the SMIP is solved by the progressive hedging algorithm. Here, the case studies of the IEEE 33-bus and 123-bus test systems demonstrate the effectiveness of the proposed algorithm in reducing the expected energy not served (EENS), which is a critical criterion of resilience.

42 ENGINEERING↗

Measurement of groomed event shape observables in deep-inelastic electron-proton scattering at HERA

The H1 Collaboration at HERA reports the first measurement of groomed event shape observables in deep inelastic electron-proton scattering (DIS) at $\sqrt{s} =319$ GeV, using data recorded between the years 2003 and 2007 with an integrated luminosity of 351 pb -1 . Event shapes provide incisive probes of perturbative and non-perturbative QCD. Grooming techniques have been used for jet measurements in hadronic collisions; this paper presents the first application of grooming to DIS data. The analysis is carried out in the Breit frame, utilizing the novel Centauro jet clustering algorithm that is designed for DIS event topologies. Events are required to have squared momentum-transfer $Q^2 > 150$ GeV 2 and inelasticity $0.2< y < 0.7$. We report measurements of the production cross section of groomed event 1-jettiness and groomed invariant mass for several choices of grooming parameter. Monte Carlo model calculations and analytic calculations based on Soft Collinear Effective Theory are compared to the measurements.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

MatLab Package for Whispering Gallery Mode Data

Dye-doped whispering gallery mode resonator (WGMR) microspheres yield highly structured emission spectra that are extremely sensitive to their environment and are of intense interest for use in a variety of sensing applications. Efforts to leverage the unique sensitivities of WGMRs have relied on stringent experimental requirements to correlate specific spectral shifts/changes to an analyte/stimulus such as 1) precise positional knowledge, 2) reference spectra for each microsphere, and 3) high mechanical stability. Consequently, these can hinder adequate mixing or incorporation of analytes and creates challenges for remote sensing. The MATLAB codes provided here are to be used in conjunction with a continuous flow technique for measuring WGM spectra of dye-doped microspheres suspended in solution. One MATLAB script, smooths the data, automatically baseline corrects it to isolate the whispering gallery modes (WGM) from the unwanted bulk emission, and assesses the similarity of each spectrum to aid in selecting a set of unique WGM spectra for further analysis. The next script is designed to analyze WGM spectra to determine the size of the resonator and the refractive index (RI) of its local environment without a priori knowledge of the individual microsphere. The final script allows the user to cluster spheres based on the product of their RI and radius and the contrast ratio of the RI of the sphere material and that of its environment by using a shared nearest neighbor spectral clustering algorithm.

Lilley, Laura↗

BiG-SLiCE 2 v1.0.0

BiG-SLiCE was originally an open source Python-based command line bioinformatics software that offers a highly scalable clustering analysis on biosynthetic gene clusters (BGC) data. It allows a simultaneous analysis of millions of BGCs, exceeding the capability of other existing tools (around one hundred thousands). As a tradeoff, the clustering accuracy is relatively lower and sometimes fall short in corner cases and specific BGC classes such as the RiPPs (Ribosomally-translated, Post-translationally modified Peptides). In BiG-SLiCE V2 (developed in LBNL), the clustering algorithm has been significantly improved to deliver a much accurate result even for RiPPs and other previous corner case classes. Moreover, the speed of the overall pipeline has been improved by 50-100%. Finally, additional features were implemented to support downstream analyses of BiG-SLiCE results, such as customized tabular (TSV/CSV) and columnar (Parquet) outputs.

Kautsar, Satria↗

Differentially Private DBSCAN

SAND2022-15185 O Differentially Private DBSCAN is a differential private clustering algorithm used for running experiments and conducting data analysis. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy's National Nuclear Security Administration under contract DE-NA0003525.

Smith, MichaelReed↗

Frosted Tracks

SAND2025-01893O Frosted Tracks is a software tool to group trajectories according to sequences of their behavior. The goal is to start with a very large number of trajectories and identify groups that exhibit similar behavior patterns. The application combines TICC and Metric DBSCAN clustering algorithms for behavioral segmentation and labeling of air/sea trajectory data. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Dalbey, Keith↗

Coastal typologies and surface and subsurface characteristics of the Alaskan Beaufort Sea Coast

This dataset was generated to classify the Alaskan Beaufort Sea Coast (ABSC) into a set of distinct coastal typologies, to understand the surface and subsurface characteristics and variability of the ABSC, and to quantify relationships between these characteristics and historical rates of shoreline change. This geospatial dataset contains two csv files of points along the ABSC at a 50 m spacing, one for points sheltered by a barrier island and one for points exposed to the open ocean. Each point has a lat/lon location, and we have attributed to each point average values for elevation, historical long-term shoreline change rates, shoreline orientation, landcover, mean annual ground temperature, geomorphic unit, lithology, geology, ecological landscape unit, maximum thaw settlement potential, massive ice content, and segregated ice content. Each point is also assigned to one of 16 coastal typologies, determined by a hierarchical clustering algorithm on the elevation, shoreline change, orientation, and ground temperature data. There are 9 sheltered typologies and 7 exposed typologies, identified by an integer label in the last column of each csv file. The other two csv files contain the integer IDs and classes for the landcover and geomorphology datasets.

54 ENVIRONMENTAL SCIENCES↗

Global Characteristics of Observable Foreshocks for Large Earthquakes

Abstract Foreshocks are the only currently widely identified precursory seismic behavior, yet their utility and even identifiability are problematic, in part because of extreme variation in behavior. Here, we establish some global trends that help identify the expected frequency of foreshocks as well the type of earthquake most prone to foreshocks. We establish these tendencies using the global earthquake catalog of the U.S. Geological Survey National Earthquake Information Center with a completeness level of magnitude 5 and mainshocks with Mw≥7.0. Foreshocks are identified using three clustering algorithms to address the challenge of distinguishing foreshocks from background activity. The methods give a range of 15%–43% of large mainshocks having at least one foreshock but a narrower range of 13%–26% having at least one foreshock with magnitude within two units of the mainshock magnitude. These observed global foreshock rates are similar to regional values for a completeness level of magnitude 3 using the same detection conditions. The foreshock sequences have distinctive characteristics with the global composite population b-values being lower for foreshocks than for aftershocks, an attribute that is also manifested in synthetic catalogs computed by epidemic-type aftershock sequences, which intrinsically involves only cascading processes. Focal mechanism similarity of foreshocks relative to mainshocks is more pronounced than for aftershocks. Despite these distinguishing characteristics of foreshock sequences, the conditions that promote high foreshock productivity are similar to those that promote high aftershock productivity. For instance, a modestly higher percentage of interplate mainshocks have foreshocks than intraplate mainshocks, and reverse faulting events slightly more commonly have foreshocks than normal or strike-slip-faulting mainshocks. The western circum-Pacific is prone to having slightly more foreshock activity than the eastern circum-Pacific.

Geochemistry & Geophysics↗

Connecting the Radiative Influences of Aerosol upon the Mass Flux Profiles of Shallow Cumuli across the Southeast Atlantic Ocean Basin and its Boundaries (Final Report)

The Atlantic Ocean covers approximately 25% of Earth’s surface and the atmosphere above it is home to a complex array of clouds and aerosols that have important influences on regional and global weather and climate. These influences must be accurately depicted in short range, medium range, seasonal, and climate forecast models. Conditions over the tropical Atlantic are particularly complex due to continental scale plumes of dust from the Sahara Desert and smoke from agricultural burning in Africa that drift across the Atlantic Ocean basin toward the Americas. These plumes are often found meandering in the lower atmosphere above shallow tropical clouds that form above the ocean surface, presumably mingling with these clouds on occasion due to convective mixing processes. Elevated dust and smoke particles absorb incoming sunlight and substantially warm the marine atmosphere in the layer in which they are present. This warming may alter the thermal stability of the marine atmosphere and may throttle or enhance the development of clouds, change their internal structure and the rate at which they precipitate. Alternatively, it may isolate the lower atmosphere from drier layers above enabling water vapor to accumulate near the ocean surface potentially leading to the development of deeper convection. Our study was organized around the principal concept of determining the impact of African smoke and dust plumes upon cloud development at ASI and understanding how a popular shallow convection parameterization used in models responds to the presence of this aerosol. We analyzed observations collected during the US Department of Energy (DOE) Layered Atlantic Smoke Interactions with Clouds (LASIC) campaign using the Atmospheric Radiation Measurement Program’s Mobile Facility #1 (AMF-1), which was deployed to Ascension Island (ASI) in Southeastern Atlantic for a one-year period. Ascension Island is immediately downwind from an African source of these plumes, but far enough removed to enable the lower atmosphere to have reacted to their presence. A main goal of our study was to compute radiative heating rate profiles over Ascension Island and along the trajectory from the biomass burning regions along coastal Africa to Ascension Island. To compute the radiative heating rate due to aerosols and clouds, we employed observed profiles of temperature, humidity, and clouds from LASIC alongside aerosol optical properties from the Modern-era Retrospective analysis for Research and Applications Version 2 (MERRA-2), as input for the Rapid Radiation Transfer Model (RRTM). Radiative heating was also assessed across the southeast Atlantic Ocean using an ensemble of back trajectories from the Hybrid Single Particle Lagrangian Integrated Trajectory (HYSPLIT) model. We were successful in this effort and the resulting publication is already being cited despite its relatively short lifetime in the literature. The second part of our study involved the process-level representation of the clouds observed over the Southeastern Atlantic in models. Our initial task, upon which the balance of this portion of the study depended, was to evaluate the representativeness of the clouds observed by AMF-1. What orographic influence did Ascension Island have on the measured cloud structure? We set out to answer this question by trying to separate observations that were clearly indicative of orographic forcing from those that were representative of open ocean. The siting of AMF-1, which was 340-m above sea-level on the slope of a steep escarpment, proved demonstrably problematic for cloud process measurements. Despite considerable effort using multiple approaches including artificial intelligence (AI), we were unable to successfully compensate for the orographic forcing present at the AMF-1 site on ASI and, hence, produce process-level summaries of the convective mass flux representative of open ocean over the Southeastern Atlantic. Even though the project has officially ended, we are making a final attempt using two new types of clustering algorithm (i.e., AI) on the recommendation of a former student, who is an expert in this area. Since we have only recently begun to test these new AI schemes, the recommendations from the second part of our project outlined below are based on our experience at the time of this report.

54 ENVIRONMENTAL SCIENCES↗

Calibrating LArPix for TinyTPC

LArTPCs provide sensitivity to GeV signals, such as accelerator neutrinos and part of the supernova neutrino spectrum. TinyTPC is a LArTPC test stand for R&D of LAr doping to expand the reach of LArTPCs down to the 1-10 MeV range, which would substantially enhance the flagship analyses of experiments like DUNE, while enabling low energy analyses. We aim to dope LAr with Xe and pho- tosensitive dopants to expand the LArTPC range by converting hard-to-detect scintillation light to efficiently detected ionization charge. A critical element of the data analysis in TinyTPC is calibrating the readout. This poster will cover the calibration of TinyTPC, a pixelated liquid argon detector, where we find the distance a muon travels through each pixel. We then calculate the energy loss of muons traveling through the detector from cosmic data. We can reconstruct the path of the particles through the TPC using a density-based clustering algorithm designed to sort straight cosmic muon tracks from low energy radioactive decay curled paths and electronic noise

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗