Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “ensemble clustering”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

A Semi-supervised Hybrid Machine Learning Framework for the Qualification of Resistance Spot Welds

• Industries requiring high structural integrity, including automotive, aerospace, and construction, place considerable significance on weld quality classification. • The inspection normally involves human expertise through predefined quality metrics that are subjective, error-prone, and time-intensive • The challenge to classification model development is the scarcity of labeled data and imbalanced distributions in the data that are labeled. • This work develops a new hybrid methodology that achieves clustering using KMeans++ together with supervised classification to overcome these challenges. • The ensemble-based classifiers were identified as optimal, with accuracy enhancements of up to 8% using the pseudo-labeled dataset. • The work provides practical insight into feature engineering and machine learning integration in industrial quality assurance applications.

Rogers, Jeremy K. [Savannah River National Laborat↗

The Parallel System for Integrating Impact Models and Sectors (pSIMS)

We present a framework for massively parallel climate impact simulations: the parallel System for Integrating Impact Models and Sectors (pSIMS). This framework comprises a) tools for ingesting and converting large amounts of data to a versatile datatype based on a common geospatial grid; b) tools for translating this datatype into custom formats for site-based models; c) a scalable parallel framework for performing large ensemble simulations, using any one of a number of different impacts models, on clusters, supercomputers, distributed grids, or clouds; d) tools and data standards for reformatting outputs to common datatypes for analysis and visualization; and e) methodologies for aggregating these datatypes to arbitrary spatial scales such as administrative and environmental demarcations. By automating many time-consuming and error-prone aspects of large-scale climate impacts studies, pSIMS accelerates computational research, encourages model intercomparison, and enhances reproducibility of simulation results. We present the pSIMS design and use example assessments to demonstrate its multi-model, multi-scale, and multi-sector versatility.

crop modeling↗

Synergistic Retrievals of Ice in High Clouds From Elastic Backscatter Lidar, Ku-band Radar and Submillimeter Wave Radiometer Observations

In this study, we investigate the synergy of elastic backscatter lidar, Ku-band radar, and sub-millimeter-wave radiometer measurements in the retrieval of ice from satellite observations. The synergy is analyzed through the generation of a large dataset of IceWater Content (IWC) profiles and simulated lidar, radar and radiometer observations. The characteristics of the instruments e.g. frequencies, sensitivities, etc. are set based on the expected characteristics of instruments of the Atmosphere Observing System (AOS) mission. A hold-out validation methodology is used to assess the accuracy of the IWC profiles retrieved from various combinations of observations from the three instruments. Specifically, the IWC and associated observations are randomly divided into two datasets, one for training and the other for evaluation. The training dataset is used to train the retrieval algorithm, while the evaluation dataset is used to assess the retrieval performance. The dataset of IWC profiles is derived from CloudSat reflectivity and CALIOP lidar observations. The retrieval of the ice water content IWC profiles from the computed observations is achieved in two steps. In the first step, a class, out of 18 potential classes characterized by different vertical distribution of IWC, is estimated from the observations. The 18 classes are predetermined based on the k-Means clustering algorithm. In the second step, the IWC profile is estimated using an Ensemble Kalman Smoother (EKS) algorithm that uses the estimated class as a priori information. The results of the study show that the synergy of lidar, radar, and radiometer observations is significant in the retrieval of the IWC profiles. Nevertheless, it should be mentioned that this synergy was found under idealized conditions, and additional work might be required to materialize it in practice. The inclusion of the lidar backscatter observations in the retrieval process has a larger impact on the retrieval performance than the inclusion of the radar observations. As ice clouds have a significant impact on atmospheric radiative processes, this work is relevant to ongoing efforts to reduce uncertainties in climate analyses and projections.

Mircea Grecu↗

Bespoke Liquid/Liquid Interfaces (Final Technical Report)

The goal of DE-SC0001815 was to advance the basic science of liquid:liquid interface formation, to develop a deeper understanding of the mechanisms of phase separation and the essential relationships between solution composition, organization and dynamics that underlie the kinetic regime of solvent extraction. This included learning how interfacial organization and dynamics alters the properties of the primary coordination sphere of ions and the free energy of transport of ions complexes across a phase boundary. We relied primarily upon classical molecular dynamics studies to determine the equilibrium ensembles of these complex systems, but also utilized ab-initio MD and cluster-based density functional theory (DFT) calculations when more detailed investigation of the electronic structure was needed. We continued development of graph-theory based analyses to elucidate hierarchical correlations and expanded into geometric topology methods to quantify the collectively organized structures that can organize at a liquid/liquid interface during solute transport. One of the main conclusions was from the observation of two distinct mechanisms for solute transport - those that derive from amplifications of interfacial heterogeneity and surface roughness, and those wherein surface roughness has been dampened and instead collectively organized macrostructures work to bring solutes into the organic phase. It was our aim to create a concrete chemical model of the underlying driving forces behind interfacial primary and secondary structure formation and to map out the energetic features of solute transport so that tailored liquid/liquid can be developed that have characteristic kinetic features associated with mass transport.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Directing reaction pathways via in situ control of active site geometries in PdAu single-atom alloy catalysts

Abstract The atomic scale structure of the active sites in heterogeneous catalysts is central to their reactivity and selectivity. Therefore, understanding active site stability and evolution under different reaction conditions is key to the design of efficient and robust catalysts. Herein we describe theoretical calculations which predict that carbon monoxide can be used to stabilize different active site geometries in bimetallic alloys and then demonstrate experimentally that the same PdAu bimetallic catalyst can be transitioned between a single-atom alloy and a Pd cluster phase. Each state of the catalyst exhibits distinct selectivity for the dehydrogenation of ethanol reaction with the single-atom alloy phase exhibiting high selectivity to acetaldehyde and hydrogen versus a range of products from Pd clusters. First-principles based Monte Carlo calculations explain the origin of this active site ensemble size tuning effect, and this work serves as a demonstration of what should be a general phenomenon that enables in situ control over catalyst selectivity.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Recovery of Deep Moonquake Focal Mechanisms

Deep moonquakes are clustered not only in space but also in time: their recurrence times correspond to the durations of the anomalistic and draconic months, with some clusters preferring one of the two periods, while others are active with both periods. A key constraint for the understanding of the connection between the orbital motion of the Moon and its seismic activity is the focal mechanism: the orientation of the fault surface on which failure occurs during the quake. Due to the small aperture of the Apollo seismic network and the strong scattering of seismic waves within the lunar crust, the evaluation of P wave first motions to constrain the strike and dip of the fault planes is not feasible. Instead we evaluate the amplitude ratios of P and S waves. Seismograms are rotated into the P-SV-SH coordinate frame and amplitudes are determined as averages over short time windows after the arrival to reduce the impact of the scattering coda, which is independent of the source orientation. We allow for reversals of the fault motion, as observed for some clusters in previous studies, by taking into account the absolute amplitude only, without sign. An empirical site correction factor is applied to correct for amplitude distortions in the crust. We construct ensembles of fault plane solutions using an exhaustive grid search by accepting all orientations that reproduce the measured amplitude ratios within the observed standard deviations. Since all events of a given cluster are supposed to share the same fault plane, the combination of the individual inversion results further constrains the orientation. We evaluate 106 events from 25 different moonquake clusters. The most active cluster A001 contributes 37 events, while others contribute 1 to 9 events per cluster. Comparison of fault orientations with the variation of the tidal stress results in preferred orientations.

Weber, Renee C.↗

Cosmological inference from an emulator based halo model. II. Joint analysis of galaxy-galaxy weak lensing and galaxy clustering from HSC-Y1 and SDSS

Here, we present high-fidelity cosmology results from a blinded joint analysis of galaxy-galaxy weak lensing (ΔΣ) and projected galaxy clustering (w p ) measured from the Hyper Suprime-Cam Year-1 (HSC-Y1) data and spectroscopic Sloan Digital Sky Survey (SDSS) galaxy catalogs in the redshift range 0.15 < z < 0.7. We define luminosity-limited samples of SDSS galaxies to serve as the tracers of w p in three spectroscopic redshift bins, and as the lens samples for ΔΣ. For the ΔΣ measurements, we select a single sample of 4×10 6 source galaxies over 140 deg 2 from HSC-Y1 with photometric redshifts (photo z) greater than 0.75, enabling a better handle of photo- z errors by comparing the ΔΣ amplitudes for the three lens redshift bins. The deep, high-quality HSC-Y1 data enable significant detections of the ΔΣ signals, with integrated signal-to-noise ratio S/N ~ 15 in the range 3 ≤ R/[h –1 Mpc] ≤ 30 for the three lens samples, despite the small area coverage. For cosmological parameter inference, we use an input galaxy-halo connection model built on the dark emulator package (which uses an ensemble set of high-resolution N-body simulations and enables fast, accurate computation of the clustering observables) with a halo occupation distribution that includes nuisance parameters to marginalize over modeling uncertainties. We model the ΔΣ and wp measurements on scales from R≃3 and 2h –1 Mpc , respectively, up to 30 h –1 Mpc (therefore excluding the baryon acoustic oscillations information) assuming a flat Λ CDM cosmology, marginalizing over about 20 nuisance parameters and demonstrating the robustness of our results to them. With various tests using mock catalogs described in Miyatake et al. [preceding paper, Phys. Rev. D 106, 083519 (2022)], we show that any bias in the clustering amplitude S 8 ≡ σ 8 (Ω m /0.3) 0.5 due to uncertainties in the galaxy-halo connection is less than ~ 50 % of the statistical uncertainty of S 8 , unless the assembly biaseffect is unexpectedly large. Our best-fit models have S 8 = 0.795$_{-0.042}^{+0.049}$ (mode and 68% credible interval) for the flat Λ CDM model; we find tighter constraints on the quantity S 8 (α = 0.17)≡σ 8 (Ω m /0.3) 0.17 = 0.745$_{-0.031}^{+0.039}$.

79 ASTRONOMY AND ASTROPHYSICS↗

A high spectral resolution VLA search for H I absorption towards A496, A1795, and A2584

In this paper, we present the results of a Very Large Array (VLA) search for H I absorption with high spectral resolution (1.6 km/s) towards A496, A1795, A2584, and A2597. These observations are well matched to the properties of cold, optically thick H I clouds, where the line width is given by the width of an individual cloud rather than the dispersion in an ensemble of clouds. We do not detect any H I absorption with narrow linewidths in these clusters. Our limits mainly apply to clouds which are larger than a few tenths parsec-i.e., if the clouds are much smaller than the background radio source and have a low covering factor in velocity space, they could still escape detection. The estimated limits on column density (for clouds in this regime of parameter space) are 2-3 orders of magnitude less than the 10(exp 21)/sq cm required to explain the x-ray absorption seen in some cooling flow clusters. The combination of our high spectral resolution H I absorption searches with the existing lower spectral resolution H I absorption searches and the searches for H I emission makes it unlikely that atomic hydrogen is the dominant component of the cold x-ray absorbing gas in the inter-cloud medium (ICM).

O'Dea, Christopher P.↗

Electrocatalytic Hydrogen Evolution at Full Atomic Utilization over ITO-Supported Sub-nano-Pt n Clusters: High, Size-Dependent Activity Controlled by Fluxional Pt Hydride Species

A combination of density functional theory (DFT) and experiments with atomically size-selected Pt n clusters deposited on indium-tin oxide (ITO) electrodes was used to examine the effects of applied potential and Pt n size on the electrocatalytic activity of Pt n (n = 1, 4, 7, 8) for the hydrogen evolution reaction (HER). Activity is found to be negligible for isolated Pt atoms on ITO, increasing rapidly with Pt n size, such that Pt 7 /ITO and Pt 8 /ITO have roughly double the activity per Pt atom compared to atoms in the surface layer of polycrystalline Pt. Both DFT and experiment find that hydrogen under-potential deposition (H upd ) results in Pt n /ITO (n = 4, 7, 8) adsorbing ~2 H atoms/Pt atom at the HER threshold potential, equal to ca. double the Hupd observed for Pt bulk or nanoparticles. Here, the cluster catalysts under electrocatalytic conditions are hence best described as a Pt hydride compound, significantly departing from a metallic Pt cluster. The exception is Pt 1 /ITO, where H adsorption at the HER threshold potential is energetically unfavorable. Theory combines global optimization with grand canonical approaches for the influence of potential, uncovering that several metastable structures contribute to HER, changing with the applied potential. It is hence critical to include reactions of the ensemble of energetically accessible Pt n H x /ITO structures to correctly predict the activity vs. Pt n size and applied potential. For the small clusters, spillover of H ads from the clusters to the ITO support is significant, resulting in a competing channel for loss of H ads , particularly at slow potential scan rates.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Accurate Modeling of Bromide and Iodide Hydration with Data-Driven Many-Body Potentials

Ion–water interactions play a central role in determining the properties of aqueous systems in a wide range of environments. However, a quantitative understanding of how the hydration properties of ions evolve from small aqueous clusters to bulk solutions and interfaces remains elusive. Here, we introduce the second generation of data-driven many-body energy (MBnrg) potential energy functions (PEFs) representing bromide–water and iodide–water interactions. The MB-nrg PEFs use permutationally invariant polynomials to reproduce two-body and three-body energies calculated at the coupled cluster level of theory, and implicitly represent all higher-body energies using classical many-body polarization. Further, a systematic analysis of the hydration structure of small Br - (H 2 O) n and I - (H 2 O) n clusters demonstrates that the MBnrg PEFs predict interaction energies in quantitative agreement with “gold standard” coupled cluster reference values. Importantly, when used in molecular dynamics simulations carried out in the isothermal-isobaric ensemble for single bromide and iodide ions in liquid water, the MB-nrg PEFs predict extended X-ray absorption fine structure (EXAFS) spectra that accurately reproduce the experimental spectra, which thus allows for characterizing the hydration structure of the two ions with high level of confidence.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Supercomputer-Based Ensemble Docking Drug Discovery Pipeline with Application to Covid-19

In this work, we present a supercomputer-driven pipeline for in silico drug discovery using enhanced sampling molecular dynamics (MD) and ensemble docking. Ensemble docking makes use of MD results by docking compound databases into representative protein binding-site conformations, thus taking into account the dynamic properties of the binding sites. We also describe preliminary results obtained for 24 systems involving eight proteins of the proteome of SARS-CoV-2. The MD involves temperature replica exchange enhanced sampling, making use of massively parallel supercomputing to quickly sample the configurational space of protein drug targets. Using the Summit supercomputer at the Oak Ridge National Laboratory, more than 1 ms of enhanced sampling MD can be generated per day. We have ensemble docked repurposing databases to 10 configurations of each of the 24 SARS-CoV-2 systems using AutoDock Vina. Comparison to experiment demonstrates remarkably high hit rates for the top scoring tranches of compounds identified by our ensemble approach. We also demonstrate that, using Autodock-GPU on Summit, it is possible to perform exhaustive docking of one billion compounds in under 24 h. Finally, we discuss preliminary results and planned improvements to the pipeline, including the use of quantum mechanical (QM), machine learning, and artificial intelligence (AI) methods to cluster MD trajectories and rescore docking poses.

60 APPLIED LIFE SCIENCES↗

Intermittent Criticality Multi‐Scale Processes Leading to Large Slip Events on Rough Laboratory Faults

Abstract We discuss data of three laboratory stick‐slip experiments on Westerly Granite samples performed at elevated confining pressure and constant displacement rate on rough fracture surfaces. The experiments produced complex slip patterns including fast and slow ruptures with large and small fault slips, as well as failure events on the fault surface producing acoustic emission bursts without externally‐detectable stress drop. Preparatory processes leading to large slips were tracked with an ensemble of ten seismo‐mechanical and statistical parameters characterizing local and global damage and stress evolution, localization and clustering processes, as well as event interactions. We decompose complex spatio‐temporal trends in the lab‐quake characteristics and identify persistent effects of evolving fault roughness and damage at different length scales, and local stress evolution approaching large events. The observed trends highlight labquake localization processes on different spatial and temporal scales. The preparatory process of large slip events includes smaller events marked by confined bursts of acoustic emission activity that collectively prepare the fault surface for a system‐wide failure by conditioning the large‐scale stress field. Our results are consistent overall with an evolving process of intermittent criticality leading to large failure events, and may contribute to improved forecasting of large natural earthquakes.

Geochemistry & Geophysics↗

Glueballs at physical pion mass

Glueballs are investigated through gluonic operators on two $ N_f=2+1 $ RBC/UKQCD gauge ensembles at the physical pion mass. The statistical errors of glueball correlation functions are considerably reduced through the cluster decomposition error reduction (CDER) method. The Bethe-Salpeter wave functions are obtained for the scalar, tensor, and pseudoscalar glueballs by using spatially extended glueball operators defined through the gauge potential $ A_\mu(x) $ in the Coulomb gauge. These wave functions exhibit similar features of non-relativistic two-gluon systems and are used to optimize the signals of the related correlation functions at the early time regions, where the ground state masses are extracted. These masses are close to those from the quenched approximation and indicate the possible existence of glueballs at the physical point. The resonance feature of glueballs and the mixing with conventional mesons and multi-hadron systems should be considered in a more systematic lattice study.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Artificial Diversity and Defense Security (ADDSec)

Artificial Diversity and Defense Security (ADDSec) machine learning algorithms are used to classify and cluster threats so that an appropriate response can be initiated as a mitigation strategy. The package includes an ensemble of machine learning algorithms such as Support Vector Machines, naïve bayes, logistic regression, and random forest that evolve with the data to recognize anomalous behavior at the host and network levels. Inputs into the machine learning algorithms include end host system calls, system utilization, packet captures, and syslog messages. The machine learning algorithms can be retrained based on user defined intervals or on the number of packets received. ADDSEC's threat responses include Internet Protocol (IP) Address randomization, application port number randomization, and application library randomization. The IP randomization implementation is built on top of a Software Defined Networking (SDN) framework. The SDN controller installs flows on each of the SDN switches with randomized source and destination IP addresses. The application port numbers are randomized using iptables. The application library randomization is created with a LLVM compiler. All randomization schemes are transparent to the endpoints on the network. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525. SAND2021-3379 O

Cox, RebeccaE.↗

Computing Aerodynamic Performance of a 2D Iced Airfoil: Blocking Topology and Grid Generation

The ice accrued on airfoils can have enormously complicated shapes with multiple protruded horns and feathers. In this paper, several blocking topologies are proposed and evaluated on their ability to produce high-quality structured multi-block grid systems. A transition layer grid is introduced to ensure that jaggedness on the ice-surface geometry do not to propagate into the domain. This is important for grid-generation methods based on hyperbolic PDEs (Partial Differential Equations) and algebraic transfinite interpolation. A 'thick' wrap-around grid is introduced to ensure that grid lines clustered next to solid walls do not propagate as streaks of tightly packed grid lines into the interior of the domain along block boundaries. For ice shapes that are not too complicated, a method is presented for generating high-quality single-block grids. To demonstrate the usefulness of the methods developed, grids and CFD solutions were generated for two iced airfoils: the NLF0414 airfoil with and without the 623-ice shape and the B575/767 airfoil with and without the 145m-ice shape. To validate the computations, the computed lift coefficients as a function of angle of attack were compared with available experimental data. The ice shapes and the blocking topologies were prepared by NASA Glenn's SmaggIce software. The grid systems were generated by using a four-boundary method based on Hermite interpolation with controls on clustering, orthogonality next to walls, and C continuity across block boundaries. The flow was modeled by the ensemble-averaged compressible Navier-Stokes equations, closed by the shear-stress transport turbulence model in which the integration is to the wall. All solutions were generated by using the NPARC WIND code.

Chi, X.↗

Near-term heatwave risk in HighResMIP models across different temperature zones of West Africa

This study projects near-future (2031–2050) changes in heatwave (HW) risk across West Africa (WA) using an ensemble of eight high-resolution global climate models from the High-Resolution Model Intercomparison Project under a high-emission scenario. Using K-means clustering, we divided WA into four unique temperature zones and examined projected changes in extreme temperatures, HW occurrence and magnitude. Our results indicate a statistically significant increase in future HW events across most parts of WA, although considerable spread exists over the region and among individual models. The most pronounced increases are evident in the Sahel/Sahara and the Guinea Highlands subregions, with an ensemble mean increase of ∼10 HW events per year. In contrast, the lowest increase in HW events is projected in central WA, with increases ranging between 1 and 5 events per year. Similarly, the magnitude of HW events is projected to increase in most models, with Sahel/Sahara exhibiting the largest increases. Additionally, projections suggest that the strongest HWs will become more frequent, particularly in northern and southwestern WA. These findings highlight significant spatial heterogeneity in future HW risk across WA, emphasizing the need for targeted adaptation strategies.

HighResMIP↗

Dial-A-Cluster User Manual

The Dial-A-Cluster (DAC) model allows interactive visualization of multivariate time series data. A multivariate time series dataset consists of an ensemble of data points, where each data point consists of a set of time series curves. The example of a DAC dataset used in this guide is a collection of 100 cities in the United States, where each city collects a year's worth of weather data, including daily temperature, humidity, and wind speed measurements.

97 MATHEMATICS AND COMPUTING↗

Integrated machine learning-molecular dynamics framework for electrolyte property prediction

Electrochemical stability windows determine the operating range of battery electrolytes, yet accurate prediction remains challenging because stability emerges from statistical ensembles of local solvation environments rather than single ground-state molecular structures. Traditional density functional theory calculations on energy-minimized clusters cannot capture the thermal variations in local coordination environments and geometries that govern decomposition, while SMILES-based machine learning methods lack explicit representation of three-dimensional solvation structure and ion pairing. Here, we introduce a structure-aware machine learning framework that predicts frontier orbital energies (HOMO and LUMO) directly from molecular dynamics-sampled solvation configurations, achieving sub-0.6 eV accuracy at computational costs 3–4 orders of magnitude lower than first-principles methods. Across twelve representative battery electrolytes, we demonstrate that solvent-separated and contact ion pairs exhibit strong size- and local chemistry dependent electronic stability, with variations in coordination shifts of HOMO or LUMO level by 2–3 eV, and that extended solvation structure and partially desolvated environment further modulate stability by up to 3 eV. By encoding the statistical nature of electrochemical failure through ensemble sampling of explicit solvation geometries, our approach enables high-throughput screening and rational design of next-generation battery electrolytes with mechanistic understanding of structure–property relationships.

Energy - Storage↗