Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “ensemble clustering”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Dark Energy Survey Year 6 results: Clustering redshifts and importance sampling of self-organized-maps 𝑛⁡(𝑧) realizations for 3 × 2 ⁢pt samples

This work is part of a series establishing the redshift framework for the 3 × 2 ⁢pt analysis of the Dark Energy Survey Year 6 (DES Y6). For DES Y6, photometric redshift distributions are estimated using self-organizing maps (SOMs), calibrated with spectroscopic and many-band photometric data. To overcome limitations from color-redshift degeneracies and incomplete spectroscopic coverage, we enhance this approach by incorporating clustering-based redshift constraints (clustering-z, or WZ) from angular cross-correlations with BOSS and eBOSS galaxies and eBOSS quasar samples. We define a WZ likelihood and apply importance sampling to a large ensemble of SOM-derived 𝑛⁡(𝑧) realizations, selecting those consistent with the clustering measurements to produce a posterior sample for each lens and source bin. The analysis uses angular scales corresponding to 1.5–5 Mpc to optimize signal-to-noise ratio while mitigating modeling uncertainties and marginalizes over redshift-dependent galaxy bias and other systematics informed by the N-body simulation CARDINAL . While a sparser spectroscopic reference sample limits WZ constraining power at 𝑧 >1.1, particularly for source bins, we demonstrate that combining SOM with WZ improves redshift accuracy and enhances the overall cosmological constraining power of DES Y6. As a result, we estimate an improvement in 𝑆 8 of approximately 10% for cosmic shear and 3 ×2⁢pt analysis, primarily due to the WZ calibration of the source samples.

Cosmological parameters↗

Adversarial Ensemble Modeling of Multi-modal Mechanical Properties for Iron-Based Alloys

Mechanical properties of alloys are controlled by their microstructure; and microstructure evolution is controlled by internal and external stressors. Chemical complexity during the alloy processing may result in heterogeneity and observation of multimodal performance patterns. Thus, to ensure the desired performance of stressed components it is important to understand the origin and mechanisms of such behavior. Adversarial ensemble modeling was introduced here to explain multimodal mechanical properties of the iron-based alloys. The modeling results showed that the areas of a single mechanism predominance were contiguous across the alloy compositions clustered by similarity, with sharp and persistent boundaries separating the single-mechanism domains. The marginal compositions resulted in increased competition between the adversarial models but only led to the multimodal behavior when the competing models diverged. Transparency of the clustering method allowed explicit interpretation of the chemistries leading to such competition. The adversarial ensemble was interpreted through transition from ductile to brittle microstructure.

36 MATERIALS SCIENCE↗

Coherent motion induced fluctuations in the primary transition region of a plane shear layer

The naturally occurring large scale motions in a single stream shear layer (that is initiated from a fully turbulent boundary layer) are made evident by the induced velocities in the entrainment region beyond the active shear layer. The distinctive attributes of these induced motions are particularly evident in the Michigan State Univerity Free Shear Flow Facility since the total test section length (3m) is nominally the same as the location of the first, fully formed, coherent motion, ca/x theta (0) = 400 (or 2.5 m). Hence, detailed studies of the induced motions can be executed. Individual coherent motions are identified by the induced velocity signatures and conditional-ensemble statistics are used to represent the irrotational field properties. Clusters of such motions exist; some of their properties are substantially different from the unconditionally averaged values.

Foss, J. F.↗

Theoretical assessments of Pd–PdO phase transformation and its impacts on H 2 O 2 synthesis and decomposition pathways

The direct synthesis of H 2 O 2 from O 2 and H 2 provides a green pathway to produce H 2 O 2 , a popular industrial oxidant. Here, in this study, we theoretically investigate the effects of Pd oxidation states, coordination environments, and particle sizes on primary H 2 O 2 selectivities, assessed by calculating the ratio of rate constants for the formation of H 2 O 2 (via OOH* reduction; k O–H ) and the decomposition of OOH* (via O–O cleavage; k O–O ). For Pd metals, the k O–H /k O–O ratio decreased from 10 -4 for Pd(111) to 10 -10 for the Pd 13 cluster at 300 K, indicating poorer H 2 O 2 selectivity as Pd particle size decreases and low primary selectivities for H 2 O 2 overall. As the oxygen chemical potential increases and metals form surface and bulk oxides, the perturbation of Pd–Pd ensemble sites by lattice O atoms results in selectivities that become dramatically higher than unity. For instance, at 300 K, the k O–H /k O–O ratio increases significantly from 10 -4 to 10 9 to 10 16 as Pd(111) oxidizes to Pd 5 O 4 /Pd(111) and to PdO(100), respectively. In contrast, such selectivity enhancements are not observed for surface and bulk oxides that persistently contain rows of more metallic, undercoordinated Pd–Pd ensemble sites, such as PdO(101)/Pd(100) and PdO(101). These Pd–Pd ensembles are also absent when smaller Pd nanoparticles fully oxidize, indicating that smaller PdO clusters can be more selective for H 2 O 2 synthesis. These trends for primary H 2 O 2 selectivities were found to inversely correlate with trends for H 2 O 2 decomposition rates via O–O bond cleavage, demonstrating that catalysts with high primary H 2 O 2 selectivity can also hinder H 2 O 2 decomposition. Ab initio thermodynamic calculations are used to estimate the thermodynamically favored phase among Pd, PdO/Pd and PdO in O 2 , H 2 O 2 /H 2 O, and O 2 /H 2 environments. These results are combined to show that smaller Pd nanoparticles are more prone to be oxidized at lower oxygen chemical potentials, upon which they become more selective than larger Pd particles for H 2 O 2 synthesis.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Selecting representative geological realizations to model subsurface CO 2 storage under uncertainty

Carbon capture and storage (CCS) is one of the quickest and most effective solutions for reducing carbon emissions. The majority of subsurface storage occurs in saline aquifers, for which geological information is lacking which in turn results in geological uncertainty. To evaluate uncertainty in CO 2 injection projections, the use of multiple geological realizations (GRs) has been practiced very commonly. In this approach, hundreds or thousands of high-resolution GRs is used that quickly becomes computationally expensive. This issue can be addressed with representative geological realizations (RGRs) that preserve the uncertainty domain of the ensemble GRs. Here, in this study, we propose the use of unsupervised machine learning (UML) frameworks, including dissimilarity measurement, dimensionality reduction, clustering and sampling algorithms ta select a predetermined number of RGRs. We compare the simulation outputs of the RGR sets and the ensemble using the Kolmogorov–Smirnov (KS) test to select the best UML. The UML frameworks and their associated selection processes are evaluated using a saline aquifer with a single CO 2 injection well and 200 GRs with varying uncertain petrophysical characteristics. The best UML framework is selected to use only 5% of the GRs while maintaining the uncertainty domain of the ensemble GRs. In addition, the best UML framework is tested using a saline aquifer with three CO 2 injection wells and varied GRs. The results show that our proposed UML framework can be used to choose RGRs, capturing the whole uncertainty domain. Our approach leads to a significant reduction in the computational cost associated with scenario testing, decision-making, and development planning for CO 2 storage sites under geological uncertainty.

58 GEOSCIENCES↗

Properties of protein unfolded states suggest broad selection for expanded conformational ensembles

Much attention is being paid to conformational biases in the ensembles of intrinsically disordered proteins. However, it is currently unknown whether or how conformational biases within the disordered ensembles of foldable proteins affect function in vivo. Recently, we demonstrated that water can be a good solvent for unfolded polypeptide chains, even those with a hydrophobic and charged sequence composition typical of folded proteins. These results run counter to the generally accepted model that protein folding begins with hydrophobicity-driven chain collapse. Here we investigate what other features, beyond amino acid composition, govern chain collapse. We found that local clustering of hydrophobic and/or charged residues leads to significant collapse of the unfolded ensemble of pertactin, a secreted autotransporter virulence protein from Bordetella pertussis , as measured by small angle X-ray scattering (SAXS). Sequence patterns that lead to collapse also correlate with increased intermolecular polypeptide chain association and aggregation. Crucially, sequence patterns that support an expanded conformational ensemble enhance pertactin secretion to the bacterial cell surface. Similar sequence pattern features are enriched across the large and diverse family of autotransporter virulence proteins, suggesting sequence patterns that favor an expanded conformational ensemble are under selection for efficient autotransporter protein secretion, a necessary prerequisite for virulence. More broadly, we found that sequence patterns that lead to more expanded conformational ensembles are enriched across water-soluble proteins in general, suggesting protein sequences are under selection to regulate collapse and minimize protein aggregation, in addition to their roles in stabilizing folded protein structures.

59 BASIC BIOLOGICAL SCIENCES↗

Accelerated, scalable and reproducible AI-driven gravitational wave detection

The development of reusable artificial intelligence (AI) models for wider use and rigorous validation by the community promises to unlock new opportunities in multi-messenger astrophysics. Here we develop a workflow that connects the Data and Learning Hub for Science, a repository for publishing AI models, with the Hardware-Accelerated Learning (HAL) cluster, using funcX as a universal distributed computing service. Using this workflow, an ensemble of four openly available AI models can be run on HAL to process an entire month's worth (August 2017) of advanced Laser Interferometer Gravitational-Wave Observatory data in just seven minutes, identifying all four binary black hole mergers previously identified in this dataset and reporting no misclassifications. This approach combines advances in AI, distributed computing and scientific data infrastructure to open new pathways to conduct reproducible, accelerated, data-driven discovery. By combining a repository for artificial intelligence models and a supercomputing cluster, an entire month's worth of advanced LIGO data is analysed in just 7 min, finding all binary black hole mergers previously identified in this dataset and reporting no misclassifications.

79 ASTRONOMY AND ASTROPHYSICS↗

Companions to isolated elliptical galaxies: revisiting the Bothun-Sullivan (1977) sample using the NASA/IPAC extragalactic database

We investigate the number of physical companion galaxies for a sample of relatively isolated elliptical galaxies. The NASA/IPAC Extragalactic Database (NED) has been usedto reinvestigate the incidence of satellite galaxies for a sample of 34 elliptical galaxies, firstinvestigated by Bothun & Sullivan (1977) using a visual inspection of Palomar Sky Survey prints out to a projected search radius of 75 kpc. We have repeated their original investigation usingdata cataloged data in NED. Nine of these ellipticals appear to be members of galaxy clusters:the remaining sample of 25 galaxies reveals an average of +1.0 f 0.5 apparent companions per galaxy within a projected search radius of 75 kpc, in excess of two equal-area comparisonregions displaced by 150-300 kpc. This is nearly an order of magnitude larger than the +0.12+/- 0.42 companions/galaxy found by Bothun & Sullivan for the identical sample. Making use of published radial velocities, mostly available since the completion of the Bothun-Sullivan study,identifies the physical companions and gives a somewhat lower estimate of +0.4 companions per elliptical. This is still a factor of 3x larger than the original statistical study, but giventhe incomplete and heterogeneous nature of the survey redshifts in NED, it still yields a firmlower limit on the number (and identity) of physical companions. An expansion of the searchradius out to 300 kpc, again restricted to sampling only those objects with known redshifts in NED, gives another lower limit of 4.3 physical companions per galaxy. (Excluding fiveelliptical galaxies in the Fornax cluster this average drops to 3.5 companions per elliptical.)These physical companions are individually identified and listed, and the ensemble-averagedradial density distribution of these associated galaxies is presented. For the ensemble, the radial density distribution is found to have a fall-off consistent with p c( R^-0.5 out to approximately150 kpc. For non-Fornax cluster companions the fall-off continues out to the 300-kpc limit of thesurvey. The velocity dispersion of these companions is found to be constant with projected radial distance from the central elliptical, holding at a value of approximately +/- 300-350 km/sec overall.

NED NASA/IPAC Extragalactic Database↗

Spotlight: efficient automated global optimization in rietveld analysis of diffraction data

Performing reliable Rietveld analysis on tens or hundreds of powder diffraction datasets from parametric or time-resolved experiments often poses a bottleneck in extracting meaningful results from the data. While automated analysis of data has recently been demonstrated, high temperature annealing studies, during which phase transformations occur and lattice parameters may change due to repartitioning of elements, are prime examples where automation by a simple phase identification from a database of room temperature structures or automation by sequential refinements is likely to fail. To enable reliable, efficient, automated Rietveld analysis, we present a Python package named Spotlight , building on established Rietveld packages such as MAUD, GSAS , or GSAS-II , which extends the refinement of best fit parameters to a global optimization using an ensemble of optimizers leveraging hierarchical parallel execution on high-performance computing clusters. Spotlight further enables the efficient design of refinement plans through the iterative automated machine-learning of a surrogate for the refinement on which the global optimizations are performed until results from the surrogate converge to the response surface data. We demonstrate Spotlight with the analysis of uranium molybdenum and Ti–6Al–4V datasets, as well as in two open-source tutorials analyzing aluminium oxide and lead sulphate.

36 MATERIALS SCIENCE↗

DLSIA: Deep Learning for Scientific Image Analysis

DLSIA (Deep Learning for Scientific Image Analysis) is a Python-based machine learning library that empowers scientists and researchers across diverse scientific domains with a range of customizable convolutional neural network (CNN) architectures for a wide variety of tasks in image analysis to be used in downstream data processing. DLSIA features easy-to-use architectures, such as autoencoders, tunable U-Nets and parameter-lean mixed-scale dense networks (MSDNets). Additionally, this article introduces sparse mixed-scale networks (SMSNets), generated using random graphs, sparse connections and dilated convolutions connecting different length scales. For verification, several DLSIA-instantiated networks and training scripts are employed in multiple applications, including inpainting for X-ray scattering data using U-Nets and MSDNets, segmenting 3D fibers in X-ray tomographic reconstructions of concrete using an ensemble of SMSNets, and leveraging autoencoder latent spaces for data compression and clustering. As experimental data continue to grow in scale and complexity, DLSIA provides accessible CNN construction and abstracts CNN complexities, allowing scientists to tailor their machine learning approaches, accelerate discoveries, foster interdisciplinary collaboration and advance research in scientific image analysis.

97 MATHEMATICS AND COMPUTING↗

A Clustering-based biased Monte Carlo Approach to Protein Titration Curve Prediction

We develop and implement a novel approach to computing the ensemble averages in systems characterized by pair-wise interactions between the entities. Methods involving full enumeration of the configuration space result in exponential complexity. Sampling methods such as Markov Chain Monte Carlo (MCMC) algorithms have been proposed to tackle the exponential complexity of these problems. In certain scenarios where significant energetic coupling exists between the entities, the accuracy of the such algorithms can be diminished. We propose a strategy to improve the accuracy of the MCMC runs by taking advantage of the cluster structure in the interaction energy matrix. We propose two different schemes for performing the biased MCMC runs on the partitioned systems and show that they are valid MCMC schemes. We then apply these algorithms to the problem of computing the protonation fractions and hence the titration curves of titratable protein residues that constitute a given protein. We leverage both synthesized and real-world systems and show the improved performance of our biased MCMC methods when compared to the regular MCMC method.

Visweswara Sathanur, Arun↗

Electrocatalytic alkene epoxidation at disrupted metal ensembles in blended electrolytes

The project aims to achieve a molecular understanding of oxygen-atom transfer from water to alkenes at electrocatalytic interfaces. Molecular oxygen is the most common oxygen-atom source for epoxidations, and our group is developing sustainable routes through which epoxidation of olefins is achieved using water as the oxygen source. This route can improve the safety of the reaction while also co-producing hydrogen, demonstrating the relevance of this reaction to the energy transition. If successful in our efforts, we may enable oxygen-atom transfer reactions at the anode of water electrolyzers in the place of conventional oxygen evolution, allowing for the synthesis of sustainable value-added co-products. In this vein, we explore several approaches to acquiring high selectivity toward epoxidation over competing reactions, such as oxygen evolution. One of our aims is to allow for rational control of epoxide selectivity by disrupting contiguous metal ensembles at the surface of catalytic metal oxide nanoparticles. Specifically, we aim to synthesize single-atom, few-atom, and many-atom clusters supported on metal oxides and study the mechanism of oxygen evolution and alkene epoxidation on these materials. Thus, this approach will determine the impact of disrupting metal ensembles on the selectivity for alkene epoxidation versus oxygen evolution in blended electrolytes. Another aim is to develop a molecular-level understanding of how a blended electrolyte (i.e., a mixture of aqueous and organic solvents) influences rates of alkene epoxidation versus oxygen evolution. In other words, we are interested in understanding the catalytic influence of the solvent, as our preliminary work shows that the selectivity and reactivity of epoxidation depend strongly on the solvent composition. This investigation includes blended electrolytes and electrolytes containing redox mediator species that improve selectivity toward the desired epoxidation reaction. Overall, our proposed work will help to provide a detailed molecular-level picture of how solvents interact with substrates at the electrode-electrolyte interface, including their involvement in proton transfer reactions and screening of electric fields.

14 SOLAR ENERGY↗

Uncertainty estimation of bifurcated solutions in the Rayleigh–Bénard problem for advanced nuclear reactors applications

Multiphysics models of nuclear reactors frequently comprise nonlinear systems of equations. The nonlinear nature of these models could lead to solution bifurcations, where a small change in a certain parameter, e.g., the thermophysical properties of the coolant, can lead to a sudden change in the system’s behavior. At the point in parameter space where this happens, called a critical point, the Jacobian matrix of the model’s nonlinear operator becomes singular potentially permitting multiple solutions to coexist. In this paper, we perform uncertainty estimation (UE) in a parameter range that includes bifurcated solutions within the context of Rayleigh–Bénard problem. We perform this analysis assuming uncertain temperature difference, and tilt angle for the iterative solution algorithm with a unit Prandtl number (Pr = 1). Also, we perform this analysis under uncertain thermophysical properties for both FLiBe molten salt and liquid sodium as working fluid. We deploy two approaches to compute statistical moments for the resulting distributions of selected flow-field variables. The first approach is the blind computation of the mean and the standard deviation without any consideration of solution bifurcation, while the second approach utilizes k-means clustering to cluster each branch’s solutions together and compute separate statistical moments for each branch. The statistical distributions are obtained by perturbing the selected parameters about nominal values that correspond to a solution on one of the valid branches, and that solution is used as initial guess for the iterative solution algorithm. We found that perturbation of any parameter when its nominal value is close to its critical point always leads to branch jumping, i.e., the iterations converge to a solution on a branch different from the branch of the initial guess. This produces a statistical ensemble comprised of fundamentally different solutions leading to wrong mean values and uncertainty estimates, whereas clustering provides an efficient way to deal with this type of computation. This work is important for developing Gen IV nuclear systems because many of these systems rely on natural convection for cooling especially in accident conditions.

97 - MATHEMATICS AND COMPUTING↗

Scalar Field Comparison with Topological Descriptors: Properties and Applications for Scientific Visualization

In topological data analysis and visualization, topological descriptors such as persistence diagrams, merge trees, contour trees, Reeb graphs, and Morse–Smale complexes play an essential role in capturing the shape of scalar field data. Herein we present a state–of–the–art report on scalar field comparison using topological descriptors. We provide a taxonomy of existing approaches based on visualization tasks associated with three categories of data: single fields, time–varying fields, and ensembles. These tasks include symmetry detection, periodicity detection, key event/feature detection, feature tracking, clustering, and structure statistics. Our main contributions include the formulation of a set of desirable mathematical and computational properties of comparative measures, and the classification of visualization tasks and applications that are enabled by these measures.

97 MATHEMATICS AND COMPUTING↗

A Semi-supervised Hybrid Machine Learning Framework for the Qualification of Resistance Spot Welds

• Industries requiring high structural integrity, including automotive, aerospace, and construction, place considerable significance on weld quality classification. • The inspection normally involves human expertise through predefined quality metrics that are subjective, error-prone, and time-intensive • The challenge to classification model development is the scarcity of labeled data and imbalanced distributions in the data that are labeled. • This work develops a new hybrid methodology that achieves clustering using KMeans++ together with supervised classification to overcome these challenges. • The ensemble-based classifiers were identified as optimal, with accuracy enhancements of up to 8% using the pseudo-labeled dataset. • The work provides practical insight into feature engineering and machine learning integration in industrial quality assurance applications.

Rogers, Jeremy K. [Savannah River National Laborat↗

The Parallel System for Integrating Impact Models and Sectors (pSIMS)

We present a framework for massively parallel climate impact simulations: the parallel System for Integrating Impact Models and Sectors (pSIMS). This framework comprises a) tools for ingesting and converting large amounts of data to a versatile datatype based on a common geospatial grid; b) tools for translating this datatype into custom formats for site-based models; c) a scalable parallel framework for performing large ensemble simulations, using any one of a number of different impacts models, on clusters, supercomputers, distributed grids, or clouds; d) tools and data standards for reformatting outputs to common datatypes for analysis and visualization; and e) methodologies for aggregating these datatypes to arbitrary spatial scales such as administrative and environmental demarcations. By automating many time-consuming and error-prone aspects of large-scale climate impacts studies, pSIMS accelerates computational research, encourages model intercomparison, and enhances reproducibility of simulation results. We present the pSIMS design and use example assessments to demonstrate its multi-model, multi-scale, and multi-sector versatility.

crop modeling↗

Synergistic Retrievals of Ice in High Clouds From Elastic Backscatter Lidar, Ku-band Radar and Submillimeter Wave Radiometer Observations

In this study, we investigate the synergy of elastic backscatter lidar, Ku-band radar, and sub-millimeter-wave radiometer measurements in the retrieval of ice from satellite observations. The synergy is analyzed through the generation of a large dataset of IceWater Content (IWC) profiles and simulated lidar, radar and radiometer observations. The characteristics of the instruments e.g. frequencies, sensitivities, etc. are set based on the expected characteristics of instruments of the Atmosphere Observing System (AOS) mission. A hold-out validation methodology is used to assess the accuracy of the IWC profiles retrieved from various combinations of observations from the three instruments. Specifically, the IWC and associated observations are randomly divided into two datasets, one for training and the other for evaluation. The training dataset is used to train the retrieval algorithm, while the evaluation dataset is used to assess the retrieval performance. The dataset of IWC profiles is derived from CloudSat reflectivity and CALIOP lidar observations. The retrieval of the ice water content IWC profiles from the computed observations is achieved in two steps. In the first step, a class, out of 18 potential classes characterized by different vertical distribution of IWC, is estimated from the observations. The 18 classes are predetermined based on the k-Means clustering algorithm. In the second step, the IWC profile is estimated using an Ensemble Kalman Smoother (EKS) algorithm that uses the estimated class as a priori information. The results of the study show that the synergy of lidar, radar, and radiometer observations is significant in the retrieval of the IWC profiles. Nevertheless, it should be mentioned that this synergy was found under idealized conditions, and additional work might be required to materialize it in practice. The inclusion of the lidar backscatter observations in the retrieval process has a larger impact on the retrieval performance than the inclusion of the radar observations. As ice clouds have a significant impact on atmospheric radiative processes, this work is relevant to ongoing efforts to reduce uncertainties in climate analyses and projections.

Mircea Grecu↗

Bespoke Liquid/Liquid Interfaces (Final Technical Report)

The goal of DE-SC0001815 was to advance the basic science of liquid:liquid interface formation, to develop a deeper understanding of the mechanisms of phase separation and the essential relationships between solution composition, organization and dynamics that underlie the kinetic regime of solvent extraction. This included learning how interfacial organization and dynamics alters the properties of the primary coordination sphere of ions and the free energy of transport of ions complexes across a phase boundary. We relied primarily upon classical molecular dynamics studies to determine the equilibrium ensembles of these complex systems, but also utilized ab-initio MD and cluster-based density functional theory (DFT) calculations when more detailed investigation of the electronic structure was needed. We continued development of graph-theory based analyses to elucidate hierarchical correlations and expanded into geometric topology methods to quantify the collectively organized structures that can organize at a liquid/liquid interface during solute transport. One of the main conclusions was from the observation of two distinct mechanisms for solute transport - those that derive from amplifications of interfacial heterogeneity and surface roughness, and those wherein surface roughness has been dampened and instead collectively organized macrostructures work to bring solutes into the organic phase. It was our aim to create a concrete chemical model of the underlying driving forces behind interfacial primary and secondary structure formation and to map out the energetic features of solute transport so that tailored liquid/liquid can be developed that have characteristic kinetic features associated with mass transport.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗