Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “distributed clustering methods”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

Propagating sample variance uncertainties in redshift calibration: simulations, theory, and application to the COSMOS2015 data

ABSTRACT Cosmological analyses of galaxy surveys rely on knowledge of the redshift distribution of their galaxy sample. This is usually derived from a spectroscopic and/or many-band photometric calibrator survey of a small patch of sky. The uncertainties in the redshift distribution of the calibrator sample include a contribution from shot noise, or Poisson sampling errors, but, given the small volume they probe, they are dominated by sample variance introduced by large-scale structures. Redshift uncertainties have been shown to constitute one of the leading contributions to systematic uncertainties in cosmological inferences from weak lensing and galaxy clustering, and hence they must be propagated through the analyses. In this work, we study the effects of sample variance on small-area redshift surveys, from theory to simulations to the COSMOS2015 data set. We present a three-step Dirichlet method of resampling a given survey-based redshift calibration distribution to enable the propagation of both shot noise and sample variance uncertainties. The method can accommodate different levels of prior confidence on different redshift sources. This method can be applied to any calibration sample with known redshifts and phenotypes (i.e. cells in a self-organizing map, or some other way of discretizing photometric space), and provides a simple way of propagating prior redshift uncertainties into cosmological analyses. As a worked example, we apply the full scheme to the COSMOS2015 data set, for which we also present a new, principled SOM algorithm designed to handle noisy photometric data. We make available a catalogue of the resulting resamplings of the COSMOS2015 galaxies.

79 ASTRONOMY AND ASTROPHYSICS↗

Multiscale Modeling of Nanoparticle Precipitation in Oxide Dispersion-Strengthened Steels Produced by Laser Powder Bed Fusion

Laser Powder Bed Fusion (LPBF) enables the efficient production of near-net-shape oxide dispersion-strengthened (ODS) alloys, which possess superior mechanical properties due to oxide nanoparticles (e.g., yttrium oxide, Y-O, and yttrium-titanium oxide, Y-Ti-O) embedded in the alloy matrix. To better understand the precipitation mechanisms of the oxide nanoparticles and predict their size distribution under LPBF conditions, we developed an innovative physics-based multiscale modeling strategy that incorporates multiple computational approaches. These include a finite volume method model (Flow3D) to analyze the temperature field and cooling rate of the melt pool during the LPBF process, a density functional theory model to calculate the binding energy of Y-O particles and the temperature-dependent diffusivities of Y and O in molten 316L stainless steel (SS), and a cluster dynamics model to evaluate the kinetic evolution and size distribution of Y-O nanoparticles in as-fabricated 316L SS ODS alloys. The model-predicted particle sizes exhibit good agreement with experimental measurements across various LPBF process parameters, i.e., laser power (110–220 W) and scanning speed (150–900 mm/s), demonstrating the reliability and predictive power of the modeling approach. The multiscale approach can be used to guide the future design of experimental process parameters to control oxide nanoparticle characteristics in LPBF-manufactured ODS alloys. Additionally, our approach introduces a novel strategy for understanding and modeling the thermodynamics and kinetics of precipitation in high-temperature systems, particularly molten alloys.

Wang, Zhengming (ORCID:0000000241627112)↗

Nuclear Computational Low Energy Initiative (NUCLEI)

The NUCLEI project, as defined by the scope of work, developed, implemented and run codes for large-scale computations of many topics in low-energy nuclear physics. Physics studied include the properties of nuclei and nuclear decays, nuclear structure and reactions, and the properties of nuclear matter. The computational techniques used include Quantum Monte Carlo, Configuration Interaction, Coupled Cluster, and Density Functional methods. The research program emphasized areas of high interest to current and possible future DOE nuclear physics facilities, including ATLAS and FRIB (nuclear structure and reactions, and nuclear astrophysics), TJNAF (neutron distributions in nuclei, few body systems, and electroweak processes), NIF (thermonuclear reactions), MAJORANA and FNPB (neutrinoless double-beta decay and physics beyond the Standard Model), and LANSCE (fission studies).

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Electricity Reliability Metrics, Analysis, and Planning (CRADA Final Report)

LBNL and APPA (the team) jointly examined the extent to which differences in distribution feeder characteristics are correlated with differences in their reliability performance when exposed to three different types of natural hazards (wildlife, weather, and vegetation). The team employed data-driven approaches to quantify the relationships between various measures of feeder reliability and a suite of feeder characteristics individually and jointly via a statistically-based clustering method. The team developed suggestions on how comparisons across groupings of feeders and review of the relative contributions of the constituents of SAIFI and SAIDI could be used to help prioritize utility actions to improve reliability. However, they also caution that their suggestions require further evaluation because they are based on only one year of information from a modest number of small utilities.

24 POWER TRANSMISSION AND DISTRIBUTION↗

The emergence and transmission dynamics of HIV-1 CRF07_BC in Mainland China

A total of 1155 partial pol gene sequences of human immunodeficiency virus (HIV)-1 CRF07_BC were sampled between 1997 and 2015, spanning 13 provinces in Mainland China and risk groups [heterosexual, injecting drug users (IDU), and men who have sex with men (MSM)] to investigate the evolution, adaptation, spatiotemporal and risk group dynamics, migration patterns, and protein structure of HIV-1 CRF07_BC. Due to the unequal distribution of sequences across time, location, and risk group in the complete dataset (‘full1155’), subsampling methods were used. Maximum-likelihood and Bayesian phylogenetic analysis as well as discrete trait analysis of geographical location and risk group were carried out. To study mutations of a cluster of HIV-1 CRF07_BC (CRF07-1), we performed a comparative analysis of this cluster to the other CRF07_BC sequences (‘backbone_295’) and mapped the mutations observed in the respective protein structure. Our findings showed that HIV-1 CRF07_BC most likely originated among IDU in Yunnan Province between October 1992 to July 1993 [95 per cent hightest posterior density (HPD): May 1989–August 1995] and that IDU in Yunnan Province and MSM in Guangdong Province likely served as the viral sources during the early and more recent spread in Mainland China. We also revealed that HIV-1 CRF07-1 has been spreading for roughly 20 years and continues to cause local transmission in Mainland China and worldwide. Overall, our study sheds light on the dynamics of HIV-1 CRF07_BC distribution patterns in Mainland China. Our research may also be useful in formulating public health policies aimed at controlling acquired immune deficiency syndrome in Mainland China and globally.

60 APPLIED LIFE SCIENCES↗

Combination of cluster number counts and two-point correlations: validation on mock Dark Energy Survey

ABSTRACT We present a method of combining cluster abundances and large-scale two-point correlations, namely galaxy clustering, galaxy–cluster cross-correlations, cluster autocorrelations, and cluster lensing. This data vector yields comparable cosmological constraints to traditional analyses that rely on small-scale cluster lensing for mass calibration. We use cosmological survey simulations designed to resemble the Dark Energy Survey Year 1 (DES-Y1) data to validate the analytical covariance matrix and the parameter inferences. The posterior distribution from the analysis of simulations is statistically consistent with the absence of systematic biases detectable at the precision of the DES-Y1 experiment. We compare the χ2 values in simulations to their expectation and find no significant difference. The robustness of our results against a variety of systematic effects is verified using a simulated likelihood analysis of DES-Y1-like data vectors. This work presents the first-ever end-to-end validation of a cluster abundance cosmological analysis on galaxy catalogue level simulations.

79 ASTRONOMY AND ASTROPHYSICS↗

Signatures of Light Massive Relics on non-linear structure formation

ABSTRACT Cosmologies with Light Massive Relics (LiMRs) as a subdominant component of the dark sector are well-motivated from a particle physics perspective, and can also have implications for the σ8 tension between early and late time probes of clustering. The effects of LiMRs on the cosmic microwave background (CMB) and structure formation on large (linear) scales have been investigated extensively. In this paper, we initiate a systematic study of the effects of LiMRs on smaller, non-linear scales using cosmological N-body simulations; focusing on quantities relevant for photometric galaxy surveys. For most of our study, we use a particular model of non-thermal LiMRs but the methods developed generalizing to a large class of LiMR models – we explicitly demonstrate this by considering the Dodelson–Widrow velocity distribution. We find that, in general, the effects of LiMR on small scales are distinct from those of a ΛCDM universe, even when the value of σ8 is matched between the models. We show that weak lensing measurements around massive clusters, between ∼0.1 h−1Mpc and ∼10 h−1Mpc, should have sufficient signal-to-noise in future surveys to distinguish between ΛCDM and LiMR models that are tuned to fit both CMB data and linear scale clustering data at late times. Furthermore, we find that different LiMR cosmologies indistinguishable by conventional linear probes can be distinguished by non-linear probes if their velocity distributions are sufficiently different. LiMR models can, therefore, be best tested by jointly analyzing the CMB and late-time structure formation on both large and small scales.

79 ASTRONOMY AND ASTROPHYSICS↗

Quantitative Imaging of Cobalt Phthalocyanine Distribution on Carbon Nanotubes: A Deep Learning Approach to Catalyst Characterization

Electrochemical reduction of carbon dioxide (CO 2 ) offers a pathway to valuable products, with catalysts playing a crucial role. This study investigates the distribution of cobalt tetraaminophthalocyanine (CoPc-NH 2 ) immobilized on carbon nanotubes (CNTs), utilizing high-angle annular dark-field scanning transmission electron microscopy (HAADF-STEM) to characterize CoPc-NH 2 distribution. A challenge in the quantitative HAADF-STEM analysis is the introduction of bias from manual Co atom identification. To address this, we developed and trained a convolutional neural network (CNN) using a data set generated from images of CoPc-NH 2 /CNT samples with varying Co loadings. The CNN, implemented in TensorFlow and Keras, facilitated Co atom detections. Analysis of the CNN-generated data confirmed a correlation between Co loading and surface density, consistent with findings from UV–vis spectroscopy. Furthermore, the application of Ripley’s L(d) function highlighted the presence of slight Co atom clustering. Furthermore, this work demonstrates the utility of the combined HAADF-STEM and CNN approach for providing spatially resolved information about catalyst distribution on nonplanar supports, revealing structural details that are typically lost through other characterization methods.

HAADF-STEM↗

Image Deconvolution and Point-spread Function Reconstruction with STARRED: A Wavelet-based Two-channel Method Optimized for Light-curve Extraction

We present starred, a point-spread function (PSF) reconstruction, two-channel deconvolution, and light-curve extraction method designed for high-precision photometric measurements in imaging time series. An improved resolution of the data is targeted rather than an infinite one, thereby minimizing deconvolution artifacts. In addition, starred performs a joint deconvolution of all available data, accounting for epoch-to-epoch variations of the PSF and decomposing the resulting deconvolved image into a point source and an extended source channel. The output is a high-signal-to-noise-ratio, high-resolution frame combining all data and the photometry of all point sources in the field of view as a function of time. Of note, starred also provides exquisite PSF models for each data frame. We showcase three applications of starred in the context of the imminent LSST survey and of JWST imaging: (i) the extraction of supernovae light curves and the scene representation of their host galaxy; (ii) the extraction of lensed quasar light curves for time-delay cosmography; and (iii) the measurement of the spectral energy distribution of globular clusters in the "Sparkler," a galaxy at redshift z = 1.378 strongly lensed by the galaxy cluster SMACS J0723.3-7327. starred is implemented in jax, leveraging automatic differentiation and graphics processing unit acceleration. This enables the rapid processing of large time-domain data sets, positioning the method as a powerful tool for extracting light curves from the multitude of lensed or unlensed variable and transient objects in the Rubin-LSST data, even when blended with intervening objects.

79 ASTRONOMY AND ASTROPHYSICS↗

Earthquake Phase Association Using a Bayesian Gaussian Mixture Model

Earthquake phase association algorithms aggregate picked seismic phases from a network of seismometers into individual seismic events and play an important role in earthquake monitoring and research. Dense seismic networks and improved phase picking methods produce massive seismic phase datasets, particularly for earthquake swarms and aftershocks occurring closely in time and space, making phase association a challenging problem. Here, we present a new association method, the Gaussian Mixture Model Association (GaMMA), that combines the Gaussian mixture model with earthquake location, origin time, and magnitude estimation. We treat earthquake phase association as an unsupervised clustering problem in a probabilistic framework, where each earthquake corresponds to a cluster of P and S phases with a hyperbolic moveout of arrival times and a decay of amplitude with distance. We use the multivariate Gaussian distribution to model the collection of phase picks of an event; and the mean of the multivariate Gaussian distribution is given by the predicted arrival time and amplitude from the causative event. We carry out the pick assignment to each earthquake and determine earthquake source parameters (i.e., earthquake location, origin time, and magnitude) under the maximum likelihood criterion using the Expectation-Maximization algorithm. The GaMMA method does not require typical association steps of other algorithms, such as grid-search or supervised training. The results for both synthetic tests and for the 2019 Ridgecrest earthquake sequence show that GaMMA effectively associates phases from a temporally and spatially dense earthquake sequence while producing useful estimates of earthquake location and magnitude.

58 GEOSCIENCES↗

Detecting Rain–Snow-Transition Elevations in Mountain Basins Using Wireless Sensor Networks

Here, to provide complementary information on the hydrologically important rain–snow-transition elevation in mountain basins, this study provides two estimation methods using ground measurements from basin-scale wireless sensor networks: one based on wet-bulb temperature T wet and the other based on snow-depth measurements of accumulation and ablation. With data from 17 spatially distributed clusters (178 nodes) from two networks, in the American and Feather River basins of California’s Sierra Nevada, we analyzed transition elevation during 76 storm events in 2014–18. A T wet threshold of 0.5°C best matched the transition elevation defined by snow depth. Transition elevations using T wet in upper elevations of the basins generally agreed with atmospheric snow level from radars located at lower elevations, while radar snow level was ~100 m higher due to snow-level lowering on windward mountainsides during orographic lifting. Diurnal patterns of the difference between transition elevation and radar snow level were observed in the American basin, related to diurnal ground-temperature variations. However, these patterns were not found in the Feather basin due to complex terrain and higher uncertainties in transition-elevation estimates. The American basin tends to exhibit 100-m-higher transition elevations than does the Feather basin, consistent with the Feather basin being about 1° latitude farther north. Transition elevation averaged 155 m higher in intense atmospheric river events than in other events; meanwhile, snow-level lowering was enhanced with a 90-m-larger difference between radar snow level and transition elevation. On-the-ground continuous observations from distributed sensor networks can complement radar data and provide important ground truth and spatially resolved information on transition elevations in mountain basins.

54 ENVIRONMENTAL SCIENCES↗

Performance and power modeling and prediction using MuMMI and 10 machine learning methods

Energy-efficient scientific applications require insight into how high performance computing system features impact the applications' power and performance. This insight can result from the development of performance and power models. Here, in this article, we use the modeling and prediction tool MuMMI (Multiple Metrics Modeling Infrastructure) and 10 machine learning methods to model and predict performance and power consumption and compare their prediction error rates. We use an algorithm-based fault-tolerant linear algebra code and a multilevel checkpointing fault-tolerant heat distribution code to conduct our modeling and prediction study on the Cray XC40 Theta and IBM BG/Q Mira at Argonne National Laboratory and the Intel Haswell cluster Shepard at Sandia National Laboratories. Our experimental results show that the prediction error rates in performance and power using MuMMI are less than 10% for most cases. By utilizing the models for runtime, node power, CPU power, and memory power, we identify the most significant performance counters for potential application optimizations, and we predict theoretical outcomes of the optimizations. Based on two collected datasets, we analyze and compare the prediction accuracy in performance and power consumption using MuMMI and 10 machine learning methods.

97 MATHEMATICS AND COMPUTING↗

Predicting U 3 O 8 powder processing conditions: An AI/ML approach analyzing deep learning embeddings of SEM micrographs

High-resolution SEM images of uranium-oxide powders encode micro- and nanoscale clues to their synthesis route and calcination temperature. We trained a ResNet-50 model on 11 commercial-scale U₃O₈ classes, ammonium diuranate (ADU) or uranyl peroxide (H₂O₂) precursors calcined at temperatures ranging from 400 to 750 °C and added a 256-D projection head before the classifier to analyze the learned representation. The best of eight seeds reached 92.4 % accuracy on reserved testing data, but our focus is the structure of the embedding space rather than the accuracy and labels. We quantify class relatedness in the original 256-D space using centroid similarity and distributional distances, and we use Uniform Manifold Approximation Projection (UMAP) for visualization. ‘Unknown’ images from different preparation methods, SEM operators, and from the literature localized near the expected classes under a nearest-centroid analysis without retraining, as well as clustered in similar UMAP space. In conclusion, this embedding-centered workflow complements black-box classification by providing quantitative, similarity-based comparisons of U₃O₈ morphologies and reduces storage space by up to 98 % for image data used in millisecond vector search comparisons.

36 MATERIALS SCIENCE↗

Adaptive Hierarchical Cyber Attack Detection and Localization in Active Distribution Systems

Development of a cyber security strategy for the active distribution systems is challenging due to the inclusion of distributed renewable energy generations. Here this paper proposes an adaptive hierarchical cyber attack detection and localization framework for distributed active distribution systems via analyzing electrical waveforms. Cyber attack detection is based on a sequential deep learning model, via which even minor cyber attacks can be identified. The two-stage cyber attack localization algorithm first estimates the cyber attack sub-region, and then localize the specified cyber attack within the estimated subregion. We propose a modified spectral clustering-based network partitioning method for the hierarchical cyber attack ‘coarse’ localization. Next, to further narrow down the cyber attack location, a normalized impact score based on waveform statistical metrics is proposed to obtain a ‘fine’ cyber attack location by characterizing different waveform properties. Finally, compared with classical and state-of-art methods, a comprehensive quantitative evaluation with two case studies shows promising estimation results of the proposed framework.

42 ENGINEERING↗

Probing Heterogeneity in Bovine Enamel Composition through Nanoscale Chemical Imaging using Atom Probe Tomography

Objective: The aim of this study was to determine the heterogeneity in chemical composition of bovine enamel using atom probe tomography, and thereby evaluate the suitability of bovine enamel as a substitute for human enamel in in vitro dental research. Design: Enamel samples from extracted bovine incisor teeth were first sectioned using a diamond saw and then milled into needle-like samples (<100 nm diameter) by focused ion beam (FIB) coupled with a scanning electron microscope (SEM). The samples were then analyzed in the atom probe to acquire three-dimensional (3D) images and quantify the atomic chemistry and distribution in bovine enamel. Results: For the first time, the atomic-level composition and clustering of major constituents and impurities within bovine enamel were determined and imaged. We discovered that the chemical composition of bovine enamel is spatially inhomogeneous at the atomic scale. The average bulk Ca/P ratio, ~1.4, was in agreement with previously reported literature values from alternative conventional methods. When assessed locally at the atomic scale, the Ca/P ratio varied between 1.1 and 2.03. We also discovered that the Mg impurities were significantly segregated throughout the enamel, and such clustering influenced the variation of Ca/P ratios. The increase in Mg concentrations, near the Mg clusters, correlated with increased Ca and decreased P concentrations. Conclusion: In conclusion, the presented findings of variability in local composition should be taken into account when interpreting dental research results from bovine enamel.

59 BASIC BIOLOGICAL SCIENCES↗

Distributed Tomographic Reconstruction with Quantization

Conventional tomographic reconstruction typically depends on centralized servers for both data storage and computation, leading to concerns about memory limitations and data privacy. Distributed reconstruction algorithms mitigate these issues by partitioning data across multiple nodes, reducing server load and enhancing privacy. However, these algorithms often encounter challenges related to memory constraints and communication overhead between nodes. In this paper, we introduce a decentralized Alternating Directions Method of Multipliers (ADMM) with configurable quantization. By distributing local objectives across nodes, our approach is highly scalable and can efficiently reconstruct images while adapting to available resources. To overcome communication bottlenecks, we propose two quantization techniques based on K-means clustering and JPEG compression. Numerical experiments with benchmark images illustrate the tradeoffs between communication efficiency, memory use, and reconstruction accuracy.

Miao, Runxuan↗

Flexible silicon photonic architecture for accelerating distributed deep learning

The increasing size and complexity of deep learning (DL) models have led to the wide adoption of distributed training methods in datacenters (DCs) and high-performance computing (HPC) systems. However, communication among distributed computing units (CUs) has emerged as a major bottleneck in the training process. In this study, we propose Flex-SiPAC, a flexible silicon photonic accelerated compute cluster designed to accelerate multi-tenant distributed DL training workloads. Flex-SiPAC takes a co-design approach that combines a silicon photonic hardware platform with a tailored collective algorithm, optimized to leverage the unique physical properties of the architecture. The hardware platform integrates a novel wavelength-reconfigurable transceiver design and a micro-resonator-based wavelength-reconfigurable switch, enabling the system to achieve flexible bandwidth steering in the wavelength domain. The collective algorithm is designed to support reconfigurable topologies, enabling efficient all-reduce communications that are commonly used in DL training. The feasibility of the Flex-SiPAC architecture is demonstrated through two testbed experiments. First, an optical testbed experiment demonstrates the flexible routing of wavelengths by shuffling an array of input wavelengths using a custom-designed spatial-wavelength selective switch. Second, a four-GPU testbed running two DL workloads shows a 23% improvement in job completion time compared to a similarly sized leaf-spine topology. We further evaluate Flex-SiPAC using large-scale simulations, which show that Flex-SiPAC is able to reduce the communication time by 26% to 29% compared to state-of-the-art compute clusters under representative collective operations.

Wu, Zhenguo (ORCID:0000000322847985)↗

rustpix

rustpix is a high-performance, open-source Rust library with first-class Python bindings (via PyO3) for processing pixel-detector data in neutron imaging. It targets time-stamping detectors such as Timepix3 (TPX3) at ORNL's Spallation Neutron Source (VENUS beamline), where each detected neutron deposits charge across a cluster of pixels within a very high-rate event stream (96M+ hits/sec). rustpix parses TPX3 event data in parallel using memory-mapped I/O, offers four interchangeable clustering algorithms (ABS adjacency-based search, DBSCAN, graph/union-find connected components, and a parallel grid method), and extracts weighted, super-resolved centroids to produce neutron-event lists. A streaming architecture lets it process files larger than available memory. rustpix is distributed as a pip-installable Python package (with NumPy integration), Rust crates, a command-line tool, and an interactive GUI; it writes HDF5, Apache Arrow, and CSV; and it is designed to extend to TPX4 and other detector types. Released as open-source under the MIT License.

Zhang, Chen [Oak Ridge National Laboratory (ORNL),↗