Engineering PapersSearch

SEARCH · Engineering Papers

Results for “cluster computing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Maskman

SAND2025-04369O Maskman is a user-friendly tool designed to create hex masks, which are essential for optimizing application performance in high-performance computing environments. By converting a list of integers into binary and then hex masks, Maskman simplifies the process of setting application affinity. This ensures that software runs efficiently on specific nodes within a computing cluster. Ideal for researchers and developers, Maskman streamlines the preparation of inputs for HPC schedulers, enhancing resource management and improving overall system performance. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Pase, Douglas [Sandia National Lab. (SNL-CA), Live

Competitiveness Assessment of Decarbonizing Electricity and Process Heat Supply to a Campus with a Small Nuclear Reactor

This paper analyzes the competitiveness of siting a small nuclear reactor to support decarbonization of sites requiring tens of MW of electricity and/or process heat to support centralized heating and cooling system. This paper focuses on campuses as representative of sites with collections of buildings and research facilities with decarbonization needs represented by buildings heating, and electricity consumption by electrical loads which may include cooling via chilled water (e.g., for air conditioning and to cool down computer clusters). A nuclear reactor can be considered to decarbonize a site’s high-temperature steam generation used mostly for building heating needs, climate control, and hot water, by supplying process heat capabilities, while electricity decarbonization would be achieved mostly by the grid. However, a secondary application can be considered to maximize reactor utilization and avoid ramping down the reactor if the steam demand varies significantly throughout the year. Chilled water generation through steam-driven systems was identified as an attractive secondary option for the site analyzed, due to potential for plant design simplification, while electricity generation could be considered as well to reduce electricity purchases for a wider range of site applications. For a campus with peak 60MW thermal power demand, a small nuclear reactor with similar thermal power rating would almost eliminate CO2 emissions from steam generation and reduce electricity imports for chilled water production. A preliminary techno-economic feasibility study shows that a small nuclear reactor design that is optimized to support process heat can represent an economically feasible option when compared with other decarbonization alternatives.

Stauff, Nicolas E.

Predicting Mechanical Properties from Microstructure Images in Fiber-Reinforced Polymers Using Convolutional Neural Networks

Evaluating the mechanical response of fiber-reinforced composites can be extremely time-consuming and expensive. Machine learning (ML) techniques offer a means for faster predictions via models trained on existing input–output pairs and have exhibited success in composite research. This paper explores a fully convolutional neural network modified from StressNet, which was originally used for linear elastic materials, and extended here for a non-linear finite element (FE) simulation to predict the stress field in 2D slices of segmented tomography images of a fiber-reinforced polymer specimen. The network was trained and evaluated on data generated from the FE simulations of the exact microstructure. The testing results show that the trained network accurately captures the characteristics of the stress distribution, especially on fibers, solely from the segmented microstructure images. The trained model can make predictions within seconds in a single forward pass on an ordinary laptop, given the input microstructure, compared to 92.5 h to run the full FE simulation on a high-performance computing cluster. These results show promise in using ML techniques to conduct fast structural analysis for fiber-reinforced composites and suggest a corollary that the trained model can be used to identify the location of potential damage sites in fiber-reinforced polymers.

Sun, Yixuan (ORCID:0000000311093380)

Oak Ridge Computing Academy: An HPC cluster deployment and management pilot

The High Performance Computing Technologies (HPCT) course is a hands-on High Performance Computing (HPC) cluster deployment and management training program offered as part of the International School for Advanced Studies (SISSA) and the International Center for Theoretical Physics (ICTP) Master in High Performance Computing (MHPC) specialization. Here, this training program introduces students to key concepts in cluster configuration. which include networking, software stack provisioning, job scheduling, and monitoring. The publicly available course materials feature several examples and underlying methods that are broadly applicable to cluster deployment and management. This paper discusses the design of a new workforce development program at the Oak Ridge National Laboratory that is based on HPCT, the Oak Ridge Computing Academy (ORCA). The ORCA pilot program was hosted by the Oak Ridge Leadership Computing Facility (OLCF) in Summer 2025. As a part of this discussion, HPCT and ORCA course contents and infrastructure are outlined, ORCA participant experiences are detailed, and potential opportunities for improvement are discussed.

Education

How Well Can Quantum Embedding Method Predict the Reaction Profiles for Hydrogenation of Small Li Clusters?

Quantum computing leverages the principles of quantum mechanics in novel ways to tackle complex chemistry problems that cannot be accurately addressed using traditional quantum chemistry methods. However, the high computational cost and available number of physical qubits with high fidelity limit its application to small chemical systems. This work employed a quantum-classical framework which features a quantum active space-embedding approach to perform simulations of chemical reactions that require up to 14 qubits. This framework was applied to prototypical example metal hydrogenation reactions: the coupling between hydrogen and Li 2 , Li 3 , and Li 4 clusters. Particular attention was paid to the computation of barriers and reaction energies. The predicted reaction profiles compare well with advanced classical quantum chemistry methods, demonstrating the potential of the quantum embedding algorithm to map out reaction profiles of realistic gas-phase chemical reactions to ascertain qualitative energetic trends. Additionally, the predicted potential energy curves provide a benchmark to compare against both current and future quantum embedding approaches.

36 MATERIALS SCIENCE

Computing nuclear response functions with time-dependent coupled-cluster theory

We compute nuclear response functions by solving the time-dependent 𝐴-body Schrödinger equation, recording the time-dependent transition moment and extracting spectral information via Fourier transforms. The solution of the time-dependent many-body problem accounts for correlations on top of the mean field by taking advantage of a time-dependent formulation of coupled-cluster theory. As a validation, we focus on electric dipole transitions in 4 He and 16 O and compare moments of the response function distribution to the results of an equivalent static framework, finding negligible discrepancies. We investigate how proton and neutron densities evolve in time, and we see the traditional picture of soft and giant dipole resonances as collective oscillations of protons and neutrons emerging from our calculations in 16 O and 24 O. Furthermore, this method also allows us to investigate the behavior of the nucleus in the presence of a strong electric field. In that regime, the behavior of the system becomes chaotic. Qualitatively, the spectral information obtained in this limit is in line with previous time-dependent mean-field results.

Ab initio calculations

Using Signal Clustering Similarity for Detecting CAN Masquerade Attacks

The computer code assumes that time series representing the physical signals of the vehicle have been extracted from the CAN bus. The main input of the computer code is the multivariate time series representation of the signals in the CAN bus. The computer code cluster these time series using agglomerative hierarchical clustering from benign and attack datasets. Based on this, it generates probability distributions from the similarity of the obtained clusters based in each scenario---benign and attack---using the CluSim method (https://github.com/Hoosier-Clusters/clusim). Finally, it compares how a new data collection compares with the previous distribution to provide and probability score for an intrusion.

Moriano, Pablo

Mode-multiplexed photonic integrated vector dot-product core from inverse design

Photonic computing has the potential to harness the full degrees of freedom (DOFs) of the light field, including the wavelength, spatial mode, spatial location, phase quadrature, and polarization, to achieve a higher level of computing parallelism and scalability than digital electronic processors. While multiplexing using the wavelength and other DOFs can be readily integrated on silicon photonics platforms with compact footprints, conventional mode-division multiplexed (MDM) photonic designs occupy areas exceeding tens to hundreds of microns for a few spatial modes, significantly limiting their scalability. Here, we utilize inverse design to demonstrate an ultracompact photonic computing core that calculates vector dot products based on MDM coherent mixing. Our dot-product core integrates the functionalities of two-mode multiplexers and one multimode coherent mixer within a nominal footprint of 5 μm x 3 μm . We have experimentally demonstrated computing examples on the fabricated dot-product core, including complex number multiplication and motion estimation using optical flow. The compact dot-product core design enables large-scale on-chip integration in a parallel photonic computing primitive cluster for high-throughput scientific computing and computer vision tasks.

97 MATHEMATICS AND COMPUTING

Identifying Climate Patterns Using Clustering Autoencoder Techniques

Abstract The complexity of growing spatiotemporal resolution of climate simulations produces a variety of climate patterns under different projection scenarios. This paper proposes a new data-driven climate classification workflow via an unsupervised deep learning technique that can dimensionally reduce the vast volume of spatiotemporal numerical climate projection data into a compact representation. We aim to identify distinct zones that capture multiple climate variables as well as their future changes under different climate change scenarios. Our approach leverages convolutional autoencoders combined with k -means clustering (standard autoencoder) and online clustering based on the Sinkhorn–Knopp algorithm (clustering autoencoder) across the conterminous United States (CONUS) to capture unique climate patterns in a data-driven fashion from the Geophysical Fluid Dynamics Laboratory Earth System Model with GOLD component (GFDL-ESM2G). The developed approach compresses 70 years of GFDL-ESM2G simulation at 0.125° spatial resolution across the CONUS under multiple warming scenarios to a lower-dimensional space by a factor of 660 000 and then tested on 150 years of GFDL-ESM2G simulation data. The results show that five climate clusters capture physically reasonable and spatially stable climatological patterns matched to known climate classes defined by human experts. Results also show that using a clustering autoencoder can reduce the computational time for clustering by up to 9.2 times when compared to using a standard autoencoder. Our five unique climate patterns resulting from the deep learning–based clustering of the lower-dimensional space thereby enable us to provide insights on hydrometeorology and its spatial heterogeneity across the conterminous United States immediately without downloading large climate datasets. Significance Statement This paper presents a data-driven climate classification approach using unsupervised deep learning to dimensionally reduce climate model outputs and to identify distinct climate regions for their future changes. Our approach compresses climate information for 70 years of Geophysical Fluid Dynamics Laboratory Earth System Model data across the conterminous United States (CONUS) at 0.125° spatial resolution. The results reveal that five climate clusters capture reasonable and stable climatological patterns matched to known climate patterns. The embedded clustering process in deep learning provides ×9.2 times faster execution than the k -means clustering technique. These results give us insight about climate spatial patterns and heterogeneity of hydrological patterns across the conterminous United States without downloading large climate datasets.

Kurihana, Takuya

Parallel Programming in MCNP6

Monte Carlo N-Particle (MCNP)1 is a general-purpose Monte Carlo particle transport code developed by Los Alamos National Laboratory (LANL). To efficiently handle long simulations, MCNP version 6 (MCNP6) supports parallel execution using two primary programming models: • Shared-memory task-based threading using OpenMP (Open Multi-Processing), and • Distributed-memory calculations using MPI (Message Passing Interface). The OpenMP and MPI programming models enable MCNP6 to scale from desktop systems to high-performance computing (HPC) clusters, allowing users to run MCNP in one of three parallel modes: • OpenMP-only, • MPI-only, and • Hybrid (MPI + OpenMP). The choice of parallelization mode depends on the underlying computer architecture and the characteristics of the simulation problem.

97 MATHEMATICS AND COMPUTING

Can mesoscale models capture the effect from cluster wakes offshore?

Long wakes from offshore wind turbine clusters can extend tens of kilometers downstream, affecting the wind resource of a large area. Given the ability of mesoscale numerical weather prediction models to capture important atmospheric phenomena and mechanisms relevant to wake evolution, they are often used to simulate wakes behind large wind turbine clusters and their impact over a wider region. Yet, uncertainty persists regarding the accuracy of representing cluster wakes via mesoscale models and their wind turbine parameterizations. Here, we evaluate the accuracy of the Fitch wind farm parameterization in the Weather Research and Forecasting model in capturing cluster-wake effects using two different options to represent turbulent mixing in the planetary boundary layer. To this end, we compare operational data from an offshore wind farm in the North Sea that is fully or partially waked by an upstream array against high-resolution mesoscale simulations. In general, we find that mesoscale models accurately represent the effect of cluster wakes on front-row turbines of a downstream wind farm. However, the same models may not accurately capture cluster-wake effects on an entire downstream wind farm, due to misrepresenting internal-wake effects.

17 WIND ENERGY

Identifying Sample Provenance From SEM/EDS Automated Particle Analysis via Few-Shot Learning Coupled With Similarity Graph Clustering

Automated particle analysis (APA) provides a vast amount of compositional data via energy-dispersive X-ray spectroscopy along with size and shape data via scanning electron microscopy for individual particles in a sample. In many instances, APA data are leveraged to support identification of the source of a sample based on the detection of particles of a specific composition. Often, the particles that provide context make up a minuscule portion of the sample. Additionally, the interpretation of complex samples can be difficult due to the diversity of compositions both in the mixture and within a particle. In this work, we demonstrate a method to compute and cluster similarity graphs that describe inter-particle relationships within a sample using a multi-modal few-shot learning neural network. Here, as a proof-of-concept, we show that samples known to have been exposed to gunshot residue can be distinguished from samples occasionally mistaken for gunshot residue. Our workflow builds upon standard APA techniques and data processing methods to unveil additional information in a readily interpretable and quantitatively comparable format.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND

ELECTRONIC STRUCTURE METHODS AND PROTOCOLS WITH APPLICATION TO DYNAMICS, KINETICS AND THERMOCHEMISTRY

Hydrocarbon combustion involves the reaction dynamics of a tremendous number of species beginning with many-component fuel mixtures and proceeding via a complex system of intermediates to form primary and secondary products. Combustion conditions corresponding to new advanced engines and/or alternative fuels rely increasingly on autoignition and low-temperature-combustion chemistry. In these regimes various transient radical species such as HO2, ROO·, ·QOOH, HCO, NO2, HOCO, and Criegee intermediates play important roles in determining the detailed as well as more general dynamics. A clear understanding and accurate representation of these processes is needed for effective modeling. Given the difficulties associated with making reliable experimental measurements of these systems, computation can play an important role in developing these energy technologies. Accurate calculations have their own challenges since even within the simplest dynamical approximations such as transition state theory, the rates depend exponentially on critical barrier heights and these may be sensitive to the level of quantum chemistry. Moreover, it is well-known that in many cases it is necessary to go beyond statistical theories and consider the dynamics. Quantum tunneling, resonances, radiative transitions, and non-adiabatic effects governed by spin-orbit or derivative coupling can be determining factors in those dynamics. Building upon progress made during a period of prior support through the DOE Early Career Program, this project combines developments in the areas of potential energy surface (PES) fitting and multistate multireference quantum chemistry to allow spectroscopically and dynamically/kinetically accurate investigations of key molecular systems (such as those mentioned above), many of which are radicals with strong multireference character and have the possibility of multiple electronic states contributing to the observed dynamics. An ongoing area of investigation is to develop general strategies for robustly convergent electronic structure theory for global multichannel reactive surfaces including diabatization of energy and other relevant surfaces such as dipole transition. Combining advances in ab initio methods with automated interpolative PES fitting allows the construction of high-quality PESs (incorporating thousands of high-level data) to be done rapidly through parallel processing on high-performance computing (HPC) clusters. In addition, new methods and approaches to electronic structure theory will be developed and tested through applications. This project will explore limitations in traditional multireference calculations (e.g., MRCI) such as those imposed by internal contraction, lack of high-order correlation treatment and poor scaling. Methods such as DMRG-based extended active-space CASSCF and various Quantum Monte Carlo (QMC) methods will be applied (including VMC/DMC and FCIQMC). Insight into the relative significance of different orbital spaces and the robustness of application of these approaches on leadership class computing architectures will be gained. Synergy with other components of this research program such as automated PES fitting and multireference quantum chemistry will be used to address challenges encountered by the standard approaches to computational thermochemistry (those being single-reference quantum chemistry and perturbative treatments of the anharmonic vibrational energy, which break down for some cases of electronic structure or floppy strongly coupled vibrational modes).

74 ATOMIC AND MOLECULAR PHYSICS

Scalable Generation of High-fidelity Synthetic Population Ensembles

Used within social simulations, synthetic population ensembles enable uncertainty quantification (UQ) methods for obtaining more robust model inference and prediction. A synthetic population ensemble is a series of plausible virtual reconstructions of an area’s population at the granularity of people and residences, generated stochastically to preserve privacy of the source population survey’s respondents. In this paper, we demonstrate the production of large synthetic population ensembles for the U.S. via Oak Ridge National Laboratory’s UrbanPop framework to support modeling of high spatial resolution energy affordability metrics from nationwide social surveys in collaboration with the fusionACS project. The study involves two scenarios: creating ensembles for (1) 17 U.S. metropolitan areas in 2019 and (2) full U.S. Census Divisions in 2023, with each scenario consisting of 41 population instances (a base realization and 40 replicates). To accomplish this task at scale, we configured an integrated system within a research cloud, comprised of virtual containerizations, GPU-enhanced functionality, and orchestrated deployments of UrbanPop’s maturing Likeness Python ecosystem. Results demonstrate we maintained high-fidelity approximations of residential totals by areas of interest and the demographic characteristics of neighborhoods while reducing manual workflow burdens. Finally, we discuss plans to fine-tune and further develop our automated workflows for truly distributed job orchestration to increase computational efficiency, as well as provide an outlook for broadening applications of the ensembles.

Cluster computing

A bi-level spatiotemporal clustering approach and its application to drought extraction

We present a novel flexible bi-level spatiotemporal clustering algorithm to extract events based on their intensity and spatiotemporal structures. Our algorithm consists of using (i) a novel space-time k-means clustering to obtain spatiotemporally coherent intensity clusters, and (ii) a density-based spatial clustering of applications with noise (DBSCAN) to spatiotemporally section the intensity clusters into individual events. We discuss the development of the algorithm, the selection, tuning and meaning of the parameters within each step, as well as its validation. Finally, we apply the algorithm to a spatiotemporal drought index, standardized vapor pressure deficit drought index (SVDI), over the continental United States (US) from 1980–2021 and show that it captures historical drought events over the continental United States and their spatiotemporal extents.

17 WIND ENERGY

Microscopic origin of temperature-dependent magnetism in spin-orbit-coupled transition metal compounds

A few 4 d and 5 d transition metal compounds with various electron fillings were recently found to exhibit magnetic susceptibilities χ and magnetic moments that deviate from the well-established Kotani model. This model has been considered for decades to be the canonical expression for descriing the temperature dependence of magnetism in systems with nonnegligible spin-orbit coupling effects. In this paper, we uncover the origin of such discrepancies and determine the applicability and limitations of the Kotani model by calculating the temperature dependence of the magnetic moments of a series of 4 d (Ru-based) and 5 d (W-based) systems at different electron fillings. For this purpose, we perform exact diagonalization of -derived relativistic multiorbital Hubbard models on finite clusters and compute their magnetic susceptibilities. Comparison with experimentally measured magnetic properties indicates that contributions such as a temperature-independent χ 0 background, crystal field effects, Coulomb and Hund's couplings, and intersite interactions—not included in the Kotani model—are especially crucial for correctly describing the temperature dependence of χ and magnetic moments at various electron fillings in these systems. Based on our results, we propose a generalized approach beyond the Kotani model to accurately describe their magnetism. Published by the American Physical Society 2025

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND

A new “gold standard”: Perturbative triples corrections in unitary coupled cluster theory and prospects for quantum computing

A major difficulty in quantum simulation is the adequate treatment of a large collection of entangled particles, synonymous with electron correlation in electronic structure theory, with coupled cluster (CC) theory being the leading framework for dealing with this problem. Augmenting computationally affordable low-rank approximations in CC theory with a perturbative account of higher-rank excitations is a tractable and effective way of accounting for the missing electron correlation in those approximations. This is perhaps best exemplified by the “gold standard” CCSD(T) method, which bolsters the baseline CCSD with the effects of triple excitations using considerations from many-body perturbation theory (MBPT). Despite this established success, such a synergy between MBPT and the unitary analog of CC theory (UCC) has not been explored. In this work, we propose a similar approach wherein converged UCCSD amplitudes are leveraged to evaluate energy corrections associated with triple excitations, leading to the UCCSD[T] method. In terms of quantum computing, this correction represents an entirely classical post-processing step that improves the energy estimate by accounting for triple excitation effects without necessitating new quantum algorithm developments or increasing demand for quantum resources. The rationale behind this choice is shown to be rigorous by studying the properties of finite-order UCC energy functionals, and our efforts do not support the addition of the fifth-order contributions as in the (T) correction. We assess the performance of these approaches on a collection of small molecules and demonstrate the benefits of harnessing the inherent synergy between MBPT and UCC theories.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH

PANDORA: A Parallel Dendrogram Construction Algorithm for Single Linkage Clustering on GPU

This paper introduces Pandora, a parallel algorithm for computing dendrograms, the hierarchical cluster trees for single linkage clustering (SLC). Current parallel approaches construct dendrograms by partitioning a minimum spanning tree and removing edges. However, they struggle with skewed, hard-to-parallelize real-world dendrograms. Consequently, computing dendrograms is the sequential bottleneck in HDBSCAN*[21], a popular SLC variant. Pandora uses recursive tree contraction to address this limitation. Pandora contracts nodes to construct progressively smaller trees. It computes the smallest contracted dendrogram and expands it by inserting contracted edges. This recursive strategy is highly parallel, skew-independent, work-optimal, and well-suited for GPUs and multicores. We develop a performance portable implementation of Pandora in Kokkos[31] and evaluate its performance on multicore CPUs and multi-vendor GPUs (e.g., Nvidia, AMD) for dendrogram construction in HDBSCAN*. Multithreaded Pandora is 2.2x faster than the current best-multithreaded implementation. Our GPU version achieves 6-20x speedup on AMD GPUs and 10-37x on NVIDIA GPUs over multithreaded Pandora. Pandora removes HDBSCAN*’s sequential bottleneck, greatly boosting efficiency, particularly with GPUs.

Sao, Piyush