Engineering PapersSearch

SEARCH · Engineering Papers

Results for “parallel cluster”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Parallel Programming in MCNP6

Monte Carlo N-Particle (MCNP)1 is a general-purpose Monte Carlo particle transport code developed by Los Alamos National Laboratory (LANL). To efficiently handle long simulations, MCNP version 6 (MCNP6) supports parallel execution using two primary programming models: • Shared-memory task-based threading using OpenMP (Open Multi-Processing), and • Distributed-memory calculations using MPI (Message Passing Interface). The OpenMP and MPI programming models enable MCNP6 to scale from desktop systems to high-performance computing (HPC) clusters, allowing users to run MCNP in one of three parallel modes: • OpenMP-only, • MPI-only, and • Hybrid (MPI + OpenMP). The choice of parallelization mode depends on the underlying computer architecture and the characteristics of the simulation problem.

97 MATHEMATICS AND COMPUTING

Symbol alphabets in QCD and flag cluster algebras

The full 245-letter symbol alphabet for all planar massless two-loop six-point Feynman integrals was recently determined in arXiv:2412.19884 and arXiv:2501.01847. In a parallel mathematical development, it was shown in arXiv:2408.14956 that there is an embedding of the cluster algebra associated to the partial flag variety $\mathcal{Fl}$ $2,n-2;n$ , which describes the kinematics of n massless particles, into that of the Grassmannian Gr(n–2, 2n–4). In this paper we connect these developments by showing that most of the rational symbol letters can be expressed in terms of flag cluster variables, and that all of the algebraic symbol letters arise from infinite mutation sequences.

97 MATHEMATICS AND COMPUTING

A generalized and adaptable tensor-contraction-based cluster expansion formalism for multicomponent solids

Density functional theory (DFT)-based simulations of materials have first-principles accuracy, but are very computationally expensive. For simulating various properties of multi-component alloys, the cluster expansion (CE) technique has served as the standard workaround to improve computational efficiency. However, the standard CE technique is difficult to extend to exotic and/or low-symmetry lattices, often implemented via iteration over particular cluster types, which must be enumerated per lattice structure. In this work, we introduce the tensor cluster expansion (TCE), implemented in the open-source code tce-lib, which maps correlation functions to mixed tensor contractions, eliminating the need to iterate over cluster types and additionally making the calculation of correlation functions well-suited for massively parallel architectures like GPUs. We show that local interaction energies are an immediate consequence of the TCE formalism, yielding nearly $\mathcal{O}$(1) energy difference calculations. We then use this formalism to fit CE models for the TaW and CoNiCrFeMn systems, and use these models to respectively compute the enthalpy of mixing curve and Cowley short-range order parameters, showing excellent agreement with ground truth data.

Cluster expansion

The high level trigger and express data production at STAR

To meet the demands of the Beam Energy Scan phase-II (BES-II) program, the STAR experiment at the Relativistic Heavy Ion Collider (RHIC) developed a dual real-time framework consisting of a High Level Trigger (HLT) and an Express Data Production system (xProduction). The HLT operates online within the Data Acquisition (DAQ) chain on a dedicated multi-core CPU cluster with the option to offload compute-intensive kernels to Xeon Phi coprocessors. It uses parallelized algorithms, such as the Cellular Automaton (CA) Track Finder, to perform rapid tracking, vertexing, and event filtering. This allows it to select events of interest in real time and provide immediate feedback on detector and beam conditions. In contrast, the xProduction workflow runs concurrently and independently of the DAQ loop. It applies near offline-quality calibration and reconstruction within hours of data collection. The xProduction input is the express data stream, whose content can be enriched by HLT trigger/priority selections under DAQ/HLT resource constraints, and it uses the STAR calibration/conditions framework, incorporating online calibration/QA information when available. This enables early preliminary physics analysis, including the reconstruction of rare signals, such as hyperons and hypernuclei. It also provides collaboration-wide access to analysis-ready datasets. Together, the HLT and xProduction systems form a complementary architecture: the HLT performs online event selection while the xProduction chain delivers high-quality results within a short amount of time. This integrated framework has enabled the prompt reconstruction of the $^5_Λ$ He hypernucleus with high statistical significance and the efficient processing of hundreds of millions of heavy-ion collision events. In conclusion, its demonstrated scalability and robustness establish a model for future high-luminosity experiments requiring both online event filtering and rapid access to analysis-quality data.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS

A Faster-Than-Real-Time Framework for Reliability-Oriented Simulation of PV Inverters

Physics-of-Failure (PoF) based reliability assessment for photovoltaic (PV) inverters requires long-duration electrical and electrothermal stress histories, yet generating such stress histories with high-fidelity switching models over year long mission profiles is computationally prohibitive. Conventional methods either sacrifice modeling fidelity for speed or require runtimes that are impractical for design iteration and uncertainty studies. To address this bottleneck, this paper presents a High-Performance Computing (HPC) based simulation frame work for faster-than-real-time reliability-oriented simulation. The proposed framework integrates the Average-to-Switching (A2S) method with parallel computing techniques to accelerate switching-level waveform reconstruction. We further introduce optimization strategies, including cluster merging and sensitivity based mission profile screening, to reduce the computational burden. Evaluated using real-world mission profile inputs and a MATLAB/Simulink switching-model reference, the framework reduces the simulation time for a one-year mission from an intractable multi-year duration to approximately 7.3 minutes while maintaining low waveform error. This acceleration provides a practical reliability-oriented simulation engine that can be coupled with component-specific aging models for subsequent PV inverter PoF assessment.

High-performance Computing

Nodeman: A Node Management Tool For Hpc Clusters

NodeMan is a command line tool to manage nodes in an HPC cluster. At it's core, it is an extensible framework composed of bash scripting and GNU parallel. HPC System Administrator will find it useful in that it encapsulates desired functions and allows them to be assembled in a way familiar to administrators - through pipes. In fact, NodeMan functions can work with common command line tools as long as they use stdin/stdout. System Administrators can construct moderately complex logic and filtering on a compact command line that would normally require a substantial shell script. In the spirit of clush and pdsh, it is able to run commands remotely on nodes. Additionally, NodeMan is more flexible. For example, it can interact with IPMI and naturally processes node lists for orchestrating different tools. The library of useful pre-built functions is growing. System administrators can easily create new functions and make it their own.

Serr, ScottM

Sensor Co-design for $\textit{smartpixels}$

Pixel tracking detectors at upcoming collider experiments will see unprecedented charged-particle densities. Real-time data reduction on the detector will enable higher granularity and faster readout, possibly enabling the use of the pixel detector in the first level of the trigger for a hadron collider. This data reduction can be accomplished with a neural network (NN) in the readout chip bonded with the sensor that recognizes and rejects tracks with low transverse momentum (p$_T$) based on the geometrical shape of the charge deposition (``cluster''). To design a viable detector for deployment at an experiment, the dependence of the NN as a function of the sensor geometry, external magnetic field, and irradiation must be understood. In this paper, we present first studies of the efficiency and data reduction for planar pixel sensors exploring these parameters. A smaller sensor pitch in the bending direction improves the p$_T$ discrimination, but a larger pitch can be partially compensated with detector depth. An external magnetic field parallel to the sensor plane induces Lorentz drift of the electron-hole pairs produced by the charged particle, broadening the cluster and improving the network performance. The absence of the external field diminishes the background rejection compared to the baseline by $\mathcal{O}$(10%). Any accumulated radiation damage also changes the cluster shape, reducing the signal efficiency compared to the baseline by $\sim$ 30 - 60%, but nearly all of the performance can be recovered through retraining of the network and updating the weights. Finally, the impact of noise was investigated, and retraining the network on noise-injected datasets was found to maintain performance within 6% of the baseline network trained and evaluated on noiseless data.

Shekar, Danush [Illinois U., Chicago]

ArborX 2.0

ArborX library tackles a problem of efficiently finding geometric objects that are close in space. Variations of this problem, such as finding the nearest neighbors of a point, or finding all objects within a certain distance, are inherent components of applications in many fields. The data may be large so that solving the problem efficiently may require significant computational resources, such as multiple processors or accelerators such as general purpose GPUs. ArborX' main advantage in its ability to solve large problems efficiently utilizing a combination of distributed and on-node parallelism. ArborX can be run efficiently on a wide variety of hardware, including GPUs from different vendors, which distinguishes it from other available libraries which typically choose only few of these. The other advantage is that it supports both types of user problems: spatial problems (useful for intersections and finding objects within certain distance), and nearest neighbor problems. ArborX also supports flexible interface in its interaction with a user. Particularly, it allows a user to call user's own function on a positive match, a functionality not rarely available in other libraries. ArborX implements construction and traversal algorithms using efficient tree structures, such as bounding volume hierarchy (BVH). At its core, ArborX uses linear BVH for its low construction cost and sufficient quality. ArborX implements both spatial and nearest-neighbor traversal algorithms. ArborX also provides several clustering algorithms (minimum spanning tree, DBSCAN, HDBSCAN*), interpolation using minimum least squares and ray tracing. ArborX is written using C++, and is parallelized using the message passing interface (MPI) for the distributed communication, and the Kokkos library for on-node parallelism. This approach allows ArborX to be run on a wide variety of hardware, from common laptops and desktops to supercomputers while using the same codebase.

Prokopenko, Andrey [Oak Ridge National Laboratory

Unraveling the Determinant Mechanisms in Flow-Mediated Crystal Growth and Phase Behaviors

To uncover the critical mechanisms responsible for mesoscopic level development during flow-mediated crystal growth, we develop a semi-two-way hydrodynamic coupled structural phase-field crystal formalism (HXPFC-s2). The new formalism, inspired by previous attempts at coupling hydrodynamic and phase-field crystal (PFC) equations, allows for studying mesoscopic flow-mediated crystallization at diffusive timescales pertinent to industrial applications. Unlike previous efforts, the devised coupling to the structural PFC (XPFC) equations allows generalization to more complex crystal structures through explicit parameterization of the direct correlation function (DCF). Utilizing the HXPFC-s2 formalism, we seek to uncover the determinant physical mechanisms in crystallization under simple shear flows by comparing temperature-driven crystallization to flow-mediated crystallization under varying flow-strengths. Parallels and deviations of under-cooling and flow-strength effects on crystal growth are drawn using the crystal cluster-size and system ordering time evolutions. In doing so, we identify scaling behaviors with a Peclet-like number, Pe∼, a critical Peclet-like number, Pe∼*, and flow-field-crystal plane-dependent interactions. Our findings may be relevant for controlling crystal growth and phase behaviors in flow applications.

Willis, L. Connor (ORCID:0009000961321848)

Characterization of DESI fiber assignment incompleteness effect on 2-point clustering and mitigation methods for DR1 analysis

We present an in-depth analysis of the fiber assignment incompleteness in the Dark Energy Spectroscopic Instrument (DESI) Data Release 1 (DR1). This incompleteness is caused by the restricted mobility of the robotic fiber positioner in the DESI focal plane, which limits the number of galaxies that can be observed at the same time, especially at small angular separations. As a result, the observed clustering amplitude is suppressed in a scale-dependent manner, which, if not addressed, can severely impact the inference of cosmological parameters. We discuss the methods adopted for simulating fiber assignment on mocks and data. In particular, we introduce the fast fiber assignment (FFA) emulator, which was employed to obtain the power spectrum covariance adopted for the DR1 full-shape analysis. We present the mitigation techniques, organised in two classes: measurement stage and model stage. We then use high fidelity mocks as a reference to quantify both the accuracy of the FFA emulator and the effectiveness of the different measurement-stage mitigation techniques. This complements the studies conducted in a parallel paper for the model-stage techniques, namely the θ-cut approach. We find that pairwise inverse probability (PIP) weights with angular upweighting recover the “true” clustering in all the cases considered, in both Fourier and configuration space. Notably, we present the first ever power spectrum measurement with PIP weights from real data.

cosmological simulations

HARD: A performance portable radiation hydrodynamics code based on FleCSI framework

Hydrodynamics And Radiation Diffusion (HARD) is an open-source application for high-performance simulations of compressible hydrodynamics with radiation-diffusion coupling. Built on the FleCSI (Bergen et al., 2021 [1]) (Flexible Computational Science Infrastructure) framework, HARD expresses its computational units as tasks whose execution can be orchestrated by multiple back-end runtimes, including Legion (Bauer et al., 2012 [2]), MPI (Forum, 1994 [3]), and HPX (Kaiser et al., 2020 [4]). Node-level parallelism is handled through Kokkos (Edwards et al., 2014 [5]), providing a single-source, portable code base that runs efficiently on laptops, small homogeneous clusters, and the largest heterogeneous supercomputers currently available. To ensure scientific reliability, HARD includes a regression test suite that automatically reproduces canonical verification problems such as the Sod and LeBlanc shock tubes, and the Sedov blast wave, comparing numerical solutions against known analytical results. The project is distributed under an OSI-approved license, hosted on GitHub, and accompanied by reproducible build scripts and continuous integration workflows. This combination of performance portability, verification infrastructure, and community-focused development makes HARD a sustainable platform for advancing radiation hydrodynamics research across multiple domains.

97 MATHEMATICS AND COMPUTING

Thermodynamics and collisionality in firehose-susceptible high- β plasmas

We study the evolution of collisionless plasmas that, due to their macroscopic evolution, are susceptible to the firehose instability, using both analytic theory and hybrid-kinetic particle-in-cell simulations. We establish that, depending on the relative magnitude of the plasma β, the characteristic time scale of macroscopic evolution and the ion-Larmor frequency, the saturation of the firehose instability in high-β plasmas can result in three qualitatively distinct thermodynamic (and electromagnetic) states. By contrast with the previously identified ‘ultra-high-beta’ and ‘Alfvén-inhibiting’ states, the newly identified ‘Alfvén-enabling’ state, which is realised when the macroscopic evolution time τ exceeds the ion-Larmor frequency by a β-dependent critical parameter, can support linear Alfvén waves and Alfvénic turbulence because the magnetic tension associated with the plasma’s macroscopic magnetic field is never completely negated by anisotropic pressure forces. We characterise these states in detail, including their saturated magnetic-energy spectra. The effective collision operator associated with the firehose fluctuations is also described; we find it to be well approximated in the Alfvén-enabling state by a simple quasi-linear pitch-angle scattering operator. The box-averaged collision frequency is ν eff ∼ β/τ, in agreement with previous results, but certain subpopulations of particles scatter at a much larger (or smaller) rate depending on their velocity in the direction parallel to the magnetic field. Our findings are essential for understanding low-collisionality astrophysical plasmas including the solar wind, the intracluster medium of galaxy clusters and black hole accretion flows. We show that all three of these plasmas are in the Alfvén-enabling regime of firehose saturation and discuss the implications of this result.

astrophysical plasmas

Function, Structure, and Regulation of Nitrogen Fixation-like Metalloproteins for Nitrogen, Energy, Carbon, and Sulfur Metabolism

Nitrogenases (N 2 ases) and nitrogen fixation-like (NFL) systems play distinct roles in nitrogen, carbon, sulfur, and energy metabolism based on their fundamental differences in structure and metallocofactor identity. As new NFL systems have recently been identified and characterized, striking parallels and differences compared to N 2 ase structure, catalysis, and regulation have emerged. NFL systems use metallocofactors that span from simple [4Fe-4S] clusters to complex clusters akin to FeMo-co, previously only thought to occur in N 2 ase. This review describes the present state of knowledge on the function, structure, catalytic mechanisms, and regulation of NFL systems that perform distinct biological roles across all three domains of life. Recent advancements in N 2 ase spectroscopic techniques for probing metallocofactor structure and electronic states guide current and future work on how each NFL system catalyzes its specific biological reaction(s). Key knowledge gaps and needed areas of research for uncovering the specific metallocofactors and structural motifs that are at the heart of NFL system reaction specificity, along with how these systems are regulated, are discussed.

Bacteria

Streaming Matching and Edge Cover in Practice

Graph algorithms with polynomial space and time requirements often become infeasible for massive graphs with billions of edges or more. State-of-the-art approaches therefore employ approximate serial, parallel, and distributed algorithms to tackle these challenges. However, such approaches require storing the entire graph in memory and thus need access to costly computing resources such as clusters and supercomputers. In this paper, we present practical streaming approaches for solving massive graph problems using limited memory for two prototypical graph problems: maximum weighted matching and minimum weighted edge cover. For matching, we conduct a thorough computational study on two of the semi-streaming algorithms including a recent breakthrough result that achieves a $1/(2+\varepsilon)$-approximation of the weight while using $O( n \log W /\epsilon)$ memory (here $n$ is the number of vertices and $W$ is the maximum edge weight), designed by Paz and Schwartzman [SODA, 2017]. Empirically, we show that the semi-streaming algorithms produce matchings whose weight is close to the best $1/2$-approximate offline algorithm while requiring less time and an order-of-magnitude less memory. For minimum weighted edge cover, we develop three novel semi-streaming algorithms. Two of these algorithms require a single pass through the input graph, require $O(n \log n)$ memory, and provide a 2-approximation guarantee on the objective. We also leverage a relationship between approximate maximum weighted matching and approximate minimum weighted edge cover to develop a two-pass $3/2+\epsilon$-approximate algorithm with the memory requirement of Paz and Schwartzman's semi-streaming matching algorithm. These streaming approaches are compared against the state-of-the-art 3/2-approximate offline algorithm. The semi-streaming matching and the novel edge cover algorithms proposed in this paper can process graphs with several billions of edges in under 30 minutes using 6 GB of memory, which is at least an order of magnitude improvement from the offline (non-streaming) algorithms. For the largest graph, the best alternative offline parallel approximation algorithm (GPA+ROMA) could not finish in three hours even while employing hundreds of processors and 1 TB of memory. We also demonstrate an application of the semi-streaming algorithm by computing a matching using linearly bounded memory on item intersection graphs derived from three machine learning datasets, whereas the existing offline algorithms could not complete on one of these datasets since their memory requirements exceeded 1TB.

Ferdous, S M.

A Versatile Simulated Data Transport Layer for in Situ Workflows Performance Evaluation

In situ processing does not only allow scientific applications to face the explosion in data volume and velocity but also to address the time constraints of many simulation-analysis workflows by providing scientists with early insights about their applications at runtime. Multiple frameworks implement the concept of a data transport layer (DTL) to enable such in situ workflows. These tools are very versatile, directly or indirectly access the data generated on the same node, another node of the same compute cluster, or a completely distinct node, and allow data publishers and subscribers to run on the same computing resources or not. This versatility puts on researchers the onus of taking key decisions related to resource allocation and how to transport data to ensure the most efficient execution of their in situ workflows. However, domain scientists and workflow practitioners lack the appropriate tools to assess the respective performance of particular design and deployment options. In this paper we introduce a versatile simulated DTL designed to provide researchers with insights on the respective performance of different execution scenarios of in situ workflows. This open-source, standalone library builds on the SimGrid toolkit and can be linked to any SimGrid-based simulator. It facilitates the evaluation of the performance behavior, at scale, of different data transport configurations and the study of the effects of resource allocation strategies. We demonstrate the scalability, versatility, and accuracy of this simulated DTL by reproducing the execution of two synthetic benchmarks and of a real-world in situ workflow composed of an MPI application and a parallel data analysis. Results of simulations run on a single core show that the proposed library can simulate the interactions of tens of thousands of simulated processes deployed on two interconnected commodity clusters in a few seconds, and the execution by a thousand simulated processes of an in situ workflow in less than three minutes.

Suter, Fred [ORNL] (ORCID:0000000319021955)

Impact of Iron Species Dispersion on Fe/ZSM–5 Catalyst Performance for Methane Dehydroaromatization (MDA)

Methane dehydroaromatization (MDA) is one of the most promising technologies for directly transforming methane into aromatics. Unlike the extensively investigated Mo/ZSM-5 catalysts, the structure and, consequently, the catalytic activity of Fe/ZSM-5 are markedly influenced by the method of preparation, as shown here. In this study, we prepared 2 % and 4 % Fe/ZSM-5 catalysts via wet impregnation (WI) and incipient wetness impregnation (IWI). Characterizations (XRD, STEM, UV-Vis, NH 3 -TPD and H 2 -TPR) reveal that 2 %Fe-WI mainly possesses isolated or low-polymerized Fe species within zeolite channels, leading to a rapid activation and a higher benzene yield due to the faster reduction to iron suboxides under MDA conditions. In contrast, 2 %Fe-IWI contains bulk iron oxide aggregates, resulting in a slower activation as these aggregates transform into iron carbide through successive reduction and carbonization. Here, a deactivation kinetic study applied to the 2 % catalysts further demonstrates the quantitative relation between Fe site isolation and catalytic activity. Although both 4 % catalysts inevitably form sizable iron oxide clusters and particles due to the high Fe/Al ratio, similar trends are noted, with the WI catalysts exhibiting a shorter induction/activation period and a higher yield of benzene, paralleling observations made with 2 % catalysts.

03 NATURAL GAS

Quantum Electrodynamics Coupled-Cluster at Scale: High-Performance Implementation for Complex Systems

Coupled-cluster theory (CC) is a highly accurate and versatile method for simulating complex interactions within quantum systems. The extension of CC theory to model mixed electron-photon processes with quantum electrodynamics (QED) has improved our capability to predict cavity-modified chemistry, a field where photons are used as cost-effective and eco-friendly alternatives to catalyze/inhibit chemical reactions. However, calculations with CC methods, even without incorporating QED effects, are often prohibitively expensive. Simulations of larger systems require scalable infrastructures that exist for traditional CC methods but not for QED-CC methods. As such, we present a GPU-enabled, high-performance, open-source implementation of the quantum electrodynamics coupled-cluster method with single and double excitations (QED-CCSD) within the ExaChem quantum chemistry software package. ExaChem relies on the Tensor Algebra for Many-body Methods (TAMM) infrastructure: a parallel heterogeneous tensor library designed to achieve scalable performance on modern heterogeneous supercomputing platforms. Furthermore, we discuss theoretical foundations, algorithmic details, and numerical benchmarks to showcase the larger systems that ExaChem can simulate and how the integration of photonic degrees-of-freedom alters their ground-state properties.

Basis sets

Exploring Anomalous Photoelectron Angular Distributions in the Photoelectron Spectra of Gd 3 O 3 – : Study of Gd 3 O 2 – and Gd 3 O 3 – Using Photoelectron Spectroscopy and Density Functional Theory Calculations

Anion photoelectron (PE) spectra of lanthanide oxide clusters obtained previously have exhibited anomalous photoelectron angular distributions which were attributed to strong PE–valence electron (PEVE) interactions. Here, to further explore this effect, we have obtained the PE spectra of Gd 3 O 2 – and Gd 3 O 3 – , two clusters that have similarly complex electronic structures but contrasting symmetries. The spectra exhibit manifolds of detachment transitions at similar binding energies in a 0.5 eV window of energy. The electron affinity of Gd 3 O 2 is measured to be 1.29 ± 0.05 eV, and that of Gd 3 O 3 is 1.31 ± 0.05 eV. As seen in previous studies on lanthanide oxide cluster anions in lower than conventional oxidation states, transitions in spectra obtained lower photon energies are more congested than those obtained with higher photon energy, a signature of strong PEVE interactions. While the detachment transitions have predominantly parallel photoelectron angular distributions (PAD), the PAD varies across the manifold of transitions in the PE spectrum of Gd 3 O 3 – in a way that suggests four different subgroups of transitions. Results of calculations on Gd 3 O 2 – suggest kite or V-shape structures with antiferromagnetic coupling between one of the 4f 7 subshells with the two others. Calculations on Gd 3 O 3 – more definitively point to ring structures with a nearly isoenergetic ferromagnetically coupled high spin (24-tet) state and a dectet state in which one of the 4f 7 subshells is antiferromagnetically coupled with the other two. Taking these results as qualitative, we propose that strong mixing between the unperturbed states predicted computationally leads to overlapping transitions with different PADs.

anions