Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “weak scaling”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Massively parallel axisymmetric fluid model for streamer discharges

A highly parallelizable fluid plasma simulation tool based upon the first-order drift-diffusion equations is discussed. Atmospheric pressure plasmas have densities and gradients that require small element sizes in order to accurately simulate the plasm resulting in computational meshes on the order of millions to tens of millions of elements for realistic size plasma reactors. To enable simulations of this nature, parallel computing is required and must be optimized for the particular problem. Here, a finite-volume, electrostatic drift-diffusion implementation for low-temperature plasma is discussed. The implementation is built upon the Message Passing Interface (MPI) library in C++ using Object Oriented Programming. The underlying numerical method is outlined in detail and benchmarked against simple streamer formation from other streamer codes. Electron densities, electric field, and propagation speeds are compared with the reference case and show good agreement. Convergence studies are also performed showing a minimal space step of approximately 4 μm required to reduce relative error to below 1% during early streamer simulation times and even finer space steps are required for longer times. Additionally, strong and weak scaling of the implementation are studied and demonstrate the excellent performance behavior of the implementation up to 100 million elements on 1024 processors. Lastly, different advection schemes are compared for the simple streamer problem to analyze the influence of numerical diffusion on the resulting quantities of interest.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

A family of independent Variable Eddington Factor methods with efficient preconditioned iterative solvers

We present a family of discretizations for the Variable Eddington Factor (VEF) equations that have high-order accuracy on curved meshes and efficient preconditioned iterative solvers. The VEF discretizations are combined with the Discontinuous Galerkin transport discretization from to form effective high-order, linear transport methods. The VEF discretizations are derived by extending the unified analysis of Discontinuous Galerkin methods for elliptic problems presented by Arnold et al. to the VEF equations. This framework is used to define analogs of the interior penalty, second method of Bassi and Rebay, minimal dissipation local Discontinuous Galerkin, and continuous finite element methods. The analysis of subspace correction preconditioners, which use a continuous operator to iteratively precondition the discontinuous discretization, is extended to the case of the non-symmetric VEF system. Numerical results demonstrate that the VEF discretizations have arbitrary-order accuracy on curved meshes, preserve the thick diffusion limit, and are effective on a proxy problem from thermal radiative transfer in both outer transport iterations and inner preconditioned linear solver iterations. We demonstrate that the VEF solution converges to the S N transport solution as the mesh is refined on both problems with smooth and non-smooth behavior in angle. Parallel performance studies show that the interior penalty VEF discretization's linear solve weak scales out to 1024 processors and strong scales well on a single node. Particular attention is paid to the parallel performance of the VEF algorithm when used in combination with a parallel block Jacobi transport sweep.

97 MATHEMATICS AND COMPUTING↗

A second-order distributed memory parallel fast sweeping method for the Eikonal equation

The Eikonal equation is used to calculate wave propagation and distance fields, and due to its complexity requires numerical treatment for its solution. In this work, we present a second-order distributed memory parallel fast sweeping method. The second-order solution switches on a two-point stencil when two upwind points are available, and reverts to first-order otherwise. In all examples, the second-order method improves the solution over the first-order, allowing for significant savings in memory while achieving the same accuracy. Parallelization over distributed memory saw good weak scaling with optimal convergence. The computational time for second-order was approximately 2.5 times slower than first-order, where the largest amount of mesh points ran on 144 cores (512 GB) was ≈20 billion. The savings in memory from the second-order method combined with the distributed memory algorithm result in the ability to solve problems much larger than are possible with the serial first-order method.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Achieving performance portability in Gaussian basis set density functional theory on accelerator based architectures in NWChemEx

The numerical integration of the exchange–correlation (XC) potential is one of the primary computational bottlenecks in Gaussian basis set Kohn–Sham density functional theory (KS-DFT). To achieve optimal performance and accuracy, care must be taken in this numerical integration to preserve local sparsity as to allow for near linear weak scaling with system size. This leads to an integration scheme with several performance critical kernels which must be hand optimized for each architecture of interest. As the set of available accelerator hardware goes more diverse, a key challenge for developers of KS-DFT software is to maintain performance portability across a wide range of computational architectures. In this article, we examine a modular software design pattern which decouples the implementation details of performance critical kernels from the expression of high-level algorithmic workflows in a device-agnostic language such as C++; thus allowing for developers to target existing and emerging accelerator hardware within a single code base. We consider the efficacy of such a design pattern in the numerical integration of the XC potential by demonstrating its ability to achieve performance portability across a set of accelerator architectures which are representative of those on current and future U.S. Department of Energy Leadership Computing Facilities.

97 MATHEMATICS AND COMPUTING↗

Identifying geological structures through microseismic cluster and burst analyses complementing active seismic interpretation

At the Decatur carbon capture and storage site (IL, USA) CO 2 has been injected from 2011–2014 and from 2017 to present near the base of the Lower Mt. Simon Sandstone saline reservoir, resulting in microseismicity. Microseismicity is mainly located in the basement and distributed in distinct spatial clusters. The lack of significant impedance contrasts within the basement makes the interpretation of active-source seismic reflection data challenging, however, recent reprocessing allowed to resolve faults above and at the top of the basement. These faults generally do not coincide with the location of microseismic events and their continuation to the general depth of the seismic events cannot be assumed. This paper shows how the interpretation of the microseismicity can complement structural interpretations of active-source seismic reflection data. In particular, we analyze clusters and bursts (abrupt increases) of microseismicity, identify unresolved, smaller-scale weaknesses and extract statistical parameters. These parameters allow comparisons with the interpreted faults, and with fracture sets intercepted by boreholes. During injection at the Decatur site, the injection pressure was kept far below fracture pressure, nevertheless, seismic events were induced and spread far beyond the expected extent of the CO 2 plume. We argue that local stress transfers related to the CO 2 injection reactivated pre-existing fractures within the critically stressed basement. Finally, we conducted a slip tendency analysis for faults interpreted from active seismic, selected cluster, bursts and nodal planes from focal mechanisms to determine if the interpreted structures are optimally oriented with respect to the stress regime. Our results suggest that the orientation of fractures close to the injection well, generally shows slight deviations from the optimal orientation for slip. This might indicate either slight local deviations of the maximum horizontal stress azimuth from the average direction used in the analysis, or the lack of optimally oriented fractures at this location.

58 GEOSCIENCES↗

The Muon Smasher’s Guide

In this work, we lay out a comprehensive physics case for a future high-energy muon collider, exploring a range of collision energies (from 1 to 100 TeV) and luminosities. We highlight the advantages of such a collider over proposed alternatives. We show how one can leverage both the point-like nature of the muons themselves as well as the cloud of electroweak radiation that surrounds the beam to blur the dichotomy between energy and precision in the search for new physics. The physics case is buttressed by a range of studies with applications to electroweak symmetry breaking, dark matter, and the naturalness of the weak scale. Additionally, we make sharp connections with complementary experiments that are probing new physics effects using electric dipole moments, flavor violation, and gravitational waves. An extensive appendix provides cross section predictions as a function of the center-of-mass energy for many canonical simplified models.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Disruption thermal load mitigation with shattered pellet injection on the Joint European Torus (JET)

Disruption mitigation remains a critical, unresolved challenge for ITER. To aid in addressing this challenge, a shattered pellet injection (SPI) system was installed on JET and experiments conducted at a range of thermal energy fractions and stored energies in excess of 7 MJ. The primary goals of these experiments were to investigate the efficacy of the SPI on JET and the ability of the plasma to assimilate multiple pellets. Single pellet injections produced a saturation in total radiated energy (Wrad) with increasing injected neon content, suggesting total radiation of stored thermal energy. Further increases in injected neon quantities resulted in reduced cooling times and current quench (CQ) durations, indicating higher impurity assimilation. No significant variation in CQ duration or W rad was observed when varying the deuterium content at fixed neon quantities. Additionally, higher assimilation, inferred by shorter CQ durations, was measured when a mechanical punch was used to launch the pellets and this was attributed to a lower pellet velocity leading to higher solid content in the pellet plume and larger fragments penetrating deeper into the plasma. Radiation asymmetries averaged over the cooling time were inferred from Emis3D and ranged from 1.6 to 1.9. Asymmetries averaged over the entire disruption sequence were found to increase at higher thermal energy fractions. The radiated energy fractions decreased with increasing thermal energy fractions but this trend was eliminated when toroidal asymmetries were accounted for with Emis3D. Pure deuterium pellets were able to produce cooling times of up to 75 ms with a gradual loss in thermal stored energy of up to 80%. Experiments with multiple pellet injection indicated W rad can be increased through pellet superposition and density can be increased with an additional D2 injection without a reduction in W rad . KPRAD modelling accurately reproduced the cooling times and the CQ duration at high thermal energies. Assimilation estimates from KPRAD indicated CQ rates scale strongly whilst W rad scales weakly and saturates with assimilated neon content. Comparable W rad can be achieved with lower assimilated neon quantities as longer cooling times are attained. Thus reduced neon content can be preferential in a thermal load mitigation scheme as it may reduce radiation asymmetries and prevent flash melting.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Stau pairs from natural SUSY at high luminosity LHC

Natural supersymmetry (SUSY) with light Higgsinos is perhaps the most plausible of all weak scale SUSY models while a variety of motivations point to (right) tau sleptons as the lightest of all the sleptons. We examine a SUSY model line with rather light right staus embedded within natural SUSY. For light τ ˜ 1 of a few hundred GeV, the decays τ ˜ 1 → τ χ ˜ 1 , 2 0 and ν τ χ ˜ 1 − occur at comparable rates where the (Higgsino-like) χ ˜ 1 ± and χ ˜ 2 0 release only small visible energy: in this case, the expected τ + τ − + E T signature is diminished from the usual expectations due to the presence of the nearly invisible decay mode τ ˜ 1 → ν τ χ ˜ 1 − . However, once m τ ˜ 1 ≳ m ( b i n o ) , decays to binos such as τ ˜ 1 → τ χ ˜ 3 0 open up where χ ˜ 3 0 decays to Higgsinos plus W ± , Z 0 , and h at comparable rates. For these heavier staus, the stau pair production gives rise to diboson + E T events, which may contain 0, 1, or 2 additional hard τ leptons. From these considerations, we examine the potential for future discovery of tau-slepton pair production at a high-luminosity LHC. While we do not find a 5 σ HL-LHC discovery reach for 3000 fb − 1 , we do find a 95% CL exclusion reach, ranging between m τ ˜ 1 : 100 – 450 GeV for m χ ˜ 1 0 ∼ 100 GeV . This latter reach disappears for m χ ˜ 1 0 ≳ 200 GeV . Published by the American Physical Society 2024

Astronomy & Astrophysics↗

A multiverse outside of the swampland

A multiverse can arise from landscapes without de Sitter minima. It can be populated during a period of eternal inflation without trans-Planckian field excursions and without flat potentials. This multiverse can explain the values of the cosmological constant and of the weak scale. In the process of proving these statements, we derive a few simple, but counterintuitive results. We show that it is easy to write models of eternal inflation compatible with the distance and refined de Sitter conjectures. Secondly, tunneling transitions that move fields from a lower-energy vacuum to a higher-energy vacuum and generate baby universes are possible, and occur during eternal inflation. Finally, we relax the assumption of no de Sitter minima and show that this more standard multiverse can be populated by Coleman-de Luccia transitions in about 100 e-folds of inflation.

79 ASTRONOMY AND ASTROPHYSICS↗

Precise interpretations of traditional fine-tuning measures

We uncover two precise interpretations of traditional electroweak fine-tuning (FT) measures that were historically missed. (i) a statistical interpretation : the traditional FT measure shows the change in plausibility of a model in which a parameter was exchanged for the 𝑍 boson mass relative to an untuned model in light of the 𝑍 boson mass measurement. (ii) an information-theoretic interpretation : the traditional FT measure shows the exponential of the extra information, measured in nats, relative to an untuned model that you must supply about a parameter in order to fit the 𝑍 mass. We derive the mathematical results underlying these interpretations, and explain them using examples from weak scale supersymmetry. These new interpretations allow us to rigorously define FT in particle physics and beyond, shed fresh light on the status of extensions to the Standard Model and, lastly, allow us to precisely reinterpret historical and recent studies using traditional FT measures.

electroweak symmetry breaking↗

Advancing material modeling in hydrocodes using a concurrent finite-element and molecular dynamics multiscale framework

We present a multiscale simulation framework that couples the finite-element method with molecular dynamics. Bypassing traditional equations of state (EOS) by using in-line atomistic simulations, the method offers the advantage of incorporating detailed microscale physics not easily represented with coarse-grained models. Coupling consistency with the continuum code is ensured through the use of lifting and restriction operators, in line with heterogeneous multiscale methods. The concurrent continuum-atomistic framework is validated through comparison with experimental results and conventional EOS models, and demonstrated in a shock-driven hydrodynamic flow simulation under extreme conditions. We further evaluate the framework's usability by comparing it to state-of-the-art EOS models of deuterium. A computational performance study reveals that the atomistic EOS evaluation is a feasible alternative to conventional approaches, and demonstrates a weak scaling of 99% efficiency. These results highlight the framework's potential for large-scale multiscale modeling across a broad range of materials and conditions.

Computer science↗

Higgs potential from instantons

We propose that the Higgs potential, a key element in our understanding of nature, is partially generated by the instantons of new confining dynamics, perhaps from a hidden sector. In this picture, while the Higgs itself is a fundamental field, it controls the strength of the nonperturbative interactions that give rise to its potential. This setup can be realized in “Twin Higgs” hierarchy models, where the twin strong dynamics confines at or above the weak scale, but large quantum corrections to the Higgs potential are avoided. We examine a simple setup in which this instanton contribution is augmented by a quartic term, which is sufficient for a realistic electroweak symmetry breaking mechanism. The minimum of the potential is given by the Lambert 𝑊 0 function, with these assumptions. We discuss the predictions of this model and how it may be tested through measurements of the Higgs self coupling. Given the connection with nontrivial dynamics, one may also consider the prospects of accessing hidden sector states at colliders; this seems to be typically challenging in our setup. Symmetry restoration in the early Universe in this scenario is briefly examined. We also comment on the possible connection of our general setup with recent work on the physics of field space end points.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

OpenACC offloading of the MFC compressible multiphase flow solver on AMD and NVIDIA GPUs

GPUs are the heart of the latest generations of supercomputers. We efficiently accelerate a compressible multiphase flow solver via OpenACC on NVIDIA and AMD Instinct GPUs. Optimization is accomplished by specifying the directive clauses gang vector and collapse. Further speedups of six and ten times are achieved by packing user-defined types into coalesced multidimensional arrays and manual inlining via metaprogramming. Additional optimizations yield seven-times speedup of array packing and thirty-times speedup of select kernels on Frontier. Weak scaling efficiencies of 97% and 95% are observed when scaling to 50% of Summit and 87% of Frontier. Strong scaling efficiencies of 84% and 81% are observed when increasing the device count by a factor of 8 and 16 on V100 and MI250X hardware. The strong scaling efficiency of AMD’s MI250X increases to 92% when increasing the device count by a factor of 16 when GPU-aware MPI is used for communication.

Wilfong, Benjamin↗

Recommended conventions for reporting results from direct dark matter searches

Abstract The field of dark matter detection is a highly visible and highly competitive one. In this paper, we propose recommendations for presenting dark matter direct detection results particularly suited for weak-scale dark matter searches, although we believe the spirit of the recommendations can apply more broadly to searches for other dark matter candidates, such as very light dark matter or axions. To translate experimental data into a final published result, direct detection collaborations must make a series of choices in their analysis, ranging from how to model astrophysical parameters to how to make statistical inferences based on observed data. While many collaborations follow a standard set of recommendations in some areas, for example the expected flux of dark matter particles (to a large degree based on a paper from Lewin and Smith in 1995), in other areas, particularly in statistical inference, they have taken different approaches, often from result to result by the same collaboration. We set out a number of recommendations on how to apply the now commonly used Profile Likelihood Ratio method to direct detection data. In addition, updated recommendations for the Standard Halo Model astrophysical parameters and relevant neutrino fluxes are provided. The authors of this note include members of the DAMIC, DarkSide, DARWIN, DEAP, LZ, NEWS-G, PandaX, PICO, SBC, SENSEI, SuperCDMS, and XENON collaborations, and these collaborations provided input to the recommendations laid out here. Wide-spread adoption of these recommendations will make it easier to compare and combine future dark matter results.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Feebly-interacting particles: FIPs 2022 Workshop Report

Particle physics today faces the challenge of explaining the mystery of dark matter, the origin of matter over anti-matter in the Universe, the origin of the neutrino masses, the apparent fine-tuning of the electro-weak scale, and many other aspects of fundamental physics. Perhaps the most striking frontier to emerge in the search for answers involves new physics at mass scales comparable to familiar matter, below the GeV-scale, or even radically below, down to sub-eV scales, and with very feeble interaction strength. New theoretical ideas to address dark matter and other fundamental questions predict such feebly interacting particles (FIPs) at these scales, and indeed, existing data provide numerous hints for such possibility. A vibrant experimental program to discover such physics is under way, guided by a systematic theoretical approach firmly grounded on the underlying principles of the Standard Model. This document represents the report of the FIPs 2022 workshop, held at CERN between the 17 and 21 October 2022 and aims to give an overview of these efforts, their motivations, and the decadal goals that animate the community involved in the search for FIPs.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

A Simple, Scalable Large Deformation Solid Mechanics Implementation in the MOOSE Framework

This article describes a large deformation solid mechanics solver implemented as part of the freely available and open source MOOSE finite element simulation framework. The article documents the choices made in developing the solid mechanics framework and describes novel formulations for the gradient operator and constitutive modeling framework made to simplify implementations of different coordinate systems, stabilized gradient operators, and different constitutive model inputs and outputs. In the process, the article describes a new formulation that casts objective integration of the Cauchy stress as a linear transformation of the small stress rate. Finally, the article presents key implementation details and examines the parallel efficiency of the solid mechanics solver implemented in MOOSE. The implementation retains a good weak scaling efficiency beyond 1,000 parallel processes. The article includes a discussion of the factors limiting the parallel efficiency of implicit, large deformation solid mechanics codes on current high-performance computers, with the main current limitation being the scalability of the algebraic multigrid methods used to solve the linearized equilibrium equations.

Applied computing → Computer-aided design↗

PeleC: An adaptive mesh refinement solver for compressible reacting flows

Reacting flow simulations for combustion applications require extensive computing capabilities. Leveraging the AMReX library, the Pele suite of combustion simulation tools targets the largest supercomputers available and future exascale machines. We introduce PeleC, the compressible solver in the Pele suite, and detail its capabilities, including complex geometry representation, chemistry integration, and discretization. We present a comparison of development efforts using both OpenACC and AMReX’s C++ performance portability framework for execution on multiple GPU architectures. We discuss relevant details that have allowed PeleC to achieve high performance and scalability. PeleC’s performance characteristics are measured through relevant simulations on multiple supercomputers. The success of PeleC’s design for exascale is exhibited through demonstration of a 160 billion cell simulation and weak scaling onto 100% of Summit, an NVIDIA-based GPU supercomputer at Oak Ridge National Laboratory. Our results provide confidence that PeleC will enable future combustion science simulations with unprecedented fidelity.

97 MATHEMATICS AND COMPUTING↗

Performance portable ice-sheet modeling with MALI

High-resolution simulations of polar ice sheets play a crucial role in the ongoing effort to develop more accurate and reliable Earth system models for probabilistic sea-level projections. These simulations often require a massive amount of memory and computation from large supercomputing clusters to provide sufficient accuracy and resolution; therefore, it has become essential to ensure performance on these platforms. Many of today’s supercomputers contain a diverse set of computing architectures and require specific programming interfaces in order to obtain optimal efficiency. In an effort to avoid architecture-specific programming and maintain productivity across platforms, the ice-sheet modeling code known as MPAS-Albany Land Ice (MALI) uses high-level abstractions to integrate Trilinos libraries and the Kokkos programming model for performance portable code across a variety of different architectures. In this article, we analyze the performance portable features of MALI via a performance analysis on current CPU-based and GPU-based supercomputers. The analysis highlights not only the performance portable improvements made in finite element assembly and multigrid preconditioning within MALI with speedups between 1.26 and 1.82x across CPU and GPU architectures but also identifies the need to further improve performance in software coupling and preconditioning on GPUs. We perform a weak scalability study and show that simulations on GPU-based machines perform 1.24–1.92x faster when utilizing the GPUs. The best performance is found in finite element assembly, which achieved a speedup of up to 8.65x and a weak scaling efficiency of 82.6% with GPUs. We additionally describe an automated performance testing framework developed for this code base using a changepoint detection method. The framework is used to make actionable decisions about performance within MALI. We provide several concrete examples of scenarios in which the framework has identified performance regressions, improvements, and algorithm differences over the course of 2 years of development.

54 ENVIRONMENTAL SCIENCES↗