Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Parallel application”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 307 records · Page 17

Studies of whistler propagation along a plasma density gradient that is parallel to the magnetic field

Low frequency plasma wave generation in space is important for both scientific and practical applications. One of the most promising techniques for doing this is to directly inject whistler waves into the space environment from an antenna onboard one or more satellites. This technique has been discussed for years, but there are still open questions about the best way to generate plasma waves. So far, most theoretical [Kondrat92], lab based [Pribyl2010, Stenzel2016] and space-based experiments [DSX] have focused on studying the generation of whistler waves from an electric dipole antenna. However, a dipole antenna is very inefficient because it puts a lot of energy in waves that are not effective for most applications. Theoretical [Kondrat92] and lab experimental [Stenzel2016] results indicate that a loop antenna is much more efficient at generating whistler waves than a dipole antenna. A satellite experiment will need to be developed to demonstrate that whistler waves can be generated from a loop antenna in the space environment. The challenge is that to efficiently transmit whistler modes in the natural plasma environment of space, the loop antenna will have to be very large. For example, at L=2 (one earth radius away from the surface of the earth) a loop antenna would need a radius on the order of ~200 m to radiate efficiently, as shown in fig. 1, left. The antenna size and complexity would require a prohibitively large and expensive satellite mission. Our proposed innovation is to exploit the fact that the characteristic wavelength of whistler waves decreases in more dense plasma, which reduces the size needed for an antenna to radiate efficiently. Fortunately, a technique already exists for enhancing the local plasma density in space, called a plasma contactor [Kovaleski2001]. A plasma contactor can be used to create a local environment where the plasma density is enhanced around the satellite, which in turn reduces the size of an antenna that is needed to radiate efficiently (Fig. 1, right).

42 ENGINEERING↗

Computer Science Research Needs for Parallel Discrete Event Simulation (PDES)

Historically, scientific computing efforts have demonstrated the clear need for, and effective use of, supercomputing with traditional time-stepped simulations. Nevertheless, there are several areas in the mission spaces of the U.S. Department of Energy and other agencies waiting to tap advanced computing research using a different, discrete event style of modeling, simulation, and analysis. These span a wide spectrum of applications including energy grid resilience, urban planning and policy, transportation science, building technologies, emergency response and planning, environmental impact analysis, computational epidemiology, Internet communications, cyber security, and cyber-physical systems, to name only a few. Even within traditional scientific applications, the role of discrete event modes of execution is increasing in the form of new event-based mathematical solvers such as quantized state integration methods and discrete-continuous hybrid system solvers. Co-design of advanced supercomputing hardware systems is another area that exploits discrete event simulation at its core for effective analyses. Complex systems, entity behaviors and interconnections play a significant role in all these applications, which are mapped to large-scale models with discrete event formulations. To make advancements in all the aforementioned scientific areas, many technical aspects need to be more thoroughly studied and deeply understood in parallel discrete event simulation (PDES). The unique dynamics inherent in a discrete event modeling approach, by their very nature, intersect and influence the entire stack of the computing system, including (a) the unique nature of the instruction sets exercised in PDES workloads without a predominance of high-precision floating point operations, (b) virtual time-constrained multi-threaded execution of many logical processes per processor, (c) extremely variable and difficult to predict network traffic characteristics, (d) interfaces and inter-dependencies with machine learning and artificial intelligence codes at higher software layers, and (e) highly challenging load balancing needs, especially in effectively accounting for accelerated/extremely heterogeneous computing in current and future high-performance computing systems. Efficient and accurate parallel execution of PDES workloads is also dominated by challenges in dealing with their asynchronous concurrency fundamentally present at the model level. Conservative synchronization, optimistic/speculative synchronization, and their hybrid schemes open new questions in fundamental computer science with respect to reversibility of computation and prediction (lookahead) of behaviors inherent within model codes. On the implementation front, there are relatively few scalable, general-purpose parallel discrete event simulators in the world, and even fewer have been studied on emerging hardware platforms. To enable scientific advances using PDES, the research needs in computer science must also be pursued and met in the intersection of the algorithmic and hardware-aware aspects of scalable PDES engines. This report is aimed at capturing a computer science-oriented view of this important area of research in PDES, presenting a sample of important applications with their inherent discrete event technology elements. Needs are outlined in core areas of parallel discrete event research as well as cross-cutting directions in computer science research that positively impact scientific advancements across several important application areas. A selection of priority research opportunities in advanced computing for PDES is identified to serve as reference for key research topics and their order of importance for scientific advancements.

97 MATHEMATICS AND COMPUTING↗

ArborX: A Performance Portable Geometric Search Library

Searching for geometric objects that are close in space is a fundamental component of many applications. The performance of search algorithms comes to the forefront as the size of a problem increases both in terms of total object count as well as in the total number of search queries performed. Scientific applications requiring modern leadership-class supercomputers also pose an additional requirement of performance portability, i.e., being able to efficiently utilize a variety of hardware architectures. In this article, we introduce a new open-source C++ search library, ArborX, which we have designed for modern supercomputing architectures. Herein, we examine scalable search algorithms with a focus on performance, including a highly efficient parallel bounding volume hierarchy implementation, and propose a flexible interface making it easy to integrate with existing applications. We demonstrate the performance portability of ArborX on multi-core CPUs and GPUs and compare it to the state-of-the-art libraries such as Boost.Geometry.Index and nanoflann.

97 MATHEMATICS AND COMPUTING↗

Employing artificial intelligence to steer exascale workflows with colmena

Computational workflows are a common class of application on supercomputers, yet the loosely coupled and heterogeneous nature of workflows often fails to take full advantage of their capabilities. We created Colmena to leverage the massive parallelism of a supercomputer by using Artificial Intelligence (AI) to learn from and adapt a workflow as it executes. Colmena allows scientists to define how their application should respond to events (e.g., task completion) as a series of cooperative agents. In this paper, we describe the design of Colmena, the challenges we overcame while deploying applications on exascale systems, and the science workflows we have enhanced through interweaving AI. The scaling challenges we discuss include developing steering strategies that maximize node utilization, introducing data fabrics that reduce communication overhead of data-intensive tasks, and implementing workflow tasks that cache costly operations between invocations. These innovations coupled with a variety of application patterns accessible through our agent-based steering model have enabled science advances in chemistry, biophysics, and materials science using different types of AI. In conclusion, our vision is that Colmena will spur creative solutions that harness AI across many domains of scientific computing.

Workflows↗

Small tensor product distributed active space (STP-DAS) framework for relativistic and non-relativistic multiconfiguration calculations: Scaling from 10 9 on a laptop to 10 12 determinants on a supercomputer

Despite the power and flexibility of configuration interaction (CI) based methods in computational chemistry, their broader application is limited by an exponential increase in both computational and storage requirements, particularly due to the substantial memory needed for excitation lists that are crucial for scalable parallel computing. Here, the objective of this work is to develop a new CI framework, namely, the small tensor product distributed active space (STP-DAS) framework, aimed at drastically reducing memory demands for extensive CI calculations on individual workstations or laptops, while simultaneously enhancing scalability for extensive parallel computing. Moreover, the STP-DAS framework can support various CI-based techniques, such as complete active space (CAS), restricted active space, generalized active space, multireference CI, and multireference perturbation theory, applicable to both relativistic (two- and four-component) and non-relativistic theories, thus extending the utility of CI methods in computational research. We conducted benchmark studies on a supercomputer to evaluate the storage needs, parallel scalability, and communication downtime using a realistic exact-two-component CASCI (X2C-CASCI) approach, covering a range of determinants from 10 9 to 10 12 . Additionally, we performed large X2C-CASCI calculations on a single laptop and examined how the STP-DAS partitioning affects performance.

Complete-active space self-consistent field↗

HipBone: A performance-portable graphics processing unit-accelerated C++ version of the NekBone benchmark

We present hipBone, an open-source performance-portable proxy application for the Nek5000 (and NekRS) computational fluid dynamics applications. HipBone is a fully GPU-accelerated C++ implementation of the original NekBone CPU proxy application with several novel algorithmic and implementation improvements which optimize its performance on modern fine-grain parallel GPU accelerators. Our optimizations include a conversion to store the degrees of freedom of the problem in assembled form in order to reduce the amount of data moved during the main iteration and a portable implementation of the main Poisson operator kernel. We demonstrate near-roofline performance of the operator kernel on three different modern GPU accelerators from two different vendors. We present a novel algorithm for splitting the application of the Poisson operator on GPUs which aggressively hides MPI communication required for both halo exchange and assembly. Our implementation of nearest-neighbor MPI communication then leverages several different routing algorithms and GPU-Direct RDMA capabilities, when available, which improves scalability of the benchmark. We demonstrate the performance of hipBone on three different clusters housed at Oak Ridge National Laboratory, namely, the Summit supercomputer and the Frontier early-access clusters, Spock and Crusher. Our tests demonstrate both portability across different clusters and very good scaling efficiency, especially on large problems.

Computer Science↗

Meeting Report on the 3rd Chinese American Society for Mass Spectrometry Conference—Advancing Biological and Pharmaceutical Mass Spectrometry

Following the highly successful Chinese American Society for Mass Spectrometry (CASMS) conferences in the previous 2 years, the 3rd CASMS Conference was held virtually on August 28–31, 2023, using the Gather. Town platform to bring together scientists in the MS field. The conference offered a 4-day agenda with a scientific program consisting of two plenary lectures, and 14 parallel symposia in which a total of 70 speakers presented technological innovations and their applications in proteomics and biological MS and metabo-lipidomics and pharmaceutical MS. In addition, 16 invited speakers/panelists presented at two research-focused and three career development workshops. Moreover, 86 posters, 12 lightning talks, 3 sponsored workshops, and 11 exhibitions were presented, from which 9 poster awards and 2 lightning talk awards were selected. Furthermore, the conference featured four young investigator awardees to highlight early-career achievements in MS from our society. In conclusion, the conference provided a unique scientific platform for young scientists (i.e. graduate students, postdocs, and junior faculty/investigators) to present their research, meet with prominent scientists, learn about career development, and job opportunities (http://casms.org).

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Simulations of plasmas and fluids using anti-symmetric models

ALMA (Anti-symmetric, Large-Moment, Accelerated) is a fast, flexible, and scalable toolkit designed to solve hyperbolic conservation law systems in hybrid supercomputers. Here this manuscript describes the theoretical background and implementation of ALMA, which uses the anti-symmetric formulation of fluids to obtain simple, robust, and easily paralellizable code. Practical GPU acceleration is realized on entire applications with an overall gain factor of 2 to 4. ALMA also provides a parallel, GPU accelerated sparse solver based on geometric multigrid, capable of diagonalizing linear systems with 239 unknowns. Here we demonstrate ALMA's scaling and performance in petascale supercomputers and use standard fluid models to verify the overall approach with canonical benchmark problems.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Synthetic-domain computing and neural networks using lithium niobate integrated nonlinear phononics

Analogue computing uses the physical behaviours of devices to provide energy-efficient arithmetic operations. However, scaling up analogue computing platforms by simply increasing the number of devices leads to challenges such as device-to-device variation. Here, in this study, we report scalable analogue computing and neural networks in the synthetic frequency domain using an integrated nonlinear phononic platform on lithium niobate. This synthetic-domain computing is robust to device variations, as vectors and matrices are concurrently encoded at different frequencies within a single device, achieving a high throughput per area. Leveraging inherent nonlinearities, our device-aware neural network can perform a four-class classification task with an accuracy of 98.2%. The nonlinear phononic computing hardware also maintains consistent performance over a wide operational temperature range (characterized up to 192 °C). Our synthetic-domain computing combines single-device parallelism, inherent nonlinearity and environmental stability, and could be of use in edge computing applications in which power efficiency and environmental resilience are crucial.

Ji, Jun [Virginia Polytechnic Inst. and State Univ↗

Analyzing the Performance Trade-Off in Implementing User-Level Threads

User-level threads have been widely adopted as a means of achieving lightweight concurrent execution without the costs of OS-level threads. Nevertheless, the costs of managing user-level threads represent a performance barrier that dictates how fine grained the concurrency exposed by an application can be without incurring significant overheads; this in turn may translate into insufficient parallelism to exploit highly parallel systems. This article is a deep dive into the fundamental costs in implementing user-level threads. We first identify that one of the highest sources of fork-join overheads stems from deviations, events that incur context switching during the execution of a thread and disrupt a run-to-completion execution. We then conduct an in-depth investigation of a wide spectrum of methods with respect to how they handle deviations while covering both parent- and child-first scheduling policies. Our methodology involves a comprehensive instruction- and cache-level analysis of all methods on several modern CPU architectures. Finally, the primary finding of our evaluation is that dynamic promotion methods that assume the absence of deviation and dynamically provide context-switching support offer the best trade-off between performance and capability when the likelihood of deviation is low.

97 MATHEMATICS AND COMPUTING↗

Caffeine v0.1.0

Caffeine is the CoArray Fortran Framework of Efficient Interfaces to Network Environments. Caffeine aims to produce a parallel runtime library that will support Fortran compilers with a programming-model-agnostic application binary interface (ABI) to various lower-level communication libraries. The current version of Caffeine uses the GASNet-EX networking middleware, also developed at Berkeley Lab. On many combinations of applications and platforms, GASNet-EX outperforms the widely used Message Passing Interface (MPI). Through GASNet-EX's support for communicating between graphics processing units (GPU), GASNet-EX has features that specifically target the emerging, leading-edge exascale computing platforms.

Rouson, Damian↗

MrHyDE v.1.0

SAND2024-01324O MrHyDE, which stands for Multi-resolution Hybridized Differential Equations, is a general-purpose C++ package for the solution of coupled multiphysics and multiscale systems on massively parallel computing systems. MrHyDE is designed to enable moving beyond forward simulation for multiscale applications which includes optimization, control, uncertainty quantification, and stochastic inversion. The framework provides interfaces to several packages within the Trilinos framework and leverages automatic differentiation to enable adjoint capabilities for large-scale, gradient-based optimization. MrHyDE provides automated multiscale capabilities through a subgrid model interface and multiscale Dirichlet-to-Neumann maps. For extreme-scale applications, MrHyDE provides in situ data-compression algorithms to reduce memory requirements while maintaining performance. MrHyDE is a general-purpose, computational framework for the solution of multiscale and multiphysics applications. It uses a combination of structure-preserving, physics-compatible discretizations, fully implicit methods, multi-resolution schemes, or fully explicit methods. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

SciDAC↗

VerifyIO: Verifying Adherence to Parallel I/O Consistency Semantics

VerifyIO is a tool designed for verifying I/O consistency semantics in High-Performance Computing (HPC) applications. It addresses the challenges of ensuring correctness and portability across different I/O consistency models, such as POSIX, Commit, Session, and MPI-IO. By analyzing execution traces, detecting conflicts, and verifying synchronization adherence, VerifyIO provides actionable insights for both application developers and I/O library designers.

Wang, Chen [Lawrence Livermore National Laboratory↗

Performance Portability Evaluation of Fluid-Structure Interaction Simulations on Heterogeneous Platforms

The rapid proliferation of heterogeneous programming languages and multi-vendor hardware has underscored the critical need to evaluate the performance portability of scientific applications. In this work, we present the systematic porting and optimization of a massively parallel fluid-structure interaction code across multiple heterogeneous programming frameworks for deployment on leadership-class supercomputers from major vendors. Our analysis focuses on at-scale performance for simulations involving hundreds of millions of deformable cells, executed on a combination of CPUs and GPUs spanning thousands of nodes on exascale machines. We benchmark the performance of each implementation, highlighting the trade-offs inherent in adopting diverse programming models. Key insights regarding the portability of CUDA on multi-vendor platforms, the superior multi-core CPU performance from SYCL, and architectural considerations on performance optimization are distilled from our experience, offering guidance to other users of high performance computing based on our findings.

Martin, Aristotle [Duke University]↗

Ultra-Fast Vertical Ordering of Lamellar Block Copolymer Films on Unmodified Substrates

To utilize the full potential of block copolymer (BCP) thin films for use in technological devices ranging from ion conduction membranes, transistors to nanowire array antennas, rapid self- assembly of lamellar block copolymers (l-BCPs) with vertically oriented lamellar domains on a variety of unmodified substrates is needed. l-BCPs have an inherently larger interfacial area for transport compared to their cylindrical counterpart. Our observations demonstrate that the as-cast weakly ordered vertically oriented state of l-BCP films of polystyrene-block-poly(methyl methacrylate) (PS-b-PMMA) from directional evaporation of select solvents, act as “seed templates” for their ultra-fast evolution (~30 s) into well-ordered vertically oriented nanostructures, using a thermal gradient-based Cold Zone Annealing (CZA) technique. Furthermore, vertical lamellae are obtained on unmodified substrates, Quartz and Kapton, and the kinetics of l-BCP ordering is much faster by CZA as compared to the isotropic oven annealing. The rapid ordering kinetics of vertical l-BCPs is tested and found applicable to different molecular weights and film thicknesses ranging from 20 nm to 480 nm, which ultimately flip over to their equilibrium parallel morphology at upper limits of annealing times. This rapid ordering strategy for vertical orientation of l-BCPs using roll-to-roll compatible CZA would be highly relevant for fundamental studies of interfacial transport as well as for industrial applications from membranes to nanowires.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Scalable Parallel Measurement of Individual Nitrogen-Vacancy Centers

The nitrogen-vacancy (NV) center in diamond is a solid-state spin defect that has been widely adopted for quantum sensing and quantum information processing applications. Typically, experiments are performed either with a single isolated NV center or with an unresolved ensemble of many NV centers, resulting in a trade-off between measurement speed and spatial resolution or control over individual defects. In this work, we introduce an experimental platform that bypasses this trade-off by addressing multiple optically resolved NV centers in parallel. We perform charge- and spin-state manipulations selectively on multiple NV centers from within a larger set, and we manipulate and measure the electronic spin states of over 100 NV centers in parallel. We show that the high signal-to-noise ratio of the measurements enables the detection of shot-to-shot pairwise correlations between the spin states of 108 NV centers, corresponding to the simultaneous measurement of 5778 unique correlation coefficients. We discuss how our platform can be scaled to parallel experiments with thousands of individually resolved NV centers. These results enable parallelized high-throughput sensing experiments that retain the spatial resolution of single defects and will, thereby, help to unlock advances in applications such as single-molecule NMR and characterization of integrated circuits. In addition, our approach to multiplexing provides a natural platform for the application of recently developed correlated sensing techniques.

NV centers↗

Design-time performance modeling of compositional parallel programs

Performance models are powerful instruments for understanding the performance of parallel systems and uncovering their bottlenecks. Already during system design, performance models can help ponder alternative development options. However, creating a performance model – whether theoretically or empirically – for an entire application that does not exist yet is challenging. In this paper, we propose to generate performance models of full programs from performance models of their components using formal composition operators derived from parallel design patterns. As long as the design of the overall system follows such a pattern, its performance model can be predicted with reasonable accuracy without an actual implementation. In conclusion, we demonstrate our approach with design patterns of varying complexity, including pipeline, task pool, and eventually MapReduce, which is representative of a broad class of data-analytics applications.

97 MATHEMATICS AND COMPUTING↗

EXAGRAPH: Graph and combinatorial methods for enabling exascale applications

Combinatorial algorithms in general and graph algorithms in particular play a critical enabling role in numerous scientific applications. However, the irregular memory access nature of these algorithms makes them one of the hardest algorithmic kernels to implement on parallel systems. With tens of billions of hardware threads and deep memory hierarchies, the exascale computing systems in particular pose extreme challenges in scaling graph algorithms. The codesign center on combinatorial algorithms, ExaGraph, was established to design and develop methods and techniques for efficient implementation of key combinatorial (graph) algorithms chosen from a diverse set of exascale applications. Algebraic and combinatorial methods have a complementary role in the advancement of computational science and engineering, including playing an enabling role on each other. In this paper, we survey the algorithmic and software development activities performed under the auspices of ExaGraph from both a combinatorial and an algebraic perspective. In particular, we detail our recent efforts in porting the algorithms to manycore accelerator (GPU) architectures. We also provide a brief survey of the applications that have benefited from the scalable implementations of different combinatorial algorithms to enable scientific discovery at scale. We believe that several applications will benefit from the algorithmic and software tools developed by the ExaGraph team.

97 MATHEMATICS AND COMPUTING↗