Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Parallel application”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

Nuclear data sensitivity and uncertainty study of copper-reflected integral experiments [Slides]

This presentation touches on reducing uncertainties in intermediate-energy actinide nuclear data and this continues to be a high priority for many applications. The goal of PARADIGM (PARallel Approach of Differential and InteGral Measurements): accelerate efforts to reduce biases and uncertainties in nuclear data through improvements to the nuclear data pipeline. This presentation includes integral experiments and a summarization of existing copper nuclear data.

Cu63↗

Characterization and Optimization of the Fitting of Quantum Correlation Functions

This case study presents a characterization and optimization of an application code for extracting parton distribution functions from high energy electron-proton scattering data. Profiling this application code reveals that the phase-space density computation accounts for 93% of the overall execution time for a single iteration on a single core. When executing multiple iterations in parallel on a multicore system, the application spends 78% of its overall execution time idling due to load imbalance. We address these issues by first transforming the application code from Python to C++ and then tackling the application load imbalance via a hybrid scheduling strategy that combines dynamic and static scheduling. These techniques result in a 62% reduction in CPU idle time and a 2.46x speedup in overall execution time per node. In addition, the typically enabled power-management mechanisms in supercomputers (e.g., AMD Turbo Core, Intel Turbo Boost, and RAPL) can significantly impact intra-node scalability when more than 50% of the CPU cores are used. This finding underscores the importance of understanding system interactions with power management, as they can adversely impact application performance, and highlights the necessity of intra-node scaling tests to identify performance degradation that inter-node scaling tests might otherwise overlook.

Chuang, Pi-Yueh [Virginia Tech,Dept. of Computer S↗

AMReX and pyAMReX: Looking beyond the exascale computing project

AMReX is a software framework for the development of block-structured mesh applications with adaptive mesh refinement (AMR). AMReX was initially developed and supported by the AMReX Co-Design Center as part of the U.S. DOE Exascale Computing Project (ECP), and is continuing to grow post-ECP. In addition to adding new functionality and performance improvements to the core AMReX framework, we have also developed a Python binding, pyAMReX, that provides a bridge between AMReX-based application codes and the data science ecosystem. pyAMReX provides zero-copy application GPU data access for AI/ML, in situ analysis and application coupling, and enables rapid, massively parallel prototyping. In this paper we review the overall functionality of AMReX and pyAMReX, focusing on new developments, new functionality, and optimizations of key operations. We also summarize capabilities of ECP projects that used AMReX and provide an overview of new, non-ECP applications.

Myers, Andrew↗

The development and applications of multidimensional biomolecular spectroscopy illustrated by photosynthetic light harvesting

The parallel and synergistic developments of atomic resolution structural information, new spectroscopic methods, their underpinning formalism, and the application of sophisticated theoretical methods have led to a step function change in our understanding of photosynthetic light harvesting, the process by which photosynthetic organisms collect solar energy and supply it to their reaction centers to initiate the chemistry of photosynthesis. The new spectroscopic methods, in particular multidimensional spectroscopies, have enabled a transition from recording rates of processes to focusing on mechanism. We discuss two ultrafast spectroscopies – two-dimensional electronic spectroscopy and two-dimensional electronic-vibrational spectroscopy – and illustrate their development through the lens of photosynthetic light harvesting. Both spectroscopies provide enhanced spectral resolution and, in different ways, reveal pathways of energy flow and coherent oscillations which relate to the quantum mechanical mixing of, for example, electronic excitations (excitons) and nuclear motions. The new types of information present in these spectra provoked the application of sophisticated quantum dynamical theories to describe the temporal evolution of the spectra and provide new questions for experimental investigation. While multidimensional spectroscopies have applications in many other areas of science, we feel that the investigation of photosynthetic light harvesting has had the largest influence on the development of spectroscopic and theoretical methods for the study of quantum dynamics in biology, hence the focus of this review. We conclude with key questions for the next decade of this review.

59 BASIC BIOLOGICAL SCIENCES↗

Addressing Load Imbalance in Bioinformatics and Biomedical Applications: Efficient Scheduling across Multiple GPUs

Computational bioinformatics and biomedical applications frequently contain heterogeneously sized units of work or tasks, for instance due to variability in the sizes of biological sequences and molecules. Variable-sized workloads lead to load imbalances in parallel implementations which detract from efficiency and performance. Many modern computing resources now have multiple graphics processing units(GPUs) per computer for acceleration. These multiple GPU resources need to be used efficiently through balancing of workloads across the GPUs. OpenMP is a portable directive-based parallel programming API used ubiquitously in bioscience applications to program CPUs; recently, the use of OpenMP directives for GPU acceleration has become possible. Here, motivated by experiences with imbalanced loads in GPU-accelerated bioinformatics applications, we address the load balancing problem using OpenMP task-to-GPU scheduling combined with OpenMP GPU offloading for multiply heterogeneous workloads – loads with both variable input sizes, and simultaneously, variable convergence rates for algorithms with a stochastic component – scheduled across multiple GPUs. We aim to develop strategies which are both easy to use and have lower overheads, and may be incorporated incrementally in existing programs which already make use of OpenMP for CPU-based threading in order to make use of multi-GPU computers. We test different combinations of input size variability and convergence rate variability, and characterize the effects of these different scenarios on the performance of scheduling strategies across multiple GPUs with OpenMP. We present several dynamic scheduling solutions for different parallel patterns, explore optimizations, and provide publicly available example computational kernels to make these strategies easy to use in programs. This work will enable application developers to efficiently and easily use multiple GPUs for imbalanced workloads found in bioinformatics and biomedical applications.

Thavappiragasam, Mathialakan↗

Fast shared-memory streaming multilevel graph partitioning

In this report we show that a fast parallel graph partitioner can benefit many applications by reducing data transfers. The online methods for partitioning graphs have to be fast and they often rely on simple one-pass streaming algorithms, while the offline methods for partitioning graphs contain more involved algorithms and the most successful methods in this category belong to the multilevel approaches. In this work, we assess the feasibility of using streaming graph partitioning algorithms within the multilevel framework. Our end goal is to come up with a fast parallel offline multilevel partitioner that can produce competitive cutsize quality. We rely on a simple but fast and flexible streaming algorithm throughout the entire multilevel framework. This streaming algorithm serves multiple purposes in the partitioning process: a clustering algorithm in the coarsening, an effective algorithm for the initial partitioning, and a fast refinement algorithm in the uncoarsening. Its simple nature also lends itself easily for parallelization. The experiments on various graphs show that our approach is on the average up to 5.1x faster than the multi-threaded MeTiS, which comes at the expense of only 2x worse cutsize.

97 MATHEMATICS AND COMPUTING↗

Model Exploration of an Information-Based Healthcare Intervention Using Parallelization and Active Learning

This paper describes the application of a large-scale active learning method to characterize the parameter space of a computational agent-based model developed to investigate the impact of CommunityRx, a clinical information-based health intervention that provides patients with personalized information about local community resources to meet basic and self-care needs. Additionally, the diffusion of information about community resources and their use is modeled via networked interactions and their subsequent effect on agents' use of community resources across an urban population. A random forest model is iteratively fitted to model evaluations to characterize the model parameter space with respect to observed empirical data. We demonstrate the feasibility of using high-performance computing and active learning model exploration techniques to characterize large parameter spaces; by partitioning the parameter space into potentially viable and non-viable regions, we rule out regions of space where simulation output is implausible to observed empirical data. We argue that such methods are necessary to enable model exploration in complex computational models that incorporate increasingly available micro-level behavior data. We provide public access to the model and high-performance computing experimentation code.

97 MATHEMATICS AND COMPUTING↗

Scalable, In-situ Data Clustering Data Analysis for Extreme Scale Scientific Computing (Final Report)

The objective of this project is to address challenges in the design and development of scalable in-situ data clustering and analytics algorithms and software. Our goal is to develop parallel software consisting of a set of spatio-temporal data clustering and anomaly detection functions, both of which are very important for large-scale analysis and have wide applicability for in-situ runs as well as post-processing analysis. Our design principles for in-situ analysis consider the following: (1) identify parts of the computation can be done close to the data within the nodes, while it is still in memory; (2) extract analysis components can (and should) be performed in remote staging and analysis nodes; (3) develop error-bound approximation methods for applications tolerable for small errors; (4) identify the type of derived distributions and statistics, for spatio-temporal data, that can be kept locally in order to both accelerate computations and meet energy constraints in subsequent iterations and phases; (5) use a self-describing data format so that data can be consistent and understood among local storage (memory and SSDs) and at staging and analysis nodes, thereby providing portability and flexibility; (6) develop service-oriented functions that can schedule in-situ and post-hoc analysis tasks based on the dynamic requirements of applications. Our development focus is to produce the parallel data analysis software/library that will be scalable, reusable, extensible, and generic for applications in different disciplines. The software will be able to run in-situ with the simulations as well as post-hoc analysis. This approach will satisfy many synergistic requirements for data intensive applications executed on data coming from instruments and experiments. In particular, the proposed multilevel approach is directly applicable to perform design tradeoffs for running part of the algorithms near the instruments and the rest on remote (analysis) systems.

97 MATHEMATICS AND COMPUTING↗

Studying CPU and memory utilization of applications on Fujitsu A64FX and Nvidia Grace Superchip

ARM-based manycore CPU architectures are well-positioned to provide the rising memory throughput requirements of modern data intensive scientific applications in High Performance Computing (HPC). The Fujitsu A64FX CPU platform is based on the ARM v8.2A architecture, and is the processor of the flagship Japanese supercomputer - "Fugaku", which was previously ranked as the #1 supercomputer in the world according to the Top500 list. The Nvidia Grace superchip features 144 Neoverse V2 cores based on the ARMv9 architecture with 4x128b SVE2, providing exceptional computational power. The chip supports up to 480GB of memory, making it ideal for AI, machine learning, and scientific computing workloads. In this paper, we conduct a thorough performance exploration of a variety of parallel bandwidth-sensitive benchmarks and applications compiled with the native Fujitsu compiler on a Fugaku A64FX compute node and ARM (LLVM) Compiler on an NVIDIA Grace superchip compute node, engaging all the computational cores per cluster using OpenMP multithreading (assuming the cores can drive the available bandwidth). Our ultimate goals are to study the resource utilization of scientific applications and benchmarks on A64FX and Grace superchip, considering graph application scenarios ( GAP Benchmark suite) and eleven appli- cation proxies from the Rodinia heterogeneous benchmark suite (considering domains such as Data Mining, Bioinformatics, Fluid Dynamics, Pattern Recognition, etc.). Through exhaustive performance monitoring, we quantify the resource utilization of diverse OpenMP-based HPC applications on both the Fujitsu A64FX and the Nvidia Grace Superchip platforms.

benchmarking, Performance Analysis, High performan↗

ARENA: Asynchronous Reconfigurable Accelerator Ring to Enable Data-Centric Parallel Computing

The next generation HPC and data centers are likely to be reconfigurable and data-centric due to the trend of hardware specialization and the emergence of data-driven applications. In this work, we propose ARENA – an asynchronous reconfigurable accelerator ring architecture as a potential scenario on how the future HPC and data centers will be like. Despite using the coarse-grained reconfigurable arrays (CGRAs) as the substrate platform, our key contribution is not only the CGRA-cluster design itself, but also the ensemble of a new architecture and programming model that enables asynchronous tasking across a cluster of reconfigurable nodes, so as to bring specialized computation to the data rather than the reverse. We presume distributed data storage without asserting any prior knowledge on the data distribution. Hardware specialization occurs at runtime when a task finds the majority of data it requires are available at the present node. In other words, we dynamically generate specialized CGRA accelerators where the data reside. The asynchronous tasking for bringing computation to data is achieved by circulating the task token, which describes the dataflow graphs to be executed for a task, among the CGRA cluster connected by a fast ring network. Evaluations on a set of HPC and data-driven applications across different domains show that ARENA can provide better parallel scalability with reduced data movement (53.9 percent). Compared with contemporary compute-centric parallel models, ARENA can bring on average 4.37× speedup. The synthesized CGRAs and their task-dispatchers only occupy 2.93mm 2 chip area under 45nm process technology and can run at 800MHz with on average 759.8mW power consumption. ARENA also supports the concurrent execution of multi-applications, offering ideal architectural support for future high-performance parallel computing and data analytics systems.

97 MATHEMATICS AND COMPUTING↗

Impact of lowering potassium contamination in liquid scintillation cocktails for ultra-sensitive radiation detection

Intrinsic 40 K radioactive backgrounds from impurities of natural K in liquid scintillation cocktails have previously been demonstrated to limit their use in ultra-sensitive applications. Here, this work explores two methodologies in parallel for the reduction of 40 K backgrounds in the cocktails, and lays the groundwork for use in ultra-sensitive applications. In one method, alternative low-K liquid scintillation matrix constituents were identified and in the other, a simple purification method for single components and finished cocktails was developed. Both methods were verified via ICP-MS analysis. Liquid scintillation counting of selected purified cocktails demonstrated background reduction, improved stability, and enhanced performance. The best performing purified cocktail was also counted on a custom-built ultra-low background liquid scintillation counter, with results below the detector background.

38 RADIATION CHEMISTRY, RADIOCHEMISTRY, AND NUCLEA↗

Accelerating shared file checkpoint with local burst buffers

A data management system and method for accelerating shared file checkpointing. Written application data is aggregated in an application data file created in a local burst buffer memory at a compute node, and an associated data mapping built index to maintain information related to the offsets into a shared file at which segments of the application data is to be stored in a parallel file system, and where in the buffer those segments are located. The node asynchronously transfers a data file containing the application data and the associated data mapping index to a file server for shared file storage. The data management system and method further accelerates shared file checkpointing in which a shared file, together with a map file that specifies how the shared file is to be distributed, is asynchronously transferred to local burst buffer memories at the nodes to accelerate reading of the shared file.

Gooding, Thomas↗

Space-Time Block Preconditioning for Incompressible Flow

Parallel-in-time methods have become increasingly popular in the simulation of time-dependent numerical PDEs, allowing for the efficient use of additional message passing interface processes when spatial parallelism saturates. Most methods treat the solution and parallelism in space and time separately. In contrast, all-at-once methods solve the full space-time system directly, largely treating time as simply another spatial dimension. All-at-once methods offer a number of benefits over separate treatment of space and time, most notably significantly increased parallelism and faster time to solution (when applicable). However, the development of fast, scalable all-at-once methods has largely been limited to time-dependent (advection-)diffusion problems. This paper introduces the concept of space-time block preconditioning for the all-at-once solution of incompressible flow. By extending well-known concepts of spatial block preconditioning to the space-time setting, we develop a block preconditioner whose application requires the solution of a space-time (advection-)diffusion equation in the velocity block, coupled with a pressure Schur complement approximation consisting of independent spatial solves at each time-step, and a space-time matrix-vector multiplication. The new method is tested on four classical models in incompressible flow. Finally, the results indicate perfect scalability in refinement of spatial and temporal mesh spacing, perfect scalability in nonlinear Picard iteration count when applied to a nonlinear Navier--Stokes problem, and minimal overhead in terms of number of preconditioner applications compared with sequential time-stepping.

97 MATHEMATICS AND COMPUTING↗

Paralleling of LLC Resonant Converters

The LLC resonant converter is a popular, variable switching frequency DC-DC converter that may be controlled using two methods: charge and frequency control. In this paper, the application of LLC resonant converters to input-parallel, output-parallel system is studied. In this respect, the models of output-port I-V characteristics and small-signal output impedance of the charge controlled LLC converter are proposed. In addition, a mathematical framework is developed for droop-based paralleled DC-DC systems. Here, it distinctly identifies the output DC voltage and circulating current modes of stability, even in systems comprising of non-identical converters.

42 ENGINEERING↗

Impacts of Mode Mixity on Controlled Spalling of (100)-Oriented Germanium

Controlled spalling is a technology to prepare single-crystal thin films of semiconductors by fracture with a subsurface crack propagating nearly parallel to the substrate surface. Practical applications require uniform thickness and a smooth surface across the whole film. Both wafer-scale and patterned-stressor-defined small-area spalling of germanium substrates are conducted experimentally and numerically. River line features are observed on spalled surfaces close to lateral edges of the spall, regardless of the spall direction and the size of the spalled area. Three-dimensional finite element method modeling shows the river lines are caused by mixed mode I?+?III loading near the lateral edges of spall and predicts a spall depth variation near the lateral edges of spall due to mixed mode I?+?II loading. The absolute range of river lines increases with lateral size of spall, while the relative range of river lines decreases, consistent with variations in mode mixity.

36 MATERIALS SCIENCE↗

Performance of Julia for High Energy Physics Analyses

We argue that the Julia programming language is a compelling alternative to currently more common implementations in Python and C++ for common data analysis workflows in high energy physics. We compare the speed of implementations of different workflows in Julia with those in Python and C++. Furthermore, our studies show that the Julia implementations are competitive for tasks that are dominated by computational load rather than data access. For work that is dominated by data access, we demonstrate an application with concurrent file reading and parallel data processing.

97 MATHEMATICS AND COMPUTING↗

Towards improved speed and accuracy of laser powder bed fusion simulations via multiscale spatial representations

Due to the growing popularity of laser powder bed fusion (LPBF) as a metal additive manufacturing technique, there is a strong need to be able to accurately predict build outcomes. Full fidelity simulations of this process are not feasible due to the vast range of length and time scales inherent to it. While part-scale codes for simulating residual stress and distortion have shown reasonable predictive capability, they often neglect many aspects of the process occurring over smaller length/time scales, and thus are unable to capture effects of process parameter adjustments or the behavior of fine features. One way of capturing aspects at more refined length scales is through the use of adaptive mesh refinement (AMR). AMR allows for the process to be simulated at scales approaching the physical spatial dimensions without drastically increasing the total degrees of freedom in the simulation. This manuscript describes the implementation of an AMR algorithm within a multiphysics, parallelized finite element code, and its application to the LPBF problem. In this work, part-scale examples are provided where the use of AMR has allowed for higher fidelity thermal and thermomechanical simulations, as compared to experimental measurements. Results from these higher resolution simulations show that while AMR is a necessary component for increased accuracy in a computationally efficient manner, other improvements are also necessary, including handling of the multiple time scales inherent to the problem and the need for improved AM-specific material models.

42 ENGINEERING↗

STEPS: A Portable Numerical Simulation Toolkit for Electrical Power System Dynamic Studies

Numerical simulation is the key technique for large scale power system analysis. Redistribution of global renewable power via international interconnections requires new simulation tools to study the interconnected systems with different nominal frequencies as a whole. In this paper we introduce an open source simulation toolkit for electrical power systems (STEPS) which is hosted at Github. Its kernel is coded in C++ with major functions of power flow and electro-mechanical dynamic simulation. Flexible options are provided and configurable to improve power flow solution and dynamic simulation. Common devices and models are supported in STEPS for AC/DC hybrid system studies. Studies of interconnected systems with different nominal frequencies is supported in STEPS for research of international interconnection. Application program interfaces are provided and wrapped with Python to enable high-level interfaces for general applications. STEPS is thread safe and parallel computation is supported in both kernel and script levels to accelerate simulation. It is portable and works on Windows and GNU/Linux platforms. Cases from small to large scale systems are thoroughly tested to validate the toolkit with commercial packages as benchmarks.

42 ENGINEERING↗