Engineering Papers⌕ Search

Engineering topics

Ren, Bin

Publications and source records attributed to Ren, Bin.

Disk Evolution Study through Imaging of Nearby Young Stars (DESTINYS): A Panchromatic View of DO Tau’s Complex Kilo-astronomical-unit Environment

While protoplanetary disks are often treated as isolated systems in planet formation models, observations increasingly suggest that vigorous interactions between Class II disks and their environments are not rare. DO Tau is a T Tauri star that has previously been hypothesized to have undergone a close encounter with the HV Tau system. As part of the DESTINYS ESO Large Programme, we present new Very Large Telescope (VLT)/SPHERE polarimetric observations of DO Tau and combine them with archival Hubble Space Telescope (HST) scattered-light images and Atacama Large Millimeter/submillimeter Array (ALMA) observations of CO isotopologues and CS to map a network of complex structures. The SPHERE and ALMA observations show that the circumstellar disk is connected to arms extending out to several hundred astronomical units. HST and ALMA also reveal stream-like structures northeast of DO Tau, some of which are at least several thousand astronomical units long. These streams appear not to be gravitationally bound to DO Tau, and comparisons with previous Herschel far-IR observations suggest that the streams are part of a bridge-like structure connecting DO Tau and HV Tau. We also detect a fainter redshifted counterpart to a previously known blueshifted CO outflow. While some of DO Tau's complex structures could be attributed to a recent disk–disk encounter, they might be explained alternatively by interactions with remnant material from the star formation process. These panchromatic observations of DO Tau highlight the need to contextualize the evolution of Class II disks by examining processes occurring over a wide range of size scales.

79 ASTRONOMY AND ASTROPHYSICS↗

MICCO: An Enhanced Multi-GPU Scheduling Framework for Many-Body Correlation Functions

Calculation of many-body correlation functions is one of the critical kernels utilized in many scientific computing areas, especially in Lattice Quantum Chromodynamics (Lattice QCD). It is formalized as a sum of a large number of contraction terms each of which can be represented by a graph consisting of vertices describing quarks inside a hadron node and edges designating quark propagations at specific time intervals. Due to its computation- and memory-intensive nature, real-world physics systems (e.g., multi-meson or multi-baryon systems) explored by Lattice QCD prefer to leverage multi-GPUs. Different from general graph processing, many-body correlation function calculations show two specific features: a large number of computation-/data-intensive kernels and frequently repeated appearances of original and intermediate data. The former results in expensive memory operations such as tensor movements and evictions. The latter offers data reuse opportunities to mitigate the data-intensive nature of many-body correlation function calculations. However, existing graph-based multi-GPU schedulers cannot capture these data-centric features, thus resulting in a sub-optimal performance for many-body correlation function calculations. To address this issue, this paper presents a multi-GPU scheduling framework, MICCO, to accelerate contractions for correlation functions particularly by taking the data dimension (e.g., data reuse and data eviction) into account. This work first performs a comprehensive study on the interplay of data reuse and load balance, and designs two new concepts: local reuse pattern and reuse bound to study the opportunity of achieving the optimal trade-off between them. Based on this study, MICCO proposes a heuristic scheduling algorithm and a machine-learning-based regression model to generate the optimal setting of reuse bounds. Specifically, MICCO is integrated into a real-world Lattice QCD system, Redstar, for the first time running on multiple GPUs. The evaluation demonstrates MICCO outperforms other state-of-art works, achieving up to 2.25× speedup in synthesized datasets, and 1.49× speedup in real-world correlation functions.

Wang, Qihan↗

MemHC: An Optimized GPU Memory Management Framework for Accelerating Many-body Correlation

The many-body correlation function is a fundamental computation kernel in modern physics computing applications, e.g., Hadron Contractions in Lattice quantum chromodynamics (QCD). This kernel is both computation and memory intensive, involving a series of tensor contractions, and thus usually runs on accelerators like GPUs. Existing optimizations on many-body correlation mainly focus on individual tensor contractions (e.g., cuBLAS libraries and others). In contrast, this work discovers a new optimization dimension for many-body correlation by exploring the optimization opportunities among tensor contractions. More specifically, it targets general GPU architectures (both NVIDIA and AMD) and optimizes many-body correlation’s memory management by exploiting a set of memory allocation and communication redundancy elimination opportunities: first, GPU memory allocation redundancy: the intermediate output frequently occurs as input in the subsequent calculations; second, CPU-GPU communication redundancy: although all tensors are allocated on both CPU and GPU, many of them are used (and reused) on the GPU side only, and thus, many CPU/GPU communications (like that in existing Unified Memory designs) are unnecessary; third, GPU oversubscription: limited GPU memory size causes oversubscription issues, and existing memory management usually results in near-reuse data eviction, thus incurring extra CPU/GPU memory communications.

97 MATHEMATICS AND COMPUTING↗

COMET: A Domain-Specific Compilation of High-Performance Computational Chemistry

The computational power increases over the past decades have greatly enhanced the ability to simulate chemical reactions and understand ever more complex transformations. Tensor contractions are the fundamental computational building block of these simulations. These simulations have often been tied to one platform and restricted in generality by the interface provided to the user. The expanding prevalence of accelerators and researcher demands necessitate a more general approach which is not tied to specific hardware or requires contortion of algorithms to specific hardware platforms. In this paper we present COMET, a domain-specific programming language and compiler infrastructure for tensor contractions targeting heterogeneous accelerators. We present a system of progressive lowering through multiple layers of abstraction and optimization that achieves up to 1.98×speedup for 30 tensor contractions commonly used in computational chemistry and beyond.

Mutlu, Erdal↗

A High Performance Sparse Tensor Algebra Compiler in MLIR

Sparse tensor algebra is widely used in many applications, including scientific computing, machine learning, and data analytics. The performance of sparse tensor algebra kernels strongly depends on the intrinsic characteristics of the input tensors, hence many storage formats are designed for tensors to achieve optimal performance for particular applications/architectures, which makes it challenging to implement and optimize every tensor operation of interest on a given architecture. We propose a tensor algebra domain-specific language (DSL) and compiler framework to automatically generate kernels for mixed sparse-dense tensor algebra operations. The proposed DSL provides high-level programming abstractions that resemble the familiar Einstein notation to represent tensor algebra operations. The compiler introduces a new Sparse Tensor Algebra dialect built on top of LLVM's extensible MLIR compiler infrastructure for efficient code generation while covering a wide range of tensor storage formats. Our compiler also leverages input-dependent code optimization to enhance data locality for better performance. Our results show that the performance of automatically generated kernels outperforms the state-of-the-art sparse tensor algebra compiler, with up to 20.92x, 6.39x, and 13.9x performance improvement over state-of-the-art tensor algebra compilers, for parallel SpMV, SpMM, and TTM, respectively.

Tian, Ruiqin↗

MemXCT: Design, Optimization, Scaling, and Reproducibility of X-Ray Tomography Imaging

Here, this work extends our previous research entitled "MemXCT: Memory-centric X-ray CT Reconstruction with Massive Parallelization" that was originally published at SC19 conference (Hidayetoglu et al., 2019) with reproducibility of the computational imaging performance. X-ray computed tomography (XCT) is regularly used at synchrotron light sources to study the internal morphology of materials at high resolution. However, experimental constraints, such as radiation sensitivity, can result in noisy or undersampled measurements. Further, depending on the resolution, sample size and data acquisition rates, the resulting noisy dataset can be in the order of terabytes. Advanced iterative reconstruction techniques can produce high-quality images from noisy measurements, but their computational requirements have made their use an exception rather than the rule. We propose a novel memory-centric approach that avoids redundant computations at the expense of additional memory complexity. We develop a memory-centric iterative reconstruction system, MemXCT, that uses an optimized SpMV implementation with two-level pseudo-Hilbert ordering and multi-stage input buffering. We evaluate MemXCT on various supercomputer architectures involving KNL and GPU. MemXCT can reconstruct a large (11Kx11K) mouse brain tomogram in 10 seconds using 4096 KNL nodes (256K cores). The results presented in our original article at the SC19 were based on large-scale supercomputing resources. The MemXCT application was selected for the Student Cluster Competition (SCC) Reproducibility Challenge and evaluated on a variety of cloud computing resources by universities around the world in the SC20 conference. We summarize the results of the top-ranked SCC Reproducibility Challenge teams and identify the most pertinent measures for ensuring the reproducibility of our experiments in this article.

47 OTHER INSTRUMENTATION↗

A Decade of MWC 758 Disk Images: Where are the Spiral-Arm-Driving Planets?

Large-scale spiral arms have been revealed in scattered light images of a few protoplanetary disks. Theoretical models suggest that such arms may be driven by and corotate with giant planets, which has called for remarkable observational efforts to look for them. By examining the rotation of the spiral arms for the MWC 758 system over a 10 year timescale, we are able to provide dynamical constraints on the locations of their perturbers. We present reprocessed Hubble Space Telescope (HST)/NICMOS F110W observations of the target in 2005, and the new Keck/NIRC2 L'-band observations in 2017. MWC 758's two well-known spiral arms are revealed in the NICMOS archive at the earliest observational epoch. With additional Very Large Telescope (VLT)/SPHERE data, our joint analysis leads to a pattern speed of 0.°6(sup +3.°3, sub -0.°6)/yr at 3σ for the two major spiral arms. If the two arms are induced by a perturber on a near-circular orbit, its best-fit orbit is at 89 au (0."59), with a 3σ lower limit of 30 au (0."20). This finding is consistent with the simulation prediction of the location of an arm-driving planet for the two major arms in the system.

MWC 758↗