Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “distributed memory”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

Progress on Optimizing Wind Farms and Rotor Designs Using Adjoints

Modern wind plants are increasingly tasked with multiple performance objectives. In addition to designing plants that maximize power output and minimize the levelized cost of energy (LCOE), the design and operation of wind plants is increasingly influenced by challenges regarding grid integration of variable generation renewables. This places a growing emphasis on making wind plants more controllable and predictable. WindSE is a Reynolds-averaged Navier-Stokes (RANS) model designed around analytical gradient and adjoint methods, with the ability to capture terrain-induced effects, as shown in Figure 1. The recent addition of an unsteady solver with an actuator line method (ALM) and ongoing work to enable massively parallel optimizations gives it a unique niche to explore coupled plant-level controls and design problems. This code is an open source python package built on the FEniCS framework that utilizes fast, parallel PETSc solvers to model fluid flow throughout wind-farm scale domains. Two recent studies performed using WindSE demonstrate the capability to optimize under a wide variety of flow conditions and objective functions. In the first, we present an optimization focused on modifying the layout of a wind farm with a fixed number of turbines for maximum total power output [1]. This study highlights the ability to quickly perform simulations using the steady Navier-Stokes solver combined with rotors represented as actuator disks while also stressing the importance of capturing terrain-induced effects. Gradient-based optimization using the RANS equations is viable due to the inclusion of efficiently computed adjoint derivatives. We interpret the physical results of the optimal layout and also discuss the computational cost of scaling to larger problems. In the second study, we present the capabilities of the unsteady Navier-Stokes solver, where rotor-blade profiles represented by actuator lines are optimized to enhance wake steering effects and overall power production [2]. We quantify the wind plant performance gains obtained from this type of simultaneous control co-design optimization as compared to optimizing the blade design and yaw independently. Figure 2 shows the differences between a baseline two-turbine system and an optimized system where we fine-tune the blade chord profile. Results and challenges from each study are quickly summarized and used to motivate the current development efforts within WindSE. Current and future work is focused on enabling higher-resolution studies with more degrees of freedom through parallelization of both the simulation and optimization algorithms. We present benchmarking results to show that WindSE performs well in both weak- and strong-scaling tests and further demonstrate that the optimizer obtains the same convergence rates in both shared- and distributed-memory environments. Using larger wind farms, we can study deep-array effects within an optimization context, allowing the use of objective functions that have been previously unstudied. As an example, we present ongoing work on a blockage metric which characterizes the loss of available kinetic energy due to wake effects from multiple upstream turbines.

adjoint optimization↗

Relocation and persistence of named data elements in coordination namespace

An approach is disclosed that relocates a named data element. A request to move a name corresponding to the named data element is received from a first storage area in a Coordination Namespace to a second storage area in the Coordination Namespace. The first storage area has a first level of persistence, and the second storage area has a second level of persistence. The named data element exists in a Coordination Namespace that is allocated in a memory distributed amongst a plurality of nodes that include the local node and one or more remote nodes. The approach then creates a copy of the named data element in the second storage area.

Nair, Ravi↗

symPACK: A GPU-Capable Fan-Out Sparse Cholesky Solver

Sparse symmetric positive definite systems of equations are ubiquitous in scientific workloads and applications. Parallel sparse Cholesky factorization is the method of choice for solving such linear systems. Therefore, the development of parallel sparse Cholesky codes that can efficiently run on today’s large-scale heterogeneous distributed-memory platforms is of vital importance. Modern supercomputers offer nodes that contain a mix of CPUs and GPUs. To fully utilize the computing power of these nodes, scientific codes must be adapted to offload expensive computations to GPUs. We present symPACK, a GPU-capable parallel sparse Cholesky solver that uses one-sided communication primitives and remote procedure calls provided by the UPC++ library. We also utilize the UPC++ "memory kinds" feature to enable efficient communication of GPU-resident data. We show that on a number of large problems, symPACK outperforms comparable state-of-the-art GPU-capable Cholesky factorization codes by up to 14x on the NERSC Perlmutter supercomputer.

Bellavita, Julian↗

Quantum optical memory for entanglement distribution

Optical photons are powerful carriers of quantum information, which can be delivered in free space by satellites or in fibers on the ground over long distances. Entanglement of quantum states over long distances can empower quantum computing, quantum communications, and quantum sensing. Quantum optical memories are devices designed to store quantum information in the form of stationary excitations, such as atomic coherence, and are capable of coherently mapping these excitations to flying qubits. Quantum memories can effectively store and manipulate quantum states, making them indispensable elements in future long-distance quantum networks. Over the past two decades, quantum optical memories with high fidelities, high efficiencies, long storage times, and promising multiplexing capabilities have been developed, especially at the single-photon level. In this review, we introduce the working principles of commonly used quantum memory protocols and summarize the recent advances in quantum memory demonstrations. We also offer a vision for future quantum optical memory devices that may enable entanglement distribution over long distances.

Lei, Yisheng↗

Distributed out-of-memory NMF on CPU/GPU architectures

We propose an efficient distributed out-of-memory implementation of the non-negative matrix factorization (NMF) algorithm for heterogeneous high-performance-computing systems. The proposed implementation is based on prior work on NMFk, which can perform automatic model selection and extract latent variables and patterns from data. In this work, we extend NMFk by adding support for dense and sparse matrix operation on multi-node, multi-GPU systems. The resulting algorithm is optimized for out-of-memory problems where the memory required to factorize a given matrix is greater than the available GPU memory. Memory complexity is reduced by batching/tiling strategies, and sparse and dense matrix operations are significantly accelerated with GPU cores (or tensor cores when available). Input/output latency associated with batch copies between host and device is hidden using CUDA streams to overlap data transfers and compute asynchronously, and latency associated with collective communications (both intra-node and inter-node) is reduced using optimized NVIDIA Collective Communication Library (NCCL) based communicators. Benchmark results show significant improvement, from 32X to 76x speedup, with the new implementation using GPUs over the CPU-based NMFk. Good weak scaling was demonstrated on up to 4096 multi-GPU cluster nodes with approximately 25,000 GPUs when decomposing a dense 340 Terabyte-size matrix and an 11 Exabyte-size sparse matrix of density 10 -6 .

97 MATHEMATICS AND COMPUTING↗

Spatially Distributed Ramp Reversal Memory in VO 2

Ramp‐reversal memory has recently been discovered in several insulator‐to‐metal transition materials where a non‐volatile resistance change can be set by repeatedly driving the material partway through the transition. This study uses optical microscopy to track the location and internal structure of accumulated memory as a thin film of VO 2 is temperature cycled through multiple training subloops. These measurements reveal that the gain of insulator phase fraction between consecutive subloops occurs primarily through front propagation at the insulator‐metal boundaries. By analyzing transition temperature maps, it is found, surprisingly, that the memory is also stored deep inside both insulating and metallic clusters throughout the entire sample, making the metal‐insulator coexistence landscape more rugged. This non‐volatile memory is reset after heating the sample to higher temperatures, as expected. Diffusion of point defects is proposed to account for the observed memory writing and subsequent erasing over the entire sample surface. By spatially mapping the location and character of non‐volatile memory encoding in VO 2 , this study results enable the targeting of specific local regions in the film where the full insulator‐to‐metal resistivity change can be harnessed in order to maximize the working range of memory elements for conventional and neuromorphic computing applications.

36 MATERIALS SCIENCE↗

LATTE: open-source, high-performance traveltime computation, tomography and source location in acoustic and elastic media

Traveltime-based tomography and source location are fundamental approaches for imaging subsurface structures and understanding the spatiotemporal distribution of seismicity from local to global scales. We present an open-source, high-performance framework integrating eikonal equation solvers and adjoint-state theory for traveltime computation, velocity tomography, source location and joint tomography-location in 2-D/3-D acoustic and elastic media. We introduce novel regularization schemes based on total generalized p-variation, structural similarity and multitask machine learning to enhance the fidelity and interpretability of inverted models and source locations. Key features of our implementation also include the ability to leverage both absolute-difference and double-difference traveltime misfits for high-fidelity velocity tomography and source parameter estimation; support for traveltime computation and inversion in diverse 2-D/3-D scenarios with arbitrary source and receiver distributions; and a perturbation-based optimal step-size estimation method to reduce computational costs. In addition, our implementation employs shared-memory and distributed-memory parallelization to provide an efficient solution for traveltime computation, tomography, and source location. In conclusion, we validate the efficacy and accuracy of our approach through multiple synthetic data examples.

58 GEOSCIENCES↗

IRIS-DMEM: Efficient Memory Management for Heterogeneous Computing

This paper proposes an efficient data memory management approach for the Intelligent RuntIme System (IRIS) heterogeneous computing framework along with new data transfer policies. IRIS provides a task-based programming model for extreme heterogeneous computing (e.g., CPU, GPU, DSP, FPGA) with support for today's most important programming languages (e.g., OpenMP, OpenCL, CUDA, HIP, OpenACC). However, the IRIS framework either forces the programmer to introduce data transfer commands for each task or relies on suboptimal memory management for automatic and transparent data transfers. The work described here extends IRIS with novel heterogeneous memory handling and introduces novel data transfer policies by employing the Distributed data MEMory handler (DMEM) for efficient and optimal movement of data among the various computing resources. The proposed approach achieves performance gains of up to 7× for tiled LU factorization and tiled DGEMM (i.e., matrix multiplication) benchmarks. Moreover, this approach also reduces data transfers by up to 71% when compared to previous IRIS heterogeneous memory management handlers. This work compares the performance results of the IRIS framework's novel DMEM with the StarPU runtime and MAGMA math library for GPUs. Experiments show a performance gain of up to 1.95× over StarPU and 2.1× over MAGMA.

Miniskar, Narasinga Rao↗

Distributed gather/scatter operations across a network of memory nodes

Devices, methods, and systems for distributed gather and scatter operations in a network of memory nodes. A responding memory node includes a memory; a communications interface having circuitry configured to communicate with at least one other memory node; and a controller. The controller includes circuitry configured to receive a request message from a requesting node via the communications interface. The request message indicates a gather or scatter operation, and instructs the responding node to retrieve data elements from a source memory data structure and store the data elements to a destination memory data structure. The controller further includes circuitry configured to transmit a response message to the requesting node via the communications interface. The response message indicates that the data elements have been stored into the destination memory data structure.

97 MATHEMATICS AND COMPUTING↗

Implementing Directive-Based Deferred Execution for Effective Network Aggregation

Remote direct memory access technology provides an efficient mechanism for one-sided communication that can be leveraged to implement a distributed shared memory programming model. However, when applications generate large numbers of small, irregular messages, network congestion often arises. Existing solutions address this small message problem by facilitating message aggregation but typically require disruptive code transformations that detract from the algorithmic intent of applications, or can be limited by dependent operations on aggregated data between synchronisation points. A solution is to use a directive-assisted approach that enables compilers to transform code dependent on aggregated communication for deferred execution. This paper presents an algorithm that a compiler can use to implement and optimise deferred execution for code dependent on aggregated data, based on an "aggregation context" extension for the OpenSHMEM partitioned global address space library. This new capability addresses a key challenge of message aggregation, allowing its full potential to reduce network congestion and enhance programmability to be realised.

Welch, Aaron [ORNL]↗

pyDRESCALk

Modern data scientists are tasked to analyze ever-growing data sets with increasingly complex relationships. Tensor decompositions have come to play a central role in identifying underlying latent structures in higher-order data. The problem of fitting tensor models to different distributions is complicated by the combinations of size, dimensionality, and sparsity present in real world data. The situation demands efficient algorithms designed for shared-memory and distributed systems. This work will present new research that tackles these challenges on several different fronts, leveraging optimizations in numerical algorithms and sparse tensor representations in heterogeneous high performance computing environments.

Bhattarai, Manish↗

Axially-deformed solution of the Skyrme-Hartree-Fock-Bogoliubov equations using the transformed harmonic oscillator basis (IV) HFBTHO (v4.0): A new version of the program

We describe the new version 4.0 of the code HFBTHO that solves the nuclear Hartree-Fock-Bogoliubov problem by using the deformed harmonic oscillator basis in cylindrical coordinates. In the new version, we have implemented the restoration of rotational, particle number, and reflection symmetry for even-even nuclei. The restoration of rotational symmetry does not require using bases closed under rotation. Furthermore, we added the SeaLL1 functional and improved the calculation of the Coulomb potential. Finally, we refactored the code to facilitate maintenance and future developments.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

In-depth analysis on parallel processing patterns for high-performance Dataframes

The Data Science domain has expanded monumentally in both research and industry communities during the past decade, predominantly owing to the Big Data revolution. Artificial Intelligence (AI) and Machine Learning (ML) are bringing more complexities to data engineering applications, which are now integrated into data processing pipelines to process terabytes of data. Typically, a significant amount of time is spent on data preprocessing in these pipelines, and hence improving its efficiency directly impacts the overall pipeline performance. The community has recently embraced the concept of Dataframes as the de-facto data structure for data representation and manipulation. However, the most widely used serial Dataframes today (R, pandas) experience performance limitations while working on even moderately large data sets. We believe that there is plenty of room for improvement by taking a look at this problem from a high-performance computing point of view. In a prior publication, we presented a set of parallel processing patterns for distributed dataframe operators and the reference runtime implementation, Cylon. In this paper, we are expanding on the initial concept by introducing a cost model for evaluating the said patterns. Furthermore, we evaluate the performance of Cylon on the ORNL Summit supercomputer.

97 MATHEMATICS AND COMPUTING↗

Parallel algorithms for finding connected components using linear algebra

Finding connected components is one of the most widely used operations on a graph. Optimal serial algorithms for the problem have been known for half a century, and many competing parallel algorithms have been proposed over the last several decades under various different models of parallel computation. This paper presents a class of parallel connected-component algorithms designed using linear-algebraic primitives. These algorithms are based on a PRAM algorithm by Shiloach and Vishkin and can be designed using standard GraphBLAS operations. Here, we demonstrate two algorithms of this class, one named LACC for Linear Algebraic Connected Components, and the other named FastSV which can be regarded as LACC’s simplification. With the support of the highly-scalable Combinatorial BLAS library, LACC and FastSV outperform the previous state-of-the-art algorithm by a factor of up to 12x for small to medium scale graphs. For large graphs with more than 50B edges, LACC and FastSV scale to 4K nodes (262K cores) of a Cray XC40 supercomputer and outperform previous algorithms by a significant margin. This remarkable performance is accomplished by (1) exploiting sparsity that was not present in the original PRAM algorithm formulation, (2) using high-performance primitives of Combinatorial BLAS, and (3) identifying hot spots and optimizing them away by exploiting algorithmic insights.

97 MATHEMATICS AND COMPUTING↗

A Decentralized Approach for Modeling Organized Convection Based on Thermal Populations on Microgrids

Abstract In this study, a spectral model for convective transport is coupled to a thermal population model on a two‐dimensional horizontal “microgrid,” covering the typical gridbox size of general circulation models. The goal is to explore new ways of representing impacts of spatial organization in cumulus cloud fields. The thermals are considered the smallest building block of convection, with thermal life cycle and movement represented through binomial functions. Thermals interact through two simple rules, reflecting pulsating growth and environmental deformation. Long‐lived thermal clusters thus form on the microgrid, exhibiting scale growth and spacing that represent simple forms of spatial organization and memory. Size distributions of cluster number are diagnosed from the microgrid through an online clustering algorithm, and provided as input to a spectral multiplume eddy‐diffusivity mass flux scheme. This yields a decentralized transport system, in that the thermal clusters acting as independent but interacting nodes that carry information about spatial structure. The main objectives of this study are (a) to seek proof of concept of this approach, and (b) to gain insight into impacts of spatial organization on convective transport. Single‐column model experiments demonstrate satisfactory skill in reproducing two observed cases of continental shallow convection. Metrics expressing self‐organization and spatial organization match well with large‐eddy simulation results. We find that in this coupled system, spatial organization impacts convective transport primarily through the scale break in the size distribution of cluster number. The rooting of saturated plumes in the subcloud mixed layer plays a key role in this process.

54 ENVIRONMENTAL SCIENCES↗

Distributed Quantum Computing with Photons and Atomic Memories

The promise of universal quantum computing requires scalable single- and inter-qubit control interactions. Currently, three of the leading candidate platforms for quantum computing are based on superconducting circuits, trapped ions, and neutral atom arrays. However, these systems have strong interaction with environmental and control noises that introduce decoherence of qubit states and gate operations. Alternatively, photons are well decoupled from the environment and have advantages of speed and timing for quantum computing. Photonic systems have already demonstrated capability for solving specific intractable problems like Boson sampling, but face challenges for practically scalable universal quantum computing solutions because it is extremely difficult for a single photon to “talk” to another deterministically. Here, a universal distributed quantum computing scheme based on photons and atomic-ensemble-based quantum memories is proposed. Taking the established photonic advantages, two-qubit nonlinear interaction is mediated by converting photonic qubits into quantum memory states and employing Rydberg blockade for the controlled gate operation. Spatial and temporal scalability of this scheme is demonstrated further. Furthermore, these results show photon-atom network hybrid approach can be a potential solution to universal distributed quantum computing.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗