Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “memory spaces”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Dynamic kernel memory space allocation

A processing unit includes one or more processor cores and a set of registers to store configuration information for the processing unit. The processing unit also includes a coprocessor configured to receive a request to modify a memory allocation for a kernel concurrently with the kernel executing on the at least one processor core. The coprocessor is configured to modify the memory allocation by modifying the configuration information stored in the set of registers. In some cases, initial configuration information is provided to the set of registers by a different processing unit. The initial configuration information is stored in the set of registers prior to the coprocessor modifying the configuration information.

Gutierrez, Anthony↗

Developing And Scaling an OpenFOAM Model to Study Turbulent Flow in a HFIR Coolant Channel

Improving the understanding of how computational fluid dynamics (CFD) direct numerical simulations (DNS) of flows in the High Flux Isotope Reactor (HFIR) perform when run in parallel using the high performance computing (HPC) platform Summit at the Oak Ridge Leadership Computing Facility (OLCF) is of particular importance to boost the computational tools used to support HFIR conversion to low enriched fuel (LEU). Evaluation of scaling performance was driven by the increasing importance of graphics processing unit (GPU) usage in HPC, which is becoming the standard for modern supercomputers such as Summit. The desired results are to obtain a strong positive correlation between the computational resources dedicated to a problem and the relative speed-up of the simulation in comparison to a benchmark. This capability will allow substantially improvement in HFIR flow analytical capabilities, specifically when predicting turbulence properties at high Reynolds numbers. The study leverages previous simulation results performed with code PHASTA (finite element) on HPC platforms Cori (NERSC) and Theta (ALCF) [1] with computing options provided in the computing platform OpenFOAM (finite volume) at OLCF. Transitioning from PHASTA to OpenFOAM will (1) eliminate dependence on third-party software for mesh generation and manipulation, (2) reduce resource needs by employing modern architectures, and (3) build expertise for future modeling of HFIR-specific problems like heat transfer in involute geometry, entrance effects, flow structure in channel corners, and so on—all important issues when defining the available thermal margins in the transition to LEU. CPUs and GPUs differ significantly in their architecture and utilization, as discussed in the literature [2]. The most important differences are in the approach to computations and their memory. A single GPU contains a large quantity of cores, enabling it to perform with a much higher throughput than a CPU, but execution requires a different approach. GPU codes execute instructions using the Single-Instruction Multiple-Thread (SIMT) approach in which a single instruction is used for groups of threads called warps. A warp typically consists of 32 threads which must execute the same set of instructions, although on separate threads. Alternately, a CPU has far fewer cores that are much more flexible in their operation, excelling at quickly performing more complex serial computations. This is why GPUs have greater throughput when properly utilized. The second important difference is seen when comparing their memory spaces. Limited memory allocations and CPU–GPU communications cause a significant bottleneck in GPU-accelerated programs. Further study was required to properly take advantage of GPU resources. A comprehensive analysis of code performance and the model-specific features of turbulence constitutes the core of this work. In this study, a DNS simulation of HFIR channel turbulence was performed with the finite volume CFD code OpenFOAM v2112 and CUDA v11.0 on Red Hat Enterprise Linux v8.2. The OpenFOAM installation had AMGx integrated to enable GPU acceleration and utilizes the PETSc4FOAM library. The computational resources and the problem size were scaled on CPU and CPU + GPU architectures to gain a better understanding of the performance of a DNS problem on modern computing hardware. The study aimed to analyze the scaling of the code exclusively on CPUs and then to examine the scaling of the codes with GPU acceleration enabled. Scaling studies included CPU and GPU acceleration on a mesh of varying resolution to analyze the impact of problem size relative to computational resources. In the course of preparing the GPU configuration on Summit, mainly using the AMGX solvers, difficulties were encountered stemming from constant changes resulting from extensive ongoing development activities and the changing environment. This resulted in the inability to complete the GPU portion of the work. The code was compiled and tested, but production runs to assess acceleration were not performed because the used discretional compute time allocation expired as year-end approached. The Summit HPC platform is scheduled for decommissioning in 2024, making it unattractive for future use with Nvidia-based GPUs. Therefore, the work will be moved onto NERSC machines in FY24. An application was prepared and submitted, and sufficient node-hours were awarded to continue the research in the next calendar year. This report summarizes work performed thus far, which mostly focused on CPU OpenFOAM computing.

21 SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLANTS↗

A view into FleCSI [Slides]

FleCSI is largely based around sharing of memory spaces. FleCSI facilitates parallel execution and sharing of memory across memory spaces. Execute() transforms sender-side data structures into receiver-side data structures.

97 MATHEMATICS AND COMPUTING↗

Design Space Exploration of Emerging Memory Technologies for Machine Learning Applications

Memory design space exploration methods study memory systems’ performances and limitations before implementation. The computer memory design space has grown exponentially because of the enormous growth of memory types, memory controllers, and application software. Computer simulators are commonly used for memory design space exploration. However, complex memory simulations take an enormous amount of time. Hence, in this paper, we proposed a machine learning-based design space exploration method for dynamic random-access memory and non-volatile memory systems. We applied our method to the CosmoGAN and LeNet applications to predict the following six memory response parameters: (i) bandwidth, (ii) power, (iii) average latency, (iv) average total latency, (v) memory reads, and (vi) memory writes. Our experimental results show that machine learning models can predict memory response parameter values faster than simulations. We used support vector machine, random forest, and gradient boosting machine learning models. We observed that the support vector machine provides better performance for bandwidth, average latency, and average total latency. The random forest model works better for memory reads and writes. The gradient boosting model provides superior prediction performance for power. We provide a detailed discussion on learning curve characteristics, error analysis, and memory type recommendation.

Hasan, S M Shamimul↗

Consistent Space Runtime (CSPACER) v1.0

This software provides a runtime for communication-bound applications, especially with irregular communication patterns. The target programming abstraction for the runtime is the space consistency model, which defines consistency guarantees at the granularity of memory spaces. This model has relaxed consistency semantics that enables a wide range of runtime optimizations. The runtime leverages threading to accelerate communication primitives, especially collective operations. It also allows efficient pipelining of communication operations and enable constructing a consistent state of multiple unordered communication activities targeting a memory space. The runtime uses a reduced API design that decomposes complex communication primitives in traditional general-purpose runtime into a sequence of simpler steps. To improve the productivity of using this runtime, we provide communication patterns commonly used for regular scientific computing applications and irregular data analytics. These communication patterns offer skeletons for the integration with application computation to allow efficient overlap.

Ibrahim, Khaled↗

Design and Performance of Kokkos Staging Space toward Scalable Resilient Application Couplings

With the growing number of applications designed for heterogeneous HPC devices, application programmers and users are finding it challenging to compose scalable workflows as ensembles of these applications, that are portable, performant and resilient. The Kokkos C++ library has been designed to simplify this cumbersome procedure by providing an intra-application uniform programming model and portable performance. However, assembling multiple Kokkos-enabled applications into a complex workflow is still a challenge. Although Kokkos enables a uniform programming model, the inter-application data exchange still remains a challenge from both performance and software development cost perspectives. In order to address this issue, we propose Kokkos data staging memory space, an extension of Kokkos' data abstraction (memory space) for heterogeneous computing systems. This new abstraction allows to express data on a virtual shared-space for multiple Kokkos applications, thus extending Kokkos to support inter-application data exchange to build an efficient application workflow. Additionally, we study the effectiveness of asynchronous data layout conversions for applications requiring different memory access patterns for the shared data. Our preliminary evaluation with a synthetic benchmark indicate the effectiveness of this conversion adapted to three different scenarios representing access frequency and use patterns of the shared data.

97 MATHEMATICS AND COMPUTING↗

Performance of CUDA Unified Memory in CMS Heterogeneous Pixel Reconstruction

The management of separate memory spaces of CPUs and GPUs brings an additional burden to the development of software for GPUs. To help with this, CUDA unified memory provides a single address space that can be accessed from both CPU and GPU. The automatic data transfer mechanism is based on page faults generated by the memory accesses. This mechanism has a performance cost, that can be with explicit memory prefetch requests. Various hints on the inteded usage of the memory regions can also be given to further improve the performance. The overall effect of unified memory compared to an explicit memory management can depend heavily on the application. In this paper we evaluate the performance impact of CUDA unified memory using the heterogeneous pixel reconstruction code from the CMS experiment as a realistic use case of a GPU-targeting HEP reconstruction software. We also compare the programming model using CUDA unified memory to the explicit management of separate CPU and GPU memory spaces.

Kortelainen, Matti J.↗

Enabling Scalable and Extensible Memory-mapped Datastores in Userspace

Exascale workloads are expected to incorporate data-intensive processing in close coordination with traditional physics simulations. These emerging scientific, data-analytics and machine learning applications need to access a wide variety of datastores in flat files and structured databases. Programmer productivity is greatly enhanced by mapping datastores into the application process's virtual memory space to provide a unified “in-memory” interface. Currently, memory mapping is provided by system software primarily designed for generality and reliability. However, scalability at high concurrency is a formidable challenge on exascale systems. Also, there is a need for extensibility to support new datastores potentially requiring HPC data transfer services. In this article, we present UMap , a scalable and extensible userspace service for memory-mapping datastores. Furthermore, through decoupled queue management, concurrency aware adaptation, and dynamic load balancing, UMap enables application performance to scale even at high concurrency. We evaluate UMap in data-intensive applications, including sorting, graph traversal, database operations, and metagenomic analytics. Our results show that UMap as a userspace service outperforms an optimized kernel-based service across a wide range of intra-node concurrency by 1.22-1.9 × . We performed two case studies to demonstrate UMap 's extensibility. First, a new datastore residing in remote memory is incorporated into UMap as an application-specific plugin. Second, we present a persistent memory allocator Metall built atop UMap for unified storage/memory.

97 MATHEMATICS AND COMPUTING↗

Symmetry-protected self-correcting quantum memory in three space dimensions

Whether self-correcting quantum memories can exist at nonzero temperature in a physically reasonable setting remains a great open problem. Furthermore, it has recently been argued that symmetry-protected topological (SPT) systems in three space dimensions subject to a strong constraint—that the quantum dynamics respect a 1-form symmetry—realize such a quantum memory. We illustrate how this works in Walker-Wang codes, which provide a specific realization of these desiderata. In this setting we show that it is sufficient for the 1-form symmetry to be enforced on a subvolume of the system. This strongly suggests that the SPT character of the state is not essential. We confirm this by constructing an explicit example with a trivial (paramagnetic) bulk that realizes a self-correcting quantum memory. We therefore show that the enforcement of a 1-form symmetry on a measure-zero subvolume of a three-dimensional system can be sufficient to stabilize a self-correcting quantum memory at nonzero temperature.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

A MultiGPU Performance-Portable Solution for Array Programming Based on Kokkos

Today, multiGPU nodes are widely used in high-performance computing and data centers. However, current programming models do not provide simple, transparent, and portable support for automatically targeting multiple GPUs within a node on application areas of array programming. In this paper, we describe a new application programming interface based on the Kokkos programming model to enable array computation on multiple GPUs in a transparent and portable way across both NVIDIA and AMD GPUs. We implement different variations of this technique to accommodate the exchange of stencils (array boundaries) among different GPU memory spaces, and we provide autotuning to select the proper number of GPUs, depending on the computational cost of the operations to be computed on arrays, that is completely transparent to the programmer. We evaluate our multiGPU extension on Summit (#5 TOP500), with six NVIDIA V100 Volta GPUs per node, and Crusher that contains identical hardware/software as Frontier (#1 TOP500), with four AMD MI250X GPUs, each with 2 Graphics Compute Dies (GCDs)for a total of 8 GCDs per node. We also compare the performance of this solution against the use of MPI + Kokkos, which is the cur-rent de facto solution for multiple GPUs in Kokkos. Our evaluation shows that the new Kokkos solution provides good scalability for many GPUs and a faster and simpler solution (from a programming productivity perspective) than MPI + Kokkos.

Valero Lara, Pedro↗

Homomorphic data compression for real time photon correlation analysis

The construction of highly coherent X-ray sources, combined with next-generation detectors that are larger and faster, has enabled new research opportunities across the scientific landscape. Among the techniques that benefit most from these advancements is X-ray photon correlation spectroscopy (XPCS), where faster acquisition unlocks the ability to study faster dynamics within samples. However, faster acquisition on larger detectors also introduces unprecedented challenges for online data processing and offline data storage. Such challenges are particularly prominent for XPCS, where real time analyses require simultaneous calculation of all the previously acquired data in the time series. We present a homomorphic compression scheme to effectively reduce the computational time and memory space required for XPCS analysis. Leveraging similarities in the mathematical expression between a matrix-based compression algorithm and the correlation calculation, our approach allows direct operation on the compressed data without their decompression. The offline compression scheme extends storage capacity by a factor of 40 while preserving key features in the lossy compressed data. Meanwhile, the online compression scheme reduces the computational time to below 1 ms, enabling real time calculation of the correlation functions at kHz framerate. Our demonstration of a homomorphic compression of scientific data provides an effective solution to the big data challenge at coherent light sources. Beyond the example shown in this work, the framework can be extended to facilitate real-time operations directly on a compressed data stream for other techniques.

36 MATERIALS SCIENCE↗

UNITY: Unified Memory and Storage Space

UNITY is a 36-month project focused on providing design and evaluate a new distributed storage paradigm that unifies the traditionally distinct application views of memory- and file-based data storage into a single scalable and resilient environment. The project is a collaboration among Oak Ridge National Laboratory (ORNL, Lead institution), Los Alamos National Labs (LANL), and Georgia Tech (GT). The main contributions of the GT team have been around development of low level systems software for best leveraging the capabilities of new types of persistent memory technologies, and for development of methods for intelligent data management across memory/storage substrates with heterogeneous components. GT contributed the Phoenix library for optimized checkpoint/restart for HPC I/O for systems with non-volatile memory (NVM), the NVStream library for NVM-specialized streaming I/O for HPC workflows, the CoMerge, Mnemo and Kleio solutions for intelligent data management on NVM-based systems. These contributions result in significant improvements in both application performance and system efficiency.

97 MATHEMATICS AND COMPUTING↗

Soft error-mitigating semiconductor design system and associated methods

A soft error-mitigating semiconductor design system and associated methods that tailor circuit design steps to mitigate corruption of data in storage elements (e.g., flip flops) due to Single Events Effects (SEEs). Required storage elements are automatically mapped to triplicated redundant nodes controlled by a voting element that enforces majority-voting logic for fault-free output (i.e., Triple Modular Redundancy (TMR)). Storage elements are also optimally positioned for placement in keeping with SEE-tolerant spacing constraints. Additionally, clock delay insertion (employing either a single global clock or clock triplication) in the TMR specification may introduce useful skew that protects against glitch propagation through the designed device. The resultant layout generated from the TMR configuration may relax constraints imposed on register transfer level (RTL) engineers to make rad-hard designs, as automation introduces TMR storage registers, memory element spacing, and clock delay/triplication with minimal designer input.

Miryala, Sandeep↗

Soft error-mitigating semiconductor design system and associated methods

A soft error-mitigating semiconductor design system and associated methods that tailor circuit design steps to mitigate corruption of data in storage elements (e.g., flip flops) due to Single Events Effects (SEEs). Required storage elements are automatically mapped to triplicated redundant nodes controlled by a voting element that enforces majority-voting logic for fault-free output (i.e., Triple Modular Redundancy (TMR)). Storage elements are also optimally positioned for placement in keeping with SEE-tolerant spacing constraints. Additionally, clock delay insertion (employing either a single global clock or clock triplication) in the TMR specification may introduce useful skew that protects against glitch propagation through the designed device. The resultant layout generated from the TMR configuration may relax constraints imposed on register transfer level (RTL) engineers to make rad-hard designs, as automation introduces TMR storage registers, memory element spacing, and clock delay/triplication with minimal designer input.

Miryala, Sandeep↗

GAHLS: an optimized graph analytics based high level synthesis framework

The urgent need for low latency, high-compute and low power on-board intelligence in autonomous systems, cyber-physical systems, robotics, edge computing, evolvable computing, and complex data science calls for determining the optimal amount and type of specialized hardware together with reconfigurability capabilities. With these goals in mind, we propose a novel comprehensive graph analytics based high level synthesis (GAHLS) framework that efficiently analyzes complex high level programs through a combined compiler-based approach and graph theoretic optimization and synthesizes them into message passing domain-specific accelerators. This GAHLS framework first constructs a compiler-assisted dependency graph (CaDG) from low level virtual machine (LLVM) intermediate representation (IR) of high level programs and converts it into a hardware friendly description representation. Next, the GAHLS framework performs a memory design space exploration while account for the identified computational properties from the CaDG and optimizing the system performance for higher bandwidth. The GAHLS framework also performs a robust optimization to identify the CaDG subgraphs with similar computational structures and aggregate them into intelligent processing clusters in order to optimize the usage of underlying hardware resources. Finally, the GAHLS framework synthesizes this compressed specialized CaDG into processing elements while optimizing the system performance and area metrics. Evaluations of the GAHLS framework on several real-life applications (e.g., deep learning, brain machine interfaces) demonstrate that it provides 14.27× performance improvements compared to state-of-the-art approaches such as LegUp 6.2.

97 MATHEMATICS AND COMPUTING↗

On-the-fly response function generation method for composite coarse mesh

The hybrid stochastic deterministic transport code COMET, based on the incident response expansion theory, is used to model reactor cores with high fidelity and formidable computational speed. COMET models a reactor core using a library of incident flux response expansion coefficients that are pre computed for all the unique lattice cells (e.g., fuel assemblies, reflector blocks, etc.) in the core. In order to further improve its computational efficiency in pre-calculating the response library a new response function generation method is developed to compute the response functions for the composite coarse meshes made of a smaller set of unique lattices on the fly within the COMET's deterministic transport core sweep. The efficiency is achieved by eliminating a number of unique lattices that can be made up from the reduced set of unique meshes on the fly. The numerical process consists of the following steps. First, the boundary condition on composite coarse mesh boundaries is projected onto the expansion basis to compute the incident flux moments on external surfaces of all the basic (reduced set of unique) coarse meshes. Secondly, the deterministic sweeping solver in COMET is used to converge on the outgoing/incoming flux expansion moments crossing interfaces between the basic coarse meshes. Thirdly, the response functions for the composite coarse meshes are constructed as a superposition on the fly. The new response function generation method was tested on 88 composite coarse meshes consisting of CANDU fuel bundles and moderator blocks. It was found that response functions generated by the new method agree very well with those generated by direct Monte Carlo calculations. The average and maximum relative differences in the surface-to-surface response coefficients computed by the two methods are 0.10% and 0.20%, respectively. Similarly, the average and maximum relative differences in the response fission densities are 0.13% and 0.43%, respectively. These discrepancies are within one standard deviation of the stochastic uncertainties. The new method is five times faster than the original direct Monte Carlo method. The size of the response function library for the new method is five times smaller than that for the original method, leading to significantly less requirement for the computer hard drive space and memory. (authors)

21 SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLANTS↗

Frequency and time multiplexing for multiphoton state generation

One of the primary challenges with photonic quantum information processing is that single-photon states are created and heralded probabilistically, rather than being created on demand. One solution to this problem is multiplexing , which is attempting probabilistic generation of a single photon at multiple locations, times, or frequencies, and switching one successfully generated photon into an output mode. Motivated by the development of frequency-based photonic quantum information processing, we consider multiplexing using frequency and time as our options to use to get many photon generation opportunities. We devise an approach for generating multiphoton states, with photons populating multiple frequency modes in the same spatiotemporal mode. This method uses a variable-length optical delays to manipulate the temporal mode of the photons, and spaced fiber Bragg grating (FBG) reflectors to jointly manipulate the frequency and temporal modes of the photons. Experimental progress toward implementing this multiplexing scheme is proceeding along 2 fronts, each with different ways to achieve the variable-length optical delay. First, we will implement this scheme using a free-space optical quantum memory with multiple discretely adjustable free-space delays. For this implementation method, I calculate multiphoton generation rates, accounting for loss, that are realistically achievable with commercially available hardware. This work will appear in a theory paper, currently in preparation. Second, we will implement this multiplexing scheme in an integrated manner: the variable-length optical delay will happen on-chip using a long-lifetime Q-switchable Fabry-Perot cavity. The joint manipulation of the frequency and temporal modes of the photons will still happen off-chip in fiber with FBG reflectors. I report on experimental progress in constructing integrated Q-switchable Fabry-Perot cavities.

97 MATHEMATICS AND COMPUTING↗