Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “DRAM”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

30 records · Page 2

Understanding Strong Scaling on GPUs Using Empirical Performance Saturation Size

The roofline model provides a concise overview of the maximum performance capabilities of a given computer system through a combination of peak memory bandwidth and compute performance rates. The increasing complexity of scheduling and cache in recent GPUs, however, has introduced complicated performance variability that is not captured by arithmetic intensity alone. This work examines the effect of problem size and GPU launch configurations on roofline performance for V100, A100, MI100, and MI250X graphics processing units. We introduce an extended roofline model that takes problem size into account, and find that strong scaling on GPUs can be characterized by saturation problem sizes as additional key metrics. Saturation problem sizes break up a plot of GPU performance vs. problem size into three distinct performance regimes– size-limited, cache-bound, and DRAM-bound. With our extended roofline model, we are able to provide a robust view of these performance regimes across recent GPU architectures.

Eberius, David↗

Flexible and Effective Object Tiering for Heterogeneous Memory Systems

Computing platforms that package multiple types of memory, each with their own performance characteristics, are quickly becoming mainstream. To operate efficiently, heterogeneous memory architectures require new data management solutions that are able to match the needs of each application with an appropriate type of memory. As the primary generators of memory usage, applications create a great deal of information that can be useful for guiding memory tiering, but the community still lacks tools to collect, organize, and leverage this information effectively. To address this gap, this work introduces a novel software framework that collects and analyzes object-level information to guide memory tiering. Using this framework, this study evaluates and compares the impact of a variety of data tiering choices, including how the system prioritizes objects for faster memory as well as the frequency and timing of migration events. The results, collected on a modern Intel platform with conventional DRAM as well as non-volatile RAM, show that guiding data tiering with object-level information can enable significant performance and efficiency benefits compared to standard hardware- and software-directed data tiering strategies.

Kammerdiener, Brandon↗

cuAlign: Scalable Network Alignment on GPU Accelerators

Given two graphs, the objective of network alignment is to find the best one-to-one mapping of vertices in one graph (??) to vertices in the other (??), such that the number of overlaps is maximized. We say that edges(??, ??) ???and(??', ??') ??? are overlapped if ?? is mapped to ??' and ?? is mapped to??'. Network alignment is an important optimization problem with several applications in bioinformatics, computer vision and ontology matching. Since it is an NP-hard problem, efficient heuristics and scalable implementations are necessary. In this work, we introduce a new framework that combines the concepts of intra-network proximity using vertex embedding,Belief Propagation (BP) and approximate weighted matching, and provides qualitative improvements up to22%over state-of-the-art approaches. We also provide scalable implementations on GPU accelerators, demonstrating up to19×speedup for Belief Propagation and 3× speedup for approximate weighted matching relative to previous multithreaded implementation. A combination of combinatorial and algebraic kernels within the network alignment algorithm poses significant hurdles for parallelization. Load imbalance and irregular DRAM traffic limit achievable performance on GPUs. Our novel approach identifies and exploits unique structural proper-ties of the BP-based algorithm and employs code fusion to reduce data movement between different steps of the algorithm. Using a diverse set of inputs, we demonstrate qualitative improvements of our algorithms, and performance gains of our GPU-accelerated implementation. We believe that our work will enable algorithmic improvements and practical applications of network alignment.

Xiang, Lizhi↗

Vector-Matrix Multiplication Engine for Neuromorphic Computation with a CBRAM Crossbar Array [Slides]

The core function of many neural network algorithms is the dot product, or vector matrix multiply (VMM) operation. Crossbar arrays utilizing resistive memory elements can reduce computational energy in neural algorithms by up to five orders of magnitude compared to conventional CPUs. Moving data between a processor, SRAM, and DRAM dominates energy consumption. By utilizing analog operations to reduce data movement, resistive memory crossbars can enable processing of large amounts of data at lower energy than conventional memory architectures.

97 MATHEMATICS AND COMPUTING↗

Chemistry of Titanium Deposition Precursors for Area-Selective Deposition of Functionalized Silicon [Posters]

Area-selective atomic layer deposition (AS-ALD) is an appealing bottom-up fabrication technique that can produce atomic-scale device features, overcoming challenges in current industrial techniques such as edge alignment errors. TiCI 4 is a common thermal ALD precursor for Ti0 2 thin films, which are appealing candidates for DRAM capacitors due to their excellent dielectric constants. Hydrogen and chlorine termination passivate the Si surface, allowing for selective deposition of TiCI 4 onto HO-terminated areas. However, selectivity loss occurs after several ALD cycles. Ti oxide nucleates onto surface defects on Cl- and H-Si resists. Previously, the use of H-Si as an ALD resist has been studied extensively, but less work has focused on chemical forces driving nucleation, especially for Cl-Si. Here, formation of defect nuclei was investigated with selectivity loss during Ti0 2 ALD with TiCI 4 and water on the (100) and (111) crystal surfaces of hydrogenated, chlorinated, and oxidized Si.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

IceNet for FireBox - A Berkeley Warehouse-Scale Computer

Berkeley’s FireBox is a next-generation warehouse-scale computer (WSC) that utilizes the energy-efficiency and bandwidth density of integrated silicon-photonic interconnects to enable a new high-bandwidth and low-latency network fabric connecting thousands of compute nodes to petabytes of DRAM and Flash storage. The high bandwidth, low latency and high connectivity of FireBox’s WSC network fabric (IceNet) will enable dramatic improvements in the overall system energy efficiency enabling fine-grain power control on system resources (processors, links and memory/storage components). IceNet is a special 3-stage photonic Clos network architected to achieve ultra-low-latency connectivity between processor nodes and memory, drastically cutting down on the energy wasted in resource idling (processors and memory stalled due to pending network requests). This is achieved by integration of the first and last switch stages into processor/memory hub clients and by heavy over-provisioning of the high-radix middle switches (FlareSwitches). A key hardware component developed in this program is an active laser power management photonic integrated circuits called LightSpark. It interacts with the FlareSwitch and provides laser power to a subset of occupied switch ports, increasing the utilization of laser light in the photonic network by an order of magnitude. In addition to guiding the laser power where it is needed, the laser-power management module enables both wavelength and laser redundancy, significantly increasing the robustness of the system. The goal of the IceNet fabric is to enable communication between 1000s of processor nodes and PBs of memory/storage with <100ns latency, <10pJ/b wall-plug energy cost at multiple Pb/s of available connectivity bandwidth. These metrics represent two-orders of magnitude improvement with respect to the status of current data-center technology.

42 ENGINEERING↗

Finding the missing pieces: filling gaps that impede the translation of omics data into models

High-throughput omics technologies such as DNA sequencing have made the sequencing and computational assembly of microbial genomes recovered from the environment relatively routine. Computational inference of the protein products encoded by these genomes, and the associated biochemical functions, should enable the accurate prediction and modeling of microbial metabolism, organismal interactions, and ecosystem processes. However, a lack of scalable, probabilistic protein annotation tools limits the full potential of modeling for understanding the metabolism and biogeochemical cycles of microbial communities. Our approach to improve inference of protein annotations and metabolic models relied on learning from and emulating expert manual curation, leveraging software engineering and data science best practices to scale up the throughput and accuracy of annotations and metabolic model construction, building software to objectively evaluate different annotation strategies, and more closely linking the protein annotation and metabolic model inference process. Outcomes of this research include several improved or new computational tools, including DRAM (Distilled and Refined Annotation of Metabolism) for annotating microbial genomes with protein function and metabolic traits, CAMPER (Curated Annotations for Microbial Polyphenol Enzymes and Reactions) for annotating key polyphenol metabolisms, EC-Bench for comprehensive and unbiased benchmarking of annotation tools, and several apps available via the DOE Systems Biology Knowledgebase (KBase) for building genome-scale metabolic models. We demonstrate that these tools allow us to scalably annotate and understand thousands of genomes for microbial communities from a variety of systems and test cases, including rivers, thawing permafrost, and gut microbiomes. All of these computational tools are available as open-source software, with most broadly and easily accessible to the scientific community via KBase apps.

59 BASIC BIOLOGICAL SCIENCES↗

Luteolibacter sp. strain Populi

Luteolibacter sp. strain Populi is bacterium from the phylum Verrucomicrobiota, isolated from the rhizosphere of a black cottonwood tree, Populus trichocarpa, from the Cascade mountains in Washington. Its 6.6 Mb chromosome was completely sequenced using Oxford Nanopore long-reads and is predicted to encode 5301 proteins and 60 RNAs. The bacteria was isolated from the rhizosphere of a mature Populus trichocarpa from the Tieton riverwatershed of Washington state, USA (Lat: 46°42’9” N, Lon: 120°25 39’36” W). A rhizosphere sample (fine roots and adhering soil) was used to obtain a microbial fraction by centrifugation on Histodenz (12) and stained with 5µM Syto59 (Thermo Fisher Scientific Inc). A Cytopeia Influx cell sorter (BD, Franklin Lakes, NJ) was used to sort and array single cells (100 per plate) based on forward-side scatter and fluorescence intensity on asparagine-glucose nutrient agar (ATCC medium 184). The Luteolibacter sp. Populi genome sequence has been deposited in GenBank under the accession number CP161812. A draft genome annotated with Prokka and DRAM is available in this Narrative as Luteolibacter_sp_Prokka.240711.

59 BASIC BIOLOGICAL SCIENCES↗

Genome collection processing for “Conserved upper thermal limits and small safety margins in soil copiotrophic bacteria”

We extracted the genomic DNA of 400 randomly selected isolates using a Quick-DNA Microprep Kit (Zymo Research D3020) according to the manufacturer’s protocol. We then submitted the extracted gDNA samples for short-read Illumina sequencing (200 Mbp) at SeqCoast Genomics (Portsmouth, NH, USA). After preprocessing the sequences using Trimmommatic (Bolger et al. 2014), we assembled the genomes using SPADES (Bankevich et al. 2012) and checked the quality of each assembly using QUAST (Gurevich et al. 2013). We processed the genome assemblies using a KBase (v1.4.0) pipeline (Allen et al. 2017; Arkin et al. 2018). Briefly, we used DRAM (v0.1.2) with default settings to annotate the genome assemblies. We then evaluated genome quality and possible contamination levels using CheckM (v1.0.18) (Parks et al. 2015) and retained genomes with completeness above 98% and contamination below 5% (n = 354), following the authors' guidelines. We then obtained taxonomic assignments for all remaining isolates using the Genome Taxonomy Database tool GTDB-Tk (v2.3.2, database version r214) (Chaumeil et al. 2019). We constructed a phylogenetic tree using the tool SpeciesTree (v2.2.0). We then trimmed the tree (using Trim SpeciesTree to GenomeSet- v1.4.0), retaining only tips within our collection with measured thermal performance.

59 BASIC BIOLOGICAL SCIENCES↗

DENOVA: Deduplication Extended NOVA File System

This paper shows mathematically and experimentally that inline deduplication is not suitable for file systems on ultra-low latency Intel Optane DC PM devices in terms of performance, and proposes DeNova, an offline deduplication specially designed for log-structured NVM file systems such as NOVA. DeNova offers high-performance and low-latency I/O processing and executes deduplication in the background without interfering with foreground I/Os. DeNova employs DRAM-free persistent deduplication metadata, favoring CPU cache line, and ensures failure consistency on any system failure. We implement DeNova in the NOVA file system. Evaluation with DeNova confirms a negligible performance drop of baseline NOVA of less than 1%, while gaining high storage space savings. Extensive experiments show DeNova is failure consistent in all failure scenario cases.

Khan, Awais↗

Evaluating HPC Kernels for Processing in Memory

Memory subsystems contribute significantly to the performance and energy efficiency of high-performance computing (HPC) applications. Traditional memory technologies with conventional organization (e.g., DRAM) are struggling to keep up with the increasing memory requirements of modern applications. Techniques such as multilayer cache hierarchy and out-of-order execution are still falling short of mitigating the penalty incurred by memory accesses. Processing-in-memory (PIM), which involves moving memory-intensive kernels to memory for execution instead of bringing the data to the processing unit, is emerging as a promising technique. PIM has recently received traction among computer architecture researchers, and the increasing research activity surrounding this technique indicates its potential to alleviate main memory performance bottlenecks. In this paper, we characterize and identify memory-intensive HPC kernels, perform a first-order evaluation of the PIM technique for selected HPC kernels, quantify performance deviation, and analyze the key factors that affect PIM efficiency.

Asifuzzaman, Kazi↗

Architectural Implications of Neural Network Inference for High Data-Rate, Low-Latency Scientific Applications

With more scientific fields relying on neural networks (NNs) to process data incoming at extreme throughputs and latencies, it is crucial to develop NNs with all their parameters stored on-chip. In many of these applications, there is not enough time to go off-chip and retrieve weights. Even more so, off-chip memory such as DRAM does not have the bandwidth required to process these NNs as fast as the data is being produced (e.g., every 25 ns). As such, these extreme latency and bandwidth requirements have architectural implications for the hardware intended to run these NNs: 1) all NN parameters must fit on-chip, and 2) codesigning custom/reconfigurable logic is often required to meet these latency and bandwidth constraints. In our work, we show that many scientific NN applications must run fully on chip, in the extreme case requiring a custom chip to meet such stringent constraints.

Weng, Olivia↗