Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “DRAM”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Understanding Strong Scaling on GPUs Using Empirical Performance Saturation Size

The roofline model provides a concise overview of the maximum performance capabilities of a given computer system through a combination of peak memory bandwidth and compute performance rates. The increasing complexity of scheduling and cache in recent GPUs, however, has introduced complicated performance variability that is not captured by arithmetic intensity alone. This work examines the effect of problem size and GPU launch configurations on roofline performance for V100, A100, MI100, and MI250X graphics processing units. We introduce an extended roofline model that takes problem size into account, and find that strong scaling on GPUs can be characterized by saturation problem sizes as additional key metrics. Saturation problem sizes break up a plot of GPU performance vs. problem size into three distinct performance regimes– size-limited, cache-bound, and DRAM-bound. With our extended roofline model, we are able to provide a robust view of these performance regimes across recent GPU architectures.

Eberius, David↗

Flexible and Effective Object Tiering for Heterogeneous Memory Systems

Computing platforms that package multiple types of memory, each with their own performance characteristics, are quickly becoming mainstream. To operate efficiently, heterogeneous memory architectures require new data management solutions that are able to match the needs of each application with an appropriate type of memory. As the primary generators of memory usage, applications create a great deal of information that can be useful for guiding memory tiering, but the community still lacks tools to collect, organize, and leverage this information effectively. To address this gap, this work introduces a novel software framework that collects and analyzes object-level information to guide memory tiering. Using this framework, this study evaluates and compares the impact of a variety of data tiering choices, including how the system prioritizes objects for faster memory as well as the frequency and timing of migration events. The results, collected on a modern Intel platform with conventional DRAM as well as non-volatile RAM, show that guiding data tiering with object-level information can enable significant performance and efficiency benefits compared to standard hardware- and software-directed data tiering strategies.

Kammerdiener, Brandon↗

cuAlign: Scalable Network Alignment on GPU Accelerators

Given two graphs, the objective of network alignment is to find the best one-to-one mapping of vertices in one graph (??) to vertices in the other (??), such that the number of overlaps is maximized. We say that edges(??, ??) ???and(??', ??') ??? are overlapped if ?? is mapped to ??' and ?? is mapped to??'. Network alignment is an important optimization problem with several applications in bioinformatics, computer vision and ontology matching. Since it is an NP-hard problem, efficient heuristics and scalable implementations are necessary. In this work, we introduce a new framework that combines the concepts of intra-network proximity using vertex embedding,Belief Propagation (BP) and approximate weighted matching, and provides qualitative improvements up to22%over state-of-the-art approaches. We also provide scalable implementations on GPU accelerators, demonstrating up to19×speedup for Belief Propagation and 3× speedup for approximate weighted matching relative to previous multithreaded implementation. A combination of combinatorial and algebraic kernels within the network alignment algorithm poses significant hurdles for parallelization. Load imbalance and irregular DRAM traffic limit achievable performance on GPUs. Our novel approach identifies and exploits unique structural proper-ties of the BP-based algorithm and employs code fusion to reduce data movement between different steps of the algorithm. Using a diverse set of inputs, we demonstrate qualitative improvements of our algorithms, and performance gains of our GPU-accelerated implementation. We believe that our work will enable algorithmic improvements and practical applications of network alignment.

Xiang, Lizhi↗

Vector-Matrix Multiplication Engine for Neuromorphic Computation with a CBRAM Crossbar Array [Slides]

The core function of many neural network algorithms is the dot product, or vector matrix multiply (VMM) operation. Crossbar arrays utilizing resistive memory elements can reduce computational energy in neural algorithms by up to five orders of magnitude compared to conventional CPUs. Moving data between a processor, SRAM, and DRAM dominates energy consumption. By utilizing analog operations to reduce data movement, resistive memory crossbars can enable processing of large amounts of data at lower energy than conventional memory architectures.

97 MATHEMATICS AND COMPUTING↗

Chemistry of Titanium Deposition Precursors for Area-Selective Deposition of Functionalized Silicon [Posters]

Area-selective atomic layer deposition (AS-ALD) is an appealing bottom-up fabrication technique that can produce atomic-scale device features, overcoming challenges in current industrial techniques such as edge alignment errors. TiCI 4 is a common thermal ALD precursor for Ti0 2 thin films, which are appealing candidates for DRAM capacitors due to their excellent dielectric constants. Hydrogen and chlorine termination passivate the Si surface, allowing for selective deposition of TiCI 4 onto HO-terminated areas. However, selectivity loss occurs after several ALD cycles. Ti oxide nucleates onto surface defects on Cl- and H-Si resists. Previously, the use of H-Si as an ALD resist has been studied extensively, but less work has focused on chemical forces driving nucleation, especially for Cl-Si. Here, formation of defect nuclei was investigated with selectivity loss during Ti0 2 ALD with TiCI 4 and water on the (100) and (111) crystal surfaces of hydrogenated, chlorinated, and oxidized Si.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

IceNet for FireBox - A Berkeley Warehouse-Scale Computer

Berkeley’s FireBox is a next-generation warehouse-scale computer (WSC) that utilizes the energy-efficiency and bandwidth density of integrated silicon-photonic interconnects to enable a new high-bandwidth and low-latency network fabric connecting thousands of compute nodes to petabytes of DRAM and Flash storage. The high bandwidth, low latency and high connectivity of FireBox’s WSC network fabric (IceNet) will enable dramatic improvements in the overall system energy efficiency enabling fine-grain power control on system resources (processors, links and memory/storage components). IceNet is a special 3-stage photonic Clos network architected to achieve ultra-low-latency connectivity between processor nodes and memory, drastically cutting down on the energy wasted in resource idling (processors and memory stalled due to pending network requests). This is achieved by integration of the first and last switch stages into processor/memory hub clients and by heavy over-provisioning of the high-radix middle switches (FlareSwitches). A key hardware component developed in this program is an active laser power management photonic integrated circuits called LightSpark. It interacts with the FlareSwitch and provides laser power to a subset of occupied switch ports, increasing the utilization of laser light in the photonic network by an order of magnitude. In addition to guiding the laser power where it is needed, the laser-power management module enables both wavelength and laser redundancy, significantly increasing the robustness of the system. The goal of the IceNet fabric is to enable communication between 1000s of processor nodes and PBs of memory/storage with <100ns latency, <10pJ/b wall-plug energy cost at multiple Pb/s of available connectivity bandwidth. These metrics represent two-orders of magnitude improvement with respect to the status of current data-center technology.

42 ENGINEERING↗

Finding the missing pieces: filling gaps that impede the translation of omics data into models

High-throughput omics technologies such as DNA sequencing have made the sequencing and computational assembly of microbial genomes recovered from the environment relatively routine. Computational inference of the protein products encoded by these genomes, and the associated biochemical functions, should enable the accurate prediction and modeling of microbial metabolism, organismal interactions, and ecosystem processes. However, a lack of scalable, probabilistic protein annotation tools limits the full potential of modeling for understanding the metabolism and biogeochemical cycles of microbial communities. Our approach to improve inference of protein annotations and metabolic models relied on learning from and emulating expert manual curation, leveraging software engineering and data science best practices to scale up the throughput and accuracy of annotations and metabolic model construction, building software to objectively evaluate different annotation strategies, and more closely linking the protein annotation and metabolic model inference process. Outcomes of this research include several improved or new computational tools, including DRAM (Distilled and Refined Annotation of Metabolism) for annotating microbial genomes with protein function and metabolic traits, CAMPER (Curated Annotations for Microbial Polyphenol Enzymes and Reactions) for annotating key polyphenol metabolisms, EC-Bench for comprehensive and unbiased benchmarking of annotation tools, and several apps available via the DOE Systems Biology Knowledgebase (KBase) for building genome-scale metabolic models. We demonstrate that these tools allow us to scalably annotate and understand thousands of genomes for microbial communities from a variety of systems and test cases, including rivers, thawing permafrost, and gut microbiomes. All of these computational tools are available as open-source software, with most broadly and easily accessible to the scientific community via KBase apps.

59 BASIC BIOLOGICAL SCIENCES↗

Luteolibacter sp. strain Populi

Luteolibacter sp. strain Populi is bacterium from the phylum Verrucomicrobiota, isolated from the rhizosphere of a black cottonwood tree, Populus trichocarpa, from the Cascade mountains in Washington. Its 6.6 Mb chromosome was completely sequenced using Oxford Nanopore long-reads and is predicted to encode 5301 proteins and 60 RNAs. The bacteria was isolated from the rhizosphere of a mature Populus trichocarpa from the Tieton riverwatershed of Washington state, USA (Lat: 46°42’9” N, Lon: 120°25 39’36” W). A rhizosphere sample (fine roots and adhering soil) was used to obtain a microbial fraction by centrifugation on Histodenz (12) and stained with 5µM Syto59 (Thermo Fisher Scientific Inc). A Cytopeia Influx cell sorter (BD, Franklin Lakes, NJ) was used to sort and array single cells (100 per plate) based on forward-side scatter and fluorescence intensity on asparagine-glucose nutrient agar (ATCC medium 184). The Luteolibacter sp. Populi genome sequence has been deposited in GenBank under the accession number CP161812. A draft genome annotated with Prokka and DRAM is available in this Narrative as Luteolibacter_sp_Prokka.240711.

59 BASIC BIOLOGICAL SCIENCES↗

Genome collection processing for “Conserved upper thermal limits and small safety margins in soil copiotrophic bacteria”

We extracted the genomic DNA of 400 randomly selected isolates using a Quick-DNA Microprep Kit (Zymo Research D3020) according to the manufacturer’s protocol. We then submitted the extracted gDNA samples for short-read Illumina sequencing (200 Mbp) at SeqCoast Genomics (Portsmouth, NH, USA). After preprocessing the sequences using Trimmommatic (Bolger et al. 2014), we assembled the genomes using SPADES (Bankevich et al. 2012) and checked the quality of each assembly using QUAST (Gurevich et al. 2013). We processed the genome assemblies using a KBase (v1.4.0) pipeline (Allen et al. 2017; Arkin et al. 2018). Briefly, we used DRAM (v0.1.2) with default settings to annotate the genome assemblies. We then evaluated genome quality and possible contamination levels using CheckM (v1.0.18) (Parks et al. 2015) and retained genomes with completeness above 98% and contamination below 5% (n = 354), following the authors' guidelines. We then obtained taxonomic assignments for all remaining isolates using the Genome Taxonomy Database tool GTDB-Tk (v2.3.2, database version r214) (Chaumeil et al. 2019). We constructed a phylogenetic tree using the tool SpeciesTree (v2.2.0). We then trimmed the tree (using Trim SpeciesTree to GenomeSet- v1.4.0), retaining only tips within our collection with measured thermal performance.

59 BASIC BIOLOGICAL SCIENCES↗

On the Routh approximation technique and least squares errors

A new method for calculating the coefficients of the numerator polynomial of the direct Routh approximation method (DRAM) using the least square error criterion is formulated. The necessary conditions have been obtained in terms of algebraic equations. The method is useful for low frequency as well as high frequency reduced-order models.

Aburdene, M. F.↗

Reed Solomon codes for error control in byte organized computer memory systems

A problem in designing semiconductor memories is to provide some measure of error control without requiring excessive coding overhead or decoding time. In LSI and VLSI technology, memories are often organized on a multiple bit (or byte) per chip basis. For example, some 256K-bit DRAM's are organized in 32Kx8 bit-bytes. Byte oriented codes such as Reed Solomon (RS) codes can provide efficient low overhead error control for such memories. However, the standard iterative algorithm for decoding RS codes is too slow for these applications. Some special decoding techniques for extended single-and-double-error-correcting RS codes which are capable of high speed operation are presented. These techniques are designed to find the error locations and the error values directly from the syndrome without having to use the iterative algorithm to find the error locator polynomial.

Lin, S.↗

Error control for reliable digital data transmission and storage systems

A problem in designing semiconductor memories is to provide some measure of error control without requiring excessive coding overhead or decoding time. In LSI and VLSI technology, memories are often organized on a multiple bit (or byte) per chip basis. For example, some 256K-bit DRAM's are organized in 32Kx8 bit-bytes. Byte oriented codes such as Reed Solomon (RS) codes can provide efficient low overhead error control for such memories. However, the standard iterative algorithm for decoding RS codes is too slow for these applications. In this paper we present some special decoding techniques for extended single-and-double-error-correcting RS codes which are capable of high speed operation. These techniques are designed to find the error locations and the error values directly from the syndrome without having to use the iterative alorithm to find the error locator polynomial. Two codes are considered: (1) a d sub min = 4 single-byte-error-correcting (SBEC), double-byte-error-detecting (DBED) RS code; and (2) a d sub min = 6 double-byte-error-correcting (DBEC), triple-byte-error-detecting (TBED) RS code.

Costello, D. J., Jr.↗

Decoding of DBEC-TBED Reed-Solomon codes

A problem in designing semiconductor memories is to provide some measure of error control without requiring excessive coding overhead or decoding time. In LSI and VLSI technology, memories are often organized on a multiple bit (or byte) per chip basis. For example, some 256 K bit DRAM's are organized in 32 K x 8 bit-bytes. Byte-oriented codes such as Reed-Solomon (RS) codes can provide efficient low overhead error control for such memories. However, the standard iterative algorithm for decoding RS codes is too slow for these applications. The paper presents a special decoding technique for double-byte-error-correcting, triple-byte-error-detecting RS codes which is capable of high-speed operation. This technique is designed to find the error locations and the error values directly from the syndrome without having to use the iterative algorithm to find the error locator polynomial.

Deng, Robert H.↗

Spread Of Charge From Ion Tracks In Integrated Circuits

Single-event upsets (SEU's) propagate to adjacent cells in integrated memory circuits. Findings of experiments in lateral transport of electrical-charge carriers from ion tracks in 256K dynamic randon-access memories (DRAM's). As dimensions of integrated circuits decrease, vulnerability to SEU's increases. Understanding gained enables design of less vulnerable circuits.

Zoutendyk, John A.↗

Lateral charge transport from heavy-ion tracks in integrated circuit chips

A 256K DRAM has been used to study the lateral transport of charge (electron-hole pairs) induced by direct ionization from heavy-ion tracks in an IC. The qualitative charge transport has been simulated using a two-dimensional numerical code in cylindrical coordinates. The experimental bit-map data clearly show the manifestation of lateral charge transport in the creation of adjacent multiple-bit errors from a single heavy-ion track. The heavy-ion data further demonstrate the occurrence of multiple-bit errors from single ion tracks with sufficient stopping power. The qualitative numerical simulation results suggest that electric-field-funnel-aided (drift) collection accounts for single error generated by an ion passing through a charge-collecting junction, while multiple errors from a single ion track are due to lateral diffusion of ion-generated charge.

Zoutendyk, J. A.↗

Characterization of multiple-bit errors from single-ion tracks in integrated circuits

The spread of charge induced by an ion track in an integrated circuit and its subsequent collection at sensitive nodal junctions can cause multiple-bit errors. The authors have experimentally and analytically investigated this phenomenon using a 256-kb dynamic random-access memory (DRAM). The effects of different charge-transport mechanisms are illustrated, and two classes of ion-track multiple-bit error clusters are identified. It is demonstrated that ion tracks that hit a junction can affect the lateral spread of charge, depending on the nature of the pull-up load on the junction being hit. Ion tracks that do not hit a junction allow the nearly uninhibited lateral spread of charge.

Zoutendyk, J. A.↗

Multiple-Bit Errors Caused By Single Ions

Report describes experimental and computer-simulation study of multiple-bit errors caused by impingement of single energetic ions on 256-Kb dynamic random-access memory (DRAM) integrated circuit. Studies illustrate effects of different mechanisms for transport of charge from ion tracks to various elements of integrated circuits. Shows multiple-bit errors occur in two different types of clusters about ion tracks causing them.

Zoutendyk, John A.↗

Non-volatile, high density, high speed, Micromagnet-Hall effect Random Access Memory (MHRAM)

The micromagnetic Hall effect random access memory (MHRAM) has the potential of replacing ROMs, EPROMs, EEPROMs, and SRAMs because of its ability to achieve non-volatility, radiation hardness, high density, and fast access times, simultaneously. Information is stored magnetically in small magnetic elements (micromagnets), allowing unlimited data retention time, unlimited numbers of rewrite cycles, and inherent radiation hardness and SEU immunity, making the MHRAM suitable for ground based as well as spaceflight applications. The MHRAM device design is not affected by areal property fluctuations in the micromagnet, so high operating margins and high yield can be achieved in large scale integrated circuit (IC) fabrication. The MHRAM has short access times (less than 100 nsec). Write access time is short because on-chip transistors are used to gate current quickly, and magnetization reversal in the micromagnet can occur in a matter of a few nanoseconds. Read access time is short because the high electron mobility sensor (InAs or InSb) produces a large signal voltage in response to the fringing magnetic field from the micromagnet. High storage density is achieved since a unit cell consists only of two transistors and one micromagnet Hall effect element. By comparison, a DRAM unit cell has one transistor and one capacitor, and a SRAM unit cell has six transistors.

Wu, Jiin C.↗