Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Memory systems”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11

Distributed Multi-GPU Community Detection on Exascale Computing Platforms

Community detection is a fundamental operation in graph mining, and by uncovering hidden structures and patterns within complex systems it helps solve fundamental problems pertaining to social networks, such as information diffusion, epidemics, and recommender systems. Scaling graph algorithms for massive networks becomes challenging on modern distributed-memory multi-GPU (Graphics Processing Unit) systems due to limitations such as irregular memory access patterns, load imbalances, higher communication-computation ratios, and cross-platform support. We present a novel algorithm HiPDPL-GPU (Distributed Parallel Louvain) to address these challenges. We conduct experiments involving different partitioning techniques to achieve an optimized performance of HiPDPL-GPU on the two largest supercomputers: Frontier and Summit. Remarkably, HiPDPL-GPU processes a graph with 4.2 billion edges in less than 3 minutes using 1024 GPUs. Qualitatively, the performance of HiPDPL-GPU is similar or better compared to other state-of-the-art CPU- and GPU-based implementations. While prior GPU implementations have predominantly employed CUDA, our first-of-its-kind implementation for community detection is cross-platform, accommodating both AMD and NVIDIA GPUs.

Sattar, Naw Safrin↗

A PIPS + SrI 2 (Eu) detector for atmospheric radioxenon monitoring

The PIPS–SrI 2 (Eu) is a prototype atmospheric radioxenon detection system designed at Oregon State University in support of international efforts towards monitoring clandestine nuclear weapon testing activities. This detector aims to address some shortcomings found in currently deployed beta–gamma atmospheric radioxenon detection systems, such as lackluster energy resolution and memory effect, by employing modern detection materials and readout. The system uses a PIPSBox, a silicon-based gas cell, for electron detection, and a pair of ultrabright, D-shaped SrI 2 (Eu) scintillators coupled to silicon photomultipliers for photon detection. A custom eight-channel digital pulse processor equipped with a field programmable gate-array (FPGA) identifies electron–photon coincidences between the volumes in near real-time. Gas samples of the four radioxenon isotopes of interest were independently measured with the PIPS–SrI 2 (Eu) detection system to determine energy resolution and efficiency. Application of FPGA-based coincidence discrimination in near real-time reduced the ambient background count rate by 95.85 ± 0.04%. Using parameters from the Xenon International gas processing unit and assuming a blank sample and zero memory effect the minimum detectable concentrations (MDCs) for the isotopes were calculated to be 0.12 ± 0.03, 0.27 ± 0.05, 0.15 ± 0.02, and 1.00 ± 0.08 mBq/m 3 air for 131m Xe, 133 Xe, 133m Xe, and 135 Xe, respectively. These MDC estimates compare well with other radioxenon detection systems employed in the International Monitoring System (IMS) and indicate that the PIPS–SrI 2 (Eu) is in compliance with the Comprehensive Nuclear Test-Ban-Treaty Organization (CTBTO) sensitivity requirement of ≤ 1 mBq/m 3 for 133 Xe.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Human Factors Design for Particle Accelerator Control Room Interfaces

Fermilab, the birthplace of many scientific discoveries in physics and particle accelerator sciences, is in the midst of a widescale modernization effort. The Accelerator Control Operations Research Network (ACORN project’s goal is to modernize the accelerator control system by replacing end-of-life power supplies and enhance future operations of the Fermilab accelerator complex with megawatt particle beams. Within ACORN, opportunities for process improvement concerning software development, human-system interface design, and task performance are also being considered. Human factors researchers from Idaho National Laboratory in collaboration with usability experts from Fermilab, are currently investigating human-centered design improvements for the accelerator control system. For example, substantial tribal knowledge and memory recall are required to effectively operate the accelerator system. This contributes to high cognitive workload and potential burnout of accelerator operators. Developing guidance for consistent visual and functional design enables a more intuitive interaction and relieves operators of cognitive burden. Additionally, developing more intuitive and integrated interfaces can also lead to improved accelerator efficacy by empowering operators with greater understanding and control of the systems. The challenge in developing such interfaces is in designing for a wide variety of user goals, system specifications, and level of experience in users. The challenges need to be met while e also considering the maintainability of the control system. The purpose of this paper is to detail the human factors process and design within the ACORN project, describe results gathered thus far, and discuss the larger implications for this work.

43 PARTICLE ACCELERATORS↗

Ab Initio-Based Bond Order Potential for Arsenene Polymorphs Developed via Hierarchical Reinforcement Learning

Arsenene, a less-explored two-dimensional material, holds the potential for applications in wearable electronics, memory devices, and quantum systems. This study introduces a bond-order potential model with Tersoff formalism, the ML-Tersoff, which leverages multireward hierarchical reinforcement learning (RL), trained on an ab initio data set. This data set covers a spectrum of properties for arsenene polymorphs, enhancing our understanding of its mechanical and thermal behaviors without the complexities of traditional models requiring multiple parameter sets. Our RL strategy utilizes decision trees coupled with a hierarchical reward strategy to accelerate convergence in high-dimensional continuous search spaces. Unlike the Stillinger-Weber approach, which demands separate formalisms for buckled and puckered forms, the ML-Tersoff model concurrently captures multiple properties of the two polymorphs by effectively representing the local environment, thereby avoiding the need for different atomic types. Here, we apply the ML model to understand the mechanical and thermal properties of the arsenene polymorphs and nanostructures. We observe an inverse relationship between the critical strain and temperature in arsenene. Thermal conductivity calculations in nanosheets show good agreement with ab initio data, reflecting a decrease in thermal conductivity attributable to increased anharmonic effects at higher temperatures. We also apply the model to predict the thermal behavior of arsenene nanotubes.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

A Comparison of Linear Solvers for Resolving Flow in Three-Dimensional Discrete Fracture Networks

We compare various methods for resolving steady flow within three-dimensional discrete fracture networks, including direct methods, Krylov subspace methods with and without preconditioning, and multi-grid methods. We compared the performance of the methods based on compute times and scaling of the solution as a function of the number of grid nodes and log-variance of the hydraulic aperture. The methods are applied to three test cases: (a) variable density of networks with a truncated power-law distribution of fracture lengths, (b) a fixed network composed of monodisperse fracture sizes but varied permeability/aperture heterogeneity, (c) and a network based on field site in Nevada, US. We chose these cases to allow us to study the impact of the mesh size and flow properties, as well as to demonstrate our conclusions on a large-scale, realistic problem (more than 40 million mesh nodes). A direct solution using Cholesky factorization outperformed other methods for every example but was closely followed in performance by some algebraic multigrid (AMG) preconditioned Krylov subspace methods. Among the Krylov methods, conjugate gradients (CG) with an AMG preconditioner performs the best. Generally, Cholesky factorization is recommended, but CG with an AMG preconditioner may be suitable for very large problems beyond 40 million nodes where the entire linear system cannot reside in memory.

58 GEOSCIENCES↗

Sheath transitions in a cylindrical filament discharge: Axisymmetric 1D3V PIC-MCC simulations

We present the first nonplanar hot cathode discharge simulations that capture the role of the trapped-ions plasma, elucidating new phenomena unobservable in planar geometric discharges. A discharge struck between a single emitting wire filament cathode and a bounding anode is simulated in cylindrical geometry using an axisymmetric (radial) particle-in-cell Monte-Carlo collisions code. Operating the discharge near its ionization energy threshold can lead to the formation of a two plasma mode (TPM). One plasma forms in the conventional upstream region through electron impact ionization of background neutrals. A second plasma, whose global effect on the discharge was not previously well understood, forms downstream through the trapping of cold ions in the potential well of the filament’s virtual cathode, a process enabled by ion-neutral charge exchange collisions. Three space charge regions intersperse the electrode gap—an emissive sheath between the cathode filament and trapped-ions plasma, a double layer between the two plasmas, and a classical sheath between the upstream plasma and the outer anode. Simulations exhibit mode transitions and quenching instabilities that transform the discharge between the TPM and other single-plasma sheath modes that include classical (temperature-limited), space charge limited, and inverse (anode glow) modes. The transitions are explained via “aid-and-compete” dynamics wherein the growth of one plasma enhances growth in the other while concurrently exhibiting expansion dynamics antagonistic to each other. The system exhibits strong hysteresis memory during the mode transitions. Improved understanding and control of these sheath mode transitions are expected to benefit plasma applications with hot cathodes.

Electrical hysteresis↗

Er : Li Nb O 3 with High Optical Coherence Enabling Optical Thickness Control

Integrated photonics capable of incorporating rare-earth ions with high optical coherence is desirable for realizing efficient quantum transducers, compact quantum memories, and hybrid quantum systems. Here we describe a photonic platform based on the SmartCut erbium-doped lithium niobate thin film, and explore its stable optical transitions at telecom wavelength in a dilution refrigerator. Optical coherence time of up to 180μs, rivaling the value of bulk crystals, is achieved in optical ridge waveguides and ring resonators. With this integrated platform, we demonstrate tunable light-ion interaction and flexible control of optical thickness by exploiting long waveguides, whose lengths are in principle variable. This unique ability to obtain high optical density using low-concentration ions further leads to the observation of multiecho pulse trains in centimeter-long waveguides. Our results establish a promising photonic platform for quantum information processing with rare-earth ions.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

MemXCT: Design, Optimization, Scaling, and Reproducibility of X-Ray Tomography Imaging

Here, this work extends our previous research entitled "MemXCT: Memory-centric X-ray CT Reconstruction with Massive Parallelization" that was originally published at SC19 conference (Hidayetoglu et al., 2019) with reproducibility of the computational imaging performance. X-ray computed tomography (XCT) is regularly used at synchrotron light sources to study the internal morphology of materials at high resolution. However, experimental constraints, such as radiation sensitivity, can result in noisy or undersampled measurements. Further, depending on the resolution, sample size and data acquisition rates, the resulting noisy dataset can be in the order of terabytes. Advanced iterative reconstruction techniques can produce high-quality images from noisy measurements, but their computational requirements have made their use an exception rather than the rule. We propose a novel memory-centric approach that avoids redundant computations at the expense of additional memory complexity. We develop a memory-centric iterative reconstruction system, MemXCT, that uses an optimized SpMV implementation with two-level pseudo-Hilbert ordering and multi-stage input buffering. We evaluate MemXCT on various supercomputer architectures involving KNL and GPU. MemXCT can reconstruct a large (11Kx11K) mouse brain tomogram in 10 seconds using 4096 KNL nodes (256K cores). The results presented in our original article at the SC19 were based on large-scale supercomputing resources. The MemXCT application was selected for the Student Cluster Competition (SCC) Reproducibility Challenge and evaluated on a variety of cloud computing resources by universities around the world in the SC20 conference. We summarize the results of the top-ranked SCC Reproducibility Challenge teams and identify the most pertinent measures for ensuring the reproducibility of our experiments in this article.

47 OTHER INSTRUMENTATION↗

rmon (Resource monitor) [SWR-24-128]

The resource monitor application provides monitoring, collection, and visualization of resource utilization in compute nodes. This package contains utilities to monitor system resource utilization (CPU, memory, disk, network). Here are the ways you can use it: -Monitor resource utilization for a compute node for a given set of resource types and process IDs. -Start a process and monitor its resource utilization. -Monitor resource utilization for a compute node asynchronously with the ability to dynamically change the resource types and process IDs being monitored. -Produce JSON reports of aggregated metrics. -Produce interactive HTML plots of the statistics.

Thom, Daniel [National Renewable Energy Laboratory↗

arco (Assembled Resource-Constrained Optimization) [SWR-26-030]

Arco (Assembled Resource-Constrained Optimization) is a memory-smart optimization DSL and solver for LP and MIP problems on constrained hardware. The software is an optimization framework built around a KDL-based domain-specific language and a CLI compiler/solver. You write optimization models in .kdl files, and the arco CLI compiles, validates, inspects, and solves them. Language bindings (Python today, more planned) provide programmatic access to the same engine. Built for harder optimization problems on constrained resources, Arco is intentional about every allocation, careful with stack and heap behavior, and relentless about minimizing memory usage so more systems can run real workloads. Arco is built primarily for internal use within our organization. You are welcome to try it, but we make no guarantees about API stability or robustness at this stage

Sanchez Perez, Pedro Andres [National Laboratory o↗

Leveraging Spin-Orbit Coupling in Ge/SiGe Heterostructures for Quantum Information Transfer

Hole spin qubits confined to lithographically - defined lateral quantum dots in Ge/SiGe heterostructures show great promise. On reason for this is the intrinsic spin - orbit coupling that allows all - electric control of the qubit. That same feature can be exploited as a coupling mechanism to coherently link spin qubits to a photon field in a superconducting resonator, which could, in principle, be used as a quantum bus to distribute quantum information. The work reported here advances the knowledge and technology required for such a demonstration. We discuss the device fabrication and characterization of different quantum dot designs and the demonstration of single hole occupation in multiple devices. Superconductor resonators fabricated using an outside vendor were found to have adequate performance and a path toward flip-chip integration with quantum devices is discussed. The results of an optical study exploring aspects of using implanted Ga as quantum memory in a Ge system are presented.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Parallel Runtime Interface for Fortran (PRIF) Specification (Rev. 0.3)

This document specifies an interface to support the parallel features of Fortran, named the Parallel Runtime Interface for Fortran (PRIF). PRIF is a proposed solution in which the runtime library is responsible for coarray allocation, deallocation and accesses, image synchronization, atomic operations, events, and teams. In this interface, the compiler is responsible for transforming the invocation of Fortran-level parallel features into procedure calls to the necessary PRIF procedures. The interface is designed for portability across shared- and distributed-memory machines, different operating systems, and multiple architectures. Implementations of this interface are intended as an augmentation for the compiler's own runtime library. With an implementation-agnostic interface, alternative parallel runtime libraries may be developed that support the same interface. One benefit of this approach is the ability to vary the communication substrate. A central aim of this document is to define a parallel runtime interface in standard Fortran syntax, which enables us to leverage Fortran to succinctly express various properties of the procedure interfaces, including argument attributes.

97 MATHEMATICS AND COMPUTING↗

Parallel Runtime Interface for Fortran (PRIF) Specification (Rev. 0.5)

This document specifies an interface to support the parallel features of Fortran, named the Parallel Runtime Interface for Fortran (PRIF). PRIF is a proposed solution in which the runtime library is primarily responsible for implementing coarray allocation, deallocation and accesses, image synchronization, atomic operations, events, teams and collective subroutines. In this interface, the compiler is responsible for transforming the invocation of Fortran-level parallel features into procedure calls to the necessary PRIF subroutines. The interface is designed for portability across shared- and distributed-memory machines, different operating systems, and multiple architectures. Implementations of this interface are intended as an augmentation for the compiler's own runtime library. With an implementation-agnostic interface, alternative parallel runtime libraries may be developed that support the same interface. One benefit of this approach is the ability to vary the communication substrate. A central aim of this document is to define a parallel runtime interface in standard Fortran syntax, which enables us to leverage Fortran to succinctly express various properties of the procedure interfaces, including argument attributes.

97 MATHEMATICS AND COMPUTING↗

Parallel Runtime Interface for Fortran (PRIF) Design Document (Rev. 0.2)

This design document proposes an interface to support the parallel features of Fortran, named the Parallel Runtime Interface for Fortran (PRIF). PRIF is a proposed solution in which the runtime library is responsible for coarray allocation, deallocation and accesses, image synchronization, atomic operations, events, and teams. In this interface, the compiler is responsible for transforming the invocation of Fortran-level parallel features into procedure calls to the necessary PRIF procedures. The interface is designed for portability across shared- and distributed-memory machines, different operating systems, and multiple architectures. Implementations of this interface are intended as an augmentation for the compiler's own runtime library. With an implementation-agnostic interface, alternative parallel runtime libraries may be developed that support the same interface. One benefit of this approach is the ability to vary the communication substrate. A central aim of this document is to define a parallel runtime interface in standard Fortran syntax, which enables us to leverage Fortran to succinctly express various properties of the procedure interfaces, including argument attributes.

97 MATHEMATICS AND COMPUTING↗

Parallel Runtime Interface for Fortran (PRIF) Specification (Rev. 0.6)

This document specifies an interface to support the multi-image parallelism features of Fortran, named the Parallel Runtime Interface for Fortran (PRIF). PRIF is a solution in which a runtime library is primarily responsible for implementing coarray allocation, deallocation and accesses, image synchronization, atomic operations, events, teams and collective subroutines. The Fortran compiler is responsible for transforming the invocation of Fortran-level multi-image parallelism features into procedure calls to the necessary PRIF subroutines. The interface is designed for portability across shared- and distributed-memory machines, different operating systems, and multiple architectures. Implementations of this interface are intended as an augmentation for the compiler's own runtime library. With an implementation-agnostic interface, alternative parallel runtime libraries may be developed that support the same interface. One benefit of this approach is the ability to vary the communication substrate. A central aim of this document is to define a parallel runtime interface in standard Fortran syntax, which enables us to leverage Fortran to succinctly express various properties of the procedure interfaces, including argument attributes.

97 MATHEMATICS AND COMPUTING↗

Parallel Runtime Interface for Fortran (PRIF) Specification (Rev. 0.4)

This document specifies an interface to support the parallel features of Fortran, named the Parallel Runtime Interface for Fortran (PRIF). PRIF is a proposed solution in which the runtime library is responsible for coarray allocation, deallocation and accesses, image synchronization, atomic operations, events, and teams. In this interface, the compiler is responsible for transforming the invocation of Fortran-level parallel features into procedure calls to the necessary PRIF procedures. The interface is designed for portability across shared- and distributed-memory machines, different operating systems, and multiple architectures. Implementations of this interface are intended as an augmentation for the compiler's own runtime library. With an implementation-agnostic interface, alternative parallel runtime libraries may be developed that support the same interface. One benefit of this approach is the ability to vary the communication substrate. A central aim of this document is to define a parallel runtime interface in standard Fortran syntax, which enables us to leverage Fortran to succinctly express various properties of the procedure interfaces, including argument attributes.

97 MATHEMATICS AND COMPUTING↗

Optically connected memory for disaggregated data centers

Recent advances in integrated photonics enable the implementation of reconfigurable, high-bandwidth, and low energy-per-bit interconnects in next-generation data centers. We propose and evaluate an Optically Connected Memory (OCM) architecture that disaggregates the main memory from the computation nodes in data centers. OCM is based on micro-ring resonators (MRRs), and it does not require any modification to the DRAM memory modules. We calculate energy consumption from real photonic devices and integrate them into a system simulator to evaluate performance. Here, our results show that (1) OCM is capable of interconnecting four DDR4 memory channels to a computing node using two fibers with 1.02 pJ energy-per-bit consumption and (2) OCM performs up to 5.5× faster than a disaggregated memory with 40G PCIe NIC connectors to computing nodes.

97 MATHEMATICS AND COMPUTING↗