Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Memory systems”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 307 records · Page 17

Magnetic hysteresis experiments performed on quantum annealers

While quantum annealers have emerged as versatile and controllable platforms for experimenting on correlated spin systems, the important phenomenology of magnetic memory and hysteresis remain unexplored on hardware designed to escape metastable states via quantum tunneling. Here, we present the first general protocol to experiment on magnetic hysteresis on programmable quantum annealers and implement it on three D-Wave superconducting qubit quantum annealers, using up to thousands of spins, for both ferromagnetic and disordered Ising models, and across different graph topologies. We observe hysteresis loops whose area depends nonmonotonically on quantum fluctuations, exhibiting both expected and unexpected features, such as disorder-induced steps and nonmonotonicities. Our work establishes quantum annealers as a platform for probing nonequilibrium emergent magnetic phenomena, thereby broadening the role of analog quantum computers into foundational questions in condensed matter physics.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

HydraGNN v3.0

New or improved capabilities included in v3.0 release are as follows: 1. Enhancement in message passing layers through generalization of the class inheritance to enable the inclusion of a broader set of message passing policies Inclusion of equivariant message passing layers from the original implementations of: SchNet (https://pubs.aip.org/aip/jcp/article/148/24/241722/962591/SchNet-A-deep-learning-architecture-for-molecules); DimeNet++ (https://arxiv.org/abs/2011.14115); EGNN models (https://arxiv.org/pdf/2102.09844.pdf) 2. Restructuring of class inheritance for data management 3. Support of DDStore https://github.com/ORNL/DDStore capabilities for improved distributed data parallelism on large volumes of data that cannot fit on intra-node memory capacities 4. Large-scale system support for OLCF-Crusher and OLCF-Frontier

Lupo Pasini, Massimiliano [Oak Ridge National Labo↗

Privacy-Preserving Control of Partitioned Energy Resources

Distributed energy resources are an increasingly important part of the electric grid. We examine the problem of partitioning a distributed energy resource among many users while providing privacy to them. In this model, clients can send requests to a server, the server can verify that the requests are valid and aggregate them, but it cannot see the actual values in the requests. Without privacy, each user is forced to reveal their daily schedule or energy use. Energy resources add a novel challenge that prior systems do not address: they require verifying limits on private power (a rate over time) and energy (a sum) values. Furthermore, the cryptographic mechanisms must run on embedded energy control systems. We describe Weft, a novel cryptographic system that verifies both power (rate) and energy (integral) constraints on private client values and aggregates them. The key insight behind the approach is to rely on additively homomorphic secret shares, which allows servers to compute sums from rates. We present 3 cryptographic proof systems with different system trade-off for embedded systems: bit-splitting proofs minimize memory use, sorting proofs minimize computation, and commitment proofs minimize network communication. Using bit-splitting proofs, it takes an IoT client using a CortexM microcontroller 4 minutes of compute time to privately control its share of an energy resource for a day at 20s granularity.

Laufer, Evan↗

MICCO: An Enhanced Multi-GPU Scheduling Framework for Many-Body Correlation Functions

Calculation of many-body correlation functions is one of the critical kernels utilized in many scientific computing areas, especially in Lattice Quantum Chromodynamics (Lattice QCD). It is formalized as a sum of a large number of contraction terms each of which can be represented by a graph consisting of vertices describing quarks inside a hadron node and edges designating quark propagations at specific time intervals. Due to its computation- and memory-intensive nature, real-world physics systems (e.g., multi-meson or multi-baryon systems) explored by Lattice QCD prefer to leverage multi-GPUs. Different from general graph processing, many-body correlation function calculations show two specific features: a large number of computation-/data-intensive kernels and frequently repeated appearances of original and intermediate data. The former results in expensive memory operations such as tensor movements and evictions. The latter offers data reuse opportunities to mitigate the data-intensive nature of many-body correlation function calculations. However, existing graph-based multi-GPU schedulers cannot capture these data-centric features, thus resulting in a sub-optimal performance for many-body correlation function calculations. To address this issue, this paper presents a multi-GPU scheduling framework, MICCO, to accelerate contractions for correlation functions particularly by taking the data dimension (e.g., data reuse and data eviction) into account. This work first performs a comprehensive study on the interplay of data reuse and load balance, and designs two new concepts: local reuse pattern and reuse bound to study the opportunity of achieving the optimal trade-off between them. Based on this study, MICCO proposes a heuristic scheduling algorithm and a machine-learning-based regression model to generate the optimal setting of reuse bounds. Specifically, MICCO is integrated into a real-world Lattice QCD system, Redstar, for the first time running on multiple GPUs. The evaluation demonstrates MICCO outperforms other state-of-art works, achieving up to 2.25× speedup in synthesized datasets, and 1.49× speedup in real-world correlation functions.

Wang, Qihan↗

IRIS-GNN: Leveraging Graph Neural Networks for Scheduling on Truly Heterogeneous Runtime Systems

The diversity of accelerators in computer systems poses significant challenges for software developers, such as managing vendor-specific compiler toolchains, code fragmentation requiring different kernel implementations, and performance portability issues. To address these, the Intelligent Runtime System (IRIS) was developed. IRIS works across various systems, from smartphones to supercomputers, enabling automatic performance scaling based on available accelerators. It introduces abstract tasks for seamless execution transitions between accelerators while ensuring memory consistency and task dependencies. Although IRIS simplifies system details, optimal dynamic scheduling still requires user input to understand workload structures. To address this, we introduce a new scheduling policy for IRIS, termed IRIS-GNN, which is the first IRIS hybrid policy that operates in conjunction with the dynamic policies. This policy employs a Graph-Neural Network (GNN) to conduct Graph Classification of any task graphs submitted to IRIS. This GNN analyzes the structure and attributes of the task graph, categorizing it as either locality, concurrency, or mixed. This classification subsequently guides the selection of the dynamic policy used by IRIS. We provide a comparison of the performance of IRIS-GNN against the complete spectrum of IRIS’s dynamic policies, assess the overhead introduced by the GNN within this scheduling framework, and ultimately explore its practical application in real-world scenarios.

Johnston, Beau↗

Microstructure-sensitive mechanical behavior of an additively manufactured psuedoelastic shape memory alloy

The additive manufacturing of shape memory alloys into complex geometries enables fabrication of advanced functional systems across a variety of fields and domains. This work presents results focused on the mechanical behavior of additively manufactured shape memory pseudoelastic NiTi. The deformation induced solid state phase transformation from austenite to martensite allows this system to accommodate large recoverable strains. This deformation behavior is fundamentally driven by crystal-scale transformation physics. Laser powder bed fusion processing reveals that the resulting microstructure, both grain morphology and crystallographic texture, is strongly dependent on the manufacturing processing history. Exhaustive mechanical testing demonstrates that these microstructural factors strongly impact both tensile and cyclic stress–strain behavior. Cyclic dissipative behavior, however, is similar across all tested microstructures following an initial transient period. Remarkably, analysis of spatial strain fields during tensile loading reveals two distinctly different localization “modes”. The first is initiation of localized deformation bands which continuously propagate through the tensile bar during loading. In the second mode localization is observed but lacks propagation; instead additional localization cites nucleate during subsequent loading. The latter phenomena is suspected to be driven by grain-scale deformation physics as the localized band morphologies coincide with grain morphologies. These phenomena strongly impact the resulting aggregate stress–strain behavior. Hence, manufacturers and designers of psuedoelastic functional components must at the very least consider the potential variability in properties when considering additive manufacturing processing. More ideally the process–structure–property relations can be used to further tailor and optimize final functional performance.

Additive manufacturing↗

Fast and Scalable Sparse Triangular Solver for Multi-GPU Based HPC Architectures

Designing efficient and scalable sparse linear algebra kernels on modern multi-GPU based HPC systems is a daunting task due to significant irregular memory references and workload imbalance across the GPUs. This is particularly the case for \textit{Sparse Triangular Solver (SpTRSV)} which introduces additional two-dimensional computation dependencies among subsequent computation steps. Dependency information is exchanged and shared among GPUs, thus warrant for efficient memory allocation, data partitioning, and workload distribution as well as fine-grained communication and synchronization support. In this work, we demonstrate that directly adopting unified memory can adversely affect the performance of SpTRSV on multi-GPU architectures, despite linking via fast interconnect like NVLinks and NVSwitches. Alternatively, we employ the latest NVSHMEM technology based on Partitioned Global Address Space programming model to enable efficient fine-grained communication and drastic synchronization overhead reduction. Furthermore, to handle workload imbalance, we propose a malleable task-pool execution model which can further enhance the utilization of GPUs. By applying these techniques, our experiments on the NVIDIA multi-GPU supernode V100-DGX-1 and DGX-2 systems demonstrate that our design can achieve on average 3.53x (up to 9.86x) speedup on a DGX-1 system and 3.66x (up to 9.64x) speedup on a DGX-2 system with 4-GPUs over the Unified-Memory design. The comprehensive sensitivity and scalability studies also show that the proposed zero-copy SpTRSV is able to fully utilize the computing and communication resources of the multi-GPU system.

Xie, Chenhao↗

Distributed out-of-memory NMF on CPU/GPU architectures

We propose an efficient distributed out-of-memory implementation of the non-negative matrix factorization (NMF) algorithm for heterogeneous high-performance-computing systems. The proposed implementation is based on prior work on NMFk, which can perform automatic model selection and extract latent variables and patterns from data. In this work, we extend NMFk by adding support for dense and sparse matrix operation on multi-node, multi-GPU systems. The resulting algorithm is optimized for out-of-memory problems where the memory required to factorize a given matrix is greater than the available GPU memory. Memory complexity is reduced by batching/tiling strategies, and sparse and dense matrix operations are significantly accelerated with GPU cores (or tensor cores when available). Input/output latency associated with batch copies between host and device is hidden using CUDA streams to overlap data transfers and compute asynchronously, and latency associated with collective communications (both intra-node and inter-node) is reduced using optimized NVIDIA Collective Communication Library (NCCL) based communicators. Benchmark results show significant improvement, from 32X to 76x speedup, with the new implementation using GPUs over the CPU-based NMFk. Good weak scaling was demonstrated on up to 4096 multi-GPU cluster nodes with approximately 25,000 GPUs when decomposing a dense 340 Terabyte-size matrix and an 11 Exabyte-size sparse matrix of density 10 -6 .

97 MATHEMATICS AND COMPUTING↗

Independent malware detection architecture

A system and method (referred to as the system) detect malware by training a rule-based model, a functional based model, and a deep learning-based model from a memory snapshot of a malware free operating state of a monitored device. The system extracts a feature set from a second memory snapshot captured from an operating state of the monitored device and processes the feature set by the rule-based model, the functional-based model, and the deep learning-based model. The system identifies identifying instances of malware on the monitored device without processing data identifying an operating system of the monitored device, data associated with a prior identification of the malware, data identifying a source of the malware, data identifying a location of the malware on the monitored device, or any operating system specific data contained within the monitored device.

Smith, Jared M.↗

IRIS-DMEM: Efficient Memory Management for Heterogeneous Computing

This paper proposes an efficient data memory management approach for the Intelligent RuntIme System (IRIS) heterogeneous computing framework along with new data transfer policies. IRIS provides a task-based programming model for extreme heterogeneous computing (e.g., CPU, GPU, DSP, FPGA) with support for today's most important programming languages (e.g., OpenMP, OpenCL, CUDA, HIP, OpenACC). However, the IRIS framework either forces the programmer to introduce data transfer commands for each task or relies on suboptimal memory management for automatic and transparent data transfers. The work described here extends IRIS with novel heterogeneous memory handling and introduces novel data transfer policies by employing the Distributed data MEMory handler (DMEM) for efficient and optimal movement of data among the various computing resources. The proposed approach achieves performance gains of up to 7× for tiled LU factorization and tiled DGEMM (i.e., matrix multiplication) benchmarks. Moreover, this approach also reduces data transfers by up to 71% when compared to previous IRIS heterogeneous memory management handlers. This work compares the performance results of the IRIS framework's novel DMEM with the StarPU runtime and MAGMA math library for GPUs. Experiments show a performance gain of up to 1.95× over StarPU and 2.1× over MAGMA.

Miniskar, Narasinga Rao↗

Charged particle tracking in real-time using a full-mesh data delivery architecture and associative memory techniques

We present a flexible and scalable approach to address the challenges of charged particle track reconstruction in real-time event filters (Level-1 triggers) in collider physics experiments. The method described here is based on a full-mesh architecture for data distribution and relies on the Associative Memory approach to implement a pattern recognition algorithm that quickly identifies and organizes hits associated to trajectories of particles originating from particle collisions. We describe a successful implementation of a demonstration system composed of several innovative hardware and algorithmic elements. The implementation of a full-size system relies on the assumption that an Associative Memory device with the sufficient pattern density becomes available in the future, either through a dedicated ASIC or a modern FPGA. We demonstrate excellent performance in terms of track reconstruction efficiency, purity, momentum resolution, and processing time measured with data from a simulated LHC-like tracking detector.

47 OTHER INSTRUMENTATION↗

Quantum Memristors in Frequency-Entangled Optical Fields

A quantum memristor is a passive resistive circuit element with memory, engineered in a given quantum platform. It can be represented by a quantum system coupled to a dissipative environment, in which a system–bath coupling is mediated through a weak measurement scheme and classical feedback on the system. In quantum photonics, such a device can be designed from a beam splitter with tunable reflectivity, which is modified depending on the results of measurements in one of the outgoing beams. Here, we show that a similar implementation can be achieved with frequency-entangled optical fields and a frequency mixer that, working similarly to a beam splitter, produces state superpositions. We show that the characteristic hysteretic behavior of memristors can be reproduced when analyzing the response of the system with respect to the control, for different experimentally attainable states. Since memory effects in memristors can be exploited for classical and neuromorphic computation, the results presented in this work could be a building block for constructing quantum neural networks in quantum photonics, when scaling up.

36 MATERIALS SCIENCE↗

Extracting energy from ocean thermal and salinity gradients to power unmanned underwater vehicles: State of the art, current limitations, and future outlook

Thermal gradient energy-generation technologies for powering unmanned underwater vehicles (UUVs) or autonomous sensing systems in the ocean are mainly in the research development phase or commercially available at a limited scale, and salinity-gradient energy-generation technologies have not been adequately researched yet. The demand for self-powered UUVs suitable for long-term deployments has been growing, and further research related to small-scale ocean gradient energy systems is needed. In this study, we conducted a comprehensive review about harvesting energy from ocean thermal or salinity gradients for powering UUVs, focusing on gliders and profiling floats. Thermal gradient energy systems for UUVs based on phase change materials (PCM) cannot provide the energy required for powering autonomous sensing systems because of the systems' low energy conversion efficiency. Besides reducing energy consumption by developing more efficient electrical-mechanical systems, enhancing the thermal conductivity of the PCMs may help address this challenge by increasing the power generation rate of the UUVs. Several other emerging technologies, such as thermoelectric generators, shape memory alloys, and small-scale thermodynamic cycle systems, have shown potential for powering UUVs, but they are still only at the laboratory testing or conceptual design phase. The most advanced power generation technologies based on salinity gradients, reverse electrodialysis and pressure-retarded osmosis, are still not economically viable for large-scale deployment, mainly because of the high cost of the components required to operate in harsh saline environments. Our feasibility evaluation showed that existing salinity gradient power generation technologies are not directly feasible for powering UUVs in the open ocean.

16 TIDAL AND WAVE POWER↗

Diagnosing and Destroying Non-Markovian Noise

Nearly every protocol used to analyze the performance of quantum information processors is based on an assumption that the errors experienced by the device during logical operations are constant in time and are insensitive to external contexts. This assumption is pervasive, rarely stated, and almost always wrong. Quantum devices that do behave this way are termed "Markovian:' but nearly every system we have ever probed has displayed drift or crosstalk or memory effects they are all non-Markovian. Strong non-Markovianity introduces spurious effects in characterization protocols and violates assumptions of the fault-tolerance threshold theorems. This SAND report details a three year laboratory-directed research and development (LDRD) project entitled, "Diagnosing and Destroying non-Markovian Noise in Quantum Information Processors." This program was initiated to build tools to study non-Markovian dynamics and quantum systems and develop robust methodologies for eliminating it. The program achieved a number of notable successes, including the first statistically rigorous protocol for identifying and characterizing drift in quantum systems, a formalism for modeling memory effects in quantum devices, and the successful suppression of drift in a Sandia trapped-ion quantum processor.

97 MATHEMATICS AND COMPUTING↗

Electrochemical Random-Access Memory: Progress, Perspectives, and Opportunities

Non-von Neumann computing using neuromorphic systems based on analogue synaptic and neuronal elements has emerged as a potential solution to tackle the growing need for more efficient data processing, but progress toward practical systems has been stymied due to a lack of materials and devices with the appropriate attributes. Recently, solid state electrochemical ion-insertion, also known as electrochemical random access memory (ECRAM) has emerged as a promising approach to realize the needed device characteristics. ECRAM is a three terminal device that operates by tuning electronic conductance in functional materials through solid-state electrochemical redox reactions. This mechanism can be considered as a gate-controlled bulk modulation of dopants and/or phases in the channel. Early work demonstrating that ECRAM can achieve nearly ideal analogue synaptic characteristics has sparked tremendous interest in this approach. More recently, the realization that electrochemical ion insertion can be used to tune the electronic properties of many types of materials including transition metal oxides, layered two-dimensional materials, organic and coordination polymers, and that the changes in conductance can span orders of magnitude has further attracted interest in ECRAM as the basis for analogue synaptic elements for inference accelerators as well as for dynamical devices that can emulate a wide range of neuronal characteristics for implementation in analogue spiking neural networks. At its core, ECRAM shares many fundamental aspects with rechargeable batteries, where ion insertion materials are used extensively for their ability to reversibly store charge and energy. Computing applications, however, present drastically different requirements: systems will require many millions of devices, scaled down to tens of nanometers, all while achieving reliable electronic-state tuning at scaled-up rates and endurances, and with minimal energy dissipation and noise. Further, in this review, we discuss the history, basic concepts, recent progress, as well as the challenges and opportunities for different types of ECRAM, broadly grouped by their primary mobile ionic charge carrier, including Li, protons, and oxygen vacancies.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Feasibility Study of Millimeter Wave Radars For Safeguards Applications

Containment and surveillance are fundamental measures in nuclear safeguards. Techniques such as video surveillance and laser curtain for containment provide effective monitoring in areas where maintaining continuity of knowledge is required. These systems, however, can be susceptible to loss of monitoring capabilities under certain environmental conditions such as poor visibility (i.e. low light conditions, smoke, fog, etc.) or extended power loss past the duration that the backup power system is designed for. Brookhaven National Laboratory has been investigating the feasibility of millimeter waves (mmWave) as a new perimeter seal in which radio frequency waves in the range of 60-64 GHz are used to detect and monitor objects of interest. Signals in this frequency range are not susceptible to environmental conditions. For proof-of-concept tests, mmWave sensors from Texas Instruments (TI), specificallyIWR6843, are used in a test bed at BNL's Waste Management facility to simulate the operations at nuclear facilities. The unique design of TI mmWave sensors requires less memory and power consumption compared to counterpart systems. These devices are capable of exporting 3D point-cloud data, which is visualized graphically and compared to videos recorded at the same time to validate the performance of the mmWave sensor. A set of experiments were planned to test the feasibility of the mmWave in this application, including monitoring static containers in a storage area and detecting intrusions at the boundaries of the area. In addition, the experiments also identify potential blind spots relative to sensor position and utilize multiple operating sensors simultaneously to reduce or eliminate such blind spots. The optimal positioning of multiple sensors was determined for the experimental room configuration. In this paper, we will discuss the details of this novel perimeter sealing concept and present the test results.

98 NUCLEAR DISARMAMENT, SAFEGUARDS, AND PHYSICAL P↗

Machine-learning Kohn–Sham potential from dynamics in time-dependent Kohn–Sham systems

Abstract The construction of a better exchange-correlation potential in time-dependent density functional theory (TDDFT) can improve the accuracy of TDDFT calculations and provide more accurate predictions of the properties of many-electron systems. Here, we propose a machine learning method to develop the energy functional and the Kohn–Sham potential of a time-dependent Kohn–Sham (TDKS) system is proposed. The method is based on the dynamics of the Kohn–Sham system and does not require any data on the exact Kohn–Sham potential for training the model. We demonstrate the results of our method with a 1D harmonic oscillator example and a 1D two-electron example. We show that the machine-learned Kohn–Sham potential matches the exact Kohn–Sham potential in the absence of memory effect. Our method can still capture the dynamics of the Kohn–Sham system in the presence of memory effects. The machine learning method developed in this article provides insight into making better approximations of the energy functional and the Kohn–Sham potential in the TDKS system.

97 MATHEMATICS AND COMPUTING↗

Development of Segregated Thermal-Hydraulics Solvers in MOOSE

The simulation of fluid flows is an essential part of the design and analysis of nuclear systems. Algorithms able to simulate flows at different fidelity levels are available in the Multiphysics Object-Oriented Simulation Environment (MOOSE) and MOOSE-based applications such as Pronghorn \cite{novak2018pronghorn}, Pronghorn-Subchannel, RELAP-7, and SAM. Currently, significant effort is being invested in the development of coarse-mesh Computational Fluid Dynamics (CFD) capabilities within MOOSE and Pronghorn for the simulation of Generation IV nuclear reactors. Traditionally, the solution algorithms in MOOSE have relied on Newton or quasi-Newton methods (such as the preconditioned Jacobian-free Newton-Krylov method) where residuals and Jacobians (or approximations thereof) are constructed. Both Newton and quasi-Newton methods require the solution of a linear system at each nonlinear Newton iteration with the Jacobian as the system matrix. The Jacobian contains blocks originating from all variables in the problem (i.e., for thermal-hydraulics at least pressure, velocities, and temperature). Due to the formulation of the problem in a general multiphysics setting on unstructured mesh, creating a good preconditioner for the linear system can be challenging, thus many fluid applications have utilized direct solver-based methods such as LU factorization. However, with increasing system size and complexity in multi-dimensional problems, the direct solution of linear systems becomes computationally expensive both in execution time and and memory. For this reason, recent effort has focused on adapting segregated solution algorithms for CFD problems in MOOSE. These algorithms use fixed-point iteration between segregated systems whose assembly and preconditioning are easier those of the monolithic system. Initial results show that the segregated solution algorithm outperforms the monolithic approach in terms of memory usage and for large 3D problems in terms of CPU time as well.

42 ENGINEERING↗