Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Memory systems”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 271 records · Page 15

Jamming, fragility and pinning phenomena in superconducting vortex systems

Abstract We examine driven superconducting vortices interacting with quenched disorder under a sequence of perpendicular drive pulses. As a function of disorder strength, we find four types of behavior distinguished by the presence or absence of memory effects. The fragile and jammed states exhibit memory, while the elastic and pinning dominated regimes do not. In the fragile regime, the system organizes into a pinned state during the first pulse, flows during the second perpendicular pulse, and then returns to a pinned state during the third pulse which is parallel to the first pulse. This behavior is the hallmark of the fragility proposed for jamming in particulate matter. For stronger disorder, we observe a robust jamming state with memory where the system reaches a pinned or reduced flow state during the perpendicular drive pulse, similar to the shear jamming of granular systems. We show signatures of the different states in the spatial vortex configurations, and find that memory effects arise from coexisting elastic and pinned components of the vortex assembly. The sequential perpendicular driving protocol we propose for distinguishing fragile, jammed, and pinned phases should be general to the broader class of driven interacting particles in the presence of quenched disorder.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

Coherently coupled quantum oscillators for quantum reservoir computing

Here, we analyze the properties of a quantum system composed of two coherently coupled quantum oscillators and show through simulations that it fulfills the two properties required for reservoir computing: non-linearity and fading memory. We first show that the basis states of this system apply a set of nonlinear transformations on the input signals and thus can implement neurons. We then show that the system exhibits a fading memory that can be controlled by its dissipation rates. Finally we show that a strong coupling between the oscillators is important in order to ensure complex dynamics and to populate a number of basis state neurons that is exponential in the number of physical devices.

Dudas, Julien↗

Feature Classification for Control System Devices

Control systems are used to automate industrial processes, smart grids, and smart cities. Unfortunately, cyber attacks on control systems are on the rise. Additionally, control systems lack the plethora of tools available for commodity systems for forensic investigation. An important step towards the proper forensic investigation is to analyze device memory. To assist in identifying features of device memory, we present a machine learning-based technique that integrates ontology information for feature classification in a control system device’s memory.

ahmed mithu, M Rayhan↗

Redirecting data to improve page locality in a scalable data fabric

A data processing system includes a host processor, a local memory coupled to the host processor, a plurality of remote memory media, and a scalable data fabric coupled to the host processor and to the plurality of remote memory media. The scalable data fabric includes a filter for storing information indicating a location of data that is stored by the data processing system. The host processor includes a hardware sequencer coupled to the filter for selectively moving data stored by the filter to the local memory.

97 MATHEMATICS AND COMPUTING↗

Precisely computing phonons via irreducible derivatives

Computing phonons from first principles is typically considered a solved problem, yet inadequacies in existing techniques continue to yield deficient results in systems with sensitive phonons. Here, in this study, we circumvent this issue using the lone irreducible derivative (LID) and bundled irreducible derivative (BID) approaches to computing phonons via finite displacements, where the former optimizes precision via energy derivatives and the latter provides the most efficient algorithm using force derivatives. A condition number optimized basis for BID is derived which guarantees the minimum amplification of error. Additionally, a hybrid LID-BID approach is formulated, in which select irreducible derivatives computed using LID replace BID results. We illustrate our approach on two prototypical systems with sensitive phonons: the shape memory alloy AuZn and metallic lithium. Comparing our resulting phonons in the aforementioned crystals to calculations in the literature reveals nontrivial inaccuracies. Our approaches can be fully automated, making them well suited for both niche systems of interest and high-throughput approaches.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

Toward Performance Portable Programming for Heterogeneous System-on-Chips: Case Study with Qualcomm Snapdragon SoC

Future heterogeneous Domain-Specific System-on-Chips (DSSoC) will be extraordinarily complex in terms of processors, memory hierarchies, and interconnection networks.To manage this complexity, architects, system software designers, and application developers need programming technologies that are flexible, accurate, efficient, and productive. These technologies will need to be as independent of any one specific architecture as is practical, because the sheer dimensionality and scale of the complexity will not allow porting and optimizing applications foreach given DSSoC. To address these issues, we are developing Cosmic Castle, a performance portable programming toolchain for streaming applications on heterogeneous architectures. The primary focus of Cosmic Castle is on enabling efficient and performant code generation through the smart compiler and intelligent runtime system. This paper presents the preliminary evaluation of our ongoing work toward Cosmic Castle. Specifically, we detail our code porting efforts and evaluate various benchmarks on the Qualcomm Snapdragon SoC using tools developed through Cosmic Castle.

Cabrera, Anthony↗

Cycle accurate and cycle reproducible memory for an FPGA based hardware accelerator

A method, system and computer program product are disclosed for using a Field Programmable Gate Array (FPGA) to simulate operations of a device under test (DUT). The DUT includes a device memory having a number of input ports, and the FPGA is associated with a target memory having a second number of input ports, the second number being less than the first number. In one embodiment, a given set of inputs is applied to the device memory at a frequency Fd and in a defined cycle of time, and the given set of inputs is applied to the target memory at a frequency Ft. Ft is greater than Fd and cycle accuracy is maintained between the device memory and the target memory. In an embodiment, a cycle accurate model of the DUT memory is created by separating the DUT memory interface protocol from the target memory storage array.

Asaad, Sameh W.↗

Generative memory for lifelong machine learning

Techniques are disclosed for training machine learning systems. An input device receives training data comprising pairs of training inputs and training labels. A generative memory assigns training inputs to each archetype task of a plurality of archetype tasks, each archetype task representative of a cluster of related tasks within a task space and assigns a skill to each archetype task. The generative memory generates, from each archetype task, auxiliary data comprising pairs of auxiliary inputs and auxiliary labels. A machine learning system trains a machine learning model to apply a skill assigned to an archetype task to training and auxiliary inputs assigned to the archetype task to obtain output labels corresponding to the training and auxiliary labels associated with the training and auxiliary inputs assigned to the archetype task to enable scalable learning to obtain labels for new tasks for which the machine learning model has not previously been trained.

Nadamuni Raghavan, Aswin↗

Combustion control using spiking neural networks

A system that controls a combustion engine stores network vectors in a memory that represent diverse and distinct spiking neural networks. The system decodes the network vectors and trains and evaluates the spiking neural networks. The system duplicates selected network vectors and crosses-over the duplicated network vectors that represent modified spiking neural networks. The system mutates the crossed-over duplicated network vectors by randomly modifying one or more portions of the crossing-over duplicated network vectors. The system meter exhaust gas into an intake manifold when an engine temperature exceeds a threshold, an engine load exceeds a threshold, an engine's rotation-per-minute rate exceeds a threshold, and a fuel flow exceeds a threshold. The system modifies fuel flow into an engine's combustion chamber on a cycle-to-cycle basis by the trained spiking neural network.

Puente, Brian P. Maldonado↗

Sparse Approximate Multifrontal Factorization with Composite Compression Methods

This article presents a fast and approximate multifrontal solver for large sparse linear systems. In a recent work by Liu et al., we showed the efficiency of a multifrontal solver leveraging the butterfly algorithm and its hierarchical matrix extension, HODBF (hierarchical off-diagonal butterfly) compression to compress large frontal matrices. The resulting multifrontal solver can attain quasi-linear computation and memory complexity when applied to sparse linear systems arising from spatial discretization of high-frequency wave equations. To further reduce the overall number of operations and especially the factorization memory usage to scale to larger problem sizes, in this article we develop a composite multifrontal solver that employs the HODBF format for large-sized fronts, a reduced-memory version of the nonhierarchical block low-rank format for medium-sized fronts, and a lossy compression format for small-sized fronts. This allows us to solve sparse linear systems of dimension up to 2.7 × larger than before and leads to a memory consumption that is reduced by 70% while ensuring the same execution time. The code is made publicly available in GitHub.

97 MATHEMATICS AND COMPUTING↗

Study of interconnect errors, network congestion, and applications characteristics for throttle prediction on a large scale HPC system

Today’s High Performance Computing (HPC) systems contain thousand of nodes which work together to provide performance in the order of petaflops. The performance of these systems depends on various components like processors, memory, and interconnect. Among all, interconnect plays a major role as it glues together all the hardware components in an HPC system. A slow interconnect can impact a scientific application running on multiple processes severely as they rely on fast network messages to communicate and synchronize frequently. Unfortunately, the HPC community lacks a study that explores different interconnect errors, congestion events and applications characteristics on a large-scale HPC system. In our previous work, we process and analyze interconnect data of the Titan supercomputer to develop a thorough understanding of interconnects faults, errors, and congestion events. In this work, we first show how congestion events can impact application performance. We then investigate application characteristics interaction with interconnect errors and network congestion to predict applications encountering congestion with more than 90% accuracy.

97 MATHEMATICS AND COMPUTING↗

Dataset for manuscript "Rotational Memory Function of SPC/E water"

Memory effect are essential for dynamics of condensed materials and are responsible for non-exponential relaxation of correlation functions of dynamic variables through the memory function entering the memory equation. Memory functions of dipole rotations for polar liquids have never been calculated. We present here calculations of memory functions and single-dipole rotations and of the overall system dipole moment for SPC/E water measured by dielectric spectroscopy. The memory functions for single-particle and collective dynamics turn out to be nearly identical. This result validates theories of dielectric spectroscopy in terms of single-particle time correlation function and the connection between the collective and single-particle relaxation times in terms of the Kirkwood factor. The dataset includes single particle and system dipole moments, including their time-dependence.

74 ATOMIC AND MOLECULAR PHYSICS↗

Compute in‐Memory with Non‐Volatile Elements for Neural Networks: A Review from a Co‐Design Perspective

Abstract Deep learning has become ubiquitous, touching daily lives across the globe. Today, traditional computer architectures are stressed to their limits in efficiently executing the growing complexity of data and models. Compute‐in‐memory (CIM) can potentially play an important role in developing efficient hardware solutions that reduce data movement from compute‐unit to memory, known as the von Neumann bottleneck. At its heart is a cross‐bar architecture with nodal non‐volatile‐memory elements that performs an analog multiply‐and‐accumulate operation, enabling the matrix‐vector‐multiplications repeatedly used in all neural network workloads. The memory materials can significantly influence final system‐level characteristics and chip performance, including speed, power, and classification accuracy. With an over‐arching co‐design viewpoint, this review assesses the use of cross‐bar based CIM for neural networks, connecting the material properties and the associated design constraints and demands to application, architecture, and performance. Both digital and analog memory are considered, assessing the status for training and inference, and providing metrics for the collective set of properties non‐volatile memory materials will need to demonstrate for a successful CIM technology.

36 MATERIALS SCIENCE↗

Software-Hardware Co-design of Heterogeneous SmartNIC System for Recommendation Models Inference and Training

Deep Learning Recommendation Models (DLRMs) are critical applications in various domains and have evolved as one of the single largest machine learning applications. Trillions of DLRM parameters exceed the on-chip memory capacity of GPUs. Large-scale multi-node systems are required for distributed DLRM inference and training, which suffer from the all-to-all communication bottleneck, mainly limiting the scalability of ever-growing DLRMs. In recent years, SmartNICs have evolved with coupled computation and communication capabilities providing opportunities for a powerful heterogeneous device in the system. However, there isn't such a distributed system that fully leverages the abundant smartNIC resources that resolve the scalability issue of DLRMs. In this work, we proposed a software-hardware co-design of a heterogeneous smartNIC system that resolves the communication bottleneck of distributed DLRMs, mitigates the memory bandwidth pressure, and improves computation efficiency. We provide a set of smartNIC designs of cache systems (including local cache and remote cache) and smartNIC computation kernels which reduce data movement, relieve memory lookup intensity, and improve the GPU's computation efficiency. In addition, we propose a graph algorithm that improves the data locality of queries within batches which optimizes the overall system performance with higher data reuse. Our evaluation shows that our system achieves 2.1x latency speedup for inference and 1.6x throughput speedup for training.

Guo, Anqi↗

Tunable shear thickening, aging, and rejuvenation in suspensions of shape-memory-endowed liquid crystalline particles

The morphological features of particles, notably shape anisotropy, critically influence the rheological properties of dense suspensions, spanning both natural and engineered systems. This work explores the potential of using shape memory particles to dynamically regulate suspension fluid flow through controllable shape transformations. First, we synthesize shape-memory particles with programmable anisotropy from liquid crystal elastomers, such that the stiffness and shapes of the particles can be tuned by manipulating temperature. Our findings reveal that suspensions from such particles exhibit significant tunability in shear thickening behavior, transitioning from discontinuous shear thickening to a Newtonian-like response within a narrow temperature range of 60 ° C. This capability to modulate rheological responses in situ presents an approach for addressing processing challenges in many applications where control over flow behavior is paramount. Furthermore, we also show that suspensions composed of these anisotropic particles can undergo physical aging, and evolve into a glassy state. This state can be escaped upon activation of the shape memory effect. This reversibility underscores the potential for using such materials to engineer systems that can enter or come out of kinetic arrest by leveraging internal mechanical responses to external stimuli. The insights gained here not only broaden our understanding of the interplay between particle geometry and suspension dynamics but also pave the way for leveraging ensembles of stimuli-responsive objects to precisely control collective behaviors in many-body systems.

Science & Technology - Other Topics↗

Exploring Architectural-Aware Affinity Policies in Modern HPC Runtimes

Modern commodity and High-Performance Computing (HPC) systems are evolving with complex CPU architectures. These architectures now feature higher core and NUMA domain counts and implement features such as hyperthreading. When considering significant differences in hardware configurations, library availability, and hardware-tailored system/software stacks, which could substantially vary from one system to another, performance portability is hard to achieve. Throughout the years, this trend resulted in an increasingly high burden on application developers to fine-tune their workloads for each architecture. This work explores how hardware-dependent aspects such as locality/process/thread affinity affect performance in modern CPU architectures. We focus our study on the Global Memory and Threading (GMT) distributed runtime system as a representative of Partitioned Global Address Space (PGAS) software stacks commonly adopted for productivity. In particular, to appreciate performance implications, we evaluate GMT’s thread affinity policies, and, introduce two new ones which exploit architectural awareness. Finally, we explore alternative NUMA configurations via different process bindings and perform a scalability study on three HPC clusters with varying CPU architectures and NUMA layouts. Our analysis indicates that more complex architectures are more affected by affinity and binding policies and highlights the importance of setting proper runtime configurations to achieve superior performance.

Di Dio Lavore, Ian↗

Experimental Characterization of OpenMP Offloading Memory Operations and Unified Shared Memory Support

The OpenMP specification recently introduced support for unified shared memory, allowing implementation to leverage underlying system software to provide a simpler GPU offloading model where explicit mapping of variables is optional. Support for this feature is becoming more available in different OpenMP implementations on several hardware platforms. A deeper understanding of the different implementation’s execution profile and performance is crucial for applications as they consider the performance portability implications of adopting a unified memory offloading programming style. This work introduces a benchmark tool to characterize unified memory support in several OepnMP compilers and runtimes, with emphasis on identifying discrepancies between different OpenMP implementations as to how they various memory allocation strategies interact with unified shared memory. The benchmark tool is used to characterize OpenMP compilers on three leading High Performance Computing platforms supporting different CPU and device architectures. The benchmark tool is used to assess the impact of enabling unified shared memory on the performance of memory-bound code, highlighting implementation differences that should be accounted for when applications consider performance portability across platforms and compilers.

Elwasif, Wael↗

Comprehensive assessment of deep reinforcement learning approaches for economic dispatch in nuclear-driven microgrids

As the electrical grid integrates more variable renewable energy sources such as wind and solar, the demand for distributed and flexible systems to address this increased variability becomes critical. Nuclear-driven microgrids provide a promising solution by offering stable generation to complement intermittent renewables, ensuring grid reliability and operating efficiency. This paper proposes a recurrent deep reinforcement learning framework for optimal economic dispatch in a nuclear-powered microgrid integrating renewable energy sources, small modular reactors, battery storage systems, and balance-of-plant dynamics. A three-agent control architecture is developed, where demand and renewable energy agents act as forecasters, and a reinforcement learning-based dispatch agent performs real-time energy allocation. A nonlinear programming formulation is first used to generate an optimal baseline for benchmarking. The proposed dispatch controller, based on Proximal Policy Optimization enhanced with Long Short-Term Memory networks, exploits temporal correlations in system dynamics by taking advantage of the time series used as inputs to improve policy robustness under uncertainty. Comparative analysis against established deep reinforcement learning methods, including Proximal Policy Optimization with a feedforward architecture, Soft Actor-Critic, and Twin Delayed Deep Deterministic Policy Gradient, demonstrates superior performance. Numerical results indicate that the proposed controller achieves a 0.39% cost reduction relative to the nonlinear programming benchmark and outperforms other learning-based methods by generating additional revenue of up to 0.35%. All reinforcement learning controllers compute dispatch actions in less than 0.3 s, resulting in a computational speedup of more than three orders of magnitude over the nonlinear programming baseline. The findings of this paper highlight their applicability for real-time operation and control in nuclear-integrated microgrids under volatile operating conditions.

24 POWER TRANSMISSION AND DISTRIBUTION↗