Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “limited memory”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Is random access memory random?

Most software is contructed on the assumption that the programs and data are stored in random access memory (RAM). Physical limitations on the relative speeds of processor and memory elements lead to a variety of memory organizations that match processor addressing rate with memory service rate. These include interleaved and cached memory. A very high fraction of a processor's address requests can be satified from the cache without reference to the main memory. The cache requests information from main memory in blocks that can be transferred at the full memory speed. Programmers who organize algorithms for locality can realize the highest performance from these computers.

Denning, P. J.↗

Fluid/Structure Interaction Studies of Aircraft Using High Fidelity Equations on Parallel Computers

Abstract Aeroelasticity which involves strong coupling of fluids, structures and controls is an important element in designing an aircraft. Computational aeroelasticity using low fidelity methods such as the linear aerodynamic flow equations coupled with the modal structural equations are well advanced. Though these low fidelity approaches are computationally less intensive, they are not adequate for the analysis of modern aircraft such as High Speed Civil Transport (HSCT) and Advanced Subsonic Transport (AST) which can experience complex flow/structure interactions. HSCT can experience vortex induced aeroelastic oscillations whereas AST can experience transonic buffet associated structural oscillations. Both aircraft may experience a dip in the flutter speed at the transonic regime. For accurate aeroelastic computations at these complex fluid/structure interaction situations, high fidelity equations such as the Navier-Stokes for fluids and the finite-elements for structures are needed. Computations using these high fidelity equations require large computational resources both in memory and speed. Current conventional super computers have reached their limitations both in memory and speed. As a result, parallel computers have evolved to overcome the limitations of conventional computers. This paper will address the transition that is taking place in computational aeroelasticity from conventional computers to parallel computers. The paper will address special techniques needed to take advantage of the architecture of new parallel computers. Results will be illustrated from computations made on iPSC/860 and IBM SP2 computer by using ENSAERO code that directly couples the Euler/Navier-Stokes flow equations with high resolution finite-element structural equations.

Guruswamy, Guru↗

A randomized sketching trust-region secant method for low-memory dynamic optimization

The numerical solution of dynamic optimization problems is often limited by the memory required to store the state trajectory, which is used to evaluate the objective function and its derivatives. Recently, [R. Muthukumar et al., SIAM Journal on Optimization 31(2), pp. 1242–1275 (2021)] introduced a trust-region method for dynamic optimization that employs randomized sketching to compress the state trajectory, resulting in inexact derivative computations. By adaptively learning the sketch rank, the trust-region algorithm achieves rigorous convergence guarantees. Here, we extend this approach to use secant Hessian approximations. Due to the randomness introduced by the sketch, the traditional secant update formulae can produce poor Hessian approximations. In particular, the difference of two gradients, computed from two different sketches, may be inconsistent. To overcome this, we employ a sketched approximation of the Hessian application, in lieu of computing the gradient difference. We numerically demonstrate the improved stability of this approach on an example from PDE-constrained optimization.

dynamic optimization↗

Design Space Exploration of Emerging Memory Technologies for Machine Learning Applications

Memory design space exploration methods study memory systems’ performances and limitations before implementation. The computer memory design space has grown exponentially because of the enormous growth of memory types, memory controllers, and application software. Computer simulators are commonly used for memory design space exploration. However, complex memory simulations take an enormous amount of time. Hence, in this paper, we proposed a machine learning-based design space exploration method for dynamic random-access memory and non-volatile memory systems. We applied our method to the CosmoGAN and LeNet applications to predict the following six memory response parameters: (i) bandwidth, (ii) power, (iii) average latency, (iv) average total latency, (v) memory reads, and (vi) memory writes. Our experimental results show that machine learning models can predict memory response parameter values faster than simulations. We used support vector machine, random forest, and gradient boosting machine learning models. We observed that the support vector machine provides better performance for bandwidth, average latency, and average total latency. The random forest model works better for memory reads and writes. The gradient boosting model provides superior prediction performance for power. We provide a detailed discussion on learning curve characteristics, error analysis, and memory type recommendation.

Hasan, S M Shamimul↗

Performance analysis and kernel size study of the Lynx real-time operating system

This paper analyzes the Lynx real-time operating system (LynxOS), which has been selected as the operating system for the Space Station Freedom Data Management System (DMS). The features of LynxOS are compared to other Unix-based operating system (OS). The tools for measuring the performance of LynxOS, which include a high-speed digital timer/counter board, a device driver program, and an application program, are analyzed. The timings for interrupt response, process creation and deletion, threads, semaphores, shared memory, and signals are measured. The memory size of the DMS Embedded Data Processor (EDP) is limited. Besides, virtual memory is not suitable for real-time applications because page swap timing may not be deterministic. Therefore, the DMS software, including LynxOS, has to fit in the main memory of an EDP. To reduce the LynxOS kernel size, the following steps are taken: analyzing the factors that influence the kernel size; identifying the modules of LynxOS that may not be needed in an EDP; adjusting the system parameters of LynxOS; reconfiguring the device drivers used in the LynxOS; and analyzing the symbol table. The reductions in kernel disk size, kernel memory size and total kernel size reduction from each step mentioned above are listed and analyzed.

Liu, Yuan-Kwei↗

Programmable control means for providing safe and controlled medication infusion

An implantable programmable infusion pump (IPIP) is disclosed and generally includes: a fluid reservoir filled with selected medication; a pump for causing a precise volumetric dosage of medication to be withdrawn from the reservoir and delivered to the appropriate site within the body; and, a control means for actuating the pump in a safe and programmable manner. The control means includes a microprocessor, a permanent memory containing a series of fixed software instructions, and a memory for storing prescription schedules, dosage limits and other data. The microprocessor actuates the pump in accordance with programmable prescription parameters and dosage limits stored in the memory. A communication link allows the control means to be remotely programmed. The control means incorporates a running integral dosage limit and other safety features which prevent an inadvertent or intentional medication overdose. The control means also monitors the pump and fluid handling system and provides an alert if any improper or potentially unsafe operation is detected.

Fischell, Robert E.↗

Memory and subjective workload assessment

Recent research suggested subjective introspection of workload is not based upon specific retrieval of information from long term memory, and only reflects the average workload that is imposed upon the human operator by a particular task. These findings are based upon global ratings of workload for the overall task, suggesting that subjective ratings are limited in ability to retrieve specific details of a task from long term memory. To clarify the limits memory imposes on subjective workload assessment, the difficulty of task segments was varied and the workload of specified segments was retrospectively rated. The ratings were retrospectively collected on the manipulations of three levels of segment difficulty. Subjects were assigned to one of two memory groups. In the Before group, subjects knew before performing a block of trials which segment to rate. In the After group, subjects did not know which segment to rate until after performing the block of trials. The subjective ratings, RTs (reaction times) and MTs (movement times) were compared within group, and between group differences. Performance measures and subjective evaluations of workload reflected the experimental manipulations. Subjects were sensitive to different difficulty levels, and recalled the average workload of task components. Cueing did not appear to help recall, and memory group differences possibly reflected variations in the groups of subjects, or an additional memory task.

Staveland, L.↗

HAM: Hotspot-Aware Manager for Improving Communications with 3D-Stacked Memory

merging High-Performance Computing (HPC) workloads, such as graph analytics, machine learning, and big data science, are data-intensive. Data-intensive workloads usually present fine-grained memory accesses with limited or no data locality, and thus incur frequent cache misses and low utilization of memory bandwidth. 3D-stacked memory devices such as Hybrid Memory Cube (HMC) and High Bandwidth Memory (HBM) can provide significantly higher bandwidth than conventional memory modules. However, the traditional interfaces and optimization methods for JEDEC DDR devices do not allow to fully exploit the potential performance of 3D-stacked memory with the massive amount of irregular memory accesses of data-intensive applications. In this paper, we propose a novel Hotspot-Aware Manager (HAM) infrastructure for 3D-stacked memory devices capable of optimizing memory access streams via request aggregation, hotspot detection, and in-memory prefetching. %and an associated hotspot-aware page policy. We present the HAM design and implementation, and simulate it on a system using RISC-V embedded cores with attached HMC devices. We extensively evaluate HAM with over 12 benchmarks and applications representing diverse irregular memory access patterns. The results show that, on average, HAM reduces redundant requests by 37.51\% and increases the prefetch buffer hit rate by 4.2 times, compared to a baseline streaming prefetcher. On the selected benchmark set, HAM provides performance gains of 21.81\% in average (up to 34.28\%) and power savings of 35.07\% over a standard 3D-stacked memory.

Wang, Xi↗

Mapping coastal vegetation, land use and environmental impact from ERTS-1

The author has identified the following significant results. Digital analysis of ERTS-1 imagery was used in an attempt to map and inventory the significant ecological communities of Delaware's coastal zone. Eight vegetation and land use discrimination classes were selected: (1) Phragmites communis (giant reed grass); (2) Spartina alterniflora (salt marsh cord grass); (3) Spartina patens (salt marsh hay); (4) shallow water and exposed mud; (5) deep water (greater than 2 m); (6) forest; (7) agriculture; and (8) exposed sand and concrete. Canonical analysis showed the following classification accuracies: Spartina alterniflora, exposed sand, concrete, and forested land - 94% to 100%; shallow water - mud and deep water - 88% and 93% respectively; Phragmites communis 83%; Spartina patens - 52%. Classification accuracy for agriculture was very poor (51%). Limitations of time and available class-memory space resulted in limiting the analysis of agriculture to very gross identification of a class which actually consists of many varied signature classes. Abundant ground truth was available in the form of vegetation maps compiled from color and color infrared photographs. It is believed that with further refinement of training set selection, sufficiently accurate results can be obtained for all categories.

Klemas, V.↗

The Exploitation of Data Reduction for Visualization

The disparity between the computational speed and storage bandwidth, as demonstrated in Figure 1, is a well known problem that grows with each successive generation. The visualization community is principally responding to this issue by using in situ to reduce which data must be written to storage. However, other communities are taking different, possibly complementary approaches. In particular, data compression is a common general approach to reduce storage demands. Data compression technologies are typically not designed with post processing in mind. The principal metrics measured are compression ratio, the improved bandwidth to storage, and the error introduced. It is assumed that data is inflated to its full size before any post processing can happen. Although when talking about bandwidth disparities, HPC’s dirty little secret is that no part of the memory nor interconnect hardware is increasing at the rate of computation. For example, the Summit supercomputer has a peak computation rate almost 10 times its predecessor, Titan, but only about 4 times the memory, less than twice the aggregate memory bandwidth, and almost no improvement in the interconnect bisection bandwidth. Naively inflating data for post processing does not help with limitations in the memory and interconnect systems.

97 MATHEMATICS AND COMPUTING↗

A DRAM compiler algorithm for high performance VLSI embedded memories

In many applications, the limited density of the embedded SRAM does not allow integrating the memory on the same chip with other logic and functional blocks. In such cases, the embedded DRAM provides the optimum combination of very high density, low power, and high performance. For ASIC's to take full advantage of this design strategy, an efficient and highly reliable DRAM compiler must be used. The embedded DRAM architecture, cell, and peripheral circuit design considerations and the algorithm of a high performance memory compiler are presented .

Eldin, A. G.↗

Advances in relaxation and memory effects of magnetic nanoparticles for biomedical applications

Functionalized magnetic nanoparticles are pivotal in magnetic resonance imaging, computed tomography, controlled drug delivery, and hyperthermia treatments due to their exceptional magnetic relaxation and functional properties. The magnetic core composition and structure significantly affects the complex magnetic properties of these nanoparticles necessitating a thorough examination of magnetism fundamentals related to these systems. One important aspect is the ability of magnetic nanoparticles to retain previous magnetic state configurations known as memory effect, primarily governed by domain structure and magnetic anisotropy. Despite its relevance to advanced applications, comprehensive studies on magnetic relaxation and memory effects remain limited. Here, the present review aims to bridge this gap by investigating relaxation mechanisms, synthesis strategies, and applications, fostering further innovation. It investigates the memory effects and their dependence on particle composition and morphology along with key synthesis techniques for large-scale production in industrial adoption. Structured into focused sections on magnetic properties and their influence on biomedical and technological applications, this review provides essential insights into memory effects, magneto-relaxation mechanisms, influencing factors, and both experimental and theoretical methodologies. It also delves into computational modelling and AI-driven design, which are revolutionizing the prediction, discovery, and optimization of materials with tailored properties.

36 MATERIALS SCIENCE↗

Beyond Binary: Automated PLC Memory Forensics through RGB Image Analysis and Deep Learning

The introduction of Industry 4.0 and the evolution of industrial control systems (ICS) to adopt Internet-based technologies enhanced productivity, but have inadvertently increased their vulnerability to cyber-based malicious attacks. When an ICS system is compromised, security analysts need to identify the root cause quickly to start the recovery process and develop mitigation strategies to safeguard against future instances. Memory forensics is critical in the analysis process to ascertain what occurred. To date, approaches to analyze the persistent memory in ICS devices are limited, and almost nonexistent for volatile memory. This paper proposes an automated methodology, COMA, for PLC memory dump analysis using computer vision and deep learning techniques. Specifically, COMA converts the sequences of bytes in a PLC memory dump to RGB pixels and creates a deep learning model that learns the underlying patterns and features of pre-labeled forensic artifacts in images and segments them into distinct regions. COMA then uses the trained model to automatically segment new memory images and extract forensic artifacts. We evaluate COMA on a Schneider Electric Modicon M221 PLC involving two cyber-based attack scenarios: (i) code injection and (ii) code modification. The empirical results show that COMA can successfully detect attack artifacts in memory dumps in both scenarios.

Asmar Awad, Rima↗

Improving scalability of electronic structure code for molecular simulations in the presence of environment

A scalable density functional electronic code with Gaussian basis set, called UTEP-NRLMOL, is developed to perform simulations of molecular systems in the presence of the environment with particular attention to the memory requirements. In the electronic structure calculations, the memory and computation time are proportional to the number of atoms. Memory requirements for density functional calculations scale as N*N, where N is the number of atoms. While the recent advances in HPC offer platforms with large numbers of cores, the limited amount of memory available on a given node and poor scalability of the electronic structure codes hinder their efficient usage of these platforms. We have introduced new scaling and parallelization paradigms using MPI-3 shared-memory functionality combined with usage of sparse algebra and storage of matrices in sparse format. This extends the range of applicability of the UTEP-NRLMOL code to large systems over 10,000 atoms, or using up to 67,000 basis functions, and making use of HPC architectures using over 6,000 processors utilizing all available cores. We have also interfaced code with effective fragment potential and polarizable continuum model libraries. The code was used in simulations of several applications which are published in reputed scientific journals.

74 ATOMIC AND MOLECULAR PHYSICS↗

Neptune encounter - Guidance and control's finest hour

The Voyager 2 technology had to be stretched significantly to confront the challenges posed by the environment of Neptune. Little was known about Neptune's characteristics because of its great distance from the sun; and its reduced light levels required longer exposure time. Thus, it was necessary to reduce the spacecraft limit cycle; this was made more difficult by the trajectory selected, which caused high spacecraft-to-target rates (these rates dictated the use of image motion compensation). Another challenge faced by Voyager 2 was the need to perform all the required image motion compensation with a limited quantity of attitude control computer and command computer memory, together with a limited quantity of tape recorder resources. This made necessary a new concept of real-time image motion compensation.

Miller, Maurine↗

A scalable superconducting nanowire memory array with row–column addressing

Scalable superconducting memory is required for the development of low-energy superconducting computers and fault-tolerant quantum computers. Conventional superconducting logic-based memory cells possess a large footprint that limits scaling; nanowire-based superconducting memory cells, although more compact, have high error rates, which hinders integration into large arrays. Here we report a 4 × 4 superconducting nanowire memory array that is designed for scalable row–column operations and has a functional density of 2.6 Mbit cm −2 . Each memory cell is based on a nanowire loop consisting of two temperature-dependent superconducting switches and a variable kinetic inductor. The arrays operate at 1.3 K, where we implement and characterize multiflux quanta state storage and destructive read-out. By optimizing the write- and read-pulse sequences, we minimize bit errors and maximize operating margins. We achieve a minimum bit error rate of 10 −5 . Here, we also use circuit-level simulations to understand the memory cell’s dynamics, performance limits and stability under varying pulse amplitudes.

Electrical and electronic engineering↗

Distributed Multi-GPU Community Detection on Exascale Computing Platforms

Community detection is a fundamental operation in graph mining, and by uncovering hidden structures and patterns within complex systems it helps solve fundamental problems pertaining to social networks, such as information diffusion, epidemics, and recommender systems. Scaling graph algorithms for massive networks becomes challenging on modern distributed-memory multi-GPU (Graphics Processing Unit) systems due to limitations such as irregular memory access patterns, load imbalances, higher communication-computation ratios, and cross-platform support. We present a novel algorithm HiPDPL-GPU (distributed parallel Louvain) to address these challenges. We conduct experiments involving different partitioning techniques to achieve optimized performance of HiPDPL-GPU on the two largest supercomputers: Frontier and Summit. Remarkably, HiPDPL-GPU processes a graph with 4.2 billion edges in less than 3 minutes using 1024 GPUs. Qualitatively performance of HiPDPL-GPU is similar or better compared to other state-of-the-art CPU- and GPU-based implementations. While prior GPU implementations have predominantly employed CUDA, our first-of-its-kind implementation for community detection is cross-platform, accommodating both AMD and NVIDIA GPUs.

graph algorithms, high performance comptuing↗