Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “data movement”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12

Enabling Highly Efficient Capsule Networks Processing Through A PIM-Based Architecture Design

In recent years, the CNNs have achieved great successes in the image processing tasks, e.g., image recognition and object detection. Unfortunately, traditional CNN's classication is found to be easily misled by increasingly complex image features due to the usage of pooling operations, hence unable to preserve accurate position and pose information of the objects. To address this challenge, a novel neural network structure called Capsule Network has been proposed, which introduces equivariance through capsules to signicantly enhance the learning ability for image segmentation and object detection. Due to its requirement of performing a high volume of matrix operations, CapsNets have been generally accelerated on modern GPU platforms that provide highly optimized software library for common deep learning tasks. However, based on our performance characterization on modern GPUs, CapsNets exhibit low effciency due to the special program and execution features of their routing procedure, including massive unshareable intermediate variables and intensive syn- chronizations, which are very dicult to optimize at software level. To address these challenges, we propose a hybrid computing architecture design named PIM-CapsNet. It preserves GPU's on-chip computing capability for accelerating CNN types of layers in CapsNet, while pipelining with an off-chip in-memory acceleration solution that effectively tackles routing procedure's ineffciency by leveraging the processing-in-memory capability of today's 3D stacked memory. Using routing procedure's inherent parallellization feature, our design enables hierarchical improvements on CapsNet inference effciency through minimizing data movement and maximizing parallel processing in memory. Evaluation results demonstrate that our proposed design can achieve substantial improvement on both performance and energy savings for CapsNet inference, with almost zero accuracy loss. The results also suggest good performance scalability in optimizing the routing procedure with increasing network size.

Zhang, Xingyao↗

MEPHESTO: Modeling Energy-Performance in Heterogeneous SoCs and Their Trade-Offs

Integrated shared memory heterogeneous architectures are pervasive because they satisfy the diverse needs of mobile, autonomous, and edge computing platforms. Although specialized processing units (PUs) that share a unified system memory improve performance and energy efficiency by reducing data movement, they also increase contention for this memory since the PUs interact with each other. Prior work has investigated performance degradation due to memory contention, but few have studied the relationship of power and energy to memory contention. Moreover, a comprehensive solution that models memory contention for kernel placement on contemporary heterogeneous systems on chip (SoCs) in response to energy and performance has been largely unaddressed.This paper presents MEPHESTO, a novel and holistic approach for managing this balance. The authors characterize applications and PUs in terms of two memory contention factors - time factors and power factors - to achieve the desired trade-off between energy and performance for collocated kernel execution on heterogeneous systems. The authors believe that this investigation is the first to combine all of these factors and present a simple knob-based approach that expresses the target trade-off. The approach is evaluated on a diverse integrated shared memory heterogeneous system with a CPU, GPU, and programmable vision accelerator. By using an empirical model for memory contention that provides up to 92% accuracy, the kernel collocation approach can provide a near-optimal ordering and placement based on the user-defined, energy-performance trade-off parameter. Moreover, the dynamic programming-based heuristics provide up to 30% better energy or 20% performance benefits when compared with the greedy approaches commonly employed by previous studies.

Alaul haque monil, Mohammad↗

Effectively Using Remote I/O For Work Composition in Distributed Workflows

Distributed scientific workflows are becoming more important with the interest in incorporating AI into their loops. A critical programming and performance question is how to compose workflow tasks when data is produced on one system but must be consumed on another. Since the dominant technique is composition with remote I/O, this paper explores its performance expectations. We describe BigFlowSim, a workflow I/O simulator that captures key implementation choices for remote I/O, including intensity, reuse, locality, access pattern, and data movement.With BigFlowSim, we generate a synthetic benchmark. We quantify the effects of each parameter with a performance sensitivity study. We explain trends in terms of data movement reduction and show that, under certain conditions, it is possible to establish a total order among most parameters. We apply these insights to a high energy physics workflow, Belle II Monte Carlo and simulate several I/O optimizations. Speedups range from 5% to 2×, without changing compute time.

Friese, Ryan D.↗

Arrangements for communicating data in a computing system using multiple processors

Systems and methods for reducing data movement in a computer system. The systems and methods use information or knowledge about the structure of an algorithm, operations to be executed at a receiving processing unit, variables or subsets or groups of variables in a distributed algorithm, or other forms of contextual information, for reducing the number of bits transmitted from at least one transmitting processing unit to at least one receiving processing unit or storage device.

Gonzalez, Juan Guillermo↗

Arrangements for communicating and processing data in a computing system

Systems and methods for reducing data movement in a computer system. The systems and methods use information or knowledge about the structure of an algorithm, operations to be executed at a receiving processing unit, variables or subsets or groups of variables in a distributed algorithm, or other forms of contextual information, for reducing the number of bits transmitted from at least one transmitting processing unit to at least one receiving processing unit or storage device.

Gonzalez, Juan Guillermo↗

Cache management based on access type priority

Systems, apparatuses, and methods for cache management based on access type priority are disclosed. A system includes at least a processor and a cache. During a program execution phase, certain access types are more likely to cause demand hits in the cache than others. Demand hits are load and store hits to the cache. A run-time profiling mechanism is employed to find which access types are more likely to cause demand hits. Based on the profiling results, the cache lines that will likely be accessed in the future are retained based on their most recent access type. The goal is to increase demand hits and thereby improve system performance. An efficient cache replacement policy can potentially reduce redundant data movement, thereby improving system performance and reducing energy consumption.

Yin, Jieming↗

Eye/Brain/Task Testbed And Software

Eye/brain/task (EBT) testbed records electroencephalograms, movements of eyes, and structures of tasks to provide comprehensive data on neurophysiological experiments. Intended to serve continuing effort to develop means for interactions between human brain waves and computers. Software library associated with testbed provides capabilities to recall collected data, to process data on movements of eyes, to correlate eye-movement data with electroencephalographic data, and to present data graphically. Cognitive processes investigated in ways not previously possible.

Janiszewski, Thomas↗

Manipulator system performance measurement

Free-flying teleoperator technology for satellite servicing is reported. The rationale for a series of manipulator system tests is presented. Data are reported on movement time in a fine positioning task using two different manipulator systems. The movement time data showed reliable effects of movement direction and index of difficulty. These data were considered to be baseline performance measures for the manipulator system used and modifications to the control and visual systems were suggested for future testing.

Kirkpatrick, M., III↗

Application of Optical Measurement Techniques During Stages of Pregnancy: Use of Phantom High Speed Cameras for Digital Image Correlation (D.I.C.) During Baby Kicking and Abdomen Movements

Paired images were collected using a projected pattern instead of standard painting of the speckle pattern on her abdomen. High Speed cameras were post triggered after movements felt. Data was collected at 120 fps -limited due to 60hz frequency of projector. To ensure that kicks and movement data was real a background test was conducted with no baby movement (to correct for breathing and body motion).

Gradl, Paul↗

Motion analysis report

Human motion analysis is the task of converting actual human movements into computer readable data. Such movement information may be obtained though active or passive sensing methods. Active methods include physical measuring devices such as goniometers on joints of the body, force plates, and manually operated sensors such as a Cybex dynamometer. Passive sensing de-couples the position measuring device from actual human contact. Passive sensors include Selspot scanning systems (since there is no mechanical connection between the subject's attached LEDs and the infrared sensing cameras), sonic (spark-based) three-dimensional digitizers, Polhemus six-dimensional tracking systems, and image processing systems based on multiple views and photogrammetric calculations.

Badler, N. I.↗

Numerical Models For Control Of Robots

Algorithm develops numerical models of kinematics of robots for use in directing movements of robots. Based on empirical data. Predicts movements from previous measurements of actual movements. Replaces analytical or iterative models used commonly.

Waggener, Mary S.↗

Decoding Golden Eagle Movement Behavior from High-Resolution, Variable-Rate Telemetry Data Through Bayesian Filtering

The recent advances in animal tracking technology have enabled the collection of a vast amount of in situ data regarding the movement of wildlife at high spatiotemporal resolution. These data are usually available at variable time resolutions and contains noise (error) originating from GPS fixes. Decoding movement characteristics, particularly of flying animals, from telemetry data while handling these factors is a challenging yet important task for conservation purposes. Typically, this task is broken into two subtasks: resampling, and model calibration. The resampling subtask converts the variable rate positional data into a constant time interval data, while the model calibration subtask uses the resampled data to tune time-invariant parameters of the proposed models. For telemetry data at high temporal resolutions (order of 1 second), it is very challenging to decouple noise from actual movements using interpolation-based resampling techniques. Any errors introduced during resampling can significantly alter the the calibration and prediction attributes of the movement model. We address this problem through a unified Bayesian state-space framework that can handle both the resampling and calibration tasks in a single step. In addition, we use the speed and heading of the bird from telemetry data to regularize the position information of the bird. We use a Kalman filtering approach to include these nonlinearly related motion parameters within the state space framework. We cross-validated to quantify how this inclusion affects the model performance in estimating true bird movements. The relationship between the true state of the bird and environmental and topographical covariates is then represented parametrically. These parameters are then tuned using stochastic sampling strategies like Markov Chain Monte Carlo (MCMC). We use the telemetry data collected from golden eagles in the western USA to demonstrate the applicability of this approach to build a predictive, probabilistic movement model. Our preliminary results show that this approach provides improved predictive performance in terms of capturing higher-order motion parameters such as angular and horizontal accelerations, which may have simpler and more direct relationships with environmental covariates than corresponding speeds. In this talk, we will demonstrate how this state-space approach benefits the prediction capabilities of a movement model in simulating golden eagle paths through a wind power plant in Wyoming given certain atmospheric conditions. The model outcomes are aimed at informing mitigation strategies that can minimize the potential for collisions of golden eagles with wind turbines.

Bayesian methods↗

IRIS-DMEM: Efficient Memory Management for Heterogeneous Computing

This paper proposes an efficient data memory management approach for the Intelligent RuntIme System (IRIS) heterogeneous computing framework along with new data transfer policies. IRIS provides a task-based programming model for extreme heterogeneous computing (e.g., CPU, GPU, DSP, FPGA) with support for today's most important programming languages (e.g., OpenMP, OpenCL, CUDA, HIP, OpenACC). However, the IRIS framework either forces the programmer to introduce data transfer commands for each task or relies on suboptimal memory management for automatic and transparent data transfers. The work described here extends IRIS with novel heterogeneous memory handling and introduces novel data transfer policies by employing the Distributed data MEMory handler (DMEM) for efficient and optimal movement of data among the various computing resources. The proposed approach achieves performance gains of up to 7× for tiled LU factorization and tiled DGEMM (i.e., matrix multiplication) benchmarks. Moreover, this approach also reduces data transfers by up to 71% when compared to previous IRIS heterogeneous memory management handlers. This work compares the performance results of the IRIS framework's novel DMEM with the StarPU runtime and MAGMA math library for GPUs. Experiments show a performance gain of up to 1.95× over StarPU and 2.1× over MAGMA.

Miniskar, Narasinga Rao↗

Lunar Relay Satellite Network for Space Exploration: Architecture, Technologies and Challenges

NASA is planning a series of short and long duration human and robotic missions to explore the Moon and then Mars. A key objective of these missions is to grow, through a series of launches, a system of systems infrastructure with the capability for safe and sustainable autonomous operations at minimum cost while maximizing the exploration capabilities and science return. An incremental implementation process will enable a buildup of the communication, navigation, networking, computing, and informatics architectures to support human exploration missions in the vicinities and on the surfaces of the Moon and Mars. These architectures will support all space and surface nodes, including other orbiters, lander vehicles, humans in spacesuits, robots, rovers, human habitats, and pressurized vehicles. This paper describes the integration of an innovative MAC and networking technology with an equally innovative position-dependent, data routing, network technology. The MAC technology provides the relay spacecraft with the capability to autonomously discover neighbor spacecraft and surface nodes, establish variable-rate links and communicate simultaneously with multiple in-space and surface clients at varying and rapidly changing distances while making optimum use of the available power. The networking technology uses attitude sensors, a time synchronization protocol and occasional orbit-corrections to maintain awareness of its instantaneous position and attitude in space as well as the orbital or surface location of its communication clients. A position-dependent data routing capability is used in the communication relay satellites to handle the movement of data among any of multiple clients (including Earth) that may be simultaneously in view; and if not in view, the relay will temporarily store the data from a client source and download it when the destination client comes into view. The integration of the MAC and data routing networking technologies would enable a relay satellite system to provide end-to-end communication services for robotic and human missions in the vicinity, or on the surface of the Moon with a minimum of Earth-based operational support.

Bhasin, Kul B.↗

Non-Repudiation for Drone-Related Data

Concepts for the management of Uncrewed Aircraft Systems (UAS) at scale rely on the exchange of data amongst multiple stakeholders. Even as these concepts vary across nations and industries, the movement of data between entities is a common theme. While there is universal agreement on the necessity of appropriate cybersecurity measures to address data communication, there has been minimal focus on the feasibility of implementing non-repudiation solutions for UAS systems. This means that data exchanged in support of UAS operations are open to “attack” via parties that may deny sending or receiving certain data, which can weaken the effectiveness and acceptability of these systems. This paper highlights the current and future need for non-repudiation, supported by references to multiple international organizations, and an approach to implementing non-repudiation leveraging open standards.

drone↗

Understanding the Effects of Spaceflight on Head-trunk Coordination during Walking and Obstacle Avoidance

Prolonged exposure to spaceflight conditions results in a battery of physiological changes, some of which contribute to sensorimotor and neurovestibular deficits. Upon return to Earth, functional performance changes are tested using the Functional Task Test (FTT), which includes an obstacle course to observe post‐flight balance and postural stability, specifically during turning. The goal of this study was to quantify changes in movement strategies during turning events by observing the latency between head‐and‐trunk coordinated movements. It was hypothesized that subjects experiencing neurovestibular adaptations would exhibit head‐to‐trunk locking ('en bloc' movement) during turning, exhibited by a decrease in latency between head and trunk movement. FTT data samples were collected from ISS missions. Samples were analyzed three times pre‐exposure, immediately postexposure (1 day post) and 2‐to‐3 times during recovery from the microgravity environment. Two 3D inertial measurements units (XSens MTx) were attached to subjects, one on the head and one on the upper back. This study focused primarily on the yaw movements about the subject's center of rotation. Time differences (latency) between head and trunk movement were calculated at two points on the obstacle course: the first turn to enter the obstacle course (approximately 90⁰ turn) and averaged across a slalom obstacle portion, consisting of three turns (approximately three 90⁰ turns). Preliminary analysis of the data shows a trend toward decreasing head‐to‐trunk movement latency during postflight ambulation in slalom turning after reintroduction to Earth gravity in ISS astronauts. It is clear that changes in movement strategies are adopted during exposure to the microgravity environment and upon reintroduction to a gravity environment. Most ISS subjects exhibit symptoms of neurovestibular changes ('en bloc head and trunk movement) which may impact their ability to perform post‐flight functional tasks.

Madansingh, S.↗