Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “data movement”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

UPC++ v1.0 Programmer’s Guide, Revision 2023.3.0

UPC++ is a C++ library that supports Partitioned Global Address Space (PGAS) programming. It is designed for writing efficient, scalable parallel programs on distributed-memory parallel computers. The key communication facilities in UPC++ are one-sided Remote Memory Access (RMA) and Remote Procedure Call (RPC). The UPC++ control model is single program, multiple-data (SPMD), with each separate constituent process having access to local memory as it would in C++. The PGAS memory model additionally provides one-sided RMA communication to a global address space, which is allocated in shared segments that are distributed over the processes. UPC++ also features Remote Procedure Call (RPC) communication, making it easy to move computation to operate on data that resides on remote processes. UPC++ was designed to support exascale high-performance computing, and the library interfaces and implementation are focused on maximizing scalability. In UPC++, all communication operations are syntactically explicit, which encourages programmers to consider the costs associated with communication and data movement. Moreover, all communication operations are asynchronous by default, encouraging programmers to seek opportunities for overlapping communication latencies with other useful work. UPC++ provides expressive and composable abstractions designed for efficiently managing aggressive use of asynchrony in programs. Together, these design principles are intended to enable programmers to write applications using UPC++ that perform well even on hundreds of thousands of cores.

97 MATHEMATICS AND COMPUTING↗

UPC++ v1.0 Programmer’s Guide, Revision 2020.3.0

UPC++ is a C++11 library that provides Partitioned Global Address Space (PGAS) programming. It is designed for writing parallel programs that run efficiently and scale well on distributed-memory parallel computers. The PGAS model is single program, multiple-data (SPMD), with each separate constituent process having access to local memory as it would in C++. However, PGAS also provides access to a global address space, which is allocated in shared segments that are distributed over the processes. UPC++ provides numerous methods for accessing and using global memory. In UPC++, all operations that access remote memory are explicit, which encourages programmers to be aware of the cost of communication and data movement. Moreover, all remote-memory access operations are by default asynchronous, to enable programmers to write code that scales well even on hundreds of thousands of cores.

97 MATHEMATICS AND COMPUTING↗

UPC++ v1.0 Programmer’s Guide, Revision 2022.9.0

UPC++ is a C++ library that supports Partitioned Global Address Space (PGAS) programming. It is designed for writing efficient, scalable parallel programs on distributed-memory parallel computers. The key communication facilities in UPC++ are one-sided Remote Memory Access (RMA) and Remote Procedure Call (RPC). The UPC++ control model is single program, multiple-data (SPMD), with each separate constituent process having access to local memory as it would in C++. The PGAS memory model additionally provides one-sided RMA communication to a global address space, which is allocated in shared segments that are distributed over the processes. UPC++ also features Remote Procedure Call (RPC) communication, making it easy to move computation to operate on data that resides on remote processes. UPC++ was designed to support exascale high-performance computing, and the library interfaces and implementation are focused on maximizing scalability. In UPC++, all communication operations are syntactically explicit, which encourages programmers to consider the costs associated with communication and data movement. Moreover, all communication operations are asynchronous by default, encouraging programmers to seek opportunities for overlapping communication latencies with other useful work. UPC++ provides expressive and composable abstractions designed for efficiently managing aggressive use of asynchrony in programs. Together, these design principles are intended to enable programmers to write applications using UPC++ that perform well even on hundreds of thousands of cores.

97 MATHEMATICS AND COMPUTING↗

Enabling Highly Efficient Capsule Networks Processing Through A PIM-Based Architecture Design

In recent years, the CNNs have achieved great successes in the image processing tasks, e.g., image recognition and object detection. Unfortunately, traditional CNN's classication is found to be easily misled by increasingly complex image features due to the usage of pooling operations, hence unable to preserve accurate position and pose information of the objects. To address this challenge, a novel neural network structure called Capsule Network has been proposed, which introduces equivariance through capsules to signicantly enhance the learning ability for image segmentation and object detection. Due to its requirement of performing a high volume of matrix operations, CapsNets have been generally accelerated on modern GPU platforms that provide highly optimized software library for common deep learning tasks. However, based on our performance characterization on modern GPUs, CapsNets exhibit low effciency due to the special program and execution features of their routing procedure, including massive unshareable intermediate variables and intensive syn- chronizations, which are very dicult to optimize at software level. To address these challenges, we propose a hybrid computing architecture design named PIM-CapsNet. It preserves GPU's on-chip computing capability for accelerating CNN types of layers in CapsNet, while pipelining with an off-chip in-memory acceleration solution that effectively tackles routing procedure's ineffciency by leveraging the processing-in-memory capability of today's 3D stacked memory. Using routing procedure's inherent parallellization feature, our design enables hierarchical improvements on CapsNet inference effciency through minimizing data movement and maximizing parallel processing in memory. Evaluation results demonstrate that our proposed design can achieve substantial improvement on both performance and energy savings for CapsNet inference, with almost zero accuracy loss. The results also suggest good performance scalability in optimizing the routing procedure with increasing network size.

Zhang, Xingyao↗

MEPHESTO: Modeling Energy-Performance in Heterogeneous SoCs and Their Trade-Offs

Integrated shared memory heterogeneous architectures are pervasive because they satisfy the diverse needs of mobile, autonomous, and edge computing platforms. Although specialized processing units (PUs) that share a unified system memory improve performance and energy efficiency by reducing data movement, they also increase contention for this memory since the PUs interact with each other. Prior work has investigated performance degradation due to memory contention, but few have studied the relationship of power and energy to memory contention. Moreover, a comprehensive solution that models memory contention for kernel placement on contemporary heterogeneous systems on chip (SoCs) in response to energy and performance has been largely unaddressed.This paper presents MEPHESTO, a novel and holistic approach for managing this balance. The authors characterize applications and PUs in terms of two memory contention factors - time factors and power factors - to achieve the desired trade-off between energy and performance for collocated kernel execution on heterogeneous systems. The authors believe that this investigation is the first to combine all of these factors and present a simple knob-based approach that expresses the target trade-off. The approach is evaluated on a diverse integrated shared memory heterogeneous system with a CPU, GPU, and programmable vision accelerator. By using an empirical model for memory contention that provides up to 92% accuracy, the kernel collocation approach can provide a near-optimal ordering and placement based on the user-defined, energy-performance trade-off parameter. Moreover, the dynamic programming-based heuristics provide up to 30% better energy or 20% performance benefits when compared with the greedy approaches commonly employed by previous studies.

Alaul haque monil, Mohammad↗

Effectively Using Remote I/O For Work Composition in Distributed Workflows

Distributed scientific workflows are becoming more important with the interest in incorporating AI into their loops. A critical programming and performance question is how to compose workflow tasks when data is produced on one system but must be consumed on another. Since the dominant technique is composition with remote I/O, this paper explores its performance expectations. We describe BigFlowSim, a workflow I/O simulator that captures key implementation choices for remote I/O, including intensity, reuse, locality, access pattern, and data movement.With BigFlowSim, we generate a synthetic benchmark. We quantify the effects of each parameter with a performance sensitivity study. We explain trends in terms of data movement reduction and show that, under certain conditions, it is possible to establish a total order among most parameters. We apply these insights to a high energy physics workflow, Belle II Monte Carlo and simulate several I/O optimizations. Speedups range from 5% to 2×, without changing compute time.

Friese, Ryan D.↗

Arrangements for communicating data in a computing system using multiple processors

Systems and methods for reducing data movement in a computer system. The systems and methods use information or knowledge about the structure of an algorithm, operations to be executed at a receiving processing unit, variables or subsets or groups of variables in a distributed algorithm, or other forms of contextual information, for reducing the number of bits transmitted from at least one transmitting processing unit to at least one receiving processing unit or storage device.

Gonzalez, Juan Guillermo↗

Arrangements for communicating and processing data in a computing system

Systems and methods for reducing data movement in a computer system. The systems and methods use information or knowledge about the structure of an algorithm, operations to be executed at a receiving processing unit, variables or subsets or groups of variables in a distributed algorithm, or other forms of contextual information, for reducing the number of bits transmitted from at least one transmitting processing unit to at least one receiving processing unit or storage device.

Gonzalez, Juan Guillermo↗

Cache management based on access type priority

Systems, apparatuses, and methods for cache management based on access type priority are disclosed. A system includes at least a processor and a cache. During a program execution phase, certain access types are more likely to cause demand hits in the cache than others. Demand hits are load and store hits to the cache. A run-time profiling mechanism is employed to find which access types are more likely to cause demand hits. Based on the profiling results, the cache lines that will likely be accessed in the future are retained based on their most recent access type. The goal is to increase demand hits and thereby improve system performance. An efficient cache replacement policy can potentially reduce redundant data movement, thereby improving system performance and reducing energy consumption.

Yin, Jieming↗

Decoding Golden Eagle Movement Behavior from High-Resolution, Variable-Rate Telemetry Data Through Bayesian Filtering

The recent advances in animal tracking technology have enabled the collection of a vast amount of in situ data regarding the movement of wildlife at high spatiotemporal resolution. These data are usually available at variable time resolutions and contains noise (error) originating from GPS fixes. Decoding movement characteristics, particularly of flying animals, from telemetry data while handling these factors is a challenging yet important task for conservation purposes. Typically, this task is broken into two subtasks: resampling, and model calibration. The resampling subtask converts the variable rate positional data into a constant time interval data, while the model calibration subtask uses the resampled data to tune time-invariant parameters of the proposed models. For telemetry data at high temporal resolutions (order of 1 second), it is very challenging to decouple noise from actual movements using interpolation-based resampling techniques. Any errors introduced during resampling can significantly alter the the calibration and prediction attributes of the movement model. We address this problem through a unified Bayesian state-space framework that can handle both the resampling and calibration tasks in a single step. In addition, we use the speed and heading of the bird from telemetry data to regularize the position information of the bird. We use a Kalman filtering approach to include these nonlinearly related motion parameters within the state space framework. We cross-validated to quantify how this inclusion affects the model performance in estimating true bird movements. The relationship between the true state of the bird and environmental and topographical covariates is then represented parametrically. These parameters are then tuned using stochastic sampling strategies like Markov Chain Monte Carlo (MCMC). We use the telemetry data collected from golden eagles in the western USA to demonstrate the applicability of this approach to build a predictive, probabilistic movement model. Our preliminary results show that this approach provides improved predictive performance in terms of capturing higher-order motion parameters such as angular and horizontal accelerations, which may have simpler and more direct relationships with environmental covariates than corresponding speeds. In this talk, we will demonstrate how this state-space approach benefits the prediction capabilities of a movement model in simulating golden eagle paths through a wind power plant in Wyoming given certain atmospheric conditions. The model outcomes are aimed at informing mitigation strategies that can minimize the potential for collisions of golden eagles with wind turbines.

Bayesian methods↗

IRIS-DMEM: Efficient Memory Management for Heterogeneous Computing

This paper proposes an efficient data memory management approach for the Intelligent RuntIme System (IRIS) heterogeneous computing framework along with new data transfer policies. IRIS provides a task-based programming model for extreme heterogeneous computing (e.g., CPU, GPU, DSP, FPGA) with support for today's most important programming languages (e.g., OpenMP, OpenCL, CUDA, HIP, OpenACC). However, the IRIS framework either forces the programmer to introduce data transfer commands for each task or relies on suboptimal memory management for automatic and transparent data transfers. The work described here extends IRIS with novel heterogeneous memory handling and introduces novel data transfer policies by employing the Distributed data MEMory handler (DMEM) for efficient and optimal movement of data among the various computing resources. The proposed approach achieves performance gains of up to 7× for tiled LU factorization and tiled DGEMM (i.e., matrix multiplication) benchmarks. Moreover, this approach also reduces data transfers by up to 71% when compared to previous IRIS heterogeneous memory management handlers. This work compares the performance results of the IRIS framework's novel DMEM with the StarPU runtime and MAGMA math library for GPUs. Experiments show a performance gain of up to 1.95× over StarPU and 2.1× over MAGMA.

Miniskar, Narasinga Rao↗

Characterizing the performance of node-aware strategies for irregular point-to-point communication on heterogeneous architectures

Supercomputer architectures are trending toward higher computational throughput due to the inclusion of heterogeneous compute nodes. These multi-GPU nodes increase on-node computational efficiency, while also increasing the amount of data to be communicated and the number of potential data flow paths. In this work, we characterize the performance of irregular point-to-point communication with MPI on heterogeneous compute environments through performance modeling, demonstrating the limitations of standard communication strategies for both device-aware and staging-through-host communication techniques. Presented models suggest staging communicated data through host processes then using node-aware communication strategies for high inter-node message counts. Notably, the models also predict that node-aware communication utilizing all available CPU cores to communicate inter-node data leads to the most performant strategy when communicating with a high number of nodes. Furthermore, model validation is provided via a case study of irregular point-to-point communication patterns in distributed sparse matrix–vector products. Importantly, we include a discussion on the implications model predictions have on communication strategy design for emerging supercomputer architectures.

97 MATHEMATICS AND COMPUTING↗

Machine learning-based analysis of COVID-19 pandemic impact on US research networks

Here in this study we explore how fallout from the changing public health policy around COVID-19 has changed how researchers access and process their science experiments. Using a combination of techniques from statistical analysis and machine learning, we conduct a retrospective analysis of historical network data for a period around the stay-at-home orders that took place in March 2020. Our analysis takes data from the entire ESnet infrastructure to explore DOE high-performance computing (HPC) resources at OLCF, ALCF, and NERSC, as well as User sites such as PNNL and JLAB. We look at detecting and quantifying changes in site activity using a combination of t-Distributed Stochastic Neighbor Embedding (t-SNE) and decision tree analysis. Our findings bring insights into the working patterns and impact on data volume movements, particularly during late-night hours and weekends.

97 MATHEMATICS AND COMPUTING↗

Optimizing aircraft flows at airports using data driven predicted capabilities

A method for safe and efficient use of airport runway capacity includes receiving, at an air traffic control system at an airport, airport data related to movement areas of the airport, time data related to a time period, aircraft data related to a plurality of aircraft expected to operate into and out of the airport during the time period, and environmental data related to environmental conditions predicted for the airport during the time period. The method further includes computing a probability distribution for inter-aircraft spacing by applying the airport data, the time data, the aircraft data, and the environmental data to a trained Bayesian network, producing the probability distribution for the inter-aircraft spacing as an output observation of the trained Bayesian network, and, using the probability distribution and a confidence value, identifying an inter-aircraft spacing value for the plurality of aircraft expected to operate into and out of the airport during the time period.

Sweet, Douglas↗

Infrared intrusion detection system (IRIDS)

A system and method for intrusion detection includes an imager directed towards an object in an interior space. The imager is in data communication with a computer. The computer is arranged to process digital three-dimensional image data received from the imager and programmed to execute a change detection algorithm in response to the processed three-dimensional data to determine movement of the object. The computer generates an alarm output in response to detecting movement of the object above a predetermined threshold. The method includes providing an imager directed towards an object in an interior space; receiving Time of Flight signals by the imager; processing digital three-dimensional image data received from the imager; and executing a change detection algorithm in response to the processed three-dimensional data to determine movement of the object.

Russell, John L.↗

Understanding collective human movement dynamics during large-scale events using big geosocial data analytics

Conventional approaches for modeling human mobility pattern often focus on human activity and movement dynamics in their regular daily lives and cannot capture changes in human movement dynamics in response to large-scale events. With the rapid advancement of information and communication technologies, many researchers have adopted alternative data sources (e.g., cell phone records, GPS trajectory data) from private data vendors to study human movement dynamics in response to large-scale natural or societal events. Big geosocial data such as georeferenced tweets are publicly available and dynamically evolving as real-world events are happening, making it more likely to capture the real-time sentiments and responses of populations. However, precisely-geolocated geosocial data is scarce and biased toward urban population centers. In this research, we developed a big geosocial data analytical framework for extracting human movement dynamics in response to large-scale events from publicly available georeferenced tweets. The framework includes a two-stage data collection module that collects data in a more targeted fashion in order to mitigate the data scarcity issue of georeferenced tweets; in addition, a variable bandwidth kernel density estimation(VB-KDE) approach was adopted to fuse georeference information at different spatial scales, further augmenting the signals of human movement dynamics contained in georeferenced tweets. To correct for the sampling bias of georeferenced tweets, we adjusted the number of tweets for different spatial units (e.g., county, state) by population. To demonstrate the performance of the proposed analytic framework, we chose an astronomical event that occurred nationwide across the United States, i.e., the 2017 Great American Eclipse, as an example event and studied the human movement dynamics in response to this event. Finally, this analytic framework can easily be applied to other types of large-scale events such as hurricanes or earthquakes.

54 ENVIRONMENTAL SCIENCES↗