Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “partitioned algorithm”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Automated Instrumentation, Monitoring and Visualization of PVM Programs Using AIMS

We present views and analysis of the execution of several PVM (Parallel Virtual Machine) codes for Computational Fluid Dynamics on a networks of Sparcstations, including: (1) NAS Parallel Benchmarks CG and MG; (2) a multi-partitioning algorithm for NAS Parallel Benchmark SP; and (3) an overset grid flowsolver. These views and analysis were obtained using our Automated Instrumentation and Monitoring System (AIMS) version 3.0, a toolkit for debugging the performance of PVM programs. We will describe the architecture, operation and application of AIMS. The AIMS toolkit contains: (1) Xinstrument, which can automatically instrument various computational and communication constructs in message-passing parallel programs; (2) Monitor, a library of runtime trace-collection routines; (3) VK (Visual Kernel), an execution-animation tool with source-code clickback; and (4) Tally, a tool for statistical analysis of execution profiles. Currently, Xinstrument can handle C and Fortran 77 programs using PVM 3.2.x; Monitor has been implemented and tested on Sun 4 systems running SunOS 4.1.2; and VK uses XIIR5 and Motif 1.2. Data and views obtained using AIMS clearly illustrate several characteristic features of executing parallel programs on networked workstations: (1) the impact of long message latencies; (2) the impact of multiprogramming overheads and associated load imbalance; (3) cache and virtual-memory effects; and (4) significant skews between workstation clocks. Interestingly, AIMS can compensate for constant skew (zero drift) by calibrating the skew between a parent and its spawned children. In addition, AIMS' skew-compensation algorithm can adjust timestamps in a way that eliminates physically impossible communications (e.g., messages going backwards in time). Our current efforts are directed toward creating new views to explain the observed performance of PVM programs. Some of the features planned for the near future include: (1) ConfigView, showing the physical topology of the virtual machine, inferred using specially formatted IP (Internet Protocol) packets: and (2) LoadView, synchronous animation of PVM-program execution and resource-utilization patterns.

Mehra, Pankaj↗

Mimas: Preliminary Evidence For Amorphous Water Ice from VIMS

We have conducted a statistical clustering analysis (1,2) on a mosaic of VIMS data cubes obtained on February 13, 2010, for Saturn s satellite Mimas. Seven VIMS cubes were geometrically projected and re-sampled to a common spatial resolution. The clustering technique consists of a partitioning algorithm coupled to a criterion that prevents sub-optimal solutions and tests for the influence of random noise in the measurements. The clustering technique is agnostic about the meaning of the clusters, and scientific interpretation requires their a posteriori evaluation. The preliminary results yielded five clusters, demonstrating that spectral variability across Mimas surface is statistically significant. The ratios of the means calculated for each of the clusters show structure within the 1.6- micron water ice band, as well as the shape and the central wavelength of the strong ice band at 2 micron, that map spatially in patterns apparently related to the topography of Mimas, in particular certain regions in and around Herschel crater. The mean spectra of the five clusters, show similarities with laboratory spectra of amorphous and crystalline H2O ice (3) that are suggestive of the presence of an amorphous ice component in certain regions of Mimas, notably on the central peak of Herschel, on the crater floor, and in faults surrounding the crater. This may represent a mixture of both ice phases, or perhaps a layer of amorphous ice on a base of crystalline ice. Another possible occurrence of amorphous ice appears southwest of Herschel, close to the south pole.

Cruikshank, Dale P.↗

Dynamic Airspace Configuration

In air traffic management systems, airspace is partitioned into regions in part to distribute the tasks associated with managing air traffic among different systems and people. These regions, as well as the systems and people allocated to each, are changed dynamically so that air traffic can be safely and efficiently managed. It is expected that new air traffic control systems will enable greater flexibility in how airspace is partitioned and how resources are allocated to airspace regions. In this talk, I will begin by providing an overview of some previous work and open questions in Dynamic Airspace Configuration research, which is concerned with how to partition airspace and assign resources to regions of airspace. For example, I will introduce airspace partitioning algorithms based on clustering, integer programming optimization, and computational geometry. I will conclude by discussing the development of a tablet-based tool that is intended to help air traffic controller supervisors configure airspace and controllers in current operations.

airspace↗

A Spatiotemporal Indexing Approach for Efficient Processing of Big Array-Based Climate Data with MapReduce

Climate observations and model simulations are producing vast amounts of array-based spatiotemporal data. Efficient processing of these data is essential for assessing global challenges such as climate change, natural disasters, and diseases. This is challenging not only because of the large data volume, but also because of the intrinsic high-dimensional nature of geoscience data. To tackle this challenge, we propose a spatiotemporal indexing approach to efficiently manage and process big climate data with MapReduce in a highly scalable environment. Using this approach, big climate data are directly stored in a Hadoop Distributed File System in its original, native file format. A spatiotemporal index is built to bridge the logical array-based data model and the physical data layout, which enables fast data retrieval when performing spatiotemporal queries. Based on the index, a data-partitioning algorithm is applied to enable MapReduce to achieve high data locality, as well as balancing the workload. The proposed indexing approach is evaluated using the National Aeronautics and Space Administration (NASA) Modern-Era Retrospective Analysis for Research and Applications (MERRA) climate reanalysis dataset. The experimental results show that the index can significantly accelerate querying and processing (10 speedup compared to the baseline test using the same computing cluster), while keeping the index-to-data ratio small (0.0328). The applicability of the indexing approach is demonstrated by a climate anomaly detection deployed on a NASA Hadoop cluster. This approach is also able to support efficient processing of general array-based spatiotemporal data in various geoscience domains without special configuration on a Hadoop cluster.

big data↗

Quantifying Spatial Drought Propagation Potential in North America Using Complex Network Theory

Droughts have a dominant three-dimensional (3-D) spatiotemporal structure typically spanning hundreds of kilometers and often lasting for months to years. Here, we introduced a novel framework to explore the 3-D structure of the evolution of droughts based on network theory concepts. The proposed framework is applied to identify critical source regions responsible for large-scale drought onsets during 1901–2014 for the North American continent using the Standardized Precipitation Evaporation Index (SPEI). We built a spatial network connecting the drought onset timings for the North American continent. Using a spatially weighted network partitioning algorithm, the whole continent is then classified into regional spatial drought networks (RSN), where droughts are more likely to propagate within these regional systems. Finally, a customized network metric was applied to identify locations (source regions) where the drought onsets further propagate to other areas within the regional spatial network. Our results indicated that the West coast, Texas coastal region, and Southeastern Arkansas as major source regions through which atmospheric drought propagates to Western, South Central, and Eastern North America. The formation of drought source regions are due to presence of high pressure ridges and anomalous wind patterns. Furthermore, our results indicate that the drought propagation from these source regions may be due to inadequate moisture transport. The proposed framework can help to develop an early warning detection system for droughts and other spatially extensive extreme events such as heatwaves and floods.

Goutam Konapala↗

Data-driven estimation of energy consumption for electric bus under real-world driving conditions

Reliable and accurate estimation of an electric bus’s instantaneous energy consumption is critical in evaluating energy impacts of planning and control of electric bus operations. In this study, we developed machine learning-based long short-term memory (LSTM) and artificial neural network (ANN) models to estimate 1 Hz energy consumption of electric buses based on continuous monitoring data of electric buses in Chattanooga, Tennessee, in 2019 and 2020. We propose a data-partitioning algorithm to separate energy charging and discharging modes before applying data-driven estimation models. Here, a K-fold cross-validation-based model selection process was conducted to identify the optimal model structure and input variables in terms of prediction accuracy. The estimation results show the predicted mean absolute percentage error rates of LSTM and ANN models were 3% and 5%, respectively. We compared the proposed models with existing models in the literature based on the same testing data to demonstrate the predictability of our models.

Artificial neural network↗

GSplit: Scaling Graph Neural Network Training on Large Graphs via Split-Parallelism

Graph neural networks (GNNs), an emerging class of machine learning models for graphs, have gained popularity for their superior performance in various graph analytical tasks. Mini-batch training is commonly used to train GNNs on large graphs, and data parallelism is the standard approach to scale mini-batch training across multiple GPUs. Data parallel approaches contain redundant work as subgraphs sampled by different GPUs contain significant overlap. To address this issue, we introduce a hybrid parallel mini-batch training paradigm called Split parallelism. Split parallelism avoids redundant work by splitting the sampling, loading, and training of each mini-batch across multiple GPUs. Split parallelism, however, introduces communication overheads that can be more than the savings from removing redundant work. We further present a lightweight partitioning algorithm that probabilistically minimizes these overheads. We implement spllit parllelism in GSplit and show that it outperforms state-of-the-art mini-batch training systems like DGL, Quiver, and P3.

Lim, Seung-Hwan [ORNL] (ORCID:0000000194616866)↗

A parallel evolutionary multiple-try metropolis Markov chain Monte Carlo algorithm for sampling spatial partitions

We develop an Evolutionary Markov Chain Monte Carlo (EMCMC) algorithm for sampling spatial partitions that lie within a large, complex, and constrained spatial state space. Our algorithm combines the advantages of evolutionary algorithms (EAs) as optimization heuristics for state space traversal and the theoretical convergence properties of Markov Chain Monte Carlo algorithms for sampling from unknown distributions. Local optimality information that is identified via a directed search by our optimization heuristic is used to adaptively update a Markov chain in a promising direction within the framework of a Multiple-Try Metropolis Markov Chain model that incorporates a generalized Metropolis-Hastings ratio. We further expand the reach of our EMCMC algorithm by harnessing the computational power afforded by massively parallel computing architecture through the integration of a parallel EA framework that guides Markov chains running in parallel.

97 MATHEMATICS AND COMPUTING↗

CARPE DIEM: Coupled Algorithms for Robust Partitioning of Equations for the Dynamic Interactions of Evolving Materials

From aircraft design to non-proliferation, technical and policy decisions are becoming increasingly reliant on simulation of complex, real-world, multi-physics systems involving multiple interacting domains. As the power of computers has grown, deficiencies associated with traditional low-order-accurate mechanisms for inter-domain coupling have become increasingly apparent. The CARPE DIEM project addressed such deficiencies in the context of fluid-structure interaction (FSI) by developing new algorithms and simulation techniques that are based on a novel and rigorous mathematical approach and that are designed for efficiency on modern high-performance computing platforms.

36 MATERIALS SCIENCE↗

Extraction and classification of objects in multispectral images

Presented here is an algorithm that partitions a digitized multispectral image into parts that correspond to objects in the scene being sensed. The algorithm partitions an image into successively smaller rectangles and produces a partition that tends to minimize a criterion function. Supervised and unsupervised classification techniques can be applied to partitioned images. This partition-then-classify approach is used to process images sensed from aircraft and the ERTS-1 satellite, and the method is shown to give relatively accurate results in classifying agricultural areas and extracting urban areas.

Robertson, T. V.↗

Multi-time-step integration using nodal partitioning

An algorithm is presented which integrates different groups of nodes of a finite element mesh with different time steps and different integrators. Since the nodal groups are updated independently no unsymmetric systems need be solved. Stability is demonstrated by showing that an energy norm of the solution decreases after every update if the time step is less than a given critical value. The element eigenvalue inequality theorem is used to give the critical time step in terms of element eigenvalues.

Smolinski, P.↗

Parallel adaptive mesh refinement within the PUMAA3D Project

To enable the solution of large-scale applications on distributed memory architectures, we are designing and implementing parallel algorithms for the fundamental tasks of unstructured mesh computation. In this paper, we discuss efficient algorithms developed for two of these tasks: parallel adaptive mesh refinement and mesh partitioning. The algorithms are discussed in the context of two-dimensional finite element solution on triangular meshes, but are suitable for use with a variety of element types and with h- or p-refinement. Results demonstrating the scalability and efficiency of the refinement algorithm and the quality of the mesh partitioning are presented for several test problems on the Intel DELTA.

Freitag, Lori↗

Parallel algorithms for placement and routing in VLSI design

The computational requirements for high quality synthesis, analysis, and verification of very large scale integration (VLSI) designs have rapidly increased with the fast growing complexity of these designs. Research in the past has focused on the development of heuristic algorithms, special purpose hardware accelerators, or parallel algorithms for the numerous design tasks to decrease the time required for solution. Two new parallel algorithms are proposed for two VLSI synthesis tasks, standard cell placement and global routing. The first algorithm, a parallel algorithm for global routing, uses hierarchical techniques to decompose the routing problem into independent routing subproblems that are solved in parallel. Results are then presented which compare the routing quality to the results of other published global routers and which evaluate the speedups attained. The second algorithm, a parallel algorithm for cell placement and global routing, hierarchically integrates a quadrisection placement algorithm, a bisection placement algorithm, and the previous global routing algorithm. Unique partitioning techniques are used to decompose the various stages of the algorithm into independent tasks which can be evaluated in parallel. Finally, results are presented which evaluate the various algorithm alternatives and compare the algorithm performance to other placement programs. Measurements are presented on the parallel speedups available.

Brouwer, Randall Jay↗

Evaluating the performance of multicomputer configurations

Steps to optimize the performance of a multicomputer system (MCS) are discussed. Three aspects are emphasized: (1) the interconnection scheme that ties all the processors together, (2) the scheduling and mapping of the algorithm on the architecture, and (3) the mechanism for detecting parallelism and partitioning the algorithm into modules which achieve computational speedup when run on an MCS. Mapping and scheduling issues are addressed, and an application example is given.

Agrawal, D. P.↗

Algorithms for parallel flow solvers on message passing architectures

The purpose of this project has been to identify and test suitable technologies for implementation of fluid flow solvers -- possibly coupled with structures and heat equation solvers -- on MIMD parallel computers. In the course of this investigation much attention has been paid to efficient domain decomposition strategies for ADI-type algorithms. Multi-partitioning derives its efficiency from the assignment of several blocks of grid points to each processor in the parallel computer. A coarse-grain parallelism is obtained, and a near-perfect load balance results. In uni-partitioning every processor receives responsibility for exactly one block of grid points instead of several. This necessitates fine-grain pipelined program execution in order to obtain a reasonable load balance. Although fine-grain parallelism is less desirable on many systems, especially high-latency networks of workstations, uni-partition methods are still in wide use in production codes for flow problems. Consequently, it remains important to achieve good efficiency with this technique that has essentially been superseded by multi-partitioning for parallel ADI-type algorithms. Another reason for the concentration on improving the performance of pipeline methods is their applicability in other types of flow solver kernels with stronger implied data dependence. Analytical expressions can be derived for the size of the dynamic load imbalance incurred in traditional pipelines. From these it can be determined what is the optimal first-processor retardation that leads to the shortest total completion time for the pipeline process. Theoretical predictions of pipeline performance with and without optimization match experimental observations on the iPSC/860 very well. Analysis of pipeline performance also highlights the effect of uncareful grid partitioning in flow solvers that employ pipeline algorithms. If grid blocks at boundaries are not at least as large in the wall-normal direction as those immediately adjacent to them, then the first processor in the pipeline will receive a computational load that is less than that of subsequent processors, magnifying the pipeline slowdown effect. Extra compensation is needed for grid boundary effects, even if all grid blocks are equally sized.

Vanderwijngaart, Rob F.↗