Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “partitioned algorithm”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12

Intelligent System Partitioning for Agent-Based Security Constrained Optimal Power Flow

This project developed scalable, computationally efficient algorithms to solve realistic large-scale power system optimization problems as part of a larger series of competitions run by ARPA-E. These problems are important because the secure and reliable operation of the power grid, especially under increased uncertainty and variability, is growing increasingly challenging. The economic feasibility of the proposed methods developed by our team is quite low, considering it’s a purely software-based solution to operate power grids more efficiently. The technical effectiveness, as evidenced by our performance in the competition, balances heuristics and approximations to provide a tradeoff between speed and accuracy.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Partitioning problems in parallel, pipelined and distributed computing

The problem of optimally assigning the modules of a parallel program over the processors of a multiple computer system is addressed. A Sum-Bottleneck path algorithm is developed that permits the efficient solution of many variants of this problem under some constraints on the structure of the partitions. In particular, the following problems are solved optimally for a single-host, multiple satellite system: partitioning multiple chain structured parallel programs, multiple arbitrarily structured serial programs and single tree structured parallel programs. In addition, the problems of partitioning chain structured parallel programs across chain connected systems and across shared memory (or shared bus) systems are also solved under certain constraints. All solutions for parallel programs are equally applicable to pipelined programs. These results extend prior research in this area by explicitly taking concurrency into account and permit the efficient utilization of multiple computer architectures for a wide range of problems of practical interest.

Bokhari, S.↗

A parallel row-based algorithm with error control for standard-cell replacement on a hypercube multiprocessor

A new row-based parallel algorithm for standard-cell placement targeted for execution on a hypercube multiprocessor is presented. Key features of this implementation include a dynamic simulated-annealing schedule, row-partitioning of the VLSI chip image, and two novel new approaches to controlling error in parallel cell-placement algorithms; Heuristic Cell-Coloring and Adaptive (Parallel Move) Sequence Control. Heuristic Cell-Coloring identifies sets of noninteracting cells that can be moved repeatedly, and in parallel, with no buildup of error in the placement cost. Adaptive Sequence Control allows multiple parallel cell moves to take place between global cell-position updates. This feedback mechanism is based on an error bound derived analytically from the traditional annealing move-acceptance profile. Placement results are presented for real industry circuits and the performance is summarized of an implementation on the Intel iPSC/2 Hypercube. The runtime of this algorithm is 5 to 16 times faster than a previous program developed for the Hypercube, while producing equivalent quality placement. An integrated place and route program for the Intel iPSC/2 Hypercube is currently being developed.

Sargent, Jeff Scott↗

Partitioning problems in parallel, pipelined, and distributed computing

The problem of optimally assigning the modules of a parallel program over the processors of a multiple-computer system is addressed. A sum-bottleneck path algorithm is developed that permits the efficient solution of many variants of this problem under some constraints on the structure of the partitions. In particular, the following problems are solved optimally for a single-host, multiple-satellite system: partitioning multiple chain-structured parallel programs, multiple arbitrarily structured serial programs, and single-tree structured parallel programs. In addition, the problem of partitioning chain-structured parallel programs across chain-connected systems is solved under certain constraints. All solutions for parallel programs are equally applicable to pipelined programs. These results extend prior research in this area by explicitly taking concurrency into account and permit the efficient utilization of multiple-computer architectures for a wide range of problems of practical interest.

Bokhari, Shahid H.↗

Three-Dimensional High-Lift Analysis Using a Parallel Unstructured Multigrid Solver

A directional implicit unstructured agglomeration multigrid solver is ported to shared and distributed memory massively parallel machines using the explicit domain-decomposition and message-passing approach. Because the algorithm operates on local implicit lines in the unstructured mesh, special care is required in partitioning the problem for parallel computing. A weighted partitioning strategy is described which avoids breaking the implicit lines across processor boundaries, while incurring minimal additional communication overhead. Good scalability is demonstrated on a 128 processor SGI Origin 2000 machine and on a 512 processor CRAY T3E machine for reasonably fine grids. The feasibility of performing large-scale unstructured grid calculations with the parallel multigrid algorithm is demonstrated by computing the flow over a partial-span flap wing high-lift geometry on a highly resolved grid of 13.5 million points in approximately 4 hours of wall clock time on the CRAY T3E.

Mavriplis, Dimitri J.↗

Randomized Cholesky Preconditioning for Graph Partitioning Applications

A graph is a mathematical representation of a network; we say it consists of a set of vertices, which are connected by edges. Graphs have numerous applications in various fields, as they can model all sorts of connections, processes, or relations. For example, graphs can model intricate transit systems or the human nervous system. However, graphs that are large or complicated become difficult to analyze. This is why there is an increased interest in the area of graph partitioning, reducing the size of the graph into multiple partitions. For example, partitions of a graph representing a social network might help identify clusters of friends or colleagues. Graph partitioning is also a widely used approach to load balancing in parallel computing. The partitioning of a graph is extremely useful to decompose the graph into smaller parts and allow for easier analysis. There are different ways to solve graph partitioning problems. For this work, we focus on a spectral partitioning method which forms a partition based upon the eigenvectors of the graph Laplacian (details presented in Acer, et. al.). This method uses the LOBPCG algorithm to compute these eigenvectors. LOBPCG can be accelerated by an operator called a preconditioner. For this internship, we evaluate a randomized Cholesky (rchol) preconditioner for its effectiveness on graph partitioning problems with LOBPCG. We compare it with two standard preconditioners: Jacobi and Incomplete Cholesky (ichol). This research was conducted from August to December 2021 in conjunction with Sandia National Laboratories.

97 MATHEMATICS AND COMPUTING↗

Computational algorithms for increased control of depth-viewing volume for stereo three-dimensional graphic displays

Three-dimensional pictorial displays incorporating depth cues by means of stereopsis offer a potential means of presenting information in a natural way to enhance situational awareness and improve operator performance. Conventional computational techniques rely on asymptotic projection transformations and symmetric clipping to produce the stereo display. Implementation of two new computational techniques, as asymmetric clipping algorithm and piecewise linear projection transformation, provides the display designer with more control and better utilization of the effective depth-viewing volume to allow full exploitation of stereopsis cuing. Asymmetric clipping increases the perceived field of view (FOV) for the stereopsis region. The total horizontal FOV provided by the asymmetric clipping algorithm is greater throughout the scene viewing envelope than that of the symmetric algorithm. The new piecewise linear projection transformation allows the designer to creatively partition the depth-viewing volume, with freedom to place depth cuing at the various scene distances at which emphasis is desired.

Williams, Steven P.↗

Computer code for controller partitioning with IFPC application: A user's manual

A user's manual for the computer code for partitioning a centralized controller into decentralized subcontrollers with applicability to Integrated Flight/Propulsion Control (IFPC) is presented. Partitioning of a centralized controller into two subcontrollers is described and the algorithm on which the code is based is discussed. The algorithm uses parameter optimization of a cost function which is described. The major data structures and functions are described. Specific instructions are given. The user is led through an example of an IFCP application.

Schmidt, Phillip H.↗

PLA realizations for VLSI state machines

A major problem associated with state assignment procedures for VLSI controllers is obtaining an assignment that produces minimal or near minimal logic. The key item in Programmable Logic Array (PLA) area minimization is the number of unique product terms required by the design equations. This paper presents a state assignment algorithm for minimizing the number of product terms required to implement a finite state machine using a PLA. Partition algebra with predecessor state information is used to derive a near optimal state assignment. A maximum bound on the number of product terms required can be obtained by inspecting the predecessor state information. The state assignment algorithm presented is much simpler than existing procedures and leads to the same number of product terms or less. An area-efficient PLA structure implemented in a 1.0 micron CMOS process is presented along with a summary of the performance for a controller implemented using this design procedure.

Gopalakrishnan, S.↗

Solving finite element equations on concurrent computers

This paper discusses the development of a concurrent algorithm for the solution of systems of equations arising in finite element applications. The approach is based on a hybrid of direct elimination method and preconditioned conjugate iteration. Two different preconditioners are used; diagonal scaling and a concurrent implementation of incomplete LU factorization. First, an automatic procedure is used to partition the finite element mesh into sub-structures. The particular mesh partition is chosen to minimize an estimate of the cost for evaluating the solution using this algorithm on a concurrent computer. These procedures are implemented in a finite element program on the JPL/CalTech MARK III hypercube computer. An overview of the structure of this program is presented. The performance of the solution method is demonstrated with the aid of a number of numerical test runs, and its advantages for concurrent implementations are discussed. Efficiency and speed-up factors over sequential machines for the numerical examples are highlighted.

Nour-Omid, B.↗

Quantum Simulation of the First-Quantized Pauli-Fierz Hamiltonian

We provide an explicit recursive divide-and-conquer approach for simulating quantum dynamics and derive a discrete first-quantized nonrelativistic QED Hamiltonian based on the many-particle Pauli-Fierz Hamiltonian. We apply this recursive divide-and-conquer algorithm to this Hamiltonian and compare it to a concrete simulation algorithm that uses qubitization. Our divide-and-conquer algorithm, using lowest-order Trotterization, scales for fixed grid spacing as O ~ ( Λ N 2 η 2 t 2 / ϵ ) for grid size N , η particles, simulation time t , field cutoff Λ , and error ϵ . Our qubitization algorithm scales as O ~ ( N ( η + N ) ( η + Λ 2 ) t log ( 1 / ϵ ) ) . This shows that even a naive partitioning and low-order splitting formula can yield, through our divide-and-conquer formalism, superior scaling to qubitization for large Λ . We compare the relative costs of these two algorithms on systems that are relevant for applications such as the spontaneous emission of photons and the photoionization of electrons. We observe that for different parameter regimes, one method can be favored over the other. Finally, we give new algorithmic and circuit-level techniques for gate optimization, including a new way of implementing a group of multicontrolled- X gates that can be used for better analysis of circuit cost. Published by the American Physical Society 2024

Mukhopadhyay, Priyanka (ORCID:0000000164639100)↗

Fast Solution in Sparse LDA for Binary Classification

An algorithm that performs sparse linear discriminant analysis (Sparse-LDA) finds near-optimal solutions in far less time than the prior art when specialized to binary classification (of 2 classes). Sparse-LDA is a type of feature- or variable- selection problem with numerous applications in statistics, machine learning, computer vision, computational finance, operations research, and bio-informatics. Because of its combinatorial nature, feature- or variable-selection problems are NP-hard or computationally intractable in cases involving more than 30 variables or features. Therefore, one typically seeks approximate solutions by means of greedy search algorithms. The prior Sparse-LDA algorithm was a greedy algorithm that considered the best variable or feature to add/ delete to/ from its subsets in order to maximally discriminate between multiple classes of data. The present algorithm is designed for the special but prevalent case of 2-class or binary classification (e.g. 1 vs. 0, functioning vs. malfunctioning, or change versus no change). The present algorithm provides near-optimal solutions on large real-world datasets having hundreds or even thousands of variables or features (e.g. selecting the fewest wavelength bands in a hyperspectral sensor to do terrain classification) and does so in typical computation times of minutes as compared to days or weeks as taken by the prior art. Sparse LDA requires solving generalized eigenvalue problems for a large number of variable subsets (represented by the submatrices of the input within-class and between-class covariance matrices). In the general (fullrank) case, the amount of computation scales at least cubically with the number of variables and thus the size of the problems that can be solved is limited accordingly. However, in binary classification, the principal eigenvalues can be found using a special analytic formula, without resorting to costly iterative techniques. The present algorithm exploits this analytic form along with the inherent sequential nature of greedy search itself. Together this enables the use of highly-efficient partitioned-matrix-inverse techniques that result in large speedups of computation in both the forward-selection and backward-elimination stages of greedy algorithms in general.

Moghaddam, Baback↗

Methodologies and systems for heterogeneous concurrent computing

Heterogeneous concurrent computing is gaining increasing acceptance as an alternative or complementary paradigm to multiprocessor-based parallel processing as well as to conventional supercomputing. While algorithmic and programming aspects of heterogeneous concurrent computing are similar to their parallel processing counterparts, system issues, partitioning and scheduling, and performance aspects are significantly different. In this paper, we discuss critical design and implementation issues in heterogeneous concurrent computing, and describe techniques for enhancing its effectiveness. In particular, we highlight the system level infrastructures that are required, aspects of parallel algorithm development that most affect performance, system capabilities and limitations, and tools and methodologies for effective computing in heterogeneous networked environments. We also present recent developments and experiences in the context of the PVM system and comment on ongoing and future work.

Sunderam, V. S.↗

Improved Multi-Partition Method for Line-Based Iteration Schemes

Regular 3-dimensional multi-partitioning has been shown to be an efficient domain decomposition method for the parallelization of ADI-type algorithms on MIMD architectures. This paper discusses further improvements that can be made to the scheme that increase the granularity and reduce the communication density. These improvements, which are illustrated by simulation and parallel benchmark results, make multi-partitioning the method of choice on systems with relatively poor communication capabilities, such as networks of workstations, or on massively parallel machines with very fast processors, such as the IBM SP2.

Smith, Merritt H.↗

Profile Images and Annotations for Vehicle Re-identification Algorithms (PRIMAVERA)

This dataset contains 636,246 profile images of vehicles representing 13,963 unique vehicles. The data was collected by a set of roadside sensors over the course of three years. Each time a vehicle passed by one of the sensors, a series of images was collected. The images were processed to detect and localize each vehicle, and a license plate reader collocated with the sensor was used to provide a unique ID for the vehicle. Actual license plate numbers have been obfuscated by replacing with an arbitrary numerical ID for each vehicle. After localizing the vehicle in each image, the original RGB image was rotated, scaled, and shifted to produce a new RGB image of size 234x234 pixels such that the outermost two wheels are located at predetermined pixel locations in the image. In this way, all vehicle images are aligned to one another. This registration process occasionally results in a portion of certain vehicles being cutoff at the edges of the image. The dataset has been partitioned into two sets called training and validation. The two partitions no common vehicles, i.e., a vehicle present in one partition is guaranteed not to be present in the other. In this way, an algorithm can be validated against a set of new vehicles that were not seen during the training process. The training set contains 543,926 images from 64,440 vehicle passes representing 11,918 unique vehicles, while the validation set contains 92,320 images from 10,991 vehicle passes representing 2,045 unique vehicles. Vehicle images are organized by directories corresponding to unique vehicles. The file naming scheme is as follows: veh_{vehID}_tr_{passID}_{frameID}_{elevation}_{timeofday}.jpg where {vehID} is the vehicle ID (unique across the entire dataset), {passID} is an identifier for each tracked vehicle pass (unique across the entire dataset), {frameID} is the index of the frame within the given vehicle pass starting at 0, {elevation} is a two-letter string indicating whether the sensor was elevated (el) or at ground-level (gl), and {timeofday} is a two-letter string indicating whether the image was captured during daytime (dt) or nighttime (nt).

image↗

Survivable algorithms and redundancy management in NASA's distributed computing systems

The design of survivable algorithms requires a solid foundation for executing them. While hardware techniques for fault-tolerant computing are relatively well understood, fault-tolerant operating systems, as well as fault-tolerant applications (survivable algorithms), are, by contrast, little understood, and much more work in this field is required. We outline some of our work that contributes to the foundation of ultrareliable operating systems and fault-tolerant algorithm design. We introduce our consensus-based framework for fault-tolerant system design. This is followed by a description of a hierarchical partitioning method for efficient consensus. A scheduler for redundancy management is introduced, and application-specific fault tolerance is described. We give an overview of our hybrid algorithm technique, which is an alternative to the formal approach given.

Malek, Miroslaw↗

Unconditionally stable concurrent procedures for transient finite-element analysis

A family of algorithms was outlined which would appear to be particularly well-suited for implementation in a parallel environment. This is due to the fact that for any partition of the mesh each subdomain in the partition can be processed over a time step simultaneously and independently of the rest. The method eliminates the need for assembling and factorizing large global arrays while retaining the unconditional stability properties of the algorithms used at the local level. To critically appraise the proposed methodology, two limiting cases were considered: element-by-element mesh partitions, and coarse mesh partitions. It was concluded that while the proposed methodology can be useful in sequential machines, it would appear to be promising as it bears on computation. It should also be emphasized that extensions of the method to nonlinear problems are possible.

Ortiz, Michael↗