Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “hypercube”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Performance of a parallel algorithm for standard cell placement on the Intel Hypercube

A parallel simulated annealing algorithm for standard cell placement that is targeted to run on the Intel Hypercube is presented. A tree broadcasting strategy that is used extensively in our algorithm for updating cell locations in the parallel environment is presented. Studies on the performance of our algorithm on example industrial circuits show that it is faster and gives better final placement results than the uniprocessor simulated annealing algorithms.

Jones, Mark↗

Method and apparatus for implementing a maximum-likelihood decoder in a hypercube network

A method and a structure to implement maximum-likelihood decoding of convolutional codes on a network of microprocessors interconnected as an n-dimensional cube (hypercube). By proper reordering of states in the decoder, only communication between adjacent processors is required. Faster and more efficient operation is enabled, and decoding of large constraint length codes is feasible using standard VLSI technology.

Pollara-Bozzola, Fabrizio↗

Method and apparatus for implementing a traceback maximum-likelihood decoder in a hypercube network

A method and a structure to implement maximum-likelihood decoding of convolutional codes on a network of microprocessors interconnected as an n-dimensional cube (hypercube). By proper reordering of states in the decoder, only communication between adjacent processors is required. Communication time is limited to that required for communication only of the accumulated metrics and not the survivor parameters of a Viterbi decoding algorithm. The survivor parameters are stored at a local processor's memory and a trace-back method is employed to ascertain the decoding result. Faster and more efficient operation is enabled, and decoding of large constraint length codes is feasible using standard VLSI technology.

Pollara-Bozzola, Fabrizio↗

Computer program to minimize prediction error in models from experiments with 16 hypercube points and 0 to 6 center points

A previous report described a backward deletion procedure of model selection that was optimized for minimum prediction error and which used a multiparameter combination of the F - distribution and an order statistics distribution of Cochran's. A computer program is described that applies the previously optimized procedure to real data. The use of the program is illustrated by examples.

Holms, A. G.↗

Experiences with hypercube operating system instrumentation

The difficulties in conceptualizing the interactions among a large number of processors make it difficult both to identify the sources of inefficiencies and to determine how a parallel program could be made more efficient. This paper describes an instrumentation system that can trace the execution of distributed memory parallel programs by recording the occurrence of parallel program events. The resulting event traces can be used to compile summary statistics that provide a global view of program performance. In addition, visualization tools permit the graphic display of event traces. Visual presentation of performance data is particularly useful, indeed, necessary for large-scale parallel computers; the enormous volume of performance data mandates visual display.

Reed, Daniel A.↗

Concurrent hypercube system with improved message passing

A network of microprocessors, or nodes, are interconnected in an n-dimensional cube having bidirectional communication links along the edges of the n-dimensional cube. Each node's processor network includes an I/O subprocessor dedicated to controlling communication of message packets along a bidirectional communication link with each end thereof terminating at an I/O controlled transceiver. Transmit data lines are directly connected from a local FIFO through each node's communication link transceiver. Status and control signals from the neighboring nodes are delivered over supervisory lines to inform the local node that the neighbor node's FIFO is empty and the bidirectional link between the two nodes is idle for data communication. A clocking line between neighbors, clocks a message into an empty FIFO at a neighbor's node and vica versa. Either neighbor may acquire control over the bidirectional communication link at any time, and thus each node has circuitry for checking whether or not the communication link is busy or idle, and whether or not the receive FIFO is empty. Likewise, each node can empty its own FIFO and in turn deliver a status signal to a neighboring node indicating that the local FIFO is empty. The system includes features of automatic message rerouting, block message transfer and automatic parity checking and generation.

Peterson, John C.↗

Concurrent Image Processing Executive (CIPE). Volume 2: Programmer's guide

This manual is intended as a guide for application programmers using the Concurrent Image Processing Executive (CIPE). CIPE is intended to become the support system software for a prototype high performance science analysis workstation. In its current configuration CIPE utilizes a JPL/Caltech Mark 3fp Hypercube with a Sun-4 host. CIPE's design is capable of incorporating other concurrent architectures as well. CIPE provides a programming environment to applications' programmers to shield them from various user interfaces, file transactions, and architectural complexities. A programmer may choose to write applications to use only the Sun-4 or to use the Sun-4 with the hypercube. A hypercube program will use the hypercube's data processors and optionally the Weitek floating point accelerators. The CIPE programming environment provides a simple set of subroutines to activate user interface functions, specify data distributions, activate hypercube resident applications, and to communicate parameters to and from the hypercube.

Williams, Winifred I.↗

Optimal cube-connected cube multiprocessors

Many CFD (computational fluid dynamics) and other scientific applications can be partitioned into subproblems. However, in general the partitioned subproblems are very large. They demand high performance computing power themselves, and the solutions of the subproblems have to be combined at each time step. The cube-connect cube (CCCube) architecture is studied. The CCCube architecture is an extended hypercube structure with each node represented as a cube. It requires fewer physical links between nodes than the hypercube, and provides the same communication support as the hypercube does on many applications. The reduced physical links can be used to enhance the bandwidth of the remaining links and, therefore, enhance the overall performance. The concept and the method to obtain optimal CCCubes, which are the CCCubes with a minimum number of links under a given total number of nodes, are proposed. The superiority of optimal CCCubes over standard hypercubes was also shown in terms of the link usage in the embedding of a binomial tree. A useful computation structure based on a semi-binomial tree for divide-and-conquer type of parallel algorithms was identified. It was shown that this structure can be implemented in optimal CCCubes without performance degradation compared with regular hypercubes. The result presented should provide a useful approach to design of scientific parallel computers.

Sun, Xian-He↗

Concurrent Image Processing Executive (CIPE)

The design and implementation of a Concurrent Image Processing Executive (CIPE), which is intended to become the support system software for a prototype high performance science analysis workstation are discussed. The target machine for this software is a JPL/Caltech Mark IIIfp Hypercube hosted by either a MASSCOMP 5600 or a Sun-3, Sun-4 workstation; however, the design will accommodate other concurrent machines of similar architecture, i.e., local memory, multiple-instruction-multiple-data (MIMD) machines. The CIPE system provides both a multimode user interface and an applications programmer interface, and has been designed around four loosely coupled modules; (1) user interface, (2) host-resident executive, (3) hypercube-resident executive, and (4) application functions. The loose coupling between modules allows modification of a particular module without significantly affecting the other modules in the system. In order to enhance hypercube memory utilization and to allow expansion of image processing capabilities, a specialized program management method, incremental loading, was devised. To minimize data transfer between host and hypercube a data management method which distributes, redistributes, and tracks data set information was implemented.

Lee, Meemong↗

Concurrent Image Processing Executive (CIPE). Volume 1: Design overview

The design and implementation of a Concurrent Image Processing Executive (CIPE), which is intended to become the support system software for a prototype high performance science analysis workstation are described. The target machine for this software is a JPL/Caltech Mark 3fp Hypercube hosted by either a MASSCOMP 5600 or a Sun-3, Sun-4 workstation; however, the design will accommodate other concurrent machines of similar architecture, i.e., local memory, multiple-instruction-multiple-data (MIMD) machines. The CIPE system provides both a multimode user interface and an applications programmer interface, and has been designed around four loosely coupled modules: user interface, host-resident executive, hypercube-resident executive, and application functions. The loose coupling between modules allows modification of a particular module without significantly affecting the other modules in the system. In order to enhance hypercube memory utilization and to allow expansion of image processing capabilities, a specialized program management method, incremental loading, was devised. To minimize data transfer between host and hypercube, a data management method which distributes, redistributes, and tracks data set information was implemented. The data management also allows data sharing among application programs. The CIPE software architecture provides a flexible environment for scientific analysis of complex remote sensing image data, such as planetary data and imaging spectrometry, utilizing state-of-the-art concurrent computation capabilities.

Lee, Meemong↗

Multiphase complete exchange on Paragon, SP2 and CS-2

The overhead of interprocessor communication is a major factor in limiting the performance of parallel computer systems. The complete exchange is the severest communication pattern in that it requires each processor to send a distinct message to every other processor. This pattern is at the heart of many important parallel applications. On hypercubes, multiphase complete exchange has been developed and shown to provide optimal performance over varying message sizes. Most commercial multicomputer systems do not have a hypercube interconnect. However, they use special purpose hardware and dedicated communication processors to achieve very high performance communication and can be made to emulate the hypercube quite well. Multiphase complete exchange has been implemented on three contemporary parallel architectures: the Intel Paragon, IBM SP2 and Meiko CS-2. The essential features of these machines are described and their basic interprocessor communication overheads are discussed. The performance of multiphase complete exchange is evaluated on each machine. It is shown that the theoretical ideas developed for hypercubes are also applicable in practice to these machines and that multiphase complete exchange can lead to major savings in execution time over traditional solutions.

Bokhari, Shahid H.↗

Comparison of multiobjective optimization methods for the $\mathrm{LCLS-II}$ photoinjector

Particle accelerators are among some of the largest science experiments in the world and can consist of thousands of components with a wide variety of input ranges. These systems can easily become unwieldy optimization problems during design and operations studies. Starting in the early 2000s, searching for better beam dynamics configurations became synonymous with heuristic optimization methods in the accelerator physics community. Genetic algorithms and particle swarm optimization are currently the most widely used. These algorithms can take thousands of simulation evaluations to find optimal solutions for one machine prototype. For large facilities such as the Linac Coherent Light Source (LCLS) and others, this equates to a limited exploration of many possible design configurations. In this paper, the LCLS-II photoinjector is optimized with three optimization algorithms. All optimizations were started from both a uniform random and Latin hypercube sample. In all cases, the optimizations started from Latin hypercube samples outperformed optimizations started from uniform samples. All three algorithms were able to optimize the photoinjector, with the model-based methods approximating the Pareto front in fewer simulation evaluations. This work, in combination with previous optimization observations, indicates objective penalties have a strong impact on the efficiency of such methods. In general, we recommend heuristic methods for initial optimizations and model-based methods when information about the objective space is available.

43 PARTICLE ACCELERATORS↗

Multipole groups and fracton phenomena on arbitrary crystalline lattices

Multipole symmetries are of interest in multiple contexts, from the study of fracton phases, to nonergodic quantum dynamics, to the exploration of new hydrodynamic universality classes. However, prior explorations have focused on continuum systems or hypercubic lattices. In this work, we systematically explore multipole symmetries on arbitrary crystal lattices. We explain how, given a crystal structure (specified by a space group and the occupied Wyckoff positions), one may systematically construct all consistent multipole groups. We focus on two-dimensional crystal structures for simplicity, although our methods are general and extend straightforwardly to three dimensions. We classify the possible multipole groups on all two-dimensional Bravais lattices, and on the Kagome and breathing Kagome crystal structures to illustrate the procedure on general crystal lattices. Using Wyckoff positions, we provide an in-principle classification of all possible multipole groups in any space group. We explain how, given a valid multipole group, one may construct a consistent lattice Hamiltonian and a low-energy field theory. We then explore the physical consequences, beginning by generalizing certain results originally obtained on hypercubic lattices to arbitrary crystal structures. Next, we identify two apparently novel phenomena: an emergent, robust subsystem symmetry on the triangular lattice, and an exact multipolar symmetry on the breathing Kagome lattice that does not include conservation of charge (monopole), but instead conserves a vector charge. This makes clear that there is new physics to be found by exploring the consequences of multipolar symmetries on arbitrary lattices, and this work provides the map for the exploration thereof, as well as guiding the search for emergent multipolar symmetries and the attendant exotic phenomena in real materials based on nonhypercubic lattices.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Behavioral Ensemble CLM5 Hydrological Parameter Sets

This repository contains hydrological parameter sets derived using the hybrid regionalization method for three distinct streamflow signatures: Streamflow Signatures: Q10: Represents low flow, indicating the nonexceedance probability of 0.1 for daily streamflow. Q90: Represents high flow, with a nonexceedance probability of 0.9 for daily streamflow. Qmean: Indicates the mean annual flow. Parameters for 464 CAMELS Basins: CAMELS_1000_parameters.csv: Contains 1,000 ensemble parameter sets generated using the Latin hypercube sampling method for CLM5, encompassing 15 hydrological parameters. CAMELS_q10_behavioral_parameter_num.csv: Provides the behavioral ensemble parameter sets for the Q10 streamflow signature for each basin. The associated ID number refers to entries in the CAMELS_1000_parameters.csv file. A minimum of 10 ensemble parameter sets are available for each basin. CAMELS_q90_behavioral_parameter_num.csv: Similar to the above file but for the Q90 streamflow signature. CAMELS_qmean_behavioral_parameter_num.csv: Corresponds to the Qmean streamflow signature, similar to the previous files. Parameters for 50,629 1/8° CONUS Land Grid Cells: CONUS_350_parameters.csv: Contains 350 ensemble parameter sets derived using the Latin hypercube sampling method for CLM5's 15 hydrological parameters within 1/8° CONUS land grid cells. CONUS_q10_behavioral_parameter_num.csv: Holds the behavioral ensemble parameter sets for the Q10 streamflow signature, organized for each grid cell. The ID number relates to entries in CONUS_350_parameters.csv. A minimum of 10 ensemble parameter sets are provided for each grid cell. CONUS_q90_behavioral_parameter_num.csv: Similar to the above file but focusing on the Q90 streamflow signature. CONUS_qmean_behavioral_parameter_num.csv: Corresponds to the Qmean streamflow signature, following a similar structure to the previous files.

Yan, Hongxiang↗

Reduction of the effects of the communication delays in scientific algorithms on message passing MIMD architectures

The efficient implementation of algorithms on multiprocessor machines requires that the effects of communication delays be minimized. The effects of these delays on the performance of a model problem on a hypercube multiprocessor architecture is investigated and methods are developed for increasing algorithm efficiency. The model problem under investigation is the solution by red-black Successive Over Relaxation YOUN71 of the heat equation; most of the techniques described here also apply equally well to the solution of elliptic partial differential equations by red-black or multicolor SOR methods. Methods for reducing communication traffic and overhead on a multiprocessor are identified and results of testing these methods on the Intel iPSC Hypercube reported. Methods for partitioning a problem's domain across processors, for reducing communication traffic during a global convergence check, for reducing the number of global convergence checks employed during an iteration, and for concurrently iterating on multiple time-steps in a time-dependent problem. Empirical results show that use of these models can markedly reduce a numewrical problem's execution time.

Saltz, J. H.↗

Concurrent Cholesky factorization of positive definite banded Hermitian matrices

First, the Cholesky factorization is extended to cover uniformly partitioned banded positive definite matrices of rank n which may be real symmetric or Hermitian. Then, two stratagems are given for the use of the algorithm in concurrent machines where the number of processing elements is less than required to factor the matrix in as few serial steps as possible, and where uniformly high efficiency is expected from all processing elements. Expressions are given for the efficiency factor e appearing in the speed-up expression q = eN, and these are specialized for the N node hypercube machine as a function of partition size s, the number N of processing elements of the hypercube machine, and the cost mu of interelement transmission relative to computation. It is shown that the efficiency factor e is inversely proportional to mu/s, and that e is almost independent of N when N is large and mu/s = 0. The task is completed in n/s serial steps with no limit on n. The half bandwidth b of the matrix is 2 Ns.

Utku, S.↗

Spectral element methods: Algorithms and architectures

Spectral element methods are high-order weighted residual techniques for partial differential equations that combine the geometric flexibility of finite element methods with the rapid convergence of spectral techniques. Spectral element methods are described for the simulation of incompressible fluid flows, with special emphasis on implementation of spectral element techniques on medium-grained parallel processors. Two parallel architectures are considered: the first, a commercially available message-passing hypercube system; the second, a developmental reconfigurable architecture based on Geometry-Defining Processors. High parallel efficiency is obtained in hypercube spectral element computations, indicating that load balancing and communication issues can be successfully addressed by a high-order technique/medium-grained processor algorithm-architecture coupling.

Fischer, Paul↗