Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “partitioned algorithm”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Dynamics of Quantum Adiabatic Evolution Algorithm for Number Partitioning

We have developed a general technique to study the dynamics of the quantum adiabatic evolution algorithm applied to random combinatorial optimization problems in the asymptotic limit of large problem size n. We use as an example the NP-complete Number Partitioning problem and map the algorithm dynamics to that of an auxiliary quantum spin glass system with the slowly varying Hamiltonian. We use a Green function method to obtain the adiabatic eigenstates and the minimum exitation gap, gmin = O(n2(sup -n/2)), corresponding to the exponential complexity of the algorithm for Number Partitioning. The key element of the analysis is the conditional energy distribution computed for the set of all spin configurations generated from a given (ancestor) configuration by simultaneous flipping of a fixed number of spins. For the problem in question this distribution is shown to depend on the ancestor spin configuration only via a certain parameter related to the energy of the configuration. As the result, the algorithm dynamics can be described in terms of one-dimensional quantum diffusion in the energy space. This effect provides a general limitation of a quantum adiabatic computation in random optimization problems. Analytical results are in agreement with the numerical simulation of the algorithm.

Smelyanskiy, Vadius↗

Dynamics of Quantum Adiabatic Evolution Algorithm for Number Partitioning

We have developed a general technique to study the dynamics of the quantum adiabatic evolution algorithm applied to random combinatorial optimization problems in the asymptotic limit of large problem size n. We use as an example the NP-complete Number Partitioning problem and map the algorithm dynamics to that of an auxiliary quantum spin glass system with the slowly varying Hamiltonian. We use a Green function method to obtain the adiabatic eigenstates and the minimum excitation gap. g min, = O(n 2(exp -n/2), corresponding to the exponential complexity of the algorithm for Number Partitioning. The key element of the analysis is the conditional energy distribution computed for the set of all spin configurations generated from a given (ancestor) configuration by simultaneous flipping of a fixed number of spins. For the problem in question this distribution is shown to depend on the ancestor spin configuration only via a certain parameter related to 'the energy of the configuration. As the result, the algorithm dynamics can be described in terms of one-dimensional quantum diffusion in the energy space. This effect provides a general limitation of a quantum adiabatic computation in random optimization problems. Analytical results are in agreement with the numerical simulation of the algorithm.

Smelyanskiy, V. N.↗

A parallel algorithm for switch-level timing simulation on a hypercube multiprocessor

The parallel approach to speeding up simulation is studied, specifically the simulation of digital LSI MOS circuitry on the Intel iPSC/2 hypercube. The simulation algorithm is based on RSIM, an event driven switch-level simulator that incorporates a linear transistor model for simulating digital MOS circuits. Parallel processing techniques based on the concepts of Virtual Time and rollback are utilized so that portions of the circuit may be simulated on separate processors, in parallel for as large an increase in speed as possible. A partitioning algorithm is also developed in order to subdivide the circuit for parallel processing.

Rao, Hariprasad Nannapaneni↗

A Linear-Complexity Tensor Butterfly Algorithm for Compressing High-Dimensional Oscillatory Integral Operators

This paper presents a multilevel tensor compression algorithm called tensor butterfly algorithm for efficiently representing large-scale and high-dimensional oscillatory integral operators, including Green's functions for wave equations and integral transforms such as Radon transforms and Fourier transforms. The proposed algorithm leverages a tensor extension of the so-called complementary low-rank property of existing matrix butterfly algorithms. The algorithm partitions the discretized integral operator tensor into subtensors of multiple levels and factorizes each subtensor at the middle level as a Tucker-type interpolative decomposition, whose factor matrices are formed in a multilevel fashion. For a d-dimensional (d > 1) integral operator discretized into a 2d-mode tensor with n2d entries, the overall CPU time and memory requirement scale as O(nd), in stark contrast to the O(nd log n) complexity of existing matrix algorithms such as matrix butterfly algorithms and fast Fourier transforms (FFTs), where n is the number of points per direction. When comparing with other tensor algorithms such as quantized tensor train (QTT), the proposed algorithm also shows superior CPU and memory performance for tensor contraction. Remarkably, the tensor butterfly algorithm can efficiently model high-frequency Green's function interactions between two unit cubes, each spanning 512 wavelengths per direction, which represents problems of scale over 512× larger than that existing butterfly algorithms can handle, with the same amount of computation resources. On the other hand, for a problem representing 64 wavelengths per direction, which is the largest size existing algebraic matrix algorithms can handle, our tensor butterfly algorithm exhibits 200x speedups and 30× memory reduction compared with existing ones. Moreover, the tensor butterfly algorithm also permits O(nd)-complexity FFTs and Radon transforms up to d = 6 dimensions.

Kielstra, P Michael↗

Evaluation of a dual processor implementation for a fault inferring nonlinear detection system

The design of a modified fault inferring nonlinear detection system (FINDS) algorithm for a dual-processor configured flight computer is described. The algorithm was changed in order to divide it into its translational dynamics and rotational kinematics and to use it for parallel execution on the flight computer. The FINDS consists of: (1) a no-fail filter (NFF), (2) a set of test-of-mean detection tests, (3) a bank of first order filters to estimate failure levels in individual sensors, and (4) a decision function. NFF filter performance using flight recorded sensor data is analyzed using a filter autoinitialization routine. The failure detection and isolation capability of the partitioned algorithm is evaluated. A multirate implementation for the bias-free and bias filter gain and covariance matrices is discussed.

Godiwala, P. M.↗

Aboveground and belowground contributions to ecosystem respiration in a temperate deciduous forest

In this study, we developed a three-way carbon dioxide (CO 2 ) flux-partitioning algorithm that separates net ecosystem exchange (NEE) into aboveground plant respiration (R above ), belowground root and soil respiration (R below ), and gross primary production (GPP). We applied this algorithm to a coupled dataset of continuous chamber-measured soil respiration and eddy covariance (EC)-measured NEE of CO 2 in an oak-hickory (Quercus-Carya) deciduous broadleaf forest from 2006 to 2015. We found that on annual time scale, R below dominated over R above with the former accounting for 66.9–86.4% and the latter 13.6–33.1%, of the total ecosystem respiration (R eco ). The ratio of R below to R above varied seasonally, ranging from 1.77 to 7.25 in growing season, and 1.02 to 4.57 in non-growing season. The temperature sensitivity (E 0 ) of R below was significantly higher than that of R above , and E 0 of R eco responded differently to air and soil temperature. Over the whole study period, annual mean R above , R below , and GPP were 243, 806, and 1170 g C m –2 , respectively, with annual R eco accounting for 89.6% of GPP, of which 68.8% was lost as R below and 20.8% lost as R above , and leaving only 10% of the carbon fixation in ecosystems. Furthermore, these estimates, however, did not consider potential light inhibition of leaf respiration. If we accept the presence of light inhibition, then the daytime three-way partitioning method would underestimate annual R above by 20.4% whereas the nighttime method would overestimate R above by 23.9% and GPP by 4.7%, compared with estimates accounting for light inhibition in leaves.

54 ENVIRONMENTAL SCIENCES↗

A Data-Driven Approach to Nation-Scale Building Energy Modeling

In 2019, 125 million U.S. residential and commercial buildings consumed $412 billion in energy bills. These buildings currently consume 40% of the nation's primary energy, 73% of electricity, 80% of energy during peak electric grid use, and responsible for 39% of greenhouse gas emissions [14]. Urban-scale building energy modeling has grown significantly in the past decade, allowing individual campuses or communities of buildings to be modeled, simulated, and cost-effective solutions for intelligent management to be identified and implemented. While traditionally limited to individual counties and usually less than 2,000 buildings, the Automatic Building Energy Modeling (AutoBEM) soft-ware suite has been developed to process unconventional, nation-scale data sources to generate unique OpenStudio and EnergyPlus models of each building. Through the use of High Performance Computing (HPC) resources, every U.S. building has been simulated. This paper showcases the data layout, node partitioning, algorithmic approaches, and analytic results that were used to create, share, and analyze 124.4 million U.S. building models.

Berres, Andy↗

Vertically Resolved Convective–Stratiform Echo-Type Identification and Convectivity Retrieval for Vertically Pointing Radars

Using data from the airborne HIAPER Cloud Radar (HCR), a partitioning algorithm (ECCO-V) that provides vertically resolved convectivity and convective versus stratiform radar-echo classification is developed for vertically pointing radars. The algorithm is based on the calculation of reflectivity and radial velocity texture fields that measure the horizontal homogeneity of cloud and precipitation features. The texture fields are translated into convectivity, a numerical measure of the convective or stratiform nature of each data point. The convective–stratiform classification is obtained by thresholding the convectivity field. Subcategories of low, mid-, and high stratiform, shallow, mid-, deep, and elevated convective, and mixed echoes are introduced, which are based on the melting-layer and divergence-level altitudes. As the algorithm provides vertically resolved classifications, it is capable of identifying different types of vertically layered echoes, and convective features that are embedded in stratiform cloud layers. Its robustness was tested on data from four HCR field campaigns that took place in different meteorological and climatological regimes. The algorithm was adapted for use in spaceborne and ground-based radars, proving its versatility, as it is adaptable not only to different radar types and wavelengths, but also different research applications.

54 ENVIRONMENTAL SCIENCES↗

Processing Particle Data Flows with SmartNICs

Many distributed applications implement complex data flows and need a flexible mechanism for routing data between producers and consumers. Recent advances in programmable network interface cards, or SmartNICs, represent an opportunity to offload data-flow tasks into the network fabric, thereby freeing the hosts to perform other work. System architects in this space face multiple questions about the best way to leverage SmartNICs as processing elements in data flows. In this paper, we advocate the use of Apache Arrow as a foundation for implementing data-flow tasks on SmartNICs. We report on our experiences adapting a partitioning algorithm for particle data to Apache Arrow and measure the on-card processing performance for the BlueField-2 SmartNIC. Our experiments confirm that the BlueField-2’s (de)compression hardware can have a significant impact on in-transit workflows where data must be unpacked, processed, and repacked.

97 MATHEMATICS AND COMPUTING↗

Sequential square root filtering and smoothing of discrete linear systems

A square root information filter/smoother is derived using recursive least-squares arguments. The combined filter/smoother algorithm has the following attributes: (1) it has a square root structure, which enhances numerical accuracy; (2) filter and smoother mechanizations are identical in form, facilitating implementation of the smoother; and (3) storage and computation requirements are modest compared with other smoothing algorithms. Partitioning the results to separate bias parameters provides further computational economies and reduction of storage requirements.

Bierman, G. J.↗

A compact high-speed parallel multiplication scheme

This paper discusses a compact, fast, parallel multiplication scheme of the generation-reduction type using generalized Dadda-type pseudoadders for reduction and m x m multipliers for generation. The implications of present and future LSI are considered, a partitioning algorithm is presented, and the results obtained for a 24 x 24-bit implementation are discussed.

Stenzel, W. J.↗

The ERODYN and QRPIG computer programs

The role of the ERODYN computer program in providing error analyses involving orbital, geodetic, and geophysical parameters is discussed. It was designed to operate as a companion program to the GEODYN orbit determination and parameter estimating program. The Q R Partitioned Eigenvalue/Eigenvector analysis program (QRPIG) is designed to process symmetric matrices with an out-of-core partitioning algorithm for the eigenvectors and eigenvalues of the matrices.

Felsentreger, T. L.↗

Simulating a small turboshaft engine in real-time multiprocessor simulator (RTMPS) environment

A Real-Time Multiprocessor Simulator (RTMPS) has been developed at NASA Lewis Research Center. The RTMPS uses parallel microprocessors to achieve computing speeds needed for real-time engine simulation. This report describes the use of the RTMPS system to simulate a small turboshaft engine. The process of programming the engine equations and distributing them over one, two, and four processors is discussed. Steady-state and transient results from the RTMPS simulation are compared with results from a main-frame-based simulation. Processor execution times and the associated execution time savings for the two and four processor cases are presented using actual data obtained from the RTMPS system. Included is a discussion of why the minimum achievable calculation time for the turboshaft engine model was attained using four processors. Finally, future enhancements to the RTMPS system are discussed including the development of a generalized partitioning algorithm to automatically distribute the system equations among the processors in optimum fashion.

Milner, E. J.↗

Parallel processing for digital picture comparison

In picture processing an important problem is to identify two digital pictures of the same scene taken under different lighting conditions. This kind of problem can be found in remote sensing, satellite signal processing and the related areas. The identification can be done by transforming the gray levels so that the gray level histograms of the two pictures are closely matched. The transformation problem can be solved by using the packing method. Researchers propose a VLSI architecture consisting of m x n processing elements with extensive parallel and pipelining computation capabilities to speed up the transformation with the time complexity 0(max(m,n)), where m and n are the numbers of the gray levels of the input picture and the reference picture respectively. If using uniprocessor and a dynamic programming algorithm, the time complexity will be 0(m(3)xn). The algorithm partition problem, as an important issue in VLSI design, is discussed. Verification of the proposed architecture is also given.

Cheng, H. D.↗

Efficient bulk-loading of gridfiles

This paper considers the problem of bulk-loading large data sets for the gridfile multiattribute indexing technique. We propose a rectilinear partitioning algorithm that heuristically seeks to minimize the size of the gridfile needed to ensure no bucket overflows. Empirical studies on both synthetic data sets and on data sets drawn from computational fluid dynamics applications demonstrate that our algorithm is very efficient, and is able to handle large data sets. In addition, we present an algorithm for bulk-loading data sets too large to fit in main memory. Utilizing a sort of the entire data set it creates a gridfile without incurring any overflows.

Leutenegger, Scott T.↗

PLUM: Parallel Load Balancing for Unstructured Adaptive Meshes

Dynamic mesh adaption on unstructured grids is a powerful tool for computing large-scale problems that require grid modifications to efficiently resolve solution features. Unfortunately, an efficient parallel implementation is difficult to achieve, primarily due to the load imbalance created by the dynamically-changing nonuniform grid. To address this problem, we have developed PLUM, an automatic portable framework for performing adaptive large-scale numerical computations in a message-passing environment. First, we present an efficient parallel implementation of a tetrahedral mesh adaption scheme. Extremely promising parallel performance is achieved for various refinement and coarsening strategies on a realistic-sized domain. Next we describe PLUM, a novel method for dynamically balancing the processor workloads in adaptive grid computations. This research includes interfacing the parallel mesh adaption procedure based on actual flow solutions to a data remapping module, and incorporating an efficient parallel mesh repartitioner. A significant runtime improvement is achieved by observing that data movement for a refinement step should be performed after the edge-marking phase but before the actual subdivision. We also present optimal and heuristic remapping cost metrics that can accurately predict the total overhead for data redistribution. Several experiments are performed to verify the effectiveness of PLUM on sequences of dynamically adapted unstructured grids. Portability is demonstrated by presenting results on the two vastly different architectures of the SP2 and the Origin2OOO. Additionally, we evaluate the performance of five state-of-the-art partitioning algorithms that can be used within PLUM. It is shown that for certain classes of unsteady adaption, globally repartitioning the computational mesh produces higher quality results than diffusive repartitioning schemes. We also demonstrate that a coarse starting mesh produces high quality load balancing, at a fraction of the cost required a fine initial mesh. Results indicate that our parallel load balancing strategy will remain viable on large numbers of processors.

Oliker, Leonid↗

Automated Instrumentation, Monitoring and Visualization of PVM Programs Using AIMS

We present views and analysis of the execution of several PVM codes for Computational Fluid Dynamics on a network of Sparcstations, including (a) NAS Parallel benchmarks CG and MG (White, Alund and Sunderam 1993); (b) a multi-partitioning algorithm for NAS Parallel Benchmark SP (Wijngaart 1993); and (c) an overset grid flowsolver (Smith 1993). These views and analysis were obtained using our Automated Instrumentation and Monitoring System (AIMS) version 3.0, a toolkit for debugging the performance of PVM programs. We will describe the architecture, operation and application of AIMS. The AIMS toolkit contains (a) Xinstrument, which can automatically instrument various computational and communication constructs in message-passing parallel programs; (b) Monitor, a library of run-time trace-collection routines; (c) VK (Visual Kernel), an execution-animation tool with source-code clickback; and (d) Tally, a tool for statistical analysis of execution profiles. Currently, Xinstrument can handle C and Fortran77 programs using PVM 3.2.x; Monitor has been implemented and tested on Sun 4 systems running SunOS 4.1.2; and VK uses X11R5 and Motif 1.2. Data and views obtained using AIMS clearly illustrate several characteristic features of executing parallel programs on networked workstations: (a) the impact of long message latencies; (b) the impact of multiprogramming overheads and associated load imbalance; (c) cache and virtual-memory effects; and (4significant skews between workstation clocks. Interestingly, AIMS can compensate for constant skew (zero drift) by calibrating the skew between a parent and its spawned children. In addition, AIMS' skew-compensation algorithm can adjust timestamps in a way that eliminates physically impossible communications (e.g., messages going backwards in time). Our current efforts are directed toward creating new views to explain the observed performance of PVM programs. Some of the features planned for the near future include: (a) ConfigView, showing the physical topology of the virtual machine, inferred using specially formatted IP (Internet Protocol) packets; and (b) LoadView, synchronous animation of PVM-program execution and resource-utilization patterns.

Mehra, Pankaj↗