Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “hardware algorithms”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11

BoBa

BoBa is a C++ software library for working with large matrices, tensors, and tensor decompositions. The library provides tools for dense matrix and tensor operations, tensor decompositions, and tensor decomposition methods that support modern CPU and GPU architectures. It includes portable abstractions for linear algebra, tensor algebra, and multidimensional computation. BoBa is intended for scientific computing applications that involve large multidimensional data sets or high dimensional mathematical models. Its capabilities support tasks such as data compression, linear algebra, efficient numerical computation, and the development of scalable algorithms for heterogeneous hardware. Tutorials, tests, and example applications are included to help users learn and apply the library.

Yao, Jin [Lawrence Livermore National Laboratory (↗

An HPC benchmark survey and taxonomy for characterization

The field of High-Performance Computing (HPC) is defined by providing computing devices with highest performance for a variety of demanding scientific users. The tight co-design relationship between HPC providers and users propels the field forward, paired with technological improvements, achieving continuously higher performance and resource utilization. A key device for system architects, architecture researchers, and scientific users are benchmarks, allowing for well-defined assessment of hardware, software, and algorithms. Many benchmarks exist in the community, from individual niche benchmarks testing specific features, to large-scale benchmark suites for whole procurements. We survey the available HPC benchmarks, summarizing them in table form with key details and concise categorization, also through an interactive website. For categorization, we present a benchmark taxonomy for well-defined characterization of benchmarks.

Benchmarking↗

VQE method: a short survey and recent developments

Abstract The variational quantum eigensolver (VQE) is a method that uses a hybrid quantum-classical computational approach to find eigenvalues of a Hamiltonian. VQE has been proposed as an alternative to fully quantum algorithms such as quantum phase estimation (QPE) because fully quantum algorithms require quantum hardware that will not be accessible in the near future. VQE has been successfully applied to solve the electronic Schrödinger equation for a variety of small molecules. However, the scalability of this method is limited by two factors: the complexity of the quantum circuits and the complexity of the classical optimization problem. Both of these factors are affected by the choice of the variational ansatz used to represent the trial wave function. Hence, the construction of an efficient ansatz is an active area of research. Put another way, modern quantum computers are not capable of executing deep quantum circuits produced by using currently available ansatzes for problems that map onto more than several qubits. In this review, we present recent developments in the field of designing efficient ansatzes that fall into two categories—chemistry–inspired and hardware–efficient—that produce quantum circuits that are easier to run on modern hardware. We discuss the shortfalls of ansatzes originally formulated for VQE simulations, how they are addressed in more sophisticated methods, and the potential ways for further improvements.

Fedorov, Dmitry A. (ORCID:0000000316598580)↗

Advances in Mixed Precision Algorithms: 2021 Edition

Over the last year, the ECP xSDK-multiprecision effort has made tremendous progress in developing and deploying new mixed precision technology and customizing the algorithms for the hardware deployed in the ECP flagship supercomputers. The effort also has succeeded in creating a cross-laboratory community of scientists interested in mixed precision technology and now working together in deploying this technology for ECP applications. In this report, we highlight some of the most promising and impactful achievements of the last year. Among the highlights we present are: Mixed precision IR using a dense LU factorization and achieving a 1.8× speedup on Spock; results and strategies for mixed precision IR using a sparse LU factorization; a mixed precision eigenvalue solver; Mixed Precision GMRES-IR being deployed in Trilinos, and achieving a speedup of 1.4× over standard GMRES; compressed Basis (CB) GMRES being deployed in Ginkgo and achieving an average 1.4× speedup over standard GMRES; preparing hypre for mixed precision execution; mixed precision sparse approximate inverse preconditioners achieving an average speedup of 1.2×; and detailed description of the memory accessor separating the arithmetic precision from the memory precision, and enabling memory-bound low precision BLAS 1/2 operations to increase the accuracy by using high precision in the computations without degrading the performance. We emphasize that many of the highlights presented here have also been submitted to peer-reviewed journals or established conferences, and are under peer-review or have already been published.

97 MATHEMATICS AND COMPUTING↗

Advances in Mixed Precision Algorithms: 2021 Edition

Over the last year, the ECP xSDK-multiprecision effort has made tremendous progress in developing and deploying new mixed precision technology and customizing the algorithms for the hardware deployed in the ECP flagship supercomputers. The effort also has succeeded in creating a cross-laboratory community of scientists interested in mixed precision technology and now working together in deploying this technology for ECP applications. In this report, we highlight some of the most promising and impactful achievements of the last year. Among the highlights we present are • Mixed precision IR using a dense LU factorization and achieving a 1.8× speedup on Spock; • Results and strategies for mixed precision IR using a sparse LU factorization; • A mixed precision eigenvalue solver; • Mixed Precision GMRES-IR being deployed in Trilinos, and achieving a speedup of 1.4× over standard GMRES; • Compressed Basis (CB) GMRES being deployed in Ginkgo and achieving an average 1.4× speedup over standard GMRES; • Preparing hypre for mixed precision execution; • Mixed precision sparse approximate inverse preconditioners achieving an average speedup of 1.2×; • Detailed description of the memory accessor separating the arithmetic precision from the memory precision, and enabling memory-bound low precision BLAS 1/2 operations to increase the accuracy by using high precision in the computations without degrading the performance. We emphasize that many of the highlights presented here have also been submitted to peer-reviewed journals or established conferences, and are under peer-review or have already been published.

97 MATHEMATICS AND COMPUTING↗

Extended Low Load Boiler Operation to Improve Performance and Economics of an Existing Coal Fired Power Plant (Final Report)

The overall goal is to improve the performance and economics of existing coal fired power plants by extending low load boiler operation to lower loads than is currently achievable. The objective of this program is to develop and validate sensor hardware and analytical algorithms to lower plant operating expenses (OPEX) for the currently operating pulverized coal utility boiler fleet. Coal fired utility boilers are increasingly under grid dispatch pressure. In some cases, the coal fired cost of generation is noncompetitive with respect to natural gas generation and subsidized renewable sources. To remain profitable and remain fully compliant with existing environmental regulations, the installed coal fired fleet must find technologies which allow it to move into a more flexible cyclic load dispatch model. Today the installed coal fired utility fleet must be cost of generation competitive, fully emissions compliant, and responsive to the variability inherent in renewable energy generation sources. In the Phase I of the project, GE Steam Power, Inc. (GE) performed modeling of different operating scenarios for low load operation using an existing full plant dynamic model developed for a 660MW steam power plant. Sensors and analytic algorithms to enable a stable and steady coal supply for low load pulverizer operation were identified and tested at the Pulverizer Development Facility (PDF) at GE’s Clean Energy Center in Bloomfield, Connecticut. Sensors and analytic algorithms to enable stable combustion for low load operation were identified and tested at the 15 MWth Industrial Scale Burner facility (ISBF) at GE’s Clean Energy Center. A concept was developed to test the sensors and control algorithms, down selected after testing, at a full-scale coal fired power plant. A budget estimate was then developed, and the concept was implemented at an existing utility power plant. The specific objectives of the experimental work were to: • Identify and select sensors and analytic algorithms for monitoring coal pulverizer operation at lower loads to provide stable operation and appropriate coal fineness at lower coal throughput; Identify and select sensors and analytic algorithms for a Boiler Flame Stability Monitor to better balance air and fuel at each burner. This enables a reduction in a coal boiler’s safe low load power level while maintaining stable flame characteristics; Develop a concept in Phase I for low load operation of a full-scale power plant and develop a budget estimate for testing and execute the test plan at an existing plant in Phase II; Validate the capability of the extended low load boiler system to extend the minimum load operating point in a safe and reliable manner on an existing full-scale utility boiler. At the completion of this experimental study, GE has developed a set of sensors and analytic algorithms, down selected after testing, that have the potential to enable safe low load operation of a utility boiler. GE has also identified a host site for testing these identified sensors and analytic algorithms. GE has generated a full set of deliverables that provide sufficient information to proceed with the next step of testing at a host site. This includes a potential host site and budget estimate for concept testing at host site. In the Phase II of the project, a series of field tests were completed to validate the extended low load boiler operation, which consisted of detailed engineering, installation, commissioning, and testing the additional sensors and analytics for the coal-fired combustion system on an existing full-scale utility boiler. The optimization work has been supported by the host plant and endorsed by their engineering and operation staff.

01 COAL, LIGNITE, AND PEAT↗

Integration of a real-time orientation measurement system for a real-time evaluator (RTE) to measure the position and orientation of crane-lifted components

Prefabrication of building components holds the potential to revolutionize the construction industry. Prefabrication consists of manufacturing building components, modules, and other elements in a factory to be shipped and installed on a construction site. Prefabricated components have been produced for various applications including precast concrete panels for new construction and exterior wall retrofits. The manufacturing process has seen much innovation in recent years; however, the installation process has seen minimal advancements. A real-time evaluator (RTE) was developed to reduce the installation cost of prefabricated components by reducing installation time, decreasing rework, and improving accuracy. The RTE uses off-the-shelf hardware and novel algorithms to assist erectors with component installation by measuring the real-time positions of connections and prefabricated components, providing installation guidance through a graphical user interface, and monitoring the accumulated installation errors. An overview of the RTE and the proposed workflow is presented. Previous on-site demonstrations provided valuable feedback from users on the potential areas for improvement of the system. One common request was real-time measurement of component orientation during lifting, a process that previously required that the component remain stationary while the laser tracker cycled through target prisms. This paper will present the incorporation and testing of a real-time orientation measurement system as it was implemented into the RTE, allowing for measurement of component orientation during movement.

Selvakumar, Balaji [ORNL]↗

Rasterization with Data-Parallel Primitives

Parallel rasterization can suffer from race conditions during fragment generation, which is traditionally addressed by using specialized hardware accessible via vendor graphics APIs. Unfortunately, graphics APIs are increasingly problematic on high-performance computers, either because they are not provided or because of concerns about dependencies with in situ visualization. In response, we present a hardware-agnostic rasterization algorithm that handles race conditions using only data-parallel primitives (DPPs), enabling efficient rendering on HPC systems without graphics API dependencies and aligning with recent efforts to deliver visualization software with DPPs. Our evaluation consists of three phases: (1) evaluating portability across different CPU and GPU architectures, (2) evaluating competitiveness with a community standard, and (3) evaluating performance across varying workloads and available parallelism. The supporting experiments run on both AMD and NVIDIA GPUs, considering data sets as large as 460 million triangles and 160 million pixels. While performance generally falls short of graphics API baselines, it achieves interactive frame rates on most workloads. As a result, we conclude our approach is a viable solution for rasterization on high-performance computers since our approach is portably performant across different architectures without the need for specialized vendor support.

Buckley, Makani [University of Oregon] (ORCID:0009↗

IBIS - A geographic information system based on digital image processing and image raster datatype

IBIS (Image Based Information System) is a geographic information system which makes use of digital image processing techniques to interface existing geocoded data sets and information management systems with thematic maps and remotely sensed imagery. The basic premise is that geocoded data sets can be referenced to a raster scan that is equivalent to a grid cell data set. The first applications (St. Tammany Parish, Louisiana, and Los Angeles County) have been restricted to the design of a land resource inventory and analysis system. It is thought that the algorithms and the hardware interfaces developed will be readily applicable to other Landsat imagery.

Bryant, N. A.↗

Radiometric Accuracy Assessment of LANDSAT-4 Multispectral Scanner (MSS) Data

The LANDSAT-4 mission has unique characteristics relative to previous LANDSAT missions. The spacecraft is new; the orbit is lower with a more frequent repeat cycle; and the ground processing facility consists of new hardware with different algorithms being applied. How some of these changes affect the character of the radiometric data quality is explored. Banding effects; radiometric differences between LANDSAT 3 and 4; and the woodgrain pattern observed visually in the images are considered.

Alford, W. L.↗

Autonomous Integrated Receive System (AIRS) requirements definition. Volume 3: Performance and simulation

The autonomous and integrated aspects of the operation of the AIRS (Autonomous Integrated Receive System) are discussed from a system operation point of view. The advantages of AIRS compared to the existing SSA receive chain equipment are highlighted. The three modes of AIRS operation are addressed in detail. The configurations of the AIRS are defined as a function of the operating modes and the user signal characteristics. Each AIRS configuration selection is made up of three components: the hardware, the software algorithms and the parameters used by these algorithms. A comparison between AIRS and the wide dynamics demodulation (WDD) is provided. The organization of the AIRS analytical/simulation software is described. The modeling and analysis is for simulating the performance of the PN subsystem is documented. The frequence acquisition technique using a frequency-locked loop is also documented. Doppler compensation implementation is described. The technological aspects of employing CCD's for PN acquisition are addressed.

Chie, C. M.↗

Shuttle-attached Antenna Flight Experiment Definition Study (FEDS)

The control algorithms, techniques, and hardware which would be required to support whether flight experiments of large space structures control are assessed for a 55-meter diameter wrap-rib reflector with a three degree-of-freedom gimbal. Strowman requirements were established for geometry, mass property, and elastic mode identification as well as for control and slewing. A five-body simulation of the Shuttle and test article was built with the ALLFLEX computer program. A maximum likelihood estimator, the flight experiment timeline, and the LSS control development test plan are discussed.

Hannan, G. J.↗

A two-dimensional intensified photodiode array for imaging spectroscopy

The Johns Hopkins University is currently developing an instrument to fly aboard NASA's Space Shuttle as a Spartan payload in the late 1980s. This Spartan free flyer will obtain spatially resolved spectra of faint extended emission line objects in the wavelength range 750-1150 A at about 2-A resolution. The use of two-dimensional photon counting detectors will give simultaneous coverage of the 400 A spectral range and the 9 arc-minute spatial resolution along the spectrometer slit. The progress towards the flight detector is reported here with preliminary results from a laboratory breadboard detector, and a comparison with the one-dimensional detector developed for the Hopkins Ultraviolet Telescope. A hardware digital centroiding algorithm has been successfully implemented. The system is ultimately capable of 15-micron resolution in two dimensions at the image plane and can handle continuous counting rates of up to 8000 counts/s.

Tennyson, P. D.↗

Computational mechanics - Advances and trends; Proceedings of the Session - Future directions of Computational Mechanics of the ASME Winter Annual Meeting, Anaheim, CA, Dec. 7-12, 1986

The papers contained in this volume provide an overview of the advances made in a number of aspects of computational mechanics, identify some of the anticipated industry needs in this area, discuss the opportunities provided by new hardware and parallel algorithms, and outline some of the current government programs in computational mechanics. Papers are included on advances and trends in parallel algorithms, supercomputers for engineering analysis, material modeling in nonlinear finite-element analysis, the Navier-Stokes computer, and future finite-element software systems.

Noor, Ahmed K.↗

A post-processing system for automated rectification and registration of spaceborne SAR imagery

An automated post-processing system has been developed that interfaces with the raw image output of the operational digital SAR correlator. This system is designed for optimal efficiency by using advanced signal processing hardware and an algorithm that requires no operator interaction, such as the determination of ground control points. The standard output is a geocoded image product (i.e. resampled to a specified map projection). The system is capable of producing multiframe mosaics for large-scale mapping by combining images in both the along-track direction and adjacent cross-track swaths from ascending and descending passes over the same target area. The output products have absolute location uncertainty of less than 50 m and relative distortion (scale factor and skew) of less than 0.1 per cent relative to local variations from the assumed geoid.

Curlander, John C.↗

Transmission delays in hardware clock synchronization

Various methods, both with software and hardware, have been proposed to synchronize a set of physical clocks in a system. Software methods are very flexible and economical but suffer an excessive time overhead, whereas hardware methods require no time overhead but are unable to handle transmission delays in clock signals. The effects of nonzero transmission delays in synchronization have been studied extensively in the communication area in the absence of malicious or Byzantine faults. The authors show that it is easy to incorporate the ideas from the communication area into the existing hardware clock synchronization algorithms to take into account the presence of both malicious faults and nonzero transmission delays.

Shin, Kang G.↗

Advanced computing

Advanced concepts in hardware, software and algorithms are being pursued for application in next generation space computers and for ground based analysis of space data. The research program focuses on massively parallel computation and neural networks, as well as optical processing and optical networking which are discussed under photonics. Also included are theoretical programs in neural and nonlinear science, and device development for magnetic and ferroelectric memories.

Source record↗

Phase noise in pulsed Doppler lidar and limitations on achievable single-shot velocity accuracy

The smaller sampling volumes afforded by Doppler lidars compared to radars allows for spatial resolutions at and below some sheer and turbulence wind structure scale sizes. This has brought new emphasis on achieving the optimum product of wind velocity and range resolutions. Several recent studies have considered the effects of amplitude noise, reduction algorithms, and possible hardware related signal artifacts on obtainable velocity accuracy. We discuss here the limitation on this accuracy resulting from the incoherent nature and finite temporal extent of backscatter from aerosols. For a lidar return from a hard (or slab) target, the phase of the intermediate frequency (IF) signal is random and the total return energy fluctuates from shot to shot due to speckle; however, the offset from the transmitted frequency is determinable with an accuracy subject only to instrumental effects and the signal to noise ratio (SNR), the noise being determined by the LO power in the shot noise limited regime. This is not the case for a return from a media extending over a range on the order of or greater than the spatial extent of the transmitted pulse, such as from atmospheric aerosols. In this case, the phase of the IF signal will exhibit a temporal random walk like behavior. It will be uncorrelated over times greater than the pulse duration as the transmitted pulse samples non-overlapping volumes of scattering centers. Frequency analysis of the IF signal in a window similar to the transmitted pulse envelope will therefore show shot-to-shot frequency deviations on the order of the inverse pulse duration reflecting the random phase rate variations. Like speckle, these deviations arise from the incoherent nature of the scattering process and diminish if the IF signal is averaged over times greater than a single range resolution cell (here the pulse duration). Apart from limiting the high SNR performance of a Doppler lidar, this shot-to-shot variance in velocity estimates has a practical impact on lidar design parameters. In high SNR operation, for example, a lidar's efficiency in obtaining mean wind measurements is determined by its repetition rate and not pulse energy or average power. In addition, this variance puts a practical limit on the shot-to-shot hard target performance required of a lidar.

Mcnicholl, P.↗