Engineering PapersSearch

Engineering topics

Watson, Andrew B.

Publications and source records attributed to Watson, Andrew B..

At least 37 records · Page 2

Entropy Masking

This paper details two projects that use the World Wide Web (WWW) for dissemination of curricula that focus on remote sensing. 1) Presenting grade-school students with the concepts used in remote sensing involves educating the teacher and then providing the teacher with lesson plans. In a NASA-sponsored project designed to introduce students in grades 4 through 12 to some of the ideas and terminology used in remote sensing, teachers from local grade schools and middle schools were recruited to write lessons about remote sensing concepts they could use in their classrooms. Twenty-two lessons were produced and placed in seven modules that include: the electromagnetic spectrum, two- and three-dimensional perception, maps and topography, scale, remote sensing, biotic and abiotic concepts, and landscape chi rise. Each lesson includes a section that evaluates what students have learned by doing the exercise. The lessons, instead of being published in a workbook and distributed to a limited number of teachers, have been placed on a WWW server, enabling much broader access to the package. This arrangement also allows for the lessons to be modified after feedback from teachers accessing the package. 2) Two-year colleges serve to teach trade skills, prepare students for enrollment in senior institutions of learning, and more and more, retrain students who have college degrees in new technologies and skills. A NASA-sponsored curriculum development project is producing a curriculum using remote sensing analysis an Earth science applications. The project has three major goals. First, it will implement the use of remote sensing data in a broad range of community college courses. Second, it will create curriculum modules and classes that are transportable to other community colleges. Third, the project will be an ongoing source of data and curricular materials to other community colleges. The curriculum will have these course pathways to a certificate; a) a Science emphasis, b) an Arts and Letters emphasis, and c) a Computer Science emphasis Each pathway includes course work in remote sensing, geographical information systems (GIS), computer science, Earth science, software and technology utilization, and communication. Distribution of products from this project to other two-year colleges will be accomplished using the WWW.

Watson, Andrew B.

The Search for Optimal Visual Stimuli

In 1983, Watson, Barlow and Robson published a brief report in which they explored the relative visibility of targets that varied in size, shape, spatial frequency, speed, and duration (referred to subsequently here as WBR). A novel aspect of that paper was that visibility was quantified in terms of threshold contrast energy, rather than contrast. As they noted, this provides a more direct measure of the efficiency with which various patterns are detected, and may be more edifying as to the underlying detection machinery. For example, under certain simple assumptions, the waveform of the most efficiently detected signal is an estimate of the receptive field of the visual system's most efficient detector. Thus one goal of their experiment Basuto search for the stimulus that the 'eye sees best'. Parenthetically, the search for optimal stimuli may be seen as the most general and sophisticated variant of the traditional 'subthreshold summation' experiment, in which one measures the effect upon visibility of small probes combined with a base stimulus.

Watson, Andrew B.

Perceptually Lossless Wavelet Compression

The Discrete Wavelet Transform (DWT) decomposes an image into bands that vary in spatial frequency and orientation. It is widely used for image compression. Measures of the visibility of DWT quantization errors are required to achieve optimal compression. Uniform quantization of a single band of coefficients results in an artifact that is the sum of a lattice of random amplitude basis functions of the corresponding DWT synthesis filter, which we call DWT uniform quantization noise. We measured visual detection thresholds for samples of DWT uniform quantization noise in Y, Cb, and Cr color channels. The spatial frequency of a wavelet is r 2(exp -1), where r is display visual resolution in pixels/degree, and L is the wavelet level. Amplitude thresholds increase rapidly with spatial frequency. Thresholds also increase from Y to Cr to Cb, and with orientation from low-pass to horizontal/vertical to diagonal. We propose a mathematical model for DWT noise detection thresholds that is a function of level, orientation, and display visual resolution. This allows calculation of a 'perceptually lossless' quantization matrix for which all errors are in theory below the visual threshold. The model may also be used as the basis for adaptive quantization schemes.

Watson, Andrew B.

A Rorschach Test for Visual Classification Strategies

Contemporary models of pattern, detection and discrimination often employ template matching, but there have been few direct tests of this proposition. Adopting a method developed by Ahumada, we have analyzed how human observers discriminate between two letters of the alphabet ('c' and 'x'). The stimulus consisted of a one degree tall letter plus a four degree field of static white noise, both displayed for 16 frames at a 67 Hz frame rate. Our font and display dimensions approximated those of Solomon and Pelli. The observer identified the letter presented. A QUEST staircase varied letter contrast to maintain a 75% correct rate. For each trial, we preserved the information required to reconstruct the noise field. Possible trial categories based on (signal, response) pairs are: (c,c), (c,x), (x,c), (x,x). Noise fields were averaged separately for each category, and a final classification image was obtained by averaging the four mean images after inverting the sign of categories in which x was the response. If the observer employs a template, it should be revealed in the classification image. The lowpass-filtered classification image derived from 2048 responses of one observer is shown here, along with the corresponding ideal template. An approximation to the ideal template can be seen appropriately located within the classification image. We have also simulated and will discuss the classification images expected from various discrimination models in this experimental context. The construction of classification images appears to be a powerful tool for studying classification strategies used by human observers. Like a Rorschach test, it surreptitiously discovers the inner desires of the visual system.

Watson, Andrew B.

DCTune Perceptual Optimization of Compressed Dental X-Rays

In current dental practice, x-rays of completed dental work are often sent to the insurer for verification. It is faster and cheaper to transmit instead digital scans of the x-rays. Further economies result if the images are sent in compressed form. DCTune is a technology for optimizing DCT (digital communication technology) quantization matrices to yield maximum perceptual quality for a given bit-rate, or minimum bit-rate for a given perceptual quality. Perceptual optimization of DCT color quantization matrices. In addition, the technology provides a means of setting the perceptual quality of compressed imagery in a systematic way. The purpose of this research was, with respect to dental x-rays, 1) to verify the advantage of DCTune over standard JPEG (Joint Photographic Experts Group), 2) to verify the quality control feature of DCTune, and 3) to discover regularities in the optimized matrices of a set of images. We optimized matrices for a total of 20 images at two resolutions (150 and 300 dpi) and four bit-rates (0.25, 0.5, 0.75, 1.0 bits/pixel), and examined structural regularities in the resulting matrices. We also conducted psychophysical studies (1) to discover the DCTune quality level at which the images became 'visually lossless,' and (2) to rate the relative quality of DCTune and standard JPEG images at various bitrates. Results include: (1) At both resolutions, DCTune quality is a linear function of bit-rate. (2) DCTune quantization matrices for all images at all bitrates and resolutions are modeled well by an inverse Gaussian, with parameters of amplitude and width. (3) As bit-rate is varied, optimal values of both amplitude and width covary in an approximately linear fashion. (4) Both amplitude and width vary in systematic and orderly fashion with either bit-rate or DCTune quality; simple mathematical functions serve to describe these relationships. (5) In going from 150 to 300 dpi, amplitude parameters are substantially lower and widths larger at corresponding bit-rates or qualities. (6) Visually lossless compression occurs at a DCTune quality value of about 1. (7) At 0.25 bits/pixel, comparative ratings give DCTune a substantial advantage over standard JPEG. As visually lossless bit-rates are approached, this advantage of necessity diminishes. We have concluded that DCTune optimized quantization matrices provide better visual quality than standard JPEG. Meaningful quality levels may be specified by means of the DCTune metric. Optimized matrices are very similar across the class of dental x-rays, suggesting the possibility of a 'class-optimal' matrix. DCTune technology appears to provide some value in the context of compressed dental x-rays.

Watson, Andrew B.

Perceptual Image Compression in Telemedicine

The next era of space exploration, especially the "Mission to Planet Earth" will generate immense quantities of image data. For example, the Earth Observing System (EOS) is expected to generate in excess of one terabyte/day. NASA confronts a major technical challenge in managing this great flow of imagery: in collection, pre-processing, transmission to earth, archiving, and distribution to scientists at remote locations. Expected requirements in most of these areas clearly exceed current technology. Part of the solution to this problem lies in efficient image compression techniques. For much of this imagery, the ultimate consumer is the human eye. In this case image compression should be designed to match the visual capacities of the human observer. We have developed three techniques for optimizing image compression for the human viewer. The first consists of a formula, developed jointly with IBM and based on psychophysical measurements, that computes a DCT quantization matrix for any specified combination of viewing distance, display resolution, and display brightness. This DCT quantization matrix is used in most recent standards for digital image compression (JPEG, MPEG, CCITT H.261). The second technique optimizes the DCT quantization matrix for each individual image, based on the contents of the image. This is accomplished by means of a model of visual sensitivity to compression artifacts. The third technique extends the first two techniques to the realm of wavelet compression. Together these two techniques will allow systematic perceptual optimization of image compression in NASA imaging systems. Many of the image management challenges faced by NASA are mirrored in the field of telemedicine. Here too there are severe demands for transmission and archiving of large image databases, and the imagery is ultimately used primarily by human observers, such as radiologists. In this presentation I will describe some of our preliminary explorations of the applications of our technology to the special problems of telemedicine.

Watson, Andrew B.

Perceptually-Based Adaptive JPEG Coding

An extension to the JPEG standard (ISO/IEC DIS 10918-3) allows spatial adaptive coding of still images. As with baseline JPEG coding, one quantization matrix applies to an entire image channel, but in addition the user may specify a multiplier for each 8 x 8 block, which multiplies the quantization matrix, yielding the new matrix for the block. MPEG 1 and 2 use much the same scheme, except there the multiplier changes only on macroblock boundaries. We propose a method for perceptual optimization of the set of multipliers. We compute the perceptual error for each block based upon DCT quantization error adjusted according to contrast sensitivity, light adaptation, and contrast masking, and pick the set of multipliers which yield maximally flat perceptual error over the blocks of the image. We investigate the bitrate savings due to this adaptive coding scheme and the relative importance of the different sorts of masking on adaptive coding.

Watson, Andrew B.

Image data compression having minimum perceptual error

A method for performing image compression that eliminates redundant and invisible image components is described. The image compression uses a Discrete Cosine Transform (DCT) and each DCT coefficient yielded by the transform is quantized by an entry in a quantization matrix which determines the perceived image quality and the bit rate of the image being compressed. The present invention adapts or customizes the quantization matrix to the image being compressed. The quantization matrix comprises visual masking by luminance and contrast techniques and by an error pooling technique all resulting in a minimum perceptual error for any given bit rate, or minimum bit rate for a given perceptual error.

Watson, Andrew B.

Visibility of Wavelet Quantization Noise

The Discrete Wavelet Transform (DWT) decomposes an image into bands that vary in spatial frequency and orientation. It is widely used for image compression. Measures of the visibility of DWT quantization errors are required to achieve optimal compression. Uniform quantization of a single band of coefficients results in an artifact that is the sum of a lattice of random amplitude basis functions of the corresponding DWT synthesis filter, which we call DWT uniform quantization noise. We measured visual detection thresholds for samples of DWT uniform quantization noise in Y, Cb, and Cr color channels. The spatial frequency of a wavelet is r 2(exp)-L , where r is display visual resolution in pixels/degree, and L is the wavelet level. Amplitude thresholds increase rapidly with spatial frequency. Thresholds also increase from Y to Cr to Cb, and with orientation from low-pass to horizontal/vertical to diagonal. We describe a mathematical model to predict DWT noise detection thresholds as a function of level, orientation, and display visual resolution. This allows calculation of a "perceptually lossless" quantization matrix for which all errors are in theory below the visual threshold. The model may also be used as the basis for adaptive quantization schemes.

Watson, Andrew B.

A comparison of Image Quality Models and Metrics Predicting Object Detection

Many models and metrics for image quality predict image discriminability, the visibility of the difference between a pair of images. Some image quality applications, such as the quality of imaging radar displays, are concerned with object detection and recognition. Object detection involves looking for one of a large set of object sub-images in a large set of background images and has been approached from this general point of view. We find that discrimination models and metrics can predict the relative detectability of objects in different images, suggesting that these simpler models may be useful in some object detection and recognition applications. Here we compare three alternative measures of image discrimination, a multiple frequency channel model, a single filter model, and RMS error.

Rohaly, Ann Marie

A visual detection model for DCT coefficient quantization

The discrete cosine transform (DCT) is widely used in image compression and is part of the JPEG and MPEG compression standards. The degree of compression and the amount of distortion in the decompressed image are controlled by the quantization of the transform coefficients. The standards do not specify how the DCT coefficients should be quantized. One approach is to set the quantization level for each coefficient so that the quantization error is near the threshold of visibility. Results from previous work are combined to form the current best detection model for DCT coefficient quantization noise. This model predicts sensitivity as a function of display parameters, enabling quantization matrices to be designed for display situations varying in luminance, veiling light, and spatial frequency related conditions (pixel size, viewing distance, and aspect ratio). It also allows arbitrary color space directions for the representation of color. A model-based method of optimizing the quantization matrix for an individual image was developed. The model described above provides visual thresholds for each DCT frequency. These thresholds are adjusted within each block for visual light adaptation and contrast masking. For given quantization matrix, the DCT quantization errors are scaled by the adjusted thresholds to yield perceptual errors. These errors are pooled nonlinearly over the image to yield total perceptual error. With this model one may estimate the quantization matrix for a particular image that yields minimum bit rate for a given total perceptual error, or minimum perceptual error for a given bit rate. Custom matrices for a number of images show clear improvement over image-independent matrices. Custom matrices are compatible with the JPEG standard, which requires transmission of the quantization matrix.

Ahumada, Albert J., Jr.

Motion-Contrast Sensitivity: Visibility of Motion Gradients of Various Spatial Frequencies

The purpose of our experiments was to estimate basic sensitivity to motion gradients and to evaluate the evidence for second-order integration and differentiation of motion signals. We measured sensitivity to spatially sinusoidal contrast modulation between two oppositely moving bandpass-filtered noise images. The motion-contrast sensitivity function, defined as the inverse of threshold modulation amplitude as a function of modulation spatial frequency, was bandpass in shape with declines at both highest and lowest frequencies. The functions for three noise spatial frequencies had approximately the same shape when modulation frequency was expressed as a fraction of noise frequency. We compared the data with a model in which linear motion filters, whose outputs are squared or rectified, are followed by a second stage of excitatory or inhibitory pooling. The data are consistent with a model in which (1) all excitatory pooling occurs at the linear stage and (2) the second stage contains a large inhibitory pooling area, with a radius approximately eight times that of the linear receptive field.

Watson, Andrew B.

Image-adapted visually weighted quantization matrices for digital image compression

A method for performing image compression that eliminates redundant and invisible image components is presented. The image compression uses a Discrete Cosine Transform (DCT) and each DCT coefficient yielded by the transform is quantized by an entry in a quantization matrix which determines the perceived image quality and the bit rate of the image being compressed. The present invention adapts or customizes the quantization matrix to the image being compressed. The quantization matrix comprises visual masking by luminance and contrast techniques and by an error pooling technique all resulting in a minimum perceptual error for any given bit rate, or minimum bit rate for a given perceptual error.

Watson, Andrew B.

Visibility of DCT Basis Functions: Effects of Contrast Masking

Current strategies for compressing images neglect to accommodate certain aspects of the human visual system. One such neglected aspect is contrast masking. Here we provide measurements of contrast masking within the framework of the discrete cosine transform (DCT). We describe how to optimize DCT-based compression with respect to contrast masking.

Soloman, Joshua A.

Visibility of DCT Quantization Error: Effects of Display Resolution

As part of a program of research to understand the visibility of DCT quantization errors and thereby design optimal quantizers, we measured visibility of DCT quantization error as a function of display resolution in pixels/degree. Visibilities are consistent with a model incorporating effects of block size and spatial pooling.

Watson, Andrew B.

A Modular, Portable Model of Image Fidelity

There is a persistent need for a trustworthy model of perceptual image fidelity, especially in applications such as image compression and display design. A fidelity model provides a measure of the visual discriminability of two images. Ahumada has previously shown that the existing fidelity models may be categorized according to their inclusion of various canonical properties, such as a contrast sensitivity function, spatial frequency channels, etc. This suggests that research would be aided by the availability of a modular model, in which these components could be easily inserted or removed. A further impediment to research in this area has been that most models are written in low-level languages and are consequently large, non-portable, and difficult to understand, modify, and maintain. We therefore believe research would also be aided by models written in high-level languages. To serve both of these purposes, and to honor our conference host for his lifetime dedication to the problem of image quality. Global brightness and its effect on perceptual image quality. We offer a modular model written in the high-level language Mathematica. We will demonstrate this model and show how it may be modified.

Watson, Andrew B.

Perceptual Optimization of DCT Color Quantization Matrices

Many image compression schemes employ a block Discrete Cosine Transform (DCT) and uniform quantization. Acceptable rate/distortion performance depends upon proper design of the quantization matrix. In previous work, we showed how to use a model of the visibility of DCT basis functions to design quantization matrices for arbitrary display resolutions and color spaces. Subsequently, we showed how to optimize greyscale quantization matrices for individual images, for optimal rate/perceptual distortion performance. Here we describe extensions of this optimization algorithm to color images.

Watson, Andrew B.

Contrast Gain Control Model Fits Masking Data

We studied the fit of a contrast gain control model to data of Foley (JOSA 1994), consisting of thresholds for a Gabor patch masked by gratings of various orientations, or by compounds of two orientations. Our general model includes models of Foley and Teo & Heeger (IEEE 1994). Our specific model used a bank of Gabor filters with octave bandwidths at 8 orientations. Excitatory and inhibitory nonlinearities were power functions with exponents of 2.4 and 2. Inhibitory pooling was broad in orientation, but narrow in spatial frequency and space. Minkowski pooling used an exponent of 4. All of the data for observer KMF were well fit by the model. We have developed a contrast gain control model that fits masking data. Unlike Foley's, our model accepts images as inputs. Unlike Teo & Heeger's, our model did not require multiple channels for different dynamic ranges.

Watson, Andrew B.