Engineering PapersSearch

SEARCH · Engineering Papers

Results for “PHONEME”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

Techniques for decoding speech phonemes and sounds: A concept

Techniques studied involve conversion of speech sounds into machine-compatible pulse trains. (1) Voltage-level quantizer produces number of output pulses proportional to amplitude characteristics of vowel-type phoneme waveforms. (2) Pulses produced by quantizer of first speech formants are compared with pulses produced by second formants.

Lokerson, D. C.

(abstract) Synthesis of Speaker Facial Movements to Match Selected Speech Sequences

We are developing a system for synthesizing image sequences the simulate the facial motion of a speaker. To perform this synthesis, we are pursuing two major areas of effort. We are developing the necessary computer graphics technology to synthesize a realistic image sequence of a person speaking selected speech sequences. Next, we are developing a model that expresses the relation between spoken phonemes and face/mouth shape. A subject is video taped speaking an arbitrary text that contains expression of the full list of desired database phonemes. The subject is video taped from the front speaking normally, recording both audio and video detail simultaneously. Using the audio track, we identify the specific video frames on the tape relating to each spoken phoneme. From this range we digitize the video frame which represents the extreme of mouth motion/shape. Thus, we construct a database of images of face/mouth shape related to spoken phonemes. A selected audio speech sequence is recorded which is the basis for synthesizing a matching video sequence; the speaker need not be the same as used for constructing the database. The audio sequence is analyzed to determine the spoken phoneme sequence and the relative timing of the enunciation of those phonemes. Synthesizing an image sequence corresponding to the spoken phoneme sequence is accomplished using a graphics technique known as morphing. Image sequence keyframes necessary for this processing are based on the spoken phoneme sequence and timing. We have been successful in synthesizing the facial motion of a native English speaker for a small set of arbitrary speech segments. Our future work will focus on advancement of the face shape/phoneme model and independent control of facial features.

graphics audio digitizing morphing enunciation pho

Voice intelligibility in satellite mobile communications

An amplitude control technique is reported that equalizes low level phonemes in a satellite narrow band FM voice communication system over channels having low carrier to noise ratios. This method presents at the transmitter equal amplitude phonemes so that the low level phonemes, when they are transmitted over the noisey channel, are above the noise and contribute to output intelligibility. The amplitude control technique provides also for squelching of noise when speech is not being transmitted.

Wishna, S.

Comparison of voice types for helicopter voice warning systems

Three related studies were conducted to compare different types of human voice warnings. In the first study, a comparison of three LPC-encoded voices, human female, human male, and phoneme-synthesized, by the criteria of pilot flight task performance showed no differences due to the voice type. In the second study, pilots' preferences were investigated, by comparing preference for direct synthesized speech to the LPC-encoded human female speech and to LPC-encoded synthesized speech. Most pilots were found to prefer direct synthesized speech over both LPC-encoded human female speech and the LPC-encoded synthesized speech. In the third study, phonetically balanced (PB) words heard in simulated helicopter noise were used to compare the intelligibility of direct synthesized and LPC-encoded phoneme-synthesized speech types. PB word intelligibility was found to be better for direct synthesized speech than for the LPC-encodes synthesized speech.

Simpson, C. A.

Speech Acquisition and Automatic Speech Recognition for Integrated Spacesuit Audio Systems

A voice-command human-machine interface system has been developed for spacesuit extravehicular activity (EVA) missions. A multichannel acoustic signal processing method has been created for distant speech acquisition in noisy and reverberant environments. This technology reduces noise by exploiting differences in the statistical nature of signal (i.e., speech) and noise that exists in the spatial and temporal domains. As a result, the automatic speech recognition (ASR) accuracy can be improved to the level at which crewmembers would find the speech interface useful. The developed speech human/machine interface will enable both crewmember usability and operational efficiency. It can enjoy a fast rate of data/text entry, small overall size, and can be lightweight. In addition, this design will free the hands and eyes of a suited crewmember. The system components and steps include beam forming/multi-channel noise reduction, single-channel noise reduction, speech feature extraction, feature transformation and normalization, feature compression, model adaption, ASR HMM (Hidden Markov Model) training, and ASR decoding. A state-of-the-art phoneme recognizer can obtain an accuracy rate of 65 percent when the training and testing data are free of noise. When it is used in spacesuits, the rate drops to about 33 percent. With the developed microphone array speech-processing technologies, the performance is improved and the phoneme recognition accuracy rate rises to 44 percent. The recognizer can be further improved by combining the microphone array and HMM model adaptation techniques and using speech samples collected from inside spacesuits. In addition, arithmetic complexity models for the major HMMbased ASR components were developed. They can help real-time ASR system designers select proper tasks when in the face of constraints in computational resources.

Huang, Yiteng

Intelligibility improvement of analog communication systems using an amplitude control technique.

An amplitude control technique has been employed for use with analog voice communication systems, which improves low-level phoneme reception and eliminates the received noise between words and syllables. Tests were conducted on a narrow-band frequency-modulation simplex voice communication channel employing the amplitude control technique. Presented for both the modified rhyme word tests and the phonetically balanced word tests are a series of graphical plots of the tests' score distribution, mean, and standard deviation as a function of received carrier-to-noise power density ratio. At low received carrier-to-noise power density ratios, a significant improvement in the intelligibility was obtained. A voice intelligibility improvement of more than 2 dB was obtained for the modified rhyme test words, and a voice intelligibility improvement in excess of 4 dB was obtained for the phonetically balanced word tests.

Wishna, S.

A new VOX technique for reducing noise in voice communication systems

A VOX technique for reducing noise in voice communication systems is described which is based on the separation of voice signals into contiguous frequency-band components with the aid of an adaptive VOX in each band. It is shown that this processing scheme can effectively reduce both wideband and narrowband quasi-periodic noise since the threshold levels readjust themselves to suppress noise that exceeds speech components in each band. Results are reported for tests of the adaptive VOX, and it is noted that improvements can still be made in such areas as the elimination of noise pulses, phoneme reproduction at high-noise levels, and the elimination of distortion introduced by phase delay.

Morris, C. F.

Synthesized speech rate and pitch effects on intelligibility of warning messages for pilots

In civilian and military operations, a future threat-warning system with a voice display could warn pilots of other traffic, obstacles in the flight path, and/or terrain during low-altitude helicopter flights. The present study was conducted to learn whether speech rate and voice pitch of phoneme-synthesized speech affects pilot accuracy and response time to typical threat-warning messages. Helicopter pilots engaged in an attention-demanding flying task and listened for voice threat warnings presented in a background of simulated helicopter cockpit noise. Performance was measured by flying-task performance, threat-warning intelligibility, and response time. Pilot ratings were elicited for the different voice pitches and speech rates. Significant effects were obtained only for response time and for pilot ratings, both as a function of speech rate. For the few cases when pilots forgot to respond to a voice message, they remembered 90 percent of the messages accurately when queried for their response 8 to 10 sec later.

Simpson, C. A.

Alerting prefixes for speech warning messages

A major question posed by the design of an integrated voice information display/warning system for next-generation helicopter cockpits is whether an alerting prefix should precede voice warning messages; if so, the characteristics desirable in such a cue must also be addressed. Attention is presently given to the results of a study which ascertained pilot response time and response accuracy to messages preceded by either neutral cues or the cognitively appropriate semantic cues. Both verbal cues and messages were spoken in direct, phoneme-synthesized speech, and a training manipulation was included to determine the extent to which previous exposure to speech thus produced facilitates these messages' comprehension. Results are discussed in terms of the importance of human factors research in cockpit display design.

Bucher, N. M.

Learning to read aloud: A neural network approach using sparse distributed memory

An attempt to solve a problem of text-to-phoneme mapping is described which does not appear amenable to solution by use of standard algorithmic procedures. Experiments based on a model of distributed processing are also described. This model (sparse distributed memory (SDM)) can be used in an iterative supervised learning mode to solve the problem. Additional improvements aimed at obtaining better performance are suggested.

Joglekar, Umesh Dwarkanath

Syntactic error modeling and scoring normalization in speech recognition

The objective was to develop the speech recognition system to be able to detect speech which is pronounced incorrectly, given that the text of the spoken speech is known to the recognizer. Research was performed in the following areas: (1) syntactic error modeling; (2) score normalization; and (3) phoneme error modeling. The study into the types of errors that a reader makes will provide the basis for creating tests which will approximate the use of the system in the real world. NASA-Johnson will develop this technology into a 'Literacy Tutor' in order to bring innovative concepts to the task of teaching adults to read.

Olorenshaw, Lex

Linguistics

Studies of Latvian morphophonemics, metrics of Arabic poetry, and intersection of regular languages and languages generated by transformational grammars

LANGUAGE