Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “speech recognition”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

Autonomy Voice Assistant for NPAS (NASA Platform for Autonomous Systems)

A prototype voice interaction system, Autonomy Voice Assistant (AVA), is described in this paper. AVA is designed to seamlessly integrate into the NASA Platform for Autonomous Systems (NPAS), an autonomy software platform, and to enable an operator to interact with NPAS autonomy applications through voice conversations. By integrating VA with NPAS, a major enhancement to NPAS applications is facilitated, enabling interaction through natural language expressions. An AVA prototype has been designed incorporating two principles:(1) self-containment (no external data or computations required), and (2) a readily modifiable, reconfigurable, and flexible architecture. By using voice messages in an NPAS application, an additional layer of user interface capability is enabled, thereby enhancing a user’s overall experience. Advancements, over the past several decades in speech recognition and natural language processing technologies has made it possible for AVA to implement robust messaging capabilities while still being lightweight. The main objective of incorporating a voice assistant like AVA is to augment the number and effectiveness of interactions a user has with a system that typically uses mouse-based interaction, while simultaneously enriching the user experience and providing heightened system awareness.

Lucian Murdock↗

Understanding and Verifying Neural Networks

Deep Neural Networks (DNNs) have gained immense popularity in recent times and have widespread use in applications such as image classification, sentiment analysis, speech recognition and also in safety-critical applications such as autonomous driving. However, they suffer limitations such as lack of explainability and robustness which raise safety and security concerns in their usage. Further, the complex structure and large input spaces of DNNs act as an impediment to thorough verification and testing. The SafeDNN project at the Robust Software Engineering (RSE) group at NASA aims at exploring techniques to ensure that systems that use deep neural networks are safe, robust and interpretable. In this talk, I will be presenting our technique Prophecy that automatically infers formal properties of deep neural network models. The tool extracts patterns based on neuron activations as preconditions that imply certain desirable output properties of the model. I would be highlighting case studies that use Prophecy in obtaining explanations for network decisions, understanding correct and incorrect behavior, providing formal guarantees wrt safety and robustness, and debugging neural network models. We have applied the tool on image classification networks, neural network controllers providing turn advisories in unmanned aircrafts, regression models used for autonomous center-line tracking in aircrafts and neural network object detectors

Deep Neural Networks↗

Intelligent Response and Interaction System (IRIS) - FY21 Closeout Report

In the second year, the IRIS team developed the components necessary for successful offline deployment of the IRIS services. This includes custom automated speech recognition training on NASA audio data, and in-house development and integration of online and offline conversational services. Lastly, the team worked on integrating the IRIS technology with stakeholders and projects that have a strong need for voice interaction.

Aly Shehata↗

Speech variability effects on recognition accuracy associated with concurrent task performance by pilots

In the present study of the responses of pairs of pilots to aircraft warning classification tasks using an isolated word, speaker-dependent speech recognition system, the induced stress was manipulated by means of different scoring procedures for the classification task and by the inclusion of a competitive manual control task. Both speech patterns and recognition accuracy were analyzed, and recognition errors were recorded by type for an isolated word speaker-dependent system and by an offline technique for a connected word speaker-dependent system. While errors increased with task loading for the isolated word system, there was no such effect for task loading in the case of the connected word system.

Simpson, C. A.↗

Speech therapy and voice recognition instrument

Characteristics of electronic circuit for examining variations in vocal excitation for diagnostic purposes and in speech recognition for determiniog voice patterns and pitch changes are described. Operation of the circuit is discussed and circuit diagram is provided.

Cohen, J.↗

Characterization of Response Times based on Voice Communication and Traffic Surveillance Data

A barrier to the integration of remotely piloted aircraft operations in the U.S. National Airspace System is the latency of voice communications between the air traffic controller and the remote pilot, and the latency of communication between the aircraft and the remote pilot. The latency can be substantial especially when satellite-based beyond-radio-line-of-sight communication and relay through the aircraft are employed. This study uses voice recordings of controller-pilot communications and aircraft track data to establish a baseline of pilot readback latencies and maneuver detection delays in the current piloted operations. A machine learning pipeline was developed to parse the contents of the air traffic control clearances including the callsigns using natural language processing. After manually validating the results obtained using the pipeline, the average pilot readback latency was found to be about 0.6 seconds. The average latency between the end of maneuver (inferred from track data), initiated by the pilot in response to the clearance, and the end of clearance was found to be about 176 seconds for altitude change commands, 69 seconds for heading change commands, and 182 seconds for speed change commands. The average latency between the beginning of maneuver and the end of clearance was found to be about 17 seconds for altitude change commands, 17seconds for heading change commands, and 25 seconds for speed change commands.

controller-pilot communication, communication late↗

Characterization of Response Times Based on Voice Communication and Traffic Surveillance Data

A barrier to the integration of remotely piloted aircraft operations in the U.S. National Airspace System is the latency of voice communications between the air traffic controller and the remote pilot, and the latency of communication between the aircraft and the remote pilot. The latency can be substantial especially when satellite-based beyond-radio-line-of-sight communication and relay through the aircraft are employed. This study uses voice recordings of controller-pilot communications and aircraft track data to establish a baseline of pilot readback latencies and maneuver detection delays in the current piloted operations. A machine learning pipeline was developed to parse the contents of the air traffic control clearances including the callsigns using natural language processing. After manually validating the results obtained using the pipeline, the average pilot readback latency was found to be about 0.6 seconds. The average latency between the end of maneuver (inferred from track data), initiated by the pilot in response to the clearance, and the end of clearance was found to be about 176 seconds for altitude change commands, 69 seconds for heading change commands, and 182 seconds for speed change commands. The average latency between the beginning of maneuver and the end of clearance was found to be about 17 seconds for altitude change commands, 17seconds for heading change commands, and 25 seconds for speed change commands.

controller-pilot communication↗

Robotics control using isolated word recognition of voice input

A speech input/output system is presented that can be used to communicate with a task oriented system. Human speech commands and synthesized voice output extend conventional information exchange capabilities between man and machine by utilizing audio input and output channels. The speech input facility is comprised of a hardware feature extractor and a microprocessor implemented isolated word or phrase recognition system. The recognizer offers a medium sized (100 commands), syntactically constrained vocabulary, and exhibits close to real time performance. The major portion of the recognition processing required is accomplished through software, minimizing the complexity of the hardware feature extractor.

Weiner, J. M.↗

Integrated voice and visual systems research topics

A series of studies was performed to investigate factors of helicopter speech and visual system design and measure the effects of these factors on human performance, both for pilots and non-pilots. The findings and conclusions of these studies were applied by the U.S. Army to the design of the Army's next generation threat warning system for helicopters and to the linguistic functional requirements for a joint Army/NASA flightworthy, experimental speech generation and recognition system.

Williams, Douglas H.↗

Experimental Test-Bed for Intelligent Passive Array Research

This document describes the test-bed designed for the investigation of passive direction finding, recognition, and classification of speech and sound sources using sensor arrays. The test-bed forms the experimental basis of the Intelligent Small-Scale Spatial Direction Finder (ISS-SDF) project, aimed at furthering digital signal processing and intelligent sensor capabilities of sensor array technology in applications such as rocket engine diagnostics, sensor health prognostics, and structural anomaly detection. This form of intelligent sensor technology has potential for significant impact on NASA exploration, earth science and propulsion test capabilities. The test-bed consists of microphone arrays, power and signal distribution modules, web-based data acquisition, wireless Ethernet, modeling, simulation and visualization software tools. The Acoustic Sensor Array Modeler I (ASAM I) is used for studying steering capabilities of acoustic arrays and testing DSP techniques. Spatial sound distribution visualization is modeled using the Acoustic Sphere Analysis and Visualization (ASAV-I) tool.

Solano, Wanda M.↗

Acoustical and Intelligibility Test of the Vocera(Copyright) B3000 Communication Badge

To communicate with each other or ground support, crew members on board the International Space Station (ISS) currently use the Audio Terminal Units (ATU), which are located in each ISS module. However, to use the ATU, crew members must stop their current activity, travel to a panel, and speak into a wall-mounted microphone, or use either a handheld microphone or a Crew Communication Headset that is connected to a panel. These actions unnecessarily may increase task times, lower productivity, create cable management issues, and thus increase crew frustration. Therefore, the Habitability and Human Factors and Human Interface Branches at the NASA Johnson Space Center (JSC) are currently investigating a commercial-off-the-shelf (COTS) wireless communication system, Vocera(C), as a near-term solution for ISS communication. The objectives of the acoustics and intelligibility testing of this system were to answer the following questions: 1. How intelligibly can a human hear the transmitted message from a Vocera(c) badge in three different noise environments (Baseline = 20 dB, US Lab Module = 58 dB, Russian Module = 70.6 dB)? 2. How accurate is the Vocera(C) badge at recognizing voice commands in three different noise environments? 3. What body location (chest, upper arm, or shoulder) is optimal for speech intelligibility and voice recognition accuracy of the Vocera(C) badge on a human in three different noise environments?

Archer, Ronald↗

Bio-Inspired Neural Model for Learning Dynamic Models

A neural-network mathematical model that, relative to prior such models, places greater emphasis on some of the temporal aspects of real neural physical processes, has been proposed as a basis for massively parallel, distributed algorithms that learn dynamic models of possibly complex external processes by means of learning rules that are local in space and time. The algorithms could be made to perform such functions as recognition and prediction of words in speech and of objects depicted in video images. The approach embodied in this model is said to be "hardware-friendly" in the following sense: The algorithms would be amenable to execution by special-purpose computers implemented as very-large-scale integrated (VLSI) circuits that would operate at relatively high speeds and low power demands.

Duong, Tuan↗

Systems concept for speech technology application in general aviation

The application potential of voice recognition and synthesis circuits for general aviation, single-pilot IFR (SPIFR) situations is examined. The viewpoint of the pilot was central to workload analyses and assessment of the effectiveness of the voice systems. A twin-engine, high performance general aviation aircraft on a cross-country fixed route was employed as the study model. No actual control movements were considered and other possible functions were scored by three IFR-rated instructors. The SPIFR was concluded helpful in alleviating visual and manual workloads during take-off, approach and landing, particularly for data retrieval and entry tasks. Voice synthesis was an aid in alerting a pilot to in-flight problems. It is expected that usable systems will be available within 5 yr.

North, R. A.↗

Digital signal processing algorithms for automatic voice recognition

The current digital signal analysis algorithms are investigated that are implemented in automatic voice recognition algorithms. Automatic voice recognition means, the capability of a computer to recognize and interact with verbal commands. The digital signal is focused on, rather than the linguistic, analysis of speech signal. Several digital signal processing algorithms are available for voice recognition. Some of these algorithms are: Linear Predictive Coding (LPC), Short-time Fourier Analysis, and Cepstrum Analysis. Among these algorithms, the LPC is the most widely used. This algorithm has short execution time and do not require large memory storage. However, it has several limitations due to the assumptions used to develop it. The other 2 algorithms are frequency domain algorithms with not many assumptions, but they are not widely implemented or investigated. However, with the recent advances in the digital technology, namely signal processors, these 2 frequency domain algorithms may be investigated in order to implement them in voice recognition. This research is concerned with real time, microprocessor based recognition algorithms.

Botros, Nazeih M.↗

Automatic voice recognition using traditional and artificial neural network approaches

The main objective of this research is to develop an algorithm for isolated-word recognition. This research is focused on digital signal analysis rather than linguistic analysis of speech. Features extraction is carried out by applying a Linear Predictive Coding (LPC) algorithm with order of 10. Continuous-word and speaker independent recognition will be considered in future study after accomplishing this isolated word research. To examine the similarity between the reference and the training sets, two approaches are explored. The first is implementing traditional pattern recognition techniques where a dynamic time warping algorithm is applied to align the two sets and calculate the probability of matching by measuring the Euclidean distance between the two sets. The second is implementing a backpropagation artificial neural net model with three layers as the pattern classifier. The adaptation rule implemented in this network is the generalized least mean square (LMS) rule. The first approach has been accomplished. A vocabulary of 50 words was selected and tested. The accuracy of the algorithm was found to be around 85 percent. The second approach is in progress at the present time.

Botros, Nazeih M.↗

Inverse Text Normalization of Air Traffic Control System Command Center Planning Telecon Transcriptions

We present a hybrid neural network and rule-based Inverse Text Normalization (ITN) method for domains containing unique technical phraseology, specifically Air Traffic Control System Command Center (ATCSCC) planning telecon audio transcriptions. The ATCSCC hosts bi-hourly planning telephone conferences (or planning telecons) to ensure smooth operations within the National Airspace (NAS). Access to both live and post meeting transcripts of this speech audio would enable quick review of meetings. Provided speech transcripts, ITN is the process of converting unformatted raw Automated Speaker Recognition (ASR) model transcripts into a human (expert) readable written form. Our hybrid ITN framework utilizes a fine-tuned Bidirectional Encoder Representations from Transformers neural network to format conversational English, and rule-based methods to format domain-specific aviation text. With an overall Punctuation Error Rate (PER) of 25.56 and Word Error Rate with Punctuation and Capitalization (WER PC) of 5.47, we show that this method has vast potential in being applied to ATCSCC planning telecon audio and other audio/text based data available in ATM.

ATM↗