Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “speech recognition”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

Characterization of Response Times Based on Voice Communication and Traffic Surveillance Data

A barrier to the integration of remotely piloted aircraft operations in the U.S. National Airspace System is the latency of voice communications between the air traffic controller and the remote pilot, and the latency of communication between the aircraft and the remote pilot. The latency can be substantial especially when satellite-based beyond-radio-line-of-sight communication and relay through the aircraft are employed. This study uses voice recordings of controller-pilot communications and aircraft track data to establish a baseline of pilot readback latencies and maneuver detection delays in the current piloted operations. A machine learning pipeline was developed to parse the contents of the air traffic control clearances including the callsigns using natural language processing. After manually validating the results obtained using the pipeline, the average pilot readback latency was found to be about 0.6 seconds. The average latency between the end of maneuver (inferred from track data), initiated by the pilot in response to the clearance, and the end of clearance was found to be about 176 seconds for altitude change commands, 69 seconds for heading change commands, and 182 seconds for speed change commands. The average latency between the beginning of maneuver and the end of clearance was found to be about 17 seconds for altitude change commands, 17seconds for heading change commands, and 25 seconds for speed change commands.

controller-pilot communication↗

Robotics control using isolated word recognition of voice input

A speech input/output system is presented that can be used to communicate with a task oriented system. Human speech commands and synthesized voice output extend conventional information exchange capabilities between man and machine by utilizing audio input and output channels. The speech input facility is comprised of a hardware feature extractor and a microprocessor implemented isolated word or phrase recognition system. The recognizer offers a medium sized (100 commands), syntactically constrained vocabulary, and exhibits close to real time performance. The major portion of the recognition processing required is accomplished through software, minimizing the complexity of the hardware feature extractor.

Weiner, J. M.↗

Neural-based time series forecasting of loss of coolant accidents in nuclear power plants

During the last few years, deep learning in neural networks has demonstrated impressive successes in the areas of computer vision, speech and image recognition, text generation, and many others. However, sensitive engineering areas such as nuclear engineering benefited less from these efficient techniques. In this work, deep learning expert systems are utilized to model and predict time series progression of a design-basis nuclear accident, featuring a loss of coolant accident. Two major findings are accomplished in this work. First, the ability to train expert systems with high accuracy, which could help nuclear power plant operators to figure out plant responses during the accident. Second, building fast, efficient, and accurate deep models to simulate nuclear phenomena, which could be valuable to nuclear computational science. In this work, large amount of time series data is obtained from simulation tools by simulating different conditions of the base-case/nominal accident scenario. Four critical outputs/responses are monitored during the accident (e.g. temperature, pressure, break flow rate, water level). Two approaches are adopted in this work. The first approach is to use feedforward deep neural networks (DNN) to fit all time steps and outputs in a single model. The second approach is to use long short-term memory (LSTM) to fit all time steps together for each reactor response separately. Both DNN and LSTM demonstrate very good performance in predicting the test and base-case scenarios, with accuracy as low as 92% and as high as 99%, where these test scenarios are unknown to the expert systems and are not included in the model training. In addition, both approaches demonstrate a significant reduction in computational costs, as the deep expert system is able to accurately predict the accident 100,000 times faster than the original simulation tool. Given sufficient data, the methodology adopted in this study demonstrates that DNN/LSTM expert systems can be used as a decision support system to model advanced time series phenomena within nuclear power plants with high accuracy and negligible computational costs.

42 ENGINEERING↗

Quantum machine learning with differential privacy

Abstract Quantum machine learning (QML) can complement the growing trend of using learned models for a myriad of classification tasks, from image recognition to natural speech processing. There exists the potential for a quantum advantage due to the intractability of quantum operations on a classical computer. Many datasets used in machine learning are crowd sourced or contain some private information, but to the best of our knowledge, no current QML models are equipped with privacy-preserving features. This raises concerns as it is paramount that models do not expose sensitive information. Thus, privacy-preserving algorithms need to be implemented with QML. One solution is to make the machine learning algorithm differentially private, meaning the effect of a single data point on the training dataset is minimized. Differentially private machine learning models have been investigated, but differential privacy has not been thoroughly studied in the context of QML. In this study, we develop a hybrid quantum-classical model that is trained to preserve privacy using differentially private optimization algorithm. This marks the first proof-of-principle demonstration of privacy-preserving QML. The experiments demonstrate that differentially private QML can protect user-sensitive information without signficiantly diminishing model accuracy. Although the quantum model is simulated and tested on a classical computer, it demonstrates potential to be efficiently implemented on near-term quantum devices [noisy intermediate-scale quantum (NISQ)]. The approach’s success is illustrated via the classification of spatially classed two-dimensional datasets and a binary MNIST classification. This implementation of privacy-preserving QML will ensure confidentiality and accurate learning on NISQ technology.

97 MATHEMATICS AND COMPUTING↗

Integrated voice and visual systems research topics

A series of studies was performed to investigate factors of helicopter speech and visual system design and measure the effects of these factors on human performance, both for pilots and non-pilots. The findings and conclusions of these studies were applied by the U.S. Army to the design of the Army's next generation threat warning system for helicopters and to the linguistic functional requirements for a joint Army/NASA flightworthy, experimental speech generation and recognition system.

Williams, Douglas H.↗

Experimental Test-Bed for Intelligent Passive Array Research

This document describes the test-bed designed for the investigation of passive direction finding, recognition, and classification of speech and sound sources using sensor arrays. The test-bed forms the experimental basis of the Intelligent Small-Scale Spatial Direction Finder (ISS-SDF) project, aimed at furthering digital signal processing and intelligent sensor capabilities of sensor array technology in applications such as rocket engine diagnostics, sensor health prognostics, and structural anomaly detection. This form of intelligent sensor technology has potential for significant impact on NASA exploration, earth science and propulsion test capabilities. The test-bed consists of microphone arrays, power and signal distribution modules, web-based data acquisition, wireless Ethernet, modeling, simulation and visualization software tools. The Acoustic Sensor Array Modeler I (ASAM I) is used for studying steering capabilities of acoustic arrays and testing DSP techniques. Spatial sound distribution visualization is modeled using the Acoustic Sphere Analysis and Visualization (ASAV-I) tool.

Solano, Wanda M.↗

Acoustical and Intelligibility Test of the Vocera(Copyright) B3000 Communication Badge

To communicate with each other or ground support, crew members on board the International Space Station (ISS) currently use the Audio Terminal Units (ATU), which are located in each ISS module. However, to use the ATU, crew members must stop their current activity, travel to a panel, and speak into a wall-mounted microphone, or use either a handheld microphone or a Crew Communication Headset that is connected to a panel. These actions unnecessarily may increase task times, lower productivity, create cable management issues, and thus increase crew frustration. Therefore, the Habitability and Human Factors and Human Interface Branches at the NASA Johnson Space Center (JSC) are currently investigating a commercial-off-the-shelf (COTS) wireless communication system, Vocera(C), as a near-term solution for ISS communication. The objectives of the acoustics and intelligibility testing of this system were to answer the following questions: 1. How intelligibly can a human hear the transmitted message from a Vocera(c) badge in three different noise environments (Baseline = 20 dB, US Lab Module = 58 dB, Russian Module = 70.6 dB)? 2. How accurate is the Vocera(C) badge at recognizing voice commands in three different noise environments? 3. What body location (chest, upper arm, or shoulder) is optimal for speech intelligibility and voice recognition accuracy of the Vocera(C) badge on a human in three different noise environments?

Archer, Ronald↗

Bio-Inspired Neural Model for Learning Dynamic Models

A neural-network mathematical model that, relative to prior such models, places greater emphasis on some of the temporal aspects of real neural physical processes, has been proposed as a basis for massively parallel, distributed algorithms that learn dynamic models of possibly complex external processes by means of learning rules that are local in space and time. The algorithms could be made to perform such functions as recognition and prediction of words in speech and of objects depicted in video images. The approach embodied in this model is said to be "hardware-friendly" in the following sense: The algorithms would be amenable to execution by special-purpose computers implemented as very-large-scale integrated (VLSI) circuits that would operate at relatively high speeds and low power demands.

Duong, Tuan↗

Systems concept for speech technology application in general aviation

The application potential of voice recognition and synthesis circuits for general aviation, single-pilot IFR (SPIFR) situations is examined. The viewpoint of the pilot was central to workload analyses and assessment of the effectiveness of the voice systems. A twin-engine, high performance general aviation aircraft on a cross-country fixed route was employed as the study model. No actual control movements were considered and other possible functions were scored by three IFR-rated instructors. The SPIFR was concluded helpful in alleviating visual and manual workloads during take-off, approach and landing, particularly for data retrieval and entry tasks. Voice synthesis was an aid in alerting a pilot to in-flight problems. It is expected that usable systems will be available within 5 yr.

North, R. A.↗

Digital signal processing algorithms for automatic voice recognition

The current digital signal analysis algorithms are investigated that are implemented in automatic voice recognition algorithms. Automatic voice recognition means, the capability of a computer to recognize and interact with verbal commands. The digital signal is focused on, rather than the linguistic, analysis of speech signal. Several digital signal processing algorithms are available for voice recognition. Some of these algorithms are: Linear Predictive Coding (LPC), Short-time Fourier Analysis, and Cepstrum Analysis. Among these algorithms, the LPC is the most widely used. This algorithm has short execution time and do not require large memory storage. However, it has several limitations due to the assumptions used to develop it. The other 2 algorithms are frequency domain algorithms with not many assumptions, but they are not widely implemented or investigated. However, with the recent advances in the digital technology, namely signal processors, these 2 frequency domain algorithms may be investigated in order to implement them in voice recognition. This research is concerned with real time, microprocessor based recognition algorithms.

Botros, Nazeih M.↗

Many but not all deep neural network audio models capture brain responses and exhibit correspondence between model stages and brain regions

Models that predict brain responses to stimuli provide one measure of understanding of a sensory system and have many potential applications in science and engineering. Deep artificial neural networks have emerged as the leading such predictive models of the visual system but are less explored in audition. Prior work provided examples of audio-trained neural networks that produced good predictions of auditory cortical fMRI responses and exhibited correspondence between model stages and brain regions, but left it unclear whether these results generalize to other neural network models and, thus, how to further improve models in this domain. We evaluated model-brain correspondence for publicly available audio neural network models along with in-house models trained on 4 different tasks. Most tested models outpredicted standard spectromporal filter-bank models of auditory cortex and exhibited systematic model-brain correspondence: Middle stages best predicted primary auditory cortex, while deep stages best predicted non-primary cortex. However, some state-of-the-art models produced substantially worse brain predictions. Models trained to recognize speech in background noise produced better brain predictions than models trained to recognize speech in quiet, potentially because hearing in noise imposes constraints on biological auditory representations. The training task influenced the prediction quality for specific cortical tuning properties, with best overall predictions resulting from models trained on multiple tasks. The results generally support the promise of deep neural networks as models of audition, though they also indicate that current models do not explain auditory cortical responses in their entirety.

59 BASIC BIOLOGICAL SCIENCES↗

Automatic voice recognition using traditional and artificial neural network approaches

The main objective of this research is to develop an algorithm for isolated-word recognition. This research is focused on digital signal analysis rather than linguistic analysis of speech. Features extraction is carried out by applying a Linear Predictive Coding (LPC) algorithm with order of 10. Continuous-word and speaker independent recognition will be considered in future study after accomplishing this isolated word research. To examine the similarity between the reference and the training sets, two approaches are explored. The first is implementing traditional pattern recognition techniques where a dynamic time warping algorithm is applied to align the two sets and calculate the probability of matching by measuring the Euclidean distance between the two sets. The second is implementing a backpropagation artificial neural net model with three layers as the pattern classifier. The adaptation rule implemented in this network is the generalized least mean square (LMS) rule. The first approach has been accomplished. A vocabulary of 50 words was selected and tested. The accuracy of the algorithm was found to be around 85 percent. The second approach is in progress at the present time.

Botros, Nazeih M.↗

Inverse Text Normalization of Air Traffic Control System Command Center Planning Telecon Transcriptions

We present a hybrid neural network and rule-based Inverse Text Normalization (ITN) method for domains containing unique technical phraseology, specifically Air Traffic Control System Command Center (ATCSCC) planning telecon audio transcriptions. The ATCSCC hosts bi-hourly planning telephone conferences (or planning telecons) to ensure smooth operations within the National Airspace (NAS). Access to both live and post meeting transcripts of this speech audio would enable quick review of meetings. Provided speech transcripts, ITN is the process of converting unformatted raw Automated Speaker Recognition (ASR) model transcripts into a human (expert) readable written form. Our hybrid ITN framework utilizes a fine-tuned Bidirectional Encoder Representations from Transformers neural network to format conversational English, and rule-based methods to format domain-specific aviation text. With an overall Punctuation Error Rate (PER) of 25.56 and Word Error Rate with Punctuation and Capitalization (WER PC) of 5.47, we show that this method has vast potential in being applied to ATCSCC planning telecon audio and other audio/text based data available in ATM.

ATM↗

Inverse Text Normalization of Air Traffic Control System Command Center Planning Telecon Transcriptions

We present a hybrid neural network and rule-based Inverse Text Normalization (ITN) method for domains containing unique technical phraseology, specifically Air Traffic Control System Command Center (ATCSCC) planning telecon audio transcriptions. The ATCSCC hosts bi-hourly planning telephone conferences (or planning telecons) to ensure smooth operations within the National Airspace (NAS). Access to both live and post meeting transcripts of this speech audio would enable quick review of meetings. Provided speech transcripts, ITN is the process of converting unformatted raw Automated Speaker Recognition (ASR) model transcripts into a human (expert) readable written form. Our hybrid ITN framework utilizes a fine-tuned Bidirectional Encoder Representations from Transformers neural network to format conversational English, and rule-based methods to format domain-specific aviation text. With an overall Punctuation Error Rate (PER) of 25.56 and Word Error Rate with Punctuation and Capitalization (WER PC) of 5.47, we show that this method has vast potential in being applied to ATCSCC planning telecon audio and other audio/text based data available in ATM.

ATM↗

Man/computer communication in a space environment

The present work reports on a study of the technology required to advance the state of the art in man/machine communications. The study involved the development and demonstration of both hardware and software to effectively implement man/computer interactive channels of communication. While tactile and visual man/computer communications equipment are standard methods of interaction with machines, man's speech is a natural media for inquiry and control. As part of this study, a word recognition unit was developed capable of recognizing a minimum of one hundred different words or sentences in any one of the currently used conversational languages. The study has proven that efficiency in communication between man and computer can be achieved when the vocabulary to be used is structured in a manner compatible with the rigid communication requirements of the machine while at the same time responsive to the informational needs of the man.

Hodges, B. C.↗

A Model Based Approach to Extract Health Information from Textual Data

In current nuclear power plants (NPPs) a large amount of condition-based data is being generated and stored to assess and monitor component health and performance. The format of this data can be either numeric (e.g., pump vibration data) or textual (e.g., condition report which assess component health). While assessing component health from numeric data can be performed with a large variety of methods, the extraction of information from textual data still remains a challenge. Natural language processing (NLP) methods are starting to be deployed in current NPPs mainly to filter out incident reports (IRs) that are not safety related by employing supervised machine learning methods. However, these methods do not really provide the quantitative information that might be contained in IRs. This paper presents an approach to extract information from textual data (e.g., from IRs, maintenance reports) that is based on NLP data analytics methods coupled with model-based system engineer (MBSE) models. NLP methods are employed to perform syntactic and semantic analyses. Syntactic analysis analyzes the grammatical structure of a sentence; such analysis includes: part of speech (POS) tagging (i.e., identification of grammatic elements of each string - e.g., nouns, verbs), named entity recognition (i.e., identification of text entities - e.g., names, dates, events), and relation extraction (e.g., coreference resolution). On the other hand, semantic analysis is designed to analyze the logic structure of a sentence. Through a specific set of rules, our methods can identify whether a sentence contains health information of a component (e.g., degraded performance, anomaly behavior) or the causal relationship between two events (i.e., a cause-effect pair). An innovative element of our approach is that semantic analysis relies on MBSE models to identify links between textual elements. MBSE are diagrams designed to represent system and component dependencies (from both a form and functional point of view). In our approach, MBSE models emulate system engineer knowledge about component/system architecture. This paper presents in detail how the integration of NLP methods and MBSE models is performed. Few analysis examples focusing on centrifugal pumps are presented.

97 - MATHEMATICS AND COMPUTING↗