Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “speech communication”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Communication as group process mediator of aircrew performance

Considering recent operating experience as a group-level input factor, aspects of the communication process between crewmembers (captain and first officer) were explored as a possible mediator to performance. Communication patterns were defined by a speech-act typology adapted for the flight-deck setting and distinguished crews that had previously flown together (FT) from those that had not flown together (NFT). A more open communication channel and greater first officer participation in task-related topics was shown by FT crews, while NFT crews engaged in more nontask discourse.

Kanki, Barbara G.↗

The feasibility of miniaturizing the versatile portable speech prosthesis: A market survey of commercial products

The feasibility of a miniature versatile portable speech prosthesis (VPSP) was analyzed and information on its potential users and on other similar devices was collected. The VPSP is a device that incorporates speech synthesis technology. The objective is to provide sufficient information to decide whether there is valuable technology to contribute to the miniaturization of the VPSP. The needs of potential users are identified, the development status of technologies similar or related to those used in the VPSP are evaluated. The VPSP, a computer based speech synthesis system fits on a wheelchair. The purpose was to produce a device that provides communication assistance in educational, vocational, and social situations to speech impaired individuals. It is expected that the VPSP can be a valuable aid for persons who are also motor impaired, which explains the placement of the system on a wheelchair.

Walklet, T.↗

Collaboration in Controller-Pilot Communication

Like other forms of dialogue, air traffic control (ATC) communication is an act of collaboration between two or more people. Collaboration progresses more or less smoothly depending on speaker and listener strategies. For example, we have found that the way controllers organize and deliver messages influences how easily pilots understand these messages, which in turn determines how much time and effort is needed to successfully complete the transaction. In this talk, I will introduce a collaborative framework for investigating controller-pilot communication and then describe a set of studies that investigate ATC communication from two complementary directions. First, we focused on the impact of ATC message factors (e.g., length, speech rate) on the cognitive processes involved in ATC: communication. Second, we examined pilot factors that influence the amount of cognitive resources available for these communication processes. These studies also illustrate how the collaborate framework can help analyze the impact of proposed visual data link systems on ATC communication. Examining the joint effects of communication medium, message factors, and pilot/controller factors on performance should help improve air safety and communication efficiency. Increased efficiency is important for meeting the growing demands on the National Air System.

Morrow, Daniel↗

A study and experiment plan for digital mobile communication via satellite

The viability of mobile communications is examined within the context of a frequency division multiple access, single channel per carrier satellite system emphasizing digital techniques to serve a large population of users. The intent is to provide the mobile users with a grade of service consistant with the requirements for remote, rural (perhaps emergency) voice communications, but which approaches toll quality speech. A traffic model is derived on which to base the determination of the required maximum number of satellite channels to provide the anticipated level of service. Various voice digitalization and digital modulation schemes are reviewed along with a general link analysis of the mobile system. Demand assignment multiple access considerations and analysis tradeoffs are presented. Finally, a completed configuration is described.

Jones, J. J.↗

Characterization of Response Times based on Voice Communication and Traffic Surveillance Data

A barrier to the integration of remotely piloted aircraft operations in the U.S. National Airspace System is the latency of voice communications between the air traffic controller and the remote pilot, and the latency of communication between the aircraft and the remote pilot. The latency can be substantial especially when satellite-based beyond-radio-line-of-sight communication and relay through the aircraft are employed. This study uses voice recordings of controller-pilot communications and aircraft track data to establish a baseline of pilot readback latencies and maneuver detection delays in the current piloted operations. A machine learning pipeline was developed to parse the contents of the air traffic control clearances including the callsigns using natural language processing. After manually validating the results obtained using the pipeline, the average pilot readback latency was found to be about 0.6 seconds. The average latency between the end of maneuver (inferred from track data), initiated by the pilot in response to the clearance, and the end of clearance was found to be about 176 seconds for altitude change commands, 69 seconds for heading change commands, and 182 seconds for speed change commands. The average latency between the beginning of maneuver and the end of clearance was found to be about 17 seconds for altitude change commands, 17seconds for heading change commands, and 25 seconds for speed change commands.

controller-pilot communication, communication late↗

Characterization of Response Times Based on Voice Communication and Traffic Surveillance Data

A barrier to the integration of remotely piloted aircraft operations in the U.S. National Airspace System is the latency of voice communications between the air traffic controller and the remote pilot, and the latency of communication between the aircraft and the remote pilot. The latency can be substantial especially when satellite-based beyond-radio-line-of-sight communication and relay through the aircraft are employed. This study uses voice recordings of controller-pilot communications and aircraft track data to establish a baseline of pilot readback latencies and maneuver detection delays in the current piloted operations. A machine learning pipeline was developed to parse the contents of the air traffic control clearances including the callsigns using natural language processing. After manually validating the results obtained using the pipeline, the average pilot readback latency was found to be about 0.6 seconds. The average latency between the end of maneuver (inferred from track data), initiated by the pilot in response to the clearance, and the end of clearance was found to be about 176 seconds for altitude change commands, 69 seconds for heading change commands, and 182 seconds for speed change commands. The average latency between the beginning of maneuver and the end of clearance was found to be about 17 seconds for altitude change commands, 17seconds for heading change commands, and 25 seconds for speed change commands.

controller-pilot communication↗

Enhancing Air Traffic Control Planning with Automatic Speech Recognition

The decisions made during the Federal Aviation Administration Air Traffic Control System Command Center's planning teleconferences hold significant sway over the National Airspace System. Held every two hours, these teleconferences convene air traffic managers and stakeholders from across the nation to discuss airspace conditions, weather, and constraints, leading to the formulation and adjustment of traffic management initiatives. Given the critical nature of these decisions, the need for accurate and efficient record-keeping is paramount. In recent years, the application of automatic speech recognition has gained popularity across diverse industries, including aviation. While traditional applications focus on transcribing air traffic control communication, this paper explores a unique application of automatic speech recognition by converting the audio from planning teleconferences into text transcriptions. This innovative approach addresses key challenges in the field, presenting potential benefits for quality assurance, real-time participation, and downstream natural language processing tasks. A notable breakthrough in the machine learning community, namely the transformer neural network architecture, forms the backbone of the proposed solution in this paper. The transformer architecture's role in this research represents a paradigm shift in the efficiency of automatic speech recognition models. By reducing the amount of in-domain training data required, this architecture allows for the fine-tuning of such models like Whisper, originally pretrained on vast English speech datasets. The adaptability of the transformer architecture proves invaluable in capturing the nuances of aviation terminology and specific language used in planning teleconferences. Leveraging the Whisper model as a baseline, our research details the fine-tuning and validation using a dataset comprising 20 hours of meticulously transcribed planning teleconferences. Notably, the baseline pretrained Whisper model exhibited a word error rate of 18.77%. Through the fine-tuning process, the model achieved a substantial improvement, demonstrating an impressive performance with a reduced word error rate of 6.82%. This substantial decrease in WER not only highlights the effectiveness of the transformer architecture but also emphasizes the practical advancements achieved through the application of automatic speech recognition in this specific domain. The utilization of automatic speech recognition in planning teleconferences in this work introduces several novelties. Firstly, the creation of text transcriptions offers a valuable tool for quality assurance and facilitates the efficient review of teleconferences. This is an important aspect of the proposed solution, given the time-sensitive and high-stakes nature of decisions made during these meetings. Furthermore, text-searchable transcriptions provide a streamlined approach for locating and validating critical information, potentially saving hours of manual effort in searching through audio recordings. Moreover, our research identifies a key use case for external facilities and stakeholders. In situations where attendance at the planning teleconference is not feasible, having access to text transcriptions in real-time or shortly after the teleconference ends, proves to be a time-saving and informative resource. This feature enhances collaboration and ensures that stakeholders can stay abreast of important discussions and decisions even in their absence. Despite the efficiency gains facilitated by the transformer architecture in automatic speech recognition technology, it is essential to acknowledge the human factors in data creation. Subject matter experts play a crucial role in accurately transcribing planning teleconferences due to the specificity and complexity of the information discussed. The research dataset, consisting of 20 hours of transcribed planning teleconferences, forms the foundation for fine-tuning and validating the Whisper model. The achieved word error rate of 6.82% demonstrates promising advancements, particularly in recognizing essential aviation terminology within the teleconferences. In conclusion, this paper presents a comprehensive exploration of the application of automatic speech recognition in Air Traffic Control System Command Center planning teleconferences, leveraging the transformer architecture for enhanced efficiency. The novel contributions lie in the improved accessibility of decision-making records, real-time participation opportunities for external stakeholders, and the potential for downstream natural language processing advancements. As the aviation industry continues to evolve, the integration of automatic speech recognition technologies holds the promise of revolutionizing decision-making processes and contributing to the overall safety and efficiency of air traffic management.

ATM↗

Biomedical technology transfer. Applications of NASA science and technology

Ongoing projects described address: (1) intracranial pressure monitoring; (2) versatile portable speech prosthesis; (3) cardiovascular magnetic measurements; (4) improved EMG biotelemetry for pediatrics; (5) ultrasonic kidney stone disintegration; (6) pediatric roentgen densitometry; (7) X-ray spatial frequency multiplexing; (8) mechanical impedance determination of bone strength; (9) visual-to-tactile mobility aid for the blind; (10) Purkinje image eyetracker and stabilized photocoalqulator; (11) neurological applications of NASA-SRI eyetracker; (12) ICU synthesized speech alarm; (13) NANOPHOR: microelectrophoresis instrument; (14) WRISTCOM: tactile communication system for the deaf-blind; (15) medical applications of NASA liquid-circulating garments; and (16) hip prosthesis with biotelemetry. Potential transfer projects include a person-portable versatile speech prosthesis, a critical care transport sytem, a clinical information system for cardiology, a programmable biofeedback orthosis for scoliosis a pediatric long-bone reconstruction, and spinal immobilization apparatus.

Harrison, D. C.↗

The effect of simultaneous exposure on the attention selection and integration of segments and lexical tones by Urdu-Cantonese bilingual speakers

In the perceptual learning of lexical tones, an automatic and robust attention-to-phonology system enables native tonal listeners to adapt to acoustically non-optimal speech, such as phonetic conflicts in daily communications. Previous tone research reveals that non-native listeners who do not linguistically employ lexical tones in their mother tongue may find it challenging to attend to the tonal dimension or integrate it with the segmental features. However, it is unknown whether the attentional interference initially caused by a maternal attentional system would continue influencing the non-optimal tone perception for simultaneous bilingual teenagers. From an endpoint in the age of language acquisition, we investigate whether the tone-specific attention mechanism developed by the Urdu-Cantonese simultaneous bilinguals is automatic enough to assist them in adapting to a phonetically-conflicting environment. Three groups of teenagers engaged in a four-condition ABX task: Urdu-Cantonese simultaneous bilinguals, Cantonese native listeners, and Urdu-speaking, late learners of Cantonese. The results showed that although the simultaneous bilinguals could phonologically process Cantonese tones in a Cantonese-like way under a conflict-free listening condition, they still failed in adapting to the phonetic conflicts, especially the segment-induced ones. It thus demonstrated that the simultaneous exposure and years of regular education in Hong Kong local schools still could not automatically guarantee simultaneous bilingual processing of Cantonese tones. In interpreting the findings, it hypothesized that, except for simultaneous exposure, the development of a tone-specific attention mechanism is also likely to be L1-inhibitory, tone experience-driven, and language-specific for simultaneous bilinguals.

Ning, Jinghong↗

Noise and blast

Noise and blast environments are described, providing a definition of units and techniques of noise measurement and giving representative booster-launch and spacecraft noise data. The effects of noise on hearing sensitivity and performance are reviewed, and community response to noise exposure is discussed. Physiological, or nonauditory, effects of noise exposure are also treated, as are design criteria and methods for minimizing the noise effects of hearing sensitivity and communications. The low level sound detection and speech reception are included, along with subjective and behavioral responses to noise.

Hodge, D. C.↗

Automatic Speech Acquisition and Recognition for Spacesuit Audio Systems

NASA has a widely recognized but unmet need for novel human-machine interface technologies that can facilitate communication during astronaut extravehicular activities (EVAs), when loud noises and strong reverberations inside spacesuits make communication challenging. WeVoice, Inc., has developed a multichannel signal-processing method for speech acquisition in noisy and reverberant environments that enables automatic speech recognition (ASR) technology inside spacesuits. The technology reduces noise by exploiting differences between the statistical nature of signals (i.e., speech) and noise that exists in the spatial and temporal domains. As a result, ASR accuracy can be improved to the level at which crewmembers will find the speech interface useful. System components and features include beam forming/multichannel noise reduction, single-channel noise reduction, speech feature extraction, feature transformation and normalization, feature compression, and ASR decoding. Arithmetic complexity models were developed and will help designers of real-time ASR systems select proper tasks when confronted with constraints in computational resources. In Phase I of the project, WeVoice validated the technology. The company further refined the technology in Phase II and developed a prototype for testing and use by suited astronauts.

Ye, Sherry↗

Communication variations and aircrew performance

Crew-related communication variations and their effects on performance are examined. The communication analysis involves evaluating the performance of 18 pilots to a high-fidelity full-mission simulation. Initiating speech consists of four categories: commands, questions, observations, and dysfluencies. Response speech is coded as: reply, acknowledgements, and zero response. A standard form of communication has been adopted which should aid in the coordination process and enhance crew performance.

Kanki, Barbara G.↗

Pilot Workload and Speech Analysis: A Preliminary Investigation

Prior research has questioned the effectiveness of speech analysis to measure the stress, workload, truthfulness, or emotional state of a talker. The question remains regarding the utility of speech analysis for restricted vocabularies such as those used in aviation communications. A part-task experiment was conducted in which participants performed Air Traffic Control read-backs in different workload environments. Participant's subjective workload and the speech qualities of fundamental frequency (F0) and articulation rate were evaluated. A significant increase in subjective workload rating was found for high workload segments. F0 was found to be significantly higher during high workload while articulation rates were found to be significantly slower. No correlation was found to exist between subjective workload and F0 or articulation rate.

Bittner, Rachel M.↗

Orthogonal transform feasibility study

The application of various orthogonal transformations to communication was investigated, with particular emphasis placed on speech and visual signal processing. The fundamentals of the one- and two-dimensional orthogonal transforms and their application to speech and visual signals are treated in detail.

Robinson, G. S.↗

Mobile Agents: A Distributed Voice-Commanded Sensory and Robotic System for Surface EVA Assistance

A model-based, distributed architecture integrates diverse components in a system designed for lunar and planetary surface operations: spacesuit biosensors, cameras, GPS, and a robotic assistant. The system transmits data and assists communication between the extra-vehicular activity (EVA) astronauts, the crew in a local habitat, and a remote mission support team. Software processes ("agents"), implemented in a system called Brahms, run on multiple, mobile platforms, including the spacesuit backpacks, all-terrain vehicles, and robot. These "mobile agents" interpret and transform available data to help people and robotic systems coordinate their actions to make operations more safe and efficient. Different types of agents relate platforms to each other ("proxy agents"), devices to software ("comm agents"), and people to the system ("personal agents"). A state-of-the-art spoken dialogue interface enables people to communicate with their personal agents, supporting a speech-driven navigation and scheduling tool, field observation record, and rover command system. An important aspect of the engineering methodology involves first simulating the entire hardware and software system in Brahms, and then configuring the agents into a runtime system. Design of mobile agent functionality has been based on ethnographic observation of scientists working in Mars analog settings in the High Canadian Arctic on Devon Island and the southeast Utah desert. The Mobile Agents system is developed iteratively in the context of use, with people doing authentic work. This paper provides a brief introduction to the architecture and emphasizes the method of empirical requirements analysis, through which observation, modeling, design, and testing are integrated in simulated EVA operations.

Clancey, William J.↗

Call sign intelligibility improvement using a spatial auditory display

A spatial auditory display was used to convolve speech stimuli, consisting of 130 different call signs used in the communications protocol of NASA's John F. Kennedy Space Center, to different virtual auditory positions. An adaptive staircase method was used to determine intelligibility levels of the signal against diotic speech babble, with spatial positions at 30 deg azimuth increments. Non-individualized, minimum-phase approximations of head-related transfer functions were used. The results showed a maximal intelligibility improvement of about 6 dB when the signal was spatialized to 60 deg or 90 deg azimuth positions.

Begault, Durand R.↗

Man/computer communication in a space environment

The present work reports on a study of the technology required to advance the state of the art in man/machine communications. The study involved the development and demonstration of both hardware and software to effectively implement man/computer interactive channels of communication. While tactile and visual man/computer communications equipment are standard methods of interaction with machines, man's speech is a natural media for inquiry and control. As part of this study, a word recognition unit was developed capable of recognizing a minimum of one hundred different words or sentences in any one of the currently used conversational languages. The study has proven that efficiency in communication between man and computer can be achieved when the vocabulary to be used is structured in a manner compatible with the rigid communication requirements of the machine while at the same time responsive to the informational needs of the man.

Hodges, B. C.↗

Clarissa Spoken Dialogue System for Procedure Reading and Navigation

Speech is the most natural modality for humans use to communicate with other people, agents and complex systems. A spoken dialogue system must be robust to noise and able to mimic human conversational behavior, like correcting misunderstandings, answering simple questions about the task and understanding most well formed inquiries or commands. The system aims to understand the meaning of the human utterance, and if it does not, then it discards the utterance as being meant for someone else. The first operational system is Clarissa, a conversational procedure reader and navigator, which will be used in a System Development Test Objective (SDTO) on the International Space Station (ISS) during Expedition 10. In the present environment one astronaut reads the procedure on a Manual Procedure Viewer (MPV) or paper, and has to stop to read or turn pages, shifting focus from the task. Clarissa is designed to read and navigate ISS procedures entirely with speech, while the astronaut has his eyes and hands engaged in performing the task. The system also provides an MPV like graphical interface so the procedure can be read visually. A demo of the system will be given.

Hieronymus, James↗