Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “lock latency”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

Accelerating Parallel Applications in Cloud Platforms via Adaptive Time-Slice Control

Cloud platforms can provide flexible and cost-effective environments for parallel applications. However, the resource over-commitment issues, i.e., cloud providers often provide much more executable virtual CPUs than available physical CPUs, still impede the synchronization operations of parallel applications, causing severe performance degradation. Existing methods optimize parallel applications by promoting the priorities of involved VMs. They cannot fully explore the performance of parallel applications, because they ignore the time-slice requirements of different phases of parallel applications. Furthermore, non-parallel applications experience unsatisfied performance because of low scheduling priorities. Given empirical analysis on time-slices of virtual machines (VMs), we find that shortening time-slices can mitigate synchronization overhead which incurs during communication phases, while over-short time-slices cause frequent cache misses in computation phases. Accordingly, we propose an Adaptive Time-slice Control (ATC) mechanism. ATC first detects the phases of parallel applications based on lock latency or cache misses. Then, ATC shortens time-slices during communication phases and prolongs time-slices during computation phases for parallel applications, and sets a uniform time-slice for non-parallel applications. Finally, we evaluate ATC using seven well-known benchmarks with 25+ applications. Experiments show that ATC obtains 1.5-75x performance gain for running parallel applications than state-of-the-art solutions, with nearly unaffected impact on non-parallel applications.

97 MATHEMATICS AND COMPUTING↗

Fast Offset Laser Phase-Locking System

Figure 1 shows a simplified block diagram of an improved optoelectronic system for locking the phase of one laser to that of another laser with an adjustable offset frequency specified by the user. In comparison with prior systems, this system exhibits higher performance (including higher stability) and is much easier to use. The system is based on a field-programmable gate array (FPGA) and operates almost entirely digitally; hence, it is easily adaptable to many different systems. The system achieves phase stability of less than a microcycle. It was developed to satisfy the phase-stability requirement for a planned spaceborne gravitational-wave-detecting heterodyne laser interferometer (LISA). The system has potential terrestrial utility in communications, lidar, and other applications. The present system includes a fast phasemeter that is a companion to the microcycle-accurate one described in High-Accuracy, High-Dynamic-Range Phase-Measurement System (NPO-41927), NASA Tech Briefs, Vol. 31, No. 6 (June 2007), page 22. In the present system (as in the previously reported one), beams from the two lasers (here denoted the master and slave lasers) interfere on a photodiode. The heterodyne photodiode output is digitized and fed to the fast phasemeter, which produces suitably conditioned, low-latency analog control signals which lock the phase of the slave laser to that of the master laser. These control signals are used to drive a thermal and a piezoelectric transducer that adjust the frequency and phase of the slave-laser output. The output of the photodiode is a heterodyne signal at the difference between the frequencies of the two lasers. (The difference is currently required to be less than 20 MHz due to the Nyquist limit of the current sampling rate. We foresee few problems in doubling this limit using current equipment.) Within the phasemeter, the photodiode-output signal is digitized to 15 bits at a sampling frequency of 40 MHz by use of the same analog-to-digital converter (ADC) as that of the previously reported phasemeter. The ADC output is passed to the FPGA, wherein the signal is demodulated using a digitally generated oscillator signal at the offset locking frequency specified by the user. The demodulated signal is low-pass filtered, decimated to a sample rate of 1 MHz, then filtered again. The decimated and filtered signal is converted to an analog output by a 1 MHz, 16-bit digital-to-analog converters. After a simple low-pass filter, these analog signals drive the thermal and piezoelectric transducers of the laser.

Shaddock, Daniel↗

Understanding the Effects of Spaceflight on Head-trunk Coordination During Walking and Obstacle Avoidance

Prolonged exposure to spaceflight conditions results in a battery of physiological changes, some of which contribute to sensorimotor and neurovestibular deficits. Upon return to Earth, functional performance changes are tested using the Functional Task Test (FTT), which includes an obstacle course to observe post-flight balance and postural stability, specifically during turning. The goal of this study was to quantify changes in movement strategies during turning events by observing the latency between head-and-trunk coordinated movements. It was hypothesized that subjects experiencing neurovestibular adaptations would exhibit head-to-trunk locking ('en bloc' movement) during turning, exhibited by a decrease in latency between head and trunk movement. FTT data samples were collected from 13 ISS astronauts and 26 male 70-day head down tilt bed rest subjects, including bed rest controls (10 BRC) and bed rest exercisers (16 BRE). Samples were analyzed three times pre-exposure, immediately post-exposure (0 or 1 day post) and 2-to-3 times during recovery from the unloading environment. Two 3D inertial measurements units (XSens MTx) were attached to subjects, one on the head and one on the upper back. This study focused primarily on the yaw movements about the subject's center of rotation. Time differences (latency) between head and trunk movement were averaged across a slalom obstacle portion, consisting of three turns (approximately three 60⁰ turns). All participants were grouped as 'decreaser' or 'increaser,' relating to their change in head-to-trunk movement latency between pre- and post- environmental adaptation measures. Space flight unloading (ISS) showed a bimodal response between the 'increaser' and 'decreaser' group, while both bed rest control (BRC) and bed rest exercise (BRE) populations showed increased preference towards a 'decreaser' categorization, displaying greater head-trunk locking. It is clear that changes in movement strategies are adopted during exposure to an unloading environment. These results further the understanding of vestibular-somatosensory convergence and support the use of bed rest as an exclusionary model to better understand sensorimotor changes in space flight.

Madansingh, S.↗

Understanding the Effects of Spaceflight on Head-trunk Coordination during Walking and Obstacle Avoidance

Prolonged exposure to spaceflight conditions results in a battery of physiological changes, some of which contribute to sensorimotor and neurovestibular deficits. Upon return to Earth, functional performance changes are tested using the Functional Task Test (FTT), which includes an obstacle course to observe post‐flight balance and postural stability, specifically during turning. Aims: To quantify changes in movement strategies during turning events by observing the latency between head‐andtrunk coordinated movement. Hypothesis: It is hypothesized that subjects experiencing neurovestibular adaptations will exhibit head‐to‐trunk locking ('en bloc' movement) during turning, exhibited by a decrease in latency between head and trunk movement. Sample: FTT data samples were collected from Shuttle and ISS missions. Samples were analyzed three times pre exposure, immediately post‐exposure (0 or 1 day post) and 2‐to‐3 times during recovery from the microgravity environment. Methods: Two 3D inertial measurements units (XSens MTx) were attached to subjects, one on the head and one on the upper back. This study focused primarily on the yaw movements about the subject's center of rotation. Time differences (latency) between head and trunk movement were calculated at two points: the first turn (Fturn) to enter the obstacle course (approximately 90⁰ turn) and averaged across a slalom obstacle portion, consisting of three turns (approximately three 90⁰ turns). Results: Preliminary analysis of the data shows a trend toward decreasing head‐to‐trunk movement latency during post‐flight ambulation, after reintroduction to Earth gravity in Shuttle and ISS astronauts. Conclusion: It is clear that changes in movement strategies are adopted during exposure to the microgravity environment and upon reintroduction to a gravity environment. Some subjects exhibit symptoms of neurovestibular neuropathy ('en bloc movement') that may impact their ability to perform post‐flight functional tasks.

Madansingh, S.↗

Understanding the Effects of Spaceflight on Head-trunk Coordination during Walking and Obstacle Avoidance

Prolonged exposure to spaceflight conditions results in a battery of physiological changes, some of which contribute to sensorimotor and neurovestibular deficits. Upon return to Earth, functional performance changes are tested using the Functional Task Test (FTT), which includes an obstacle course to observe post‐flight balance and postural stability, specifically during turning. The goal of this study was to quantify changes in movement strategies during turning events by observing the latency between head‐and‐trunk coordinated movements. It was hypothesized that subjects experiencing neurovestibular adaptations would exhibit head‐to‐trunk locking ('en bloc' movement) during turning, exhibited by a decrease in latency between head and trunk movement. FTT data samples were collected from ISS missions. Samples were analyzed three times pre‐exposure, immediately postexposure (1 day post) and 2‐to‐3 times during recovery from the microgravity environment. Two 3D inertial measurements units (XSens MTx) were attached to subjects, one on the head and one on the upper back. This study focused primarily on the yaw movements about the subject's center of rotation. Time differences (latency) between head and trunk movement were calculated at two points on the obstacle course: the first turn to enter the obstacle course (approximately 90⁰ turn) and averaged across a slalom obstacle portion, consisting of three turns (approximately three 90⁰ turns). Preliminary analysis of the data shows a trend toward decreasing head‐to‐trunk movement latency during postflight ambulation in slalom turning after reintroduction to Earth gravity in ISS astronauts. It is clear that changes in movement strategies are adopted during exposure to the microgravity environment and upon reintroduction to a gravity environment. Most ISS subjects exhibit symptoms of neurovestibular changes ('en bloc head and trunk movement) which may impact their ability to perform post‐flight functional tasks.

Madansingh, S.↗

The suitability of the ILLIAC IV architecture for image processing

The major architectural features of the ILLIAC IV large scale, array processor are summarized along with their applicability to image processing. Several image processing algorithms are considered, including multispectral classification, texture feature extraction, two-dimensional Fourier transform, and synthetic aperture radar processing. The basic parallelism of the ILLIAC IV (64 processing elements acting in lock-step) is usually fully utilized by the image processing applications. The major architectural aspect of the system with respect to image processing is the relatively small local scratch-pad memory and the long latency time to access the main storage device. The major precision used for the image processing applications is the 32-bit floating point, given a choice of 8-bit integers and 64-bit floating point.

Stevenson, D. K.↗

Electron cyclotron emission detection of neoclassical tearing modes for control for ITER

Successful operation of ITER requires control of magnetic instabilities including neoclassical tearing modes (NTMs) that can degrade confinement and lead to disruption. Low latency detection by electron cyclotron emission (ECE) diagnostics has been demonstrated in a few current experiments. Using a synthetic diagnostic, we demonstrate low latency NTM detection for ITER with plasmas described by ITER IMAS database scenarios and with realistic limitations imposed on the instrumentation by these high temperature scenarios. 2/1 NTMs are detected 430 ms after magnetic island seeding and before island locking. The radiometer configuration was optimized using simulation, and the smallest detectable island size was explored. Island sizes of ∼3 cm are detectable at the 2/1 surface. The simulated signals incorporate recent physics models for island growth and rotation, which show early locking and continued island growth after locking and before disruption. This work determines limits for ITER ECE spatial resolution imposed by relativistic broadening of channels, which informs hardware design. Real-time detection is demonstrated in hardware that is required by ITER, including on an NI PXI-7853R FPGA system. Development of a synthetic diagnostic and details of the hardware will be discussed.

Cyclotron radiation↗

NASA Tech Briefs, June 2014

Topics include: Real-Time Minimization of Tracking Error for Aircraft Systems; Detecting an Extreme Minority Class in Hyperspectral Data Using Machine Learning; KSC Spaceport Weather Data Archive; Visualizing Acquisition, Processing, and Network Statistics Through Database Queries; Simulating Data Flow via Multiple Secure Connections; Systems and Services for Near-Real-Time Web Access to NPP Data; CCSDS Telemetry Decoder VHDL Core; Thermal Response of a High-Power Switch to Short Pulses; Solar Panel and System Design to Reduce Heating and Optimize Corridors for Lower-Risk Planetary Aerobraking; Low-Cost, Very Large Diamond-Turned Metal Mirror; Very-High-Load-Capacity Air Bearing Spindle for Large Diamond Turning Machines; Elevated-Temperature, Highly Emissive Coating for Energy Dissipation of Large Surfaces; Catalyst for Treatment and Control of Post-Combustion Emissions; Thermally Activated Crack Healing Mechanism for Metallic Materials; Subsurface Imaging of Nanocomposites; Self-Healing Glass Sealants for Solid Oxide Fuel Cells and Electrolyzer Cells; Micromachined Thermopile Arrays with Novel Thermo - electric Materials; Low-Cost, High-Performance MMOD Shielding; Head-Mounted Display Latency Measurement Rig; Workspace-Safe Operation of a Force- or Impedance-Controlled Robot; Cryogenic Mixing Pump with No Moving Parts; Seal Design Feature for Redundancy Verification; Dexterous Humanoid Robot; Tethered Vehicle Control and Tracking System; Lunar Organic Waste Reformer; Digital Laser Frequency Stabilization via Cavity Locking Employing Low-Frequency Direct Modulation; Deep UV Discharge Lamps in Capillary Quartz Tubes with Light Output Coupled to an Optical Fiber; Speech Acquisition and Automatic Speech Recognition for Integrated Spacesuit Audio Systems, Version II; Advanced Sensor Technology for Algal Biotechnology; High-Speed Spectral Mapper; "Ascent - Commemorating Shuttle" - A NASA Film and Multimedia Project DVD; High-Pressure, Reduced-Kinetics Mechanism for N-Hexadecane Oxidation; Method of Error Floor Mitigation in Low-Density Parity-Check Codes; X-Ray Flaw Size Parameter for POD Studies; Large Eddy Simulation Composition Equations for Two-Phase Fully Multicomponent Turbulent Flows; Scheduling Targeted and Mapping Observations with State, Resource, and Timing Constraints;

Source record↗

A high performance totally ordered multicast protocol

This paper presents the Reliable Multicast Protocol (RMP). RMP provides a totally ordered, reliable, atomic multicast service on top of an unreliable multicast datagram service such as IP Multicasting. RMP is fully and symmetrically distributed so that no site bears un undue portion of the communication load. RMP provides a wide range of guarantees, from unreliable delivery to totally ordered delivery, to K-resilient, majority resilient, and totally resilient atomic delivery. These QoS guarantees are selectable on a per packet basis. RMP provides many communication options, including virtual synchrony, a publisher/subscriber model of message delivery, an implicit naming service, mutually exclusive handlers for messages, and mutually exclusive locks. It has commonly been held that a large performance penalty must be paid in order to implement total ordering -- RMP discounts this. On SparcStation 10's on a 1250 KB/sec Ethernet, RMP provides totally ordered packet delivery to one destination at 842 KB/sec throughput and with 3.1 ms packet latency. The performance stays roughly constant independent of the number of destinations. For two or more destinations on a LAN, RMP provides higher throughput than any protocol that does not use multicast or broadcast.

Montgomery, Todd↗

Analysis and visualization of single-trial event-related potentials

In this study, a linear decomposition technique, independent component analysis (ICA), is applied to single-trial multichannel EEG data from event-related potential (ERP) experiments. Spatial filters derived by ICA blindly separate the input data into a sum of temporally independent and spatially fixed components arising from distinct or overlapping brain or extra-brain sources. Both the data and their decomposition are displayed using a new visualization tool, the "ERP image," that can clearly characterize single-trial variations in the amplitudes and latencies of evoked responses, particularly when sorted by a relevant behavioral or physiological variable. These tools were used to analyze data from a visual selective attention experiment on 28 control subjects plus 22 neurological patients whose EEG records were heavily contaminated with blink and other eye-movement artifacts. Results show that ICA can separate artifactual, stimulus-locked, response-locked, and non-event-related background EEG activities into separate components, a taxonomy not obtained from conventional signal averaging approaches. This method allows: (1) removal of pervasive artifacts of all types from single-trial EEG records, (2) identification and segregation of stimulus- and response-locked EEG components, (3) examination of differences in single-trial responses, and (4) separation of temporally distinct but spatially overlapping EEG oscillatory activities with distinct relationships to task events. The proposed methods also allow the interaction between ERPs and the ongoing EEG to be investigated directly. We studied the between-subject component stability of ICA decomposition of single-trial EEG epochs by clustering components with similar scalp maps and activation power spectra. Components accounting for blinks, eye movements, temporal muscle activity, event-related potentials, and event-modulated alpha activities were largely replicated across subjects. Applying ICA and ERP image visualization to the analysis of sets of single trials from event-related EEG (or MEG) experiments can increase the information available from ERP (or ERF) data. Copyright 2001 Wiley-Liss, Inc.

Non-NASA Center↗

Custom Accessors: Enabling Scalable Data Ingestion, (Re-)Organization, and Analysis on Distributed Systems

The emerging class of high velocity and high volume data analytic workflows comprise interwoven data ingestion, organization, and processing stages, with ingestion and organization steps often contributing comparable or even higher computational costs than actual processing steps. Since complex workflows consist of a variety of phases that view and use data differently, being able to construct efficient, scalable, distributed data structures (arrays, vectors, sets, maps, and multi-maps) is essential and requires custom methods to extend and shrink containers, analyze and position data, and, maintain globallyconsistent meta-data. In this paper, we propose a novel datastructure access paradigm based on the concept of Accessors. At a high level, accessors are customizable callable objects that can modify the behavior of insert, read, update, and delete operations for distributed containers while preserving atomicity guarantees. Accessors provide a very clean and natural way to implement a variety of programming patterns, e.g., conditional insertion/deletion and cascading computations, which would be otherwise hard (or even impossible) to express in parallel and distributed settings without using locks. We demonstrate the practicality and usefulness of our approach with two representative use cases and study the performance of these applications on a distributed High-Performance Computing system. Our analysis highlights that our proposed abstraction allows for an effective overlapping and concurrent execution of different workflow steps (e.g., data ingestion and analysis), which in a conventional analytics pipeline would execute sequentially, contributing cumulatively to the overall latency.

Castellana, Vito G. [BATTELLE (PACIFIC NW LAB)] (O↗

NASA Tech Briefs, September 2009

opics covered include: Filtering Water by Use of Ultrasonically Vibrated Nanotubes; Computer Code for Nanostructure Simulation; Functionalizing CNTs for Making Epoxy/CNT Composites; Improvements in Production of Single-Walled Carbon Nanotubes; Progress Toward Sequestering Carbon Nanotubes in PmPV; Two-Stage Variable Sample-Rate Conversion System; Estimating Transmitted-Signal Phase Variations for Uplink Array Antennas; Board Saver for Use with Developmental FPGAs; Circuit for Driving Piezoelectric Transducers; Digital Synchronizer without Metastability; Compact, Low-Overhead, MIL-STD-1553B Controller; Parallel-Processing CMOS Circuitry for M-QAM and 8PSK TCM; Differential InP HEMT MMIC Amplifiers Embedded in Waveguides; Improved Aerogel Vacuum Thermal Insulation; Fluoroester Co-Solvents for Low-Temperature Li+ Cells; Using Volcanic Ash to Remove Dissolved Uranium and Lead; High-Efficiency Artificial Photosynthesis Using a Novel Alkaline Membrane Cell; Silicon Wafer-Scale Substrate for Microshutters and Detector Arrays; Micro-Horn Arrays for Ultrasonic Impedance Matching; Improved Controller for a Three-Axis Piezoelectric Stage; Nano-Pervaporation Membrane with Heat Exchanger Generates Medical-Grade Water; Micro-Organ Devices; Nonlinear Thermal Compensators for WGM Resonators; Dynamic Self-Locking of an OEO Containing a VCSEL; Internal Water Vapor Photoacoustic Calibration; Mid-Infrared Reflectance Imaging of Thermal-Barrier Coatings; Improving the Visible and Infrared Contrast Ratio of Microshutter Arrays; Improved Scanners for Microscopic Hyperspectral Imaging; Rate-Compatible LDPC Codes with Linear Minimum Distance; PrimeSupplier Cross-Program Impact Analysis and Supplier Stability Indicator Simulation Model; Integrated Planning for Telepresence With Time Delays; Minimizing Input-to-Output Latency in Virtual Environment; Battery Cell Voltage Sensing and Balancing Using Addressable Transformers; Gaussian and Lognormal Models of Hurricane Gust Factors; Simulation of Attitude and Trajectory Dynamics and Control of Multiple Spacecraft; Integrated Modeling of Spacecraft Touch-and-Go Sampling; Spacecraft Station-Keeping Trajectory and Mission Design Tools; Efficient Model-Based Diagnosis Engine; and DSN Simulator.

Source record↗

Responses evoked from man by acoustic stimulation

Clicks and other acoustic stimuli evoke time-locked responses from the brain of man. The properties of the waves recordable within the interval from 1 to 10 msec after the stimuli strike the eardrum are discussed along with factors influencing the waves in the 100 to 500 msec epoch. So-called brainstem responses from a normal young adult are considered. No waves were observed for clicks to weak to be heard. With increasing stimulus strength the waves become larger in amplitude and their latency shortens.

Galambos, R.↗

Cooperative Data Sharing: Simple Support for Clusters of SMP Nodes

Libraries like PVM and MPI send typed messages to allow for heterogeneous cluster computing. Lower-level libraries, such as GAM, provide more efficient access to communication by removing the need to copy messages between the interface and user space in some cases. still lower-level interfaces, such as UNET, get right down to the hardware level to provide maximum performance. However, these are all still interfaces for passing messages from one process to another, and have limited utility in a shared-memory environment, due primarily to the fact that message passing is just another term for copying. This drawback is made more pertinent by today's hybrid architectures (e.g. clusters of SMPs), where it is difficult to know beforehand whether two communicating processes will share memory. As a result, even portable language tools (like HPF compilers) must either map all interprocess communication, into message passing with the accompanying performance degradation in shared memory environments, or they must check each communication at run-time and implement the shared-memory case separately for efficiency. Cooperative Data Sharing (CDS) is a single user-level API which abstracts all communication between processes into the sharing and access coordination of memory regions, in a model which might be described as "distributed shared messages" or "large-grain distributed shared memory". As a result, the user programs to a simple latency-tolerant abstract communication specification which can be mapped efficiently to either a shared-memory or message-passing based run-time system, depending upon the available architecture. Unlike some distributed shared memory interfaces, the user still has complete control over the assignment of data to processors, the forwarding of data to its next likely destination, and the queuing of data until it is needed, so even the relatively high latency present in clusters can be accomodated. CDS does not require special use of an MMU, which can add overhead to some DSM systems, and does not require an SPMD programming model. unlike some message-passing interfaces, CDS allows the user to implement efficient demand-driven applications where processes must "fight" over data, and does not perform copying if processes share memory and do not attempt concurrent writes. CDS also supports heterogeneous computing, dynamic process creation, handlers, and a very simple thread-arbitration mechanism. Additional support for array subsections is currently being considered. The CDS1 API, which forms the kernel of CDS, is built primarily upon only 2 communication primitives, one process initiation primitive, and some data translation (and marshalling) routines, memory allocation routines, and priority control routines. The entire current collection of 28 routines provides enough functionality to implement most (or all) of MPI 1 and 2, which has a much larger interface consisting of hundreds of routines. still, the API is small enough to consider integrating into standard os interfaces for handling inter-process communication in a network-independent way. This approach would also help to solve many of the problems plaguing other higher-level standards such as MPI and PVM which must, in some cases, "play OS" to adequately address progress and process control issues. The CDS2 API, a higher level of interface roughly equivalent in functionality to MPI and to be built entirely upon CDS1, is still being designed. It is intended to add support for the equivalent of communicators, reduction and other collective operations, process topologies, additional support for process creation, and some automatic memory management. CDS2 will not exactly match MPI, because the copy-free semantics of communication from CDS1 will be supported. CDS2 application programs will be free to carefully also use CDS1. CDS1 has been implemented on networks of workstations running unmodified Unix-based operating systems, using UDP/IP and vendor-supplied high- performance locks. Although its inter-node performance is currently unimpressive due to rudimentary implementation technique, it even now outperforms highly-optimized MPI implementation on intra-node communication due to its support for non-copy communication. The similarity of the CDS1 architecture to that of other projects such as UNET and TRAP suggests that the inter-node performance can be increased significantly to surpass MPI or PVM, and it may be possible to migrate some of its functionality to communication controllers.

DiNucci, David C.↗