Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “context switch”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Context Switching with Multiple Register Windows: A RISC Performance Study

Although previous studies have shown that a large file of overlapping register windows can greatly reduce procedure call/return overhead, the effects of register windows in a multiprogramming environment are poorly understood. This paper investigates the performance of multiprogrammed, reduced instruction set computers (RISCs) as a function of window management strategy. Using an analytic model that reflects context switch and procedure call overheads, we analyze the performance of simple, linearly self-recursive programs. For more complex programs, we present the results of a simulation study. These studies show that a simple strategy that saves all windows prior to a context switch, but restores only a single window following a context switch, performs near optimally.

Konsek, Marion B.↗

Ferroelectric FET-based context-switching FPGA enabling dynamic reconfiguration for adaptive deep learning machines

Field programmable gate array (FPGA) is widely used in the acceleration of deep learning applications because of its reconfigurability, flexibility, and fast time-to-market. However, conventional FPGA suffers from the trade-off between chip area and reconfiguration latency, making efficient FPGA accelerations that require switching between multiple configurations still elusive. Here, we propose a ferroelectric field-effect transistor (FeFET)–based context-switching FPGA supporting dynamic reconfiguration to break this trade-off, enabling loading of arbitrary configuration without interrupting the active configuration execution. Leveraging the intrinsic structure and nonvolatility of FeFETs, compact FPGA primitives are proposed and experimentally verified. The evaluation results show our design shows a 63.0%/74.7% reduction in a look-up table (LUT)/connection block (CB) area and 82.7%/53.6% reduction in CB/switch box power consumption with a minimal penalty in the critical path delay (9.6%). Besides, our design yields significant time savings by 78.7 and 20.3% on average for context-switching and dynamic reconfiguration applications, respectively.

97 MATHEMATICS AND COMPUTING↗

Fast Context Switching in Real-Time Propositional Reasoning

The trend to increasingly capable and affordable control processors has generated an explosion of embedded real-time gadgets that serve almost every function imaginable. The daunting task of programming these gadgets is greatly alleviated with real-time deductive engines that perform all execution and monitoring functions from a single core model, Fast response times are achieved using an incremental propositional deductive database (an LTMS). Ideally the cost of an LTMS's incremental update should be linear in the number of labels that change between successive contexts. Unfortunately an LTMS can expend a significant percentage of its time working on labels that remain constant between contexts. This is caused by the LTMS's conservative approach: a context switch first removes all consequences of deleted clauses, whether or not those consequences hold in the new context. This paper presents a more aggressive incremental TMS, called the ITMS, that avoids processing a significant number of these consequences that are unchanged. Our empirical evaluation for spacecraft control shows that the overhead of processing unchanged consequences can be reduced by a factor of seven.

Nayak, P. Pandurang↗

Efficient Ada multitasking on a RISC register window architecture

This work addresses the problem of reducing context switch overhead on a processor which supports a large register file - a register file much like that which is part of the Berkeley RISC processors and several other emerging architectures (which are not necessarily reduced instruction set machines in the purest sense). Such a reduction in overhead is particularly desirable in a real-time embedded application, in which task-to-task context switch overhead may result in failure to meet crucial deadlines. A storage management technique by which a context switch may be implemented as cheaply as a procedure call is presented. The essence of this technique is the avoidance of the save/restore of registers on the context switch. This is achieved through analysis of the static source text of an Ada tasking program. Information gained during that analysis directs the optimized storage management strategy for that program at run time. A formal verification of the technique in terms of an operational control model and an evaluation of the technique's performance via simulations driven by synthetic Ada program traces are presented.

Kearns, J. P.↗

Analyzing the Performance Trade-Off in Implementing User-Level Threads

User-level threads have been widely adopted as a means of achieving lightweight concurrent execution without the costs of OS-level threads. Nevertheless, the costs of managing user-level threads represent a performance barrier that dictates how fine grained the concurrency exposed by an application can be without incurring significant overheads; this in turn may translate into insufficient parallelism to exploit highly parallel systems. This article is a deep dive into the fundamental costs in implementing user-level threads. We first identify that one of the highest sources of fork-join overheads stems from deviations, events that incur context switching during the execution of a thread and disrupt a run-to-completion execution. We then conduct an in-depth investigation of a wide spectrum of methods with respect to how they handle deviations while covering both parent- and child-first scheduling policies. Our methodology involves a comprehensive instruction- and cache-level analysis of all methods on several modern CPU architectures. Finally, the primary finding of our evaluation is that dynamic promotion methods that assume the absence of deviation and dynamically provide context-switching support offer the best trade-off between performance and capability when the likelihood of deviation is low.

97 MATHEMATICS AND COMPUTING↗

Algorithms and Libraries

This exploratory study initiated our inquiry into algorithms and applications that would benefit by latency tolerant approach to algorithm building, including the construction of new algorithms where appropriate. In a multithreaded execution, when a processor reaches a point where remote memory access is necessary, the request is sent out on the network and a context--switch occurs to a new thread of computation. This effectively masks a long and unpredictable latency due to remote loads, thereby providing tolerance to remote access latency. We began to develop standards to profile various algorithm and application parameters, such as the degree of parallelism, granularity, precision, instruction set mix, interprocessor communication, latency etc. These tools will continue to develop and evolve as the Information Power Grid environment matures. To provide a richer context for this research, the project also focused on issues of fault-tolerance and computation migration of numerical algorithms and software. During the initial phase we tried to increase our understanding of the bottlenecks in single processor performance. Our work began by developing an approach for the automatic generation and optimization of numerical software for processors with deep memory hierarchies and pipelined functional units. Based on the results we achieved in this study we are planning to study other architectures of interest, including development of cost models, and developing code generators appropriate to these architectures.

Dongarra, Jack↗

Testing SOAR tools in use

Investigations within Security Operation Centers (SOCs) are tedious as they rely on manual efforts to query diverse data sources, overlay related logs, correlate the data into information, and then document results in a ticketing system. Security Orchestration, Automation, and Response (SOAR) tools are a relatively new technology that promise, with appropriate configuration, to collect, filter, and display needed diverse information; automate many of the common tasks that unnecessarily require SOC analysts’ time; facilitate SOC collaboration; and, in doing so, improve both efficiency and consistency of SOCs. There has been no prior research to test SOAR tools in practice; hence, understanding and evaluation of their effect is nascent and needed. Here, in this paper, we design and administer the first hands-on user study of SOAR tools, involving 24 participants and six commercial SOAR tools. Our contributions include the experimental design, itemizing six characteristics of SOAR tools, and a methodology for testing them. We describe configuration of a cyber range test environment, including network, user, and threat emulation; a full SOC tool suite; and creation of artifacts allowing multiple representative investigation scenarios to permit testing. We present the first research results on SOAR tools. Concisely, our findings are that: per-SOC SOAR configuration is extremely important; SOAR tools increase efficiency and reduce context switching, although with potentially decreased ticketing accuracy/completeness; user preference is slightly negatively correlated with their performance with the tool; internet dependence varies widely among SOAR tools; and balance of automation with assisting decision making is preferred by senior participants. We deliver a public user- and tool-anonymized and -obfuscated version of the data.

97 MATHEMATICS AND COMPUTING↗

UMap: An application-oriented user level memory mapping library

Exploiting the prominent role of complex memories in exascale node architecture, the UMap page fault handler offers new capabilities to access large memory-mapped data sets directly. UMap provides flexible configuration options to customize page handling to each application, including analysis of massive observational and simulation data sets. The high-performance design features I/O decoupling, dynamic load balancing, and application-level controls. Page faults triggered by application threads and processes accessing data mapped to a UMapp’ed region are handled via the Linux userfaultfd protocol, an asynchronous message-oriented kernel-user communication mechanism that avoids the context switch penalty of traditional signal fault handlers. UMap is fully open source. In this paper, we give an overview of the UMap library architecture, its extensible plugin architecture, and the use/performance of UMap in emerging heterogeneous memory hierarchies such as near-node Non-volatile Memory (NVM) and network attached memories. We highlight new capabilities in two pagefault management plugins, the NetworkStore and SparseStore. We demonstrate the integration between UMap and multiple ECP products including Caliper, Metall, ZFP, Mochi, and Ripples.

97 MATHEMATICS AND COMPUTING↗

Using C to build a satellite scheduling expert system: Examples from the Explorer platform planning system

Recently, many expert systems were developed in a LISP environment and then ported to the real world C environment before the final system is delivered. This situation may require that the entire system be completely rewritten in C and may actually result in a system which is put together as quickly as possible with little regard for maintainability and further evolution. With the introduction of high performance UNIX and X-windows based workstations, a great deal of the advantages of developing a first system in the LISP environment have become questionable. A C-based AI development effort is described which is based on a software tools approach with emphasis on reusability and maintainability of code. The discussion starts with simple examples of how list processing can easily be implemented in C and then proceeds to the implementations of frames and objects which use dynamic memory allocation. The implementation of procedures which use depth first search, constraint propagation, context switching and a blackboard-like simulation environment are described. Techniques for managing the complexity of C-based AI software are noted, especially the object-oriented techniques of data encapsulation and incremental development. Finally, all these concepts are put together by describing the components of planning software called the Planning And Resource Reasoning (PARR) shell. This shell was successfully utilized for scheduling services of the Tracking and Data Relay Satellite System for the Earth Radiation Budget Satellite since May 1987 and will be used for operations scheduling of the Explorer Platform in November 1991.

Mclean, David R.↗

Using C to build a satellite scheduling expert system: Examples from the Explorer Platform planning system

A C-based artificial intelligence (AI) development effort which is based on a software tools approach is discussed with emphasis on reusability and maintainability of code. The discussion starts with simple examples of how list processing can easily be implemented in C and then proceeds to the implementations of frames and objects which use dynamic memory allocation. The implementation of procedures which use depth first search, constraint propagation, context switching, and blackboard-like simulation environment are described. Techniques for managing the complexity of C-based AI software are noted, especially the object-oriented techniques of data encapsulation and incremental development. Finally, all these concepts are put together by describing the components of planning software called the Planning And Resource Reasoning (PARR) Shell. This shell was successfully utilized for scheduling services of the Tracking and Data Relay Satellite System for the Earth Radiation Budget Satellite since May of 1987 and will be used for operations scheduling of the Explorer Platform in Nov. of 1991.

Mclean, David R.↗

Context-specific adaptation of saccade gain

Previous studies established that vestibular reflexes can have two adapted states (e.g., gain) simultaneously, and that a context cue (e.g., vertical eye position) can switch between the two states. The present study examined this phenomenon of context-specific adaptationfor horizontal saccades, using a variety of contexts. Our overall goal was to assess the efficacy of different context cues in switching between adapted states. A standard double-step paradigm was used to adapt saccade gain. In each experiment, we asked for a simultaneous gain decrease in one context and gain increase in another context, and then determined if a change in the context would invoke switching between the adapted states. Horizontal eye position worked well as a context cue: saccades with the eyes deviated to the right could be made to have higher gains while saccades with the eyes deviated to the left could be made to have lower gains. Vertical eye position was less effective. This suggests that the more closely related a context cue is to the response being adapted, the more effective it is. Roll tilt of the head, and upright versus supine orientations, were somewhat effective in context switching; these paradigms contain orientation of gravity with respect to the head as part of the context.

NASA Discipline Neuroscience↗

Performance Study of CXL Memory Topology

This paper presents a comprehensive evaluation of the performance impact of various Compute Express Link (CXL) memory topologies, with a particular emphasis on CXL switches, in the context of High- Performance Computing (HPC) and Large Language Model (LLM) inference workloads. Our study unveils significant performance variations across different topologies, demonstrating that certain configurations yield superior performance for specific workloads. These findings underscore the critical importance of tailored topol- ogy selection in optimizing system performance. Additionally, we address the inherent challenges associated with integrating CXL switches, including overhead considerations and routing complex- ities. Our research highlights the necessity for thorough evalua- tion methodologies to fully leverage CXL technology’s potential in contemporary computing environments. These insights provide valuable guidance for system architects and data center operators in designing and optimizing CXL-based infrastructures for diverse workload requirements.

CXL, memory, Artificial Intelligence (AI), HPC↗

Context-specific adaptation of saccade gain in parabolic flight

Previous studies established that vestibular reflexes can have two adapted states (e.g., gains) simultaneously, and that a context cue (e.g., vertical eye position) can switch between the two states. Our earlier work demonstrated this phenomenon of context-specific adaptation for saccadic eye movements: we asked for gain decrease in one context state and gain increase in another context state, and then determined if a change in the context state would invoke switching between the adapted states. Horizontal and vertical eye position and head orientation could serve, to varying degrees, as cues for switching between two different saccade gains. In the present study, we asked whether gravity magnitude could serve as a context cue: saccade adaptation was performed during parabolic flight, which provides alternating levels of gravitoinertial force (0 g and 1.8 g). Results were less robust than those from ground experiments, but established that different saccade magnitudes could be associated with different gravity levels.

Clinical Trial↗

The Selection of Q-Switch for a 350mJ Air-borne 2-micron Wind Lidar

In the process of designing a coherent, high energy 2micron, Doppler wind Lidar, various types of Q-Switch materials and configurations have been investigated for the oscillator. Designing an oscillator with a relatively low gain laser material, presents challenges related to the management high internal circulating fluence due to high reflective output coupler. This problem is compounded by the loss of hold-off. In addition, the selection has to take into account the round trip optical loss in the resonator and the loss of hold-off. For this application, a Brewster cut 5mm aperture, fused silica AO Q-switch is selected. Once the Q-switch is selected various rf frequencies were evaluated. Since the Lidar has to perform in single longitudinal and transverse mode with transform limited line width, in this paper, various seeding configurations are presented in the context of Q-Switch diffraction efficiency. The master oscillator power amplifier has demonstrated over 350mJ output when the amplifier is operated in double pass mode and higher than 250mJ when operated in single pass configuration. The repetition rate of the system is 10Hz and the pulse length 200ns.

Petros, Mulugeta↗

Context-specific adaptation of the gain of the oculomotor response to lateral translation using roll and pitch head tilts as contexts

Previous studies established that vestibular and oculomotor behaviors can have two adapted states (e.g., gain) simultaneously, and that a context cue (e.g., vertical eye position) can switch between the two states. The present study examined this phenomenon of context-specific adaptation for the oculomotor response to interaural translation (which we term "linear vestibulo-ocular reflex" or LVOR even though it may have extravestibular components). Subjects sat upright on a linear sled and were translated at 0.7 Hz and 0.3 gpeak acceleration while a visual-vestibular mismatch paradigm was used to adaptively increase (x2) or decrease (x0) the gain of the LVOR. In each experimental session, gain increase was asked for in one context, and gain decrease in another context. Testing in darkness with steps and sines before and after adaptation, in each context, assessed the extent to which the context itself could recall the gain state that was imposed in that context during adaptation. Two different contexts were used: head pitch (26 degrees forward and backward) and head roll (26 degrees or 45 degrees, right and left). Head roll tilt worked well as a context cue: with the head rolled to the right the LVOR could be made to have a higher gain than with the head rolled to the left. Head pitch tilt was less effective as a context cue. This suggests that the more closely related a context cue is to the response being adapted, the more effective it is.

Non-NASA Center↗

Next generation communications satellites: Multiple access and network studies

Following an overview of issues involved in the choice of promising system architectures for efficient communication with multiple small inexpensive Earth stations serving hetergeneous user populations, performance evaluation via analysis and simulation for six SS/TDMA (satellite-switched/time-division multiple access) system architectures is discussed. These configurations are chosen to exemplify the essential alternatives available in system design. Although the performance evaluation analyses are of fairly general applicability, whenever possible they are considered in the context of NASA's 30/20 GHz studies. Packet switched systems are considered, with the assumption that only a part of transponder capacit is devoted to packets, the integration of circuit and packet switched traffic being reserved for further study. Three types of station access are distinguished: fixed (FA), demand (DA), and random access (RA). Similarly, switching in the satellite can be assigned on a fixed (FS) or demand (DS) basis, or replaced by a buffered store-and-forward system (SF) onboard the satellite. Since not all access/switching combinations are practical, six systems are analyzed in detail: three FS SYSTEMS, FA/FS, DA/ES, RA/FS; one DS system, DA/DS; and two SF systems, FA/SF, DA/SF. Results are presented primarily in terms of delay-throughput characteristics.

Stern, T. E.↗

REC protein family expansion by the emergence of a new signaling pathway

This report presents multi-genome evidence that REC protein family expansion occurs when the emergence of new pathways gives rise to functional discordance. Specificity between residues in REC domain containing response regulators with paired histidine kinases is under negative purifying selection, constrained by the presence of other bacterial two-component systems signaling cascades that share sequence and structural identity. Presuming that the two-component systems can evolve by neutral amino acid changes (neutral drift) when purifying evolutionary constraints are relaxed, how might the REC protein family expand by amino acid changes when these constraints remain intact? Using an unsupervised machine learning approach to observe the sequence landscape of REC domains across long phylogenetic distances, we find that within-gene recombination, a subcategory of gene conversion, switched the effector domain and, consequently, the regulatory context of a duplicated response regulator from transcriptional regulation by σ54 to that by σ70. We determined that the recombined response regulator diverged from its parent by episodic diversifying selection and neutral drift. Functional experiments of the parent of recombined response regulators in a model Pseudomonas putida KT2440 model system revealed that the parent and recombined response regulators sense and respond to different carboxylic acids. Finally, a residue-switching experiment using structural predictions and functional characterization suggests that the new residues in the recombined regulator could form a new interaction interface and mediate condition-specific phosphotransfer. Overall, our study finds that genetic perturbations can create conditions of functional discordance, whereby the REC protein family can evolve by episodic diversifying selection.

59 BASIC BIOLOGICAL SCIENCES↗

Smart Hand Offs and Earthdata Search

Often a science user will discover data of interest in a general-purpose discovery tool like Earthdata Search. At that point it might be beneficial for the science user to switch to a more specialized tool. Smart handoffs allow the user to switch from one tool to another with their context (collection, spatial and temporal constraints) intact. This provides an efficient means of traversing tools and services. The idea of smart handoffs could be expanded to commercial search engines such as Google, visualization tools like State of the Ocean and Giovanni and domain-specific search tools like Earthdata Search.

Giovanni↗