Engineering PapersSearch

SEARCH · Engineering Papers

Results for “attention”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Contingent attentional capture or delayed allocation of attention?

Under certain circumstances, external stimuli will elicit an involuntary shift of spatial attention, referred to as attentional capture. According to the contingent involuntary orienting account (Folk, Remington, & Johnston, 1992), capture is conditioned by top-down factors that set attention to respond involuntarily to stimulus properties relevant to one's behavioral goals. Evidence for this comes from spatial cuing studies showing that a spatial cuing effect is observed only when cues have goal-relevant properties. Here, we examine alternative, decision-level explanations of the spatial cuing effect that attribute evidence of capture to postpresentation delays in the voluntary allocation of attention, rather than to on-line involuntary shifts in direct response to the cue. In three spatial cuing experiments, delayed-allocation accounts were tested by examining whether items at the cued location were preferentially processed. The experiments provide evidence that costs and benefits in spatial cuing experiments do reflect the on-line capture of attention. The implications of these results for models of attentional control are discussed.

Attention

Moving attention - Evidence for time-invariant shifts of visual selective attention

Two experiments measured the time to shift spatial selective attention across the visual field to targets 2 or 10 deg from central fixation. A central arrow cued the most likely target location. The direction of attention was inferred from reaction times to expected, unexpected, and neutral locations. The development of a spatial attentional set with time was examined by presenting target probes at varying times after the cue. There were no effects of distance on the time course of the attentional set. Reaction times for far locations were slower than for near, but the effects of attention were evident by 150 msec in both cases. Spatial attention does not shift with a characteristic, fixed velocity. Rather, velocity is proportional to distance, resulting in a movement time that is invariant over the distances tested.

Remington, R.

Method of encouraging attention by correlating video game difficulty with attention level

A method of encouraging attention in persons such as those suffering from Attention Deficit Disorder is provided by correlating the level of difficulty of a video game with the level of attention in a subject. A conventional video game comprises a video display which depicts objects for interaction with a player and a difficulty adjuster which increases the difficulty level, e.g., action speed and/or evasiveness of the depicted object, in a predetermined manner. The electrical activity of the brain is measured at selected sites to determine levels of awareness, e.g., activity in the beta, theta, and alpha states. A value is generated based on this measured electrical signal which is indicative of the level of awareness. The difficulty level of the game is increased as the awareness level value decreases and is decreased as this awareness level value increases.

Pope, Alan T.

JVP Flash Attention (jvp_flash_attention) v0.0.4

A Flash Attention Triton kernel with support for second-order derivatives, such as Jacobian-Vector Products (JVPs) and Hessian-Vector Products (HVPs).

Morehead, Alex [Lawrence Berkeley National Laborat

Graph-Based Attention Mechanisms for Solving the AC Optimal Power Flow Problem in Electrical Power Networks

With the increasing complexity and data availability in modern power systems, learning-based approaches to AC Optimal Power Flow (AC OPF) have garnered significant attention. In particular, the structure of smart grids lends itself naturally to graph-based representations, where Graph Neural Networks (GNNs) can capture spatial and relational dependencies. This paper investigates attention-based GNN architectures tailored to heterogeneous graph representations of electric grids. We evaluate two major paradigms: relational attention, which distinguishes between edge types during message passing, and meta-path attention, which captures high-level semantics through multi-hop, typed paths. Using a large corpus of public AC OPF scenarios, we benchmark representative models of each type of attention. Our results demonstrate the benefits of heterogeneous attention-based models in accurately capturing grid dynamics; heterogeneous attention models achieve superior performance in both standard and perturbed settings. The findings highlight the importance of semantic-aware architectures for improving prediction robustness and interpretability in power system applications.

Trigui, Ali [Qubit Engineering Inc.]

Extended attention span training system

Attention Deficit Disorder (ADD) is a behavioral disorder characterized by the inability to sustain attention long enough to perform activities such as schoolwork or organized play. Treatments for this disorder include medication and brainwave biofeedback training. Brainwave biofeedback training systems feed back information to the trainee showing him how well he is producing the brainwave pattern that indicates attention. The Extended Attention Span Training (EAST) system takes the concept a step further by making a video game more difficult as the player's brainwaves indicate that attention is waning. The trainee can succeed at the game only by maintaining an adequate level of attention. The EAST system is a modification of a biocybernetic system that is currently being used to assess the extent to which automated flight management systems maintain pilot engagement. This biocybernetic system is a product of a program aimed at developing methods to evaluate automated flight deck designs for compatibility with human capabilities. The EAST technology can make a contribution in the fields of medical neuropsychology and neurology, where the emphasis is on cautious, conservative treatment of youngsters with attention disorders.

Pope, Alan T.

Involuntary attentional capture by abrupt onsets

Five experiments were carried out to examine the extent to which brief abrupt-onset visual stimuli involuntarily capture spatial attention. A fundumantal limitation on the conscious control of spatial attention is demonstrated. Data obtained reveal conditions under which the control of spatial attention is completely involuntary: attention is captured by an irrelevant event despite subjects' intentions to ignore the event. The paradigm used provided strong incentives to ignore the distracting abrupt onset, but these were insufficient to prevent capture. Results suggest that voluntary control of attention is limited to focusing attention in advance on locations, objects, or properties of interest. Under appropriate conditions, spatial attention can be involantarily drawn to abrupt-onset events despite the intention of subjects' to ignore them.

Remington, Roger W.

Enhancing Unknown Waveform Detection by Learning Intra and Inter-domain Dependencies with Advanced Attention Fusion Mechanisms

Detection of unknown waveforms in mission-critical communications is a crucial area of interest for the Department of Energy (DoE). Traditional methods and recent deep learning-based approaches often assume that the training set includes all possible classes, which is impractical for detecting new waveforms. This limitation gives rise to the problem of open-set recognition (OSR), which involves correctly identifying known classes while detecting and rejecting unknown or unseen classes. To address this limitation, we propose a novel dual-domain complex-valued neural architecture that jointly processes time-domain and frequency-domain signal representations using transformer mechanisms. A transformer model is a deep learning architecture that uses self-attention mechanisms to process and learn relationships in sequential data. Our model employs a cosine similarity loss to extract domain-specific features and incorporates a transformer architecture in the latent space to weigh the importance of different features from the time and frequency domains. The transformer layer includes stacked self-attention and cross-attention modules to learn intra-domain and inter-domain dependencies, creating a more holistic signal representation. An attention-based fusion module intelligently combines the time and frequency-domain features using multi-head attention, enabling the network to learn the optimal feature for each domain in each input signal. Quantitative results demonstrate the impact of these architectural choices on overall performance, showing significant improvement after incorporating self and cross-attention modules and using complex attention fusion over simple weighted fusion. Our ongoing work will focus on addressing the limitations of threshold-based OSR methods by developing a novel generative framework that integrates a conditional diffusion probabilistic model (DPM). DPM is a generative framework that learns to synthesize complex data by reversing a gradual noising process using a neural network trained to denoise step-by-step. Our goal is to leverage the inherent strengths of DPMs for identifying unknown signals more robustly. One primary advantage of using a DPM is its ability to provide a more reliable anomaly score based on the model's reconstruction error, rather than relying solely on classifier confidence. Additionally, the iterative denoising process of DPMs makes this approach naturally resilient to low Signal-to-Noise Ratio (SNR) conditions, where traditional methods often fail. By implementing this generative framework, we aim to enhance the model's capability to accurately detect unknown waveforms and maintain performance in challenging environments.

99 - GENERAL AND MISCELLANEOUS

Deformable phrase level attention: A flexible approach for improving AI based medical coding

Objective: Improving the AI-driven automated medical encoding of clinical text plays a vital role in gathering information on the occurrence of diseases to improve population-level health. This work presents a novel attention mechanism designed to enhance text classification models and ensure appropriate classification of medical concepts in unstructured electronic health records. Materials and Methods: We developed a deformable, phrase-level attention mechanism to identify important lexical word-level and contextual phrase-level information from clinical text documents. We evaluated conventional and transformer-based deep learning models that we extended with our attention mechanism on the extraction of critical cancer information (e.g., site, subsite, laterality, histology, behavior) from 629,908 electronic pathology reports and on the automated medical encoding of 52,722 hospital discharge summaries. Results: Transformer-based models with the deformable, phrase-level attention mechanism achieved the best performance on the extraction of critical cancer information from pathology reports. Conventional- and transformer-based models show similar or better performance than their baseline counterparts on the automated medical encoding of clinical documents. Discussion: The addition of phrase-level information allowed models extended with our proposed method to outperform standard word-level attention. Our method showed favorable properties for the real-world application in terms of model robustness and phenotyping. These results indicate that our method is promising for automated data harmonization for common data models. Conclusion: This work proposes a novel deformable, phrase-level attention mechanism that enhances text classification models in the extraction of medical concepts from clinical text documents. We demonstrate strong performances on two clinical text datasets and showcase real-world deployability of our method.

Automated medical encoding

RingX: Scalable Parallel Attention for Long-Context Learning on HPC

The attention mechanism has become foundational for remarkable AI breakthroughs since the introduction of the Transformer, driving the demand for increasingly longer context to power frontier models such as large-scale reasoning language models and high-resolution image/video generators. However, its quadratic computational and memory complexities present substantial challenges. Current state-of-the-art parallel attention methods, such as ring attention, are widely adopted for long-context training but utilize a point-to-point communication strategy that fails to fully exploit the capabilities of modern HPC network architectures. In this work, we propose ringX, a scalable family of parallel attention methods optimized explicitly for HPC systems. By enhancing workload partitioning, refining communication patterns, and improving load balancing, ringX achieves up to 3.4 × speedup compared to conventional ring attention on the Frontier supercomputer. Optimized for both bi-directional and causal attention mechanisms, ringX demonstrates its effectiveness through training benchmarks of a Vision Transformer (ViT) on a climate dataset and a Generative Pre-Trained Transformer (GPT) model, Llama3 8B. Our method attains an end-to-end training speedup of approximately 1.5 × in both scenarios. To our knowledge, the achieved 38% model FLOPs utilization (MFU) for training Llama3 8B with a 1M-token sequence length on 4,096 GPUs represents one of the highest training efficiencies reported for long-context learning on HPC systems. Our code implementation is available at https://github.com/jqyin/ringX-attention.

Yin, Junqi [ORNL] (ORCID:0000000338435520)

When does global attention help: a unified empirical study on atomistic graph learning

Graph neural networks (GNNs) are widely used as surrogates for costly experiments and first-principles simulations to study the behavior of compounds at atomistic scale, and their architectural complexity is constantly increasing to enable the modeling of complex physics. While most recent GNNs combine more traditional message passing neural networks (MPNNs) layers to model short-range interactions with more advanced graph transformers (GTs) with global attention mechanisms to model long-range interactions, it is still unclear when global attention mechanisms provide real benefits over well-tuned MPNN layers due to inconsistent implementations, features, or hyperparameter tuning. We introduce the first unified, reproducible benchmarking framework–built on HydraGNN–that enables seamless switching among four controlled model classes: MPNN, MPNN with chemistry/topology encoders, GPS-style hybrids of MPNN with global attention, and fully fused localglobal models with encoders. Using seven diverse open-source datasets for benchmarking across regression and classification tasks, we systematically isolate the contributions of message passing, global attention, and encoder-based feature augmentation. Our study shows that encoder-augmented MPNNs form a robust baseline, while fused localglobal models yield the clearest benefits for properties governed by long-range interaction effects. We further quantify the accuracycompute trade-offs of attention, reporting its overhead in memory. Together, these results establish the first controlled evaluation of global attention in atomistic graph learning and provide a reproducible testbed for future model development.

Equivariant graph neural networks

The division of attention and the human auditory evoked potential

The sensitivity of the scalp-recorded, auditory evoked potential to selective attention was examined while subjects responded to stimuli presented to one ear (focused attention) and to both ears (divided attention). The amplitude of the N1 component was found to be largest to stimuli in the ear upon which attention was to be focused, smallest to stimuli in the ear to be ignored, and intermediate to stimuli in both ears when attention was divided. The results are interpreted as supporting a capacity model of attention.

Hink, R. F.

Characteristics of covert and overt visual orienting: Evidence from attentional and oculomotor capture

Five visual search experiments found oculomotor and attentional capture consistent with predictions of contingent orienting, contrary to claims that oculomotor capture is purely stimulus driven. Separate saccade and attend-only conditions contained a color target appearing either singly, with an onset or color distractor, or both. In singleton mode, onsets produced oculomotor and attentional capture. In feature mode, capture was absent or greatly reduced, providing evidence for top-down modulation of both types of capture. Although attentional capture by color abstractors was present throughout, oculomotor capture by color occurred only when accompanied by transient change, providing evidence for a dissociation between oculomotor and attentional capture. Oculomotor and attentional capture appear to be mediated by top-down attentional control settings, but transient change may be necessary for oculomotor capture. ((c) 2003 APA, all rights reserved).

Clinical Trial

Attentional Modulation of Eye Torsion Responses

Eye movements generally have both reflexive and voluntary aspects, but torsional eye movements are usually thought of as a reflexive response to image rotation around the line of sight (torsional OKN) or to head roll (torsional VOR). In this study we asked whether torsional responses could be modulated by attention in a case where two stimuli rotated independently, and whether attention would influence the latency of responses. The display consisted of rear-projected radial "pinwheel" gratings, with an inner annulus segment extending from the center to 22 degrees eccentricity, and an outer annulus segment extending from 22 degrees out to 45 degrees eccentricity. The two segments rotated around the center in independent random walks, stepping randomly 4 degrees clockwise or counterclockwise at 60 Hz. Subjects were asked to attend to one or the other while keeping fixation steady at the center of the display. To encourage attention on one or the other segment of the display, subjects were asked to move a joystick in synchrony with the back and forth rotations of one part of the image while ignoring the other. Eye torsion was recorded with the scleral search coil technique, sampled at 500 Hz. All four subjects showed roughly 50% stronger torsion responses to the attended compared to unattended segments. Latency varied from 100 to 150 msec across subjects and was unchanged by attention. These findings suggest that attention can influence eye movement responses that are not typically under voluntary control.

attention

Attentional Modulation of Eye Torsion Responses

Eye movements generally have both reflexive and voluntary aspects, but torsional eye movements are usually thought of as a reflexive response to image rotation around the line of sight (torsional OKN) or to head roll (torsional VOR). In this study we asked whether torsional responses could be modulated by attention in a case where two stimuli rotated independently, and whether attention would influence the latency of responses. The display consisted of rear-projected radial pinwheel gratings, with an inner annulus segment extending from the center to 22 degrees eccentricity, and an outer annulus segment extending from 22 degrees out to 45 degrees eccentricity. The two segments rotated around the center in independent random walks, stepping randomly 4 degrees clockwise or counterclockwise at 60 Hz. Subjects were asked to attend to one or the other while keeping fixation steady at the center of the display. To encourage attention on one or the other segment of the display, subjects were asked to move a joystick in synchrony with the back and forth rotations of one part of the image while ignoring the other. Eye torsion was recorded with the scleral search coil technique, sampled at 500 Hz. All four subjects showed roughly 50 stronger torsion responses to the attended compared to unattended segments. Latency varied from 100 to 150 msec across subjects and was unchanged by attention. These findings suggest that attention can influence eye movement responses that are not typically under voluntary control.

ocular torsion

Attention-based explainability for structure–property relationships

Machine learning methods are emerging as a universal paradigm for constructing correlative structure–property relationships in materials science based on multimodal characterization. However, this necessitates the development of methods for the physical interpretability of the resulting correlative models. Here, we demonstrate the potential of attention-based neural networks for revealing structure–property relationships and the underlying physical mechanisms, using the ferroelectric properties of PbTiO3 thin films as a case study. Through the analysis of attention scores, we disentangle the influence of distinct domain patterns on the polarization switching process. The attention-based Transformer model is explored both as a direct interpretability tool and as a surrogate for explaining representations learned via unsupervised machine learning, enabling the identification of physically grounded correlations. We compare attention-derived interpretability scores with classical SHapley Additive exPlanations analysis and show that, in contrast to applications in natural language processing, attention mechanisms in materials science exhibit high efficiency in highlighting meaningful structural features.

Slautin, Boris [Independent Researcher]