Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “tokenization”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

183 records · Page 11

Patch Hierarchical Attention Transformer for Efficient Particle Jet Tagging

Real-time jet tagging is critical for identifying short-lived particle decays in the high-throughput detectors of the Large Hadron Collider, where real-time trigger systems responsible for deciding which collision events to store impose strict latency and accuracy constraints. While transformer architectures achieve the highest jet tagging accuracy when compute is unconstrained, their quadratic self-attention cost makes inference restrictive on trigger budget. Existing efficient variants reduce the computational cost, but hinder the classification performance. To address this limitation, we introduce the Patch Hierarchical Attention Transformer (PHAT-JeT), which combines two mechanisms: a physics-inspired geometric message-passing module that encodes local detector-plane structure, and a hierarchical patch-based attention scheme that computes exact attention within small particle groups while preserving global context through lightweight patch-token communication. Within a restricted budget, PHAT-JeT achieves state-of-the-art accuracy and background rejection among all resource-constrained jet tagging models on four benchmarks (\textsc{hls4ml}, JetClass, Top Tagging, and Quark--Gluon). Our code is available at https://github.com/aaronw5/PHAT-JeT.

Wang, Aaron [Illinois U., Chicago] (ORCID:00000003↗

Dispersion Loss Counteracts Embedding Condensation and Improves Generalization in Small Language Models

Large language models (LLMs) achieve remarkable performance through ever-increasing parameter counts, but scaling incurs steep computational costs. To better understand LLM scaling, we study representational differences between LLMs and their smaller counterparts, with the goal of replicating the representational qualities of larger models in smaller models. We observe a geometric phenomenon which we term embedding condensation, where token embeddings collapse into a narrow cone-like subspace in some language models. Through systematic analyses across multiple Transformer families, we show that small models such as GPT2 and Qwen3-0.6B exhibit severe condensation, whereas larger models such as GPT2-x1 and Qwen3-32B are more resistant to this phenomenon. Additional observations show that embedding condensation is not reliably mitigated by knowledge distillation from larger models. To fight against it, we formulate a dispersion loss that explicitly encourages embedding dispersion during training. Experiments demonstrate that it mitigates condensation, recovers dispersion patterns seen in larger models, and yields performance gains across 10 benchmarks. We believe this work offers a principled path toward improving smaller Transformers without additional parameters.

Xiao, Xi [ORNL] (ORCID:0009000009316982)↗

CASTLE: Conflict Analysis Strategy Testing Laboratory Environment v.1.0.0

SAND2024-01743O The Conflict Analysis Strategy Testing Laboratory Environment (CASTLE) is a software framework that enables and simplifies building a novel, turn-based strategy game in which it can define its own rules, maps, pieces, and interactions. The software is for novice to experienced programmers with some knowledge of Unity3D, a tool used in game production. CASTLE includes a library of common game mechanics used for strategic wargames and traditional board games, such as cards, tokens, dice, and grid maps. It follows design principles popularized by the video game industry and uses singletons for managing portions of the code. CASTLE builds on Unity's component-based design and can respond to engine events during execution. Among the numerous user-friendly features: Build games quickly and cost-effectively Network in real-time and apply data to new games developed on the framework Host multiple participants online Connect rule- or machine learning-based agents to a CASTLE game to serve as opponents or to simulate games Collect data collection from players and in-game behaviors Create a survey to gather demographics or opinions from players Store data locally or save it to an external database through Representational State Transfer (REST) functions CASTLE, which was prototyped using Microsoft Azure, is also designed for easily distributing online games using popular cloud services. The multiplayer functionality includes an agent interface, allowing developers to construct AI players that can substitute for humans in any of the games. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Fabian, Nathan↗