Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “tokenization”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

88 records · Page 5

Extracting Material Property Measurement Data from Scientific Articles

Machine learning-based prediction of material properties is often hampered by the lack of sufficiently large training datasets. The majority of such measurement data is embedded in scientific literature and the ability to automatically extract these data is essential to support the development of reliable property prediction methods. In this work, we describe a methodology for an automatic property extraction framework using material solubility as the target property. We create an annotated dataset containing tags for solubility-related entities using a combination of regular expressions and manual tagging. We then compare five entity recognition models leveraging both token-level and span-level architectures on the task of classifying solute names, solubility values, and solubility units. Additionally, we explore a novel pretraining approach that leverages automated chemical name and quantity extraction tools to generate large datasets that do not rely on intensive manual effort. Finally, we perform an analysis to identify the causes of classification errors.

Panapitiya, Gihan U.↗

A Review of Technologies that can Provide a 'Root of Trust' for Operational Technologies

The supply chain attack pathway is being increasingly used by adversaries to bypass security controls and gain unauthorized access to sensitive networks and equipment (e.g., Critical Digital Assets). Cyber-attacks targeting supply chain generally aim to compromise the environments, products, or services of vendors and suppliers to inject, add, or substitute authentic software and hardware with malicious elements. These malicious elements are deemed to be authentic as they arise from the vendor or supplier (i.e., the supply chain). This research aims at providing a survey of technologies that have the potential to reduce exposure of sensitive networks and equipment to these attacks, thereby improving tamper resistance. The recent advances in the performance and capabilities of these technologies in recent years has increased their potential applications to reduce or mitigate exposure of the supply chain attack pathway. The focus being on providing an analysis of the benefits and disadvantages of smart cards, secure tokens, and elements to provide root of trust. This analysis provides evidence that these roots of trust can increase the technical capability of equipment and networks to authenticate changes to software and configuration thereby increasing resilience to some supply chain attacks, such as those related to logistics and ICT channels, but not development environment attacks.

21 SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLANTS↗

Simple, Secure, Internet Delivery of MOOSE-based Applications

Application packaging and distribution are the final steps for delivering software to end-users; both are frequently neglected when creating scientific software. Commercial businesses rely on electronic distribution systems that have rendered disk drives obsolete. Still, national laboratories continue to rely heavily on removable media to distribute and limit access to controlled applications. With increasing concerns of unauthorized copying of sensitive applications, a modern distribution system that utilizes cryptographically secure communication and authentication protocols has been developed. This new distribution system will secure the chain of custody for nuclear software while simultaneously simplifying access to these tools. This report summarizes four primary advancements made toward the secure distribution of Nuclear Energy Advanced Modeling and Simulation (NEAMS)-developed, Multiphysics Object Oriented Simulation Environment (MOOSE)-based applications: application installation, package distribution, automated package building, and distribution of documentation. NEAMS is currently developing more than ten separate applications based on the open-source MOOSE Framework. Distribution of these applications has primarily been accomplished by distributing source code, with end-users compiling the applications themselves. This work created a mechanism where MOOSE applications can be installed in a similar way to any other software. This allows both administrators and end-users simplified access to runnable executables. With this new installation capability, it was then possible to rethink distribution. A new, secure capability for delivering MOOSE-based applications over the internet has been created. This system requires unique cryptographic tokens for authentication, greatly securing the custody chain for software. Once granted access, installation of any NEAMS code can be accomplished with these terminal commands: "conda install ncrc" "ncrc install ncrc-bison." After these two commands (and authenticating) the BISON application will be securely down- loaded from Idaho National Laboratory (INL)’s servers, installed, and ready to use. To enable this new distribution capability to be successful, the open-source Continuous Integration, Verification, Enhancement, and Testing (CIVET) Continuous Integration (CI) capability was augmented to add Continuous Delivery (CD). CD enables the automated building and packaging of MOOSE-based applications as they are modified by development teams, ensuring that our customers can obtain up-to-date versions of the software at any time. The need for instruction on how to use these applications was addressed through modifications to the MOOSE documentation system. The MooseDocs capability, which enables robust documentation of MOOSE-based applications, has been extended to allow both for the installation of documentation and the packaging of documentation with installed applications. Together, these enhancements form the core of a new, secure distribution mechanism for nuclear simulation tools. In concert with the Nuclear Computational Resource Center (NCRC), NEAMS- developed applications will now be straightforward to obtain securely.

97 MATHEMATICS AND COMPUTING↗

Hypothesis testing via AI: Generating physically interpretable models of scientific data with machine learning (Full Technical Report)

Deep learning has demonstrated an exceptional ability to solve complex tasks (an engineering success); however, it has done so at the expense of the ability to generate new knowledge (a scientific failure). We propose an alternative framework—entitled Deep Symbolic Regression (DSR)—in which artificial neural networks (NNs) rapidly generate hypotheses about physical relationships among inputs. This framework bypasses the need to interpret an NN altogether, while still leveraging the representational power of deep learning. The resulting models are tractable mathematical expressions, which are inherently and readily human interpretable and can provide insights into underlying physical phenomena. Further, we fold this methodology into the scientific process by allowing the scientist to directly integrate a priori knowledge and beliefs to accelerate learning. We demonstrate this methodology on symbolic regression—the problem of rediscovering underlying expressions describing a dataset—and achieve state-of-the-art performance across a wide variety of symbolic regression problems. Further, we generalize our DSR framework to apply to the more general class of symbolic optimization problems, in which one seeks to optimize a sequence of symbols or “tokens” under a black-box reward function. Examples of other symbolic optimization problems include neural architecture search and computational antibody design. Our generalized tool, Deep Symbolic Optimization (DSO), has been demonstrated on the task of learning symbolic control policies for reinforcement learning environments, and has been adopted as an enabling capability for computational antibody design.

97 MATHEMATICS AND COMPUTING↗

2024 OES-Environmental 2024 State of the Science Report, Chapter 5: Stakeholder Engagement for Marine Renewable Energy

Stakeholder engagement is a critical piece of any new development project that affects public or private interests. Effective, thoughtful engagement and participatory activities early in the planning process of a project can help planners and project developers understand local concerns, adjust designs to avoid negative environmental impacts, select the best site for a project, answer questions, reduce delay, enhance opportunities and benefits, and build support for a project (Cuppen et al. 2016; Portman 2009; Wiersma & DevineWright 2014). On the other hand, cursory or inadequate engagement that is viewed as “checking the box” or tokenism is unlikely to be effective, and can result in project failures, diminished trust, strong opposition, or costly, drawn-out processes (Butcher & MacLennan 2020; Garard & Kowarsch 2017; Gill & Rand 2022; Jolivet & Heiskanen 2010; Pizzi et al. 2021; Sterling et al. 2017).

16 TIDAL AND WAVE POWER↗

IRI Technology Landscape – A survey of re-usable components and methodologies

This document describes technical implementation details on network access schemes connecting API-driven workflows to supercomputer centers. API-driven workflows are a central theme in connected computing, since they bring the terminal-mainframe' access pattern present since the 1970s up to the task of interfacing with modern web browser technologies. Both security (HTTPS/TLS/IPSec/VPNs/public key cryptography/digital signatures) and network protocol stacks (HTTP-REST APIs, tokens, gRPC, SRTP) have evolved to the point where implementing API-driven workflows is possible using stable, secure off-the-shelf software.

97 MATHEMATICS AND COMPUTING↗

Efficient Distributed Sequence Parallelism for Transformer-Based Image Segmentation

We introduce an efficient distributed sequence parallel approach for training transformer-based deep learning image segmentation models. The neural network models are comprised of a combination of a Vision Transformer encoder with a convolutional decoder to provide image segmentation mappings. The utility of the distributed sequence parallel approach is especially useful in cases where the tokenized embedding representation of image data are too large to fit into standard computing hardware memory. To demonstrate the performance and characteristics of our models trained in sequence parallel fashion compared to standard models, we evaluate our approach using a 3D MRI brain tumor segmentation dataset. We show that training with a sequence parallel approach can match standard sequential model training in terms of convergence. Furthermore, we show that our sequence parallel approach has the capability to support training of models that would not be possible on standard computing resources.

Lyngaas, Isaac↗

Optimizing Distributed Training on Frontier for Large Language Models

Large language models (LLMs) have demonstrated remarkable success as foundational models, benefiting various downstream applications through fine-tuning. Loss scaling studies have demonstrated the superior performance of larger LLMs compared to their smaller counterparts. Nevertheless, training LLMs with billions of parameters poses significant challenges and requires considerable computational resources. For example, training a one trillion parameter GPT-style model on 20 trillion tokens requires a staggering 120 million exaflops. This research explores efficient distributed training strategies to extract this computation from Frontier, the world's first exascale supercomputer. We enable and investigate various model and data parallel training techniques, such as tensor parallelism, pipeline parallelism, and sharded data parallelism, to facilitate training a trillion-parameter model on Frontier. We empirically assess these techniques and their associated parameters to determine their impact on memory footprint, communication latency, and GPU's computational efficiency. We analyze the complex interplay among these techniques and find a strategy to combine them to achieve high throughput through hyperparameter tuning. We have identified efficient strategies for training large LLMs of varying sizes through empirical analysis and hyperparameter tuning. For 22 Billion, 175 Billion, and 1 Trillion parameters, we achieved GPU throughputs of 38.38%, 36.14%, and 31.96%, respectively. For the training of the 175 Billion parameter model and the 1 Trillion parameter model, we achieved 100% weak scaling efficiency on 1024 and 3072 Mi250X GPUs, respectively. We also achieved strong scaling efficiencies of 89% and 87% for these two models. We trained these models only tens of iterations instead of training till completion.

Yin, Junqi↗

SULI Oral Presentation

Furthering our understanding of the prevalence and severity of issues that customers face when charging their electric vehicles (EVs) is crucial in order to improve the charging experience across the United States. This project utilizes web-scraping, machine leaning (ML), and natural language processing (NLP) techniques to analyze and categorize user-generated reviews. Selenium was used to build a data collection tool that can scrape vast amounts of user review data from the PlugShare website. Sentiment analysis was employed on this dataset in order to filter out negative reviews for further analysis. NLP techniques such as tokenization and word embedding were then used to convert user-written comments into a numerical format that a ML model can interpret. Multiple ML approaches are currently being explored in order to identify and categorize the charging issues being talked about in each review. Ultimately, the results from the ML model will be visualized and explained in a report on customer pain points to be delivered to the ChargeX Consortium, therefore revealing specific areas for improvement in the customer charging experience.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

SULI Oral Presentation

Furthering our understanding of the prevalence and severity of issues that customers face when charging their electric vehicles (EVs) is crucial in order to improve the charging experience across the United States. This project utilizes web-scraping, machine leaning (ML), and natural language processing (NLP) techniques to analyze and categorize user-generated reviews. Selenium was used to build a data collection tool that can scrape vast amounts of user review data from the PlugShare website. Sentiment analysis was employed on this dataset in order to filter out negative reviews for further analysis. NLP techniques such as tokenization and word embedding were then used to convert user-written comments into a numerical format that a ML model can interpret. Multiple ML approaches are currently being explored in order to identify and categorize the charging issues being talked about in each review. Ultimately, the results from the ML model will be visualized and explained in a report on customer pain points to be delivered to the ChargeX Consortium, therefore revealing specific areas for improvement in the customer charging experience.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

BETTER Together

The Standard Energy Efficiency Data (SEED) and Building Efficiency Targeting Tool for Energy Retrofits (BETTER) platforms are both developed by the Department of Energy and work better together. SEED is a database to manage building characteristics and performance data from a variety of sources. BETTER provides simple energy efficiency measure analyses based on high level data about the building or portfolio of buildings. A demonstration of each platform and their integration will be provided. The inputs for BETTER are building type, floor area, location, utility data, and whether PV shall be included in the analysis. The BETTER analysis can be manually set up through the web application or data can be uploaded with a BuildingSync XML file either directly or through the API. SEED can be the source of this data and the data can be sent to BETTER through the SEED application after the BETTER API token has been entered. The benefit of utilizing SEED is that it has connections to many other sources of data such as ENERGY STAR Portfolio Manager, Audit Template, and Salesforce. Therefore, it is likely that a user of SEED will already have the required inputs for BETTER in SEED already and can create BETTER analyses across their whole portfolio in a couple mouse clicks. This is a major time savings and enables decision makers an easy path to identify buildings that should undergo more detailed audits or retrofit pathways.

ASHRAE↗

Adaptive Patching for High-resolution Image Segmentation with Transformers

Attention-based models are proliferating in the space of image analytics, including segmentation. The standard method of feeding images to transformer encoders is to divide the images into patches and then feed the patches to the model as a linear sequence of tokens. For high-resolution images, e.g. microscopic pathology images, the quadratic compute and memory cost prohibits the use of an attention-based model, if we are to use smaller patch sizes that are favorable in segmentation. The solution is to either use custom complex multi-resolution models or approximate attention schemes. We take inspiration from Adapative Mesh Refinement (AMR) methods in HPC by adaptively patching the images, as a pre-processing step, based on the image details to reduce the number of patches being fed to the model, by orders of magnitude. This method has a negligible overhead, and works seamlessly with any attention-based model, i.e. it is a pre-processing step that can be adopted by any attention-based model without friction. We demonstrate superior segmentation quality over SoTA segmentation models for realworld pathology datasets while gaining a geomean speedup of 6.9× for resolutions up to 64K2, on up to 2, 048 GPUs.

Zhang, Enzhi↗

Improving Cyber Situational Understanding

Effective cybersecurity operations require the ability to analyze large amounts of information to assess security risks and formulate defensive strategies against adversaries. This has become more complex in recent years as the sprawl and interconnectivity of devices grows through implementation of virtualization, cloud computing, and Internet of Things (IoT). The amount of data and analysis required for effective cybersecurity command and control decisions far exceeds humans’ capacity to perform manually. We characterize the analysis problem as cyber situational understanding. The research presented to improve cyber situational understanding focuses on vulnerability analysis and threat intelligence. Regarding vulnerabilities, entities must analyze and plan work for between thousands and tens of thousands of software vulnerabilities annually. Entities heavily use network firewalls to limit vulnerability exposure. As a result, some of these vulnerabilities permit exposure to adversarial exploitation, whereas others are inaccessible and therefore present negligible risk of exploitation. Distinguishing between high and low risk software vulnerabilities requires a deep understanding of the vulnerability, network firewall protection, and characteristics of the targeted device. This problem is solved by extracting network service features from vulnerability data features using both machine-learning and natural language processing. Then, the network firewall topology is parsed to determine which vulnerabilities are reachable by adversaries. Ultimately, a state-based safety analysis ascertains which vulnerabilities are unsafe. A related vulnerability analysis problem occurs in cybersecurity operations when associating an entity’s hardware and software assets to public vulnerability databases. Assets often reveal hardware and software through installation artifacts and network service identification, and entities store these artifacts in inventory databases. However, software and hardware vendors apply a standard Common Platform Enumeration (CPE) naming convention when publicly reporting vulnerabilities. Associating these two datasets often requires many hours to days of manual inspection. The proposed solution automates the mapping approach of human analysts using fuzzy matching techniques, natural language processing, and, ultimately, machine learning to present a small set of recommendations for mapping the two datasets. The result significantly reduces human analysis time and reduces the occurrence of false positives in vulnerability notifications. Finally, cyber threat intelligence (CTI) requires associating cyber observable artifacts, such as IP addresses, URIs, and file hashes, with cyber threat tactics, techniques, and procedures. Unfortunately, most CTI data is compartmentalized across multiple organizations and cannot be shared due to the legal and reputational risk with cyber threat being associated with the entity. The approach to solving this problem inovlves using a distributed ledger with anonymous token spending and authentication. This allows a consortium of semi-trusted entities to share the workload of curating CTI for a threat sharing community’s cooperative benefit.

Huff, Philip↗

Patch Hierarchical Attention Transformer for Efficient Particle Jet Tagging

Real-time jet tagging is critical for identifying short-lived particle decays in the high-throughput detectors of the Large Hadron Collider, where real-time trigger systems responsible for deciding which collision events to store impose strict latency and accuracy constraints. While transformer architectures achieve the highest jet tagging accuracy when compute is unconstrained, their quadratic self-attention cost makes inference restrictive on trigger budget. Existing efficient variants reduce the computational cost, but hinder the classification performance. To address this limitation, we introduce the Patch Hierarchical Attention Transformer (PHAT-JeT), which combines two mechanisms: a physics-inspired geometric message-passing module that encodes local detector-plane structure, and a hierarchical patch-based attention scheme that computes exact attention within small particle groups while preserving global context through lightweight patch-token communication. Within a restricted budget, PHAT-JeT achieves state-of-the-art accuracy and background rejection among all resource-constrained jet tagging models on four benchmarks (\textsc{hls4ml}, JetClass, Top Tagging, and Quark--Gluon). Our code is available at https://github.com/aaronw5/PHAT-JeT.

Wang, Aaron [Illinois U., Chicago] (ORCID:00000003↗

Dispersion Loss Counteracts Embedding Condensation and Improves Generalization in Small Language Models

Large language models (LLMs) achieve remarkable performance through ever-increasing parameter counts, but scaling incurs steep computational costs. To better understand LLM scaling, we study representational differences between LLMs and their smaller counterparts, with the goal of replicating the representational qualities of larger models in smaller models. We observe a geometric phenomenon which we term embedding condensation, where token embeddings collapse into a narrow cone-like subspace in some language models. Through systematic analyses across multiple Transformer families, we show that small models such as GPT2 and Qwen3-0.6B exhibit severe condensation, whereas larger models such as GPT2-x1 and Qwen3-32B are more resistant to this phenomenon. Additional observations show that embedding condensation is not reliably mitigated by knowledge distillation from larger models. To fight against it, we formulate a dispersion loss that explicitly encourages embedding dispersion during training. Experiments demonstrate that it mitigates condensation, recovers dispersion patterns seen in larger models, and yields performance gains across 10 benchmarks. We believe this work offers a principled path toward improving smaller Transformers without additional parameters.

Xiao, Xi [ORNL] (ORCID:0009000009316982)↗

CASTLE: Conflict Analysis Strategy Testing Laboratory Environment v.1.0.0

SAND2024-01743O The Conflict Analysis Strategy Testing Laboratory Environment (CASTLE) is a software framework that enables and simplifies building a novel, turn-based strategy game in which it can define its own rules, maps, pieces, and interactions. The software is for novice to experienced programmers with some knowledge of Unity3D, a tool used in game production. CASTLE includes a library of common game mechanics used for strategic wargames and traditional board games, such as cards, tokens, dice, and grid maps. It follows design principles popularized by the video game industry and uses singletons for managing portions of the code. CASTLE builds on Unity's component-based design and can respond to engine events during execution. Among the numerous user-friendly features: Build games quickly and cost-effectively Network in real-time and apply data to new games developed on the framework Host multiple participants online Connect rule- or machine learning-based agents to a CASTLE game to serve as opponents or to simulate games Collect data collection from players and in-game behaviors Create a survey to gather demographics or opinions from players Store data locally or save it to an external database through Representational State Transfer (REST) functions CASTLE, which was prototyped using Microsoft Azure, is also designed for easily distributing online games using popular cloud services. The multiplayer functionality includes an agent interface, allowing developers to construct AI players that can substitute for humans in any of the games. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Fabian, Nathan↗