Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “tokenization”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10

Research Supporting Satellite Communications Technology

This report describes the second year of research effort under the grant Research Supporting Satellite Communications Technology. The research program consists of two major projects: Fault Tolerant Link Establishment and the design of an Auto-Configurable Receiver. The Fault Tolerant Link Establishment protocol is being developed to assist the designers of satellite clusters to manage the inter-satellite communications. During this second year, the basic protocol design was validated with an extensive testing program. After this testing was completed, a channel error model was added to the protocol to permit the effects of channel errors to be measured. This error generation was used to test the effects of channel errors on Heartbeat and Token message passing. The C-language source code for the protocol modules was delivered to Goddard Space Flight Center for integration with the GSFC testbed. The need for a receiver autoconfiguration capability arises when a satellite-to-ground transmission is interrupted due to an unexpected event, the satellite transponder may reset to an unknown state and begin transmitting in a new mode. During Year 2, we completed testing of these algorithms when noise-induced bit errors were introduced. We also developed and tested an algorithm for estimating the data rate, assuming an NRZ-formatted signal corrupted with additive white Gaussian noise, and we took initial steps in integrating both algorithms into the SDR test bed at GSFC.

Horan Stephen↗

CORBASec Used to Secure Distributed Aerospace Propulsion Simulations

The NASA Glenn Research Center and its industry partners are developing a Common Object Request Broker (CORBA) Security (CORBASec) test bed to secure their distributed aerospace propulsion simulations. Glenn has been working with its aerospace propulsion industry partners to deploy the Numerical Propulsion System Simulation (NPSS) object-based technology. NPSS is a program focused on reducing the cost and time in developing aerospace propulsion engines. It was developed by Glenn and is being managed by the NASA Ames Research Center as the lead center reporting directly to NASA Headquarters' Aerospace Technology Enterprise. Glenn is an active domain member of the Object Management Group: an open membership, not-for-profit consortium that produces and manages computer industry specifications (i.e., CORBA) for interoperable enterprise applications. When NPSS is deployed, it will assemble a distributed aerospace propulsion simulation scenario from proprietary analytical CORBA servers and execute them with security afforded by the CORBASec implementation. The NPSS CORBASec test bed was initially developed with the TPBroker Security Service product (Hitachi Computer Products (America), Inc., Waltham, MA) using the Object Request Broker (ORB), which is based on the TPBroker Basic Object Adaptor, and using NPSS software across different firewall products. The test bed has been migrated to the Portable Object Adaptor architecture using the Hitachi Security Service product based on the VisiBroker 4.x ORB (Borland, Scotts Valley, CA) and on the Orbix 2000 ORB (Dublin, Ireland, with U.S. headquarters in Waltham, MA). Glenn, GE Aircraft Engines, and Pratt & Whitney Aircraft are the initial industry partners contributing to the NPSS CORBASec test bed. The test bed uses Security SecurID (RSA Security Inc., Bedford, MA) two-factor token-based authentication together with Hitachi Security Service digital-certificate-based authentication to validate the various NPSS users. The test bed is expected to demonstrate NPSS CORBASec-specific policy functionality, confirm adequate performance, and validate the required Internet configuration in a distributed collaborative aerospace propulsion environment.

Blaser, Tammy M.↗

NASA Taxonomies for Searching Problem Reports and FMEAs

Many types of hazard and risk analyses are used during the life cycle of complex systems, including Failure Modes and Effects Analysis (FMEA), Hazard Analysis, Fault Tree and Event Tree Analysis, Probabilistic Risk Assessment, Reliability Analysis and analysis of Problem Reporting and Corrective Action (PRACA) databases. The success of these methods depends on the availability of input data and the analysts knowledge. Standard nomenclature can increase the reusability of hazard, risk and problem data. When nomenclature in the source texts is not standard, taxonomies with mapping words (sets of rough synonyms) can be combined with semantic search to identify items and tag them with metadata based on a rich standard nomenclature. Semantic search uses word meanings in the context of parsed phrases to find matches. The NASA taxonomies provide the word meanings. Spacecraft taxonomies and ontologies (generalization hierarchies with attributes and relationships, based on terms meanings) are being developed for types of subsystems, functions, entities, hazards and failures. The ontologies are broad and general, covering hardware, software and human systems. Semantic search of Space Station texts was used to validate and extend the taxonomies. The taxonomies have also been used to extract system connectivity (interaction) models and functions from requirements text. Now the Reconciler semantic search tool and the taxonomies are being applied to improve search in the Space Shuttle PRACA database, to discover recurring patterns of failure. Usual methods of string search and keyword search fall short because the entries are terse and have numerous shortcuts (irregular abbreviations, nonstandard acronyms, cryptic codes) and modifier words cannot be used in sentence context to refine the search. The limited and fixed FMEA categories associated with the entries do not make the fine distinctions needed in the search. The approach assigns PRACA report titles to problem classes in the taxonomy. Each ontology class includes mapping words - near-synonyms naming different manifestations of that problem class. The mapping words for Problems, Entities and Functions are converted to a canonical form plus any of a small set of modifier words (e.g. non-uniformity NOT + UNIFORM.) The report titles are parsed as sentences if possible, or treated as a flat sequence of word tokens if parsing fails. When canonical forms in the title match mapping words, the PRACA entry is associated with the corresponding Problem, Entity or Function in the ontology. The user can search for types of failures associated with types of equipment, clustering by type of problem (e.g., all bearings found with problems of being uneven: rough, irregular, gritty ). The results could also be used for tagging PRACA report entries with rich metadata. This approach could also be applied to searching and tagging failure modes, failure effects and mitigations in FMEAs. In the pilot work, parsing 52K+ truncated titles (the test cases that were available), has resulted in identification of both a type of equipment and type of problem in about 75% of the cases. The results are displayed in a manner analogous to Google search results. The effort has also led to the enrichment of the taxonomy, adding some new categories and many new mapping words. Further work would make enhancements that have been identified for improving the clustering and further reducing the false alarm rate. (In searching for recurring problems, good clustering is more important than reducing false alarms). Searching complete PRACA reports should lead to immediate improvement.

Malin, Jane T.↗

Improving Realism in Reduced Gravity Simulators

Since man was first determined to walk on the moon, simulating the lunar environment became a priority. Providing an accurate reduced gravity environment is crucial for astronaut training and hardware testing. This presentation will follow the development of reduced gravity simulators to a final comparison of environments between the currently used systems. During the Apollo program era, multiple systems were built and tested, with several NASA centers having their own unique device. These systems ranged from marionette-like suspension devices where the subject laid on his side, to pneumatically driven offloading harnesses, to parabolic flights. However, only token comparisons, if any, were made between systems. Parabolic flight allows the entire body to fall at the same rate, giving an excellent simulation of reduced gravity as far as the biomechanics and physical perceptions are concerned. While the effects are accurate, there is limited workspace, limited time, and high cost associated with these tests. With all mechanical offload systems only the parts of the body that are actively offloaded feel any reduced gravity effects. The rest of the body still feels the full effect of gravity. The Partial Gravity System (Pogo) is the current ground-based offload system used to training and testing at the NASA Johnson Space Center. The Pogo is a pneumatic type system that allows for offloaded motion in the z-axis and free movement in the x-axis, but has limited motion in the y-axis. The pneumatic system itself is limited by cylinder stroke length and response time. The Active Response Gravity Offload System (ARGOS) is a next generation groundbased offload system, currently in development, that is based on modern robotic manufacturing lines. This system is projected to provide more z-axis travel and full freedom in both the x and y-axes. Current characterization tests are underway to determine how the ground-based offloading systems perform, how they compare to parabolic flights, and which of the systems is preferable for specific uses. These tests were conducted with six degree of freedom robots and manual inputs. Initial results show a definitive difference in abilities of the two offload systems.

Cowley, Matthew↗

Exploring Discretization Error in Simulation-Based Aerodynamic Databases

This work examines the level of discretization error in simulation-based aerodynamic databases and introduces strategies for error control. Simulations are performed using a parallel, multi-level Euler solver on embedded-boundary Cartesian meshes. Discretization errors in user-selected outputs are estimated using the method of adjoint-weighted residuals and we use adaptive mesh refinement to reduce these errors to specified tolerances. Using this framework, we examine the behavior of discretization error throughout a token database computed for a NACA 0012 airfoil consisting of 120 cases. We compare the cost and accuracy of two approaches for aerodynamic database generation. In the first approach, mesh adaptation is used to compute all cases in the database to a prescribed level of accuracy. The second approach conducts all simulations using the same computational mesh without adaptation. We quantitatively assess the error landscape and computational costs in both databases. This investigation highlights sensitivities of the database under a variety of conditions. The presence of transonic shocks or the stiffness in the governing equations near the incompressible limit are shown to dramatically increase discretization error requiring additional mesh resolution to control. Results show that such pathologies lead to error levels that vary by over factor of 40 when using a fixed mesh throughout the database. Alternatively, controlling this sensitivity through mesh adaptation leads to mesh sizes which span two orders of magnitude. We propose strategies to minimize simulation cost in sensitive regions and discuss the role of error-estimation in database quality.

Aftosmis, Michael J.↗

ANTLR Tree Grammar Generator and Extensions

A computer program implements two extensions of ANTLR (Another Tool for Language Recognition), which is a set of software tools for translating source codes between different computing languages. ANTLR supports predicated- LL(k) lexer and parser grammars, a notation for annotating parser grammars to direct tree construction, and predicated tree grammars. [ LL(k) signifies left-right, leftmost derivation with k tokens of look-ahead, referring to certain characteristics of a grammar.] One of the extensions is a syntax for tree transformations. The other extension is the generation of tree grammars from annotated parser or input tree grammars. These extensions can simplify the process of generating source-to-source language translators and they make possible an approach, called "polyphase parsing," to translation between computing languages. The typical approach to translator development is to identify high-level semantic constructs such as "expressions," "declarations," and "definitions" as fundamental building blocks in the grammar specification used for language recognition. The polyphase approach is to lump ambiguous syntactic constructs during parsing and then disambiguate the alternatives in subsequent tree transformation passes. Polyphase parsing is believed to be useful for generating efficient recognizers for C++ and other languages that, like C++, have significant ambiguities.

Craymer, Loring↗

MSLICE Sequencing

MSLICE Sequencing is a graphical tool for writing sequences and integrating them into RML files, as well as for producing SCMF files for uplink. When operated in a testbed environment, it also supports uplinking these SCMF files to the testbed via Chill. This software features a free-form textural sequence editor featuring syntax coloring, automatic content assistance (including command and argument completion proposals), complete with types, value ranges, unites, and descriptions from the command dictionary that appear as they are typed. The sequence editor also has a "field mode" that allows tabbing between arguments and displays type/range/units/description for each argument as it is edited. Color-coded error and warning annotations on problematic tokens are included, as well as indications of problems that are not visible in the current scroll range. "Quick Fix" suggestions are made for resolving problems, and all the features afforded by modern source editors are also included such as copy/cut/paste, undo/redo, and a sophisticated find-and-replace system optionally using regular expressions. The software offers a full XML editor for RML files, which features syntax coloring, content assistance and problem annotations as above. There is a form-based, "detail view" that allows structured editing of command arguments and sequence parameters when preferred. The "project view" shows the user s "workspace" as a tree of "resources" (projects, folders, and files) that can subsequently be opened in editors by double-clicking. Files can be added, deleted, dragged-dropped/copied-pasted between folders or projects, and these operations are undoable and redoable. A "problems view" contains a tabular list of all problems in the current workspace. Double-clicking on any row in the table opens an editor for the appropriate sequence, scrolling to the specific line with the problem, and highlighting the problematic characters. From there, one can invoke "quick fix" as described above to resolve the issue. Once resolved, saving the file causes the problem to be removed from the problem view.

Crockett, Thomas M.↗

Licklider Transmission Protocol Implementation

This software is an implementation of the Licklider Transmission Protocol (LTP), a communications protocol intended to support the Bundle Protocol in Delay-Tolerant Network (DTN) operations. LTP is designed to provide retransmission-based reliability over links characterized by extremely long message round-trip times and/or frequent interruptions in connectivity. Communication in interplanetary space is the most prominent example of this sort of environment, and LTP is principally aimed at supporting long-haul reliable transmission over deep-space RF links. Like any reliable transport service employing ARQ (Automatic Repeat re-Quests), LTP is stateful. In order to assure the reception of a block of data it has sent, LTP must retain for possible retransmission all portions of that block which might not have been received yet. In order to do so, it must keep track of which portions of the block are known to have been received so far, and which are not, together with any additional information needed for purposes of retransmitting part, or all, of the block. Long round-trip times mean substantial delay between the transmission of a block of data and the reception of an acknowledgement from the block s destination, signaling arrival of the block. If LTP postponed transmission of additional blocks of data until it received acknowledgement of the arrival of all prior blocks, valuable opportunities to use what little deep space transmission bandwidth is available would be forever lost. For this reason, LTP is based in part on a notion of massive state retention. Any number of requested transmission conversations (sessions) may be concurrently in flight at various displacements along the link between two LTP engines, and the LTP engines must necessarily retain transmission status and retransmission resources for all of them. Moreover, if any of the data of a given block are lost en route, it will be necessary to retain the state of that transmission during an additional round trip while the lost data are retransmitted; even multiple retransmission cycles may be necessary. LTP's possible multiplicity of sessions per association makes it necessary for each segment of application data to include an additional demultiplexing token: a session ID that uniquely identifies the session in which the segment was issued and, implicitly, the block of data being conveyed by this session. This software comprises a prototype implementation developed by Johns Hopkins University APL in cooperation with JPL, together with adaptations that improve the robustness, correctness, and operability of that implementation.

Burleigh, Scott C.↗

Translating MAPGEN to ASPEN for MER

This software translates MAPGEN (Europa and APGEN) domains to ASPEN, and the resulting domain can be used to perform planning for the Mars Exploration Rover (MER). In other words, this is a conversion of two distinct planning languages (both declarative and procedural) to a third (declarative) planning language in order to solve the problem of faithful translation from mixed-domain representations into the ASPEN Modeling Language. The MAPGEN planning system is an example of a hybrid procedural/declarative system where the advantages of each are leveraged to produce an effective planner/scheduler for MER tactical planning. The adaptation of the planning system (ASPEN) was investigated, and, with some translation, much of the procedural knowledge encoding is amenable to declarative knowledge encoding. The approach was to compose translators from the core languages used for adapting MAGPEN, which consists of Europa and APGEN. Europa is a constraint- based planner/scheduler where domains are encoded using a declarative model. APGEN is also constraint-based, in that it tracks constraints on resources and states and other variables. Domains are encoded in both constraints and code snippets that execute according to a forward sweep through the plan. Europa and APGEN communicate to each other using proxy activities in APGEN that represent constraints and/or tokens in Europa. The composition of a translator from Europa to ASPEN was fairly straightforward, as ASPEN is also a declarative planning system, and the specific uses of Europa for the MER domain matched ASPEN s native encoding fairly closely. On the other hand, translating from APGEN to ASPEN was considerably more involved. On the surface, the types of activities and resources one encodes in APGEN appear to match oneto- one to the activities, state variables, and resources in ASPEN. But, when looking into the definitions of how resources are profiled and activities are expanded, one sees code snippets that access various information available during planning for the moment in time being planned to decide at the time what the appropriate profile or expansion is. APGEN is actually a forward (in time) sweeping discrete event simulator, where the model is composed of code snippets that are artfully interleaved by the engine to produce a plan/schedule. To solve this problem, representative code is simulated as a declarative series of task expansions. Predominantly, three types of procedural models were translated: loops, if statements, and code blocks. Loops and if statements were handled using controlled task expansion, and code blocks were handled using constraint networks that maintained the generation of results based on what the order of execution would be for a procedural representation. One advantage with respect to performance for MAPGEN is the use of APGEN s GUI. This GUI is written in C++ and Motif, and performs very well for large plans.

Rabideau, Gregg R.↗

Marve: Measurement Context Extraction from Text

We propose Marve, a system for extracting measurement values, units, and related words from natural language text. Marve uses conditional random fields (CRF) to identify measurement values and units, followed by a rule-based system to find related entities, descriptors and modifiers within a sentence. Sentence tokens are represented by an undirected graphical model, and rules are based on part-of-speech and word dependency patterns connecting values and units to contextual words. Marve is unique in its focus on measurement context and early experimentation demonstrates Marve’s ability to generate high-precision extractions with strong recall. We also discuss Marve’s role in justifying NASA JPL’s proposed HyspIRI mission, a hyper spectral infrared imaging satellite that will study the world’s ecosystems. In general, our work with HyspIRI demonstrates the value of semantic measurement extractions in characterizing quantitative discussion contained in large corpuses of natural language text. These extractions accelerate broad-cross cu ing literature surveys and expose researchers and scientists new algorithmic approaches and experimental nuances. They also facilitate identification of scientific opportunities enabled by HyspIRI leading to more informed scientific investment and research.

Mattmann, Chris A.↗

Building Sustainable Capacity within Under-Represented Communities Through A Deep Partnership Model

Despite varied programmatic inducements and the documented benefits of Diversity Equity & Inclusion (DE&I), participation rates of underrepresented communities within NASA’s planetary sciences activities have remained low and nearly unchanged over the last decade. The lack of success of various efforts may be attributable to factors identified in social science research. First, the proportionality of under-represented groups has been shown to be critical to achieving the positive impacts of DE&I. When under-represented groups make up less than ~30%, group dynamics such as assimilation or tokenism routinely appear, with destructive and counter-productive results. Additionally, if participation only occurs within limited segments of an organization, benefits may be limited. Proportionality must extend across all levels or layers of the organization for the benefits of DE&I to accrue. Second, the inclusion of “Third Spaces” in the planning and support of DE&I efforts is critical to the long-term success or “stickiness” of DE&I activities. Third Spaces are those beyond home (First Space) and work (Second Space) where informal, community-building activities occur. Examples include conferences, after-work activities, and weekend social engagements. Without explicitly acknowledging and including “Third Space” interactions, DE&I efforts can fall short of their long-term goals. In 2014, the Laboratory for Atmospheric and Space Physics (LASP) at the University of Colorado Boulder began a partnership with the United Arab Emirates to develop and fly a space mission to Mars. The scientific mission had an explicit goal of training scientists and engineers, and creating a sustainable ecosystem that would spark and accelerate the UAE’s space economy. To accomplish this, LASP and the UAE set up a program that matched scientists and engineers in one-on-one mentorship relationships, at all levels of the organization. Over the course of the roughly 4-year project, those partnering relationships fostered not only strong professional relationships and knowledge transfer, but also personal friendships that extended into “Third Spaces.” The UAE/LASP partnership now continues through the co-development of an asteroid mission currently in definition phase. In collaboration with NASA Ames Research Center, LASP is developing the conceptual frame work to extend this deep partnership model to domestic opportunities. By implementing a planetary science mission that employs an analogous one-on-one partnership model across all elements of the organization, this project will demonstrate how proportionality and Third Space relationships can enable the development of localized, sustainable communities that are explicitly and deeply engaged with the NASA ecosystem.

Capacity Building↗

Gender Diversity in Heliophysics

Science is conducted by people. When those people do not feel safe in their workplace, they will struggle to produce quality science. The American scientific community has traditionally been dominated by cisgender white men–cisgender meaning that their gender aligns with the one assigned to them at birth. Individuals who are not part of this dominant demographic group have historically been excluded from scientific debate. However, the demographic landscape is changing rapidly [e.g., Jones (2022)], and organizations must ensure early career scientists of all identities feel accepted so they can achieve their goals in the field. Heliophysics describes the confluence and interaction of historically delineated scientific disciplines, including plasma, solar, and space physics. The scientific architecture of our field is founded on collaboration between people with diverse interests, backgrounds, skill sets, and ways of approaching problems. It should follow that the cohort of heliophysicists is at least as diverse as our research problems. A framing often referred to as “the business case” for diversity holds that perspectives different than our own enrich the ways in which we solve problems and communicates the positive outcomes for diverse working groups Starck et al. (2021). However, this rationale is insufficient in scope and uncompassionate in motivation; the safety of marginalized individuals is just as important as the achievements of a group. From the expectations that marginalized people outperform in order to prove themselves to the tokenization of their inclusion in an otherwise normative space, the “business case” for diversity is often harmful to historically marginalized individuals Haacker et al. (2022). The primary motivation for a diverse constituency of heliophysicists ought to be equity. Only by accepting the authentic selves of our fellow heliophysicists can we create an environment in which they have the mental and emotional safety necessary to do their best work. This white paper focuses on a particular axis of identity which the authors believe lacks visibility within heliophysics: gender expansion. It begins with definitions, explains the current landscape, and suggests actions toward a better future. The authors seek to shed light on these issues so that we can work together as a community to create a more inclusive, safe, and welcoming space for people of all identities.

M. Kenny↗

SULI Oral Presentation

Furthering our understanding of the prevalence and severity of issues that customers face when charging their electric vehicles (EVs) is crucial in order to improve the charging experience across the United States. This project utilizes web-scraping, machine leaning (ML), and natural language processing (NLP) techniques to analyze and categorize user-generated reviews. Selenium was used to build a data collection tool that can scrape vast amounts of user review data from the PlugShare website. Sentiment analysis was employed on this dataset in order to filter out negative reviews for further analysis. NLP techniques such as tokenization and word embedding were then used to convert user-written comments into a numerical format that a ML model can interpret. Multiple ML approaches are currently being explored in order to identify and categorize the charging issues being talked about in each review. Ultimately, the results from the ML model will be visualized and explained in a report on customer pain points to be delivered to the ChargeX Consortium, therefore revealing specific areas for improvement in the customer charging experience.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

SULI Oral Presentation

Furthering our understanding of the prevalence and severity of issues that customers face when charging their electric vehicles (EVs) is crucial in order to improve the charging experience across the United States. This project utilizes web-scraping, machine leaning (ML), and natural language processing (NLP) techniques to analyze and categorize user-generated reviews. Selenium was used to build a data collection tool that can scrape vast amounts of user review data from the PlugShare website. Sentiment analysis was employed on this dataset in order to filter out negative reviews for further analysis. NLP techniques such as tokenization and word embedding were then used to convert user-written comments into a numerical format that a ML model can interpret. Multiple ML approaches are currently being explored in order to identify and categorize the charging issues being talked about in each review. Ultimately, the results from the ML model will be visualized and explained in a report on customer pain points to be delivered to the ChargeX Consortium, therefore revealing specific areas for improvement in the customer charging experience.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

BETTER Together

The Standard Energy Efficiency Data (SEED) and Building Efficiency Targeting Tool for Energy Retrofits (BETTER) platforms are both developed by the Department of Energy and work better together. SEED is a database to manage building characteristics and performance data from a variety of sources. BETTER provides simple energy efficiency measure analyses based on high level data about the building or portfolio of buildings. A demonstration of each platform and their integration will be provided. The inputs for BETTER are building type, floor area, location, utility data, and whether PV shall be included in the analysis. The BETTER analysis can be manually set up through the web application or data can be uploaded with a BuildingSync XML file either directly or through the API. SEED can be the source of this data and the data can be sent to BETTER through the SEED application after the BETTER API token has been entered. The benefit of utilizing SEED is that it has connections to many other sources of data such as ENERGY STAR Portfolio Manager, Audit Template, and Salesforce. Therefore, it is likely that a user of SEED will already have the required inputs for BETTER in SEED already and can create BETTER analyses across their whole portfolio in a couple mouse clicks. This is a major time savings and enables decision makers an easy path to identify buildings that should undergo more detailed audits or retrofit pathways.

ASHRAE↗

Adaptive Patching for High-resolution Image Segmentation with Transformers

Attention-based models are proliferating in the space of image analytics, including segmentation. The standard method of feeding images to transformer encoders is to divide the images into patches and then feed the patches to the model as a linear sequence of tokens. For high-resolution images, e.g. microscopic pathology images, the quadratic compute and memory cost prohibits the use of an attention-based model, if we are to use smaller patch sizes that are favorable in segmentation. The solution is to either use custom complex multi-resolution models or approximate attention schemes. We take inspiration from Adapative Mesh Refinement (AMR) methods in HPC by adaptively patching the images, as a pre-processing step, based on the image details to reduce the number of patches being fed to the model, by orders of magnitude. This method has a negligible overhead, and works seamlessly with any attention-based model, i.e. it is a pre-processing step that can be adopted by any attention-based model without friction. We demonstrate superior segmentation quality over SoTA segmentation models for realworld pathology datasets while gaining a geomean speedup of 6.9× for resolutions up to 64K2, on up to 2, 048 GPUs.

Zhang, Enzhi↗

Improving Cyber Situational Understanding

Effective cybersecurity operations require the ability to analyze large amounts of information to assess security risks and formulate defensive strategies against adversaries. This has become more complex in recent years as the sprawl and interconnectivity of devices grows through implementation of virtualization, cloud computing, and Internet of Things (IoT). The amount of data and analysis required for effective cybersecurity command and control decisions far exceeds humans’ capacity to perform manually. We characterize the analysis problem as cyber situational understanding. The research presented to improve cyber situational understanding focuses on vulnerability analysis and threat intelligence. Regarding vulnerabilities, entities must analyze and plan work for between thousands and tens of thousands of software vulnerabilities annually. Entities heavily use network firewalls to limit vulnerability exposure. As a result, some of these vulnerabilities permit exposure to adversarial exploitation, whereas others are inaccessible and therefore present negligible risk of exploitation. Distinguishing between high and low risk software vulnerabilities requires a deep understanding of the vulnerability, network firewall protection, and characteristics of the targeted device. This problem is solved by extracting network service features from vulnerability data features using both machine-learning and natural language processing. Then, the network firewall topology is parsed to determine which vulnerabilities are reachable by adversaries. Ultimately, a state-based safety analysis ascertains which vulnerabilities are unsafe. A related vulnerability analysis problem occurs in cybersecurity operations when associating an entity’s hardware and software assets to public vulnerability databases. Assets often reveal hardware and software through installation artifacts and network service identification, and entities store these artifacts in inventory databases. However, software and hardware vendors apply a standard Common Platform Enumeration (CPE) naming convention when publicly reporting vulnerabilities. Associating these two datasets often requires many hours to days of manual inspection. The proposed solution automates the mapping approach of human analysts using fuzzy matching techniques, natural language processing, and, ultimately, machine learning to present a small set of recommendations for mapping the two datasets. The result significantly reduces human analysis time and reduces the occurrence of false positives in vulnerability notifications. Finally, cyber threat intelligence (CTI) requires associating cyber observable artifacts, such as IP addresses, URIs, and file hashes, with cyber threat tactics, techniques, and procedures. Unfortunately, most CTI data is compartmentalized across multiple organizations and cannot be shared due to the legal and reputational risk with cyber threat being associated with the entity. The approach to solving this problem inovlves using a distributed ledger with anonymous token spending and authentication. This allows a consortium of semi-trusted entities to share the workload of curating CTI for a threat sharing community’s cooperative benefit.

Huff, Philip↗

Patch Hierarchical Attention Transformer for Efficient Particle Jet Tagging

Real-time jet tagging is critical for identifying short-lived particle decays in the high-throughput detectors of the Large Hadron Collider, where real-time trigger systems responsible for deciding which collision events to store impose strict latency and accuracy constraints. While transformer architectures achieve the highest jet tagging accuracy when compute is unconstrained, their quadratic self-attention cost makes inference restrictive on trigger budget. Existing efficient variants reduce the computational cost, but hinder the classification performance. To address this limitation, we introduce the Patch Hierarchical Attention Transformer (PHAT-JeT), which combines two mechanisms: a physics-inspired geometric message-passing module that encodes local detector-plane structure, and a hierarchical patch-based attention scheme that computes exact attention within small particle groups while preserving global context through lightweight patch-token communication. Within a restricted budget, PHAT-JeT achieves state-of-the-art accuracy and background rejection among all resource-constrained jet tagging models on four benchmarks (\textsc{hls4ml}, JetClass, Top Tagging, and Quark--Gluon). Our code is available at https://github.com/aaronw5/PHAT-JeT.

Wang, Aaron [Illinois U., Chicago] (ORCID:00000003↗