Engineering Papers⌕ Search

Engineering topics

Bridges, Robert

Publications and source records attributed to Bridges, Robert.

AI ATAC 1: An Evaluation of Prominent Commercial Malware Detectors

This work presents an evaluation of six prominent commercial endpoint malware detectors, a network malware detector, and a file-conviction algorithm from a cyber technology vendor. The evaluation was administered as the first of the Artificial I ntelligence Applications t o Autonomous Cybersecurity (AI ATAC) prize challenges, funded by / completed in service of the US Navy. The experiment employed 100K files (50/50% benign/malicious) with a stratified distribution of file types, including ~1K zero-day program executables (increasing experiment size two orders of magnitude over previous work). We present an evaluation process of delivering a file to a fresh virtual machine donning the detection technology, waiting 90s to allow static detection, then executing the file and waiting another period for dynamic detection; this allows greater fidelity in the observational data than previous experiments, in particular, resource and time-to-detection statistics. To execute all 800K trials (100K files × 8 tools), a software framework is designed to choreograph the experiment into an automated, time-synced, and reproducible workflow with substantial parallelization. Software with base classes for this framework are provided. A cost-benefit model was configured to integrate the tools’ detection statistics into a comparable quantity by simulating costs of use. This provides a ranking methodology for cyber competitions and a lens for reasoning about the varied statistical results. The results provide insights on state of commercial malware detection.

Bridges, Robert↗

A Privacy-Aware Federated Learning Framework for Distributed Energy Resource Analytics in Constrained Environments

To be resilient against extreme weather events, the rural communities in Puerto Rico are leveraging distributed energy resources (DER). However, computing frameworks sup-porting the grid in critical decision-making are still largely centralized. Sensitive consumer data are transmitted over the Internet or cellular networks to a secondary or tertiary node. It guarantees better situational awareness at the cost of a wider attack surface, jeopardizing user privacy, as more DER come online. Cloud, Edge, and Fog computing all require data aggregation at some level. This paper introduces a privacy-aware federated learning framework that leverages the Fog model by pushing analytics all the way to the DER and load assets. These local models train on individual asset data and transmit only learned parameters (such as weights) over secure communications to a global decision-maker. By abstracting personally identifiable consumer data without impacting decision optimality, this framework better aligns with distributed power generation paradigm.

Sundararajan, Aditya↗

Detecting CAN Masquerade Attacks with Signal Clustering Similarity

Vehicular Controller Area Networks (CANs) are susceptible to cyber attacks of different levels of sophistication. Fabrication attacks are the easiest to administer—an adversary simply sends (extra) frames on a CAN—but also the easiest to detect because they disrupt frame frequency. To overcome time-based detection methods, adversaries must administer masquerade attacks by sending frames in lieu of (and therefore at the expected time of) benign frames but with malicious payloads. Research efforts have proven that CAN attacks, and masquerade attacks in particular, can affect vehicle functionality. Examples include causing unintended acceleration, deactivation of vehicle’s brakes, as well as steering the vehicle. We hypothesize that masquerade attacks modify the nuanced correlations of CAN signal time series and how they cluster together. Therefore, changes in cluster assignments should indicate anomalous behavior. We confirm this hypothesis by leveraging our previously developed capability for reverse engineering CAN signals (i.e., CAN-D [Controller Area Network Decoder]) and focus on advancing the state of the art for detecting masquerade attacks by analyzing time series extracted from raw CAN frames. Specifically, we demonstrate that masquerade attacks can be detected by computing time series clustering similarity using hierarchical clustering on the vehicle’s CAN signals (time series) and comparing the clustering similarity across CAN captures with and without attacks. We test our approach in a previously collected CAN dataset with masquerade attacks (i.e., the ROAD dataset) and develop a forensic tool as a proof of concept to demonstrate the potential of the proposed approach for detecting CAN masquerade attacks.

Moriano Salazar, Pablo↗

Soar Survey Scoring

ORNL conducted an experiment to test COTS (commercial off the shelf) SOAR tools. Vendors interested in entering their tool into the experiment were asked to submit information videos about their tool for phase 1 to down-select which tools should make it through to the testing phase. SOC operators were assigned to watch a subset of the videos such that each video was assigned to the same number of operators. After watching all the videos, operators completed surveys which asked them to grade basic aspects about the tool as best as they could based on the provided video, and asked them to rank the order of their interest in the vendors they were assigned to watch videos on. The problem solved was how do we aggregate the results of the surveys and each of the operator's ranked lists into an overall ranking of all the tools. In addition, it was desirable to be able to work on this code before final results were collected from operators, so code to generate fake data was developed to output in the same output as the survey results.

Bridges, Robert↗

D2U: Data Driven User Emulation for the Enhancement of Cyber Testing, Training, and Data Set Generation

Whether testing intrusion detection systems, conducting training exercises, or creating data sets to be used by the broader cybersecurity community, realistic user behavior is a critical component of a cyber range. Existing methods either rely on network level data or replay recorded user actions to approximate real users in a network. Our work is the first to produce generative models trained on actual user data (sequences of application usage) collected on endpoints. Once trained to the user's behavioral data, these models can generate novel sequences of actions %that appear to come from the same distribution as the training data. These sequences of actions are then fed to our custom software via configuration files, which replicate those behaviors on end devices. Notably, our models are platform agnostic and could generate behavior data for any emulation software package. In this paper we present our model generation process, software architecture, and an initial evaluation of the fidelity of our models. Our software is currently deployed in a cyber range to help evaluate the efficacy of defensive cyber technologies. We suggest additional ways that the cyber community as a whole can benefit from more realistic user behavior emulation. The data used to train our model, as well as sample configuration files produced by the model, are available at [redacted].

Oesch, T↗

What Clinical Trials Can Teach Us About the Development of More Resilient AI for Cybersecurity

Policy-mandated, rigorously administered scientific testing is needed to provide transparency into the efficacy of artificial intelligence-–based (AI-based) cyber defense tools for consumers and to prioritize future research and development. In this work, we propose a model that is informed by our experience, urged forward by massive-scale cyberattacks, and inspired by parallel developments in the biomedical field and the unprecedentedly fast development of new vaccines to combat global pathogens.

97 MATHEMATICS AND COMPUTING↗