Engineering Papers⌕ Search

Engineering topics

Verma, Miki

Publications and source records attributed to Verma, Miki.

D2U: Data Driven User Emulation for the Enhancement of Cyber Testing, Training, and Data Set Generation

Whether testing intrusion detection systems, conducting training exercises, or creating data sets to be used by the broader cybersecurity community, realistic user behavior is a critical component of a cyber range. Existing methods either rely on network level data or replay recorded user actions to approximate real users in a network. Our work is the first to produce generative models trained on actual user data (sequences of application usage) collected on endpoints. Once trained to the user's behavioral data, these models can generate novel sequences of actions %that appear to come from the same distribution as the training data. These sequences of actions are then fed to our custom software via configuration files, which replicate those behaviors on end devices. Notably, our models are platform agnostic and could generate behavior data for any emulation software package. In this paper we present our model generation process, software architecture, and an initial evaluation of the fidelity of our models. Our software is currently deployed in a cyber range to help evaluate the efficacy of defensive cyber technologies. We suggest additional ways that the cyber community as a whole can benefit from more realistic user behavior emulation. The data used to train our model, as well as sample configuration files produced by the model, are available at [redacted].

Oesch, T↗

Time-Based CAN Intrusion Detection Benchmark

Modern vehicles are complex cyber-physical systems made of hundreds of electronic control units (ECUs) that communicate over controller area networks (CANs). This inherited complexity has expanded the CAN attack surface by the injection of malicious messages that vary their time-based characteristics. To detect these malicious messages, time-based intrusion detection systems (IDS) have been proposed. However, time-based IDS are usually trained and tested on low-fidelity datasets with unrealistic labeled attacks. This makes difficult the task of evaluating, comparing, and validating IDS. Here we detail and benchmark four time-based IDS in a dataset with real and advanced attacks. We found that methods with strong assumptions regarding the distribution of inter-arrival times have lower performance than distribution agnostic based methods. In particular, distribution agnostic based methods outperform distribution based methods at least on $55\%$ in area under the precision-recall (AUC-PR) curve. Our results expand the body of knowledge of CAN time-based IDS by providing details of these methods and reporting their results when tested on datasets with real and advanced attacks. We describe limitations, open challenges, and how lessons learnt from this research can inform the design of deployable time-based IDS in modern vehicles.

Blevins, Deborah↗