Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “malware”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Experimental Validation of a Command and Control Traffic Detection Model

Network intrusion detection systems (NIDS) are commonly used to detect malware communications, including command-and-control (C2) traffic from botnets. NIDS performance assessments have been studied for decades, but mathematical modeling has rarely been used to explore NIDS performance. This paper details a mathematical model that describes a NIDS performing packet inspection and its detection of malware's C2 traffic. Here, the paper further describes an emulation testbed and a set of cyber experiments that used the testbed to validate the model. These experiments included a commonly used NIDS (Snort) and traffic with contents from a pervasive malware (Emotet). Results are presented for two scenarios: a nominal scenario and a “stressed” scenario in which the NIDS cannot process all incoming packets. Model and experiment results match well, with model estimates mostly falling within 95 % confidence intervals on the experiment means. Model results were produced 70-3000 times faster than the experimental results. Consequently, the model's predictive capability could potentially be used to support decisions about NIDS configuration and effectiveness that require high confidence results, quantification of uncertainty, and exploration of large parameter spaces. Furthermore, the experiments provide an example for how emulation testbeds can be used to validate cyber models that include stochastic variability.

mathematical model↗

Toward the Detection of Polyglot Files

Standardized file types play a key role in the development and use of computer software. However, it is possible to confound standardized file type processing by creating a file that is valid in multiple file types. The resulting polyglot (many languages) file can confuse file type identification, allowing elements of the file to evade analysis. This is especially problematic for malware detection systems that rely on file type identification for feature extraction. Although work has been done to identify file types using more comprehensive methods than file signatures, accurate identification of polyglot files remains an open problem. Since malware detection systems routinely perform file type-specific feature extraction, polyglot files need to be filtered out prior to ingestion by these systems. Otherwise, malicious content could pass through undetected. To address the problem of polyglot detection we assembled a data set using the mitra tool. We then evaluated the performance of the most commonly used file identification tools, including file, polydet, binwalk, and TrID. Our analysis demonstrates that existing file type detection tools fail to provide reliable polyglot detection. We then evaluated the ability of a range of machine and deep learning models to detect polyglot files. The most performant models were MalConv2 and Catboost, which demonstrated the highest recall on our data set with 95.16% and 95.45%, respectively. These models outperformed existing methods and could be incorporated into a malware detector’s file processing pipeline to filter out potentially malicious polyglots before file type-dependent feature extraction takes place.

Koch, Luke↗

Toward the Detection of Polyglot Files

Standardized file types play a key role in the development and use of computer software. However, it is possible to confound standardized file type processing by creating a file that is valid in multiple file types. The resulting polyglot (many languages) file can confuse file type identification, allowing elements of the file to evade analysis. This is especially problematic for malware detection systems that rely on file type identification for feature extraction. Although work has been done to identify file types using more comprehensive methods than file signatures, accurate identification of polyglot files remains an open problem. Since malware detection systems routinely perform file type-specific feature extraction, polyglot files need to be filtered out prior to ingestion by these systems. Otherwise, malicious content could pass through undetected. To address the problem of polyglot detection we assembled a data set using the mitra tool. We then evaluated the performance of the most commonly used file identification tools, including file, polydet, binwalk, and TrID. Our analysis demonstrates that existing file type detection tools fail to provide reliable polyglot detection. We then evaluated the ability of a range of machine and deep learning models to detect polyglot files. The most performant models were MalConv2 and Catboost, which demonstrated the highest recall on our data set with 95.16% and 95.45%, respectively. These models outperformed existing methods and could be incorporated into a malware detector’s file processing pipeline to filter out potentially malicious polyglots before file type-dependent feature extraction takes place.

Koch, Lucas↗

The Evolution of Volatile Memory Forensics

The collection and analysis of volatile memory is a vibrant area of research in the cybersecurity community. The ever-evolving and growing threat landscape is trending towards fileless malware, which avoids traditional detection but can be found by examining a system’s random access memory (RAM). Additionally, volatile memory analysis offers great insight into other malicious vectors. It contains fragments of encrypted files’ contents, as well as lists of running processes, imported modules, and network connections, all of which are difficult or impossible to extract from the file system. For these compelling reasons, recent research efforts have focused on the collection of memory snapshots and methods to analyze them for the presence of malware. However, to the best of our knowledge, no current reviews or surveys exist that systematize the research on both memory acquisition and analysis. We fill that gap with this novel survey by exploring the state-of-the-art tools and techniques for volatile memory acquisition and analysis for malware identification. For memory acquisition methods, we explore the trade-offs many techniques make between snapshot quality, performance overhead, and security. For memory analysis, we examined the traditional forensic methods used, including signature-based methods, dynamic methods performed in a sandbox environment, as well as machine learning-based approaches. We summarize the currently available tools, and suggest areas for more research.

Nyholm, Hannah↗

AI-based Detection and Defense Against Cyberattacks in Distributed Energy Resources

This study will provide comprehensive artificial intelligence (AI)-based solution tools for network security, malware prevention, and sensor data anomaly detection for distributed energy resource (DER) research, development, and demonstration. DER technologies are energy systems (e.g., solar panels, wind turbines, and energy storage systems) that are often connected to the internet and thus vulnerable to cyberattacks. Cybersecurity should be of primary concern for DERs, which is why we propose an integrated multi-layer cyber-defense system for DERs. This system encompasses risk assessments, network security, malware prevention, and detection of anomalies in the sensor data. Implementation of a comprehensive risk assessment with an overview of the model architecture should be the primary step, and should include the potential impact of experiencing, at a given time, one or more cyberattacks on the system. The second step is to ensure that the network security includes firewalls, intrusion detection, and malware prevention. The third step is to provide solution tools that enable sensor data anomaly detection for DERs. By incorporating these considerations into DER research, development, and demonstration, organizations can help ensure the safety and security of their systems and protect against potential cyberattacks.

20 FOSSIL-FUELED POWER PLANTS↗

Systems and methods for binary code analysis

Human-readable (HR) code may be derived from a binary. The HR code may be configured to have statistical properties suitable for machine-learned (ML) translation. The HR code may comprise source code, intermediate code, assembly code, or the like. A machine-learned translator may be configured to translate the HR code into labels comprising semantic information pertaining to respective functions of the binary, such as a function name, role, or the like. Execution of the binary may be blocked in response to translating the HR code to a label associated with malware, such as cryptocurrency mining malware or the like. Conversely, the binary may be permitted to proceed to execution in response to determining that the translation is free from labels indicative of malware.

Anderson, Matthew W.↗

idaholab/cape2stix

This software allows for the conversion, extraction, and transformation of malware behavior data from "Malware Configuration And Payload Extraction" (CAPEv2) sandbox reports, to Structured Threat Information eXpression (STIX). This allows for further analysis to be performed, sharing of threat data, and transit to a graph database.

Cutshaw, Michael [Idaho National Laboratory (INL),↗

Precursor Analysis Report: Industroyer Targeting Ukraine Electric Power Transport Utility (Ukrenergo) 2016

The Industroyer Targeting Ukraine Electric Power Transport Utility (Ukrenergo) 2016 Precursor Analysis Report leverages publicly available information about the December 2016 cyber attack against the Ukrainian Ukrenergo electric transmission utility and catalogs anomalous observables for each technique employed in the attack. This analysis is based upon the methodology of the Cybersecurity for the Operational Technology Environment (CyOTE) program. Industroyer is a modular malware framework designed to deploy several Industrial Control System (ICS) protocol-specific attack payloads to disrupt electricity distribution. Adversaries deployed Industroyer within the target network on a Microsoft Windows endpoint capable of directly manipulating or communicating with ICS. Industroyer abuses the functionality of a targeted ICS’s legitimate control system to achieve its intended impact. Adversaries likely first gained access to Ukrenergo enterprise networks in early 2016 after a successful spearphishing campaign against organizations in the electric power sector. Adversaries then began capturing credentials beginning on 1 December 2016. This allowed access to the ICS environment at the Pivnichna electric transmission substation outside Kyiv through a device dual-homed on the Information Technology (IT) and ICS networks. Adversaries conducted discovery, targeting, and access to this device using information and previously captured credentials from compromised enterprise IT machines. Finally, the adversaries deployed and launched the Industroyer malware just before midnight on 17 December. By midnight, Ukrenergo had lost control of a targeted substation, resulting in electric power outages for over an hour in the city of Kyiv and the Kyiv region. Researchers and analysts identified 31 unique techniques (used in a sequence of 33 steps) utilized during the attack with a total of 846 observables using MITRE ATT&CK® for Industrial Control Systems. The CyOTE program assesses observables accompanying techniques used prior to the triggering event to identify opportunities to detect malicious activity. If observables accompanying the attack techniques are perceived and investigated prior to the triggering event, earlier comprehension of malicious activity can take place. Twenty-nine of the identified techniques used during the Industroyer cyber attack were precursors to the triggering event. Analysis identified 548 observables associated with these precursor techniques, 353 of which were assessed to have an increased likelihood of being perceived in the 300 days preceding the triggering event. The response and comprehension time could have been reduced if the observables had been identified earlier. The information gathered in this report contributes to a library of observables tied to a repository of artifacts, data sources, and technique detection references for practitioners and developers to support the comprehension of indicators of attack. Asset owners and operators can use these products if they experience similar observables or to prepare for comparable scenarios.

45 MILITARY TECHNOLOGY, WEAPONRY, AND NATIONAL DEF↗

Sequence-Based Anomaly Detection in Critical Infrastructure Networks

United States critical infrastructure faces new cyber threats from adversarial nation-state actors in the form of malware-free attacks. Traditional cybersecurity techniques use rules-based methods to identify indicators of compromise on networks, often missing these sophisticated attacks. Our approach leverages multiple state of the art machine learning models in a pipeline to identify abnormal network events through sequential analysis. We combine both device and packet-level information into individual events to characterize anomalous network actions. The model is trained and tested on real network traffic from the Idaho National Lab High Performance Computing (HPC) with greater than 98% precision. It is capable of flagging malicious tactics used by adversaries in malware-free attacks, severe changes to the network, and abnormal user activity by network devices.

99 - GENERAL AND MISCELLANEOUS↗

Evidence-based Graph Adversary Mapping (EGRAM) [Poster]

Cybersecurity companies such as CrowdStrike, Dragos, Microsoft and Unit 42 categorize Advanced Persistent Threats (APTs) using their own naming schemes. As a result, these APTs are mapped to different malware sources and campaigns, all from differing sources, leading to inconsistent mapping. Inconsistent mapping causes confusion and adds further obscurity around these groups, making it difficult to track and mitigate APT cyberattacks. The Evidence-based Graph Adversary Mapping (EGRAM) tool remediates the mapping challenge by collecting, updating and converting adversary data and their sources into a valid, codified STIX v2.1 bundle which is then stored in a Neo4j graph database. It utilizes graph traversal methods and centrality analysis to generate actionable information as a Structured Threat Intelligence Graph (STIG), based on user queries. EGRAM exists as Python code and a Jupyter Notebook that acts as a searchable, evidence-based, source of intelligence for APT groups’ artifacts and cyber campaigns.

24 - POWER TRANSMISSION AND DISTRIBUTION↗

SeqMask: Behavior Extraction Over Cyber Threat Intelligence Via Multi-Instance Learning

Abstract Identification and extraction of Tactics, Techniques and Procedures (TTPs) for Cyber Threat Intelligence (CTI) restore the full picture of cyber attacks and guide the analysts to assess the system risk. Existing frameworks can hardly provide uniform and complete processing mechanisms for TTPs information extraction without adequate knowledge background. A multi-instance learning approach named SeqMask is proposed in this paper as a solution. SeqMask extracts behavior keywords from CTI evaluated by the semantic impact, and predicts TTPs labels by conditional probabilities. Still, the framework has two mechanisms to determine the validity of keywords. One using expert experience verification. The other verifies the distortion of the classification effect by blocking existing keywords. In the experiments, SeqMask reached 86.07% and 73.99% in F1 scores for TTPs classifications. For the top 20% of keywords, the expert approval rating is 92.20%, where the average repetition of keywords whose scores between 100% and 90% is 60.02%. Particularly, when the top 65% of the keywords were blocked, the F1 decreased to about 50%; when removing the top 50%, the F1 was under 31%. Further, we also validate the possibility of extracting TTPs from full-size CTI and malware whose F1 are improved by 2.16% and 0.81%.

Ge, Wenhan↗

AI-based Cyber Event OSINT via Twitter Data

Open-Source Intelligence (OSINT) is largely regarded as a necessary component for cybersecurity intelligence gathering to secure network systems. With the advancement of artificial intelligence (AI) and increasing usage of social media, like Twitter, we have a unique opportunity to obtain and aggregate information from social media. In this study, we propose an AI-based scheme capable of automatically pulling information from Twitter, filtering out security-irrelevant tweets, performing natural language analysis to correlate the tweets about each cybersecurity event (e.g., a malware campaign), and validating the information. This scheme has many applications, such as providing a means for security operators to gain insight into ongoing events and helping them prioritize vulnerabilities to deal with. To give examples of the possible uses, we present three case studies demonstrating the event discovery and investigation processes.

Dale, Dakota↗

Accelerating Random Forest Classification on GPU and FPGA

Random Forests (RFs) are a commonly used machine learning method for classification and regression tasks spanning a variety of application domains, including bioinformatics, business analytics, and software optimization. While prior work has focused primarily on improving performance of the training of RFs, many applications, such as malware identification, cancer prediction, and banking fraud detection, require fast RF classification. In this work, we accelerate RF classification on GPU and FPGA. In order to provide efficient support for large datasets, we propose a hierarchical memory layout suitable to the GPU/FPGA memory hierarchy. We design three RF classification code variants based on that layout, and we investigate GPU- and FPGA-specific considerations for these kernels. Our experimental evaluation, performed on an Nvidia Xp GPU and on a Xilinx Alveo U250 FPGA accelerator card using publicly available datasets on the scale of millions of samples and tens of features, covers various aspects. First, we evaluate the performance benefits of our hierarchical data structure over the standard compressed sparse row (CSR) format. Second, we compare our GPU implementation with cuML, a machine learning library targeting Nvidia GPUs. Third, we explore the performance/accuracy tradeoff resulting from the use of different tree depths in the RF. Finally, we perform a comparative performance analysis of our GPU and FPGA implementations. Our evaluation shows that for high accuracy targets, our GPU implementation yields 5-9x speedup over CSR, and up to a 2x speedup over cuML.

FPGA, Xilinx FPGA, GPU, Random Forest classificati↗

Binary Driller

Binary Driller (BD) is a visualization tool which uses the data produced from the Troglodyte tool developed on the Deep Learning Malware project. Binary Driller performs function matching using the provided function embeddings (function representations), then displays the matches for each function in a layout that mimics the location of each function within the binary.

Cutsha, Michael↗