Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “malware”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Secure boot, trusted boot and remote attestation for ARM TrustZone-based IoT Nodes

With the extensive application of IoT techniques, IoT devices have become ubiquitous in daily lives. Meanwhile, attacks against IoT devices have emerged to compromise IoT devices by tampering with system pre-installed programs or injecting new malware. To mitigate these attacks, integrity enforcement of IoT systems has been proposed. The integrity of an IoT device system includes load-time integrity and runtime integrity. In this paper, we design an IoT system based on ARM TrustZone to enforce the system integrity. First, we establish the root of trust and propose a hybrid booting approach consisting of both secure boot and trusted boot to enforce the system load-time integrity. Second, we investigate a paging-based process integrity measurement method to measure the NW processes and conduct remote attestation based on the measurement results ensuring the NW runtime process integrity. We implement an IoT prototype system on a NXP i.MX6Q SABRE SD development board to assess its feasibility. Finally, real-world experiment results demonstrate that our prototype introduces negligible performance overhead to the original system.

97 MATHEMATICS AND COMPUTING↗

SeqMask: Behavior Extraction Over Cyber Threat Intelligence Via Multi-Instance Learning

Abstract Identification and extraction of Tactics, Techniques and Procedures (TTPs) for Cyber Threat Intelligence (CTI) restore the full picture of cyber attacks and guide the analysts to assess the system risk. Existing frameworks can hardly provide uniform and complete processing mechanisms for TTPs information extraction without adequate knowledge background. A multi-instance learning approach named SeqMask is proposed in this paper as a solution. SeqMask extracts behavior keywords from CTI evaluated by the semantic impact, and predicts TTPs labels by conditional probabilities. Still, the framework has two mechanisms to determine the validity of keywords. One using expert experience verification. The other verifies the distortion of the classification effect by blocking existing keywords. In the experiments, SeqMask reached 86.07% and 73.99% in F1 scores for TTPs classifications. For the top 20% of keywords, the expert approval rating is 92.20%, where the average repetition of keywords whose scores between 100% and 90% is 60.02%. Particularly, when the top 65% of the keywords were blocked, the F1 decreased to about 50%; when removing the top 50%, the F1 was under 31%. Further, we also validate the possibility of extracting TTPs from full-size CTI and malware whose F1 are improved by 2.16% and 0.81%.

Ge, Wenhan↗

AI-based Cyber Event OSINT via Twitter Data

Open-Source Intelligence (OSINT) is largely regarded as a necessary component for cybersecurity intelligence gathering to secure network systems. With the advancement of artificial intelligence (AI) and increasing usage of social media, like Twitter, we have a unique opportunity to obtain and aggregate information from social media. In this study, we propose an AI-based scheme capable of automatically pulling information from Twitter, filtering out security-irrelevant tweets, performing natural language analysis to correlate the tweets about each cybersecurity event (e.g., a malware campaign), and validating the information. This scheme has many applications, such as providing a means for security operators to gain insight into ongoing events and helping them prioritize vulnerabilities to deal with. To give examples of the possible uses, we present three case studies demonstrating the event discovery and investigation processes.

Dale, Dakota↗

Accelerating Random Forest Classification on GPU and FPGA

Random Forests (RFs) are a commonly used machine learning method for classification and regression tasks spanning a variety of application domains, including bioinformatics, business analytics, and software optimization. While prior work has focused primarily on improving performance of the training of RFs, many applications, such as malware identification, cancer prediction, and banking fraud detection, require fast RF classification. In this work, we accelerate RF classification on GPU and FPGA. In order to provide efficient support for large datasets, we propose a hierarchical memory layout suitable to the GPU/FPGA memory hierarchy. We design three RF classification code variants based on that layout, and we investigate GPU- and FPGA-specific considerations for these kernels. Our experimental evaluation, performed on an Nvidia Xp GPU and on a Xilinx Alveo U250 FPGA accelerator card using publicly available datasets on the scale of millions of samples and tens of features, covers various aspects. First, we evaluate the performance benefits of our hierarchical data structure over the standard compressed sparse row (CSR) format. Second, we compare our GPU implementation with cuML, a machine learning library targeting Nvidia GPUs. Third, we explore the performance/accuracy tradeoff resulting from the use of different tree depths in the RF. Finally, we perform a comparative performance analysis of our GPU and FPGA implementations. Our evaluation shows that for high accuracy targets, our GPU implementation yields 5-9x speedup over CSR, and up to a 2x speedup over cuML.

FPGA, Xilinx FPGA, GPU, Random Forest classificati↗

Binary Driller

Binary Driller (BD) is a visualization tool which uses the data produced from the Troglodyte tool developed on the Deep Learning Malware project. Binary Driller performs function matching using the provided function embeddings (function representations), then displays the matches for each function in a layout that mimics the location of each function within the binary.

Cutsha, Michael↗

OGhidra

This is an AI driven Binary Analysis tool. It uses locally hosted Agentic AI's to automatically reverse engineer binaries to find malware and vulnerabilities.

Wang, Enoch↗

Science and Engineering of Cybersecurity by Uncertainty quantification and Rigorous Experimentation (SECURE) (Final Report)

This report summarizes the activities performed as part of the Science and Engineering of Cybersecurity by Uncertainty quantification and Rigorous Experimentation (SECURE) Grand Challenge LDRD project. We provide an overview of the research done in this project, including work on cyber emulation, uncertainty quantification, and optimization. We present examples of integrated analyses performed on two case studies: a network scanning/detection study and a malware command and control study. We highlight the importance of experimental workflows and list references of papers and presentations developed under this project. We outline lessons learned and suggestions for future work.

97 MATHEMATICS AND COMPUTING↗

TF9 Dataset Analysis

Incident Overview: In the time between November 2, 2019 and November 11, 2019, WheelByte was plagued by breaches in security. These insecurities led to breaches in customer data, company data, and even the death of an employee, Matthew Swift. They have launched an investigation into the company’s computer systems in hopes to find the root cause. We have been provided with the following artifacts from WheelByte: memory images, disk images, network packet captures, and emails. We have found multiple cyber-system attacks against WheelByte. Our investigation lasted from July 13th - August 3rd, 2023. WheelByte allowed us to look at any and every file, and there were no restrictions on what we could or could not use in our investigation. By the end of our investigation, we have been able to deduce who is behind the attack, what they have done, and why they did it. A company that is closely related to WheelByte is called Slyde. Slyde sells electric scooters and it is known that the Chief Executive Officer (CEO) of Slyde, Kimberly Holmes, sees WheelByte as a threat to business, as Wheelbyte sells electric skateboards. We have been able to deduce that Slyde is likely behind many of the malicious attacks. We have seen exfiltration addresses to Slyde domains, along with other Slyde information within their malware. We can see lots of traffic to and from Slyde Internet Protocol (IP) addresses. This may be an attempt to cripple WheelByte’s productivity to remove Slyde’s competitor from the market.

97 MATHEMATICS AND COMPUTING↗

Energy Delivery Systems with Verifiable Trustworthiness (Final Report)

Energy Delivery Systems (EDS) must be verified to be free from intrusive and malicious software. One way of verifying this software is to perform device scans to detect malicious code. Because it is possible to have “fileless” malware that exists only in device (volatile) memory, offline scanning and even many forms of online scanning is insufficient for detection. This project (“Verify”) addresses this need by performing direct sampling of memory during device operation to detect unexpected or modified software while not interfering with device operation. The Verify project provides a proof-of-concept of detection by random sampling combined with remote software- and timing-based attestation methods for robust detection of in-memory threats. An external review of Verify was performed by our partner, General Electric (GE), and a summary of their findings is provided.

97 MATHEMATICS AND COMPUTING↗

FORESTR: Finding, Organizing, Representing, Explaining, Summarizing, and Thinning Random forests

Random forests have become popular models used for data driven predictions. As a result, random forests are currently used or being considered for high-consequence mission applications in national security, such as the prediction of yield from optical signals and malware detection. While random forests may provide accurate predictions, the complexity of the algorithm causes a lack of interpretability. Random forests are an ensemble of regression or decision trees. Individual regression and decision trees are interpretable, but ensembles are inherently difficult to interpret due to the compilation of many models. We aim to increase the interpretability of random forests by finding patterns in the ensemble of trees that can be used to “thin” (or remove) trees. As a starting point, in this report, we develop a new distance metric for quantifying the similarity between trees based on their topologies (i.e., shapes). We base the metric on a novel distance metric for graphs that is a proper mathematical distance, is invariant to transformations, has registration between graphs, and computes topological evolutions between graphs. We use the tree distance metric to compute tree statistics such as a “mean tree” and to identify clusters of trees. We apply the developed methodology to a toy dataset and a mission relevant product inspection dataset to demonstrate how the metric can provide insight into random forests. Furthermore, we discuss the limitations of the approach and ideas for future research into how the metric could be used as a thinning tool to develop less complex models.

97 MATHEMATICS AND COMPUTING↗

Countering Weapons of Mass Destruction Office (CWMD) Data Categorization Study: Chemical, Biological, Radiological, and Nuclear (CBRN) Detection Device Data

Pacific Northwest National Laboratory (PNNL) seeks to address critical questions related to chemical, biological, radiological, and nuclear (CBRN) detection devices. This research aims to enhance the security and understanding of these devices by investigating various aspects of their identification, communication, and functionality. The primary focus is on network security, malware detection, device identification, and intelligence gathering. CBRN data can be categorized in various ways depending on the purpose of CBRN detection devices and the specific context of the applications for analysis. Criteria that can be used to assist in this effort include but are not limited to data type, data protocol, source/destination, application, time, security, and content. This study will inform additional paths for data classification, data profiling, data mapping, and data modeling. This will help the Countering Weapons of Mass Destruction Office (CWMD) better understand their data and make informed decisions based on the insights gained from this study and their application. The CBRN Data Categorization study will include the identification of 5–10 different CBRN detection devices with unique characteristics for assessing and analyzing the data that is being produced by and transmitted from these devices.

45 MILITARY TECHNOLOGY, WEAPONRY, AND NATIONAL DEF↗

Resilience Through Data-Driven, Intelligent Designed Control: A Formal Methods Approach

The PNNL and GTRI team developed a strategy to integrate temporal logic rule specification for detection of cyber-intrusion in the source code and control algorithms of CPS using advanced cyber-data. The GTRI team utilized its capabilities in rule synthesis and temporal logic specifications for software assurance and verification to detect and predict impact of cyber-intrusions and malware in the computational and control algorithms of cyber-physical systems. The team also developed a testing and verification approach that could be used to validate the suggested approach against a realistic use-case CPS showcasing improvements in system impact prediction performance. Temporal logic offers a compact expression of events in absolute and relative time and has a formalized translation to state machines. As such, temporal logic rules can feasibly be synthesized to any system as a rule engine, with the process being formally verified to be correct. The goal here is to utilize temporal logic rules to detect cyber-attacks and manipulations in the computational algorithms and provide real-time software assurance and verification guarantees.

97 MATHEMATICS AND COMPUTING↗

Precursor Analysis Report: Doppelpaymer Ransomware Attack on Petroleos Mexicanos (PEMEX) 2019

The DoppelPaymer Ransomware Attack on Petroleos Mexicanos (PEMEX) 2019 Precursor Analysis Report leverages publicly available information about the PEMEX cyber attack and catalogs anomalous observables for each technique employed in the attack. This analysis is based upon the methodology of the Cybersecurity for the Operational Technology Environment (CyOTE) program. The 2019 DoppelPaymer ransomware attack on PEMEX, Mexico’s nationalized petroleum corporation, highlights a unique threat that ransomware and cybercriminal extortion poses to Operational Technology (OT) environments in critical infrastructure. The incident began with an employee downloading commodity malware that allowed adversaries to gain initial access to PEMEX’s enterprise environment. After conducting privilege escalation, tool ingress, and data exfiltration, the adversaries deployed DoppelPaymer ransomware throughout the PEMEX enterprise environment, resulting in the company having to take dozens of systems offline for at least several days. Although PEMEX stated that their operations were not affected, the data exfiltrated from PEMEX was made available for download on DoppelPaymer’s leak site, as well as on other illicit criminal forums. This stolen data included not only company information, but also sensitive OT-specific configuration data. This incident showcases how cybercriminal exfiltration and posting of sensitive OT architecture documentation can pose security concerns for the targeted organization for years due to the long lifespan of OT assets and architectures. Researchers and analysts identified 18 unique techniques utilized during the attack with a total of 190 observables using MITRE ATT&CK® for Industrial Control Systems. The CyOTE program assesses observables accompanying techniques used prior to the triggering event to identify opportunities to detect malicious activity. If observables accompanying the attack techniques are perceived and investigated prior to the triggering event, earlier comprehension of malicious activity can take place. Fifteen of the identified techniques used during the DoppelPaymer ransomware attack were precursors to the triggering event. Analysis identified 163 observables associated with these precursor techniques, 34 of which were assessed to have an increased likelihood of being perceived in the 60 days preceding the triggering event. The response and comprehension time could have been reduced if the observables had been identified earlier. The information gathered in this report contributes to a library of observables tied to a repository of artifacts, data sources, and technique detection references for practitioners and developers to support the comprehension of indicators of attack. Asset owners and operators can use these products if they experience similar observables or to prepare for comparable scenarios.

45 MILITARY TECHNOLOGY, WEAPONRY, AND NATIONAL DEF↗

Automated Threat Information Generation

Manual malware analysis is time consuming and is not scalable. This poster describes tools that INL has developed that automate and codify threat information.

99 GENERAL AND MISCELLANEOUS↗

Trends in Cybersecurity Threats to Clean Energy

As deployments of clean energy generation and storage assets continue to grow, the increased attack surface creates a greater risk for cyber threats, but is clean energy truly a target for cyber adversaries? This poster will present research on the trends in cyber incidents that have affected clean energy companies and assets as well as the trends in disclosed and exploited vulnerabilities. From a series of ransomware attacks on European wind manufacturers, to vulnerabilities exploited in solar assets to turn controllers into botnets, to attacks on communication infrastructure that have resulted in extended outages of remote control and monitoring, we explore the techniques used and the impacts to the clean energy sector. Key takeaways include understanding of how OT-focused malware is becoming more flexible and more destructive, how known vulnerabilities are being exploited, the growing number of IT and OT attacks that use built in tools and functionalities. Additionally, we highlight the presumed motivations and targeted sectors for various identified cyber adversaries. Viewers will leave with an understanding of how recent headlines fit into the development of cyberattack trends and what preventions they may need to take to protect against increasingly popular tactics.

14 SOLAR ENERGY↗

Open Source Intelligence for Cybersecurity Events via Twitter Data

Open-Source Intelligence (OSINT) is largely regarded as a necessary component for cybersecurity intelligence gathering to secure network systems. With the advancement of artificial intelligence (AI) and increasing usage of social media, like Twitter, we have a unique opportunity to obtain and aggregate information from social media. In this study, we propose an AI-based scheme capable of automatically pulling information from Twitter, filtering out security-irrelevant tweets, performing natural language analysis to correlate the tweets about each cybersecurity event (e.g., a malware campaign), and validating the information. This scheme has many applications, such as providing a means for security operators to gain insight into ongoing events and helping them prioritize vulnerabilities to deal with. To give examples of the possible uses, we present three case studies demonstrating the event discovery and investigation processes. We also examine the potential of OSINT for identifying the network protocols associated with specific events, which can aid in the mitigation procedures by informing operators if the vulnerability is exploitable given their system’s network configurations.

Dale, Dakota↗

Automated Generation of Graph-based Cyber Threat Intel

With the advancement of AI technology and tools, specifically in the cybersecurity domain, both cyber defenders and threat actors are continuously adapting the use of these capabilities to expedite their operations. With this phenomenon, threat intelligence that is up to date, refreshable, and has relevant context to a specific threat becomes more and more important as it enables cybersecurity professionals to gain insight into relevant data and relationships to guide their operations. This project enables users to frequently aggregate threat intelligence from various sources, such as vendor vulnerability advisories affecting critical infrastructure, malware reports, and adversary writeups into a centralized, standardized database. The project utilizes the Structured Threat Intelligence eXpression (STIX) for a standardized, shareable threat intelligence data format and Neo4j as a graph database solution to store STIX nodes and relationships. Initial results of the project include datasets of over 8,000 nodes and 20,000 relationships extracted from over 500 data sources that have been released within the past month.

Threat Intelligence↗

General-Purpose Unsupervised Cyber Anomaly Detection via Non-Negative Tensor Factorization

Distinguishing malicious anomalous activities from unusual but benign activities is a fundamental challenge for cyber defenders. Prior studies have shown that statistical user behavior analysis yields accurate detections by learning behavior profiles from observed user activity. These unsupervised models are able to generalize to unseen types of attacks by detecting deviations from normal behavior, without knowledge of specific attack signatures. However, approaches proposed to date based on probabilistic matrix factorization are limited by the information conveyed in a two-dimensional space. Non-negative tensor factorization, on the other hand, is a powerful unsupervised machine learning method that naturally models multi-dimensional data, capturing complex and multi-faceted details of behavior profiles. Herein, our new unsupervised statistical anomaly detection methodology matches or surpasses state-of-the-art supervised learning baselines across several challenging and diverse cyber application areas, including detection of compromised user credentials, botnets, spam e-mails, and fraudulent credit card transactions.

97 MATHEMATICS AND COMPUTING↗