Engineering Papers⌕ Search

Engineering topics

Casalnuovo, Casey

Publications and source records attributed to Casalnuovo, Casey.

A Framework for Evaluating the Implementation Cost of Attacks on Large Language Models

Large Language Models (LLMs) have been increasingly proposed as a method to enhance productivity in tasks that involve language and code. However, these models are large, complex, and their capabilities are not easily understood and controlled, meaning that their adoption opens many possibilities for new cyberattacks and misuse. Numerous attacks on LLMs have been reported and summarized in literature reviews, but we found existing reviews lacking in understanding the implementation cost of the attacks - i.e., how much effort would an attacker need in terms of coding, expertise, and resources to adopt attacks presented in the literature. Therefore, we divide existing attacks on LLMs into a taxonomy, and define a cost evaluation framework to determine the cost of the attack. An attack’s cost can be 1) estimated from reading the publication about the attack or 2) determined by implementing the attack from that publication. We provide an example evaluation of a couple jailbreaking frameworks based on experiments, and then apply the more lightweight cost estimate to a representative selection of attacks across the taxonomy we define. We discuss the relative difficulty of the attacks and also highlight defenses that have attempted to mitigate these attacks and assess their effectiveness.

97 MATHEMATICS AND COMPUTING↗

Evaluation of the Self Retrieval Augmented Generation Technique on Common Security Advisory Framework Data

This small experimental report evaluates a variation of Retrieval Augmented Generation (RAG), called Self-RAG. This method uses a generative language model that incorporates retrieved facts into its generation and is explicitly trained to be able to determine whether retrieved information is enough to answer the input query, with a user-defined threshold for confidence. We performed an experiment using data from the publicly available CISA Common Security Advisory Framework (CSAF) repository (https://github.com/cisagov/CSAF) as the database of facts to be used in retrieval. Qualitative results from the experiment demonstrate that the Self-RAG method has some ability to provide reasonable answers to queries that are in the dataset and will often ignore irrelevant information when asked outside of domain questions (e.g., general facts). In settings with deliberately confusing questions (the question is within domain, but asks about a fabricated advisory), it was able to refuse 40% of the time without further adjustments to the original framework. While this performance is not sufficient for current practical use, further improvements to data formatting, disambiguating results, and leveraging threshold values could improve performance significantly. However, evaluating this will require more extensive evaluations on larger datasets and potentially better models.

97 MATHEMATICS AND COMPUTING↗

Mini Report: LLMs for Vulnerability Repair in Code

Software vulnerability repair is a notoriously difficult task that is both time consuming and labor intensive. While research into this area has a long history, the recent successes of large language models (LLMs) across many tasks have also spurred efforts to leverage LLM capabilities for automated software vulnerability repair. Currently, there are limitations in the capabilities of LLMs to fix bugs and insufficiently addressed problems in the evaluations of these studies may cause performance to not transfer when they are used in practice. Additionally, most research in the area treats finding and fixing bugs as separate concerns - how to best combine all the subtasks involved in removing vulnerabilities from code remains an open question. In this report, we summarize our findings and opinions on the current state of the art in LLM-assisted code vulnerability repair, highlighting current unresolved problems in the field as well as potential applications and future research.

97 MATHEMATICS AND COMPUTING↗

Mini Report: Jailbreaking Attacks and Defenses

Overall, jailbreaking defenses are unreliable, and no defenses proposed thus far can completely stop such attacks in any verifiable way. Even manual attacks can trivially bypass some claimed ‘defenses’, and over the course of 2023 automated attacks on prior models have been adapted to work on LLMs while other attacks draw on ideas such as fuzz testing. At best, some defenses can make jailbreaks more difficult, but with the developing landscape of attacks existing papers have not robustly evaluated how effective they are against all these methods. However, our opinion is that given the way these large generative models are trained and ‘aligned’ to stated goals of safety via fine tuning, it will be exceedingly difficult if not impossible to eliminate the possibility of jailbreaking attacks. A major barrier is that the feature space of LLMs is not sufficiently understood in a way where guarantees can be made about the outputs. Barring major changes, the expectation around jailbreaking defenses should be that they can mitigate misuse, but not verifiably prevent it. However, one defense we believe merits further investigation depends on the fact as automated attacks produce text, they need an automated way to identify a successful jailbreak - a judgment model. We have seen some attempts to repurpose models like these to defend against jailbreaks, but the evaluations are small scale and not robust.

97 MATHEMATICS AND COMPUTING↗

Mini Report: LLMs on Code Translation

The goal of this mini report is to give opinions on the current state of the art in large language model (LLM) assisted code translation, which aims to replicate the behavior of a program written in one programming language to another. The first section provides our overall sense of literature and suggested directions of research, and the second section summarizes a few of the most relevant papers in greater depth.

97 MATHEMATICS AND COMPUTING↗