Resilience to Disinformation: Quantify and Improve Organizational Security Postures.
Abstract not provided.
SEARCH · Engineering Papers
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Abstract not provided.
Foreign disinformation campaigns undermine national security. Various supervised language modeling techniques in NLP can help to understand and dismantle these campaigns, but they rely heavily on large, labeled (often by humans) datasets. This work provides a solution to this problem in the form of an active learning (AL) framework, which is used to generate labeled datasets and leverage human input for detecting disinformation. The developed AL framework utilizes task adaptive pretraining to fully leverage the unlabeled data and boost the performance of the classifier used for labeling. A disinformation rhetoric metric was developed to measure the presence of common rhetorical techniques used in text that are meant to deceive, for both the classifier and human to use in the task of identifying disinformation. This metric was combined with an uncertainty criterion to create a hybrid acquisition method for AL, and this hybrid method was tested alongside other acquisition functions. A sophisticated and robust stopping strategy was developed to signal the AL process should terminate, saving human time from being wasted on iterations that would not significantly benefit classifier performance.
Social media has enabled a new era of manipulation in the information and cognitive domains. Deceptive content—misleading, falsified, and fabricated—is routinely created and spread in the modern social media environment with the intent to create confusion and widen political and social divides, and exploit the societal conflict exacerbated by these divides in the real-world (aka physical domain). Such disinformation campaigns demonstrate a threat to the integrity of economic, political, cultural, public health, and national security institutions around the world. In this work we overview our artificial intelligence (AI) capabilities to detect, describe, and defend against information operations on Twitter as an example social platform to understand the influence of misleading and falsified content diffusion and better enable those charged with defending against such manipulation to enable responsive parties to counter it. We first present novel linguistically-informed deep learning (DL) models for misinformation and disinformation detection, and present an in-depth linguistic analysis of psycho-linguistic markers across broad deception categories. We then demonstrate how our models perform in the multilingual and multimodal setting and categorize falsified and misleading content based on the intent to deceive. We also provide a large-scale analysis to describe user behavior and spread patterns while engaging with deceptive content and report novel findings about the immediate diffusion of deceptive content by characterizing the vulnerable sub-populations and their demographics, and explicitly measuring speed and scale of deception spread to uncover who shares deceptive content, how quickly, how much, and how evenly. In addition, we measure audience reactions to misinformation and disinformation at scale, distinguishing the reactions of users identified as bots versus humans. Finally, we take advantage of deep translation and generation models to create unique solutions for real-time defense against digital deception and discuss how to apply causal inference to prescribe and intervene into strategic communications jointly across information, cognitive, and physical domains.
Both human subject experiments and computational, modeling and simulations have been used to study detection of deception. This work aims to combine these two methods by integrating empirically-derived information (from human subject experiments) into agent-based models to generate novel insights into the complex problems of detection of disinformation content. Computational experiments are used to simulate across multiple scenarios for evaluation and decision-making regarding the validity of potentially deceptive scientific documents. Factors influencing the human agent behaviors in the model were identified through a human subject experiment that was conducted to evaluate and characterize decision making related to disinformation discernment. Correlation and regression analyses were used to translate insights from the human subjects experiment to inform the parameterization of agent features and scenario development. Three scenarios were evaluated with the agent-based models to help evaluate the replicability of the simulations (validation analysis) and assess the influence of human agent and document features (sensitivity analyses). A replication of the human participant experiment demonstrated that the agent-based simulations compare favorably to empirical findings. The agent-based modeling was then used to conduct sensitivity analysis on the accuracy of deception detection as a function of document proportions and human agent features. Results indicate that precision values are adversely impacted when the proportion of deceptive documents is lower in the overall sample, whereas recall values are more sensitive to changes in human agent features. These findings indicate important nuances in accuracy evaluations that should be further considered (including consideration of potential alternate metrics) in future agent-based models of disinformation. Additional areas for future exploration include extension of simulations to consider other ways to align the agent-based model design with psychological theory and inclusion of agent-agent interactions, especially as it pertains to sharing of scientific information within an organizational context.
Disinformation has become a problem in various spaces, from social media to organizations. Given the lack of evidence and transparency surrounding organizational disinformation, we aim to understand issue of information exposure effect and source trust from the lens of traditional information consumption online. Using an online experiment and multi-level mixed-effects modeling, we find that the exposure effect exists even over only two exposures to a headline, as long as participants are not shown fact-checking. Fact-checking can be effective in reducing the impact of the exposure effect. Additionally, we find that participants trusted sources significantly more before seeing the headlines. The possible presence of disinformation reduces trust in sources, even when considered highly neutral and reputable. These findings can inform simulations of agent interactions with misinformation as well as future experimental designs focused on evaluations of belief and trust based on recognizability of a source.
To date, disinformation research has focused largely on the production of false information ignoring the suppression of select information. We term this alternative form of disinformation information suppression. Information suppression occurs when facts are withheld with the intent to mislead. In order to detect information suppression, we focus on understanding the actors who withhold information. In this research, we use knowledge of human behavior to find signatures of different gatekeeping behaviors found in text. Specifically, we build a model to classify the different types of edits on Wikipedia using the added text alone and compare a human-informed feature engineering approach to a featureless algorithm. Being able to computationally distinguish gatekeeping behaviors is a first step towards identifying when information suppression is occurring.
In recent times, disinformation has spread rapidly through social media and news sites, biasing our (moral) judgements of other people and groups. “Deepfakes”, a new type of AI-generated media, represent a powerful new tool for spreading disinformation online. Although Deepfaked images, videos, and audio may appear genuine, they are actually hyper-realistic fabrications that enable one to digitally control what another person says or does. Given the recent emergence of this technology, we set out to examine the psychological impact of Deepfaked online content on viewers. Across seven preregistered studies (N = 2558) we exposed participants to either genuine or Deepfaked content, and then measured its impact on their explicit (self-reported) and implicit (unintentional) attitudes as well as behavioral intentions. Results indicated that Deepfaked videos and audio have a strong psychological impact on the viewer, and are just as effective in biasing their attitudes and intentions as genuine content. Many people are unaware that Deepfaking is possible; find it difficult to detect when they are being exposed to it; and most importantly, neither awareness nor detection serves to protect people from its influence. All preregistrations, data and code available at osf.io/f6ajb.
With the increasing use of automated, machine learning-driven tools and the downstream impact that algorithmic judgements can have, it is critical to develop models that are robust to evolving or manipulated inputs. Evaluating the reliability of multimodal models across linguistic variations to understand model susceptibility to intentional linguistic adversarial attacks as well as natural linguistic variations is essential in this pursuit. We present extensive analysis of model robustness and susceptibility to linguistic variations in the setting of deceptive news detection, a difficult classification task that is an increasingly important problem to solve with the impact of misinformation spread online. We evaluate the effectiveness of incorporating adversarial defense strategies and measure model susceptibility to state-of-the-art adversarial attacks using two types of linguistic attacks — character and word perturbations. We consider two multiclass prediction tasks — a 3-way classification of tweets as trustworthy, propaganda, or disinformation; and a 4-way classification as clickbait, hoax, satire, or conspiracy — and compare the performance of three embeddings that have been state-of-the-art for several NLP tasks — GloVe, ELMo, and BERT — to highlight consistent trends in susceptibility, high confidence misclassifications, and high impact failures. We find that character or mixed ensemble models are the most effective defense mechanisms and that character perturbations are a more effective attack than word perturbations for deception classification.
Organizations play a key role in supporting various societal functions, ranging from environmental governance to the manufacturing of goods. Here, the behaviors of organization are impacted by various influences, including information, technology, authority, economic leverage, historical experiences, and external factors, such as regulations. This paper introduces a generalized framework, focused on the relative structure of an organization (tight vs. loose), that can be used to understand how different influence pathways can impact decision-making within differently structured organizations. This generalized framework is then translated into a modeling and simulation platform to support and assess implications of these structural differences in resilience to disinformation (measured by organizational behaviors of timeliness and inclusion of quality information) using a systems dynamics approach Preliminary results indicate that a tightly structured organization may be less timely at processing information but could be more resilient against using poor quality information in organizational decisions compared to a loosely structured organization. Ongoing work is underway to understand the robustness of these findings and to validate current model design activities with empirical insights.
Scientists, social scientists, risk communicators, and many others are often thrust into a crisis situation where they need to interact with a range of stakeholders, including governmental personnel (tribal, U.S. federal, state, local), local residents, and other publics, as well as other scientists and other risk communicators in situations where information is incomplete and evolving. This paper provides: (1) an overall framework for thinking about communication during crises, from acute to chronic, and local to widespread, (2) a template for the types of ecological information needed to address public and environmental concerns, and (3) examples to illustrate how this information will aid risk communicators. The main goal is providing an approach to the knowledge needed by communicators to address the challenges of protecting ecological resources during an environmental crisis, or for an on-going, chronic environmental issue. To understand the risk to these ecological resources, it is important to identify the type of event, whether it is acute or chronic (or some combination of these), what receptors are at risk, and what stressors are involved (natural, biological, chemical, radiological). For ecological resources, the key information a communicator needs for a crisis is whether any of the following are present: threatened or endangered species, species of special concern, species groups of concern (e.g., neotropical bird migrants, breeding frogs in vernal ponds, rare plant assemblages), unique or rare habitats, species of commercial and recreational interest, and species/habitats of especial interest for medicinal, cultural, or religious activities. Communication among stakeholders is complicated with respect to risk to ecological receptors because of differences in trust, credibility, empathy, perceptions, world view valuation of the resources, and in many cases, a history of misinformation, disinformation, or no information. Exposure of salmon spawning in the Columbia River to hexavalent chromium from the Hanford Site is used as an example of communication challenges with different stakeholders, including Native Americans with Tribal Treaty rights to the land.
On March 30th and 31st, 2022, the University of Texas at Austin (UT) Office of the Vice President for Research (OVPR) hosted Sandia National Laboratories (Sandia) for “Sandia Day at UT Austin” to understand the status of the strategic partnership and explore opportunities for partnership growth. The event brought together more than 115 UT and Sandia participants including executive leadership, researchers, faculty, staff, and students. Sandia Day primarily consisted of a half-day leadership meeting, a research poster session and networking event, and three break-out sessions focused on strategic priority areas: Microelectronics, Energy and Climate Security, and High-Performance and Edge Computing. Appendix A contains the full Sandia Day agenda. Additional meetings and workshops (adjunct meetings) were held in conjunction with Sandia Day to maximize partnership exploration. Adjunct meetings were Hypersonics, Decarbonization, Disinformation, and Battery Workshops. A summary of Sandia Day events, sessions, and meetings follows.
Widespread integration of social media into daily life has fundamentally changed the way society communicates, and, as a result, how individuals develop attitudes, personal philosophies, and worldviews. The excess spread of disinformation and misinformation due to this increased connectedness and streamlined communication has been extensively studied, simulated, and modeled. Less studied is the interaction of many pieces of misinformation, and the resulting formation of attitudes. We develop a framework for the simulation of attitude formation based on exposure to multiple cognitions. We allow a set of cognitions with some implicit relational topology to spread on a social network, which is defined with separate layers to specify online and offline relationships. An individual’s opinion on each cognition is determined by a process inspired by the Ising model for ferromagnetism. We conduct experimentation using this framework to test the effect of topology, connectedness, and social media adoption on the ultimate prevalence of and exposure to certain attitudes.
ARES was in part motivated by the determination of President’s Council of Advisors on Science and Technology (PCAST) on May 13th, 2023 that published a set of inquiries: In an era in which convincing images, audio, and text can be generated with ease on a massive scale, how can we ensure reliable access to verifiable, trustworthy information? How can we be certain that a particular piece of media is genuinely from the claimed source? What technologies, policies, and infrastructure can be developed to detect and counter AI-generated disinformation? In an effort to automatically analyze and patch/optimize code the work in this report describes various neural Machine Learning (ML) analysis engine implementations to assist in situations where source code is deficient or completely lacking to decompile (lift) binary code to ’C’. The goal is to gradually reduce human intervention. To this end, two Large Language Model (LLM) variants (Code LLama 2, LLama 3.1 and Starcoder1, Starcoder 2) where finetuned with ’before/after’ code pairs on the OpenBLAS library. LLama trained on the lowering process, Starcoder trained on the lifting process with National Security Agency’s (NSA) open-source Ghidra decompiler assist. The inferencing test results indicate correctness for only very short sequences for Starcoder 2. Moving forward, the experiments conclude with a set of recommendations of required resources and technologies
The emerging multipolar international security environment represents a fundamental restructuring of global nuclear balance of power to include two nuclear peer competitors, growing non-peer nuclear threats, and concerns of nuclear latency from both allies and adversaries. Conflicts in the grey zone, cyber operations, mis- and disinformation campaigns, and emerging disruptive technologies like drones, and hypersonic missiles are becoming more prevalent. These present a risk of cross-domain and multi-domain conflicts that may not follow known escalatory patterns. In order to prepare for the new deterrence environment, it is critical to have quantitative and qualitative understandings of these cross-domain conflicts, their potential for escalation, and which systems they may impact. To that end, our team created a Multi-Layer Network (MLN) model of ‘integrated deterrence’ where instruments of national power are modeled as individual network graph layers that include efforts from all domains. We then evaluate the potential for escalation against escalation scenarios. Analysis of the escalation scenarios is then used to identify insights of potential risk and escalation within integrated deterrence.
With the ever-increasing pace of research and high volume of scholarly communication, scholars face a daunting task. Not only must they keep up with the growing literature in their own and related fields, scholars increasingly also need to rebut pseudo-science and disinformation. These needs have motivated an increasing focus on computational methods for enhancing search, summarization, and analysis of scholarly documents. However, the various strands of research on scholarly document processing remain fragmented. To reach out to the broader NLP and AI/ML community, pool distributed efforts in this area, and enable shared access to published research, we held the 2nd Workshop on Scholarly Document Processing (SDP) at NAACL 2021 as a virtual event (https://sdproc.org/2021/). The SDP workshop consisted of a research track, three invited talks, and three Shared Tasks (LongSumm 2021, SCIVER, and 3C). The program was geared towards the application of NLP, information retrieval, and data mining for scholarly documents, with an emphasis on identifying and providing solutions to open challenges.
Massive digital disinformation is one of the main risks of modern society. Hundreds of models and linguistic analyses have been done to compare and contrast misleading and credible content online. However, most models do not remove the confounding factor of a topic or narrative when training, so the resulting models learn a clear topical separation for misleading versus credible content. We study the feasibility of using two strategies to disentangle the topic bias from the models to understand and explicitly measure linguistic and stylistic properties of content from misleading versus credible content. First, we develop conditional generative models to create news content that is characteristic of different credibility levels. We perform multi-dimensional evaluation of model performance on mimicking both the style and linguistic differences that distinguish news of different credibility using machine translation metrics and classification models. We show that even though generative models are able to imitate both the style and language of the original content, additional conditioning on both the news category and the topic leads to reduced performance. In a second approach, we perform deception style ``transfer" by translating deceptive content into the style of credible content and vice versa. Extending earlier studies, we demonstrate that, when conditioned on a topic, deceptive content is shorter, less readable, more biased, and more subjective than credible content, and transferring the style from deceptive to credible content is more challenging than the opposite direction.