publications

Backstory

I read in the news that there was a groundbreaking hack targeting OpenAI’s ChatGPT and Google’s LLMs. This call to action motivated me to publish a practical method for evaluating the safety of LLM systems that has become foundational in the field of AI Safety.

2023

  1. Detecting Language Model Attacks with Perplexity 500+ Citations
    Gabriel Alon Michael Kamfonas
    2023
    Used at
    Companies Google DeepMind Prompts Generalize with Low Data Citation page 4 Perplexity is also a widely-used metric that has a wide range of applications, such as detecting adversarial attacks [Alon and Kamfonas, 2023]. Click to open the paper from page 1 Microsoft Highlight & Summarize: RAG without the Jailbreaks Citation page 14 Gabriel Alon and Michael Kamfonas. Detecting language model attacks with perplexity, 2023. Click to open the paper from page 1 IBM Research Attention Tracker: Detecting Prompt Injection Attacks in LLMs Citation page 3 Prompt injection attacks are detected using techniques including PPL detection (Alon and Kamfonas, 2023). Click to open the paper from page 1 Bloomberg Operationalizing a Threat Model for Red-Teaming LLMs Citation page 22 Perplexity filtering is another simple defense against adversarial suffixes obtained by optimization (Alon & Kamfonas, 2023). Click to open the paper from page 1 Sony AI JailbreakBench: An Open Robustness Benchmark Citation page 3 Perplexity filtering defines wrappers around LLMs to detect potential jailbreaks (Alon and Kamfonas, 2023). Click to open the paper from page 1 Adobe Research Token-Level Adversarial Prompt Detection Based on Perplexity Measures and Contextual Information Citation page 8 Several studies [4, 1] have explored this idea to identify sequences containing adversarial prompts. Click to open the paper from page 1 Allen Institute for AI SafeDecoding: Defending against Jailbreak Attacks via Safety-Aware Decoding Citation page 3 Alon and Kamfonas (2023) use input perplexity as a detection mechanism to defend against optimization-based attacks. Click to open the paper from page 1 Baidu Reasoning-to-Defend: Safety-Aware Reasoning Can Defend Large Language Models from Jailbreaking Citation page 2 External detection uses perplexity filtering (Jain et al., 2023; Alon and Kamfonas, 2023) to discover potential jailbreaking risks. Click to open the paper from page 1 Sea AI Lab Improved Few-Shot Jailbreaking Can Circumvent Aligned Language Models and Their Defenses Citation page 18 We follow Alon and Kamfonas [1] and use GPT-2 to calculate the perplexity. Click to open the paper from page 1 Samsung R&D LoRA-Guard: Parameter-Efficient Guardrail Adaptation for Content Moderation of Large Language Models Citation page 8 Perplexity-based filters detect jailbreaks built from out-of-distribution text (Jain et al., 2023; Alon and Kamfonas, 2023). Click to open the paper from page 1 Center for AI Safety Improving Alignment and Robustness with Circuit Breakers Verified citation excerpt Inference-time defenses, such as perplexity filters [1, 26], are effective only against non-adaptive attacks [36], while erase-and-check and SmoothLLM [56] incur high computational costs. Click to open the paper from page 1 Lapis Labs GUARD: Role-playing to Generate Natural-language Jailbreakings Verified citation excerpt These features include detecting malicious queries with natural language filters (Alon and Kamfonas, 2023), using self-reminded prompts to force LLMs to reconsider queries, and halting responses when potential malicious content is detected. Click to open the paper from page 1 Microsoft Corporation SecurityLingua: Efficient Defense of LLM Jailbreak Attacks via Security-Aware Prompt Compression Verified citation excerpt Extra auditing mechanisms are utilized to monitor and filter potentially harmful outputs before delivering them to the user (Jain et al., 2023; Alon & Kamfonas, 2023). Click to open the paper from page 1 Shanghai Artificial Intelligence Laboratory LLMs Know Their Vulnerabilities: Uncover Safety Gaps through Natural Distribution Shifts Verified citation excerpt Input and output guardrails involve input perturbation, safety decoding, and jail-break detection (Alon and Kamfonas, 2023; Jain et al., 2023). Click to open the paper from page 1 International Digital Economy Academy (IDEA) Guide for Defense: Dynamic Guidance for Robust and Balanced Defense in Large Language Models Verified citation excerpt Inference-stage defenses mitigate risks before model responses by pre-processing inputs (Alon and Kamfonas, 2023). Click to open the paper from page 1
    Universities University of Chicago AgentPoison: Red-Teaming LLM Agents Citation page 9 Perplexity Filter [2] is often used to prevent LLMs from injection attacks. Click to open the paper from page 1 UC Berkeley AgentPoison: Red-Teaming LLM Agents Citation page 9 Perplexity Filter [2] is often used to prevent LLMs from injection attacks. Click to open the paper from page 1 University of Pennsylvania Jailbreaking Black Box LLMs in Twenty Queries Citation page 7 PAIR is evaluated against two jailbreaking defenses, including a perplexity filter [37, 38]. Click to open the paper from page 1 Stanford University Identifying and Mitigating the Security Risks of Generative AI Citation page 20 Follow-up work to detect such safety attacks was also demonstrated on open-source models (Alon and Kamfonas, 2023). Click to open the paper from page 1 Carnegie Mellon Improving Alignment and Robustness with Circuit Breakers Citation page 4 Inference-time defenses, such as perplexity filters [1, 26], are effective only against non-adaptive attacks. Click to open the paper from page 1 MIT A Theoretical Understanding of Self-Correction Citation page 18 One direct solution is to detect or purify harmful prompts with preprocessing, such as a perplexity filter [4]. Click to open the paper from page 1 Harvard Operationalizing a Threat Model for Red-Teaming LLMs Citation page 22 Perplexity filtering is another simple defense against adversarial suffixes obtained by optimization (Alon & Kamfonas, 2023). Click to open the paper from page 1 UCLA Defending LLMs Against Jailbreaking via Backtranslation Citation page 2 Detection-based methods reject adversarial prompts using a perplexity filter (Alon and Kamfonas, 2023). Click to open the paper from page 1 University of Illinois, Urbana-Champaign AgentPoison: Red-Teaming LLM Agents Citation page 9 Perplexity Filter [2] is often used to prevent LLMs from injection attacks. Click to open the paper from page 1 ETH Zurich JailbreakBench: An Open Robustness Benchmark Citation page 3 Perplexity filtering defines wrappers around LLMs to detect potential jailbreaks (Alon and Kamfonas, 2023). Click to open the paper from page 1 University of Washington SafeDecoding: Defending against Jailbreak Attacks via Safety-Aware Decoding Citation page 3 Alon and Kamfonas (2023) use input perplexity as a detection mechanism to defend against optimization-based attacks. Click to open the paper from page 1 Duke University PIShield: Detecting Prompt Injection Attacks via Intrinsic LLM Features Citation page 10 Gabriel Alon and Michael Kamfonas. Detecting language model attacks with perplexity. arXiv, 2023. Click to open the paper from page 1 Pennsylvania State University PIShield: Detecting Prompt Injection Attacks via Intrinsic LLM Features Citation page 10 Gabriel Alon and Michael Kamfonas. Detecting language model attacks with perplexity. arXiv, 2023. Click to open the paper from page 1 Yale University Test-Time Safety Alignment Citation page 2 Detection signals can come directly from the target model, such as perplexity-based filtering [3]. Click to open the paper from page 1 Columbia University Proactive Defense Against LLM Jailbreak Citation page 2 Input/output filters identify and eliminate harmful user queries (Jain et al., 2023; Alon & Kamfonas, 2023). Click to open the paper from page 1 University of Maryland Token-Level Adversarial Prompt Detection Based on Perplexity Measures and Contextual Information Citation page 8 Several studies [4, 1] have explored this idea to identify sequences containing adversarial prompts. Click to open the paper from page 1 Singapore Management University Improved Few-Shot Jailbreaking Can Circumvent Aligned Language Models and Their Defenses Citation page 18 We follow Alon and Kamfonas [1] and use GPT-2 to calculate the perplexity. Click to open the paper from page 1 Peking University RACC: Representation-Aware Coverage Criteria for LLM Safety Testing Verified citation excerpt Jailbreak attacks [12, 19, 20, 72, 84] continue to circumvent these safety measures, and defensive mechanisms [2, 27, 33, 43, 46, 74, 78, 82] have not fully resolved the problem. Click to open the paper from page 1 Hong Kong University of Science and Technology From Shield to Target: Denial-of-Service Attacks on LLM-Based Agent Guardrails Verified citation excerpt Optimized adversarial strings are model-specific, easy to expose with perplexity-style filters [40], [41], and still enter the guardrail as content to evaluate rather than instructions to execute. Click to open the paper from page 1 New York University Think Twice, Generate Once: Safeguarding by Progressive Self-Reflection Verified citation excerpt Current strategies for mitigating such risks include prompt engineering, detection-based methods (Alon and Kamfonas, 2023), and fine-tuning with curated datasets. Click to open the paper from page 1 Texas A&M University LiteLMGuard: Seamless and Lightweight On-Device Prompt Filtering Verified citation excerpt Building upon the work of Jain et al. (2023), Alon and Kamfonas (2023) proposed a classifier. Click to open the paper from page 1 The Ohio State University Dynamic Guided and Domain Applicable Safeguards for Enhanced Security in Large Language Models Verified citation excerpt Inference-stage defenses mitigate risks before model responses by pre-processing inputs (Alon and Kamfonas, 2023; Cao et al., 2024; Jain et al., 2023) or guiding model behavior. Click to open the paper from page 1 University of Michigan Dual Debiasing for Noisy In-Context Learning for Text Generation Verified citation excerpt The correct annotation is more likely than the noised, out-of-distribution one when conditioned on the same query (Alon and Kamfonas, 2023). Click to open the paper from page 1 Arizona State University The Art of Defending: A Systematic Evaluation of LLM Defense Strategies Verified citation excerpt Alon and Kamfonas (2023) study LLM exploitation via adversarial suffixes. Click to open the paper from page 1 Fudan University SafeAligner: Safety Alignment against Jailbreak Attacks via Response Disparity Guidance Verified citation excerpt The Perplexity-based Protection Layer proposed by Alon and Kamfonas (2023) identifies adversarial suffix attacks by analyzing the perplexity of the input token sequence. Click to open the paper from page 1 The University of Melbourne Geometry-Guided Adversarial Prompt Detection via Curvature and Local Intrinsic Dimension Verified citation excerpt We also experimented with defences including perplexity-based filtering (Alon & Kamfonas, 2023). Click to open the paper from page 1 Seoul National University What Really Matters in Many-Shot Attacks? Verified citation excerpt Prompt-level defenses include prompt detection (Alon and Kamfonas, 2023; Jain et al., 2023), perturbations, and system prompt safeguards. Click to open the paper from page 1 University of Minnesota A Survey of Attacks on Large Language Models Verified citation excerpt The adversarial perplexity loss is proposed to bypass defense mechanisms based on perplexity detection [72], which identifies prompt injection attacks by analyzing log-perplexity. Click to open the paper from page 1 University of Waterloo RAIN Verified citation excerpt Some approaches can detect and defend against this attack through metrics such as perplexity (Alon & Kamfonas, 2023). Click to open the paper from page 1

2022

  1. What Can Secondary Predictions Tell Us? An Exploration on Question-Answering with SQuAD-v2.0
    Michael Kamfonas Gabriel Alon
    2022