publications
Backstory
I read in the news that there was a groundbreaking hack targeting OpenAI’s ChatGPT and Google’s LLMs. This call to action motivated me to publish a practical method for evaluating the safety of LLM systems that has become foundational in the field of AI Safety.
2023
- Detecting Language Model Attacks with Perplexity 500+ Citations2023Used atCompanies Google DeepMind Prompts Generalize with Low Data Citation page 4
Perplexity is also a widely-used metric that has a wide range of applications, such as detecting adversarial attacks [Alon and Kamfonas, 2023].
Click to open the paper from page 1 Microsoft Highlight & Summarize: RAG without the Jailbreaks Citation page 14Gabriel Alon and Michael Kamfonas. Detecting language model attacks with perplexity, 2023.
Click to open the paper from page 1 IBM Research Attention Tracker: Detecting Prompt Injection Attacks in LLMs Citation page 3Prompt injection attacks are detected using techniques including PPL detection (Alon and Kamfonas, 2023).
Click to open the paper from page 1 Bloomberg Operationalizing a Threat Model for Red-Teaming LLMs Citation page 22Perplexity filtering is another simple defense against adversarial suffixes obtained by optimization (Alon & Kamfonas, 2023).
Click to open the paper from page 1 Sony AI JailbreakBench: An Open Robustness Benchmark Citation page 3Perplexity filtering defines wrappers around LLMs to detect potential jailbreaks (Alon and Kamfonas, 2023).
Click to open the paper from page 1 Adobe Research Token-Level Adversarial Prompt Detection Based on Perplexity Measures and Contextual Information Citation page 8Several studies [4, 1] have explored this idea to identify sequences containing adversarial prompts.
Click to open the paper from page 1 Allen Institute for AI SafeDecoding: Defending against Jailbreak Attacks via Safety-Aware Decoding Citation page 3Alon and Kamfonas (2023) use input perplexity as a detection mechanism to defend against optimization-based attacks.
Click to open the paper from page 1 Baidu Reasoning-to-Defend: Safety-Aware Reasoning Can Defend Large Language Models from Jailbreaking Citation page 2External detection uses perplexity filtering (Jain et al., 2023; Alon and Kamfonas, 2023) to discover potential jailbreaking risks.
Click to open the paper from page 1 Sea AI Lab Improved Few-Shot Jailbreaking Can Circumvent Aligned Language Models and Their Defenses Citation page 18We follow Alon and Kamfonas [1] and use GPT-2 to calculate the perplexity.
Click to open the paper from page 1 Samsung R&D LoRA-Guard: Parameter-Efficient Guardrail Adaptation for Content Moderation of Large Language Models Citation page 8Perplexity-based filters detect jailbreaks built from out-of-distribution text (Jain et al., 2023; Alon and Kamfonas, 2023).
Click to open the paper from page 1 Center for AI Safety Improving Alignment and Robustness with Circuit Breakers Verified citation excerptInference-time defenses, such as perplexity filters [1, 26], are effective only against non-adaptive attacks [36], while erase-and-check and SmoothLLM [56] incur high computational costs.
Click to open the paper from page 1 Lapis Labs GUARD: Role-playing to Generate Natural-language Jailbreakings Verified citation excerptThese features include detecting malicious queries with natural language filters (Alon and Kamfonas, 2023), using self-reminded prompts to force LLMs to reconsider queries, and halting responses when potential malicious content is detected.
Click to open the paper from page 1 Microsoft Corporation SecurityLingua: Efficient Defense of LLM Jailbreak Attacks via Security-Aware Prompt Compression Verified citation excerptExtra auditing mechanisms are utilized to monitor and filter potentially harmful outputs before delivering them to the user (Jain et al., 2023; Alon & Kamfonas, 2023).
Click to open the paper from page 1 Shanghai Artificial Intelligence Laboratory LLMs Know Their Vulnerabilities: Uncover Safety Gaps through Natural Distribution Shifts Verified citation excerptInput and output guardrails involve input perturbation, safety decoding, and jail-break detection (Alon and Kamfonas, 2023; Jain et al., 2023).
Click to open the paper from page 1 International Digital Economy Academy (IDEA) Guide for Defense: Dynamic Guidance for Robust and Balanced Defense in Large Language Models Verified citation excerptInference-stage defenses mitigate risks before model responses by pre-processing inputs (Alon and Kamfonas, 2023).
Click to open the paper from page 1Universities University of Chicago AgentPoison: Red-Teaming LLM Agents Citation page 9Perplexity Filter [2] is often used to prevent LLMs from injection attacks.
Click to open the paper from page 1 UC Berkeley AgentPoison: Red-Teaming LLM Agents Citation page 9Perplexity Filter [2] is often used to prevent LLMs from injection attacks.
Click to open the paper from page 1 University of Pennsylvania Jailbreaking Black Box LLMs in Twenty Queries Citation page 7PAIR is evaluated against two jailbreaking defenses, including a perplexity filter [37, 38].
Click to open the paper from page 1 Stanford University Identifying and Mitigating the Security Risks of Generative AI Citation page 20Follow-up work to detect such safety attacks was also demonstrated on open-source models (Alon and Kamfonas, 2023).
Click to open the paper from page 1 Carnegie Mellon Improving Alignment and Robustness with Circuit Breakers Citation page 4Inference-time defenses, such as perplexity filters [1, 26], are effective only against non-adaptive attacks.
Click to open the paper from page 1 MIT A Theoretical Understanding of Self-Correction Citation page 18One direct solution is to detect or purify harmful prompts with preprocessing, such as a perplexity filter [4].
Click to open the paper from page 1 Harvard Operationalizing a Threat Model for Red-Teaming LLMs Citation page 22Perplexity filtering is another simple defense against adversarial suffixes obtained by optimization (Alon & Kamfonas, 2023).
Click to open the paper from page 1 UCLA Defending LLMs Against Jailbreaking via Backtranslation Citation page 2Detection-based methods reject adversarial prompts using a perplexity filter (Alon and Kamfonas, 2023).
Click to open the paper from page 1 University of Illinois, Urbana-Champaign AgentPoison: Red-Teaming LLM Agents Citation page 9Perplexity Filter [2] is often used to prevent LLMs from injection attacks.
Click to open the paper from page 1 ETH Zurich JailbreakBench: An Open Robustness Benchmark Citation page 3Perplexity filtering defines wrappers around LLMs to detect potential jailbreaks (Alon and Kamfonas, 2023).
Click to open the paper from page 1 University of Washington SafeDecoding: Defending against Jailbreak Attacks via Safety-Aware Decoding Citation page 3Alon and Kamfonas (2023) use input perplexity as a detection mechanism to defend against optimization-based attacks.
Click to open the paper from page 1 Duke University PIShield: Detecting Prompt Injection Attacks via Intrinsic LLM Features Citation page 10Gabriel Alon and Michael Kamfonas. Detecting language model attacks with perplexity. arXiv, 2023.
Click to open the paper from page 1 Pennsylvania State University PIShield: Detecting Prompt Injection Attacks via Intrinsic LLM Features Citation page 10Gabriel Alon and Michael Kamfonas. Detecting language model attacks with perplexity. arXiv, 2023.
Click to open the paper from page 1 Yale University Test-Time Safety Alignment Citation page 2Detection signals can come directly from the target model, such as perplexity-based filtering [3].
Click to open the paper from page 1 Columbia University Proactive Defense Against LLM Jailbreak Citation page 2Input/output filters identify and eliminate harmful user queries (Jain et al., 2023; Alon & Kamfonas, 2023).
Click to open the paper from page 1 University of Maryland Token-Level Adversarial Prompt Detection Based on Perplexity Measures and Contextual Information Citation page 8Several studies [4, 1] have explored this idea to identify sequences containing adversarial prompts.
Click to open the paper from page 1 Singapore Management University Improved Few-Shot Jailbreaking Can Circumvent Aligned Language Models and Their Defenses Citation page 18We follow Alon and Kamfonas [1] and use GPT-2 to calculate the perplexity.
Click to open the paper from page 1 Peking University RACC: Representation-Aware Coverage Criteria for LLM Safety Testing Verified citation excerptJailbreak attacks [12, 19, 20, 72, 84] continue to circumvent these safety measures, and defensive mechanisms [2, 27, 33, 43, 46, 74, 78, 82] have not fully resolved the problem.
Click to open the paper from page 1 Hong Kong University of Science and Technology From Shield to Target: Denial-of-Service Attacks on LLM-Based Agent Guardrails Verified citation excerptOptimized adversarial strings are model-specific, easy to expose with perplexity-style filters [40], [41], and still enter the guardrail as content to evaluate rather than instructions to execute.
Click to open the paper from page 1 New York University Think Twice, Generate Once: Safeguarding by Progressive Self-Reflection Verified citation excerptCurrent strategies for mitigating such risks include prompt engineering, detection-based methods (Alon and Kamfonas, 2023), and fine-tuning with curated datasets.
Click to open the paper from page 1 Texas A&M University LiteLMGuard: Seamless and Lightweight On-Device Prompt Filtering Verified citation excerptBuilding upon the work of Jain et al. (2023), Alon and Kamfonas (2023) proposed a classifier.
Click to open the paper from page 1 The Ohio State University Dynamic Guided and Domain Applicable Safeguards for Enhanced Security in Large Language Models Verified citation excerptInference-stage defenses mitigate risks before model responses by pre-processing inputs (Alon and Kamfonas, 2023; Cao et al., 2024; Jain et al., 2023) or guiding model behavior.
Click to open the paper from page 1 University of Michigan Dual Debiasing for Noisy In-Context Learning for Text Generation Verified citation excerptThe correct annotation is more likely than the noised, out-of-distribution one when conditioned on the same query (Alon and Kamfonas, 2023).
Click to open the paper from page 1 Arizona State University The Art of Defending: A Systematic Evaluation of LLM Defense Strategies Verified citation excerptAlon and Kamfonas (2023) study LLM exploitation via adversarial suffixes.
Click to open the paper from page 1 Fudan University SafeAligner: Safety Alignment against Jailbreak Attacks via Response Disparity Guidance Verified citation excerptThe Perplexity-based Protection Layer proposed by Alon and Kamfonas (2023) identifies adversarial suffix attacks by analyzing the perplexity of the input token sequence.
Click to open the paper from page 1 The University of Melbourne Geometry-Guided Adversarial Prompt Detection via Curvature and Local Intrinsic Dimension Verified citation excerptWe also experimented with defences including perplexity-based filtering (Alon & Kamfonas, 2023).
Click to open the paper from page 1 Seoul National University What Really Matters in Many-Shot Attacks? Verified citation excerptPrompt-level defenses include prompt detection (Alon and Kamfonas, 2023; Jain et al., 2023), perturbations, and system prompt safeguards.
Click to open the paper from page 1 University of Minnesota A Survey of Attacks on Large Language Models Verified citation excerptThe adversarial perplexity loss is proposed to bypass defense mechanisms based on perplexity detection [72], which identifies prompt injection attacks by analyzing log-perplexity.
Click to open the paper from page 1 University of Waterloo RAIN Verified citation excerptSome approaches can detect and defend against this attack through metrics such as perplexity (Alon & Kamfonas, 2023).
Click to open the paper from page 1Every reviewed citing paper for which a citation excerpt is available.
2022
- What Can Secondary Predictions Tell Us? An Exploration on Question-Answering with SQuAD-v2.02022