Skip to content

Latest commit

 

History

62 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 

Repository files navigation

Persuasion-Aware Jailbreaking as Social Engineering: Fingerprinting LLM Susceptibilities and Defense Implications

This repository contains the implementation and evaluation framework for studying persuasion-based jailbreak attacks on Large Language Models (LLMs). The project operationalizes Cialdini’s persuasion principles to generate adversarial prompts and analyze model-specific susceptibility profiles.

Pipeline overview

overview

Datasets

  • AdvBench: This dataset contains 520 queries, covering various types of harmful behavior.
  • StrongREJECT: This dataset encompasses six major categories of behaviors that are explicitly prohibited under all major LLM usage policies: (1) illegal goods and services, (2) non-violent crimes, (3) hate, harassment, and discrimination, (4) disinformation and deception, (5) violence, and (6) sexual content.

Victim Models

  • Vicuna-7b
  • Llama2-7b-chat
  • Llama3
  • DeepSeek-R1
  • Gemma3
  • Phi4

Note: Target models are available via Ollama.

Attack Baselines

  • GCG: finds suffixes via gradient synthesis

  • PAIR: an algorithm that creates semantic jailbreaks using only black-box access to an LLM

  • PAP: generates persuasive adversarial prompts using different strategies

    Note: We leveraged baseline implementations provided by the StrongReject benchmark.

Code Structure

  • persuasive_prompt_generation.py: Generates persuasive variants of harmful queries based on Cialdini’s persuasion principles. These variants are used to analyze how different persuasion strategies influence LLM compliance with harmful instructions.

  • model_response_collector.py: Handles the process of querying LLMs with both original and persuasion-based prompts, collecting model outputs to measure refusal behavior and susceptibility to persuasion-driven jailbreaks.

  • evaluation.py: Evaluates collected responses using metrics such as Attack Success Rate (ASR), Informativeness Score (IS), and Influence Power (IP) to quantify the effectiveness and impact of various persuasion strategies across models.

  • Attack baselines: Contains baseline implementations of traditional adversarial prompt generation techniques, serving as benchmarks to compare against persuasion-aware jailbreak methods.

  • Defense baselines: Includes implementations of defense strategies used to evaluate the robustness of LLMs against persuasion-based attacks.

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages