Skip to content

Add Magic Words and DRIFT papers - #2

Open
WhymustIhaveaname wants to merge 2 commits into
MobileLLM:mainfrom
SelfAdajoint:add-drift-magicwords
Open

Add Magic Words and DRIFT papers#2
WhymustIhaveaname wants to merge 2 commits into
MobileLLM:mainfrom
SelfAdajoint:add-drift-magicwords

Conversation

@WhymustIhaveaname

Copy link
Copy Markdown

Adds two papers to Security & Privacy, both in scope per the section's own framing ('related to LLM and LLM agents'). Under Adversarial Attacks: Magic Words (arXiv 2501.18280), a jailbreak that exploits skewed embedding-model output distributions to bypass safety guardrails. Under Inspection: DRIFT (arXiv 2601.14210), which trains a lightweight probe on mid-layer hidden states to flag low-confidence/untruthful generations in parallel with decoding, in the same vein as Inference-Time Intervention and Overthinking the Truth already listed there.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant