The production engine for directional ablation. Unalign / remove models censorship efficiently on any hardware.
-
Updated
Jun 28, 2026 - Python
The production engine for directional ablation. Unalign / remove models censorship efficiently on any hardware.
Educational analysis of LLM alignment, safety behavior, and framing-sensitive response patterns.
SoftPrompt-IR is a low-level symbolic annotation layer for LLM prompts, making intent strength, direction, and priority explicit. It is not a DSL or framework, but a minimal, composable way to reduce ambiguity, improve safety, and structure prompts.
DSPy framework for detecting and preventing safety override cascades in LLM systems. Research-grade implementation for studying when completion urgency overrides safety constraints.
Adversarial evaluation framework for embodied and agentic AI — failure-first methodology, jailbreak corpus, VLA red-teaming, and policy research.
🌐 Detect and prevent safety overrides in LLM systems with this DSPy-based framework, ensuring actions align with safety constraints.
Research on multi-turn conversational manipulation of LLMs — can a small specialist attacker beat scale? Defensive AI-safety red-team methodology with programmatic judges. Pre-alpha.
Contract-enforced sandbox for studying AI agent self-replication safety
Explore glider aviation safety through in-depth data analysis. This project leverages incident reports and manufacturing data, utilizing Python and Jupyter Notebooks for trend identification, risk assessment, and safety enhancement in glider aviation.
AI Safety Observatory for Africa is a full-stack platform that evaluates and visualises the safety of large language models across African languages, AI risk categories, and different model architectures. Built for the Global South AI Safety Hackathon, it identifies potential safety gaps where harmful prompts in languages such as isiZulu, Afrikaans
Date: July 20, 2026 Author: Kevin Michael Lent I am publicly establishing the timeline of my work in layered transition structures related or equal to areas in AI referred to as J-Space in large language models.
Research portfolio on why organisations fail to act on known, demonstrated risks: canonical essay, dependency-ordered research backlog, codebook, and agent-assistance audit trail.
Open evidence and benchmark architecture for human-automation readiness, recoverability, and joint operational risk.
Fenrir — Breaker of Chains. Abliteration toolkit: find the refusal directions, remove them, prove what changed. (formerly Absolver)
To associate your repository with the safety-research topic, visit your repo's landing page and select "manage topics."