Exploratory four-probe experiment on attribution faithfulness in multi-factor LLM reasoning. Personal research notes across eight models, not a validated benchmark.
benchmark attribution ai-safety explainability legal-nlp mechanistic-interpretability llm-evaluation llm-auditing trustworthy-ml llm-faithfulness
-
Updated
Apr 14, 2026 - Python