-
Notifications
You must be signed in to change notification settings - Fork 10
Expand file tree
/
Copy pathcontribute.astro
More file actions
161 lines (151 loc) · 12.1 KB
/
Copy pathcontribute.astro
File metadata and controls
161 lines (151 loc) · 12.1 KB
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
---
export const prerender = true;
import Base from "../layouts/Base.astro";
import SideNav from "../components/SideNav.astro";
import BackHome from "../components/BackHome.astro";
import { PROMPT_VERSION } from "@defipunkd/prompts";
---
<Base title="Contribute · DeFiPunk'd" description="DEFI@home — help DeFiPunk'd assess DeFi protocols by running a pinned prompt in an LLM of your choice.">
<SideNav contentMaxWidth={800} links={[
{ href: "#the-flow", label: "The flow" },
{ href: "#pinned-prompt", label: "Pinned prompt" },
{ href: "#evidence", label: "Evidence" },
{ href: "#unknown", label: "Saying unknown" },
{ href: "#sharing-your-chat", label: "Sharing chat" },
{ href: "#batch", label: "Batch submissions" },
{ href: "#quorum", label: "Quorum" },
{ href: "#weighting", label: "Weighting" },
{ href: "#reconcile", label: "Reconcile" },
]} />
<main class="page">
<BackHome />
<h1>DEFI@home — contribute an assessment</h1>
<p class="lede">
DEFI@home assessments are produced by contributors, not crawlers. You run a pinned prompt through the LLM of your choice (Claude, ChatGPT, Gemini, etc.) and submit the JSON output as a pull request. A quorum bot merges your submission once at least 3 independent runs from different models agree — that agreement is what turns a grade into a published claim.
</p>
<h2 id="the-flow">The flow</h2>
<ol>
<li>Open any protocol's page and find the <strong>Audit a dimension yourself · DEFI@home</strong> section under the Risk analysis cards. Click <strong>Copy prompt</strong> on one slice, or use <strong>Audit all</strong> to run the five risk dimensions in one LLM session.</li>
<li>The prompt includes the snapshot timestamp, the protocol's chains, GitHub repos, audit links, and the current <code>prompt_version</code> already pinned in. Paste it into Claude, ChatGPT, Gemini, or any LLM with browsing / tool use.</li>
<li>A slice prompt returns a single JSON object matching the <a href="https://github.com/guil-lambert/defipunkd/blob/main/data/schema/slice-assessment.v2.json" rel="noreferrer" target="_blank" class="link">slice-assessment</a> schema. <strong>Audit all</strong> returns a single JSON array containing exactly five normal slice objects. The schema file is named v2 but is forward-compatible: prompts emit <code>schema_version: 4</code> today and the validator accepts older supported versions too. No markdown fences outside the single JSON code block, no prose outside the JSON.</li>
<li>Click <strong>Submit run ↗</strong> on the same row. It opens GitHub's new-file interface pre-pointed at <code>data/submissions/<slug>/<slice>/</code> for one slice or <code>data/submissions/<slug>/all/</code> for the combined array. Paste your JSON, commit, open the PR.</li>
<li>CI validates your JSON against the schema. The quorum bot compares it against other submissions; overlapping grades + citations land an accepted assessment into <code>data/assessments/</code>, which the site reads on next build.</li>
</ol>
<h2 id="pinned-prompt">Why the prompt is pinned</h2>
<p>
The prompt you see on the protocol page bakes in the DeFiLlama snapshot timestamp, analysis date, and prompt version. Pinning these fixes the evidence set, so consensus across re-runs is meaningful and re-verifiable rather than dependent on any single LLM being deterministic — a contributor running the same prompt a week from now works from the same inputs.
</p>
<p>
The current pin is <code>PROMPT_VERSION = {PROMPT_VERSION}</code>. Older versions still merge, just at a lower weight (see the <a href="#weighting" class="link">weighting table</a>).
</p>
<p>
Every submission records the model used (<code>claude-opus-4-7</code>, <code>gpt-5</code>, etc.) and the commit SHA of any repo it cites. A reviewer can re-open the same PR two months later and re-verify every citation.
</p>
<h2 id="evidence">What counts as evidence</h2>
<ul>
<li>Public block explorers (Etherscan, Basescan, Arbiscan, etc.) for addresses in the protocol's contract set. Slices that touch on-chain state require at least one block-explorer URL.</li>
<li>Commits in the protocol's linked GitHub repositories, cited with a commit SHA.</li>
<li>Audit PDFs or reports linked from DeFiLlama or the protocol's docs.</li>
<li>DeFiLlama's pinned fields — but only for category / chain lists, never for risk assessment.</li>
</ul>
<p>
Anything outside that list (Twitter threads, Medium posts, Discord screenshots) does not count. The review is a PR against a git repo: if a reviewer can't independently re-fetch the URL later, the evidence isn't evidence. Adding <code>fetched_at</code> timestamps to evidence entries earns a small weight bonus, since they make later re-verification cheaper.
</p>
<h2 id="unknown">When to say <code>unknown</code></h2>
<p>
If you cannot find a signal after checking the sources above, submit <code>grade: "unknown"</code> with at least one entry in <code>unknowns[]</code> describing what you looked for, prefixed with the checklist code. "C3: couldn't find the timelock delay on the proxy admin" is useful to the next contributor; a guessed grade is a defect. Listing <code>unknowns[]</code> alongside a non-unknown grade is also valued — it earns a small bonus because it shows you completed the checklist honestly rather than papering over gaps.
</p>
<h2 id="sharing-your-chat">Sharing your chat (recommended)</h2>
<p>
The <code>chat_url</code> field in the JSON output is for a <strong>publicly-readable</strong> share URL of the conversation that produced the assessment. The quorum bot gives extra weight to submissions where the reasoning chain is independently re-readable.
</p>
<p>
<strong>Important:</strong> the default share link your LLM may offer is usually the <em>private</em> kind, requiring viewers to be logged into the same account. You need to explicitly enable public sharing. Per-platform:
</p>
<ul>
<li><strong>Claude (claude.ai):</strong> Click the share icon at the top of the conversation → toggle "Share publicly" → copy the resulting <code>https://claude.ai/share/...</code> URL.</li>
<li><strong>ChatGPT:</strong> Click "Share" → "Create public link" → copy the <code>https://chatgpt.com/share/...</code> URL. ("Anyone with the link" is the right setting.)</li>
<li><strong>Gemini (gemini.google.com):</strong> Click the three-dot menu on the response → "Share & export" → "Create public link" → copy the URL.</li>
<li><strong>Local LLMs / API:</strong> No public-share option exists. Leave <code>chat_url</code> as <code>null</code>; the submission is still accepted, just without the public-share weight bonus.</li>
</ul>
<p>
Paste the public URL into the <code>chat_url</code> field of your JSON before opening the PR. The prompt explicitly tells the LLM to leave this field as <code>null</code> — only you can produce a public-share URL, since it requires a user-side toggle the LLM cannot perform.
</p>
<h2 id="batch">Batch submissions</h2>
<p>
A slice-directory submission file can hold either a single JSON object or an <strong>array</strong> of objects with the same <code>slug</code> and <code>slice</code> but different <code>model</code> values. Useful when one contributor runs the same prompt through several models in one sitting — each entry in the array is scored independently and counts toward quorum separately. The <code>all</code> directory is reserved for <strong>Audit all</strong>: one array with exactly five objects, one for each risk slice. Naming convention for batch files: <code>models-<date>.json</code>.
</p>
<h2 id="quorum">Quorum and autorun</h2>
<p>
Three GitHub Actions run the semi-automated pipeline:
</p>
<ul>
<li><strong>validate-submission</strong> — every submission PR is schema-checked, format-cleaned (markdown-wrapped URLs auto-stripped where inner == outer), and labeled. Failures block the PR with a structured comment.</li>
<li><strong>quorum</strong> — runs daily at 06:00 UTC and after every submission push to main. Computes consensus per (slug, slice): <strong>strong</strong> = weight share ≥60% with ≥3 submissions, <strong>weak</strong> = weight share ≥50% with ≥2 submissions. On consensus, opens a PR into <code>data/assessments/</code>. The first submission for a slug lazily opens a persistent aggregation issue (one issue per protocol, not per slice) where per-slice consensus and dissent are tracked as comments.</li>
<li><strong>autorun (third voice)</strong> — runs Mondays at 04:00 UTC, picks (slug, slice) pairs stuck at 1–2 submissions ordered by TVL, and runs the same pinned prompt through the Anthropic API. Default model is <code>claude-sonnet-4-6</code>. Submissions get the model name suffixed with <code>(autorun)</code> so their weight is transparent.</li>
</ul>
<h2 id="weighting">How submissions are weighted</h2>
<p>
Every submission gets a base weight of 1.0, then adjusted by the factors below. The total weight per grade decides which grade wins quorum. The highest-weight submission for the winning grade becomes the canonical rationale in the merged assessment.
</p>
<table class="weights">
<thead>
<tr><th>Factor</th><th>Adjustment</th></tr>
</thead>
<tbody>
<tr><td>Base</td><td>1.0</td></tr>
<tr><td>Public <code>chat_url</code> share link</td><td>+0.3</td></tr>
<tr><td>Block-explorer URLs in evidence</td><td>+0.1 each, capped at +0.3</td></tr>
<tr><td><code>fetched_at</code> on evidence entries</td><td>+0.05 each, capped at +0.2</td></tr>
<tr><td>Non-empty <code>unknowns[]</code> with a non-unknown grade</td><td>+0.15</td></tr>
<tr><td>Autorun submission (model suffixed <code>(autorun)</code>)</td><td>+0.2</td></tr>
<tr><td>Older <code>prompt_version</code></td><td>−0.2 per version behind current</td></tr>
<tr><td>Snapshot mismatch (different <code>snapshot_generated_at</code> than current)</td><td>−0.1</td></tr>
<tr><td>Non-thinking model (no <code>thinking</code> in name and not opus / o-series / gemini-3-pro)</td><td>×0.2 (5× penalty)</td></tr>
<tr><td>Hallucination-prone model (claude-haiku-4-5, gemini-3-flash-preview, gpt ≤ 5.3)</td><td>×0.05 (20× penalty)</td></tr>
<tr><td>Floor</td><td>weight ≥ 0.1 (≥ 0.02 non-thinking, ≥ 0.0025 hallucination-prone)</td></tr>
</tbody>
</table>
<h2 id="reconcile">Reconcile — the master file</h2>
<p>
After quorum merges per-slice assessments, a fourth scheduled action — <strong>reconcile</strong>, Mondays at 06:00 UTC — runs Claude Sonnet over the merged assessments plus the raw submissions and writes a synthesized verdict to <code>data/master/<slug>.json</code>. This feeds the protocol detail page's narrative: findings, steel-man arguments per grade, verdict, noted dissent, and any flags. Reconcile does not re-grade — it consolidates what the quorum already decided into prose. If your submission lost quorum, your dissenting view still surfaces here.
</p>
</main>
</Base>
<style>
.page {
max-width: 800px;
margin: 0 auto;
padding: 2rem 1.5rem;
color: var(--text);
line-height: 1.6;
}
/* SideNav exposes a home link from this width up; hide the in-page back link there. */
@media (min-width: 1101px) { :global(.back-home) { display: none; } }
h1 { margin-bottom: 0.5rem; }
.lede { color: var(--text-muted); margin-top: 0; }
h2 {
color: var(--text);
border-bottom: 1px solid var(--surface-raised);
padding-bottom: 0.5rem;
margin-top: 2rem;
}
.link, a { color: var(--accent-link); }
code, strong { color: var(--text); }
table.weights {
width: 100%;
border-collapse: collapse;
margin: 1rem 0;
font-size: 0.95rem;
}
table.weights th, table.weights td {
text-align: left;
padding: 0.5rem 0.75rem;
border-bottom: 1px solid var(--surface-raised);
vertical-align: top;
}
table.weights th {
color: var(--text-muted);
font-weight: 600;
}
</style>