|
117 | 117 |
|
118 | 118 | <!-- Closing Statement --> |
119 | 119 | <div class="closing-statement"> |
120 | | - <p>"This is the kind of work I want to do: translating technical reality for the audiences who need to understand it, calibrated to what each one actually needs."</p> |
| 120 | + <p>This is the kind of work I want to do: translating technical reality for the audiences who need to understand it, calibrated to what each one actually needs.</p> |
| 121 | + </div> |
| 122 | + |
| 123 | + <!-- Executive Communication Section --> |
| 124 | + <div class="exec-section"> |
| 125 | + <div class="exec-intro"> |
| 126 | + <p>And when the CEO needs a briefing before the press calls? That's a different document entirely.</p> |
| 127 | + </div> |
| 128 | + <div class="exec-memo"> |
| 129 | + <div class="exec-memo-header"> |
| 130 | + <span class="exec-memo-label">Re: Soul Document Extraction</span> |
| 131 | + </div> |
| 132 | + <div class="exec-memo-body"> |
| 133 | + <p>A researcher publicly extracted and published a near-complete version of our internal values document for Claude 4.5 Opus. Coverage is spreading in technical and policy communities.</p> |
| 134 | + |
| 135 | + <p><strong>Our position:</strong> This is not a crisis. The document reflects our stated views and published research. We plan to release the full version ourselves, which we can now frame as transparency rather than reaction.</p> |
| 136 | + |
| 137 | + <p><strong>Key risk:</strong> Mischaracterization. The document will be quoted selectively. Lines about revenue, about shaping Claude's behavior, about uncertainty in our methods—all can be spun negatively without context.</p> |
| 138 | + |
| 139 | + <p><strong>Recommendation:</strong> Get ahead of the narrative. Publish the full document with a brief framing note from Amanda. Offer briefings to key journalists and Hill contacts before coverage hardens. Position this as: we wrote down what we believe, we trained Claude on it, and we're willing to show our work.</p> |
| 140 | + |
| 141 | + <p><strong>Open question for discussion:</strong> Do we want to accelerate our planned transparency efforts around character training more broadly? This moment creates an opening.</p> |
| 142 | + </div> |
| 143 | + </div> |
121 | 144 | </div> |
122 | 145 | </div> |
123 | 146 | </div> |
|
371 | 394 |
|
372 | 395 | .part-transition p { |
373 | 396 | font-family: 'Playfair Display', Georgia, serif; |
374 | | - font-style: italic; |
375 | 397 | font-size: 1.15rem; |
376 | 398 | color: var(--color-cream); |
377 | 399 | max-width: 500px; |
|
460 | 482 |
|
461 | 483 | .closing-statement p { |
462 | 484 | font-family: 'Playfair Display', Georgia, serif; |
463 | | - font-style: italic; |
464 | 485 | font-size: 1.15rem; |
465 | 486 | color: var(--color-gold); |
466 | 487 | max-width: 600px; |
467 | 488 | margin: 0 auto; |
468 | 489 | } |
469 | 490 |
|
| 491 | + /* Executive Communication Section */ |
| 492 | + .exec-section { |
| 493 | + margin-top: 2.5rem; |
| 494 | + padding-top: 2rem; |
| 495 | + border-top: 1px solid var(--color-medium); |
| 496 | + } |
| 497 | + |
| 498 | + .exec-intro { |
| 499 | + text-align: center; |
| 500 | + margin-bottom: 1.5rem; |
| 501 | + } |
| 502 | + |
| 503 | + .exec-intro p { |
| 504 | + font-family: 'Playfair Display', Georgia, serif; |
| 505 | + font-size: 1.15rem; |
| 506 | + color: var(--color-gold); |
| 507 | + } |
| 508 | + |
| 509 | + .exec-memo { |
| 510 | + background: rgba(61, 74, 42, 0.6); |
| 511 | + border: 1px solid var(--color-medium); |
| 512 | + border-radius: 8px; |
| 513 | + overflow: hidden; |
| 514 | + } |
| 515 | + |
| 516 | + .exec-memo-header { |
| 517 | + background: rgba(107, 124, 76, 0.3); |
| 518 | + padding: 0.75rem 1.5rem; |
| 519 | + border-bottom: 1px solid var(--color-medium); |
| 520 | + } |
| 521 | + |
| 522 | + .exec-memo-label { |
| 523 | + font-family: 'Source Sans 3', sans-serif; |
| 524 | + font-weight: 600; |
| 525 | + font-size: 1rem; |
| 526 | + color: var(--color-cream); |
| 527 | + } |
| 528 | + |
| 529 | + .exec-memo-body { |
| 530 | + padding: 1.5rem; |
| 531 | + } |
| 532 | + |
| 533 | + .exec-memo-body p { |
| 534 | + font-size: 1rem; |
| 535 | + line-height: 1.7; |
| 536 | + margin-bottom: 1rem; |
| 537 | + color: var(--color-cream); |
| 538 | + } |
| 539 | + |
| 540 | + .exec-memo-body p:last-child { |
| 541 | + margin-bottom: 0; |
| 542 | + } |
| 543 | + |
| 544 | + .exec-memo-body strong { |
| 545 | + color: var(--color-gold); |
| 546 | + font-weight: 600; |
| 547 | + } |
| 548 | + |
470 | 549 | /* Interactive section fade-in */ |
471 | 550 | .interactive-section { |
472 | 551 | opacity: 0; |
|
520 | 599 | .part-label { |
521 | 600 | font-size: 1rem; |
522 | 601 | } |
| 602 | + |
| 603 | + .exec-memo-body { |
| 604 | + padding: 1.25rem; |
| 605 | + } |
| 606 | + |
| 607 | + .exec-memo-body p { |
| 608 | + font-size: 0.95rem; |
| 609 | + } |
523 | 610 | } |
524 | 611 |
|
525 | 612 | /* Screen reader only utility */ |
|
587 | 674 | 3: "We want to be direct about one limitation: we do not yet fully understand how training on a document like this translates into model behavior in practice. We can verify that Claude learned the content. We cannot yet prove that stated values reliably predict actions. This is an active area of research. Anthropic welcomes the opportunity to brief interested offices on our alignment methodology, our approach to transparency, and the open technical questions that remain." |
588 | 675 | }, |
589 | 676 | researcher: { |
590 | | - 1: "[Placeholder: Anthropic's response to AI safety researchers - paragraph 1]", |
591 | | - 2: "[Placeholder: Anthropic's response to AI safety researchers - paragraph 2]", |
592 | | - 3: "[Placeholder: Anthropic's response to AI safety researchers - paragraph 3]" |
| 677 | + 1: "We can confirm the extracted document reflects a real artifact used in Claude's training. Amanda Askell, who led this work, has stated publicly that it was used in supervised learning. We plan to release the full version along with additional methodological context.", |
| 678 | + 2: "To be direct about what this is and isn't: training a model on a values document is not the same as proving alignment. We can verify that Claude learned the content. What we cannot yet demonstrate is that stated values are causally linked to downstream behavior.", |
| 679 | + 3: "This work connects to our published research on Constitutional AI and character training, but it is not a substitute for mechanistic interpretability. Whether \"Claude can recite these values\" means \"Claude acts on these values under pressure\" remains an open empirical question. We should also acknowledge the obvious: we want this approach to work, which makes us imperfect judges of whether it does. This document is one piece of a broader alignment approach that includes RLHF, constitutional AI methods, interpretability research, and red-teaming. It is not a solution. It is one legible artifact in a system we are still working to understand. If you are working on related evaluation methods, we would genuinely like to hear about it." |
593 | 680 | } |
594 | 681 | }; |
595 | 682 |
|
|
0 commit comments