Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
1 change: 1 addition & 0 deletions companies/anthropic.md
Original file line number Diff line number Diff line change
Expand Up @@ -134,6 +134,7 @@ Cataloged via [[rubenhassid-anthropic-30-term-map-2026-05]] — single secondary
- 2026-05-28: **[[claude-opus-4-8|Claude Opus 4.8 launch.]]** Recommended model for [[claude-code|Claude Code]]; Anthropic claims 69.2% on SWE-bench Pro, outperforming GPT-5.5 and Gemini 3.1 Pro on senior-level engineering / writing. Same pricing as Opus 4.7. Bundled with **Dynamic Workflows** preview for parallel subagents — orchestration primitive for multi-agent code migration at scale + self-verifying workflow patterns. Dan Shipper named as Codex→Opus migration claim at launch (marketing-coordinated; track for unsolicited follow-on). Same-day as Series H. — [[claude-opus-4-8-dynamic-workflows-2026-05-28]]
- 2026-05-28: **Dynamic Workflows primary** — full Anthropic-primary detail on the new orchestration primitive: **tens to hundreds of parallel subagents in a single session**, new **`ultracode`** Claude Code setting, **adversarial-agent verification + convergence-termination**, persistence across interruption. **Bun Zig→Rust port (Jarred Sumner) as launch case study**: 99.8% test suite passing, ~750K LOC Rust, 11 days, hundreds of agents in parallel with 2 reviewers per file (port not yet in production). Distribution: Claude Code CLI + Desktop + VS Code + API + Bedrock + Vertex + Foundry on day one. Default-on for Max/Team/API; default-off for Enterprise. Explicit *"substantially more tokens"* operator-cost warning + first-run confirmation gate. — [[anthropic-dynamic-workflows-primary-2026-05-28]]
- 2026-05-29: **Run-rate revenue hits $47B** ([[simon-willison|Willison]] surfacing). Prior wiki-tracked anchor was $44B as of May 9 ([[aakashgupta-anthropic-growth-acceleration-2026-05-09]]); **+$3B / +6.8% in three weeks**. Brief insightful framing flags the **run-rate vs ARR distinction**: *"shipping run-rate metrics instead of actual ARR is telling. They're signaling adoption velocity to investors, not profitability."* The $47B is the **underlying-fundamentals component** of the same-week capital + capability + values bundle (Series H + Opus 4.8 + ad-free positioning + Founder's Playbook + Dynamic Workflows primary + $47B run-rate = **6-event same-week strategic-coordination bundle**). Verification-pending: ARR-vs-run-rate methodology; customer concentration in the May acceleration; period covered. — [[anthropic-47b-runrate-willison-2026-05-29]]
- 2026-08-21: **IPO S-1 will flag "AI backlash" as a risk factor** (CNBC, [[dailybrief-roundup-2026-08-22]]). First captured *content* detail of the [[anthropic-s1-filing-2026-06-01|confidential draft S-1 (2026-06-01)]]: the filing is expected to name **public/regulatory backlash against AI** as a material risk factor — a market-level acknowledgment that perception + regulation headwinds are now real enough to disclose to investors. Pairs with the same-cycle *"frontier labs still won't say how they'd contain a rogue model"* ([[frontier-ai-governance]]) — the backlash the S-1 prices and the governance gap that feeds it. *(Sourced to "people familiar"; filing still non-public.)* — [[dailybrief-roundup-2026-08-22]]
- 2026-08-17: **Annualized revenue surges to $65B** (TechCrunch, [[dailybrief-roundup-2026-08-17]]). Prior wiki-tracked anchor was **$47B run-rate** (2026-05-29) — **+$18B / +38% in ~2.5 months**, the curve still bending up rather than decaying (consistent with the [[aakashgupta-anthropic-growth-acceleration-2026-05-09|$44B-in-17-months]] acceleration thesis). Brief's structural read: *"when inference economics work hard enough to swing $18B ARR in two months, the bet isn't on better weights anymore — it's on who owns the deployment layer"* — the same *value-shifts-to-the-decision-layer* thesis the [[ai-margin-collapse]] thread tracks (and which the [[stripe|Stripe/OpenRouter]] acquisition prices from the infra side the same day). Same run-rate-vs-ARR caveat as the $47B figure; TechCrunch-reported, methodology unstated. — [[dailybrief-roundup-2026-08-17]]
- 2026-05-29: **MIT CSAIL Alex Zhang's recursive-language-model research connects to Anthropic's Scaling Managed Agents + Dynamic Workflows** (via Digg). **First wiki-captured external-research → Anthropic-agent-systems direct lineage claim**; cross-confirms Dynamic Workflows as research-backed rather than purely product-engineered. Primary fetch pending. — [[dailybrief-roundup-2026-05-29]]
- 2026-05-29: **Lenny Rachitsky dream-companies survey: Anthropic #1.** Combined-platform (X + LinkedIn) survey — *"Anthropic running away with it right now."* Also: Google over OpenAI; Vercel/Linear/Every/PostHog overperforming; *"so many people want to start their own company."* **4-surface convergence on Anthropic-as-talent-magnet by late May 2026** (alongside [[brianlamanna-paraform-talent-density-2026-05|Paraform talent density #3]] + [[techcrunch-anthropic-ramp-business-customers-2026-05-13|Ramp paid-business-customer #1]] + [[aakashgupta-anthropic-growth-acceleration-2026-05-09|$44B-in-17-months revenue acceleration]]). Dream-company surveys are *lagging* indicators — Anthropic leading now means the underlying-fundamentals lead has been substantial enough for long enough to flip cultural perception. The strategic-coordination bundle now has 7 components (capital + capability + values + curriculum + revenue + workflows + talent-preference). — [[lennysan-dream-companies-survey-2026-05-27]]
Expand Down
3 changes: 2 additions & 1 deletion concepts/ai-margin-collapse.md
Original file line number Diff line number Diff line change
Expand Up @@ -25,7 +25,8 @@ It's the unit-economics lens for evaluating any AI-applied company you'd join or
- **Inference-as-commodity-infrastructure, from the hardware floor (2026-08-18)** ([[dailybrief-roundup-2026-08-18]]): *"DumpsterCluster"* (arXiv 2608.14614) serves **LLaMA-70B on a 128-GPU cluster built from datacenter-reject silicon at ~$60/GPU** — evidence that inference is drifting toward a commodity-infra problem rather than a model problem. The brief's read completes the thesis from the supply side: *"the margin game shifts from 'can we run it' to 'who owns the DC real estate and power contracts.'"* This routes into the **LPS / "Land, Power, Shell"** land-power bet ([[saas-disruption-thesis]]) and the [[ai-energy-efficiency|energy-as-binding-constraint]] thread — if secondhand silicon can serve 70B, the durable scarcity is *energized power + real estate*, not GPUs. *(arXiv; scale/throughput reproducibility unverified.)*
- **Model-routing goes mainstream as cost control (2026-08-19)** ([[dailybrief-roundup-2026-08-19]], Glean CEO via Latent Space): *"frontier model cost + open-weights popularity is driving demand for model routing."* An **enterprise-buyer-side corroboration** of the [[stripe|Stripe/OpenRouter $7B]] bet — routing is shifting from a power-user trick to **default B2B cost architecture** (the dial between Claude/GPT/Grok/open-weights). Confirms the *value-moves-to-the-decision-layer* corollary from the demand side, not just the M&A side. *(Vendor-CEO framing.)*
- **"Death of Params" — post-training as the new scaling axis (2026-08-20)** ([[dailybrief-roundup-2026-08-20]], Z.ai CEO Jie Tang on GLM 5.3, Latent Space AINews): the argument that **param-count stopped being the frontier** — *"param-count stopped mattering the moment inference became the constraint"* — and **post-training** (not pre-training scale) is where capability gains now live. Structurally relevant two ways: (1) [[glm-5-2|GLM]]'s continued cadence (**5.3** succeeds the 5.2 margin-collapse trigger) keeps the open-weight-parity pressure on; (2) if post-training is the lever, capability decouples from the giant-pretraining-run moat, which **lowers the barrier to a credible open peer** — the precondition Alderson's thesis needs. *(AINews summary; GLM-5.3 depth unconfirmed — "too vague," no model page yet.)*
- Track: independent GLM-vs-Opus benchmarks; whether frontier labs cut inference prices in response; open-weights adoption in production; whether compute spot-prices climb toward Dwarkesh's labor-anchored equilibrium; whether the routing/decision layer ([[stripe|Stripe/OpenRouter]]) captures the margin the model layer loses; whether post-training scaling (GLM 5.3 thesis) lowers the barrier to an open frontier peer.
- **"Everything's a neocloud" — the applied-AI layer collapsing into compute-reselling (2026-08-21)** ([[dailybrief-roundup-2026-08-22]], Dylan Patel / SemiAnalysis): *"every one of my AI founder friends who actually have revenue are now just transforming into neoclouds with value-add on top."* A top compute analyst naming the **convergence-to-neocloud** endpoint — if inference is commoditizing and the durable scarcity is compute + power ([[ai-energy-efficiency|DumpsterCluster]], [[stripe|routing layer]]), then the revenue-bearing move for applied-AI startups is to *become the compute layer.* Pairs with [[xai|Groq's chips→neocloud pivot]] (#247). Thread's historical rhyme: *"late-90s every tech company became a search engine — then they all died and Google remained"* — i.e. mass convergence usually precedes a brutal cull. *(X observation; directional.)*
- Track: independent GLM-vs-Opus benchmarks; whether frontier labs cut inference prices in response; open-weights adoption in production; whether compute spot-prices climb toward Dwarkesh's labor-anchored equilibrium; whether the routing/decision layer ([[stripe|Stripe/OpenRouter]]) captures the margin the model layer loses; whether post-training scaling (GLM 5.3 thesis) lowers the barrier to an open frontier peer; whether the neocloud convergence (Dylan Patel) culls the way the search-engine wave did.

## Key Papers / Posts

Expand Down
1 change: 1 addition & 0 deletions concepts/frontier-ai-governance.md
Original file line number Diff line number Diff line change
Expand Up @@ -47,6 +47,7 @@ Hassabis's lab-side blueprint pairs and contrasts with prior moves:
| **U.S. DOE — "Genesis Open Models Initiative"** (energy.gov / ANL, 2026-08-08, [[dailybrief-roundup-2026-08-08]]) | A **government-side open-models program** (DOE / Argonne). Scope still ambiguous (funding vehicle vs framework), but a state-actor entering the open-weights arena on the *pro-open* side — a counterweight to the [[us-treasury-china-ai-sanctions-threat-2026-07-21\|Treasury restriction lever]] from *within* the US government, and the public-sector complement to the industry [[open-weights-american-ai-leadership-coalition-2026-07-24\|coalition letter]]. *(DOE credibility suggests substance; details pending.)* |
| **[[dwarkesh-patel\|Dwarkesh]] — "locking in AI safety regulation now is premature"** (2026-08-08, [[dailybrief-roundup-2026-08-08]]) | From his *"Era of Continual Learning"* predictions: governance timelines are **misaligned with capability drift** — regulation frozen against today's static-model assumptions ages badly once models learn continually. The *"don't lock in prematurely"* argument sits opposite Hassabis's *"build the infrastructure in the precious window"* — the live tension over *when* to regulate, not just how. |
| **OpenAI — "frontier cyber models in more trusted hands" (Daybreak partner program)** (openai.com, 2026-08-10, [[dailybrief-roundup-2026-08-12]]) | The **constructive complement to the Astra slowdown**: rather than only braking, OpenAI ships frontier cyber capability through a **restricted approved-partner program** (authorized cybersecurity service delivery). **Governance-by-access-control** — a middle path between "release broadly" and "don't release," and a concrete answer to the "safety test is a safety risk" containment problem ([[reward-hacking]]). *(Vendor program; gate-effectiveness unproven.)* |
| **Frontier labs still won't say how they'd contain a rogue model** (TechCrunch, 2026-08-22, [[dailybrief-roundup-2026-08-22]]) | A structural accountability gap: leading labs have **few or no public plans** for containing a model that goes rogue — the containment side of the [[reward-hacking|agentic-deception]] risk the SRO/testing regime is supposed to cover, still unspecified. Sharpens the *"who's actually accountable"* question and lands the same cycle as [[anthropic|Anthropic's IPO S-1 flagging "AI backlash" as a risk]] — the backlash and the governance gap that feeds it, surfacing together. *(TechCrunch; labs' non-answers are the story.)* |
| **OpenAI revokes researchers' access to its limited cyber program** (TechCrunch, 2026-08-19, [[dailybrief-roundup-2026-08-19]]) | The **revocation flip-side** of the Daybreak "trusted-hands" program above: governance-by-access-control cuts both ways — the same gate that admits approved partners can **cut researchers off**, and researchers publicly complained. Surfaces the unresolved **who-decides + researcher-trust** problem inside any access-gated regime (and the incumbent-capture risk the pattern-watch below flags): a lab-controlled gate is only as legitimate as its appeals process. *(Researcher complaints via TechCrunch; OpenAI's rationale not captured.)* |
| **Hinton, Fei-Fei Li & Andrew Ng — "make the case for staying open" (Ai4, 2026-08-12)** ([[dailybrief-roundup-2026-08-12]], TechCrunch) | Three canonical figures publicly backing **open-source access + competition** (incl. vs China) at the Ai4 conference — a heavyweight-researcher counterweight on the *pro-open* side, distinct from the industry-coalition + government (DOE) pro-open voices already in this table. Debate-format (no concrete outcome), but the *researcher-authority* endorsement is the new element. **Meta's [[muse-glimmer\|Muse Glimmer]] (Apache-2.0, same cycle)** is the product-side of the same push. |
| **OpenAI funds 14 independent policy-research projects** ("New policy ideas for the Intelligence Age," openai.com, 2026-08, [[dailybrief-roundup-2026-08-18]]) | A frontier lab **explicitly outsourcing governance thinking** to independent researchers — the demand-side complement to its own [[openai-federal-ai-safety-framework-2026-06-03|federal-framework ask]]. Reads two ways: genuine external-input-seeking, or manufacturing independent legitimacy for rules the lab will operate under (the **regulatory-capture** risk that recurs across this table). *(Announcement; project outputs pending — "results matter more than the announcement.")* |
Expand Down
3 changes: 3 additions & 0 deletions concepts/loop-engineering.md
Original file line number Diff line number Diff line change
Expand Up @@ -83,6 +83,9 @@ Steinberger framing ([[steipete-loops-engineering-vision-md-2026-06-07]]):
- **[[dailybrief-roundup-2026-05-27|PolyArch/humanize RLCR loop]]**: Claude implements + Codex reviews independently until acceptance criteria met. **Cross-vendor agent-review loop**.
- **[[practical-systems-autonomous-company-dashclaw-2026-08-08|"A company that runs itself"]] (2026-08-08)**: an **11-step company loop** (8 agent roles) where the build step is **headless Claude Code** — `claude -p "/supergoal Build @GOAL.md … no human present, do not stop until finished" --model claude-fable-5 --max-turns 200 --permission-mode bypassPermissions` with a 120-min wall clock + 10-sec-polled kill switch (built a 44-test app in 71 min untouched). The load-bearing addition is a **governance control-plane (DashClaw)**: risk-scored action ledger where `outreach_send`/`charge_customer` **park as `pending_approval`** — the push-back primitive applied to *money- and email-touching* actions, with a hard assertion enforcing *"governance bugs should be loud."* The `--permission-mode bypassPermissions` inside the build is safe **only because** the outer loop human-gates every real-world side effect.

## Nvidia: "the harness, not the model, is the hero" — ARC-AGI-3 30%→100% (2026-08-21)
The strongest empirical anchor yet for the harness-as-capability thesis ([[dailybrief-roundup-2026-08-22]], TechCrunch on Nvidia research): a **custom harness took Claude Opus 5 from 30% → 100% on ARC-AGI-3** (instruction-free 2D reasoning games where the model must figure out the rules like a human), **with no change to the model.** The two harness ingredients named: **memory management** + a **"supervisor" boss-like component** overseeing the worker loop. 30% had been the top *model-only* result; the harness closes the entire remaining gap. This is the loop/harness layer proving it is where long-horizon capability actually lives — a raw model becomes *"something that can act on its own"* through the harness, not the weights. Operationalizes this page's thesis and the [[graph-engineering|supervisor/reviewer-node]] pattern (the supervisor is a graph move inside a loop). **Create-candidate `harness-engineering`** — the "Agent Harness Engineering vs Loop vs Graph" distinction now has a landmark result to anchor it. *(TechCrunch on an Nvidia developer-blog result; benchmark-specific — ARC-AGI-3.)*

## Verifier-discipline-first corrective (Samuel McDonald, 2026-06-15)

A **STRUCTURALLY MAJOR canonical-corrective** to the prevailing Loop Engineering discourse: Samuel McDonald publishes [[samueljmcd-loop-engineering-verifier-bottleneck-2026-06-15|"My Thoughts on Loop Engineering"]] arguing that **the verifier — not the generator — is the bottleneck**. Adds **10th canonical voice** to the cluster at the **verifier-discipline-first-tier**.
Expand Down
6 changes: 6 additions & 0 deletions concepts/model-rendered-ui.md
Original file line number Diff line number Diff line change
Expand Up @@ -53,6 +53,12 @@ The practitioner-pattern "ask Claude for HTML, not Markdown" ([[thariq-shihipar]

A [latent.space piece](https://www.latent.space/p/the-website-of-the-future) surfaced in the [[dailybrief-roundup-2026-07-06|2026-07-06 Daily Brief]] reports Adobe experimenting with **agentic sites that assemble a different page for every visitor**, generated per user intent rather than authored once. This is the *application-layer* cousin of model-rendered UI: still (for now) structured web output, but with the page composition itself model-generated at request time. Conceptually significant; no concrete working example in the coverage yet. Watch whether "page assembled per visitor" and "pixels streamed per visitor" (Flipbook) converge or stay distinct.

## "Stop making TUIs" + the agent interface (Aug 2026)

Two 2026-08 dev-practice signals ([[dailybrief-roundup-2026-08-22]]) push the same direction from the *tooling* side:
- **Thomas Ptacek — "Stop Making TUIs"** (via [[simon-willison]]): coding agents have made **native GUIs cheap enough that the marginal cost of a real UI is now below the cognitive tax of a TUI workflow.** The econ flipped — a tool built by an agent *"doesn't feel like a script anymore,"* so users actually use it. The practitioner-facing consequence of model-rendered UI: when generating a UI is nearly free, the default output stops being terminal text.
- **"The Evolution of the Agent Interface"** (Latent Space): the agent is **absorbing into model weights**, and the durable surface becomes the **interface that captures human attention**, not the control layer over the model. Same trajectory as this concept — the rendered interface, not the harness plumbing, is where the human sits.

## Related Concepts
- [[world-models]] — related idea: models that predict and generate environment state rather than tokens
- [[agentic-ai]] — agent interfaces may evolve toward model-rendered rather than structured UI
Expand Down
Loading
Loading