You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
Default 12 GiB VM boot takes ~60s; require <2s without weakening gates #319
This issue is the single entry point for resuming the whole task, not just the boot fix. Read this handoff and the linked phase issues. The source branches below are pushed and their exact remote SHAs were verified. Local Sprinty artifacts, clone contents and old agent process handles are not required to recover the scope.
Suggested restart instruction: “Resume #319 on its pushed integration branch. Complete the 13-phase #289/#291 plan, beginning with under-two-second production-default VM boot and unchanged authoritative gates.”
Full user goal and dependencies
Complete Capsem 0.7, polish its clients, then refactor, qualify and land Inspect. Reuse existing implementations/package identities. Work through each issue in isolation and verify its outcome before dependent work begins. These are issue boundaries, not mandatory single-PR boundaries. Service changes regenerate affected clients, including Rust; client polish remains separately reviewable.
Existing Cargo-fixture lease implementation plus two-line compiler-environment correction. Corrected native owner verdict pending; do not integrate or mark proven yet.
Separate fixture documentation correction; combining its doc blob with corrected fixture branch reproduces source-tested tree b3ed4c36efc2457964b977b4ee1436951218456b. No corrected native pass claimed.
Isolated lazy existing-session handles, parent/sibling close ownership and tests. 338 SDK tests, 98.92% branch coverage, language/generated checks and local wheel/sdist clean-install proof passed. Still version0.6.3; not a final0.7 client/publication or service/OAuth proof. Held on runtime/service dependencies.
Independent image qualification completed, but it does not establish whole runtime/release readiness: default-size Kingslanding subsequently failed as detailed below. ARM pins were preserved. No merge or release/package publication occurred.
Accepted requirements and unresolved decisions
Runtime/performance: Production-default 12GiB/4CPU fresh startup must be under2s. User authorized both guest wiping settings off and rejected THP/prefault workaround. Preserve existing authoritative gates/deadlines/default RAM. Independently prove admission, create, exec, files, restart, fork and cleanup before Inspect. New host builds/gates take its genuine locks; do not transfer old-host PID/lease ownership or infer native verdicts from a snapshot.
Service/cache: Catalog resolution and admission stay service-owned. Installed Rust BlobCache owns truthful availability, inventory, disk usage, prefetch and exact removal; clients never inspect directories. Preserve bytes required by sessions, including independent session image shares, workspace, overlay, persistent state and forks. Development repository Python cache and installed OCI cache are distinct scopes. Polling paths stay in milliseconds.
The existing reviewed-but-UNAPPROVED service proposal adds four method/path pairs: GET /images/cache, POST /images/cache/remove, POST /vms/managed/{request_id}/claim, POST /vms/managed/{request_id}/close, plus extensions to existing image/pull/create DTOs. Proposed cache states: unknown/missing/partial/ready; remove uses preview/apply with generation-bound plan token; managed creation uses UUID request identity, hashed256-bit ownership capability, finite300s lease/60s renewal and bounded24h tombstones. Close is a real owned shutdown/cleanup completion barrier. These numbers/routes are proposals requiring explicit public-contract review, not implemented or silently approved by the user's boot/commit/push instructions.
Google credentials: Use existing broker for consent, secure durable storage, refresh, disconnect and explicit per-session grants; hermetic acceptance fixtures; no tokens in task files/logs. Proposed connection/authorization and session-grant HTTP surfaces remain UNAPPROVED/unimplemented. Scope/capability catalog, real Google client/enrollment registration and consumer agy eligibility require deliberate resolution. Do not borrow a vendor OAuth identity or substitute ADC/API-key authentication for consumer agy acceptance. Google consent does not automatically grant a workload; reconnect must not restore/broaden grants.
SDKs: Friendly typed APIs for images/cache/sessions/credentials, lazy existing-session handles and managed ephemeral ownership. Centralize discovery, path validation and typed failures in SDKs instead of duplicating in Inspect. Connection/handle close must preserve unrelated/attached/named sessions; managed cleanup only owns explicitly created ephemeral resources.
Agent/UI clients: Qualify actual Claude Code and agy startup, authentication, MCP and lifecycle. Claude Desktop requires its authenticated GUI path on supported platforms; unsupported combinations stay explicit. Browser verification plus rebuilt desktop binary, supported-host tray actions and real gateway TUI flows are required.
Packages: Retain PyPI capsem, npm @capsem/sdk and @capsem/mcp; Inspect identity inspect-capsem-sandbox. Existing source builds do not establish registry publication authority/workflows or dependency resolution. SDK/MCP manifests still spell0.6.3 while Cargo is0.7.0: choose explicit client version/cohort policy, regenerate/update locks together. Establish verified repository-owned publication workflows and package/scope authority, accept exact wheel/sdist/tarballs in clean outside-workspace environments, publish only those bytes, verify registry hashes and fresh installs. npm tests that link repository node_modules are not clean-install proof. Do not repeat historical registry metadata as current ownership/version evidence. Package acceptance design is preparation only, not implemented: reuse owned/exported packages products with actual mutation/lease protection, exact source/cohort receipts and mandatory installed probes. Publish Inspect only after its own acceptance.
Inspect: Preserve all proposed #291 VM/workload and Dockerfile/Compose features and Pierre's authorship. Split modules/tests below source ceilings. Enforce regular-file/no-follow transfers and byte limits; cleanup only owned sessions; host environment/path/build-context grants are evaluator-controlled. Acceptance is unconditional in the gate. Do not replace full Dockerfile semantics with a reduced instruction interpreter. Existing guest runtime lacks a qualified Docker/BuildKit builder contract; evaluate a separately service-authorized native guest builder with bounded namespaces, cgroups, storage and network, then qualify actual multistage/build args/target/COPY/heredoc and Compose behavior. Host Docker is not implicitly authorized; do not attach its socket, use ambient env/.env/binds, edit global admission or spoof another workload's capability profile.
Workflow: Re-read pushed AGENTS.md, RELEASE.md and relevant skills. Work in isolated owned branches/worktrees; no main/other-session edits, stash, force-push, private shared-Cargo targets or pattern killing. Bound direct diagnostics and respect machine locks; just test/release entrypoints own their bounds. Focused tests, fast checks, Citadel and real black-box acceptance at runtime/image/credential crossings; exact final authoritative proof before merge/release. Sprinty remains the progress ledger: resume copied ledger only if present/correct; otherwise reconstruct the original goal/dependencies/evidence from these remote issues rather than shrinking the goal. No phase/item is complete merely because a draft or narrow suite passed.
Immediate restart sequence
Fetch the integration branch and read the detailed boot bug below plus linked phase issues.
Verify new-host KVM/vhost-vsock and environment, then obtain fresh core/Clippy/build proof for the committed flags-off source.
Measure public create through completed first exec at omitted production defaults and explicit2GiB controls, twice each; preserve stage/RSS/failure evidence and owned cleanup. Diagnose residual cost until <2s is demonstrated.
Finish whole Kingslanding and relevant authoritative runtime/installed proof. Complete the fixture native correction before carrying it; apply the separate doc correction without claiming source-only checks are native evidence.
Continue phases in the table. Public service/OAuth interfaces and unresolved enrollment/version/publication decisions need explicit contract review; draft source branches remain reviewable and separate.
GitHub reported8 Dependabot alerts on the default branch during push (7high,1moderate). Actual packages/current advisory status were not triaged in this session; run the mandatory audit and resolve real findings before final proof, without suppression/bypass assumptions.
Detailed boot bug and initial draft record
The following report was created before the local commit/push; its full patch and measured baseline remain the recovery evidence. The remote branch table above supersedes its historical “uncommitted/not pushed” status.
Default VM startup is far above the required under 2 seconds. On the observed Linux/KVM host, a fresh default-sized VM took about 60.6 seconds, mostly clearing guest RAM before userspace. This also causes the existing 30-second readiness deadline to expire during runtime acceptance.
This is a bug and a durable implementation handoff for the 0.7 runtime work in #289 / #299. Complete this before dependent client/Inspect qualification (#291; phase issues #300–#311). Do not rely on a machine clone, local Sprinty files, or an old agent's live process handles to recover the work.
Required outcome and authorized direction
The maintainer requires very fast startup for hundreds of production VMs: under 2 seconds with the existing authoritative gates passing. They explicitly authorized disabling guest heap wiping on allocation and free, and rejected the proposed huge-page/prefault workaround.
Set effective guest init_on_alloc=0 init_on_free=0 on both architectures.
Explicit zero values are essential: both config/docker/image/kernel/defconfig.arm64 and defconfig.x86_64 currently set CONFIG_INIT_ON_ALLOC_DEFAULT_ON=y. Simply removing the cmdline flags leaves allocation wiping enabled.
Preserve fresh anonymous host-mapping zero initialization, checkpoint contents on resume, other host boundaries, guest binary permissions and read-only rootfs.
Keep production defaults 12 GiB / 4 CPUs, existing readiness deadlines, and authoritative gate coverage.
Do not declare success from a 2 GiB guest, a warm resume, a kernel-only number, or an extended timeout.
No global THP/sysctl changes or eager prefaulting are part of this fix. The abandoned foundation memory-advice draft was restored out of the working tree.
Reproduction and observed evidence
Source inspected: local integration branch integration/0.7-clients-inspect, HEAD 36ce307575228a766e70f8ca8613abe56da80a15. This branch/head and the draft below were not pushed. If unavailable on a new machine, reconcile the current #289 implementation with current main in an isolated branch and review/reapply the draft against that source.
Host: 16 logical CPUs, approximately 63 GiB RAM, Intel Xeon 2.80 GHz, nested KVM with usable /dev/kvm. Build/gate queue time is separate from the VM timings below.
An owned native diagnostic completed with exit 0, booting twice for each RAM/CPU combination and retrieving dmesg successfully for every sample:
RAM
CPUs
Run 1: observed startup seconds
Run 2
2 GiB
2
11.132
11.032
2 GiB
4
11.279
11.022
12 GiB
2
60.573
60.718
12 GiB
4
60.577
60.577
These old diagnostic timings observe a process-log state through polling; they are not final public create-to-completed-exec acceptance measurements.
The baseline guest cmdline has init_on_alloc=1 init_on_free=1. Representative dmesg:
2 GiB:
[0.191824] mem auto-init: clearing system memory may take some time...
[10.067297] SLUB: ...
12 GiB / 2 CPU:
[1.419019] mem auto-init: clearing system memory may take some time...
[59.485432] SLUB: ...
The dominant delay is the eager guest RAM sweep. The existing host mapping remains lazy private anonymous memory (MAP_PRIVATE | MAP_ANONYMOUS | MAP_NORESERVE), with 4 KiB pages and no anonymous huge pages in the observed baseline. There is no before/after resident-memory comparison yet; do not invent one.
Native gate run 20261005-233445-179b66-test-kingslanding on source 36ce30757522 exited 1: 58 collected, 1 passed and 5 failed before configured fail-fast stopped the remaining cases. The debug-image/tool test passed. Egress, admission follow-up, catalog, read-only image-share and iperf cases all failed at the same vm-ready 30-second timeout before their workload probes. These failures do not establish a policy bypass or a network/ledger defect.
One preserved failure recorded kernel 62,220 ms plus guest initialization 360 ms. The agent subsequently reached readiness after the deadline and received shutdown.
Whole Citadel: 1,392 passed in 134.99 seconds, native exit 0 with warnings treated as errors.
Formatting and scoped Ruff passed.
Still unfinished: full core Rust tests, all-target Clippy, boot measurements with the changed flags, whole Kingslanding, and exact final source/installed-package proof. The compiler checks were genuinely queued behind another machine's gate; no compiler or runtime pass was inferred. Old-host jobs/leases/PIDs are not ownership or proof on a clone. Never cancel another holder or reuse copied PID identities.
Remaining work and acceptance
Recover/review the draft below against the actual current source; preserve isolated worktree ownership.
Complete core fmt, configured Clippy/tests and source guards. Rebuild the owning runtime binaries and stage the updated guest diagnostics through existing owners.
Measure a fresh public create request through the first completed successful exec, using a cached permitted image where applicable, with no readiness retries hiding failures. Include the production-default request that omits RAM/CPU overrides, plus explicit 2 GiB controls, twice each.
Record actual configured RAM/CPUs, kernel/guest/host/handshake/workload/readiness-delivery stages and resident memory. Preserve all samples, failure evidence and owned cleanup before evaluating the final under-two-second assertion.
If flags-off remains above two seconds, diagnose the residual stages and fix them; do not mark this issue complete merely because the RAM sweep disappeared.
Verify lifecycle, restart, fork, exec, file access and cleanup through existing runtime gates; run whole Kingslanding without optional live-test flags or removed assertions.
Run the final repository-authoritative checks and exact-source installed/package proof required by the enclosing plan; do not publish or merge unproven artifacts.
Existing boot/lifecycle test helpers use 2 GiB / 2 CPUs (tests/helpers/constants.py), whereas the service production defaults are 12 GiB / 4 CPUs. Preserve those existing gates and also prove the production-default target.
Residual performance leads, not proven fixes
Subtracting only the observed clearing interval from the 12 GiB baselines leaves approximately 2.507–2.541 seconds. This is an estimate, not a flags-off measurement or an achieved target. Post-clear kernel work reaches /init in approximately 140–149 ms. Guest initialization is 360–390 ms, including audit about 110 ms, overlay 50–60 ms and network 50–60 ms; Python venv creation already runs in the background.
Readiness delivery can add up to 500 ms through container-create polling; VM readiness polling caps at 50 ms. Separate actual ready timestamps from HTTP response/exec completion before changing this path.
Every baseline CPU logged that fast string operations were disabled. A read-only audit found that guest Linux 6.18.44 intel.c checks IA32_MISC_ENABLE bit 0 and disables REP_GOOD/ERMS when it is clear; KVM's reset default lacks that bit. Current Capsem cold boot in crates/capsem-core/src/hypervisor/kvm/boot_x86_64.rs does not initialize MSRs; set_msrs appears in checkpoint restore and its selected MSR list omits 0x1a0. This is a strong unmeasured hypothesis. Measure the flags-off change first, then inspect actual guest capabilities/MSR state before selecting another fix; account for checkpoint semantics if MSR setup changes.
Complete current draft for recovery
This is an uncommitted, partially verified draft, not a landed patch or release proof. Review against the recovered source rather than blindly applying to a different tree.
Seven-file patch, including the new untracked source guard
diff --git a/CHANGELOG.md b/CHANGELOG.md
index af269a936..c752a0f0f 100644
--- a/CHANGELOG.md+++ b/CHANGELOG.md@@ -100,10 +100,9 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
line-discipline autoload, no ICMP redirects or source routing, and SYN
cookies. Yama and SYN cookies are now built into the kernel for this.
-- Guest VMs now boot with `init_on_free=1`, so freed kernel heap memory is- zeroed as well as new allocations, and with `oops=panic`, so a kernel oops- ends the VM instead of leaving it running on a kernel in an unknown state.- capsem-doctor checks both are in effect.+- Guest VMs now boot with `oops=panic`, so a kernel oops ends the VM instead+ of leaving it running on a kernel in an unknown state. capsem-doctor checks+ this is in effect.
- The guest kernel is now built from allnoconfig, so it enables nothing the
defconfig does not pin. It used to fill every unpinned option with the
@@ -339,6 +338,11 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
### Fixed
+- VM startup no longer wipes guest heap memory on allocation or free. This+ removes the eager boot-time pass over all VM RAM, including the default+ 12 GiB. Fresh anonymous host memory still starts zeroed; existing checkpoint+ contents are preserved on resume.+
- Gate reports and summaries mark unfinished journals as incomplete instead
of successful, and interrupted lock waits retain their cancellation cause.
diff --git a/crates/capsem-core/src/hypervisor/kvm/fdt/tests.rs b/crates/capsem-core/src/hypervisor/kvm/fdt/tests.rs
index c197980af..c5b1000ca 100644
--- a/crates/capsem-core/src/hypervisor/kvm/fdt/tests.rs+++ b/crates/capsem-core/src/hypervisor/kvm/fdt/tests.rs@@ -79,7 +79,7 @@ fn fdt_eight_cpus() {
fn fdt_with_long_cmdline() {
let mut config = minimal_config();
config.cmdline =
- "console=hvc0 root=/dev/vda ro init_on_alloc=1 slab_nomerge page_alloc.shuffle=1 capsem.storage=virtiofs"+ "console=hvc0 root=/dev/vda ro init_on_alloc=0 init_on_free=0 slab_nomerge page_alloc.shuffle=1 capsem.storage=virtiofs"
.to_string();
let blob = build_fdt(&config).unwrap();
assert!(!blob.is_empty());
@@ -182,7 +182,7 @@ fn fdt_full_config() {
ram_base: memory::RAM_BASE,
ram_size: 4 * 1024 * 1024 * 1024,
cpu_count: 4,
- cmdline: "console=hvc0 root=/dev/vda ro init_on_alloc=1 slab_nomerge".to_string(),+ cmdline: "console=hvc0 root=/dev/vda ro init_on_alloc=0 init_on_free=0 slab_nomerge".to_string(),
initrd_start: 0x1_3000_0000,
initrd_end: 0x1_3500_0000,
virtio_devices: vec![
diff --git a/crates/capsem-core/src/hypervisor/kvm/memory/tests.rs b/crates/capsem-core/src/hypervisor/kvm/memory/tests.rs
index a14b5e6a1..496023c05 100644
--- a/crates/capsem-core/src/hypervisor/kvm/memory/tests.rs+++ b/crates/capsem-core/src/hypervisor/kvm/memory/tests.rs@@ -102,6 +102,19 @@ fn guest_memory_new_valid() {
assert!(!mem.as_ptr().is_null());
}
+#[test]+fn fresh_guest_memory_is_zeroed_without_reusing_a_previous_guests_bytes() {+ let size = 2 * 1024 * 1024;+ {+ let old = GuestMemory::new(size).unwrap();+ old.write_at(0, &vec![0xa5; size as usize]).unwrap();+ }+ let fresh = GuestMemory::new(size).unwrap();+ let mut bytes = vec![0xff; size as usize];+ fresh.read_at(0, &mut bytes).unwrap();+ assert!(bytes.iter().all(|byte| *byte == 0));+}+
#[test]
fn guest_memory_new_zero_fails() {
assert!(GuestMemory::new(0).is_err());
diff --git a/crates/capsem-core/src/vm/config.rs b/crates/capsem-core/src/vm/config.rs
index 0caf30b9c..7ed8815a1 100644
--- a/crates/capsem-core/src/vm/config.rs+++ b/crates/capsem-core/src/vm/config.rs@@ -16,8 +16,9 @@ const MAX_RAM: u64 = 16 * 1024 * 1024 * 1024; // 16 GB
/// - `loglevel=4 quiet`: kernel warnings and errors still reach the serial
/// console the test fixtures keep. At `loglevel=1` a guest whose every
/// VSOCK link ended in one millisecond left a console that said nothing.
-/// - `init_on_alloc=1 init_on_free=1`: heap pages are zeroed when allocated-/// and when freed, so neither stale nor freed data reaches a later reader.+/// - `init_on_alloc=0 init_on_free=0`: guest heap wiping is disabled, including+/// kernels with allocation wiping enabled by default. Fresh host memory is+/// already zero-initialized; guest boot must not eagerly touch all VM RAM.
/// - `slab_nomerge`: every slab cache stays separate, so an overflow in one
/// object type cannot reach objects of another merged into its cache.
/// - `page_alloc.shuffle=1`: page allocation order is randomized.
@@ -26,7 +27,7 @@ const MAX_RAM: u64 = 16 * 1024 * 1024 * 1024; // 16 GB
/// - `random.trust_cpu=1`: the CPU's RNG seeds the pool at boot.
macro_rules! guest_kernel_flags {
() => {
- "root=/dev/vda ro loglevel=4 quiet init_on_alloc=1 init_on_free=1 slab_nomerge \+ "root=/dev/vda ro loglevel=4 quiet init_on_alloc=0 init_on_free=0 slab_nomerge \
page_alloc.shuffle=1 oops=panic random.trust_cpu=1"
};
}
diff --git a/crates/capsem-core/src/vm/config/tests.rs b/crates/capsem-core/src/vm/config/tests.rs
index ce602e296..d8b52f330 100644
--- a/crates/capsem-core/src/vm/config/tests.rs+++ b/crates/capsem-core/src/vm/config/tests.rs@@ -597,8 +597,6 @@ fn kernel_cmdline_carries_every_hardening_flag() {
for flag in [
"root=/dev/vda",
"ro",
- "init_on_alloc=1",- "init_on_free=1",
"slab_nomerge",
"page_alloc.shuffle=1",
"oops=panic",
@@ -610,3 +608,14 @@ fn kernel_cmdline_carries_every_hardening_flag() {
#[cfg(not(target_arch = "x86_64"))]
assert!(flags.contains(&"console=hvc0"));
}
++#[test]+fn kernel_cmdline_disables_guest_heap_wiping_even_with_kernel_defaults_on() {+ let flags: Vec<&str> = KERNEL_CMDLINE.split_whitespace().collect();+ for name in ["init_on_alloc", "init_on_free"] {+ let prefix = format!("{name}=");+ let settings: Vec<_> = flags.iter().copied().filter(|flag| flag.starts_with(&prefix)).collect();+ let expected = format!("{name}=0");+ assert_eq!(settings, [expected.as_str()], "{KERNEL_CMDLINE}");+ }+}diff --git a/guest/artifacts/diagnostics/test_sandbox.py b/guest/artifacts/diagnostics/test_sandbox.py
index 54cea90e6..37efb6912 100644
--- a/guest/artifacts/diagnostics/test_sandbox.py+++ b/guest/artifacts/diagnostics/test_sandbox.py@@ -389,25 +389,33 @@ def test_no_kallsyms():
# -- Kernel cmdline hardening --
-def test_init_on_alloc():- """Kernel cmdline must include init_on_alloc=1 for heap zeroing."""+def test_guest_heap_wiping_is_explicitly_disabled():+ """Explicit zeros override kernel defaults without touching every RAM page."""
result = run("cat /proc/cmdline")
assert result.returncode == 0
- assert "init_on_alloc=1" in result.stdout, f"init_on_alloc=1 not in cmdline: {result.stdout}"+ flags = result.stdout.split()+ for name in ("init_on_alloc", "init_on_free"):+ values = [flag for flag in flags if flag.startswith(f"{name}=")]+ assert values == [f"{name}=0"], result.stdout-def test_heap_is_zeroed_on_alloc_and_on_free():- """init_on_alloc=1 and init_on_free=1 are in effect, not only on the cmdline.+def test_guest_heap_wiping_is_off_in_the_running_kernel():+ """Both guest heap wiping settings are off, not only on the cmdline.
The kernel reports what it applied in one boot line, e.g.
- "mem auto-init: stack:off, heap alloc:on, heap free:on".+ "mem auto-init: stack:all(zero), heap alloc:off, heap free:off".
"""
result = run("dmesg")
assert result.returncode == 0, f"dmesg failed: {result.stderr}"
lines = [line for line in result.stdout.splitlines() if "mem auto-init:" in line]
- # The kernel may add an informational line, e.g. that clearing memory at- # boot takes time; the settings line is the one naming both heap modes.- assert any("heap alloc:on" in line and "heap free:on" in line for line in lines), lines+ assert any("heap alloc:off" in line and "heap free:off" in line for line in lines), lines+++def test_guest_boot_does_not_eagerly_clear_system_memory():+ """Sparse VM memory must not be faulted in by an eager boot-time wipe."""+ result = run("dmesg")+ assert result.returncode == 0, f"dmesg failed: {result.stderr}"+ assert "clearing system memory may take some time" not in result.stdout, result.stdout
def test_an_oops_panics_the_vm():
diff --git a/tests/citadel/test_guest_boot_memory.py b/tests/citadel/test_guest_boot_memory.py
new file mode 100644
index 000000000..59c82252c
--- /dev/null+++ b/tests/citadel/test_guest_boot_memory.py@@ -0,0 +1,45 @@+"""Fast, sparse VM boot does not wipe guest heap pages eagerly."""++import re+from pathlib import Path++import pytest++ROOT = Path(__file__).resolve().parents[2]+GUEST_BOOT_MEMORY_RATIONALE = """\+Guest heap wiping is explicitly disabled for fast, sparse VM startup.+Wiping 12 GiB took a minute before userspace. Fresh anonymous host memory+already starts zeroed; guest alloc/free wiping must not touch all RAM at boot.+Explicit zeros override CONFIG_INIT_ON_ALLOC_DEFAULT_ON in both kernels.+"""+++def assert_wiping_disabled(flags: str) -> None:+ for name in ("init_on_alloc", "init_on_free"):+ values = re.findall(rf"(?<!\S){name}=([^\s\\]+)", flags)+ assert values == ["0"], f"{name} values {values}: {GUEST_BOOT_MEMORY_RATIONALE}"+++def test_the_real_kernel_flags_disable_guest_heap_wiping():+ source = (ROOT / "crates/capsem-core/src/vm/config.rs").read_text()+ macro = source.split("macro_rules! guest_kernel_flags", 1)[1].split("};", 1)[0]+ assert_wiping_disabled(macro.replace('"', " "))+++@pytest.mark.parametrize(+ "flags",+ [+ "init_on_alloc=1 init_on_free=0",+ "init_on_alloc=0 init_on_free=1",+ "init_on_free=0",+ "init_on_alloc=0",+ "init_on_alloc=0 init_on_alloc=1 init_on_free=0",+ ],+)+def test_missing_enabled_or_overridden_wiping_flags_are_refused(flags):+ with pytest.raises(AssertionError, match="Guest heap wiping"):+ assert_wiping_disabled(flags)+++def test_explicit_zero_flags_are_accepted():+ assert_wiping_disabled("root=/dev/vda ro init_on_alloc=0 init_on_free=0 slab_nomerge")
Implementation checkpoint: committed locally as 2ad763800d59a779cec0feab81c18e4464c12c5c (fix(core): disable eager guest memory wiping at boot) on integration/0.7-clients-inspect. All seven reviewed files are included, with Elie Bursztein github@elie.net as author and committer. The integration worktree is clean.
The issue body contains the complete patch for remote recovery; this commit is currently local. Source verification remains 1,392 passing Citadel guards plus focused checks. Native core tests and Clippy remain queued on the original host. Flags-off boot timing, the under-two-second target, whole Kingslanding and exact final installed/source proof remain unfinished. This issue stays open.
Resume the boot fix on integration/0.7-clients-inspect. The fixture, schema and Python SDK drafts remain separate; their dependency and verification holds remain in force. Native compiler/fixture jobs on the original host have not supplied terminal pass evidence, and under-two-second production-default boot is still unproven. Fetch the remote source and obtain new-host KVM/build/runtime proof; never treat copied PIDs or locks as ownership.
Start here after moving machines
This issue is the single entry point for resuming the whole task, not just the boot fix. Read this handoff and the linked phase issues. The source branches below are pushed and their exact remote SHAs were verified. Local Sprinty artifacts, clone contents and old agent process handles are not required to recover the scope.
Suggested restart instruction: “Resume #319 on its pushed integration branch. Complete the 13-phase #289/#291 plan, beginning with under-two-second production-default VM boot and unchanged authoritative gates.”
Full user goal and dependencies
Complete Capsem 0.7, polish its clients, then refactor, qualify and land Inspect. Reuse existing implementations/package identities. Work through each issue in isolation and verify its outcome before dependent work begins. These are issue boundaries, not mandatory single-PR boundaries. Service changes regenerate affected clients, including Rust; client polish remains separately reviewable.
Pushed source checkpoints
All five worktrees were clean at handoff. Fetch these branches; preserve their isolation and pending dependency holds.
2ad763800d59a779cec0feab81c18e4464c12c5c390c61f70dc4c41d2205fdb287c0263a19eb5d431f88166e8f1492e498c9df5514b21faba9a83224b3ed4c36efc2457964b977b4ee1436951218456b. No corrected native pass claimed.54191c8c0a959b93b144aa0feafe1c0ecb76fd72ba76171fd3aee1f5de72827ea25e924ea1a0f6e4The integration branch carries qualified published amd64 image pins:
sha256:aa823bfde4361684dc3f496bd591b2d10c64f8a2c2654204b98f57d479786b3d;sha256:5cd02c592445a0223a1284b9b3f96c3ab94ae4842ddb7e5fc2882b0304834a2b.Independent image qualification completed, but it does not establish whole runtime/release readiness: default-size Kingslanding subsequently failed as detailed below. ARM pins were preserved. No merge or release/package publication occurred.
Accepted requirements and unresolved decisions
Runtime/performance: Production-default 12GiB/4CPU fresh startup must be under2s. User authorized both guest wiping settings off and rejected THP/prefault workaround. Preserve existing authoritative gates/deadlines/default RAM. Independently prove admission, create, exec, files, restart, fork and cleanup before Inspect. New host builds/gates take its genuine locks; do not transfer old-host PID/lease ownership or infer native verdicts from a snapshot.
Service/cache: Catalog resolution and admission stay service-owned. Installed Rust BlobCache owns truthful availability, inventory, disk usage, prefetch and exact removal; clients never inspect directories. Preserve bytes required by sessions, including independent session image shares, workspace, overlay, persistent state and forks. Development repository Python cache and installed OCI cache are distinct scopes. Polling paths stay in milliseconds.
The existing reviewed-but-UNAPPROVED service proposal adds four method/path pairs:
GET /images/cache,POST /images/cache/remove,POST /vms/managed/{request_id}/claim,POST /vms/managed/{request_id}/close, plus extensions to existing image/pull/create DTOs. Proposed cache states: unknown/missing/partial/ready; remove uses preview/apply with generation-bound plan token; managed creation uses UUID request identity, hashed256-bit ownership capability, finite300s lease/60s renewal and bounded24h tombstones. Close is a real owned shutdown/cleanup completion barrier. These numbers/routes are proposals requiring explicit public-contract review, not implemented or silently approved by the user's boot/commit/push instructions.Google credentials: Use existing broker for consent, secure durable storage, refresh, disconnect and explicit per-session grants; hermetic acceptance fixtures; no tokens in task files/logs. Proposed connection/authorization and session-grant HTTP surfaces remain UNAPPROVED/unimplemented. Scope/capability catalog, real Google client/enrollment registration and consumer agy eligibility require deliberate resolution. Do not borrow a vendor OAuth identity or substitute ADC/API-key authentication for consumer agy acceptance. Google consent does not automatically grant a workload; reconnect must not restore/broaden grants.
SDKs: Friendly typed APIs for images/cache/sessions/credentials, lazy existing-session handles and managed ephemeral ownership. Centralize discovery, path validation and typed failures in SDKs instead of duplicating in Inspect. Connection/handle close must preserve unrelated/attached/named sessions; managed cleanup only owns explicitly created ephemeral resources.
Agent/UI clients: Qualify actual Claude Code and agy startup, authentication, MCP and lifecycle. Claude Desktop requires its authenticated GUI path on supported platforms; unsupported combinations stay explicit. Browser verification plus rebuilt desktop binary, supported-host tray actions and real gateway TUI flows are required.
Packages: Retain PyPI
capsem, npm@capsem/sdkand@capsem/mcp; Inspect identityinspect-capsem-sandbox. Existing source builds do not establish registry publication authority/workflows or dependency resolution. SDK/MCP manifests still spell0.6.3 while Cargo is0.7.0: choose explicit client version/cohort policy, regenerate/update locks together. Establish verified repository-owned publication workflows and package/scope authority, accept exact wheel/sdist/tarballs in clean outside-workspace environments, publish only those bytes, verify registry hashes and fresh installs. npm tests that link repository node_modules are not clean-install proof. Do not repeat historical registry metadata as current ownership/version evidence. Package acceptance design is preparation only, not implemented: reuse owned/exported packages products with actual mutation/lease protection, exact source/cohort receipts and mandatory installed probes. Publish Inspect only after its own acceptance.Inspect: Preserve all proposed #291 VM/workload and Dockerfile/Compose features and Pierre's authorship. Split modules/tests below source ceilings. Enforce regular-file/no-follow transfers and byte limits; cleanup only owned sessions; host environment/path/build-context grants are evaluator-controlled. Acceptance is unconditional in the gate. Do not replace full Dockerfile semantics with a reduced instruction interpreter. Existing guest runtime lacks a qualified Docker/BuildKit builder contract; evaluate a separately service-authorized native guest builder with bounded namespaces, cgroups, storage and network, then qualify actual multistage/build args/target/COPY/heredoc and Compose behavior. Host Docker is not implicitly authorized; do not attach its socket, use ambient env/.env/binds, edit global admission or spoof another workload's capability profile.
Workflow: Re-read pushed AGENTS.md, RELEASE.md and relevant skills. Work in isolated owned branches/worktrees; no main/other-session edits, stash, force-push, private shared-Cargo targets or pattern killing. Bound direct diagnostics and respect machine locks;
just test/release entrypoints own their bounds. Focused tests, fast checks, Citadel and real black-box acceptance at runtime/image/credential crossings; exact final authoritative proof before merge/release. Sprinty remains the progress ledger: resume copied ledger only if present/correct; otherwise reconstruct the original goal/dependencies/evidence from these remote issues rather than shrinking the goal. No phase/item is complete merely because a draft or narrow suite passed.Immediate restart sequence
GitHub reported8 Dependabot alerts on the default branch during push (7high,1moderate). Actual packages/current advisory status were not triaged in this session; run the mandatory audit and resolve real findings before final proof, without suppression/bypass assumptions.
Detailed boot bug and initial draft record
The following report was created before the local commit/push; its full patch and measured baseline remain the recovery evidence. The remote branch table above supersedes its historical “uncommitted/not pushed” status.
Default VM startup is far above the required under 2 seconds. On the observed Linux/KVM host, a fresh default-sized VM took about 60.6 seconds, mostly clearing guest RAM before userspace. This also causes the existing 30-second readiness deadline to expire during runtime acceptance.
This is a bug and a durable implementation handoff for the 0.7 runtime work in #289 / #299. Complete this before dependent client/Inspect qualification (#291; phase issues #300–#311). Do not rely on a machine clone, local Sprinty files, or an old agent's live process handles to recover the work.
Required outcome and authorized direction
The maintainer requires very fast startup for hundreds of production VMs: under 2 seconds with the existing authoritative gates passing. They explicitly authorized disabling guest heap wiping on allocation and free, and rejected the proposed huge-page/prefault workaround.
init_on_alloc=0 init_on_free=0on both architectures.config/docker/image/kernel/defconfig.arm64anddefconfig.x86_64currently setCONFIG_INIT_ON_ALLOC_DEFAULT_ON=y. Simply removing the cmdline flags leaves allocation wiping enabled.Reproduction and observed evidence
Source inspected: local integration branch
integration/0.7-clients-inspect, HEAD36ce307575228a766e70f8ca8613abe56da80a15. This branch/head and the draft below were not pushed. If unavailable on a new machine, reconcile the current #289 implementation with current main in an isolated branch and review/reapply the draft against that source.Host: 16 logical CPUs, approximately 63 GiB RAM, Intel Xeon 2.80 GHz, nested KVM with usable
/dev/kvm. Build/gate queue time is separate from the VM timings below.An owned native diagnostic completed with exit 0, booting twice for each RAM/CPU combination and retrieving dmesg successfully for every sample:
These old diagnostic timings observe a process-log state through polling; they are not final public create-to-completed-exec acceptance measurements.
The baseline guest cmdline has
init_on_alloc=1 init_on_free=1. Representative dmesg:The dominant delay is the eager guest RAM sweep. The existing host mapping remains lazy private anonymous memory (
MAP_PRIVATE | MAP_ANONYMOUS | MAP_NORESERVE), with 4 KiB pages and no anonymous huge pages in the observed baseline. There is no before/after resident-memory comparison yet; do not invent one.Native gate run
20261005-233445-179b66-test-kingslandingon source36ce30757522exited 1: 58 collected, 1 passed and 5 failed before configured fail-fast stopped the remaining cases. The debug-image/tool test passed. Egress, admission follow-up, catalog, read-only image-share and iperf cases all failed at the samevm-ready30-second timeout before their workload probes. These failures do not establish a policy bypass or a network/ledger defect.One preserved failure recorded kernel 62,220 ms plus guest initialization 360 ms. The agent subsequently reached readiness after the deadline and received shutdown.
Local evidence, if still available, is under:
cache/target/gate-runs/20261005-233445-179b66-test-kingslanding/run.jsonl;.sprinty/artifacts/egress-acceptance-failure-evidence.json;.sprinty/artifacts/kernel-resource-diagnostic/;cache/worktrees/7ef1980f25ede27a, including archived service/process/serial/guest logs.The facts and complete draft below are embedded here because those local paths are not guaranteed to survive.
Implementation state: source checked, runtime unverified
The seven-file uncommitted draft below:
Recorded source verification:
Still unfinished: full core Rust tests, all-target Clippy, boot measurements with the changed flags, whole Kingslanding, and exact final source/installed-package proof. The compiler checks were genuinely queued behind another machine's gate; no compiler or runtime pass was inferred. Old-host jobs/leases/PIDs are not ownership or proof on a clone. Never cancel another holder or reuse copied PID identities.
Remaining work and acceptance
Existing boot/lifecycle test helpers use 2 GiB / 2 CPUs (
tests/helpers/constants.py), whereas the service production defaults are 12 GiB / 4 CPUs. Preserve those existing gates and also prove the production-default target.Residual performance leads, not proven fixes
Subtracting only the observed clearing interval from the 12 GiB baselines leaves approximately 2.507–2.541 seconds. This is an estimate, not a flags-off measurement or an achieved target. Post-clear kernel work reaches /init in approximately 140–149 ms. Guest initialization is 360–390 ms, including audit about 110 ms, overlay 50–60 ms and network 50–60 ms; Python venv creation already runs in the background.
Readiness delivery can add up to 500 ms through container-create polling; VM readiness polling caps at 50 ms. Separate actual ready timestamps from HTTP response/exec completion before changing this path.
Every baseline CPU logged that fast string operations were disabled. A read-only audit found that guest Linux 6.18.44 intel.c checks
IA32_MISC_ENABLEbit 0 and disables REP_GOOD/ERMS when it is clear; KVM's reset default lacks that bit. Current Capsem cold boot incrates/capsem-core/src/hypervisor/kvm/boot_x86_64.rsdoes not initialize MSRs;set_msrsappears in checkpoint restore and its selected MSR list omits0x1a0. This is a strong unmeasured hypothesis. Measure the flags-off change first, then inspect actual guest capabilities/MSR state before selecting another fix; account for checkpoint semantics if MSR setup changes.Complete current draft for recovery
This is an uncommitted, partially verified draft, not a landed patch or release proof. Review against the recovered source rather than blindly applying to a different tree.
Seven-file patch, including the new untracked source guard