| title | FFN runtime | |||||||
|---|---|---|---|---|---|---|---|---|
| kind | module | |||||||
| status | draft | |||||||
| owners |
|
|||||||
| primary_code_paths |
|
|||||||
| related_code_paths |
|
|||||||
| depends_on |
|
|||||||
| validation_paths |
|
|||||||
| upstream_refs |
|
|||||||
| verified_platform_refs |
|
|||||||
| related_issues |
|
|||||||
| last_reviewed | 2026-08-27 |
This document is the primary design for connector-driven FFN lifecycle orchestration: daemon startup, empty KV-cache behavior, scheduler rejection, compute handoff, error propagation, and shutdown. Transport semantics belong to connector contracts; CUDA/Ascend graph, stream, profiler, native-op, and build mechanisms belong to execution platforms.
FFN consumes plugin configuration, connector transport, model-side compute, platform mechanisms, and EngineCore compatibility. No shared module may depend on the FFN worker or runner implementation.
FFN is launched as a vllm serve process, and AFD config normalization selects
its role-specific worker when worker_cls="auto". It does not serve requests.
Attention and FFN may be started in either order. Send API traffic only to
Attention.
| Platform | Worker | Model runner | Current connectors |
|---|---|---|---|
| CUDA | afd_plugin.v1.worker.AFDFFNWorker |
GPUFFNModelRunner |
P2pNcclAFDConnector |
| NPU | afd_plugin.v1.worker.npu.AFDNPUFFNWorker |
AFDNPUFFNModelRunner |
CAMP2pAFDConnector, CAMAsyncAFDConnector |
GPU and NPU runtimes use separate internal class paths. AFDNPUFFNModelRunner
inherits vLLM-Ascend NPUModelRunner directly instead of inheriting the GPU
GPUFFNModelRunner. Shared AFD semantics are kept in config, connector,
metadata, validation, and small helper functions rather than through a
cross-device inheritance chain.
CUDA launch shape:
vllm serve <model> \
--additional-config '{"afd":{"role":"ffn","connector":"P2pNcclAFDConnector","host":"127.0.0.1","port":1239,"num_attention_ranks":1,"num_ffn_ranks":1}}'Ascend launch shape:
VLLM_PLUGINS=ascend,afd vllm serve <model> \
--additional-config '{"afd":{"role":"ffn","connector":"CAMP2pAFDConnector","host":"127.0.0.1","port":1239,"num_attention_ranks":1,"num_ffn_ranks":1}}'The common FFN initialization sequence is:
vLLM worker construction
-> validate AFD config, role, connector, and selected worker class
-> validate the paired V1 or V2 deployment contract
-> initialize the matching upstream device worker
-> construct the AFD FFN model runner
-> derive the role-local AFD rank from DP/TP ranks when needed
-> load the role-aware FFN model
-> initialize an empty KV-cache surface
-> initialize the connector
-> start the background FFN loop
FFN does not own request KV blocks. Both workers return an empty KV-cache spec,
and both runners no-op KV-cache initialization. compile_or_warm_up_model()
returns 0.0; warmup and graph capture, when supported, are driven later by
connector metadata. Because the FFN EngineCore skips that upstream warmup path,
the NPU worker performs vLLM-Ascend's CPU-binding step once immediately before
starting its connector daemon when enable_cpu_binding is enabled. The
EngineCore compatibility patch keeps FFN daemon mode out of upstream
scheduler/KV-cache startup assumptions and selects the daemon busy loop. See
compatibility and patches.
The worker owns the daemon thread, shutdown event, and captured loop error. The model runner owns the model, connector, profiler, and graph cache. The connector owns transport/process-group resources.
FFN remains connector-driven when the upstream configuration selects ModelRunnerV2. It does not adopt native V2 request state or scheduler-driven execution:
- CUDA temporarily substitutes
GPUFFNModelRunnerat the native V2 construction seam insideWorker.init_device(); the existing AFD runner is still the runtime object. - Ascend always constructs
AFDNPUFFNModelRunner. When V2 is selected, its connector-driven forward enters the upstream MRV2 forward-context shape so Ascend-specific state is installed inadditional_kwargs. - Both workers run the same V2 topology and feature validator as the paired Attention role before constructing communication resources.
The V2 pair therefore requires the synchronous platform connector,
compute_gate_on_attention=false, PP/PCP/DCP size 1, configured role ranks
equal to DP x TP, static EP, a registered AFD model, and no DBO or ubatching.
CUDA uses P2pNcclAFDConnector; Ascend uses CAMP2pAFDConnector. These limits
describe the pair even though only Attention inherits a native V2 runner.
The worker selects one of two FFN step paths from the optional
connector.control_plane interface:
| Selection state | Connectors | Worker behavior |
|---|---|---|
control_plane is not None |
P2pNcclAFDConnector, CAMP2pAFDConnector |
Call control_plane.recv_dp_metadata_list(), then profile, warm, capture, replay, or execute its stage map. |
control_plane is None |
CAMAsyncAFDConnector (NPU only) |
Block directly on a connector work item; no separate DP-metadata control plane. |
The connector-driven path exists only on Ascend. GPU FFN supports
control-plane-driven connectors exclusively: GPUFFNModelRunner asserts
connector.control_plane is not None at construction, and the GPU daemon loop
raises NotImplementedError if a connector without a control plane is ever
installed.
flowchart TD
START["FFN daemon loop"] --> CONTROL_PLANE{"connector.control_plane"}
CONTROL_PLANE -->|is not None| CONTROL["Receive AFDControlPayload"]
CONTROL --> FLAGS{"Profile, warmup, capture, replay, or eager?"}
FLAGS --> GRAPH["Prepare graph state or eager context"]
GRAPH --> RECEIVE["Receive Attention payload"]
CONTROL_PLANE -->|is None| WORK["Block on AFDAsyncFFNWorkItem"]
WORK --> CONTEXT["Build minimal forward context"]
RECEIVE --> COMPUTE["Role-aware FFN compute"]
CONTEXT --> COMPUTE
COMPUTE --> SEND["Send result to Attention"]
SEND --> START
FFN initialize_from_config(...)
-> initialize empty KV-cache surface
-> initialize connector
-> start daemon thread
daemon thread:
-> connector.control_plane.recv_dp_metadata_list()
-> inspect stage metadata plus profile/warmup/capture flags
-> capture/warm matching graph, or execute FFN forward
-> synchronize the current accelerator
-> repeat
The Ascend worker checks connector.control_plane. When it is None, as for
CAM async, the worker calls execute_connector_driven_step() instead of a
control-plane receive. For each layer, the runner:
- receives a normalized
AFDAsyncFFNWorkItemfromconnector.recv_ffn_work_item(...); CAM metadata supplies the actual layer index and compact expert interval; counts.sum() supplies this chunk’s routed token count, and the connector slices tensors from operator capacity down to those counts; - builds a single-stage forward context sized to the work item's token count,
with
dp_metadata = None; - installs the work item's
AFDTransferContext.metadataasafd_metadata; - calls the role-aware FFN compute, forwarding the routed MoE compute
payloads (
group_list,dynamic_scales) from the work item'sAFDAsyncTransferStateonAFDTransferContext.states; - returns the routed outputs through
connector.send_ffn_work_item_output(...), which also handles the floating zero-routed-token placeholder required by CAM combine-send. A layer is complete only when the received interval ends at the final local expert; earlier intervals continue the receive loop.
CAM async is eager-only and does not use the FFN graph-control path; the graph cache is keyed by DP metadata, which does not exist without a control plane.
The current runner contract for one control payload is:
- Update connector state from the stage-indexed DP metadata.
- Build the minimal vLLM forward context required by model-side MoE compute.
- Iterate model layers and sorted stage ids.
- Receive an
AFDA2FTransferPayloadfor the current layer/stage, which carries the hidden states plus anAFDTransferContext(transfer metadata and the backendAFDTransferState). - Install stage DP metadata and transfer metadata at
ForwardContext.additional_kwargs["afd_metadata"]. - Set vLLM's current MoE layer index when the upstream context exposes the layer list.
- Compute FFN output and send it to Attention with the same layer/stage
identity, passing the same
AFDTransferContextback intosend_ffn_output()so the connector can reuse its receive-time state.
On Ascend, is_profile is also forwarded into the minimal forward context so
vLLM-Ascend preserves its balanced dummy-MoE and communication state during
the split memory-profile run. CUDA carries the field in the common envelope
but does not need a separate FFN profile-context branch.
On CUDA, GPUFFNModelRunner is a plugin-owned minimal runner. It invokes
model.compute_ffn_output(hidden_states, layer_idx) when provided and
otherwise passes hidden states through; production AFD model paths are
expected to provide the compute method. On Ascend, AFDNPUFFNModelRunner
directly calls the role-aware model. The control-plane-driven path (CAMP2P)
does not thread per-transfer payload fields through compute_ffn_output:
CAMP2P's CAMP2PTransferState (operator sizes, the A2E-returned
atten_batch_size, x_active_mask, and HCCL endpoint name) rides on
AFDTransferContext.states and the forward context between the receive and
send phases.
Ascend translates AFD-level token counts to the DP-level token-count vector expected by the upstream forward context. When TP expands Attention rank counts beyond DP metadata width, token counts are replicated per TP rank and then aggregated for each FFN rank.
The FFN process uses vLLM's model loader, while plugin-owned model integration constructs FFN MLP/expert components plus shared components needed by the upstream lifecycle, without Attention modules. FFN runners do not sample tokens. GPU LoRA mutation methods return unsupported/no-op results; the NPU runner rejects sampling explicitly.
Detailed model construction and weight ownership remain in model integration.
For connectors with a control plane, the graph cache key is derived from each stage's token counts. Ascend also includes A/F topology when it must aggregate Attention counts to FFN counts. The shared dispatch states are eager, warmup, capture, and replay:
- warmup runs FFN compute after applying control state;
- capture applies connector control state before entering the formal device graph so control-plane work is not replayed;
- replay runs an existing graph for the matching key;
- a missing graph falls back to eager execution.
The Attention control payload is the source of warmup/capture flags. Device graph objects, pools, supported modes, and exact cache behavior are specified by execution platforms.
execute_model() on either worker raises immediately if the native scheduler
tries to drive FFN work. The runner also rejects a normal execution call that
lacks its connector metadata/work item.
The daemon catches failures, stores the original exception, logs it, and makes
the error visible through raise_ffn_loop_error_if_any(). Re-entering startup
or completing shutdown therefore surfaces an earlier background failure.
Shutdown ordering is:
signal daemon event
-> runner stops profiler and closes connector
-> join daemon thread with a bounded timeout
-> surface stored loop error
-> delegate remaining worker/runner shutdown upstream
Connector receive calls may be blocking; the connector's close() behavior
is responsible for releasing its communication resources so shutdown can
complete.
The following RFC candidate remains non-normative while this document is draft:
ROLE-INV-001(FFN part): FFN remains connector-driven and rejects scheduler execution.
The optional control-plane and connector work-item surfaces, and the current
AFDTransferContext/AFDTransferState transfer payload shape, are not stable
contracts.
Changes must be compared with the pinned vLLM worker and EngineCore behavior and, for Ascend, with the tested runtime evidence. Run the FFN runner, EngineCore patch, NPU runtime, and serving tests listed in the metadata. Control-plane selection or work-item changes also require connector and CAM async model E2E coverage.
Current shared limits are the supported vLLM release, connector-driven FFN, and registered role-aware model integrations. V1 native DBO accepts exactly two ubatches. A V2-paired FFN rejects DBO/ubatching and uses the same AFD FFN runner with a V2-compatible construction and forward-context seam. CAM async instead uses eager connector work items and may enable its distinct two-stage MoE pipeline. Platform-specific limits are centralized in execution platforms.
Issue #107 completed the optional control-plane split. Connector metadata ownership, transfer state separation, and the public shape of the connector work-item interface remain open in #88 and #105.