Extend the single-GPU example into a two-node pipeline-parallel run that splits llama3.2:1b across two rented vast.ai GPUs and closes the autoregressive decode loop over iroh.
- topology: add linear-chain helpers where each stage derives its neighbours locally from `STAGE`/`NUM_STAGES`, registering `pp-entry`/`pp-exit`/`pp-stage-{i}` SWIM names
- messages: add `StageActivation` (bf16 hidden-state hand-off carrying position/seq_len/is_prefill) and `NextToken` (sampled-token feedback with a `done` flag) that close the autoregressive loop between stage 0 and stage 1
- stage_actor: add `Stage0Actor` (tokenize -> embed_and_forward -> prefill activation; decode_step on each NextToken) and `Stage1Actor` (forward_and_sample -> NextToken back; emit InferenceResponse on EOS/max_tokens)
- vastai: fork the client and add `create_pipeline_instances` (rents one instance per stage, threading `STAGE`/`NUM_STAGES`, best-effort destroys on partial failure) and `destroy_all_instances`
- pp_tinygrad_worker.py: per-stage worker slicing `model.blk[start:end]` in stub and real (GGUF) modes, plus new `pp_gpu_node`/`pp_smoke_run` binaries and ROADMAP/SPEC/TEST_SPEC docs
- reuse: build on the single-GPU example's iroh transport and process bridge unchanged; add actor/codec/topology/integration test suites
Signed-off-by: Zachery Aaron Shores-Chmielewski <zacheryasc@gmail.com>
30 lines
701 B
TOML
30 lines
701 B
TOML
[workspace]
|
|
|
|
[package]
|
|
name = "pipeline-parallel-inference"
|
|
version = "0.1.0"
|
|
edition = "2024"
|
|
publish = false
|
|
|
|
[dependencies]
|
|
swactor = { path = "../..", features = ["transport", "serde", "std"] }
|
|
swactor-process = { path = "../../crates/process" }
|
|
serde = { version = "1", features = ["derive"] }
|
|
serde_json = "1"
|
|
reqwest = { version = "0.12", features = ["json"] }
|
|
tokio = { version = "1", features = ["full"] }
|
|
distribution = { path = "../../crates/distribution", features = ["iroh"] }
|
|
iroh = "0.96"
|
|
urlencoding = "2"
|
|
base64 = "0.22"
|
|
|
|
[[bin]]
|
|
name = "pp-gpu-node"
|
|
path = "src/bin/pp_gpu_node.rs"
|
|
|
|
[[bin]]
|
|
name = "pp-smoke-run"
|
|
path = "src/bin/pp_smoke_run.rs"
|
|
|
|
[dev-dependencies]
|
|
wiremock = "0.6"
|