Add a complete single-GPU distributed-inference example that rents a vast.ai GPU, boots a worker container, and runs a prompt end-to-end over iroh/SWIM. - examples/single-gpu-inference: add the `single_gpu_inference` orchestrator binary that starts a local iroh node, waits for the remote gpu-node to register the `"inference"` SWIM name, then sends an `InferenceRequest` and prints the response - examples/single-gpu-inference: add the `gpu_node` binary that joins the cluster via `SEED_ADDR`, spawns an `InferenceActor` over `tinygrad_worker.py`, and registers the `"inference"` bridge - inference_actor: bridge swactor messaging to a Python child process via stdin/stdout JSON, with `ProcessBridge`/`RequestBridge` adapters that satisfy the single-`Incoming` actor constraint - iroh_transport: add `IrohActorTransport` that sends `WireEnvelope`s over iroh QUIC uni-streams (connection-cached against early close), plus wire encode/decode and an inbound drain helper - vastai: add a vast.ai REST client (`find_offer` with reliability/cuda/geo filters excluding CN, `create_instance`, `wait_for_running`, `destroy_instance`) parameterised by a mockable `base_url` - worker/docs/tests: ship `tinygrad_worker.py`/`echo_worker.py` (newline-JSON, `--stub`/`--model` defaulting to llama3.2:1b), a Dockerfile, Makefile, SPEC, and actor/codec/cluster/integration/vastai test suites Signed-off-by: Zachery Aaron Shores-Chmielewski <zacheryasc@gmail.com>
27 lines
918 B
Docker
27 lines
918 B
Docker
FROM nvidia/cuda:12.6.3-devel-ubuntu24.04
|
|
|
|
# Install Python 3, pip, and CUDA runtime compiler (tinygrad compiles kernels via NVRTC)
|
|
RUN apt-get update && \
|
|
apt-get install -y --no-install-recommends \
|
|
python3 \
|
|
python3-venv \
|
|
python3-pip \
|
|
ca-certificates && \
|
|
rm -rf /var/lib/apt/lists/*
|
|
|
|
# Install tinygrad and numpy
|
|
RUN python3 -m pip install --no-cache-dir --break-system-packages \
|
|
tinygrad==0.12.0 \
|
|
numpy
|
|
|
|
# Copy the gpu-node binary and tinygrad worker
|
|
# Build context should be the workspace root:
|
|
# docker build -f examples/single-gpu-inference/Dockerfile -t <tag> .
|
|
COPY examples/single-gpu-inference/target/release/gpu-node /usr/local/bin/gpu-node
|
|
COPY examples/single-gpu-inference/tinygrad_worker.py /usr/local/share/tinygrad_worker.py
|
|
|
|
# Enable CUDA backend for tinygrad
|
|
ENV CUDA=1
|
|
ENV WORKER_SCRIPT=/usr/local/share/tinygrad_worker.py
|
|
|
|
CMD ["gpu-node"]
|