swactor-development-history/datastore/PROTOCOL.md

471 lines
22 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# Swactor Datastore Protocol Specification
**Version:** 0.2.0 (MVP)
**Status:** Draft
## 1. Overview
The Swactor Datastore is a distributed personal file/blob storage protocol for small trusted clusters (laptop, phone, browser). It provides content-hash-first addressing with immutable content-addressed objects, replicated via a Kademlia-based metadata DHT.
### Design Principles
- **Content-hash-first addressing** — every object is identified by `blake3(blob_bytes)`. This is the primary key for all operations.
- **Immutable content-addressed objects** — content hashes are unique identifiers. There are no write conflicts by construction.
- **Names are metadata** — optional flat strings attached to objects, not keys. Multiple objects can share a name; distinguished by content hash.
- **Separation of data and metadata** — chunks are large opaque blobs; metadata is small, gossiped, and queryable.
- **Crash-safe** — fsync before acknowledge on all writes.
- **Actor-based** — three actor types coordinate via message passing within the swactor runtime.
- **Transport-agnostic** — protocol messages defined as `NetworkMessage` types; MVP uses iroh (QUIC + NAT hole-punch + encryption).
- **Pluggable storage** — `StorageBackend` trait abstracts I/O for filesystem (MVP), IndexedDB (browser), etc.
## 2. Terminology
| Term | Definition |
|------|-----------|
| **Object** | Content-addressed blob identified by `blake3(blob_bytes)`. May carry an optional human-readable name as metadata. |
| **Blob** | The raw byte content of an object. |
| **Chunk** | A fixed-size (1 MB default) slice of a blob, identified by its blake3 content hash. |
| **Manifest** | An ordered list of `ChunkRef`s describing how to reassemble an object from chunks. Stored under the object's content hash. |
| **ContentHash** | 32-byte blake3 digest. Primary identifier for blobs and DHT key. |
| **ObjectEntry** | Metadata record: content hash, optional name, owner node, tags. |
| **DHT overlay** | A Kademlia distributed hash table for object metadata, separate from the actor directory DHT. Keys are `blake3(blob_bytes)`. |
| **StorageBackend** | Trait abstracting chunk and manifest I/O for pluggable backends (filesystem, IndexedDB, etc.). |
| **Node** | A device running the swactor runtime with a datastore actor set (BlobStoreActor + MetadataActor). |
## 3. Data Model
### 3.1 ContentHash
```
ContentHash = blake3(data)[0..32] // 32 bytes
```
- **Hashing algorithm:** blake3 — 2-3x faster than sha256, tree-hashable (parallel hashing of large chunks), same 32-byte output. Supports streaming hashing for large blobs via `blake3::Hasher`.
- **Display:** first 8 bytes as hex + ellipsis (e.g. `a1b2c3d4e5f6a7b8…`).
- **XOR distance:** bitwise XOR of the 32-byte arrays, used for Kademlia routing in the metadata DHT.
### 3.2 ObjectEntry
```
ObjectEntry {
content_hash: ContentHash, // blake3(entire_blob) — primary identifier
name: Option<String>, // Optional flat string, not a path
node_id: NodeId, // Node that stores the object
tags: BTreeMap<String, String>, // User-defined key-value tags
size_bytes: u64, // Total object size
created_at: u64, // Wall-clock creation time (informational)
}
```
No conflict resolution is needed — content hashes are unique identifiers. Storing the same blob twice is a no-op (same content hash). Different blobs always have different content hashes.
### 3.3 ObjectManifest
```
ObjectManifest {
content_hash: ContentHash, // blake3(entire_blob) — NOT the hash of this manifest
chunks: Vec<ChunkRef>, // Ordered list of chunks
total_size: u64, // Total object size in bytes
chunk_size: u32, // Fixed chunk size used (e.g. 1MB)
content_type: Option<String>, // MIME type
}
ChunkRef {
hash: ContentHash, // Content hash of chunk data
offset: u64, // Byte offset in original object
size: u32, // Actual size (last chunk may be smaller)
}
```
The `content_hash` field is `blake3(entire_blob)`, computed via a streaming hasher alongside chunking. The manifest is stored and looked up using this content hash as the key.
### 3.4 Storage Backend
The `StorageBackend` trait abstracts all chunk and manifest I/O:
```rust
pub trait StorageBackend: Send {
fn write_chunk(&mut self, hash: &ContentHash, data: &[u8]) -> Result<(), io::Error>;
fn read_chunk(&self, hash: &ContentHash) -> Result<Option<Vec<u8>>, io::Error>;
fn delete_chunk(&mut self, hash: &ContentHash) -> Result<(), io::Error>;
fn has_chunk(&self, hash: &ContentHash) -> bool;
fn list_chunks(&self) -> Vec<ContentHash>;
fn write_manifest(&mut self, manifest: &ObjectManifest) -> Result<(), io::Error>;
fn read_manifest(&self, content_hash: &ContentHash) -> Result<Option<ObjectManifest>, io::Error>;
fn delete_manifest(&mut self, content_hash: &ContentHash) -> Result<(), io::Error>;
}
```
#### MVP: FilesystemBackend
Two-level directory sharding to avoid huge directories:
```
{storage_path}/
├── chunks/
│ └── {hex[0..2]}/
│ └── {hex[2..4]}/
│ └── {full_hex_hash} # Raw chunk bytes
└── manifests/
└── {hex[0..2]}/
└── {hex[2..4]}/
└── {full_hex_hash} # JSON-serialized ObjectManifest
```
Example: chunk with hash `abcdef12...` is stored at `chunks/ab/cd/abcdef12...`.
All writes are fsynced before acknowledging.
## 4. Content Addressing
### 4.1 Chunking Algorithm
Fixed-size chunking (MVP):
1. Read the input file in `chunk_size` byte blocks (default: 1,048,576 = 1 MB).
2. For each block, compute `ContentHash::of(block)`.
3. Store each chunk via the `StorageBackend`.
4. Build a `Vec<ChunkRef>` with sequential offsets.
5. Compute `content_hash = blake3(entire_blob)` using a streaming hasher fed alongside chunking.
6. Create the `ObjectManifest` with this `content_hash` and store it via the `StorageBackend` keyed by `content_hash`.
The last chunk may be smaller than `chunk_size`.
### 4.2 Reassembly
1. Read the `ObjectManifest` (by its `content_hash`).
2. For each `ChunkRef` in order, read the chunk by `hash`.
3. Concatenate all chunk data to reconstruct the original blob.
4. Verify: `blake3(reassembled) == content_hash` (optional integrity check).
## 5. Metadata DHT
> **Status:** Types and routing table logic exist in the `distribution` crate. `MetadataActor` has a dissemination queue (`enqueue`/`take_pending`) and peer-to-peer replication via `SetPeers` + `DisseminateTick`. Verified in local multi-node simulation. Full Kademlia iterative lookup (FIND_VALUE with α-parallel queries) is not yet implemented — dissemination is epidemic/gossip-style.
### 5.1 Overlay Design
The metadata DHT is a **separate Kademlia overlay** from the actor directory. It stores `ObjectEntry` records keyed by `blake3(blob_bytes)` — the content hash of the entire blob.
This separation ensures:
- Object metadata routing doesn't interfere with actor discovery.
- Different replication factors can be used (objects may be stored on fewer nodes).
- The DHT can be independently tuned for the metadata workload.
### 5.2 Key Mapping
```
DHT key = blake3(blob_bytes) = entry.content_hash
```
### 5.3 Store Flow
When storing object metadata:
1. Use `key = entry.content_hash`.
2. Find the `k` closest nodes to `key` in the metadata DHT routing table.
3. Send `StoreObjectRequest { entry }` to each of the `k` closest nodes.
### 5.4 Lookup Flow
When looking up object metadata:
1. Send `FindObjectRequest { content_hash }` to the `α` closest known nodes.
2. Each node responds with either `Found(ObjectEntry)` or `Closer(Vec<(NodeId, SocketAddr)>)`.
3. Continue querying closer nodes until convergence.
No merge step is needed — content hashes are unique identifiers.
## 6. Protocol Flows
### 6.1 PUT — Store an Object
```
User MetadataActor BlobStoreActor
│ │ │
│─── PutObject ────────────>│ │
│ │ │
│ │ (chunk the file, │
│ │ stream blake3 hash) │
│ │ │
│ │─── WriteChunk ────────>│
│ │<── ChunkStored ────────│ (repeat for each chunk)
│ │ │
│ │─── WriteManifest ─────>│
│ │<── ManifestStored ─────│
│ │ │
│ │ (create ObjectEntry, │
│ │ store in local index,│
│ │ enqueue for DHT │
│ │ dissemination) │
│ │ │
│<── PutOk {content_hash} ─│ │
```
### 6.2 GET — Retrieve an Object (Local)
```
User MetadataActor BlobStoreActor
│ │ │
│─── GetObject ────────────>│ │
│ {content_hash} │ │
│ │ (lookup content_hash │
│ │ in local index) │
│ │ │
│<── GetOk { entry, │ │
│ manifest } ──────│ │
│ │
│ (for each chunk in manifest) │
│───────────── ReadChunk ───────────────────────────>│
│<────────────── ChunkOk ───────────────────────────│
│ │
│ (reassemble chunks into original file) │
```
### 6.3 GET — Retrieve an Object (Remote)
> **Status:** `TransferActor` state machine is functional and stores received chunks to the local `BlobStoreActor`. Chunks must be fed externally (via `ChunkReceived` messages). Automatic chunk pulling from remote nodes is not yet implemented — the test harness or a future network adapter plays the "pull" role. Verified in multi-node simulation.
```
User MetadataActor TransferActor Remote BlobStore
│ │ │ │
│─ GetObject ──>│ │ │
│ {content_hash}│ │ │
│ │ (not in local │ │
│ │ index; DHT │ │
│ │ lookup) │ │
│ │ │ │
│ │─ StartDownload ───>│ │
│ │ │ │
│ │ │── GetChunkRequest ─>│
│ │ │<─ GetChunkResponse ─│
│ │ │ │
│ │ │ (repeat for each │
│ │ │ chunk) │
│ │ │ │
│ │<─ TransferComplete │ │
│ │ │ │
│<── GetOk ────│ │ (stops self) │
```
### 6.4 DELETE — Remove an Object
```
User MetadataActor
│ │
│─── DeleteObject ─────────>│
│ {content_hash} │
│ │
│ │ (remove from local )
│ │ (index, best-effort )
│ │ (notify DHT peers )
│ │
│<── DeleteOk │
│ {content_hash} ────────│
```
Chunk data is **not** immediately deleted. Unreferenced chunks are cleaned up during GC sweeps (see Section 10).
### 6.5 LIST — List Objects (Local)
```
User MetadataActor
│ │
│─── ListLocal ────────────>│
│ {name_filter} │
│ │
│ │ (filter local index )
│ │ (by name substring )
│ │
│<── ListOk { entries } ───│
```
### 6.6 LIST — List Objects (Swarm-Wide)
> **Status:** `ListSwarm` currently delegates to `ListLocal` (returns local entries only). Fan-out to peer MetadataActors is not yet wired. Swarm-wide listing is verified in simulation by querying each node and merging results in the test harness.
```
User MetadataActor Remote MetadataActors
│ │ │
│─ ListSwarm ──>│ │
│ {name_filter} │ │
│ │── ListObjectsRequest ─>│ (fan-out to all known
│ │<─ ListObjectsResponse ─│ alive nodes)
│ │ │
│ │ (merge all results, │
│ │ deduplicate by │
│ │ content hash) │
│ │ │
│<── ListOk ───│ │
```
## 7. Actor Architecture
> **Status:** All three actor types are fully implemented and tested. `DatastoreNode` coordinator routes commands to internal actors. 83+ tests across 8 test files verify single-node operations. Multi-node dissemination and cross-node transfers verified in simulation.
### 7.1 BlobStoreActor
**Responsibility:** Chunk and manifest I/O via `StorageBackend` trait.
- **State:** `Box<dyn StorageBackend>`
- **Lifecycle:** Long-lived, one per node.
- **Guarantees:** Delegates to backend; filesystem backend fsyncs before acknowledging.
**Message types:** `BlobStoreMsg` (see `messages.rs`)
### 7.2 MetadataActor
**Responsibility:** Object metadata index, DHT routing.
- **State:** Local object index (`HashMap<ContentHash, ObjectEntry>`), manifest cache, dissemination queue.
- **Lifecycle:** Long-lived, one per node.
- **Coordinates with:** BlobStoreActor (for manifest storage), remote MetadataActors (DHT operations).
**Message types:** `MetadataMsg` (see `messages.rs`)
### 7.3 TransferActor
**Responsibility:** Downloading an object (all its chunks) from a remote node.
- **State:** Manifest, pending/received chunk sets, retry counts.
- **Lifecycle:** Ephemeral — spawned per download, self-terminates on completion/failure/cancel.
- **Coordinates with:** Remote BlobStoreActor (chunk requests), local BlobStoreActor (chunk storage).
**Message types:** `TransferMsg` (see `messages.rs`)
## 8. Wire Protocol
> **Status:** All message types are defined with `NetworkMessage` implementations and stable type tags. Serialization is JSON (serde). No transport integration yet — messages are passed directly via actor addresses in simulation.
### 8.1 Message Types
All inter-node messages implement `NetworkMessage` with a stable `type_tag()`:
| Message | type_tag | Direction |
|---------|----------|-----------|
| `GetChunkRequest` | `swactor_datastore::GetChunkRequest` | requester → holder |
| `GetChunkResponse` | `swactor_datastore::GetChunkResponse` | holder → requester |
| `StoreObjectRequest` | `swactor_datastore::StoreObjectRequest` | writer → DHT nodes |
| `FindObjectRequest` | `swactor_datastore::FindObjectRequest` | reader → DHT nodes |
| `FindObjectResponse` | `swactor_datastore::FindObjectResponse` | DHT node → reader |
| `GetManifestRequest` | `swactor_datastore::GetManifestRequest` | requester → holder |
| `GetManifestResponse` | `swactor_datastore::GetManifestResponse` | holder → requester |
| `ListObjectsRequest` | `swactor_datastore::ListObjectsRequest` | requester → remote node |
| `ListObjectsResponse` | `swactor_datastore::ListObjectsResponse` | remote node → requester |
`FindObjectRequest` contains a `content_hash` field (the `blake3(blob_bytes)` key).
### 8.2 Serialization
MVP: serde JSON for all messages. Binary format (bincode or msgpack) planned for later to reduce overhead, especially for `GetChunkResponse` which carries large payloads.
### 8.3 Framing
Messages are framed over iroh QUIC streams:
- Each request/response pair uses a single bidirectional stream.
- Message format: `[4-byte length (big-endian)][JSON payload]`.
## 9. Naming
Names are **optional flat strings** — human-readable labels attached to objects as metadata.
- Names are not keys. The content hash is the only primary identifier.
- Multiple objects can share the same name. They are distinguished by content hash.
- Names are simple strings (e.g. `"vacation.jpg"`, `"backup-2024-01"`). No path hierarchy, no separators enforced.
- No conflict resolution is needed — different content always produces different content hashes.
## 10. Garbage Collection
> **Status:** Fully implemented. `MetadataActor::gc_tick()` builds a referenced chunk set from all local manifests and sends `GcUnreferenced` to `BlobStoreActor`. Verified with 6 GC-specific tests including deduplication safety, interval gating, and empty-store edge case.
### 10.1 Entry Removal
Deleting an object:
1. Remove the `ObjectEntry` from the local index.
2. Best-effort notify DHT peers to remove their replicas.
3. Remove the local manifest.
### 10.2 Chunk Reference Counting
Unreferenced chunk cleanup:
1. Build a referenced set: union of all chunk hashes from all local manifest entries.
2. Send `BlobStoreMsg::GcUnreferenced { referenced }` to the BlobStoreActor.
3. BlobStoreActor diffs its chunk list against the referenced set and deletes unreferenced chunks.
**Safety:** A chunk may be referenced by multiple objects (deduplication). Only delete when zero references remain.
### 10.3 GC Schedule
- `MetadataActor` runs `gc_tick()` every tick. Actual GC sweep happens every `gc_interval` ticks (default: 1000).
- Chunk GC is triggered less frequently (order of minutes) to avoid overhead.
## 11. Failure Modes
### 11.1 Node Offline
- **Metadata persists** in the DHT (replicated to k-closest nodes). Lookups succeed as long as any replica is alive.
- **Chunk fetches fail** if the only copy is on the offline node. The TransferActor retries once, then reports failure.
- **Recovery:** When the node comes back, its metadata is re-disseminated (anti-entropy).
### 11.2 Transfer Interrupted
- **Partial state:** Some chunks may be written to the local BlobStoreActor before the transfer fails.
- **Cleanup:** Partially downloaded chunks are not harmful — they're content-addressed and may be useful for future downloads. Unreferenced chunks are cleaned up by GC.
- **Retry:** The user can retry the GET, and only missing chunks need to be fetched (future optimization).
### 11.3 DHT Inconsistency
- **Stale metadata:** A node may serve an outdated ObjectEntry. Anti-entropy dissemination ensures replicas converge.
- Content addressing eliminates write conflicts — storing the same content hash twice is idempotent.
### 11.4 Disk Full
- `StorageBackend::write_chunk` fails with an I/O error, which is propagated back to the requester as `DatastoreResponse::Error`.
- No partial writes — fsync ensures atomicity (filesystem backend).
## 12. CLI Interface
> **Status:** Command types defined in `src/cli.rs`. Parser, dispatcher, and `[[bin]]` target not yet implemented. Planned for a follow-up session.
```
swactor-store put <local-path> [--name <label>] [--tag key=value...]
Store a local file as a distributed object.
Returns the content hash of the stored object.
--name sets an optional human-readable label.
swactor-store get <content-hash>[@<node>] [--output <local-path>]
Retrieve an object by content hash. Fetches from the specified node or discovers via DHT.
--output defaults to the object's name (if set) in the current directory.
swactor-store delete <content-hash>
Remove an object from the local index and notify DHT peers.
swactor-store list [--name <substring>] [--node <node-name>] [--all]
List objects. --name filters by name substring. --all queries all nodes (swarm-wide). Default is local.
swactor-store status
Show node info: identity, chunk count, storage usage.
```
## 13. Browser API
> **Status:** Not started. Separate milestone.
WASM-exposed functions for browser integration:
```
list_objects(name_filter: Option<String>) -> Vec<ObjectEntry>
List objects visible to this node, optionally filtered by name.
get_object(content_hash: ContentHash) -> Result<Vec<u8>, Error>
Download and reassemble an object by content hash.
put_object(data: Vec<u8>, name: Option<String>) -> Result<ContentHash, Error>
Chunk, store, and register an object. Returns the content hash.
delete_object(content_hash: ContentHash) -> Result<(), Error>
Remove an object from the local index.
get_node_status() -> NodeStatus
Node identity, chunk count, connected peers.
```
These map directly to the MetadataActor message types. The WASM runtime handles serialization across the JS/Rust boundary.