Toolkit for interaction with vastai instances. Eventual framework for low price spot distributed training.
1.6 KiB
1.6 KiB
dtrain
CLI for distributed training on vast.ai. Provision GPU nodes, deploy scripts, and monitor training with simple commands.
Why
Training on multiple GPUs across multiple machines shouldn't require clicking through web UIs or juggling SSH sessions. dtrain wraps the vast.ai CLI to provide a streamlined workflow: rent nodes, deploy code, monitor progress, tear down.
Requirements
- Python 3
- vast.ai CLI configured with your API key
- SSH key registered with vast.ai
Usage
# Find available GPUs (sorted by price)
dtrain search --vram 8
# Rent 2 cheapest instances with 8GB+ VRAM
dtrain rent -n 2 --vram 8
# Check instance status (wait for SSH to be ready)
dtrain status
# Deploy and run your training script
dtrain deploy --script train.py
# Monitor training progress
dtrain ps
# Run arbitrary commands on all nodes
dtrain run "nvidia-smi"
dtrain run "tail -20 /workspace/train.log"
# Tear down when done
dtrain destroy
Commands
| Command | Description |
|---|---|
search |
Find GPU offers, sorted by price |
rent |
Rent instances by ID or auto-select N cheapest |
status |
Show instance info and SSH commands |
ps |
Show training status (log age + last output) |
deploy |
Copy script to nodes and run in background |
run |
Run command on all nodes |
destroy |
Tear down all instances |
How it works
- Scripts are deployed to
/workspace/and run vianohupwith output to<script>.log dtrain psreads the most recent log file to show what's happening- All state comes from vast.ai - dtrain doesn't store anything locally