vastai-reference-architecture

Implement Vast.ai reference architecture for GPU compute workflows. Use when designing ML training pipelines, structuring GPU orchestration, or establishing architecture patterns for Vast.ai applications. Trigger with phrases like "vastai architecture", "vastai design pattern", "vastai project structure", "vastai ml pipeline".

1,868 stars

byjeremylongshore

View on GitHub Installation ↓

Best use case

vastai-reference-architecture is best used when you need a repeatable AI agent workflow instead of a one-off prompt.

Teams using vastai-reference-architecture should expect a more consistent output, faster repeated execution, less prompt rewriting.

When to use this skill

You want a reusable workflow that can be run more than once with consistent structure.

When not to use this skill

You only need a quick one-off answer and do not need a reusable workflow.
You cannot install or maintain the underlying files, dependencies, or repository context.

Installation

Claude Code / Cursor / Codex

$curl -o ~/.claude/skills/vastai-reference-architecture/SKILL.md --create-dirs "https://raw.githubusercontent.com/jeremylongshore/claude-code-plugins-plus-skills/main/plugins/saas-packs/vastai-pack/skills/vastai-reference-architecture/SKILL.md"

Manual Installation

Download SKILL.md from GitHub
Place it in .claude/skills/vastai-reference-architecture/SKILL.md inside your project
Restart your AI agent — it will auto-discover the skill

How vastai-reference-architecture Compares

Feature / Agent	vastai-reference-architecture	Standard Approach
Platform Support	Not specified	Limited / Varies
Context Awareness	High	Baseline
Installation Complexity	Unknown	N/A

Frequently Asked Questions

What does this skill do?

Where can I find the source code?

You can find the source code on GitHub using the link provided at the top of the page.

Related Guides

Best AI Skills for Claude

Explore the best AI skills for Claude and Claude Code across coding, research, workflow automation, documentation, and agent operations.

ChatGPT vs Claude for Agent Skills

Compare ChatGPT and Claude for AI agent skills across coding, writing, research, and reusable workflow execution.

Cursor vs Codex for AI Workflows

Compare Cursor and Codex for AI coding workflows, repository assistance, debugging, refactoring, and reusable developer skills.

SKILL.md Source

# Vast.ai Reference Architecture

## Overview
Production architecture for GPU compute workflows on Vast.ai. Covers the three-tier pattern (orchestrator, GPU workers, artifact storage), job queue design, and fault-tolerant training pipelines.

## Prerequisites
- Vast.ai account with CLI
- Cloud storage (S3, GCS, or MinIO) for artifacts
- Understanding of ML training pipelines

## Instructions

### Architecture: Three-Tier GPU Compute

```
┌─────────────────────────────────────────────────┐
│  ORCHESTRATOR (your server / CI / cloud function) │
│  - Job queue management                          │
│  - Instance provisioning via Vast.ai API         │
│  - Status monitoring and auto-recovery           │
│  - Cost tracking and budget enforcement          │
└───────────────┬─────────────────────────────────┘
                │ Vast.ai REST API
┌───────────────▼─────────────────────────────────┐
│  GPU WORKERS (Vast.ai rented instances)          │
│  - Training / inference execution                │
│  - Checkpoint saving to cloud storage            │
│  - Health reporting back to orchestrator         │
│  - Graceful shutdown on SIGTERM (spot preemption)│
└───────────────┬─────────────────────────────────┘
                │ S3 / GCS / MinIO
┌───────────────▼─────────────────────────────────┐
│  ARTIFACT STORAGE (persistent)                   │
│  - Model checkpoints                             │
│  - Training logs and metrics                     │
│  - Dataset cache                                 │
│  - Final model artifacts                         │
└─────────────────────────────────────────────────┘
```

### Project Structure

```
ml-pipeline/
  orchestrator/
    job_queue.py         # Job definition and scheduling
    provisioner.py       # Vast.ai instance lifecycle
    monitor.py           # Status polling and auto-recovery
    cost_tracker.py      # Budget enforcement
  worker/
    Dockerfile           # GPU worker image
    train.py             # Training entry point
    checkpoint.py        # Cloud storage checkpoint manager
    health.py            # Report status back to orchestrator
  config/
    gpu_profiles.yaml    # GPU selection criteria per job type
    budgets.yaml         # Cost limits per team/project
  scripts/
    deploy.py            # CLI for launching jobs
    cost_report.py       # Spending analysis
```

### GPU Profile Configuration

```yaml
# config/gpu_profiles.yaml
profiles:
  dev-test:
    gpu_name: RTX_4090
    num_gpus: 1
    max_dph: 0.25
    reliability_min: 0.90
    max_duration_hours: 2

  training-standard:
    gpu_name: A100
    num_gpus: 1
    max_dph: 2.00
    reliability_min: 0.98
    max_duration_hours: 24

  training-distributed:
    gpu_name: H100_SXM
    num_gpus: 4
    max_dph: 4.00
    reliability_min: 0.99
    max_duration_hours: 48

  inference-batch:
    gpu_name: RTX_4090
    num_gpus: 1
    max_dph: 0.15
    reliability_min: 0.95
    max_duration_hours: 4
```

### Checkpoint Manager Pattern

```python
import boto3, os, json, time

class CheckpointManager:
    def __init__(self, bucket, prefix, interval_steps=500):
        self.s3 = boto3.client("s3")
        self.bucket = bucket
        self.prefix = prefix
        self.interval = interval_steps

    def save(self, model, optimizer, step, metrics):
        if step % self.interval != 0:
            return
        checkpoint = {
            "model_state": model.state_dict(),
            "optimizer_state": optimizer.state_dict(),
            "step": step, "metrics": metrics,
            "timestamp": time.time(),
        }
        path = f"{self.prefix}/checkpoint-{step}.pt"
        torch.save(checkpoint, f"/tmp/checkpoint-{step}.pt")
        self.s3.upload_file(f"/tmp/checkpoint-{step}.pt", self.bucket, path)

    def load_latest(self):
        objects = self.s3.list_objects_v2(Bucket=self.bucket, Prefix=self.prefix)
        if not objects.get("Contents"):
            return None
        latest = max(objects["Contents"], key=lambda o: o["LastModified"])
        self.s3.download_file(self.bucket, latest["Key"], "/tmp/latest.pt")
        return torch.load("/tmp/latest.pt")
```

## Output
- Three-tier architecture (orchestrator, GPU workers, artifact storage)
- Project structure for ML pipeline on Vast.ai
- GPU profile configuration per job type
- Checkpoint manager with cloud storage integration

## Error Handling
| Error | Cause | Solution |
|-------|-------|----------|
| Orchestrator loses track of instance | API timeout | Implement heartbeat from worker |
| Checkpoint upload fails | S3 permissions | Verify credentials on GPU instance |
| Worker can't reach orchestrator | No public IP | Use polling model (worker pulls jobs) |
| Budget exceeded | No cost controls | Implement profile-based max_duration_hours |

## Resources
- [Vast.ai REST API](https://vast.ai/developers/api)
- [PyTorch Distributed](https://pytorch.org/tutorials/intermediate/ddp_tutorial.html)

## Next Steps
For multi-environment configuration, see `vastai-multi-env-setup`.

## Examples

**Simple pipeline**: Orchestrator searches for offers matching `training-standard` profile, provisions instance, uploads data via SCP, runs training, saves checkpoints to S3, destroys instance.

**Fault-tolerant training**: Worker saves checkpoint every 500 steps to S3. On preemption, orchestrator provisions replacement and worker resumes from latest checkpoint.

Related Skills

workhuman-reference-architecture

1868

from jeremylongshore/claude-code-plugins-plus-skills

Workhuman reference architecture for employee recognition and rewards API. Use when integrating Workhuman Social Recognition, or building recognition workflows with HRIS systems. Trigger: "workhuman reference architecture".

wispr-reference-architecture

1868

from jeremylongshore/claude-code-plugins-plus-skills

Wispr Flow reference architecture for voice-to-text API integration. Use when integrating Wispr Flow dictation, WebSocket streaming, or building voice-powered applications. Trigger: "wispr reference architecture".

windsurf-reference-architecture

1868

from jeremylongshore/claude-code-plugins-plus-skills

Implement Windsurf reference architecture with optimal project structure and AI configuration. Use when designing workspace configuration for Windsurf, setting up team standards, or establishing architecture patterns that maximize Cascade effectiveness. Trigger with phrases like "windsurf architecture", "windsurf project structure", "windsurf best practices", "windsurf team setup", "optimize for cascade".

windsurf-architecture-variants

1868

from jeremylongshore/claude-code-plugins-plus-skills

Choose workspace architectures for different project scales in Windsurf. Use when deciding how to structure Windsurf workspaces for monorepos, multi-service setups, or polyglot codebases. Trigger with phrases like "windsurf workspace strategy", "windsurf monorepo", "windsurf project layout", "windsurf multi-service", "windsurf workspace size".

webflow-reference-architecture

1868

from jeremylongshore/claude-code-plugins-plus-skills

Implement Webflow reference architecture — layered project structure, client wrapper, CMS sync service, webhook handlers, and caching layer for production integrations. Trigger with phrases like "webflow architecture", "webflow project structure", "how to organize webflow", "webflow integration design", "webflow best practices".

vercel-reference-architecture

1868

from jeremylongshore/claude-code-plugins-plus-skills

Implement a Vercel reference architecture with layered project structure and best practices. Use when designing new Vercel projects, reviewing project structure, or establishing architecture standards for Vercel applications. Trigger with phrases like "vercel architecture", "vercel project structure", "vercel best practices layout", "how to organize vercel project".

vercel-architecture-variants

1868

from jeremylongshore/claude-code-plugins-plus-skills

Choose and implement Vercel architecture blueprints for different scales and use cases. Use when designing new Vercel projects, choosing between static, serverless, and edge architectures, or planning how to structure a multi-project Vercel deployment. Trigger with phrases like "vercel architecture", "vercel blueprint", "how to structure vercel", "vercel monorepo", "vercel multi-project".

veeva-reference-architecture

1868

from jeremylongshore/claude-code-plugins-plus-skills

Veeva Vault reference architecture for REST API and clinical operations. Use when working with Veeva Vault document management and CRM. Trigger: "veeva reference architecture".

vastai-webhooks-events

1868

from jeremylongshore/claude-code-plugins-plus-skills

Build event-driven workflows around Vast.ai instance lifecycle events. Use when monitoring instance status changes, implementing auto-recovery, or building event-driven GPU orchestration. Trigger with phrases like "vastai events", "vastai instance monitoring", "vastai status changes", "vastai lifecycle events".

vastai-upgrade-migration

1868

from jeremylongshore/claude-code-plugins-plus-skills

Upgrade Vast.ai CLI, migrate API versions, and handle breaking changes. Use when upgrading vastai CLI, detecting deprecations, or migrating between API versions. Trigger with phrases like "upgrade vastai", "vastai migration", "vastai breaking changes", "update vastai CLI".

vastai-security-basics

1868

from jeremylongshore/claude-code-plugins-plus-skills

Apply Vast.ai security best practices for API keys and instance access. Use when securing API keys, hardening SSH access to GPU instances, or auditing Vast.ai security configuration. Trigger with phrases like "vastai security", "vastai secrets", "secure vastai", "vastai API key security", "vastai ssh security".

vastai-sdk-patterns

1868

from jeremylongshore/claude-code-plugins-plus-skills

Apply production-ready Vast.ai SDK patterns for Python and REST API. Use when implementing Vast.ai integrations, refactoring SDK usage, or establishing coding standards for GPU cloud operations. Trigger with phrases like "vastai SDK patterns", "vastai best practices", "vastai code patterns", "idiomatic vastai".