castai-performance-tuning

Optimize CAST AI autoscaler performance, node provisioning speed, and API efficiency. Use when nodes take too long to provision, autoscaler is not reacting fast enough, or optimizing API call patterns for multi-cluster dashboards. Trigger with phrases like "cast ai performance", "cast ai slow", "cast ai node provisioning", "cast ai autoscaler speed".

25 stars

byComeOnOliver

View on GitHub Installation ↓

Best use case

castai-performance-tuning is best used when you need a repeatable AI agent workflow instead of a one-off prompt.

Teams using castai-performance-tuning should expect a more consistent output, faster repeated execution, less prompt rewriting.

When to use this skill

You want a reusable workflow that can be run more than once with consistent structure.

When not to use this skill

You only need a quick one-off answer and do not need a reusable workflow.
You cannot install or maintain the underlying files, dependencies, or repository context.

Installation

Claude Code / Cursor / Codex

$curl -o ~/.claude/skills/castai-performance-tuning/SKILL.md --create-dirs "https://raw.githubusercontent.com/ComeOnOliver/skillshub/main/skills/jeremylongshore/claude-code-plugins-plus-skills/castai-performance-tuning/SKILL.md"

Manual Installation

Download SKILL.md from GitHub
Place it in .claude/skills/castai-performance-tuning/SKILL.md inside your project
Restart your AI agent — it will auto-discover the skill

How castai-performance-tuning Compares

Feature / Agent	castai-performance-tuning	Standard Approach
Platform Support	Not specified	Limited / Varies
Context Awareness	High	Baseline
Installation Complexity	Unknown	N/A

Frequently Asked Questions

What does this skill do?

Where can I find the source code?

You can find the source code on GitHub using the link provided at the top of the page.

SKILL.md Source

# CAST AI Performance Tuning

## Overview

Tune CAST AI for faster node provisioning, more responsive autoscaling, and efficient API usage. Covers headroom configuration, instance family selection, and API caching for multi-cluster dashboards.

## Prerequisites

- CAST AI Phase 2 (full automation) enabled
- Understanding of workload scheduling patterns
- Access to autoscaler policy configuration

## Instructions

### Step 1: Optimize Node Provisioning Speed

```bash
# Configure headroom for proactive scaling (avoids waiting for pending pods)
curl -X PUT -H "X-API-Key: ${CASTAI_API_KEY}" \
  -H "Content-Type: application/json" \
  "https://api.cast.ai/v1/kubernetes/clusters/${CASTAI_CLUSTER_ID}/policies" \
  -d '{
    "enabled": true,
    "unschedulablePods": {
      "enabled": true,
      "headroom": {
        "enabled": true,
        "cpuPercentage": 15,
        "memoryPercentage": 15
      }
    }
  }'
```

Headroom pre-provisions spare capacity so pods schedule immediately instead of waiting 2-5 minutes for new nodes.

### Step 2: Instance Family Optimization

```hcl
# Terraform: Prefer instance families with fast launch times
resource "castai_node_template" "fast_launch" {
  cluster_id = castai_eks_cluster.this.id
  name       = "fast-launch-workers"

  constraints {
    spot                  = true
    use_spot_fallbacks    = true
    fallback_restore_rate_seconds = 300

    # Newer instance types launch faster and have better availability
    instance_families {
      include = ["m6i", "m7i", "c6i", "c7i", "r6i", "r7i"]
    }

    # Enable spot diversity for faster provisioning
    spot_diversity_price_increase_limit_percent = 25

    architectures = ["amd64"]
  }
}
```

### Step 3: Evictor Tuning for Faster Consolidation

```bash
# Reduce empty node delay for dev/staging (faster downscale)
helm upgrade castai-evictor castai-helm/castai-evictor \
  -n castai-agent \
  --reuse-values \
  --set evictor.aggressiveMode=true \
  --set evictor.cycleInterval=120

# For production, use non-aggressive with longer intervals
# --set evictor.aggressiveMode=false
# --set evictor.cycleInterval=600
```

### Step 4: API Performance for Multi-Cluster Dashboards

```typescript
import { LRUCache } from "lru-cache";

const cache = new LRUCache<string, unknown>({ max: 100, ttl: 60_000 });

interface ClusterSummary {
  id: string;
  name: string;
  savings: number;
  savingsPercent: number;
  nodeCount: number;
  spotPercent: number;
}

async function getClusterSummary(clusterId: string): Promise<ClusterSummary> {
  const cacheKey = `summary:${clusterId}`;
  const cached = cache.get(cacheKey) as ClusterSummary | undefined;
  if (cached) return cached;

  const [cluster, savings, nodes] = await Promise.all([
    castaiGet(`/v1/kubernetes/external-clusters/${clusterId}`),
    castaiGet(`/v1/kubernetes/clusters/${clusterId}/savings`),
    castaiGet(`/v1/kubernetes/external-clusters/${clusterId}/nodes`),
  ]);

  const spotNodes = nodes.items.filter(
    (n: { lifecycle: string }) => n.lifecycle === "spot"
  ).length;

  const summary: ClusterSummary = {
    id: clusterId,
    name: cluster.name,
    savings: savings.monthlySavings,
    savingsPercent: savings.savingsPercentage,
    nodeCount: nodes.items.length,
    spotPercent: nodes.items.length > 0
      ? (spotNodes / nodes.items.length) * 100
      : 0,
  };

  cache.set(cacheKey, summary);
  return summary;
}

// Aggregate across all clusters
async function getDashboardData(
  clusterIds: string[]
): Promise<ClusterSummary[]> {
  return Promise.all(clusterIds.map(getClusterSummary));
}
```

### Step 5: Workload Autoscaler Tuning

```yaml
# Faster resource adjustment with shorter cooldown
# (use with caution in production)
metadata:
  annotations:
    autoscaling.cast.ai/cpu-headroom: "10"     # Lower headroom = tighter fit
    autoscaling.cast.ai/memory-headroom: "15"
    autoscaling.cast.ai/apply-type: "immediate" # Apply without waiting
```

## Performance Benchmarks

| Metric | Default | Tuned |
|--------|---------|-------|
| Node provision time | 3-5 min | 1-3 min (with headroom) |
| Empty node removal | 5 min | 2 min (aggressive evictor) |
| Workload resize | 5 min cooldown | Immediate |
| API response (cached) | 200ms | <5ms |

## Error Handling

| Issue | Cause | Solution |
|-------|-------|----------|
| Headroom over-provisioning | Percentage too high | Reduce to 5-10% |
| Aggressive evictor causing disruptions | PDB not set | Add PodDisruptionBudgets |
| Cache stale data | TTL too long | Reduce cache TTL to 30s |
| Instance type unavailable | Too narrow constraints | Add more instance families |

## Resources

- [Autoscaler Settings](https://docs.cast.ai/docs/autoscaler-settings)
- [Workload Autoscaler Annotations](https://docs.cast.ai/docs/workload-autoscaler-annotations-reference)
- [Node Configuration](https://docs.cast.ai/docs/node-configuration)

## Next Steps

For cost optimization strategies, see `castai-cost-tuning`.

Related Skills

validating-performance-budgets

from ComeOnOliver/skillshub

Validate application performance against defined budgets to identify regressions early. Use when checking page load times, bundle sizes, or API response times against thresholds. Trigger with phrases like "validate performance budget", "check performance metrics", or "detect performance regression".

tuning-hyperparameters

from ComeOnOliver/skillshub

Optimize machine learning model hyperparameters using grid search, random search, or Bayesian optimization. Finds best parameter configurations to maximize performance. Use when asked to "tune hyperparameters" or "optimize model". Trigger with relevant phrases based on skill purpose.

analyzing-query-performance

from ComeOnOliver/skillshub

This skill enables Claude to analyze and optimize database query performance. It activates when the user discusses query performance issues, provides an EXPLAIN plan, or asks for optimization recommendations. The skill leverages the query-performance-analyzer plugin to interpret EXPLAIN plans, identify performance bottlenecks (e.g., slow queries, missing indexes), and suggest specific optimization strategies. It is useful for improving database query execution speed and resource utilization.

providing-performance-optimization-advice

from ComeOnOliver/skillshub

Provide comprehensive prioritized performance optimization recommendations for frontend, backend, and infrastructure. Use when analyzing bottlenecks or seeking improvement strategies. Trigger with phrases like "optimize performance", "improve speed", or "performance recommendations".

profiling-application-performance

from ComeOnOliver/skillshub

Execute this skill enables AI assistant to profile application performance, analyzing cpu usage, memory consumption, and execution time. it is triggered when the user requests performance analysis, bottleneck identification, or optimization recommendations. the... Use when optimizing performance. Trigger with phrases like 'optimize', 'performance', or 'speed up'.

performance-testing

from ComeOnOliver/skillshub

This skill enables Claude to design, execute, and analyze performance tests using the performance-test-suite plugin. It is activated when the user requests load testing, stress testing, spike testing, or endurance testing, and when discussing performance metrics such as response time, throughput, and error rates. It identifies performance bottlenecks related to CPU, memory, database, or network issues. The plugin provides comprehensive reporting, including percentiles, graphs, and recommendations.

detecting-performance-regressions

from ComeOnOliver/skillshub

This skill enables Claude to automatically detect performance regressions in a CI/CD pipeline. It analyzes performance metrics, such as response time and throughput, and compares them against baselines or thresholds. Use this skill when the user requests to "detect performance regressions", "analyze performance metrics for regressions", or "find performance degradation" in a CI/CD environment. The skill is also triggered when the user mentions "baseline comparison", "statistical significance analysis", or "performance budget violations". It helps identify and report performance issues early in the development cycle.

performance-lighthouse-runner

from ComeOnOliver/skillshub

Performance Lighthouse Runner - Auto-activating skill for Frontend Development. Triggers on: performance lighthouse runner, performance lighthouse runner Part of the Frontend Development skill category.

performance-baseline-creator

from ComeOnOliver/skillshub

Performance Baseline Creator - Auto-activating skill for Performance Testing. Triggers on: performance baseline creator, performance baseline creator Part of the Performance Testing skill category.

optimizing-cache-performance

from ComeOnOliver/skillshub

Execute this skill enables AI assistant to analyze and improve application caching strategies. it optimizes cache hit rates, ttl configurations, cache key design, and invalidation strategies. use this skill when the user requests to "optimize cache performance"... Use when optimizing performance. Trigger with phrases like 'optimize', 'performance', or 'speed up'.

aggregating-performance-metrics

from ComeOnOliver/skillshub

This skill enables Claude to aggregate and centralize performance metrics from various sources. It is used when the user needs to consolidate metrics from applications, systems, databases, caches, queues, and external services into a central location for monitoring and analysis. The skill is triggered by requests to "aggregate metrics", "centralize performance metrics", or similar phrases related to metrics aggregation and monitoring. It facilitates designing a metrics taxonomy, choosing appropriate aggregation tools, and setting up dashboards and alerts.

fathom-cost-tuning

from ComeOnOliver/skillshub

Optimize Fathom API usage and plan selection. Trigger with phrases like "fathom cost", "fathom pricing", "fathom plan".