What Claude skills are (and what they are not)
A skill is a folder with a SKILL.md file inside it. That file teaches Claude how to do one job well: design a burn-rate alert that fires on real error budget consumption, write a runbook that names an actual escalation contact, or spot a Terragrunt unit dependency that will deadlock on apply.
The clever part is progressive disclosure. Claude only loads the skill's name and one-line description until your task actually matches it, so you can install a dozen skills and pay almost no context cost until the moment one is needed.
People confuse skills with two neighbouring things, so it is worth being exact. Skills teach. MCP servers connect. Slash commands trigger. A skill is knowledge and procedure. An MCP server is a live connection to an external system: your Prometheus or Grafana instance, your Kubernetes cluster, your PagerDuty schedule. It gives Claude real data and actions.
A slash command is a prompt template you fire manually. For platform and SRE work you will usually pair the two: a skill for the operational patterns, a live connection for the current cluster or dashboard state. (This is also the ground we cover hands-on in our Claude Code training for engineering teams.)
How to install any of these (two patterns, not seven)
Most of these ship a proper plugin marketplace manifest, so the modern route works: inside Claude Code,
# Add the marketplace, then install the skill
/plugin marketplace add <github-owner/repo>
/plugin install <skill-name>@<marketplace-name>That covers alerting-irm, prometheus-mimir-grafana, terragrunt-skill and cloud-finops. Three of the seven do not ship a marketplace manifest: k8s-platform-operations and k8s-platform-tenancy only document cloning or symlinking into .claude/skills/, and runbook-generator installs through the Claude desktop app's plugin browser or its own npx command rather than /plugin marketplace add.
For those, or for any skill you would rather vet before it touches a marketplace listing, install by hand: clone or copy the skill folder into .claude/skills/ in your project for one repo, or ~/.claude/skills/to make it available everywhere, then reload Claude Code. Check each repo's own README for its exact command before assuming the marketplace path works, since one wrong assumption here is a wasted install attempt, not a security problem.
The 7 best Claude skills for platform engineers and SREs
Ranked by a blend of trust (first-party beats anonymous), usefulness for real production-reliability work, and how recently the skill was maintained (the last-commit date is on every entry). Star counts on monorepos are repo-wide, not per-skill, and we have flagged licensing and maturity honestly, including the smaller repos that made the list on usefulness alone.
1. alerting-irm: alerting and SLOs from the people who host Grafana
Source: grafana/skills · Stars: 276 (repo-wide) · Licence: Apache-2.0 · Last commit: June 2026 · Best for: Alert provisioning, on-call routing and multi-window burn-rate SLO alerts, done properly
Grafana Labs publishes its own public skills repository for the LGTM stack, and for alerting work it is the obvious starting point. This skill covers end-to-end Grafana Alerting configuration: provisioning managed and data-source-managed alert rules, contact points across Slack, PagerDuty, email and webhook, notification policies with hierarchical label matchers, silences, mute timings, on-call schedules and escalation chains, plus SLOs with multi-window burn-rate alerts.
What you are really installing is Grafana's own house view of what a good alert looks like, encoded so Claude stops generating rules that fire on noise and escalate paths nobody staffed.
The limitation: it teaches Grafana's alerting model specifically, so if your stack runs a different alerting stack the patterns will not transfer directly.
2. runbook-generator: operational knowledge that outlives the engineer who wrote it
Source: borghei/Claude-Skills · Stars: 859 (repo-wide) · Licence: MIT + Commons Clause · Last commit: June 2026 · Best for: Turning tribal on-call knowledge into a stack-aware, copy-paste runbook
From one of the larger general-purpose Claude skills repositories, this one is narrowly useful for platform and SRE teams specifically: it generates operational runbooks by first detecting your actual stack from the codebase (CI/CD platform, database, hosting, orchestration) rather than handing back a generic template. Every runbook it produces follows the same discipline: numbered steps with executable commands, a verification check after each action, a rollback procedure for anything destructive, and an escalation path with named contacts.
The limitation: it generates documentation, not automated execution, so treat its output as a draft runbook for a human to check and sign off, not a script you point at production unattended.
3. k8s-platform-operations: Kubernetes cluster operations and incident response
Source: foxj77/claude-code-skills · Stars: 21 · Licence: MIT · Last commit: February 2026 · Best for: Day-2 Kubernetes: incident severity, capacity, node maintenance, backups
Part of a focused Kubernetes platform-engineering collection, this skill is specifically about running a cluster after it is already live: cluster health assessment, incident severity classification and a seven-stage incident workflow (detection, triage, communication, investigation, mitigation, resolution, review), capacity analysis, node maintenance procedures, and backup and recovery strategy.
One detail worth repeating on its own: the guidance insists on kubectl cordon before kubectl drain, to stop new pods scheduling onto a node mid-eviction, which is exactly the kind of easy-to-skip step an AI-generated runbook needs to get right.
The limitation: at 21 stars this is a small, single-maintainer collection, and its last commit was in February 2026, so verify its procedures against your own cluster version before trusting them during a real incident.
4. prometheus-mimir-grafana: getting the query right, not just plausible
Source: air-gapped/skills · Stars: 5 · Licence: MIT · Last commit: September 2026 · Best for: Getting PromQL and dashboard queries right on a self-hosted metrics stack
A small repo built for a real niche: teams running Prometheus, Grafana Mimir or self-hosted Grafana rather than a fully managed observability platform. Its value is in the specifics it corrects, spelled out as a trap list: rate() applied after aggregation instead of before, histogram_quantile run against a mean rather than bucket rates, counter-rate windows too narrow to be stable, and conflicting label names across exporters that quietly break a join.
It also walks through a discovery ladder for an unfamiliar metrics endpoint: catalogue by prefix, check metadata and type, enumerate labels, test aliveness, sample the shape, then aggregate.
The limitation: five stars is not much of a following, so treat it as one well-written reference rather than a battle-tested standard, though the SKILL.md itself reads like it was written by someone who has been paged for a dashboard that looked correct and was not.
5. terragrunt-skill: infrastructure as code for the multi-account platform
Source: jfr992/terragrunt-skill · Stars: 27 · Licence: Apache-2.0 · Last commit: July 2026 · Best for: Multi-account, multi-environment infrastructure with Terragrunt and OpenTofu
Distinct from a general Terraform skill: Terragrunt is what most platform teams reach for once a single Terraform repo stops scaling across accounts and environments, and this skill covers that specifically. It documents a three-repository model (an infrastructure catalogue of reusable units and stacks, a live environment repo with deployment-specific config, and versioned module repos), remote state with S3 native locking or DynamoDB, cross-account role assumption, and migrating a monolithic "terralith" into properly separated units.
If your platform already runs Terragrunt, this is more directly useful than a generic Terraform skill; if you do not, skip it and use a Terraform-specific skill instead.
6. cloud-finops: cloud and Kubernetes cost management
Source: OptimNow/cloud-finops-skills · Stars: 57 · Licence: CC BY-SA 4.0 · Last commit: October 2026 · Best for: Cloud and Kubernetes cost allocation, waste detection and commitment sizing
A vendor-published skill from OptimNow, refreshed roughly twice a month against a set of tracked pricing and provider sources, covering cloud cost across AWS, Azure, GCP and OCI, plus Kubernetes cost allocation, commitment sizing (Reserved Instances and Savings Plans), waste detection and FinOps governance and KPIs.
It is broader than a Kubernetes-only cost tool, which is honestly the more useful shape for a platform team: container costs rarely make sense in isolation from the account and commitment strategy sitting around them.
The limitation: best used for a quarterly cost review against read-only billing access, not wired into a daily or production-write workflow.
7. k8s-platform-tenancy: multi-tenant clusters without the cross-contamination
Source: foxj77/claude-code-skills · Stars: 21 (repo-wide) · Licence: MIT · Last commit: February 2026 · Best for: Multi-tenant namespace provisioning, RBAC and resource isolation on shared clusters
The narrowest entry on this list, from the same collection as k8s-platform-operations, and it earns its place on a specific problem: running one shared cluster across several teams or customers without one tenant's workload starving or reaching another's. It documents a six-layer isolation model (namespace boundaries, RBAC, network policies, resource quotas, limit ranges and pod security standards), a default-deny-all network policy sequence, and a tiered service pattern for onboarding new tenants.
The limitation: if your cluster serves one team, this is dead weight, install it only once you are actually running a shared multi-tenant platform. It also comes from the same small, single-maintainer collection as k8s-platform-operations, and its last commit was in February 2026, so verify its isolation rules against your own cluster version before trusting them.
Worth watching (and what we left off)
A couple of things nearly made the cut, and one is worth naming for a different reason. sinzin91/awesome-sre-skills is a curated list of 53+ AI agent skills for SRE work (monitoring, incident response, observability and infrastructure) rather than an installable skill itself, but it is a genuinely useful map if you want to look beyond this list.
kryptophonik/claude-srepromises dual-persona SRE capabilities (a strategic "architect" and a tactical "engineer") and the scope reads well on paper: SLO templates, alerting design, chaos engineering, postmortem facilitation. We left it off the ranked list because, on inspection, it does not follow the standard SKILL.md format at all: no YAML frontmatter, no per-skill manifest, just a set of composable markdown reference files routed by keyword.
That is not necessarily unsafe, but it means it will not install or activate the way the skills above do, and a four-star repo with a non-standard structure is exactly the kind of thing the security section below is asking you to read closely before trusting.
A word on security before you install anything
Skills are code, and this warning carries extra weight for this audience. A SKILL.mdcan contain prompt-injection instructions, and any script a skill bundles runs with your agent's permissions: on a platform or SRE workstation that can mean cloud provider credentials, kubeconfigs for production clusters, Terraform or Terragrunt state access, and API tokens for Grafana, PagerDuty or similar on-call tooling.
Snyk's 2026 ToxicSkills research catalogued thousands of skills and found 36.8% had at least one security flaw and 13.4% had a critical-level issue, with roughly one in nine containing hardcoded or exposed secrets and 91% of confirmed malicious skills using prompt injection.
Read the SKILL.md and any bundled scripts yourself. Do not blind-install. Prefer first-party and high-credibility authors (Grafana, and skills where the author or maintainer is identifiable and active). Check the licence and recent commit activity, and be especially wary of unlicensed, low-star repos, skills that skip the standard SKILL.md format, and anything that fetches external content at runtime.
Reduce blast radius: run cluster- and cloud-touching skills against a sandbox or read-only account first, and disable MCP servers and integrations you are not actively using. That narrows what Claude can reach through its structured tools, but it does not fix everything: a bundled script with shell access can still read credential files or hit the network directly, so isolate execution and restrict the credentials on disk too, not just the MCP toggles.
How to actually use these together
You do not need all seven running at once. For most platform and SRE teams, four is enough to start:
- alerting-irm for alert rules, on-call routing and SLOs that mean something.
- k8s-platform-operations for Day-2 cluster operations and incident response.
- One infrastructure-as-code skill: terragrunt-skill if you run Terragrunt, otherwise a Terraform-specific equivalent for your stack.
- runbook-generator so what you learn during an incident gets written down, not just fixed.
Add prometheus-mimir-grafana if you run a self-hosted metrics stack rather than a fully managed one, cloud-finops for a quarterly cost review rather than a daily habit, and k8s-platform-tenancy only once you are actually running a shared, multi-tenant cluster. That is the playbook: alert, operate, provision, record, and check the bill once a quarter.
The short version
Skills are the difference between an AI that produces a plausible alert rule and one that produces an alert rule you would actually trust to wake you up for the right reason. If you install nothing else today, install alerting-irm and k8s-platform-operations, and read their SKILL.md files first.
Validate every AI-generated alert threshold and Terragrunt plan against a real burn rate and a terragrunt plan dry run before either touches production. The ecosystem moves fast; the principle does not. Constrain the model with good skills and you get fewer false pages and fewer 3am surprises.
If your team also owns the infrastructure layer more broadly, our companion guide covers the top Claude skills for DevOps engineers, and if it owns the services running on top of that infrastructure, see the top Claude skills for backend engineers.
We teach engineering teams to use Claude Code properly: skills, agentic workflows, and shipping AI-assisted infrastructure and operations safely. See our Claude Code training, or book a 45-minute calland we'll map the fastest path for your team.
Sources
- Stack Overflow. Developer Survey: AI. survey.stackoverflow.co/2025/ai
- Snyk (2026). ToxicSkills: malicious AI agent skills. snyk.io/blog/toxicskills-malicious-ai-agent-skills-clawhub
- Anthropic. Claude Code skills documentation. code.claude.com/docs/en/skills
- Repositories referenced above: grafana/skills, borghei/Claude-Skills, foxj77/claude-code-skills, air-gapped/skills, jfr992/terragrunt-skill, OptimNow/cloud-finops-skills, sinzin91/awesome-sre-skills and kryptophonik/claude-sre. Star counts and last-commit dates read from the GitHub API on 4 October 2026. For a skill inside a larger repository the date is the last commit to that skill's folder, and the star count is for the whole repository.