Top Claude Skills for ML/AI Engineers (2026 Guide)

Claude Code will happily write you a training script, a RAG pipeline, or an eval harness. The hard part is getting it to write the version that survives contact with real data, real users, and a real production incident review. That is what skills are for. Here are the ten best Claude skills for ML and AI engineers right now: what each one does, who it suits, and how to install them without handing your Hugging Face, W&B or cloud GPU credentials to a stranger.

80%+
of AI projects fail to reach production or deliver a return, roughly twice the failure rate of non-AI IT projects
RAND Corporation, "Why AI Projects Fail and How They Can Succeed", 2025
52%
of AI engineering teams run offline evals on their agents, even though 89% have already added observability
LangChain & Datadog, State of Agent Engineering 2026
36.8%
of catalogued skills had at least one security flaw, so read before you install
Snyk ToxicSkills research, 2026
~1 in 9
catalogued skills contained hardcoded or exposed secrets: a real risk on a credential-heavy ML/AI workstation
Snyk ToxicSkills research, 2026

What Claude skills are (and what they are not)

A skill is a folder with a SKILL.mdfile inside it. That file teaches Claude how to do one job well: structure a reproducible training run, chunk documents for a RAG pipeline the right way, or write an eval that would actually catch a regression instead of rubber-stamping it. The clever part is progressive disclosure. Claude only loads the skill's name and one-line description until your task actually matches it. So you can install thirty skills and pay almost no context cost until the moment one is needed.

People confuse skills with two neighbouring things, so it is worth being exact. Skills teach. MCP servers connect. Slash commands trigger.A skill is knowledge and procedure. An MCP server is a live connection to an external system: your W&B project, your vector database, your model registry. It gives Claude real data and actions. A slash command is a prompt template you fire manually. For ML/AI work you will usually pair the two: a skill for the patterns, a connection (a platform CLI, an MCP server) for the live experiment or deployment state. (This is also the ground we cover hands-on in our Claude Code training for engineering teams.)

How to install any of these (one pattern, not ten)

Every skill below installs the same way, so here is the pattern once. The modern route is plugin marketplaces. Inside Claude Code:

# Add the marketplace, then install the skill
/plugin marketplace add <github-owner/repo>
/plugin install <skill-name>

The exact marketplace and skill-set names vary by repo, so check each one's README. You can also install by hand. Clone or copy the skill folder into .claude/skills/ in your project for one repo, or ~/.claude/skills/ to make it available everywhere, then reload Claude Code. That is the whole mechanism. For the rest of this guide we focus on what each skill is good for, not on repeating install steps.

The 10 best Claude skills for ML/AI engineers

Ranked by a blend of trust (first-party beats anonymous), usefulness for real day-to-day ML/AI work, and how actively maintained the skill is in mid-2026. Star counts on monorepos are repo-wide, not per-skill, and we have flagged licensing honestly.

1. huggingface/skills: the Hub, from the source

Official · first party · Apache-2.0

Source: huggingface/skills · Stars: ~11.1k · Licence: Apache-2.0 · Updated: September 2026 · Best for: Hub operations, dataset creation, fine-tuning and evaluation, one official skill set

Hugging Face publishes its own skill collection for Claude Code, and it is the obvious starting point for anything touching the Hub: dataset creation and handling, LLM and vision model fine-tuning (including a dedicated TRL training skill), Spaces deployment, running models locally, and community-driven evaluation. What you are really installing is Hugging Face's own conventions for the ecosystem most ML/AI engineers already live in, encoded so Claude stops guessing at Hub API calls and dataset schemas.

2. wandb/skills: experiment tracking that matches what you actually run

Official · first party · Apache-2.0

Source: wandb/skills · Stars: ~69 · Licence: Apache-2.0 · Updated: September 2026 · Best for: Experiment tracking, sweeps and Weave tracing wired to the platform you already run

Weights & Biases' own skill for the platform: automatic logging of training metrics, hyperparameter sweeps, and Weave tracing for agent and LLM applications. If your team already lives in W&B dashboards, this is the skill that makes Claude write training code that shows up there correctly the first time, rather than a run you have to patch afterwards to make comparable to the last one.

3. mlflow/skills: tracing and eval for GenAI apps

Official · first party · Apache-2.0

Source: mlflow/skills · Stars: ~79 · Licence: Apache-2.0 · Updated: September 2026 · Best for: Instrumenting GenAI apps with MLflow tracing, debugging traces and running automated evals

MLflow's own skill collection, built for instrumenting GenAI applications with tracing, debugging traces when something goes wrong in production, and running automated evaluations against an observability layer you can actually query. It complements W&B rather than competing with it directly: wandb/skills suits training-run tracking, this suits tracing and evaluating what a deployed LLM or agent actually did in the wild.

4. nemo-retriever: production RAG at NVIDIA scale

Official · first party · NVIDIA-verified, signed skill cards

Source: NVIDIA/skills · Stars: ~3.4k (repo-wide) · Licence: Apache-2.0 (code) / CC-BY-4.0 (docs) · Updated: September 2026 · Best for: Deploying NeMo Retriever locally and answering questions against your own corpus, on the NVIDIA stack

One entry from NVIDIA's official, NVIDIA-verified skills catalogue: it deploys the NeMo Retriever library locally, extracts information from your own corpus, and answers questions against it. Each skill in the catalogue ships with a signed skill card and governance metadata, an unusually strong trust signal for this list. The honest limitation is scope: it is built for teams already on, or evaluating, the NVIDIA NeMo stack, not a general-purpose RAG starting point for everyone.

5. mle-workflow: the ML lifecycle, not a one-off notebook

Community · MIT · very popular monorepo

Source: affaan-m/everything-claude-code · Stars: ~267k (repo-wide) · Licence: MIT · Updated: September 2026 · Best for: The full ML lifecycle: data contracts, reproducible training, promotion criteria, serving and rollback

From the most-starred Claude skills monorepo in the ecosystem, this skill pushes Claude toward the practices that separate a maintainable ML system from a script that only the person who wrote it can rerun: explicit data contracts, reproducible training pipelines, measurable promotion criteria before a model ships, and serving and rollback plans. It is the skill to reach for when you are hardening something that started life as exploratory work into a system someone else will operate.

6. pytorch-patterns: training code without the footguns

Community · MIT · very popular monorepo

Source: affaan-m/everything-claude-code · Stars: ~267k (repo-wide) · Licence: MIT · Updated: September 2026 · Best for: Idiomatic PyTorch: device-agnostic code, reproducibility, training loops and data pipelines

The companion training skill from the same monorepo: device-agnostic code, reproducibility (seeding, deterministic ops), model architecture and data-loading patterns, checkpointing, and performance optimisation, with explicit shape management called out as a core principle. It also documents the anti-patterns directly, which matters more than the happy-path examples, since that is what an unsupervised model tends to reproduce.

7. rag-architect: retrieval systems built to be evaluated, not just demoed

Community · MIT · popular monorepo

Source: Jeffallan/claude-skills · Stars: ~11.6k (repo-wide) · Licence: MIT · Updated: September 2026 · Best for: Chunking, embeddings, vector store config, hybrid search, reranking and retrieval evaluation

Designs and implements production-grade RAG systems: chunking documents, generating embeddings, configuring vector stores, building hybrid search pipelines, applying reranking, and evaluating retrieval quality rather than eyeballing a handful of example queries. The evaluation step is the part most hand-rolled RAG builds skip, and it is the reason a retrieval pipeline that looked great in a demo quietly degrades once real queries hit it.

8. ml-pipeline: orchestration and feature engineering, wired together

Community · MIT · popular monorepo

Source: Jeffallan/claude-skills · Stars: ~11.6k (repo-wide) · Licence: MIT · Updated: September 2026 · Best for: Kubeflow/Airflow orchestration, feature engineering and a versioned, containerised model lifecycle

Covers the infrastructure layer around a model: Kubeflow and Airflow orchestration, MLflow-backed experiment tracking, feature engineering pipelines, and automated model lifecycle management with proper versioning. It provides code templates and validation workflows for reproducible, containerised ML systems, the part of the job that has nothing to do with model architecture and everything to do with whether last month's training run can be reproduced today.

9. senior-ml-engineer: the production concerns nobody wants to own

Community · MIT · very popular monorepo

Source: alirezarezvani/claude-skills · Stars: ~26.4k (repo-wide) · Licence: MIT · Updated: September 2026 · Best for: Productionising models, feature stores, drift monitoring and cost-aware LLM API integration

Guidance for productionising models: deployment to services like Kubernetes, feature store architecture, monitoring for performance and data drift, RAG pipelines, and cost-aware LLM API integration with retry logic. It reads as the skill for the engineer who inherits a model after the research is done and has to keep it working, priced sensibly, and honest about when it starts drifting.

10. promptfoo-evals: eval suites that live in your repo, not a spreadsheet

Official · first party · bundled with the eval tool itself

Source: promptfoo/promptfoo · Stars: ~25.4k · Licence: MIT · Updated: September 2026 · Best for: Writing and maintaining eval suites: test cases, providers, deterministic and model-graded assertions

Bundled directly in the promptfoo tool's own repository, this skill writes and maintains evaluation suites: creating test cases, choosing providers, writing both deterministic and model-graded assertions, and running evaluations as part of a maintainable config rather than a one-off script. It explicitly excludes adversarial red team work, which promptfoo covers separately, so treat this as the correctness-and-quality half of evaluation, not a security tool.

Worth watching (and what we left off)

A couple of things nearly made the cut. llm-eval-skills (harshrathod0585) is a tight three-skill set covering RAG-triad evaluation, G-Eval-style custom judges and model-selection benchmarking. It is genuinely well-scoped, but it is a brand-new, unstarred single-author repo, so audit it fully before running it against anything real and treat it as one to watch rather than a default pick.

The wider NVIDIA/skills catalogue also ships dedicated fine-tuning and RLHF skills (the nemo-automodel-* and nemo-rl-* families, covering GRPO, DPO and SFT training) beyond the retrieval skill listed above. They are excellent if you are fully committed to the NeMo/Megatron-Core stack, but too narrow a dependency for a general ML/AI engineering list.

A word on security before you install anything

Skills are code, and this warning carries extra weight for this audience. A SKILL.mdcan contain prompt-injection instructions, and any script a skill bundles runs with your agent's permissions: on an ML/AI workstation that can mean Hugging Face tokens, W&B and MLflow API keys, cloud GPU credentials, and the API keys for whichever model providers you call. Snyk's 2026 ToxicSkills research catalogued thousands of skills and found 36.8% had at least one security flaw and 13.4% had a critical-level issue, with roughly one in nine containing hardcoded or exposed secrets and 91% of confirmed malicious skills using prompt injection.

Before you install

Read the SKILL.mdand any bundled scripts yourself. Do not blind-install. Prefer first-party and high-credibility authors (Hugging Face, Weights & Biases, MLflow, NVIDIA). Check the licence and recent commit activity, and be especially wary of unlicensed, low-star repos and skills that fetch external content at runtime. Reduce blast radius: run cloud-touching or training skills against a sandbox project first, and disable MCP servers and integrations you are not actively using, because a skill can only reach what you have left switched on.

How to actually use these together

You do not need all ten. Every installed skill adds to context on every turn whether Claude uses it that turn or not, so treat this as a menu, not a checklist. The strongest loadout for most ML/AI engineers is five skills covering the core of the job:

  1. huggingface/skills for Hub operations, dataset handling and fine-tuning, the layer most ML/AI work actually starts from.
  2. One experiment-tracking skill: wandb/skills or mlflow/skills, not both, so training runs and production traces are comparable instead of scattered.
  3. mle-workflow so training and serving follow data contracts and promotion criteria rather than a notebook that only runs on one laptop.
  4. rag-architect once retrieval is part of the system, for chunking, vector store config and, critically, evaluation.

Add promptfoo-evals the moment anything ships to real users, so quality is measured rather than assumed, and reach for pytorch-patterns or ml-pipeline once you are deep enough into custom training or orchestration for their specifics to matter. That is the playbook: source data, track it, ship it properly, evaluate it honestly.

The short version

Skills are the difference between an AI that produces a training script which runs once on your machine and one that produces a system you could hand to another engineer, or defend in a post-incident review when a model starts drifting. If you install nothing else today, install huggingface/skills and one experiment-tracking skill, and read their SKILL.md files first. Given that only 52% of teams run offline evals despite 89% having observability, closing that specific gap with promptfoo-evals or a similar skill is probably the single highest-leverage addition on this list. The ecosystem moves fast; the principle does not. Constrain the model with good skills and you get fewer models that looked fine in the demo and quietly failed in production.

If your team also owns the infrastructure these models get deployed onto, our companion guide covers the top Claude skills for DevOps engineers, and if the same engineers own the API layer in front of a model, see the top Claude skills for backend engineers.

Work with us

We teach engineering teams to use Claude Code properly: skills, agentic workflows, and shipping AI-assisted code safely. See our Claude Code training, or book a 45-minute calland we'll map the fastest path for your team.

Sources

  1. RAND Corporation (2025). Why AI Projects Fail and How They Can Succeed. rand.org/pubs/research_reports/RRA2680-1
  2. LangChain & Datadog (2026). State of Agent Engineering. langchain.com/state-of-agent-engineering
  3. Snyk (2026). ToxicSkills: malicious AI agent skills. snyk.io/blog/toxicskills-malicious-ai-agent-skills-clawhub
  4. Anthropic. Claude Code skills documentation. code.claude.com/docs/en/skills
  5. Repositories referenced above: huggingface/skills, wandb/skills, mlflow/skills, NVIDIA/skills, affaan-m/everything-claude-code, Jeffallan/claude-skills, alirezarezvani/claude-skills and promptfoo/promptfoo, plus harshrathod0585/llm-eval-skills. Star counts and dates verified via GitHub, 25 September 2026.

Claude Skills for ML/AI Engineers: Frequently Asked Questions

A Claude skill is a folder containing a SKILL.md file (plus any optional scripts or reference docs) that teaches Claude Code how to do a specific job, for example fine-tune a Hugging Face model correctly or write a retrieval evaluation that catches a regression. Claude only loads a skill's name and description until your task matches it, so you can install dozens of skills without filling up the context window. The same SKILL.md format also works in Codex, Cursor and Gemini CLI.

Skills teach, MCP connects. A skill is bundled knowledge and procedure: how to structure a reproducible training run, what a sane RAG evaluation actually measures. An MCP server is a live connection to an external system, such as your Weights & Biases project, your vector database or your model registry, giving Claude real data and actions. In practice ML/AI engineers combine both: a skill for the patterns, plus a platform MCP server or CLI for the live experiment or deployment state.

A five-skill loadout covers most of the job: huggingface/skills for Hub operations, dataset handling and fine-tuning; one experiment-tracking skill, wandb/skills or mlflow/skills, not both, to see and share what actually happened in a run; mle-workflow (everything-claude-code) so training and serving follow a real lifecycle rather than a one-off notebook; rag-architect if you are building retrieval systems; and promptfoo-evals so shipped behaviour is measured, not assumed.

Treat them as untrusted code, and be stricter than most developers need to be: an ML/AI workstation typically holds Hugging Face, W&B, OpenAI or Anthropic API keys and cloud GPU credentials, so the blast radius is bigger. Snyk's 2026 ToxicSkills research found 36.8% of catalogued skills had at least one security flaw, 13.4% had a critical-level issue, and roughly one in nine contained hardcoded or exposed secrets. Read every SKILL.md and bundled script before installing, prefer first-party or high-credibility authors, and be wary of unlicensed, low-star repos.

The modern way is through plugin marketplaces: run "/plugin marketplace add <github-repo>" inside Claude Code, then "/plugin install <skill-name>". You can also install manually by cloning the skill folder into ".claude/skills/" in your project (or "~/.claude/skills/" for global use). Restart or reload Claude Code and the skill becomes available.

Observability first if you have to choose, but do not stop there. Industry data backs this up: LangChain and Datadog's State of Agent Engineering 2026 found 89% of teams have added observability for their agents, yet only 52% run offline evals. That gap is exactly where regressions hide, since a trace tells you what happened, not whether it was correct. Wire up tracing (MLflow, Weave via wandb/skills) so you can see failures, then add promptfoo-evals or a similar eval skill so you catch them before a user does.

Want your team using Claude Code properly?

We run hands-on Claude Code trainingfor engineering teams: skills, agentic workflows, and how to ship AI-assisted systems safely. Book a free 45-minute call and we'll map the fastest path for your team.

Book a free call →