What Claude skills are (and what they are not)
A skill is a folder with a SKILL.mdfile inside it. That file teaches Claude how to do one job well: structure a reproducible training run, chunk documents for a RAG pipeline the right way, or write an eval that would actually catch a regression instead of rubber-stamping it. The clever part is progressive disclosure. Claude only loads the skill's name and one-line description until your task actually matches it. So you can install thirty skills and pay almost no context cost until the moment one is needed.
People confuse skills with two neighbouring things, so it is worth being exact. Skills teach. MCP servers connect. Slash commands trigger.A skill is knowledge and procedure. An MCP server is a live connection to an external system: your W&B project, your vector database, your model registry. It gives Claude real data and actions. A slash command is a prompt template you fire manually. For ML/AI work you will usually pair the two: a skill for the patterns, a connection (a platform CLI, an MCP server) for the live experiment or deployment state. (This is also the ground we cover hands-on in our Claude Code training for engineering teams.)
How to install any of these (one pattern, not ten)
Every skill below installs the same way, so here is the pattern once. The modern route is plugin marketplaces. Inside Claude Code:
# Add the marketplace, then install the skill
/plugin marketplace add <github-owner/repo>
/plugin install <skill-name>The exact marketplace and skill-set names vary by repo, so check each one's README. You can also install by hand. Clone or copy the skill folder into .claude/skills/ in your project for one repo, or ~/.claude/skills/ to make it available everywhere, then reload Claude Code. That is the whole mechanism. For the rest of this guide we focus on what each skill is good for, not on repeating install steps.
The 10 best Claude skills for ML/AI engineers
Ranked by a blend of trust (first-party beats anonymous), usefulness for real day-to-day ML/AI work, and how actively maintained the skill is in mid-2026. Star counts on monorepos are repo-wide, not per-skill, and we have flagged licensing honestly.
1. huggingface/skills: the Hub, from the source
Source: huggingface/skills · Stars: ~11.1k · Licence: Apache-2.0 · Updated: September 2026 · Best for: Hub operations, dataset creation, fine-tuning and evaluation, one official skill set
Hugging Face publishes its own skill collection for Claude Code, and it is the obvious starting point for anything touching the Hub: dataset creation and handling, LLM and vision model fine-tuning (including a dedicated TRL training skill), Spaces deployment, running models locally, and community-driven evaluation. What you are really installing is Hugging Face's own conventions for the ecosystem most ML/AI engineers already live in, encoded so Claude stops guessing at Hub API calls and dataset schemas.
2. wandb/skills: experiment tracking that matches what you actually run
Source: wandb/skills · Stars: ~69 · Licence: Apache-2.0 · Updated: September 2026 · Best for: Experiment tracking, sweeps and Weave tracing wired to the platform you already run
Weights & Biases' own skill for the platform: automatic logging of training metrics, hyperparameter sweeps, and Weave tracing for agent and LLM applications. If your team already lives in W&B dashboards, this is the skill that makes Claude write training code that shows up there correctly the first time, rather than a run you have to patch afterwards to make comparable to the last one.
3. mlflow/skills: tracing and eval for GenAI apps
Source: mlflow/skills · Stars: ~79 · Licence: Apache-2.0 · Updated: September 2026 · Best for: Instrumenting GenAI apps with MLflow tracing, debugging traces and running automated evals
MLflow's own skill collection, built for instrumenting GenAI applications with tracing, debugging traces when something goes wrong in production, and running automated evaluations against an observability layer you can actually query. It complements W&B rather than competing with it directly: wandb/skills suits training-run tracking, this suits tracing and evaluating what a deployed LLM or agent actually did in the wild.
4. nemo-retriever: production RAG at NVIDIA scale
Source: NVIDIA/skills · Stars: ~3.4k (repo-wide) · Licence: Apache-2.0 (code) / CC-BY-4.0 (docs) · Updated: September 2026 · Best for: Deploying NeMo Retriever locally and answering questions against your own corpus, on the NVIDIA stack
One entry from NVIDIA's official, NVIDIA-verified skills catalogue: it deploys the NeMo Retriever library locally, extracts information from your own corpus, and answers questions against it. Each skill in the catalogue ships with a signed skill card and governance metadata, an unusually strong trust signal for this list. The honest limitation is scope: it is built for teams already on, or evaluating, the NVIDIA NeMo stack, not a general-purpose RAG starting point for everyone.
5. mle-workflow: the ML lifecycle, not a one-off notebook
Source: affaan-m/everything-claude-code · Stars: ~267k (repo-wide) · Licence: MIT · Updated: September 2026 · Best for: The full ML lifecycle: data contracts, reproducible training, promotion criteria, serving and rollback
From the most-starred Claude skills monorepo in the ecosystem, this skill pushes Claude toward the practices that separate a maintainable ML system from a script that only the person who wrote it can rerun: explicit data contracts, reproducible training pipelines, measurable promotion criteria before a model ships, and serving and rollback plans. It is the skill to reach for when you are hardening something that started life as exploratory work into a system someone else will operate.
6. pytorch-patterns: training code without the footguns
Source: affaan-m/everything-claude-code · Stars: ~267k (repo-wide) · Licence: MIT · Updated: September 2026 · Best for: Idiomatic PyTorch: device-agnostic code, reproducibility, training loops and data pipelines
The companion training skill from the same monorepo: device-agnostic code, reproducibility (seeding, deterministic ops), model architecture and data-loading patterns, checkpointing, and performance optimisation, with explicit shape management called out as a core principle. It also documents the anti-patterns directly, which matters more than the happy-path examples, since that is what an unsupervised model tends to reproduce.
7. rag-architect: retrieval systems built to be evaluated, not just demoed
Source: Jeffallan/claude-skills · Stars: ~11.6k (repo-wide) · Licence: MIT · Updated: September 2026 · Best for: Chunking, embeddings, vector store config, hybrid search, reranking and retrieval evaluation
Designs and implements production-grade RAG systems: chunking documents, generating embeddings, configuring vector stores, building hybrid search pipelines, applying reranking, and evaluating retrieval quality rather than eyeballing a handful of example queries. The evaluation step is the part most hand-rolled RAG builds skip, and it is the reason a retrieval pipeline that looked great in a demo quietly degrades once real queries hit it.
8. ml-pipeline: orchestration and feature engineering, wired together
Source: Jeffallan/claude-skills · Stars: ~11.6k (repo-wide) · Licence: MIT · Updated: September 2026 · Best for: Kubeflow/Airflow orchestration, feature engineering and a versioned, containerised model lifecycle
Covers the infrastructure layer around a model: Kubeflow and Airflow orchestration, MLflow-backed experiment tracking, feature engineering pipelines, and automated model lifecycle management with proper versioning. It provides code templates and validation workflows for reproducible, containerised ML systems, the part of the job that has nothing to do with model architecture and everything to do with whether last month's training run can be reproduced today.
9. senior-ml-engineer: the production concerns nobody wants to own
Source: alirezarezvani/claude-skills · Stars: ~26.4k (repo-wide) · Licence: MIT · Updated: September 2026 · Best for: Productionising models, feature stores, drift monitoring and cost-aware LLM API integration
Guidance for productionising models: deployment to services like Kubernetes, feature store architecture, monitoring for performance and data drift, RAG pipelines, and cost-aware LLM API integration with retry logic. It reads as the skill for the engineer who inherits a model after the research is done and has to keep it working, priced sensibly, and honest about when it starts drifting.
10. promptfoo-evals: eval suites that live in your repo, not a spreadsheet
Source: promptfoo/promptfoo · Stars: ~25.4k · Licence: MIT · Updated: September 2026 · Best for: Writing and maintaining eval suites: test cases, providers, deterministic and model-graded assertions
Bundled directly in the promptfoo tool's own repository, this skill writes and maintains evaluation suites: creating test cases, choosing providers, writing both deterministic and model-graded assertions, and running evaluations as part of a maintainable config rather than a one-off script. It explicitly excludes adversarial red team work, which promptfoo covers separately, so treat this as the correctness-and-quality half of evaluation, not a security tool.
Worth watching (and what we left off)
A couple of things nearly made the cut. llm-eval-skills (harshrathod0585) is a tight three-skill set covering RAG-triad evaluation, G-Eval-style custom judges and model-selection benchmarking. It is genuinely well-scoped, but it is a brand-new, unstarred single-author repo, so audit it fully before running it against anything real and treat it as one to watch rather than a default pick.
The wider NVIDIA/skills catalogue also ships dedicated fine-tuning and RLHF skills (the nemo-automodel-* and nemo-rl-* families, covering GRPO, DPO and SFT training) beyond the retrieval skill listed above. They are excellent if you are fully committed to the NeMo/Megatron-Core stack, but too narrow a dependency for a general ML/AI engineering list.
A word on security before you install anything
Skills are code, and this warning carries extra weight for this audience. A SKILL.mdcan contain prompt-injection instructions, and any script a skill bundles runs with your agent's permissions: on an ML/AI workstation that can mean Hugging Face tokens, W&B and MLflow API keys, cloud GPU credentials, and the API keys for whichever model providers you call. Snyk's 2026 ToxicSkills research catalogued thousands of skills and found 36.8% had at least one security flaw and 13.4% had a critical-level issue, with roughly one in nine containing hardcoded or exposed secrets and 91% of confirmed malicious skills using prompt injection.
Read the SKILL.mdand any bundled scripts yourself. Do not blind-install. Prefer first-party and high-credibility authors (Hugging Face, Weights & Biases, MLflow, NVIDIA). Check the licence and recent commit activity, and be especially wary of unlicensed, low-star repos and skills that fetch external content at runtime. Reduce blast radius: run cloud-touching or training skills against a sandbox project first, and disable MCP servers and integrations you are not actively using, because a skill can only reach what you have left switched on.
How to actually use these together
You do not need all ten. Every installed skill adds to context on every turn whether Claude uses it that turn or not, so treat this as a menu, not a checklist. The strongest loadout for most ML/AI engineers is five skills covering the core of the job:
- huggingface/skills for Hub operations, dataset handling and fine-tuning, the layer most ML/AI work actually starts from.
- One experiment-tracking skill: wandb/skills or mlflow/skills, not both, so training runs and production traces are comparable instead of scattered.
- mle-workflow so training and serving follow data contracts and promotion criteria rather than a notebook that only runs on one laptop.
- rag-architect once retrieval is part of the system, for chunking, vector store config and, critically, evaluation.
Add promptfoo-evals the moment anything ships to real users, so quality is measured rather than assumed, and reach for pytorch-patterns or ml-pipeline once you are deep enough into custom training or orchestration for their specifics to matter. That is the playbook: source data, track it, ship it properly, evaluate it honestly.
The short version
Skills are the difference between an AI that produces a training script which runs once on your machine and one that produces a system you could hand to another engineer, or defend in a post-incident review when a model starts drifting. If you install nothing else today, install huggingface/skills and one experiment-tracking skill, and read their SKILL.md files first. Given that only 52% of teams run offline evals despite 89% having observability, closing that specific gap with promptfoo-evals or a similar skill is probably the single highest-leverage addition on this list. The ecosystem moves fast; the principle does not. Constrain the model with good skills and you get fewer models that looked fine in the demo and quietly failed in production.
If your team also owns the infrastructure these models get deployed onto, our companion guide covers the top Claude skills for DevOps engineers, and if the same engineers own the API layer in front of a model, see the top Claude skills for backend engineers.
We teach engineering teams to use Claude Code properly: skills, agentic workflows, and shipping AI-assisted code safely. See our Claude Code training, or book a 45-minute calland we'll map the fastest path for your team.
Sources
- RAND Corporation (2025). Why AI Projects Fail and How They Can Succeed. rand.org/pubs/research_reports/RRA2680-1
- LangChain & Datadog (2026). State of Agent Engineering. langchain.com/state-of-agent-engineering
- Snyk (2026). ToxicSkills: malicious AI agent skills. snyk.io/blog/toxicskills-malicious-ai-agent-skills-clawhub
- Anthropic. Claude Code skills documentation. code.claude.com/docs/en/skills
- Repositories referenced above: huggingface/skills, wandb/skills, mlflow/skills, NVIDIA/skills, affaan-m/everything-claude-code, Jeffallan/claude-skills, alirezarezvani/claude-skills and promptfoo/promptfoo, plus harshrathod0585/llm-eval-skills. Star counts and dates verified via GitHub, 25 September 2026.