Skill Constellations maps how agent skills are copied across GitHub

Skill Constellations maps how agent skills are copied across GitHub

Agent skills are SKILL.md instructions and scripts that AI coding agents such as Claude Code and Codex run with the permissions of their user. Developers share them by copying them between repositories. The authors argue this makes skills a software supply chain without a registry, versions or provenance, so the origin of a copied skill, the reach of a security fix and the repositories that deserve review are all unknown. They add that earlier studies, which record which repositories hold a skill at a single point in time, cannot show who copied it from whom.

Their contribution is the first dated copy network of agent skills. It is built from the git history of every SKILL.md in GitSkills and covers 2,193,119 skill adoptions across GitHub. An interactive viewer comes with it, and the paper lists a project website at https://fahdseddik.github.io/Skill-Constellations/.

The network yields two findings. First, a few repositories are the source of almost all copies, and GitHub stars do not identify them. Second, skill copies almost never change with their source, so a fix made at the source rarely reaches them.

The authors then fit a model of which repositories others copy from and use it to rank repositories for audit. Reviewing the 100 repositories the model ranks highest prevents 14.9% of later adoptions of high-risk skills, against 0.5% for the 100 most starred repositories. The authors frame this as a short list that security engineers can check before a skill spreads. Their recommendation to platforms is to distribute versioned references rather than copies.

Key facts

  • The paper presents the first dated copy network of agent skills, built from the git history of every SKILL.md in GitSkills and covering 2,193,119 skill adoptions across GitHub, with an interactive viewer.
  • A few repositories are the source of almost all copies, and GitHub stars do not identify them.
  • Skill copies almost never change with their source, so a fix at the source rarely reaches them.
  • Reviewing the 100 repositories the authors' model ranks highest prevents 14.9% of later adoptions of high-risk skills, against 0.5% for the 100 most starred.
  • The authors say platforms should distribute versioned references rather than copies.

Why it matters

Agent skills run with the permissions of the user, and the paper argues that they spread by copying with no registry, versions or provenance. That leaves three things unknown: where a copied skill came from, how far a security fix travels, and which repositories deserve review. A dated copy network is the first attempt in the paper's framing to answer who copied from whom, which single-snapshot studies cannot do.

Who it affects

Developers who use AI coding agents such as Claude Code and Codex and copy skills between repositories are the people exposed to this supply chain. Security engineers are the audience for the ranked list of repositories to check before a skill spreads. Platforms are the target of the recommendation to distribute versioned references instead of copies.

How to use it

The paper ships an interactive viewer alongside the network and lists a project website at https://fahdseddik.github.io/Skill-Constellations/. Security engineers can treat the model's ranking as a short audit list, starting from the repositories others copy from rather than the ones with the most stars. Platform builders can take the authors' advice to move from copies to versioned references.

How solid is it

The work is empirical: the network comes from the git history of every SKILL.md in GitSkills and covers 2,193,119 adoptions across GitHub. The headline comparison, 14.9% of later high-risk adoptions prevented by auditing the model's top 100 versus 0.5% for the 100 most starred, is the authors' own result as stated in the abstract, which this retelling rests on.

Risks and caveats

The abstract does not define 'high-risk' skills or say how high-risk status was determined. The 14.9% versus 0.5% result concerns later adoptions prevented by review; the abstract does not say how many repositories or skills were actually found malicious or vulnerable, and no actual security incident or exploit is reported. The abstract does not specify what the model is or its features, and it does not state the time span of the 2,193,119 adoptions. It gives no figure beyond 'almost never' and 'rarely' for how often copies track their source or fixes propagate.

“Developers share skills by copying them between repositories, which makes them a software supply chain without a registry, versions or provenance.”

— Skill Constellations paper abstract