RepoDaily · 2026-08-26 · AI model / Agent framework

Ponytail: The Agent Skill That Makes Claude Code Write 54% Less Code

#6 AI model / Agent framework JavaScript +944 DietrichGebert/ponytail Open repository

An MIT-licensed JavaScript skill that turns coding agents into a 'lazy senior dev' - benchmarked at -54% LOC, -20% cost, and 100% safety retention on a real FastAPI + React repo.

Repo typeAI model / Agent framework
Best forDevelopers running Claude Code, OpenCode, Qoder, or pi whose agents routinely produce oversized diffs on ordinary feature tickets
Risk levelMedium - a prompt-level behavior layer from a single maintainer, MIT-licensed, ships an uninstall script
Time to evaluateAbout an hour: install the npm package, run one ticket with and without the skill, compare the git diffs

Primary question: Can a prompt-only skill reliably shrink your agent's diffs on your repo, not just on the author's benchmark repo?

91/100

RepoDaily adoption score

RepoDaily rates this as 91/100 (strong) for adoption: evidence, installation path, production risk, differentiation, license clarity, and AI/agent fit are scored from the article sources and adoption notes.

Directional score from RepoDaily sources and adoption notes, not a benchmark.Risk: Medium
96Evidence quality

4 source(s) across 3 source category/categories, plus a RepoDaily-specific evidence module when available.

92Installability

5 workflow step(s), 5 next-action step(s), and 1 command/install signal(s) were detected.

68Maintenance confidence

Trending momentum is +944 stars, with maintenance/release/issue signals counted when present.

93Production readiness

Risk is marked medium, with 5 security note(s) and 4 explicit skip condition(s).

100Differentiation

3 opportunity lens item(s), 4 alternative(s), and 3 type-specific section(s) support differentiation.

82License clarity

License source or license wording is present.

100Agent / AI fit

7 AI/agent-related signal(s) were detected in the article text and metadata.

Project overview

Ponytail is a JavaScript package, @dietrichgebert/ponytail v4.9.0, that installs a persona into your coding agent: the long-ponytailed senior developer in oval glasses who has been at the company longer than version control. The README's one-line pitch is 'He says nothing. He writes one line. It works.' The canonical example asks for a date picker: a stock agent installs flatpickr, writes a wrapper component, adds a stylesheet, and starts a discussion about timezones. Ponytail's agent answers with <input type="date"> and the comment <!-- ponytail: browser has one -->.

What separates Ponytail from most prompt-engineering repositories is that it prints its measurement method next to its claims. The benchmark ran headless Claude Code sessions against tiangolo's full-stack-fastapi-template, a real FastAPI plus React repository, across twelve feature tickets, with the same agent running with and without the skill, four runs per task on Haiku 4.5, and scored everything on the git diff the session left behind. The mean result: 54% fewer lines of code, 22% fewer tokens, 20% lower cost, and time at 73% of the no-skill baseline.

The README also polices its own history. An earlier single-shot benchmark reported 80-94% as a flat figure; the current text states that against a fair agentic baseline that range is the per-task ceiling, not the average. Gains reach 94% where agents over-build, as with the date picker, and fall to near zero where the code is already minimal. A separate adversarial safety tier is the detail that earned attention: the baseline, a 'caveman' prompt, and Ponytail all scored 100%, while a bare 'write one-liners' prompt dropped to 95%. Ponytail keeps every safety guard.

Distribution is unusually tidy for the category. It is MIT-licensed, published to npm, and the README badge claims compatibility with 20 agents. package.json backs that breadth with shipped directories for OpenCode, Qoder, the Qoder plugin format, and a pi extension, alongside hooks/, skills/, AGENTS.md, an uninstall script, Spanish and Korean README translations, an examples/ directory of before-and-after 'survivors', and a reproduction setup under benchmarks/.

Problem it solves

  • Agents over-build routine UI: the README's date-picker story pulls in flatpickr, a wrapper component, a stylesheet, and a timezone debate for a control the browser already ships.
  • Every extra line is paid for three times: tokens spent generating it, review time reading it, and maintenance carrying it.
  • The obvious fix backfires: a bare 'write one-liners' prompt scored 95% on the adversarial safety tier, and the 'caveman' arm pushed tokens, cost, and time above the no-skill baseline.
  • Bloated output drifts away from platform primitives, so codebases accumulate wrappers around things browsers, React, or FastAPI already provide.

How it works

  1. Install @dietrichgebert/ponytail (v4.9.0) from npm; the entry point is .opencode/plugins/ponytail.mjs, also exported under ./plugin.
  2. The package drops AGENTS.md, hooks/, and skills/ plus agent-specific directories (.opencode/, .qoder/, .qoder-plugin/, pi-extension/) into place, so one install targets multiple agent platforms.
  3. The skill content biases the agent toward what already exists: platform primitives, standard-library calls, and the shortest diff that satisfies the ticket.
  4. Unlike a bare one-liner prompt, Ponytail retains the agent's existing safety guards: the adversarial tier scored it 100% versus 95% for the yagni-oneliner arm.
  5. Judge results on the git diff the session leaves behind, the same way the author's benchmark is scored, and remove everything with scripts/uninstall.js if the experiment disappoints.

Integration surface: what the npm package actually installs

  • Package: @dietrichgebert/ponytail v4.9.0, MIT, published to npm with public access.
  • Entry point ./.opencode/plugins/ponytail.mjs, re-exported at both "." and "./plugin".
  • Published files: AGENTS.md, hooks/, skills/, .opencode/, .qoder/, .qoder-plugin/, pi-extension/, scripts/uninstall.js, assets/, LICENSE.
  • pi integration is declared natively as extensions ./pi-extension/index.js and skills ./skills; npm keywords opencode-plugin, pi-package, and qoder match that layout.
  • GitHub topics add claude, claude-code, claude-code-plugin, and cursor-rules; the README badge claims 20 supported agents.
  • The test script chains node --test tests/*.test.js with npm test --prefix pi-extension and npm test --prefix ponytail-mcp, three tested components in one repo.

Try-it path: from README examples to your own A/B test

  • Skim examples/ for the survivors, starting with the date picker reduced to <input type="date">.
  • Read benchmarks/results/2026-06-18-agentic.md, the full writeup behind the -54% mean.
  • Open benchmarks/ for the reproduction setup: 12 tickets, n=4, Haiku 4.5, git-diff scoring.
  • Install v4.9.0 from npm, run one routine ticket with and without the skill, and diff the two outputs.
  • If it is not for you, scripts/uninstall.js ships in the package files list for clean removal.
  • A waitlist banner at ponytail.dev/soon signals more than the current package is planned.

Maintenance risk and how far the numbers travel

Everything points to a single maintainer: the author field in package.json is Dietrich Gebert, bugs route to GitHub issues, and the package already sits at 4.9.0 in a space where agent plugin interfaces keep moving. That pace is normal for the category, but updates track one person's attention.

The benchmark is self-reported and deliberately narrow: one repository (full-stack-fastapi-template), one model (Haiku 4.5), four runs per task. The README is upfront that gains span from 94% on over-built tasks down to near zero on already-minimal code, so your result depends on how much your agent currently over-builds.

The offsets are concrete: MIT license, a public npm package, an uninstall script in the published files, and benchmark artifacts you can rerun instead of trusting.

Who should pay attention?

Good fit if

  • You run Claude Code, OpenCode, Qoder, or pi and routinely receive diffs far larger than the ticket asked for.
  • Your stack has native primitives the agent keeps wrapping: browser controls like <input type="date">, framework utilities, standard-library functions.
  • Per-token cost is a line item: the -22% tokens and -20% cost in the benchmark came purely from writing less code.
  • You can run a with/without A/B on a handful of tickets and judge the git diff yourself.

Skip for now if

  • Codebases that are already minimal; the README concedes gains there are near zero.
  • Domains without safe one-line equivalents, where brevity would hide real complexity.
  • Locked-down environments that forbid third-party hooks or skills inside coding agents.
  • Buyers who need independently verified numbers; every figure here is author-run.

Risks and cautions

Medium

MIT-licensed with an uninstall script, but it is a single-maintainer prompt layer whose headline numbers come from author-run benchmarks on one repo and one model.

  • Sole maintainer (Dietrich Gebert per package.json) at version 4.9.0 in a fast-shifting agent-plugin space.
  • Benchmark scope: one repo (full-stack-fastapi-template), Haiku 4.5, n=4 per task, self-scored on git diffs.
  • It deliberately changes agent output; validate on your own tickets before rolling it out anywhere shared.
  • Mitigations: MIT license, public npm package, scripts/uninstall.js, and reproducible artifacts under benchmarks/.
  • LICENSE is MIT, Copyright (c) 2026 DietrichGebert, with the standard AS-IS, no-warranty disclaimer.
  • Safety retention is a measured claim, not a slogan: the adversarial tier scored baseline, caveman, and Ponytail at 100% and the bare yagni-oneliner prompt at 95%.
  • The package ships hooks/ and skills/ that run inside your agent; this review covers only README, LICENSE, and package.json, so audit those directories before use in sensitive repos.
  • scripts/uninstall.js is part of the published files, so removal is scripted rather than manual.
  • No telemetry or network behavior is documented in the source pack; undocumented does not mean absent, so check hooks/ locally.

Alternatives to compare

ApproachWhen to useTrade-off
Hand-written AGENTS.md / CLAUDE.md brevity rules
you want the same write-less pressure with zero third-party packagesFree; a few versioned lines
ESLint complexity and size rules
you would rather enforce diff discipline with CI gates than prompt nudges on JavaScriptFree, MIT-licensed
Biome
you want a single fast Rust-based linter and formatter capping JS/TS outputFree, open source
First-party agent configuration
you prefer tuning settings your agent vendor already ships over adding a pluginIncluded in existing agent subscriptions

What this trend reveals

Re-score Ponytail on your own repository

Every published number comes from one template repo and one model; your gain tracks your agent's current over-build rate, anywhere from 94% down to near zero.

Follow the setup in benchmarks/, swap in your repo, and run the same with/without comparison across a dozen tickets scored on git diff.

Package your house style with the same layout

package.json documents the recipe: skills/ plus AGENTS.md plus per-agent directories in one npm package, with an uninstall script included.

Copy the files layout (AGENTS.md, hooks/, skills/, per-agent folders), publish internally, and A/B agent output before and after.

Bank token savings where run volume is highest

The benchmark attributes -22% tokens and -20% cost purely to shorter code; the effect multiplies with the number of agent runs you execute.

Log tokens and cost per ticket for one week, enable Ponytail, and compare against the same ticket mix.

Best next action

Run one real ticket twice and score the git diff

The author's method is replicable in an afternoon: same agent, same ticket, with and without the skill, judged on the diff left behind.

  1. Install @dietrichgebert/ponytail v4.9.0 from npm into your agent.
  2. Pick a ticket your agent historically over-builds, such as a date picker or a settings form.
  3. Run it without the skill and record LOC, tokens, cost, and runtime; commit nothing.
  4. Run the same ticket with Ponytail enabled and record the same four numbers.
  5. Confirm tests and safety guards still pass, then compare your deltas against benchmarks/results/2026-06-18-agentic.md.

RepoDaily verdict

Ponytail packages one genuinely useful pressure, write less, into a clean MIT-licensed npm skill, and it measures itself more honestly than most: the old 80-94% flat claim is corrected to a per-task ceiling, near-zero gains are admitted, and safety retention is tested rather than asserted. The 944 stars this period ride on the date-picker one-liner as much as on the numbers. Read -54% as 'on a repo where the agent over-builds', run your own with/without A/B on the git diff, and keep scripts/uninstall.js close. If your agent's diffs routinely dwarf the ticket, this is one of the cheapest experiments available.

Sources