Primary question: Can a prompt-only skill reliably shrink your agent's diffs on your repo, not just on the author's benchmark repo?
RepoDaily adoption score
RepoDaily rates this as 91/100 (strong) for adoption: evidence, installation path, production risk, differentiation, license clarity, and AI/agent fit are scored from the article sources and adoption notes.
4 source(s) across 3 source category/categories, plus a RepoDaily-specific evidence module when available.
5 workflow step(s), 5 next-action step(s), and 1 command/install signal(s) were detected.
Trending momentum is +944 stars, with maintenance/release/issue signals counted when present.
Risk is marked medium, with 5 security note(s) and 4 explicit skip condition(s).
3 opportunity lens item(s), 4 alternative(s), and 3 type-specific section(s) support differentiation.
License source or license wording is present.
7 AI/agent-related signal(s) were detected in the article text and metadata.
Project overview
Ponytail is a JavaScript package, @dietrichgebert/ponytail v4.9.0, that installs a persona into your coding agent: the long-ponytailed senior developer in oval glasses who has been at the company longer than version control. The README's one-line pitch is 'He says nothing. He writes one line. It works.' The canonical example asks for a date picker: a stock agent installs flatpickr, writes a wrapper component, adds a stylesheet, and starts a discussion about timezones. Ponytail's agent answers with <input type="date"> and the comment <!-- ponytail: browser has one -->.
What separates Ponytail from most prompt-engineering repositories is that it prints its measurement method next to its claims. The benchmark ran headless Claude Code sessions against tiangolo's full-stack-fastapi-template, a real FastAPI plus React repository, across twelve feature tickets, with the same agent running with and without the skill, four runs per task on Haiku 4.5, and scored everything on the git diff the session left behind. The mean result: 54% fewer lines of code, 22% fewer tokens, 20% lower cost, and time at 73% of the no-skill baseline.
The README also polices its own history. An earlier single-shot benchmark reported 80-94% as a flat figure; the current text states that against a fair agentic baseline that range is the per-task ceiling, not the average. Gains reach 94% where agents over-build, as with the date picker, and fall to near zero where the code is already minimal. A separate adversarial safety tier is the detail that earned attention: the baseline, a 'caveman' prompt, and Ponytail all scored 100%, while a bare 'write one-liners' prompt dropped to 95%. Ponytail keeps every safety guard.
Distribution is unusually tidy for the category. It is MIT-licensed, published to npm, and the README badge claims compatibility with 20 agents. package.json backs that breadth with shipped directories for OpenCode, Qoder, the Qoder plugin format, and a pi extension, alongside hooks/, skills/, AGENTS.md, an uninstall script, Spanish and Korean README translations, an examples/ directory of before-and-after 'survivors', and a reproduction setup under benchmarks/.
Why it is trending now
- 944 stars in this trend period and the #6 spot on the 2026-08-26 list.
- Numbers with a printed method: -54% LOC mean, -22% tokens, -20% cost, and time at 73% of baseline, from 12 real feature tickets on full-stack-fastapi-template (Haiku 4.5, n=4).
- The date-picker one-liner (<input type="date">) is a shareable hook that communicates the whole idea in one screenshot.
- The badge claims it works with 20 agents; topics and npm keywords cover claude-code-plugin, cursor-rules, opencode-plugin, pi-package, and qoder.
- Unusual candor: the README downgrades its own earlier 80-94% claim to a per-task ceiling and admits gains are near zero on already-minimal code.
- MIT license, a public npm package, scripts/uninstall.js, and an open reproduction directory lower the cost of trying it to minutes.
Problem it solves
- Agents over-build routine UI: the README's date-picker story pulls in flatpickr, a wrapper component, a stylesheet, and a timezone debate for a control the browser already ships.
- Every extra line is paid for three times: tokens spent generating it, review time reading it, and maintenance carrying it.
- The obvious fix backfires: a bare 'write one-liners' prompt scored 95% on the adversarial safety tier, and the 'caveman' arm pushed tokens, cost, and time above the no-skill baseline.
- Bloated output drifts away from platform primitives, so codebases accumulate wrappers around things browsers, React, or FastAPI already provide.
How it works
- Install @dietrichgebert/ponytail (v4.9.0) from npm; the entry point is .opencode/plugins/ponytail.mjs, also exported under ./plugin.
- The package drops AGENTS.md, hooks/, and skills/ plus agent-specific directories (.opencode/, .qoder/, .qoder-plugin/, pi-extension/) into place, so one install targets multiple agent platforms.
- The skill content biases the agent toward what already exists: platform primitives, standard-library calls, and the shortest diff that satisfies the ticket.
- Unlike a bare one-liner prompt, Ponytail retains the agent's existing safety guards: the adversarial tier scored it 100% versus 95% for the yagni-oneliner arm.
- Judge results on the git diff the session leaves behind, the same way the author's benchmark is scored, and remove everything with scripts/uninstall.js if the experiment disappoints.
Integration surface: what the npm package actually installs
- Package: @dietrichgebert/ponytail v4.9.0, MIT, published to npm with public access.
- Entry point ./.opencode/plugins/ponytail.mjs, re-exported at both "." and "./plugin".
- Published files: AGENTS.md, hooks/, skills/, .opencode/, .qoder/, .qoder-plugin/, pi-extension/, scripts/uninstall.js, assets/, LICENSE.
- pi integration is declared natively as extensions ./pi-extension/index.js and skills ./skills; npm keywords opencode-plugin, pi-package, and qoder match that layout.
- GitHub topics add claude, claude-code, claude-code-plugin, and cursor-rules; the README badge claims 20 supported agents.
- The test script chains node --test tests/*.test.js with npm test --prefix pi-extension and npm test --prefix ponytail-mcp, three tested components in one repo.
Try-it path: from README examples to your own A/B test
- Skim examples/ for the survivors, starting with the date picker reduced to <input type="date">.
- Read benchmarks/results/2026-06-18-agentic.md, the full writeup behind the -54% mean.
- Open benchmarks/ for the reproduction setup: 12 tickets, n=4, Haiku 4.5, git-diff scoring.
- Install v4.9.0 from npm, run one routine ticket with and without the skill, and diff the two outputs.
- If it is not for you, scripts/uninstall.js ships in the package files list for clean removal.
- A waitlist banner at ponytail.dev/soon signals more than the current package is planned.
Maintenance risk and how far the numbers travel
Everything points to a single maintainer: the author field in package.json is Dietrich Gebert, bugs route to GitHub issues, and the package already sits at 4.9.0 in a space where agent plugin interfaces keep moving. That pace is normal for the category, but updates track one person's attention.
The benchmark is self-reported and deliberately narrow: one repository (full-stack-fastapi-template), one model (Haiku 4.5), four runs per task. The README is upfront that gains span from 94% on over-built tasks down to near zero on already-minimal code, so your result depends on how much your agent currently over-builds.
The offsets are concrete: MIT license, a public npm package, an uninstall script in the published files, and benchmark artifacts you can rerun instead of trusting.
Who should pay attention?
Good fit if
- You run Claude Code, OpenCode, Qoder, or pi and routinely receive diffs far larger than the ticket asked for.
- Your stack has native primitives the agent keeps wrapping: browser controls like <input type="date">, framework utilities, standard-library functions.
- Per-token cost is a line item: the -22% tokens and -20% cost in the benchmark came purely from writing less code.
- You can run a with/without A/B on a handful of tickets and judge the git diff yourself.
Skip for now if
- Codebases that are already minimal; the README concedes gains there are near zero.
- Domains without safe one-line equivalents, where brevity would hide real complexity.
- Locked-down environments that forbid third-party hooks or skills inside coding agents.
- Buyers who need independently verified numbers; every figure here is author-run.
Risks and cautions
MIT-licensed with an uninstall script, but it is a single-maintainer prompt layer whose headline numbers come from author-run benchmarks on one repo and one model.
- Sole maintainer (Dietrich Gebert per package.json) at version 4.9.0 in a fast-shifting agent-plugin space.
- Benchmark scope: one repo (full-stack-fastapi-template), Haiku 4.5, n=4 per task, self-scored on git diffs.
- It deliberately changes agent output; validate on your own tickets before rolling it out anywhere shared.
- Mitigations: MIT license, public npm package, scripts/uninstall.js, and reproducible artifacts under benchmarks/.
- LICENSE is MIT, Copyright (c) 2026 DietrichGebert, with the standard AS-IS, no-warranty disclaimer.
- Safety retention is a measured claim, not a slogan: the adversarial tier scored baseline, caveman, and Ponytail at 100% and the bare yagni-oneliner prompt at 95%.
- The package ships hooks/ and skills/ that run inside your agent; this review covers only README, LICENSE, and package.json, so audit those directories before use in sensitive repos.
- scripts/uninstall.js is part of the published files, so removal is scripted rather than manual.
- No telemetry or network behavior is documented in the source pack; undocumented does not mean absent, so check hooks/ locally.
Alternatives to compare
| Approach | When to use | Trade-off |
|---|---|---|
Hand-written AGENTS.md / CLAUDE.md brevity rules | you want the same write-less pressure with zero third-party packages | Free; a few versioned lines |
ESLint complexity and size rules | you would rather enforce diff discipline with CI gates than prompt nudges on JavaScript | Free, MIT-licensed |
Biome | you want a single fast Rust-based linter and formatter capping JS/TS output | Free, open source |
First-party agent configuration | you prefer tuning settings your agent vendor already ships over adding a plugin | Included in existing agent subscriptions |
What this trend reveals
Re-score Ponytail on your own repository
Every published number comes from one template repo and one model; your gain tracks your agent's current over-build rate, anywhere from 94% down to near zero.
Follow the setup in benchmarks/, swap in your repo, and run the same with/without comparison across a dozen tickets scored on git diff.
Package your house style with the same layout
package.json documents the recipe: skills/ plus AGENTS.md plus per-agent directories in one npm package, with an uninstall script included.
Copy the files layout (AGENTS.md, hooks/, skills/, per-agent folders), publish internally, and A/B agent output before and after.
Bank token savings where run volume is highest
The benchmark attributes -22% tokens and -20% cost purely to shorter code; the effect multiplies with the number of agent runs you execute.
Log tokens and cost per ticket for one week, enable Ponytail, and compare against the same ticket mix.
RepoDaily verdict
Ponytail packages one genuinely useful pressure, write less, into a clean MIT-licensed npm skill, and it measures itself more honestly than most: the old 80-94% flat claim is corrected to a per-task ceiling, near-zero gains are admitted, and safety retention is tested rather than asserted. The 944 stars this period ride on the date-picker one-liner as much as on the numbers. Read -54% as 'on a repo where the agent over-builds', run your own with/without A/B on the git diff, and keep scripts/uninstall.js close. If your agent's diffs routinely dwarf the ticket, this is one of the cheapest experiments available.