Primary question: Can one localhost router safely hold your provider keys, auto-failover across 34 free tiers, and keep you under every cap?
RepoDaily adoption score
RepoDaily rates this as 91/100 (strong) for adoption: evidence, installation path, production risk, differentiation, license clarity, and AI/agent fit are scored from the article sources and adoption notes.
5 source(s) across 2 source category/categories, plus a RepoDaily-specific evidence module when available.
5 workflow step(s), 5 next-action step(s), and 3 command/install signal(s) were detected.
Trending momentum is +398 stars, with maintenance/release/issue signals counted when present.
Risk is marked medium, with 6 security note(s) and 4 explicit skip condition(s).
3 opportunity lens item(s), 5 alternative(s), and 4 type-specific section(s) support differentiation.
License source or license wording is present.
9 AI/agent-related signal(s) were detected in the article text and metadata.
Project overview
FreeLLMAPI is a TypeScript self-hosted aggregator that pools the free tiers of 34 LLM providers — 474 model families across 635 free endpoints — behind a single OpenAI-compatible /v1 endpoint. The README's pitch is arithmetic: every serious AI lab now offers a free tier of a few million tokens a month and a few thousand requests a day, which is a toy on its own, but stacked together they total roughly 7.4 billion tokens of monthly inference capacity. Beyond the free catalog, you can register any custom OpenAI-compatible chat, embedding, image, or audio endpoint.
The router picks the best available model for each request, fails over to the next provider when one is rate-limited, and tracks per-key usage so every free-tier cap stays respected. The model catalog updates itself from a signed feed: new free models, quota changes, and compatibility fixes land without a git pull. Free installs receive that catalog as a monthly snapshot — a model reaches them 30 days after it joins the live feed — while a $19/year premium router gets it the same day.
The reason this lands in the security-tool category is that the asset under management is your credential pile. Provider keys are encrypted at rest with AES-256-GCM in SQLite behind an ENCRYPTION_KEY, dashboard accounts use scrypt password hashing with session tokens, and the server binds 127.0.0.1 by default. SECURITY.md is unusually candid about what this is: a single-user, trusted-network tool with no multi-tenant auth by design, no bug bounty, and one maintainer.
Distribution is broad for a hobby project: a prebuilt Docker image on ghcr.io, an Electron desktop app for macOS and Windows, an Android app on Google Play, and even an experimental Termux install that uses Node's built-in SQLite driver. The dashboard ships 60 locales with a full Simplified Chinese translation, and the whole thing is MIT licensed, copyright 2026 Tashfeen Ahmed.
Why it is trending now
- 398 stars in the trend period put it at rank 16 on 2026-08-27.
- The headline math — roughly 7.4B free tokens per month across 34 providers and 635 endpoints — turns dozens of toy tiers into one usable inference budget.
- One /v1 endpoint with automatic failover removes the per-provider switching cost that kills most free-tier stacking attempts.
- AES-256-GCM encrypted key storage plus a localhost-by-default binding give the aggregator a security story most key-juggling scripts never had.
- Wide reach: ghcr.io Docker image, macOS/Windows desktop releases, a Google Play app, and documented recipes for Claude Code, Codex CLI, Cline, Continue, Aider, opencode, and Cursor.
Problem it solves
- Every serious AI lab ships a free tier, but each is capped at a few million tokens a month and a few thousand requests a day — useless alone for sustained work.
- Juggling dozens of provider keys means manual rate-limit bookkeeping and scattered plaintext .env files.
- Most clients and coding agents speak only the OpenAI chat format, while some tooling expects an Anthropic Messages surface.
- Free tiers change constantly — models appear, quotas shift — so a hand-maintained catalog rots quickly.
- Rotating keys safely requires encryption at rest and per-key usage tracking, which ad-hoc scripts rarely implement.
How it works
- Deploy: pull the ghcr.io image with Docker Compose per docs/install.md, install the desktop app, or run npm run dev — server on :3001, dashboard on :5173.
- Add provider keys through the dashboard; they are stored AES-256-GCM encrypted in SQLite and decrypted only via ENCRYPTION_KEY.
- Point any OpenAI-compatible client at the single /v1 endpoint using the unified freellmapi-… bearer key.
- The router selects a model per request using auto:* routing strategies and fails over to the next provider when one is rate-limited.
- Per-key usage tracking keeps each account inside its free-tier cap, and the signed catalog feed delivers model, quota, and compatibility updates without a git pull.
Product demo and interface preview




Security posture: what the policy actually promises
SECURITY.md is the most informative file in the repo because it draws hard lines. In scope: the /v1 proxy, /api/* admin routes, and /mcp endpoint; dashboard authentication (scrypt, session tokens) and the unified freellmapi-… API key; key handling including AES-256-GCM encryption at rest, ENCRYPTION_KEY usage, encrypted DB backups, and key import/export; the Electron desktop app's local data directory; and Docker packaging (Dockerfile, docker-compose.yml, install scripts) including anything that leaks secrets into image layers or logs. Premium license key validation and the signed catalog feed are also in scope.
Out of scope by design: vulnerabilities in upstream providers, and anything that requires the operator to deliberately expose the server to the public internet. The tool binds 127.0.0.1 by default, HOST_BIND=0.0.0.0 is a documented opt-in with a warning attached, and there is no multi-tenant auth. The policy states plainly that “I put it on a public IP and someone used my quota” is expected behavior, not a vulnerability.
The hardening guidance restates the two rules that matter most: protect ENCRYPTION_KEY and .env, because they decrypt every stored provider key and losing the key means losing the keys; and set ENCRYPTION_KEY explicitly in production rather than relying on the generated dev fallback. Docker users are told to re-pull :latest to stay on the patched 0.6.x line.
Command surface: the exact dev loop
- npm install, then npm run dev — server on :3001, dashboard on :5173, both with hot reload.
- npm run db:migration:up applies all schema migrations; migrations are file-per-migration under server/src/db/migrations/ and previously applied files must not be edited.
- npm run db:migration:create --name=add_embedding_index scaffolds a new migration; npm run db:migration:down rolls back.
- npm test runs the server vitest suite (plus client tests if present); every PR must include a test and keep the suite green.
- Running npm run check:i18n from client/ validates the 60 dashboard locales against en.json, the translation source of truth.
- Bootstrap scripts scripts/dev-bootstrap.sh and scripts/dev-bootstrap.ps1 install dependencies only when package-lock.json has changed and create .env when it is absent.
- Website-side assets include docs/index.html, install.sh (Unix Docker bootstrap), install.ps1 (PowerShell bootstrap), and success.html.
Maintenance risk: version lines and one pair of hands
- SECURITY.md says it directly: “This is a hobby project maintained by one person” — acknowledgement within a few days, explicitly not a 24-hour SLA.
- Only 0.6.x (the current line) receives security fixes; 0.5.x is best effort and 0.4.x and older get nothing. There are no backported patch releases, so Docker users must re-pull :latest.
- There is no bug bounty — the offer is credit in the release notes and a genuine thank-you.
- Free installs receive the catalog as a monthly snapshot, so a new model lands 30 days after it joins the live feed; same-day updates cost $19 per year at freellmapi.co.
- Contribution rules push back on unreviewed model output: “No invented facts. Provider rate limits, model ids, and endpoints must be verified against the provider” — a wrong rate limit in the catalog ships to everyone.
Integration surface: what talks to what
- Inbound: one OpenAI-compatible /v1 endpoint, an Anthropic Messages surface, an /mcp endpoint, and /api/* admin routes.
- Documented client recipes cover Claude Code, Codex CLI, Cline, Continue, Aider, opencode, and Cursor, plus editor autocomplete and a Context Handoff feature.
- Modalities: chat completions, streaming, tool calling, vision, Gemini Google Search grounding, embeddings, and image/audio via custom OpenAI-compatible providers.
- Extras: prompt compression with per-request controls and custom tool-output filters, and a Fetch Relay that routes provider HTTP requests through a user-controlled, streaming application-layer relay.
- Platforms: the ghcr.io Docker image, an Electron desktop app for macOS/Windows, a Google Play Android app, and an experimental Termux install using Node's built-in SQLite driver.
Adoption checklist before you trust it with keys
- Confirm you are running 0.6.x — older release lines stop receiving security fixes entirely.
- Set ENCRYPTION_KEY explicitly in .env and back it up somewhere safe; losing it means losing every stored provider key.
- Leave HOST_BIND unset so the server stays on 127.0.0.1; if you must widen it, put a reverse proxy with TLS and auth in front.
- Verify the models your stack depends on already appear in your catalog snapshot — free installs lag the live feed by up to 30 days.
- Check that every provider you add meets the project's own bar, stated in CONTRIBUTING.md: tiers that are genuinely free to start using without a credit card.
- If you self-build from source, run npm test; CONTRIBUTING.md expects the vitest suite green on every PR.
Who should pay attention?
Good fit if
- Solo developers and students who want free inference capacity for prototypes without adding a credit card anywhere.
- Coding-agent users — docs/clients.md ships recipes for Claude Code, Codex CLI, Cline, Continue, Aider, opencode, and Cursor.
- Self-hosters who keep services on localhost or a trusted LAN and can keep a Docker image current.
- Anyone who wants chat, embeddings, image, and audio behind one OpenAI-compatible endpoint.
Skip for now if
- Teams needing multi-tenant access control — the project ships none by design.
- Anyone planning to expose the endpoint to the public internet; that scenario is declared out of security scope.
- Organizations requiring SLA-backed vulnerability handling — support is one maintainer, best effort, with no bug bounty.
- Users who need new models same-day without paying; free installs wait up to 30 days.
Risks and cautions
The engineering is real and the security documentation is unusually honest, but a credential-holding service maintained by one person on a single supported release line is a medium risk for anything beyond personal use.
- SECURITY.md self-describes as a hobby project maintained by one person, with acknowledgement in days rather than hours and no bug bounty.
- Security fixes land only on 0.6.x; 0.5.x is best effort, 0.4.x and older are unsupported, and there are no backported patches.
- The service stores every provider key you own, so a leaked ENCRYPTION_KEY or .env compromises all of them at once.
- Free catalog updates lag the live feed by 30 days, so quota or compatibility fixes can arrive late.
- Provider API keys are encrypted at rest with AES-256-GCM in SQLite; ENCRYPTION_KEY is the single decryption secret and has only a generated dev fallback by default.
- Dashboard accounts use scrypt password hashing with session tokens; API access uses one unified freellmapi-… bearer key.
- Network default is 127.0.0.1; HOST_BIND=0.0.0.0 is an opt-in with a warning, and the docs recommend a reverse proxy with TLS and auth if exposure widens.
- Encrypted database backups and key import/export are explicitly in vulnerability scope, as are secrets leaking into Docker image layers or logs.
- Reporting goes through GitHub private vulnerability reporting or [email protected] with “[freellmapi security]” in the subject; public issues are discouraged until a fix is out.
- The Electron desktop app's local data directory is in scope, while upstream provider bugs and deliberately public deployments are excluded.
Alternatives to compare
| Approach | When to use | Trade-off |
|---|---|---|
LiteLLM | You want a widely deployed OpenAI-compatible proxy with provider routing and spend controls, maintained as a larger project. | Open source (MIT) with paid enterprise options |
LocalAI | You would rather run open-weight models on your own hardware than aggregate cloud free tiers. | Open source (MIT) |
OpenRouter | You want a hosted, zero-maintenance multi-model gateway and will pay per token. | Usage-based pricing with some free models |
Portkey | You need a managed AI gateway with observability and guardrails for a team. | Commercial tiers |
Direct provider SDKs | You use one or two providers and do not need failover or a unified endpoint. | Whatever free tier each provider grants |
What this trend reveals
Free capacity as a real budget line
7.4 billion tokens a month is enough to run prototypes, agents, and side projects at zero marginal cost — if the caps are respected automatically. The router's per-key usage tracking is the piece that makes the number usable rather than theoretical.
Route one real workload through /v1 for a week, then read the per-key usage stats and count how many provider caps were approached.
A cost floor for coding agents
The docs include recipes for Claude Code, Codex CLI, Cline, Continue, Aider, opencode, and Cursor, plus an /mcp endpoint — and agents are the heaviest token consumers most developers run.
Point one agent at /v1 with an auto: routing strategy and compare its weekly token spend against your current paid key.
Catalog freshness as the paid wedge
The project monetizes latency, not features: free installs see a model 30 days after the live feed, premium sees it day one for $19/year. That split tells you exactly what the money buys.
Check whether the specific models your stack depends on already appear in the monthly snapshot before deciding premium matters.
RepoDaily verdict
FreeLLMAPI is the rare aggregator that treats your provider keys as the primary asset: encrypted at rest, localhost by default, and documented with a security policy that names its own limits. Run it as designed — single-user, on a trusted machine, on 0.6.x — and 34 free tiers behave like one generous endpoint; ask it to be a shared service and it will politely refuse.