RepoDaily · 2026-08-27 · Security tool

FreeLLMAPI: One Local /v1 Router for 34 Free LLM Providers and 635 Endpoints

#16 Security tool TypeScript +398 tashfeenahmed/freellmapi Open repository

A TypeScript self-hosted router stacks ~7.4B free tokens a month across 635 endpoints behind one /v1 API, with AES-256-GCM key storage and an explicitly single-user security model.

Repo typeSecurity tool
Best forSolo developers and coding-agent users who want every free LLM tier behind one local OpenAI-compatible endpoint, with provider keys encrypted at rest
Risk levelMedium — holds real credentials, single maintainer, security fixes only on 0.6.x
Time to evaluate1–2 hours: Docker Compose up, add keys, point one client at /v1

Primary question: Can one localhost router safely hold your provider keys, auto-failover across 34 free tiers, and keep you under every cap?

91/100

RepoDaily adoption score

RepoDaily rates this as 91/100 (strong) for adoption: evidence, installation path, production risk, differentiation, license clarity, and AI/agent fit are scored from the article sources and adoption notes.

Directional score from RepoDaily sources and adoption notes, not a benchmark.Risk: Medium
98Evidence quality

5 source(s) across 2 source category/categories, plus a RepoDaily-specific evidence module when available.

100Installability

5 workflow step(s), 5 next-action step(s), and 3 command/install signal(s) were detected.

62Maintenance confidence

Trending momentum is +398 stars, with maintenance/release/issue signals counted when present.

96Production readiness

Risk is marked medium, with 6 security note(s) and 4 explicit skip condition(s).

100Differentiation

3 opportunity lens item(s), 5 alternative(s), and 4 type-specific section(s) support differentiation.

82License clarity

License source or license wording is present.

90Agent / AI fit

9 AI/agent-related signal(s) were detected in the article text and metadata.

Project overview

FreeLLMAPI is a TypeScript self-hosted aggregator that pools the free tiers of 34 LLM providers — 474 model families across 635 free endpoints — behind a single OpenAI-compatible /v1 endpoint. The README's pitch is arithmetic: every serious AI lab now offers a free tier of a few million tokens a month and a few thousand requests a day, which is a toy on its own, but stacked together they total roughly 7.4 billion tokens of monthly inference capacity. Beyond the free catalog, you can register any custom OpenAI-compatible chat, embedding, image, or audio endpoint.

The router picks the best available model for each request, fails over to the next provider when one is rate-limited, and tracks per-key usage so every free-tier cap stays respected. The model catalog updates itself from a signed feed: new free models, quota changes, and compatibility fixes land without a git pull. Free installs receive that catalog as a monthly snapshot — a model reaches them 30 days after it joins the live feed — while a $19/year premium router gets it the same day.

The reason this lands in the security-tool category is that the asset under management is your credential pile. Provider keys are encrypted at rest with AES-256-GCM in SQLite behind an ENCRYPTION_KEY, dashboard accounts use scrypt password hashing with session tokens, and the server binds 127.0.0.1 by default. SECURITY.md is unusually candid about what this is: a single-user, trusted-network tool with no multi-tenant auth by design, no bug bounty, and one maintainer.

Distribution is broad for a hobby project: a prebuilt Docker image on ghcr.io, an Electron desktop app for macOS and Windows, an Android app on Google Play, and even an experimental Termux install that uses Node's built-in SQLite driver. The dashboard ships 60 locales with a full Simplified Chinese translation, and the whole thing is MIT licensed, copyright 2026 Tashfeen Ahmed.

Problem it solves

  • Every serious AI lab ships a free tier, but each is capped at a few million tokens a month and a few thousand requests a day — useless alone for sustained work.
  • Juggling dozens of provider keys means manual rate-limit bookkeeping and scattered plaintext .env files.
  • Most clients and coding agents speak only the OpenAI chat format, while some tooling expects an Anthropic Messages surface.
  • Free tiers change constantly — models appear, quotas shift — so a hand-maintained catalog rots quickly.
  • Rotating keys safely requires encryption at rest and per-key usage tracking, which ad-hoc scripts rarely implement.

How it works

  1. Deploy: pull the ghcr.io image with Docker Compose per docs/install.md, install the desktop app, or run npm run dev — server on :3001, dashboard on :5173.
  2. Add provider keys through the dashboard; they are stored AES-256-GCM encrypted in SQLite and decrypted only via ENCRYPTION_KEY.
  3. Point any OpenAI-compatible client at the single /v1 endpoint using the unified freellmapi-… bearer key.
  4. The router selects a model per request using auto:* routing strategies and fails over to the next provider when one is rate-limited.
  5. Per-key usage tracking keeps each account inside its free-tier cap, and the signed catalog feed delivers model, quota, and compatibility updates without a git pull.

Product demo and interface preview

Feature overview diagram of the FreeLLMAPI router
Feature overview — The README's feature map shows how routing, failover, and encrypted key storage fit together before you install anything. README.md image
Screenshot of the FreeLLMAPI desktop app
FreeLLMAPI desktop app — The desktop build packages the same dashboard for macOS and Windows, so key management does not require a browser tab or Docker. README.md image
Comparison table of FreeLLMAPI against OpenRouter, LiteLLM, and Portkey
Feature comparison against OpenRouter, LiteLLM, and Portkey — The project's own comparison table frames the choice between self-hosted free-tier pooling and managed gateways like OpenRouter, LiteLLM, and Portkey. README.md image
Stacked chart of roughly 7.4 billion free tokens per month across 34 providers
The free tier, stacked — ~7.4B tokens of free inference per month across 34 providers — This stacked view is the core pitch: no single free tier is usable, but 34 of them add up to about 7.4 billion tokens a month. README.md image

Security posture: what the policy actually promises

SECURITY.md is the most informative file in the repo because it draws hard lines. In scope: the /v1 proxy, /api/* admin routes, and /mcp endpoint; dashboard authentication (scrypt, session tokens) and the unified freellmapi-… API key; key handling including AES-256-GCM encryption at rest, ENCRYPTION_KEY usage, encrypted DB backups, and key import/export; the Electron desktop app's local data directory; and Docker packaging (Dockerfile, docker-compose.yml, install scripts) including anything that leaks secrets into image layers or logs. Premium license key validation and the signed catalog feed are also in scope.

Out of scope by design: vulnerabilities in upstream providers, and anything that requires the operator to deliberately expose the server to the public internet. The tool binds 127.0.0.1 by default, HOST_BIND=0.0.0.0 is a documented opt-in with a warning attached, and there is no multi-tenant auth. The policy states plainly that “I put it on a public IP and someone used my quota” is expected behavior, not a vulnerability.

The hardening guidance restates the two rules that matter most: protect ENCRYPTION_KEY and .env, because they decrypt every stored provider key and losing the key means losing the keys; and set ENCRYPTION_KEY explicitly in production rather than relying on the generated dev fallback. Docker users are told to re-pull :latest to stay on the patched 0.6.x line.

Command surface: the exact dev loop

  • npm install, then npm run dev — server on :3001, dashboard on :5173, both with hot reload.
  • npm run db:migration:up applies all schema migrations; migrations are file-per-migration under server/src/db/migrations/ and previously applied files must not be edited.
  • npm run db:migration:create --name=add_embedding_index scaffolds a new migration; npm run db:migration:down rolls back.
  • npm test runs the server vitest suite (plus client tests if present); every PR must include a test and keep the suite green.
  • Running npm run check:i18n from client/ validates the 60 dashboard locales against en.json, the translation source of truth.
  • Bootstrap scripts scripts/dev-bootstrap.sh and scripts/dev-bootstrap.ps1 install dependencies only when package-lock.json has changed and create .env when it is absent.
  • Website-side assets include docs/index.html, install.sh (Unix Docker bootstrap), install.ps1 (PowerShell bootstrap), and success.html.

Maintenance risk: version lines and one pair of hands

  • SECURITY.md says it directly: “This is a hobby project maintained by one person” — acknowledgement within a few days, explicitly not a 24-hour SLA.
  • Only 0.6.x (the current line) receives security fixes; 0.5.x is best effort and 0.4.x and older get nothing. There are no backported patch releases, so Docker users must re-pull :latest.
  • There is no bug bounty — the offer is credit in the release notes and a genuine thank-you.
  • Free installs receive the catalog as a monthly snapshot, so a new model lands 30 days after it joins the live feed; same-day updates cost $19 per year at freellmapi.co.
  • Contribution rules push back on unreviewed model output: “No invented facts. Provider rate limits, model ids, and endpoints must be verified against the provider” — a wrong rate limit in the catalog ships to everyone.

Integration surface: what talks to what

  • Inbound: one OpenAI-compatible /v1 endpoint, an Anthropic Messages surface, an /mcp endpoint, and /api/* admin routes.
  • Documented client recipes cover Claude Code, Codex CLI, Cline, Continue, Aider, opencode, and Cursor, plus editor autocomplete and a Context Handoff feature.
  • Modalities: chat completions, streaming, tool calling, vision, Gemini Google Search grounding, embeddings, and image/audio via custom OpenAI-compatible providers.
  • Extras: prompt compression with per-request controls and custom tool-output filters, and a Fetch Relay that routes provider HTTP requests through a user-controlled, streaming application-layer relay.
  • Platforms: the ghcr.io Docker image, an Electron desktop app for macOS/Windows, a Google Play Android app, and an experimental Termux install using Node's built-in SQLite driver.

Adoption checklist before you trust it with keys

  • Confirm you are running 0.6.x — older release lines stop receiving security fixes entirely.
  • Set ENCRYPTION_KEY explicitly in .env and back it up somewhere safe; losing it means losing every stored provider key.
  • Leave HOST_BIND unset so the server stays on 127.0.0.1; if you must widen it, put a reverse proxy with TLS and auth in front.
  • Verify the models your stack depends on already appear in your catalog snapshot — free installs lag the live feed by up to 30 days.
  • Check that every provider you add meets the project's own bar, stated in CONTRIBUTING.md: tiers that are genuinely free to start using without a credit card.
  • If you self-build from source, run npm test; CONTRIBUTING.md expects the vitest suite green on every PR.

Who should pay attention?

Good fit if

  • Solo developers and students who want free inference capacity for prototypes without adding a credit card anywhere.
  • Coding-agent users — docs/clients.md ships recipes for Claude Code, Codex CLI, Cline, Continue, Aider, opencode, and Cursor.
  • Self-hosters who keep services on localhost or a trusted LAN and can keep a Docker image current.
  • Anyone who wants chat, embeddings, image, and audio behind one OpenAI-compatible endpoint.

Skip for now if

  • Teams needing multi-tenant access control — the project ships none by design.
  • Anyone planning to expose the endpoint to the public internet; that scenario is declared out of security scope.
  • Organizations requiring SLA-backed vulnerability handling — support is one maintainer, best effort, with no bug bounty.
  • Users who need new models same-day without paying; free installs wait up to 30 days.

Risks and cautions

Medium

The engineering is real and the security documentation is unusually honest, but a credential-holding service maintained by one person on a single supported release line is a medium risk for anything beyond personal use.

  • SECURITY.md self-describes as a hobby project maintained by one person, with acknowledgement in days rather than hours and no bug bounty.
  • Security fixes land only on 0.6.x; 0.5.x is best effort, 0.4.x and older are unsupported, and there are no backported patches.
  • The service stores every provider key you own, so a leaked ENCRYPTION_KEY or .env compromises all of them at once.
  • Free catalog updates lag the live feed by 30 days, so quota or compatibility fixes can arrive late.
  • Provider API keys are encrypted at rest with AES-256-GCM in SQLite; ENCRYPTION_KEY is the single decryption secret and has only a generated dev fallback by default.
  • Dashboard accounts use scrypt password hashing with session tokens; API access uses one unified freellmapi-… bearer key.
  • Network default is 127.0.0.1; HOST_BIND=0.0.0.0 is an opt-in with a warning, and the docs recommend a reverse proxy with TLS and auth if exposure widens.
  • Encrypted database backups and key import/export are explicitly in vulnerability scope, as are secrets leaking into Docker image layers or logs.
  • Reporting goes through GitHub private vulnerability reporting or [email protected] with “[freellmapi security]” in the subject; public issues are discouraged until a fix is out.
  • The Electron desktop app's local data directory is in scope, while upstream provider bugs and deliberately public deployments are excluded.

Alternatives to compare

ApproachWhen to useTrade-off
LiteLLM
You want a widely deployed OpenAI-compatible proxy with provider routing and spend controls, maintained as a larger project.Open source (MIT) with paid enterprise options
LocalAI
You would rather run open-weight models on your own hardware than aggregate cloud free tiers.Open source (MIT)
OpenRouter
You want a hosted, zero-maintenance multi-model gateway and will pay per token.Usage-based pricing with some free models
Portkey
You need a managed AI gateway with observability and guardrails for a team.Commercial tiers
Direct provider SDKs
You use one or two providers and do not need failover or a unified endpoint.Whatever free tier each provider grants

What this trend reveals

Free capacity as a real budget line

7.4 billion tokens a month is enough to run prototypes, agents, and side projects at zero marginal cost — if the caps are respected automatically. The router's per-key usage tracking is the piece that makes the number usable rather than theoretical.

Route one real workload through /v1 for a week, then read the per-key usage stats and count how many provider caps were approached.

A cost floor for coding agents

The docs include recipes for Claude Code, Codex CLI, Cline, Continue, Aider, opencode, and Cursor, plus an /mcp endpoint — and agents are the heaviest token consumers most developers run.

Point one agent at /v1 with an auto: routing strategy and compare its weekly token spend against your current paid key.

Catalog freshness as the paid wedge

The project monetizes latency, not features: free installs see a model 30 days after the live feed, premium sees it day one for $19/year. That split tells you exactly what the money buys.

Check whether the specific models your stack depends on already appear in the monthly snapshot before deciding premium matters.

Best next action

Stand it up on localhost and point one coding agent at it for a week

The cheapest decisive test is a single-machine deployment with two or three provider keys you already own, keeping every default that matters: localhost binding, an explicit ENCRYPTION_KEY, and one documented client recipe.

  1. Start the router with Docker Compose per docs/install.md, or run npm run dev for the server on :3001 and dashboard on :5173.
  2. Set ENCRYPTION_KEY explicitly in .env and back it up; leave HOST_BIND unset so it stays on 127.0.0.1.
  3. Add keys for two or three providers with genuinely free, no-credit-card tiers and confirm they store encrypted.
  4. Configure one documented client — the Claude Code or Codex CLI recipe in docs/clients.md — against the /v1 endpoint.
  5. After a week, review per-key usage and failover behavior, then decide whether the 30-day catalog lag justifies the $19/year tier.

RepoDaily verdict

FreeLLMAPI is the rare aggregator that treats your provider keys as the primary asset: encrypted at rest, localhost by default, and documented with a security policy that names its own limits. Run it as designed — single-user, on a trusted machine, on 0.6.x — and 34 free tiers behave like one generous endpoint; ask it to be a shared service and it will politely refuse.

Sources