RepoDaily · 2026-08-15 · Dataset / Public directory

Needle 2: a 45M-parameter tool-calling model in a 14MB binary

#8 Dataset / Public directory Python +661 cactus-compute/needle Open repository

Cactus Compute ships Needle 2, a 45M-parameter tool-calling model compressed into one 14MB binary that holds a session near 28MB of RAM, with pip-installable inference, LoRA tuning, and export.

Repo typeDataset / Public directory
Best forPython developers adding schema-constrained tool calling or structured extraction on phones, wearables, smart-home hubs, and robots
Risk levelMedium — version 2.0.0, a license-field mismatch, and single-vendor weights
Time to evaluate1-2 hours for the quickstart and one real tool

Primary question: Does a 14MB, 2-bit model meet your accuracy bar for tool calls and structured extraction?

90/100

RepoDaily adoption score

RepoDaily rates this as 90/100 (strong) for adoption: evidence, installation path, production risk, differentiation, license clarity, and AI/agent fit are scored from the article sources and adoption notes.

Directional score from RepoDaily sources and adoption notes, not a benchmark.Risk: Medium
100Evidence quality

4 source(s) across 4 source category/categories, plus a RepoDaily-specific evidence module when available.

100Installability

7 workflow step(s), 5 next-action step(s), and 3 command/install signal(s) were detected.

65Maintenance confidence

Trending momentum is +661 stars, with maintenance/release/issue signals counted when present.

93Production readiness

Risk is marked medium, with 5 security note(s) and 5 explicit skip condition(s).

100Differentiation

3 opportunity lens item(s), 4 alternative(s), and 4 type-specific section(s) support differentiation.

82License clarity

License source or license wording is present.

78Agent / AI fit

5 AI/agent-related signal(s) were detected in the article text and metadata.

Project overview

Needle 2, from Cactus Compute, is an open 45M-parameter model built for three jobs: tool calling, device use, and structured extraction. The entire model ships as a single 14MB binary with the weights baked in, and the README reports that a full session holds near 28MB of RAM. The team compressed it to CQ2-bit with their Cactus Quants tooling and packaged it with its own inference engine, so there are no separate model files to download, version, and manage.

The GitHub repository is the Python package: inference, LoRA fine-tuning, and export. Installing is one line — pip install cactus-needle — after which the inference engine is fetched once from Hugging Face and cached. The README says there is nothing else to build, and points to doc/apis.md for offline setup on air-gapped devices. The weights also live at huggingface.co/Cactus-Compute/needle2 for anyone who wants the artifacts directly.

On the benchmarks the project publishes, Needle 2 trades wins with FunctionGemma 270M, LFM2.5 230M, and Apple FM while being 5x to 70x smaller, running 2-bit weights against their f16. Those are vendor-reported numbers, so treat them as a starting point rather than a conclusion — but the size gap is the whole pitch. Models in this class usually cost hundreds of megabytes in f16.

Architecturally it is a Simple Attention Network, a dense small-model recipe with a Hadamard MLP in place of the FFN, GQA attention, engram key-value memory, and multi-lane hyper-connections, documented in a paper at arXiv:2607.18363. The design details matter less than the constraints they buy: bounded memory, constrained output, and a calibrated confidence score on every response.

Problem it solves

  • Tool-calling models small enough for phones and wearables are rare; most need hundreds of megabytes or a server round-trip.
  • Distributing a model usually means shipping and versioning separate weight files alongside your app.
  • Long conversations on small models blow past memory budgets as the KV cache grows.
  • Unconstrained small models emit malformed JSON or call tools with invented arguments.
  • Autonomous action needs a know-when-to-escalate signal, which raw logits do not provide.
  • Large tool catalogues overwhelm tiny context windows.

How it works

  1. Install with pip install cactus-needle (Python >= 3.9 per pyproject.toml). The first run fetches the inference engine from Hugging Face once and caches it.
  2. Describe tools by decorating functions with @needle.tool: the signature supplies argument types, the docstring becomes the tool description.
  3. Construct needle.Needle(tools=[...]) and call agent.run(query): the model picks the call, Needle executes your function, feeds the result back, and returns the final response with executed results attached under results.
  4. For structured extraction, pass a Pydantic model to needle.extract() and get a typed object back — the README demos an Invoice schema with vendor, total, and due_date fields.
  5. At decode time, a byte-level grammar compiled from your schemas constrains every token, and the retrieval head narrows the grammar to the top five tools for the turn.
  6. Confidence comes from a learned, calibrated head on every response, so high-confidence calls can run autonomously and low-confidence ones route to a human or a larger model.
  7. The same package covers LoRA fine-tuning and export, matching pyproject.toml's description of 'inference, LoRA finetuning, and build'.

Product demo and interface preview

Simple Attention Network architecture
Simple Attention Network architecture — The README's architecture figure shows how the Hadamard MLP, GQA attention, and engram key-value memory are arranged in each block. README.md image

Architecture read: what fits in 14MB

Needle 2 is a Simple Attention Network: a dense small-model recipe that swaps the usual FFN for a Hadamard MLP, uses GQA attention, adds engram key-value memory, and connects layers with multi-lane hyper-connections. The README links the design and ablations to arXiv:2607.18363. Each block carries its own update rule: the Walsh-Hadamard transform H is a fixed orthonormal matrix applied in n log n time with no weights to read; the (k, v) rows are gathered from hashed n-gram tables; routing logits are normalized into a doubly-stochastic matrix P by Sinkhorn iteration; the a, b, g and sigma gates are learned and input-dependent. Attention and MLP residuals are sandwich-normed and gated, and engram sites fire at two layers.

Two mechanisms do the practical work. Memory stays bounded through a 256-token sliding window with the tools pinned as KV sinks — that is why the session holds near 28MB regardless of conversation length. Output stays valid because decoding is constrained by a byte-level grammar compiled from the schemas you declare (text in, JSON out), and the retrieval head renders only the top five tools per turn with the grammar constrained to that subset. CQ2-bit quantization via Cactus Quants is what collapses 45M parameters into the 14MB binary.

The confidence gate is a learned, calibrated head rather than a heuristic: every response carries a score you can threshold. That single number is the line between 'the model guessed' and 'the device decided'.

Try-it path: quickstart in one sitting

  • pip install cactus-needle — package name and version 2.0.0 from pyproject.toml, requires Python >= 3.9.
  • The first run fetches the inference engine from Hugging Face once and caches it; nothing else to build. For air-gapped devices, the README points to doc/apis.md for offline setup.
  • The README's minimal tool loop: decorate get_weather(city: str) with @needle.tool, build needle.Needle(tools=[get_weather]), call agent.run("what's it like in Lagos right now?"), and read the results key — the example returns [{'city': 'Lagos', 'temp_c': 27, 'sky': 'clear'}].
  • Extraction path: define class Invoice(BaseModel) with vendor, total, and due_date, then call needle.extract(text) for a typed object.
  • The README is blunt that describing tools well 'is the whole game' — signatures give the types, docstrings are the descriptions, so budget time for writing them.
  • Optional extras from pyproject.toml: gpu installs jax[cuda12]; metal pins jax==0.4.38, jaxlib==0.4.38, jax-metal, flax==0.10.2, optax==0.2.4; test adds pytest and pydantic.

Maintenance read: verify before you depend on it

  • Version 2.0.0 — a fresh major version line, so expect churn in the Python API surface.
  • A license mismatch worth one email to the maintainers: the LICENSE file is MIT (Copyright (c) 2026 Cactus Compute) while pyproject.toml declares license = { text = "Apache-2.0" }. Both are permissive, but pick after it is resolved.
  • The dependency chain is JAX-heavy: huggingface_hub, numpy, jax, jaxlib, flax>=0.10.2, optax, sentencepiece. The metal extra pins exact versions (jax==0.4.38), signaling tight coupling to specific JAX releases.
  • Weights and engine come from Hugging Face (Cactus-Compute/needle2) on first run; air-gapped installs rely on doc/apis.md, which was not part of this source pack.
  • Tests exist — pyproject.toml sets testpaths to tests and defines a slow marker for end-to-end JAX build/finetune tests — but the source pack includes no CI configuration, issue tracker history, or changelog.
  • The comparisons against FunctionGemma 270M, LFM2.5 230M, and Apple FM are the vendor's own; reproduce them on your own tools before committing.

Integration surface: what your code touches

  • Python API: the @needle.tool decorator, the Needle class taking a tools list, run() for the tool-calling loop returning results, and extract() accepting a Pydantic model for typed output.
  • A CLI exists: pyproject.toml registers the console script needle = needle.cli:main.
  • Shipped assets per package-data: needle.model carries *.model and *.vocab files; needle.playground carries index.html, app.js, and style.css — a bundled browser playground directory.
  • Runtime footprint: a single 14MB engine binary, a session near 28MB RAM, inference with no network after setup, and a 256-token sliding window.

Who should pay attention?

Good fit if

  • On-device prototypes — phones, wearables, smart-home hubs, robots — where a ~28MB RAM session fits the budget.
  • Products that need tool calls constrained to declared schemas rather than free-form text.
  • Edge or air-gapped deployments: one-time fetch, cached engine, inference with no network.
  • Python services doing high-volume structured extraction where a hosted-LLM bill is the alternative.
  • Builders who need a confidence score to route between autonomous action and human escalation.

Skip for now if

  • Open-ended generation, long-form writing, or general chat — the model is tuned for tool calls and extraction.
  • Product requirements that need one unambiguous license today; resolve the MIT vs Apache-2.0 mismatch first.
  • Tool-calling jobs where the top-five-per-turn retrieval limit starves the model of tools it actually needs.
  • Non-Python runtimes — this repository is the Python package, and the source pack shows no other bindings.
  • Anything demanding f16-precision small models with headroom over the 2-bit trade.

Risks and cautions

Medium

A permissive-licensed, genuinely small model with a real Python surface, but at version 2.0.0 from one vendor, with a license-field mismatch, a JAX-pinned dependency chain, and vendor-reported benchmarks.

  • The LICENSE file says MIT; pyproject.toml says Apache-2.0 — both permissive, but the inconsistency should be resolved before legal review.
  • Version 2.0.0 sits on a fresh major version line; expect API churn.
  • The metal extras pin jax==0.4.38, jaxlib==0.4.38, flax==0.10.2, and optax==0.2.4 — tight coupling to specific releases.
  • Weights and engine are fetched from Hugging Face on first run; the source pack shows no signing or provenance details.
  • All benchmark claims (vs FunctionGemma 270M, LFM2.5 230M, Apple FM) come from the vendor's own README.
  • Inference does no network after the one-time engine fetch, per the README — favorable for devices that should not phone home.
  • Decoding is constrained by a byte-level grammar compiled from your schemas, which limits what tokens the model can emit at all.
  • The self-contained 14MB binary means no separate weight files to tamper with or mismatch — one artifact to hash and pin.
  • Review doc/apis.md before an air-gapped rollout; the source pack does not cover artifact verification or a security policy.
  • Tool execution still runs your Python: treat model-chosen arguments as untrusted input to any tool with side effects.

Alternatives to compare

ApproachWhen to useTrade-off
FunctionGemma 270M
You want a function-calling model with more capacity and can carry f16 weightsOpen weights; larger memory footprint
LFM2.5 230M
Mobile-class on-device inference where a bigger small model still fitsOpen weights; f16 at 5x or more of Needle's size
Apple FM
You ship inside Apple's platform stack and its model familyDistributed through Apple's channels
llama.cpp with GGUF small models
You need general-purpose local generation and broad runtime support rather than schema-constrained tool callsFree, MIT-licensed; you supply and quantize the model

What this trend reveals

Wearables that actually call tools

A session pinned near 28MB RAM fits devices where a 230M-parameter f16 model does not; declare the device's own functions as tools and the grammar keeps every call valid.

Port the engine to the target device, define 5-10 real tools, and measure tool-call accuracy and peak RAM on-device.

Edge extraction without a cloud bill

extract() with a Pydantic model turns unstructured text into typed objects locally, and the 14MB binary makes per-unit distribution trivial.

Run extraction over a few hundred of your real documents and score field-level accuracy against your current pipeline.

Confidence-gated automation

The calibrated confidence head yields one thresholdable number per response: act above it, escalate below it.

Replay logged requests, sweep thresholds, and plot escalation rate against error rate before wiring it to real actions.

Best next action

Run the quickstart, then swap in one real tool

One sitting is enough to test the core claims on your machine: install, replay the README's weather example, then replace it with one real tool from your product and watch the grammar and the confidence score behave.

  1. pip install cactus-needle and let the first run fetch and cache the engine from Hugging Face.
  2. Copy the README's get_weather example: @needle.tool, needle.Needle(tools=[...]), then read agent.run(...)['results'].
  3. Replace it with one real function from your codebase, docstring included.
  4. Try needle.extract() on a real document with a Pydantic model.
  5. Log the confidence scores and check the returned JSON shape against your schema.

RepoDaily verdict

Needle 2's pitch is unusually concrete — 45M parameters, one 14MB binary, ~28MB RAM sessions, grammar-constrained JSON, calibrated confidence — and the Python package makes those claims checkable in an afternoon. The medium risk sits in maturity, not design: a 2.0.0 version line, a MIT-vs-Apache-2.0 license mismatch, JAX version pins, and vendor-only benchmarks. If your device budget matches the memory ceiling, run the quickstart before believing the frontier chart.

Sources