Show HN: Needle2: 14MB agentic LLM for phones, wearables, smart home and robots — editorial image
Signal Ledger illustration · Generated
Technology

Show HN: Needle2: 14MB agentic LLM for phones, wearables, smart home and robots

The update

We previously released Cactus Needle, a 14MB agentic LLM for tool call, device use, and structured extraction for phones, wearables, smart homes, small robots and microcontrollers. We got really great feedback here, and have now incorporated the suggestions to release Needle 2.

Why this counts

The whole model is a single 14MB binary that runs a full session in 28MB of RAM; 45m parameters at 2bit compression. Needle hits 500 tokens/sec decode speed on a Raspberry Pi 5, sits between 400-1,500 tokens/sec on VR devices like Meta Quest 3S and Apple Vision Pro, and ranges 300-700 on sub-$200 phones such as the Samsung A-Series.

What precedes this

On the tool call and mobile device use benchmarks, Needle 2 trades wins with closest small models like LFM2

Where this fits in Signal Ledger

Related coverage from the Technology desk.

The lens

Needle hits 500 tokens/sec decode speed on a Raspberry Pi 5, sits between 400-1,500 tokens/sec on VR devices like Meta Quest 3S and Apple Vision Pro, and ranges 300-700 on sub-$200 phones such as the Samsung A-Series. 5 230M and Apple Foundation Model, at 5x to 70x smaller, both at f16 vs Needle 2 at 2bit.

Source note

Hacker News reporting: https://cactuscompute.com/needle

Read the original reporting