signalHacker News Show HN2026-09-18
Show HN: Cactus Needle 3: 8-29MB automation models can match DeepSeek V4 Flash
Cactus Needle 3 is an 8-29MB automation model (25-121M parameters at 2-bit) optimized for tool calls and structured JSON output, with the 20-layer version scoring 86.0 on Mobile Actions versus LFM2.5 1.2B (82.4), Qwen3.5 0.8B (76.0), and Apple's on-device model (57.6) at f16. It runs up to 4k tokens/sec decode and 10k prefill on a Raspberry Pi 5, supports multilingual input, and reaches DeepSeek v4 Flash-grade performance on narrow tasks with only 4L fine-tuning. The model ships with monarch Hadamard MLP layers at O(d√d) compute, calibrated confidence scores, regex-based grounding triggers, and runs across desktop, mobile, embedded, and WebAssembly platforms.
- why now
- New tiny models match DeepSeek
- topic
- AI Tech & New Models
- source
- Hacker News Show HN
score
score 8 out of 10. 0-10: how dense the facts are, multiplied by how much you can do with them after reading. 8+ means the topic's evidence bar is met: benchmarks and availability for a new model, amount and investors for a funding round, revenue figures for a solo-money story. Below 5 an item does not enter the digest. A press release scores 3 or less, a reprint loses 2, anything older than 14 days loses 1, a headline that misleads loses 3.
read the source