Needle 2: 14MB Agentic LLM for Edge Devices

HenryNdubuaku· August 10, 2026 View original

Key takeaways

  • Needle 2 is a highly efficient 14MB LLM designed for edge devices.
  • It supports tool calling, structured extraction, and runs on low-power hardware.
  • The model can be fine-tuned for specific tasks and offers a confidence score for hybrid deployments.
  • It opens up new possibilities for AI on billions of connected devices.

Who benefits

Consumer ElectronicsIoTRoboticsAutomotiveHealthcare

Summary

Cactus has released Needle 2, a 14MB agentic LLM designed for phones, wearables, smart homes, and small robots. It offers efficient tool calling and structured extraction, running on low-power devices with high decode speeds and minimal RAM.

Cactus has launched Needle 2, an ultra-compact 14MB agentic Large Language Model specifically engineered for resource-constrained edge devices. This model, which operates within 28MB of RAM, targets a vast market of over 21 billion connected IoT devices, including budget smartphones, wearables, smart home devices, and small robots, many of which lack dedicated NPUs or powerful GPUs. Needle 2 excels in tasks like tool calling, device interaction, and structured data extraction, achieving impressive decode speeds of up to 1,500 tokens/sec on devices like the Meta Quest 3S and Raspberry Pi 5. Its efficiency is attributed to a novel architecture based on Simple Attention Networks, consuming significantly less power per token compared to other small LLMs. The model can be fine-tuned quickly on a Mac or PC using an automated data-generation pipeline, allowing for custom task optimization. It also incorporates a learned confidence score, enabling developers to decide whether to act on the model's output or escalate to a larger cloud-based model for higher-stakes scenarios.

Why it matters

This model significantly lowers the barrier for deploying advanced AI capabilities on a wide range of low-cost, low-power edge devices, enabling new applications in IoT, consumer electronics, and robotics.

How to implement this in your domain

  1. 1Integrate Needle 2 into existing IoT device firmware or mobile applications for on-device AI processing.
  2. 2Utilize the Python package to fine-tune Needle 2 with custom tool vocabularies or structured extraction schemas relevant to specific product functionalities.
  3. 3Implement a hybrid AI strategy by using Needle 2 for initial processing on-device and escalating tasks with low confidence scores to larger cloud LLMs.
  4. 4Explore its application for text classification or summarization by defining appropriate output schemas.

Original post by HenryNdubuaku

"Hey HN, Henry from Cactus here! We previously released Cactus Needle, a 14MB agentic LLM for tool call, device use, and structured extraction for phones, wearables, smart homes, small robots and microcontrollers. We got really great feedback here, and have now incorporated the su…"

View on X

Originally posted by HenryNdubuaku on X · view source

Want to go deeper?

Turn these trends into skills with Learnijoy's hands-on AI & tech courses.

Explore courses