The Essentials

Card 1

Runs 24/7 with unlimited tokens.

Leave AI running in the background, without counting tokens or watching the bill.

Card 2

Pocket-sized and fully offline.

Your AI goes wherever you go, from long flights to remote field work.

Card 3

Runs 24/7 with unlimited tokens.

Leave AI running in the background, without counting tokens or watching the bill.

Card 4

Pocket-sized and fully offline.

Your AI goes wherever you go, from long flights to remote field work.

Up to 120 billion parameters
Fits in Your Palm
Models once reserved for servers are now brought to your pocket, driven by PowerInfer. Powerinfer v1 was the first infra to run 175B-parameter models on a consumer RTX 4090, reaching 90% of A100 performance and up to 11.69 times faster inference than prior methods. The new generation we adopt pushes even further.

Tested,
Not Claimed

Gemma 4 E4B (Int4)
Qwen 3.6 35B A3B (Int4)
Qwen3-coder-next 80B (Int4)
GPT-OSS 120B (Int4)
4K0Decode tok/s
64K0Decode tok/s
128K0Decode tok/s
0Prefill tok/s
0Prefill tok/s
0Prefill tok/s
TiinyOS
Made Intuitively Simple
From reasoning to coding, transcription to creativity. Ready out of the box, and always fresh with continuous SOTA updates. Run multiple models at once. Switch instantly. Or let them collaborate in the same workflow.

Model Strore

One Library for Every Task

From reasoning to coding, transcription to creativity. Ready out ofthe box, and always fresh with continuous SOTA updates. Runmultiple models at once. Switch instantly. Or let them collaboratein the same workflow.
From reasoning to coding, transcription to creativity. Ready out ofthe box, and always fresh with continuous SOTA updates. Runmultiple models at once. Switch instantly. Or let them collaboratein the same workflow.
View all models
Model Store

Straight to Work with Zero Config

Skip the setup. Pick an agent, handling everything from vibe coding and data analysis to automation.

Code Agent
Discover and install agents for writing, research, coding, analysis, and more. Or bring your own — Tiiny is open to any agent you upload.
Writer Agent
Discover and install agents for writing, research, coding, analysis, and more. Or bring your own — Tiiny is open to any agent you upload.
Research Agent
Discover and install agents for writing, research, coding, analysis, and more. Or bring your own — Tiiny is open to any agent you upload.
Analyst Agent
Discover and install agents for writing, research, coding, analysis, and more. Or bring your own — Tiiny is open to any agent you upload.
Creative Agent
Discover and install agents for writing, research, coding, analysis, and more. Or bring your own — Tiiny is open to any agent you upload.
Code Agent Writer Agent Research Agent Analyst Agent Creative Agent
Models

As Easy as Your Favorite Apps

Chat Chat UI

The Experience You
Already Know

Ask questions, upload files, and get instant answers—with the smoothness of a cloud AI service.

Search Search UI

Bridge local compute
with live world

Optionally connect to the web for real-time information retrieval and dynamic fact-checking.

Task Task UI

Complex problems solved
in one conversation

Simply state your goal. Tiiny coordinates the right agents to handle the execution with step-by-step status tracking.

Get 24/7 autonomous assistance

Link your daily tools to automate replies, manage schedules, and get briefings. Leverage your existing workflows via native MCP support anytime.

Scheduled Automation

Set it and forget it. Easily schedule recurring tasks to run in the background, freeing you for what matters.

Tiiny Vault

Own Your AI Brain

The private knowledge base built exclusively for you, with all data stored on-device. Preset your preferences and let it evolve with every interaction via manual edits, live chat requests, or automated summaries.

No Rebuilding, No Lock-in

Your Data
For Your Eyes Only
Device access is solely yours. All data becomes instantly unreadable if the drive is removed. Zero-compromise security for proprietary code, research, and knowledge.
Security Encryption
  • 100% on-device processing.
  • AES-256 full-disk encryption.
  • User-controlled keys.

Stay Chill, Outrun Boundaries

Up to 3x Energy Efficiency

Let every single watt deliver 3x tokens per second. Tiiny's 30W TDP slays the power-hungry monster of local LLM deployments. Same intelligence, stress-free running.

Tiiny AI Pocket

30W Active Running Power

Eval Decode Tok/s Tok/s·W
Gemma 4 E4B (Int4) 35-42 1.17-1.4
Qwen 3.6 35B A3B (Int4) 28-35 0.93-1.17
GPT OSS 120B (Int4) 15-20 0.5-0.67

AMD RyzenTM AI Halo 128GB

30W Active Running Power

Eval Decode Tok/s Tok/s·W
Gemma 4 E4B (Int4) 50-70 0.42-0.58
Qwen 3.6 35B A3B (Int4) 35-50 0.29-0.42
GPT OSS 120B (Int4) 35-55 0.29-0.46

Tiiny AI Pocket

30W Active Running Power

Eval Decode Tok/s Tok/s·W
Gemma 4 E4B (Int4) 60-80 0.43-0.57
Qwen 3.6 35B A3B (Int4) 50-65 0.38-0.46
GPT-OSS 120B (Int4) 62-80 0.44-0.57

Heavy Workloads,
Silently Handled

Tiiny AI Pocket delivers smooth performance at under 35 dB, made possible by an ultra-thin vapor chamber, dual fans, and an integrated fin-and-fan cooling design engineered to eliminate localized heat.

Cooling System

Packed Tight,
Travel Light

Product Exploded View

Specifications

LLMs
GPT-OSS-120B Llama3.1-8B
Qwen3-30b
Avg. Output Speed
18 - 40 tokens/s
Processor / NPU
CPU(arm v9.2) + NPU
30 INT8 TOPS
dNPU
160 INT8 TOPS
Memory
80GB LPDDR5X @6400MT/s
Storage
1TB PCIe 4.0 SSD
Wireless
Wi-Fi 802.11ax · BT 5.3 w/LE
Interfaces
Type-C × 3
Power
TDP 30W (65W adapter required)
Weight
300g
Dimensions
142 × 80 × 22 mm
Compatible System
macOS & Windows

FAQ

The estimated shipping date is August 2026. We are currently in the middle of production and feature optimization. We have set our delivery timeline for August to ensure we have ample time to complete essential regulatory certifications (such as the FCC) and deliver a thoroughly tested, high-quality product to you.

We have already started processing the $100 cashback. Because our own payment system is still being set up, and Shopify's anti-fraud/AML restrictions prevent us from refunding more than the original transaction amount, we temporarily have to use bank transfering as a workaround. You can find the link to submit your payout information in the relevant Kickstarter Update.

Please note that sharing your bank information is completely optional. If you have any privacy concerns, you can wait for alternative methods (like PayPal), which we expect to be available around late August.

In short, you have two options:
1. Receive it now: Use the temporary bank transfer option.
2. Keep details private: Wait for our alternative payout methods to launch later this summer.

Since our Kickstarter campaign has officially concluded, we plan to launch Tiiny on our official website this September with an MSRP of $1,999.
In the meantime, you can head over to https://tiiny.ai/ to join our waitlist, which will lock in an exclusive launch discount for you.

Tiiny is designed specifically for running LLMs and agents, and it works alongside your computer. It comes with its own interface (TiinyOS) and SDK, so you use it through a client rather than as a completely standalone device.

Yes, it is designed to work fully offline. You can run models, chat, generate images, and process your local files without any internet connection.

Yes. Although we don't currently have a dedicated client for Linux like we do for macOS or Windows, you can still run Tiiny AI Pocket on Linux via TiinySDK.

Tiiny AI Pocket supports running open-source models up to 120B (int4) parameters.

Tiiny is well suited for MoE-based large LLMs, with inference speed depending on model size and context length.
Under typical real world deployment with int4 quantization:
• 20B models: ~25 to 40 tokens/s
• 70B models: ~14 to 23 tokens/s
• 120B models: ~10 to 20 tokens/s

Note: Lower speeds correspond to 32k long context; higher speeds reflect 4k short context.

Right now you can expect around ~64K to 128K context in practice, depending on the model and setup.

Yes. You can use our conversion tool to adapt your own models into a Tiiny-compatible format. The conversion tool will be fully open-source and available on GitHub. While our team is manually adapting existing frameworks (like Qwen, Llama, etc.) to the ONNX pipeline, the open-source release enables the community to contribute and support additional architectures independently.

Not right now — Tiiny doesn't support multi-device clustering / distributed inference yet.

Tiiny is not optimized for heavy video generation or diffusion workloads, and performance will vary depending on the model and configuration. Results may not meet expectations for such tasks. For use cases primarily focused on video generation, a dedicated GPU system would be more suitable than Tiiny.

All your private data is stored locally on your Tiiny's internal SSD — it never leaves the device unless you choose to move it. All data is also end-to-end encrypted, secured with a unique encryption key that only you possess, created during device setup.

No. Most features in TiinyOS will never send your data outside Tiiny and your computer, except for the following two optional online features:
1. Web Search: Search engines utilize your prompts for data retrieval.
2. Email Binding: We collect your device ID and email address solely to enable password resets via verification codes.

Yes. TiinyOS includes a built-in toolkit that lets you easily back up data to external drives or your PC, export it in standard formats, or permanently delete everything in a few steps.

Cart

loading