tabehod.ai / 食べ放題 / all you can eat
Unmetered agentic coding for you and 50 friends. (I need more friends.)
Pool money with your friends, rent a node, and serve an open-weight coding model on it. Each member gets an API key with a minimum guaranteed number of concurrent sessions and an encrypted prefix cache. The source and the infrastructure code are public.
Models
| Model | Shape | Hardware | Seat / month |
|---|---|---|---|
| Kimi K3 | 2.8T MoE, 104B active, 1M context, MXFP4 | 1 x 8xB300 (2.3 TB HBM) | $1,000 |
| Qwen3.8-27B | 27B dense, 262k context, FP8, Apache 2.0 | 1-2 x H200 or B200 | $100 |
K3 is the planner tier. Qwen3.8-27B is the worker tier. Want another model? Send me a PR to add whatever you (and your friends) want. Or fork the repo and start your own cohort.
"It costs how much?"
The 1.56 TB K3 checkpoint fits on one commodity node tier. An 8xB200 node has 1.5 TB of HBM and does not fit. Hopper nodes fit across four machines but lack native MXFP4 and run dequantization-bound (16xH200 measured at 16.8 tok/s single stream).
| Option | Nodes | HBM | Headroom after weights | Spot / month |
|---|---|---|---|---|
| 8xB300 | 1 | 2.3 TB | ~740 GB | ~$34k |
| 16xB200 | 2 | 3.1 TB | ~1.5 TB | ~$62k |
| 32xH100 | 4 | 2.6 TB | n/a | ~$68k |
| 32xH200 | 4 | 4.5 TB | n/a | ~$77k |
node 8xB300 secure, $7.89/GPU-hr x 8 x 730 h = $46.1k/mo seats 50 x $1,000 = $50.0k/mo fees Stripe ~ $1.5k/mo margin $50.0k - $46.1k - $1.5k ~ $2.4k/mo
Spot pricing (~$34k) leaves real margin for operating costs and overhead (cluster spin-up time). Secure on-demand pricing leaves almost none. The gap closes by pricing at $1,200, growing the cohort to 60, or landing a reserved rate.
Per seat
Decode on K3 is bound by memory bandwidth. Past a batch of about 60, every step touches nearly all 896 experts, so step cost stops growing with batch size.
step floor ~1.5 TB weights / ~64 TB/s HBM ~ 25 ms per stream 1 / 25 ms ~ 30-40 tok/s aggregate 2,000 tok/GPU-s x 8 GPUs (vLLM Pareto) ~ 16k tok/s per month 16k tok/s x 2.6M s ~ 42B output tok per seat 42B / 50 ~ 840M output tok list value 840M x $15/MTok > $12k
We are constrained by concurrency, not memory.
Heavy user consumption
Target members (me and my friends) have several $200/month subscriptions and have to juggle them between 5h and 7d limits. Ain't nobody got time for that.
one agent-hour 30 tok/s x 50% duty x 3,600 s ~ 54k output tok
20:1 input mix ~ 1.1M input tok
at 90% cache hit ~ $1.40
at 70% cache hit ~ $2.30
one engineer 6-10 agents x 6-8 h x 22 days ~ 800-1,750 agent-h
x $1.40-2.30 ~ $1.1k-4k/mo
fifty of them 50 x 54k x 800-1,750 agent-h ~ 2-7.5B output tok/mo
against ~42B node capacity ~ 5-18% utilization
It isn't totally about saving money. It's about unmetered usage.
Five guaranteed slots
Most agent wall-clock is tool execution. (You are unit testing, right?) An agent that is running tests or reading files holds a session slot but no decode step. Measured duty is about 30%.
knee aggregate stops scaling at ~ 200-300 active streams (guess) slots 300 streams / 30% duty ~ 900 open sessions per member 900 / 50, with headroom = 5 published, queue behind agents 5 slots / 30% duty ~ 8-10 agents kept busy
Felt latency is time-to-first-token after each tool result, so the prefix-cache hit rate is the user experience. Rate limits are per key: a slot semaphore plus a prompt-token rate, so one member's burst does not evict everyone else's cache.
KV pool bottleneck
K3 has 93 layers. 69 are linear attention (KDA) with a fixed ~54 MB state per request and no per-token cache. 24 are gated MLA with per-token KV. From a 16xH200 report (7.95 GB pool holding 308,608 tokens):
MLA KV 7.95 GB / 308,608 tok ~ 26 KB/token pool ~740 GB headroom, ~500 GB for KV ~ 19M resident tokens 700k session 700k x 26 KB ~ 18 GB eviction 18 GB over PCIe Gen5 at 50-60 GB/s ~ 0.3 s to host RAM cache miss whole-node prefill of 700k tokens ~ 30-90 s
Resident long contexts are cheap. Evicted ones are ruinous. So the KV tier writes through to host RAM (8xB300 nodes carry 2-4 TB) and eviction becomes a storage event, not a recompute.
The 19M figure holds only with data-parallel attention plus expert parallelism. Tensor-parallel attention replicates MLA KV across 8 ranks and divides the pool by 8. Three known inconsistencies remain in the v1 model: 5 slots x 128k exceeds a 380k per-member budget; a 500 x 48k sweep needs 24M tokens against the 19M pool; the KV size was derived from BF16 and serving is FP8.
$100 tier
Qwen3.8-27B is dense, so decode reads the whole ~27 GB of FP8 weights every step. One H200 has ~4.8 TB/s.
step floor 27 GB / 4.8 TB/s ~ 6 ms aggregate batched, conservative ~ 3-6k output tok/s per GPU worker demand 75% of agent-hours, ~30% duty ~ 1-2.7B output tok/mo hardware 1-2 x H200 or B200 ~ $3-6k/mo seats 50 x $100 = $5k/mo
In the planner/worker split, the 27B does 75% of agent-hours and K3 (or hosted Qwen3.8-Max at $2/$0.25/$6) does the 25% that needs a bigger model.
The vendor claims 27B beats Qwen3.7-Plus on real-world coding. No SWE-bench table has been published. It is being tested against qwen3.6-27b on real delegated briefs before any routing changes.
Encrypted per-key cache
Prefix-cache hashes are salted per tenant, so one member's cached prompt cannot be hit or timing-probed by another member's request. Disk-tier KV blobs are encrypted under a key derived from the member's API key. Revoke the key and the blobs are unreadable. Pre-load your codebase once and reuse it across sessions.
This protects members from each other. Analytics record counts, token totals, latency, and cache-hit ratio.
TBD
The real reserved price for an 8xB300 node. Aggregate tokens/second at 200-500 concurrent requests on Blackwell. KDA-state offload under load. Contention under many simultaneous evictions. Currently planning a concurrency sweep at 50, 100, 200, 300, and 500, and the published slot count is that cap divided by 50.