tabehod.ai / 食べ放題 / all you can eat

Unmetered agentic coding for you and 50 friends. (I need more friends.)

Pool money with your friends, rent a node, and serve an open-weight coding model on it. Each member gets an API key with a minimum guaranteed number of concurrent sessions and an encrypted prefix cache. The source and the infrastructure code are public.

Models

Model Shape Hardware Seat / month
Kimi K3 2.8T MoE, 104B active, 1M context, MXFP4 1 x 8xB300 (2.3 TB HBM) $1,000
Qwen3.8-27B 27B dense, 262k context, FP8, Apache 2.0 1-2 x H200 or B200 $100

K3 is the planner tier. Qwen3.8-27B is the worker tier. Want another model? Send me a PR to add whatever you (and your friends) want. Or fork the repo and start your own cohort.

"It costs how much?"

The 1.56 TB K3 checkpoint fits on one commodity node tier. An 8xB200 node has 1.5 TB of HBM and does not fit. Hopper nodes fit across four machines but lack native MXFP4 and run dequantization-bound (16xH200 measured at 16.8 tok/s single stream).

Option Nodes HBM Headroom after weights Spot / month
8xB300 1 2.3 TB ~740 GB ~$34k
16xB200 2 3.1 TB ~1.5 TB ~$62k
32xH100 4 2.6 TB n/a ~$68k
32xH200 4 4.5 TB n/a ~$77k
node     8xB300 secure, $7.89/GPU-hr x 8 x 730 h  = $46.1k/mo
seats    50 x $1,000                              = $50.0k/mo
fees     Stripe                                   ~ $1.5k/mo
margin   $50.0k - $46.1k - $1.5k                  ~ $2.4k/mo

Spot pricing (~$34k) leaves real margin for operating costs and overhead (cluster spin-up time). Secure on-demand pricing leaves almost none. The gap closes by pricing at $1,200, growing the cohort to 60, or landing a reserved rate.

Per seat

Decode on K3 is bound by memory bandwidth. Past a batch of about 60, every step touches nearly all 896 experts, so step cost stops growing with batch size.

step floor   ~1.5 TB weights / ~64 TB/s HBM        ~ 25 ms
per stream   1 / 25 ms                              ~ 30-40 tok/s
aggregate    2,000 tok/GPU-s x 8 GPUs (vLLM Pareto) ~ 16k tok/s
per month    16k tok/s x 2.6M s                     ~ 42B output tok
per seat     42B / 50                               ~ 840M output tok
list value   840M x $15/MTok                        > $12k

We are constrained by concurrency, not memory.

Heavy user consumption

Target members (me and my friends) have several $200/month subscriptions and have to juggle them between 5h and 7d limits. Ain't nobody got time for that.

one agent-hour   30 tok/s x 50% duty x 3,600 s     ~ 54k output tok
                 20:1 input mix                    ~ 1.1M input tok
                 at 90% cache hit                  ~ $1.40
                 at 70% cache hit                  ~ $2.30
one engineer     6-10 agents x 6-8 h x 22 days     ~ 800-1,750 agent-h
                 x $1.40-2.30                      ~ $1.1k-4k/mo
fifty of them    50 x 54k x 800-1,750 agent-h      ~ 2-7.5B output tok/mo
                 against ~42B node capacity        ~ 5-18% utilization

It isn't totally about saving money. It's about unmetered usage.

Five guaranteed slots

Most agent wall-clock is tool execution. (You are unit testing, right?) An agent that is running tests or reading files holds a session slot but no decode step. Measured duty is about 30%.

knee        aggregate stops scaling at          ~ 200-300 active streams (guess)
slots       300 streams / 30% duty              ~ 900 open sessions
per member  900 / 50, with headroom             = 5 published, queue behind
agents      5 slots / 30% duty                  ~ 8-10 agents kept busy

Felt latency is time-to-first-token after each tool result, so the prefix-cache hit rate is the user experience. Rate limits are per key: a slot semaphore plus a prompt-token rate, so one member's burst does not evict everyone else's cache.

KV pool bottleneck

K3 has 93 layers. 69 are linear attention (KDA) with a fixed ~54 MB state per request and no per-token cache. 24 are gated MLA with per-token KV. From a 16xH200 report (7.95 GB pool holding 308,608 tokens):

MLA KV       7.95 GB / 308,608 tok                ~ 26 KB/token
pool         ~740 GB headroom, ~500 GB for KV     ~ 19M resident tokens
700k session 700k x 26 KB                         ~ 18 GB
eviction     18 GB over PCIe Gen5 at 50-60 GB/s   ~ 0.3 s to host RAM
cache miss   whole-node prefill of 700k tokens    ~ 30-90 s

Resident long contexts are cheap. Evicted ones are ruinous. So the KV tier writes through to host RAM (8xB300 nodes carry 2-4 TB) and eviction becomes a storage event, not a recompute.

The 19M figure holds only with data-parallel attention plus expert parallelism. Tensor-parallel attention replicates MLA KV across 8 ranks and divides the pool by 8. Three known inconsistencies remain in the v1 model: 5 slots x 128k exceeds a 380k per-member budget; a 500 x 48k sweep needs 24M tokens against the 19M pool; the KV size was derived from BF16 and serving is FP8.

$100 tier

Qwen3.8-27B is dense, so decode reads the whole ~27 GB of FP8 weights every step. One H200 has ~4.8 TB/s.

step floor   27 GB / 4.8 TB/s                     ~ 6 ms
aggregate    batched, conservative                ~ 3-6k output tok/s per GPU
worker demand 75% of agent-hours, ~30% duty      ~ 1-2.7B output tok/mo
hardware     1-2 x H200 or B200                   ~ $3-6k/mo
seats        50 x $100                            = $5k/mo

In the planner/worker split, the 27B does 75% of agent-hours and K3 (or hosted Qwen3.8-Max at $2/$0.25/$6) does the 25% that needs a bigger model.

The vendor claims 27B beats Qwen3.7-Plus on real-world coding. No SWE-bench table has been published. It is being tested against qwen3.6-27b on real delegated briefs before any routing changes.

Encrypted per-key cache

Prefix-cache hashes are salted per tenant, so one member's cached prompt cannot be hit or timing-probed by another member's request. Disk-tier KV blobs are encrypted under a key derived from the member's API key. Revoke the key and the blobs are unreadable. Pre-load your codebase once and reuse it across sessions.

This protects members from each other. Analytics record counts, token totals, latency, and cache-hit ratio.

TBD

The real reserved price for an 8xB300 node. Aggregate tokens/second at 200-500 concurrent requests on Blackwell. KDA-state offload under load. Contention under many simultaneous evictions. Currently planning a concurrency sweep at 50, 100, 200, 300, and 500, and the published slot count is that cap divided by 50.