The financial outputs use two different thresholds on purpose. Operating cash is private billing minus the recurring costs entered here. Simple payback divides capex by positive operating cash. The capital-recovery target additionally asks the service to recover capex within the selected number of years. A plan can therefore make positive operating cash and have a finite payback while still missing a faster target. Financing costs, tax, working capital, depreciation, residual value and any staffing, software, insurance or network cost not entered above are excluded.
Observed coding workloads vary sharply: one 12,000-developer study reported a 51M-token monthly median and about 380M at P90, while an agent-task study found up to 30× variation between runs. The higher presets represent continuously operating agent-worker equivalents, not ordinary human seats. Start there, then replace them with your own exports.Review the developer workload study →Review the agent-task variance study →
Each model starts from a complete replica shape, not a model-size speed multiplier alone. gpt‑oss‑120b, MiniMax M3, Qwen3.5 and DeepSeek V4 Pro now use 8K/1K single-node InferenceX profiles; Nemotron 3 Super and Ultra, GLM‑5.2 and Kimi K2.7 use matched Lambda runs; DeepSeek V3.2 uses a GPUStack sweep; and Kimi K3 keeps its throughput–interactivity curve. The MiniMax and Qwen memory gates include both data-parallel checkpoint copies in their measured four-GPU cells. Mistral Small 4, Mistral Large 3 and DeepSeek V4 Flash remain explicitly labelled reference-topology proxies. A complete proxy cell is no longer divided by all eight GPUs merely because several small cells pack into one node.Inspect the InferenceX profiles →Nemotron 3 Super measured profile →Nemotron 3 Ultra measured profile →DeepSeek V3.2 measured sweep →GLM‑5.2 measured profile →Kimi K2.7 measured profile →Kimi K3 throughput curve →
Nemotron 3 Ultra needs a special translation: its public run uses 8,192 input / 65,536 output tokens. The published input rate is therefore the observed token share of a decode-heavy saturation test, not an independently saturated prefill ceiling. The preset retains that observed value for context but uses measured decode throughput, interactivity and memory as its hard capacity gates; replace the profile with a mixed-shape sweep before procurement.
The service targets separate user-visible output speed from total server throughput. Independent API measurements checked on 1 August 2026 observed about 32 output tok/s for Kimi K3, 66 for GPT-5.6 Sol at maximum effort and 74 for Claude Fable 5. The 30 and 70 targets are parity anchors, 100 is a fast interactive target and 15 is for background agents; none is a latency SLA. gpt‑oss‑120b, MiniMax M3, Qwen3.5 and DeepSeek V4 Pro load measured points from their InferenceX concurrency sweeps where available, while Kimi K3 uses its approximate published curve. High-volume workloads also reserve the parallel active requests implied by monthly output volume and measured per-request speed; the per-request KV assumption is multiplied by that request count. Every single-point benchmark and topology proxy changes aggregate throughput through a conservative cross-model envelope calibrated to those 8K/1K sweeps: 0.70×, 1.00×, 1.50× and 1.90× the normalized frontier point. Those values are planning estimates, not measurements for the selected model; whole-platform rounding or a tighter memory gate can still leave two adjacent targets at the same purchase count.Review the measured Kimi K3 API speed →Review the measured GPT-5.6 Sol API speed →Review the measured Claude Fable 5 API speed →
Prefix reuse changes the self-hosted prefill path, not output decoding. The simulator counts uncached input at full prefill work and cached input at the editable cache-hit allowance. Moonshot reports above 90% cache hits for Kimi K3 coding traffic, so its coding presets start at 90%; that provider result is not a promise for a different scheduler or workload. Cache lookup, state transfer, retention and eviction still consume memory and network capacity.Review Moonshot’s K3 cache claim and direct pricing →Review vLLM’s hybrid prefix-cache design →
Every preset assigns at least 20% of tokens to output; coding profiles use a 50/50 split so billed reasoning is not hidden. Replace workload, prefix reuse, cache-read work, aggregate throughput, concurrency, interactivity and per-request memory assumptions with telemetry from the same ISL/OSL test shape.
The same-model OpenRouter rate remains the grounding benchmark even when another API alternative is selected. Rates are stored as planning inputs checked on 1 August 2026 rather than fetched live; recheck the linked model page before quoting. Direct API comparisons exclude tool calls, cache writes, cache storage, long-context or regional uplifts and negotiated discounts. All prices exclude VAT. Review the platform evidence →