⚡️ DeepSeek V4 Flash, 14 providers, 12,795 calls. The cheapest option isn't always cheapest.
Prompt caching flips the math completely. On DeepSeek's own endpoint, a repeated 100k-token prompt with a 100-token output costs ~15x less than cold. Your "budget" provider might not look so cheap once your prompts are warm.
Open models give you freedom to shop around. But that freedom's only useful if you actually measure. Someone did. Source

inference.academy
DeepSeek V4 Flash across 14 providers: cost, speed and caching
DeepSeek V4 Flash sent down 14 serving routes in 12,795 controlled calls at 1K, 10K and 100K input: cost, first-token latency, generation speed, cache share and failure rate per route.
September 7, 2026 274 2