Guido Appenzeller
@appenz
It's kind of ridiculous that you can get a model that is competitive with early 2026 frontier models for under 1.5 cents per million tokens.
LLmflation is real. And go @relace_ai!
LLmflation is real. And go @relace_ai!
Eitan Borgnia@EBorgnia · Sep 24Relace is now also the cheapest provider of DeepSeek v4.1 Flash on OpenRouter.
Agent traces are 96%+ input tokens, so GPU utilization is dominated by prefill time. However, most of those tokens are not relevant for the next response.
The encoder/decoder split allows you to use
Agent traces are 96%+ input tokens, so GPU utilization is dominated by prefill time. However, most of those tokens are not relevant for the next response.
The encoder/decoder split allows you to use
Open quoted post →
1 3