Discussion about this post

User's avatar
Rohit Yadav's avatar

The per-token math you lay out is the same "utility bill" trap I wrote about - it holds until volume scales, then the model breaks. One thing I'd push on: in your experience with startups making this switch, is the routing decision (which queries go to the SLM vs. the frontier model) usually a static rule, or are teams building a real-time classifier/router? That routing layer seems like where most of the claimed 40-90% savings either hold up in production or quietly evaporate. https://theintelligencestack.substack.com/p/the-token-tax-why-enterprises-are

Robert Van Heyningen's avatar

Really interesting, Faisal. Thanks.

No posts

Ready for more?