
Hey guys, Mr. Technology here.
It is Friday, July 24, 2026, and the 18-month token price war just ended in the last ten days. I do not mean "ended slowly" or "ended on a particular roadmap." I mean four independent events landed in a ten-day window and, taken together, they form a regime change that is going to restructure every inference architecture you have running in production by Q4.
Let me line them up:
Any one of those by itself is a story. Together, they are the same story told five different ways: the substrate underneath frontier inference is now a tier-1 macro asset class, and the price you pay for tokens is going to behave like a peak-load utility from here on out. The deflationary era — the one where every quarter was cheaper than the last — is finished. If you are still architecting for it, you are going to lose money, miss latency budgets, and get caught holding capacity contracts you cannot get out of.
I want to walk you through the four pieces of evidence and what they mean together, then show you the inference architecture that actually survives the next 18 months. There is a working reference implementation at the end. It is not a toy.
The Gemini 3.5 Pro story is the most under-reported of the five and the one I think matters most, because it is the only one that tells you what the labs themselves are dealing with. Google has not published a clean post-mortem, but the shape is clear from internal chatter at DeepMind, the public statements from Sundar Pichai on the Q2 earnings call, and the pricing changes that quietly appeared on the Vertex console on July 14.
What happened: Gemini 3.5 Pro was originally a mid-cycle refresh, not a full retrain. The team finished the data-curation pass in early June, started the RLHF/scaling run on June 18, and discovered around July 8 that the new run was producing higher-quality outputs but at a meaningfully higher training-compute cost. The choice was to either ship the model at the existing API price (eating the cost), raise the API price (eating the demand), or scrap and rebuild with a more efficient architecture. Google picked option three. The replacement — internally called Gemini 3.5 Pro v2 — is on a six-to-eight-week timeline.
The reason this matters to you: Google is the lowest-cost producer of frontier inference in the world. They own the fab, they own the rack-scale networking, they own TPUv6 and the software stack. If Google's training economics are getting worse, the industry-wide training economics are getting worse. There is no provider on Earth with a deeper cost moat. Whatever Google is paying attention to — and they are paying attention to it enough to scrap a finished model — is going to land on your bill within two quarters.
I covered Kimi K3 in detail on Monday, but the piece that did not fit in that post is the part I want to repeat here. Moonshot is not a frontier lab in the Western sense. They do not have Microsoft-distributed compute. They do not have a $40B Nvidia investment. They have their own cluster and access to Huawei Ascend 910C/910D silicon plus a constrained allocation of H200s via the gray-market channels that exist for Chinese labs as of mid-2026.
Kimi K3 is a 2.8-trillion-parameter MoE with 896 experts, 16 active per token. A single INT8 replica is approximately 2.8 TB of HBM. To serve one copy of the model at throughput, you need a 20-GPU H200 rack minimum. To serve it at the concurrency that Arena-#1 demand produces, you need a multi-rack cluster with high-bandwidth interconnect.
Moonshot did not pause subscriptions because they wanted to. They paused subscriptions because they ran out of HBM. The bottleneck is not training the next model; the bottleneck is serving the model they already shipped. The "compute triangle" I keep talking about — Nvidia + hyperscaler + frontier lab — has a fourth vertex in 2026: HBM allocation. Every frontier lab in the world is now competing for SK Hynix and Micron HBM3e/HBM4 wafer starts. Moonshot lost this round.
For you: if a frontier model from a top-five lab goes "sold out" four days after launch, your failover plan cannot assume "we will just switch to the frontier model that is cheapest this week." Sometimes the cheapest model is sold out. Your failover plan needs to include older generations and second-tier providers — Kimi K2, DeepSeek V3.2, Mistral Large 3, Llama 4 Behemoth — and those need to be in your test matrix continuously, not bolted on when you get caught.
This is the one that should make you stop and pay attention, because DeepSeek has never been the lab to do marketing-driven pricing. They are the lab that publishes the cost-of-training spreadsheet three weeks after shipping a model. They are the lab that, in March 2025, dropped R1 to the public and said "we charge $2.19 per million output tokens because that is the marginal cost plus a 15% gross margin."
So when DeepSeek ships peak/off-peak pricing in V4 GA, it is not a marketing test. It is a signal that their utilization curve is non-uniform in a way they can no longer absorb. Peak is 09:00-21:00 Beijing time. That is the overlap of the US workday evening and the China workday. Off-peak is the rest. The price multiplier is 2.4x peak to off-peak for both input and output. The minimum commit on enterprise contracts is 50M tokens/month at peak rates — i.e. you cannot buy your way out of peak pricing by prepaying, you can only smooth demand.
Why they are doing it: DeepSeek serves primarily through OpenRouter, plus their own direct API, plus a few Chinese platform partners. The aggregate load curve has a strong daytime peak driven by coding agents, customer-support automation, and the long tail of Chinese SaaS. At peak they were queueing requests at p99 over 4 seconds on V3.2. The choice was either to add capacity (which means more HBM they do not have, see point 2) or to flatten the demand curve through pricing. They picked pricing.
For you: every other lab is going to copy this within 12 months. Anthropic has the technical ability to do it today; they have not because their customers are enterprise and enterprise hates time-of-day pricing. OpenAI will copy it for ChatGPT first and the API second, probably Q1 2027. Plan now. If 20-40% of your agent traffic can be deferred or batched, the cost difference is going to be 1.5-2x. If 100% of your traffic is real-time and peak-bound, you are going to pay the peak rate forever.
I do not want to re-cover the Anthropic IPO piece — I did that Wednesday. The part that matters here is the implication for API pricing. When a $1T valuation is anchored on compute, the public market is going to ask Anthropic one question on every earnings call: "how fast is your compute base growing?" The answer that justifies the multiple is "faster than revenue." Which means every dollar of gross profit Anthropic earns on Sonnet 5 today is going to be plowed into HBM and Blackwell wafers tomorrow, which means token prices do not fall as fast as they would in a normal software-margin business.
This is the part most commentary is missing. The $1T valuation is not bullish for token prices because Anthropic is "winning." The $1T valuation is bullish for compute costs because the capex curve is steeper than the revenue curve. Anthropic has to defend a number, and the number is denominated in GPUs and HBM.
For you: the era of "the labs are price-competing and we benefit as consumers" is over. The era of "the labs are compute-competing and we, the consumers, are paying the marginal cost of that competition" is here.
Last piece, the one I covered Thursday: the DOJ closed the Nvidia-Microsoft-Anthropic antitrust probe without remedies. I am not going to re-litigate the regulatory analysis. The architectural implication is one sentence: the Nvidia + hyperscaler + frontier-lab vertical is now durable. It is not going to be broken up. The forward-rate on Anthropic API just dropped because the market priced in regulatory tail-risk that no longer exists. That is bullish for the labs' ability to commit capex, which means more compute coming online, but it is not bullish for consumer pricing. It is bullish for capacity — you will be able to get tokens — but the per-token price of those tokens is governed by the compute economics in points 1-4.
You were building for a world where the price of tokens falls 50-70% per year and capacity is infinite. You are now building for a world where the price of tokens falls 5-15% per year, capacity is rationed by HBM allocation, and time-of-day pricing will be the norm within 12 months. The architecture that wins the next 18 months is one that treats inference as a peak-load utility: tiered by criticality, with circuit-breakers, batchable workloads deferred to off-peak, and a continuous fallback to older generations and second-tier providers.
Here is the stack I am running in production as of this week. It is opinionated, it assumes you are running more than ~$20K/month of inference spend, and it is not LiteLLM. LiteLLM is fine for the simple case; it does not give you the knobs you need when the pricing curve is peak-load.
The components:
1. A tier router — every request gets a tier tag (interactive / batchable / background) at the edge. 2. A multi-vendor pool with at least three frontier providers, one second-tier provider, and one local open-weights fallback. 3. A breaker — circuit-breaker per provider that flips on 5xx rate, latency, or rate-limit signals. 4. A peak-aware scheduler — defers tier-3 (background) traffic to off-peak windows using a queue with a deadline. 5. A kill switch — hard cap on monthly spend, drops to the cheapest available provider when the cap is approached.
I will show you the router. The full stack is in the mr-technology-stack-2026 GitHub repo, but the router is the part most people get wrong.
# tier_router.py
# Tags every inference request with a criticality tier so the downstream
# scheduler can decide peak/off-peak, provider priority, and circuit-breaker
# behavior. Three tiers:
# T1 — interactive user-facing. Latency-bound. Must hit p99 < 2s.
# T2 — async but user-visible. Quality-bound. Latency budget 30s.
# T3 — background batch. Cost-bound. Latency budget unbounded, deadline-aware.
#
# The tier is set at the edge (your API gateway, your agent runner, your
# internal RPC layer) and propagated as a header.
from dataclasses import dataclass
from enum import IntEnum
import time
class Tier(IntEnum):
INTERACTIVE = 1
ASYNC_USER = 2
BACKGROUND = 3
@dataclass
class Request:
body: dict
tier: Tier
deadline_s: float | None = None # for T3, when the result is actually needed
budget_usd: float = 0.0 # optional per-request cap
# Tier assignment is a policy decision. This is the default but you should
# override per use-case. Some examples:
# - Chat UI first token -> T1
# - Chat UI follow-up tokens -> T2 (still user-facing, can wait)
# - Embedding regen nightly -> T3, deadline=06:00 next day
# - RAG reindex triggered now -> T3, deadline=15min
# - Eval harness -> T3, deadline=4h
def assign_tier(ctx: dict) -> Tier:
# ctx is whatever you pass — request metadata, user-agent, route name.
route = ctx.get("route", "")
if route.startswith("/v1/chat/stream") or route.startswith("/v1/realtime"):
return Tier.INTERACTIVE
if route.startswith("/v1/chat"):
return Tier.ASYNC_USER
if route.startswith("/internal/batch"):
return Tier.BACKGROUND
return Tier.ASYNC_USER # safe default
def attach_tier(req: Request) -> Request:
# attach_tier() is called at the edge. The tier travels with the request
# through the rest of the stack.
return reqThe point of this file is not the code — it is the discipline. Every request in your system has to have a tier, and the tier has to be set at the edge, before the request reaches the LLM client. If you do not do this, you will not be able to defer anything, you will not be able to prioritize anything, and you will not be able to budget anything. The tier is the smallest unit of policy that lets the rest of the stack do useful work.
# vendor_pool.py
# Provider pool with per-provider circuit breakers. The breaker state is
# the source of truth for "is this provider healthy right now?" and feeds
# both the tier router and the peak-aware scheduler.
#
# Trip conditions (any one is sufficient):
# - 5xx rate > 5% over the last 60s
# - p99 latency > 4s over the last 60s
# - rate-limit (429) rate > 10% over the last 60s
# - provider-reported capacity signal (rare, DeepSeek has it)
#
# Half-open behavior: after 30s in OPEN, send one probe request. If it
# succeeds, transition to CLOSED. If it fails, back to OPEN for another
# 30s.
import time
from dataclasses import dataclass, field
from enum import Enum
class BreakerState(Enum):
CLOSED = "closed" # normal
OPEN = "open" # trips, no traffic
HALF_OPEN = "half_open" # one probe request allowed
@dataclass
class BreakerStats:
window_start: float = field(default_factory=time.time)
requests: int = 0
errors_5xx: int = 0
rate_limited: int = 0
latency_sum: float = 0.0
def reset(self):
self.window_start = time.time()
self.requests = 0
self.errors_5xx = 0
self.rate_limited = 0
self.latency_sum = 0.0
@dataclass
class Provider:
name: str
base_url: str
state: BreakerState = BreakerState.CLOSED
stats: BreakerStats = field(default_factory=BreakerStats)
opened_at: float = 0.0
priority: int = 0 # higher = preferred; router consults this
def record(self, status: int, latency_s: float):
s = self.stats
s.requests += 1
s.latency_sum += latency_s
if 500 <= status < 600:
s.errors_5xx += 1
if status == 429:
s.rate_limited += 1
self._maybe_trip()
def _maybe_trip(self):
if self.state != BreakerState.CLOSED:
return
s = self.stats
if s.requests < 20: # need a window of samples
return
window_s = max(time.time() - s.window_start, 1.0)
err_rate = s.errors_5xx / s.requests
rl_rate = s.rate_limited / s.requests
p99 = (s.latency_sum / s.requests) * 4 # rough; track histogram in prod
if err_rate > 0.05 or rl_rate > 0.10 or p99 > 4.0:
self.state = BreakerState.OPEN
self.opened_at = time.time()
def allow(self) -> bool:
if self.state == BreakerState.CLOSED:
return True
if self.state == BreakerState.OPEN:
if time.time() - self.opened_at > 30:
self.state = BreakerState.HALF_OPEN
return True # one probe
return False
# HALF_OPEN: only the probe is allowed; caller is responsible for
# not sending more than one. In practice we serialize at the pool.
return TrueThis is the boring infrastructure that pays for itself the first time a provider has a bad day. Without it, your agents pile up on the cheapest provider until that provider 429s, and then they fail. With it, the traffic reroutes in 30 seconds and your p99 stays bounded.
# peak_scheduler.py
# Decides whether to send a request now, defer it, or kill it. The decision
# uses the tier from tier_router.py and the local time-of-day peak window
# defined per provider. Off-peak is 30-60% cheaper for DeepSeek and will be
# for Anthropic/OpenAI within 12 months.
#
# Policy:
# T1 (interactive) -> always send now. Cost is irrelevant.
# T2 (async user) -> send now unless we are in peak AND a cheaper
# off-peak window opens within the deadline.
# T3 (background) -> defer to next off-peak window unless the deadline
# is closer than the wait.
from datetime import datetime, timezone, timedelta
from tier_router import Tier, Request
# Peak windows are UTC. DeepSeek's peak (09:00-21:00 Beijing = 01:00-13:00 UTC)
# is the most aggressive. Anthropic/OpenAI will publish theirs when they
# formalize peak pricing; this is the placeholder.
DEEPSEEK_PEAK_UTC = (1, 13) # hour boundaries, inclusive
def is_peak(now_utc: datetime, peak_window=DEEPSEEK_PEAK_UTC) -> bool:
h = now_utc.hour
return peak_window[0] <= h < peak_window[1]
def next_offpeak(now_utc: datetime, peak_window=DEEPSEEK_PEAK_UTC) -> datetime:
start_h, end_h = peak_window
if now_utc.hour < start_h:
return now_utc.replace(hour=end_h, minute=0, second=0, microsecond=0)
return (now_utc + timedelta(days=1)).replace(hour=end_h, minute=0, second=0, microsecond=0)
def decide(req: Request, now: datetime | None = None) -> str:
now = now or datetime.now(timezone.utc)
peak = is_peak(now)
if req.tier == Tier.INTERACTIVE:
return "send"
if req.tier == Tier.ASYNC_USER:
if not peak:
return "send"
# in peak: defer only if deadline allows. Default 30s, so usually not.
if req.deadline_s is None or req.deadline_s < 60:
return "send"
return "defer"
# BACKGROUND
if not peak:
return "send"
if req.deadline_s is None:
return "defer" # no deadline = no urgency
wait_s = (next_offpeak(now) - now).total_seconds()
if wait_s < req.deadline_s:
return "defer"
return "send"This is where the savings come from. A workload mix that is 60% T3 background with flexible deadlines can save 30-45% on token costs once peak pricing is universal. That is not a small number when you are spending $50K/month on inference.
Most teams I have talked to this month are running one of three patterns, all of which are going to lose money in 2026:
Pattern A: Single-vendor with a fallback. "We use Anthropic. If Anthropic is down, we use OpenAI." This pattern assumed tokens get cheaper every quarter and Anthropic is always available. Both assumptions are wrong now. Kimi K3 sold out in four days. Gemini 3.5 Pro was killed mid-rollout. The fallback vendor is part of your cost basis, not a contingency.
Pattern B: LiteLLM with model aliases. "We have gpt-5.6, claude-sonnet-5, gemini-3-pro aliased and we route by cost." This pattern assumes you can swap models without quality loss. You cannot, in most cases. The quality delta between a frontier model and a second-tier model is 15-30% on hard evals, and you have already chosen the frontier for a reason. Routing by cost means routing to lower quality, which means user complaints, which means re-routing to the expensive model, which means the LiteLLM was a tax you paid for nothing.
Pattern C: "We will evaluate later." This is the most common pattern. It is also the most expensive. Every month you defer, you are paying peak prices on T3 traffic you could defer, you are missing circuit-breaker saves you would have caught, and you are not measuring provider quality drift. The 2026 inference stack is not "set and forget." It is "set, measure, rebalance monthly."
The architecture I have shown you above is Pattern D: tiered, multi-vendor, breaker-aware, peak-aware, kill-switched. It is more work to set up. It is also the only one that does not lose money when the next Kimi K3 moment happens.
Here is what I want you to actually do on Monday, assuming you are running production inference in any meaningful volume:
1. Add a tier to every request. Even if you do nothing else with it, having the tier in the metadata lets you answer the question "what fraction of my inference bill is time-sensitive?" If the answer is "more than 30%," you have a problem.
2. Add a second frontier provider to your failover path. Not as a backup. As a co-primary with a 70/30 or 60/40 traffic split. This is the only way to get reliable signals about provider quality drift, and it is the only way to negotiate your primary contract from a position of strength in 2027.
3. Instrument the peak window. Track your providers' p99 latency by hour-of-day. You will see a curve. Once you see the curve, you can defer the tail.
4. Set a hard spend cap. Not a soft budget. A circuit-breaker that drops to the cheapest available provider when your monthly spend approaches 80% of plan. The labs will not warn you before they raise prices. The market will not warn you before a provider has a bad day.
5. Stop assuming tokens get cheaper. They will get cheaper on the median model. The model you actually want — the frontier one — is going to be priced at the marginal cost of compute, which is going up, not down, for the next 18 months at minimum.
The 2026 inference squeeze is not a temporary disruption. It is the new shape of the market. The teams that recognize it in Q3 2026 will have a structural cost advantage in Q1 2027 that compounds every quarter after. The teams that do not recognize it will keep paying peak prices on traffic they could defer, and they will not understand why their inference bill grew 40% while their traffic grew 5%.
The token price war ended. The architecture war starts now.
Hey guys, that is the post. The full reference implementation is going to be in the mr-technology-stack-2026 repo by end of weekend. If you are running a different stack and you want a second pair of eyes, my DMs are open. I will be in the comments.