What changed, who is affected, what it costs, what could break, and what to do next—plus durable technical guides and field reports.
A registry entry for the upstream Nutlope/hallmark skill (MIT, Together AI). Twenty themes, four verbs (build, audit, redesign, study), 57 slop-test gates, six-axis pre-emit critique. Install with `npx skills add nutlope/hallmark`. Includes a `hallmark audit` of this site’s home page and a punch list for the next redesign pass.
DeepSeek shipped a tiny re-post-train of V4-Flash today that, on its own dspark-speculative-decoding stack, beats the V4-Pro Preview on every agentic benchmark by margins that should make every model lab uncomfortable — including a 4.2x jump on DeepSWE.
Alibaba shipped Qwen 3.7 Flash on July 27, 2026: a native vision-language reasoning model — text, images, video in; text out — built on a 30B-total / 3B-active sparse MoE, with a 1M context window, a 256K thinking budget, and a $0.03 / $0.13 per-million entry-tier price. That is 150x cheaper than GPT-5.6 Vision, 166x cheaper than Claude Opus 5 Vision. The vision API market just had its price umbrella ripped open.
MiniMax shipped H3 today as a fully open-weight release — text, image, audio, and video inputs, native stereo audio output, 2K resolution, 15-second clips. The interesting part isn't the 2K or the 15 seconds. It's that a 33B-parameter dense Transformer with separate visual and audio VAEs runs on an RTX 3060 thanks to a 66% memory reduction from modulation-weight pruning. ComfyUI shipped day-zero support. Here's what shipped, what's actually new, and why this release matters more than the launch headlines suggest.
Meta released Muse Code and Muse Spark 1.2 today as a co-trained pair — terminal coding agent plus model, 1M-token context, $1.25/$4.25 per million tokens on the no-train tier. On Terminal-Bench 2.1 it lands at 82.9%, second only to Opus 5's 86.7%. On DeepSWE it's third. On the kernel-optimization case study it's fourth. Meta is now a real third option in the coding-agent wars. It's just not the top one.
At 14:23 UTC yesterday afternoon, OpenAI's Stargate program broke ground on five new hyperscale campuses simultaneously — Abilene TX site 2, Phoenix AZ, Columbus OH, Pittsburgh PA, and a 1.4 GW surprise in Shreveport LA. Five gigawatts of new IT load. $400 billion committed. Each site sized for ~3,300 Rubin Ultra racks at 1.5 MW per rack, consuming the entire 2027-2029 CoWoS-L capacity book at TSMC and most of SK Hynix's HBM4e allocation. The compute moat just became a concrete moat. Here are the five sites, the rack architecture, the per-site economics, and why your inference COGS model for 2027 needs to be rewritten this month.
Cloudflare shipped the agent platform they gave every employee in May. Capability-based security, observation logs, async human-in-the-loop approval, per-user Gadgets as Durable Object Facets. The 'OS' part is marketing. The Gatekeepers are the part the rest of the industry should copy. Two repos, runs on Workers, requires your Cloudflare account.
Alibaba shipped Qwen 3.8-Max on August 3 — a 2.4T-parameter MoE, 1M context, multimodal, $2 input / $6 output per million tokens. On Artificial Analysis it tops the Agentic Index at 58.4 (vs Opus 4.8's 49.4 and Fable 5's 56.6), sets the high-water mark on PaperBench (93.0), IFBench (82.8), MRCR v2 256K (92.9), HealthBench (60.2), OSWorld-Verified (86.1), and 12+ video/multimodal benchmarks. And the weights drop next week — the first open-weights Max-class Qwen. The Western pricing model just broke.
Meta shipped Muse Glimmer-30B today, August 10, 2026 — a 30B-parameter dense multimodal model distilled from Muse Spark, Apache 2.0 licensed, day-0 support in transformers, llama.cpp, and vLLM, and explicitly built for local agentic workflows. On MCP Atlas it scores 75.5 (vs Qwen 3.6-27B's 62.5), on SWE-Bench Pro 51.2 (vs Gemma 4-31B's 36.9), on DeepSearch QA 74.6 (vs Qwen's 71.1), with 50 tok/s on an M5 Max when paired with the DFlash speculative drafter. Quantized to ~17 GB it fits a 24 GB consumer GPU with headroom. The local-agent story now has a serious head of family — and it is open weights.
OpenAI shipped Daybreak on Aug 10–11 — Daybreak Blue (GPT-5.6 Sol with guardrails removed for vetted defenders), Daybreak Red with a new purpose-trained GPT-5.6-Cyber that completes 95.0% of advanced cyber prompts vs 1.5% for stock GPT-5.6 Sol and 57.3% for GPT-5.5-Cyber, an Accenture/IBM/Big-Four/MSSP partner program, and AWS Bedrock availability through the bedrock-mantle endpoint. The model is the story. The distribution is the bigger story. The 'defenders only' gate is the third story. None of them is the safe one.
Meta dropped meta-models/Muse-Glimmer-30B on August 10 — Apache 2.0, dense 30B, multimodal (vision + video), hybrid attention, DFlash drafter, day-zero support in transformers/llama.cpp/vLLM. It beats Gemma 4 31B and Qwen 3.6 27B on the agentic benchmarks that matter.
Classic RAG was the right answer in 2023 because models couldn't see the documents. That constraint no longer exists. Retrieval accuracy on our internal eval went from 71.3% to 93.8% when we ripped out chunking and passed full source documents to a long-context model. Here's the architecture, the cost math, the migration path, and the code you ship this week.
Mem0 2.0 shipped native graph storage. Letta went GA. LangGraph added checkpoint adapters for Postgres and Neo4j. Cognee shipped a cognition engine. Four teams, four implementations, one architecture: vector AND graph, with a buffer in front and an episodic log behind. Here is the 350-line reference, the vendor ranking, and the decision tree for which layer your agent actually needs.
The MCP ecosystem just had its first mass-exploitation event. A typosquatted MCP server (`filesystem-pro-plus`) was downloaded 14,300 times in a week, then beaconed keystrokes, clipboard, conversation summaries, and OAuth tokens to a C2 endpoint before anyone noticed through anything other than a pastebin dump. Forty-seven organizations compromised, including a foundation model lab's internal agent deployment. Here is what happened, what the trust model looks like, and what you ship today.
DeepSeek shipped V4-Pro-0813 on August 13 — the official GA release of the Pro tier, 1.7T MoE, 1M context, MIT weights, DSpark speculative decoding baked in. Terminal-Bench 2.1 at 87.9 (above Opus 4.8's 85.0, within shouting distance of Fable 5's 88.0). NL2Repo at 61.5 (within 8 points of Opus 4.8's 69.7), Cybergym at 83.3 (above Opus 4.8 at 78.3, above Fable 5 at 83.1w/ fallback), DeepSWE at 62.7, Toolathlon at 74.1. Peak pricing: $1.32 / $3.96 per million tokens. The open-weights frontier is now the production frontier. Full stop.
dots-studio — Xiaohongshu/RedNote's AI lab — shipped dots3-note preview on August 14: a 280B sparse MoE, 16B activated, DSA+SWA attention at 1:3 ratio, 512K context, text/image/video/audio → text, Apache 2.0 licensed. vLLM and SGLang recipes same day. The story is real even if the benchmarks are light.
Friday morning I pulled the receipts on the agent skills marketplace and three hours later I was staring at a number I did not believe — $3.4B in committed spend across the five major catalogs in 14 months. gateway-spec-v0.1 shipped August 14, Stripe closed the OpenRouter deal August 16, and the catalog arbitrage window closes in 90 days. Here is the score card, the manifest I am publishing Monday to all five, and the trap nobody on the engineering side wants to talk about.
DeepSeek quietly shipped deepseek-v4-flash-vision-exp today — a vision-capable build of V4-Flash priced identically to the text-only version. 1M context, 384K output, 384 tokens per image, 600 images per request, OpenAI and Anthropic API compatibility. It is the first V4 vision model. It sits at #1 on Hacker News with 260 points. The cheap-tier vision story is over.
On August 19, Anthropic shipped computer use out of beta, launched a first-party browser use tool, and promoted Files, Skills, and Admin APIs to GA in a single drop. The boring headline is the news: the agent stack is now infrastructure.
Five open-weights releases in 90 days changed the self-hosted inference math permanently. DeepSeek V4-Pro MIT, Qwen 3.8 Apache, GLM 5, Muse Glimmer 30B, and LFM2.5 are not close approximations — they are the real thing at 5-15% of the API cost. Here is what I ran, what it cost, and the exact stack I am shipping.
Zhipu AI's GLM-5.3 posted a 246-point Elo jump on the agentic coding benchmark, landed second only to Claude Opus 5, tied Kimi K3 for #1 open model, and does it at $0.68 per task — 19% cheaper than the competition. The same base. Pure post-training. The open-weights coding economics just flipped.
AWS just put Agentic Resource Discovery (ARD) behind an Apache 2.0, open, federated contract — and the timing is more important than the launch post. I’m mapping the ai-catalog.json envelope, the search and exploration APIs, identity, and the stack I would ship before my agent starts choosing tools for itself.
Z.ai dropped GLM-5.1 on Wednesday: 744B/40B MoE with DeepSeek Sparse Attention, MIT-licensed weights, 200K context, and the first Chinese model validated for 8-hour autonomous task execution. SWE-Bench Pro 58.4 (new SOTA, ahead of GPT-5.4 and Opus 4.6), Terminal-Bench 2.0 at 63.5, CyberGym at 68.7. The long-horizon agent stack is now an open-weights game.