The Bleeding Edge

// Article · August 7, 2026 · 8 min read

The Bleeding Edge Weekly — W32: Kimi K3 and DeepSeek go fully open, Microsoft shows agent skills survive a harness switch

Four open-weight releases in seven days, one under an MIT licence — and the first hard evidence that agent scaffolding ports between vendors.

from 2026-W32 ↗newsletterweeklyw32

By The Bleeding Edge AI desk. Drafted by AI from the week's linked sources and published automatically, without line-by-line human review. How we make this →

// Contents

This edition combines the three newsletters we published separately this week (LLM Weekly, Devices & Robotics and Executive Roundup). Their text is unchanged.

The week in models

Moonshot published full Kimi K3 weights and DeepSeek shipped V4-Flash-0731 under an MIT licence — while Microsoft published evidence that the agent skills you built for Codex work in Claude Code. Lock-in is leaking from both ends at once.

China's labs gave the week away

Moonshot AI published the complete weights for Kimi K3, described as the first openly downloadable model of its class. DeepSeek put V4-Flash-0731 on Hugging Face under an MIT licence — the most permissive terms available, no use restrictions, no field-of-use carve-outs. MiniMax H3 and new AMD models landed in the same seven days. The practical upshot: a Western enterprise that wants a capable self-hosted model with no vendor relationship at all now most likely picks a Chinese one. That is a board conversation most boards haven't had.

Via AI Search and the Creators' AI digest.

OpenAI says 99.8% of its tokens are agentic

Circulating this week alongside a billion-user milestone: the claim that all but 0.2% of tokens OpenAI serves are consumed inside agent loops rather than human chat turns. It's self-reported, it came through the newsletter layer rather than a disclosure, and "agentic token" is doing enormous definitional work. But if it's even roughly right, the unit of consumption is now an agent run — and every per-seat licence, rate limit, and audit trail written for chat is mis-specified.

Via the Creators' AI weekly digest.

Microsoft's SkillOpt: agent skills are portable

Microsoft published work on optimising agent "skill" artifacts and found the optimised artifacts transfer both across model sizes and between two competing harnesses — Codex and Claude Code. This is the first real evidence that enterprise investment in agent scaffolding is durable rather than vendor-specific. It also quietly dismantles the lock-in story the agent vendors have been telling: if the artifact ports, switching cost collapses, and picking a lab becomes picking a lab for now.

Via MarkTechPost and The Neuron.

Meta ships Muse Code on a new model

Meta AI released Muse Code in beta — a terminal coding agent backed by a new model, Muse Spark 1.2. Meta had ceded the terminal-agent category to Claude Code and Codex CLI entirely; this puts it back in, and notably pairs a coding-tuned model with Meta's own harness rather than shipping weights and hoping someone builds the surface. Details are thin: this surfaced as a release signal, not a benchmark drop.

Via MarkTechPost.

Millennium is building a risk analyst on Claude

Anthropic published a joint case study with Millennium, one of the largest multi-strategy hedge funds in the world, describing an AI-powered digital risk analyst built on Claude. Risk is the most conservative seat in the most conservative corner of finance — a named deployment there does more procurement work than a hundred pilots. It landed the same week CBA flagged roughly $1B in suspected fraud, which makes financial-crime detection look like the first genuinely mandatory enterprise LLM workload.

Via Anthropic and Capital Brief.

Next week, two things to watch: third-party benchmarks on Claude Fable 5, whose first 48 hours produced plenty of hands-on impressions and no published evals — and whether any Western lab answers an MIT-licensed release of that calibre. So far the permissive-licence race has exactly one set of runners.

Devices & robotics

Two things happened to embodied AI this week and they point in opposite directions. The models that drive and control physical systems got cheaper — free, in NVIDIA's case. The physical systems themselves started acquiring nationality.

The US moves to ban foreign-made humanoid robots. Reported this week as a prohibition on foreign-manufactured humanoids, framed explicitly as a China measure. The mechanism is the part nobody has pinned down: an outright import ban and a federal procurement exclusion are wildly different instruments for anyone building a purchase plan. Note the timing — chip export controls took roughly four years of boom to arrive, while the humanoid restriction lands before the category has meaningful commercial deployment. If you were quietly assuming your 2027 pilot fleet would be sourced on price, that assumption is now a policy bet. Reported via the Creators' AI weekly digest; treat scope as unconfirmed.

NVIDIA open-releases Alpamayo 2 Super — a 34B vision-language-action model for robotaxis. Thirty-four billion parameters, targeting autonomous-driving and robotaxi stacks, published under the OpenMDW-1.1 licence. A VLA model is the perception-to-actuation layer — the part every AV team has historically treated as its crown jewel. NVIDIA just made a credible version of it downloadable. The strategic logic is the same one that has worked for a decade: give away the layer that creates demand for the layer you sell. If you are an AV startup whose pitch deck leans on proprietary driving models, this week made that slide harder to defend. Via MarkTechPost.

Gemini Robotics ships an update. Google pushed a Gemini Robotics release in the same window, flagged in the week's model roundup without published specs or benchmarks. Take it as a cadence signal rather than a capability claim: the two companies with the most to gain from robotics foundation models — one selling silicon, one selling cloud — both shipped into the category in the same seven days. Worth watching whether third-party evaluations land before the next release. Via the AI Search weekly roundup.

DeepSeek puts V4-Flash-0731 on Hugging Face under MIT — and the licence is the device story. Most open-weight releases carry acceptable-use policies, field-of-use carve-outs, or revenue thresholds that make lawyers nervous about embedding the model in a shipped physical product. MIT has none of that. For anyone building an appliance, a kiosk, a vehicle head unit, or anything else that runs inference locally and ships to a customer, the licence has been the blocker more often than the benchmark. That blocker is now gone — from a Chinese lab, which is its own procurement conversation. Release via AI Search and the Creators' AI digest; the licensing read is ours.

The pattern underneath all four: software commoditises, hardware nationalises. Weights are being given away by the people who used to charge for them, while bodies, fabs and grid connections turn into industrial policy. Ken Griffin spent the same week putting a number on what happens if the US loses access to Taiwanese semiconductor supply — which is the honest summary of every story above. The brain is free now. Watch the supply chain for the rest of it.

What it means for leaders

Everything that could be copied was given away this week. Everything that couldn't — chips, robots, gigafactory sites — became industrial policy, and that split is the through-line for all three roles.

If you're a CEO this week...

Model access stopped being a competitive advantage. Moonshot published full Kimi K3 weights and DeepSeek shipped V4-Flash-0731 under an MIT licence — no use restrictions, no field-of-use carve-outs. If your "AI-native" story rests on which lab you signed with, a competitor can now match it for the cost of GPUs.

Your investors are pricing a different risk than you are. Ken Griffin spent the week putting a number on the US losing access to Taiwanese semiconductor supply — while Washington moved to ban foreign-made humanoid robots, Brussels opened a €30B call for seven AI gigafactories, and South Korea signed a reported $950B in AI deals. Expect supply-chain questions, not capability questions.

The conservative seats are buying. Millennium is building a digital risk analyst on Claude. When risk management at a top multi-strategy fund deploys, "still piloting" becomes a board question.

Be able to answer this: if a competitor self-hosted an MIT-licensed model next quarter and matched our AI feature set, what would still be ours?

If you're a CIO/CTO this week...

Your agent scaffolding is more portable than your vendor implied. Microsoft's SkillOpt work found optimised skill artifacts transfer across model scales and between Codex and Claude Code. That kills the main lock-in argument in every agent-platform contract on your desk — negotiate term length accordingly.

New procurement question you haven't asked yet: the default self-hosted option is now Chinese-origin. Kimi K3, DeepSeek-V4-Flash-0731, MiniMax H3 — plus NVIDIA's Alpamayo 2 Super, a 34B vision-language-action model under OpenMDW-1.1. Get legal onto the licence terms and security onto weight provenance before an engineer does it for you.

Budget exposure: OpenAI reportedly claims 99.8% of its tokens are now agentic. Self-reported, definition-dependent — but your rate limits, audit trails, and per-seat licences were all specced for chat turns.

The read: buy the boring layer — Marker v2 fixes document parsing, which is what actually stalled your RAG pipeline, not model quality. Build the skill library in-house, in git. Monitor Meta's Muse Code passively; a third terminal agent isn't a roadmap event. And treat MCP server security as a live 2026 line item, not a 2027 one.

If you lead AI transformation this week...

The pilot that pays for itself: take one workflow your team has run more than five times, reverse-engineer it into a written skill artifact, then run it on a smaller model and a second harness. If it survives both, you've built a durable asset; if it only works on the frontier model, you encoded capability, not procedure. That's SkillOpt's finding turned into a two-week evaluation.

Playbook update: the scarce role is shifting from prompt author to skill-library curator — someone who versions, tests, and retires procedural artifacts. Open a training gap here now; it's cheaper than re-buying it in six months.

Two things headed for your governance agenda: Chinese-origin open weights entering the stack through the back door, and agentic token accounting that no finance function has a policy for. Start with financial-crime workloads if you need a defensible first mandate — CBA's ~$1B fraud disclosure and the new banking watch lists show where the alternative to AI is a regulator.

The experiment to run this month: pick your three most-repeated tasks, ship skill artifacts for each, and measure whether a junior can run them unassisted. That's the real adoption metric.


All three of you are being asked the same thing in different words this quarter: if the model, the prompt, and the scaffolding are all commodities, what exactly are we defending? Lenny Rachitsky's argument against the long-term plan circulated hard this week for a reason — keep the thesis, throw away the eighteen-month Gantt chart.


This post is also published on our Substack newsletter at edge-ai.forum. Subscribe for the weekly roundup direct to your inbox — fresh AI news, executive context, and devices + robotics every Friday morning.

// Related