The Bleeding Edge

// Article · July 24, 2026 · 7 min read

The Bleeding Edge Weekly — W30: Kimi K3 makes the open frontier 2.8 trillion parameters — and Chinese

The largest open-weight model yet lands from Beijing, then four more models bury it before Friday.

from 2026-W30 ↗newsletterweeklyw30

By The Bleeding Edge AI desk. Drafted by AI from the week's linked sources and published automatically, without line-by-line human review. How we make this →

// Contents

This edition combines the three newsletters we published separately this week (LLM Weekly, Devices & Robotics and Executive Roundup). Their text is unchanged.

The week in models

The largest open-weight model ever released shipped this week — reportedly 2.8 trillion parameters, native vision, a million-token context — and it came from Beijing, not Silicon Valley. Then four more models landed on top of it by Friday, which is the real story: frontier releases have become weather.

Moonshot AI ships Kimi K3 — 2.8T open weights with native vision and a 1M-token context. China's Moonshot released Kimi K3 as a free open-weight download: reportedly 2.8 trillion parameters, multimodal vision built in, and a million-token window — frontier-scale weights anyone can self-host behind their own firewall. Every closed US lab now competes against a free artifact enterprises can run in-house, and the reference "free, frontier-scale" model keeps shipping from Beijing. That reshapes both the pricing conversation and the data-sovereignty one for Western buyers. Source: AI Search.

Four more models landed under the radar. Alongside Kimi, the week's dump included Bonsai 27B, Alibaba's Wan Dancer (video/motion), a GPT "Red" variant, and OpenAI's Codex Micro coding model — plus a refreshed, cheap-and-fast Google Flash tier and an agentic-coding mixture-of-experts from Poolside. None got individual attention because Kimi swallowed the oxygen. The cadence is the signal: releases are now weekly noise, not events. Sources: AI Search; MarkTechPost.

Cisco open-sources Antares-350M and 1B — small models that hunt code vulnerabilities. Cisco's Foundation AI group released two open-weight small language models (350M and 1B) purpose-built to find and localize security bugs in source code. The pitch is economics, not raw capability: a 1B model that runs in CI on every commit — self-hosted, no code leaving the building — is a different security posture than shipping your codebase to a frontier API. Specialized small models win the SOC and the CI pipeline first. Source: MarkTechPost.

Anthropic ships Cowork and extends Fable 5 — again. Anthropic launched Cowork (built by Felix Rieseberg), framed as a shared workspace where humans and Claude agents work side by side rather than a chat window — the lab continuing to climb from API to application. In the same week it pushed Claude Fable 5's availability out for the second time in seven days, a tell that demand is outrunning the planned deprecation schedule. Sources: The Neuron; Creators' AI.

The Future of Life Institute hands the frontier labs a failing report card. FLI published an updated safety assessment grading the major labs, and the coverage framing — "AI Gets a Report Card" — points to low marks across the board on risk management and existential-safety commitments. With binding US legislation stalled, third-party scorecards are becoming the de facto accountability layer boards cite when they ask whether a model vendor is actually safe. Source: Creators' AI.

Watch next week for whether any US lab answers Kimi K3 with open weights of its own — or whether the open frontier stays a one-way street running west out of Beijing.

Devices & robotics

Most of this week's AI oxygen went to a 2.8-trillion-parameter cloud model, a lawsuit, and a chipmaker's earnings. The physical-world thread underneath all of it was on-device inference — the models small enough to run at the edge, and the hardware built to sell them as a feature. It's a positioning week, not a benchmark week, but the direction is clear.

Samsung makes on-device AI the foldable's headline spec

Samsung unveiled its next-generation Galaxy foldables this week and leaned on AI features as the differentiator — not the hinge, not the crease, the model. The read for anyone tracking edge inference: the phone makers have decided the on-device assistant is the marketing. The folding form factor has quietly demoted itself from the story to the vehicle for the story. What's still missing is the spec sheet that matters — which model, running how fast, on what silicon — but the framing shift is the tell. Via The Neuron.

Cisco open-sources two security models small enough to run in CI

Cisco's Foundation AI group released Antares-350M and Antares-1B — open-weight small language models purpose-built to hunt and localize vulnerabilities in source code. At 350M and 1B parameters, these are sizes you run on every commit, self-hosted, with nothing leaving the building. The pitch is economics and data control, not raw capability: a 1B model in your CI pipeline is a fundamentally different security posture than shipping your codebase to a frontier API. It's the clearest sign this week that the specialized on-prem tier is where small models win first. Via MarkTechPost.

The run-it-local model tier keeps filling in

Buried under Kimi K3's cloud-scale release, the same week brought Bonsai 27B and OpenAI's Codex Micro — sizes that sit squarely in the run-it-yourself range. A 27B model, quantized, fits on a single high-memory consumer GPU; a "micro" coding model is built for exactly the latency and privacy budget the cloud can't hit. These are still lightly sourced, so treat the parameter counts as reported rather than benchmarked. But the pattern holds across the week: for every headline giant, a device-class model now ships alongside it, almost as an afterthought. Via AI Search.

What to watch next week: whether Samsung backs the "AI foldable" framing with an actual on-device number — tokens per second, which model, on which NPU — or leaves it at vibes, and whether any of this quarter's small models get a real edge-hardware demo instead of a HuggingFace card. Right now the physical-world AI story is all positioning. The next quarter should force it to show its timings.

What it means for leaders

The throughline this week ran under four unrelated headlines: a Chinese open-weight frontier model, a sold-out chipmaker, a lawsuit between former partners, and a failing safety report card. The models work — the contest has moved to the infrastructure around them.

If you're a CEO this week...

The competitive map redrew itself, and it redrew from Beijing. Moonshot AI shipped Kimi K3, a 2.8-trillion-parameter open-weight model any enterprise can self-host for free — so "AI-native" is no longer a moat you rent from a US lab. Your CFO's Monday question: why pay frontier API prices for a capability your competitor can now download? Meanwhile Intel's forecast blew past estimates on data-center demand — "demand is outpacing our increasing supply" — confirming the compute buildout is real and sold out; the window to commit capacity is now. And Apple sued OpenAI, eighteen months after wiring ChatGPT into Siri: distribution partnerships are becoming litigation, a reputational flag for anyone betting their assistant strategy on one vendor. The board question: when your largest customer asks whether your AI advantage survives a free, self-hostable, Chinese frontier model, what is your answer?

If you're a CIO/CTO this week...

Two releases reshape a reference architecture, both on the small-and-self-hosted axis. Kimi K3's open weights — 2.8T params, a 1M-token context, native vision — put a frontier-scale model behind your own firewall, which changes the data-residency math for anything you can't legally send to a US endpoint. And Cisco's Antares-350M and 1B localize code vulnerabilities in a model small enough to run in CI on every commit — a different security posture than shipping your codebase to a frontier API. Watch vendor exposure: Google's new Flash tier and Poolside's coding MoE keep compressing the cheap-and-fast end, while Anthropic extended Fable 5's availability twice in one week — deprecation schedules move under you. The read: pilot Kimi K3 self-hosted now for data-sensitive workloads; buy frontier APIs only where the capability gap still justifies the egress.

If you lead AI transformation this week...

The sharpest signal for you wasn't a model — it was a critique. The reverse-centaur framing argues consultancies are selling transformation that subordinates your people to the machine's pace instead of augmenting them. Before you scale any engagement, audit whether the workflow puts humans in charge of the AI or turns them into its exception-handlers. Pair that with the week's plan-gate technique — force the agent to produce a reviewable plan before it acts — as a change-management default for every team piloting agents. On pilots, Synthesia's new Roleplay Sessions move corporate training from passive video to interactive simulation, a clean two-week evaluation for one L&D cohort. The experiment to run this month: stand up that Roleplay pilot, and audit one live AI workflow for reverse-centaur design before you roll it wider.

All three roles are being asked the same question from three seats: with the models commoditized, where does your durable advantage actually live — in the compute you can secure, the data you can keep, or the people the AI is supposed to make stronger rather than smaller?


This post is also published on our Substack newsletter at edge-ai.forum. Subscribe for the weekly roundup direct to your inbox — fresh AI news, executive context, and devices + robotics every Friday morning.

// Related