Moonshot’s Kimi K2.6 is open weight, and its headline feature is an agent swarm. The vendor’s technical blog says the swarm scales to 300 sub-agents across 4,000 coordinated steps, up from 100 sub-agents and 1,500 steps in K2.5. Those are Moonshot’s own numbers. Launch chatter rounds them up to a thousand agents. I could not find a thousand in anything Moonshot itself published, so I am ignoring it and sticking to 300. The weights are on Hugging Face: https://huggingface.co/moonshotai/Kimi-K2.6
The pitch is easy to like. Context windows end, even generous ones, so slice the work and hand each agent a slice. Wall-clock time collapses. Every agent works with a small, fresh context instead of one long loop drowning in its own history. It sounds like the answer to the context problem.
I think it moves the problem. Every interesting defect I have met lives at a boundary. The auth module assumes something the billing module never promised. The migration handles every case except the one the API actually sends. A sliced agent never sees a boundary. Its slice is coherent, and coherence is cheap when you draw the borders yourself. Merge three hundred coherent slices and the seams between them are averaged, not checked. Nobody ran the path that crosses from one slice into another, because that path belonged to no agent.
Then there is the reviewer. Someone has to hold the whole design to judge three hundred outputs, and that someone is you. The context bottleneck did not disappear. It relocated to the one human who signs off. A swarm that outproduces its reviewer just builds a bigger queue of unchecked work.
Moonshot’s own showcase reads like a tell. The product page leads with documents, slides, and spreadsheets completed in a single run. Those are tasks where the seams are thin. One section of a report rarely breaks another the way one module breaks another. I buy swarms for work that is genuinely separable. Software mostly is not.
So here is my boring answer. Fewer agents, one shared written plan. The plan names the interfaces, the invariants, and what done means, in words every worker can read. Parallelize against the plan instead of the problem statement, and check each result against it before anything merges. The plan is the shared context that no window can fit.
I have not run the swarm myself, so this is an argument from experience, not a benchmark. But I have watched human teams parallelize exactly this way. It ends with an integration phase that costs everything the parallel phase saved. The plan is cheaper up front. Give me one good agent reading from it over a thousand agents without one.
