Working Notes on Agent Systems/Brad Zhang

@teach_fireworks / X longform

So the question is, who will be the third child of the Yusan family? 🤔 On Claude's side, Sonnet 5, Fable 5

So the question is, who will be the third child of the Yusan family? 🤔 On Claude's side, Sonnet 5, Fable 5, and Mythos 5 have already rolled out the model echelon, and...

July 10, 2026 · 4 min read

So the question is, who will be the third child of the Yusan family? 🤔 On Claude's side, Sonnet 5, Fable 5
Figure 1 / source image

So the question is, who will be the third child of the Yusan family? 🤔 On Claude's side, Sonnet 5, Fable 5, and Mythos 5 have already rolled out the model echelon, and Claude Code has developed a mature development workflow through code understanding, modification, testing, PR submission, and sub-agent collaboration. OpenAI continues to advance from the long task execution of GPT-5.5 to GPT-5.6. The three dimensions of Sol, Terra, and Luna cover flagship capabilities, cost balance, and high throughput scenarios. Coupled with the parallel Agent workspace of Codex, the model and product have formed a complete closed loop. Therefore, there is almost no suspense in the short term for the first two seats. They no longer occupy only the model list, but also the entrances that developers actually use every day. To judge the next "big three", you must look at at least five things: model upper limit, agent tool chain, developer mentality, ecological distribution, and the comprehensive cost of completing a real task. Many say Gemini is falling behind. My judgment is a little more conservative: Gemini is more like the developers' minds have stalled, and the underlying capabilities have not collapsed. The distribution network composed of Gemini

  • 5, Antigravity, Google Cloud, Workspace and Android is still the most complete set of all candidates. The real problem is that the main product line is not focused enough, and the tools and brands have been adjusted many times, making it difficult to form a clear and stable cognitive anchor like Claude Code and Codex. What may really change the seat rankings is GLM
  • 2. 1M context, long-range tasks, coding agents, engineering specification compliance, coupled with price and Chinese developer ecosystem, GLM
  • 2 is no longer just "a good performer in domestic models", it has begun to have the product foundation to hit the third pole in the world. What is lacking now is mainly the global developer tool chain, enterprise-level distribution and brand mentality. It can take several months to catch up with model capabilities, but it often takes several years to accumulate ecological position. The model upper limit and iteration speed of Grok 4.5 are worthy of vigilance. Behind this are the data and distribution advantages brought by X, SpaceXAI and Cursor training collaboration. However, its developer tool stack is still under construction. It is currently more like a highly capable model and has not yet formed a stable enough productivity ecosystem. Where the Qwen 3.7 Plus excels is in its coverage. Multi-modality, Alibaba Cloud, extensive development framework adaptation, and Qwen’s long-term accumulation of open source influence make it difficult for it to fall out of the first echelon. It currently lacks a super portal like Claude Code or Codex to consolidate these scattered advantages into a unified developer experience. Kimi K2.7 Code is a very sharp knife. It aims at long-term software engineering tasks. Officially, it reduces the number of thinking tokens by about 30% compared to K2.6, and it is already equipped with Kimi Code CLI. It may be a surprise in the Coding category, but to take the third overall position, it also needs to complete its general capabilities, corporate ecology and global distribution. Therefore, if you must make a bet today: the overall third place is still Gemini; the one most likely to complete the change of position on the Coding Agent track is GLM
  • 2. Qwen is the most stable, Grok is the most variable, and Kimi is the sharpest. The definition of Yusanjia has also changed. What everyone is competing for is three sets of productivity systems, not just three model rankings. Once Benchmark first, it is difficult to trade for a long-term chair. Models, Agents, CLI/IDE, tool protocols, cloud, billing and enterprise governance, whoever twists these things into a closed loop first can truly secure the third seat. The third player in the industry will be very interesting to watch in the next few months. The size of the models has basically reached the T level. Scaling Law can still run for at least five years. This is already a consensus among AI prototype factories. Where are the variables? It is high-quality data and ecological expansion. Then back-train the model.

Visual summary

Article argument map

Generated from the post's content graph

FORMATTOPICCAPABILITYMARKETcoverscoverscoverscoverscoverssignalssignalssignalsFORMATarticle featureTOPICagentsTOPICtechnical distributionTOPICmemoryTOPICretrievalTOPICevaluationCAPABILITYagent workflowCAPABILITYdeveloper toolingCAPABILITYevaluationCAPABILITYproduct surface
Mermaid outline
flowchart LR
  format-article["article feature"]
  topic-agents["agents"]
  topic-technical-distribution["technical distribution"]
  topic-memory["memory"]
  topic-retrieval["retrieval"]
  topic-evaluation["evaluation"]
  capability-agent-workflow["agent workflow"]
  capability-developer-tooling["developer tooling"]
  capability-evaluation["evaluation"]
  capability-product-surface["product surface"]
  format-article -->|covers| topic-agents
  format-article -->|covers| topic-technical-distribution
  format-article -->|covers| topic-memory
  format-article -->|covers| topic-retrieval
  format-article -->|covers| topic-evaluation
  format-article -->|signals| capability-agent-workflow
  format-article -->|signals| capability-developer-tooling
  format-article -->|signals| capability-evaluation

Visual structure

Essay structure map

Built from summary and key paragraph positions

So the question is, who will be the third child of the Yusan family? 🤔 On Claude's s...THESISSo the question is,who will be the thirdchild of the Yusanfamily? 🤔 On Claude'sSIGNALSo the question is,who will be the thirdchild of the Yusanfamily? 🤔 On Claude'sOPERATOR5.2. 1M context,long-range tasks,coding agents,engineeringIMPLICATION5.2. Qwen is the moststable, Grok is themost variable, andKimi is the sharpest.
Mermaid outline
flowchart LR
  thesis["So the question is, who will be the third child of the Yusan family? 🤔 On Claude's side, Sonnet 5, Fable 5..."]
  signal["So the question is, who will be the third child of the Yusan family? 🤔 On Claude's side, Sonnet 5, Fable 5..."]
  operator["5.2. 1M context, long-range tasks, coding agents, engineering specification compliance, coupled with price..."]
  implication["5.2. Qwen is the most stable, Grok is the most variable, and Kimi is the sharpest. The definition of Yusanj..."]
  thesis -->|frames| signal
  signal -->|develops| operator
  operator -->|lands in| implication

Source: View the original post