Working Notes on Agent Systems/Brad Zhang

@teach_fireworks / X longform

Traditional large language models are often limited to short context windows such as 8K, 32K or 128K.

Traditional large language models are often limited to short context windows such as 8K, 32K or 128K. Expanding the context window to millions of tokens is not just a...

August 13, 2025 · 2 min read

Traditional large language models are often limited to short context windows such as 8K, 32K or 128K. Expanding the context window to millions of tokens is not just a linear improvement, but brings an essential breakthrough in capabilities. When a model can process hundreds of pages of text simultaneously, it no longer needs to segment information into small pieces to understand, but can grasp the overall structure of the material and the complex connections between content. This ability enables AI to transform from processing fragments to processing complete information systems, thereby achieving deeper understanding and reasoning. It is foreseeable that other large model manufacturers should gradually follow up the basic capabilities of ultra-long context tokens in

  • The competition among the current basic model manufacturers is very interesting. On the one hand, every major model manufacturer must have some unique skills, such as those who draw pictures, write code, produce reports, open source, reason, etc. On the other hand, whenever a phenomenon-level capability is recognized by users, it will almost always follow up within a short period of time. More and more popular functions are delivered directly to users as factory capabilities of the model. This phenomenon seems to be getting worse, and no one wants to fall behind. I remember what Ultraman said before. Those working in the application layer should not think about relying on temporary engineering capabilities to make up for the so-called shortcomings of large models. They should think about how to make full use of the latest model basic capabilities. This sentence has been well verified by Manus. Now, after Manus became popular around the world in the early stages of internal testing, it has a valuation of US$500 million. @manusai This is basically the way the wind is blowing. The application side needs to input more information and requires large models to accurately use the reasoning of large models in different scenarios to give better and longer answers. The shorter the complete life cycle, the faster the speed, and the higher the quality of the results, the better the experience it brings to users. The better the experience, the more determined they will be to renew monthly or annual memberships, and the business story will make sense.

Visual summary

Article argument map

Generated from the post's content graph

FORMATTOPICCAPABILITYMARKETcoverscoverscoverssignalssignalssignalssignalsFORMATlongform noteTOPICai workbenchesTOPICmemoryTOPICretrievalCAPABILITYAI-native workbenchCAPABILITYdeveloper toolingCAPABILITYevaluationMARKETopen-source builders
Mermaid outline
flowchart LR
  format-long_post["longform note"]
  topic-ai-workbenches["ai workbenches"]
  topic-memory["memory"]
  topic-retrieval["retrieval"]
  capability-ai-native-workbench["AI-native workbench"]
  capability-developer-tooling["developer tooling"]
  capability-evaluation["evaluation"]
  market-open-source-builders["open-source builders"]
  format-long_post -->|covers| topic-ai-workbenches
  format-long_post -->|covers| topic-memory
  format-long_post -->|covers| topic-retrieval
  format-long_post -->|signals| capability-ai-native-workbench
  format-long_post -->|signals| capability-developer-tooling
  format-long_post -->|signals| capability-evaluation
  format-long_post -->|signals| market-open-source-builders

Visual structure

Essay structure map

Built from summary and key paragraph positions

Traditional large language models are often limited to short context windows such as...THESISTraditional largelanguage models areoften limited to shortcontext windows suchSIGNALTraditional largelanguage models areoften limited to shortcontext windows suchOPERATOR2025. The competitionamong the currentbasic modelmanufacturers is veryIMPLICATION2025. The competitionamong the currentbasic modelmanufacturers is very
Mermaid outline
flowchart LR
  thesis["Traditional large language models are often limited to short context windows such as 8K, 32K or 128K. Expan..."]
  signal["Traditional large language models are often limited to short context windows such as 8K, 32K or 128K. Expan..."]
  operator["2025. The competition among the current basic model manufacturers is very interesting. On the one hand, eve..."]
  implication["2025. The competition among the current basic model manufacturers is very interesting. On the one hand, eve..."]
  thesis -->|frames| signal
  signal -->|develops| operator
  operator -->|lands in| implication

Source: View the original post