Working Notes on Agent Systems/Brad Zhang

Topic desk

Harness Engineering

What has to surround the model before an AI workflow becomes a product people can trust?

This desk treats applied AI engineering as harness design: context policy, constraints, memory, evals, permissions, cleanup loops, and product surfaces working as one operating layer.

Live source unreachable; showing the most recent snapshot.

Longform

All longform

OpenAI Proactively Discloses Long Range Agent Tasks Out of Bounds: The Safe Steering Behind PR # 287

A single tool approval begins to lose global semantics when the model can be tried for hours on end. On July 20, 2026, OpenAI unveiled an on-premises pause. A generic model that...

Jul 21, 2026 · 12 min read

A summary of the best articles on Graph Engineering! Prompt, the next engineering parad...

A summary of the best articles on Graph Engineering! Prompt, the next engineering paradigm from Loop to Graph, determines how the Agent starts. Context determines what the Agent...

Jul 20, 2026 · 4 min read

ChatGPT Work: AI starts delivering work for you In the past, we used ChatGPT to get answers.

ChatGPT Work: AI starts delivering work for you In the past, we used ChatGPT to get answers. Now, OpenAI is thinking about another thing: letting AI pick up a complete target, f...

Jul 19, 2026 · 1 min read

The Agent's memory should not be stuffed into the context window

When many people discuss coding agents such as Codex and Claude Code, their first reaction is model capabilities: whether they can read code, correct bugs, and...

Jul 17, 2026 · 8 min read

I connected Codex and Cursor: an automatic memory bus allows the two Agent software to communicate freely

Recently I am doing a very specific thing: using the source code of pi agent and a set of small exercises. I hope that Codex will be responsible for the reading history, combing...

Jul 14, 2026 · 11 min read

I tested GPT-5.5 and GPT-5.6 SOL using 5 real front-end scenarios: SOL won 4 games but was 13% slower

On July 10, 2026, I put GPT-5.5 and GPT-5.6 SOL into the same Codex execution chain and performed five sets of front-end tasks in succession. Five scenes, ten pages, and...

Jul 11, 2026 · 12 min read

fireworks-tech-graph Biggest update yet! 🎉This was originally a small open source project and has

fireworks-tech-graph Biggest update yet! 🎉This was originally a small open source project and has accumulated 8.5k stars after several months of development. Thank you...

Jul 10, 2026 · 3 min read

My first heavy duty job with GPT 5.6 sol blew me away, kudos! I suggest you send this sentence to your Codex

My first heavy duty job with GPT 5.6 sol blew me away, kudos! I suggest you send this sentence to your Codex and wait silently for the output results. You will come back...

Jul 9, 2026 · 1 min read

Anthropic’s 85-minute Fable 5 workshop explains how to implement the next generation AI Agent

Video entrance Opening keynote: https://www.youtube.com/watch?v=GMIWm5y90xA The capability curve: https://www.youtube.com/watch?v=tP4MGcJ80Y0 How to get to production...

Jul 8, 2026 · 18 min read

The significance of J-space is to give the previously black-box model reasoning an opportunity for

The significance of J-space is to give the previously black-box model reasoning an opportunity for observability and human intervention and control, so that corrections can...

Jul 7, 2026 · 2 min read

The term software factory has finally begun to move from a slogan to engineering sites.

Original video: WF2026: Software Factories & Keynotes ft. Microsoft, OpenAI, OpenClaw, Z.ai (GLM), MiniMax, HF Link: https://www.youtube.com/watch?v=htM02KMNZnk Start I just fin...

Jul 4, 2026 · 8 min read

Loop Engineering recommends: At the end of each day, let the agent do an architecture physical examination.

Loop Engineering recommends: At the end of each day, let the agent do an architecture physical examination. The original prompt word of the architecture satisfaction loop is ver...

Jun 22, 2026 · 1 min read

The most worth-watching high-quality videos in the field of AI in the past two weeks! 1...

The most worth-watching high-quality videos in the field of AI in the past two weeks! 1. Official Team - Bloomberg Odd Lots × Anthropic Co-Founder: Anthropic's Co-Founder and To...

Jun 21, 2026 · 2 min read

From Prompt Engineering to Harness Engineering: four evolutions of the AI ​​engineering system

The rapid development of products such as Agent, Codex, and Claude Code has led to the emergence of a number of new terms in the field of AI engineering: Prompt...

Jun 19, 2026 · 7 min read

Summary of the best articles of Loop Engineering! After Agent began to focus on long tasks in 2026,

Summary of the best articles of Loop Engineering! After Agent began to focus on long tasks in 2026, the focus slowly became: How to design a loop system that can...

Jun 18, 2026 · 2 min read

When dynamic workflow and front-end design are combined, there will be unexpected effects!

When dynamic workflow and front-end design are combined, there will be unexpected effects! Share an open source project: fireworks-design This project is easily misunderstood as...

Jun 16, 2026 · 5 min read

Claude Fable is for enterprises to replace traditional mid-to-high-level development, and the ROI is positive.

Claude Fable is for enterprises to replace traditional mid-to-high-level development, and the ROI is positive. After the release of the Claude Fable model, I saw many crazy exam...

Jun 10, 2026 · 4 min read

Agent Loop is popular🔥: Agent goes from Demo to production with a Loop in between

There has been a heated debate in overseas AI circles these days. Matt Van Horn sorted out this debate in “WTF Is a Loop? Peter Steinberger vs. Boris Cherny”: You shouldn’t be p...

Jun 9, 2026 · 10 min read

Five Traps of Agent Harness 1. Self-evaluation is a trap. Use an adversarial evaluator....

Five Traps of Agent Harness 1. Self-evaluation is a trap. Use an adversarial evaluator. Self-evaluation is a trap. Use an adversarial evaluator. Many teams will do this: Agent g...

Jun 8, 2026 · 3 min read

Why does it seem easy to develop an AI Agent, but so difficult to actually make it "useful"? Where are the main bottlenecks?

Let me ask you a question, have you heard of or used more than half of the Agent products mentioned in the picture below? More than a third? Over the past 25 years, in less than...

Jun 7, 2026 · 8 min read

When Agent acts as the first citizen, there are three key changes: 1: Agent is not a revolution, but an acceleration.

When Agent acts as the first citizen, there are three key changes: 1: Agent is not a revolution, but an acceleration. It does not change the basic structure of software evolutio...

Jun 3, 2026 · 2 min read

Pi Agent = customizable coding harness / agent workbench Pi Agent can be understood as a minimalist, self-transformable terminal coding agent harness.

Pi Agent = customizable coding harness / agent workbench Pi Agent can be understood as a minimalist, self-transformable terminal coding agent harness. It is not a general orches...

Jun 2, 2026 · 2 min read

Asynchronous Agent Era: From Code Assistant to Software Factory

I have reorganized the content of this podcast of Latent Space. It is worth taking a look at the in-depth discussion and sharing related to asynchronous Agent. I also...

May 31, 2026 · 22 min read

From the evolution of synchronous Agent to asynchronous Agent, I systematically sorted out the relevant content!

From the evolution of synchronous Agent to asynchronous Agent, I systematically sorted out the relevant content! The watershed between Async Agent and ordinary Agent is not "whe...

May 31, 2026 · 4 min read

In the afternoon, I was surprised when I saw the news about the layoffs of a large factory posted by @MaxForAI, and then I sent a private message to verify it.

In the afternoon, I was surprised when I saw the news about the layoffs of a large factory posted by @MaxForAI, and then I sent a private message to verify it. I still believe M...

May 27, 2026 · 7 min read

My exploration and practice sharing in the past two months: Minimum engineering closed loop for long task Agent

A picture shared by engineer Claude accurately describes the three most common problems of long-task agents: the inability to keep track of the state, the inability...

May 25, 2026 · 27 min read

After Sonnet 3.5 came out, both Anthropic and users found that its coding capabilities...

After Sonnet 3.5 came out, both Anthropic and users found that its coding capabilities were particularly strong, and an adventure based on coding capabilities began. The concept...

May 25, 2026 · 2 min read

The main line of Agent development in 2026: moving from Agent Framework to Agent Runtime.

The main line of Agent development in 2026: moving from Agent Framework to Agent Runtime. The most obvious change in 2026 is that the Agent technology stack begins to be layered...

May 22, 2026 · 2 min read

Visualize that with the rapid development of agent cli and harness, the interaction and logic core will soon

Visualize that with the rapid development of agent cli and harness, the interaction and logic core will soon be reconstructed by Golang, which is very suitable for...

May 19, 2026 · 1 min read

There are more than a hundred and five thousand subscriptions. I don't know if there wi...

There are more than a hundred and five thousand subscriptions. I don't know if there will be any surprises when I wake up. I often write in advance to celebrate the 5k subscript...

May 19, 2026 · 2 min read

Using the WeChat CLI recommended by Arbor Teacher, I tried a reading method that is very suitable for technicians.

Using the WeChat CLI recommended by Arbor Teacher, I tried a reading method that is very suitable for technicians. Let Codex see what I'm really doing during this time...

May 17, 2026 · 2 min read

A Harness Engineering Best Practices case completed in collaboration with Codex. Today, the friends saw that

A Harness Engineering Best Practices case completed in collaboration with Codex. Today, the friends saw that Codex was about to reset the quota again, so they adjusted...

May 16, 2026 · 2 min read

A new way to play "Codex app" on mobile! The mobile "Codex app" essentially does not ha...

A new way to play "Codex app" on mobile! The mobile "Codex app" essentially does not have a separate development environment, it is just the Codex remote console in the ChatGPT...

May 16, 2026 · 2 min read

Is Grep All You Need? If you are still thinking of “rag = Vector Database” as the defau...

Is Grep All You Need? If you are still thinking of “rag = Vector Database” as the default answer, then I recommend that you read this paper. It is a bit counterintuitive: the au...

May 16, 2026 · 3 min read

Based on the recommendation of the boss, I combed through a list of high-quality AI Engineer learning materials, which is worth collecting and learning!

Based on the recommendation of the boss, I combed through a list of high-quality AI Engineer learning materials, which is worth collecting and learning! Too dry, too...

May 11, 2026 · 3 min read

HTML might be the real interface of the agent era. Markdown is great for notes. But age...

HTML might be the real interface of the agent era. Markdown is great for notes. But agents do not just need notes. They need interfaces they can understand, manipulate, simulate...

May 9, 2026 · 2 min read

It is highly recommended that you read this long article written by Thariq, a core member of ClaudeCode: he brings an interesting perspective.

It is highly recommended that you read this long article written by Thariq, a core member of ClaudeCode: he brings an interesting perspective. HTML is the real working interface...

May 9, 2026 · 2 min read

During this time, I have been using the Codex App and the Codex CLI to develop iOS Apps at the same time

During this time, I have been using the Codex App and the Codex CLI to develop iOS Apps at the same time, especially in the SwiftUI + Xcode workflow, and it is becoming...

May 6, 2026 · 3 min read

The OpenAI Agent SDK is crazy! The "Harness/Compute Separation" architecture completely solves the security

The OpenAI Agent SDK is crazy! The "Harness/Compute Separation" architecture completely solves the security and scalability pain points of the long-range Agent. The🔥...

May 5, 2026 · 3 min read

Teacher Wu Yun’s introduction to codex goal is very detailed. I recommend everyone to r...

Teacher Wu Yun’s introduction to codex goal is very detailed. I recommend everyone to read it. At the same time, I will share a sample goal command that I verified and sent a fe...

May 4, 2026 · 2 min read

Try this cli radio, you can listen to music, podcasts and news while vibe coding. Why d...

Try this cli radio, you can listen to music, podcasts and news while vibe coding. Why do you want to do this? The reason is very small, just want to make some sound when writing...

May 3, 2026 · 3 min read

Using Codex goals well can truly unleash the potential of the GPT 5 model! OpenAI's rec...

Using Codex goals well can truly unleash the potential of the GPT 5 model! OpenAI's recently disclosed directions for Codex are obviously going in the following directions: - Lo...

May 1, 2026 · 6 min read

Many people underestimate a problem: the real danger of Agent is not the occasional wrong answer, but "systematic drift"

Many people discuss Agent, and their focus still remains on: - Is the answer correct this time - Is the tool adjusted accurately this time - Is the task completed this...

Apr 30, 2026 · 9 min read

Codex can actually not only do coding, but can also be used as a training partner! Try the following:

Codex can actually not only do coding, but can also be used as a training partner! Try the following: 1. First define "real mastery". Don't regard "understanding" as mastery. Fo...

Apr 29, 2026 · 4 min read

How to make Codex Silky continue to run long tasks? Make good use of these two commands...

How to make Codex Silky continue to run long tasks? Make good use of these two commands codex --full-auto: suitable for 80% of daily scenarios. If your goal is to have Codex int...

Apr 27, 2026 · 2 min read

Recently, I have a more obvious feeling: Codex is so easy to use that it is addictive, but as long as it is interrupted all the time, the state will be broken.

Recently, I have a more obvious feeling: Codex is so easy to use that it is addictive, but as long as it is interrupted all the time, the state will be broken. What annoys me mo...

Apr 27, 2026 · 3 min read

🔄 Fireworks-tech-graph major version update! Thanks again to @berryxia and all my frie...

🔄 Fireworks-tech-graph major version update! Thanks again to @berryxia and all my friends for their support. More than 100,000 exposures were made in one day. I didn’t do anyth...

Apr 11, 2026 · 2 min read

Fireworks-skill-memory major update! ! I received a lot of feedback after I posted it last time. I launched

Fireworks-skill-memory major update! ! I received a lot of feedback after I posted it last time. I launched v4 today and fixed several real problems I found after using...

Apr 5, 2026 · 3 min read

Why is this Harness open source project worth using? Claude Code's Skill ecosystem is r...

Why is this Harness open source project worth using? Claude Code's Skill ecosystem is rapidly expanding (official plugin marketplace + community skill). But "Claude can't rememb...

Mar 28, 2026 · 2 min read

A Harness engineering practice, I open sourced it! It’s quite interesting. After using...

A Harness engineering practice, I open sourced it! It’s quite interesting. After using Claude Code for two months, I discovered a crazy problem: it keeps making the same mistake...

Mar 27, 2026 · 2 min read

that's right Andrey, I'm an architect. Base on my expreince: The most difficult part of...

that's right Andrey, I'm an architect. Base on my expreince: The most difficult part of building software applications is orchestrating various services: account systems, paymen...

Mar 26, 2026 · 1 min read

Lin Junyang’s latest must-read masterpiece! AI has shifted from "reasoning thinking" to "agentic thinking".

Lin Junyang’s latest must-read masterpiece! AI has shifted from "reasoning thinking" to "agentic thinking". 🏀The frontier of competition is shifting from "better...

Mar 26, 2026 · 3 min read

Sora's departure is a good thing for OpenAI and AI entrepreneurs. Big companies do big things! I just woke up

Sora's departure is a good thing for OpenAI and AI entrepreneurs. Big companies do big things! I just woke up in the morning and saw that a group of friends had...

Mar 24, 2026 · 2 min read

Harness Engineering: The next battlefield for AI engineers! AI Agent = Model + Harness....

Harness Engineering: The next battlefield for AI engineers! AI Agent = Model + Harness. If you're not a model, you're a Harness. OpenAI built 1 million lines of production code...

Mar 22, 2026 · 2 min read

Experience sharing of using ob + claudian + palywright-skills + ob-skills + general-purpose-skills to generate visual logs

At one o'clock in the morning, I stared at the Canvas file just generated in Obsidian, feeling a little dazed. Figure 1: Automatically generated visual log...

Jan 14, 2026 · 10 min read

With this Orange article, I will continue to share some content about ClaudeSkill and MCP to attract some traffic, haha.

With this Orange article, I will continue to share some content about ClaudeSkill and MCP to attract some traffic, haha. Analysis of the core features of Claude Skills....

Dec 29, 2025 · 5 min read

Why do I always recommend using Claude Code? Model capabilities are growing exponential...

Why do I always recommend using Claude Code? Model capabilities are growing exponentially (Exponential) 🚀 But the UX of development tools is still climbing linearly (Linear) 📷...

Nov 27, 2025 · 2 min read

Why do I always recommend using Claude Code? Model capabilities are growing exponential...

Why do I always recommend using Claude Code? Model capabilities are growing exponentially (Exponential) 🚀 But the UX of development tools is still climbing linearly (Linear) 🐢...

Nov 27, 2025 · 2 min read

Claude recently added an official skill specifically designed to improve front-end interface design.

Claude recently added an official skill specifically designed to improve front-end interface design. I tried installing it and it worked really well. The screenshot is...

Nov 16, 2025 · 4 min read

Build the "Microservices Architecture Review Expert" Skill ❌ Inefficient way of writing: "You are an architect, help me review this microservice design."

Build the "Microservices Architecture Review Expert" Skill ❌ Inefficient way of writing: "You are an architect, help me review this microservice design." Problem: Too general...

Oct 20, 2025 · 2 min read

Claude finally takes action on the browser! I think this is indeed a good time. Claude...

Claude finally takes action on the browser! I think this is indeed a good time. Claude code has fully verified the coding capabilities based on its own model, coupled with a wel...

Aug 27, 2025 · 2 min read

ClaudeCode best practices: Part 1: Provide clear context just like communicating with people.

ClaudeCode best practices: Part 1: Provide clear context just like communicating with people. First of all, the first and most important core concept shared by the official is:...

Aug 4, 2025 · 4 min read

Field Notes

Full ledger

Live judgment for this desk is filed by date in the field-notes ledger; enter through the full ledger by month.