Topic desk
Harness Engineering
What has to surround the model before an AI workflow becomes a product people can trust?
This desk treats applied AI engineering as harness design: context policy, constraints, memory, evals, permissions, cleanup loops, and product surfaces working as one operating layer.
Longform
All longformOpenAI Proactively Discloses Long Range Agent Tasks Out of Bounds: The Safe Steering Behind PR # 287
A single tool approval begins to lose global semantics when the model can be tried for hours on end. On July 20, 2026, OpenAI unveiled an on-premises pause. A generic model that...
A summary of the best articles on Graph Engineering! Prompt, the next engineering parad...
A summary of the best articles on Graph Engineering! Prompt, the next engineering paradigm from Loop to Graph, determines how the Agent starts. Context determines what the Agent...
ChatGPT Work: AI starts delivering work for you In the past, we used ChatGPT to get answers.
ChatGPT Work: AI starts delivering work for you In the past, we used ChatGPT to get answers. Now, OpenAI is thinking about another thing: letting AI pick up a complete target, f...
The Agent's memory should not be stuffed into the context window
When many people discuss coding agents such as Codex and Claude Code, their first reaction is model capabilities: whether they can read code, correct bugs, and...
I connected Codex and Cursor: an automatic memory bus allows the two Agent software to communicate freely
Recently I am doing a very specific thing: using the source code of pi agent and a set of small exercises. I hope that Codex will be responsible for the reading history, combing...
I tested GPT-5.5 and GPT-5.6 SOL using 5 real front-end scenarios: SOL won 4 games but was 13% slower
On July 10, 2026, I put GPT-5.5 and GPT-5.6 SOL into the same Codex execution chain and performed five sets of front-end tasks in succession. Five scenes, ten pages, and...
fireworks-tech-graph Biggest update yet! 🎉This was originally a small open source project and has
fireworks-tech-graph Biggest update yet! 🎉This was originally a small open source project and has accumulated 8.5k stars after several months of development. Thank you...
My first heavy duty job with GPT 5.6 sol blew me away, kudos! I suggest you send this sentence to your Codex
My first heavy duty job with GPT 5.6 sol blew me away, kudos! I suggest you send this sentence to your Codex and wait silently for the output results. You will come back...
Anthropic’s 85-minute Fable 5 workshop explains how to implement the next generation AI Agent
Video entrance Opening keynote: https://www.youtube.com/watch?v=GMIWm5y90xA The capability curve: https://www.youtube.com/watch?v=tP4MGcJ80Y0 How to get to production...
The significance of J-space is to give the previously black-box model reasoning an opportunity for
The significance of J-space is to give the previously black-box model reasoning an opportunity for observability and human intervention and control, so that corrections can...
The term software factory has finally begun to move from a slogan to engineering sites.
Original video: WF2026: Software Factories & Keynotes ft. Microsoft, OpenAI, OpenClaw, Z.ai (GLM), MiniMax, HF Link: https://www.youtube.com/watch?v=htM02KMNZnk Start I just fin...
Loop Engineering recommends: At the end of each day, let the agent do an architecture physical examination.
Loop Engineering recommends: At the end of each day, let the agent do an architecture physical examination. The original prompt word of the architecture satisfaction loop is ver...
The most worth-watching high-quality videos in the field of AI in the past two weeks! 1...
The most worth-watching high-quality videos in the field of AI in the past two weeks! 1. Official Team - Bloomberg Odd Lots × Anthropic Co-Founder: Anthropic's Co-Founder and To...
From Prompt Engineering to Harness Engineering: four evolutions of the AI engineering system
The rapid development of products such as Agent, Codex, and Claude Code has led to the emergence of a number of new terms in the field of AI engineering: Prompt...
Summary of the best articles of Loop Engineering! After Agent began to focus on long tasks in 2026,
Summary of the best articles of Loop Engineering! After Agent began to focus on long tasks in 2026, the focus slowly became: How to design a loop system that can...
When dynamic workflow and front-end design are combined, there will be unexpected effects!
When dynamic workflow and front-end design are combined, there will be unexpected effects! Share an open source project: fireworks-design This project is easily misunderstood as...
Claude Fable is for enterprises to replace traditional mid-to-high-level development, and the ROI is positive.
Claude Fable is for enterprises to replace traditional mid-to-high-level development, and the ROI is positive. After the release of the Claude Fable model, I saw many crazy exam...
Agent Loop is popular🔥: Agent goes from Demo to production with a Loop in between
There has been a heated debate in overseas AI circles these days. Matt Van Horn sorted out this debate in “WTF Is a Loop? Peter Steinberger vs. Boris Cherny”: You shouldn’t be p...
Five Traps of Agent Harness 1. Self-evaluation is a trap. Use an adversarial evaluator....
Five Traps of Agent Harness 1. Self-evaluation is a trap. Use an adversarial evaluator. Self-evaluation is a trap. Use an adversarial evaluator. Many teams will do this: Agent g...
Why does it seem easy to develop an AI Agent, but so difficult to actually make it "useful"? Where are the main bottlenecks?
Let me ask you a question, have you heard of or used more than half of the Agent products mentioned in the picture below? More than a third? Over the past 25 years, in less than...
When Agent acts as the first citizen, there are three key changes: 1: Agent is not a revolution, but an acceleration.
When Agent acts as the first citizen, there are three key changes: 1: Agent is not a revolution, but an acceleration. It does not change the basic structure of software evolutio...
Pi Agent = customizable coding harness / agent workbench Pi Agent can be understood as a minimalist, self-transformable terminal coding agent harness.
Pi Agent = customizable coding harness / agent workbench Pi Agent can be understood as a minimalist, self-transformable terminal coding agent harness. It is not a general orches...
Asynchronous Agent Era: From Code Assistant to Software Factory
I have reorganized the content of this podcast of Latent Space. It is worth taking a look at the in-depth discussion and sharing related to asynchronous Agent. I also...
From the evolution of synchronous Agent to asynchronous Agent, I systematically sorted out the relevant content!
From the evolution of synchronous Agent to asynchronous Agent, I systematically sorted out the relevant content! The watershed between Async Agent and ordinary Agent is not "whe...
In the afternoon, I was surprised when I saw the news about the layoffs of a large factory posted by @MaxForAI, and then I sent a private message to verify it.
In the afternoon, I was surprised when I saw the news about the layoffs of a large factory posted by @MaxForAI, and then I sent a private message to verify it. I still believe M...
My exploration and practice sharing in the past two months: Minimum engineering closed loop for long task Agent
A picture shared by engineer Claude accurately describes the three most common problems of long-task agents: the inability to keep track of the state, the inability...
After Sonnet 3.5 came out, both Anthropic and users found that its coding capabilities...
After Sonnet 3.5 came out, both Anthropic and users found that its coding capabilities were particularly strong, and an adventure based on coding capabilities began. The concept...
The main line of Agent development in 2026: moving from Agent Framework to Agent Runtime.
The main line of Agent development in 2026: moving from Agent Framework to Agent Runtime. The most obvious change in 2026 is that the Agent technology stack begins to be layered...
Visualize that with the rapid development of agent cli and harness, the interaction and logic core will soon
Visualize that with the rapid development of agent cli and harness, the interaction and logic core will soon be reconstructed by Golang, which is very suitable for...
There are more than a hundred and five thousand subscriptions. I don't know if there wi...
There are more than a hundred and five thousand subscriptions. I don't know if there will be any surprises when I wake up. I often write in advance to celebrate the 5k subscript...
Using the WeChat CLI recommended by Arbor Teacher, I tried a reading method that is very suitable for technicians.
Using the WeChat CLI recommended by Arbor Teacher, I tried a reading method that is very suitable for technicians. Let Codex see what I'm really doing during this time...
A Harness Engineering Best Practices case completed in collaboration with Codex. Today, the friends saw that
A Harness Engineering Best Practices case completed in collaboration with Codex. Today, the friends saw that Codex was about to reset the quota again, so they adjusted...
A new way to play "Codex app" on mobile! The mobile "Codex app" essentially does not ha...
A new way to play "Codex app" on mobile! The mobile "Codex app" essentially does not have a separate development environment, it is just the Codex remote console in the ChatGPT...
Is Grep All You Need? If you are still thinking of “rag = Vector Database” as the defau...
Is Grep All You Need? If you are still thinking of “rag = Vector Database” as the default answer, then I recommend that you read this paper. It is a bit counterintuitive: the au...
Based on the recommendation of the boss, I combed through a list of high-quality AI Engineer learning materials, which is worth collecting and learning!
Based on the recommendation of the boss, I combed through a list of high-quality AI Engineer learning materials, which is worth collecting and learning! Too dry, too...
HTML might be the real interface of the agent era. Markdown is great for notes. But age...
HTML might be the real interface of the agent era. Markdown is great for notes. But agents do not just need notes. They need interfaces they can understand, manipulate, simulate...
It is highly recommended that you read this long article written by Thariq, a core member of ClaudeCode: he brings an interesting perspective.
It is highly recommended that you read this long article written by Thariq, a core member of ClaudeCode: he brings an interesting perspective. HTML is the real working interface...
During this time, I have been using the Codex App and the Codex CLI to develop iOS Apps at the same time
During this time, I have been using the Codex App and the Codex CLI to develop iOS Apps at the same time, especially in the SwiftUI + Xcode workflow, and it is becoming...
The OpenAI Agent SDK is crazy! The "Harness/Compute Separation" architecture completely solves the security
The OpenAI Agent SDK is crazy! The "Harness/Compute Separation" architecture completely solves the security and scalability pain points of the long-range Agent. The🔥...
Teacher Wu Yun’s introduction to codex goal is very detailed. I recommend everyone to r...
Teacher Wu Yun’s introduction to codex goal is very detailed. I recommend everyone to read it. At the same time, I will share a sample goal command that I verified and sent a fe...
Try this cli radio, you can listen to music, podcasts and news while vibe coding. Why d...
Try this cli radio, you can listen to music, podcasts and news while vibe coding. Why do you want to do this? The reason is very small, just want to make some sound when writing...
Using Codex goals well can truly unleash the potential of the GPT 5 model! OpenAI's rec...
Using Codex goals well can truly unleash the potential of the GPT 5 model! OpenAI's recently disclosed directions for Codex are obviously going in the following directions: - Lo...
Many people underestimate a problem: the real danger of Agent is not the occasional wrong answer, but "systematic drift"
Many people discuss Agent, and their focus still remains on: - Is the answer correct this time - Is the tool adjusted accurately this time - Is the task completed this...
Codex can actually not only do coding, but can also be used as a training partner! Try the following:
Codex can actually not only do coding, but can also be used as a training partner! Try the following: 1. First define "real mastery". Don't regard "understanding" as mastery. Fo...
How to make Codex Silky continue to run long tasks? Make good use of these two commands...
How to make Codex Silky continue to run long tasks? Make good use of these two commands codex --full-auto: suitable for 80% of daily scenarios. If your goal is to have Codex int...
Recently, I have a more obvious feeling: Codex is so easy to use that it is addictive, but as long as it is interrupted all the time, the state will be broken.
Recently, I have a more obvious feeling: Codex is so easy to use that it is addictive, but as long as it is interrupted all the time, the state will be broken. What annoys me mo...
🔄 Fireworks-tech-graph major version update! Thanks again to @berryxia and all my frie...
🔄 Fireworks-tech-graph major version update! Thanks again to @berryxia and all my friends for their support. More than 100,000 exposures were made in one day. I didn’t do anyth...
Fireworks-skill-memory major update! ! I received a lot of feedback after I posted it last time. I launched
Fireworks-skill-memory major update! ! I received a lot of feedback after I posted it last time. I launched v4 today and fixed several real problems I found after using...
Why is this Harness open source project worth using? Claude Code's Skill ecosystem is r...
Why is this Harness open source project worth using? Claude Code's Skill ecosystem is rapidly expanding (official plugin marketplace + community skill). But "Claude can't rememb...
A Harness engineering practice, I open sourced it! It’s quite interesting. After using...
A Harness engineering practice, I open sourced it! It’s quite interesting. After using Claude Code for two months, I discovered a crazy problem: it keeps making the same mistake...
that's right Andrey, I'm an architect. Base on my expreince: The most difficult part of...
that's right Andrey, I'm an architect. Base on my expreince: The most difficult part of building software applications is orchestrating various services: account systems, paymen...
Lin Junyang’s latest must-read masterpiece! AI has shifted from "reasoning thinking" to "agentic thinking".
Lin Junyang’s latest must-read masterpiece! AI has shifted from "reasoning thinking" to "agentic thinking". 🏀The frontier of competition is shifting from "better...
Sora's departure is a good thing for OpenAI and AI entrepreneurs. Big companies do big things! I just woke up
Sora's departure is a good thing for OpenAI and AI entrepreneurs. Big companies do big things! I just woke up in the morning and saw that a group of friends had...
Harness Engineering: The next battlefield for AI engineers! AI Agent = Model + Harness....
Harness Engineering: The next battlefield for AI engineers! AI Agent = Model + Harness. If you're not a model, you're a Harness. OpenAI built 1 million lines of production code...
Experience sharing of using ob + claudian + palywright-skills + ob-skills + general-purpose-skills to generate visual logs
At one o'clock in the morning, I stared at the Canvas file just generated in Obsidian, feeling a little dazed. Figure 1: Automatically generated visual log...
With this Orange article, I will continue to share some content about ClaudeSkill and MCP to attract some traffic, haha.
With this Orange article, I will continue to share some content about ClaudeSkill and MCP to attract some traffic, haha. Analysis of the core features of Claude Skills....
Why do I always recommend using Claude Code? Model capabilities are growing exponential...
Why do I always recommend using Claude Code? Model capabilities are growing exponentially (Exponential) 🚀 But the UX of development tools is still climbing linearly (Linear) 📷...
Why do I always recommend using Claude Code? Model capabilities are growing exponential...
Why do I always recommend using Claude Code? Model capabilities are growing exponentially (Exponential) 🚀 But the UX of development tools is still climbing linearly (Linear) 🐢...
Claude recently added an official skill specifically designed to improve front-end interface design.
Claude recently added an official skill specifically designed to improve front-end interface design. I tried installing it and it worked really well. The screenshot is...
Build the "Microservices Architecture Review Expert" Skill ❌ Inefficient way of writing: "You are an architect, help me review this microservice design."
Build the "Microservices Architecture Review Expert" Skill ❌ Inefficient way of writing: "You are an architect, help me review this microservice design." Problem: Too general...
Claude finally takes action on the browser! I think this is indeed a good time. Claude...
Claude finally takes action on the browser! I think this is indeed a good time. Claude code has fully verified the coding capabilities based on its own model, coupled with a wel...
ClaudeCode best practices: Part 1: Provide clear context just like communicating with people.
ClaudeCode best practices: Part 1: Provide clear context just like communicating with people. First of all, the first and most important core concept shared by the official is:...
Field Notes
Full ledgerLive judgment for this desk is filed by date in the field-notes ledger; enter through the full ledger by month.
What really limits the Agent is the same model of Harness. If you change the set of Har...
The model benchmark is very high. Why does it still happen that "the code has been chan...
Production-level agents require an infrastructure that can run for a long time, have controlled access, and leave evidence.