OpenAI this week released a new model aimed at agentic coding, with long-horizon autonomous work as the headline feature. According to the official announcement, the model can work unsupervised for hours, closing the loop from requirement breakdown to code and test fixes.
Key points
- New record on SWE-bench Verified, with long-task success rates nearly doubling
- A new Agent API supports mid-task human intervention during tool calls
- Available to ChatGPT subscribers in Codex; API billed by usage
Our take
Coding has been the fastest-landing LLM use case of the past year, bar none. The keyword of this generation is “long horizon” — whoever can reliably turn an afternoon request into an evening delivery will truly enter developers’ workflows. For everyone else: the barrier to programming is vanishing, but the ability to describe requirements just became more valuable.
Compiled from public reports; opinions are our own.