OpenAI's product manager discusses the "third era" of AI: Agents will become collaborators during their growth phase, and corporate competition will shift to "who can steer the ship better."

OpenAI's product manager discusses the "third era" of AI: Agents will become collaborators during their growth phase, and corporate competition will shift to "who can steer the ship better."

If chat is the first era of AI products, and agents are the second era, then what might the third era be?

Tara Seshan, product lead for OpenAI Codex and ChatGPT Work, offers the answer: an AI colleague who can work long-term and collaborate with you repeatedly.

In an episode of Lenny's Podcast released on August 30, Seshan discussed how AI is transforming knowledge work from "rowing" to "steering": agents are taking on more and more execution, while humans are responsible for setting direction, judging results, and correcting course. Once execution capabilities become widespread, what truly differentiates individuals from businesses will no longer be "whether they can do it," but rather their judgment, taste, and ambition.

This isn't an officially released product roadmap from OpenAI, but rather a visionary assessment of future work methods from someone at the forefront of product development. More compelling questions arise: Why are programming agents beginning to enter knowledge work? Why can't companies just purchase a single chat application? And why, the more capable AI becomes, the more important managers become?

The difference between AI colleagues and chatbots

Seshan divides AI products into three stages: the first stage is chatting, the second stage is collaborating with agents, and the next stage is "AI colleagues who work continuously".

A typical chatbot interaction involves a question and an answer; an agent can continuously execute actions around a goal; an AI colleague is more like a team member: it works independently for a period of time, receives feedback from humans, and then continues to move forward, with both sides viewing and synchronizing their work at different paces. Seshan calls this "let the agent cook," where humans don't need to specify every step but rather provide direction at a higher level of abstraction.

This relationship needs to evolve from a "one person, one agent" model to multi-person collaboration. Currently, OpenAI employees sometimes exchange screenshots of Codex threads on Slack, explaining how a certain number was derived. However, screenshots don't represent a natural collaborative interface. The future challenge is: after my agent completes its analysis, can my colleagues' agents review it, and can we integrate people and multiple agents into the same task, rather than creating isolated, individual automation?

However, sustained operation depends on more intelligent models. Seshan used a vivid analogy: locking a new colleague in a room without giving them Google Docs, Slack, and the company database won't help, no matter how smart they are. Cloud-based agents also require enterprise data, third-party systems, cloud infrastructure, and reliability guarantees. A so-called "AI colleague" is first and foremost a working environment with appropriate context, permissions, and tools.

From rowing to steering, human work moves upwards.

"Rowing" means personally completing specific tasks, while "steering" means deciding where the boat goes and adjusting the direction based on feedback.

Seshan argues that the loops agents navigate will become increasingly longer, and the point of human control will shift from completing a single line of code to deliverables, goals, and even higher levels. However, taking the helm is not simply about writing down a goal and waiting for the outcome. Data can provide direction, but many crucial choices still rely on human judgment, intuition, and proactive visions of the future: not because another path is unfeasible, but because people want the product and the world to become a certain way.

She likened software to film, not real estate. In real estate, more capital investment often translates to more assets; however, a high film budget doesn't guarantee success. Software, too, requires expression, choices, and taste from its creators. When every company can utilize similar models and agents, the tools themselves become homogenized, and the individual's "proposal" becomes the differentiator.

This impacts businesses not just on reducing the amount of execution work in each role, but also on the need to change performance metrics and organizational division of labor. In the past, rewards were given to those who skillfully completed tasks; in the future, it will be necessary to identify who can ask better questions, set higher-quality goals, determine when agents deviate from their course, and hold them accountable for the final results.

With execution becoming cheaper, ambition becomes a new bottleneck.

Seshan observed that the most effective AI users are not just those who automate repetitive tasks, but those who expand their capabilities. In the past, people who were proficient in product, design, and engineering were rare "unicorns"; now, one person can quickly generate designs, build prototypes, and deduce pricing models and different scenarios, and tasks that were previously beyond their capabilities are now within their reach.

Therefore, the constraint has shifted from execution capability to "whether we can think bigger." She suggests that managers constantly ask themselves: Is there a more ambitious version? Can we try it faster? Can we scale it up 10 times? In her view, enhancing the ambition of team members is becoming an important task for product managers.

Several popular phrases within OpenAI aptly illustrate this culture: "Is this maximally accelerated?" and "Are you mainlining it yet?" The former questions speed, while the latter demands that product teams use their products intensively, shortening feedback loops through real-world work.

Even more radical is the product planning. Seshan stated that designing for today's model capabilities will fail, as will designing for what we imagine a year from now; OpenAI is trying to target model capabilities two or three months from now. Product teams need to stay close to the research roadmap while "making way" to avoid limiting new capabilities with product structures from the old model era.

This approach is suitable for cutting-edge model manufacturers, but cannot be directly copied by ordinary enterprises. The two-to-three-month window relies on internal information between the R&D and model teams; the notion that "ambition becomes a bottleneck" is also the judgment of the vendor's leader, rather than a universally validated rule. What enterprises can learn from this is to reduce heavy investment in temporary limitations that are about to be internalized by the model, and instead focus resources on their own processes, data, permissions, and business acceptance.

Why does ChatGPT Work "hide" Codex?

This interview also explains OpenAI's product strategy of pushing programming agents into knowledge work.

According to Seshan, ChatGPT's Work mode uses Codex's execution capabilities at its core, but removes programming-oriented interfaces such as the work tree. Users can let it build complex financial models and other knowledge-based work tasks; if the same task is proposed in Codex, its capabilities are not weaker, the main difference lies in what interface and technical details are displayed to the user during execution.

OpenAI 's ideal state is where users no longer understand the differences between Chat, Work, Codex, Model, and Harness. Humans simply describe the task, and the system automatically selects the appropriate model and execution framework. The current multiple entry points represent a transitional form as the product migrates from chat to agent-based.

The real challenge lies in the fact that knowledge work cannot simply adopt the acceptance methods of programming agents. Whether code passes tests can often be judged from its output; even if a strategic report or financial analysis appears complete, it cannot be deemed "90% correct" solely based on the final document. Users also need to examine the process, inputs, references, and intermediate work to understand how the conclusions were reached.

Therefore, knowledge work agents cannot simply deliver answers; they must also involve people in the journey of forming those answers: seeing which data was used, whether citations support the conclusions, how key assumptions changed, and where human judgment was needed. For B2B products, this means that traceability, process collaboration, and business acceptance are not additional features, but rather the thresholds from the Coding Agent to the core enterprise workflow.

The company needs to provide more than just an account.

For an AI colleague to work continuously and be accepted into a company, at least four conditions must be met.

First, there's the context. Can it understand documents, emails, meetings, and business data within its authorized scope, instead of requiring employees to re-expose the context each time? Second, there's the operational capability: which systems to call, under whose identity, and how to roll back after a write failure? Third, there's the collaboration mechanism: when do people check, when does the agent pause, and how do other employees and agents take over? Finally, there's accountability: who verifies the output quality, and who approves high-risk decisions?

This will also change the competitive landscape for B2B AI providers. While models will provide general capabilities, the truly hard-to-replicate value will lie more in connectors, identity and permissions, enterprise context, long-term task reliability, process evidence, and team collaboration. An agent that can "demonstrate everything" is not necessarily more valuable than an agent that can reliably enter specific processes, accept audits, and be subject to human intervention.

When selecting AI assistants, companies shouldn't just test how impressive the results are on a single test. A more effective approach is to select real-world historical tasks, run the product under realistic permissions, data conflicts, and anomaly conditions, and record end-to-end completion rates, human intervention, traceability, failure recovery, and the cost of successful tasks. The value of AI assistants lies not in never needing humans, but in their ability to handle a wider range of tasks with fewer human interventions.

What cannot be outsourced is "writing for thinking."

When the agent took over execution, Seshan still reserved a clearly defined artificial area for himself: writing for thinking.

She divides writing into two categories. Weekly reports, status updates, and format conversions belong to "writing as a report," which can be delegated to the model as much as possible; product direction, strategic judgments, and controversial viewpoints belong to "writing as a reflection," which she insists on doing herself, because outlining, writing, revising, and accepting criticism from others are themselves processes of clarifying ideas.

She summarizes her approach as: "Starting with myself, and ending with myself." AI can assist in research, supplement data, or challenge viewpoints along the way, but it won't generate the initial conclusions for her. Meanwhile, OpenAI's internal sharing of results is shifting from "documents" to "prototypes and results"—mocks, not docs; prototypes, not docs. Long documents can now be easily generated and no longer inherently represent depth of thought; interactive prototypes and experimental results are more likely to drive decision-making.

This precisely represents the two sides of work in the AI era: while execution is delegated to agents, judgment should not be diminished; while reporting can be automated, the process of forming opinions must still be retained. Otherwise, companies may acquire more documents but lose the people who can truly steer the ship.

The third era truly tests organizations.

"Continuously working AI colleagues" are still the direction of evolution described by Seshan, rather than a unified product that has been maturely delivered. ChatGPT Work, Codex, and multi-agent collaboration are still in the process of product integration and interaction exploration.

But the direction is clear enough: AI products are moving from answering questions to executing tasks, and then to long-term participation in work. The model's capabilities determine how fast it can go, the corporate context and infrastructure determine whether it can be put on board, and human judgment, responsibility, and ambition determine where the ship will sail.

The key to the third era of AI is not how many agents a company has, but whether it can establish a new way of collaboration: enabling AI to work continuously while allowing humans to truly take the helm.

This article comes from the WeChat official account " AI Native Lab " , which continuously analyzes real-world AI implementation cases and shares enterprise AI practices and methodologies.

Risk Warning and DisclaimerInvesting involves risk; please exercise caution. This article does not constitute personal investment advice and does not take into account the specific investment objectives, financial situation, or needs of individual users. Users should consider whether any opinions, views, or conclusions in this article are suitable for their specific circumstances. Any investment decisions made based on this information are at your own risk.