OpenAI launched a data intelligence agent that supposedly transforms enterprise data into analytics and dashboards in a single sentence, but did not disclose its accuracy benchmarks.

OpenAI launched a data intelligence agent that supposedly transforms enterprise data into analytics and dashboards in a single sentence, but did not disclose its accuracy benchmarks.

OpenAI has taken another step into the enterprise data analytics market.

On September 10th local time, OpenAI announced the launch of a new “Data agent” in ChatGPT Work. Users no longer need to write SQL or manually manage multiple data sources. They can simply ask questions in natural language, and the agent can connect to the enterprise’s authorized data and documents, investigate what changes have occurred in the data, generate analytics and interactive dashboards, and perform follow-up actions after obtaining user approval.

However, the product launch also left a key question unanswered: OpenAI has not yet released publicly available accuracy or retrieval accuracy benchmarks for external customers. This means that while companies can see which data it can connect to and what tasks it can accomplish, they currently lack an independent, quantifiable, publicly available metric to determine just how accurate it is.

OpenAI turns data analysis into dialogue by "interrogating" 600PB of data with a single sentence.

OpenAI states that the goal of data agents is to enable employees to directly "ask questions about the data."

For example, a user can ask why a certain business metric has changed, requesting the system to identify influencing factors, compare different time periods or customer groups, and then generate a report needed by management. The AI will investigate relevant data itself, rather than simply returning a SQL query or a number.

It can also directly transform analysis results into interactive dashboards, including charts, metrics, and further editable visualizations, and supports team members in sharing, refreshing, and continuing to ask follow-up questions.

Prior to the official announcement of the new intelligent agent, OpenAI's internal data agent had already been operating and covering massive amounts of data. According to VentureBeat, OpenAI's internal data agent has served more than 3,500 users, covering more than 600 PB of data and approximately 70,000 datasets.

This also explains why OpenAI emphasizes that this product was not developed from scratch, but rather a further productization of the data analysis workflows already in use internally.

From Database to Dashboard: Connecting Enterprise Data, Documents, and BI Tools

Compared to traditional "natural language database lookup" tools, OpenAI places greater emphasis on the integration of data intelligence agents into an enterprise's existing toolchain.

Currently supported data platforms include Amazon Redshift, Google BigQuery, ClickHouse, Databricks, MongoDB, Snowflake, and Datadog; it can also include files and documents from Google Drive and SharePoint in the analysis.

After the analysis is completed, the data agent can also work in conjunction with dashboards and BI tools such as Omni, Oracle BI, Power BI, Sigma, Tableau, and ThoughtSpot.

OpenAI claims that its data agent can also leverage established business terminology, metric definitions, custom calculation methods, and data relationships within an enterprise. In other words, it attempts to understand not just "what's in this table," but how the enterprise internally defines business metrics such as revenue, customers, and orders.

In addition, administrators can predetermine which data connections and roles can be used, and data queries continue to follow the enterprise's original table-level, row-level, and column-level permissions.

Furthermore, data intelligence agents go beyond simply "providing answers." They can suggest who to contact next, which teams need to be involved, and share results or execute user-approved actions through connected tools, based on the analysis results.

First for personal use: Over 3,500 people within OpenAI already rely on it to retrieve data.

OpenAI states that data agents originate from the company's own data usage needs.

VentureBeat, citing Arpan Shah, OpenAI's corporate technology lead, reported that about a year ago, only a very small number of data issues within OpenAI could be resolved end-to-end without human intervention; now, company employees can handle large numbers of data issues on their own.

The internal version served over 3,500 users, covering approximately 70,000 datasets and over 600 PB of data. OpenAI also stated that almost all product team members and over two-thirds of the GTM (Marketing & Sales) teams are currently using data agents.

This is also a significant change in OpenAI's product positioning: it aims to transform tasks that previously required data analysts and engineers into tasks that ordinary employees can complete using natural language.

The biggest question: It can answer questions, but how accurate is it really?

However, for enterprise-level data products, "whether it can be done" is only the first hurdle; "how accurately it can be done" is the more critical issue.

OpenAI did not release a publicly available accuracy or retrieval accuracy benchmark for external customers in this announcement. VentureBeat stated that OpenAI did conduct internal comparative tests, including comparing the results of its data agents with the company's existing data tools, but OpenAI did not disclose specific results or provide a figure for external customers to compare against.

This does not mean that data intelligence agents have low accuracy, but rather that enterprise customers currently lack a public, unified quantitative standard to judge their performance in complex enterprise data environments.

This is particularly noteworthy because enterprise data is often far more complex than a typical Q&A: answers may be scattered across multiple data tables, BI dashboards, documents, and even Slack messages, and may involve different definitions of the same metric by different teams.

Meanwhile, competitors are actively using benchmarks to demonstrate their data intelligence capabilities.

For example, Databricks' recent Adaptive Instructed-Retriever study claims that its system achieves response quality comparable to Claude Sonnet 5, GPT-5.6 Luna, and DeepSeek-V4-Flash in specific enterprise retrieval tasks, with an average response time of approximately 5.8 seconds. It should be noted that these results are from Databricks' own testing and have not been independently verified.

Therefore, as companies begin to delegate more data queries to intelligent agents, who can prove "accuracy" with publicly available and reproducible metrics may be just as important as who can connect to more data sources.

OpenAI is pushing data intelligence into enterprises so that employees don't have to rely on analysts for everything.

From OpenAI's product design perspective, its ultimate goal is not simply to launch a "more powerful SQL assistant," but to make data intelligence agents a new entry point between enterprise employees and data.

Salespeople can directly analyze customer and sales data, marketing personnel can track marketing metrics, and management can ask agents to investigate business changes and generate dashboards, instead of waiting for data analysts to prepare reports every time.

This also makes the data intelligence agent a further step for ChatGPT Work into enterprise workflows: from answering questions, to investigating data, creating analysis, generating visualizations, and then executing follow-up actions.

On the other end, enterprise data infrastructure vendors such as Databricks and Snowflake are also accelerating their expansion into intelligent agents. The future competition in enterprise data intelligence may no longer be about "who has a better big model," but rather who can more accurately understand enterprise data, adhere to access control systems, and truly translate analytical results into business actions.

For OpenAI, this release addresses the issue of "making data accessible to everyone"; however, whether it can further prove "that everyone can confidently trust the answers" remains to be seen, and publishing the accuracy benchmark is still the next challenge it needs to face.

Risk Warning and DisclaimerInvesting involves risk; please exercise caution. This article does not constitute personal investment advice and does not take into account the specific investment objectives, financial situation, or needs of individual users. Users should consider whether any opinions, views, or conclusions in this article are suitable for their specific circumstances. Any investment decisions made based on this information are at your own risk.