I gave ChatGPT’s Data Agent a customer churn dataset, deliberately avoided telling it what to find, and let it investigate. It analyzed thousands of records, built a model, created visualizations, and produced an interactive report. Then I audited it.
The first result looked convincing.
That turned out to be part of the problem.
One of the most interesting errors wasn’t a fabricated statistic or an obviously broken chart. The number itself was real, but the analysis had attached it to the wrong business meaning.
As AI tools get better at moving from raw data to polished analysis, the question is becoming less:
Can AI analyze business data?
And more:
How do we know when AI-generated analysis is trustworthy enough to use?
I wanted to test that.
So I gave ChatGPT’s Data Agent an IBM Telco customer churn dataset containing 7,043 customer records and a deliberately open-ended instruction: investigate the data, identify patterns, quantify the findings, create useful visualizations, identify limitations, and distinguish correlation from causation—but don’t assume what the important patterns should be.
Then I put the resulting analysis through multiple rounds of verification before publishing it.
The result was more nuanced than either “AI nailed it” or “AI hallucinated.”
Most of the core analysis held up. Some of it didn’t.
And the failure taught me something important about what human oversight may look like as AI becomes increasingly capable of analytical work.
The experiment: give the AI the data, not the answer
I wanted to avoid designing an experiment where the prompt effectively told the AI what conclusions to discover.
So I started with the dataset and a deliberately open-ended instruction.

The prompt began:
“I’m evaluating how well you can independently investigate a business dataset.”
I asked it to understand the dataset, identify important patterns and customer segments associated with churn, quantify those findings, create useful visualizations, identify limitations, and distinguish correlation from causation.
One instruction was particularly important:
“Do not assume a conclusion in advance and do not begin with a predetermined list of churn drivers.”
In other words, I didn’t ask whether contract type was driving churn. I didn’t ask whether new customers were more likely to leave. And I didn’t give it a checklist of findings I expected it to reproduce.
I gave it specific analytical tasks, but I didn’t tell it what those tasks should uncover.
The Codex task identified the model configuration for this run as GPT-5.6 Sol Light. For clarity, I ran the experiment using the Data plugin inside Codex with an uploaded dataset. OpenAI also makes Data available in ChatGPT Work, where it can connect to company data sources, business context, and BI tools. This article evaluates the workflow I actually tested, not every way Data can be used.
Then I let it work.
It did much more than generate a few charts
This is where the experiment became more interesting than I expected.
The system didn’t simply return a paragraph summarizing the spreadsheet. It analyzed all 7,043 customer records, generated Python analysis code, created structured report data, built visualizations, and assembled the results into an interactive web report.
The report highlighted an overall churn rate of 26.5%, representing 1,869 churned customers.
It identified large differences across contract types. Month-to-month customers had an observed churn rate of 42.7%, compared with 2.8% among customers on two-year contracts.
It found a similarly large tenure gradient. Customers in months 0–5 showed 54.3% observed churn, compared with 9.6% among customers with 48 months of tenure or more.
It also built a churn-ranking model whose corrected five-fold out-of-fold ROC AUC was approximately 0.86.
If you don’t spend your days thinking about machine-learning metrics, here’s the practical interpretation: an AUC of 0.86 means the model gives a randomly selected churner a higher risk score than a randomly selected non-churner about 86% of the time.
That’s useful discrimination. It is not the same as correctly predicting whether any individual customer will churn.

At first glance, this looked like a successful experiment.
But I didn’t want to judge the system based on how professional the dashboard looked. I wanted to know whether the analysis underneath it survived scrutiny.
So instead of publishing it, I audited it.
The audit found something more interesting than a hallucinated number
The audit found that one of the report’s highlighted customer segments was mislabeled.
The original report associated this result with month-to-month fiber customers without tech support:
- 1,774 customers
- 1,031 churners
- 58.1% churn
The numbers were real.
But they didn’t describe the segment the report said they described.
Those figures actually belonged to month-to-month fiber customers where:
Online Security = No
The correct figures for the Tech Support = No segment were:
- 1,796 customers
- 1,033 churners
- 57.5% churn
The correction log traced the problem to the way compound segment labels moved through the analysis pipeline: the values survived, while the field names were discarded. That allowed two different attributes containing the value “No” to become ambiguous.
The system hadn’t hallucinated 58.1%—it had calculated a valid statistic and attached the wrong semantic identity to it.
That’s what made the mistake so interesting.
An obviously impossible number invites skepticism. A plausible number inside a polished dashboard can be much harder to notice.
The two findings weren’t independent, either
There was another useful qualification.
The Online Security and Tech Support segments overlapped substantially: 1,524 of the 1,774 customers in the Online Security = No segment were also in the Tech Support = No segment—about 85.9%.
So these shouldn’t be interpreted as two completely independent customer populations. They’re largely the same customers viewed through two different service attributes.
That doesn’t make either calculation wrong. It changes how confidently we should interpret them as separate business signals.

The audit went beyond the headline error
The segment label wasn’t the only issue.
The report originally said “after 48 months” even though its filter included month 48. That became the more precise “48 months or more.”
The audit also found a methodological issue with the model evaluation. Preprocessing had initially been fitted using the full dataset before cross-validation.
Why does that matter?
Because fitting preprocessing on the full dataset before cross-validation can allow information from the validation portion to influence how the training data is transformed. Even when that leakage is subtle, it weakens the separation that cross-validation is supposed to provide.
The correction moved encoding and scaling inside each training fold, so preprocessing for each evaluation was learned without using that fold’s validation data.
After the correction, the five-fold out-of-fold ROC AUC was 0.855331, still approximately 0.86.
So the methodology changed in a meaningful way without materially changing the headline model result.
Several interpretations were tightened too.
Predictors were no longer broadly described as “pre-outcome,” because the dataset didn’t contain a historical prediction timestamp proving when every field would have been available in a real deployment.
Feature importance was clarified as the reduction in fitted-training ROC AUC when individual features were shuffled—not proof of statistical, business, or causal importance.
And language suggesting that churn risk was “targetable” was replaced with the narrower observation that the top 20% of customers ranked retrospectively by the model contained 968 of the 1,869 observed churners, or 51.8%.
That describes concentration inside this dataset. It doesn’t prove future performance or that intervening on those customers would prevent churn.
The audit also surfaced less dramatic data-quality details. There were 11 blank or non-numeric Total Charges values, the dataset contained high-cardinality fields such as City, and some variables were closely related to others.
Those weren’t the headline of the experiment, but they’re exactly the kinds of details a real analytical review should notice.
The overall assessment wasn’t that the analysis had failed completely. It was:
VALID WITH MATERIAL QUALIFICATIONS
Most of the core descriptive findings survived.
The first report still wasn’t ready to publish.
I wanted a second, independent calculation
The first audit had an advantage: it could interrogate the analysis and implementation that produced the report.
But I didn’t want the original system to be the only check on its own numbers.
So I added a second verification step.
I opened a separate environment in Perplexity, supplied the original source dataset, and asked it to independently reconstruct the quantitative claims.
This was not a second code-level audit. The independent verifier didn’t have the original implementation files.
Instead, it served a different purpose: recalculate the reported numbers from the original source data and see whether the quantitative claims held up.

The independent verification reproduced the major numbers:
- 7,043 customers.
- 1,869 churners.
- 26.5% overall churn.
- 42.7% month-to-month churn.
- 2.8% two-year churn.
- 54.3% churn during months 0–5.
It also independently reproduced the two segment definitions:
- Online Security = No → 1,774 customers / 1,031 churners / 58.1%
- Tech Support = No → 1,796 customers / 1,033 churners / 57.5%
That meant the two checks were doing different jobs.
The first interrogated the implementation and methodology. The second independently challenged the reported numbers.
Together, they supported both sides of the story:
Most of the quantitative analysis held up.
And the semantic labeling error was real.
Because the second verifier didn’t have the original implementation files, I wouldn’t describe this as an independent code audit.
It was an independent data-level verification of the reported claims.
Before accepting the correction, I manually filtered the source spreadsheet and checked the disputed segment counts and churn rates myself.

Finding an error isn’t the same as fixing the system
At this point, the easiest response would have been to change 58.1% to 57.5% on the dashboard.
But the underlying problem wasn’t the percentage. It was the mechanism that allowed a valid statistic to acquire the wrong label.
So instead of feeding the original system the verifier’s entire answer and telling it to imitate the correction, I supplied the independently confirmed issues and required it to recalculate the affected results from the source data.
The compound-segment logic was changed so labels retained both their field names and values. The model preprocessing was moved inside each training fold. Interpretations that exceeded the evidence were narrowed. Source fields were separated from fields created during analysis: 33 original fields + 2 derived analysis fields.
Then the report was regenerated.
That produced the corrected distinction:
- Online Security = No → 58.1%
- Tech Support = No → 57.5%

And then the final QA still failed
After those corrections, I ran a publication-readiness check.
It still failed.
The calculations weren’t wrong again. This time, the problem was visible provenance.
The report showed that the model’s top 20% contained 51.8% of observed churners, but it didn’t expose the underlying counts: 968 of 1,869.
It also distinguished 33 source fields from 2 derived fields internally, but that provenance wasn’t visible to the reader.
So I blocked publication again, added those details, rebuilt the report, and reran the checks.
Only then did the publication check return:
PUBLICATION READINESS: PASS
This changed what I thought was most interesting about the experiment.
It wasn’t just that AI could build the report.
It was that the workflow could include explicit gates preventing a polished artifact from automatically becoming a published artifact.
The final report survived—but with narrower claims

The corrected report still contains several substantial findings.
Month-to-month customers represent 55.0% of the customer base but account for 88.6% of observed churn. Their churn rate is 42.7%, compared with 2.8% among two-year customers.
Customers in their first six months show 54.3% observed churn, compared with 9.6% among customers with 48 months of tenure or more.
Fiber customers show 41.9% observed churn, compared with 19.0% for DSL and 7.4% among customers without internet service.
Electronic-check customers show 45.3% observed churn, compared with 15.2% among automatic credit-card customers.
And the corrected ranking model retains approximately 0.86 five-fold out-of-fold ROC AUC.
Those are potentially useful signals.
They aren’t proof that any of those characteristics cause churn, and they don’t establish that acting on them would improve retention.
That’s a narrower conclusion than the first polished report invited.
It’s also a more defensible one.
What this experiment changed for me
I started by asking whether ChatGPT’s Data Agent could analyze a real business-style dataset.
It clearly could.
From one relatively neutral prompt, it moved from raw customer data to exploratory analysis, segmentation, modeling, charts, business interpretation, and an interactive web report.
Most of the central quantitative findings survived independent reconstruction.
That’s impressive.
But the first polished artifact still contained a subtle semantic error and several places where the methodology or language needed tightening.
That’s the part I think matters most.
Polish is not proof.
As AI-generated analytical work becomes more sophisticated, the human role may increasingly move away from manually producing every calculation and toward designing the system around the analysis:
- What question are we actually asking?
- What assumptions did the system make?
- Which claims matter enough to verify independently?
- What evidence would change our conclusion?
- What has to be true before this is allowed to influence a decision?
For organizations experimenting with AI analysis, my takeaway isn’t to avoid these tools.
It’s to build verification into the workflow before the analysis begins.
Define what “done” means. Require evidence for claims that could influence a decision. Separate descriptive findings from causal claims. Use independent verification where the stakes justify it. And treat the first polished output as a draft, not a deliverable.
The workflow I ended up with looked like this:
AI analysis → self-audit → independent verification → source-grounded correction → publication QA → human decision
Not every spreadsheet needs that level of scrutiny.
The amount of verification should match the consequences of being wrong.
But if an analysis could affect customers, money, strategy, or a published claim, a professional-looking dashboard shouldn’t be the thing that earns our trust.
Evidence should.
So, can ChatGPT’s Data Agent pass an audit?
For this experiment, a simple yes or no would miss the point.
The initial analysis did not pass unchanged.
Most of its central quantitative findings survived scrutiny, but verification found a real semantic error and several methodological or interpretive issues that needed correction.
After those corrections, another QA gate found additional provenance issues. Those were corrected too.
Then the report passed the publication-readiness check.
So the result I care about isn’t:
“The AI was right.”
It’s:
The workflow was capable of discovering where the AI was wrong before the artifact was published.
That seems like a much more useful standard.
Because the future of AI-assisted analytics probably isn’t going to depend on systems that never make mistakes.
It will depend on whether we build processes capable of finding consequential mistakes before we trust the output.
See the corrected report
I published the final interactive artifact so you can inspect the result yourself:
View the published Customer Churn Risk Report
This is the corrected report, not the initial version discussed earlier in the experiment.
Methodology and limitations
This was a practical experiment, not a benchmark or peer-reviewed evaluation of AI data-analysis systems.
I tested one dataset using one configuration: the Data capability in Codex with GPT-5.6 Sol Light. The result should not be generalized to every model, account, Data configuration or surface, dataset, or analytical task.
The source contained 7,043 customer records and 33 original fields. Two additional fields—Tenure band and Monthly charge band—were derived during analysis.
The dataset is cross-sectional and lacks a historical prediction timestamp, so this experiment cannot establish that every model feature would have been available before churn in a real prospective deployment.
The reported model performance wasn’t temporally or externally validated. The experiment didn’t test calibration, fairness, production performance, or whether acting on the model’s rankings would improve customer retention.
Observed associations should not be interpreted as causal effects.
The independent verification reconstructed quantitative claims from the source dataset but didn’t have the original implementation files available for code-level inspection.
And this experiment evaluates the workflow I actually ran. It doesn’t establish the maximum capability of ChatGPT’s Data Agent or GPT-5.6.


