This case study is for anyone building, evaluating, or relying on AI systems that turn business data into explanations, dashboards, or recommendations.
I built InsightForge as an AI-assisted business intelligence prototype that combines deterministic analytics with generative interpretation.
The core idea was simple:
Calculate business facts first, then ask the language model to interpret them.
Pandas handles deterministic calculations. Structured facts are then passed to Azure OpenAI through LangChain for natural-language interpretation, while Streamlit provides the user-facing interface.
The application accepts CSV or Excel data, detects relevant fields, generates dashboard metrics and visualizations, exposes supporting data, and lets the user ask questions about the dataset.
The architecture produced a functioning prototype and reduced the risk that the language model would invent the underlying arithmetic.
But once I had it working, a more important question emerged:
Was grounding the model in calculated facts enough to make the resulting analysis trustworthy?
The answer turned out to be more complicated.
The project eventually became less about preventing an LLM from inventing numbers and more about preserving what those numbers actually mean as they move from source data to metrics, interpretation, and action.
InsightForge v1 in Action
InsightForge v1 in action. The historical prototype loads demonstration sales data, detects its schema, generates dashboard metrics and visualizations, answers natural-language questions from calculated facts, exposes supporting data, and enables processed-data export.
Validation note: The public sanitized release independently reproduced the deterministic workflow. Re-validating a live Azure OpenAI provider request was outside the release-validation scope.
The demonstration dataset contains 2,500 dated sales observations with product, region, sales, age, gender, and satisfaction attributes.
The historical demo shows InsightForge:
- loading and detecting the structure of uploaded data,
- calculating dashboard metrics,
- generating revenue and time-series visualizations,
- answering natural-language analytical questions,
- exposing the structured facts used in Analysis,
- and allowing processed-data export.
The runtime behavior was real.
The audit that followed asked a different question:
What did those outputs actually establish?
Before getting into what the audit revealed, it helps to make the implemented architecture explicit—what InsightForge calculated deterministically, what it passed to the language model, and where the boundary between those responsibilities sat.
What I Built
InsightForge v1 is a Streamlit-based analytical application that processes uploaded CSV or Excel data with Pandas, detects relevant schema fields, performs deterministic calculations, creates structured facts, and uses Azure OpenAI through LangChain to translate those facts into natural-language interpretations.
The prototype detects fields associated with concepts such as date, revenue or sales, quantity, price, order ID, product, region, customer identifier, age, gender, segment, and satisfaction.
The detected schema then drives dashboard metrics, visualizations, and the structured fact object used by the generative layer.
The interface is organized into four main areas:
- Overview
- Visualizations
- Analysis
- Data
The Overview presents key metrics and an AI-generated summary.
The Visualizations tab generates revenue and time-series charts.
The Analysis tab lets the user ask a natural-language question.
The Data tab exposes processed records and provides a CSV download.
ARCHITECTURE A · WHAT INSIGHTFORGE V1 ACTUALLY DID

The Analysis path is straightforward:
facts = format_facts(df, schema)
response = explain(query, facts)
with st.expander("📊 Supporting Data"):
st.json(facts)
That small excerpt captures an important part of the v1 architecture:
1. deterministic code creates the structured facts,
2. those facts are passed into the interpretation layer,
3. and the same facts can be inspected by the user.
There is, however, an important boundary:
In InsightForge v1, the user's question does not determine which new Pandas calculation is executed.
Each Analyze action rebuilds the same predefined format_facts() structure.
The question changes what the model is asked to discuss.
It does not dynamically select a new analytical calculation.
The prototype also generates a sequential internal __order_id__ when no source order identifier is detected. That behavior is useful for record handling, but its business meaning becomes important later in the audit.
The architecture nonetheless established a valuable separation:
Deterministic code performs the calculations, while the language model interprets the resulting facts.
That separation solved a real problem: it reduced the amount of numerical reasoning assigned to the language model.
What Grounding Solved
The original InsightForge design followed a useful principle:
Calculate before you generate.
Where deterministic code can calculate something reliably, deterministic code should retain responsibility for that calculation.
For example, regional revenue can be calculated directly with Pandas:
revenue_by_region = (
df.groupby(region_col)[revenue_col]
.sum()
.sort_values(ascending=False)
)
That calculation happens before the model is asked to explain anything.
format_facts() builds structured values such as:
- total and average revenue,
- revenue by region,
- revenue by product,
- top region,
- top product,
- order counts when an order field exists,
- customer counts when a validated customer identifier exists,
- gender-based observation counts,
- age statistics,
- quantity totals,
- satisfaction statistics,
- satisfaction by gender,
- satisfaction by region,
- satisfaction by product,
- and date ranges.
The model then receives the user's question together with those structured calculated facts and generates a natural-language interpretation.
EVIDENCE 01 · HISTORICAL V1

A factual question such as:
“what region has the highest revenue?”
produced:
West — 361,383.0
That result aligned with the deterministic regional aggregation.
EVIDENCE 02 · SANITIZED RELEASE

That design gave the prototype four useful properties.
Deterministic calculation authority.
The application did not rely on the model to derive supported revenue, age, or satisfaction aggregates.
Structured evidence.
The Analysis path passed calculated facts rather than automatically transmitting the entire dataframe.
Numeric grounding.
The model was instructed to use the supplied numerical evidence rather than invent values.
Inspectability.
The user could inspect the same fact structure used in the Analysis flow.
Those are meaningful strengths.
But they solve a narrower problem than I initially assumed.
Grounding can constrain the numeric evidence a model receives. It cannot, by itself, prove that a metric has the business meaning assigned to it, that it was calculated at the scope the question requires, or that an interpretation supports the recommendation that follows.
That became clear when I audited the system from the source data forward.
What Grounding Alone Didn’t Solve
1. Data Grain and Entity Meaning
Problem
The demonstration dataset contains seven source fields:
- Date
- Product
- Region
- Sales
- Customer_Age
- Customer_Gender
- Customer_Satisfaction
It does not contain:
- a persistent customer identifier,
- or a source order identifier.
The safest description of the row grain is therefore:
Each row is a dated sales observation with associated product, region, demographic, and satisfaction attributes.
A row contains a gender value.
That does not establish a unique customer.
A row contains a sales value.
That does not establish a distinct business order.
Evidence
EVIDENCE 03 · SANITIZED RELEASE

When no source order identifier is found, the prototype creates one:
if ORDER_COL is None:
# This sequential internal identifier does not establish that each
# source row represents a verified distinct order.
df["__order_id__"] = np.arange(1, len(df) + 1, dtype=int)
ORDER_COL = "__order_id__"
That identifier is useful as technical record identity.
UNIQUE RECORD ID ≠ VERIFIED BUSINESS ORDER ID
The historical Overview can count those generated IDs and display:
Total Orders: 2,500
The arithmetic is internally consistent.
The source data, however, does not establish that the 2,500 observations represent 2,500 verified business orders.
Why it matters
This exposed a broader lesson:
The same principle applies to gender-based counts.
Without a persistent customer identifier, counts of observations carrying Male or Female values do not establish the number of distinct male or female customers.
The problem is not incorrect arithmetic.
It is semantic overreach.
2. Metric Semantics
Related calculations can look interchangeable even when they represent different business concepts.
Consider:
- total_records
- customer_count_by_gender
- total_customers
Those are not the same metric.
total_records can be derived from row count.
A grouped gender count can describe how many observations carry each gender attribute.
total_customers, however, implies a count of customer entities.
Without a customer identifier, the source data does not establish a unique-customer count.
This problem is not solely generated by the language model.
The deterministic fallback logic itself can attach customer-oriented language to grouped observations.
That means:
Why it matters
A formula is not a complete metric definition.
A defensible metric should carry a contract such as:
- Name
- Business definition
- Entity / population
- Source fields
- Required grain
- Aggregation rule
- Allowed dimensions / filters
- Time period
- Known limitations
The key question is not only:
“Can I calculate this?”
It is also:
“Can I defend what I call it?”
That distinction becomes even more important once a language model can turn a convenient metric label into a confident business explanation.
3. Evidence Scope and Intersections
Problem
The historical demo included the question:
“give me recommendations to boost male satisfaction score”
That historical wording should be preserved because it is the actual prompt.
But analytically, the question concerns satisfaction among observations carrying the Male value.
Those individual aggregates were grounded in deterministic calculations.
The underlying dataset contained dimensions that could support those calculations.
The v1 Analysis execution path simply did not calculate them dynamically for the question.
Evidence
EVIDENCE 04 · HISTORICAL V1

The supplied evidence effectively contained:
overall satisfaction by product
plus
overall satisfaction by gender
plus
overall satisfaction by region
That does not automatically establish:
satisfaction among male observations by product
or
satisfaction among male observations by region
Why it matters
The recommendation may still be a useful hypothesis.
But the supplied evidence did not establish the narrower relationship implied by parts of the recommendation.
That led to one of the strongest lessons in the audit:
If a recommendation depends on a narrower population than the evidence supplied to the model, the analytical layer should calculate evidence at that narrower scope first.
Even then, calculating additional intersections is not enough by itself. The system should also surface relevant subgroup size, missing-data conditions, period coverage, and analytical warnings before treating a subgroup pattern as decision-ready.
That produces a stronger design principle:
The question should determine which validated deterministic calculation is executed—not merely which pre-calculated facts the model discusses.
4. Recommendations and Evidentiary Distance
Once I saw the intersection issue, I started separating analytical claims by how far they travel beyond directly observed evidence.
The sequence I now use is:
For example:
Product A has lower satisfaction for a defined subgroup.
This subgroup appears less satisfied with Product A than with other products.
A Product A experience may be contributing to lower satisfaction.
Investigate the Product A experience and test a targeted intervention.
Each stage introduces a stronger claim.
A deterministic calculation may establish the first.
The second requires interpretation.
The third introduces a possible explanation.
The fourth proposes an action.
This does not mean recommendations must be proven before they can be useful.
A recommendation can still be a reasonable next step.
But the system should make visible where evidence ends and where inference begins.
The core principle is:
That is why I would distinguish future outputs such as:
The point is not to eliminate judgment.
It is to make evidentiary distance visible.
5. Visualization Context
The audit also surfaced an issue that did not depend on the language model at all.
InsightForge creates its sales trend by aggregating revenue into monthly periods.
The demonstration dataset runs from January 1, 2022 through November 4, 2028.
October 2028 is a completed month.
November 2028 contains observations only through November 4.
The monthly aggregation is correct.
The comparison context is not equivalent.
EVIDENCE 05 · HISTORICAL V1

The final point can therefore suggest a sharp deterioration if the viewer interprets it as directly comparable with completed monthly periods.
Why it matters
My practical rule would be:
Do not present a partial reporting period as directly comparable to completed periods without making the difference visible.
Possible treatments include:
- marking the period as partial,
- excluding it from completed-period comparisons,
- comparing month-to-date values with equivalent elapsed periods,
- or exposing a visible completeness indicator and data-through date.
The sample data also contains dates later than the 2025 historical recording.
InsightForge is not forecasting those values.
The chart simply plots the supplied demonstration records.
That distinction matters because analytical presentation needs to preserve context just as carefully as analytical calculation.
The Problems Were Connected
Individually, each finding looked like a different analytical problem.
Together, they pointed to the same deeper issue:
Meaning could weaken at multiple points as evidence moved through the system.
The source data determined what evidence existed.
The data grain determined what entities the system was entitled to count.
Metric semantics determined what those calculations were allowed to mean.
Evidence scope determined whether the calculation actually matched the question being asked.
AI interpretation introduced claims, hypotheses, and recommendations that could travel beyond the directly calculated evidence.
And presentation context could change how even a correct result was understood.
Grounding the language model sat inside that sequence.
It did not replace the rest of it.
Most of these responsibilities existed before an LLM entered the system. Data grain, metric semantics, evidence scope, and presentation context matter in conventional analytics too.
AI adds another interpretation layer.
It does not remove the analytical responsibilities underneath it.
Once I saw those responsibilities as a connected sequence rather than isolated checks, I needed a practical way to reason about whether analytical meaning was surviving from one stage to the next.
That became the AI Analytics Trust Chain.
FROM AUDIT FINDINGS TO A REUSABLE FRAMEWORK
The AI Analytics Trust Chain

I use the AI Analytics Trust Chain as a practical reasoning framework for examining how analytical meaning is preserved as data becomes metrics, evidence, interpretation, and eventually action.
Here, trust does not mean certainty.
It means being able to inspect and defend how evidence became a metric, interpretation, or proposed action.
It is not a mathematical proof of trust.
It is not a numerical scoring system.
It does not guarantee that an analytical output is correct.
And I am not presenting it as an industry standard.
Its central principle is:
A weakness at any stage can weaken every stage after it.
Grounding controls what evidence the model receives.
The Trust Chain asks whether that evidence remained analytically defensible before and after the model boundary.
The practical test I use is:
Can I defend how meaning was preserved from the available evidence to the claim or action being presented?
The chain has six stages.
1. Source Data
Ask: What evidence is actually available?
Check:
- provenance and upstream transformations,
- available fields and time coverage,
- missing values, identifiers, and source limitations.
Carry forward: provenance, fields, coverage, identifier availability, missing-data conditions, authorization context, and known limitations.
The goal is not to prove that the source data is perfect.
It is to prevent downstream analysis from assuming evidence the source does not contain.
2. Data Grain
Ask: What does one record represent?
Check:
- the real-world event or entity represented by a row,
- whether the same entity can appear multiple times,
- which identifiers come from the source and which were generated internally.
Carry forward: declared unit of observation, source identifiers, uniqueness assumptions, and unresolved grain limitations.
Grain determines what entities the analytical system is entitled to count.
3. Metric Semantics
Ask: What does this calculation mean?
Check:
- business definition,
- population or entity,
- required grain,
- aggregation rule,
- dimensions, filters, and time window.
Carry forward: metric definition, population, source fields, aggregation rule, scope, time period, completeness status, and limitations.
A formula is not a complete metric definition.
And:
The next stage should receive more than a number and a convenient label.
4. Evidence Scope
Ask: Did I calculate evidence at the scope this question requires?
Check:
- population,
- metrics and dimensions,
- required intersections,
- filters and period,
- subgroup size, missingness, and relevant warnings.
Carry forward: question-specific calculation, population, filters, dimensions, period, supporting values, subgroup size, and analytical warnings.
The question should determine which validated deterministic calculation is executed—not merely which pre-calculated facts the model discusses.
5. AI Interpretation
Ask: What is the model claiming?
Check:
- whether the statement is fact, inference, hypothesis, or recommendation,
- whether material factual claims trace back to supplied evidence,
- whether the model changed the population, period, or scope,
- whether limitations and uncertainty remain visible.
Carry forward: claim type, claim-to-evidence mapping, analytical scope, assumptions, limitations, and hypotheses introduced during interpretation.
The purpose is not merely to produce fluent language. It is to preserve the relationship between the language and the evidence supporting it.
6. Action
Ask: How far does the action extend beyond the evidence?
Check:
- which evidence directly supports the action,
- which parts of the reasoning remain hypotheses,
- whether additional validation is needed,
- consequences, reversibility, and appropriate human review.
Carry forward: supporting evidence, assumptions, limitations, validation needs, consequence context, and the proposed action.
The final stage is Action, not “AI Decision.”
The model may help communicate or propose an action.
It should not be treated as the authority that determines whether its own recommendation is sufficiently supported.
The purpose of the chain is not to automate trust. It is to make the basis for trust more inspectable.
What Should Survive the Entire Chain
A downstream claim should still be traceable backward:
| STAGE | QUESTION |
|---|---|
| SOURCE | Where did the evidence come from? |
| GRAIN | What did the records represent? |
| METRIC | What exactly was calculated? |
| SCOPE | For which population, dimensions, filters, and period? |
| INTERPRETATION | What claim was made, and which evidence supports it? |
| ACTION | What is being proposed, and what remains uncertain? |
Presentation context remains cross-cutting rather than becoming a seventh stage.
A partial reporting period, missing-data warning, subgroup size, source limitation, or other important condition may affect how evidence should be interpreted throughout the chain.
The responsibility is not simply to calculate correctly.
It is to carry the conditions needed to interpret the calculation correctly as the result moves downstream.
How I Would Evolve InsightForge
The v1 architecture established a useful boundary:
deterministic code calculated a predefined set of facts, and the language model interpreted them.
I would preserve that separation.
But I would make the analytical system between the source data and the model much more explicit.
At a high level, the next design would move toward:
The biggest change would be the role of the user's question.
Question + Predefined Fact Set → AI Interpretation
The language model can help interpret intent and communicate results.
DESIGN BOUNDARY
The analytical system retains authority over what calculation is executed, what the result means, and what evidence accompanies it.
That points toward a richer evidence structure.
I think of that proposed mechanism as an Evidence Package.
Before interpretation, it could preserve information such as:
BEFORE INTERPRETATION
- RESULT
- calculated result
- METRIC MEANING
- metric definition
- POPULATION
- population or entity
- SCOPE
- dimensions and filters
- TIME CONTEXT
- time period and period completeness
- PROVENANCE + LINEAGE
- source provenance and calculation lineage
- ACCESS CONTEXT
- access context and authorization boundaries
- SUPPORTING EVIDENCE
- supporting calculations
- LIMITATIONS
- limitations or warnings
Provenance answers where the evidence originated.
Calculation lineage records the transformations, filters, and calculations that produced the result.
After AI interpretation, the package could be extended with:
AFTER INTERPRETATION
- CLAIM TYPE
- fact, inference, hypothesis, or recommendation
- CLAIM-TO-EVIDENCE MAPPING
- claim-to-evidence mapping
- ASSUMPTIONS
- assumptions introduced during interpretation
- HYPOTHESES REQUIRING VALIDATION
- hypotheses requiring validation
- REMAINING LIMITATIONS / UNCERTAINTY
- remaining limitations or uncertainty
That claim-to-evidence mapping matters because grounding should remain inspectable after generation, not only before it.
A reader should be able to determine which calculation, population, period, and warning support each material statement.

The proposed architecture introduces five broad layers:
1. Access & Governed Data
2. Data Trust
3. Controlled Analytics
4. Evidence & Generative Interpretation
5. Evidence-Aware Presentation
with cross-cutting concerns such as security, managed secrets, auditability, observability, evaluation, versioning, and deployment controls.
InsightForge v2 has not been built.
The follow-up article presents the architecture I would design based on what I learned from v1.
The next question is no longer:
How do I add more AI to the prototype?
It is:
How do I design the analytical system so that every downstream claim remains traceable to evidence with an explicit meaning, scope, and limitation?
That is the direction I would take InsightForge next.
Calculate before you generate—but validate what the calculation means before you trust what the generation says about it.
What I Took Away From the Project
InsightForge began with a practical question:
How can I combine deterministic business analytics with a language model without assigning the arithmetic to the model?
The first answer was useful:
Calculate before you generate.
The audit exposed the more important question.
A language model can receive correct numbers and still inherit:
- unclear data grain,
- weak metric semantics,
- incomplete evidence scope,
- missing comparison context,
- or a recommendation that travels farther than the evidence supports.
For me, that is the more useful definition of grounded AI analytics.
NOT SIMPLY:
“Did the model use supplied facts?”
THE STRONGER TEST:
“Can I defend how meaning was preserved from the available evidence to the claim or action being presented?”
That is the standard I would use to evaluate the next version of InsightForge.
And it is the standard I would increasingly bring to any AI system expected to turn business data into explanations, recommendations, or decisions.
Build and Inspect InsightForge for Yourself!
AI-Assisted Business Intelligence & Analytics
The sanitized public package includes the Streamlit application, walkthrough notebook, demonstration data, dependencies, verification script, and implementation notes.
It preserves the historical v1 architecture—including its limitations—so readers can inspect the implementation behind this case study.
The package contains:
README.mdLICENSEinsightforge_demo.pyinsightforge_walkthrough.ipynbsample_sales_data.csvrequirements.txtverify_release.py
The demonstration dataset contains the original source fields:
- Date
- Product
- Region
- Sales
- Customer_Age
- Customer_Gender
- Customer_Satisfaction
The sanitized application also preserves the v1 behavior of generating an internal __order_id__ when no source order identifier is present.
That behavior is retained intentionally rather than silently implementing the semantic changes proposed for a future design.
InsightForge's deterministic analytics remain usable without Azure OpenAI. Azure configuration is required to reproduce the prototype's generative summaries and natural-language interpretations.
Validation boundaries
Validated
- deterministic CSV workflow,
- deterministic XLSX workflow,
- clean-environment release verification.
Not revalidated in the final release QA
- a new live Azure OpenAI provider invocation,
- full notebook end-to-end execution.
Not included
- the proposed future architecture described above.
Have fun exploring it, and let me know if you find it useful.
If you’re building AI systems that turn business data into explanations or recommendations, I’d be interested to hear what evidence, governance, or interpretation challenges you’re working through.
Sources & updates
Reviewed September 30, 2026. Deterministic CSV/XLSX release workflows were revalidated; a new live Azure OpenAI invocation and full notebook run were outside final release QA.



