The Wrong Number Problem

Deep dive: how to never act on an AI number you can't trust

Why this matters

AI hands you numbers that look finished. Clean tables, a confident tone, a tidy "you're up 12%." Looking finished and being correct are not the same thing, and a wrong number never announces itself. It just sits in a nice chart while you reorder 5,000 of the wrong SKU, kill a product that was actually winning, or raise a budget on a margin that was never real.

The fix is not a smarter model. It is a habit: before a number drives a decision, prove you can trust it. This is that habit, built into a system you can run on anything. An Amazon P&L, a PPC dashboard, a cash forecast, a Shopify export, a sponsor's audience figures. It does not assume what you are looking at. It reads what you give it and attacks it.

Part 1: The eight ways a number lies

This is the whole game. Every wrong number you will ever be handed is one of these eight. Learn to smell them and you are most of the way there.

  1. Ghost number. A figure with no source behind it. Nothing you can open says where 4,200 came from.
  2. Time mismatch. Two numbers compared or divided that cover different periods. Thirty days of ad spend laid over a 34-day sales window.
  3. Apples and oranges. Things combined that are measured differently. Gross revenue minus net costs, two currencies added, per-item next to a total.
  4. Missing-cost total. A profit, margin, or "net" that dropped a cost. The margin that forgot freight, or fees, or the coupon.
  5. Double count. The same money or units counted twice. A refund subtracted once, then its fee subtracted again.
  6. Flattering frame. A decline dressed as growth by picking a kind comparison. Up from last month, down 30 percent from last year, and only the first half gets shown.
  7. Guess in a suit. An estimate, forecast, or assumption written as if it were measured. "Lead time 30 days" stated like a fact when nobody timed it.
  8. Sources that disagree. Two reports that should match and don't. The tool says 900 units sold, the source of record says 1,050, and the summary just picked one.

Part 2: The system (run this before any money decision)

Run the full system when a real decision rides on the numbers: a reorder, a keep-or-kill, setting or raising a budget, a price change, or sending the numbers to a partner, lender, or buyer. For a quick gut-check, use the 90-second version at the end.

It is three stages on purpose. The most important move is Stage 2: you run the attack in a separate chat from the one that built the summary. A reviewer that did not write the work has no ego in defending it, which is exactly why it catches more. That is the "two independent reviewers" idea, made practical.

Stage 1: Map it

You are a skeptical numbers auditor. Below are numbers an AI produced or
summarized for me, and I may act on them. Do not improve or reword them.
First, map them.

List every distinct number, total, percentage, and claim. For each one, state:
- what it measures
- where it came from (which file, report, or input)
- what time period it covers
- what unit it is in (currency, count, percent, per-item or total)

If any of those four is missing for a number, mark it UNVERIFIED. If you cannot
tell what a number means or which numbers relate to each other, ask me before
going further. Do not guess to fill a gap.

Output a clean table, then a short list of every UNVERIFIED number and every
question you need answered.

Numbers below:

Stage 2: Attack it (run in a SEPARATE chat or a different model)

You are a hostile reviewer. You did not build the summary below and you assume
it contains at least one wrong number. Do not make it look better. Find the
problems.

Test every number against these eight failure modes:
1. Ghost number: no source behind it.
2. Time mismatch: two numbers compared or divided across different periods.
3. Apples and oranges: things combined that are measured differently (gross vs
   net, per-item vs total, two currencies, two definitions of one word).
4. Missing-cost total: a profit, margin, net, or total that dropped a cost.
5. Double count: the same money or units counted twice.
6. Flattering frame: a decline shown as growth by choosing a kind comparison.
7. Guess in a suit: an estimate, forecast, or assumption written as a measured
   fact.
8. Sources that disagree: two inputs that should match but don't.

For each problem: name the number, the failure mode, the evidence, and the
corrected or flagged figure. If a number is clean, say so. Do not fix anything
in place. List the findings.

Summary and numbers below:

Stage 3: Reconcile and rule

Here are the findings from two independent reviews of the same numbers.
Combine them.

1. Reconcile. Where the two reviews disagree, show both and say which is better
   supported. Where two source files disagree on the same figure, show both
   numbers with their date ranges. Never silently pick one.
2. Score each key number on five checks: Sourced, Dated, Unit clear, Complete
   (no missing cost), Reconciled (matches other sources). A number that fails
   any one of the five is not safe to act on.
3. Verdict. State my decision back to me, then GO or NO-GO, then the single
   most important thing to fix first if NO-GO.

Findings below:

The Decision Gate Card

Keep this. Fill one row per number that feeds your decision. All five Yes means trust it. A single No means fix it before you bet inventory or budget on it.

NumberSourced?Dated?Unit clear?Complete?Reconciled?Safe to act on?
Net profit / unit
ACOS or ad efficiency
Units to reorder
(add your own)

Part 3: Worked example

The AI summary said:

Net profit was 8,400 dollars last month on SKU-A. It is your winner. Scale the ad budget.

What the audit caught:

  1. Missing-cost total. The COGS used the factory price only. Add inbound freight at 0.62 per unit across 5,000 units and that is 3,100 dollars of real cost left out of "profit."
  2. Time mismatch. The sales figure covered 34 days. The ad spend covered 30. The ad efficiency number that justified "scale" was built on two different windows.
  3. Corrected. Real net profit is closer to 5,300 dollars, and the ad efficiency was understated, so the product is thinner than it looked.

Verdict: NO-GO on scaling the budget until freight is inside COGS and both figures use the same dates. One audit just stopped you from pouring money into a product you half-misread. (Illustration, not live data.)

Part 4: Make it permanent

Do not rebuild this every time.

  1. Fast. Paste all three prompts into a saved note, or into a ChatGPT Project or Claude Project, so you run them in two clicks instead of retyping.
  2. Pro. Save Stage 2 as a custom GPT or a saved skill that runs the eight-mode attack on anything you paste. Then the audit is one command, and "did I check this" stops being a question.

Part 5: The 90-second version (low stakes)

When the decision is small and you just want a fast sanity pass:

Before I trust these numbers: list every number with its source, date, and unit,
and mark anything missing one of those as UNVERIFIED. Then check for a missing
cost, a date mismatch between two figures, and any two sources that disagree.
Give me a one-line GO or NO-GO and the first thing to fix.

Numbers:

The operating rule

A number you cannot trace is a guess in a suit. Do not bet inventory or ad spend on a guess. Make AI show its sources, its dates, and its math, then decide.