Cancelled and test orders: small numbers that move a month

Both sit in every export and neither is trade. What each one does to revenue, order counts and average order value if it is left in.

Shopify1 Sep 20266 min read

Ibrahim Ölmez

Founder, nouz

Every raw order export contains rows that are not trade. Test orders placed while somebody was checking the checkout, and cancelled orders that never shipped and never earned. Both are trivially small in count and neither is harmless, because the numbers they distort are averages and ratios rather than totals: one cancelled bulk order can move a month's average order value visibly, and a handful of test rows can flatter a conversion rate. Filtering them is a one-line decision that has to be made once, deliberately, at the top of every report.

  • A test order is not trade and belongs nowhere in a statement.
  • A cancelled order never shipped and never earned, so it inflates counts and averages if left in.
  • The damage is concentrated in ratios: average order value, conversion, revenue per visitor.
  • Cancelled is not the same as refunded: one never happened, the other happened and was reversed.

Cancelled is not refunded

The distinction matters because the two are booked completely differently. A cancelled order never shipped, so there is nothing to reverse: it simply leaves the count of gross orders. A refunded order did ship, consumed a parcel and a fee, and is reversed on the day the refund was issued, which is why it stays in the order count and appears again in returns.

Blending them produces a statement that either double counts a cancellation or loses a refund's costs, and both errors are quiet enough to survive for months.

Where the distortion actually lands

Totals barely move; ratios move a lot. Leaving one cancelled €4.000 wholesale order in a month of 700 retail orders adds nothing you can spend and raises the gross AOV by roughly six euros, which is the sort of change somebody will try to explain with a merchandising theory.

Test orders do the same to conversion. A dozen test checkouts on a slow day is a visible improvement in a percentage that decides whether somebody keeps a landing page, and none of it happened.

Where these rows come from

Test orders arrive in bursts, usually around a checkout change, a new payment method or a theme release, which is exactly when somebody is also looking closely at the numbers. Cancellations cluster differently: stock that turned out not to exist, a customer who changed their mind before dispatch, a fraud check that failed, or a wholesale enquiry that became an invoice instead of an order.

Knowing the pattern is useful beyond filtering. A rising cancellation rate is a signal about stock accuracy or fraud pressure rather than about demand, and it is worth watching as its own small number rather than being quietly removed and forgotten.

Filter once, at the top

The place to do it is the first step of the bridge from gross sales down to what you kept, before any percentage is computed, which is exactly where a gross to net revenue calculator starts. Filtering later means every intermediate figure was computed on rows that should not have been there.

It is also one of the ordinary reasons why your Shopify numbers never match: two reports with different filters on the same month are not disagreeing about arithmetic, they are answering slightly different questions, and only one of them was asked deliberately.

The edge cases worth a decision

Partially cancelled orders are the awkward middle: one line removed, the rest shipped. The order happened, so it stays in the count and in revenue at what actually shipped, and the cancelled line simply never existed commercially. Treating the whole order as cancelled because part of it was is how a real shipment disappears from a month.

Fraud cancellations deserve their own note. An order cancelled because it was fraudulent never earned anything, but it may already have consumed picking time or even a parcel, and those costs are real even though the revenue is not. Counting the order out while leaving its costs in is the honest treatment, and it makes the cost of fraud visible instead of hiding it in a variance.

And draft or manually created orders behave like ordinary orders once they are paid, which is right, but they often carry unusual prices for wholesale or replacements. If you use them heavily, splitting them out is worth the effort before drawing any conclusion from an average.

The house rules worth adopting

  • Exclude test orders everywhere, without exception. There is no report that benefits from them.
  • Exclude cancelled orders from revenue and from order counts; they never happened commercially.
  • Keep refunded orders in the count and reverse them in returns on the day the refund was issued.
  • Mark test orders clearly at the moment you create them, so nobody has to identify them later by memory.
  • If a large order is cancelled, check the month's averages before drawing conclusions from them.

Written by

Ibrahim ÖlmezFounder, nouz

Builds the P&L engine behind nouz. Writes about the costs that decide whether a Shopify store is actually profitable.