← Back to Trends & Insights

Methods/Measurement/Model Quality

Grade Your Own Forecasts: Post-Open Accuracy as a Growth Advantage

Almost every multi-unit brand approves sites with a forecast. Almost none of them go back and check whether the forecast was right. That single missing discipline is why some brands get better at picking sites and others just accumulate leases.

Updated  ·  9 min read

Here is a question worth asking your real estate team this week: of the stores you opened two years ago, how many landed within 15% of the sales forecast the committee approved? If the answer takes more than a day to produce, or arrives as an anecdote rather than a number, you have found something more important than any individual site decision. You have a forecast that nobody grades, which means you have an opinion wearing a spreadsheet.

Forecast accuracy is the one part of site selection that almost nobody runs as a formal discipline. Deals get underwritten with impressive rigor and then, once the store opens, the forecast is quietly retired. The lease is signed, the P&L rolls into the district, and the number that justified a ten-year commitment is never compared to reality. This is how models stagnate for a decade and how a CFO ends up with no basis to trust the next request for capital.

In short

A forecast without a scorecard is an opinion. Measure absolute percentage error at maturity, not at month three; look at the direction and dispersion of error, not just the average, because a model that is right on average while wildly wrong per store is dangerous. Separate forecast error from execution error before you change anything. Then feed every opened store back in as a training analog, so each decision makes the next one sharper.

The Case

Why an Ungraded Forecast Is Just an Opinion

A forecast makes a falsifiable claim: this store will do roughly this much volume once it stabilizes. That claim is either right or wrong, and the answer arrives for free in your own point-of-sale data. Choosing not to look keeps the claim unfalsifiable, which is exactly what makes it an opinion rather than a model output.

The costs compound quietly. Your model never learns, because it never receives feedback about which assumptions were wrong. Your team’s calibration drifts, because the loudest lesson from any opening is whoever tells the best story about it. And when the CFO asks why the committee should believe the next forecast, the honest answer is that nobody checked the last forty. A scorecard turns that conversation into a review of evidence — and two chains running the same tools will diverge fast, because only one knows where its own judgment fails.

What to Measure

The Four Numbers That Belong on the Scorecard

1. Absolute percentage error at maturity

The headline metric is the absolute difference between forecast and actual, divided by the forecast. The word doing the work there is maturity. New stores ramp, and how fast they ramp has little to do with whether the forecast level was right. Comparing month-three sales to a stabilized forecast measures your ramp assumptions, not your site model — which is why brands that grade early conclude their model is pessimistic and then spend two years over-forecasting. Pick the maturity point your brand actually believes in, usually twelve to twenty-four months and consistent by format, and use trailing-twelve-month sales. Our guide to the new store ramp curve to maturity covers how to set that threshold defensibly.

2. Direction of error, not just size

Absolute error tells you how wrong you were; signed error tells you which way. Track both. A portfolio averaging 12% absolute error with a signed error near zero is noisy but unbiased. The same 12% with a signed error of +9% means you are systematically optimistic, and every deal approved near the threshold was approved on inflated numbers. Bias is more dangerous than noise, because it shows up in the rent you agreed to pay.

3. Dispersion, because averages hide the failures

This is the metric most scorecards omit and the one that matters most. A model that is right on average while wildly wrong store by store produces confidence without reliability. Look at the whole distribution:

4. Error against the decision, not the latest forecast

Grade the forecast the committee approved, on the date it approved it — not a number revised after the store opened. Re-forecasting is legitimate operational planning and worthless as accuracy measurement.

Attribution

Separating Forecast Error From Execution Error

A store that missed because the model was wrong about the trade area is a different problem from a store that missed because the GM turned over twice in eighteen months. Both show up as the same negative variance, and conflating them is how good models get corrupted: you teach the algorithm to distrust perfectly good sites because a few were operated badly. Before any miss changes the model, run it through an attribution review with operations in the room:

Only what survives that filter is forecast error, and only forecast error should drive model changes. The split itself is a finding: a meaningful share of underperformers turn out to be fine sites that needed a different operator.

A model that is right on average and wrong on every store is not a model. It is an average.
Diligence

The Honest Questions to Ask a Vendor

Everyone selling a forecast quotes an accuracy figure. Very few survive contact with these questions. Listen for specifics rather than adjectives.

Before you believe an accuracy claim
  • Measured on which stores, and how many? A number from 300 openings across formats means something; a number from nine flagship stores does not.
  • At what maturity was it measured? Early-month accuracy on a ramping store is a different and much easier claim.
  • Were those stores held out of training, or is this in-sample fit? A model can describe stores it has already seen almost perfectly and still fail on the next one.
  • What does the distribution look like? Ask for the share within 10% and 20%, and for the worst five misses.

The held-out question separates real validation from a demo: any model with enough parameters can fit stores it was trained on. Your own scorecard runs that test automatically every time you open a store, which is why an internal accuracy program is stronger evidence than any vendor case study. Our overview of how AI revenue forecasting with mobile data works explains what these models are actually inferring, and site selection software covers how the categories differ on transparency.

Blind Spots

What Error Patterns Reveal About Your Model

Residuals are the most valuable proprietary dataset a growing brand owns, and they usually sit unexamined. Sort misses by site characteristic and the noise resolves into structure. The patterns below are illustrative, but versions of them show up constantly:

That last point matters more as local intent grows. Semrush data for 2026 puts “near me” keyword variations at roughly 7.1 million US searches per month, up 29% between Q1 2025 and Q1 2026, with “near me tonight” (+41%) and “near me open now” (+38%) fastest-growing — per Semrush’s keyword volume research. If your model was calibrated on analogs from an era when discovery worked differently, some residuals may be measuring that shift rather than anything about the real estate.

Each named pattern is a fix: urban over-prediction becomes a capture-rate adjustment, under-predicted co-tenancy becomes a variable, low-awareness optimism becomes a market-entry ramp factor. Model improvement is not a rebuild; it is a sequence of specific corrections earned from specific misses.

The Loop

Building the Feedback Loop

The mechanics are unremarkable, which is why the discipline is rare — it fails on process, not sophistication. Four things have to be true.

That last step is where the compounding happens: fifty graded stores make a materially better model than fifty ungraded ones — same real estate, entirely different institutional knowledge. At Locate we treat the loop as part of the work rather than an afterthought, because the analysis that recommends a site and the post-open grade that tests it belong in the same system.

Bottom Line

Brands That Grade Themselves Get Better

The advantage here is not a better tool. It is a habit that turns every opening into information. Brands that grade their forecasts learn which site types they misjudge, which misses were really operating problems, and how much confidence a given forecast deserves — and they can prove it to a CFO. Brands that do not make the same mistake at store 40 that they made at store 4, and call the accumulating leases a growth strategy.

Start with one cohort. Pull every store that opened twenty-four months ago, find the approved forecast, compute the error, and sort the misses into model versus execution. It is usually an afternoon of work, and it will change what your next real estate committee argues about. For a second set of eyes, or a forecast designed from the start to be graded, talk to Locate.

FAQ

Common Questions

How accurate should a new store sales forecast be?
For a mature, well-calibrated model working with clean analog data, absolute percentage error in the 10–15% range at maturity is a reasonable expectation for most of your openings, with the distribution mattering as much as the average. What you should be suspicious of is any claim in the low single digits, or any accuracy number quoted without saying which stores it covers, at what point after opening it was measured, and whether those stores were held out of the model's training data.
When should you measure forecast error for a new store?
At maturity, not at month three. New stores ramp, and the ramp curve varies by format, market familiarity, and marketing spend, so an early-month comparison mostly measures how fast the store ramped rather than whether the forecast level was right. Measure at the point your brand considers a store stabilized — often 12 to 24 months — and compare trailing-twelve-month sales to the forecast the committee actually approved.
What is the difference between forecast error and execution error?
Forecast error means the model was wrong about the location: the trade area, the competitive draw, the co-tenancy, the capture rate. Execution error means the location was fine and the store underperformed for operating reasons — management turnover, staffing gaps, a delayed build-out, a marketing launch that never happened. They demand opposite responses. Forecast error should change the model; execution error should not, or you will teach the model to distrust perfectly good sites.
What questions should I ask a vendor about their forecast accuracy claims?
Four: On which stores was this measured, and how many? At what maturity? Were those stores held out of the training data, or is this in-sample fit? And what does the error distribution look like — not just the mean, but the tails? A vendor who can answer all four with specifics is telling you something real. A vendor who quotes a single accuracy percentage with no denominator is quoting marketing.
How do error patterns improve a site selection model?
Residuals are the most valuable dataset you own. When you sort misses by site type, you usually find structure rather than noise: systematic over-prediction of urban infill, under-prediction of strong co-tenancy, consistent optimism in markets where you have no brand awareness. Each pattern is a named blind spot that can be corrected with a variable, a weighting change, or better analogs — and each opened store then becomes a new training analog for the next decision.

The right location changes everything.

Get In Touch