Here is a question worth asking your real estate team this week: of the stores you opened two years ago, how many landed within 15% of the sales forecast the committee approved? If the answer takes more than a day to produce, or arrives as an anecdote rather than a number, you have found something more important than any individual site decision. You have a forecast that nobody grades, which means you have an opinion wearing a spreadsheet.
Forecast accuracy is the one part of site selection that almost nobody runs as a formal discipline. Deals get underwritten with impressive rigor and then, once the store opens, the forecast is quietly retired. The lease is signed, the P&L rolls into the district, and the number that justified a ten-year commitment is never compared to reality. This is how models stagnate for a decade and how a CFO ends up with no basis to trust the next request for capital.
A forecast without a scorecard is an opinion. Measure absolute percentage error at maturity, not at month three; look at the direction and dispersion of error, not just the average, because a model that is right on average while wildly wrong per store is dangerous. Separate forecast error from execution error before you change anything. Then feed every opened store back in as a training analog, so each decision makes the next one sharper.
Why an Ungraded Forecast Is Just an Opinion
A forecast makes a falsifiable claim: this store will do roughly this much volume once it stabilizes. That claim is either right or wrong, and the answer arrives for free in your own point-of-sale data. Choosing not to look keeps the claim unfalsifiable, which is exactly what makes it an opinion rather than a model output.
The costs compound quietly. Your model never learns, because it never receives feedback about which assumptions were wrong. Your team’s calibration drifts, because the loudest lesson from any opening is whoever tells the best story about it. And when the CFO asks why the committee should believe the next forecast, the honest answer is that nobody checked the last forty. A scorecard turns that conversation into a review of evidence — and two chains running the same tools will diverge fast, because only one knows where its own judgment fails.
The Four Numbers That Belong on the Scorecard
1. Absolute percentage error at maturity
The headline metric is the absolute difference between forecast and actual, divided by the forecast. The word doing the work there is maturity. New stores ramp, and how fast they ramp has little to do with whether the forecast level was right. Comparing month-three sales to a stabilized forecast measures your ramp assumptions, not your site model — which is why brands that grade early conclude their model is pessimistic and then spend two years over-forecasting. Pick the maturity point your brand actually believes in, usually twelve to twenty-four months and consistent by format, and use trailing-twelve-month sales. Our guide to the new store ramp curve to maturity covers how to set that threshold defensibly.
2. Direction of error, not just size
Absolute error tells you how wrong you were; signed error tells you which way. Track both. A portfolio averaging 12% absolute error with a signed error near zero is noisy but unbiased. The same 12% with a signed error of +9% means you are systematically optimistic, and every deal approved near the threshold was approved on inflated numbers. Bias is more dangerous than noise, because it shows up in the rent you agreed to pay.
3. Dispersion, because averages hide the failures
This is the metric most scorecards omit and the one that matters most. A model that is right on average while wildly wrong store by store produces confidence without reliability. Look at the whole distribution:
- What share of stores landed within 10% of forecast? Within 20%?
- How many missed by more than 30% in either direction, and what do they have in common?
- What is the median error, and how far is it from the mean? A big gap says a few catastrophic misses are hiding inside a comfortable average.
- Are the tails symmetric, or is every large miss on the same side?
4. Error against the decision, not the latest forecast
Grade the forecast the committee approved, on the date it approved it — not a number revised after the store opened. Re-forecasting is legitimate operational planning and worthless as accuracy measurement.
Separating Forecast Error From Execution Error
A store that missed because the model was wrong about the trade area is a different problem from a store that missed because the GM turned over twice in eighteen months. Both show up as the same negative variance, and conflating them is how good models get corrupted: you teach the algorithm to distrust perfectly good sites because a few were operated badly. Before any miss changes the model, run it through an attribution review with operations in the room:
- Staffing and leadership: did the store open with a trained manager, and did that manager stay? Two turnovers in year one is an execution story.
- Build and opening: did it open on time, at the planned size and layout, with the signage package the forecast assumed?
- Marketing: did the launch program run at the planned spend? Forecasts routinely assume awareness-building that quietly gets cut.
- Site conditions on the ground: did the anchor go dark, did a road project block access for a season? Real-world changes to the site, not model failures, but log them separately from either bucket.
- Portfolio effects: did you open something nearby that drew from it? A miss caused by your own new store is a cannibalization modeling gap, which is a model problem — a specific and fixable one.
Only what survives that filter is forecast error, and only forecast error should drive model changes. The split itself is a finding: a meaningful share of underperformers turn out to be fine sites that needed a different operator.
A model that is right on average and wrong on every store is not a model. It is an average.
The Honest Questions to Ask a Vendor
Everyone selling a forecast quotes an accuracy figure. Very few survive contact with these questions. Listen for specifics rather than adjectives.
- →Measured on which stores, and how many? A number from 300 openings across formats means something; a number from nine flagship stores does not.
- →At what maturity was it measured? Early-month accuracy on a ramping store is a different and much easier claim.
- →Were those stores held out of training, or is this in-sample fit? A model can describe stores it has already seen almost perfectly and still fail on the next one.
- →What does the distribution look like? Ask for the share within 10% and 20%, and for the worst five misses.
The held-out question separates real validation from a demo: any model with enough parameters can fit stores it was trained on. Your own scorecard runs that test automatically every time you open a store, which is why an internal accuracy program is stronger evidence than any vendor case study. Our overview of how AI revenue forecasting with mobile data works explains what these models are actually inferring, and site selection software covers how the categories differ on transparency.
What Error Patterns Reveal About Your Model
Residuals are the most valuable proprietary dataset a growing brand owns, and they usually sit unexamined. Sort misses by site characteristic and the noise resolves into structure. The patterns below are illustrative, but versions of them show up constantly:
- Systematic over-prediction of urban infill. Dense daytime population looks enormous in the data, but urban capture rates run lower than the model assumes: competition is denser, visits are shorter, and a big residential count includes people who will never walk in.
- Under-prediction of co-tenancy effects. Models tend to treat neighbors as competition or ignore them. The lift a genuinely complementary neighbor generates is real and frequently missed — see anchor tenants and co-tenancy.
- Optimism in low-awareness markets. Forecasts built on analogs from established markets implicitly assume brand awareness you do not yet have. Stores in new regions miss low, and the miss shrinks as the market matures — a ramp assumption, not a site assumption.
- Digital demand mismatch. Category demand varies more by market than most models allow. According to Semrush’s keyword research data, its keyword database spans 26.7 billion keywords across 142 geographic databases with volume reported down to city and region level — a cheap way to check whether local interest in your category supports the forecast, or explains a systematic miss in one metro.
That last point matters more as local intent grows. Semrush data for 2026 puts “near me” keyword variations at roughly 7.1 million US searches per month, up 29% between Q1 2025 and Q1 2026, with “near me tonight” (+41%) and “near me open now” (+38%) fastest-growing — per Semrush’s keyword volume research. If your model was calibrated on analogs from an era when discovery worked differently, some residuals may be measuring that shift rather than anything about the real estate.
Each named pattern is a fix: urban over-prediction becomes a capture-rate adjustment, under-predicted co-tenancy becomes a variable, low-awareness optimism becomes a market-entry ramp factor. Model improvement is not a rebuild; it is a sequence of specific corrections earned from specific misses.
Building the Feedback Loop
The mechanics are unremarkable, which is why the discipline is rare — it fails on process, not sophistication. Four things have to be true.
- Freeze the forecast in the deal record. At approval, store the number, its date, the analogs used, and the key assumptions as an immutable record on the site. Most accuracy programs die here: two years later nobody can find what was promised.
- Set an automatic grading date. A calendared obligation with a named owner, run the same way for winners and losers. Grading only the failures is how you learn that your model is pessimistic.
- Review as a cohort, quarterly. Individual stores are anecdotes; patterns live at the cohort level. Feed the output back into your new store sales forecasting process as an explicit, documented adjustment.
- Promote every opened store to an analog. A graded store with known actuals and clean attribution is a better analog than anything you can buy, because it is your brand, your format, your operating model.
That last step is where the compounding happens: fifty graded stores make a materially better model than fifty ungraded ones — same real estate, entirely different institutional knowledge. At Locate we treat the loop as part of the work rather than an afterthought, because the analysis that recommends a site and the post-open grade that tests it belong in the same system.
Brands That Grade Themselves Get Better
The advantage here is not a better tool. It is a habit that turns every opening into information. Brands that grade their forecasts learn which site types they misjudge, which misses were really operating problems, and how much confidence a given forecast deserves — and they can prove it to a CFO. Brands that do not make the same mistake at store 40 that they made at store 4, and call the accumulating leases a growth strategy.
Start with one cohort. Pull every store that opened twenty-four months ago, find the approved forecast, compute the error, and sort the misses into model versus execution. It is usually an afternoon of work, and it will change what your next real estate committee argues about. For a second set of eyes, or a forecast designed from the start to be graded, talk to Locate.
Common Questions
- How accurate should a new store sales forecast be?
- For a mature, well-calibrated model working with clean analog data, absolute percentage error in the 10–15% range at maturity is a reasonable expectation for most of your openings, with the distribution mattering as much as the average. What you should be suspicious of is any claim in the low single digits, or any accuracy number quoted without saying which stores it covers, at what point after opening it was measured, and whether those stores were held out of the model's training data.
- When should you measure forecast error for a new store?
- At maturity, not at month three. New stores ramp, and the ramp curve varies by format, market familiarity, and marketing spend, so an early-month comparison mostly measures how fast the store ramped rather than whether the forecast level was right. Measure at the point your brand considers a store stabilized — often 12 to 24 months — and compare trailing-twelve-month sales to the forecast the committee actually approved.
- What is the difference between forecast error and execution error?
- Forecast error means the model was wrong about the location: the trade area, the competitive draw, the co-tenancy, the capture rate. Execution error means the location was fine and the store underperformed for operating reasons — management turnover, staffing gaps, a delayed build-out, a marketing launch that never happened. They demand opposite responses. Forecast error should change the model; execution error should not, or you will teach the model to distrust perfectly good sites.
- What questions should I ask a vendor about their forecast accuracy claims?
- Four: On which stores was this measured, and how many? At what maturity? Were those stores held out of the training data, or is this in-sample fit? And what does the error distribution look like — not just the mean, but the tails? A vendor who can answer all four with specifics is telling you something real. A vendor who quotes a single accuracy percentage with no denominator is quoting marketing.
- How do error patterns improve a site selection model?
- Residuals are the most valuable dataset you own. When you sort misses by site type, you usually find structure rather than noise: systematic over-prediction of urban infill, under-prediction of strong co-tenancy, consistent optimism in markets where you have no brand awareness. Each pattern is a named blind spot that can be corrected with a variable, a weighting change, or better analogs — and each opened store then becomes a new training analog for the next decision.