Retail Catchment Analysis Without Proprietary Footfall Data

Retail Catchment Analysis Without Proprietary Footfall Data

E
By Etzal Earth
12 min read

Footfall data is sold, not published. Mobile location panels, card transaction panels and sensor networks all live behind commercial contracts, and for most teams the licence is either unaffordable or unavailable in the market they care about. The question is what can be built without it.

A useful amount, as it turns out, provided the output is labelled correctly. Open road networks, open places data, transit feeds and gridded population support a defensible model of where a store's customers plausibly come from, how many of them there might be, and who else is competing for them. What they do not support is a count of people walking past the door, and the failure mode of this whole discipline is quietly presenting a model as if it were that count.

The distinction between a modelled catchment and a measured one is the entire subject. Everything below is either a technique for making the model less wrong or a technique for making sure nobody mistakes it for a measurement.

What a catchment actually is

A catchment is the area from which a site draws its customers. That definition is deceptively simple, because a catchment is not a boundary, it is a gradient of probability, and it is not a property of the site alone.

Three properties get lost when a catchment is drawn as a line on a map. It is probabilistic: the chance that a household at a given location shops at this site declines with travel cost rather than dropping to zero at an edge. It is competitive: the same household is inside the catchment of every competing site it can reach, and its choice depends on the alternatives, not only on the distance to this one. And it is format-specific: a supermarket, a coffee shop and a furniture showroom have different catchments from the identical location, because people will travel different distances for different purposes.

That last point kills the idea of a single reusable catchment per site. A 5 minute walk catchment describes the trade area of a convenience format. A 20 minute drive catchment describes a destination format. Applying the wrong one produces a number that is not merely imprecise, it is answering a different question.

The practical consequence is that a catchment model needs a stated purchase mission before it needs any data: what someone is coming for, how they are travelling, and how far they would reasonably go for it. Those three answers determine the mode, the threshold, and the decay assumption. If they are not written down, they will be chosen implicitly by whoever picks the radius.

The circle is almost always wrong

The radius circle survives because it is easy to compute and easy to draw. It assumes that travel cost is proportional to straight-line distance in every direction, which is true only on an open plain with no barriers.

Real catchments are shaped by the network. A one-way system pushes the catchment asymmetrically. A dual carriageway with no crossing for several hundred metres cuts the catchment in half along that line while leaving the circle intact. A river, a rail corridor, a walled development, a steep slope, a motorway junction that is only accessible from one direction: all of these move large amounts of population from inside the circle to outside the reachable area, and they do it without leaving any trace in the circle-based number.

The error has a systematic component that matters commercially. Barriers depress land value, so sites near barriers are common in retail property. The circle overestimates catchment most severely exactly where a business is most likely to be considering a cheap site, which means it flatters the weakest candidates in a shortlist. It is not a neutral approximation.

There is one case where a circle is defensible: an early, coarse filter over a long list of locations where the intent is to eliminate obviously unsuitable sites cheaply, and where the results are never reported as catchment sizes. Any use where the number reaches a decision maker needs a network-based catchment.

Isochrones, and what they assume

An isochrone is the area reachable within a given travel time from a point by a given mode. Computing one requires a routable network, a mode profile, a speed assumption and a departure time. Each of those is a modelling choice and each one changes the answer.

  • Mode. Walk, cycle, drive and public transport produce different shapes, not scaled versions of the same shape. Walking uses paths and crossings that driving cannot; driving uses roads that walking should not.
  • Speed. A walk isochrone built on a flat constant speed ignores slope, crossing waits and stairs. In hilly terrain the downhill side of a walk catchment can be substantially larger than the uphill side, and a constant-speed model shows them as equal.
  • Time. A free-flow drive isochrone is a best case. If the trade is concentrated at peak hours, the free-flow catchment describes a situation the customer never experiences.
  • Network completeness. The isochrone can only use edges that exist in the data. A missing footbridge or an unmapped path through a housing estate shrinks the modelled catchment. An access-restricted road tagged as public inflates it.

The threshold itself deserves scepticism. Cutting at 10 minutes implies that someone at 9 minutes 50 seconds is a customer and someone at 10 minutes 10 is not. A distance decay treatment is closer to reality: weight each location by a declining function of travel cost rather than including or excluding it. The weighting function is another assumption, and unless it has been fitted to real customer origins it is judgement, but it produces a more truthful shape than a hard edge and it makes the catchment comparable between sites with different network densities.

Competition and cannibalization inside the catchment

A catchment computed for a site in isolation systematically overstates its trade, because it counts every reachable household as if the site were the only option.

Competitive allocation is the correction. The general form, familiar from spatial interaction modelling, is that the probability a location chooses a given store rises with the store's attractiveness and falls with the travel cost to it, normalized across all the stores that location can reach. Attractiveness in a proprietary model comes from floor area, brand and range. In an open-data model, proxies have to stand in: building footprint area as a crude size proxy, category and brand tags from open places data, sometimes the number of tagged features in a complex.

Two honest cautions apply. Open places data has uneven category completeness, so a competitor set assembled from it may be missing independent operators, which are systematically less well recorded than chains. And attractiveness proxies are weak: footprint area does not distinguish a full-range store from a warehouse. The allocation is still worth doing, because a model that ignores competition entirely is wrong in a much larger way, but the competitor set should be reviewed by someone who knows the area before the output is trusted.

Cannibalization is the same computation applied to the operator's own network. When a new site is added, the model reallocates demand across all sites including the existing ones, and the difference between the network total before and after is the genuinely incremental trade. A new site that looks strong on its own catchment can be almost entirely cannibalistic if it sits inside the catchment of two existing sites. Presenting site-level potential without a network view is the standard way a portfolio grows without growing.

Footfall proxies, and their honest limits

Where no measured footfall exists, several open signals correlate with pedestrian activity. Each is a proxy, and the word proxy has to survive to the output.

  • Road classification and street network structure. Junction density, block size and the mix of street types describe how walkable a fabric is. Fine-grained networks with many junctions carry more pedestrian movement than coarse ones with long blocks.
  • Connectivity and through-movement. Network centrality measures, which rank street segments by how often they lie on efficient routes between other places, give a purely topological estimate of where movement concentrates. This is a well established technique in street network analysis and it uses nothing but the graph.
  • POI density and mix. A concentration of complementary uses generates trips. A street with food, services and retail together sustains more pedestrian presence than a single-use frontage, and open places data can measure that mix even where it cannot measure volume.
  • Transit access. Stop locations and service frequency where an open feed exists. A high-frequency stop is a pedestrian generator, and its absence from the model is a visible hole in the estimate.
  • Residential and workplace density. Gridded population for residents; workplace population is rarely open at useful resolution and is often the largest missing input in a city centre model.

The limits need stating in the same breath. None of these is a count, and their relationship to actual volumes is not fixed: the same centrality value corresponds to different pedestrian volumes in different cities, in different parts of the same city, and at different times of day. They have almost no temporal resolution, so they cannot distinguish a street that is busy at lunchtime from one that is busy in the evening, which is a first-order distinction for food formats. And they measure potential rather than realized activity: a highly connected street with vacant frontage will score well and carry nobody.

A composite index built from these signals is an ordering, not a measurement. It can say that one location is likely busier than another within the same city. It cannot say how many people pass, and any output that renders it as an implied count has crossed the line.

The ecological fallacy

Every area statistic in this method describes an area, and the temptation is to attribute it to the individuals inside. That inference is invalid, and it is the most common analytical error in catchment work.

A grid cell with a high average income does not mean the households near the site are affluent; it means the average across the cell is high, and averages hide distributions. An area with a high proportion of young adults does not mean the customer base is young, because the people who actually visit are a self-selected subset of the residents. A catchment described as containing a certain demographic profile is describing the residential population of a modelled area, not the customers.

The fallacy compounds when profiles are used to justify format decisions. A store range chosen for the average resident of a catchment can be wrong for every actual customer if the catchment contains two distinct populations, which mixed neighbourhoods routinely do. Aggregating them into a mean produces a description of a person who does not exist.

The mitigation is mostly linguistic discipline backed by geometric care. Report area statistics as properties of the area. Use the finest geographic unit the data honestly supports rather than aggregating up, since coarser units mix more distinct populations. Where a distribution is available, report a distribution rather than a mean. And treat any statement about who the customers are, as opposed to who lives nearby, as a claim requiring evidence the model does not contain.

Validating where real numbers do exist

A model built entirely from proxies can still be tested, because most operators have some real numbers somewhere, and a few validation points constrain a model far more than none.

Existing store performance is the strongest available anchor. Fitting or checking the model against known turnover across an existing estate reveals whether the catchment definition, the decay function and the attractiveness proxies produce a sensible ordering. The ordering matters more than the absolute values: if the model ranks existing stores in roughly the order their performance ranks, it has some claim to ranking candidates.

Loyalty or delivery postcodes, where they exist, are more direct still, because they describe actual customer origins. They are biased toward the customers who signed up or ordered online, so they should not be treated as the full origin distribution, but they can be used to check the shape and extent of the modelled catchment: whether the real origins fall inside the isochrone, and how the density of origins decays with travel time. That decay curve, measured, is worth far more than any assumed one.

Manual counts are cheap and undervalued. A few hours of counting at a location, repeated across a handful of sites and time windows, gives a small set of anchor points that can be regressed against the proxy index. A proxy with a known relationship to a handful of real counts in the same city is a different instrument from a proxy with no anchor at all.

Validation has to be reported alongside the model, including when it fails. A model whose proxy index has no relationship to the counts that were taken is a model that should not be used for that city, and finding that out is a successful outcome of validation rather than an embarrassment.

Communicating a modelled catchment

The last risk is presentational. A catchment map is a persuasive object: a smooth polygon on a clean basemap looks like a survey result, and the smoother it looks the more it will be believed.

A few practices keep the reading honest. Show the network the isochrone was built from, so the shape is visibly a consequence of streets rather than an authored boundary. Use a graded surface rather than a hard-edged polygon where distance decay is in the model, because the gradient communicates probability and the edge communicates certainty. Label every derived figure with its basis: modelled, proxy, or measured, with measured reserved for numbers that were actually counted. State the mode, the threshold, the time assumption and the date of the network extract next to the map, since a catchment is not reproducible without them.

The failure this discipline exists to prevent is specific and it is not hypothetical. A modelled catchment population becomes a slide, the slide becomes a sales forecast, the forecast becomes a lease, and by the time anyone asks where the number came from, its provenance has been lost through four hands. The only reliable defence is that the number carried its own qualification everywhere it went, in the field name, in the label, and in the rounding. Anything looser, and the model will eventually be quoted as a count.