Site Selection with Open Data: A Practical Method

Site Selection with Open Data: A Practical Method

E
By Etzal Earth
12 min read

Site selection is usually presented as a ranking problem. Six candidate locations go in, a score comes out for each, and the highest number wins. That framing is convenient and it is also the reason most site scoring systems produce results nobody trusts a second time.

The scoring is the easy part. The hard parts are assembling a genuinely comparable picture of each candidate from sources with uneven coverage, deciding how much each factor matters when nobody has calibrated those weights against outcomes, and preventing the arithmetic from punishing a site simply because less is known about it. All three problems are invisible in the final number, which is exactly why the final number gets believed.

What follows is a method that works with open data, states its own limits, and produces output a human can argue with. Arguing with the reasoning is the point. A score a reviewer can only accept or reject is not a decision tool, it is an oracle.

The question a site decision actually asks

Before any data is pulled, it is worth writing down the specific question, because generic site scoring answers no question at all.

A convenience store, a distribution depot, a clinic, and a mobile tower each need a different definition of a good site, and the differences are not about weights, they are about which factors exist at all. The convenience store cares about pedestrian flow past the door and the trade area within a few minutes on foot. The depot cares about heavy vehicle access, turning radius at the entrance, and drive time to a delivery cluster during the hours it actually delivers. The clinic cares about population within a travel time that a sick person can manage on the transport they have. The tower cares about line of sight and power.

Two further things belong in the written question. First, the decision the score has to support: choosing between a shortlist, filtering a long list down to a shortlist, or defending a choice already made. These need different levels of rigour, and only the first two are honest uses. Second, the disqualifiers, meaning the conditions that rule a site out regardless of how well it scores elsewhere. A site inside a flood-prone zone or without a legal use class for the intended activity should be excluded, not given a low score on one factor and averaged back up by strong scores on others. Compensatory scoring hides hard constraints, and hard constraints are where site decisions actually go wrong.

Building a comparable picture from open sources

The available open sources cluster into five groups, and each carries a different kind of unreliability.

  • Access. Road network from OpenStreetMap, giving classification, connectivity, junction density, and one-way or turn restrictions where they have been mapped. Public transport stops and routes where a feed exists. This is usually the strongest open layer in urban areas.
  • Catchment. Gridded population data, administrative population where the census is published, and building footprints as a density proxy. Resolution is the constraint: a gridded surface at coarse resolution smears a dense settlement across the countryside next to it.
  • Competition. Points of interest from OpenStreetMap or an open places dataset, categorized by type. Completeness varies enormously by country, by category, and by how commercially interesting the area is to mappers.
  • Footfall proxies. Road class, junction density, POI clustering, transit access. These are indirect and should be labelled as such at every stage.
  • Physical constraints. Elevation and slope from an open DEM, land cover, flood or hazard zoning where a public layer exists, parcel geometry where a cadastre is open.

Comparability is the requirement that gets skipped. A picture assembled from different sources for different candidates is not a comparison, it is five separate descriptions with a number stapled on. If one candidate sits in a city with an open transport feed and another does not, the transit factor cannot simply be scored high for the first and low for the second. The correct handling is to record the factor as unavailable for the second candidate, which the scoring layer then has to deal with honestly.

The same rule applies to vintage. Comparing a site described by a survey from one year against a site described by a survey from five years earlier will produce a difference, and the difference will partly be time rather than place. Recording the age of each input per candidate is cheap and it prevents the most embarrassing category of error, which is discovering that the winning site won because its data was newer.

Straight lines are not distances

The single most common technical error in site selection is the radius circle. Draw a 1 kilometre circle around each candidate, count the population inside, call it the catchment. It is fast, it is easy to explain, and it is wrong nearly everywhere that matters.

A circle assumes travel is equally possible in all directions at equal cost. Real movement is constrained by the network, and the network is shaped by exactly the features that make a site good or bad. A river with a single bridge, a rail corridor, a limited-access road, a hillside, a gated estate: each of these puts large amounts of population inside the circle and far outside the reachable area. The error is not random. It is systematically largest at sites near barriers, and sites near barriers are common because barriers are where land is cheap.

Network distance fixes the geometry. An isochrone, meaning the set of locations reachable within a stated travel time by a stated mode, is the correct catchment shape. Building one requires a routable network, a mode, a speed profile, and a time of day. Each of those is a modelling choice that should be recorded with the result.

  • Mode matters more than time. A 10 minute walk isochrone and a 10 minute drive isochrone from the same point describe different businesses.
  • Time of day matters where congestion is significant. A free-flow drive time isochrone is an upper bound on reach, and for a site chosen for peak-hour trade it is the wrong bound.
  • Barriers only appear if the network data includes them. If a footbridge is unmapped, the walk isochrone is wrong in the conservative direction. If a private road is mapped without access tags, it is wrong in the optimistic direction.

Population inside an isochrone still has to be computed carefully. Gridded population intersected with an irregular polygon needs area weighting, and area weighting assumes population is uniform inside each cell, which it is not. The estimate is usable, the false precision is not: reporting a catchment of 43,812 people from a coarse gridded surface clipped to a modelled isochrone claims a resolution the inputs never had.

Weights are opinion, and pretending otherwise is the real problem

Once factors are measured, they get combined, and combination means weights. Access 30 percent, catchment 30 percent, competition 20 percent, constraints 20 percent. Those numbers look like parameters. In almost every implementation they are opinions, chosen in a meeting, never validated against a single outcome.

That is not automatically illegitimate. Expert judgement encoded as weights is a reasonable starting point when no outcome data exists, and for a new format in a new market no outcome data does exist. What is illegitimate is presenting the result as if the weights were derived. A weighted score built on uncalibrated weights inherits all the uncertainty of those weights and then hides it behind two decimal places.

Calibration is possible when there is outcome data: a portfolio of existing sites with known performance lets weights be fitted rather than asserted, and the fitted weights frequently contradict the assumed ones. Absent that, three practices keep the system honest.

State the provenance of the weights in the output, in one line, next to the score. Assumed weights should say assumed.

Run the ranking under alternative weightings and report whether the winner changes. If moving access from 30 to 40 percent reorders the top three, the ranking is a statement about the weights, not about the sites. That fact belongs in the deliverable.

Keep the factor count small enough that a human can hold the whole model. Twenty factors at four percent each feel rigorous and are unreviewable. Six to eight factors, each of which someone can defend out loud, produce better decisions because they can actually be challenged.

The trap: less observed scores lower

Here is the failure that quietly ruins comparisons. A candidate in a well mapped commercial district returns rich data for every factor. A candidate in an outer suburb returns fewer POIs, sparser building footprints, no transit feed, and a thinner road network with missing attributes. If missing inputs are scored as zero or low, the second site is penalized for being less observed, and the model has confused the quality of the map with the quality of the place.

The direction of this bias is predictable. Open mapping coverage correlates with density, with wealth, and with how long a mapping community has been active in an area. A site selection model that scores gaps as absences will systematically prefer central, affluent, well mapped locations, and will do so with the appearance of objectivity. If the whole purpose of the exercise was to find an underserved area worth entering, the model is now steering directly away from the answer.

It is worth separating three things a low value can mean, because they call for different handling: the factor was measured and is genuinely poor, the factor was not measurable here, or the factor was measured against a source whose coverage in this area is known to be weak. The first is a finding. The second is a gap. The third is somewhere between and should be treated as a gap with a note.

Redistribute weight, do not score a gap as zero

The mechanical fix is to give unknown a real representation in the arithmetic instead of substituting a number.

When a factor cannot be evaluated for a candidate, remove its weight from that candidate's denominator and rescale the remaining weights so they still sum to one. The score stays on the same nominal scale, and the missing factor neither helps nor harms. Alongside the score, carry a completeness figure: the fraction of the intended total weight that was actually satisfied. A score of 71 at full completeness and a score of 71 at 60 percent completeness are not the same claim, and the completeness figure is what stops them being read as the same claim.

Two rules keep redistribution from being abused. Set a floor: below some completeness threshold, refuse to produce a score at all and return the reason. A site evaluated on two of eight factors should not appear in a ranking, because the redistributed score is dominated by whichever factors happened to be available. And never redistribute across a disqualifier. If the flood zoning layer is missing for a candidate, that is not a factor to be rescaled away, it is an unverified hard constraint, and the correct output is that the site cannot be cleared.

Redistribution also has a reporting obligation. The output should name which factors were dropped for which candidates, because that list is often the most actionable part of the whole analysis. It tells the team exactly which site visit or which data purchase would change the decision, and that is worth more than the ranking.

Show the contributions, not just the number

A score presented alone can only be accepted or rejected. A score presented with its contributions can be argued with, and an argument is what improves a decision.

The minimum useful presentation gives, per candidate, each factor's raw measured value in its natural units, the normalized value, the weight applied, and the resulting contribution to the total. Population within a 10 minute walk of 18,000, normalized to 0.62, weighted 0.30, contributing 0.186. A reviewer can look at that line and say the walk time is wrong for this format, or that the normalization ceiling is set too low, or that 18,000 sounds implausible for that neighbourhood. Every one of those objections improves the model. None of them is available if the output is 71.

Normalization deserves its own visibility, because it silently encodes assumptions. Scaling a factor against the minimum and maximum across the candidate set means a score is relative to the shortlist and will change if a candidate is added or removed. Scaling against an absolute reference means the score is stable but the reference is another uncalibrated opinion. Both are usable. Which one is in force must be stated, because the two produce different answers to the question of whether a site is good in absolute terms or merely the best of these six.

The same transparency applies to any place where a value was inferred rather than measured. An inferred value in the contribution table should be marked as inferred at the point it appears, not in a footnote. Marks that survive to the point of reading are the only marks that do anything.

What the score can and cannot decide

A well built open-data site score is a filter and a structuring device. It narrows a long list, forces the criteria to be written down, and makes the reasoning inspectable. Those are real and they are worth the effort.

It is not a measurement of expected performance, and it should not be presented as one. The factors are proxies, the weights are usually assumed, and the catchment is modelled from a network graph that is itself incomplete. Between two candidates separated by a few points, the score has nothing to say, and the honest output for that pair is that they are indistinguishable on the available evidence, with a list of what would distinguish them.

The most useful thing the method produces is often not the ranking. It is the record of what could not be determined for each site: the missing transit feed, the POI category with poor local coverage, the hazard layer that does not exist for that province. That list is a work plan. The ranking is a hypothesis, and the difference between a team that treats it as a hypothesis and a team that treats it as an answer shows up years later, in the sites that were never visited because a number said they were fourth.