Most spatial analysis is about finding things. Where are the clinics, which roads connect to the port, how many buildings sit inside the flood extent. Those questions are tractable because the evidence is positive: a feature either appears in the data or it does not, and when it appears you can inspect it.
Gap analysis inverts the question, and the inversion is much harder than it looks. Finding the districts with no health facility, the settlements with no all-weather road access, the neighbourhoods with no water point means reasoning about absence, and absence in an open dataset has two completely different causes that produce identical evidence. Either the facility does not exist, or the facility exists and nobody recorded it.
Everything about the value of a gap analysis depends on separating those two. Get it right and the output steers investment toward places that genuinely lack service. Get it wrong and the output steers investment away from places that merely lack data, which is usually the same set of places that lack service, compounding the original problem rather than correcting it.
Absence is the hardest thing to observe
A dataset records what someone chose to record. It contains no explicit statement about what was looked for and not found, because recording non-findings is expensive, unrewarding, and outside the workflow of almost every data collection process that exists.
The consequence is that no open geospatial dataset can distinguish, from the data alone, between an area that was surveyed and found empty and an area that was never surveyed. Both appear as a region with no features. The distinction lives outside the feature data, in metadata about the survey process, and that metadata is usually either absent, incomplete, or expressed at a resolution too coarse to apply.
This is not a defect of any particular source. It is a structural property of feature-based data collection. A community mapping project stores the buildings that were traced, not the tiles that were checked and found to contain nothing. A national facility register stores the facilities that were registered, not the villages that were visited and found to have none. Turning a record of presences into a claim about absences requires an additional input in every case, and the whole discipline of gap analysis is really the discipline of finding that additional input.
Two kinds of empty, and why the distinction decides everything
A real gap is an area where the service does not exist. An unsurveyed area is an area where the data collection process has not reached. These are different objects and they call for opposite responses.
A real gap is an investment opportunity or a service failure, depending on who is looking. It supports a decision: build here, extend the route, deploy the mobile unit. An unsurveyed area supports exactly one decision, which is to go and look, or to find another source that has looked. Sending a construction budget to an unsurveyed area on the assumption that it is empty is how a programme discovers, expensively, that the facility was already there.
The distinction also determines whether the analysis has any meaning at all. If the proportion of the study area that is unsurveyed is large and unknown, the gap map is a map of survey coverage wearing the costume of a needs assessment. It will correlate with need, because unsurveyed areas are often underserved, and that correlation is what makes the error so durable: the output looks plausible to everyone who reviews it, including people who know the region.
The practical test to apply before any gap analysis is whether the study area's survey coverage can be characterized independently of the features themselves. If it cannot, the analysis can still be run, but its output is a list of candidates for investigation rather than a list of gaps, and it must be labelled that way from the first slide onward.
Coverage bias runs in a predictable direction
Coverage in open geospatial data is not randomly distributed. It concentrates where mapping is easy, valuable, and populated by people with the time, connectivity and tooling to do it.
The pattern is consistent across sources. Urban areas are better mapped than rural ones. Wealthier areas are better mapped than poorer ones. Formal settlements are better mapped than informal ones. Commercial data has its own version of the same bias, since collection effort follows commercial value, and the areas with the least commercial value are precisely the areas a gap analysis is looking at.
Imagery-derived layers reduce some of this and introduce their own. Building footprints extracted automatically from satellite imagery cover large areas uniformly, which is a genuine improvement, but detection performance varies with roof material, building size, density and image quality. Small structures with vegetated or low-contrast roofs are missed more often, and those are concentrated in exactly the places already under-recorded. Uniform coverage of the imagery does not mean uniform accuracy of the extraction.
The direction of the bias means the errors do not cancel. A gap analysis run naively over an open dataset will find the most gaps in the poorest, most rural, least connected areas, and it will find them whether or not the services are there. Any result that ranks areas by apparent deprivation and happens to reproduce the ranking of mapping coverage should be treated as unvalidated until the two are shown to be separable.
Triangulating absence from independent sources
The only reliable way to raise confidence in an absence is agreement between sources whose failure modes are unrelated.
Independence is the operative word. Two datasets that both derive from the same imagery, or one that was seeded by importing the other, are not independent, and their agreement adds almost nothing. Independence has to be assessed on provenance, not on the fact that the files came from different websites. A surprising amount of open geospatial data shares ancestry, and an import performed years ago can make two sources agree in a way that looks like corroboration and is actually an echo.
Useful independent combinations exist. A community-mapped feature layer and a government facility register are collected by different people through different processes, so their agreement about the absence of a clinic in a district is meaningful. A settlement layer derived from imagery and a population grid derived from census disaggregation come from different chains of evidence about where people are. Night lights indicate electrified activity through a mechanism entirely unrelated to whether anyone mapped a road.
The procedure is straightforward once independence is established. For each candidate gap, check every available independent source for evidence of presence. Where all independent sources agree on absence, confidence is high. Where sources disagree, the area is a conflict, which is more interesting than either a confirmed gap or a confirmed presence, because it usually means a real feature exists in a form one process did not capture. Where only one source has coverage, the absence is unverified and should be reported as such rather than being folded into the gap count.
The output of triangulation is therefore not a binary. It is three classes, and the third class, unverified, is the one that most analyses erase.
Proxies for presence
Where no direct source covers an area, indirect evidence can still discriminate between an empty place and an unobserved one, because most infrastructure leaves traces in signals that were collected for other reasons.
- Settlement signal. If a population grid or an imagery-derived settlement layer shows people living somewhere, then the absence of any recorded feature is suspicious rather than informative. A populated area with zero mapped features is a coverage gap until proven otherwise.
- Access signal. Roads and tracks lead somewhere. A road network that terminates at a cluster of buildings implies a destination, and destinations usually have functions. A settlement with mapped access and no mapped services is a stronger gap candidate than one with no mapped anything.
- Activity signal. Night light emissions indicate electrified activity. Persistent light where the feature layer is empty is evidence of something the feature layer missed.
- Structural signal. Building footprints from automated extraction cover areas the community has not reached. A large structure with a distinctive footprint in an otherwise residential pattern is a candidate for a school, a warehouse or a clinic.
- Named-place signal. Gazetteers and administrative place lists record settlements independently of whether anything inside them was mapped. A named settlement with no features is almost always a coverage gap.
None of these confirms a facility. Their value is asymmetric: they are much better at raising doubt about an apparent absence than at confirming one. That asymmetry is useful, because the expensive error in gap analysis is the false gap, and proxies for presence attack exactly that error.
A composite screen falls out of this. Any candidate gap that sits inside a populated, accessible, lit, structurally built-up area with a name should be demoted from gap to unverified, on the grounds that a place with that much human activity and no recorded features is more likely under-mapped than under-served.
The policy consequence of getting it wrong
The reason to be careful is not methodological purity. Gap analyses are used to allocate resources, and a wrong gap map moves money.
The failure mode has a shape. Areas with poor data coverage are disproportionately poor and rural. A naive gap analysis marks them as having many gaps. Attention and budget flow toward them. So far this looks like a success, and in a rough sense it may be one. The damage appears at the next iteration, when the analysis is used the other way: to identify areas that already have adequate service, or to verify whether an intervention worked, or to allocate between two underserved regions with different mapping histories. In every one of those uses, the better-mapped region shows more existing service and gets less investment, and the difference is an artefact of who did the mapping.
There is a second, subtler consequence. When a gap analysis is wrong in a place where local officials know the ground truth, it destroys the credibility of the whole method for that audience. A district head who is told their district has no secondary school, when there are three, will discount every subsequent output from the same system, including the correct ones. The cost of a false gap is not confined to the resource it misdirects.
A useful discipline is to write down, before the analysis runs, which decisions the output will be allowed to inform. Prioritizing where to send a verification team is a low-risk use that tolerates false gaps. Deciding not to fund a region because it appears already served is a high-risk use that a data-derived gap map cannot support without ground validation.
Reporting confidence in the absence
The deliverable should assert an absence only where the evidence supports it, and should express what the evidence is everywhere else. That means the output carries more than a list of gaps.
Per gap, four properties are worth carrying explicitly. The sources checked, named individually, because a gap confirmed by two independent sources is a different object from one confirmed by a single layer. The coverage status of each source in that specific area, since a source may be excellent nationally and absent in one province. The proxies for presence that were evaluated, and whether any of them fired. And a resulting confidence class, with the classes defined in terms of the evidence rather than as arbitrary labels.
Three classes are usually enough: confirmed gap, where multiple independent sources with local coverage agree on absence and no presence proxy fired; probable gap, where a single covered source shows absence, or where sources agree but proxies raise doubt; unverified, where no source has usable coverage in that area. The third class must be published, not hidden, because its size is the honest measure of how much of the study area the analysis could speak to at all.
Aggregate reporting follows the same rule. A count of gaps at national or provincial level should be accompanied by the proportion of the area that was unverified. A statement that a region contains a certain number of unserved settlements, when a large share of its settlements could not be assessed, is a statement about the assessed subset, and the sentence should say so.
The map deserves particular care, because a map with blank areas reads as empty and a map with dots reads as full. Rendering unverified areas as a distinct visual class, rather than leaving them white, is the difference between a reader seeing no data here and a reader seeing nothing here. That single rendering choice probably prevents more misinterpretation than any amount of accompanying text.
Fieldwork is the resolution step, and it should be targeted
No amount of remote analysis converts an unverified area into a confirmed gap. The conversion requires somebody to look, whether by field visit, by phone call to a local authority, or by manual inspection of high-resolution imagery where the feature type is visually identifiable.
The right way to use that budget is to let the analysis rank the uncertainty rather than the gaps. The highest value verification targets are the conflicts, where independent sources disagree, and the unverified areas with the strongest presence proxies, because those are the places where the analysis is most likely to be wrong in the expensive direction. Sending a team to confirm a gap that three independent sources already agree on is a use of resources that produces almost no information.
Verification also has to flow back. A field visit that finds a facility is a new record, and if it is contributed to the open source it improves the next analysis for everyone, including the next organization to run the same query. Gap analyses that consume open data without returning anything to it leave the coverage bias exactly where they found it, and then run again next year against the same map.
The honest position at the end of a gap analysis is narrower than the one most reports take. The analysis has produced a ranked set of hypotheses about where service is missing, with an explicit statement of how much of the area it could not speak to. That is genuinely useful and it is not the same as a map of what is not there. The distance between those two claims is the distance between a tool that improves decisions and one that launders a survey budget into a fact.