Where to Find Open Building Footprint Data, and What It Costs You

Where to Find Open Building Footprint Data, and What It Costs You

E
By Etzal Earth
11 min read

A building footprint is a closed polygon drawn around the outline of a structure. It is the simplest object in applied geospatial work, and it is the one that most often ships wrong, because a footprint layer that looks correct in the three neighbourhoods someone spot-checked can be systematically broken across the region the analysis actually depends on.

There is more open footprint data available now than at any earlier point, and it arrives from sources with almost nothing in common. Some of it was drawn by people tracing imagery by hand. Some was extracted by models run across continental image mosaics. Some was released by national mapping agencies that have maintained cadastral registers for a century. These are not three flavours of the same product. They fail in different directions, and choosing between them is mostly a question of which failure the use case can survive.

The cost in the title is not money. All of these are free to obtain. The cost is the assumptions inherited with the geometry, the completeness that cannot be observed from inside the data, and the licence obligations that follow the polygons into whatever gets built on top of them.

The three families, and why the distinction matters

Community-mapped footprints come from people tracing structures from aerial or satellite imagery, sometimes with local knowledge, sometimes with none. OpenStreetMap is the dominant example. The unit of production is a human decision, so the geometry tends to be sensible: a person looking at a terrace of houses usually knows it is several houses and splits them, because that is what a person sees.

Machine-extracted footprints come from segmentation models run over large imagery mosaics, then vectorised into polygons. Several organisations have released continental or near-global layers built this way. The unit of production is a pixel classification, so coverage is enormous and uniform in effort, but the geometry carries the model's confusions rather than a human's.

National releases come from mapping agencies, cadastral authorities and municipal open data portals. The unit of production is a legal or administrative record, often maintained for taxation or planning rather than for mapping. Where these exist they are usually the most accurate footprints available, and they usually stop dead at the national border, sometimes at a municipal one.

The important consequence is that these three families have uncorrelated error profiles. A structure missing from one is not necessarily missing from the others, and a shape that is wrong in one is rarely wrong in the same way in another. That is useful for validation, and it is exactly what makes naive merging dangerous.

What machine extraction actually gets wrong

Machine-extracted footprints fail in recognisable, repeatable ways, and knowing them is the difference between using the data and being surprised by it.

Merged terraces are the most common. A model segmenting a row of attached houses sees one continuous roof mass with no reliable visual break at the party walls, and emits one polygon. The building count for the street collapses. Anything downstream that assumes one polygon equals one address, one household or one premises now carries an error that is not random: it is concentrated in exactly the dense residential fabric where most people live.

Missed small structures are the mirror image. Outbuildings, informal extensions, market stalls, single-room dwellings and anything close to the resolution limit of the source imagery drop below the detection threshold. In settlements built from small structures, this is not a rounding error in the count, it is most of the count.

Phantom buildings appear from shadow and from anything that looks like a roof to a model trained mostly on roofs. Dark shadow cast onto open ground can carry the texture and shape signature of a low building. Shipping containers, large vehicles, solar arrays, tarpaulins, water tanks and the flat tops of walls all produce plausible detections. These are hardest to catch precisely because they occur where verification is least convenient.

Geometry itself degrades in specific ways. Corners get rounded by the raster to vector step, then over-regularised by the polygon simplifier that runs afterwards, producing a shape that is neither faithful to the roof nor faithful to a rectangle. Courtyards and light wells get filled in. Roof overhangs are included, so the polygon describes the roof rather than the ground footprint, which matters when the number feeds a floor area calculation.

Registration offset is subtle and consequential. If the source imagery was not perfectly orthorectified, tall buildings lean away from the sensor nadir, and the extracted polygon lands displaced from the true ground position by an amount that grows with height and with distance from the image centre. The shape can be right and the position wrong.

Completeness varies by region, and the variation is systematic

The single most misleading property of open footprint data is that missing buildings look exactly like empty land. Nothing in a polygon layer announces its own incompleteness.

Community mapping follows contributors. Contributors follow population with a heavy weighting toward wealth, connectivity, disaster response events and mapping campaigns. A city that hosted a humanitarian mapping effort can be better mapped than a wealthier city that never needed one. Coverage is therefore lumpy at a scale of kilometres, not smoothly declining with remoteness.

Machine extraction follows imagery. Where the available imagery is old, cloudy, low resolution or simply not in the mosaic the model was run over, extraction produces nothing, and the gap has a shape: it follows persistent cloud belts, image acquisition schedules and licensing boundaries.

Extraction quality also follows building style. Models trained predominantly on one architectural vocabulary underperform on others: flat roofs of the same tone as the surrounding ground, thatch, corrugated metal that saturates the sensor, structures shaded by dense canopy, and settlements whose plan geometry does not resemble the training regions.

All three of these mechanisms correlate with income, latitude and cloud climatology. That is what makes the gaps systematic. An analysis comparing two cities is not comparing two cities, it is comparing two sampling regimes, and the poorer, cloudier, less mapped one will read as less built, less dense and less economically active for reasons that have nothing to do with the ground.

The defensive habit is simple to state and rarely practised: before comparing anything across regions, measure the data density in each region and treat a large difference as a reason to stop, not as a finding.

Attributes are where the disappointment lives

Most people who go looking for footprint data are not really after outlines. They want height, use type, number of storeys, construction material, age or occupancy. Those attributes are usually absent.

Height is present in a minority of records, usually where a national agency derived it from a surface model or where a local community tagged it. Where it appears in machine-derived layers it is typically itself an estimate, and it should be read as an estimate with a method behind it rather than a measurement.

Building type is worse. A polygon tagged as residential may have been tagged from a default assumption, from a land use layer intersected with the footprint, or from an actual survey. Those are radically different evidential positions carrying the same word. Mixed use buildings, which are the norm in older urban cores, are forced into a single category by most schemas.

Age and material are close to nonexistent outside specific national registers. Anything claiming continental coverage of construction year is almost certainly modelling it, and modelled age is fine as long as nobody downstream treats it as a record.

The practical rule is to check, per record, whether an attribute came with the geometry or was joined on later. A joined attribute inherits the error of the join as well as its own, and joins between footprint layers and address or parcel layers are exactly where silent duplication and misassignment happen.

Licensing is not a footnote

Open footprint sources differ in licence in ways that determine what can be built.

Share-alike licences, of which ODbL is the one most often encountered, require that derived databases be released under the same terms. What counts as a derived database rather than a produced work is a genuinely hard question, and it turns on how tightly the source data is bound into the output. A map image rendered from the data and a database that embeds the geometry sit differently.

Permissive licences typically require attribution and little else. Public domain dedications require nothing, though attribution remains good practice and good defence.

National releases carry national terms, which can include restrictions on commercial use, on redistribution, or on combining with other datasets. These are frequently written for a pre-API world and do not map cleanly onto serving data through an endpoint.

The failure mode to avoid is mixing families without tracking provenance per record. Once a share-alike layer has been conflated into a permissive one and the origin of individual polygons has been lost, the entire result inherits the strictest obligation present, and there is no way to unpick it later. Provenance per feature, carried all the way through the pipeline, is the only version of this that survives contact with a legal review.

Choosing: coverage or accuracy

The choice becomes straightforward once the question is framed as which error is fatal.

Choose coverage, meaning machine-extracted continental layers, when the analysis is statistical and operates over many structures at once. Density surfaces, settlement extent, change over large areas and screening exercises all tolerate individual polygons being wrong because the errors partially cancel and the conclusion does not rest on any single feature.

Choose accuracy, meaning national releases or well-mapped community data, when a decision attaches to an individual building. Insurance on a specific address, a planning judgement, a site assessment, anything where someone will point at one polygon and act on it. Here a merged terrace or a phantom structure is not noise, it is the whole answer being wrong.

Choose community data specifically when semantics matter more than geometric precision: what the building is called, what it contains, how it relates to the street. That information exists only where people put it there.

A useful middle path is to use a wide-coverage layer as the base and a higher-quality layer as an override where it exists, keeping both identifiers on the feature. This preserves the ability to answer the question that matters later, which is not what the geometry says but where it came from.

Validating before trusting

Validation of footprint data does not require a labelled benchmark. It requires a stratified sample and some patience.

Sample by settlement type rather than uniformly. Uniform sampling over an area lands almost everywhere in whatever land cover dominates, which is usually not built up, and produces an accuracy figure that reflects empty land. Stratify by dense core, suburban, peri-urban, informal, industrial and rural, and evaluate each separately. The variation between strata is normally larger than the headline figure, and the headline figure hides the exact place the data is weakest.

Check three things per sample tile. Count agreement: does the number of structures match what is visible in imagery. Shape agreement: do the outlines follow the roof lines. Position agreement: are the polygons displaced consistently in one direction, which points at an orthorectification or datum problem rather than at extraction quality.

Cross-family comparison is the cheapest signal available. Where two independently produced layers agree, confidence is reasonable. Where they disagree, something is wrong in at least one of them, and the disagreement map itself is often more informative than either source. Disagreement clusters geographically, and those clusters are where any manual effort should go.

The failure that actually ships

The footprint failure that reaches production is rarely a wrong polygon. It is a count.

Someone joins a footprint layer to a population raster, or to a business register, or to a hazard extent, and produces a number: buildings exposed, premises in catchment, structures at risk. The number is computed correctly. The polygons underneath it merged every terrace in the dense district into one feature, missed most of the informal settlement entirely, and included forty shadows on the edge of town as buildings. The number arrives in a report with no method note, no provenance and three significant figures.

Nobody catches it, because the map looks right. Footprint layers always look right at the zoom level people review them at.

The habit that prevents this is to require, for any building count, an accompanying statement of which layer produced it, when that layer was built, what its known completeness behaviour is in that region, and what the count would have been under the other available source. If the two sources give counts that differ by a factor that would change the decision, the honest output is a range, not a number.