Elevation data looks like the most solid ground in geospatial work. Terrain does not move much, the files are just grids of numbers, and every provider ships something global and free. That impression of solidity is why elevation is where a lot of quietly wrong analysis originates.
The failures are rarely dramatic. A flood model produces a plausible looking inundation map that is wrong in a specific, systematic way. A viewshed says a tower covers a village that it does not cover. A slope map exaggerates steepness along one direction. In each case the arithmetic is correct and the input was not what the analyst assumed.
Most of these trace back to two things: not knowing whether the model includes trees and buildings, and not knowing what the heights are measured from.
What a digital elevation model is
A digital elevation model is a raster where each cell holds one number representing height at that location. That is the whole data structure. Everything interesting is in what the number means.
The first question is what surface the number describes. A digital surface model records the top of whatever is there: bare ground in an open field, the top of the canopy in a forest, the roof of a building in a city. A digital terrain model records the ground itself, with vegetation and structures removed. The generic term covers both, which is exactly why the generic term causes so much trouble.
The second question is how the number was produced, because that determines the failure pattern.
- Radar interferometry compares phase between two radar observations to derive height. It works through cloud and at night, which is why it produced the first consistent near global datasets. Over vegetation the radar signal penetrates partway into the canopy, so the resulting height is neither the treetop nor the ground but something in between, varying with vegetation type and moisture.
- Optical stereo derives height from parallax between images of the same place taken from different angles. It needs cloud free views and it needs texture, so it struggles over uniform surfaces such as snow, sand, and still water, and it fails in permanent shadow.
- Airborne laser scanning fires pulses and times the returns. Because some pulses reach the ground through gaps in the canopy, it can produce both a surface and a genuine terrain model from the same acquisition. It is far more accurate than the global products and it is only available where someone paid to fly it.
- Photogrammetry from drone imagery is the local, cheap version of optical stereo, with the same fundamental limitation that it sees the top of things.
The widely used open global datasets divide along those lines. SRTM, the radar interferometry mission that produced the first near global model, is still ubiquitous, still cited across a great deal of published work, and still adequate for many purposes, and a reprocessed version of it addresses some of the original's voids and noise. The Copernicus DEM, derived from a later dedicated radar satellite pair, offers better consistency and fewer gaps and has become the sensible default for global work. Optical stereo global models such as the ALOS World 3D derivatives and the ASTER global model behave differently again, with characteristic noise and artifacts rather than the smoothness of radar products. In parallel, a growing number of countries publish national laser scanning products, and where those exist they are almost always the right answer.
The important pattern: the near global products from radar and optical stereo are surface models. Terrain models at global scale are produced by estimating and removing vegetation and buildings, which is a modeling step with its own error, not a measurement.
The distinction that silently corrupts analysis
Using a surface model where a terrain model is required does not produce an obvious error. It produces a result that looks normal and is wrong in a consistent direction.
Flood modeling is the clearest case. Water flows downhill across the ground, under tree canopy and around buildings. In a surface model, a stand of trees is an elevated ridge many meters high and a building is a solid block. Route water across that and the trees act as a dam, the building footprints act as barriers, and the water gets directed along whatever gaps exist, typically streets, which happens to look convincing because water really does flow along streets. The map looks right and the reasoning behind it is wrong, so it fails where it matters: in an area of scattered structures or dense vegetation, the model will hold water back from places that would actually flood and channel it into places that would not.
The direction of the error is also unhelpful. Buildings, whose interiors are the thing a flood analysis is meant to protect, are the highest points in a surface model, so a naive depth calculation reports zero depth exactly at the structures of interest.
Viewshed analysis fails in the mirror image way. Line of sight computed over a terrain model treats the world as if buildings and forests are transparent, so a mast in a city appears to see the whole city. Computed over a surface model it treats canopy as opaque, which is closer to right for radio and for visual line of sight, but the canopy height in a global radar surface model is an underestimate of true canopy, so the obstruction is systematically too short. Neither product is correct on its own, which is why serious visibility work uses a terrain model plus explicit obstruction layers for buildings and vegetation rather than trusting a single grid to encode everything.
Other silent corruptions follow the same pattern. Slope computed over a surface model is dominated by canopy edges and roof lines rather than by ground gradient, which corrupts erosion, accessibility, and construction suitability analysis. Solar potential is one of the few cases where a surface model is the right choice, since roof geometry is precisely the thing being measured. Drainage and watershed delineation over a surface model produces spurious basins wherever a canopy gap forms an apparent depression.
The rule is simple to state and often skipped. Decide whether the phenomenon interacts with the ground or with the top of things, then pick the model accordingly, then verify from the metadata that the file actually is what its name suggests. Many redistributed elevation files carry names that describe neither their source nor their surface type.
Resolution, voids, and what fills them
Grid spacing is the number people compare, and it is the least informative of the properties that matter.
A finer grid does not mean finer real detail. Many published products are resampled from coarser sources, and a resampled grid has smaller cells containing interpolated values, which is not the same as more information. Vertical accuracy, meaning how close a height is to reality, is independent of horizontal spacing, and a coarse grid with good vertical accuracy is more useful for most hydrology than a fine grid with noisy heights.
Voids are the other half of the story. Radar interferometry fails in specific geometries: steep slopes facing away from the sensor, deep valleys, radar shadow behind mountains, and low return surfaces such as calm water and dry sand. Optical stereo fails under permanent cloud, in shadow, and over textureless surfaces. The result is holes, and the holes are not randomly distributed. They cluster in exactly the terrain that matters most for landslide, avalanche, and mountain hydrology work.
Products handle voids in one of several ways, and the choice propagates into results. Some leave them as a no data value, which is honest and forces the consumer to decide. Some interpolate across them, producing smooth surfaces that are plausible and unmeasured. Some fill them from another source, which is usually the most defensible option but introduces a seam where two datasets with different characteristics meet.
The practical consequences are worth stating.
- A no data value read as a number is catastrophic. Common sentinel values are large negatives, and a slope calculation that treats one as a real height produces a cliff. Always check the declared no data value and mask it explicitly.
- Filled areas should be tracked. If a product ships a source mask indicating which cells were measured and which were filled, that mask belongs in the analysis, not in the discard pile.
- Water bodies are frequently flattened or edited in a post process, so apparent perfect flatness over a lake is an edit, not a measurement.
Vertical datum, and why two sources disagree
Two elevation products can report different heights for the same point and both be correct, because height is meaningless without a stated reference surface.
Two references are common. The ellipsoid is a smooth mathematical approximation of the Earth's shape, and it is what satellite positioning natively measures against. The geoid is a model of the surface of equal gravitational potential that mean sea level approximates, and it is what people mean by height above sea level. The two differ because Earth's mass is unevenly distributed, and the separation between them varies from place to place by amounts that are large enough to matter, reaching tens of meters in some regions.
The practical consequences appear immediately in mixed pipelines. A satellite receiver typically outputs ellipsoidal height. A global elevation product typically publishes orthometric height above a geoid model. Subtracting one from the other without conversion produces an offset that is constant locally and therefore easy to mistake for a real feature or a calibration issue.
There are further complications. Different geoid models exist, and successive versions of the same model differ. National vertical datums are frequently defined against a local tide gauge and do not match global models, sometimes by a meter or more, which matters enormously for flood work where the difference between safe and inundated is often less than that. Tidal datums used in coastal engineering are different again.
The discipline that prevents most of these errors: record the vertical reference of every elevation input alongside the data, refuse to combine heights whose references are unknown, and convert explicitly rather than assuming. When a coastal analysis produces results offset by a suspiciously uniform amount, the vertical datum is the first thing to check.
Derived products inherit every flaw
Almost nothing uses raw elevation directly. The useful outputs are derivatives, and each amplifies a different weakness.
Slope and aspect are computed from differences between neighboring cells, so they amplify noise. A vertical error of a fraction of a meter that is invisible in the elevation map becomes visible speckle in the slope map, and the amplification depends on cell size, so the same terrain yields different slope statistics at different resolutions. Any threshold expressed in degrees is resolution dependent and has to be stated together with the grid it was computed on.
Hillshade is a visualization rather than an analysis product, but it is the best artifact detector available. Processing seams, resampling patterns, void fills, and stripes from the acquisition geometry all become obvious under low angle illumination while staying invisible in a color ramp. Rendering a hillshade before trusting a new elevation file is the cheapest quality check there is.
Hydrology and visibility derivatives are the least forgiving
Flow direction, accumulation, and watershed delineation punish every flaw at once. They assume water can always move downhill, so any local depression, real or artificial, traps flow. Real pipelines therefore run a depression filling or breaching step, which is a substantial modification of the data. Road embankments and bridges cause a specific and common failure: an embankment appears as a continuous barrier because the culvert or the underpass is smaller than a cell, so the modeled drainage diverges from the real drainage at exactly the points where infrastructure crosses water.
Viewshed depends on the observer height, the target height, the earth curvature correction, and an atmospheric refraction assumption, and results change materially with each. A viewshed published without those parameters is not reproducible.
Contours are interpolation. They look authoritative while being a smoothed representation of a grid that was itself a sample of terrain, and their spacing implies a precision the data may not have.
Choosing a source by the question
The instinct to reach for the highest advertised resolution is the wrong default. A better procedure runs through four questions.
First, does the phenomenon interact with the ground or with the top of things? Water, vehicles, and pedestrians interact with the ground. Radio propagation, shadows, wind, and solar exposure interact with the top. This single question decides between terrain and surface, and it decides more than resolution does.
Second, what vertical accuracy does the decision require? Deciding whether a mountain road is steep tolerates several meters of error. Deciding whether a plot floods at a given sea level tolerates far less, and if the required accuracy is below what any global product provides, the honest answer is that the question cannot be answered from global open data and needs local survey or national laser scanning.
Third, is consistency or peak accuracy more important? A study spanning several countries is better served by one consistent global product, even a mediocre one, than by a patchwork of superior national datasets with different datums, vintages, and definitions. Seams between sources create artifacts that look like findings. For a single site, the best local data wins.
Fourth, what does the terrain do to the acquisition method? In steep terrain, expect radar voids and layover. Under dense canopy, expect surface model bias and unreliable terrain estimation. In flat, wet, low lying areas, expect the vertical error of a global product to be comparable to the elevation differences that determine the outcome, which is the situation where global elevation data is least able to answer the question being asked of it.
The closing note is that last case, because it is both the most common and the most ignored. Flood exposure in a river delta is exactly the question people most want to answer with free elevation data, and a delta is exactly where the relief across kilometers can be smaller than the vertical uncertainty of the grid. The correct output there is not a map of who floods. It is a statement that the available data cannot resolve the difference, together with what would be required to resolve it. Publishing the map anyway, because a map is what was asked for, is how elevation data does its quiet damage.