Most people meet free satellite imagery through a screenshot: a green field, a brown scar where a forest used to be, a city glowing in false color. The screenshot is the end of a long chain of decisions, and almost every mistake made with open Earth observation data happens somewhere in that chain rather than in the final picture.
Sentinel-2 is the workhorse of that chain. It is an optical imaging mission in the European Copernicus program, its data is free and openly licensed, and it is the default answer whenever someone asks where to get repeated pictures of the same piece of ground without paying for them. It is also routinely misunderstood, because the phrase "satellite imagery" makes people think of the imagery they have seen in mapping apps, which is aerial photography and very high resolution commercial capture, not this.
The gap between those two things is where the trouble starts, and it is worth closing before opening a scene.
What the mission actually provides
The mission consists of more than one identical satellite flying in the same orbital plane, offset so the ground is revisited more often than a single spacecraft could manage. Each carries a multispectral imager that sweeps a wide swath beneath it and records reflected sunlight in a set of separate wavelength ranges, called bands. Three consequences follow from that design.
It is passive and optical. It measures reflected sunlight, so it sees nothing at night and nothing through cloud. Radar missions do not share that limitation, which is why serious monitoring pipelines pair optical and radar rather than choosing one.
It is wide swath and moderate resolution. The design trades detail for coverage and frequency, and the trade is the whole point: the value is not in any single scene, it is in having a comparable scene of the same place over and over for years.
It is systematically acquired. The satellites are not tasked by customers pointing them at targets. They image continuously along their orbit to a fixed plan, so the archive is dense and uniform rather than clustered around whoever was paying attention. For change detection that matters more than resolution.
Why more than visible light matters
A normal camera records three broad bands that approximate human vision. This imager records those and several more, spread across the near infrared and the shortwave infrared, plus a handful of narrow bands positioned specifically to characterize the atmosphere rather than the ground.
Healthy vegetation is the clearest case. Chlorophyll absorbs strongly in the red and the internal structure of leaf tissue scatters near infrared very efficiently, so a vigorous leaf is dark in red and extremely bright in near infrared, a contrast with no equivalent in visible color. A stressed, dying, or harvested plant loses that contrast long before it looks different in a true color image, which is why vegetation monitoring runs on infrared rather than on green pixels.
Water behaves in the opposite direction. It absorbs near infrared and shortwave infrared almost completely, so open water goes nearly black in those bands while staying ambiguous in visible color, where a dark roof, a shadow, and a pond are hard to tell apart. Flood mapping leans on this. Bare soil, concrete, and asphalt reflect more strongly in the shortwave infrared than vegetation does, which is the basis for separating built surfaces from green ones. The narrow bands aimed at water vapor and cirrus signatures are there for a different reason again: they give atmospheric correction something to work with.
Resolution, and what a pixel really contains
Not all bands share the same ground sampling. The visible and near infrared bands are the sharpest, the red edge and shortwave infrared bands come on a coarser grid, and the atmospheric bands are coarser still. Any workflow that stacks bands has to resample to a common grid, and the direction of that resampling is a real decision: upsampling a coarse band does not create detail, it only makes the array shapes agree.
What matters more than the number is what a pixel is. A pixel is not a small photograph. It is one integrated measurement of everything within its footprint, and its value is the mixture. A pixel straddling a road, a hedge, and a field returns a number that belongs to none of them. This is the mixed pixel problem, the largest source of naive error in moderate resolution analysis, and it has practical consequences:
- Linear features narrower than a pixel do not disappear, they contaminate. A path through a forest will not appear as a path, it will shift the values of the pixels it crosses, and people then read the shift as a change in the forest.
- Small objects cannot be counted. If a thing is smaller than the footprint, its presence moves a number but never creates a countable shape.
- Boundaries are fuzzy by construction. A field edge is a gradient a pixel or two wide, so any area statistic computed by counting pixels inside a polygon carries an error term proportional to that polygon's perimeter. Small parcels are measured proportionally worse than large ones.
- Geolocation is good but not perfect. Between two dates the same feature can sit fractionally differently, and for per pixel time series over small objects that shift can dominate the signal.
Revisit cadence and the clouds that eat it
The revisit interval is the headline people quote: the constellation returns to the same ground track every few days at mid latitudes, and more often at high latitudes and in the overlaps between adjacent swaths. That describes acquisitions. It does not describe usable observations.
Cloud is the dominant filter. Over a persistently cloudy tropical region, months can pass without a clear view of a given field. Over an arid region the same nominal cadence makes almost every scene usable. Effective revisit is a property of the place and the season rather than of the satellite, and measuring it is the first step before promising anyone a monitoring product.
The second order effect catches people repeatedly. Cloud is not randomly distributed in time. It correlates with the rainy season and with monsoon onset, meaning exactly the periods that tend to be most interesting. A time series built only from clear observations is a biased sample of the year that overrepresents dry conditions. If a vegetation index appears to drop every wet season, the first hypothesis should be the sampling, not the vegetation.
Cadence also interacts with the thing being observed. A harvest that happens between two clear observations weeks apart is invisible as an event and appears only as a state change, with no way to date it precisely from optical data alone. Any claim about when something happened carries an uncertainty at least as wide as the gap between usable scenes, and honest reporting states that interval rather than the date of the later scene.
Cloud masking is a probability, not a fact
Every serious pipeline masks cloud. Products ship a scene classification layer labeling pixels as cloud, cloud shadow, snow, water, vegetation, bare soil and so on, and separate cloud probability layers exist too. None are exact, and their failure modes are specific enough to plan for.
Thin cirrus is the hardest case. It attenuates and scatters without producing an obvious white patch, so it often passes the mask while still shifting reflectance values. A subtle change across a whole scene, in the same direction for every land cover type, is usually undetected cirrus.
Bright surfaces get confused with cloud. Fresh snow, salt flats, bright sand, white roofs, and greenhouses are misclassified regularly, so in an industrial area full of metal roofing, aggressive masking deletes exactly the pixels of interest, which then reads as missing data rather than as bias.
Cloud shadow is worse than cloud. Shadow detection works by projecting cloud geometry using sun angle and an assumed cloud height, and the assumption is often wrong. Unmasked shadow darkens pixels and drags vegetation indices down, producing false detections of stress or clearing. When a change map shows narrow curved features unrelated to anything in the landscape, look for shadow. Masks also need dilating, since pixels beside a cloud are contaminated by adjacency, and buffering by a few pixels costs data while buying reliability.
Treat a mask as a layer of uncertainty, not a cleanup step. A useful pipeline records, for every pixel and date, whether the value was observed clear, observed and flagged, or absent, and carries that record into the final output. Once a masked value has been silently replaced with an interpolated one, no downstream consumer can tell measurement from guess.
Raw and corrected products are different measurements
Two product levels matter for beginners, and confusing them produces results that look fine and are not comparable.
The first is top of atmosphere reflectance: what the instrument measured, converted to a physical reflectance value and orthorectified so the pixels land in the right place. It includes everything the atmosphere did to the light on the way down and back up, meaning scattering by air molecules, absorption by water vapor and ozone, and haze from aerosols. The second is surface reflectance, sometimes called bottom of atmosphere, where a correction model estimates the atmospheric contribution at the moment of acquisition and removes it, approximating what a sensor at ground level would have measured.
For a single pretty image the distinction does not matter. For anything comparative it matters enormously, because atmospheric effects vary between dates, between seasons, and across a single wide scene. Two top of atmosphere scenes of the same field a month apart differ partly because the field changed and partly because the air did. The correction is an estimate rather than a truth, and it degrades over bright surfaces, over water, under heavy aerosol load, and in steep terrain, but an imperfect correction is still far better than none for time series work.
Terrain adds another layer. On a slope, incoming sunlight depends on the angle between the surface and the sun, so north and south facing slopes with identical vegetation return different reflectance. Where relief is significant, topographic correction belongs in the pipeline, otherwise the map of vegetation vigor is partly a map of aspect.
Indices, explained by mechanism rather than by formula
Spectral indices are simple arithmetic combinations of bands, popular because they are cheap, interpretable, and partly self correcting: a ratio cancels some illumination effects that scale all bands together.
The vegetation family contrasts a band where chlorophyll absorbs against a band where leaf structure scatters, so high values mean a lot of photosynthetically active canopy per unit area. The mechanism explains the limits directly. The index saturates in dense canopy, because once the ground is fully covered, more leaves barely change the contrast, which is why it separates bare from vegetated far better than good forest from great forest. It is also sensitive to soil brightness in sparse cover, where the background dominates the signal, which is why arid land work uses variants designed to suppress the soil line.
The water family contrasts a visible band against a band that water absorbs. It works well for open water and badly for shallow, turbid, or vegetated water, where sediment and emergent plants push the values back toward land.
The built up family contrasts shortwave infrared against near infrared, because impervious surfaces are relatively brighter in the shortwave. It is genuinely useful and genuinely blunt: bare soil, dry ground, and quarries look built up, and a freshly cleared construction site scores much like a finished one. Anyone using such an index as a proxy for urbanization needs to check what fraction of the high values are simply dry earth.
Every index is a hypothesis about which physical property dominates the difference between two bands. When something else dominates, the index reports it without complaint. Knowing the mechanism tells you the failure mode; memorizing the formula does not.
A worked example: did this field get harvested
Take a practical question: one field, and the need to know whether it was harvested in the last month. The naive approach takes the latest scene, computes a vegetation index, sees a low value, and calls it harvest. That is unsupported, because a low vegetation index has several causes: harvest, drought failure, flooding, undetected cloud shadow, or a scene whose atmospheric correction failed.
A defensible approach differs in structure rather than in cleverness. Pull every acquisition over the field for the season. Apply cloud and shadow masks with a buffer, and record how many clear observations survive and where they fall in time. Aggregate to the field polygon eroded inward by a pixel or two, so boundary pixels are excluded. Read the shape of the trajectory rather than one value: harvest drops abruptly from a plateau to a low, stable level, drought stress declines slowly, and flooding shows a simultaneous jump in the water index that harvest does not. Then state the answer with the date range imposed by the gap between the last high observation and the first low one, plus how many clear observations supported it. That is the same arithmetic with the uncertainty kept visible, which is usually what separates analysis that survives scrutiny from analysis that does not.
The limits, stated plainly
Resolution sets a hard floor and no amount of processing raises it. Cars cannot be counted. Signs, license plates, and people are not merely hard to see, they are absent from the measurement. Individual small buildings in dense settlements are not reliably separable. Two similar crops at the same growth stage often differ by less than the noise, which is why crop type classification depends on the shape of a whole season rather than on any single scene.
Pansharpening and super resolution models produce images that look sharper. They do not add information that was never recorded. A generative upsampler invents plausible detail, and plausible detail presented as observation is fabrication.
What the mission gives instead is consistency over time, global coverage, and a spectral range that makes physical properties measurable rather than merely visible. The right instinct when facing a spatial problem is therefore not to ask for the highest resolution available, but to ask which physical property would distinguish the outcomes, whether any band is sensitive to it, and how often the sky over that place is clear enough to look. If the answer to the last question is "rarely", the honest response is to change methods or widen the stated uncertainty, not to publish the one clear scene that happened to be available.