AIS and ADS-B: Tracking Ships and Aircraft with Open Feeds

AIS and ADS-B: Tracking Ships and Aircraft with Open Feeds

E
By Etzal Earth
12 min read

A live map of ships and aircraft is the most convincing thing in open geospatial data. Thousands of icons move in real time, each with an identity, a speed, a heading, a destination. It feels like watching the world directly.

It is not that. Both systems behind those maps are cooperative: the vehicle broadcasts its own position and identity, a receiver hears the broadcast, and a network shares it. Nothing on those maps is a detection in the radar sense. Every icon is a claim made by the thing it represents, relayed by whoever happened to be listening.

That single property drives every analytical mistake made with this data, and the most damaging one is simple. An empty area on the map means nothing was heard there. It does not mean nothing was there.

What the two systems are

AIS, the Automatic Identification System, is a maritime broadcast system. Vessels transmit on VHF marine frequencies, and other vessels and shore stations receive them. It was designed for collision avoidance and traffic management, meaning its purpose is for a ship to be visible to the ships and authorities around it, not to be visible to the world. Carriage is mandatory for certain classes of vessel under international rules, with lighter requirements for smaller craft. Class A transponders, carried by larger commercial vessels, transmit more often and carry more information than the Class B units common on smaller boats.

ADS-B, Automatic Dependent Surveillance Broadcast, is the aviation equivalent. An aircraft determines its own position from satellite navigation and broadcasts it periodically, along with identity and status. Two link technologies are in general use, one on the frequency shared with existing secondary radar transponders and used worldwide, and another on a separate frequency used mainly for general aviation in the United States. The name is worth reading literally: automatic, meaning it transmits without being interrogated; dependent, meaning it depends on the aircraft's own navigation solution; and broadcast, meaning it is sent to anyone in range rather than to a specific recipient.

Both are unencrypted and unauthenticated. That was a design choice appropriate to their purpose, and it has two consequences. Anyone with a cheap radio receiver can decode them, which is why large volunteer networks exist and why open feeds are available at all. And nothing in the protocol verifies that a transmitted position or identity is true.

Cooperative means absence proves nothing

A radar detects an object because energy bounces off it. A cooperative system knows about an object because the object announced itself. The distinction sounds academic until it decides an analysis.

Reasons a vehicle can be genuinely present and absent from the feed:

  • It is not required to carry a transponder. Small fishing boats, pleasure craft, many general aviation aircraft, and vessels below the size thresholds are simply outside the mandate.
  • The equipment is switched off, failed, or misconfigured. Transponders fail like any other electronics, and installation errors that produce a wrong identity code or no position are common enough to be a known category.
  • The operator turned it off deliberately. This happens for reasons ranging from security, in areas with piracy risk, to sanction evasion and illegal fishing.
  • Military and state aircraft and vessels are frequently exempt, and are also filtered out by many public services even when they do broadcast.
  • Nobody was listening. This is the largest category by far, and it is the one that gets forgotten.

An absence in this data is therefore ambiguous across at least five explanations, and only one of them is "no traffic". Any pipeline that converts an empty result into a count of zero has thrown away that ambiguity at the first step.

Coverage geometry decides what you see

Both systems are line of sight radio. That fact, plus a small amount of geometry, explains most of the shape of the coverage map.

For aircraft, the horizon distance from a ground receiver grows with the square root of altitude, so a high altitude aircraft can be heard from a long way off while an aircraft on approach vanishes behind terrain and buildings well before it lands. This is why coverage of cruise traffic looks excellent while low altitude coverage is patchy, and why a receiver on a hill is worth several in a valley. Terrain shadowing produces persistent wedges of no coverage that stay in the same place forever, and any map of receiver density inherits them.

For ships, the geometry is worse in one specific way: both the transmitter and the receiver are near sea level. Terrestrial AIS reception is therefore limited to a coastal band, with range varying with antenna height and atmospheric conditions. Beyond that band, the ocean is dark unless someone is listening from space.

Satellite reception changes the picture and introduces its own distortions. A satellite sees a huge area at once, which is exactly the problem: in a busy shipping lane, many vessels transmit on the same channels within the satellite's footprint and their messages collide, so the busiest waters are where satellite reception degrades most. Coverage also comes in passes rather than continuously, so a vessel's satellite track is a series of samples with gaps between them, and the gaps get read as loitering or as diversion when they are only orbital mechanics.

The consequence for analysis is that observed density is a product of real density and receiver density, and the second term varies by orders of magnitude between a European coastline and a remote ocean. Comparing traffic between two regions without normalizing for coverage compares receiver networks, not traffic.

What is measured and what is asserted

Both systems mix two very different classes of information in the same message stream, and treating them as equally reliable is a persistent error.

Derived from the vehicle's own navigation system:

  • Position, from satellite navigation.
  • Speed over ground and course over ground, derived from successive positions.
  • For ships, heading from a compass where one is interfaced, which differs from course over ground whenever there is current or wind.
  • For aircraft, barometric altitude, geometric altitude, and vertical rate. Barometric altitude depends on a pressure setting and is not the same as height above the ground or above the ellipsoid, which catches people who compare it to terrain data directly.

Typed in by a human or configured once at installation:

  • For ships, the destination field, the estimated time of arrival, the vessel name, the reported draught, and the cargo or vessel type.
  • For aircraft, the callsign entered by the crew.

The self reported fields are exactly the ones analysts most want, and they are the least reliable. Destination fields contain abbreviations, informal port names, old entries never updated after the last voyage, and occasional jokes. Draught is set manually and is often left at whatever it was after the last loading. Vessel type is a coded field, and codes get chosen carelessly.

There is a third category that matters: fields derived by the aggregator rather than the vehicle. Many feeds add a resolved vessel or aircraft registry entry, an inferred flag state, a normalized destination, an interpolated position between reports, or a computed "last seen" status. These are useful and they are inferences, and they should never be presented with the same confidence as a received position. An interpolated point on a map looks identical to an observed one unless the interface distinguishes them, and users will not assume the distinction exists.

The static identifiers deserve their own note. Ships are identified by an MMSI, which is assigned by administrations and can be reassigned when a vessel changes flag, and separately by an IMO number that is intended to persist for the hull's life. Aircraft carry a fixed address assigned to the airframe, which persists across registration changes in some jurisdictions and not others. Building a history on the wrong identifier merges two vehicles or splits one.

Spoofing, gaps, and other deliberate behavior

Because nothing is authenticated, the content of a message is only as trustworthy as the sender chooses to make it, and manipulation is well documented as a category.

Identity manipulation. Transmitting another vessel's identifier, or one that belongs to no vessel, is straightforward. The result is a track that looks ordinary and belongs to a fiction. Two positions reported for the same identifier at the same time, far apart, is the simplest detector and it works.

Position manipulation. Broadcasting a false position produces a vessel that appears to be somewhere it is not. Some cases are crude, with tracks that jump, cross land, or move faster than the vessel type allows. Others are consistent enough to require kinematic checks against plausible speed, turn rate, and the physical navigability of the water.

Interference with the navigation source. Satellite navigation jamming and spoofing occur in various regions, and their effect on a cooperative system is direct: the transponder reports what its receiver tells it, faithfully broadcasting a wrong position. A characteristic symptom is many vehicles in one area reporting positions arranged in an implausible pattern, such as a tight circle, or all reporting the same position.

Intentional gaps. Turning a transponder off produces a hole in a track. Gaps are a well known signal in maritime analysis, associated with illicit transfers between vessels and with fishing in closed areas, and they are also produced by ordinary equipment failure, by leaving coverage, and by satellite pass timing. A gap is evidence that something should be looked at. It is not evidence of what happened.

The general rule for all of these: cross checks come from physics and from independent data. A track that is impossible for the vehicle type is wrong regardless of how well formed the messages are. A vessel reported in a port with no berth for its size is wrong. A track that satellite imagery contradicts is wrong. None of this can be resolved inside the feed itself, because the feed only knows what was said.

Deduplication across receiver networks

Open aircraft and ship data mostly comes from community networks: volunteers running receivers who share what they hear. This produces excellent coverage in populated areas and creates a specific engineering problem, because a single transmission is heard by many receivers and arrives many times.

Deduplication at the message level uses the vehicle identifier plus the message content and timestamp. The complications are practical. Receiver clocks disagree, so the same transmission arrives with different timestamps and a tolerance window is needed. The window has to be wider than clock skew and narrower than the reporting interval, which is a real tension when reporting intervals are short. Some networks forward messages with the time of relay rather than the time of reception, which shifts everything by an unknown amount.

Aggregation across networks is harder than within one. Different providers apply their own filtering, smoothing, and interpolation before publishing, so the same underlying transmission emerges with slightly different values from two sources. Naive merging then produces a track that zigzags between two versions of the truth. Preferring one source per vehicle per time window, and switching sources only at explicit boundaries, produces cleaner tracks than blending.

Counting is where deduplication failures become visible. If the question is how many aircraft were over a region, and the pipeline counts messages or counts source records rather than distinct vehicles, the answer scales with receiver density. Regions with many volunteers appear busier. The correct unit is always the distinct vehicle over a defined time window, and the window has to be stated, because the number of distinct vehicles in an hour and in a day differ enormously.

A related trap is the retained ghost. A vehicle heard once and never again will sit on a map indefinitely unless an expiry rule removes it. Expiry has to differ by context: an aircraft not heard for several minutes has probably left coverage or landed, while a vessel not heard for hours may simply be beyond terrestrial range and awaiting a satellite pass. A single global timeout produces either a map full of ghosts or a map that erases real vessels.

The trap, stated directly

The analytical failure that this data produces most often is the inference from a sparse map to a quiet place.

It appears in commercial work as an underestimate of activity at ports and airfields in regions with thin receiver coverage. It appears in environmental work as an underestimate of fishing effort where small vessels are not required to transmit. It appears in security work as the assumption that a clear stretch of water is empty, when what is clear is the listening.

The defense is procedural rather than clever. Every count derived from cooperative broadcast data should carry, alongside the number, a statement of the coverage from which it was derived and the classes of vehicle that are not obliged to appear in it. Where a coverage model exists, whether from receiver locations, from terrain, or from observed detection rates for vehicles known to be present, the count should be qualified by it. Where no coverage model exists, the honest output is a lower bound, described as a lower bound.

Stating a lower bound feels weaker than stating a number. It is the stronger position, because it is the one that survives the first time someone points out the ship that was there and never appeared on the map.