Estimating Economic Activity Without Official Statistics

Estimating Economic Activity Without Official Statistics

E
By Etzal Earth
12 min read

Official economic statistics are excellent and they arrive with three constraints that limit what can be done with them. They are slow, often published quarterly or annually with a lag of months. They are coarse, reported at national or at best provincial level in most of the world. And in a large number of places, for a large number of quantities, they simply do not exist at the granularity anyone needs.

That gap is why proxies exist. Anyone trying to answer whether a district is growing, whether a corridor of investment is producing anything, whether a shock has hit a specific area, or how one town compares to another next year, is going to be working with indirect measurements or working with nothing.

Proxies are useful and they are also where careless analysis produces confident, harmful, and unfalsifiable claims. What follows is an account of the main families, the mechanism that makes each of them work, the point where each one breaks, and the discipline that keeps the output defensible.

Why proxies exist at all

Three distinct gaps drive proxy use, and they call for different things.

Timeliness. A quarterly figure published two months after the quarter ends is months behind the ground when it arrives. For anything responsive, a rougher measurement available weekly beats a precise one available late, provided the roughness is stated.

Granularity. National accounts are constructed for national purposes. Sub national breakdowns, where they exist, are usually apportioned rather than measured, which means the district level detail in an official series may itself be a model. A proxy measured directly at district level is not necessarily worse than an official figure disaggregated by assumption.

Absence. Informal economic activity is by construction underrepresented in official statistics, and in many economies it is the majority of activity. No amount of statistical refinement recovers what was never measured. Proxies that observe physical activity see some of what surveys do not.

Which gap is being filled matters, because a proxy that is good for timeliness may be poor for granularity and useless for capturing informality.

Night lights

Satellite measurements of artificial light at night are the most used economic proxy, because the mechanism is intuitive: economic activity consumes energy, and a portion of that energy leaves as light visible from orbit. Lit area and light intensity track settlement, industry, transport infrastructure, and commerce.

The proxy works best for large changes over meaningful time spans. Electrification of a previously unlit area is unmistakable. Industrial development along a corridor shows up. Conflict, mass displacement, and severe economic collapse produce clear darkening. These are strong signals and they are hard to fake.

The biases are well documented in the field and they are structural.

  • Saturation. Bright urban cores reach the top of the sensor range and further growth produces no additional signal, so intensification in an already bright center is invisible.
  • Blooming and scatter. Light spreads in the atmosphere and across the sensor, so bright sources appear larger than they are and dim sources near bright ones are swallowed.
  • Composition. Light depends on what kind of activity, not just how much. A logistics yard with floodlights outshines a far more valuable office. A gas flare outshines a city. Agriculture emits almost nothing.
  • Technology change. More efficient lighting changes emitted light per unit of activity and changes the spectrum in ways that affect what a sensor records. A place can get richer and darker at once.
  • Sensor changes. Long time series stitched from different instruments carry discontinuities that look like economic events.

The rule that follows is that night lights are a measure of lighting, which correlates with activity under conditions that vary by place and by decade. Used for change within one place over a few years, they are informative. Used to rank the wealth of two different countries, they are measuring lighting policy as much as economics.

Built-up growth

Construction is capital formation made visible. Detecting new built up area from imagery, or growth in building footprint area and height, measures investment that has been physically realized.

The mechanism is direct and the lag is short: buildings appear when they appear, and imagery revisits frequently enough to catch them within weeks or months. Building growth captures a form of activity that night lights miss entirely, since a new unlit warehouse is invisible to one and obvious to the other.

Its limits are equally direct. Construction is lumpy and lagging: it reflects investment decisions made years earlier, so a burst of building can coincide with a downturn already underway. It says nothing about occupancy, and vacant construction is common enough in speculative markets that built up growth can signal misallocation rather than activity. It misses the intensification of use inside existing buildings, which is where much growth happens in mature areas. And detection is sensitive to imagery conditions, so a change between two dates may be cloud, season, or sensor rather than the world.

The productive use is as a complement rather than a substitute. Built up growth with night light growth suggests occupied new development. Built up growth without it suggests something built and not yet in use, which is itself a finding worth reporting as a finding rather than smoothing into an index.

Business density and churn

Counts of businesses, and more importantly changes in those counts, come from points of interest datasets, open business registries where they are published, and map contributions. The mechanism is that commercial activity requires premises, and premises get recorded.

Density of businesses per unit area, or per resident, indicates commercial intensity. Churn, the rate at which businesses appear and disappear, is often more informative: high churn with stable counts suggests a competitive, fluid local economy, while a sharp rise in closures without corresponding openings is an early distress signal that no official statistic will report for months.

The biases here are more severe than for the remote sensing proxies, and they are the biases of who records data rather than of physics.

  • Recording is voluntary in open map data and commercially motivated elsewhere. A district full of businesses that do not advertise online appears empty.
  • Closure is recorded far less reliably than opening. Nobody has an incentive to update a listing for a business that no longer exists, so counts drift upward regardless of reality and measured churn is asymmetric.
  • Formality bias. Registries capture registered businesses, so in economies with substantial informal commerce the count measures formalization as much as activity, and a rise can reflect a registration drive.
  • Category inconsistency. Classification schemes differ across sources and change over time, so a shift in the mix can be a taxonomy artifact.

Business data needs the coverage discipline most, because its gaps correlate directly with the economic characteristics being estimated.

Movement: transport and shipping

Vehicle, vessel, and aircraft movement are measurements of economic exchange in progress. Vessel tracking through automatic identification systems shows port calls, dwell times, and route volumes. Aircraft tracking shows passenger and freight connectivity. Road traffic, where sensed, shows commercial vehicle activity.

Movement data has a property the other proxies lack: it is close to real time and individually verifiable, in the sense that a specific ship arriving at a specific berth is a discrete, checkable event rather than a statistical inference. Port throughput estimated from vessel calls, weighted by capacity and dwell, is a strong indicator of trade activity.

Its limits are about coverage and about what it does not see. Reception depends on receiver networks, which are dense near wealthy coastlines and sparse elsewhere, so apparent traffic partly maps the receivers. Transmission can be switched off. Vessel capacity is not cargo carried, and estimating load state from draft observations is itself a model. And the family is blind to domestic overland trade in places without instrumented roads, which is most places.

Movement proxies are strongest for tradeable, high value, infrastructure dependent activity, and weakest for exactly the local informal economy that the other proxies also miss.

Change over time beats comparison across places

The single most important discipline in proxy work: these measures are far more reliable for tracking one place against its own past than for comparing two places against each other.

The reason is that every bias listed above is roughly constant within one place over a short period and wildly different between places. If a district's lighting technology, sensor coverage, mapping culture, formalization rate, and economic composition are approximately stable, a change in the proxy is mostly a change in the underlying activity. Across two districts all of those differ, and the difference in the proxy mixes real difference with bias difference in a way that cannot be separated.

This shapes how results should be phrased. Activity in this district rose over three years on a lighting proxy is a defensible claim with caveats. This district is more economically active than that one is not, and no methodological sophistication makes it so.

The corollary is that even within one place, the proxy measures a rate of change and not a level. There is no reliable conversion from light output to currency, and any pipeline that produces an estimated GDP figure from a satellite image has introduced a calibration that is doing all of the work and deserves all of the scrutiny.

Combining proxies without laundering the uncertainty

Multiple independent proxies pointing the same direction is worth more than any one of them. It is not proof, and the difference matters.

Two proxies agree either because both are tracking the underlying activity or because they share a bias. Night lights and built up growth are not independent: both respond to construction, both are remote sensed, both are affected by the same atmospheric and seasonal conditions. Business counts and points of interest density from the same underlying map are not independent at all. Genuine independence means different mechanisms and different failure modes, such as a remote sensed measure alongside a movement measure alongside an administrative record.

The construction to avoid is the composite index that averages several proxies into a single number. It feels rigorous and it destroys the only useful information in the ensemble, which is the disagreement. When two proxies diverge, that divergence is a finding: usually one mechanism has broken, or something structural is happening, such as building without occupancy or activity without formalization. Averaging turns the most informative signal in the dataset into noise around a mean.

Report the ensemble as an ensemble. Each proxy, its direction, its magnitude, its known biases in this context, and an explicit statement of whether they agree. Where they agree, call it consistent evidence. Where they disagree, say which one is more likely to be reliable here and why, and if that cannot be said, report the disagreement and stop.

The ethics of inferring wealth

Estimating economic activity is estimating how well people are doing, at a resolution fine enough to identify neighborhoods and sometimes buildings. The uses are not neutral.

Estimates like these feed decisions about where infrastructure goes, where aid is targeted, where credit is priced, where insurance is offered, and where enforcement concentrates. Being wrong in one direction denies a place resources it needs. Being wrong in the other attracts attention a community did not ask for. Both errors fall disproportionately on places with the least data, which are also the places least able to contest a number produced elsewhere.

Three practices reduce the harm without pretending it away. Do not publish estimates at a resolution finer than the method supports: a model calibrated on district level data does not become a building level model by being rendered at building level. State the direction of the error, not just its size, because if the method systematically underestimates informal activity then a low estimate in an informal economy is more likely a measurement failure than a finding. And give the subject a route to contest it, which means an estimate about a place should be explainable to the people who live there in terms of what was observed rather than in terms of a model weight.

Validating against official data to calibrate where there is none

Official statistics are the ground truth for a proxy, and the point of validating against them is not to reproduce them. If the proxy only worked where official figures exist, it would be redundant.

The purpose is calibration transfer. Fit the relationship between the proxy and the official measure in places where both exist, characterize the residual error, understand what makes the relationship vary, and then apply the relationship where only the proxy exists, carrying the residual uncertainty with it.

That transfer is valid only to the extent that the target place resembles the calibration set in the ways that matter. Calibrating a lighting to activity relationship on urban districts in one country and applying it to rural districts in another is transferring a relationship across two dimensions it was never tested on. The honest procedure states the calibration domain and flags every application outside it.

Practically, this means holding out places for validation rather than fitting everywhere, testing across the boundaries expected to matter such as urban and rural or formal and informal, and re-validating whenever new official data lands. It also means accepting a poor result. A proxy that does not track official figures where they exist has no claim on credibility where they do not, however plausible its mechanism.

The estimate that should not be shipped

The output that causes damage is the single currency figure per small area, published with no interval, no method note, and no coverage field, because it is the output that looks most like an official statistic and is trusted like one.

The version worth shipping is a direction, a magnitude with an interval, the proxies that produced it, their disagreement where it exists, the calibration domain, and a list of the activity the method cannot see. It is harder to chart. It is also the only form in which someone downstream can tell whether the number applies to their question, which is the entire purpose of producing it.