Live public-transit positions & on-time performance, tracked over time.
Pick one or more cities above.
Each city's grade is its on-time performance: 90%+ = A, 80% = B, 70% = C, 60% = D, below = F. The "Issues" column separately flags feed problems (e.g. feed down). A city with no gradeable data is left ungraded — never given a letter. Methodology
Loading…
Loading…
How evenly vehicles are spaced on each route — the typical observed gap between vehicles and how variable it is. Approximate, reconstructed from live vehicle positions; both directions combined. Methodology
Loading live data… (a cold start can take up to ~20 s)
How often transit actually runs on time, ranked by route, from accumulated public GTFS-Realtime history. On-time = within the configured band of schedule. Methodology
Loading live data… (a cold start can take up to ~20 s)
How Global Transit Stats turns public GTFS-Realtime feeds into the numbers on this site — and the limits of those numbers.
We read each city's public GTFS-Realtime feed (the open standard agencies publish for live vehicle positions and, sometimes, trip delays) plus their static GTFS schedule. We do not use any private or paid feed. A city is only as good as its agency's public data.
Most transit agencies publish a timetable (static GTFS); a smaller set also publish live vehicle positions. Cities labelled "Timetable only" are ones whose agency publishes no live tracking — we show their routes, stops and today's schedule, but there are no moving dots and no on-time statistics (you can't grade punctuality without observing vehicles). This is a fact about the agency's data publication, not a gap in the product: the moment an agency starts publishing a live feed, the city is eligible for the live tier. A live-tracked city with zero vehicles in service right now (overnight, holidays) temporarily leaves the live list and returns with its next vehicle — the list reflects what is verifiably moving now.
Each tracked vehicle's delay is sampled continuously: an observation is recorded whenever its delay moves by ≥30 s, plus a 5-minute heartbeat while it holds steady (so stable and volatile trips are both represented without duplicate readings). Each recorded observation is graded on time, late, or early from its delay in seconds (on-time band configurable; default ±2 min). When a feed carries no explicit delay, we derive one by comparing the observed position/time against the static schedule ("schedule-diff"). Feeds whose realtime trip IDs don't match their static schedule show "schedule data pending" rather than a fabricated number.
The on-time percentage for any window is: on-time % = n_on_time ÷ n_obs, computed only on observed data. There is no imputation of missing hours.
Numbers are always computed for a window (5 min / 1 h / 24 h / 7 / 30 days / 3 / 6 / 9 months / 1 / 5 years / all time). Short windows read raw observations; long windows read daily rollups. A fixed window shows only cities with enough history to cover it (shorter-history cities are listed separately, never blank); all time uses each city's full history, so it can mix cities of different ages. History accumulates forward from our first poll — it cannot be backfilled, so long windows are empty until enough time passes.
Daily rollups use each observation's agency-local calendar date, not an exact GTFS service date. Tracked limitation: overnight trips with stop times beyond 24:00 remain on the local calendar date because an exact trip service date is not currently derived.
National feeds (e.g. the Netherlands, Norway) report vehicles far beyond one city. "City" scope bounds results to the city's geographic box; "network" scope includes the whole feed.
Storage change, 26 Jul 2026: cities sharing one national feed used to record the shared network's observations once per city; since 26 Jul 2026 those shared observations are recorded once in total. Network-scope observation counts for windows spanning that date therefore step down — the removed rows were duplicate copies of the same observations, so percentages (on-time %, medians, percentiles) are unaffected, and city-scope numbers are unchanged entirely.
For a small set of high-frequency networks we also measure headway — the gap between consecutive vehicles on the same route — and flag bunching (two vehicles arriving nearly together, which riders experience as a doubled wait). Headways are derived from the same live positions as everything else, per route and direction, and compared against the scheduled gap where a service calendar exists. It's computed for a deliberately small city set because it is expensive to derive; the directory lists which cities currently have it.
Every on-time percentage is a sample, so we publish it with a 95% Wilson score interval (the "±" you see beside scores). Wilson intervals stay honest at small sample sizes and extreme rates where the usual normal approximation breaks: a rate from 40 observations shows visibly wider bounds than one from 4,000. Leaderboards rank by the conservative lower bound of that interval, not the point estimate — so a route or city with a lucky-but-thin sample cannot out-rank one that is well measured. "Lowest measured on-time" lists symmetrically use the upper bound (confidently low, not unluckily sampled).
Missing data biases a score while leaving it looking confident: if we observed only half a window — whether our monitoring was down or the agency's feed was — the surviving observations may not represent the whole service. Every score therefore carries a data coverage figure: healthy monitoring probes observed ÷ probes expected for the window (calibrated against gold-standard cities, so ~100% is achievable in practice). A city below 50% coverage is marked PROVISIONAL · NOT RANKED — its numbers stay visible, but it is excluded from leaderboards and cross-city rankings rather than flattering or punishing the board on partial data. Below ~33% coverage we withhold the letter grade entirely.
The problem with a plain average. Suppose a city's feed is down every weekday rush hour for a month — exactly when trains run late. The observations we do collect (off-peak, weekends) look fine: say 85% on-time. Report that as the city's score and it's not just misleading — it's systematically misleading in the best possible direction. The gap doesn't look like a gap; it looks like a data point.
This is not hypothetical. Feed outages correlate with congested periods (high vehicle counts stress agency systems). Our gaps and the agency's gaps are not random noise — they cluster at the worst moments.
What we do instead. Every on-time figure is paired with a data-completeness ratio. Below 50% coverage the score is marked provisional and excluded from rankings — a number from half the window can't fairly represent the full window. The hourly-stratification step (see below) goes further: it reweights observed hours by how much service the schedule planned for each hour, so a gap at 8 am counts differently than a gap at 3 am.
Systemic outages are recorded explicitly. When our poller is down, every city is affected at once. That event is logged; affected report cards show the annotation so readers can see whose gap it was, not just that a gap exists.
When an hour of a city's data is missing we attribute it: "ours" means our pipeline wasn't looking (poller, probes, or host down — a systemic outage hits every city at once and is recorded in an explicit outage log you'll see annotated on affected report cards); "source" means we were probing and the agency's feed itself was down. Both reduce coverage identically — a gap is a gap — but the attribution is published so you can tell whose gap it was. Feed up with nothing to observe (e.g. no overnight service) is neither: it costs no coverage. Attribution is deliberately conservative against us: any gap we cannot positively attribute — including hours that predate our probe monitoring (before 24 Jun 2026) — is counted as ours, never as the agency's.
A plain on-time % treats every observation equally, so a monitoring gap at rush hour quietly removes the hardest hours from the average. The coverage-adjusted figure on report cards recomputes the score by local hour of day and recombines the hours weighting each by scheduled service (how much service the agency planned that hour, from its published schedule — falling back to the city's own long-run hourly profile where no service calendar exists). An hour observed on some days stands in for its missing days within the same hour; an hour we never observed is excluded and reported as "unobserved scheduled service" with the ours/source split — never silently averaged away. This hourly history covers our full collection span (rebuilt from the raw archive back to 17 May 2026); the ours/source attribution is probe-verified from 24 Jun 2026 onward and conservative (counted as ours) before that.
Every ~2 minutes we record whether each city's feed is responding and how fresh it is. The daily data-quality grade (A–F) combines feed uptime, freshness, and whether on-time data is gradeable, using weights and bands defined in configuration. Each grade carries a confidence (how much data we had) and a list of issues. A grade measures the public data, not the agency; missing data can make good service invisible. A city with no data gets no grade — we never invent one.
Where a city's schedule includes a service calendar and its realtime trip IDs match it, we estimate the share of scheduled trips that were actually observed. This is inference, not ground truth: a feed outage can look like cancellations, so each figure carries a confidence tied to that day's feed uptime, and feeds whose IDs don't match their schedule are marked unsupported rather than given a misleading number.
Illustrative example — not real data. Numbers are chosen to make the logic clear.
Setup. Imagine a city — call it Exampleville — with a transit feed we've been polling for 30 days. We're computing its on-time percentage for the most recent 7-day window.
Step 1: collect observations. During those 7 days our poller collects observations continuously. Each observation records a vehicle's delay in seconds at that moment. Out of an expected 14,400 probes across the week, we successfully collected 13,200 — the other 1,200 were lost when our pipeline was restarting one night (a systemic outage, attributed to us).
Coverage ratio: 13,200 ÷ 14,400 = 91.7%. This city is above the 50% threshold, so it is eligible for rankings.
Step 2: grade each observation. Of the 13,200 collected observations, the poller classifies each one:
| Grade | Delay range | Count |
|---|---|---|
| Early | < −120 s (more than 2 min early) | 660 |
| On time | −120 s to +120 s (within ±2 min) | 9,240 |
| Late | > +120 s (more than 2 min late) | 3,300 |
Step 3: compute the on-time percentage.
on-time % = n_on_time ÷ n_obs = 9,240 ÷ 13,200 = 70.0%
This is the raw on-time percentage, computed only on observed data. No imputation.
Step 4: apply the Wilson confidence interval. With 13,200 observations and a 70% rate, the 95% Wilson interval is approximately ±0.8 percentage points. We report this as 70.0% ± 0.8%. For a ranking we use the lower bound, 69.2%.
Step 5: show the coverage annotation. Because coverage is 91.7% (above 50%), no provisional badge is shown. The coverage percentage appears beside the score so readers can judge the data completeness themselves. Had coverage been 40%, the score would be marked PROVISIONAL · NOT RANKED and excluded from leaderboards — the 70% figure might look fine, but it's drawn from an incomplete week.
What the coverage-adjusted score adds. The 70.0% above weights all 13,200 observations equally. The coverage-adjusted version groups observations by local hour, notes that our outage happened between 2 am and 5 am (low-service hours), and reweights. Because the missing hours were low-traffic, the adjustment is small — 70.3% — and the score is marked as coverage-adjusted so the reader can see both figures.
What this shows. The 70.0% figure is an honest count of what was observed. The coverage annotation explains the data completeness. The Wilson interval makes sampling uncertainty visible. Together, these prevent a partial week from masquerading as a full one.
Raw observations are stored losslessly. Every observation record is preserved in full-fidelity daily archives (Parquet format, compressed), with a verified offsite mirror. Rollups — the pre-aggregated summaries used for fast display — are derived from raw data and rebuildable from scratch if needed. They store additive counters (n_obs, n_on_time, delay sums), never pre-computed averages, so any time window recombines exactly and correctly weighted by actual observation counts. The raw archive is never pruned without first verifying the Parquet copy; that gate is permanent and unconditional.
A city only appears once its data is complete enough (routes loaded, labels resolved, in-area vehicles, gradeable delays). Incomplete cities are hidden automatically and reappear when their data improves.
These are transparent, reproducible measures from open feeds — not official agency metrics. Please cite "Global Transit Stats, from public GTFS-Realtime feeds" and link the specific city/route page.
Last updated: 2026-07-29
Built from this city's live data. Copy it, then paste into your email or a public comment — edit the bracketed fields first.