Part V — The fix · Chapter 19
What Honest Measurement Looks Like
Working chapter of Why can’t Toronto move? — the report’s summary page uses only claims that passed our receipt check. Figures below marked ⚠️ are still in the re-verification queue, labelled honestly rather than hidden. How that works: check our work.
On the very same TTC data, on-time performance reads 20.7% or 39.6% depending only on which definition you use — and the gap between those two numbers is the whole case for a scoreboard that can't be gamed.
Toronto doesn't lack an accountability instrument for its transit system's performance — the TTC publishes a monthly CEO's Report with real strengths: it states missed targets in plain running text, not just a red cell in a table, and it reports safety metrics that got worse with the same visibility as ones that improved. What Toronto lacks is a definition of "on time" rigorous enough that a good number actually means something, and a public habit of running the same honest check every month regardless of whether that month was good.
This project's own 55-day analysis of TTC's real-time GPS data, covering 1.89 million trips, put a number on exactly how much the definition matters. Run the identical dataset, the identical buses, the identical 55 days, through three different on-time windows, and the on-time rate moves from 20.7% to 39.6% to 71.1% — purely by widening the tolerance. Under the tight, rider-useful standard (no more than 30 seconds early, no more than 2 minutes late), barely one trip in five counts as genuinely on time. Under the TTC's own official standard (one minute early to five minutes late), that same data reads as roughly two trips in five. Under a looser ±5-minute band that no agency should use as a headline, seven in ten trips pass. None of the three descriptions is false. Only one of them is useful to someone standing at a stop.
The optimal standard exposes something the looser windows hide entirely: early-running. A bus that leaves a stop before its scheduled time strands a waiting rider exactly as surely as a late one does, but both the TTC's own standard and the ±5-minute band are generous enough toward early departures that this failure mode barely registers. On route 119 Torbarrie, more than 30 seconds early on 92% of trips; on 503 Kingston Road, 86%. A single on-time percentage can't see this kind of failure — it can only be seen by reporting early-running and lateness as two separate numbers, which is exactly what this project's data shows Toronto currently doesn't do.
A second finding sharpens the point further: measurement location matters almost as much as tolerance. The arrival-based figure this project computed at the last observed stop (39.6% under the TTC-standard window) sits well below the roughly 75% bus on-time performance the TTC's own CEO's Report states for what should be a comparable standard — a gap most plausibly explained by the TTC measuring departure punctuality at terminals rather than arrival, and by the roughly 60-second resolution of the public data feed. The direction of that gap matters less than what it proves: two defensible ways of applying "the same" standard can land 35 points apart. Publish the standard's name without the measurement point, and the number is not yet meaningful.
None of this is uniquely a TTC problem — it's the generic failure mode of transit-agency self-reporting everywhere. New York's MTA quietly redefined its headline wait-time metric around 2018 after Wait Assessment proved too permissive; London's TfL now reports a journey-time metric and Excess Wait Time specifically because a simple on-time percentage is trivially gamed; no agency surveyed for this project — not TTC, not San Francisco Muni, not the UK's national rail standard, not Chicago's CTA — publishes a standard as rigorous as the one this project's own data demonstrates is buildable from TTC's existing GPS feed today.
The fix this project's research proposes is not a new dataset — it's a set of publishing rules applied to data TTC and the City already collect. Lock in the rigorous, rider-centric on-time window as the headline number, not a lenient one chosen because it looks better. Report the exact tolerance and the exact measurement point — terminal or all stops, arrival or departure — every single cycle, because both move the number as much as the window does. Split early-running from lateness as two permanent, separate lines, not one blended figure. Whenever a metric's definition changes, publish what the new number would have been under the old definition for at least a full trailing year, so a redefinition can never quietly launder a bad result. And treat design promises and delivered operations as two different scores, the way the international Bus Rapid Transit Standard already does, so a corridor's paper design can't stand in for what riders actually experience once it opens.
Toronto already has real transparency habits worth keeping — TTC's practice of naming a missed target in prose, not just colour-coding it, is genuinely uncommon and should survive any redesign. What's missing is a single number that means the same thing every month, defined tightly enough that improving the number requires improving the service, not just loosening the definition.
⚠️ The on-time figures above come from this project's own exploratory GTFS-Realtime computation — first-party, reproducible from raw TTC stop-level data, and independently cross-validated in several ways (its bus-speed figure matches TTC's own separately reported number) — but not yet re-verified end-to-end by an outside party. They should be read as this project's best current evidence, not as an official TTC statistic.
Receipts
Source: one of this library's internal records (Parts 1–2: design rules and proposed metric table, §1.9) and one of this library's internal records ("On-time performance — it depends enormously on the standard"). On-time figures (20.7% optimal / 39.6% TTC-standard / 71.1% ±5min, terminal measurement; early-running rates for routes 119 Torbarrie and 503 Kingston Road) are this project's own 55-day, 1.89-million-trip GTFS-Realtime computation, carried with the report's own evidentiary label: our own computation, not yet independently re-verified by a second party. ⚠️ carried forward accordingly. Note: the measurement report circulates under two filenames (one of this library's internal records and one of this library's internal records); this chapter cites the latter, and they appear to be the same document.