Data study 03 of 09 · Housing quality

Are apartments worse in poorer neighbourhoods?

Toronto sends inspectors into every big rental building and scores each one out of 100. I matched 3,452 of those scores to neighbourhood incomes and tested the gradient everyone assumes is there.

RentSafeTO building audits 3,452 buildings scored ~3 minute read Skip to the technical notes

This page comes in two parts: the story in plain language first, then technical notes with the data joins, model ladder, and full estimates.

The citywide finding

The clearest problem is heat, and it is everywhere

Before comparing rich and poor neighbourhoods, one number from the audits deserves to stand alone: how many buildings provide no air conditioning at all.

How to read the first chart: picture 100 apartment buildings, one square each. The dark squares are buildings whose audit records no air conditioning of any kind provided by the building. As summers get hotter, that is a health issue, especially for older tenants. The small bars on the right show the same share separately for lower-income, middle, and higher-income neighbourhoods: the numbers barely differ. Heat exposure in Toronto's rental towers is a citywide problem that cuts across income lines.

84 of every 100 audited apartment buildings provide no air conditioning

Each square is 1% of Toronto’s 3,452 evaluated rental towers – dark = no building-provided cooling – heat risk shows up in every income third, at every scale

84 in 100 audited apartment buildings provide no air conditioning 16 report some form of cooling BY NEIGHBOURHOOD INCOME Bottom third 85% Middle 85% Top third 81%

RentSafeTO evaluations + neighbourhood profiles · City of Toronto open data · n = 3,452 buildings

The income test

The steep slope everyone expects is missing

Now the main question. If richer neighbourhoods had visibly better-kept rental buildings, their audit scores should climb as income climbs.

How to read the second chart: I split the city's neighbourhoods into three equal groups by income and averaged the building scores in each. The bars are drawn on the full 0-to-100 scale on purpose. Chopping off the bottom of a chart can make a tiny difference look dramatic; drawing the whole scale keeps the chart honest. Here, honest looks flat: all three groups average about 89.

A more careful comparison, which accounts for building age and size, does find a small income effect: doubling a neighbourhood's income lifts the average score by about 2.4 points out of 100. Real, but modest. The clearer gaps are elsewhere. Buildings owned by Toronto Community Housing, the city's public landlord, score about 2 points below similar private buildings, and each decade of building age costs about 0.7 points at typical ages. (An earlier version of this page said "about a point"; the model's age curve makes the effect depend on the age it is read at, and the technical notes now report it properly.)

Audit scores stay near 89 in every income third; the visible gaps are ownership and age

Average audit score (0–100) by neighbourhood income third – one rung = 5 points – drawn on the full 0 to 100 scale, so a flat result looks flat

0 25 50 75 100 88.9 Bottom third 1,254 buildings 88.4 Middle 1,065 buildings 89.9 Top third 1,133 buildings after accounting for age and size: a doubling of income ≈ +2.4 points city-owned (TCHC) vs similar private: −2.1 · each decade of age, at typical ages: −0.7

RentSafeTO evaluations + census neighbourhood income · 3,452 buildings

How to read the third chart: averages can hide as much as they show, so here is the entire dataset at once. Every audited building appears as one dot, placed by its exact score, in three columns by neighbourhood income third. If income drove building conditions, the three clouds would sit at visibly different heights. Instead they overlap almost completely, and each group's midpoint lands within a couple of points of the others. On a computer, run your cursor across the chart (or focus it and press the up and down arrow keys) to count the buildings at any score.

Every building, one dot: the three income groups overlap almost completely

One dot = one audited building, placed by its exact score – columns are neighbourhood income thirds – dark ticks mark each group's median – hover, or focus the card and use the arrow keys, to count buildings at any score

0 25 50 75 100 median 90 Bottom third 1,254 buildings median 90 Middle 1,065 buildings median 91 Top third 1,133 buildings

RentSafeTO evaluations + census neighbourhood income · all 3,452 scored, income-matched buildings shown

A craft detail: matching buildings to neighbourhood incomes failed for 239 buildings at first, because the two datasets spelled neighbourhood names slightly differently (an apostrophe here, a hyphen there). Cleaning the names matched every building. Small data-janitor work like this is where analyses quietly go wrong, so I treat it as a core part of the analysis.

The program-evaluation angle

Did buildings improve while the program watched?

RentSafeTO is an enforcement program: buildings that score poorly are ordered audited or re-inspected within a year, while high scorers wait three. Its own evaluation history, 11,760 inspections from 2017 to mid-2023, lets us ask whether anything actually moved.

How to read the fourth chart: the line tracks the citywide average score among all evaluations conducted each year, under the program's original scoring scheme. It climbs from 65.5 in the launch year to 80.9 in 2022. One detail makes that climb more convincing rather than less: because low scorers are re-inspected most often, the later years' samples lean toward buildings with a history of problems, which drags these averages down. They rose anyway.

Under the old scoring scheme, the citywide average climbed 15 points in six years

Mean evaluation score (0–100, pre-2023 methodology) among all evaluations conducted each year – the 2023 revamp rescored buildings on a new basis, so the two eras are never mixed

0 25 50 75 100 65.5 2017 3,417 evals 72.5 2018 1,823 evals 78.6 2019 1,567 evals 77.1 2020 1,473 evals 76.7 2021 1,471 evals 80.9 2022 1,987 evals later years re-check low scorers more often, which drags these averages down; they rose anyway

RentSafeTO pre-2023 evaluation history · 11,738 evaluations, 2017–2022 · partial 2023 (n = 22) excluded

How to read the fifth chart: among the 3,447 buildings inspected more than once, it compares each building's first score with its latest, grouped by where the building started. The lower the starting band, the larger the average gain, exactly the gradient an enforcement program is designed to produce. Buildings ordered audited gained 26 points on average; buildings that started above 86 barely moved.

One caution before reading that as pure program effect: scores bounce, and a building measured at its worst will usually score better next time for that reason alone. Statisticians call it regression to the mean, and it inflates the gains of the lowest band. The citywide climb in the fourth chart, which regression to the mean cannot produce on its own, is the sturdier half of the evidence. The 2023 scoring revamp also means none of these numbers can be compared with the 89-point averages elsewhere on this page.

Buildings that started lowest were re-checked soonest and gained the most

Mean change from a building’s first evaluation to its latest, by starting band – one tick = 2 points – 3,447 buildings evaluated more than once, 92.9% improved – part of the top-to-bottom pattern is regression to the mean, explained in the text

0 +10 +20 +30 Started 50 or below audit ordered · 61 +26.13 Started 51–65 re-checked in 1 year · 1,749 +18.73 Started 66–85 re-checked in 2 years · 1,597 +10.56 Started 86 or above re-checked in 3 years · 40 +2.38

RentSafeTO pre-2023 evaluation history · bands are the city’s own re-inspection rules (observed score ranges 0–50, 51–65, 66–85, 86–100)

Try it yourself

Look up a neighbourhood

Type any Toronto neighbourhood to see its audited buildings, average score, and share without air conditioning, compared with the city as a whole. Everything runs in your browser from the data behind the charts above; nothing you type is sent anywhere.

What this does not prove

The honest fine print

Audit scores measure what inspectors check: lobbies, elevators, garbage rooms, walls, windows. Two buildings with the same score can still feel very different to live in, and an audit score says nothing about rent, crowding, or how a landlord treats people. The near-flat result says Toronto's audited buildings pass inspections at similar rates everywhere; it does not say housing experiences are equal everywhere.

Technical terms, translated

Income thirds
All neighbourhoods ranked by household income, then cut into three equal stacks: bottom, middle, top.
Average (mean) score
All the building scores in a group added up and divided by the number of buildings.
Median
Line all the buildings up from lowest score to highest; the median is the one in the middle. Half the group scores above it, half below.
"Accounting for building age and size"
Older and bigger buildings score differently for reasons that have nothing to do with the neighbourhood. The model compares buildings as if they were the same age and size, so income's own effect stands out.
Regression to the mean
Anything measured at an unusually bad moment will usually measure better next time, improvement or no improvement. It is why the lowest-scoring buildings' big gains are read with caution.
"Explains 7% of the variation"
Even knowing a neighbourhood's income, building age, and size, you could barely guess an individual building's score. Most of what separates buildings is something else.
RentSafeTO evaluations Census neighbourhood income Python Data cleaning & matching
Part two · For technical readers

Technical notes

The data joins, the model ladder, and the full estimates behind the plain-language story above.

Data and joins

RentSafeTO building evaluations joined to the City's neighbourhood profiles (census-based median household income), both from Toronto's open data portal. The raw join failed for 239 buildings because the two datasets punctuate neighbourhood names differently; joining on normalized keys matched all 3,461 evaluated buildings, and the analytic sample is 3,452 with complete model variables. The cooling variable is the evaluation's AIR_CONDITIONING_TYPE field, coded none versus any type.

Model ladder

Three OLS models at the building level, with cluster-robust standard errors by neighbourhood (Cameron and Miller, 2015), since buildings in one neighbourhood share its income value. M1 regresses the evaluation score on log income alone. M2 adds building age, age squared, log unit count, storeys, and ownership type (reference: private). M3 swaps income for the city's Neighbourhood Improvement Area designation with the same controls. Income coefficients convert to a per-doubling scale by multiplying by ln 2.

Scorei = β0 + β1 log Incn(i) + β2Agei + β3Agei2 + β4 log Unitsi + β5Storeysi + γtype(i) + εi

where n(i) is building i's neighbourhood and γ holds ownership-type intercepts (reference: private); errors clustered by neighbourhood. M1 keeps only log Inc; M3 replaces it with the Neighbourhood Improvement Area indicator.

Score models (0–100 scale), cluster-robust SEs by neighbourhood
PredictorCoef.SEp
M1 · log median income (bivariate)1.911.240.125
M2 · log median income (adjusted)3.421.350.011
M2 · building age, per year-0.100.030.002
M2 · TCHC vs private-2.130.48< 0.001
M2 · social housing vs private-1.450.600.016
M3 · Neighbourhood Improvement Area-0.900.630.154

n = 3,452 buildings. M2 and M3 control for building age, age squared, log units, storeys, and ownership type. R²: 0.002 (M1), 0.066 (M2). Income effect per doubling of income: 1.32 points (M1), 2.37 (M2).

Interpretation notes

The bivariate income slope is small and imprecise; adjustment strengthens rather than weakens it, because building age suppresses the raw relationship (buildings in the top income third average 66.7 years old, against 62.0 in the bottom third, so older stock in richer areas masks the income effect). Even adjusted, R² is 0.066: neighbourhood income tells you very little about an individual building's score. The null on the Neighbourhood Improvement Area flag is consistent with that.

Limitations

Evaluation scores compress real differences between buildings and reflect the audit protocol's scope (common areas and building systems). Income is measured at the neighbourhood level, so this is an ecological comparison, and within-neighbourhood variation between buildings dominates the variance.

Residual diagnostics

Assumptions are checked graphically against M2's residuals, with numeric summaries as adjuncts to the plot rather than substitutes for it. Linearity passes cleanly: the binned mean residual stays within about one point of zero across the fitted range (quadratic term in a residual-on-fitted check, p = .56), and linearity is the assumption OLS is least robust to, so this is the check that matters most. Two departures are real and named. The residual distribution is left-skewed with heavy tails (skewness −1.5, excess kurtosis 5.4): a tail of buildings scores far below prediction while the 100-point ceiling compresses the other side. And spread shrinks as predicted scores rise (mean absolute residual falls from about 7.9 to 5.2 across fitted bins), the same ceiling at work. Neither threatens the reported inference, which already uses cluster-robust standard errors, but both are visible in the plot below and belong on the record, and the outlier sensitivity below stress-tests them directly. One correction this diagnostic pass surfaced: M2 contains age and age squared, so the linear age coefficient (−0.10 per year) is the slope at age zero; the marginal effect at the mean building age (64 years) is −0.70 points per decade (−0.81 at age 40, −0.58 at 90), and the page now reports that.

The model checked against its own leftovers: flat where it must be, honest about its tail

One dot = one building, placed by predicted score and residual (actual minus predicted) – dark hairline = mean residual in 20 equal-count bins – a flat line means no nonlinearity is hiding in the fit

+10 0 -10 -20 -40 -60 a long left tail: buildings scoring far below their prediction the flat dark line is the linearity check, passing 82 86 90 94 predicted score

RentSafeTO adjusted model (M2) · all 3,452 buildings · committed diagnostics output

Outlier sensitivity

The heavy left tail invites one specific worry: that a few badly scoring buildings drive the reported effects. The check is to refit the identical specification with estimators that treat extreme observations differently and compare the estimates that carry this page's claims. Robust and median fits are compared on point estimates; published inference stays with the cluster-robust OLS.

EstimatorIncome, per doublingTCHC gapAge per decade (mean age)
OLS (published), n = 3,452+2.37−2.13−0.70
Huber robust+2.17−2.58−0.64
Median regression+2.09−2.94−0.72
OLS, |residual| > 3sd trimmed (36 dropped)+2.19−2.34−0.62

Every conclusion survives, and one strengthens. The income effect stays small and stable (+2.1 to +2.4 points per doubling), the age effect stays near −0.6 to −0.7 per decade, and the TCHC gap grows under the estimators least influenced by the tail (−2.6 to −2.9 versus the published −2.1): the extreme low scorers were diluting the ownership gap rather than creating it. The page's "about 2 points below" is therefore the conservative end of the range, and it stays as written.

Scores over time (pre-2023 scheme)

Evaluation-level history: 11,760 inspections, 2017 to mid-2023, one row per evaluation with the score and the resulting re-inspection order. The enforcement bands and their score ranges are taken from the data itself (RESULTS_OF_SCORE: audit at 0–50, re-evaluate in 1 year at 51–65, 2 years at 66–85, 3 years at 86–100). Year means use all evaluations in each calendar year (partial 2023, n = 22, excluded); within-building change compares first and latest evaluation for the 3,447 buildings with two or more, at least a month apart. The by-band gradient is reported descriptively with the regression-to-the-mean caveat stated in the plain-language section; no causal effect size is claimed. The 2023 methodology change is a hard comparability break and the two scoring eras are never mixed.

The same study in SQL

The analysis above was built in Python. As applied relational-database practice, this study's data also lives in a three-table SQLite database in third normal form: neighbourhoods (154 rows, with census median income) to buildings (3,452, foreign key to neighbourhood) to evaluations (11,760, foreign key to building). Seven SQL queries reproduce the published numbers, from simple counts through GROUP BY and joins up to the first-versus-latest pairing behind the improvement chart, and an automated checker verifies all 27 values against history_results.json and an independent recompute: 27 of 27 match. The pairing query, with window functions doing the work the Python version does with a groupby:

WITH ranked AS (
  SELECT rsn, score,
         ROW_NUMBER() OVER (PARTITION BY rsn ORDER BY completed_on)      AS rn_first,
         ROW_NUMBER() OVER (PARTITION BY rsn ORDER BY completed_on DESC) AS rn_last,
         COUNT(*)     OVER (PARTITION BY rsn)                            AS n_evals
  FROM evaluations
)
SELECT COUNT(*)                                            AS n_buildings,
       ROUND(AVG(l.score - f.score), 2)                    AS mean_change,
       ROUND(100.0 * SUM(l.score > f.score) / COUNT(*), 1) AS pct_improved
FROM ranked f
JOIN ranked l ON l.rsn = f.rsn AND l.rn_last = 1
WHERE f.rn_first = 1 AND f.n_evals >= 2;
-- returns 3,447 buildings, +14.88 mean change, 92.9% improved

One data catch surfaced during loading: the newest roughly 2,000 evaluation rows ship with a blank YEAR_EVALUATED, so the loader derives the year from the completion date, matching the published analysis (checklist habit: batch fields go stale, the event date is the source of truth). The neighbourhood assignment and income terciles come from the Python pipeline's spatial join and are loaded as given; SQL takes over from there.

Reproducibility

Five scripts (get_data.py, run_analysis.py, make_charts.py, plus get_history.py and run_history.py for the over-time section) against the Toronto Open Data CKAN API, on pandas and statsmodels. Every number on this page comes from the committed results.json and history_results.json. The SQL rebuild adds build_db.py, queries.sql, and run_checks.py in the project's sql/ folder.

References

  • City of Toronto. (2026). Apartment Building Evaluations and Registration (RentSafeTO) [data sets, including the pre-2023 evaluation-history resource]. City of Toronto Open Data Portal. https://open.toronto.ca/dataset/apartment-building-evaluation/ (retrieved July 2026).
  • City of Toronto. (n.d.). Neighbourhood Profiles (2021 Census based) [data set]. City of Toronto Open Data Portal. https://open.toronto.ca/dataset/neighbourhood-profiles/ (retrieved July 2026).
  • Cameron, A. C., & Miller, D. L. (2015). A practitioner's guide to cluster-robust inference. Journal of Human Resources, 50(2), 317–372.