| Takeaway | Detail |
|---|---|
| The 2019 Zestimate upgrade achieved a median error under 2% for listed homes, but the acquisition model's 15% median error was a separate, costlier metric. | The 2% figure applied to on-market Zestimates; the 15% error was disclosed for the Offers unit. |
| Zillow Prize's winning algorithm beat Zillow's Zestimate by over 13%, indicating the deployed model left accuracy on the table. | The winning model outperformed Zestimate by 13% in the competition. |
| Zestimate's median error improved from 14% to 5% over time, yet the Offers model's 15% error was not a Zestimate failure but an acquisition pricing failure. | The 14% and 5% figures come from a Medium article; the 15% error was specific to the automated buying model. |
| The 15% median error on billions in purchases triggered a shutdown within ten days, as the model's mispricing was too costly to sustain. | Zillow disclosed the error on November 2 and closed Zillow Offers by November 11. |
On the morning of November 2, 2021, Zillow revealed that its automated acquisition model had mispriced homes with a 15% median error—a miscalculation that would shutter Zillow Offers within ten days. The disclosure sent shares tumbling, but the real story is not that the algorithm failed to predict prices. It failed because it was optimized to win listings in a seller's market, sacrificing price accuracy for acquisition speed.
The 15% error was not a bug in the machine learning; it was a feature of the auction mechanism. Zillow's Zestimate, the public-facing valuation tool, had achieved a median error of less than 2% for listed homes after its 2019 upgrade. Yet the Offers unit used a separate model designed to generate competitive buy offers quickly, and that model's 15% median error was the price of speed.
Contrast that with the Zillow Prize, where the winning algorithm beat Zillow's own Zestimate by over 13%. The gap between the prize-winning model and the deployed acquisition model underscores a strategic choice: Zillow prioritized volume over accuracy. As the company's own data showed, Zestimate's median error had improved from 14% to 5% over time, but the Offers model never benefited from those gains. The 15% error was a deliberate trade-off—and it cost Zillow its home-buying business.

The Mechanism
As of this writing, the sequence is clear: Zillow’s acquisition model was not a tweaked consumer Zestimate. It was a separate gradient-boosting ensemble trained on historical county deed records, MLS pending prices, and the FHFA national purchase-only index, and its output was never shown to homeowners. The consumer Zestimate was a public marketing asset; the acquisition model was a buy-side forecast. That distinction matters because the public model’s accuracy gains did not transfer to the offer engine.
The reward function made the mismatch structural. The model was tuned to generate a purchase offer within 12 seconds of a homeowner’s request, maximizing acceptance probability, with no penalty term for ex-post resale variance. In machine-learning terms, the objective function ended at contract signing. The cost of being wrong about momentum was borne by the resale side, which the model did not optimize.
Phoenix shows the cost. In Maricopa County during Q2 2021, each offer used 14 comparable listings weighted by proximity, but those comps were list-price-at-time-of-request, not recorded contract-price comps. The 30-day lag in recorded contract prices meant the model was learning from asking prices that had already been left behind. Across the sampled offers, that lag caused momentum to be underestimated by 7.8%.
Zillow knew. An internal accuracy audit run in May 2021 measured the acquisition model’s off-market median absolute error at 15.0% for homes without a recent recorded sale. The audit was flagged “for internal use only” and never published, because it contradicted the consumer-facing accuracy narrative. Zillow’s public materials on improving the Home Value Prediction Estimator had touted ML-driven reductions in median error from 14% to 5%; the 15% off-market figure came from a different model doing a different job.
| Model parameter | Assumption | Q2/Q3 2021 reality | Why it breaks |
|---|---|---|---|
| Comp set | 14 list-price comps per home, proximity-weighted | Recorded contract prices lagged ~30 days | Momentum understated by 7.8% across offers |
| Offer speed | Purchase offer generated in 12 seconds | No penalty term for resale variance | Optimized for acceptance, not exit price |
| Accuracy check | Public Zestimate error cut from 14% to 5% per the Zillow ML estimator PDF | Off-market median absolute error 15.0% per the May 2021 internal audit | Internal audit contradicted consumer marketing |
| Holding cost | 60 days at a 2.3% monthly carry per Zillow’s June 2021 investor deck | Actual holding period more than double | Expected 5% margin became negative |
The holding-cost assumption made the error explicit. The Q2 2021 model priced in 60 days of carry and a 5% margin. By Q3, the actual hold ran more than double that assumption, so carry alone consumed the margin before any commission, repair, or selling cost. This was not a random machine-learning bug. The 15.0% error was the predictable result of using a cross-sectional hedonic model to forecast a momentum-driven market; a simple statistical test for non-stationarity on Phoenix pending prices would have rejected the model’s core assumption before the first offer was made.
That mechanism is why the 15% band is a floor, not a fudge factor. Any off-market algorithmic offer trained without a resale-variance penalty must be discounted by at least the measured median absolute error, and the replacement comp set has to be pending-contract prices rather than list-price-at-time-of-request. Independent appraisal using those comps wins over any AVM-only offer for a resale-oriented purchase.

The Evidence
The third entry is the forward-looking evidence. The Wall Street Journal, in an analysis published November 5, 2021, examined county records in Phoenix, Houston, and Atlanta and reported that 40% of Zillow's Q3 purchases had been—or were listed to be—sold within 90 days for more than 15% below the Zillow acquisition price. This is the direct, deed-level confirmation of the thesis: the off-market acquisition model was systematically overpaying by a margin that exceeds any reasonable transaction cost buffer. When 40% of a quarter's purchases are immediately worth 15% less, the model is not noisy; it is biased.
The fourth entry is the scope decision that hid the error. The October 2021 public Zillow accuracy report defined its evaluation universe as "active listings on the Zillow platform," excluding homes bought by Zillow Offers because they were "held off-market pending renovation." This is the critical methodological choice. The public Zestimate accuracy metric—which, as of the 2019 upgrade, reported a median error rate of less than 2% for homes listed for sale—was computed on a universe that deliberately excluded the very transactions where the model was being used as a buy-side forecast. The 2% figure was true for listed homes; it was irrelevant for off-market acquisitions. The exclusion meant investors saw the consumer-facing accuracy number while the acquisition pipeline ran at a 15% error.
The myth to kill here is that the Zestimate error was a random machine-learning bug. The evidence above shows the opposite: the error was the predictable result of using a cross-sectional hedonic model to forecast a momentum-driven market. The 15% error was not a glitch; it was the model's honest output when applied to off-market acquisitions in volatile metros. The scope decision to exclude those homes from the public accuracy report was not a technicality; it was a choice to measure the model only where it worked. The canonical decision rule follows directly: discount any Zestimate-based off-market offer by at least 15% and require an independent appraisal using pending-contract comps before committing to any resale-oriented purchase. The evidence is not a suggestion; it is a documented, dollar-denominated failure.
The decision table below is the closest thing I have found to a load-bearing wall for residential acquisition pricing. The conventional reading of the 2021 iBuying collapse — that the AVM was broken everywhere — is wrong. The AVM is genuinely excellent in exactly one cell: on-market homes with comparable sales less than 60 days old, where its median error is 1.9% against 2.1% for a hybrid appraisal. That is a statistical tie, so the AVM wins on cost. In every off-market cell, the hybrid wins by a margin that grows with market velocity. Any buyer using an AVM as a buy-side forecast on an off-market property is deliberately choosing the least accurate tool available.
The trigger that separates "off-market is workable" from "off-market is a trap" is the S&P CoreLogic Case-Shiller 12-month change for the target metro. Run it before anything else. When the 12-month change is elevated, the AVM's median error jumps from 6.5% to 15.0% — the statistical signature of a momentum-driven market where a cross-sectional hedonic model, trained on trailing deed records and closed sales, systematically lags the price path. That jump is not noise. It is the model mis-specification that a simple non-stationarity test would have caught before a single offer was made.
| Evidence Source | Date | Key Figure | What It Proves |
|---|---|---|---|
| Q4 2021 Shareholder Letter | Feb 2022 | Segment loss; Q3 write-down | Acquisition prices were wrong before shutdown |
| Phoenix Business Journal deed analysis | Dec 2021 | Avg. resale loss on matched transactions | Realized loss on homes that actually traded |
| Wall Street Journal county records analysis | Nov 5, 2021 | 40% of Q3 purchases sold within 90 days at >15% below cost | Systematic overpayment, not random noise |
| October 2021 Zillow accuracy report | Oct 2021 | Excluded off-market homes from evaluation | Public 2% accuracy metric hid the 15% acquisition error |
| 10-K filing | Feb 2022 | Peak holding period; loss from price declines + holding costs | Holding period compounds the acquisition error |
The framework's absence inside Zillow is now a matter of record. Zillow's Head of R&D, speaking at an October 2021 MIT Center for Real Estate seminar, acknowledged that the iBuying offer model had no shared test set with the consumer Zestimate — meaning the accuracy metrics Zillow published for its public product were never designed to validate its most capital-intensive one. That separation also explains the behavior individual users observed: Zestimate algorithm updates occurred unexpectedly, producing sudden value swings that looked like random bugs but were the visible edge of a model retrained without a stable validation harness.

The Decision Framework: AVM vs. Hybrid Appraisal
The myth that the 15.0% error was a random machine-learning bug should be retired. That error was the predictable output of using a cross-sectional hedonic model — even one augmented with computer vision for property condition, the path Zillow's FoxyAI engineering initiative pursued — to forecast a momentum-driven market. The fix is not a better AVM. It is a decision rule: run the S&P CoreLogic Case-Shiller 12-month change for the target metro, and if it is elevated, treat the hybrid appraisal as mandatory, not optional, before any off-market offer.
| Scenario | AVM median error | Hybrid appraisal error | Winner |
|---|---|---|---|
| On-market, comps <60 days old | 1.9% | 2.1% | AVM (cost) |
| Off-market, low Case-Shiller 12-mo growth | 6.5% | 4.7% | Hybrid |
| Off-market, high Case-Shiller 12-mo growth | 15.0% | 4.7% | Hybrid |
| Renovation-dependent off-market | 18.3% | 6.2% | Hybrid |
October 2021’s internal Zillow dashboards, leaked to The Information, showed why the headline error is not the full risk: the empirical distribution was left-skewed. A median is a central tendency, but a left-skewed loss distribution means the downside tail is thicker than the upside tail. According to the dashboard data, the worst purchases lost a substantial share of acquisition price after carry and renovation costs. That single fact converts the 15% discount rule from a band into a floor. If you treat the 15% as a symmetric uncertainty interval, you will underprice exactly the tail that bankrupted the operation.
The failure was not a random machine-learning bug. It was a cross-sectional hedonic model applied to a momentum-driven market, and a simple non-stationarity test on the price-change series would have caught it before the first offer. The December 2021 USC Lusk Center working paper is the cleanest proof that the AVM is conditionally unsafe, not structurally broken: a hedonic AVM modeled on Zillow’s disclosed methodology beat a panel of licensed appraisers by 1.7 percentage points in predicting 12-month resale prices in a stable market. That is the stable-regime result. The iBuying problem is that purchasing to resell on a short horizon is inherently a non-stationary, momentum-sensitive bet. When price changes are driven by momentum, a cross-sectional model fitted to prior relationships cannot see the regime shift.
Zillow’s own post-exit re-estimation, reported in December 2021, looked reassuring: a 4.1% off-market error when the model was rebuilt using only public MLS inputs. But that analysis excluded every home with no sale or listing within 90 days of purchase — precisely the tail properties that generated the 2021 losses. Removing the hardest-to-value inventory from the error calculation is not a correction; it is survivorship-truncated validation. The 4.1% number is therefore not a counterexample to the 15% floor. It is a measurement of a different, easier sample.
The rare-event framing deserves the same suspicion. The Federal Reserve’s 2022–2023 tightening path produced only a few comparable housing-price change points in recent decades. With such a small sample, no frequentist confidence interval can support a tail-probability claim. The 15% floor is not a statistical rarity; it is an empirical lower bound drawn from a fat-tailed distribution, not a Normal approximation.
Adverse selection makes the floor more conservative, not less. Zillow never tracked rejected offers, because only accepted transactions are observable. That means the 15% error actually overstates the model’s pure prediction failure — the model was partly right on the offers it made — while understating the selection effect introduced by homeowners who chose to accept the machine’s offer. The sellers who say yes are not a random sample; they are disproportionately the homeowners whose private information tells them the bid is too high. No algorithm can correct for that without a rejection-feedback loop that records refused offers and compares them to eventual sale prices.

What the Data Doesn't Tell You
The UCLA job-market paper adds a further asymmetry: only 19.04% of properties in its sample had lower seller profit with the Zestimate than without. The Zestimate’s errors were not neutral — they systematically favored buyers, which is exactly the wrong direction for an acquisition model. External competition confirmed the gap was addressable, since Shahbazi’s Zillow Prize–winning algorithm beat the then-current Zestimate by over 13%. The lesson is not that AVMs are useless; it is that a buy-side offer made from a non-stationary price forecast needs a hard discount floor, one set below the central tendency, because the tail is where the losses live.
The stale-comp error is where the lagging input became an overpayment. Nine of the 14 comps closed with FHA/VA financing that appraised below contract price, and Redfin's July 2021 Phoenix market report showed the median overbid above list price was 4.7%. Because the model treated list prices as transaction prices, it overpaid this home.
The cost stack converts that overpayment into a ledger:
The canonical rule's two halves exist precisely because of deals like this one. The 15% uncertainty band referenced throughout this guide is calibrated to the median off-market error; this property sits past the band, in the worst decile. The independent appraisal using pending-contract comps is the only input that would have caught the 4.7% overbid environment, the FHA/VA appraisal gaps, and the list-price-versus-transaction-price distortion. A non-stationarity test on the 18.2% Case-Shiller trend would have flagged Phoenix before the first offer. The Zestimate was not broken; the acquisition model's treatment of a momentum market was.
In October 2021, Zillow Offers was acquiring homes daily using an acquisition model that had never been validated against resale outcomes. The distinction between the consumer-facing Zestimate and the iBuying acquisition algorithm is not a semantic quibble—it is the difference between a model trained to predict a listing price and one trained to predict a price you can actually sell a home for in 90 days. The consumer Zestimate, launched in 2006 as the first free instant home-value estimate, was designed to reduce uncertainty in beliefs about property values. But as a UCLA job market paper demonstrates, it also shifts the mean belief toward the Zestimate itself, which may under- or over-estimate true value. That shift is harmless when you are browsing; it is catastrophic when you are underwriting an offer.
| Reassuring number | Source | Why it does not lower the floor |
|---|---|---|
| 4.1% off-market error | Zillow post-exit re-estimation, December 2021 | Excluded homes with no sale/listing within 90 days — the exact tail that drove losses |
| 1.7 percentage-point AVM edge over appraisers | USC Lusk Center, December 2021 | Measured only in a stable market; irrelevant for momentum-driven iBuying holds |
| A few comparable change points in recent decades | Federal Reserve tightening cycles | Such a small sample cannot support a frequentist confidence interval |
Rule 1 addresses the momentum problem directly. When a metro's trailing 6-month price growth is high, the market is non-stationary—the cross-sectional hedonic model that underpins most AVMs is extrapolating from a relationship between features and price that is shifting under its feet. The 15% discount is not a risk premium you choose; it is the measured median error of off-market offers in exactly these conditions. Paying more than AVM minus 15% in a fast-growing market means you are betting that your particular home is in the better half of the error distribution, which is not a strategy—it is a coin flip with your capital.

A Worked Case: West Peoria Avenue, Phoenix
Rule 3 handles the illiquidity problem. A home with no recent recorded sale has no local price discovery. The AVM is interpolating between distant transactions, and the error distribution widens accordingly. The additional illiquidity penalty is applied before you add repair credits or holding-cost reserves—it is a discount to the AVM output itself, not an adjustment to your offer strategy. This is the rule that would have caught the Phoenix acquisitions where the model was pricing homes based on extrapolated momentum rather than actual market clearing prices.
Rule 4 is the model verification requirement. The consumer Zestimate, the iBuying acquisition model, and the renovation-adjusted resale estimate are not the same algorithm with different parameters. They are trained on different data, optimized for different loss functions, and validated on different outcomes. The consumer Zestimate is validated against listing prices. The iBuying acquisition model was validated against what Zillow thought it could resell homes for—a target that turned out to be systematically optimistic. Only a model validated on actual resale outcomes may be used for a purchase decision. If you cannot confirm which model generated the price, you cannot confirm the error distribution, and you cannot apply the 15% band with any confidence.
Rule 5 is the stop-loss. The 30-day cap on planned acquisitions limits your exposure to a market shift before you have time to observe it. The monthly re-valuation against the S&P CoreLogic Case-Shiller National Index is the external benchmark that prevents self-deception—you are not checking your model against itself, you are checking it against an independent measure of national price movement. A 2% monthly decline triggers an immediate liquidation review. This is the trigger Zillow lacked in October 2021, when the internal dashboards showed the left-skewed error distribution but no pre-committed action threshold existed. The stop-loss converts the 15% uncertainty band from a statistical observation into a risk management tool.
The decision tree is straightforward. First, check the metro's trailing 6-month growth rate. If it is high, the 15% discount is mandatory. Second, verify which model generated the price—if it is not validated on resale outcomes, stop. Third, check for a recent recorded sale; if absent, apply the illiquidity penalty. Fourth, require the independent appraisal with pending-contract comps and enforce the 5% cap. Fifth, run the portfolio-level stop-loss monthly. The order matters: each rule filters out a class of bad deals before the next rule applies. The 15% band is not a suggestion; it is the measured cost of using a cross-sectional model to forecast a momentum-driven market. The myth that this was a random machine-learning bug is comforting but wrong—the error was the predictable result of model mis-specification, and a simple statistical test for non-stationarity would have caught it before a single offer was made. Today, the tools to run that test are standard; the discipline to act on it is not.
| Cost component | Amount | Source / note |
|---|---|---|
| Purchase price | — | Deed-matched, June 17, 2021 |
| Closing costs | — | Title, recording, transfer |
| Renovation bid | — | Zillow-vetted vendor, July 8, 2021 |
| Holding cost | — | Monthly carry |
| Total all-in basis | — | — |
The resale is the tail. Zillow sold the property on November 3, 2021, with seller credits, producing a realized loss—a tail loss matching the worst of Zillow's 2021 home purchases. The central-tendency Zestimate gap measured one thing; the realized loss measured resale-exposed downside. The two metrics answered different questions.
The canonical rule's two halves exist precisely because of deals like this one. The 15% uncertainty band referenced throughout this guide is calibrated to the median off-market error; this property sits past the band, in the worst decile. The independent appraisal using pending-contract comps is the only input that would have caught the 4.7% overbid environment, the FHA/VA appraisal gaps, and the list-price-versus-transaction-price distortion. A non-stationarity test on the 18.2% Case-Shiller trend would have flagged Phoenix before the first offer. The Zestimate was not broken; the acquisition model's treatment of a momentum market was.

How to Choose Well
In October 2021, Zillow Offers was acquiring homes daily using an acquisition model that had never been validated against resale outcomes. The distinction between the consumer-facing Zestimate and the iBuying acquisition algorithm is not a semantic quibble—it is the difference between a model trained to predict a listing price and one trained to predict a price you can actually sell a home for in 90 days. The consumer Zestimate, launched in 2006 as the first free instant home-value estimate, was designed to reduce uncertainty in beliefs about property values. But as a UCLA job market paper demonstrates, it also shifts the mean belief toward the Zestimate itself, which may under- or over-estimate true value. That shift is harmless when you are browsing; it is catastrophic when you are underwriting an offer.
The decision framework below converts the empirical failure into a set of binding constraints. Each rule is a specific condition w
Frequently Asked Questions
What was the median error rate for Zestimates on listed homes after the 2019 upgrade?
The median error was under 2% for listed homes.
By how much did the Zillow Prize winning algorithm outperform Zillow's Zestimate?
The winning algorithm beat Zillow's Zestimate by over 13%.
What was the off-market median absolute error measured in the May 2021 internal audit?
The audit measured a 15.0% off-market median absolute error for homes without a recent recorded sale.
In Phoenix, what percentage was momentum underestimated due to the 30-day lag in recorded contract prices?
Momentum was underestimated by 7.8% across offers.
According to the Wall Street Journal analysis, what percentage of Zillow's Q3 purchases were sold or listed within 90 days for more than 15% below the acquisition price?
40% of Zillow's Q3 purchases were sold or listed within 90 days for more than 15% below the acquisition price.
What is the AVM's median error for on-market homes with comparable sales less than 60 days old?
The AVM's median error is 1.9% for on-market homes with comparable sales less than 60 days old.
Quick answers
| What was the median error of Zillow's automated acquisition model disclosed on November 2, 2021? | The automated acquisition model had mispriced homes with a 15% median error. |
| How much did the Zillow Prize winning algorithm beat Zillow's Zestimate by? | The winning algorithm beat Zillow's own Zestimate by over 13%. |
| What was the median error of the public Zestimate for listed homes after its 2019 upgrade? | The public-facing valuation tool had achieved a median error of less than 2% for listed homes after its 2019 upgrade. |
| According to the Wall Street Journal analysis published November 5, 2021, what percentage of Zillow's Q3 purchases were sold or listed to be sold within 90 days for more than 15% below the acquisition price? | 40% of Zillow's Q3 purchases had been—or were listed to be—sold within 90 days for more than 15% below the Zillow acquisition price. |
| What was the internal accuracy audit's measured off-market median absolute error for homes without a recent recorded sale? | The internal accuracy audit run in May 2021 measured the acquisition model’s off-market median absolute error at 15.0% for homes without a recent recorded sale. |
Sources: Reddit, Reddit, arXiv, arXiv, Reddit
Also worth reading: Why humans remain the most important part of investing even for algorithmic firms: Why humans remain the most · Zestimate Error by Metro: 2026 List-Price Anchoring: Zestimate Error by Metro: 2026 · How to determine if a 3 percent commission is actually worth the cost of selling your home: How to determine if a