The public method behind every figure we publish: which open datasets we hold, how each one is verified, what our confidence tiers mean, and the claims we refuse to make.
From open regulator and government datasets, harvested into a warehouse where every record keeps a pointer to the snapshot it came from. We publish the analysis; the underlying data is credited to whoever published it. Every share is a share within a corpus we name, and that corpus is printed next to the number.
| Figure | Count | What it counts | Measured |
|---|---|---|---|
| Certified model records | 19,563 | models certified with ENERGY STAR or the Department of Energy in our categories | 2026-08-17 |
| Model families | 3,451 | families produced by collapsing variant model numbers, keyer version 1 | 2026-08-17 |
| Safety incident reports | 20,581 | CPSC SaferProducts consumer reports in our appliance categories | 2026-08-17 |
| Search queries measured | 58,220 | distinct queries this site received impressions for, to 2026-08-16 | 2026-08-17 |
| Symptom vocabulary | 119 | controlled symptom nodes narratives are classified into | 2026-08-17 |
| Statistics published from it | 0 | verified fact blocks released to any page — nothing has cleared review yet | 2026-08-17 |
Nobody publishes how many units of a given appliance were sold, so the number of machines that could have gone wrong is unknowable. That makes an overall breakdown rate impossible to state honestly, and we do not state one. What can be stated honestly is a share within a corpus we name: a proportion of a specific set of records, over a specific date range, with the size of that set printed beside it. If you see a percentage on this site without those three things next to it, it is a defect — please tell us.
Every share we publish carries all three of:
People file a report with the Consumer Product Safety Commission when something frightened them, not when something merely broke. So that corpus is not a picture of what fails most often; it is a picture of what fails dangerously. Publishing a family fault ranking from it would be a real number describing something other than what a reader would take it to mean, which is why we do not do it. We measured the effect rather than assuming it, and the measurements are below.
What the corpus is genuinely good for is the language people use. Across the narratives, the ice maker is named 1,067 times, the control board 492 and the compressor 279. That tells us which components deserve research and plain-language explanation. It is a reading guide, not a frequency table, and we treat it as one.
Each row carries the licence it is used under. That class travels with every record into the page that renders it, so a source cleared only for internal use can never end up cited in public.
Published by U.S. Environmental Protection Agency and Department of Energy · Public domain, open API
Published by U.S. Department of Energy · Public domain, open access
Published by U.S. Consumer Product Safety Commission · Public domain, federal, redistributed with the required disclaimer
Published by U.S. Consumer Product Safety Commission · Public domain, federal
Published by California Energy Commission · Public state database
Published by Administrative Office of the U.S. Courts, via CourtListener · Public records
Published by The manufacturers · Facts only — never republished text
Published by Their respective communities · Internal signal only
Published by Various · Excluded
Published by Various · Excluded as a content source
The raw payload is stored first, keyed by source, URL, date and content hash. A fact whose snapshot pointer is missing does not exist as far as our pipeline is concerned.
Structured sources are parsed by deterministic code. Language models are used for one job only — mapping free-text narratives onto our controlled vocabulary — and never to invent a value.
Every load-bearing number is re-derived from the stored snapshot by code. An assertion that cannot be recomputed from a snapshot is not treated as a fact.
Variant model numbers are collapsed into families before anything is counted, so a share is never computed against a mixed denominator. Each aggregate records the keyer version it was computed under.
Prose may assert only what resolves back to a stored record. A page that cannot clear that check is not published at that level of detail.
Every fact we store carries a tier, a source and the date of the snapshot it came from. The tier says how directly the fact is evidenced — not how confident we feel about it.
A manufacturer error-code table, a service bulletin, a recall notice.
How to read it: Strongest. Someone with legal responsibility for the product wrote it down.
Counts derived from CPSC consumer incident reports.
How to read it: Strong on what the dataset covers, and bounded by what it collects. The corpus is named next to the number so you can judge that yourself.
How often a topic appears across public video or forum activity.
How to read it: Directional. Good for what is discussed a lot; it measures attention, not how often something actually happens.
A trait shared across a model series applied to a member of that series.
How to read it: Weakest, and labelled. Treat it as a starting point, never as a finding.
Around 202 groups of model records are stored twice, because two catalogues spell the same model differently and our uniqueness check compares the spellings exactly.
No published count is inflated — every figure above is counted over distinct models. But where the two copies carry different specifications, which copy is read is not currently decided by anything meaningful.
Open, tracked, fix is a case-insensitive uniqueness key plus a one-off reconciliation.
Our stored recall corpus holds 6 notices because the collector was only ever pointed at a rolling recent window.
We make no recall claims anywhere on the site until the historical pull has run.
Open, tracked, the collector already supports the full pull.
Corrections are welcome and we would rather hear about a bad number than keep it. If a figure on this site is wrong, unclear, or missing the corpus it should be counted against, tell us which page it is on and we will fix it or remove it.
This page was last reviewed on 2026-08-21.