Data Practice
When a Clean Dataset Still Misleads
By Naliaka Chepkoech · Aug 6, 2026 · 6 min read
A dataset can pass every visual check a reporter knows to run and still lead a story astray. The columns line up, the totals reconcile, there are no obvious blank cells staring back at you — and yet the figures describe something narrower, or different, than the headline you're about to write around them. The gap isn't in the formatting. It's in what the numbers were built to capture in the first place.
Most published datasets, whether from a government agency, a research body or a private survey firm, are collected to answer a specific question, under a specific method, over a specific window of time. None of that context travels automatically with the file. A table of "reported cases" might genuinely mean cases that were reported to a particular office, using a particular form, during years when that office was fully staffed — and quietly mean something thinner in years when it wasn't. Reading the numbers without reading the collection notes is the single most common way a factual story ends up making an unfactual claim.
Start with what the count was measuring
Before touching a single row, it's worth finding the methodology note, however brief, and asking what population or activity the dataset was actually built to track. A number described as "national" might be an aggregate of regional submissions with very different completeness. A trend line described as "annual" might reflect a reporting calendar that shifted midway through the series. None of this invalidates the data — it simply means the finding you draw from it needs the same boundary the data itself has.
"The cleanest-looking column in a spreadsheet is often the one that got summarised the most before it reached you."
Missing values are information, not noise
A blank cell or a suspiciously round number (a suspiciously frequent zero, or a repeated placeholder like 999) usually means something was not collected, not that nothing happened. Suppression for privacy, non-response, and simple gaps in submission all produce different patterns, and distinguishing between them changes what you can responsibly say. A region with consistently missing figures is not the same as a region with consistently low figures, even though a quick chart might make them look identical.
When joining two datasets — say, a health record set and a population estimate — the same caution applies to the join itself. Matching on a name, a ward, or a facility code across sources collected at different times or by different bodies can quietly create false pairs, especially where administrative boundaries have been redrawn. A join that looks tidy in a spreadsheet can still be wrong in the field.
None of this argues against using data — it argues for carrying its limits alongside its conclusions. A finding stated as "the dataset shows X, collected under Y conditions, with Z known gaps" is a stronger sentence than a bare number, and it's one that holds up when someone checks your sourcing. That habit, more than any particular tool, is what separates responsible data use from decorative data use.
Want to work through this with your own dataset?
The Public Data Handling Session is built around material relevant to your beat.
▸ Get In Touch