Every day, social media, the news and even the government sell you loud stories and extreme opinions. Digital transformation experts say that quick commerce is killing kirana stores. Enthusiastic startup bros say that gig work is liberating India's poor and uneducated. Captains of industry believe that the only way to make it in this market is to work 90 hours every week. Influencers disguised as amateur economists say that welfare measures are making people lazy, and almost everyone in the world says that AI is going to change everything.
The only thing more exhausting than listening to these stories is trying to debunk them. The programmer Alberto Brandolini coined the "Bullshit Asymmetry Principle", which states that the amount of energy needed to refute bullshit is an order of magnitude bigger than that needed to produce it. Fighting misinformation and propaganda is not only an incredibly hard, uphill task, it is practically Sisyphean. No one person can win against the flood of tiny attention spans, clickbait and virality.
But, in this talk, I'll convince you that being a Sisyphus is not only liberating, but also incredibly rewarding. Moreover, as FOSS and open-data enthusiasts, we are uniquely positioned to fight misinformation. India quietly publishes some of the richest microdata in the world — and the National Sample Survey, vast as it is, is only a fraction of what's out there. Armed with this data and a Jupyter notebook, we can dismantle even the loudest stories. It turns out, for instance, that the dreaded link between free public healthcare and "wasteful" spending on tobacco comes down to a difference of about two rupees a month — while the genuine effects of welfare hide in places no headline thinks to look, like an LPG subsidy quietly turning into a child's tuition.
For two years I have been doing exactly that in my newsletter. This talk is about the data-driven process of investigating plausible-sounding claims that may end up being dubious: taking the prior everyone repeats, getting public evidence, updating your own beliefs and finally publishing not just the results, but also all the data and code to support reproducibility.
I'll be taking specific examples that have dominated the popular discourse around technology, the economy and culture — sometimes with a clean answer, and sometimes with a more uncomfortable one. Asking what a 90-hour week would actually cost a salaried, white-collar Indian — in sleep, in time with family, in the unpaid care work that women already shoulder — settles the argument far better than any motivational LinkedIn post. And once in a while the most honest finding is that the data refuses to answer at all: our labour surveys have no way to even define a "gig worker", which should give anyone celebrating their liberation some pause. But threaded through all the examples is the unglamorous, reusable part: where this open data actually lives, why it's so painful to use (it arrives as disconnected tables of cryptic codes that you have to stitch together and decode before any of it means a thing), how to wrangle it with the FOSS stack, and how to publish an analysis others can check — because a debunking you can't reproduce is just another opinion.
This is neither a statistics lecture nor a series of hot takes. It's an argument that good-faith, reproducible, open-data analysis is civic infrastructure — and exercising it is our civic duty.