Is Data Crawling Actually Interesting to Anyone?

In the last few years if there’s something I’ve become very familiar with it’s granular economic data in the US (usually place based but sometimes demographic in the broadest sense of the word). I find that modern economics of any bent is, when talking about data, tend to be excessively focused on snapshots (one or two stats) that are usually relatively aggregate in their structure (example: the unemployment rate). The thing is that there’s actually a lot of stats out there. Job migration, local unemployment, capitalist income percentage, income percentiles, industry composition, firm formation and hiring by size… All of these things are (up to a point) publicly available data, but the sources are often diffuse and no normal human being has the combination of nerdy technical expertise and desire to compile it in a loose-but-coherent structure.

I’ve been toying around with the idea of starting a youtube channel like this for a while. Would anyone actually find a youtube channel interesting if it went through economic issues of the last twenty years: “Here’s what happened to areas that had housing booms at the outbreak of the great recession. This is how Covid affected cheap-housing areas. This is what has actually happened to American income earners at different percentiles, who they are, where they live, and what they do…” and so on.

I am deeply aware of the Austrian critiques of real world statistics, both theoretical and practical. Theoretically all data is just data and is silent to the point of being misleading without theory. Many essential variables, like genuine human preference rankings or willingness to pay, are totally unobservable even under ideal practical conditions. Practically, the unemployment rate doesn’t tell you anything about skills mismatch. Survey data is subject to any number of respondent errors. Some errors could potentially be politically motivated…

To this the best I could do is to qualify any results that I present. Where all the data indicates the same thing from different sources or is backed up by theory the more probable the result is. Things like the local unemployment rate or survey data are more politically motivated and open to error than something boring involving establishments and people who could be culpable for tax evasion (annual census data on the number of firms in a county taken from tax data is kind of obscure and dangerous to lie about/hard to and therefore open to very little sampling error).

I couldn’t tackle much in the way of explicitly Austrian questions with this in a compelling way with most of the data I’m looking at (money supply aggregates are national, differences in state policies are often minor and often correlated with other policies or conditions (New York City can tax the hell out of people and they won’t leave because it’s New York, and we can’t observe the counter-factual anyway)), so maybe this is the wrong group of (theoretically existent) econ nerds to ask, but anyway, is this in any way an interesting idea or is it just a good way to put people to sleep?

I reckon having the this type of data would be potentially very interesting even if it isn’t explicitly Austrian since Austrians could use them as interesting data points to illustrate theories or plausible narratives based from theory.

there’s actually a lot of stats out there

Statistics are not automatically a bad thing, even in Austrian theory. It’s just that statistics are usually produced by people with an agenda (to produce that stat) and that’s what makes most stats worthless. The government has an interest in making itself look good. But when the government gathers stats that have nothing to do with itself (or gather them in a way that inadvertently tells the truth about itself), they can actually be valuable. Answering the question, “Why should I believe this data tells us anything about reality?” is the first step to making credible arguments based on stats.

I recommend reading More Sex Is Safer Sex by economist Steven Landsburg. He is not an Austrian. But his methods of using stats to draw useful conclusions well illustrates the general points of credible, stats-based reasoning. See also Thomas Sowell’s Basic Economics, many of his arguments are quite solid and he doens’t just blindly trust any published statistic as though the government is an unquestionable source of truth. Another book that makes interesting use of stats is Freakonomics (and its sequels). While Levitt is far more credulous of official statistics than I could ever be, his style of reasoning illustrates a variety of sound reasoning techniques that utilize statistical data. Yet another book worth mentioning in this overall category is The Black Swan by Nassim Taleb. There are many others worth mentioning, but these books are a great complement to the standard Austrian reading-list because Austrian warnings against stats-based arguments can (inadvertently) come off as strictly prohibitive. Statistics are not inherently bad; they’re just usually worthless at best, and positively misleading, at worst.

All of these things are (up to a point) publicly available data, but the sources are often diffuse and no normal human being has the combination of nerdy technical expertise and desire to compile it in a loose-but-coherent structure.

This is one of the areas where LLMs can actually be useful (slogging through heaps of finnicky data).

“Here’s what happened to areas that had housing booms at the outbreak of the great recession. This is how Covid affected cheap-housing areas. This is what has actually happened to American income earners at different percentiles, who they are, where they live, and what they do…” and so on.

I think this is practically blue-sky market because very few people have the chops to wrangle the truth out of the data. The data is there, yes, but it’s a lot of work to find real signal in the noise. Statistics fallacies are the bread and butter of most commentators and audiences eat them up like candy. You’d need to both do the stats, and explain why you are not blindly going with “what the stats seem to suggest”. Perhaps in the tone of a Veritasium video. “You might think that XYZ implies ABC, however, things are not so simple…” Take the audience on a journey of discovery. Teach them how to use the principle of adversarial witness (organizations rarely exaggerate their weaknesses, or fail to advertise their strengths) to sift the stats with real signal from those which are just noise, as well as how to compel the stats to tell the truth even when they are trying not to. Common-mode signals can reveal truth even when it is being actively concealed by multiple parties who are telling independent lies on top of the same underlying truth (thus inadvertently canceling out each other’s lies and revealing the real truth). And so on, and so forth. Using these principles, and more, you can tease the truth even out of some adversarial statistics, such as those published strictly for the purpose of propaganda.

I am deeply aware of the Austrian critiques of real world statistics, both theoretical and practical. Theoretically all data is just data and is silent to the point of being misleading without theory. Many essential variables, like genuine human preference rankings or willingness to pay, are totally unobservable even under ideal practical conditions. Practically, the unemployment rate doesn’t tell you anything about skills mismatch. Survey data is subject to any number of respondent errors. Some errors could potentially be politically motivated…

Another way that Austrian commentators make use of real-world data is on a hypothetical basis. “If you believe in statistics, then this is also what you should believe…” In other words, use their statistics to show XYZ, but qualify it as a hypothetical – “Even your own statistics show that XYZ”. It’s not so much that you’re telling the audience to believe the statistics, as you are telling the audience that, if they believe in statistics (even if they would be better not to), then XYZ follows from their statistics. The key to this approach is not to get carried away yourself with cherry-picking or confirmation bias; it’s just a technique to try to splash some cold water in the faces of the sleeping masses and wake them from their slumber.

all the data indicates the same thing from different sources

Yes, this is especially effective for stats from truly adversarial sources, i.e. if Ukraine and Russia both agree on XYZ statistic (e.g. numbers of deaths), then that’s probably a real statistic.

(annual census data on the number of firms in a county taken from tax data is kind of obscure and dangerous to lie about/hard to and therefore open to very little sampling error).

Precisely. Where there is little incentive to lie, the stats are more likely to convey real information about the world.

I couldn’t tackle much in the way of explicitly Austrian questions with this in a compelling way with most of the data I’m looking at (money supply aggregates are national, differences in state policies are often minor and often correlated with other policies or conditions (New York City can tax the hell out of people and they won’t leave because it’s New York, and we can’t observe the counter-factual anyway)), so maybe this is the wrong group of (theoretically existent) econ nerds to ask, but anyway, is this in any way an interesting idea or is it just a good way to put people to sleep?

I think one group that is really open to Austrian thinking is the Bitcoin/crypto crowd, especially the younger ones (Gen-Z, Gen-alpha). The original investors are pretty uniformly anti-Fed but the younger crowd is less likely to understand why that matters (may seem to them like a weird superstition). Exhibiting good Austrian thinking habits (which is basically just good thinking methodology, in general) would probably be appreciated by a sizable audience of such people, and you could attend crypto conferences/etc. to promote your podcast or, at least, advertise into that segment to effect.