Bayesian Probability? Convince me...

I was wading through some game theory works when…

Bayesians argued that, as E.T. Jaynes said, that the problem addressed by their method is non-experimental or never requires verification, because the goal is different from Gauss - R. v. Mises. view of probability:

“we are estimating, from our prior information and data, the unknown constant value that the parameter had when the data were taken”,

whereas the frequency probability’s goal is

“from prior knowledge of the frequency distribution of the parameter over some large class C of repetitions of the whole experiment, the frequency distribution that it has in the subclass C(D) of cases that yield the same data D. The problems are so different that one would expect them to be solved by different procedures”.

But does subjective probability, one application of Bayesian probability, make sense? Is it really kosher per se?

Suppose a used car I’m offered, I think, is 30% lemon, 70% good, given my own thoughts about possible combinations of parts that it was repaired with: but I don’t know what class the car belongs to. So why should I think combinatorials would help me out, and make a decision using expected utility given the 30%, 70% split in the game tree. If I don’t know the class, then I don’t know if I’m aware of all the possible combinations of parts.

It seems to me Hans-Hermann Hoppe’s argument stands, even though Richard Langlois has in the past argued Bayesian probability to be appealing, perhaps, when discussing entrepreneurship (such as one person’s estimates, due to his better information, changing his decision).

It seems to me the idea of using Bayesian probability in game trees to be deceptive.

Any thoughts?

I don’t know, but you might ask on LessWrong - they love to help with Bayesian theory. This article may explain it.

I love probability theory. Both the Ludwig von Mises & Bayesian theories are correct as far as they go. What, specifically, is your question?

sidenote: I suggest that you link references to your quotes, so that others can see the full context of the quote.

@ AJ. The article was interesting, but once again the comments generated quite a bit of debate.

@ Stephen:

Q: Consider in game theory, the probabilities of the behavior of Player 2, P → action X and 1- P → action Y, are said to be supposed by Player 1. How are they arrived at by Player 1, if this is a unique event?

They are not frequency. But they cannot be combinatorial, unless Player 1 can construct deductively the purely formal structure of all Player 2’s expectations.

If we follow Kirzner, this is impossible, given some ignorance of the attributes of Player 2.

If we follow Lachmann, expectations are so elastic, that we cannot construct a formal structure to calculate our combinatorial.

If we follow standard game theory convention, the probabilities are just supposed… no answer how, if it is a Bayesian Subgame Perfect Nash Equilibrium (with information set on player 1’s reaction to Player 2), not a standard mixed Nash Equilibrium.

If we follow Mises, if there is a probability, then it can be broken into pure syllogisms, if we can define the structure of the causes into several more precise terms (but he argues we can’t do it, in which case we can’t discuss anything, because we have neither a class nor a set of syllogisms).

All that can be said, given player 1’s total ignorance of player 2, is that it is just as likely for player 2 to take action X as to not take action X and just as likely for him to take action Y as not take it. And they are combinatorial. In fact, I think all probabilities can only be measured through combinatorics.

That is exactly the problem.

This fact, which R. v. Mises in Probability, Statistics, and Truth traces to Laplace (but in my view, De Morgan–and Jevons defending De Morgan–have also credit, unfortunately, since their logic is otherwise highly superior to Whitehead and Russell and Keynes).

Suppose I know, for a fact, also, that Player 2 is biased one way. I am entirely uncertain which way, however. I cannot, then say, 50% chance biased one way, 50% chance biased another way, as then I would once again find my probabilities for each side equal, but knowing that they are unequal, I would again assert 50/50, leading me back to where I started, with the datum of lopsidedness.

In game theory, it is often said, then: 30% X and 70% Y. But how was the exact number, 30% and 70% arrived at, by the subject.

Surely, it cannot be subjective in the sense of preferences: I prefer Kiwis and you prefer Strawberries, so I say, ‘What is in the bag is 90% of the time Kiwi and 10 % of the time Strawberry; this is my belief; it is so, because I prefer it to be so.’

Obviously, this is not what is meant by subjective probability belief in the case of perfect uncertainty given only the datum of lopsidedness. There is a process, but what can it be if it is not empirical?

It’s funny, I recently bought Keynes’ and RvM’s books but haven’t read them yet. I haven’t even heard of Whitehead, Russel, or Jevons. I’m not sure if I’m qualified to have this discussion, but I’ll proceed as best as I can.

Have you been following Crovelli’s arguments recently? He wrote a paper attacking the frequentist definition of probability for the Libertarian Papers recently. He also gave a lecture at this years ASC. I agree with his criticisms, but not his conclusions.

I think (from what I understand, based on second hand information) that RvM is totally wrong about probabilities being a physical property of the phenomenon in question. Crovelli is right, that such a view is only possible in an indeterminist universe.

Crovelli goes on to define probabilities as subjective. This is clearly wrong however because probabilities are not completely arbitrary. IMO, they are an objective measure of an actor’s certainty of alternative observable outcomes of events which fit the LvM definition of class probability.

Exactly, still 50:50.

Prior experience, or imperfect (probabilistic) knowledge of the causal factors which affect the outcome.

I totally agree that its not subjective.

Are you familiar with the binomial and multinomial theorems?

Yes. The binomial theorem is the argument that Jevons gave in his two main works in defense of the 50-50 rule. Problem is that Jevons broke his own rule and explained probability after number, so he didn’t actually prove in his own logic, that the 50:50 rule is valid. He used…numbers.

But the problem is that the formal nature of the problem must be entirely developed: material propositions tell us how many outcomes are possible. If we lack the required material propositions, how can we guess what n or m is in the combinatorial? What if n and m are indeterminate given our knowledge?

On the 50:50 note, furthermore, we are back to where we started–a contradiction because we keep getting 50:50, applying our 50:50 rule to our two premises: uncertain whether Y and not- Y and certain that lopsidedness occurs:

Suppose ~ means difference and = means identity. a = not-A, b = not-B. Two symbols together means combination.

Premises: A = P = 1 - P, B = P ~ 1 - P.

Predicates: A = Ab, B = Ba

Inference: AB = AbBa = AaBb = (Aa = o)(Bb = o) = o

I’ve been trying my hand at Jayne’s 2003 textbook and prove the stuff in Jevon’s logic, but everything reduces basically to the above. Perhap’s I’m formulating the problem wrong? Edwin Jayne has great explanations, as did Jevons, but I can’t seem to prove them.

Perhaps there is another way to formulate the problem?

[[Sidenote:

He’s not correct, I would argue. Mises in Theory and History said we can always break down probable events about the real world into simple syllogisms but with two or three or more causes instead of the one attribute. Walter Wiemer and Hayek also agreed on this in a conference (Wiemer’s paper, I think, had a similar knowledge explanation of the quantum cat). The world is not deterministic, rather, because of Roger Penrose’s silver-mirror, wall, and photodetector argument that time flows one way contra Hawkings.

Alternatively, if you disagree with Penrose about his own theory, the other explanation offered by David Deutsch is that is fact that the world is probablistic in reality, but in the sense that the one attribute and only one attribute has many outcomes, all realized in multiple universes at the same time (so probability refers to our limited sensory knowledge of only one real world).]]

Reading Mises again, I am now very sure he addresses case probability in a very different way from R.v. Mises, but also contrary Bayesian probability:

Uncertainty concerns single events. Mises agreed, “Case probability is not open to any kind of numerical evaluation” (Mises 1949:113). But he argued that unmeasurable does not me cannot expect anything about single events, on the condition that they are not part of any class. Sometimes, we can certainly say something qualitative; the uncertainty being quantitative. *“*Such problems are not open to any elucidation other than that provided by understanding” (Mises 1949:115). But what did he mean by understanding? He meant: “apodictic certainty”, that is synthetic a priori—application of a synthetic a apriori material proposition to another material proposition, which gives a certain prediction—a material predicate from material premises (Mises 1949:117). The quantitative uncertainty, in those instances where case probability exists, arises purely from the fact the our material premises are indefinite or we have not combined all material premises required for case truth: class probability i.e. ‘case like-truth’, but not necessarily case truth. "Understanding, by trying to grasp what is going on in the minds of the men concerned, can approach the problem of forecasting future conditions” (Mises 1949:118).

“Understanding is always based on incomplete knowledge. We may know the motives of the acting men, the ends they are aiming at, and the means they plan to apply for the attainment of these ends. We have a definite opinion with regard to the effects to be expected from the operation of these factors. But this knowledge is defective. We cannot exclude beforehand the possibility that we have erred in the appraisal of their influence or have failed to take into consideration some factors whose interference we did lot foresee at all, or not in a correct way” (Mises 1949:112). Gambling, or acting without expectations about single events, involves class probability. Betting or, or acting with expectations about single events, involves case probability.

Case probability is either entirely unknown, or if a synthetic apriori framework exists, Mises meant we can plug in material propositions and get a certain output–not a probability in either the frequency or Bayesian sense at all–which is quantitatively indefinite, either because based on indefinite logical relations or because we lack knowledge of all material propositions to plug into the framework.

Economics = case probability. Very strange these statements or the discussion of probability prior marginal utility didn’t catch my eye as this important when I first read Human Action.

What do you think?

This issue has puzzled me, though I can’t say I have expertise in probabillity theory. I do remember when covering the subject in my applied maths classes a few years ago that the maths lecturers never spoke about “Bayesian probabillity”, but called what they were talking about degrees of belief. This was even before I was an Austrian, and I still found what they were talking about quite laughable. Probabillity as far as I knew it was confined to either complete knowledge of the distribution of known sets, or defined in the limiting tendency of a set as it’s iterations are alowed to stretch to infiinity. Obviously, this knowledge is incomplete and bears a degree of uncertainty, but so does all empirical knowledge; this doesn’t necessitate that it is somehow subjectively flawed or contradictory.

Incidentally, I came across things like the Hurcwitz criterion, without even realsing the guy was a game theorist. Maybe I’m mistaken though, and as you suspect there is a distinction between degrees of belief and Bayesian probabillity, such that the latter can be computed for the kind of events dealt with in economics in an even close to approaching rigorous way.

Also, have you read any of Kolmogorov’s work in probabillity? As far as I know he had some differences, with though is generally recognised as the intellectual successor to R. Von Mises on probabillity. I haven’t got round to reading him yet, though the guy is interesting for a whole host of other work he did too.

I remember cringing when listening to his comments on quantum mechanics and Heisenberg; I think he has grossly misunderstood the Copenhagen interpretation of quantum mechanics, though he did otherwise make an interesting argument. When I have more time, I might even write a reply to his paper, if I get the chance to study the subject more thoroughly.

Take the flipping of a coin as an example. You don’t know if the coin is fair or not. The only way to be absolutely certain of what the heads/total ratio is, is to flip it an infinite number of times. This is impossible to do. Flipping it a large number of times and observing the result can only give you an estimate. There is no way to know if how much your estimate will change if you keep flipping. How certain can you be of any ratio given your observable results? I want to suggest that there is a complete answer to this problem.

Let p represent the ratio of heads/total, n represent the number of coin flips, k the number of observations of the coin landing heads up.

The probability of making a set of observations given that p is specified, P(k=K), is provided by the binomial distribution:

If a fair coin were flipped 10 times, what is the chance that it would land on heads only once, only twice, and so on.



Times it lands on heads (k)



Probability of this occurring



0



0.000977



1



0.009766



2



0.043945



3



0.117188



4



0.205078



5



0.246094



6



0.205078



7



0.117188



8



0.043945



9



0.009766



10



0.000977

It should be obvious that counting the ratio of heads/total through random sampling is only sufficient to give an estimate of the probability of the coin landing on heads. Although a larger sample can give a more certain estimate, whether a sample is large enough is decided subjectively.

Now, for your example, imagine a case where there are two coins. The probability of the first coin landing on heads is 0.3. The probability of the second coin landing on heads is 0.7. If one of the coins were selected randomly and sampled 5 times, observing that it lands on heads 4 times, what is the likelihood that it was the 0.3 coin that was selected?

First, the likelihood of getting the observed results is calculated for each coin.



Coin



Probability of observed results



0.3



0.036757



0.7



0.200121

The probability of each coin having been drawn can be calculated indirectly by divided the likelihood of observed results for that coin by the sum of likelihoods for all the coins.

And in the case were any observations have yet to be made, n = 0 and k = 0, the median is 50:50. So even if you know that the player has a mixed strategy at the outset but don’t know which way, it’s still 50:50.

I think that when it comes to humans actions are always case probabilities. It is impossible to know how someone will act because they have a free will and learn. The conditions which produced previous actions are always different from those which will produce future actions as his means, ends, and causal knowledge change over time.

A persons behaviour fall into the category of class probabilities. How a person walks, unless they put consciously aim to change the way they walk is always the same. The way a person reads, writes, and so on are generally constant.

In poker, you can never know if a player will try to bluff. You can only read the tell tale signs.

AJ: Thanks for the link, that was a very interesting article.

thelion: I am unsure of what specific complaints you have about Bayesian probability. I find their critiques of frequentist probability to be quite solid.

So far as I understand, even the standard example given by frequentists of a fair coin having a specified probability is a statement about a persons knowledge and not a property of the dice, and therefore subjective in the Bayesian sense.

Given complete knowledge of all factors involved including force of the toss, angle of descent, air currents and what have you, the probability of heads or tails is not 50% (even assuming it cannot land on its side). The notion of probability in such a case is not even applicable, as the result is known with certainty.

Probability is only meaningful in conditions of uncertainty, and uncertainty is a property of the map, not the territory. It is in this sense that Bayesians consider probability to be subjective, and constrained by the knowledge of the subject and not inherent in the coin or dice. This is also why it makes perfect sense to say that a rigged coin, without knowing to which side it is biased, has the same probability as one that is non-biased.

Having reread Mises, I would argue:

All consideration of uncertainty fall into case probability = economics, which is non-numerical, but were the attribute could be broken up into a list of certain syllogisms, each certain in the case that the material propositions hold.

What we are uncertain about is whether we have considered all relevant material propositions. You’re right, it is a map, but not the territory.

However, since the number of material propositions is infinite (Jevons) and the list of possible material suppositions is incomplete (Godel), no numerical proposition of the uncertainty can be made, without presupposing the answer to the number of material suppositions (which is impossible).

Hence Mises argues that numerical probability, if it has to do with knowledge possessed by the actor, is impossible. The actor has a certain, non-numerical belief, on which he acts. But the actor can be mistaken.

[So then, Mises presents praxeology as the manner of determining what the actor certainly expects, but can err about, hence certain action, but consciousness of uncertainty of results.]

On the other hand, class probability deals with the actual world, not a map, since the natural sciences do not care about the action by an actor concerning the world, which is determined by expectations; they care only about the real world. So that probability is numerical. As it is an ad hoc estimate of the real world taken as the real world (in the sense of Blanshard’s theory of the idea being an unrealized object).

Bayesian and frequentist probabilities describe different situations. Frequency-as-probability is suitable when the a priori distribution is not known but empirical observations of the random variable are available. Bayesian probability is suitable for reasoning about probability when the a priori distribution is known (prior to the advent of Algorithmic Probability by Ray Solomonoff, it was possible to play skeptic about the possibility of knowing an a priori distribution but Solomonoff’s work gives an objective basis for a priori distributions). One is not better or worse except to say that you can do everything with Bayesian probability that can be done with frequentist probability but not vice-versa, so that, the attempt by modernists in mathematics to eliminate Bayesianism and replace it solely with a positivism-compatible frequentism is futile.

Bayesianism is applicable whenever the question being asked regards one’s degree of belief. Frequentism is applicable whenever the question being asked is historical in nature. Please note that most mathematicians mistakenly believe that Hume’s argument regarding the unknowability of the future based on the past has never been answered. In fact, Solomonoff’s work answers Hume’s argument quite well and the Humean objection to Bayesianism on this basis is now defunct.

Clayton -

What use is there for absolute certainty? Nothing in life is certain, at best you can have conditional certainty (certain deductions from uncertain axioms). Also, we can place bounds on our uncertainty based on the number of times we flip the coin. The more times we flip the coin, the less likely it is we have incorrectly estimated the distribution on the coin, in fact, the probability that we have incorrectly estimated the distribution on the coin decays to zero exponentially quickly.

Clayton -

Excellent posts, Clayton. You are a very clear writer. Do you read LessWrong at all? I recommend it, its a fantastic site.

+1

I’ve found LessWrong to be a perfect sister-site to mises.org, even though it’s not particularly libertarian, because it elucidates and ferrets out all the little reasons why people believe the irrational things they do (such as the need for a State). The series A Human’s Guide to Words is required reading* in my opinion, not to mention all the heuristics/biases work, the great probability discussions, and standalone articles like Generalizing from One Example that are worth their weight in gold.

It’s a pretty nice community, too, and although not libertarian, where else will you hear a left-leaning person talk about their anti-gun position in terms like this?

*Because it demystifies so many of the threads here, why people argue over minor semantic points, and many other things I always find myself wanting to point out here on the forums. It also helps explain why statist and collectivist language has so much influence on people’s views.

How so?