There are two answers to this. The short answer is that causality is a category of human knowledge - causality is when we may be said to “understand the cause of a thing.” Of course, this only pushes back the problem one step: what does it mean to understand (the causes of a thing)?
The long answer is more circuitous and goes right to induction. When we say “A causes B” we mean that whenever the condition A is satisfied, the result B must then be the case, as well. In this sense, causality is a lot like logical implication: “If A, then B” except that there is some kind of implied “time ordering” where we mean that “If A at the moment prior, then B at the moment subsequent.” The only way to reduce the latter, klunkier expression to the more elegant logical implication is if information (world state) propagates at infinite speed through the world, something that appears to be obviously false.
The long-standing problem of induction asks how we know the laws that will govern future state of the world from observing the laws that governed the past state of the world. That billiard balls have always been observed to collide instead of passing through one another says nothing about whether billiard balls must always collide instead of passing through one another. Unfortunately, many philosophers have given up on the induction problem and simply adopted a kind of “inductive nihilism”… we really don’t know the future will be like the past!
However, this view is nowadays untenable. Ray Solomonoff’s theory of induction - even though it has its origins in advanced mathematics, rather than the more traditional metaphysics - shows the way out. I’ll state the theory in slightly incorrect layman’s terms to get the idea across. I leave further research to those interested.
Consider a digital image.
The image contains information about relative levels of colors to be displayed on a digital monitor or to be printed on paper. Compression allows us to represent the image with less “storage space”… that is, we can represent the same information in less space by removing redundancies. It turns out that there are foundational connections between this fairly mundane task of data-compression and the laws of logic. Basically, removing redundancy is the same thing as “giving a reason” for something. If there is a reason for something to be this way rather than that way, then this reason manifests itself as a pattern and every pattern entails redundancies! Hence, the process of removing redundancies is like “reasoning backwards.”
Solomonoff induction is like asking: what happens if the picture is being given to us bit-by-bit over time? Can we compress the picture as we are receiving it? The answer is, yes you can! And as you compress the bits of the picture as you are receiving it over time, you are “reasoning backwards” about the patterns in the picture. Hence, you can now assign meaningful probabilities to the portions of the picture that have not been yet received!
The “inductive nihilist” on the other hand, believes that half way through receiving the digital image, we are just as likely to see randomness for the remainder of the image as we are to see the image we “inuitively” expect to see:
The problem of induction is basically like asking: “If we have received half the digital image so far, which total image is more probable, the former or the latter?” The inductive-nihilist view is that they are equally likely… after all, in a radom lottery, each picture is equally probable along with every possible picture of the same size. But when you specify certain conditions (one of which being the existence of a Universal Turing Machine), then the rules of Solomonoff induction apply and we are no longer facing an “agnostic” probability distribution… the past does, indeed, tell us something about the future.
And this is the essence of causality. If the past tells us about the future, then of course the prior moment is related to the present moment which is related to the subsequent future moment.
Final note: when we ask “why is there time, why is there an unfolding of events?” what we are really asking is why information propagates at a finite speed. This is a separate question from causality/induction and should not be confused with it.
Clayton -

