The best explanation I found of the theorem was from the book by Bertsekas and Tsitsiklis:
-> You have a resulting event B in front of you. Any of {A1,...,An} causes (all of them mutually disjoint) could have led to B. Now, you are aware of how likely each of these {A1,..,An} is of producing this B (this is P(B|A)). Bayes theorem allows you to use this information to deduce which {A1,...,An} is most likely to have been the cause given that B occurred (P(A|B)).
i.e. you use the Bayes theorem to reverse the conditional probability relationships given in the problem. The exact expression is now easily derivable using this idea and the total probability theorem.
-> You have a resulting event B in front of you. Any of {A1,...,An} causes (all of them mutually disjoint) could have led to B. Now, you are aware of how likely each of these {A1,..,An} is of producing this B (this is P(B|A)). Bayes theorem allows you to use this information to deduce which {A1,...,An} is most likely to have been the cause given that B occurred (P(A|B)).
i.e. you use the Bayes theorem to reverse the conditional probability relationships given in the problem. The exact expression is now easily derivable using this idea and the total probability theorem.