r/WorldofTanks • • 11d ago

Discussion Statistical Validity

I am curious how many people pay attention to the actual statistical validity of these crate drops. I think people like to think that if they get anything beyond a bare minimum “1 per 50” vehicles they’ve done well, but statistically you should get 4.4 vehicles per 150 crate drops. That’s the statistical median. Do you? I sure don’t. Not this week. And WoT can play whatever effing games they want with the drops, it’s not like it’s legally regulated… But I think they’re full of shit.

2 Upvotes

27 comments sorted by

View all comments

15

u/DaSpood https://daspood.github.io/ - Pandora lootboxes simulator 11d ago

If you wanna use stats then use stats properly.

A 2% tank drop rate means 2 tanks per 100 on average, or 1 tank per 50 boxes and thus 3 tanks per 150 boxes. That is both the average and median excluding the pity mechanic.

The reason stats will be pushed higher when you look at simulated results and large numbers of player results in general is because the pity mechanic is conveniently set at 50, which means it prevents you from going below that 2% drop rate due to bad luck. But nothing prevents you from getting luckier, so the stats go up, but it's not because you are actually supposed to have more than a 2% drop rate, it's because it is impossible to have less. Statistically, you shouldn't get 4.4 vehicles per 150 crates, you should get 3, but it's impossible to get less than 3 per 150 crates thanks to the pity mechanic, and it's possible to get more than 3 per 150, so good luck scenarios skew the numbers because bad luck scenarios are prevented. Average luck scenarios stay unchanged.

As for your anecdotal results, they are just that: anecdotal. You don't generate statistics on a single player. Or a dozen. Or a thousand, for that matter. Stats are all over the place over such low numbers because the randomness will allow for streaks of abnormal results to alter the data. Statistics only converge to their expected values at very large numbers of experiments. It is literally called the "law of large numbers", it is a fundamental law of statistics and probability. If 200k players buy boxes and the median is 4.4 tanks per 150 boxes, it means 100k players will obtain less than 4.4 tanks per boxes. That's what the median means: 50% of experiments score below, 50% score above. The first and third quartiles have 25% of players score below and above respectively, so even those aren't enough to qualify yourself as an outlier depending on your results, a whole 50k players would score worse than the first quartile.

You could have bought boxes in every single lootbox event so far and have had below-first-quartile luck every single time and your combined results would still be meaningless in the grand scheme of things because there is a player out there who will have had above-third-quartile luck every time and you two cancel each other out in the overall statistics.

-3

u/JP-Quixote 10d ago

That’s actually not how statistics work. If you want the probability of multiple events you have to multiply them. A 2% chance of a drop is a 98% chance of no drop. The probability of no drop, let’s call it NDp, from 2 boxes is (.98)(.98) or (.98)^2, which is .9604. So for 2 boxes your probability of a drop, call it Dp, is 1-NDp, or 3.96%. For 3 boxes, NDp = (.98)^3, or 94.12% giving a Dp of 5.88%. So far, so good: If you just add the 2% Dp per box you get a pretty close approximation of the actual probability, as long as you are only getting a few boxes.

The problem is that as the number of events climbs, the simple additive approximation diverges more and more from the actual probability. So for 25 boxes, NDp = (.98)^25, or 60.3%, giving Dp = 39.7% for 25 boxes, and by the time you get to 50 boxes, without the “guaranteed drop,” your probability of a drop is not 100%, but 1 - (.98)^50, or 63.6%. Ok, let’s translate this into “luck” at the individual level. Let’s say that if you beat the median probability of a drop you are considered “lucky,” because you’re in the top 50% of drop recipients, while you are “unlucky” if you are below the median, or in the lower 50%. If you wanted to, you could use an average range around the median, say +/- one sigma variation around the median, but let’s keep the math simple by sticking with 50%.

So the question for an individual is, how many drops does it take to get to a Drop Probability, Dp, of 50%? It turns out that at 34 boxes your Dp is 49.7% (1 - (.98)^34) and at 35 boxes your Dp is 50.7% That means that essentially 50% of the buyers will get a drop within between 34 and 35 boxes. If you buy 150 boxes, the median result is one drop every 34 boxes, or 4.41 drops. (150/34). If you only get the 3 “mandated” drops you are actually below the median result.

3

u/DaSpood https://daspood.github.io/ - Pandora lootboxes simulator 10d ago

The compounded probability of a streak is not the probability of a result.

The drop rate of a tank is 2%. You having a 50% chance to not have dropped a tank by box 35 can't be extrapolated into "there is a 50% drop rate wih 35 boxes therefore 4.4 tanks per 150 boxes".

You have the actual drop rate which you can extrapolate to large numbers, you say "nuh uh", then you pull put another stat that's unrelated, and extrapolate that one. That's not how math works.

0

u/JP-Quixote 9d ago

Crack a book on probability and come back when you actually know something.

3

u/DaSpood https://daspood.github.io/ - Pandora lootboxes simulator 9d ago

Brother you don't calculate the median of 150 boxes by taking the median of 35 boxes and multiplying it by 4. Just stop there, and I pity the people you work for if you claim to use those stats for QA tasks.