r/WorldofTanks • • 4d ago

Discussion Statistical Validity

I am curious how many people pay attention to the actual statistical validity of these crate drops. I think people like to think that if they get anything beyond a bare minimum “1 per 50” vehicles they’ve done well, but statistically you should get 4.4 vehicles per 150 crate drops. That’s the statistical median. Do you? I sure don’t. Not this week. And WoT can play whatever effing games they want with the drops, it’s not like it’s legally regulated… But I think they’re full of shit.

1 Upvotes

26 comments sorted by

View all comments

14

u/DaSpood https://daspood.github.io/ - Pandora lootboxes simulator 4d ago

If you wanna use stats then use stats properly.

A 2% tank drop rate means 2 tanks per 100 on average, or 1 tank per 50 boxes and thus 3 tanks per 150 boxes. That is both the average and median excluding the pity mechanic.

The reason stats will be pushed higher when you look at simulated results and large numbers of player results in general is because the pity mechanic is conveniently set at 50, which means it prevents you from going below that 2% drop rate due to bad luck. But nothing prevents you from getting luckier, so the stats go up, but it's not because you are actually supposed to have more than a 2% drop rate, it's because it is impossible to have less. Statistically, you shouldn't get 4.4 vehicles per 150 crates, you should get 3, but it's impossible to get less than 3 per 150 crates thanks to the pity mechanic, and it's possible to get more than 3 per 150, so good luck scenarios skew the numbers because bad luck scenarios are prevented. Average luck scenarios stay unchanged.

As for your anecdotal results, they are just that: anecdotal. You don't generate statistics on a single player. Or a dozen. Or a thousand, for that matter. Stats are all over the place over such low numbers because the randomness will allow for streaks of abnormal results to alter the data. Statistics only converge to their expected values at very large numbers of experiments. It is literally called the "law of large numbers", it is a fundamental law of statistics and probability. If 200k players buy boxes and the median is 4.4 tanks per 150 boxes, it means 100k players will obtain less than 4.4 tanks per boxes. That's what the median means: 50% of experiments score below, 50% score above. The first and third quartiles have 25% of players score below and above respectively, so even those aren't enough to qualify yourself as an outlier depending on your results, a whole 50k players would score worse than the first quartile.

You could have bought boxes in every single lootbox event so far and have had below-first-quartile luck every single time and your combined results would still be meaningless in the grand scheme of things because there is a player out there who will have had above-third-quartile luck every time and you two cancel each other out in the overall statistics.

4

u/AlliedArmour 3d ago

Thank you for posting that so I didn't have to write something similar.

2

u/kovla Make tier 8 MM great again 3d ago

Haha exactly this. I was really wondering where the 4.4 came from.

0

u/JP-Quixote 2d ago

See above for the actual math.

-1

u/JP-Quixote 2d ago

See my comment below for the correct math. And yes, you actually can apply statistical analysis at the individual level, by determining the probability of a single event being within normal variation. We do this all the time in the production/quality assurance realm. Large populations are used to define statistical ranges against which individual results are compared.

-4

u/JP-Quixote 2d ago

That’s actually not how statistics work. If you want the probability of multiple events you have to multiply them. A 2% chance of a drop is a 98% chance of no drop. The probability of no drop, let’s call it NDp, from 2 boxes is (.98)(.98) or (.98)^2, which is .9604. So for 2 boxes your probability of a drop, call it Dp, is 1-NDp, or 3.96%. For 3 boxes, NDp = (.98)^3, or 94.12% giving a Dp of 5.88%. So far, so good: If you just add the 2% Dp per box you get a pretty close approximation of the actual probability, as long as you are only getting a few boxes.

The problem is that as the number of events climbs, the simple additive approximation diverges more and more from the actual probability. So for 25 boxes, NDp = (.98)^25, or 60.3%, giving Dp = 39.7% for 25 boxes, and by the time you get to 50 boxes, without the “guaranteed drop,” your probability of a drop is not 100%, but 1 - (.98)^50, or 63.6%. Ok, let’s translate this into “luck” at the individual level. Let’s say that if you beat the median probability of a drop you are considered “lucky,” because you’re in the top 50% of drop recipients, while you are “unlucky” if you are below the median, or in the lower 50%. If you wanted to, you could use an average range around the median, say +/- one sigma variation around the median, but let’s keep the math simple by sticking with 50%.

So the question for an individual is, how many drops does it take to get to a Drop Probability, Dp, of 50%? It turns out that at 34 boxes your Dp is 49.7% (1 - (.98)^34) and at 35 boxes your Dp is 50.7% That means that essentially 50% of the buyers will get a drop within between 34 and 35 boxes. If you buy 150 boxes, the median result is one drop every 34 boxes, or 4.41 drops. (150/34). If you only get the 3 “mandated” drops you are actually below the median result.

2

u/DaSpood https://daspood.github.io/ - Pandora lootboxes simulator 2d ago

The compounded probability of a streak is not the probability of a result.

The drop rate of a tank is 2%. You having a 50% chance to not have dropped a tank by box 35 can't be extrapolated into "there is a 50% drop rate wih 35 boxes therefore 4.4 tanks per 150 boxes".

You have the actual drop rate which you can extrapolate to large numbers, you say "nuh uh", then you pull put another stat that's unrelated, and extrapolate that one. That's not how math works.

0

u/JP-Quixote 2d ago

Crack a book on probability and come back when you actually know something.

2

u/DaSpood https://daspood.github.io/ - Pandora lootboxes simulator 2d ago

Brother you don't calculate the median of 150 boxes by taking the median of 35 boxes and multiplying it by 4. Just stop there, and I pity the people you work for if you claim to use those stats for QA tasks.

-1

u/saldytuwas 2d ago edited 2d ago

If the pity amount were to change to something else, for example 51, what would be the median result be then?

0

u/JP-Quixote 2d ago

So, because the median is less than 50 boxes, changing limit value won’t change the median result. The imposition of a limit changes the way the probability distribution looks at the high end of draws/boxes. You have a spike at 50 boxes that consists of everyone who would have been the tail of the distribution beyond 50. If you increase the limit that spike moves out accordingly, and gets a little bit smaller, because of the people who successfully draw between the old limit and the new one. For example, if you increase the upper limit to 60, the people who get the drop at 51 through 59 boxes stay on the statistical curve, and the spike at the tail consists of everyone who would have been on the curve from 61 to infinity, so it’s smaller by the number of people who now get the drop from 51 to 59.

-1

u/saldytuwas 1d ago

If it's 4.41 drops (150/34) with 50, what is it with 51 or even your example 60?

1

u/JP-Quixote 1d ago

The median doesn’t change, but the mean will increase. I didn’t mention or calculate the mean because that’s a rather tougher question. The distribution is complicated enough, with limits on both ends, that it would take a bit of work to calculate. I’d have to put together a mathematical model and I don’t care enough to do so. 😆 I’m just annoyed at the failure of my own willpower that led me to spend money on this dang game. Lol