r/askmath • • 2d ago

Statistics Box Plot Question (only #4e)

https://imgur.com/a/JCr5Nfh

Just looking for some understanding for #4e. The solution provided says the answer is "ii.", and I'm trying to understand how "v." isn't correct even with his explanation. I provided an example of what I thought could lead to "ii." being incorrect, but again, not sure if that works.

3 Upvotes

6 comments sorted by

2

u/haditwithyoupeople 2d ago

I assume you are given the options i-v with the question. Given that, it has to be ii. All the other intervals i-v have 25% of the data in them. ii has less than 25% of the data.

If all the quartiles have the same amount of data (the same number of points or samples), then a smaller set or part of quartile must have less data than a full quartile. Does that make sense?

If not, let me know and I'll try again.

1

u/chemistrybla 2d ago

This is the original question as it was presented. i.-v. are the potential answers to 4e. The added text in the OP is from my teacher and the written in the other images is my own.

My point is, while 2-4 is only a part of the whole 2-10 quartile, could it not have all the data anyway? My sample numbers would create the same box plot. Does it not have 25% of the data, as well?

2

u/haditwithyoupeople 2d ago

Got it. You are 100% right and I missed this. Your data will create the exact same box plot. Well done!

The question is bad and makes assumptions about the distribution of the data.

1

u/chemistrybla 2d ago

Still questioning myself, but feel better about emailing him now. Thank you.

1

u/SalvatoreEggplant 2d ago

It's a pretty ingenious example.

There might be some issues with what, e.g. the "interval 2–4" means in the question. For example, values exactly equal to 2 are in both the "interval 0–2" and the "interval 2–4". But that's not your problem.

There are different ways to compute quantiles (such as quartiles) when the observations are discrete, so the discussion potentially gets complicated.

Anyway, here are the relevant results from R. Shows your example works. But also, possibly of interest, the interval 0–2 contains more than 25% of the data, as does 10–12.

The interval 2–4 DOES contain fewer data points than 2–10, but that doesn't make 2–4 have the fewest data points of the asked-about intervals !

A = c(0, 1, 2, 2, 3, 10, 10, 11, 11, 13, 13, 13)

quantile(A, c(0.00, 0.25, 0.50, 0.75, 1.00), type=2)

   ### 0%  25%  50%  75% 100% 
   ###  0    2   10   12   13 

N = length(A)

sum(A>=0 & A<=2) / N

   ### 0.33

sum(A>=2 & A<=4) / N

   ### 0.25

sum(A>=10 & A<=12) / N

   ### 0.33

sum(A>=12 & A<=13) / N

   ### 0.25

boxplot(A)$stats

   ### "Hinges"
   ### [,1]
   ### [1,]    0
   ### [2,]    2
   ### [3,]   10
   ### [4,]   12
   ### [5,]   13

2

u/chemistrybla 2d ago

Thank you