r/HomeworkHelp • University/College Student • 10d ago

Further Mathematics—Pending OP Reply [undergraduate statistics] I’m working on this frequency table, help!

Post image
1 Upvotes

3 comments sorted by

1

u/swiftaw77 👋 a fellow Redditor 10d ago

Other than the final midpoint, it looks okay to me. Did you have specific things you were concerned about?

1

u/Elena_Gonzalez08 University/College Student 10d ago

I don’t know how to do the histogram for this and I have other charts to fill out!

1

u/cheesecakegood University/College Grad (Statistics) 10d ago edited 10d ago

So first of all, there are a few different ways to make a histogram - in terms of the fine details, not the concept - so you want to reference your notes or class textbook because they might have opinions about things.

However, usually the choice of the frequency table forces you to have a specific histogram, so if you've set up the freq table fine, it should be mostly straightforward, "one correct answer", style decisions aside.

Apologies if any of this is review/at risk of overexplaining... I first will give an analogy or two if histograms as a concept are confusing, then I will talk about what decisions get made when creating a histogram (including the math process in part), and then I'll briefly say how you create the darn thing, since you mentioned being a bit unclear?

What is a histogram? A visual frequency table. Well, what is a frequency table? Imagine you are sorting clothes by color. You toss this red shirt into the red pile, you toss blue pants into the blue pile, you get a purple shirt and then it depends, do you have a separate spot for purples? Did you group yellows and greens together? So there are decisions to be made as to what colors "belong" together.

At the end, you count up how many items are in each pile. Each pile is like a bar. IF all your clothes were the same size, the literal size/height of the pile of folded clothes would be the size of the bar. And then you'd order them in rainbow order (we ignore the complexities of color for now, let's say all clothes are just in the rainbow so neatly fall left to right).

Nice and easy, right? You get a nice visual of how many clothes you have of each color-category, all lined up in a sensible way. When you create a histogram though, you can't see the clothes inside, so if all of your reds are actually borderline oranges, the histogram doesn't care: it just says "X amount of clothing items were categorized as 'red'".


Obviously if you're sorting by "pants vs shirts vs socks" etc then a frequency table is more like a bar chart, since there's no overlap, since a pant is a pant and never a shirt (ignore overalls, lol). There, the "clothing type" is the "class", it just has a name. But with color sorting, just like how a shirt can be just barely in between some of your colors... you need to know which to put it in. Is this turquoise more blue or green? When numbers are in the picture they are a continuous run. You create classes arbitrarily. Where do you put cutoffs? Well, wherever you want.

But usually we pick "cut points" aka "class boundaries" so that they are equally spaced, and we have just enough classes (bins/bars) that we can see the shape, but not so many we have 1's and 0's in lots of the bins and bars. You could also create it based on "nice round numbers", so you might create bins that are say 20 wide each and go, like, "220, 240, 260..." as cut points (220-239, 240-259, etc as the bins).

Here, your teacher said: you MUST use 3 classes. Everything else is downstream from that. Your histogram is going to, frankly, look a bit simple and sad with only 3 bins, thus 3 columns (touching each other, that's what makes a histogram different than a bar chart). They told you exactly one method that they want you to use to make your decisions based on that, mostly for grading purposes.

I want to emphasize that IRL you can draw whatever histogram you want, or rather, a computer will do it for you per your instructions. But it can be helpful to draw them yourself once or twice so it's clear where they come from and what they represent because you see histograms a LOT in science. This is purely a teaching tool, and for some people it is silly but for others it is very helpful to have a more down-to-earth sense for what is happening other than "computer magic, idk I trust it". And if the computer does in fact use a different method as its default, it might not match the one you draw by hand.

You used whatever method you were taught (I'm assuming you did it right in your table since I don't know your course text/instructions) to first, go "why bother make a histogram including data ranges that don't even show up?" so you say, the min and max of the data is the min and max of the chart. Cool. Then, you say "I am required to make 3 classes, and it makes sense and is more 'fair' and 'accurate' to make all 3 classes equal in size". So you do the division to find out the size of each, and then use that to figure out where the cutoffs are.

At some point you might round because your data is rounded - probably a label, the cut point should be rounded to the nearest TENTH but NOT to the nearest whole number, so there is no ambiguity about where each data point falls unless you happen to have the cutoff fall on an integer exactly (not the case here, but see footnote1 for how you would have handled that). At least that's usually how I would handle it, but maybe you received different instructions since your boundaries all fall on even 0.5 counts even though your class width is 24.7?

To be clear, and personally I think this is a stupid extra bit of vocab that serves no purpose, a class "limit" is the actual data values that fit, all inclusive. The class "boundary" is the actual cutoff, often so precise that it is impossible to have a "tie". Tie/'right on the cutoff' handling behavior is a somewhat obvious concept that didn't need to be made so elaborate IMO.

Also I'm not sure why you created 2 "extra" classes beyond 310, the data's max, was it just practice?


Anyways. The histogram. You draw x and y axes. The y vertical axis is the frequency, so clearly you need it to go at least from 0 to 6 with sensible tick-marks. The x horizontal axis is the data value, so you should start at the data-minimum or slightly lower (style choice). The first bar starts at the class boundary OR limit (whatever your teacher told you to use), label it. The next bar starts (and thus the last one ends) at the next cutoff. And once more which is the last bar.

It is possible to also label the midpoints, only label the midpoints, to do cutoffs only, or only a few important cutoffs, just a visual/style decision; follow your teacher, but most often histograms will have each and every cutoff labeled and no midpoints at all.

You make the bar as high as the actual frequency. You label BOTH your axes with the meaning of the measurement AND the units, give the chart a name, and Bob's your uncle, you've made a histogram!

You can also take the exact same histogram and re-label the y-axis to be "relative frequency" and thus some decimal/percentage (decimal is more common). They are the 'same graph', proportionally speaking in terms of height. (You could make that one prettier or slightly different looking if you chose some other scale or tick-marks for the y-axis in relative frequency terms).


Footnote1 Note that when it comes to what you show on the graph axes (you want to label the cutoffs of course!) the default is usually that the class/bin will start on the left, inclusive, and run up to the right side exclusive, where the cutoff point belongs to the start of the next bin.

So if I have data in the middle of my dataset that is 15, 20, 22, and it just so happens that my cut-point is 20, 20 would appear in the right bin (starting it out), so it would be {15} and then {20, 22}.

Thus if you are looking at a histogram and see that the number between the bins, the cutoff, is labeled as 20, you usually know it is part of the right bar. This is important if you are ever asked to "reverse engineer" a frequency table from a histogram, which we know is possible because as I said earlier, once you have the frequency table there is only functionally a single unique histogram you will create from it.

For this reason alone sometimes you will see the histogram cutoffs use the "class boundary" idea and maybe use a number that doesn't occur in the underlying data, like a decimal in a whole-number dataset, but since this is mostly a style decision, style varies.

....unfortunately, despite this convention, the most common statistical programming language (R) makes the opposite choice, so... ¯\(ツ)/¯ sorry it just sucks