Drawing an assumption before evaluating data is a really good approach. Ten points to Griffindor for that one.
They said their guess and the data differ the same way as my guess. But then they think the data is correct and they just throw out their guess. That's actually wrong. This point is very important in statistics, but you actually already learn it in 4th or 5th grade math classes. Make a guess about the end result and if your calculated results is quite different be sceptical about your calculated result. Our ability to guess is not good at getting the exact numbers right, but it's really, really good at getting the big picture. Therefore when data and guess disagree there's a really good chance the data is wrong and we must question it and not just our guess (which we can also question, but questioning the data is more important).
PS: It's quite funny how often we consume lots of data and calculate a lot and still the right answer to the question is: "Dunno yet"
At least for me that's what I learned in the last few years on HN.
The analogy doesn't hold. The point of the NYT exercise is to challenge our perceptions/biases, not challenge calculations.
Yes, in science, say, you should make a guess about the end result, based on something, like first principles, direct observation, etc. Then, if the calculated result diverges, check both.
In this case though there is no reason for the general public to have had a good picture of appropriate curve; it's a lot of hearsay and ideology which biases what we think it should look like, and it is right to trust that less than the actual data.
Isn't what you learn in 4th or 5th grade math classes that you should be sceptical about your own calculations if they diverge a lot from your estimate? When it comes to empirical data, I think the only sensible approach is to remain equally sceptical about both your guess and the way the data was compiled or modelled.
But after a few rounds of scepticism, data analysis and model critique you should get to a point where you trust the data more than your own guess. Otherwise we could just as well stop doing science at all.
Science theater is the practice of investing in research intended to provide the feeling of improved understanding while doing little or nothing to actually achieve it.
FWIW, my guess was at the 98th percentile, relatively close to the data (on the other hand, I was really bad at guessing the position of the ball in yesterday's football/soccer test).
The median guess probably suffers from a "connect the dots" bias, going from (0,0) to (100,100) through the suggested point.
Seeing how smooth the data is, it results from a large sample, and the measurements are objective, I fail to see how they could have gotten them wrong.
He's saying that, "Just because you can't guess the trend in the data (eg. the track of a soccer ball in yesterday's match) doesn't mean that something is wrong with the data; you often need to double-check both your hypothesis and your data. And in this case your hypothesis was probably at fault due to inherent bias rather than inaccuracies in the data."
Well, one thing they glossed over is the cost of tuition at the colleges in question varies. The poorest families are more likely to send their children to cheaper colleges. You might get an S curve if it was for colleges with $50K tuition, or whatever.
I agree but I would say it's not the poorest families that choose cheap colleges but lower middle-class families. People like teachers, office workers, et. al. who can't afford every college or who have to take out loans to send their kids to college are going to be price-sensitive. The poorest people generally make so little that they qualify for great financial aid without having to do academic scholarships.
Out of curiosity I drew the US wealth distribution on a linear axis a while ago, and it's really insanely packed to the left. As in everyone but a couple thousand people are in a single pixel column on the left with less than $100 million. Half the population are within a wavelength of light's distance of 0 on a linear wealth scale. The point is that wealth is nothing like a normal distribution - the very rich are so rich that even millions of dollars is nothing in comparison.
A few years ago, I made this histogram for a friend (I'm not American myself), who was interested in the modal income for the US, instead of the median income. The modal income is when you pick a US citizen at random, the expected value for their income, which can be seen as the location of the peak in this histogram:
Note that the data is from 2008, and only counts employed citizens.
I compiled the data from an official US government data/demographics/census website (I forget which one, sorry). I noticed I could query the average income for all counties split over many "occupation groups". This gave me relatively fine-grained buckets (it doesn't really matter what they were, just that they were small-ish), weighted by the number of people in it, allowing me to plot the histogram. The really proper way to build this histogram would be to bucket an actual list of income for each individual US citizen. But that list is not available, for obvious reasons.
Totally agree. My guess about the curve was wrong the same way as yours and the authors.
If that data's real then I want to know why it's tied to income and not something like wealth.
One wild theory.. quotas are based to income somehow, and that impressively straight line is because admissions officers are really good at their jobs.
Well, for one (and I'm not saying this is the explanation, just an answer to your question), FAFSA student aid eligibility is based on income, not wealth.
Actually wealth is another factor for FAFSA, it's just so seldom that a family has wealth but not income. Things like college savings accounts are actually factored against FAFSA applicants.
If your guess and data differ, wouldn't it be advisable to double check your data, collection mechanisms and your assumptions that led to your guess/estimate.
After a few goes your data needs to take precedence and one's guess dumped.
I learned also to check the original source: the data there doesn't seem to fit such a perfect straight line and does show a slight S curling at the edges.
It just shows an S curved the other way to what I drew.
Sometimes your intuition is wrong, and you can follow the wrong path for a very long time because the reality doesn't fit your intuition (the history of science is littered with examples).
But yes, the scientific method is: make a hypothesis, collect data, test your hypothesis, revise, do it again.
They said their guess and the data differ the same way as my guess. But then they think the data is correct and they just throw out their guess. That's actually wrong. This point is very important in statistics, but you actually already learn it in 4th or 5th grade math classes. Make a guess about the end result and if your calculated results is quite different be sceptical about your calculated result. Our ability to guess is not good at getting the exact numbers right, but it's really, really good at getting the big picture. Therefore when data and guess disagree there's a really good chance the data is wrong and we must question it and not just our guess (which we can also question, but questioning the data is more important).
PS: It's quite funny how often we consume lots of data and calculate a lot and still the right answer to the question is: "Dunno yet"
At least for me that's what I learned in the last few years on HN.