In my essays I really try to drive towards a positive useful point with tangible take-aways.  My success rate is in the eyes of the reader, but that is my goal.  Mostly.  OK, there might be a soupçon of snark and a dash of judgement in there as well, but that is just to bring full bloom to the aromatics.  Today, it is all about letting my id run wild and complaining about things that just annoy me when it comes to the broad topic of statistics.  For those looking for part two of my essay from Monday, rest assured that I will deliver that on Friday.  For the moment, though, I shall release the hounds upon my pet peeves.  If you have a statistical pet peeve you want to discuss, please leave it in the comments.

Before I continue, many of my examples come from sports and I considered putting endnotes on all of my sports terms.  But then I realized that I would double the essay’s length and not add any clarity.  Either you will already understand them, or my brief summary won’t likely be enough to crack the meaning.  I sincerely hope that this won’t interfere with your appreciation of my core points.

Using the word “statistics” as a cudgel

I was listening to a meeting this morning and I heard someone say, “You can look at the statistics all you want, but…” and I realized that the word “statistics” is almost always used in a negative, dismissive, or pejorative way.  The person who made the comment in this meeting was essentially saying that the actual math didn’t matter when compared to his feeling or perception or belief. 

Now I will be the first to point out that statistical analysis is the tail and not the dog.  As we will get to later, people can cherry-pick numbers to bolster their point.  People can tailor a model to generate outcomes that are useful to their point, but useless in understanding the issue.  I am completely in favor of calling bullshit on people trying to dazzle you with spurious relationships.  More on that in a minute as well.  

Refuting the math or making an argument that the math is delivering an incomplete picture are both excellent responses to statistical analysis.  Simply dismissing the math, or worse, treating the analysis with derision, feels like corporate-speak for just pointing and shouting “NERD!”  Not only does this do a disservice to the conversation at hand, but it also provides aid and comfort to those who already hate or fear math. 

Using statistics to confuse

The reason some will dismiss analysis is because there are others who are trying to confuse with analysis.  I get annoyed with people who think that showing a wall of numbers is sufficient to make a point.  This is especially the case when they will put one number out the sea to discuss, without putting context to anything else.  This often happens when I am being shown financial spreadsheets.  But it also happens when people try to bamboozle with complex-sounding statistical methods regardless if they are appropriate for the data they are using or the question they are asking.   There may be value in predicting missing data and capturing residual variance, but chances are, if you are not explaining why, you are hoping people don’t notice that you are using a lot of jargon to hide your lack of confidence in what you are speaking about.

Because of these two peeves, I don’t get angry with people who hate math.  Both sides are teaching people to fear it or distrust it.  They are both pretending it is a game, rather than a useful tool used to make informed decisions. 

Misusing the word sample

To be clear, a population is the complete set of all elements that meet a criterion or set of criteria.  A sample is a subset of that population.  (I will set aside any conversation about sample size or sample quality for another day.)  People somehow assume that all populations have to be big and all samples have to be small, but that is not the case.  The entire population of people who have been to the moon is twelve.  The average sample size of national polls can be at or above one thousand respondents. 

My peeve here is that people will often call a small population a sample.  I hear this in the context of baseball all the time and I think this is noticeable and annoying, given how attentive the sport is to statistics.  So, when I hear an announcer or blogger claim, “Player X is hitting .429 against left-handed pitchers with runners in scoring position, but that is a small sample size…” it is like nails on a chalkboard.  They did not take a sample of all the possible times that this player was in this situation.  The player was in this situation seven times and got a hit three times.  The total population was seven occurrences.  What they usually trying to say is, “Player X is hitting .429 in this situation, but since they have only been in this situation seven times, one probably should not read too much into their past performance to evaluate the current situation.” 

I hear this all the time in politics as well.  Like, for example, when people want to talk about midterm election performance for the president’s party when the president is in their second term.  There have been exactly seven occurrences of this in my lifetime and that is including (a) Nixon who was shown the door during his second term, (b) JFK/LBJ which only kind-of counts and is not technically during my lifetime, and (c) Trump, who is not serving those terms consecutively.  So, as we approach the upcoming midterms and you want to make an observation about past patterns, don’t tell me that it is a small sample; tell me if what you see is useful, given that it is based upon seven observations.

Being oblivious about context

Related to the previous is when people imply significance or importance without acknowledging context.  This seems most present in situations, like the World Cup or the Olympics or presidential elections, that don’t happen every year.  I remember that when France won the World Cup in 2018, the announcers were saying that it was the first France victory in TWENTY YEARS.  This is true.  France had previously won back in 1998, when dinosaurs still roamed the earth.  This seems momentous and seems a relief to long-suffering French fans.  Given that the World Cup only happens every four years, though, there were exactly FOUR World Cups between 1998 and 2018.  Further, given that over forty teams compete every four years, France’s performance actually seems impressive, since one could just as easily say that France has won 33% of all World Cups over the past 20 years. 

The reverse of this is also true.  Some will point to the fact that the Boston Celtics went to eleven straight NBA finals from 1959 through 1969.  This is truly impressive and not likely to be duplicated.  Mostly, though, it will not be duplicated because back when they did it, there were only FOUR teams in the Eastern Conference, and the finals simply pitted the two conference champions.  This is a far cry from the current 30 teams and four rounds of playoffs.  Given how quickly politics, sports, and everything else changes, citing patterns without providing some sort of grounding can lead one to make dubious comparisons.

Confusing statistical significance with conversational significance

This brings me to something that I have spoken of before.  Given big enough samples and wide enough gaps, many differences can be found to be statistically significant.  This, though, does not mean that they are worth talking about.  For example, in my previous professional life, I built a database of almost 500,000 inpatient interviews.  Based upon this data, I found that 61.4% of men rated the overall quality of their inpatient stay as top-box, while 60.5% of women gave the top-box score to the same question.  Based upon the sample, this difference is statistically significant at a 95% confidence interval.  The question is, though, as one sifts through the reasons why patients are satisfied or not, is a 0.9% difference based upon gender worth discussing?  Does anyone think that this is more important than, say, how well the care team communicates or provides compassion or listens?  *Spoiler Alert* NO.  Just because something is statistically significant doesn’t mean it automatically becomes relevant to the conversation. 

Statistics without logic

Earlier, I said that statistical tools are the tail and not the dog.  It can also be said that for some, statistics is the hammer in a world full of nails.  In the previous example, one can mistake significance with value.  Here, some will mistake significance with meaning.  Again, in a previous essay, I discussed the finding that, while emergency department scores were statistically significant when compared against trauma designation, the direction was not logical.  Trauma 1 designations (the best possible) score better than Trauma 2 designations, but worse than Trauma 3 designations.  Logic might dictate that they would score best of all, or worst of all, but putting them in the middle doesn’t make obvious sense.  So, presenting it as fact without explanation or any logical reason actually leads to confusing people or getting them to disregard statistics as a game.

Certainly, there are nonlinear patterns in data and some of you reading this might construct an argument to explain the pattern in Trauma designations.  The problem though, is that too often, the data is presented as useful or important, even when no one has explained WHY it is useful or important.  This usually means having a hypothesis that you are proving or disproving.  Building a hypothesis after the fact, that may or may not stand up to scrutiny or duplication is the very definition of letting the tail wag the dog.

Sculpting a dataset or statistic

I will return to sports for this one.  A year or so ago, I heard a talking head say something like, “In eight of the past twelve games, the Kansas City Chiefs have been penalized an average of ten fewer times than their opponent.”  I thought that this was an odd statistic, so I went back and looked at the data.  I saw that they were correct, but I also saw that when you looked at all twelve of the past games, the gap was only about five and not ten penalties.  This data was sculpted, then, to make a point.  Cases were eliminated not for cause, but simply because they did not align the speaker’s thesis.  They cherry-picked those cases because it made their point more impactful, albeit less accurate.

Drawing dubious conclusions

Associated with this is the implication of presented stats.  In the previous example, the person was arguing that the Kansas City Chiefs were getting preferential treatment by the officials.  This might be the case.  Of course, it also might be that they were the more disciplined team on the field.  Every year, there is a most-penalized and a least-penalized team and generally this is attributed to team discipline and attention to the way the game is officiated.  Why should I not accept this explanation for the Kansas City Chiefs and instead buy into a conspiracy theory? 

I understand the allure of this logic.  As a life-long Creighton Bluejays fan, I spent some of my formative years complaining that the NCAA tournament was rigged because my beloved Jays always had fewer trips to the free-throw line than their opponents.  Clearly this was because the other team was getting calls that my team wasn’t.  Or it was, until I realized that my team was taking a bunch of 3-point shots rather than driving to the lane, meaning that there were fewer opportunities to draw a foul.  Both explain the disparity but only one requires belief in an anti-Midwest biased shadowy cabal. 

One common theme in many of these is that “the data speaks for itself.”  But the data never speaks for themselves.  Without underlying logic or context, we are often led astray by the illusion of patterns and the laws of big numbers.  By not clarifying hypotheses or defining the universe, we can confuse ourselves with what we think we know.  Worse, we can be bamboozled by those who think that sleight of hand and an authoritative tone is enough to convince us of what is real.

Leave a comment