Well, as I write this once again around 10 pm on a Friday, it is safe to say that I am quite ready to be done with work. The wonderful ease into the semester that was the first week has unfortunately died down, as all my classes began to dive deeper into content. Coincidently, diving deeper is what this week’s newsletter is all about. This week’s reading came from Nathan Yau once again, and I decided to explore Chapter 8 of Visualize This, titled Analyzing Data Visually. This builds off of the insights from last week where we looked into creating visualizations, focusing on how to (you guessed it) analyze them. Understanding these concepts are crucial to the process of data visualization and elevates the outputs we produce from our visualizations. Let’s learn what they are.
The Process
Understanding the Data
Visualizations are awesome, as they allow us to see and understand patterns in the data that a simple spreadsheet does not allow for. But, before the truly beautiful ones can be created, we need to understand our data. Nathan Yau likes to compare this first step to a first date. When meeting a new person, you want to ask broad, general questions that allow you to understand who they are rather than deep, specific ones to see if you like the person overall. The same principle can be applied to data analysis. The first visualizations we make should be understanding summary statistics or distributions, not fiercely analyzing a specific variable. One of the best ways to do so is through boxplots.
Take the example above. We want to get a general idea of how tall basketball players are, and from our box plots we get ranges, mins, maxes, medians, and quantiles. This shows us that centers are typically the tallest types of players, while guards are typically the shortest, but almost all players are over six feet tall. Now, this makes sense, and that is a good thing. If the visualizations told me everyone was over nine feet tall and guards were on average the tallest type of player, I would be bewildered and want to check the source of my data to check its reliability. Additionally, now that we have a general sense and confidence in the data, we can continue asking more questions. For example, who is the player that is under 70 inches tall? These types of questions require data moves (our favorite), but also lead to deeper analysis, which conveniently is the natural progression in our process.
Further Insights
Now that our first date with the data went well, it follows then that we start asking more questions like on a third or fourth date. This can be focusing on a specific variable, looking into points that stick out, or going in a completely new direction. From out Sliced videos we watched this week, we got a lot of great examples of this. We will focus on DRob (who is apparently the greatest coder of all time), during his championship matchup. In his hunt to find the best model for predicting loan defaults, he created a shiny app (shown below) that allowed him to understand specific relationships of attributes with the response variable.
In the visualization, we can clearly see filters that change the bank, state, and sector. This in turn allowed him to find the explore to his heart’s content, finding best attributes for his model. Even though the visualizations are quite simple, they allowed him to see patterns and outliers he might not have gotten from a general view, which is the beauty of the deeper analysis. And, to make it even better, DRob ended up winning (but who was surprised).
The Result
After all this data analysis, we should also talk about why we need it. From all the patterns we see and the endless cycle of questions and answers we get from our visualizations, in the end we get to tell a story. Now, this story is very much up to the writer, but what we get to show very much has an impact. How we present our analysis very much matters, and we should always give appropriate context for the visualizations we create.
The graph above is from a research project I did this past summer, regarding the environment of statistics education at liberal arts education. From our analysis, we learned that the percent of introductory statistics enrollment taught outside the math department has significantly decreased in the past forty or so years. However, in the visualization, we also made sure to appropriately space out the years so the differences in time could be properly seen and show all data points for each year. This allowed our stakeholders to understand the overlying data and understand changes over time better than one line graph with only four points evenly spread out. The context we gave in turn greatly improved the story we were able to tell and create an end product that was of use. It is these types of decisions that not only elevates the work that data analysists do but also allows the best story to be told. If we do not, the whole point of the data analysis can be overshadowed be confusion and lessen the impact we have, and if that happens, then what was the point? Plus, I am sure DRob gives appropriate context, so you should too. Have a great weekend!