Concept
What does the data-to-wisdom hierarchy reveal about the role of data?
crystal1 / Evaluating a classifier
"When people ask me what data science is, here’s my go-to definition: deriving knowledge from data. But interpreting that phrase entails dissecting the difference between “knowledge” and “data,” two related but different terms. And that brings me to the data-to-wisdom hierarchy, depicted in Figure 1.1. Let’s break it down.\n\nThe real world\nUltimately, what we’re interested in is not data, but aspects of the real world – album sales and video views, stock prices and employment rates, hurricane trajectories and virus hot spots, or whatever. Data science can’t really get off the ground until some sort of data acquisition takes place that records measurements of the real world in electronic form. This sounds obvious, but it’s important to keep in mind, actually. No matter how much time we spend working with data, it’s never the data that actually matters – it’s the real-world phenomenon the data represents. It might seem strange to say that “data” is merely incidental to a data scientist, but it’s true. And I’ve definitely seen more than one data scientist get so locked on to the data that they forget this basic truth. One important observation is that decisions about exactly which data to acquire from the real world are often crucial in how things are interpreted later on. To take an example close to home, let’s say we’re gathering information on college professors so we can gauge which universities have the highest performing faculty, and how this might be changing over the years. We choose some representative set of criteria to measure for each faculty member to get a rough assessment of their performance. Let’s say we choose three things: the number of research papers the professor publishes each year, the total amount of research funding they’ve been granted, and the average student evaluation score of the courses they teach. That seems like a good first cut at assessing “faculty performance.” We then go on our merry data science way, finding correlations, making data visualizations, and drawing conclusions. This is all fine and dandy, provided we always keep in mind that it was those three qualities, and only those three, that we gathered in the first place. If our study gains any traction, and university professors find they have a vested interest in being ranked high in our yearly study, we’ll discover that they act to maximize only the categories that are being collected. We didn’t gather data on how many university committees they served on, or how many independent studies they supervised, or how many advisees they had, etc. Those metrics will inevitably become minimized in importance, because they weren’t part of what we lifted out of the real world and onto the bottom rung of our lofty chain. The moral is: what we measure matters, often more than we realize. Our country’s GDP and the Dow Jones Industrial Average are easy things to quantify, and so we often do. And thus they gain great importance in analyses of the economy. But are they actually the most important indicators? Does focusing on them leave out other, perhaps more vital, benchmarks? I’ll just leave you with that question for now.\n\nData\nHave you ever gotten blood work done, say for an annual physical? I have. I like to look over the numbers when the doctor hands me the results, just to chuckle and wonder what they all mean. To me, a non-physician, they’re all pretty much gobbledy-gook. They tell me my TBC is 4.93 x10E6/ μ L, that I have 5.7 Absolute Neutrophils, and a slightly out-of-range NT-proBNP (just 53.49 pg/mL, whatever the heck that means). When I use the word data in the context of the hierarchy, this is what I mean: recorded measurements, often (but not always) quantitative, that have not yet been interpreted. They may be very precise, but they’re also quite meaningless without the context in which to understand them. They’d even be meaningless to a physician if I didn’t provide the labels; try telling your doctor that you have 4.93 “something” and see whether he/she freaks out. The good news is that when we’re at the data stage of the hierarchy, we at least have the stuff in an electronic form so we can start to do something with it. We also often make choices at this stage about how to organize the data, choosing the appropriate type of atomic and/or aggregate data structures that we’ll discuss in detail in Chapters 3 and beyond. This will allow us to bring our analysis equipment to bear on the problem in powerful ways.\n\nInformation\nData becomes information when it informs us of something; i.e., when we know what it means. Getting large amounts of data organized, formatted, and labeled the right way are jobs for the data scientist, since turning that morass into useful knowledge is impossible without those steps. When the aspects of the real world that we’ve collected are properly structured and conceptually meaningful, we’re in business."
Related Ideas
- Why must data scientists treat their conclusions with responsibility and uncertainty?crystal1 · Evaluating a classifier