Skip to content
#Management

[2/2] Dark data (Dark Data. Why What We Don’t Know Is Even More Important Than What We Do) (Category Management)

#Management #Data #Math #Statistics

Continuing the story of the dark data I started in past postI wanted to talk about the dark data classes that David Hand, a member of the British Academy and president of the Royal Statistical Society, highlights.

1) Data that we know is missing These are known unknowns that occur when we know there are problems in the data that hide values that may have been written down. 2) Data that we do not know is missing - it's "unknown unknowns." We don't even know that we're missing some data. The book tells about the Challenger disaster, where decision makers lacked information about the behavior of O-rings in cold weather, but they did not know about it. 3) Selective facts A poor set of criteria for inclusion in the sample or an erroneous application of reasonable criteria leads to such problems. You can also take it here. p-hackingIt consists of conducting a large number of statistical tests, but telling only about those that were successful. 4) Self-selection This variant is a subtype of the previous type of dark data, namely selective facts. It manifests itself when people are given to decide for themselves what to include in the survey results and what not. The missing data may have system differences from those that were entered. 5) Unknown determining factor It's an old-fashioned story about correlation not being causal. The author immediately recalls the Simpson's paradox 6) Data that could exist (counterfactual) It is data that we would see if we had taken other actions or observed what was happening under different conditions. 7) Data that changes over time Some data may cease to be recorded outside the observation period, others because nature has changed. Time can hide data in different ways. Incorrectly defined data The definition of data may change over time to better fit its subject matter and purpose. This can cause problems in interpreting time series as the nature of the data itself changes. 9) Data synthesis - here we are talking about when we save some of their parameters instead of data: average, median, mean-square deviation and so on. So we lose some of the information from the data. 10) Measurement errors and uncertainty This type of data about the error of measurements, as well as when converting data from different formats. 11) Distortion of feedback and trickery This type of data occurs when the collected values begin to affect the original process. There are examples of inflated valuations and stock market bubbles. Plus, we can recall quantum physics, where the measurement itself affects the state of the system:) 12) Information asymmetry This type of data occurs when different participants in the interaction have their own data sets. Akerlof, Spence and Stiglitz in 2001 He received the Nobel Prize in Economics for his work on study Information Asymmetry Markets (They explored the market for used lemon cars.) 13) Intentionally obscured data This is the deliberate selection of certain facts to conceal information and manipulate facts for deception or fraud. 14) Fake and synthetic data Such data is created artificially, for example, for fraud. It is interesting that it is difficult to make high-quality synthetic data, but it is real with the help of simulation of processes. 15) Extrapolation beyond your data Data is usually used to build models. These models work within the boundaries of the data we've seen. But when we go beyond that, we get this extrapolation problem. Here is another example of the Challenger shuttle.

In general, knowing about these types of dark data is helpful, and even more useful is reading a book and hearing interesting stories of fakaps with first-hand data.

#Management #Data #Math #Statistics