Statistics – Everybody makes DATA

Statistical Model: A Map is Not the Territory

December 4, 2019 APMA1 Comment

“A good analogy is that a model is like a map, rather than the territory itself. And we all know that some maps are better than others: a simple one might be good enough to drive between cities, but we need something more detailed when walking through the countryside. “

[The Art of Statistics, David Spiegelhalter]

When you visit Disneyland, you may not need a detailed map made by satellite information. We need only a simple cartoon map which includes the relative location of all the attractions. If you are a secret agent, you need a much more detailed map to investigate. The statistical model (or even data-driven model) is the same. The fidelity of the statistical models totally depends on the purpose of the use of the statistical model and the quality of data that has been fed into it.

When making a statistical model, there is a general trade-off between bias and variance (bias-variance dilemma). If we reduce the variance, the model may fail to approximate underlying ground truth (high bias). if we reduce the bias, on the other hand, the model is vulnerable to noise, leading to failure of approximation of ground truth (overfitting and high variance). This dilemma shows that we cannot make a perfect model from data. A map is a map. it is not the territory. A menu is a menu. it is not food. A statistical model is a model. it is not ground truth. So, we don’t need to overestimate (and also underestimate) the power of the statistical model. As a map is still useful to find a right path, a statistical model is useful to understand and predict the system.

Signal and Noise: How to Understand Data?

December 2, 2019 APMALeave a comment

“In the statistical world, what we see and measure around us can be considered as the sum of a systematic mathematical idealized form plus some random contribution that cannot yet be explained.”

[The Art of Statistics, David Spiegelhalter]

The famous book by Nate Silver, The Signal and the Noise, said how to find the signal from the noises. Since the level of the signal and the noise totally depend on the quality of data, it is really hard to distinguish these perfectly from the data. Also, it requires prior knowledge, intuition, and experiences about the data. So, all statistical models have two components: (deterministic) mathematical formulation and (stochastic) residual error. Hence, when we make a statistical model for analyzing the data, we need to check what we know (mathematical form) and what we don’t know (randomness). The name “residual error” seems to refer a bad model but it is not. Of course, the large residual error may stem from the bad choice of the model but this error often stems from the lack of our knowledge, the lack of data, or the data acquisition method.

When we analyze data, we don’t need to make a perfect model (actually it is impossible due to the aforementioned issues). If we try to make an errorless model, we can be struggling with overfitting issues, leading to the worst model without any significant finding. Instead, we provide both mathematical formulations and the corresponding residual errors. That is the only thing the statistician can do. Our life is the same. We don’t self-flagellate much when our life plan fell through. This is not our mistake but randomness in our life. If the failure of the plan came from our mistake, we fail several times in a row and then we can check what we did and adjust our plan (or mindset). If not, we may make a successful comeback the next time by randomness. Hence, no matter which reason makes your plan fail, we do try more and more for the success of our life.

Trade-off: The Quality or The Quantity of Data for Better Statistics

November 27, 2019 APMALeave a comment

“When we want to use the data to draw broader conclusions about what is going on around us, then the quality of the data becomes paramount, and we need to be alert to the kind of systematic biases that can jeopardize the reliability of any claims.”

[The Art of Statistics, David Spiegelhalter]

In the age of Big Data, we can collect tremendous data from many sources that have different qualities (e.g. accuracy, resolution, or fidelity). Using all the data we can easily draw statistics about what we measured. These statistical results help us to understand what is going on by comparing the previous statistical results. However, if we want a deep understanding of hidden patterns for accurate future prediction (statistical inference), the quality of data becomes the main factor for accurate prediction; higher quality, higher accuracy. Collecting data, however, has a general trade-off between the quality and the quantity. High accurate data require expensive data acquisition costs (e.g. expansive measurements, fine-scale simulation using more computer resources) while less accurate data are relatively cheap to obtain.

As the book mentioned, the data-driven predictive model totally depends on the quality (and the quantity) of data. First, we check the accuracy (or fidelity) of data and use only the high-fidelity data to make a data-driven model for decision or prediction. Due to the aforementioned trade-off, however, we generally have a few high-fidelity data and/or many low-fidelity data. Then how to make a data-driven model? since a few high-fidelity data provide only partial information, it is hard to make an accurate model globally. the use of many low-fidelity data enables us to make a global model but it has a systematic inherent bias, leading to a wrong prediction. Hence, in data science, many researchers have focused on multi-fidelity data fusion, which enables us to make an accurate global model using both high and low fidelity data; chasing both the quality and the quantity.

Are There a Few Magic Numbers for Describing Complex Systems?

November 25, 2019 APMALeave a comment

“Large collection of numerical data are routinely summarized and communicated using a few statistics of location and spread, (…), these can take us a long way in grasping an overall pattern.”

[The Art of Statistics, David Spiegelhalter]

Can we understand all (fine-scale) patterns from a massive data set? If you were a genius, you may keep track of all the patterns. But, it is (almost) impossible to analyze all. That’s is why we employ statistics to understand and analyze a large data set and predict/estimate the future from statistical results (e.g. population, economic growth, the unemployment rate, or stock price). For example, to make a business model for kids, it is much easier to see the average birthrate in some regions rather than count the number of children in my neighborhood. Statistical approaches always provide just a few numbers to describe the complex systems. This simplification enables us to make a simple (predictive) model, leading to an efficient and optimized analytics.

I agree that a few numbers make the complex system simple and I have experienced that this simple representation gives us the proper direction to make a better decision. Then, what is the good “number (statistic)” for massive data in our hands? The average? well, but the book also said: “there is no substitute for simply looking at data properly.” Hence, we should be careful to understand the complex system using only a few statistics. Some statistics are venerable to outliers such as average. Also, we can draw the dinosaur patterns using the given mean and variance (please see my previous post). Nowadays, data-driven approaches via statistical learning (machine learning) may provide optimal numbers to describe the complex system effectively. Yet, we need to scrutinize all the statistics the data-driven models provide. However, I do expect that a data-driven AI model may find a good reparameterization of the massive data set for a better understanding in the near future.

Framing: Statistics Can Manipulate Our Thought

November 22, 2019 APMA2 Comments

“The examples in this chapter have demonstrated how the apparently simple task of calculating and communicating proportions can become a complex matter.”

[The Art of Statistics, David Spiegelhalter]

Thanks to you, the number of followers increases by 22% in November! When you see this sentence, you may think that this emerging blog is growing rapidly and there are some reasons for this success. If I wrote “4 people start to follow my blog in November”, you might have a different feeling. But both are true: my blog has 18 followers in October and now 22. Different representations in statistics can change the impact of observations, we called this positive (or negative) framing. There are many examples of positive or negative framing. For example, pharmaceutical companies want to say that a new medicine has a 95% survival rate rather than a 5% mortality rate (positive framing). Investigative journalists want to say that 3,000,000 people are suspected of tax evasion every year rather than 1% of people (negative framing). This framing also appears in the graph. Assume that we need to draw a bar chart with two bars whose values are 95 and 98, respectively. If we draw a bar chart from 0 to 100, the two bars look similar. However, if we draw a bar chart from 90 to 100, we see totally different bars on the graph.

How can we escape from this framing? Information providers should provide alternative data representations (different graphs, law data, tables) so that we can get a balanced view of the data by examining raw data. Also, we always should be skeptical when we see data. First, we should check who (and why) published statistical data; data do not lie, only presenters may lie. However, this argument does not refer to that statistics are totally crafty tricks. Statistics is still powerful to understand, analyze, and visualize data effectively. Moreover, in the age of Big Data, statistical knowledge is fast becoming the main tool to deal with big data correctly. That is, statistics are a double-edged sword; the power of statistics depends on us.

[Wrap up] Book Review: Humble PI: A Comedy of Maths Errors

October 30, 2019 APMALeave a comment

In our life, mathematics is very important for logical thinking based on evidence-based knowledge through rigorous mathematical analysis. Especially, when we predict something new, the power of mathematics overwhelms our instinct or heuristics. However, when using mathematics improperly, catastrophic results are waiting for us. In this book, the author, Matt Parker, said such an important role of mathematics and showed examples of disasters stemming from mathematical errors through exhilarating stories he has experienced.

Then, what is the role of a human in mathematics? We try to use mathematics when deciding something important. And then, we should check all the types of mathematical errors to avoid the disaster. I would like to introduce his last paragraph. “Our modern world depends on mathematics and, when things go wrong, it should serve as a sobering reminder that we need to keep an eye on the hot cheese but also remind us of all the maths which works faultlessly around us.”

The following links are some quotations from the book with my thoughts.

(1) What Number Is a Really Big Number?

(2) Please Give Math More Time to Pick up the Pieces

(3) I Don’t Count on You When You Count Numbers

(4) More Approximations, More Problems in Your Life

(5) Probably, We Are Not Independent

(6) Searching for Average Man

(7) Sometimes, Simple Mathematics is Better than Our Experiences

What Number Is a Really Big Number?

October 14, 2019 APMALeave a comment

“As humans, we are not good at judging the size of large numbers. And even when we know one is bigger than another, we don’t appreciate the size of the difference.”

[Humble Pi: A Comedy of Maths Errors, Matt Parker]

In the Stone Age, a hundred might be a sufficient number to count a herd of deer for hunting or to count gathered nuts. In the early (and mid) 20th century, a million is enough to call the rich people ‘Millionaire’ but now it is too small to count Mark Zuckerberg’s net worth (a million is still BIG money for me by-the-way). In the age of Big Data, what number is a really big number? In the 1980s, Bill Gates, the pioneer to usher in the computer age, said: “for computer memory, 640K ought to be enough for anybody.” Nobody can predict the big number correctly and this is human nature.

However, we need to estimate a certain big number for a data-driven model, for our business, or for our blogs. After unveiling a smartphone, data acquisition speed is now super fast, leading to the age of AI and Big Data. Nowadays, when we make a model, we consider its own capacity to deal with tremendous data (beyond a trillion). The recent introduction of the Internet of Things (IoT) and the autonomous vehicle will generate countless data every second. Then, we need to keep thinking about the big number again and again. That is why I am preparing for the first event for the Billionth visitor to my blog. Do you think this number is still small? It depends on your action. please visit my blog more!

[Wrap up] Book Review: How Not to Be Wrong: The Power of Mathematical Thinking

October 11, 2019October 11, 2019 APMALeave a comment

We need to focus on the book title: How not to be wrong. Why did the author, Jordan Ellenberg, not say like: How to be right? This is because mathematical thinking is not the fruit of the Tree of Knowledge. Even though we equipped ourselves with concrete mathematical thinking, we cannot get the right answer to some problems we faced in the world. However, mathematical thinking helps us to correct our view based on a popular misconception and prejudice and to understand the structure of the world more clearly.

In this book, the author presents several mathematical misconceptions (more focused on statistics) that make the wrong decision and prediction, and show how mathematical thinking can overcome such kinds of obstacles. Since mathematical thinking is the extension of common sense by other means, the author said that we need more math majors for non-mathematician such as more math majors for non-mathematician such as math major doctors, high school teachers, CEOs, and politicians.

The following links are some quotations from the book with my thoughts.

(1) Do You Want to Be a Nonlinear Thinker?

(2) The Past is in the Past: the Law of Large Numbers

(3) Improbable Things Happen All the Time

(4) Can We Predict our Future in Chaos?

(5) Make Your Problem Harder!

(6) The Triumph of Mediocrity: Do not Stumble on Your Success

(7) Everything is Connected but Not Correlated

(8) When You Meet a Mathematical Genius

Everything is Connected but Not Correlated

October 7, 2019October 14, 2019 APMALeave a comment

“Correlation is not transitive. … The non-transitivity of correlation is somehow obvious and mysterious at the same time.”

[How not to be wrong, Jordan Ellenberg]

In Hollywood, the Bacon Number of an actress/actor represents the closest connectivity to the actor, Keven Bacon through movies. Surprisingly, we observed that almost all the actresses/actors can be connected to Keven Bacon within six steps, called this: “Six Degrees of Separation” or “Small World.” This concept originally stems from “Erdős Number” in mathematics and science research, representing a collaborative distance to the mathematician, Paul Erdős. (My Erdős number is 4 by-the-way). What a small world and we feel that everybody is connected!

Sometimes, we confuse a correlation with a connection (or relation). A correlation is not transitive. Even though A and B are strongly correlated and B and C are also correlated, nobody can guarantee that A and C are correlated. However, we often think that there should be a correlation between A and C because we get used to syllogistic reasoning. Moreover, when we mixed up with causality, correlation, and relation, it’s a disaster. So, please do not make any transitivity for mutually correlated data. Also, we keep in mind that uncorrelated data can have a relationship with each other. We, you and I, are connected in the small world but we may not (or may) be correlated with each other.

Improbable Things Happen All the Time

September 27, 2019October 2, 2019 APMALeave a comment

“The universe is big, and if you’re sufficiently attuned to amazingly improbable occurrences, you’ll find them. Improbable things happen a lot.”

[How not to be wrong, Jordan Ellenberg]

You have a card deck and draw five cards from this. Surprisingly, five cards you drawn are spade A, 2, 3, 4, and 5. (Congrats! you made a straight flush). Then, you might think that this is a new card deck so it is not shuffled yet because drawing these five cards in a row might be improbable (or much lower probable). However, improbable things happen all the time. Please go to Las Vegas and check this!

When analyzing some results, we need to get used to a BIG number in our fields. Our field of interest is pretty big and you can see many improbable occurrences (we can see winners of the lottery every week). Hence, we should be careful not to make any causality from a chance occurrence. In data science, even though the data-driven model finds some patterns from Big Data, we should examine that this pattern can be made by randomness or not. (It may be improbable that millions of people read this post and like it but improbable things happen all the time!!)