Why Outlier Trimming Can Destroy Your Data Accuracy and Forecasts
Deleting extreme data points might make your charts look better but it ruins your accuracy. Learn why outlier trimming is a dangerous statistical habit.
The Temptation of Clean Data
Imagine you are an online store manager reviewing your monthly sales. Most of your orders are between fifty and eighty dollars. Then you see a single order for one hundred and fifty thousand dollars. It looks like a mistake. It is so far outside the norm that it breaks your charts. You decide to delete it because it is an outlier. You want your sales forecast to look realistic for the average customer. This seems like a smart move at first. You want clean data. You want a forecast that represents the majority of your business. But by deleting that one point you might have just deleted the most important information in your entire database. This is the danger of outlier trimming. It is a common practice in statistics that often does more harm than good.
In our latest video we dive deep into how these statistical decisions play out in the real world. You can see the full walkthrough of these concepts in the video above. Understanding the math behind outliers is the difference between a business that grows and a business that gets blindsided by reality.
What Exactly Is Outlier Trimming
Outlier trimming is a technique used to remove data points that are significantly different from the rest of the set. In many introductory statistics courses students are taught that outliers are errors. They are told that these points skew the results and should be discarded to find the true average. This is often done using a canned rule. A canned rule is a pre-determined threshold. For example a rule might say that any data point more than three standard deviations away from the mean is trash.
When you apply these rules you are essentially telling the data what it should look like. You are forcing the world to fit into a neat box. The problem is that the world is rarely neat. Data points that look like errors are often the most significant events in a system. When you trim them you are not just cleaning the data. You are rewriting the story that the data is trying to tell you.
The Mathematical Consequences of Trimming
When you remove an extreme value from a dataset several things happen to the underlying math. These changes might seem small but they have massive ripple effects on your predictions.
- The average moves. If you remove a very high value the mean of your dataset will drop. This makes your typical performance look lower than it actually is.
- The distribution shifts. Statistics relies on understanding the shape of your data. Trimming changes that shape. It often makes a distribution look more like a perfect bell curve than it really is.
- Standard deviation shrinks. This is perhaps the most dangerous change. Standard deviation or sigma measures how much your data varies. When you remove outliers you artificially lower the sigma. This gives you a false sense of certainty. You begin to believe that your results are more predictable than they actually are.
By shrinking the standard deviation you are ignoring the risk of extreme events. In the world of finance or business those extreme events are often where the most money is made or lost. A forecast that ignores outliers is a forecast that is built for a perfect world that does not exist.
Signal Versus Trash
The biggest challenge in data analysis is deciding if a weird data point is a signal or just trash. Not every outlier is important. Sometimes a data point really is just an error.
Consider two scenarios for that one hundred and fifty thousand dollar order. In the first scenario the order is a duplicate record. A glitch in the software caused the same transaction to be recorded twice. In this case the outlier is trash. It does not represent a real world event. Deleting it improves your data accuracy.
In the second scenario the order is a real renewal from your biggest corporate client. This client only orders once a year but their business accounts for a huge portion of your revenue. In this case the outlier is a signal. It is the most important piece of data you have. If you delete it your forecast will miss the renewal next year. You will be unprepared for the massive influx of work and cash flow.
Outlier trimming treats both of these scenarios the same way. The canned rule does not care about the story behind the numbers. It only cares that the number is big. This is why automated data cleaning can be so risky. It removes the human element of investigation.
The Importance of the Data Generating Story
To handle outliers correctly you must understand the data generating story. This is the real world process that created the numbers in the first place. You cannot look at a spreadsheet in a vacuum. You have to ask where the numbers came from.
- Was this a manual entry error? If a human typed an extra zero by mistake the point should be corrected or removed.
- Is this a rare but recurring event? Some industries are defined by rare events. Think of insurance claims or venture capital returns. In these fields the outliers are the entire business model.
- Does the outlier represent a shift in the market? Sometimes an outlier is the first sign of a new trend. If you delete it you will be the last person to notice that the world is changing.
Instead of blindly trimming you should investigate. Look into the specific transaction. Check the logs. Talk to the department that generated the data. If you cannot prove it is an error you should probably keep it. It is better to have a messy chart that is true than a beautiful chart that is a lie.
Rethinking Your Approach to Statistics
Statistics should be a tool for discovery rather than a tool for conformity. Many people use statistics to make their data look normal. They want everything to fit the bell curve because the bell curve is easy to understand. But the most interesting parts of life happen at the edges of that curve.
When you encounter an outlier your first instinct should be curiosity rather than frustration. That weird data point is a doorway to a deeper understanding of your system. It might reveal a flaw in your software or a massive opportunity in your market.
If you find yourself reaching for a tool to trim your data stop and ask why. Are you trying to find the truth or are you just trying to make the graph look pretty? The best analysts are the ones who embrace the outliers. They know that the most valuable insights are often hidden in the data that everyone else is trying to delete.
How do you decide which data points to keep when your results look too good to be true?
Want more from Math Unlocked?
New videos become articles here automatically. Join the community to talk about them.