Outlier Detection &Treatment - A fresh perspective

Missing value imputation and outlier treatment are two most important steps in the preliminary stage of any statistical modeling exercise. We have covered missing value imputation by various method in on of our previous articles.

Let's us try to answer few basic and few advances questions regarding outliers.

All the questions that we are going to deal in this article are:

1.  What are outliers?
2.  Why, at all, do we need to bother about outliers ?
3.  How to detect these outliers?
4.  How do we treat outliers?
5.  and last but not the least ... Is an outlier really an outlier ?


Q1.  What are outliers?

In statistics, an outlier is an observation that is numerically distant from the rest of the data.

Grubbs, defined an outlier as:
An outlying observation, or outlier, is one that appears to deviate markedly from other members of the sample in which it occurs.

What we define outlier as :
(A few)Values that are altogether different from rest of the observation and the reason for being distinct is not known.




Q2.  Why, at all, do we need to bother about outliers ?
Outliers might mislead analysts to altogether different insight regarding a data.

From a statistician point of view, it distorts the range, mean and standard deviation of the data. It also gives a wrong trend in the data as demonstrated below :

The solid red line is the actual trend line for the data cluster, but due to presence of outlier point (Xn ,Yn) , the trend line gets distorted to dashed line.

There actually exists a positive correlation between X and Y, but presence of the outlier leads analyst towards an opposite conclusion i.e. a negative correlation.

Presence of an outlier can be catastrophic to any analysis results. If any strategic decisions are taken based on such analysis results, it can be detrimental to business.






8 comments:

  1. Replies
    1. Outlier detection is an important part of the preliminary data analysis process because unusual observations can substantially affect the conclusions drawn from a dataset. This article does a good job of framing outliers not simply as extreme values, but as observations whose difference from the rest of the data may have an unknown reason, which makes investigation important before deciding how to handle them.

      The discussion of how a single outlier can distort the mean, standard deviation, trend, and even the apparent relationship between variables is particularly useful. This highlights why data preparation should be considered carefully before statistical modeling, especially when analytical results may be used to support business decisions. These concepts are closely related to practical Data Analysis Training, where understanding the quality and distribution of data is essential.

      Delete
    2. The article also raises an important question: whether an observation identified as an outlier is actually an error or represents a meaningful characteristic of the underlying data. Rather than automatically removing such observations, analysts need to understand the context and potential cause behind unusual values. This perspective is valuable when working with Machine Learning Training workflows, where preprocessing decisions can influence subsequent modeling results.

      Delete
    3. The article’s discussion of outliers is especially relevant when preparing datasets for predictive modeling because unusual observations can influence model behavior and analytical conclusions. Before selecting an algorithm, it is important to determine whether an extreme value represents an error, a genuine observation, or a meaningful pattern in the dataset. These preprocessing considerations are directly applicable to Machine Learning Projects for Final Year, where data quality and preparation can have a significant impact on the resulting model.

      Delete

Do provide us your feedback, it would help us serve your better.