Missing value imputation and outlier treatment are two most important steps in the preliminary stage of any statistical modeling exercise. We have covered missing value imputation by various method in on of our previous articles.Let's us try to answer few basic and few advances questions regarding outliers.
All the questions that we are going to deal in this article are:
1. What are outliers?
2. Why, at all, do we need to bother about outliers ?
3. How to detect these outliers?
4. How do we treat outliers?
5. and last but not the least ... Is an outlier really an outlier ?
Q1. What are outliers?
In statistics, an outlier is an observation that is numerically distant from the rest of the data.
Grubbs, defined an outlier as:
An outlying observation, or outlier, is one that appears to deviate markedly from other members of the sample in which it occurs.
Grubbs, defined an outlier as:
An outlying observation, or outlier, is one that appears to deviate markedly from other members of the sample in which it occurs.
What we define outlier as :
(A few)Values that are altogether different from rest of the observation and the reason for being distinct is not known.
Q2. Why, at all, do we need to bother about outliers ?
Outliers might mislead analysts to altogether different insight regarding a data.
From a statistician point of view, it distorts the range, mean and standard deviation of the data. It also gives a wrong trend in the data as demonstrated below :
The solid red line is the actual trend line for the data cluster, but due to presence of outlier point (Xn ,Yn) , the trend line gets distorted to dashed line.
There actually exists a positive correlation between X and Y, but presence of the outlier leads analyst towards an opposite conclusion i.e. a negative correlation.
Presence of an outlier can be catastrophic to any analysis results. If any strategic decisions are taken based on such analysis results, it can be detrimental to business.
From a statistician point of view, it distorts the range, mean and standard deviation of the data. It also gives a wrong trend in the data as demonstrated below :
The solid red line is the actual trend line for the data cluster, but due to presence of outlier point (Xn ,Yn) , the trend line gets distorted to dashed line.There actually exists a positive correlation between X and Y, but presence of the outlier leads analyst towards an opposite conclusion i.e. a negative correlation.
Presence of an outlier can be catastrophic to any analysis results. If any strategic decisions are taken based on such analysis results, it can be detrimental to business.

6AD047E90F
ReplyDeletemobil ödeme bozdurma
Evde Paketleme İşi
Bitlo Güvenilir mi
Avast Cleanup Aktivasyon Kodu
Ucuz Takipçi
Outlier detection is an important part of the preliminary data analysis process because unusual observations can substantially affect the conclusions drawn from a dataset. This article does a good job of framing outliers not simply as extreme values, but as observations whose difference from the rest of the data may have an unknown reason, which makes investigation important before deciding how to handle them.
DeleteThe discussion of how a single outlier can distort the mean, standard deviation, trend, and even the apparent relationship between variables is particularly useful. This highlights why data preparation should be considered carefully before statistical modeling, especially when analytical results may be used to support business decisions. These concepts are closely related to practical Data Analysis Training, where understanding the quality and distribution of data is essential.
The article also raises an important question: whether an observation identified as an outlier is actually an error or represents a meaningful characteristic of the underlying data. Rather than automatically removing such observations, analysts need to understand the context and potential cause behind unusual values. This perspective is valuable when working with Machine Learning Training workflows, where preprocessing decisions can influence subsequent modeling results.
DeleteThe article’s discussion of outliers is especially relevant when preparing datasets for predictive modeling because unusual observations can influence model behavior and analytical conclusions. Before selecting an algorithm, it is important to determine whether an extreme value represents an error, a genuine observation, or a meaningful pattern in the dataset. These preprocessing considerations are directly applicable to Machine Learning Projects for Final Year, where data quality and preparation can have a significant impact on the resulting model.
Delete36D85F78
ReplyDeleteÇorum
Kırklareli
Bingöl
Amasya
Çankırı
Maraş
Batman
Aydın
Osmaniye
08786BAA
ReplyDeleteAkçay
Gemlik
Oğuzeli
Erciş
Kepez
Batum
Kemalpaşa
Kandıra
Çerkezköy
530FD44C
ReplyDeleteAtakum
Esenkent
Boyabat
Akçay
Serik
İlkadım
Ortaca
Kartepe
Bafra
730667BC
ReplyDeleteKırıkhan
Kartepe
Akseki
Beydağ
Kalkan
Gölhisar
Erdemli
Bergama
Yenişehir