12.6 Outliers (Page 2/11)

Introductory statistics Page 2 / 11

Try it

Identify the potential outlier in the scatter plot. The standard deviation of the residuals or errors is approximately 8.6.

The outlier appears to be at (6, 58). The expected y value on the line for the point (6, 58) is approximately 82. Fifty-eight is 24 units from 82. Twenty-four is more than two standard deviations (2 s = (2)(8.6) = 17.2 ). So 82 is more than two standard deviations from 58, which makes (6, 58) a potential outlier.

Got questions? Get instant answers now!

Numerical identification of outliers

In [link] , the first two columns are the third-exam and final-exam data. The third column shows the predicted ŷ values calculated from the line of best fit: ŷ = –173.5 + 4.83 x . The residuals, or errors, have been calculated in the fourth column of the table: observed y value−predicted y value = y − ŷ .

s is the standard deviation of all the y − ŷ = ε values where n = the total number of data points. If each residual is calculated and squared, and the results are added, we get the SSE. The standard deviation of the residuals is calculated from the SSE as:

$s = \sqrt{\frac{S S E}{n - 2}}$

Note

We divide by ( n – 2) because the regression model involves two estimates.

Rather than calculate the value of s ourselves, we can find s using the computer or calculator. For this example, the calculator function LinRegTTest found s = 16.4 as the standard deviation of the residuals

35
–17
16
–6
–19
9
3
–1
–10
–9
–1

x	y	ŷ	y – ŷ
65	175	140	175 – 140 = 35
67	133	150	133 – 150= –17
71	185	169	185 – 169 = 16
71	163	169	163 – 169 = –6
66	126	145	126 – 145 = –19
75	198	189	198 – 189 = 9
67	153	150	153 – 150 = 3
70	163	164	163 – 164 = –1
71	159	169	159 – 169 = –10
69	151	160	151 – 160 = –9
69	159	160	159 – 160 = –1

We are looking for all data points for which the residual is greater than 2 s = 2(16.4) = 32.8 or less than –32.8. Compare these values to the residuals in column four of the table. The only such data point is the student who had a grade of 65 on the third exam and 175 on the final exam; the residual for this student is 35.

How does the outlier affect the best fit line?

Numerically and graphically, we have identified the point (65, 175) as an outlier. We should re-examine the data for this point to see if there are any problems with the data. If there is an error, we should fix the error if possible, or delete the data. If the data is correct, we would leave it in the data set. For this problem, we will suppose that we examined the data and found that this outlier data was an error. Therefore we will continue on and delete the outlier, so that we can explore how it affects the results, as a learning experience.

Compute a new best-fit line and correlation coefficient using the ten remaining points:

On the TI-83, TI-83+, TI-84+ calculators, delete the outlier from L1 and L2. Using the LinRegTTest, the new line of best fit and the correlation coefficient are:

ŷ = –355.19 + 7.39 x and r = 0.9121

The new line with r = 0.9121 is a stronger correlation than the original ( r = 0.6631) because r = 0.9121 is closer to one. This means that the new line is a better fit to the ten remaining data values. The line can better predict the final exam score given the third exam score.

Questions & Answers

how do you get the 2/50

Abba Reply

number of sport play by 50 student construct discrete data

Aminu Reply

width of the frangebany leaves on how to write a introduction

Theresa Reply

Solve the mean of variance

Veronica Reply

Step 1: Find the mean. To find the mean, add up all the scores, then divide them by the number of scores. ... Step 2: Find each score's deviation from the mean. ... Step 3: Square each deviation from the mean. ... Step 4: Find the sum of squares. ... Step 5: Divide the sum of squares by n – 1 or N.

kenneth

what is error

Yakuba Reply

Is mistake done to something

Vutshila

anas

What is the life teble

anas

Jibrin

statistics is the analyzing of data

Tajudeen Reply

what is statics?

Zelalem Reply

how do you calculate mean

Gloria Reply

diveving the sum if all values

Shaynaynay

let A1,A2 and A3 events be independent,show that (A1)^c, (A2)^c and (A3)^c are independent?

Fisaye Reply

what is statistics

Akhisani Reply

data collected all over the world

Shaynaynay

construct a less than and more than table

Imad Reply

The sample of 16 students is taken. The average age in the sample was 22 years with astandard deviation of 6 years. Construct a 95% confidence interval for the age of the population.

Aschalew Reply

Bhartdarshan' is an internet-based travel agency wherein customer can see videos of the cities they plant to visit. The number of hits daily is a normally distributed random variable with a mean of 10,000 and a standard deviation of 2,400 a. what is the probability of getting more than 12,000 hits? b. what is the probability of getting fewer than 9,000 hits?

Akshay Reply

Bhartdarshan'is an internet-based travel agency wherein customer can see videos of the cities they plan to visit. The number of hits daily is a normally distributed random variable with a mean of 10,000 and a standard deviation of 2,400. a. What is the probability of getting more than 12,000 hits

Akshay

Bright

Sorry i want to learn more about this question

Bright

Someone help

Bright

a= 0.20233 b=0.3384

Sufiyan

Shaynaynay

How do I interpret level of significance?

Mohd Reply

It depends on your business problem or in Machine Learning you could use ROC- AUC cruve to decide the threshold value

Shivam

how skewness and kurtosis are used in statistics

Owen Reply

yes what is it

Taneeya

<< Chapter < Page Page > Chapter >>

Practice FlashCards 6 Key Terms 1

Read also:

Get Jobilize Job Search Mobile App in your pocket Now!

100% Free Mobile Applications
Receive real-time job alerts and never miss the right job again

Source: OpenStax, Introductory statistics. OpenStax CNX. May 06, 2016 Download for free at http://legacy.cnx.org/content/col11562/1.18

Google Play and the Google Play logo are trademarks of Google Inc.

Notification Switch

Would you like to follow the 'Introductory statistics' conversation and receive update notifications?

Ask

	3 How does the outlier affect the best fit line? By OpenStax Read Online Course
	1 Identifying outliers By OpenStax Read Online Course
	3 Neuroscience Exam 2004 4 By David Corey Start Exam
	36 Biology 36 Sensory Systems MCQ By OpenStax Start Quiz
	31 Biology 31 Soil and Plant Nutrition MCQ By OpenStax Start Quiz
	24 AP 24 Metabolism Nutrition Essay By OpenStax Start Flashcards
	Pre Employment English Proficiency Exam By Katherina jennife... Start Quiz
	Clinical Psychology MCQ By Saylor Foundation Start Quiz
	NCE Ch 10 Professional Orientation By Anh Dao Start Quiz
	Anthropology Religion Culture By Richley Crapo Start Assignment
	8 Psychology MCQ 2011 1 Exam By John Gabrieli Start Exam
	5 Arts Society: Theater 5 By Jonathan Long Start Quiz