1.2 Data, sampling, and variation in data and sampling (Page 10/56)

Page 10 / 56

Variation in samples

It was mentioned previously that two or more samples from the same population , taken randomly, and having close to the same characteristics of the population will likely be different from each other. Suppose Doreen and Jung both decide to study the average amount of time students at their college sleep each night. Doreen and Jung each take samples of 500 students. Doreen uses systematic sampling and Jung uses cluster sampling. Doreen's sample will be different from Jung's sample. Even if Doreen and Jung used the same sampling method, in all likelihood their samples would be different. Neither would be wrong, however.

Think about what contributes to making Doreen’s and Jung’s samples different.

If Doreen and Jung took larger samples (i.e. the number of data values is increased), their sample results (the average amount of time a student sleeps) might be closer to the actual population average. But still, their samples would be, in all likelihood, different from each other. This variability in samples cannot be stressed enough.

Size of a sample

The size of a sample (often called the number of observations) is important. The examples you have seen in this book so far have been small. Samples of only a few hundred observations, or even smaller, are sufficient for many purposes. In polling, samples that are from 1,200 to 1,500 observations are considered large enough and good enough if the survey is random and is well done. You will learn why when you study confidence intervals.

Be aware that many large samples are biased. For example, call-in surveys are invariably biased, because people choose to respond or not.

Collaborative exercise

Divide into groups of two, three, or four. Your instructor will give each group one six-sided die. Try this experiment twice. Roll one fair die (six-sided) 20 times. Record the number of ones, twos, threes, fours, fives, and sixes you get in [link] and [link] (“frequency” is the number of times a particular face of the die occurs):

First experiment (20 rolls)
Face on Die	Frequency
1
2
3
4
5
6

Second experiment (20 rolls)
Face on Die	Frequency
1
2
3
4
5
6

Did the two experiments have the same results? Probably not. If you did the experiment a third time, do you expect the results to be identical to the first or second experiment? Why or why not?

Which experiment had the correct results? They both did. The job of the statistician is to see through the variability and draw appropriate conclusions.

Critical evaluation

We need to evaluate the statistical studies we read about critically and analyze them before accepting the results of the studies. Common problems to be aware of include

Problems with samples: A sample must be representative of the population. A sample that is not representative of the population is biased. Biased samples that are not representative of the population give results that are inaccurate and not valid.
Self-selected samples: Responses only by people who choose to respond, such as call-in surveys, are often unreliable.
Sample size issues: Samples that are too small may be unreliable. Larger samples are better, if possible. In some situations, having small samples is unavoidable and can still be used to draw conclusions. Examples: crash testing cars or medical testing for rare conditions
Undue influence: collecting data or asking questions in a way that influences the response
Non-response or refusal of subject to participate: The collected responses may no longer be representative of the population. Often, people with strong positive or negative opinions may answer surveys, which can affect the results.
Causality: A relationship between two variables does not mean that one causes the other to occur. They may be related (correlated) because of their relationship through a different variable.
Self-funded or self-interest studies: A study performed by a person or organization in order to support their claim. Is the study impartial? Read the study carefully to evaluate the work. Do not automatically assume that the study is good, but do not automatically assume the study is bad either. Evaluate it on its merits and the work done.
Misleading use of data: improperly displayed graphs, incomplete data, or lack of context
Confounding: When the effects of multiple factors on a response cannot be separated. Confounding makes it difficult or impossible to draw valid conclusions about the effect of each factor.

Questions & Answers

how do you get the 2/50

Abba Reply

number of sport play by 50 student construct discrete data

Aminu Reply

width of the frangebany leaves on how to write a introduction

Theresa Reply

Solve the mean of variance

Veronica Reply

Step 1: Find the mean. To find the mean, add up all the scores, then divide them by the number of scores. ... Step 2: Find each score's deviation from the mean. ... Step 3: Square each deviation from the mean. ... Step 4: Find the sum of squares. ... Step 5: Divide the sum of squares by n – 1 or N.

kenneth

what is error

Yakuba Reply

Is mistake done to something

Vutshila

anas

What is the life teble

anas

Jibrin

statistics is the analyzing of data

Tajudeen Reply

what is statics?

Zelalem Reply

how do you calculate mean

Gloria Reply

diveving the sum if all values

Shaynaynay

let A1,A2 and A3 events be independent,show that (A1)^c, (A2)^c and (A3)^c are independent?

Fisaye Reply

what is statistics

Akhisani Reply

data collected all over the world

Shaynaynay

construct a less than and more than table

Imad Reply

The sample of 16 students is taken. The average age in the sample was 22 years with astandard deviation of 6 years. Construct a 95% confidence interval for the age of the population.

Aschalew Reply

Bhartdarshan' is an internet-based travel agency wherein customer can see videos of the cities they plant to visit. The number of hits daily is a normally distributed random variable with a mean of 10,000 and a standard deviation of 2,400 a. what is the probability of getting more than 12,000 hits? b. what is the probability of getting fewer than 9,000 hits?

Akshay Reply

Bhartdarshan'is an internet-based travel agency wherein customer can see videos of the cities they plan to visit. The number of hits daily is a normally distributed random variable with a mean of 10,000 and a standard deviation of 2,400. a. What is the probability of getting more than 12,000 hits

Akshay

Bright

Sorry i want to learn more about this question

Bright

Someone help

Bright

a= 0.20233 b=0.3384

Sufiyan

Shaynaynay

How do I interpret level of significance?

Mohd Reply

It depends on your business problem or in Machine Learning you could use ROC- AUC cruve to decide the threshold value

Shivam

how skewness and kurtosis are used in statistics

Owen Reply

yes what is it

Taneeya

<< Chapter < Page Page > Chapter >>

Practice FlashCards 27 Key Terms 15

Read also:

Get Jobilize Job Search Mobile App in your pocket Now!

100% Free Mobile Applications
Receive real-time job alerts and never miss the right job again

Download PDF Learn

Source: OpenStax, Introductory statistics. OpenStax CNX. May 06, 2016 Download for free at http://legacy.cnx.org/content/col11562/1.18

Google Play and the Google Play logo are trademarks of Google Inc.

Notification Switch

Would you like to follow the 'Introductory statistics' conversation and receive update notifications?

Ask

	9 Critical evaluation By OpenStax Read Online Course
	7 Variation in samples By OpenStax Read Online Course
	1 Arts Society: Theater 1 By Jonathan Long Start Quiz
	Computer System Engineering By Samuel Madden Start Exam
	22 Biology 22 Prokaryotes Bacteria and Archaea MCQ By OpenStax Start Quiz
	5 Microbiology Final Practice By Madison Christian Start Assignment
	Spanish Lesson 4 By Anonymous User Start Quiz
	23 AP 23 Digestive System Essay By OpenStax Start Flashcards
	18 Sociology 18 Work and the Economy MCQ By OpenStax Start Quiz
	NCE Ch 08 Appraisal By Anh Dao Start Quiz
	16 AP 16 Neurological Essay Exam By OpenStax Start Flashcards
	12 AP 12 Nervous System Essay By OpenStax Start Flashcards