Monday, February 25, 2019

Statistical Error

Statistical Error- 

Type I- Reject null hypothesis when null hypothesis is true (p<.05) 

Type II- do not reject a null hypothesis when the null hypothesis is false, uses beta to express power of test and its results 

Distribution and Deviations

Deviations- 

·     Asymmetric Deviation- Skewed distributions (right is positive, left is negative) 


·     Kurtic Deviation- 

platykurtic- flat, (closer to mean) and 
leptokurtic- peaked (more extreme) 

Normal Curve- area underneath the curve can be used to find probability of finding values greater than X1 


Elementary Data Analysis (EDA)

Measure of central tendency: 

·     Arithmetic mean/average 
·     Median- middle datum of sample – (50% above, 50% below) 
·     Mode- datum that occurs most often (frequency analysis) 

Measure of dispersion: 

·     Asses variation of data around median and mode 
-     Range 
-     Variance 
-     Coefficients of variance 
-     Standard deviation
-     Standard error
Range- numerical data between minimum and maximum 

Variance- measure of variation in data 

Standard Deviation- (SD) measure of spread of data around a mean 



Standard error- describe dispersion of sample mean around the population 

Coefficient of variation- compare amount of variation among data that differs- expressed as percentage 

Synthesis- population and samples (frequency distribution/ elementary statistics) 




Sample/ Probability

Population: Universe of events (ex- patients with aids, red wing birds) 

Rarely able to obtain data from every subject- the solution is samples 

Samples- subsets drawn from a population

Techniques- 

Random- every subject has same probability of being selected (Ex- table, coin flip, drawing straws) 

Systematic- predetermined method (Ex- every 3rd, every 4th

Convince - obtain under financial constraints

Probability-  the chance or likelihood that an event will occur

Theoretical- chances an event will occur (Ex- coin flip 50%) 

Empirical- based on data collection (Ex- HTHHTTHH) 




Frequency distribution- 

Exploratory Data analysis (EDA) 

-     Used to count frequency of occurrence of each datum in a sample population 
-     Search for trends/ patterns 

Displaying results- frequency distribution 

·     Pie Chart 
·     Bar Chart 
·     Histogram 
·     Relative Frequency – (pie chart/ stacked graph) 




Scientific Method

Scientific Method- 

·     Obtain background information 
·     Ask biological questions 
·     Develop testable hypothesis 
·     Design experiment 
·     Collect data 
·     Analyze data 
·     Interpret results 
·     Answer biological questions 

·     Present results   

Data/ Measurement

Metric- 

·     Volume (ml) 
·     Distance (m) 
·     Temperature (c) 
·     weight (grams) 

Types of Data- 

·     Ratio: True zero (weight, volume, length) 
·     Interval: arbitrary zero (temp, time of day, date) 
·     Ordinal: no number value (dark vs. light) 
·     Nominal: non- numeric (colors) 

Continuous- infinite number of values between two individuals (ex: 2.7, 7.701, 2.702) 

Discrete- integers (ex: 35 seals, 22 cells) 

Accuracy- measuring device (ex: nearest .1 g) 

Precision- researcher (how close they are to repeated measurement) 

Implied Range: value between data 

Rounding rules: If x> 5, then round 

History of Statistics

17thCentury-

Graunt- Studied affairs of the state/ vital statistics of population 

Petty- economist, probability/ political/ insurance 

Pascal- Probability/ gambling 

Bernoulli- probability/ risks with vaccines 

18thCentury-

Laplace/ Gauss- normal curve 

19thCentury-

Auelelet- statistical analysis to human biology, social mechanics 

Galton- genetic regression/ variation 


20thCentury

Pearson- (father of modern statistics) natural selection using correlation, first academic department, journal helped develop chi square analysis 

Gossett- developed student T test 

Fisher- developed ANOVA and importance of experimental design 


20thCentury (later) 

Wilcoxon- studied pesticides, non-parametric equivalent of two sample tests

Kuskal, Wallis- non- parametric equivalent of the ANOVA 

Spearman- nonparametric equivalent of correlation coeffient 

Kendall- nonparametric equivalent of correlation coefficient 


Turkey- multiple comparisons procedure 

Dunnett- multiple comparisons procedure for control group 

Keuls- multiple comparisons procedure 

Computer tech- growth of investigation into new techniques 



Nonparametric Statistics

Nonparametric Statistics  Use: When data violates assumptions of parametric tests. Data transformations do not solve problems. Analyses...