ANOVA
Compare means of three or more groups.
A large F means at least one group mean differs. It does not name which group.
Biostatistics
29 formulas
Compare means of three or more groups.
A large F means at least one group mean differs. It does not name which group.
How much extra disease the exposure adds.
Subtract the incidence in the unexposed from the incidence in the exposed.
Among people who have the disease, the share who die of it.
The denominator is cases, not the whole population. That is what separates it from mortality.
Deaths from one cause in the whole population.
Often written per 1,000 people. Existing cases stay in the denominator.
Test association in a table, usually 2×2.
Expected count in a cell is (row total × column total) / grand total.
Compare spread when the means are different.
A higher CV means more scatter relative to the average.
Estimate a mean and the range likely to hold the true value.
The piece after z is the standard error. A larger sample narrows the interval. For 95%, z is about 1.96.
New cases over a period.
People who already have the disease are taken out of the denominator.
Predict an outcome from one variable.
y is the outcome. x is the predictor. b is how much y changes for each step of x.
The average of a set of values.
Add every value, then divide by how many there are.
A typical value that ignores extremes.
Middle value once the numbers are in order.
If the count is even, average the two middle values.
The most common result in a set.
The value that appears most often.
A set can have more than one mode, or none if every value appears once.
Of the negative tests, the share that are truly negative.
NPV falls when the disease is common, because more negatives are missed cases.
A symmetric bell curve around the mean.
68% within 1 SD · 95% within 2 SD · 99.7% within 3 SD
About 95% of values sit within two standard deviations of the mean.
The starting claim a test tries to reject.
No difference, or no association.
Reject it when the p-value is below the cutoff, usually 0.05.
Odds of disease in the exposed versus the unexposed.
OR = 1 means the exposure is not tied to higher odds. Read the letters from a 2×2 table.
How common a disease is at one moment.
This is a snapshot, the kind a cross-sectional study measures. It counts existing cases, not new ones.
Of the positive tests, the share that are truly positive.
PPV rises when the disease is common.
A part of a whole.
Multiply by 100 to turn the proportion into a percent.
How surprising the result is if the null hypothesis is true.
p < 0.05 → unlikely if there is truly no difference.
A small p-value is evidence against no difference. It is not the chance the result is wrong.
The distance from the smallest number to the largest.
Highest value minus the lowest value.
One pair of extremes sets it, so a single outlier can stretch it.
How many times higher the risk is in the exposed.
RR = 1 means no association. From a 2×2 table, Ie is a/(a+b) and Iu is c/(c+d).
The average squared distance from the mean.
A sample uses n − 1. Standard deviation is the square root of this.
Of the people who have the disease, the share the test calls positive.
A sensitive test misses few cases.
Of the people without the disease, the share the test calls negative.
A specific test raises few false alarms.
Typical distance of values from the mean.
On a normal curve, about 95% of values sit within 2 SD of the mean.
How much the sample mean would bounce from sample to sample.
This is the term a confidence interval multiplies by z. Larger n makes the standard error smaller.
Compare a mean when the population SD is unknown.
Use the sample SD in place of σ. The shape matches a z-score with s instead of the population SD.
How many standard deviations a value sits from the mean.
z = 1 is one SD above the mean. z = −1 is one SD below it.