Biostatistics

Formulas

29 formulas

Search

A

1.

ANOVA

Compare means of three or more groups.

A large F means at least one group mean differs. It does not name which group.

2.

Attributable risk

How much extra disease the exposure adds.

Iₑ
Incidence in the exposed
Iᵤ
Incidence in the unexposed

Subtract the incidence in the unexposed from the incidence in the exposed.

C

3.

Case-fatality rate

Among people who have the disease, the share who die of it.

The denominator is cases, not the whole population. That is what separates it from mortality.

4.

Cause-specific mortality

Deaths from one cause in the whole population.

Often written per 1,000 people. Existing cases stay in the denominator.

5.

Chi-square

High yield

Test association in a table, usually 2×2.

O
Observed count
E
Expected count

Expected count in a cell is (row total × column total) / grand total.

6.

Coefficient of variation

High yield

Compare spread when the means are different.

SD
Standard deviation
x̄
Mean

A higher CV means more scatter relative to the average.

7.

Confidence interval

High yield

Estimate a mean and the range likely to hold the true value.

x̄
Sample mean
s
Sample standard deviation
n
Sample size
SE
s / √n

The piece after z is the standard error. A larger sample narrows the interval. For 95%, z is about 1.96.

I

8.

Incidence

New cases over a period.

People who already have the disease are taken out of the denominator.

L

9.

Linear regression

Predict an outcome from one variable.

y
Dependent variable
x
Independent variable
a
Intercept
b
Slope

y is the outcome. x is the predictor. b is how much y changes for each step of x.

M

10.

Mean

High yield

The average of a set of values.

Σx
Sum of the values
n
How many values

Add every value, then divide by how many there are.

11.

Median

High yield

A typical value that ignores extremes.

Middle value once the numbers are in order.

If the count is even, average the two middle values.

12.

Mode

The most common result in a set.

The value that appears most often.

A set can have more than one mode, or none if every value appears once.

N

13.

Negative predictive value

Of the negative tests, the share that are truly negative.

TN
True negatives
FN
False negatives

NPV falls when the disease is common, because more negatives are missed cases.

14.

Normal distribution

A symmetric bell curve around the mean.

68% within 1 SD · 95% within 2 SD · 99.7% within 3 SD

About 95% of values sit within two standard deviations of the mean.

15.

Null hypothesis

The starting claim a test tries to reject.

No difference, or no association.

Reject it when the p-value is below the cutoff, usually 0.05.

O

16.

Odds ratio

High yield

Odds of disease in the exposed versus the unexposed.

a
Exposed, with disease
b
Exposed, without disease
c
Unexposed, with disease
d
Unexposed, without disease

OR = 1 means the exposure is not tied to higher odds. Read the letters from a 2×2 table.

P

17.

Point prevalence

How common a disease is at one moment.

This is a snapshot, the kind a cross-sectional study measures. It counts existing cases, not new ones.

18.

Positive predictive value

Of the positive tests, the share that are truly positive.

TP
True positives
FP
False positives

PPV rises when the disease is common.

19.

Proportion

High yield

A part of a whole.

x
The part
n
The whole

Multiply by 100 to turn the proportion into a percent.

20.

p-value

How surprising the result is if the null hypothesis is true.

p < 0.05 → unlikely if there is truly no difference.

A small p-value is evidence against no difference. It is not the chance the result is wrong.

R

21.

Range

The distance from the smallest number to the largest.

Highest value minus the lowest value.

One pair of extremes sets it, so a single outlier can stretch it.

22.

Relative risk

How many times higher the risk is in the exposed.

Iₑ
Incidence in the exposed
Iᵤ
Incidence in the unexposed

RR = 1 means no association. From a 2×2 table, Ie is a/(a+b) and Iu is c/(c+d).

S

23.

Sample variance

The average squared distance from the mean.

x̄
Mean
n − 1
Sample divisor

A sample uses n − 1. Standard deviation is the square root of this.

24.

Sensitivity

Of the people who have the disease, the share the test calls positive.

TP
True positives
FN
False negatives

A sensitive test misses few cases.

25.

Specificity

Of the people without the disease, the share the test calls negative.

TN
True negatives
FP
False positives

A specific test raises few false alarms.

26.

Standard deviation

Typical distance of values from the mean.

On a normal curve, about 95% of values sit within 2 SD of the mean.

27.

Standard error

How much the sample mean would bounce from sample to sample.

SD
Standard deviation
n
Sample size

This is the term a confidence interval multiplies by z. Larger n makes the standard error smaller.

28.

Student's t

High yield

Compare a mean when the population SD is unknown.

x̄
Sample mean
μ
Population mean
s
Sample standard deviation
n
Sample size

Use the sample SD in place of σ. The shape matches a z-score with s instead of the population SD.

Z

29.

Z-score

How many standard deviations a value sits from the mean.

x
The value
x̄
Mean
SD
Standard deviation

z = 1 is one SD above the mean. z = −1 is one SD below it.