🌍 Statistics & Probability · flashcards
Statistics & Probability Descriptive Statistics & Data Exploration Flashcards
51 question-and-answer cards covering Descriptive Statistics & Data Exploration as it is examined in Statistics & Probability. 24 of them are printed below, taken from across the deck — no signup, no paywall on the preview.
24 sample cards from the Descriptive Statistics & Data Exploration deck
Sampled from the end of the deck, so these are different cards from the ones shown on the syllabus page.
Define the coefficient of variation (CV) and give its formula.
A relative, unit-free measure of dispersion: $$\text{CV} = \frac{\sigma}{\mu} \times 100\%$$ It expresses the standard deviation as a percentage of the mean.
When is the coefficient of variation especially useful?
When comparing the variability of two or more data sets that have different units or very different means, since it is dimensionless.
State the formula for the mean absolute deviation (MAD) about the mean.
$$\text{MAD} = \frac{1}{n}\sum_{i=1}^{n} \lvert x_i - \bar{x} \rvert$$
How does the mean absolute deviation differ from the standard deviation?
MAD averages the absolute deviations from the mean, while standard deviation is based on squared deviations; MAD is less sensitive to outliers and is not differentiable at zero.
What does skewness measure?
The asymmetry of a distribution about its mean. Positive skew has a longer right tail; negative skew has a longer left tail; zero skew indicates symmetry.
Give a formula for the (population) coefficient of skewness.
$$\gamma_1 = \frac{E\left[(X-\mu)^{3}\right]}{\sigma^{3}} = \frac{1}{n}\sum_{i=1}^{n}\left(\frac{x_i - \bar{x}}{\sigma}\right)^{3}$$
What does kurtosis measure?
The 'tailedness' of a distribution — how heavy the tails are and how peaked it is relative to a normal distribution.
What is excess kurtosis, and what is its value for a normal distribution?
Excess kurtosis is kurtosis minus 3: $$\gamma_2 = \frac{E\left[(X-\mu)^{4}\right]}{\sigma^{4}} - 3$$ For a normal (mesokurtic) distribution it equals 0.
Define the terms leptokurtic, mesokurtic, and platykurtic.
Leptokurtic: heavier tails and sharper peak (excess kurtosis $>0$); mesokurtic: like the normal (excess kurtosis $=0$); platykurtic: lighter tails and flatter peak (excess kurtosis $<0$).
What is the $p$-th percentile of a data set?
The value below which approximately $p\%$ of the ordered observations fall.
What are quartiles, and which percentiles do they correspond to?
Quartiles divide ordered data into four equal parts: $Q_1$ = 25th percentile, $Q_2$ = 50th percentile (median), $Q_3$ = 75th percentile.
Distinguish quartiles, quintiles, deciles, and percentiles.
They split ordered data into equal groups: quartiles into 4, quintiles into 5, deciles into 10, and percentiles into 100.
State the formula for a z-score and what it represents.
$$z = \frac{x - \mu}{\sigma}$$ It gives the number of standard deviations an observation $x$ lies above (positive) or below (negative) the mean.
What are the mean and standard deviation of a set of standardized (z-score) values?
A standardized distribution always has mean $0$ and standard deviation $1$.
Why is standardization (converting to z-scores) useful?
It puts values from different distributions or units on a common scale, allowing direct comparison of relative positions.
According to the empirical (68-95-99.7) rule for a normal distribution, what fraction of data lies within one, two, and three standard deviations of the mean?
Approximately $68\%$ within $\mu \pm \sigma$, $95\%$ within $\mu \pm 2\sigma$, and $99.7\%$ within $\mu \pm 3\sigma$.
What is a histogram?
A graph displaying the distribution of a quantitative variable by dividing data into bins (intervals) and drawing adjacent bars whose heights (or areas) represent the frequency or relative frequency in each bin.
How does a histogram differ from a bar chart?
A histogram shows continuous quantitative data with adjacent bars (no gaps) and ordered numeric bins; a bar chart shows categorical data with separated bars in any order.
What is a frequency distribution?
A table or summary that shows the number of observations (frequency) falling into each category or class interval of a data set.
What is the difference between frequency, relative frequency, and cumulative frequency?
Frequency is the raw count in a class; relative frequency is that count divided by the total ($\frac{f}{n}$); cumulative frequency is the running total of frequencies up to and including a class.
What are the five values in the five-number summary shown by a box plot?
Minimum, first quartile $Q_1$, median $Q_2$, third quartile $Q_3$, and maximum.
State the standard (1.5 IQR) rule for identifying outliers in a box plot.
A value is an outlier if it lies below $Q_1 - 1.5 \times \text{IQR}$ or above $Q_3 + 1.5 \times \text{IQR}$.
In a box plot, what do the 'whiskers' represent?
The whiskers extend from the box to the smallest and largest data values that are not outliers (i.e. within $1.5 \times \text{IQR}$ of the quartiles); points beyond them are plotted individually as outliers.
How can the shape of a box plot indicate skewness?
If the median is closer to $Q_1$ and the upper whisker/right side is longer, the data are right-skewed; if the median is closer to $Q_3$ with a longer lower whisker, they are left-skewed; a centered median with equal whiskers suggests symmetry.
What this deck covers
The Descriptive Statistics & Data Exploration deck follows the Statistics & Probability Descriptive Statistics & Data Exploration syllabus — 6 chapters and 25 topics — so questions land on material that is genuinely examinable rather than trivia around it. That works out to roughly 8.5 cards per chapter.
Answers are written to be recallable, not just readable — averaging about 142 characters, which is long enough to carry the reasoning and short enough to say out loud.
A deck like this earns its keep on the second and third pass. Read the syllabus first so you know the shape of the subject, then use the cards to find the specific facts that have not stuck.
Descriptive Statistics & Data Exploration flashcards FAQ
How many Descriptive Statistics & Data Exploration flashcards are in this Statistics & Probability deck?
51 cards. This page previews 24 of them, sampled evenly across the deck so you can judge the difficulty before installing anything.
Are these Statistics & Probability flashcards free?
Yes. The preview here is free to read with no signup, and the full 51-card deck is free inside the Examius app.
What do the Descriptive Statistics & Data Exploration cards cover?
They follow the Statistics & Probability Descriptive Statistics & Data Exploration syllabus — 6 chapters and 25 topics — so the questions track what is actually examinable.
How should I use these flashcards?
Read the syllabus first so you know the shape of the subject, then drill the deck. Examius schedules each card with spaced repetition, so cards you keep missing come back sooner and ones you know drift further apart.