🇬🇧 Statistical Officer / Government Statistical Service (GSS) Assessment · flashcards
Statistical Officer / Government Statistical Service (GSS) Assessment Survey Methodology, Data Collection and Sampling Flashcards
55 question-and-answer cards covering Survey Methodology, Data Collection and Sampling as it is examined in Statistical Officer / Government Statistical Service (GSS) Assessment. 24 of them are printed below, taken from across the deck — no signup, no paywall on the preview.
24 sample cards from the Survey Methodology, Data Collection and Sampling deck
Sampled from the end of the deck, so these are different cards from the ones shown on the syllabus page.
Define total survey error (TSE) and its two broad components.
TSE is the total deviation of a survey estimate from the true population value, combining all error sources. Its two broad components are sampling error (from observing a sample not the whole population) and non-sampling error (all other errors), expressed via mean squared error: $MSE = \text{Variance} + \text{Bias}^2$.
List the main 'representation' (coverage/selection) errors in the total survey error framework.
Coverage error (frame omits or duplicates units), sampling error (random variation from sampling), non-response error (sampled units do not respond), and adjustment error (errors introduced by weighting/imputation).
List the main 'measurement' errors in the total survey error framework.
Specification error (concept measured differs from concept intended), measurement/response error (respondent, interviewer, instrument or mode cause incorrect answers), and processing error (editing, coding, data-entry mistakes).
Distinguish sampling error from non-sampling error and which one a larger sample reduces.
Sampling error is random variability because only a sample is observed; it decreases as $n$ increases (roughly with $\frac{1}{\sqrt{n}}$). Non-sampling error (coverage, non-response, measurement, processing) does not shrink with $n$ and can even grow with scale; bigger samples do not fix it.
What is coverage error, and distinguish undercoverage from overcoverage.
Coverage error is the mismatch between the sampling frame and the target population. Undercoverage: target units missing from the frame (e.g. people without a landline). Overcoverage: ineligible, duplicate or out-of-scope units included in the frame.
Define the unit (overall) response rate of a survey.
$$\text{Response rate} = \frac{\text{number of eligible units responding}}{\text{total number of eligible units sampled}}.$$ It is a key quality indicator; eligibility of unknown-status cases is handled using standard definitions (e.g. AAPOR).
Distinguish unit non-response from item non-response.
Unit non-response: a sampled unit provides no usable data at all (refusal, non-contact). Item non-response: a responding unit fails to answer particular questions. Unit non-response is usually addressed by weighting; item non-response by imputation.
Why is non-response bias driven by more than just the response rate?
Non-response bias depends on both the non-response rate and how different non-respondents are from respondents on the variable of interest. A high response rate with respondents like non-respondents gives little bias, while a low rate can give large bias if they differ systematically.
Name strategies used to reduce survey non-response.
Advance letters, incentives, multiple/varied contact attempts and modes, convenient timing, shorter questionnaires, well-trained interviewers, reassurance on confidentiality, refusal conversion, and effective fieldwork management.
What are MCAR, MAR and MNAR mechanisms for missing data?
MCAR (missing completely at random): missingness unrelated to any data. MAR (missing at random): missingness depends only on observed variables, so it can be corrected with those variables. MNAR (missing not at random): missingness depends on the unobserved value itself, which is the hardest to adjust for.
What is a design (base) weight and how is it calculated?
A design weight corrects for unequal selection probabilities so each sampled unit represents part of the population. It is the inverse of the selection probability: $$w_i = \frac{1}{\pi_i},$$ where $\pi_i$ is unit $i$'s probability of selection.
What is non-response weighting adjustment?
Inflating the design weights of respondents to compensate for non-respondents, often via response-propensity classes or modelled response probabilities. Respondents similar to non-respondents receive larger weights so the weighted respondent sample better represents the eligible population.
Explain calibration (post-stratification) weighting and its purpose.
Adjusting weights so weighted sample totals for auxiliary variables (e.g. age, sex, region) match known population totals from the census or administrative sources. It reduces coverage and non-response bias and can improve precision when the auxiliaries correlate with the survey variable.
Why can weighting reduce bias but increase variance?
Highly variable weights mean some units count for much more than others, increasing the variance of estimates (reflected in a larger design effect). There is a bias–variance trade-off, so extreme weights are often trimmed or capped to control variance inflation.
What is imputation and name two common methods.
Imputation fills in missing item values with plausible estimates so analysis can use complete records. Common methods: mean/ratio imputation, hot-deck (donor) imputation from similar respondents, regression/predictive imputation, and multiple imputation (which reflects imputation uncertainty).
Name the ONS labour market survey and what it measures.
The Labour Force Survey (LFS) is the main household survey of the labour market, measuring employment, unemployment and economic inactivity using International Labour Organization (ILO) definitions. It is the source of official headline employment and unemployment rates.
Which ONS survey measures household income, spending and inflation weights, and what is it used for?
The Living Costs and Food Survey (LCF) collects household expenditure and income; it provides the expenditure weights for the Consumer Prices Index (CPI/CPIH) basket and feeds analyses of household spending and poverty.
Name two further major ONS household/social surveys besides the LFS and LCF.
Examples: the Annual Population Survey (APS, a boosted LFS for local-area estimates); the Opinions and Lifestyle Survey (OPN, rapid 'omnibus' social attitudes); the Wealth and Assets Survey; and the (former) English Housing Survey/Survey of Living Conditions. Any two are acceptable.
What is the UK Census, how often is it run, and which bodies run it?
The Census is a complete count of the whole population and households, conducted every 10 years (most recently 2021 in England, Wales and Northern Ireland; Scotland ran in 2022). ONS runs it for England and Wales, NISRA for Northern Ireland and National Records of Scotland for Scotland.
What are the main uses of census data in the statistical system?
Providing benchmark population counts and small-area/sub-group detail; supplying population totals used to weight and calibrate sample surveys; constructing sampling frames; allocating public funding; and informing planning. It is the key source against which survey estimates are anchored.
What are the strengths and limitations of a traditional census?
Strengths: near-complete coverage, detailed small-area and small-group data, a benchmark for other sources. Limitations: very expensive, large respondent burden, only every 10 years so data quickly age, and still subject to coverage error (corrected via a Census Coverage Survey and estimation).
What is the difference between aggregate and linked microdata from administrative sources?
Aggregate administrative data are pre-summarised counts/totals. Linked microdata join individual-level records across sources (e.g. tax, benefits, health) at the unit level, enabling richer analysis but requiring matching keys, governance and strict confidentiality safeguards.
What is data linkage and what are deterministic versus probabilistic matching?
Data linkage joins records referring to the same unit across datasets. Deterministic matching links records on exact agreement of identifiers (e.g. NHS number). Probabilistic (fuzzy) matching uses weighted agreement across several imperfect fields to estimate match likelihood when no unique identifier exists.
What are the main quality and ethical challenges of administrative and linked data for statistics?
Coverage and definitional mismatch with the target concept, changing administrative rules over time, data quality and missingness, linkage error (false matches/missed matches), and privacy/confidentiality and legal-gateway requirements governing access and disclosure.
What this deck covers
The Survey Methodology, Data Collection and Sampling deck follows the Statistical Officer / Government Statistical Service (GSS) Assessment Survey Methodology, Data Collection and Sampling syllabus — 4 chapters and 12 topics — so questions land on material that is genuinely examinable rather than trivia around it. That works out to roughly 13.8 cards per chapter.
Answers are written to be recallable, not just readable — averaging about 271 characters, which is long enough to carry the reasoning and short enough to say out loud.
A deck like this earns its keep on the second and third pass. Read the syllabus first so you know the shape of the subject, then use the cards to find the specific facts that have not stuck.
Survey Methodology, Data Collection and Sampling flashcards FAQ
How many Survey Methodology, Data Collection and Sampling flashcards are in this Statistical Officer / Government Statistical Service (GSS) Assessment deck?
55 cards. This page previews 24 of them, sampled evenly across the deck so you can judge the difficulty before installing anything.
Are these Statistical Officer / Government Statistical Service (GSS) Assessment flashcards free?
Yes. The preview here is free to read with no signup, and the full 55-card deck is free inside the Examius app.
What do the Survey Methodology, Data Collection and Sampling cards cover?
They follow the Statistical Officer / Government Statistical Service (GSS) Assessment Survey Methodology, Data Collection and Sampling syllabus — 4 chapters and 12 topics — so the questions track what is actually examinable.
How should I use these flashcards?
Read the syllabus first so you know the shape of the subject, then drill the deck. Examius schedules each card with spaced repetition, so cards you keep missing come back sooner and ones you know drift further apart.