Similarity, data & statistics

Sampling techniques in research: which method to use and how to defend it

Which sampling method to use for your study, and how to defend the choice to your committee.

Last reviewed · 8 min read

Choose a probability sampling method (simple random, systematic, stratified, cluster or multistage) when you have, or can build, a list of your population and want to generalise your figures to it; choose a non-probability method (convenience, purposive, quota, snowball or theoretical) when no such list exists or your aim is depth rather than estimation, and name the method honestly. The choice follows from your research question and your sampling frame, not from what the last thesis in your department did.

This guide covers the terms examiners expect you to get right, the methods side by side, three situations Indian scholars keep running into, how to write the sampling paragraph, and the objections that come up at DC meetings and in examiner reports. It is about how units are chosen. For how many, see our sample size calculation guide. Last reviewed September 2026.

What are population, sampling frame and sampling unit?

Get these three right before you pick a method. Most viva questions about sampling come back to these three.

The population is the group your conclusions are meant to cover, with boundaries of who, where and when. “Nurses” is not a population. “Staff nurses in government district hospitals in Karnataka, 2025–26” is.

The sampling frame is the list you actually select from, and it is almost never identical to the population. The district teachers’ roll misses contract staff appointed last month; a college register includes students who stopped attending. Say what your frame misses. Examiners forgive an imperfect frame you described far more readily than a perfect-sounding one you didn’t.

The sampling unit is what you select: a person, household, enterprise, village or college. Multistage designs have one at each stage. India’s Periodic Labour Force Survey shows how to describe this. MoSPI’s note on its sample design calls it a stratified multi-stage design, with Census 2011 villages and Urban Frame Survey blocks as first-stage units and households as the ultimate units, selected by simple random sampling without replacement.

Probability vs non-probability sampling: which one does your study need?

In probability sampling, chance decides who is picked and every unit in the frame has a known chance of selection. That is what lets you put a margin of error on an estimate. In non-probability sampling, you, your contacts or plain availability decide. The results can be valuable but can’t be projected onto a population with stated precision.

A quick test: if your question asks about the “level of”, “proportion of” or “extent of” something in a population, you need probability sampling or a clear reason you couldn’t get it. If it asks how or why, you probably want purposive sampling, and random selection would spend your fieldwork on people with little to say. Your research design usually settles this first.

Sampling methods compared for a PhD thesis
MethodWhen it fitsWhat you needHow examiners react
Simple randomSmall population with a complete listA current list and random numbersNo objection if you show the list existed
SystematicRegisters and queues: every 5th OPD patientA random start and a fixed intervalAccepted; may ask if the list has a repeating pattern
Stratified, proportionateGroups that must appear in their true share (rural/urban, government/private)A frame showing each unit’s stratumWelcomed
Stratified, disproportionateA small group you must analyse separatelyWeights when you report overall figuresFine with weights; a problem if pooled without them
ClusterA list of groups (schools, villages) but not of peopleA list of clustersAccepted; expect questions on the larger sample needed
MultistageSpread-out populations: district to village to householdA frame and a written rule at each stageRespected if every stage is described
ConveniencePilots, or no frame and access is the limitHonest naming and modest claimsTolerated if named; rejected if called random
PurposiveQualitative work, expert panels, case selectionInclusion criteria set before fieldworkStandard; challenged if findings are generalised
QuotaNo frame, but known population sharesProportions from the Census, AISHE or similarA disciplined convenience sample, and seen as one
SnowballHidden groups: informal lenders, migrant workersSeveral starting contacts and referral recordsAccepted for such groups
TheoreticalGrounded theoryCoding alongside data collection, and memosExpected in grounded theory; odd elsewhere

Disproportionate stratified sampling (over-sampling a small group so you can analyse it) is perfectly respectable, but pooled figures then need weights, and scholars often forget. Cluster samples lose precision because students in one college share teachers, fees and catchment, so plan for more respondents.

Which sampling technique to use when there is no list? Three Indian examples

These are illustrative topics, not real studies.

MSME owners in one district

A commerce scholar is studying digital payment adoption among small manufacturers in one district. Her supervisor says “take 400 random MSMEs”. Random from what? The public Udyam Registration portal shows national totals on its dashboard, not a list you can draw from, and many small units never registered.

Her options, in the order I’d try them. First, ask the District Industries Centre or an industrial estate association what list they hold; a partial list honestly described beats none. Second, build a frame in a sample of areas: pick industrial areas or wards at random, list every unit you find there, then sample from your own list (two-stage area sampling). Third, if neither works, use quota sampling against known shares by sector and size, and call it that. What she must not do is collect 400 responses through WhatsApp groups and write “random sampling”.

Patients in a hospital

A nursing scholar studying medication adherence among diabetic outpatients has no fixed list, because patients arrive daily. Use systematic sampling from the OPD register (every k-th eligible patient from a random start), or consecutive sampling, inviting every eligible patient over a fixed period until the target is met. Consecutive sampling is still non-probability, but it removes your choice of whom to approach. Record the days and hours covered, since a Monday-morning clinic sees different patients from a Friday-evening one.

Students across affiliated colleges

An education scholar wants to measure academic stress among undergraduates in colleges affiliated to one state university. A list of 60,000 students is rarely available, but a list of colleges exists, from the university’s affiliation records or the Ministry of Education’s AISHE directory of institutions. So stratify colleges by type (government, aided, self-financing), pick colleges at random within each stratum, pick sections at random within each college, and survey every student present. Decide in advance what happens when a principal refuses (the next randomly drawn college in the same stratum) and report how many refused.

Purposive, snowball and theoretical sampling in qualitative research

Qualitative samples are chosen for what people can tell you. Palinkas and colleagues (2015), in a widely cited paper on purposeful sampling, describe it as selecting information-rich cases: people especially knowledgeable about or experienced with the phenomenon. Set your criteria before fieldwork, and name the variant where one fits (maximum variation, typical case, criterion).

Snowball sampling suits groups with no list, such as moneylenders or caregivers of children with a rare condition. Its weakness is that everyone knows everyone. Start from several unconnected contacts and keep a record of who referred whom.

Theoretical sampling belongs to grounded theory. In The Discovery of Grounded Theory (1967), Glaser and Strauss describe the analyst collecting, coding and analysing data together and deciding what to collect next, and where, from the emerging theory. You can’t fix the whole sample in your synopsis; describe the starting sample and the logic for later choices, and let your memos be the evidence. If you are doing thematic analysis rather than grounded theory, don’t claim theoretical sampling. What you have is purposive, and our thematic analysis guide covers the analysis.

How to write the sampling section of your methodology chapter

Leave out the textbook definitions. Your examiners know what stratified sampling is; they want to know what you did, clearly enough to repeat it. Our research methodology guide shows where this sits in the chapter.

  1. Define the population

    Who, where and when: owners of registered micro and small manufacturing units in Coimbatore district in 2025, not "MSMEs".

  2. Name the sampling frame

    The actual list you drew from, who keeps it, its date and what it misses. If there was no list, say so.

  3. Name the method and each stage

    The technical name, then the mechanics: how strata or clusters were formed, how random numbers were generated, what happened on refusal.

  4. Give the numbers

    Approached, agreed, usable after screening. Refer to your sample size calculation rather than repeating it.

  5. State what the sample lets you claim

    One or two sentences on how far the findings can be generalised, and to whom.

A worked example

The college study above might read like this (numbers invented for illustration):

The population comprised undergraduates enrolled in 2025–26 in arts and science colleges affiliated to the university. The frame for colleges was the university’s list of affiliated colleges (148, as on 1 June 2025); no list of students was available. A two-stage stratified cluster design was used. Colleges were stratified by management type and four were selected from each stratum using computer-generated random numbers; two colleges that declined were replaced by the next drawn college in the same stratum. In each college two second-year sections were chosen at random and all students present were invited. Of 1,020 students invited, 874 responded and 841 responses were retained after screening. Because strata of different sizes contributed equal numbers of colleges, pooled estimates were weighted by stratum size. Findings apply to second-year undergraduates in the university’s affiliated arts and science colleges.

No definitions, and no “random” anywhere it wasn’t. The final sentence is the one scholars most often leave out. If your instrument is a questionnaire, the pilot details come next; see questionnaire design.

Common examiner objections to sampling

Each of these is easy to fix before submission and tedious to fix after an examiner’s report.

  • “Randomly selected” for questionnaires handed to whoever was free in the staff room. Call it convenience sampling and the objection disappears.
  • A purposive sample of 20 people followed by claims about “entrepreneurs in India”.
  • Stratified sampling in the methodology chapter, but no stratum-wise results or weights in the analysis.
  • Cluster samples analysed as if each respondent were chosen independently, which makes standard errors look smaller than they are.
  • Two pages defining every sampling type, and one line on what the scholar actually did.
  • No response rate, or 90% for an online survey with no explanation.

The analysis has to match the sample. Inferential tests assume random selection; with a convenience sample you can still run them, as most Indian theses do, but describe the results as applying to your respondents. Our guide to choosing a statistical test takes that honesty as given.

FAQ

Questions scholars ask

Is convenience sampling acceptable in a PhD thesis?

Yes, if you name it, explain why a probability sample was not possible, and keep conclusions to the people studied. Etikan, Musa and Alkassim (2016) note that non-probability sampling is useful when randomisation is impossible.

What is the difference between purposive and convenience sampling?

Purposive sampling selects people who meet criteria set in advance. Convenience sampling takes whoever is easy to reach. One has a reason behind each inclusion; the other has only access.

What is the difference between stratified and cluster sampling?

Stratified sampling samples from every group, so each is represented. Cluster sampling selects only some groups and studies people within them. Stratifying usually improves precision; clustering usually reduces it but saves travel and permissions.

Can I use two sampling techniques in one thesis?

Yes. Multistage designs combine methods by definition, and mixed methods theses often pair a probability survey sample with a purposive interview sub-sample. Describe each separately.

Which sampling technique is best for a questionnaire survey?

Whichever your frame allows: simple random or stratified with a full list, cluster or multistage with a list of groups, quota with no list. Sample size is a separate decision, covered in our sample size calculation guide.

Keep reading

Related guides and services

Want a second pair of eyes on this?

These guides are free, with no sign-up. If you’d like a specialist to look at your own thesis, journal or data, the first consultation is free too, and we’ll tell you clearly whether you need us.

Book a free consultation
Call WhatsApp Free consultation