Key Takeaways
- A baseline survey is not a reporting formality. Baseline survey design is the decision that determines whether a project can demonstrate results at completion.
- Women’s economic empowerment is measured at the level of the individual woman, not the household. Household-level proxies hide exactly the changes these projects are designed to produce.
- Indicators should be selected on the basis that they can be repeated at endline with the same definitions, the same instrument and a comparable sample.
- If the project intends to attribute results to its activities, the comparison group must be designed into the baseline survey. It cannot be reconstructed afterwards.
- Field design decides data quality: who interviews the respondent, whether she can answer privately, what language the instrument is administered in, and when in the agricultural or migration calendar fieldwork takes place.
- Ethical protocols are not optional. Questions about income control, household decisions and safety require informed consent, privacy and a referral pathway.
- The deliverable is not only a report. Without the dataset, codebook, sampling documentation and analysis syntax, the endline cannot be compared with the baseline survey.
Introduction
Baseline survey design is one of the first technical decisions a Project Implementation Unit (PIU) has to get right, and one of the easiest to get wrong. Across Africa, Asia and Central Asia, a baseline survey for a women’s economic empowerment project is usually commissioned within the first year of implementation and usually under time pressure. The results framework already contains indicators, the financing agreement requires baseline values, and the PIU needs numbers before the first progress report.
What follows is predictable. A consultant is engaged, a survey is fielded, a report is submitted, and the numbers enter the results framework. Three or four years later, at endline, the project discovers that the baseline survey cannot support the claim it now needs to make. The indicator was measured at household level while the project worked with individual women. The sample cannot be located again. The definitions moved. There is no comparison group. The dataset was never handed over, only the report.
These failures are not caused by weak survey firms. They are baseline survey design failures, made in the first few weeks by people managing many other things at the same time.
This guide sets out what those decisions are, written from the perspective of a PIU that has to commission the survey, supervise it and defend the results later. It applies to women’s economic empowerment projects in agriculture, enterprise development, financial inclusion, skills and employment, and to the labour-force and livelihood components of larger infrastructure and rural development operations.
What a Baseline Survey Is Actually For
A baseline survey serves three separate purposes, and problems arise when a project pursues one and assumes it has achieved the others.
Reporting. It establishes the starting values for the indicators in the results framework. This is the purpose most PIUs have in mind, and it is the least demanding of the three.
Targeting and design. It describes the situation of the intended participants: what they earn, what they own, what constraints they face, what services they already use. A good baseline survey changes project design, because it usually shows that assumptions made during preparation were partly wrong.
Evaluation. It creates the basis for measuring change, and, where a comparison group exists, for attributing that change to the project rather than to everything else happening in the economy.
The third purpose imposes requirements the first does not. If the project will only ever report before-and-after values, a descriptive survey of participants is enough. If anyone will later ask whether the project caused the change, the comparison group has to be designed and surveyed at baseline. This is the single most consequential decision in the whole exercise, and it is usually made by default rather than deliberately.
The question a PIU should settle before drafting the terms of reference is simple: at completion, will we be asked what changed, or will we be asked what we caused? The answer determines the design and the budget.

Baseline Survey Design: Choosing Indicators That Survive to Endline
Economic empowerment is not a single quantity. It covers what a woman earns, what she owns, what she controls, what she can decide, and what she can do without someone’s permission. Projects that reduce it to one convenient number usually measure the part that moves least.
Two errors are common. The first is measuring at household level: household income, household assets, household food security. These are easier to collect and they are what national surveys often report, but a project can raise household income without changing anything for the women it targeted. The second is choosing sophisticated measures at baseline that the project cannot afford to repeat at endline, which leaves the indicator stranded.
Established instruments are worth examining before designing questions from scratch. The Women’s Empowerment in Agriculture Index and its later versions, including pro-WEAI, provide tested modules on decision-making, control over income, time use, group membership and mobility, with published questionnaires and calculation guidance. Even projects outside agriculture can adapt individual modules rather than invent their own. For country-level context and comparison values, the World Bank Gender Data Portal provides sex-disaggregated indicators across more than a thousand series. The World Bank’s DIME Analytics team also maintains practical guidance on questionnaire design that is worth reading before a terms of reference is drafted.
Seven measurement choices account for most of the difference between a baseline survey that holds up and one that does not.
| What projects usually measure | What to measure instead |
| Household monthly income, reported by whoever is at home | The woman’s own earnings, activity by activity, over a defined recall period |
| Whether she has access to household income | Who decides how the earnings from each activity are used |
| Household asset ownership | What she owns individually, and whether she can sell, rent or bequeath it |
| Whether she has a business | Revenue, costs, operating days, employees and whether the business is still trading |
| Whether the household has a bank account | Accounts in her own name, who controls them, and actual use in the past period |
| Hours worked, self-estimated | A structured time-use module covering paid work, unpaid care and domestic work |
| A single question about empowerment | Decision-making across defined domains, and travel to named destinations without permission |
The right-hand column costs interview time. Structured earnings recall, enterprise performance and time use are the three expensive modules, and a questionnaire cannot carry all of them alongside everything else the project wants to know. That is a choice to make deliberately at design stage, not a corner for the survey firm to cut in the field.
One domain needs separate treatment. Questions about violence or safety should not be added to a general economic survey as a few extra items. They require specialist protocols, dedicated enumerator training, privacy conditions that a routine household interview cannot guarantee, and a referral pathway for respondents who disclose. If the project genuinely needs this data, commission it properly. If it does not, leave it out rather than collecting it badly.
For each indicator finally selected, one test should be applied before the terms of reference are issued: can this be collected again at endline, with the same wording, by a different firm, within a comparable budget? If not, replace it.
Sampling When the Target Population Is Not on Any List
Sampling is where baseline survey design goes wrong quietly. Most national surveys sample from a census frame. Projects rarely have that option, because the population of interest is women who meet project criteria, and no list of them exists before the project starts.
Three approaches are used in practice, and each has consequences the PIU should understand.
Listing within selected clusters. Villages or urban blocks are selected first, often with probability proportional to size, and a quick listing exercise in each identifies eligible women. This produces a defensible sample but adds a field stage and cost.
Sampling from project application or registration lists. Cheaper and faster, but the people who applied differ systematically from those who did not. Results describe applicants, not the wider population, and this must be stated plainly in the report rather than discovered at evaluation.
Sampling from administrative registries. Cooperative membership, business registration, social protection rolls. Useful where coverage is genuinely high, misleading where informal activity dominates, which is precisely the situation in most of these projects.
Sample size should follow from what the project will need to detect, not from a round number. A survey sized to describe a population is smaller than one sized to detect a plausible change over time, which in turn is smaller than one sized to compare treated and untreated groups. Decide which of the three you are buying.
Two details are routinely omitted from terms of reference and cause disputes later: the replacement rule when a selected respondent cannot be found, and the maximum acceptable non-response rate. Specify both, and require that the final report state the achieved rates.
Field Design Decides the Data Quality
Baseline survey design does not end when the questionnaire is signed off. Four field decisions determine whether the data that comes back is usable.
Who conducts the interview
In many contexts across Asia and Africa, a woman interviewed by a male enumerator, or in the presence of her husband or mother-in-law, gives different answers about income, control and decision-making than she gives in private. This is not a cultural nicety. It is a measurement error that moves in a predictable direction and can be larger than the effect the project hopes to demonstrate.
Require female enumerators for female respondents where context demands it, and require the interview to be conducted with the respondent alone. Make the enumerator record who else was present. That variable is worth having at analysis stage.
Language and translation
Instruments drafted in English and translated once, without back-translation and without piloting in the actual dialect used in the field, produce questions enumerators then paraphrase differently in every district. Require translation into each language of administration, back-translation, and a pilot of at least 30 to 50 interviews with a written debrief. Terms like “income”, “asset”, “decide” and “own” rarely have clean single-word equivalents, and the choices made in translation should be documented so the endline uses the same ones.
Enumerator training and supervision
Training of a week or more, including practice interviews and field piloting, separates usable data from noise. Supervision arrangements should be explicit in the contract: supervisor-to-enumerator ratio, the proportion of interviews back-checked, and what happens when back-checks fail. Digital data collection using standard platforms allows validation rules, GPS stamps, interview duration records and audio audits, all of which make supervision real rather than nominal.
Timing in the calendar
Seasonality is not a detail. Agricultural income, food security, time use and even household composition shift across the year. In areas with high seasonal labour migration, who is at home in March and who is at home in September are different populations. Religious calendars, harvest periods and school terms all affect availability and answers.
The practical implication is that the endline should be conducted in the same season as the baseline survey. Record the fieldwork dates prominently, because someone commissioning the endline three years later will need them.
Ethics and Safeguarding in a Baseline Survey
Surveys that ask women about earnings, control over money and household decisions can create risk for the respondent, particularly where the answers contradict what a household member would say. Minimum standards apply regardless of country:
- Informed consent in the respondent’s language, including the right to stop at any point and to skip questions
- No implied promise that participation affects selection for project benefits, which also protects data quality
- Private interview conditions, and a protocol for postponing an interview that cannot be conducted privately
- A referral pathway to available services if a respondent discloses harm, with enumerators trained in how to respond
- Protection of personal data, with identifiers stored separately from responses and access restricted
- Any national ethical clearance or statistical office approval required for surveys in that country, confirmed before fieldwork rather than during it
Where any part of the survey touches on safety or violence, the World Health Organization’s ethical and safety recommendations for research on violence against women set the standard that donors and ethics committees expect, and they should be referenced in the terms of reference rather than left to the firm’s discretion.
These requirements belong in the terms of reference. Firms price what is written; commitments raised after contract signature are rarely delivered fully.
Timing the Baseline Survey Against Project Launch
A baseline survey measures the situation before project exposure. When procurement runs long and activities start first, the survey no longer does that, and the honest response is to record it rather than obscure it.
Where the baseline is late, note which sampled areas had already received project activities and for how long. Partial exposure documented at baseline can still be analysed. Undocumented exposure quietly invalidates the comparison later.
Where the baseline survey can be sequenced properly, the practical route is to prepare the terms of reference during project preparation, launch procurement immediately at effectiveness, and hold first disbursement of field activities in sampled areas until enumeration is complete. This requires the M&E function and the procurement function to be coordinated from the start, which is an institutional arrangement rather than a technical one.
Making the Endline Comparable
Most baseline surveys are commissioned as standalone assignments and delivered as reports. Comparability is a baseline survey design requirement, not an endline problem. The endline is procured years later, often from a different firm, often with different staff at the PIU. Comparability has to be built deliberately.
Require the following as contractual deliverables, not as optional annexes:
- The final instrument in every language of administration, as fielded
- The clean dataset in a standard format, with variable labels and value labels
- A codebook documenting every variable and how derived variables were constructed
- The analysis syntax or scripts that produced the reported figures
- Full sampling documentation: the frame, selection procedure, weights and how they were calculated
- Fieldwork documentation: dates, achieved sample, non-response and replacement rates, and quality control results
- Where a panel is intended, tracking information with respondent consent and a data protection protocol
Tracking the same respondents at endline is what makes change measurable at the individual level, and it is also where attrition threatens the whole design. Plan for it at baseline: collect the contact and location details needed to find people again, plan for a tracking budget, and expect attrition to be higher among the mobile, the young and the poor, which is to say among the people whose outcomes matter most.
Common Baseline Survey Design Mistakes PIUs Make
Treating the baseline survey as a compliance deliverable. It is procured to populate a results framework column, so nobody asks whether the design can support the claims the project will need to make at completion.
Measuring the household when the project targets the woman. The most common single error, and it is usually irreversible once the endline arrives.
Expecting attribution without a comparison group. Projects report “beneficiary incomes rose by 22 per cent” and cannot say whether incomes rose for everyone else as well.
Selecting indicators that cannot be repeated. A rich baseline followed by a cheap endline produces two surveys that cannot be compared.
Fielding the survey after activities begin. Then reporting the values as if they were pre-exposure.
Skipping the pilot. Question problems, translation problems and interview-length problems are all cheap to fix before fieldwork and expensive afterwards.
Over-long instruments. Interviews beyond about 60 to 75 minutes produce fatigue, and the last modules, often the empowerment ones, carry the worst data.
Accepting a report without the dataset. The PIU is left holding conclusions it cannot re-analyse, verify or compare.
Ignoring seasonality and migration. Fieldwork scheduled around procurement convenience rather than the calendar of the population being surveyed.
No plan for finding respondents again. Panel tracking is decided at endline, when it is already too late.
What PIUs Should Check Before Accepting a Baseline Survey Report
Design
- Are indicators measured at the individual level where the project targets individual women?
- Is every results framework indicator actually covered by the instrument?
- Is the sampling frame described, with its coverage limitations stated?
- Are sample size and precision justified against the project’s intended use of the data?
- If attribution is expected, was a comparison group surveyed on the same instrument at the same time?
Fieldwork
- Who conducted the interviews, and were female respondents interviewed by female enumerators where required?
- Were interviews conducted privately, and is the presence of others recorded?
- Was the instrument translated, back-translated and piloted, and were pilot findings acted on?
- What were the achieved sample, non-response and replacement rates?
- What quality control was applied, and what did back-checks and audio audits show?
- When did fieldwork take place, and where does that fall in the agricultural and migration calendar?
Deliverables
- Has the clean dataset been delivered, with labels, codebook and analysis syntax?
- Can the reported figures be reproduced from the dataset provided?
- Are sampling weights documented and applied correctly in the reported estimates?
- Are consent procedures, ethical approvals and data protection arrangements documented?
- Is tracking information available for a future endline, with consent obtained?
- Does the report state its own limitations, or does it present every figure with equal confidence?
Critical Questions Project Leaders Should Be Asking
- At completion, will we be asked what changed, or what we caused, and is the baseline survey designed for that answer?
- Can every indicator in our results framework be collected again in three years, at a cost we will have?
- Who will hold the dataset, the codebook and the tracking information when the current PIU staff have moved on?
- Will the baseline be complete before project activities reach the sampled areas?
- Are our survey requirements written into the terms of reference, or are we relying on the firm to propose them?
- Does anyone in the PIU have the capacity to review the methodology, or are we accepting the report on trust?
- What will we do differently in project design once the baseline findings arrive?
Building In-House Baseline Survey Design Capacity
The quality of a baseline survey is set largely by the terms of reference and by the supervision applied during fieldwork. Both sit with the client, not the consultant. A PIU that can specify sampling requirements, insist on piloting, read a methodology section critically and refuse an incomplete data deliverable gets better surveys from the same market of firms than a PIU that cannot.
This is why survey capacity belongs inside the implementing agency rather than being outsourced entirely. Ministries and agencies running successive operations benefit from standard terms of reference, a standard data deliverables list, a review checklist and a place where datasets from previous projects are actually kept. Measurement then improves across a portfolio instead of restarting with every project.
The same logic applies to how results are used afterwards. A baseline survey that is read, questioned and allowed to change targeting is worth more than a technically flawless one that is filed on receipt. For further reading on measuring outcomes rather than activities in displacement contexts, see our guide to livelihood restoration under World Bank ESS5, where the same measurement discipline determines whether restoration can be demonstrated.
Conclusion
A baseline survey for a women’s economic empowerment project is a measurement system, not a document. Baseline survey design fixes what will be counted, at what level, for whom, and against what comparison. Those choices are made early, usually quickly, and they are difficult to revisit once fieldwork has been completed.
The projects that can demonstrate results at completion are rarely those with the most elaborate baselines. They are the ones that measured individual women rather than households, chose indicators they could afford to repeat, designed the comparison they would later need, fielded the survey before activities began, and kept the dataset, the codebook and the means to find respondents again.
For PIUs, the practical takeaway is that the decisive work happens before procurement. The terms of reference are where the quality of the endline is determined, three or four years in advance.
Planning a baseline, endline or beneficiary survey? Risalat Consultants International supports government institutions, PIUs and project teams with practical capacity development in baseline survey design, research methods and data analysis, monitoring and evaluation, and gender equality and social inclusion.
Explore our research and statistics training programs or talk to our team about tailored support for your project team.








