Back to Blog Emerging Markets

Emerging Markets and Data Quality: A Framework for What You Can and Cannot Control

Abstract visualization representing data quality assessment framework for emerging markets

When a survey wave from a panel in Nairobi or Karachi comes back with quality flags, the instinct is to treat every anomaly the same way: bad response, exclude it. That reflex is understandable, but it introduces a different kind of error than the one you were trying to avoid.

Not every quality problem in emerging market panels shares a cause. Some are artifacts of the infrastructure respondents are navigating. Others are behavioral signals that indicate inauthentic or disengaged completion. Treating them identically under a single global threshold means you will wrongly exclude legitimate respondents at higher rates in regions where connectivity and device conditions are least predictable, and simultaneously under-detect certain fraud patterns that happen to resemble infrastructure noise.

The framework that guides our work at Besample separates quality problems by cause first, and by response action second. That separation changes which signals you trust, which thresholds you calibrate, and which decisions you escalate for human review.

Two Categories That Require Different Handling

Infrastructure artifacts arise from the environment the respondent is operating in, not from their intent. These include network interruptions that produce apparent timing gaps between items, device-switching mid-survey when a respondent moves from mobile data to WiFi, low-end Android devices with atypical touch latency, GPS readings degraded by dense urban canyon environments, and translation rendering delays in right-to-left scripts that inflate per-item timing compared to left-to-right equivalents.

Behavioral anomalies, by contrast, arise from intent or from disengagement. These include straightlining across a Likert scale, impossible completion speeds that no attentive respondent could achieve regardless of device, declared-versus-device geography mismatches that suggest VPN or proxy use, and systematic attention-check exploitation by respondents who have learned what the trap items look like.

The key distinction is not severity, it is cause. An infrastructure artifact can produce a response that looks suspicious but reflects a genuine respondent navigating difficult conditions. A behavioral anomaly can look unremarkable in aggregate while still indicating low-quality data.

What You Cannot Control

Fieldwork in mobile-primary panels across Sub-Saharan Africa and South Asia will produce timing distributions that look alarming on a threshold calibrated for desktop users in Western Europe. Consider a wave of fifteen thousand responses collected from a mobile panel in Lagos over three weeks. A meaningful portion of completions will show timing gaps of two to four minutes within a survey that should take eight to ten minutes overall. Some of those gaps are respondents pausing to deal with an incoming call. Some are network handoffs from 4G to 3G as the respondent moves. A few are genuine signs of distraction or disengagement. You cannot distinguish these cases using timing data alone, and you should not try to.

Similarly, geolocation precision in dense urban environments is lower than it appears. A respondent completing a survey from a Karachi apartment building may get a GPS fix that places them within a 400-meter radius rather than precisely at their address. If your geolocation validation checks against a declared city with tight tolerance, you may flag a legitimate respondent because building density degraded their fix accuracy.

What You Can Control

Answer patterns are not sensitive to infrastructure conditions. A straightliner on 2G looks identical to a straightliner on fiber: all items receive the same scale value, variance is near zero, and the pattern persists across the entire questionnaire. Pattern-based signals are the most portable across markets precisely because they detect behavioral intent, not environmental conditions.

Attention-check exploitation is similarly infrastructure-independent. A respondent who has learned to identify the reversal item in a scale battery and answer it correctly while otherwise straightlining is detectable through score-combination analysis regardless of their connectivity situation. This is worth noting because attention checks alone do not catch this pattern; you need distributional context to see it.

The practical implication is that pattern-based signals and cross-item behavioral analysis can run at the same sensitivity across all markets. Timing-based signals and geolocation signals require market-specific calibration. Building your detection framework around this distinction is more effective than applying a single sensitivity level to all signals uniformly.

Why Global Defaults Fail at the Fieldwork Level

The major survey platforms calibrated their built-in quality controls primarily against respondent populations in North America and Western Europe. A response-time floor that catches most fraudulent completions in a US-panel context will flag a significantly higher proportion of legitimate responses from a mobile-first MENA panel, where input speed naturally varies more and where device types are less homogeneous.

This is not a criticism of those platforms for their initial calibration decisions. But it means that agencies running fieldwork in MENA, Sub-Saharan Africa, South Asia, and Latin America cannot simply use the default settings and expect them to perform equivalently to what they were tested against. The default is not neutral; it is calibrated for a specific population that does not describe the panels you are fielding in.

The practical consequence we see in practice: agencies accept either a high false positive rate (over-excluding legitimate completes from mobile-primary markets) or they simply turn down their quality controls for these markets, accepting lower sensitivity as the price of not losing too much sample. Both outcomes represent a real cost.

A Layered Detection Approach

The detection framework we've built for Besample reflects this distinction. Signals are categorized by their infrastructure sensitivity, and threshold policies can be configured at the market-archetype level rather than requiring per-country settings. A mobile-first low-bandwidth archetype applies timing and geolocation calibrations appropriate for that environment while running pattern-based signals at full sensitivity. A hybrid mobile/desktop archetype gets different timing ranges but the same pattern logic.

The goal is not to lower your quality bar for markets with difficult field conditions. The goal is to apply the same quality standard using signal settings that reflect the conditions each market presents. A flag should mean the same thing, probabilistically, regardless of which market it comes from. Achieving that requires different underlying threshold configurations, not a different quality philosophy.

The Counter-Argument: Infrastructure as Camouflage

Worth addressing directly: some experienced panelists in high-volume markets have learned to use infrastructure conditions as cover for behavioral fraud. Respondents in markets where quality screening is known to use timing thresholds will sometimes deliberately introduce pauses to avoid flagging. In those environments, genuinely slow completions may be more suspicious than very fast ones.

This is real, and it argues for using timing signals in combination with pattern signals rather than as a primary detection layer. A slow completion with a flat pattern profile is more suspicious than a fast completion with natural variation. Signal combination catches this; timing thresholds alone do not.

We are not saying infrastructure noise makes timing signals worthless. We are saying timing signals without context are insufficient for emerging market quality detection, and that context means combining them with behavioral pattern analysis and treating the combination as more informative than either signal alone.

Making the Distinction Operational

The practical question for a research operations team is how to implement this framework without creating unsustainable configuration overhead. The answer is to think in terms of detection layers rather than per-signal thresholds. The first layer, pattern and behavioral analysis, runs identically across all markets and flags the clearest cases for automatic exclusion. The second layer, timing and geolocation, runs with market-appropriate calibrations and flags cases for review rather than automatic exclusion. The third layer is human review of escalated cases, where an analyst can apply contextual judgment that automated systems can't.

This layered structure does not require a different configuration for every market. It requires a small number of archetypes that map the markets you actually field in, and a clear policy for what each detection layer outputs and how escalated flags are resolved. Getting that policy written down before fieldwork opens is what separates quality assurance that holds up from quality assurance that gets revisited after delivery.

Audit your next wave with Besample

Connect your survey platform and get per-response quality scores as fieldwork runs, not after it closes.

Request Access