Back to Blog Product

Why We Built Besample: The Manual QC Problem at Scale

Abstract concept representing the founding problem behind Besample

The problem we built Besample to solve is not a new one. Research agencies that run fieldwork in emerging markets have been managing response quality manually for as long as those panels have existed. The reason it became the problem we decided to address is simpler than it might appear: manual quality checking does not scale to the volumes that modern emerging market fieldwork demands, and the cost of not scaling it shows up where it hurts most.

Not in your QC budget. In your client relationships.

What Manual QC Actually Looks Like at Scale

When a wave closes, someone on the research operations team needs to look at the data before it goes to delivery. The professional standard is to run the toplines, scan for obviously anomalous distributions, check open-text fields for coherence, and review any responses that flagged during fieldwork. For a small consumer survey of two thousand responses, this is an afternoon of work. For a multi-country wave across MENA or Sub-Saharan Africa at twenty thousand responses, the same process runs to days, requires multiple reviewers, and still only covers a fraction of the actual response file.

Manual review of a forty-thousand-response wave in a time-pressured delivery cycle is effectively a sampling exercise. You review what you can review. You develop heuristics for which segments are most likely to have quality problems and concentrate your attention there. You hope the heuristics are right, and usually they are, but occasionally they are not, and you do not find out until after delivery.

That is the specific failure mode we kept encountering. Not systematic fraud that would have been caught by any reasonable automated check. The slower, harder-to-detect pattern: straightlining respondents scattered across a wave in proportion that was below the threshold your toplines would surface, but above the threshold a client's analyst would notice when they went into the crosstabs. The wave looked fine at the summary level. It did not look fine when a careful reader spent time with it.

The Problem With Existing Automated Tools

Automated quality checking is not new. Every major survey platform has some form of it. The issue we kept running into was that the tools built by platform vendors are calibrated for the markets those platforms know best, which typically means North American and European consumer panels. When you apply those tools to fieldwork in MENA or Sub-Saharan Africa, they produce one of two outcomes: either they flag so aggressively that you lose too much sample, or they miss the quality problems that are specific to those markets because the detection logic was not designed for them.

The straightlining problem is a good example. A straightliner in a US consumer panel who marks every item 3 out of 5 on a twenty-item attitude battery is reasonably likely to be caught by a pattern-detection tool. A straightliner in a Jordanian panel who marks every item 4 out of 5, in a market where acquiescence bias is real and culturally grounded, is much harder to distinguish from a genuine high-agreement respondent using standard pattern detection. You need to know something about the expected distribution of responses in that specific cultural context to make that call confidently.

This is the gap we decided to close. Not a new category of tool, but the same category of tool built with the markets that are actually growing in the research industry at the center of its design, rather than at the periphery.

The Decision to Build

The decision to build Besample rather than to continue working around the gap was based on a simple observation: the problem was getting worse rather than better. Fieldwork volumes in MENA, Sub-Saharan Africa, APAC, and Latin America have been growing for several years, and the panel infrastructure in those markets is increasingly sophisticated. But the quality tooling did not keep pace with the volume growth. Agencies were handling larger waves with more complex cross-country designs using quality infrastructure designed for smaller, less complex projects.

We also noticed that the problem was not primarily technical. The detection signals that would catch most emerging-market quality problems already exist: response timing distributions, answer pattern analysis, geolocation validation, attention-check performance. What was missing was not new signal types but the calibration work to make those signals meaningful in specific market contexts, and the infrastructure to run that calibration systematically rather than project-by-project.

We built Besample to be a quality auditing layer that sits on top of whatever survey platform or panel provider a research agency is already using. It runs the standard detection signals, applies market-appropriate calibration, returns reason codes rather than binary pass/fail flags, and gives operations teams the context they need to make defensible decisions about borderline cases. The goal is not to replace the judgment of an experienced research manager. It is to give that manager something to work with at a scale that is actually reachable given the time and labor available.

What We Are Not Claiming

It would be easy to overstate what automated quality checking can accomplish. We do not believe it eliminates the need for human review in research operations. It does not. Borderline cases will always require human judgment, and the combination of automated detection and human escalation review is more reliable than either alone.

We also do not think pattern detection is a complete answer to panel quality problems. The structural challenges of maintaining high-quality panels in emerging markets: respondent incentive structures, literacy variation, translation quality, panel provider accountability, are not problems that a response-level audit tool solves. Those are upstream issues that require upstream interventions.

What Besample addresses is the specific problem that sits between panel recruitment and delivery: the wave has closed, the data is in hand, and the team needs to make defensible quality decisions before the deck goes out. That is the moment where the absence of properly calibrated, scalable automated checking has historically produced the most downstream cost, and it is the moment we built this tool to address.

Where We Are Now

Besample launched in 2025. We are a small team: three people, Brooklyn-based, with backgrounds across survey methodology, machine learning infrastructure, and field operations in emerging markets. We are in early access, working with research agencies that field in the markets we care most about, and using those engagements to refine the market-archetype calibrations that sit at the core of how the platform works.

The problems we describe in this article are real and ongoing. The agencies that are running fieldwork in MENA, Sub-Saharan Africa, APAC, and Latin America at volume are dealing with them now. We built Besample because we thought there was a better solution to them than the current workarounds, and we are spending this early period proving that out.

If you are running fieldwork in these markets and the manual QC problem feels familiar, we would like to talk. The contact is on our site.

Audit your next wave with Besample

Connect your survey platform and get per-response quality scores as fieldwork runs, not after it closes.

Request Access