Scott Walker

Case study

The 3.0 preference test

4 minute read

Client
Eurocamp
Year
2026
Role
Product Designer
Contribution
Designed the 3.0 templates, then designed and ran the study that tested them
Team
NL, DE and PL colleagues reviewed the translations; I coordinated that review
Platform
Web · responsive · 5 markets, 4 languages

Eurocamp's 3.0 redesign won the overall vote in every market tested — 58% to 73%, across 365 completed responses in five markets and four languages. It shipped as a six-phase gated rollout rather than a single release, because the study separated preference from comprehension and found that on one component the two pointed in opposite directions.

SHOT 01 · HERO The five paired comparisons as fielded — 2.5 left, 3.0 right, across all five templates. Job: recognition. A reader sees what the study actually was in one glance. frames/UK/COMBO.png → crop to a five-pair strip · CHECK FOR AND REMOVE THE ASK CAMILLE BUTTON
Tested May 2026 The five paired comparisons as fielded: 2.5 on the left, 3.0 on the right, across Homepage, Region, Parc, Product Cards and Content Cards. Every participant saw all five.

The brief was “does 3.0 win?” The answer to that question would have been useless on its own.

A five-market redesign does not ship on a yes. It ships market by market, template by template, and every market team needed a defensible reason to go or wait. A single global preference number would have told leadership to ship everything and told the German team nothing.

So the question I designed for was where 3.0 doesn't win, and what that costs — which meant the instrument had to survive being cut five ways, and be clean enough that a market-level reversal read as a finding rather than a study artefact.

Three constraints shaped every decision that followed, and a fourth that I couldn't design around at all:

  • Five templates in one sitting. Homepage, Region, Parc, Product Cards, Content Cards — all needed comparable data from the same person, inside a study people would finish.
  • Four languages, unmoderated. No moderator to recover a badly-worded question. Every scale anchor had to survive translation into Dutch, German and Polish.
  • A loyal sample. Recruitment ran through Eurocamp's customer database. UK came back 92% returning customers. This audience was comparing against something familiar, which is not the same as evaluating something new.
  • I designed 3.0. The person who made the templates wrote the questions about them. That is the strongest bias in the study and it isn't fixable by being careful — only by instrumenting against it.
SHOT 02 · ANNOTATED The study map — 20 blocks, 10 logic jumps, 18 questions per participant. Job: show the instrument is engineered, not assembled. Pins mark the three branch points. design/eurocamp-study-journey-map-v3.html → screenshot, then place pins
Study map v3 — the version fielded 20 blocks, 10 logic jumps, 18 questions per participant.
  1. 1Screener splits customer from prospect — the variable every later cut depends on.
  2. 2Each preference test forks on the participant's own choice. Ten of these jumps carry the whole study.
  3. 3The overall-direction question routes to one open-text path. “No preference” skips it, by design.

Both designs stay on screen, because the alternative was a hundred-question study

I wrote a sequential monadic homepage test first — Design A alone, then Design B alone, then a comparison answered from memory with neither on screen. It is the better instrument for absolute measurement — a clean per-design trust score, uncontaminated by the other option sitting next to it.

It also runs to 20 questions for one template. Five templates that way is a study nobody finishes, and the from-memory comparison introduces recall bias — the thing I was trying to avoid.

Paired side-by-side won on the trade I could defend: absolute per-design scores given up, five templates kept, selection time gained as a secondary signal. What I lost, I lost knowingly.

Everyone answers only about the design they picked

Follow-up questions branch on the participant's own choice. Pick 2.5 Parc, you answer about 2.5 Parc.

This halves the question load, removes counterfactual bias, and resolved a bias finding I'd raised against myself: the Content Cards element list had three options for 2.5 and six for 3.0, which looks like a rigged comparison until you notice no participant ever sees both lists.

The cost is real and I published it: deep-dive percentages read as “among voters who chose this variant”, never as overall measures. Every chart in the playback carries that qualifier.

Ten open-text questions was the study's biggest risk, so I deleted nine

The first build asked 10 follow-ups and every single one was a free-text box. Stated completion time 5–8 minutes; realistic time closer to 12–15. A blank box twice per section produces one-word answers, skips and abandonment — the three failures that would have made market-level analysis impossible.

I replaced them with multi-select where preference is genuinely multi-factor, five-point scales with descriptive anchors rather than numeric ones (they travel better across four languages), a 0–5 clarity scale where the concept is directional rather than a quality, and one open-text question, placed last, after the participant has a settled view.

27 questions asked became 18 per participant. Completion landed at 75%.

The instrument, before and after

First build

March 2026

27

questions asked · 12–15 min realistic

  • 10 × free-text follow-up
  • 0 × structured
  • no screener
  • no closing question

As fielded

May 2026

18

questions per participant · 75% completed

  • 1 × free-text, placed last
  • 5 × multi-select
  • 5 × Likert, descriptive anchors
  • 1 × 0–5 clarity scale
  • 2 × screener · 1 × overall direction
First build, March 2026 → as fielded, May 2026 Ten free-text boxes became multi-select, Likert, a 0–5 clarity scale, and one open question at the end.

I designed the thing I was testing, so I audited my own study against myself

Intending to be fair doesn't make an instrument neutral. So the study went through a formal bias audit before fielding, findings graded and tracked like defects.

Eleven findings, nine closed before launch. “Current” and “New” became neutral labels — “New” carries positive valence. All five rating scales were agreement statements (“This page looks trustworthy”), which invites acquiescence; all five became neutral questions with descriptive endpoints. The homepage task said “most confident”, anchoring evaluation to one dimension, and became “most like to use”. The impression-word list ran five positive to three negative: “Professional” out, “Overwhelming” in, four and four.

The one worth showing is the finding I fixed twice. The closing open-text asked only what a design gets right. My first fix added a balancing question: what would you most want to change? My second fix deleted both and replaced them with one neutral question — is there anything you'd change or improve?

Balancing a leading question with a second leading question doubles the fatigue and still frames the answer. Removing the frame was the better move, and it took writing the wrong fix first to see it.

SHOT 04 · THE REJECTED OPTION Bias audit v1 next to v2 on finding 8. Job: prove the exploration happened — a fix that was written, shipped into a document, and then reversed. archive/…Bias_Audit.docx §8 vs design/…Bias_Audit_V2.docx §8
Bias audit v1 → v2, March 2026 Two attempts at the same finding: balance the leading question, then delete it.

Reporting preference and clarity separately is what caught Content Cards

The pressure in a preference study is to produce one number per template. I reported two, and on Content Cards they diverged:

Content Cards · clarity top-box, 2.5 → 3.0

0% 20% 40% 60% PL 22% DE 16% NL 33%

2.53.0

Tested May 2026 Clarity falls in every market that preferred the new cards. The full numbers, including motivation, are in the table below.
Content Cards — preference against comprehension
Market3.0 preferenceClarity top-boxMotivation top-box
PL76%35% → 22%43% → 28%
DE72%30% → 16%40% → 16%
NL58%50% → 33%67% → 50%

People chose the 3.0 cards because they expose more content, then understood the offer less well once they had to use them. A single blended score would have shipped that component as a clean win.

SHOT 05 · THE DECISION SHOT Content Cards 2.5 and 3.0, with preference and clarity annotated on each. Job: the single most important image on the page — it carries the finding the whole case study turns on. frames/UK/2.5 - CONTENT CARDS.png + 3.0 - CONTENT CARDS.png
Tested May 2026 · gated The component that won the vote and lost the comprehension test.
  1. 12.5: image-led, three data points per card. Fewer people preferred it. More people understood it.
  2. 23.0: structured card, six data points. This density is what wins the preference vote.
  3. 3…and what costs 13–17 points of clarity once a reader has to act on the offer rather than look at it.

Prototype slot. 12 seconds of the study as participants met it — screener, one paired comparison, the branch firing. Task-shaped, not a feature tour. Ships as <video controls muted playsinline preload="metadata" poster="…"> with no autoplay, so WCAG 2.2.2 is satisfied by construction, plus a static poster as the fallback.

To build One clip, one idea, under 3 seconds to start on mobile.

The result: a rollout sequence, not a launch

3.0 won the overall direction vote in all five markets. Product Cards won 74–88% everywhere and shipped first, in week one, with no gate.

Two results moved money the other way. German customers preferred the old homepage 56/44 — the only homepage rejection anywhere — and German “Modern” rose 34 points while “Trustworthy” fell 14. Content Cards went behind a click-through comprehension study. Both got funded follow-up work instead of a launch date.

One finding survived the cut cleanly: on the UK Region page, 3.0 lost the preference vote (44%) while booking confidence went up (4.04 against 3.91). Preference flipped; confidence didn't. That kept four markets shipping Region on schedule instead of holding all five behind a UK result.

Cross-market preference · 30 cells

3.0 preference share by template and market. Cells below 50% favour the old design.
TemplateUKIEDENLPL
Homepage 73% 67% 44% 70% 51%
Parc page 77% 69% 68% 71% 67%
Region page 44% 56% 53% 62% 61%
Product cards 74% 88% 84% 72% 77%
Content cards 62% 76% 72% 58% 76%
Overall direction 62% 73% 58% 70% 59%

70%+60–6953–59tie2.5 wins

Fieldwork May–June 2026 · n=365 completers Thirty cells. Two of them favour the old design, and they fail differently: the German homepage is a blocker, the UK region page is a preference flip with booking confidence moving the other way.

Ten-week rollout · six phases

W1 W3 W5 W7 W9 W10+ 1 Product cards All markets 2 Parc page + Included / Add-ons panel All markets 3 Content cards◆ GATED After click-through validation · NL last 4 Region page IE, DE, NL, PL ship · UK follow-up in parallel 5 Homepage◆ GATED UK, IE, NL ship · DE and PL gated 6 Gated markets◆ GATED UK Region · DE Homepage · PL Homepage
Phase 1 shipped week 1 Six phases over ten weeks. Two of them gated behind follow-up research that wouldn't have been commissioned from a single blended score.

Reflection

The two results I'd most want to explain are the two I instrumented least. German and Polish homepage were the weakest scores in the study, and translation register is the most likely silent driver in both — I flagged it as a confounding variable in the playback but never audited it, so I can't separate a reaction to the design from a reaction to the copy.

I'd also fix the sample. Prospect recruitment through Prolific returned n=3 per market and was unusable, so a study that reads as a mandate is a mandate from existing customers. And section-order randomisation was still open at launch — the later templates carry whatever fatigue the earlier ones caused, and I don't know how much.

SHOT 08 · PROCESS ARTEFACT The open-items list at launch. Job: evidence of a fork, not evidence that work happened — two findings shipped without closing, and the reason. Bias_Audit_V2.docx → “Open Items Before Launch”
At launch, April 2026 What I shipped without closing, and why.