Case study
The 3.0 preference test
4 minute read
- Client
- Eurocamp
- Year
- 2026
- Role
- Product Designer
- Contribution
- Designed the 3.0 templates, then designed and ran the study that tested them
- Team
- NL, DE and PL colleagues reviewed the translations; I coordinated that review
- Platform
- Web · responsive · 5 markets, 4 languages
Eurocamp's 3.0 redesign won the overall vote in every market tested — 58% to 73%, across 365 completed responses in five markets and four languages. It shipped as a six-phase gated rollout rather than a single release, because the study separated preference from comprehension and found that on one component the two pointed in opposite directions.
The brief was “does 3.0 win?” The answer to that question would have been useless on its own.
A five-market redesign does not ship on a yes. It ships market by market, template by template, and every market team needed a defensible reason to go or wait. A single global preference number would have told leadership to ship everything and told the German team nothing.
So the question I designed for was where 3.0 doesn't win, and what that costs — which meant the instrument had to survive being cut five ways, and be clean enough that a market-level reversal read as a finding rather than a study artefact.
Three constraints shaped every decision that followed, and a fourth that I couldn't design around at all:
- Five templates in one sitting. Homepage, Region, Parc, Product Cards, Content Cards — all needed comparable data from the same person, inside a study people would finish.
- Four languages, unmoderated. No moderator to recover a badly-worded question. Every scale anchor had to survive translation into Dutch, German and Polish.
- A loyal sample. Recruitment ran through Eurocamp's customer database. UK came back 92% returning customers. This audience was comparing against something familiar, which is not the same as evaluating something new.
- I designed 3.0. The person who made the templates wrote the questions about them. That is the strongest bias in the study and it isn't fixable by being careful — only by instrumenting against it.
- 1Screener splits customer from prospect — the variable every later cut depends on.
- 2Each preference test forks on the participant's own choice. Ten of these jumps carry the whole study.
- 3The overall-direction question routes to one open-text path. “No preference” skips it, by design.
Both designs stay on screen, because the alternative was a hundred-question study
I wrote a sequential monadic homepage test first — Design A alone, then Design B alone, then a comparison answered from memory with neither on screen. It is the better instrument for absolute measurement — a clean per-design trust score, uncontaminated by the other option sitting next to it.
It also runs to 20 questions for one template. Five templates that way is a study nobody finishes, and the from-memory comparison introduces recall bias — the thing I was trying to avoid.
Paired side-by-side won on the trade I could defend: absolute per-design scores given up, five templates kept, selection time gained as a secondary signal. What I lost, I lost knowingly.
Everyone answers only about the design they picked
Follow-up questions branch on the participant's own choice. Pick 2.5 Parc, you answer about 2.5 Parc.
This halves the question load, removes counterfactual bias, and resolved a bias finding I'd raised against myself: the Content Cards element list had three options for 2.5 and six for 3.0, which looks like a rigged comparison until you notice no participant ever sees both lists.
The cost is real and I published it: deep-dive percentages read as “among voters who chose this variant”, never as overall measures. Every chart in the playback carries that qualifier.
Ten open-text questions was the study's biggest risk, so I deleted nine
The first build asked 10 follow-ups and every single one was a free-text box. Stated completion time 5–8 minutes; realistic time closer to 12–15. A blank box twice per section produces one-word answers, skips and abandonment — the three failures that would have made market-level analysis impossible.
I replaced them with multi-select where preference is genuinely multi-factor, five-point scales with descriptive anchors rather than numeric ones (they travel better across four languages), a 0–5 clarity scale where the concept is directional rather than a quality, and one open-text question, placed last, after the participant has a settled view.
27 questions asked became 18 per participant. Completion landed at 75%.
The instrument, before and after
First build
March 2026
27
questions asked · 12–15 min realistic
- 10 × free-text follow-up
- 0 × structured
- no screener
- no closing question
As fielded
May 2026
18
questions per participant · 75% completed
- 1 × free-text, placed last
- 5 × multi-select
- 5 × Likert, descriptive anchors
- 1 × 0–5 clarity scale
- 2 × screener · 1 × overall direction
I designed the thing I was testing, so I audited my own study against myself
Intending to be fair doesn't make an instrument neutral. So the study went through a formal bias audit before fielding, findings graded and tracked like defects.
Eleven findings, nine closed before launch. “Current” and “New” became neutral labels — “New” carries positive valence. All five rating scales were agreement statements (“This page looks trustworthy”), which invites acquiescence; all five became neutral questions with descriptive endpoints. The homepage task said “most confident”, anchoring evaluation to one dimension, and became “most like to use”. The impression-word list ran five positive to three negative: “Professional” out, “Overwhelming” in, four and four.
The one worth showing is the finding I fixed twice. The closing open-text asked only what a design gets right. My first fix added a balancing question: what would you most want to change? My second fix deleted both and replaced them with one neutral question — is there anything you'd change or improve?
Balancing a leading question with a second leading question doubles the fatigue and still frames the answer. Removing the frame was the better move, and it took writing the wrong fix first to see it.
Reporting preference and clarity separately is what caught Content Cards
The pressure in a preference study is to produce one number per template. I reported two, and on Content Cards they diverged:
Content Cards · clarity top-box, 2.5 → 3.0
2.53.0
| Market | 3.0 preference | Clarity top-box | Motivation top-box |
|---|---|---|---|
| PL | 76% | 35% → 22% | 43% → 28% |
| DE | 72% | 30% → 16% | 40% → 16% |
| NL | 58% | 50% → 33% | 67% → 50% |
People chose the 3.0 cards because they expose more content, then understood the offer less well once they had to use them. A single blended score would have shipped that component as a clean win.
- 12.5: image-led, three data points per card. Fewer people preferred it. More people understood it.
- 23.0: structured card, six data points. This density is what wins the preference vote.
- 3…and what costs 13–17 points of clarity once a reader has to act on the offer rather than look at it.
Prototype slot. 12 seconds of the study as participants met it — screener, one paired comparison, the branch firing.
Task-shaped, not a feature tour. Ships as <video controls muted playsinline preload="metadata" poster="…">
with no autoplay, so WCAG 2.2.2 is satisfied by construction, plus a static poster as the fallback.
The result: a rollout sequence, not a launch
3.0 won the overall direction vote in all five markets. Product Cards won 74–88% everywhere and shipped first, in week one, with no gate.
Two results moved money the other way. German customers preferred the old homepage 56/44 — the only homepage rejection anywhere — and German “Modern” rose 34 points while “Trustworthy” fell 14. Content Cards went behind a click-through comprehension study. Both got funded follow-up work instead of a launch date.
One finding survived the cut cleanly: on the UK Region page, 3.0 lost the preference vote (44%) while booking confidence went up (4.04 against 3.91). Preference flipped; confidence didn't. That kept four markets shipping Region on schedule instead of holding all five behind a UK result.
Cross-market preference · 30 cells
| Template | UK | IE | DE | NL | PL |
|---|---|---|---|---|---|
| Homepage | 73% | 67% | 44% | 70% | 51% |
| Parc page | 77% | 69% | 68% | 71% | 67% |
| Region page | 44% | 56% | 53% | 62% | 61% |
| Product cards | 74% | 88% | 84% | 72% | 77% |
| Content cards | 62% | 76% | 72% | 58% | 76% |
| Overall direction | 62% | 73% | 58% | 70% | 59% |
70%+60–6953–59tie2.5 wins
Ten-week rollout · six phases
Reflection
The two results I'd most want to explain are the two I instrumented least. German and Polish homepage were the weakest scores in the study, and translation register is the most likely silent driver in both — I flagged it as a confounding variable in the playback but never audited it, so I can't separate a reaction to the design from a reaction to the copy.
I'd also fix the sample. Prospect recruitment through Prolific returned n=3 per market and was unusable, so a study that reads as a mandate is a mandate from existing customers. And section-order randomisation was still open at launch — the later templates carry whatever fatigue the earlier ones caused, and I don't know how much.