UX Researcher
Turn raw session transcripts into findings a team will act on — and never invent a quote.
-
1 Tag by theme Claude Haiku 4.5 720ms
Open-coded 12 transcripts into 34 codes. Four clusters carry most of the weight; the rest are singletons.
-
2 Synthesise findings Claude Opus 4.8 1520ms
Nine of twelve stalled at the workspace-invite step. They all said a version of "I don’t know who to add yet" — that is one finding, not nine.
-
3 Audit the evidence GPT-5 980ms
Checked every claim against a timestamped quote. One finding — "users want templates" — rests on a single participant. Demoting it to a signal.
Each step ran on whichever model is best at that job — not one model for everything. The worker is the workflow.
Permissions & guardrails- Read: session transcripts, study protocols and prior findings — with participant PII already redacted at ingest
- Write: findings documents and the evidence log only; it never edits or deletes a source transcript
- Escalate: the researcher when a claim rests on a single participant, or when two sessions directly contradict each other
- Never: contact a participant, or surface anything that could re-identify one
- Never: paraphrase a quote it uses as evidence — every claim links to a timestamped verbatim
- 4 findings, each traced to at least 3 timestamped quotes.
- 1 demoted to a signal — single participant, flagged not dropped.
- Zero paraphrased quotes. Every line of evidence is verbatim.
84.9%
a general model gets 68.2% · +16.7pp
68 · 72.9 · 77.4 · 80.8 · 83.1 · 84.9
private rules were taught in one workspace and never leave it. general rules are true about the work, so every copy of this worker gets them.