September 02, 2026
Using Gemini Notebook to clean user research transcripts
How to use Gemini Notebook to clean user research transcripts reliably, without the summarising and rewriting problems of ordinary AI chat.
Top-down categorisation leads to weaker analysis. Take a bottom-up approach instead, using three simple rules for group size, outliers, and labels.
· By Jake McCann
Early in my career I was working with a team to analyse a couple of dozen discovery interviews. We’d decided to affinity map the data, which resulted in hundreds of sticky notes, covering several metres of wall space. To tackle that much data, we analysed it as a whole team: researchers, designers, developers and a product manager. As I was off in one corner of the wall poring over a handful of notes, the rest of the team very quickly sorted the data into about 8 groups. Job done. But this didn’t feel right; we’d spent weeks recruiting participants, running the research, and writing up the notes. Did we really do all that to just learn 8 things?
Those 8 groups were noun buckets: pre-named categories like “pain points”, “needs” and “opportunities”. You’ve probably used similar categories yourself. I have, many times. But I’ve come to find they can hold back the quality of analysis if you rely on them too heavily.
Most techniques for analysing user research data involve some form of grouping individual data points (i.e. quotes and observations). There are two ways to approach grouping: top-down or bottom-up.
What my team had done was quickly determined a set of high-level categories and then used those to group the data. They’d taken a top-down approach.
Top-down is tempting. It’s generally much quicker because you’re not having to interrogate the meaning of each note. You just need to decide if it matches an existing category. This also makes it a lot less mental effort. It’s great for things like usability testing where you might be looking for a relatively narrow set of data points.
But it has some pitfalls, especially when you’re working in more open-ended research contexts. The biggest risk is that by providing categories up front, much of the analytical process is a foregone conclusion. There’s a danger you’re sorting data into a limited set of preconceived buckets, rather than doing a meaningful analysis. There are a few reasons why this can make your analysis less effective:
The basic approach is simple enough:
The specific mechanics of this differ based on your analysis method. If you’re doing affinity mapping, this means reading one sticky note at a time and only moving them when you find another one that has a similar meaning. If you’re doing thematic analysis, this means working through highlighted passages one at a time, adding a label when a passage says something new, and only reusing an existing label when it’s genuinely the same underlying idea.
But by saying it’s simple, I don’t mean that it’s easy - it’s actually quite hard. So here are some tips to help you along.
Even if you go bottom-up, there is a risk that you just balloon up into big, vague buckets anyway. My favourite way to prevent this is to set a limit on how many data points can be in a single group. I’ve been doing this ever since reading it in Contextual Design by Karen Holzblatt and Hugh Beyer [1]:
We limit each first-level group to four notes to force the team to look deeply and make more distinctions than they would otherwise be inclined to. It pushes more of the knowledge up into the group labels.
More on those group labels in a minute.
Four items per group can be a bit too brutal. I often go for six, but feel free to experiment with what works for you and your team. The important thing is that you’re enforcing a certain level of granularity. In my experience, this is the biggest single process improvement you can make if you’re worried about coming up with broad, hard-to-action findings.
If you’re going bottom-up, you’ll eventually find you have notes that just don’t seem to have any conceptual similarity to any others. I often see people shoehorn these into groups where they don’t really fit, which is another source of vague findings.
Not every note needs to be grouped. Jiro Kawakita, the inventor of affinity mapping, called these “lone wolves” [2]. Sometimes they might be interesting observations that stand on their own. Other times they’re noise. In either case, don’t distort a group to fit them in.
Giving each group a name is the point in the process where you’re creating new knowledge about your users. Coming back to Contextual Design:
When well written, the labels tell a story about the user, structuring the problem, identifying specific issues, and organizing everything we know about that issue. The labels represent new information in an affinity.
You should be able to write a succinct, narrative summary that encapsulates every note in the group. If this proves particularly hard, it’s an indication that the group doesn’t represent a single, coherent idea.
You know you’ve done this well when you can trace each high-level finding back to the specific groups, and the individual quotes and observations, that support it.
A clear signal you’re not doing this is that you’re pouring more time and effort into writing up your findings than you did your analysis. The core narrative should already be written up in the group labels. The slide deck or report is just presentational polish. I’ve seen teams do two hours of analysis, then spend three days struggling to structure and restructure their findings and insights in a slide deck. They didn’t realise that they were still trying to analyse their data, just moving slides and text boxes around instead of sticky notes.
Doing your analysis bottom-up is slower. That’s the point. It forces you to make distinctions, hold onto context, and think carefully about what deserves to be a group. This is the stuff that analysis is made of.
Try it out on your next project. Start from the unstructured data. Cap the size of groups. Allow lone wolves. Dedicate time to encapsulating meaning in group labels. Hopefully you’ll find, like I have, that your research outputs are far more insightful than before.
(Does this sound like something you want to do, but there’s no way you can spare the time? Stakeholders need the report tomorrow morning? I’ve been there too. I’m building Stitchwork to speed up this kind of analysis: bottom-up, nuanced findings that trace back to the underlying evidence. You can join the waiting list below.)
September 02, 2026
How to use Gemini Notebook to clean user research transcripts reliably, without the summarising and rewriting problems of ordinary AI chat.
January 16, 2026
A pragmatic guide to using AI assistants in user research without compromising quality and integrity.