Work moves at the speed of a conversation
What I learned building Retinue, AI that helps you in the meeting, not just after.
Why can't your meeting be your output? Why can't things happen at the speed of conversation?
That's the question I've been working on all summer with my friend Will. The answer is Retinue: AI that helps you in the meeting, not just after. Along the way, I got our running cost down from $1–2 per user hour to $0.10–0.20. More on that below.
Stuck in meeting after meeting
I used to spend most of my week in meetings. That meant the actual work had to happen before or after them. And as more meetings got added, managing my time got harder and harder.
The worst was back-to-back meetings. Preparing in advance, understanding the context, pulling the outputs from one call and setting up the next just becomes too much.
What I needed was something that would surface my context and help me understand what was going on, even if I hadn't had the chance to read up. Something that would challenge me and give me the right questions to really dig in. And something to help me actually complete the work, because every meeting created more tasks, and it became an endless merry-go-round.
Really, I needed something that could drastically shorten the time from meeting to output.
How Retinue came about
Retinue started from a really cool piece of work Will had built. You could bring sub-agents with different personas into your call to help you. Scholar, for example, would look up research articles while you talked. That was the first version.
When we started working on this, we ran user testing with 25 people, and the feedback was really clear. People didn't want to choose between lots of different options. They just wanted the output. They wanted it to be clear, and they didn't want it to be distracting.
So that's where I started iterating. The question became: how can this give people extremely useful content, in an extremely brief and simple way, so they can get on with what they need to do?
What I learned building it
Along the way there have been so many learnings. Who knew the most complicated piece would be getting social logins to work?
The real challenge, though, was orchestration. Surfacing research live in a call takes multiple agents working together:
- A speech-to-text model transcribing the conversation in real time.
- A model deciding whether anything said is worth kicking off an internal or external search.
- A model judging whether anything that search found is important enough to show the user.
- A step that reshapes the result so it's coherent and clear on screen.
- A feedback loop that checks whether the output was actually useful, or used in the call, and fine-tunes future outputs.
And that's just for research. We built similar pipelines for questions and for actions, with a commitment checker that tracks what people agree to do in the meeting.
With that many agents in the loop, evals became essential. For me, evals sit on a triangle: latency, cost and quality. Latency obviously matters, because help that arrives after the moment has passed is useless. Quality obviously matters. And cost needs to be low. The trouble with a triangle is that you usually can't have all three.
Coming up next
In my next post, I'll dig into evals properly: how we measured all three corners of that triangle, and how I brought our cost down from $1–2 per user hour to $0.10–0.20 without giving up speed or quality.