Category: Product

  • How we run the same customer interview every quarter

    How we run the same customer interview every quarter

    Every quarter, on the second Wednesday of the last month, we sit down with three Velo customers and ask them the same eight questions we asked last quarter. Same script. Same order. Roughly the same time of day. The value is not in any one conversation. The value shows up when we diff the transcripts.

    The script we do not change

    The temptation to tweak questions is strong, and we resist it. Once you edit a question, you cannot compare answers across quarters. Our fixed eight, in order:

    • Walk us through the last thing you did in Velo this morning.
    • What did you almost use Velo for, and then use something else?
    • Which tab do you open first when you log in?
    • Who else at your company touches Velo, and how?
    • What have you built around Velo that we did not ship?
    • When was the last time Velo surprised you, good or bad?
    • If we shut down tomorrow, what would you replace us with?
    • What are you doing on the days you do not open Velo?

    Three users, chosen from a rotating pool of twelve power accounts. One customer we keep constant across the year for a longitudinal read. Two rotate.

    The diff is the deliverable

    We record with Grain, get transcripts back within the hour, and drop them into a Notion database with one row per user per quarter. On Friday we run a working session with product, design, and one engineer. The rule for that session: no talking about individual quotes until we have compared the same question across quarters.

    What we look for in the diff:

    1. Words that appear this quarter and did not appear last quarter. Q1 nobody said “audit.” In Q2, all three did.
    2. The “almost used Velo for” answer shifting toward one workflow. Two quarters in a row of the same near-miss becomes a Linear ticket in the next planning cycle.
    3. Tab-first answers drifting away from the tab we consider the home screen. That happened to Reports in Q3 last year and pushed the nav redesign.
    4. Homegrown scripts built around us. Every quarter someone shows us a Google Sheet stitched to our API. That sheet is the roadmap.

    The interview does not tell us what to build. The diff between interviews does.

    What we do with it

    By the following Monday, the working session produces a one page memo in Notion with three sections: patterns that hardened, patterns that softened, and one thing we were wrong about. That memo goes to the whole company in Slack, and the three tickets it generates land in the top of the next sprint. That is the whole loop. Eight questions, three users, one diff, one memo, three tickets.

  • Kill the quarterly roadmap, keep the direction

    Kill the quarterly roadmap, keep the direction

    Two years ago we kept a proper roadmap. Quarterly planning weeks, OKR drafts in Notion, a review meeting every other Friday, and a Linear project called Q3 Commitments that we updated on Tuesdays and lied about on Fridays. We shipped things. We also spent about six engineering weeks per quarter arguing over rows in a spreadsheet.

    We stopped doing that in March. What replaced it: a single Notion page called Next six months, in rough order. It fits on one screen. It has no dates past the current month, no confidence percentages, and no owner column.

    What the page looks like

    Three sections, always:

    1. Now. Two or three things we are building this week and next. Each links to a Linear ticket.
    2. Soon. Four to eight things we intend to start inside the next quarter. Ordered top to bottom by priority. No dates.
    3. Later, probably. Bets we think matter for the six month horizon but have not committed to. Read the room, not the calendar.

    The page has a Last edited line at the top. If it has been more than three weeks, someone owes the team an update on Slack.

    What we killed with it

    • The quarterly planning offsite. We used to lose a Thursday and half a Friday to it.
    • The OKR grading meeting. Nobody misses it.
    • The roadmap in Figma that lived in parallel to the Linear board and disagreed with it by week three.
    • The reforecast conversation, where we pretended a plan drafted in January still described reality in April.

    We kept the direction. The company still has a one paragraph statement of what we are building over the next 18 months. It sits at the top of the same page. Every commit against Now and Soon should be defensible against that paragraph.

    Why the world changing is fine

    When a competitor ships something surprising, or a big customer hits a wall, or Datadog tells us the ingestion pipeline is cracking, we open the page and reorder it. That takes an hour, not a planning cycle. The team sees the change on Slack the same day. Nobody asks whether we are missing a Q3 objective, because there is no Q3 objective to miss.

    The trade we made: we gave up the theatre of certainty. We got back the six weeks a year we used to spend rehearsing it.

  • The PRD template we abandoned

    The PRD template we abandoned

    For eighteen months, every feature at Velo started the same way. Someone opened a Notion page called “PRD Template v4,” scrolled past six pages of headings, and started typing. Problem statement. Goals. Non-goals. Success metrics. User stories. Edge cases. Rollout plan. Open questions. Appendix.

    By the time the doc reached our Thursday product review, it was 2,400 words of dutiful prose. When we asked what the user would do differently after we shipped, the room went quiet.

    What the template was hiding

    We reread twelve PRDs from Q1 and Q2. A pattern showed up in the verbs.

    • “Improve onboarding” appeared in four docs.
    • “Optimise the flow” showed up in five.
    • “Reduce friction” was in nine of the twelve.
    • “Increase engagement” closed out most success sections.

    None of these verbs pointed at anything a human could do or fail to do. They were placeholders for thinking, and the template rewarded us for placing them. Six pages of scaffolding meant six pages to fill, and the easiest way to fill six pages was to reach for words that meant almost anything.

    The reviews got worse in a specific way. Engineers would ask a sharp question on page one, and by page five we had lost the thread. Design would flag an interaction we had not considered, and someone would promise to “cover it in the appendix.” The appendix became a graveyard. In one memorable case, a Linear ticket sat in Backlog for eleven weeks because the PRD had three conflicting definitions of “activated user” and nobody wanted to be the person to pick one.

    We had built a document that made disagreement expensive to surface.

    The one-page brief

    In August we tried something smaller. Three prompts, one page, hard cap at 400 words.

    1. Problem. What is going wrong for a specific person on a specific day? Name the person by role, not persona.
    2. Target user. Who feels this problem most acutely, and which of them do we have access to for interviews within the next week?
    3. First observable behaviour. What will this person do after we ship that they cannot do or would not do today? Describe the action, not the outcome.

    That third prompt did most of the work. “Increase engagement” fails it. “The ops lead exports the reconciliation report as a CSV before their Monday standup” passes it. You can watch a person do that or not do it. You can ask them why they did not. You can count it in Datadog.

    We wrote the first brief for a billing export feature. It took forty minutes. The Thursday review took nineteen. Two engineers pushed back on the target user selection, we agreed to run three interviews before writing a line of code, and the ticket moved to In Progress the following Tuesday.

    What we gained

    The obvious thing was time. Our median PRD used to take a product manager somewhere between four and seven hours to draft, plus another two in edits after review. The brief takes forty to ninety minutes. Reviews compressed from an hour to under thirty because there was nowhere to hide vague thinking.

    The less obvious thing was that engineers started reading the doc. When something is six pages of prose, the second engineer on a project skims. When it is one page and the third bullet is a specific behaviour, they read it, and they argue. We stopped confusing “we have a PRD” with “we know what we are building.” Writing a brief did not feel like progress in the way writing a six-page PRD had felt like progress. That turned out to be honest.

    What we lost

    We lost some things we miss. The old PRD forced us to enumerate edge cases up front, and the brief does not. Two of our first four briefs shipped with a rollback plan improvised in a Slack thread the morning of launch. We have since added a lightweight “what breaks” checklist that lives next to the brief but is not part of it, and reviews now include a five-minute walkthrough of that checklist.

    We also lost some institutional memory. A six-page PRD, for all its flaws, was a decent artefact for the person joining the team six months later. Our briefs are terse enough that a new engineer opening one in GitHub cannot always reconstruct why we made the calls we did. We are experimenting with a short decisions log per feature, kept in the same Notion database, updated as we go. It is not a Figma spec and it is not a design review transcript, but it stops the “why did we do it this way” question from getting answered with a shrug.

    What we would tell you

    If your PRD template has more than three prompts, ask the last five people who filled one in whether the extra prompts changed a single decision. If the answer is no, cut them. The template is not neutral. It teaches your team what counts as thinking.

    The verbs are the tell. If your docs are full of them and nobody has flinched in a review this quarter, the template is doing the flinching for you.

  • When to open the Board, the WBS, or the Timeline

    When to open the Board, the WBS, or the Timeline

    Every project in Velo shows up in three shapes. The Board is where a card moves from Doing to Done. The WBS (work breakdown structure) is where the project gets decomposed into deliverables and sub-deliverables before anyone touches a ticket. The Timeline is what we show the CFO on Thursday when she asks whether the migration will land before Q3 close.

    We built all three because a project has three audiences: the person doing the work, the person shaping the work, and the person funding the work. They need different resolutions of the same truth.

    The Board is for the next 72 hours

    Open the Board when you want to know what is moving today. It answers: which cards are in progress, which are blocked, who owns them, and what will ship by Friday. Columns default to Backlog, Ready, Doing, Review, Done. We use the same keyboard shortcuts as Linear so nobody on the team has to retrain their fingers when they switch tools.

    Rules of thumb for when the Board is the right lens:

    • Daily standup at 9:15
    • Triaging a Slack ping from support about a production bug
    • Deciding whether to pull in a stretch card during the sprint
    • Checking WIP limits before merging another design review

    The Board hides context on purpose. You will not see when the epic started, how it maps to Q3 goals, or whether the effort estimate has drifted from the original budget. That absence is the feature. During execution, more context slows you down.

    The WBS is for the first two weeks and the ugly middle

    The WBS is where a project is born. Before we open a Board we sit with the product lead and break the initiative into deliverables, then those into sub-deliverables, then those into work packages. A migration project we ran last quarter had 4 top-level deliverables, 19 sub-deliverables, and 63 work packages by the time we stopped decomposing. Each work package eventually became one or two cards on the Board.

    The biggest mistake we see, in our own team and in every team we onboard, is this: people build a WBS in week one, move everything to the Board, and then never open the WBS again for the rest of the project.

    That is where projects die quietly. The Board tells you what is in flight. It does not tell you what you promised, what you dropped, what got scoped out, or what nobody has picked up because it sat in a sub-deliverable that never got carded. Six weeks into a project, the Board looks healthy and the actual deliverable is missing a third of its scope.

    We now run a WBS review every second Wednesday. Fifteen minutes, one question per branch of the tree: is this deliverable still on the plan, or did it fall off the Board without a decision? About one time in five we find a work package that should be a card and is not. About one time in ten we find a card doing work that is not in the WBS at all, which is a scope conversation waiting to happen.

    The Timeline is for people who do not touch the work

    The Timeline is a rolled-up Gantt with milestones, phase bands, and dependencies. It is what we send to the exec sponsor, the finance partner, and the customer stakeholder on a Tuesday-morning update email. We do not use it to plan work; we use it to communicate about work.

    A few things the Timeline does well:

    1. Shows the critical path when a dependency slips.
    2. Shows milestones against calendar dates, not sprint numbers.
    3. Shows phase overlap, so a stakeholder can see that Discovery and Design were meant to overlap by two weeks.
    4. Exports cleanly into the steering committee deck.

    If your PM lives in the Timeline, something is off. The Timeline is a read-only surface for people outside the working team. When we catch ourselves editing the Timeline directly, it is a signal that the WBS underneath is stale.

    A working rhythm

    Here is the cadence we default to on a new project, and the one we recommend when a team onboards to Velo:

    • Week 1: WBS only. No Board. No Timeline. Decompose until every leaf is estimable in a day or two.
    • Week 2: Cards get generated from work packages. Board opens. First sprint starts.
    • Every sprint: Board is the daily surface. WBS gets a 15-minute review mid-sprint. Timeline gets refreshed once, on the day before the stakeholder update.
    • Project close: We close the loop in the WBS, not the Board. A card marked Done is a promise kept only if the parent work package was in the original plan.

    Three views, one project, three different questions. The Board asks what is moving this week. The WBS asks what we promised at the start. The Timeline asks what the outside world should expect. Skip any of them and the project drifts in a way that is hard to see until it is expensive to fix.

  • What done means for a task on our team

    What done means for a task on our team

    Every team we worked on before Velo had a definition of done pinned to a wiki page nobody read. Ours did too, until a Wednesday standup in March when Priya asked whether the invoice retry work was finished, and four engineers gave four different answers. That morning cost us a customer refund and a two hour incident review. We decided the Notion page was not the problem. The definition was.

    Three tries that did not stick

    Our first attempt was a paragraph in Notion titled “shipping standards” that said tasks should be “merged, tested, and reviewed.” It read fine on the page. In practice, “tested” meant whatever the author felt like: a unit test, a manual walkthrough, or nothing if the diff was under twenty lines. We shipped a race condition in the billing worker three weeks later because the author had run the change against a fresh database and assumed that counted.

    The second attempt was a Linear checklist template with nine items. Everyone checked every box, because the boxes were reported by the author and the reviewer had no way to verify half of them without opening five other tabs. The checklist became a ritual, then a joke, then a template we quietly stopped applying to new tickets.

    The third attempt was strict: a task was done when a designated QA engineer signed off in a Slack thread. This lasted eleven days. Our QA lead, Ruth, went on holiday, and the queue backed up to forty two tickets. When she came back, half the context was gone and she had to re verify work from memory. We had traded ambiguity for a bottleneck.

    The four criteria we settled on

    After the third failure, we spent a Friday afternoon working through what we needed the definition to do. It had to be verifiable by someone other than the author, it had to survive one person being out, and it had to answer the question Priya asked in March without a debate. We landed on four criteria, in this order:

    • Works: the change does what the ticket says, verified against the acceptance criteria written before the branch was cut. If those criteria were vague, that gets fixed before the ticket moves to review, not after.
    • Tested: automated coverage exists for the new behavior, and the tests fail without the change. The reviewer runs the suite locally or points at a green CI badge tied to the merge commit.
    • Deployed: the change is live in production, not staging, not behind a flag that has never been flipped on for a real user. If the work sits behind a flag, done waits until the flag is on for the intended audience.
    • Observed: a human has confirmed the change behaves as expected in production, using logs, a Datadog dashboard, or a real user event. Not a synthetic ping. A trace of the feature being used, or a metric moving in the direction we predicted.

    The order matters. If “works” is unclear, testing the wrong thing is worse than not testing. If we skip “deployed” and call something done at merge, we hide half our incidents in the gap between main and production.

    The compromise on observed

    Observed was the criterion that almost killed the whole definition. Half the team pointed out, correctly, that internal only changes have no production traffic to watch. A new admin report, a migration script, an internal CLI: none of these throw off metrics on the customer dashboards we use for observability. Waiting for a real user event on an internal tool would mean waiting forever, or fabricating one.

    We debated dropping the criterion for internal work. We tried, for a sprint. Two internal tools broke silently and we found out from a support agent who could not load the refunds page. The criterion needed to survive.

    The compromise: for internal only changes, observed means the author or a teammate has used the feature in production for its intended purpose, with a Loom or a screenshot posted to the ticket. Not tested it. Used it. If the ticket is a migration, the observation is the query result after the migration ran. If it is a CLI, it is the terminal output from a real invocation on the real database.

    The distinction we care about is between “I believe this works” and “this has done its job for a real person, once.” The Loom feels heavy the first time. It stops feeling heavy the second time somebody catches a broken admin page before a customer does.

    How the four criteria show up in our week

    Every ticket in Linear now has four checkboxes matching the criteria. The author checks the first three. The reviewer, or on internal changes any teammate, checks observed and pastes the evidence. Our Monday planning meeting starts by pulling the list of tickets marked done in the last week and skimming the observation links. It takes eight minutes. In the six months since we adopted this, we have had two rollback situations that a proper observation caught before the on call engineer noticed. We have also had one case where the observation link was a screenshot of the wrong environment, which is a different problem, and one we are still working on.

    We do not think this definition is universal. It is what our team of eleven engineers, on a codebase with sixteen deploys a week, needs to keep the wiki page honest. If the shape of the team changes, we expect the definition to change with it.