Tag: process

  • The three-frame Figma rule

    The three-frame Figma rule

    Every Monday at our design review, someone opens Figma and scrolls. And scrolls. Twenty frames, six alternate flows, a stray Miro embed, three sticky notes with question marks. By frame nine, nobody remembers what the proposal is for.

    We made a rule. Three frames. That is the budget.

    Why three frames

    Three frames force a story: state, action, result. You cannot fake that shape. If you cannot show the current state, the moment of change, and the outcome inside three frames, the idea is not ready. It might be interesting, but it is not buildable yet.

    We use the rule for anything that lands in our Linear “Ready for eng” column. Before a ticket moves out of Design, the linked Figma file needs:

    • Frame 1: the state the user is in when this matters
    • Frame 2: the action or decision point
    • Frame 3: what changes, what breaks, what they see next

    Three frames is not a rule about fidelity. Lo-fi wireframes count. Screenshots with red boxes count. Three annotated hand drawings count. The constraint is narrative, not craft.

    What the rule caught

    The first month we tried this, our design lead pulled six proposals back to discovery. One of them was a payments refactor that had ninety-two frames and no clear before or after. When she asked the designer to pick three, he could not. The empty state was fine. The success state was fine. The middle was a fog of edge cases and dependency guesses. That fog was the real work, and it was invisible under all the surface polish.

    Now, when a designer says they cannot get under three frames, we treat it as a signal, not a failure. It tends to mean one of these:

    1. The problem is two problems glued together, and we should split the ticket
    2. The team has not agreed on which user is being served
    3. We are drawing solutions before we have named the decision

    When we break it

    We break the rule for onboarding flows, migration paths, and anything with a compliance step that legal needs to review verbatim. Those are storyboards, not proposals. The distinction matters. A storyboard documents a sequence. A proposal argues for a change. Three frames is for the arguing.

    We keep the frame counter in the corner of every design review agenda. Nobody enforces it. It sits there, and the room reads it.

  • How we run design reviews without a design lead

    How we run design reviews without a design lead

    We do not have a design lead at Velo. We have four product designers, six engineers, two PMs, and a shared Figma library that grows about eight components per week. For a long time our design reviews were a bottleneck. One person owned the calendar, the queue, and the taste. When they were on PTO, everything stalled for a week.

    Six months ago we scrapped that model. We now run a weekly design review every Thursday at 2pm, rotating facilitator, and we ship through roughly ten designs per hour. Here is how it works.

    The three questions on every draft

    Every design that enters the queue answers three questions before anyone speaks. We put them at the top of the Notion review doc, and the designer fills them in when they submit.

    1. What is the smallest unit of behavior this changes? A button state, a full flow, or a shared component?
    2. What did we consider and reject? We ask for two rejected options with one sentence each.
    3. What are we not sure about? The designer flags the specific decision they want feedback on.

    The last question does the heavy lifting. Before we added it, reviews drifted into whatever the loudest person felt like discussing. Icon choice on a screen whose real question was permissions logic. Copy on a modal whose real question was the empty state upstream. Now the designer sets the target. If someone wants to raise something outside that target, they add a comment in the Figma file and we move on.

    We also cap each entry at 250 words of context and three frames of the design itself. If it needs more, it is two designs, and it splits into two queue entries.

    Who shows up

    The room is small on purpose. Five people, always:

    • The designer whose work is under review
    • A rotating facilitator, drawn from the design pod
    • One engineer from the squad that will build the work
    • One PM, usually the one who wrote the ticket in Linear
    • One outsider from a different squad

    The outsider is the piece we fought about the longest. Early on we ran reviews with only the squad that owned the work, and every review turned into a status meeting. People agreed with each other because they had agreed with each other three days ago. Adding one person who had no context forced the designer to explain the thing they had stopped seeing.

    We rotate the outsider through a Slack workflow. Every Monday morning a bot posts in #design-review with the week’s roster, drawn from a list of eleven people across four squads. If you get tagged and cannot make it, you swap with someone else in the thread. No one has skipped in eight weeks.

    The facilitator’s job is not to have opinions. It is to keep the clock, read the three questions aloud, and cut off tangents. We rotate facilitators every week so that the role does not calcify into a proxy design lead.

    How we hit ten designs per hour

    Six minutes per design, hard cap. The facilitator sets a timer in Slack, visible to everyone. When the timer hits zero, we stop, capture the decision or the follow-up in the Notion doc, and move on.

    Six minutes sounds absurd until you try it. What we found:

    • Most designs need one or two decisions, not twelve. The three questions narrow the surface.
    • The engineer catches feasibility issues in the first ninety seconds. If a design assumes a field we do not have in our GraphQL schema, that comes out fast.
    • The PM catches scope creep. If a design has quietly grown a new setting screen, the PM flags it and we decide whether the ticket needs a split in Linear.
    • The outsider catches the assumptions everyone in the squad shares. This is the highest value input we get, and it usually shows up in the last minute.

    If a design cannot be resolved in six minutes, it goes into a longer 30 minute slot on Friday. About one in six designs ends up there. We track this in a Notion database, and if the ratio climbs above 25 percent for a month, we know something is off in how designs enter the review, and we tune the three questions.

    What we gave up

    We gave up consistency of taste. Without a single lead, our components sometimes drift. We catch this with a monthly audit that anyone on the design pod can run against the Figma library, and we open a GitHub PR to reconcile the drift. It takes about three hours.

    We also gave up the comfort of a single person who says yes or no. Decisions in our reviews are made by the designer, informed by the room. The facilitator will push back if a decision is being avoided, but no one else has veto power. This felt scary for the first month. It stopped feeling scary once we noticed the queue kept moving.

  • Why we stopped tracking hours

    Why we stopped tracking hours

    We killed the timesheet in March. Nobody misses it. What replaced it made our finance team nervous for about six weeks, and then it stopped making them nervous. Here is what happened.

    What we track now

    We stopped asking engineers, designers, and PMs to log hours against project codes in Harvest. It rewarded the wrong thing: presence. A person could sit on a Linear ticket for three days, log 24 hours to it, and produce nothing shippable. The timesheet said the work was funded. The product said otherwise.

    We now track two things per squad, weekly, in a Notion database that pulls from Linear and GitHub:

    • Shipped units of work: tickets that closed the loop, meaning code merged, feature flagged on for at least one real customer, and no rollback within seven days.
    • Cycle time: median days from “In Progress” to “Shipped” on Linear, per squad, broken out by ticket size (S, M, L).

    That is it. No story points. No hours. If a squad shipped four medium tickets in a week with a cycle time of 3.2 days, that is the artifact. We review it every Thursday in a 25 minute meeting called Ship Review.

    What finance pushed back on

    Our CFO’s first question was fair: how do we capitalize engineering costs for the R&D credit if nobody is logging hours to projects? HMRC wants a defensible allocation. Auditors want a paper trail. “The vibes were good on Thursday” is not a paper trail.

    Her second question was harder. If a contractor bills us for 40 hours and we have no internal hours to compare against, how do we know we are not being overcharged?

    What we told them

    We proposed a trade. Finance gets a defensible model. Squads keep their calendars.

    1. Every squad has a fixed roster, and every roster maps to one or two product areas in a Notion table. Payroll cost per squad is known. That is our allocation basis, reviewed quarterly.
    2. Contractors still log hours in Harvest, because they bill by the hour. Employees do not, because they do not.
    3. The Ship Review output feeds a monthly note to finance: shipped units, cycle time trend, and any squad where cycle time doubled without a headcount change. That flags stuck work faster than a timesheet ever did.

    Six months in, our R&D claim went through without a query. Cycle time on our payments squad dropped from 8.1 days to 4.4. Nobody has asked to bring hours back.

    A timesheet measures whether you showed up. A shipped ticket measures whether the customer got something.

    We know which one we would rather report on.

  • How we run the same customer interview every quarter

    How we run the same customer interview every quarter

    Every quarter, on the second Wednesday of the last month, we sit down with three Velo customers and ask them the same eight questions we asked last quarter. Same script. Same order. Roughly the same time of day. The value is not in any one conversation. The value shows up when we diff the transcripts.

    The script we do not change

    The temptation to tweak questions is strong, and we resist it. Once you edit a question, you cannot compare answers across quarters. Our fixed eight, in order:

    • Walk us through the last thing you did in Velo this morning.
    • What did you almost use Velo for, and then use something else?
    • Which tab do you open first when you log in?
    • Who else at your company touches Velo, and how?
    • What have you built around Velo that we did not ship?
    • When was the last time Velo surprised you, good or bad?
    • If we shut down tomorrow, what would you replace us with?
    • What are you doing on the days you do not open Velo?

    Three users, chosen from a rotating pool of twelve power accounts. One customer we keep constant across the year for a longitudinal read. Two rotate.

    The diff is the deliverable

    We record with Grain, get transcripts back within the hour, and drop them into a Notion database with one row per user per quarter. On Friday we run a working session with product, design, and one engineer. The rule for that session: no talking about individual quotes until we have compared the same question across quarters.

    What we look for in the diff:

    1. Words that appear this quarter and did not appear last quarter. Q1 nobody said “audit.” In Q2, all three did.
    2. The “almost used Velo for” answer shifting toward one workflow. Two quarters in a row of the same near-miss becomes a Linear ticket in the next planning cycle.
    3. Tab-first answers drifting away from the tab we consider the home screen. That happened to Reports in Q3 last year and pushed the nav redesign.
    4. Homegrown scripts built around us. Every quarter someone shows us a Google Sheet stitched to our API. That sheet is the roadmap.

    The interview does not tell us what to build. The diff between interviews does.

    What we do with it

    By the following Monday, the working session produces a one page memo in Notion with three sections: patterns that hardened, patterns that softened, and one thing we were wrong about. That memo goes to the whole company in Slack, and the three tickets it generates land in the top of the next sprint. That is the whole loop. Eight questions, three users, one diff, one memo, three tickets.

  • The PRD template we abandoned

    The PRD template we abandoned

    For eighteen months, every feature at Velo started the same way. Someone opened a Notion page called “PRD Template v4,” scrolled past six pages of headings, and started typing. Problem statement. Goals. Non-goals. Success metrics. User stories. Edge cases. Rollout plan. Open questions. Appendix.

    By the time the doc reached our Thursday product review, it was 2,400 words of dutiful prose. When we asked what the user would do differently after we shipped, the room went quiet.

    What the template was hiding

    We reread twelve PRDs from Q1 and Q2. A pattern showed up in the verbs.

    • “Improve onboarding” appeared in four docs.
    • “Optimise the flow” showed up in five.
    • “Reduce friction” was in nine of the twelve.
    • “Increase engagement” closed out most success sections.

    None of these verbs pointed at anything a human could do or fail to do. They were placeholders for thinking, and the template rewarded us for placing them. Six pages of scaffolding meant six pages to fill, and the easiest way to fill six pages was to reach for words that meant almost anything.

    The reviews got worse in a specific way. Engineers would ask a sharp question on page one, and by page five we had lost the thread. Design would flag an interaction we had not considered, and someone would promise to “cover it in the appendix.” The appendix became a graveyard. In one memorable case, a Linear ticket sat in Backlog for eleven weeks because the PRD had three conflicting definitions of “activated user” and nobody wanted to be the person to pick one.

    We had built a document that made disagreement expensive to surface.

    The one-page brief

    In August we tried something smaller. Three prompts, one page, hard cap at 400 words.

    1. Problem. What is going wrong for a specific person on a specific day? Name the person by role, not persona.
    2. Target user. Who feels this problem most acutely, and which of them do we have access to for interviews within the next week?
    3. First observable behaviour. What will this person do after we ship that they cannot do or would not do today? Describe the action, not the outcome.

    That third prompt did most of the work. “Increase engagement” fails it. “The ops lead exports the reconciliation report as a CSV before their Monday standup” passes it. You can watch a person do that or not do it. You can ask them why they did not. You can count it in Datadog.

    We wrote the first brief for a billing export feature. It took forty minutes. The Thursday review took nineteen. Two engineers pushed back on the target user selection, we agreed to run three interviews before writing a line of code, and the ticket moved to In Progress the following Tuesday.

    What we gained

    The obvious thing was time. Our median PRD used to take a product manager somewhere between four and seven hours to draft, plus another two in edits after review. The brief takes forty to ninety minutes. Reviews compressed from an hour to under thirty because there was nowhere to hide vague thinking.

    The less obvious thing was that engineers started reading the doc. When something is six pages of prose, the second engineer on a project skims. When it is one page and the third bullet is a specific behaviour, they read it, and they argue. We stopped confusing “we have a PRD” with “we know what we are building.” Writing a brief did not feel like progress in the way writing a six-page PRD had felt like progress. That turned out to be honest.

    What we lost

    We lost some things we miss. The old PRD forced us to enumerate edge cases up front, and the brief does not. Two of our first four briefs shipped with a rollback plan improvised in a Slack thread the morning of launch. We have since added a lightweight “what breaks” checklist that lives next to the brief but is not part of it, and reviews now include a five-minute walkthrough of that checklist.

    We also lost some institutional memory. A six-page PRD, for all its flaws, was a decent artefact for the person joining the team six months later. Our briefs are terse enough that a new engineer opening one in GitHub cannot always reconstruct why we made the calls we did. We are experimenting with a short decisions log per feature, kept in the same Notion database, updated as we go. It is not a Figma spec and it is not a design review transcript, but it stops the “why did we do it this way” question from getting answered with a shrug.

    What we would tell you

    If your PRD template has more than three prompts, ask the last five people who filled one in whether the extra prompts changed a single decision. If the answer is no, cut them. The template is not neutral. It teaches your team what counts as thinking.

    The verbs are the tell. If your docs are full of them and nobody has flinched in a review this quarter, the template is doing the flinching for you.

  • The two week feature flag rule

    The two week feature flag rule

    Most engineering blogs will tell you feature flags are free. Ship dark, roll out slow, keep the fallback path warm. We used to believe this too. Then we counted our flags.

    Last month we had 47 flags in LaunchDarkly. Nine were older than a quarter. Four were older than us remembering why they existed. Two of them fought each other in production and caused a payout retry loop that Datadog flagged at 2:14 on a Tuesday afternoon.

    Why the “keep it around, it’s cheap” argument is wrong

    The pitch for long lived flags is that they cost nothing. A boolean check, a config entry, a line in a dashboard. The real cost shows up somewhere else:

    • Every conditional doubles the state space a new engineer has to hold in their head when touching that file.
    • Test matrices grow multiplicatively. Two stale flags plus one new one gives you eight paths, and nobody writes eight tests.
    • Old flags rot silently. The useNewCheckout branch you shipped in March has drifted from the legacy branch it lives next to, and neither of you noticed.
    • On call gets worse. When something breaks, the first question is always “what flags are on for this customer?” and the answer takes twenty minutes to assemble.

    The flag was supposed to be a valve. Left open, it becomes a fork in the codebase you have to maintain twice.

    The rule we adopted

    Every flag has an owner and a fourteen day timer. At day fourteen, the losing branch gets deleted. Not deprecated, not marked for cleanup in Linear, deleted. If we shipped the new checkout behind a flag and it stuck, the old checkout code goes. If the rollout failed, the new code goes.

    The uncomfortable part: sometimes we delete a branch that turns out to be needed later, and we rewrite it. We accept that cost. A week of rework is cheaper than a year of dual maintenance.

    The mechanics are boring. Every Monday, a GitHub Action posts the list of aged flags to our #eng-hygiene Slack channel with the owner tagged. If the flag is still needed, the owner extends it once, by another week, and writes why in the thread. Second extensions get escalated to the Friday engineering sync.

    A flag that has been on for a month is not a flag. It is a feature you forgot to finish shipping.

    Since we started, our flag count has dropped from 47 to 12. Half the incidents we traced back to “unexpected flag interaction” have stopped happening. The rework tax has been real, and worth it.

  • How we handle mid-cycle scope creep without cancelling the cycle

    How we handle mid-cycle scope creep without cancelling the cycle

    Every two-week cycle at Velo starts with a plan we believe. By Wednesday of week one, something has usually shifted. A support pattern turns into an incident. A customer contract lands with a hard integration date. A design review surfaces a flaw that means the header refactor is bigger than we scoped. The temptation, when this happens, is to either pretend nothing changed and burn the team out, or blow up the cycle and replan from scratch. We do neither.

    Instead we run a small, boring ritual every Wednesday at 11:00 called the mid-cycle trim. It takes 30 minutes, involves the tech lead and the PM for each squad, and follows the same three moves every time. Here is what we do, and why the ritual survived four quarters when almost nothing else in our process did.

    The three moves

    We treat the cycle as a fixed container. If new work goes in, existing work comes out. That constraint is the whole game. The Wednesday trim is where we honour it.

    1. Cut the two lowest-value items still open. We pull up the Linear cycle view, sort by our internal priority score, and mark the bottom two tickets as cycle: deferred. They keep their estimates and their context; they lose their promised delivery date. If a ticket is already in review, we leave it; the cost of context-switching outweighs the saving.
    2. Split the highest-value item into a smaller slice. The biggest ticket in the cycle is usually the one hiding the most risk. We ask the engineer working on it: what is the smallest piece of this that a real customer can use on Friday of week two? That becomes the new scope. The rest becomes a follow-up ticket, linked, with the notes carried across.
    3. Log the trade in the rolling scope-change page. We keep one long Notion page per quarter titled Scope changes, Q3. Every trim adds a row: date, cycle number, what came in, what came out, who decided, one sentence on why. No approvals, no drama. A record.

    That is the whole ritual. Cut two, split one, log the trade. We finish by 11:30 and post the diff in the squad Slack channel by lunch.

    Why each move earns its place

    The cuts exist because engineers will not admit a ticket is low value while it is theirs. Naming the bottom two of the list, out loud, in the same meeting every week, removes the social cost. Nobody has to argue for their pet feature to survive. They only have to argue against a specific cut, and if they cannot, the cut stands.

    The split exists because our worst cycles were always the ones where the biggest ticket slipped by three days and dragged three smaller ones with it. Forcing a Friday-of-week-two slice on the largest ticket surfaces the risk while there is still time to react. The engineer often finds the slice is 40 percent of the work and delivers 80 percent of the value, which is a good trade even without a scope crisis.

    The log exists because we used to have the same argument every quarter about whether the team was under-delivering. Now we open the Notion page in the retro and count. Last quarter we absorbed 23 mid-cycle changes across six squads without extending a single cycle. That number ended a lot of arguments.

    What we do not do

    A few things we tried and dropped:

    • We do not require the CTO or a product lead to approve trims. The squad decides. If the squad is wrong, we catch it in retro, not in a gate.
    • We do not track “scope creep” as a metric per squad. It punishes the teams closest to customers, which is the opposite of what we want.
    • We do not roll deferred tickets automatically into the next cycle. They go back to the backlog and compete on merit. About a third never come back, which tells us the trim was correct.
    • We do not run the trim on Monday. Too early, the picture is still forming. Or Friday. Too late, the week is already lost. Wednesday is the fulcrum.

    What it looks like in practice

    Two weeks ago, the billing squad walked into Wednesday with a fresh Datadog alert showing invoice PDFs failing for 4 percent of enterprise customers. They cut a small copy update on the settings page and a nice-to-have export format. They split the ongoing tax-rate refactor: ship the US states this cycle, hold EU VAT for the next one. They logged three lines in the Notion page. The incident work landed on Thursday of week two. Nobody worked a weekend.

    That cycle looked, from the outside, like a normal cycle. That is the point. The ritual is not glamorous, and it does not solve the underlying question of why new work keeps arriving. It gives us a repeatable way to say yes to the important new thing without lying to ourselves about the cost. We recommend stealing it.

  • Why we killed the daily standup for a 15-minute doc

    Why we killed the daily standup for a 15-minute doc

    Our daily standup used to run 22 minutes on a good day. Ten engineers, one PM, one designer, all half awake, waiting their turn to recite tickets that lived in Linear anyway. We killed it in March. What replaced it is a single Notion page called Morning Doc, and it takes each of us about 90 seconds to fill in.

    How the doc works

    Every engineer, PM, and designer on the delivery pod owns a row. By 10am local time, each owner drops three lines under their name:

    • Yesterday: what shipped or moved, linked to the Linear ticket.
    • Today: the one thing they intend to close before EOD.
    • Blocker: a person, a decision, or a dependency. If none, they write none.

    That is the entire contract. No status colors, no percentages, no vibes. If your row is empty at 10:01, our engineering lead pings you once in Slack. Miss it twice in a week and you owe the pod a coffee run.

    Blockers move to threads, not meetings

    The rule we care about most: blockers do not sit in the doc. The moment someone writes one, they cross-post it into #pod-delivery as a thread with the format Blocker: [who I need] [what I need] [by when]. Whoever owns the answer replies in that thread. Our internal SLA is one hour during working time. Last month we hit it 94 percent of the time, measured with a small GitHub Action that scrapes reply timestamps.

    If a blocker sits unanswered past noon, it gets escalated to the pod lead. That has happened four times this quarter.

    What we gained, and what we gave up

    The wins are boring and real:

    1. We got roughly 110 engineer-minutes back per day across the pod.
    2. Blockers surface in writing, so Thursday retros have receipts instead of memory.
    3. Nobody performs progress for an audience. The doc rewards specificity over storytelling.

    What we gave up is real too. New hires miss the social glue of seeing faces every morning, so we kept a 20-minute Tuesday pod sync for demos and messy conversations. Designers occasionally want a live thinking-out-loud session, and we book those ad hoc in Figma instead of pretending standup was the right container.

    The morning doc is not clever. It is a shared page with a deadline and a Slack rule attached. But it turns out most of what standup gave us was the deadline, and most of what it cost us was the meeting.

  • Why our retros stopped finding the real problem

    Why our retros stopped finding the real problem

    The Friday retro ritual

    Every team we have worked on runs the same shape of retro. Sixty minutes on a Friday, three columns in a Miro board, sticky notes for what went well, what went poorly, and what to try next. Someone dot-votes. Someone else copies the top three items into a Linear ticket that nobody opens again. We used to run it that way too.

    In Q2 of last year, we shipped a release that took down billing for four hours. The retro landed the following Friday. The dot-vote surfaced “unclear on-call handoff” as the top item. We wrote a Linear ticket to rewrite the on-call runbook. Six weeks later, we shipped a different release that broke webhook delivery for two hours. The retro found the same category of problem, worded slightly differently, and produced another ticket that also went nowhere.

    The on-call runbook was fine. The problem sat upstream of it, and nobody in the room was willing to say so on a Friday afternoon with the person who owned the decision sitting three seats away.

    Why the standard retro fails

    Retros as most teams run them optimise for social comfort, not for truth. Three failure modes we kept hitting:

    • Recency bias. The team remembers Thursday’s deploy noise, not Monday’s design decision that set the deploy up to fail.
    • Consensus bias. Dot-voting rewards items several people already agree on, which selects for symptoms over root causes. Root causes are usually held by one or two people who saw them early and stayed quiet.
    • Performance bias. A live meeting is a stage. People tell the version of the story that protects the relationship, not the version that would help the next team.

    The billing incident showed us all three. The engineer who had raised a concern about the migration plan two weeks earlier did not repeat that concern in the room. Nobody wanted to spend the last hour of the week relitigating a decision that felt settled.

    The retro found something true and small. It missed the something true and large, because the format could not hold it.

    What we do instead

    We replaced the Friday retro with three artefacts, spread across the week. None of them takes more time than the meeting they replaced. We have run this shape for eleven months across two product squads and one platform squad.

    1. A written pre-mortem, filed on Wednesday

    Whoever owned the incident, feature, or sprint outcome writes a one-page pre-mortem in Notion. It is not a report of what happened. It is a written attempt to answer one question: if this failure repeats in six months, what will the story be? The author writes it alone, without review, and posts it in the squad Slack channel by end of Wednesday. It is time-boxed to forty-five minutes. Long documents mean somebody is hiding.

    2. A two-question survey, sent Thursday morning

    Every person on the squad, plus two adjacent stakeholders (usually a designer and a customer support lead), gets a Google Form with two questions:

    1. What did you see, hear, or think during this work that you did not say out loud?
    2. If you had a private ten-minute conversation with the person most responsible for the outcome, what would you ask?

    Answers are anonymised by the facilitator and pasted into the Notion doc under the pre-mortem. Response rate sits above ninety percent because the questions are specific and the form takes under five minutes.

    3. One blameless conversation, Monday at 10am

    The squad meets for thirty minutes on Monday. The pre-mortem and the survey answers are already in the room. The facilitator, who is not the tech lead, reads three or four survey answers aloud and asks the author of the pre-mortem to respond to them. No sticky notes. No dot-voting. No action items produced in the meeting itself. Proposals get added to the Notion doc during the following twenty-four hours, once people have had time to think.

    What changed

    The billing incident was one of eight retros we ran through the old format. The webhook incident was the ninth. Between month four and month eleven of the new format, we ran six retros across incidents of similar severity. Two produced Linear tickets that closed within a sprint. Three produced changes to how we scope Datadog dashboards before we ship, not after. One produced a decision to stop building a feature that two engineers privately thought would not land, and had not raised in a Friday meeting.

    The change is not that we find more problems. It is that the problems we find are the ones that matter. Three shifts explain most of it:

    • Writing before speaking gives people room to admit things they would not admit in a room.
    • Splitting the process across three days lets recency bias fade.
    • Removing the ritual of “action items produced in the meeting” removes the pressure to produce something visible, which is what pushes teams toward the easy, wrong answer.

    What we still get wrong

    The Monday conversation is fragile. If the facilitator lets it become a debate about the pre-mortem’s conclusions, it collapses back into the old format. We have had two of those in the last year. Both times, the survey answers that mattered most did not get read aloud, and the meeting ended with everyone agreeing on a symptom.

    We have also not solved the problem of what to do when the person most responsible for the outcome is the tech lead running the process. We rotate facilitation to a peer squad’s engineer in those cases, but it is not a clean answer.

    The retro, as most teams run it, is a meeting that produces the feeling of learning without the substance of it. If your Linear board carries three open “improve on-call handoff” tickets from three different retros, that is the signal.

  • The quarter we shipped no features

    The quarter we shipped no features

    Last October, three days after our Q4 planning offsite, we made a decision that felt reckless at the time. We were going to spend the entire quarter without shipping a single new feature. No new modules, no new integrations, no new dashboards. Only bugs, docs, and the internal tools our engineers had been asking for since spring.

    Our head of sales, Mira, found out on a Monday morning during our weekly go-to-market sync. She went quiet for about eight seconds, then asked whether we were serious. We were.

    What sales was worried about

    Mira had four deals in the pipeline that hinged on a specific promise: a Snowflake connector we had been talking about since June. Two of those deals were mid-market, one was a renewal expansion, and one was a competitive replacement worth around 180k in annual contract value. She pulled up the deal notes in Notion and walked us through each one.

    The fear was reasonable. If we froze features for 90 days, three things could happen:

    • Prospects would walk to competitors who kept shipping.
    • Existing customers waiting on requested features would churn at renewal.
    • The sales team would lose narrative ammunition on discovery calls.

    We agreed to review the freeze monthly. If any of those signals showed up in the data, we would call it off. Mira asked us to write down what “showed up in the data” meant, so we did: net revenue retention below 108%, gross churn above 1.4% monthly, or two consecutive weeks of stalled pipeline movement on flagged deals.

    What we shipped instead

    The engineering team split into three squads. One squad, which we called Fixit, worked exclusively through Linear tickets tagged with the “customer-reported” label. Another squad, Docs, sat with our support lead every Tuesday to identify the top ten most-hit help center pages and rewrite them. The third squad, Tooling, built the internal admin console engineers had been begging for.

    By week six we had closed 247 bugs, some of which had been open for over a year. The Datadog dashboard we cared about, the one tracking p95 API latency, dropped from 840ms to 310ms after two engineers rewrote a query planner in the reporting service. Our support team went from 34 open Zendesk tickets on any given Friday to 9.

    The internal admin console was the surprise. Before Q4, resolving a customer-reported billing issue took an engineer about 40 minutes: pull data from three tables, reconcile in a Google Sheet, patch, verify. After the tooling squad shipped the console, our support engineers were doing the same work in under 4 minutes. They did not need to page anyone.

    By the end of week ten, our on-call rotation had gone from one incident per shift to one incident every six shifts. Two engineers told me they were sleeping better. One of them had been talking about leaving.

    What happened to churn

    Here is the part nobody predicted. Gross churn went down. Not by a huge amount, but measurably: from 1.2% monthly at the start of Q4 to 0.7% by December. Net revenue retention held at 114%.

    Mira’s Snowflake deals: three of the four closed anyway. The connector question came up on discovery calls, and the answer we gave, which was that we were spending the quarter on reliability instead of new surface area, played better than we expected. One of the buyers, a VP of data at a healthcare company, told us he had never heard a vendor say that out loud. He signed in November.

    Two effects we did not model

    First, our NPS moved from 42 to 51. We got unsolicited notes in Slack from customer success managers whose accounts had stopped filing tickets. Second, our engineering hiring pipeline got healthier. Three candidates in December mentioned during their onsite loop that they had read our internal writeup about the freeze and wanted to work at a place that took reliability seriously.

    What we would do differently

    We got lucky on a few things and would not repeat every choice.

    1. We underestimated how disorienting the freeze would feel to product managers. Two of them felt sidelined for six weeks before we figured out how to give them meaningful work reviewing customer feedback and shaping the Q1 roadmap.
    2. We should have communicated the freeze to customers on day one, not week three. When we finally sent the note explaining what we were doing, the response was overwhelmingly positive. We could have banked that goodwill earlier.
    3. We did not set clear exit criteria beyond the churn and NRR thresholds. When Q1 planning arrived, some of us wanted to extend the freeze another month, and we did not have a decision framework for that conversation.

    We are not going to do this every quarter. Growth still matters, and a company that only fixes bugs is a company that gets displaced. But we now know the shape of what a deliberate pause looks like, what it costs, and what it returns. Next time we consider one, the conversation will be shorter, and the fear in the room will be smaller.