Tag: process

  • How we unbreak a Wednesday sprint without cancelling it

    How we unbreak a Wednesday sprint without cancelling it

    Every team we know has had that sprint. Monday standup looked fine. Tuesday morning a payment webhook started dropping in staging, one engineer went out sick, and the design review pushed the checkout redesign back by two days. By Wednesday afternoon the burndown on our Linear board was flat, six tickets deep, and the sprint goal read like fiction.

    We used to cancel sprints in this situation. We stopped doing that around a year ago. Cancelling costs us the retrospective, the sense of finishing something, and the muscle memory of shipping on a cadence. What replaced it is a Wednesday recovery ritual that we run in about forty minutes.

    The Wednesday triage, not another standup

    We block thirty minutes on Wednesday at 2pm called “Sprint check”. It only fires when the burndown deviates more than twenty percent from the ideal line, which our Datadog dashboard flags in a Slack channel called #eng-signals. If the sprint is on track, the meeting is cancelled by 1:45pm and nobody joins.

    When it does fire, three people attend: the engineering lead for the squad, the product manager, and whoever picked up the on-call pager that week. No designers, no wider group. The point is a fast, honest read on the remaining ten working hours across four engineers.

    The question we ask is not “can we still finish everything?” It is “what one thing, if shipped by Friday, would make this sprint worth having run?”

    That one question forces a decision that the daily standup rarely produces. Standups report status. Wednesday triage rewrites the plan.

    Cut scope in the second half

    Once the anchor ticket is named, we walk the remaining Linear tickets and sort them into three buckets. We do this on a shared Notion page titled “Sprint 47 midweek reset” with three headings and drag ticket links under each.

    1. Ship this week. The anchor ticket and anything on its critical path. Usually two or three tickets.
    2. Defer to next sprint. Work that is not blocking anyone. Move it back to the backlog with a comment explaining why.
    3. Drop, do not defer. Tickets that felt urgent on Monday and no longer do. These get closed with a short note. If they matter again, someone will reopen them.

    The third bucket is the one that saves us. Roughly a quarter of what we plan on Monday gets dropped rather than deferred, and none of it has come back to bite us in the four sprints we have tracked this pattern.

    Slice the anchor, do not shrink it

    The highest value ticket is where teams tend to lie to themselves. On Wednesday we do not promise a smaller version of the same scope. We split the ticket into two Linear issues, and the parent becomes an epic.

    Take our checkout redesign example. The original ticket read: “Ship the redesigned checkout flow with saved cards, address autofill, and Apple Pay.” By Wednesday it was clear we had ten hours of engineering work left and about twenty hours of scope. The split looked like this:

    • PAY-412: Ship the redesigned checkout behind a feature flag, five percent rollout, saved cards only. Owner: Priya. Estimate: eight hours.
    • PAY-413: Address autofill and Apple Pay under the same flag. Owner: unassigned. Moved to next sprint.

    The slice we ship on Friday is a real, running thing in production, even if it sits behind a flag at five percent. Next sprint we widen it. What we avoid is the trap of promising the whole checkout by Friday, then delivering nothing and calling it a spike.

    The rolling change log

    Every scope change goes into a single Notion page we call the Sprint Ledger. One page per sprint, appended to as things move. Each entry has a timestamp, the ticket ID, what changed, and one line of why.

    Wed 14:32  PAY-401  Dropped. Duplicated by PAY-397 already in progress.
    Wed 14:35  PAY-412  Split from PAY-388. Anchor for the week.
    Wed 14:41  ONB-215  Deferred. Blocked on design; no unblock this week.

    The ledger takes about six minutes to fill in during triage. It is read twice: once by the wider squad on Wednesday afternoon, and once by the retro facilitator on the following Monday. Nobody hunts through Slack scrollback trying to remember what changed and why.

    What we get back

    Four things, measured over ten sprints since we started running this ritual:

    • We now finish about eighty percent of our stated sprint goal by Friday, up from around fifty percent when we would grind on the original plan or cancel outright.
    • Retros focus on cause, not blame. The ledger tells us what happened; we can talk about why.
    • Product managers push back less on Wednesday cuts, because the anchor is preserved and the ledger makes the trade visible.
    • On-call load in the second half of the sprint dropped, because we stopped shipping half finished work under time pressure.

    None of this requires new tooling. Linear, Notion, Slack, one recurring calendar block, and a rule about when it fires. The hardest part is not the process; it is the willingness on Wednesday afternoon to say out loud that Monday’s plan is no longer the plan, and to write down what replaced it before the day ends.

  • What done means for a task on our team

    What done means for a task on our team

    Every team we worked on before Velo had a definition of done pinned to a wiki page nobody read. Ours did too, until a Wednesday standup in March when Priya asked whether the invoice retry work was finished, and four engineers gave four different answers. That morning cost us a customer refund and a two hour incident review. We decided the Notion page was not the problem. The definition was.

    Three tries that did not stick

    Our first attempt was a paragraph in Notion titled “shipping standards” that said tasks should be “merged, tested, and reviewed.” It read fine on the page. In practice, “tested” meant whatever the author felt like: a unit test, a manual walkthrough, or nothing if the diff was under twenty lines. We shipped a race condition in the billing worker three weeks later because the author had run the change against a fresh database and assumed that counted.

    The second attempt was a Linear checklist template with nine items. Everyone checked every box, because the boxes were reported by the author and the reviewer had no way to verify half of them without opening five other tabs. The checklist became a ritual, then a joke, then a template we quietly stopped applying to new tickets.

    The third attempt was strict: a task was done when a designated QA engineer signed off in a Slack thread. This lasted eleven days. Our QA lead, Ruth, went on holiday, and the queue backed up to forty two tickets. When she came back, half the context was gone and she had to re verify work from memory. We had traded ambiguity for a bottleneck.

    The four criteria we settled on

    After the third failure, we spent a Friday afternoon working through what we needed the definition to do. It had to be verifiable by someone other than the author, it had to survive one person being out, and it had to answer the question Priya asked in March without a debate. We landed on four criteria, in this order:

    • Works: the change does what the ticket says, verified against the acceptance criteria written before the branch was cut. If those criteria were vague, that gets fixed before the ticket moves to review, not after.
    • Tested: automated coverage exists for the new behavior, and the tests fail without the change. The reviewer runs the suite locally or points at a green CI badge tied to the merge commit.
    • Deployed: the change is live in production, not staging, not behind a flag that has never been flipped on for a real user. If the work sits behind a flag, done waits until the flag is on for the intended audience.
    • Observed: a human has confirmed the change behaves as expected in production, using logs, a Datadog dashboard, or a real user event. Not a synthetic ping. A trace of the feature being used, or a metric moving in the direction we predicted.

    The order matters. If “works” is unclear, testing the wrong thing is worse than not testing. If we skip “deployed” and call something done at merge, we hide half our incidents in the gap between main and production.

    The compromise on observed

    Observed was the criterion that almost killed the whole definition. Half the team pointed out, correctly, that internal only changes have no production traffic to watch. A new admin report, a migration script, an internal CLI: none of these throw off metrics on the customer dashboards we use for observability. Waiting for a real user event on an internal tool would mean waiting forever, or fabricating one.

    We debated dropping the criterion for internal work. We tried, for a sprint. Two internal tools broke silently and we found out from a support agent who could not load the refunds page. The criterion needed to survive.

    The compromise: for internal only changes, observed means the author or a teammate has used the feature in production for its intended purpose, with a Loom or a screenshot posted to the ticket. Not tested it. Used it. If the ticket is a migration, the observation is the query result after the migration ran. If it is a CLI, it is the terminal output from a real invocation on the real database.

    The distinction we care about is between “I believe this works” and “this has done its job for a real person, once.” The Loom feels heavy the first time. It stops feeling heavy the second time somebody catches a broken admin page before a customer does.

    How the four criteria show up in our week

    Every ticket in Linear now has four checkboxes matching the criteria. The author checks the first three. The reviewer, or on internal changes any teammate, checks observed and pastes the evidence. Our Monday planning meeting starts by pulling the list of tickets marked done in the last week and skimming the observation links. It takes eight minutes. In the six months since we adopted this, we have had two rollback situations that a proper observation caught before the on call engineer noticed. We have also had one case where the observation link was a screenshot of the wrong environment, which is a different problem, and one we are still working on.

    We do not think this definition is universal. It is what our team of eleven engineers, on a codebase with sixteen deploys a week, needs to keep the wiki page honest. If the shape of the team changes, we expect the definition to change with it.