Alcott

Loading site
Skip to content

2026-09-20 · 6 min · By Alcott Dube

How to prioritise a backlog when everything is urgent

I use weighted scoring, explicit exceptions and a fixed capacity limit to separate urgent requests from work that deserves a place in the next delivery cycle.

Stone and ceramic blocks surround a bounded brass platform with room for only a few selected pieces.

I prioritise an urgent backlog by separating genuine obligations from discretionary work, then scoring the remaining requests against agreed criteria, evidence and effort. I use that ranking to choose work within a fixed capacity limit, with every exception recorded rather than disguised as a higher score.

Separate urgent obligations from feature requests

I don't start by scoring the whole backlog. An active security incident and a request for another export format aren't competing on the same terms. Before comparing features, I separate obligations that require a response from choices about where to invest.

An obligation needs a verifiable consequence: an incident affecting customers, a statutory deadline or a contractual commitment someone can point to. A senior stakeholder wanting something by Friday isn't enough. I ask for the date, the consequence of missing it and the smallest acceptable response.

I reserve capacity for those obligations before ranking discretionary work. If they consume the available capacity, I make that visible. A scoring exercise cannot fix overcommitment. For everything else, I require a short problem statement, an affected audience, an expected outcome and an owner. Requests without those basics go back for clarification, not into the meeting.

Choose a feature prioritisation framework with explicit weights

Intercom's prioritisation model compares reach, impact, confidence and effort. I borrow its separation of potential value from evidence, but add explicit weights where stakeholders need to negotiate competing priorities. This is an adaptation, not Intercom's formula.

I agree one planning outcome first, such as increasing completed self-service tasks over the next quarter. Then I agree the weights below before showing anyone the candidate scores. Otherwise, people can adjust the rules until their preferred feature wins.

  • Outcome impact, 40%: how much the work could change the agreed customer or business measure.
  • Reach, 25%: how many relevant users would encounter the change during the same measurement period.
  • Strategic fit, 20%: how directly the work supports the current priority rather than an unrelated objective.
  • Time sensitivity, 15%: how much value disappears if the work waits one delivery cycle.
A translucent form and smaller solid stones sit on separate balance pans, contrasting apparent size with tangible weight.
I weigh expected value against the evidence and effort behind it.

Define scoring scales before stakeholders score features

I score each criterion from one to five, using written anchors. For impact, I define ranges against the chosen outcome. On a task-completion measure, one might mean less than half a percentage point of improvement, three might mean one to two points, and five might mean three or more. Those are illustrative thresholds, not universal benchmarks.

For reach, I use the share of the relevant audience: below 1%, 1% to under 5%, 5% to under 15%, 15% to under 30%, and 30% or more. I keep the denominator and time period identical across requests. Registered accounts and weekly active users aren't interchangeable.

Strategic fit scores one for unrelated work, three for an enabling dependency and five for a direct contribution. Time sensitivity scores one when waiting has little consequence, three for documented delay costs and five for an opportunity that expires next cycle. I challenge duplicated arguments: a deadline doesn't also prove impact, and executive sponsorship doesn't establish strategic fit.

Adjust priority scores for confidence and total effort

My calculation is: priority score = ((impact × 0.40) + (reach × 0.25) + (strategic fit × 0.20) + (time sensitivity × 0.15)) × confidence ÷ effort. The result is a comparison aid, not a financial forecast.

I use confidence bands of 100%, 80% and 50%, following Intercom's published model. I assign the band according to the weakest material assumption, then attach the evidence: usage records, observed task failures, customer interviews or a comparable release. Several stakeholder opinions don't become independent evidence just because they agree.

I estimate effort in person-weeks, including design, engineering, testing and release work. Four person-weeks means twenty person-days in total, not necessarily four calendar weeks. I compare similarly shaped opportunities and test an effort range when estimates are uncertain. Dividing by effort can favour tiny improvements, so I rank larger investments separately rather than letting cheap polish consume the entire plan.

Compare feature requests with a worked scoring example

Consider a hypothetical choice between signup recovery and reporting expansion. Signup recovery scores four for impact, three for reach, five for strategic fit and two for time sensitivity. With 80% confidence and four person-weeks of effort, its score is 0.73.

Reporting expansion scores five, two, three and five respectively. Its strategic fit is lower because it enables later workflow improvements rather than directly completing a task. With 50% confidence and six person-weeks of effort, its score is about 0.32. Its approaching opportunity window matters, but doesn't erase uncertain demand or higher delivery cost.

I then test the assumptions. If reporting confidence rises to 80% and a narrower version takes three person-weeks, its score becomes about 1.03. That changes the order. The useful discussion is now specific: can the team substantiate demand and shape that smaller version? I don't spend meeting time arguing about hundredths when the effort estimate could halve.

Run a prioritisation meeting without negotiating every score

I circulate the problem statements, evidence, scores and estimates before the meeting. In a 45-minute session, I spend ten minutes confirming constraints, twenty reviewing disputed assumptions and fifteen choosing work and recording decisions. I don't read the backlog aloud.

I ask stakeholders to challenge inputs, not bargain over the final number. A disagreement about reach needs audience data. A disagreement about impact needs a causal argument or a test. When evidence is missing, I lower confidence or commission a bounded investigation; I don't average everyone's preferences.

One named decision-maker owns the final selection. I record any override with its reason, the work displaced and a review date. An override can be sensible, particularly when a relationship or dependency isn't represented well by the model. Concealing it inside inflated scores makes the framework less useful next time.

Turn ranked features into a capacity-limited delivery plan

A ranked list isn't a commitment. Basecamp's Shape Up separates worthwhile ideas from the bets a team chooses to fund, and uses an appetite to constrain the solution before detailed delivery planning. I use that distinction after scoring: what version deserves the time I'm willing to spend?

I check dependencies and specialist availability before filling capacity. Ten available person-weeks won't help if every selected item needs the same engineer next week. I leave explicit space for support and uncertainty based on the team's recent workload, rather than claiming a universal buffer percentage.

After release, I compare the agreed outcome, actual reach and delivery effort with the assumptions. I schedule that check when the feature should have had enough exposure to matter. I also remove requests whose problem has disappeared. A stale backlog doesn't become more useful because every row has a score.

Questions people ask

What is the best feature prioritisation framework?

I choose a framework based on the decision it needs to support. Intercom's reach, impact, confidence and effort model is useful for comparing opportunities; explicit weights help when stakeholders disagree about which criteria matter most.

How do you prioritise features without customer data?

I label assumptions and reduce confidence rather than inventing precision. If an uncertain request could consume substantial capacity, I prioritise a small research task or technical investigation before committing to delivery.

How do you handle stakeholders who say everything is urgent?

I ask what specifically happens if each request waits one cycle. Then I show the capacity constraint and ask the decision-maker to name the work that moves out when something moves in.

How often should you reprioritise a product backlog?

I review candidates before each planning decision, not whenever someone sends a message. Between those decisions, I reopen committed work for material evidence, changed obligations or incidents, with the switching cost made explicit.

Where I checked my thinking

Start a project