65%
of PMs report working on at least one initiative per quarter that they privately rated as low-signal
11hrs
average time a PM spends per week evaluating, scoring, and defending opportunity decisions
3.2x
more likely to hit roadmap targets when evaluation uses evidence-based scoring vs. gut feel

The hardest part of product management is not deciding what to build. It is deciding what not to build. Good opportunities come across your desk constantly: a customer who will churn without feature X, a competitive gap your sales team flagged, an executive who has strong opinions about the mobile app. Each one looks urgent in isolation. None of them exist in isolation.

Most PMs evaluate opportunities one at a time — is this worth doing? — without a systematic way to compare them against everything else on the roadmap. The result is decisions that feel reasonable in the moment and get relitigated every time a stakeholder pushes back. The PM who has not defined their evaluation criteria in advance will spend more time defending decisions than making them.

The principle behind the framework: A good opportunity is not necessarily the right opportunity. The right opportunity is the one that scores highest across all five evaluation criteria when you compare it against every other candidate on your roadmap. Saying no to a good idea is not a failure of vision — it is the discipline that makes saying yes to the right idea defensible.

The 5-Criteria Opportunity Evaluation Framework

Every product initiative gets scored against five criteria before it goes on the roadmap. These are not subjective impressions — they are structured assessments with defined input types. Here is what each criterion measures and why it matters.

1

Strategic Fit

What it measures: Does this initiative connect to a declared strategic theme, an active OKR, or a documented company priority? An opportunity that scores high on strategic fit is work that advances something the organization has already committed to. One that does not connect to anything is work that advances nothing — regardless of how individually compelling it is.

How to score: Confirmed connection to a strategic theme or OKR = 5. Partial connection (supports a related goal) = 3. No confirmed connection = 1. If you cannot identify the strategic reason an initiative should be prioritized above others, strategic fit is low by definition.

2

Evidence Strength

What it measures: How well-supported is this opportunity by customer evidence? Evidence strength is not about how passionate the requester is — it is about whether the underlying problem is documented and quantified. Strong evidence means multiple data sources converge on the same problem: support tickets, churn interviews, feature usage data, and user research all pointing in the same direction.

How to score: Multi-source evidence from 3+ channels = 5. Validated by 1-2 sources, clear pattern = 3. Single data point or anecdotal = 1. Customer feedback synthesis is the discipline that converts raw signals into this kind of evidence base — without it, evidence strength defaults to low.

3

Size of the Problem

What it measures: How many users are affected, how severely, and what is the expected impact on key metrics if it is solved or not solved? Size is not just user count — it is user count multiplied by the depth of the problem for each affected user. A pain that affects 20% of users mildly scores differently than one that affects 5% of users severely enough to drive churn.

How to score: Affects a large user segment with measurable impact on a key metric = 5. Affects a meaningful segment with projected but unmeasured impact = 3. Affects a small segment with unclear impact = 1. Use your analytics to quantify — if you cannot, size is at best a 3.

4

Opportunity Cost

What it measures: What does the team not get to work on if this initiative takes a slot? This is the criterion PMs most often skip — it feels uncomfortable to articulate. But every initiative that goes on the roadmap displaces something else. Making the tradeoff explicit is what makes the decision defensible three months later when someone asks why the reporting dashboard was deprioritized.

How to score: High-value alternative is clearly displaced = 5. Some conflict with other priorities = 3. Low conflict — team has capacity or this fills a gap = 1. A score of 5 on opportunity cost does not mean the initiative should not be built — it means the tradeoff should be documented and communicated explicitly.

5

Timing

What it measures: Is now the right time to build this? Timing is about both market conditions and internal readiness. A well-evidenced opportunity built at the wrong moment produces poor results: a mobile app shipped before the core product is stable, a feature that requires behavioral change before the category is mature, a competitive response launched six months after the window closed.

How to score: Market and internal conditions are aligned = 5. Conditions are partially aligned — some dependency risk = 3. Timing is off — either too early or window has passed = 1. Document why timing scores below 5 so the decision can be revisited when conditions change rather than filed and forgotten.

The Scoring Process: 3 Steps to a Defensible Ranking

The five criteria are only useful if they are applied consistently. Here is the process that makes scoring reliable rather than a rationalization exercise where the scores come out whatever produces the ranking the PM already preferred.

1

List every candidate initiative before scoring any of them

Gather all active opportunities — everything that has been raised in the last quarter, everything in the backlog with a理由 attached, every stakeholder request, every competitive gap. The list is the portfolio. You cannot compare opportunities you have not listed. Do this first, completely, before scoring anything.

2

Score all candidates against all five criteria before reviewing results

Apply the rubric to every initiative before you look at the aggregate scores. This prevents the anchoring effect: when you see a high score on one initiative, every subsequent score gets biased toward it. Scoring all candidates sequentially against each criterion separately (all strategic fit scores first, then all evidence strength, and so on) produces more consistent results than scoring one candidate completely before moving to the next.

3

Review the ranking and explicitly override where intuition differs from the scores

The scores are an input, not a verdict. If the rubric produces a ranking that contradicts your experience or the team's capacity, that is information — not a reason to rejigger the scores. Instead, override explicitly: "The scores rank initiative B over A, but I am overriding because [specific reason]. This will be documented in the roadmap changelog." Intuition override is legitimate. Undocumented intuition override is not — it just looks like bias after the fact.

The single most important scoring rule: You must score every candidate against every other candidate, not against a hypothetical baseline. Asking "is this a 3 or a 5?" in isolation produces inconsistent scores because the threshold for a 3 shifts depending on which candidate you are looking at. The question is always: "Given the full list of candidates, does this score higher or lower than that one on this criterion?"

The Three Mistakes That Kill This Framework

Teams that adopt this framework and see poor results usually fall into one of three failure modes.

1. Skipping the opportunity cost criterion

The most common failure is treating every criterion as if it is independently good rather than treating it as part of a portfolio comparison. When opportunity cost gets skipped, PMs build a list of individually worthy initiatives without ever asking what the list displaces. The result is a roadmap that makes sense in isolation and falls apart when you notice that the top five initiatives together require 18 months of engineering time for a 6-month roadmap.

The fix is to score opportunity cost on every candidate before the meeting where the roadmap is reviewed. If the opportunity cost score is "this displaces the north star metric work by a quarter," that needs to be in the scoring sheet, not in the PM's head.

2. Accepting stakeholder scores without the underlying evidence

When a stakeholder brings an initiative forward and scores it internally before the evaluation meeting, those scores often reflect advocacy rather than assessment. A sales leader who scores their top customer's request a 5/5 on evidence strength because they believe deeply in it is not lying — they are just not using the same rubric. The PM's job is to surface the actual evidence behind each score, not to accept the stated score at face value.

The fix is to require that every score of 4 or 5 comes with a specific evidence citation: which data source, which user segment, which metric. Scores without citations get noted as "unconfirmed" and count as 3 at best until validated.

3. Running the evaluation once and treating it as permanent

The roadmap changes when circumstances change — new competitive information, a major customer loss, engineering discoveries that shift estimates. An evaluation framework that was run once at the start of a quarter and never updated produces a roadmap that becomes increasingly disconnected from reality. The teams that use this framework most effectively treat it as a live process: updated when new evidence arrives, with explicit changelog entries explaining what changed and why.

Criterion What it measures Score = 5 Score = 1
Strategic Fit Connection to declared priorities Confirmed link to OKR or strategy No confirmed connection
Evidence Strength Customer data backing the opportunity 3+ independent data sources Single anecdote
Size of Problem Scale and severity of impact Large segment, measurable metric impact Small segment, unclear impact
Opportunity Cost What is displaced if this ships High-value alternative displaced No meaningful conflict
Timing Is now the right moment? Market and internal aligned Window closed or too early

How to Say No When the Scores Say No

The framework produces a ranking. The ranking produces a conversation. The conversation produces a decision. The decision needs to be communicated — to the stakeholder whose initiative scored low, to the team that will not be working on it, to the executive who asked for it in Q1 and noticed it is not on the Q3 roadmap.

Saying no without a framework is uncomfortable. Saying no with a framework is mechanical — you are not expressing an opinion, you are reporting a score. The stakeholder who hears "this scored 7/25 because the evidence base is thin, the timing is off, and it directly conflicts with the activation work we committed to in Q3" is not being told no by you. They are being told no by the data.

The PM's job in this conversation is not to defend the no — it is to make the scoring transparent enough that the no explains itself. If a stakeholder leaves a scoring conversation still believing their initiative should be higher, one of two things is true: either the scoring inputs were wrong (and should be updated), or the stakeholder has additional evidence that was not in the original scoring sheet (and should be added).

The framework does not eliminate disagreement. It structures it — so disagreements are about evidence and criteria, not about whose intuition wins.

What to Do With the Output

The scored and ranked list becomes the roadmap input — but it is not the roadmap itself. The ranking tells you which initiatives are strongest candidates. The roadmap process takes those candidates and builds a plan that accounts for engineering capacity, dependency order, and stakeholder alignment. The roadmap process failures are the places where even a well-scored opportunity list gets destroyed by a broken process — building in isolation, conflating stakeholder requests with evidence, treating the roadmap as a commitment rather than a living plan.

The evaluation framework also feeds directly into the product decision framework — the criteria here are the structured inputs that make product decisions defensible rather than political. A PM who can show exactly how an initiative scored on five objective criteria has a document that survives executive review in a way that "I just had a feeling" does not.

Once opportunities are scored and ranked, the discovery sprint is where the top-ranked initiatives get the structured exploration time they need before engineering commit. Scoring tells you what to explore. Discovery tells you whether the score was right.

Stop evaluating opportunities on gut feel.

ChiefProduct evaluates every initiative against your declared strategy, your actual metrics, and your team capacity — so the ranking that comes out of scoring is one you can defend to anyone.

See ChiefProduct in Action →
📊
ChiefProduct
AI product manager for your team. Continuous feedback synthesis, automated metrics monitoring, backlog scoring, and OKR tracking so you can focus on the decisions that actually require judgment.