The B4 Framework — methodology

Two axes.
Four verdicts.

X · Strategic differentiation: does owning this system make you different in a way customers pay for?

Y · AI feasibility: could a competent team, with today's models, actually build and run it?

Every category's position falls out of five dimensions, scored 1 to 5. The call comes with how sure it is, and it changes with the size of the company asking.

Methodology v4.0 — August 2026

Fig. 01 — The matrix · tap a quadrant SEP 2026
AI feasibility →
Strategic differentiation →
⚒ Build 68 CATEGORIES · 4% OF THE INDEX

Own it. Build means own it. The category differentiates you competitively, and AI makes building it realistic. Build is small because there are a host of considerations and teams should really only build the software that gives them competitive advantages.

The five dimensions — scored 1 to 5

01

Specificity

Sets the X axis

How deeply the system encodes processes that are yours: data models, rules, workflows nobody else runs. Higher specificity means you need more customization, and buying generally costs more in integration and customization.

02

Strategic control

Sets the X axis

Are you going to be faster, better, or cheaper than your competitors by building a custom version of this software for your business?

03

AI feasibility

Sets the Y axis · evidence-graded

Buildability: can a competent team build and operate a custom alternative? Grounded in evidence-based research. AI feasibility doesn't score high unless teams have already built it in production, and a category whose only evidence is bolted-on AI features caps at 2.

04

Vendor value

Urgency modifier

How much value the vendor you're paying for offers. Will you use 20% of the features or 80%? It helps you determine how fast to act inside a quadrant.

05

Cost trajectory

Urgency modifier

Where the pricing is headed. Rising costs on a buildable commodity move a renewal from "someday" to "let's build it."

Score severity
1 → 5 · SINGLE HUE · GREEN MEANS ONLY "BUY"

How the verdict is decided

Methodology v4.0 — August 2026

A score is a reading, so the verdict is a range.

Five people scoring the same category do not land on the same number. They land near it. So publishing one badge per category and stopping there says more than the evidence supports.

Here is what that cost. The X axis is the average of two whole numbers, which means it lands on 3.5 a lot — and 3.5 is exactly the line between the quadrants. Roughly one category in seven sits on it. Under the old math, which side those fell on came down to how a comparison was written. Around three quarters of the index is within a single one-point move of a different badge.

So the index stopped publishing a point. Each category now carries a probability across all four calls, the call that holds the most of it, and a word for how sure that is. Same five dimensions, same 1-to-5 rubrics, same scores. What changed is the arithmetic that turns them into an answer.

01

Every score gets a band

Three of the five dimensions decide the quadrant. Each one is read as a range, not a point: one step either side of what it scored, equal weight on all three values. A dimension at 3 is treated as maybe 2, maybe 3, maybe 4. A 5 can only be wrong downward, so it bands to 4 or 5.

02

The band is enumerated, not simulated

Three dimensions, three values each, is 27 combinations. Every one of them gets placed on the map and its weight added to whichever quadrant it lands in. That produces a probability for all four calls. No sampling, no randomness — run it twice and you get the same answer to the last decimal.

03

An axis is high only at 4 or above

The line sits at 3.5, and a category has to clear it outright. Sitting exactly on the line reads as low. X is the average of two whole numbers, so it lands on 3.5 often — and that used to be settled by which way a comparison operator happened to point. Now it is settled by a rule you can read.

04

Ties fall toward the cheaper mistake

When two calls hold the same weight, the order is Buy, then Bridge, then Beware, then Build. Telling a team to build something they should have bought is the most expensive sentence this index can say, so Build is the last call a tie can produce. That is a stated principle, not an implementation detail.

05

The call comes with how sure it is

The quadrant holding the most weight is the call. Alongside it: clear when it holds 70% or more, lean between 50 and 70, split below 50. A near-call flag when the runner-up is within 15 points. Most of the index is a lean — that is the honest reading, and it was invisible before.

The org-maturity lens

"Could a team build this" depends on whose team.

A fifty-person shop and a fifty-thousand-person enterprise ask the build question with very different hands on deck. The lens reads the verdict for your team. It shifts AI feasibility one step: down at low AI maturity (no dedicated engineering), up at high (an AI-mature org with builders on staff). Every other part of the score stays put.

It works like a view setting. Three pills, switch them as often as you like, and your pick stays in your browser. Over half the index answers differently across the three maturity levels. Pro members and the MCP get the lens; everywhere else shows the medium reading.

The band and the lens do different jobs. The band says how far a reading might drift. The lens re-asks the question for a different team, and that answer carries the same confidence as any other.

Trust is built in

01

Categories, not vendors

Vendors churn, get acquired, pivot. A category's structural position moves slowly — which is what makes a score worth publishing and re-checking.

02

It re-scores itself

A nightly job scores against the quarterly publication deadline, and the AI frontier gets re-assessed every week. Every score lands as an append-only snapshot with its full five-dimension state. 1,603 and counting.

03

No vendor money

Analysts get paid by the vendors they rate. B4's one revenue source is the buyer. Only 4% of the index says build, and a framework that says "build everything" is a marketing document.

Run it on your own stack.

Search every scored category, check the dimensions, and challenge a verdict. See plans for the full directory, filters, and exports. Or connect your AI directly to the database through MCP.