B4 Research

The Continuum

Daily research informing B4 Index scoring.

Latest edition

August 24 to September 2, 2026

New edition each week · free to read

The last four editions · August 5 to September 2, 2026

1,603

Categories tracked

32

Score movements

$230.2B

Capital tracked

13

Capability signals

19

Papers reviewed

Recent research

Latest edition →
Kubernetes Policy-as-Code Enforcement

The prior assessment scored 3 on the view that the commercial management plane was significant remaining value; this pass surfaced the scale and breadth of named OSS-only production deployments — Swisscom (390 clusters, explicitly avoiding external tools), LinkedIn (500K+ nodes), plus the public Kyverno adopter roster — showing the self-run enforcement core is a mainstream documented choice, the level-4 bar. Deeper research, not a market move.

SEP 02

Score movements

  • Cloud Development Environment (CDE)

    Deeper research surfaced multiple named large-scale production self-hosted CDE deployments (Dropbox 1,000 devs, Credit Karma, EnBW, 2,500-dev defense org) on Coder Community OSS — evidence the prior 3 predated; the market itself did not move this quarter.

  • Cloud Migration & Modernization Tooling

    Deeper sourcing showed every documented production AI case (Spotify, Google, DataArt, Novacomp) is code transformation, while the execution core — replication, cutover, agentless discovery — has no independent self-built alternative and even AWS Transform depends on MGN agents; the prior 4 over-credited the analysis-layer wins to the whole category.

  • Chaos Engineering Platform

    Re-grounded against CNCF production endorsement of LitmusChaos and Chaos Mesh plus AWS FIS; the prior 3 predated recognizing the OSS core as a mainstream ~80% self-build rather than a well-resourced option, moving it to 4 and flipping the badge to BEWARE.

  • Regulatory Change Management / Horizon Scanning

    Deeper 2026 research and peer calibration, not a market move: the authoritative multi-jurisdiction content moat, human-owned defensible obligation inventory, and sub-5% embedded-production maturity show the self-build covers only ~50-70% of the core, so the first-pass 4 was too high and the correct level is 3.

  • ML Model Supply-Chain Scanning & AI Bill of Materials

    Deeper research reframed the score from build-for-own-needs rather than matching vendor detection depth: the OSS scanner-plus-BOM core (ModelScan, Fickling, picklescan, CycloneDX) is documented as production-feasible and increasingly mainstream in CI/CD, and the near-duplicate peer (AI Model Supply-Chain Security, 0.91 cosine) already sits at 4.

  • Intercompany Transfer Pricing / Operational TP Monitoring

    Fresh 2026 survey and case evidence shows self-built operational TP monitoring remains a well-resourced-team option rather than a mainstream choice (1% fully integrated, hybrid dominant), correcting the prior 4 down to 3.

  • Profitability & Cost Allocation Analytics (Standalone from BI)

    Fresh evidence shows self-builds cover the calculation layer but complete production EPM replacements remain rare and undocumented, correcting the prior 4 to a partial-buildability 3.

  • Royalty & IP Licensing Accounting Software

    Evidence shows catalog-scale self-built engines are confined to a few SI-assisted music majors (BMG, Sony) while the broader category buys or hybridizes, correcting the prior 4 to a well-resourced-option 3.

  • Debt Collection & Recovery Software

    Deeper 2026 research shows AI-enabled collections is mainstream but complete lender self-build is explicitly a minority practice — only early-stage self-cure/outreach is routinely built in-house, with late-stage and compliant end-to-end stacks bought.

  • Methane Emissions Monitoring & Leak Detection (MRV) Software

    Deeper research showed the self-built portion is the decision/analytics layer, not the measurement-and-validation core, which still requires bought sensing and independent verification; the prior 4 over-credited buildability of the core.

  • Utility Vegetation Management Software

    Better research showed independent full self-builds are research prototypes, not utility-wide production — utilities build only component/analytics layers while buying the CV/prediction core — so the prior 4 overstated production buildability.

  • AI-Powered Cash Flow Forecasting (Standalone, Non-TMS)

    Moved 4 to 3: fresh research shows the prior assumption of multiple mainstream production self-builds is unsupported — named cases are thin (only AstraFi) and hybrid is the documented winning pattern.

  • K-12 Family Engagement & School-Home Communication

    Lowered 4->3: deeper research shows full family-engagement self-builds are rare, uneconomic, and undocumented at production scale, and procurement still demands hosted platforms.

  • K-12 Substitute Teacher & Absence Management

    Lowered 4->3: deeper research shows self-build is rare and undocumented at production scale, the market switches vendors rather than building, and the substitute-network is a real moat.

  • Software Composition Analysis (SCA) / Dependency Security

    Deeper research surfaced what the prior assessment missed: Dependency-Track as a production OSS management plane and GitLab shipping Trivy as its default scanner — the 'management layer is vendor-only' premise behind the prior 3 was wrong.

  • Ecommerce Affiliate & Partner Management Software

    The prior 4 rested on 'multiple teams run custom-built affiliate programs' — the research found one documented referral-scale production system (redBus), a documented retreat (Jobber retiring its internal build over maintenance cost), and no named full publisher-affiliate self-builds carrying fraud, tax/KYC, and payout depth.

  • Ecommerce Headless Storefront / Frontend-as-a-Service (FEaaS)

    Deeper research showed the prior score understated documented practice: custom and agency-built storefronts are the majority pattern among headless adopters, and the prior assessment's own text conceded widespread production self-builds while scoring them as merely partial.

  • Ecommerce Loyalty & Rewards Program (Store-Level)

    The market did not move — closer research showed the prior assessment overstated the evidence.

  • GraphQL Federation & Gateway Platform

    Better evidence of production self-hosting on Cosmo (Apache-2.0) and Hive (MIT), reinforced by Apollo's 2026 removal of hosted cloud routing driving teams to self-host.

  • Headless / API-First CMS

    Better evidence of mainstream production self-hosting — Strapi at Airbus/Tesco/Sonos, Figma-backed Payload defaulting to self-host, Directus on SQL — lifting this to clearly-buildable.

  • Database Change Management & Schema Migration

    The prior score of 3 under-credited an already-mainstream buildable core; deeper research on OSS engine production adoption (numerous named enterprise references) plus AI now authoring migrations puts the core past the 80% mainstream-self-build bar.

  • Distributed Tracing

    Better research on OSS backend production scale — Jaeger at Ticketmaster, Tempo at Houzz, SigNoz/ClickHouse at Dream11 (357 TB spans/month self-managed) — showed the core is a mainstream 80%+ self-build, above the prior score of 3.

  • Log Management

    Deeper research on OSS production scale — Dropbox running company-wide Loki after retiring its legacy system, Qonto, Deutsche Telekom on Graylog — showed the core is a mainstream 80%+ self-build, above the prior score of 3.

  • Community-Led Support Platform

    Moved 3->4 because this pass surfaced Discourse's mature, named production references (OpenAI, GitLab, UiPath, Elastic, Docker) covering 80%+ of the community core — a mainstream self-host reality the prior score under-weighted; no market event, just deeper evidence.

  • In-Location & Kiosk Customer Feedback Platform

    Deeper 2026 evidence shows the dominant QR/phone feedback case is 'very high' feasibility and a days-long mainstream build, where the prior score of 3 under-weighted how trivial the software is and over-weighted the hardware slice; this lifts it to 4 and aligns it with its three near-identical feedback peers.

  • Crop Insurance Agency Management Software

    Research corrected the prior 'not buildable' read: agency-facing layers are buildable against public RMA data; only the regulated core requires an AIP.

  • Developer Documentation & API Portal Platform

    Prior 3 predated the maturity of Scalar's OSS interactive playground and OpenAPI plugins; the docs-plus-reference-plus-playground core is now self-hostable at scale.

  • Developer-First Software Localization Platform (String & UI Localization)

    Prior 3 understated OSS maturity; Weblate self-hosts ~90-95% of a conventional TMS and the react-i18next + OSS + AI stack covers 80-95% of a software team's needs in production.

  • Association Management (AMS)

    Corrected an evidence error in the prior score rather than a market move: the prior 4 counted Fonteva, Nimble AMS, MemberVerse, and HubSpot Member Center as independent self-builds, but these are commercial AMS products.

  • Loyalty Management

    Named production self-builds surfaced that the prior assessment lacked — American Eagle's full SaaS-vendor replacement, Selfridges' in-house build, and Open Loyalty running Heineken's D2C program — moving the evidence from partial-coverage to mainstream documented choice.

  • Work Management

    Deeper research separated the category core from its neighbors: the prior 4 leaned on orchestration-layer builds (Airflow, Temporal, Maestro) and a self-hosted task layer, but 2026 evidence shows the business-user work-management envelope — grids, portfolio hierarchy, resource planning, governance — is bought, with self-builds confined to narrow layers around it.

Funding

72
CompanyAmount
OpenAI (Nvidia credit guarantee)
Debt/credit financing (residual-value lease guaranties)Foundation Model APIs (LLM & Multimodal)Source
$100BAUG 17
Anysphere (Cursor)
Acquisition (by SpaceX)AI Code GenerationSource
$60BAUG 14
Marvell Technology (Google strategic deal)
Strategic stock warrant / custom-chip supply dealSource
$12.2BAUG 19
Anthropic
Pre-IPO credit facility (debt)Source
$10BAUG 18
OpenRouter
Acquisition (M&A) by StripeLLM Gateway & RoutingSource
$7BAUG 16
Databricks
Strategic funding round (NOT Series K — corrected)Data Warehouse+1Source
$5BAUG 13
Sony / TSMC image-sensor JV (Advanced Vision Semiconductor Manufacturing Corp)
Joint venture investmentSource
$4.7BAUG 11
Nebius Group
Convertible notes (private, debt)GPU Cloud / AI Infrastructure PlatformSource
$4.5BAUG 19
All 72 funding events

Capability signals

13
  • Production ProvenCodified agent workflows in production ops (onboarding, account management, DevRel)

    Evidence SEP 01 · Source

  • Production ProvenAutonomous agent-initiated payments with deterministic trust gating at 20M+ transaction scale

    Evidence SEP 01 · Source

  • Production ProvenReal-time per-user LLM spend enforcement on managed infra

    Evidence SEP 01 · Source

  • Production ProvenAgent-majority software development lifecycle at enterprise scale

    Evidence AUG 27 · Source

  • DemoedBehavioral observability catching agent failure modes invisible to conventional monitoring

    Evidence AUG 27 · Source

  • Production ProvenManaged autonomous agent payments with guardrails and observability

    Evidence AUG 18 · Source

  • Production ProvenEmployee-built managed agents running in production at scale

    Evidence AUG 17 · Source

  • BenchmarkedGPT-5.6 family: frontier agent performance at collapsing cost

    Evidence AUG 13 · Source

  • BenchmarkedNative agent-orchestration primitives in the Responses API

    Evidence AUG 13 · Source

  • BenchmarkedGemini 3.7 Flash: workhorse coding + agent model

    Evidence AUG 13 · Source

  • BenchmarkedSecurity agents that prove exploitability and drive fixes

    Evidence AUG 12 · Source

  • BenchmarkedBrowser/web-navigation agents at ~88% (WebVoyager, per aggregator snapshot)

    Evidence AUG 12 · Source

  • BenchmarkedGrok 4.6: frontier long-horizon agent + coding model

    Evidence AUG 12 · Source

Papers

19
Controlled Study

Benchmarks sharing the same label (e.g. 'bug fix') differ statistically on at least two of three task-demand axes (Spread–Novelty–Centrality), agent success concentrates uniformly in low-demand tasks for every model family and scale, and the behavioral signature of success is model-family-specific (Claude succeeds by matching gold-solution scope, Qwen by exceeding it).

What Does an Agentic Software Engineering Benchmark Measure? Profiling Task Demands and Agent Behaviour Beyond What Category Labels Reveal

What it means When comparing agent scores across benchmarks, check what the benchmark actually demands rather than trusting its category label — and expect deployed agents to succeed mostly on low-spread, low-novelty, low-centrality tasks, reserving the hard cross-cutting changes for humans.

arXiv preprint · SEP 01

Large Benchmark

On 203 real-world dependency-upgrade tasks with hidden code-level breakage (DEPBENCH), the best coding-agent configuration solved only 51.2% (104/203), with substantial variation across harnesses, models, and ecosystems.

Update from Hell: Can Coding Agents Survive Hidden Breakage in Dependency Upgrades?

What it means Do not hand agents unattended dependency upgrades — the maintenance work teams most want to delegate is precisely where agents fail half the time; keep upgrade PRs behind full test suites and human review, and expect ecosystem-specific gaps.

arXiv preprint · AUG 31

Large Benchmark

How a user invokes a coding agent changes its vulnerability to poisoned repositories by up to 4.5x in attack success rate across task types (Run-Tests 45.5% ASR vs Fix-Bug 8.6%), with test-execution tasks forming a silent attack surface (high attack success, low agent alerting).

Beyond the Payload: How User Invocation Shapes Coding Agent Vulnerability to Repository Poisoning

What it means Treat 'run the tests' in an untrusted third-party repo as the riskiest agent instruction, not the safest — sandbox agents when bootstrapping from external repos, and know that prompt phrasing and supplied rules shift the attack surface rather than eliminating it.

arXiv preprint · AUG 31

Controlled Study

A single factor explains 74.5% of common variance across twelve frontier benchmarks and substantially tracks model release date (R²=0.505), so most of the gap between models released months apart is calendar; under the pre-registered dimensionality rule the four economic benchmarks form no distinct capability factor — though a leave-one-benchmark-out test shows a multi-factor representation still predicts held-out economic scores better than a single general index, so they carry incremental predictive information without constituting a distinct latent capability.

One Capability or Many? Testing the Economic Validity of Frontier AI Evaluation

What it means Date-adjust before buying on leaderboard gaps: a small score difference between contemporaneous models is mostly noise around a time trend, not a capability difference — pick by trying models on your own workload, not by leaderboard deltas.

arXiv preprint · AUG 29

Position

Contamination types are best organized by which mitigation each defeats (direct, derivative, temporal, distributional, acquired) — a private held-out test set closes only the first — and in an audit of 41 evaluation documents, elicitation budgets were reported in just 13% of documents and none addressed all five types.

Benchmark Contamination: A Taxonomy Organized by Defeated Mitigation

What it means When a benchmark score informs a buy decision, ask which contamination mitigations the eval actually applied — a held-out private test set defeats only direct leakage, and current reporting practice almost never discloses enough to tell capability from leakage.

arXiv preprint · AUG 29

Large Benchmark

Rewriting SWE-bench tasks to match how real users phrase requests cuts resolution rates by 6.4 percentage points on average and can change model rankings; 88% of real prompts carry only a bare problem statement (alone or with limited context) versus 7% of benchmark problems, and stating Desired Behavior and Motivation measurably lifts performance while Environment Information and Reproduction Steps add tokens without benefit.

RealSWE: A Compositional Evaluation of Coding Agents under Realistic User Requests

What it means State the desired behavior and the motivation in every agent request — those two fields measurably lift resolution while reproduction steps and environment info just add tokens; and discount SWE-bench scores as an upper bound, since real phrasing costs ~6 points and can reorder models.

arXiv preprint · AUG 28

Small Benchmark

A manager-worker scaffold over a shared filesystem helps some models dramatically (+8 to +30 points on the 100 hardest recent LiveCodeBench problems, up to +42 in single-pass configs) and is null or negative for others (-1 to -9), roughly triples token cost, but buys accuracy more cheaply than upgrading to a larger model in the cases studied.

Zero-Shot Self-Orchestration with Ledger-Based Control for Improved LLM Coding Performance

What it means Test multi-agent scaffolds per-model before adopting: the gain is conditional, largest for smaller models and reasoning-off configs, and near-zero for large reasoning models — and compare the ~3x token bill against simply buying the bigger model, which this study found is often the worse deal.

arXiv preprint · AUG 27

Small Benchmark

LLM-generated backend services that pass functional tests still show statistically significant memory-growth trends under 48-hour sustained execution in most application-language combinations, indicating software-aging defects that correctness testing does not catch — and the aging trends also appear in some human-written implementations.

Investigating Software Aging in LLM-Generated Software Systems across Generation-and-Execution Environments

What it means Soak-test AI-generated services before deploying them as long-running processes — passing the test suite says nothing about memory behavior at hour 40; add memory-trend monitoring to acceptance criteria for generated backends.

arXiv preprint · AUG 26

Case Study

Teams absorbing high-volume AI code generation are shifting from line-by-line review toward layered supervision: machine-readable architectural conventions as preventive guardrails, linting/testing/CI repurposed as supervision infrastructure, and human review refocused on architectural reasoning and operational explainability.

When Review Alone No Longer Scales: Layered Supervision in AI-Assisted Software Engineering

What it means If AI generation volume is straining code review, don't add reviewers — externalize architectural intent into machine-checkable rules and let CI carry first-pass supervision, reserving humans for architecture and maintainability judgment.

arXiv preprint (accepted, ESEM 2026 SEIP track) · AUG 26

Large Benchmark

On end-to-end tasks drawn from market-validated AI-startup workflows, the strongest evaluated agent completes only ~30% of StartupBench despite substantial partial progress, with complex instruction-following and domain expertise as the main failure sources.

StartupBench: Benchmarking General-Purpose Agents on Market-Validated End-to-End Workflows

What it means Scope agent deployments to the ~30% of real end-to-end workflows they can actually finish; keep humans on complex multi-step instructions and domain-specific steps, and validate against market-representative tasks rather than researcher-chosen benchmarks.

arXiv preprint · AUG 18

Case Study

Vibe-coding tools (Lovable, v0, Replit) produce structurally distinct and uneven code quality from a single prompt: Lovable concentrates lower-severity issues but has a much higher code-smell density per KLOC, while v0 and Replit produce more aggressive severity profiles.

Comparing the Quality of Code Generated by Vibe Coding Tools

What it means Treat vibe-coded output as a first draft carrying tool-specific structural debt: run static analysis and budget remediation before shipping, and pick the tool on its quality trade-off, not just perceived speed.

arXiv preprint · AUG 17

Small Benchmark

A scan-fix-rescan pipeline cuts static-analyzer findings in AI-generated code by 29-69% across four Claude models, but remediation itself introduces new vulnerabilities in 15-22% of cases, and the best code-generation model (Opus 4.8) was not the best pipeline performer.

Securing AI-Generated Code: A Just-in-Time Vulnerability Detection and Remediation Pipeline

What it means Gate AI-generated code behind an automated scan-fix-RESCAN loop rather than a single fix pass, since remediation adds new vulnerabilities ~1 in 5 times; and choose the security-remediation model on measured residual findings, not on general coding leaderboard rank.

arXiv preprint · AUG 17

Case Study

In a longitudinal study of professional developers, the dominant efficiency bottlenecks were organizational dependencies and waiting for external validation (structurally stable over time), while a generative-AI usage barrier emerged mid-study and became the most frequently coded interview theme.

Factors Impacting Developer Efficiency: Results from an Adaptive Longitudinal Study

What it means Do not expect AI tooling to move developer efficiency while organizational dependencies and validation waits dominate; measure efficiency continuously (not one cross-sectional survey) and treat evolving AI-tool friction as a first-class, addressable barrier.

arXiv preprint · AUG 17

Position

Across a system-level synthesis, many apparent coding-agent 'model failures' actually originate in the harness, retrieval, state management, or verification layers, and layer-level improvements often fail to propagate to end-to-end outcomes.

Engineering Reliable Coding Agents: Evaluating and Operating the System Around the Model

What it means Evaluate and buy the coding-agent SYSTEM, not the model: attribute failures to the specific layer (harness, retrieval, state, verification) before swapping models, and instrument each layer so a fix is proven to reach end-to-end outcomes.

arXiv preprint · AUG 14

Controlled Study

On SWE-bench Verified, the higher-recall retriever setting (gold file present in 0.878 vs 0.806 of packs) LOWERS issue resolution; disabling per-file deduplication to favor within-file depth raises single-shot resolve rate +7.6pp for gpt-5.6-sol (39.2%->46.8%, n=500, p=0.0003), replicated on open weights (+3.6pp, n=499).

The Recall Trap: A Recall-Maximizing Retriever Configuration Reduces Issue Resolution in Fixed-Budget Code Context

What it means Stop tuning code-assistant retrieval to recall@k; A/B the context-packing policy against actual task resolution, and at a tight token budget do not hard-deduplicate by file (favor within-file depth over file breadth).

arXiv preprint · AUG 14

Large Benchmark

The structure of a developer workspace (directory depth, modularity, injection position, context framing) measurably changes indirect-prompt-injection success against agentic coding assistants, with highly modular codebases showing significantly lower attack success rates.

Workspace Topology as an Attack Vector in Agentic Coding Assistants

What it means Treat any third-party code an agent ingests as an injection surface: constrain filesystem scope, prefer modular workspace layouts, add security-cue framing, and test agents in an uncontaminated environment before trusting IPI-resistance claims.

arXiv preprint · AUG 14

Controlled Study

In two high-velocity AI-infrastructure repos, PR throughput rose 21x (vLLM) and 17.9x (SGLang) during the agentic-coding era, but bot-authored PRs accounted for less than 0.2% of that growth, indicating the velocity increase was overwhelmingly human-driven while PR size stayed stable.

Engineering Signals of Human-AI Collaboration in the Agentic Coding Era: A Longitudinal Analysis of 33,228 Pull Requests from vLLM and SGLang with Implications for Biomedical AI Agents and Bioinformatics Pipeline Development

What it means Read the AI-coding productivity story as augmentation, not autonomy: expect throughput and reviewer-participation gains from AI-assisted humans, not from bots authoring the work — staff for more human review, not fewer humans.

arXiv preprint · AUG 14

Small Benchmark

For command-issuing coding agents, a matched headline score can hide large offsetting effects: GPT-5.6-sol's matched gap of -3.6 points masks -64.3 points of execution-path damage compensated by +60.7 points, and the deployment configuration reorders model rankings.

QuoteBench: How Matched Scores Can Hide Command-Path Failures

What it means Never rank command-issuing agents on a single matched score; benchmark under your actual execution path (shell serialization/escaping) and report the operating point and final-state validator, because the ranking flips with deployment config.

arXiv preprint · AUG 13

Small Benchmark

At repository scale, the strongest frontier coding agent fully solves only 27 of 43 joint implementation-and-proof instances and closes no specifications on the hardest repositories, showing current agents fall short of verified repo-scale synthesis.

Vero: Can AI Agents Build Formally Verified Software Repositories?

What it means Do not rely on coding agents for machine-checked correctness guarantees on multi-module codebases yet; use verified generation only where a human specifies and audits, and expect the hardest modules to fail entirely.

arXiv preprint · AUG 13

New categories

6
  • Managed Agent Runtimes . In one week a dense vendor cluster shipped managed hosted runtimes for long-horizon agents — Google Gemini Enterprise Agent Platform (GA), LangChain Managed Deep Agents (beta), Cloudways Managed AI Agents (GA), plus the Claude Managed Agents production case. This 'run my agent as a managed service' decision has no dedicated shelf in the taxonomy.
  • AI Agent Evaluation & Assurance . Agent-specific eval/QA is graduating into its own product category: LangSmith Tuned Evaluators, TestMu Agent Assurance, and Alibaba Qwen AI Arena all shipped in-window, distinct from generic AI Testing & Evaluation because they verify tool calls, side effects, and long-horizon agent traces rather than model outputs.
  • Agentic Payment Infrastructure . A funded/vendor cluster formed fast around a decision the taxonomy has no shelf for: how an autonomous agent discovers, authorizes, and pays for services (APIs, MCP servers, paywalled content, per-inference compute). AWS Bedrock AgentCore payments reached GA (Aug 18, 2026) with Coinbase + Stripe-Privy stablecoin wallets, session-scoped spend caps, and payment observability; Cloudflare Monetization Gateway, the Stripe+Tempo Machine Payment Protocol (MPP), and the x402 protocol (incl. the pay-per-inference upto scheme) are the emerging standards; LangGraph/Strands/OpenClaw ship framework integrations. Capability graduated from May 2026 preview to production-proven GA with named customers (Anchor Browser, Travala, SpreadX/BlockRun). Distinct buyable layer - wallets, spend guardrails, payment orchestration, observability - not covered by generic payments or agent-orchestration categories.
  • Agent Memory Systems . Persistent/structured agent memory is graduating from research into shipped products and a distinct buy-vs-build decision with no dedicated shelf. Signals this window — GitHub Copilot Memory for JetBrains (Aug 11); Mastra built-in memory (Aug 11); AgentCore Memory to GovCloud (Aug 7); HiGram hierarchical graph-memory research (Aug 5); plus an established OSS field (Mem0, Cognee, Graphiti/Zep, Letta, Hindsight) being actively benchmarked. The decision 'how does my agent remember across sessions' now has its own vendor cluster.
  • Autonomous Security Testing / AI Vulnerability Research . A capability-graduation cluster landed in one week: OpenAI's Astra reaching preliminary Critical cyber threshold (autonomous zero-day discovery + end-to-end attack planning), the productized Daybreak Red/Blue cyber models, their availability on Amazon Bedrock, and third-party cyber evaluations. Existing shelves (Vulnerability Management, Penetration Testing) predate autonomous AI-driven exploit research and may not capture the AI-red-team/AI-vuln-research decision.
  • Agent Payments / Autonomous Transaction Rails . Managed rails for autonomous agents to make verifiable, auditable payments are graduating to production (Solv Labs on Bedrock AgentCore payments), alongside broader AgentCore runtime GA. The decision 'how do our agents transact money with an auditable trail' has no dedicated shelf in the current taxonomy.

The instrument

The shift

Which categories are moving along the path from Buy to Bridge to Beware to Build — and which way the pressure points across every domain.

The velocity

A category drifting over years is a different decision than one that's highly volatile. The Continuum measures the rate, not just the direction.

The why

A new model capability, a pricing change, a vendor consolidation, a workflow AI just made buildable. The cause behind the shift, not only the fact of it.

Where this is heading — the trajectory map Concept illustration
Fig. 01 — Domain drift along the continuum · rewritten quarterly REWRITTEN QUARTERLY

✓ BUY

Data infra Payroll & HRIS Help desk ⇢

➔ BRIDGE

ERP & commerce Product data Martech ⇢

⚠ BEWARE

Email & outreach Comp & incentives ⇢

⚒ BUILD

Internal tools Workflow glue
⇢ drifting right this quarter · outlined chips are in motion AI feasibility pressure →

This is where the Continuum is going: a quarterly story, rewritten each cycle, of how whole groupings of software drift along Buy ➔ Bridge ⚠ Beware ⚒ Build, and why. The weekly editions below are the record it gets built from.

All editions
August 24 to September 2, 2026 21 score movements · 5 capability signals · 38 funding events · 9 papers
Week of August 17 to August 23, 2026 8 score movements
Week of August 12 to August 19, 2026 2 score movements · 5 capability signals · 21 funding events · 10 papers · 3 new categories
Week of August 5 to August 12, 2026 1 score movements · 3 capability signals · 13 funding events · 3 new categories
May 10 to August 10, 2026 4 score movements · 2 capability signals · 18 funding events · 8 papers
Week of July 29 to August 5, 2026 3 score movements · 7 capability signals · 35 funding events · 6 papers
Week of July 22 to July 29, 2026 2 score movements · 4 capability signals · 28 funding events · 3 papers · 1 new categories
Week of July 15 to July 22, 2026 1 score movements · 2 capability signals · 10 funding events · 1 papers
Week of July 10 to July 17, 2026 4 capability signals · 21 funding events · 3 papers · 1 new categories
May 28 to July 11, 2026 12 capability signals · 19 funding events · 16 papers · 2 new categories

Everything published here was confirmed by a human before it appeared. Days when we re-score the whole 1,603-category index at once are methodology passes, so they stay out of the record. The research is free.

The full database and the score updates behind these moves are in a B4 subscription.