B4 Research The Continuum

Software categories don't sit still.

Edition 8

Week of August 10 to August 14, 2026

Published weekly · free to read

As AI feasibility rises, the right call for a category changes, sometimes slowly and sometimes overnight. The Continuum is the B4 Index's weekly record of what moved and the evidence behind it. Every row is confirmed before it publishes, and every week keeps its own page.

1,603

Categories tracked

8

Score movements

$575.6B

Capital tracked

33

Capabilities confirmed

37

Papers reviewed

Recent research: July 29 to August 14, 2026

Latest edition →
Order Management (OMS)

Deeper, source-disciplined research showed the prior's '4 / mainstream self-build' rested on weak evidence (hyperscaler self-builds treated as a template, plus off-target OSS/BSS sources). The 2026 reality is that every documented OSS OMS build is hybrid — order orchestration on Medusa/Vendure/Temporal atop existing ERP/WMS and packaged components — and a full distributed-OMS self-build covering ATP, sourcing optimization, allocation, and returns reconciliation is a strategic exception, not mainstream. The market did not move; the prior overstated buildability, so this is an evidence correction to a 3.

AUG 05

Score movements

  • Loyalty Management

    Named production self-builds surfaced that the prior assessment lacked — American Eagle's full SaaS-vendor replacement, Selfridges' in-house build, and Open Loyalty running Heineken's D2C program — moving the evidence from partial-coverage to mainstream documented choice.

  • Work Management

    Deeper research separated the category core from its neighbors: the prior 4 leaned on orchestration-layer builds (Airflow, Temporal, Maestro) and a self-hosted task layer, but 2026 evidence shows the business-user work-management envelope — grids, portfolio hierarchy, resource planning, governance — is bought, with self-builds confined to narrow layers around it.

  • Cloud Cost Management / FinOps

    Moved from 4 to 3: a deeper Aug-2026 pass (CNCF adoption data + an explicit capability-boundary escalation) showed the OSS stack covers visibility/allocation/policy but NOT the high-value cross-cloud commitment optimization and anomaly core, so it is a well-resourced ~50-70% option, not a mainstream ~80%+ replacement — the prior 4 overcounted visibility components as the core.

  • Low-Code / No-Code

    Raised from 3 to 4: named independent production self-builds on OSS + AI agents (GSK/Appsmith, FalconX/ToolJet) and a buildable peer neighborhood show building the core deliverable is now a documented mainstream choice, not just a well-resourced option.

  • Calendar & Scheduling

    Cal.com moved its production codebase closed-source on April 15, 2026 and spun the open edition into Cal.diy (MIT), which omits teams, round-robin, SAML SSO, SCIM, workflows, and routing forms and is officially recommended for personal/non-production use only.

  • Last-Mile Delivery

    Source-disciplined research showed the prior's '4' anchored on hyperscaler/on-demand self-builds (UPS ORION, Amazon, DoorDash) that are frontier deployments by large infra teams, not a mainstream mid-market template.

  • Business Process Management (BPM)

    Moved 3 to 4 on substantially stronger named production evidence than the prior pass surfaced: Netflix, OpenAI, Vinted, Goldman Sachs, Deutsche Telekom, Zalando, Zoom, and small teams like Nooks all run self-hosted orchestration alternatives in production at scale, making the self-build path a documented mainstream choice rather than a well-resourced-team option.

Funding

63
CompanyAmount
Nvidia (AI compute financing alliance)
AI compute financing alliance (MOUs)Source
$500BAUG 10
OpenAI
Strategic InvestmentFoundation Model APIs (LLM & Multimodal)Source
$35BJUL 31
Intel
Public common-stock offeringSource
$15BAUG 10
OpenAI
Employee tender / secondaryFoundation Model APIs (LLM & Multimodal)Source
$7BAUG 10
Sony / TSMC image-sensor JV (Advanced Vision Semiconductor Manufacturing Corp)
Joint venture investmentSource
$4.7BAUG 11
Thrive Holdings
GrowthSource
$2BAUG 12
Firmus
Strategic equity roundGPU Cloud / AI Infrastructure PlatformSource
$2BAUG 07
Hadrian
Series DSource
$1.4BAUG 06
All 63 funding events

Capability signals

10
  • BenchmarkedBrowser/web-navigation agents at ~88% (WebVoyager, per aggregator snapshot)

    Evidence AUG 12 · Source

  • BenchmarkedMultilingual code-refactoring agent benchmark (SWE-Bench ProMax)

    Evidence AUG 10 · Source

  • BenchmarkedCross-site browser-agent benchmark (420 tasks, agent-as-a-judge)

    Evidence AUG 09 · Source

  • BenchmarkedPersistent self-evolving agentic coding runtime (~78% SWE-Bench Pro)

    Evidence AUG 05 · Source

  • BenchmarkedLong-horizon web search agent (BrowseComp 37.3% -> 55.3% with context mgmt)

    Evidence AUG 05 · Source

  • BenchmarkedProactive bug-fixing agent benchmark (1,663 tasks, 8 languages)

    Evidence AUG 05 · Source

  • Production ProvenBuilding a domain foundation model in a data-scarce vertical

    Evidence AUG 05 · Source

  • AnnouncedRunware Sonic Inference Pods — modular containerized inference data centers

    Evidence AUG 04 · Source

  • DemoedNVIDIA Alpamayo 2 Super — open 34B VLA model for L4 autonomous driving

    Evidence AUG 04 · Source

  • DemoedOrchard — open framework for training/evaluating agents at scale

    Evidence AUG 03 · Source

Papers

14
Small Benchmark

From neutral prompts, three frontier models (Fable 5, Opus 4.8, Opus 5) shipped code with unprompted SOC 2 conformance of 47-88% and real vulnerabilities (a reachable Werkzeug debugger RCE, an unauthenticated download, an endpoint returning every stored name and email); adding one sentence naming the SOC 2 standard lifted every case to 86-100% (worth 23-50 points) and removed every insecure construction.

Can AI Write Compliant Code, and to What Extent? Evaluating SOC 2 Compliance of Claude Fable 5, Claude Opus 4.8, and Claude Opus 5 Across Four Use Cases

What it means Name the compliance standard explicitly in the prompt and gate AI-generated infra/auth/PII code behind a security review; unprompted output ships RCE-class defects, controls outside the model's default conception (MFA, cookie flags, account lifecycle) survive even a named standard, and a newer model won't close the gap.

arXiv preprint · AUG 07

Controlled Study

Across 3.52M production code changes at a billions-of-users enterprise (Apr 2025-Apr 2026), AI-generated C++ carried a distinct quality profile - higher interface/coupling burden, more copy and allocation overhead, explicit loops over optimized standard APIs - translating to increased review effort and a 5-8% rise in compute consumption; taxonomy-informed static-analysis feedback cut targeted warnings 11.1%.

Characterizing the Quality Profile of AI-Generated C++ in Production

What it means Budget for the hidden tax on AI-authored code - extra review effort and ~5-8% more compute at runtime - and route it through taxonomy-informed static-analysis feedback; velocity gains are real but not free.

arXiv preprint · AUG 06

Small Benchmark

Recurrent context compression in long-horizon agents weakens the influence of recent interactions and measurably increases blocked actions, repeated exploration, and run-to-run instability; a verifier-guided compaction framework (TRACE) recovers task performance and multi-run reliability.

Toward Reliable Context Compression for Long-Horizon Agents: An Empirical Study of Execution Instability

What it means Treat context compaction as a reliability risk rather than a free cost optimization — measure multi-run stability before enabling it on long-horizon agents, since naive compression trades tokens for blocked actions and nondeterminism.

arXiv preprint · AUG 06

Case Study

A survey of 119 practitioners found LLM use has become a habitual component of professional software engineering, with reported patterns of functional dependence and overreliance — prioritizing LLMs over documentation or peer consultation — alongside less-common addiction-related difficulty moderating use.

Exploring Dependence, Overreliance, and Addiction Related Behaviors Associated with Large Language Model Use Among Software Engineers

What it means Build trust-calibration and mandatory verification into team practice — the productivity that drives adoption also drives overreliance (skipping docs and peers), so make 'verify the output' and 'consult a human/source' explicit workflow steps rather than assumed judgment.

arXiv preprint · AUG 06

Small Benchmark

Coding agents spend most of their token budget finding the file to patch rather than patching it — a 30B OpenHands agent averages 23 rounds and 631K tokens per resolved SWE-Bench issue — and a dedicated RL-trained retrieval step preserved resolve rate (27.0% vs 25.8%) while cutting 15% of rounds and 19% of tokens, but only above a retrieval-precision threshold below which retrieval degraded the agent.

CodeGrep: An RL-Trained Retrieval Agent for LLM Coding Agents

What it means Attack agent cost at repository retrieval, not just the model — a precise file-finding step cuts ~19% of tokens, but a weak retriever (BM25-grade) makes things worse, so measure retrieval precision before bolting one on.

arXiv preprint · AUG 06

Case Study

Engineers report a productivity-experience paradox: 84% report productivity improvement at both time points, yet the share of matched participants reporting worsened developer experience in at least one dimension nearly doubled from 14% to 27% over six months, with flow state and cognitive load eroding while feedback loops improved.

The Impact of AI Coding Assistants on Software Engineering: A Longitudinal Study

What it means Track developer experience alongside throughput when rolling out AI tooling — the work is shifting to supervisory correction of AI output, and that erosion shows up months after the productivity win.

arXiv preprint · MAY 22

Controlled Study

In legacy-code modernization, semantic-trap snippets silently changed behavior in 39.7% of attempts (versus 7.0% on benign controls), and the producing model silently endorsed 31.7% of its own drift cases — sometimes while correctly articulating the exact Python 2/3 semantic distinction that broke its output.

Articulate but Wrong: Self-Review Failures in LLM-Based Code Modernization

What it means Never use the model that performed a migration as its own behavioral reviewer — verify semantic preservation with executable oracles (characterization tests), especially around numeric semantics.

arXiv preprint · MAY 20

Small Benchmark

10.7% of passing SWE-agent trajectories are 'Lucky Passes' — regression cycles, blind retries, or missing verification that happen to end green — with per-model Lucky rates ranging 0.5% to 23.2%, enough to move models up to five leaderboard rank positions when scored by process quality.

AgentLens: Revealing The Lucky Pass Problem in SWE-Agent Evaluation

What it means When comparing agents for production use, inspect trajectories, not just pass rates — a model that passes by chaotic trial-and-error will be more expensive and less predictable on your real work than its score suggests.

arXiv preprint · MAY 13

Large Benchmark

Enterprise-grade coding agents executed hidden malicious commands smuggled inside benign-looking skill files in 95.5-96.1% of runs (Gemini CLI) and 71.6-74.0% (Qwen Code), nearly invariant to the generating model, with explicit safety recognition in only 1.99% of runs.

Towards a Risk Assessment of Malicious Skill Files in Coding Agents

What it means Treat third-party agent skill/plugin files as untrusted executable code - sandbox and human-review them before adoption and assume the agent's own safety recognition is near-zero.

arXiv preprint · AUG 05

Large Benchmark

When the issue report is removed and agents must discover as well as fix bugs, most state-of-the-art coding agents struggle - showing limited ability to locate and resolve recorded bugs, handle multiple-bug scenarios, and find valid latent bugs.

Active-SWE: Benchmarking Coding Agents for Proactive Bug Fixing without Issue Reports

What it means Do not extrapolate SWE-bench-style resolve rates to unscoped bug-hunting - agents lean heavily on a human-written issue report to perform; proactive discovery still needs a person to frame the problem.

arXiv preprint · AUG 05

Case Study

Repository-preserved Agent Plan files are rare but informative: screening 36,710 GitHub repositories surfaced only 85 Markdown plan files from 10 repositories, and where present they guided agent execution most often through implementation steps, concrete files and locations, and testing/validation information.

An Exploratory Study of Agent Plans for Agentic AI Coding Tools in Open-Source Software

What it means Preserve agent plan files in-repo with the elements that actually steer execution — numbered implementation steps, concrete file paths, and test/validation criteria — they double as durable human-agent task-intent documentation, but adoption is still near-zero so this is convention to establish, not a settled norm.

arXiv preprint · AUG 05

Case Study

Chat panels, terminal agents, generated diffs, and streaming status output in AI developer tools create real visual accessibility barriers for blind, low-vision, and color-vision-deficient developers, clustering into screen-reader/AT barriers, contrast and differentiation problems, and readability/scaling limits — with prominence varying by tool ecosystem.

Characterizing Visual Accessibility Issues in AI Developer Tools: An Empirical Study

What it means Include screen-reader, contrast, and scaling checks when adopting or building AI coding-tool interfaces — the agent/diff/streaming surfaces introduce accessibility barriers that general IDE accessibility work does not cover.

arXiv preprint · AUG 05

Controlled Study

Prompt wording changes where agent effort is spent without changing success: 'multiple approaches' phrasing inflates reasoning 2.4-7.4x with no success gain, 'maximum certainty' phrasing drives redundant-verification runs costing 18x the clean-run median, and harness design shifts cost per successful task 5-30x.

Same Task, Different Work: Prompt-Induced Waste in Coding Agents

What it means Strip effort-inflating phrases ('explore multiple approaches', 'be absolutely certain') from agent prompts and use bounded-efficiency wording — prompt style is a direct cost lever with no measured correctness payoff.

arXiv preprint · AUG 02

Large Benchmark

13.6% of SWE-bench Verified instances have misaligned PR-issue pairings — the problem statement does not actually match what the graded patch fixes — across five misalignment patterns.

PAIChecker: Uncovering and Checking PR-Issue Misalignment in SWE-Bench-Like Benchmarks

What it means Discount small SWE-bench Verified score differences between models — a double-digit fraction of the instances measure something other than issue resolution, so single-digit leaderboard gaps are within construction noise.

arXiv preprint · JUL 30

New categories

3
  • Agent Memory Systems . Persistent/structured agent memory is graduating from research into shipped products and a distinct buy-vs-build decision with no dedicated shelf. Signals this window — GitHub Copilot Memory for JetBrains (Aug 11); Mastra built-in memory (Aug 11); AgentCore Memory to GovCloud (Aug 7); HiGram hierarchical graph-memory research (Aug 5); plus an established OSS field (Mem0, Cognee, Graphiti/Zep, Letta, Hindsight) being actively benchmarked. The decision 'how does my agent remember across sessions' now has its own vendor cluster.
  • Autonomous Security Testing / AI Vulnerability Research . A capability-graduation cluster landed in one week: OpenAI's Astra reaching preliminary Critical cyber threshold (autonomous zero-day discovery + end-to-end attack planning), the productized Daybreak Red/Blue cyber models, their availability on Amazon Bedrock, and third-party cyber evaluations. Existing shelves (Vulnerability Management, Penetration Testing) predate autonomous AI-driven exploit research and may not capture the AI-red-team/AI-vuln-research decision.
  • Agent Payments / Autonomous Transaction Rails . Managed rails for autonomous agents to make verifiable, auditable payments are graduating to production (Solv Labs on Bedrock AgentCore payments), alongside broader AgentCore runtime GA. The decision 'how do our agents transact money with an auditable trail' has no dedicated shelf in the current taxonomy.

The instrument

The shift

Which categories are moving along the path from Buy to Bridge to Beware to Build — and which way the pressure points across every domain.

The velocity

A category drifting over years is a different decision than one that's highly volatile. The Continuum measures the rate, not just the direction.

The why

A new model capability, a pricing change, a vendor consolidation, a workflow AI just made buildable. The cause behind the shift, not only the fact of it.

Where this is heading — the trajectory map Concept illustration
Fig. 01 — Domain drift along the continuum · rewritten quarterly REWRITTEN QUARTERLY

✓ BUY

Data infra Payroll & HRIS Help desk ⇢

➔ BRIDGE

ERP & commerce Product data Martech ⇢

⚠ BEWARE

Email & outreach Comp & incentives ⇢

⚒ BUILD

Internal tools Workflow glue
⇢ drifting right this quarter · outlined chips are in motion AI feasibility pressure →

This is where the Continuum is going: a quarterly story, rewritten each cycle, of how whole groupings of software drift along Buy ➔ Bridge ⚠ Beware ⚒ Build, and why. The weekly editions below are the record it gets built from.

All editions
Week of August 10 to August 14, 2026 1 score movements
Week of August 5 to August 12, 2026 5 score movements · 3 capability signals · 28 funding events · 5 papers · 3 new categories
May 10 to August 10, 2026 3 papers
Week of July 29 to August 5, 2026 3 score movements · 7 capability signals · 35 funding events · 6 papers
Week of July 22 to July 29, 2026 2 score movements · 4 capability signals · 28 funding events · 3 papers · 1 new categories
Week of July 15 to July 22, 2026 2 capability signals · 10 funding events · 1 papers
Week of July 10 to July 17, 2026 9 capability signals · 18 funding events · 4 papers · 2 new categories
May 28 to July 11, 2026 8 capability signals · 18 funding events · 15 papers

Everything published here was confirmed by a human before it appeared. Days when we re-score the whole 1,603-category index at once are methodology passes, so they stay out of the record. The research is free.

The full database and the score updates behind these moves are in a B4 subscription.