B4 Index / The Continuum / September 2, 2026

The Continuum

B4 Research September 2, 2026

August 24 to September 2, 2026 · 21 score movements, 5 capability signals, 40 funding events, 9 papers

1,603

Categories tracked

21

Score movements

$12.7B

Capital tracked

5

Capabilities confirmed

9

Papers reviewed

Kubernetes Policy-as-Code Enforcement

The prior assessment scored 3 on the view that the commercial management plane was significant remaining value; this pass surfaced the scale and breadth of named OSS-only production deployments — Swisscom (390 clusters, explicitly avoiding external tools), LinkedIn (500K+ nodes), plus the public Kyverno adopter roster — showing the self-run enforcement core is a mainstream documented choice, the level-4 bar. Deeper research, not a market move.

SEP 02

Score movements

  • Cloud Development Environment (CDE)

    Deeper research surfaced multiple named large-scale production self-hosted CDE deployments (Dropbox 1,000 devs, Credit Karma, EnBW, 2,500-dev defense org) on Coder Community OSS — evidence the prior 3 predated; the market itself did not move this quarter.

  • Cloud Migration & Modernization Tooling

    Deeper sourcing showed every documented production AI case (Spotify, Google, DataArt, Novacomp) is code transformation, while the execution core — replication, cutover, agentless discovery — has no independent self-built alternative and even AWS Transform depends on MGN agents; the prior 4 over-credited the analysis-layer wins to the whole category.

  • Chaos Engineering Platform

    Re-grounded against CNCF production endorsement of LitmusChaos and Chaos Mesh plus AWS FIS; the prior 3 predated recognizing the OSS core as a mainstream ~80% self-build rather than a well-resourced option, moving it to 4 and flipping the badge to BEWARE.

  • Regulatory Change Management / Horizon Scanning

    Deeper 2026 research and peer calibration, not a market move: the authoritative multi-jurisdiction content moat, human-owned defensible obligation inventory, and sub-5% embedded-production maturity show the self-build covers only ~50-70% of the core, so the first-pass 4 was too high and the correct level is 3.

  • ML Model Supply-Chain Scanning & AI Bill of Materials

    Deeper research reframed the score from build-for-own-needs rather than matching vendor detection depth: the OSS scanner-plus-BOM core (ModelScan, Fickling, picklescan, CycloneDX) is documented as production-feasible and increasingly mainstream in CI/CD, and the near-duplicate peer (AI Model Supply-Chain Security, 0.91 cosine) already sits at 4.

  • Intercompany Transfer Pricing / Operational TP Monitoring

    Fresh 2026 survey and case evidence shows self-built operational TP monitoring remains a well-resourced-team option rather than a mainstream choice (1% fully integrated, hybrid dominant), correcting the prior 4 down to 3.

  • Profitability & Cost Allocation Analytics (Standalone from BI)

    Fresh evidence shows self-builds cover the calculation layer but complete production EPM replacements remain rare and undocumented, correcting the prior 4 to a partial-buildability 3.

  • Royalty & IP Licensing Accounting Software

    Evidence shows catalog-scale self-built engines are confined to a few SI-assisted music majors (BMG, Sony) while the broader category buys or hybridizes, correcting the prior 4 to a well-resourced-option 3.

  • Debt Collection & Recovery Software

    Deeper 2026 research shows AI-enabled collections is mainstream but complete lender self-build is explicitly a minority practice — only early-stage self-cure/outreach is routinely built in-house, with late-stage and compliant end-to-end stacks bought.

  • Methane Emissions Monitoring & Leak Detection (MRV) Software

    Deeper research showed the self-built portion is the decision/analytics layer, not the measurement-and-validation core, which still requires bought sensing and independent verification; the prior 4 over-credited buildability of the core.

  • Utility Vegetation Management Software

    Better research showed independent full self-builds are research prototypes, not utility-wide production — utilities build only component/analytics layers while buying the CV/prediction core — so the prior 4 overstated production buildability.

  • AI-Powered Cash Flow Forecasting (Standalone, Non-TMS)

    Moved 4 to 3: fresh research shows the prior assumption of multiple mainstream production self-builds is unsupported — named cases are thin (only AstraFi) and hybrid is the documented winning pattern.

  • K-12 Family Engagement & School-Home Communication

    Lowered 4->3: deeper research shows full family-engagement self-builds are rare, uneconomic, and undocumented at production scale, and procurement still demands hosted platforms.

  • K-12 Substitute Teacher & Absence Management

    Lowered 4->3: deeper research shows self-build is rare and undocumented at production scale, the market switches vendors rather than building, and the substitute-network is a real moat.

  • Software Composition Analysis (SCA) / Dependency Security

    Deeper research surfaced what the prior assessment missed: Dependency-Track as a production OSS management plane and GitLab shipping Trivy as its default scanner — the 'management layer is vendor-only' premise behind the prior 3 was wrong.

  • Ecommerce Affiliate & Partner Management Software

    The prior 4 rested on 'multiple teams run custom-built affiliate programs' — the research found one documented referral-scale production system (redBus), a documented retreat (Jobber retiring its internal build over maintenance cost), and no named full publisher-affiliate self-builds carrying fraud, tax/KYC, and payout depth.

  • Ecommerce Headless Storefront / Frontend-as-a-Service (FEaaS)

    Deeper research showed the prior score understated documented practice: custom and agency-built storefronts are the majority pattern among headless adopters, and the prior assessment's own text conceded widespread production self-builds while scoring them as merely partial.

  • Ecommerce Loyalty & Rewards Program (Store-Level)

    The market did not move — closer research showed the prior assessment overstated the evidence.

  • GraphQL Federation & Gateway Platform

    Better evidence of production self-hosting on Cosmo (Apache-2.0) and Hive (MIT), reinforced by Apollo's 2026 removal of hosted cloud routing driving teams to self-host.

  • Headless / API-First CMS

    Better evidence of mainstream production self-hosting — Strapi at Airbus/Tesco/Sonos, Figma-backed Payload defaulting to self-host, Directus on SQL — lifting this to clearly-buildable.

Funding

40
CompanyRoundCategoriesAnnouncedAmount
MediaTekSourceStrategic investment (convertible bonds, by Nvidia)AUG 31$3.5B
Andreessen Horowitz (Growth Fund V)SourceVC fund raise (additional close)AUG 31$1.8B
a16z (Machine Age Fund)SourceVC fund raiseAUG 28$1.1B
Andreessen Horowitz (Machine Age Fund)SourceVC fund closeAUG 28$1.1B
Shanghai Enflame TechnologySourceIPO (STAR Market, priced)AUG 31$911M
Shanghai Enflame TechnologySourceIPOAUG 31$908M
WonderfulSourceSeries CAI Agent Frameworks & Orchestration+1SEP 02$550M
Tripo AI (VAST)SourceSeries B / B+3D Rendering & Product Visualization SoftwareSEP 01$446M
GoProSourceMajority stake sale (merger, ~90% cash-out)SEP 01$285M
InstinctSourceSeries BAUG 26$250M
OwnerSourceSeries DRestaurant Direct Online Ordering Platform+1AUG 28$240M
GatikSourceSeries DAUG 25$200M
Generalist AISourceSeries B extensionAUG 24$200M
LyteSourceSeries CSEP 02$165M
SocureSourceStrategic growth (Series E extension)Identity Verification & 2FA+1AUG 27$156M
Emerald AISourceSeries AData Center Infrastructure Management (DCIM)+1AUG 25$150M
Alice (formerly ActiveFence)SourceGrowth (unlabeled)AI / LLM Security (Runtime Guardrails & AI-SPM)+2AUG 25$140M
iPronicsSourceSeries BSEP 02$125M
HiddenLayerSourceSeries BAI / LLM Security (Runtime Guardrails & AI-SPM)+2SEP 02$100M
Stability AISourceSeries BAI Image Generation Platform+2AUG 25$76M
Physical Superintelligence (PSI)SourceSeedData Center Infrastructure Management (DCIM)SEP 01$58M
ConveoSourceSeries AAgile Market Research Platform (Self-Serve Consumer Insights)+1SEP 02$50M
AIR (AIR Security Inc.)SourceSeed (two rounds: $10M + $40M)AI / LLM Security (Runtime Guardrails & AI-SPM)+1SEP 01$50M
Deep CogitoSourceSeries AFoundation Model APIs (LLM & Multimodal)AUG 26$43M
WaferSourceSeries AGPU Workload Orchestration / SchedulingSEP 01$40M
TransfyrSourceSeedScientific Data Management System (SDMS) / FAIR Data LayerAUG 26$25M
AgentrysSourceSeed + pre-seedPCB / Electronic Design Automation (EDA)AUG 26$24.5M
EmpirikSourceSeedAIOps Event Correlation & Noise ReductionSEP 01$21M
Ringg AISourceSeries AAI Autonomous Customer Service Agent Platform+1SEP 01$15M
CliptoSourceEquity (unlabeled)Knowledge Graph & Personal Knowledge Management (PKM)+1AUG 31$15M
LegatoSourceVenture (stage undisclosed, stealth emergence)AUG 26$12M
Cloverleaf AISourceSeries ASales IntelligenceAUG 27$8M
Blue VoiceSourceSeedLegal Research+1AUG 31$6M
MultiplierSourceSeedAUG 26$6M
GuicklySourceSeedAI Governance & Compliance+1SEP 01$4.2M
ItoflowSourcePre-seedAUG 25$2.5M
IntelliciaSourcePre-Series AAgile Market Research Platform (Self-Serve Consumer Insights)SEP 02$2.5M
ConsoleSourceAcquisition (by Palo Alto Networks)IT Service Management (ITSM)SEP 01Undisclosed
SB EnergySourceIPO filing (S-1)SEP 01Undisclosed
FravitySourceacquisitionAML Case Management & SAR Filing+2AUG 27Undisclosed

Capability signals

5
  • Production ProvenCodified agent workflows in production ops (onboarding, account management, DevRel)

    Evidence SEP 01 · Source

  • Production ProvenAutonomous agent-initiated payments with deterministic trust gating at 20M+ transaction scale

    Evidence SEP 01 · Source

  • Production ProvenReal-time per-user LLM spend enforcement on managed infra

    Evidence SEP 01 · Source

  • Production ProvenAgent-majority software development lifecycle at enterprise scale

    Evidence AUG 27 · Source

  • DemoedBehavioral observability catching agent failure modes invisible to conventional monitoring

    Evidence AUG 27 · Source

Papers

9
Controlled Study

Benchmarks sharing the same label (e.g. 'bug fix') differ statistically on at least two of three task-demand axes (Spread–Novelty–Centrality), agent success concentrates uniformly in low-demand tasks for every model family and scale, and the behavioral signature of success is model-family-specific (Claude succeeds by matching gold-solution scope, Qwen by exceeding it).

What Does an Agentic Software Engineering Benchmark Measure? Profiling Task Demands and Agent Behaviour Beyond What Category Labels Reveal

What it means When comparing agent scores across benchmarks, check what the benchmark actually demands rather than trusting its category label — and expect deployed agents to succeed mostly on low-spread, low-novelty, low-centrality tasks, reserving the hard cross-cutting changes for humans.

arXiv preprint · SEP 01

Large Benchmark

On 203 real-world dependency-upgrade tasks with hidden code-level breakage (DEPBENCH), the best coding-agent configuration solved only 51.2% (104/203), with substantial variation across harnesses, models, and ecosystems.

Update from Hell: Can Coding Agents Survive Hidden Breakage in Dependency Upgrades?

What it means Do not hand agents unattended dependency upgrades — the maintenance work teams most want to delegate is precisely where agents fail half the time; keep upgrade PRs behind full test suites and human review, and expect ecosystem-specific gaps.

arXiv preprint · AUG 31

Large Benchmark

How a user invokes a coding agent changes its vulnerability to poisoned repositories by up to 4.5x in attack success rate across task types (Run-Tests 45.5% ASR vs Fix-Bug 8.6%), with test-execution tasks forming a silent attack surface (high attack success, low agent alerting).

Beyond the Payload: How User Invocation Shapes Coding Agent Vulnerability to Repository Poisoning

What it means Treat 'run the tests' in an untrusted third-party repo as the riskiest agent instruction, not the safest — sandbox agents when bootstrapping from external repos, and know that prompt phrasing and supplied rules shift the attack surface rather than eliminating it.

arXiv preprint · AUG 31

Controlled Study

A single factor explains 74.5% of common variance across twelve frontier benchmarks and substantially tracks model release date (R²=0.505), so most of the gap between models released months apart is calendar; under the pre-registered dimensionality rule the four economic benchmarks form no distinct capability factor — though a leave-one-benchmark-out test shows a multi-factor representation still predicts held-out economic scores better than a single general index, so they carry incremental predictive information without constituting a distinct latent capability.

One Capability or Many? Testing the Economic Validity of Frontier AI Evaluation

What it means Date-adjust before buying on leaderboard gaps: a small score difference between contemporaneous models is mostly noise around a time trend, not a capability difference — pick by trying models on your own workload, not by leaderboard deltas.

arXiv preprint · AUG 29

Position

Contamination types are best organized by which mitigation each defeats (direct, derivative, temporal, distributional, acquired) — a private held-out test set closes only the first — and in an audit of 41 evaluation documents, elicitation budgets were reported in just 13% of documents and none addressed all five types.

Benchmark Contamination: A Taxonomy Organized by Defeated Mitigation

What it means When a benchmark score informs a buy decision, ask which contamination mitigations the eval actually applied — a held-out private test set defeats only direct leakage, and current reporting practice almost never discloses enough to tell capability from leakage.

arXiv preprint · AUG 29

Large Benchmark

Rewriting SWE-bench tasks to match how real users phrase requests cuts resolution rates by 6.4 percentage points on average and can change model rankings; 88% of real prompts carry only a bare problem statement (alone or with limited context) versus 7% of benchmark problems, and stating Desired Behavior and Motivation measurably lifts performance while Environment Information and Reproduction Steps add tokens without benefit.

RealSWE: A Compositional Evaluation of Coding Agents under Realistic User Requests

What it means State the desired behavior and the motivation in every agent request — those two fields measurably lift resolution while reproduction steps and environment info just add tokens; and discount SWE-bench scores as an upper bound, since real phrasing costs ~6 points and can reorder models.

arXiv preprint · AUG 28

Small Benchmark

A manager-worker scaffold over a shared filesystem helps some models dramatically (+8 to +30 points on the 100 hardest recent LiveCodeBench problems, up to +42 in single-pass configs) and is null or negative for others (-1 to -9), roughly triples token cost, but buys accuracy more cheaply than upgrading to a larger model in the cases studied.

Zero-Shot Self-Orchestration with Ledger-Based Control for Improved LLM Coding Performance

What it means Test multi-agent scaffolds per-model before adopting: the gain is conditional, largest for smaller models and reasoning-off configs, and near-zero for large reasoning models — and compare the ~3x token bill against simply buying the bigger model, which this study found is often the worse deal.

arXiv preprint · AUG 27

Small Benchmark

LLM-generated backend services that pass functional tests still show statistically significant memory-growth trends under 48-hour sustained execution in most application-language combinations, indicating software-aging defects that correctness testing does not catch — and the aging trends also appear in some human-written implementations.

Investigating Software Aging in LLM-Generated Software Systems across Generation-and-Execution Environments

What it means Soak-test AI-generated services before deploying them as long-running processes — passing the test suite says nothing about memory behavior at hour 40; add memory-trend monitoring to acceptance criteria for generated backends.

arXiv preprint · AUG 26

Case Study

Teams absorbing high-volume AI code generation are shifting from line-by-line review toward layered supervision: machine-readable architectural conventions as preventive guardrails, linting/testing/CI repurposed as supervision infrastructure, and human review refocused on architectural reasoning and operational explainability.

When Review Alone No Longer Scales: Layered Supervision in AI-Assisted Software Engineering

What it means If AI generation volume is straining code review, don't add reviewers — externalize architectural intent into machine-checkable rules and let CI carry first-pass supervision, reserving humans for architecture and maintainability judgment.

arXiv preprint (accepted, ESEM 2026 SEIP track) · AUG 26

Every row here cleared the pipeline's verification before it published, and the research is free. The full database and the score updates behind it are in a B4 subscription.