B4 Index / The Continuum / September 11, 2026

The Continuum

B4 Research September 11, 2026

Week of September 3 to September 11, 2026 · 40 score movements, 2 capability signals, 33 funding events, 11 papers

1,603

Categories tracked

40

Score movements

$17.0B

Capital tracked

2

Capabilities confirmed

11

Papers reviewed

Medication Safety & Adverse Drug Event Surveillance

The prior 4 assumed multiple academic health systems and large IDNs run homegrown medication surveillance as a mainstream choice. Targeted 2026 research could not find a single named U.S. IDN or academic center publicly running a general-purpose ML ADE detector across multiple hospitals with published performance figures, and found the wider ML medication-error literature fragmented with scarce real-world evaluation. Real self-builds exist — NoHarm, the VA ADERS network, institution-level trigger and LASA detectors — but they are narrow or non-mainstream, which is the level-3 picture rather than level 4. The market did not move; the earlier read was too generous.

SEP 06

Score movements

  • Legal Billing Review & LEDES Invoice Analytics

    Research for this cycle found no named independent production self-build of the review core; the prior 4 rested on an assumption of documented builds that the record (PNC buys with managed review, Onit benchmark, consultancy modules for unnamed clients) does not support.

  • Legal Workflow Automation / No-Code Expert Systems

    This cycle's research surfaced named in-house production builds on Power Platform, Copilot Studio and Retool (KONE, Unifi, SHL, Baker McKenzie, Power Digital) that the prior assessment did not have; the market did not change, the evidence did.

  • Trademark Search & Watch / Brand Protection

    This cycle's research found no brand owner or firm operating a self-built clearance/watch core; the named cases (Specsavers, H&M, Gowling Saturn, Kilpatrick) all sit on licensed registry data.

  • Carrier Onboarding, Vetting & Compliance Monitoring

    The prior score treated FMCSA data plus anomaly detection as ~80% of the core; the named production builds at C.H.

  • Digital Freight Matching / Digital Brokerage Platform

    The prior assessment scored the matching algorithm as the core; the documented production builds show buyer-side teams (Loadshop aside) build rules and optimisation on bought liquidity, and the network half of the core has no self-built precedent.

  • Employee IT Onboarding / Offboarding Automation

    Named production builds on Okta Workflows and Entra Lifecycle Workflows (Juniper Square, Front, GitLab, Bayer 04, Wyndham, Akumin) with published outcomes show the self-build is mainstream, which the prior 3 undercounted.

  • Enterprise Service Management (ESM) Platform

    Named large-scale cross-department deployments of GLPI (Grupo DPSP, Econocom), Jira Service Management (Software AG) and Zammad, found in this pass, show the OSS and general-purpose path covering the core at 80%+, which the prior 3 undercounted.

  • Flow Analysis & Network Traffic Analytics (NetFlow/IPFIX)

    Named production deployments at scale (Free on Akvorado, Cloudflare goflow2 at 100K+ flows per second, AS286 on pmacct) and ntopng's published ClickHouse scaling show the OSS pipeline covering the core, which the prior 3 undercounted.

  • Shadow IT Discovery (Lightweight, Non-CASB)

    Research could not verify a single named independent organization running the homegrown pipeline in production; the prior's 4 relied on an uncited claim, and the rubric caps unverified self-builds at 3.

  • Restaurant Waitlist Management

    The prior score was argued from how easy the build looks rather than from documented self-builds; a fresh pass found the named production Twilio waitlists are Resy and Yelp, not operators, and no named restaurant running its own.

  • Automated Reference Checking Software

    A fresh search found exactly one named production self-build (Zapier) against a field of named vendor deployments; the prior score assumed multiple teams had shipped.

  • HR Workflow Automation & Service Delivery Platform

    Deeper research separated the two halves of the category: the named production self-builds (IBM AskHR, Mavlers, Care.com, Questco) all stop at the conversational front door and simple workflows, and none replace case management, entitlements, audit or lifecycle orchestration, so the prior 4 overstated coverage of the core.

  • Interview Intelligence Platform

    A targeted search for employer self-builds found none: every named employer buys, and the production builds cited previously (Greenhouse, Ashby, Deel on Recall.ai) are recruiting-software vendors building product, not independent teams building for their own needs, so the prior 4 misread the ladder's evidence bar.

  • Shift Marketplace & Flexible Scheduling App

    A targeted search for named self-built shift marketplaces found only Walmart; Starbucks and the hospital systems the prior assessment assumed were building turned out to run WorkJam, ShiftMed or CareRev under a branded front end, so the multiple-independent-teams bar for a 4 is not met.

  • Payer Contract Management & Reimbursement Modeling

    The prior score weighed general contract-modeling analytics feasibility without distinguishing the (buildable) reporting/scorecard layer from the (still mostly bought) contract-parsing and reimbursement-calculation core.

  • Wound Care Documentation & Imaging

    Deeper research found zero named production self-builds — the prior assessment's buildability claim rested on the existence of open-weight research models, not on any deployed provider system, so the score corrects from 4 to 2.

  • Charge Capture Software

    Prior score assumed buildability from task simplicity; targeted research found only vendor-shipped AI features and zero independent production self-builds, which the rubric caps at 2.

  • Clinical Documentation Improvement (CDI) Software

    Deeper research distinguished adjacent scribing self-builds (real, at scale) from CDI-specific self-builds (not found) — the prior score conflated the two.

  • Denial Management Software

    Named self-builds (Rush, UHS, UT Health San Antonio) confirmed real but narrow in scope, which is more consistent with level 3 than the level-4 mainstream-buildable bar.

  • Digital Patient Intake & Forms

    Deeper research found the named self-build (Body Brave/Jotform) covers forms and consent but not eligibility, payments or FHIR EHR write-back, which is short of the level-4 mainstream-replacement bar.

  • Pharma MLR / Promotional Material Review (PromoMats)

    A dedicated search for pharma self-built MLR systems returned none, while surfacing seven named vendor deployments - the prior 4 was built on an assumption of multiple independent production self-builds that the record does not contain.

  • Pharma Safety Literature Monitoring & Surveillance

    Checking the four companies the prior score leaned on found only AstraZeneca with a validated production pipeline - Novartis's work is a research publication, Pfizer's disclosed AI covers case intake rather than literature, and no Roche system surfaced at all.

  • Pharma Safety Signal Detection & Management

    A targeted search for production open-source signal engines found none, while surfacing five named 2026 commercial deployments and only one published proprietary model (Biogen) that is explicitly not established as a sole regulated engine.

  • Provider Credentialing & Enrollment Software

    A search specifically for greenfield credentialing self-builds returned none, and the strongest in-house story (Humana Dental) turns out to be insourced operations running on Verifiable's PSV API.

  • Real-World Data (RWD) / Real-World Evidence Analytics Platform

    Searching specifically for named in-house RWE platforms surfaced Bayer FOUNTAIN, J&J/Janssen ASSURE, Janssen's 400M-record ATLAS deployment and UCB's ATLAS extensions - multiple independent organisations running the analytics core in production on OMOP/OHDSI.

  • Risk Adjustment Coding & HCC Capture Platform

    The prior 4 rested on Apixio proving the AI approach and an assumption that large plans run internal abstractors in production.

  • Acuity-Based Nurse Staffing & Assignment Platform

    The prior 3 assumed full production alternatives covering real-time acuity scoring and EHR write-back were uncommon.

  • Conversational AI Patient Intake / Triage

    The prior 4 asserted multiple documented production self-builds of patient intake using GPT-4 or Claude with clinical prompt engineering.

  • Peer-to-Peer & Event Fundraising Platform

    This cycle looked for named production self-builds and found none — only a recommended hybrid architecture for organizations that already have engineers — where the prior score inferred buildability from how ordinary the components look.

  • Public Records Request / FOIA Management

    This cycle separated what the documented self-builds actually cover — portal, intake, routing, tracking, publication — from what agencies still buy, and the review/redaction/secure-release/retention half turned out to be untouched by any named self-build, which the prior score's redaction-is-mature reasoning glossed over.

  • Volunteer Management Software

    Checking the prior score's evidence this cycle showed it rested on a commercial product (SignUpGenius) and unnamed custom portals; the search for actual production self-builds returned none, and it also surfaced background screening as a vendor contract a build cannot code around.

  • 340B Drug Discount Management & Compliance

    The prior 4 inferred buildability from the presence of an active competitive analytics field; looking for actual self-builds found the hybrid pattern instead, where entities own the data layer and buy the accumulation, replenishment and audit-evidence engine.

  • Healthcare Price Transparency & Cost Estimation Tools

    The prior score credited open-source MRF parsers but treated member-facing estimation as vendor territory.

  • Payer Payment Integrity Platform (Pre & Post-Pay)

    The prior score credited mature claims ML and treated only the cross-payer dataset as the vendors remaining hold.

  • Charity Auction & Mobile Bidding Software

    The prior scored buildability from the simplicity of the engineering problem; research found the actual production self-builds are a school project, a commissioned one-off and hobby OSS, which is rung-3 evidence, not the mainstream-choice bar of rung 4.

  • Government Queue & Appointment Management

    The prior 4 rested on an unnamed claim that multiple teams run self-built queue systems.

  • Grassroots Advocacy & Public Affairs Platform

    The prior 4 assumed advocacy tools are routinely assembled from CRM plus AI.

  • Online Appointment Scheduling & Booking Links

    Moved 4->3: Cal.com moved its production code to closed-source in 2026 and the MIT community edition (Cal.diy) is now explicitly personal/non-production and stripped of teams, routing, workflows and SSO.

  • Project Time Tracking & Billable Hours

    Moved 4->3: closer research shows the billable-hours core that names this category — rate history, approvals and QuickBooks/Xero integration (weeks-to-months, no QuickBooks time-API sandbox) — meaningfully favors buying, and the prior score leaned on ‘Clockify is free at its core tier,’ which the 2026 free-tier change (billable rates now paid) undercuts.

Funding

33
CompanyRoundCategoriesAnnouncedAmount
Mistral AISourceSeries DFoundation Model APIs (LLM & Multimodal)+1SEP 08$3.5B
CognitionSourceSeries EAI Agent Frameworks & Orchestration+1SEP 08$2B
MiroSourceAcquisitionCollaborative Diagramming & Flowcharting+1SEP 10$1.4B
Alif Semiconductor (acquired by Analog Devices)SourceAcquisition (all-cash)SEP 09$1.4B
MotiveSourceGrowth financingAI Dash Cam & Video Telematics+1SEP 10$1.3B
Shanghai Enflame TechnologySourceIPO (Shanghai STAR Market)SEP 11$912M
Enflame Technology (Shanghai Enflame Technology Co., Ltd. / 燧原科技)SourceIPOSEP 11$912M
Enflame TechnologySourceIPOSEP 11$912M
Shanghai Enflame Technology (燧原科技)SourceIPOSEP 11$912M
Positron AISourceSeries CSEP 10$875M
Positron AISourceSeries C (split into a $375M Series C and a Series C-1 of up to $500M)SEP 10$875M
HarveySourceGrowth round (no letter disclosed; co-led by Diffusion and Lightspeed Venture Partners)AI Contract Review & Playbook Analysis+1SEP 09$550M
Gimlet LabsSourceSeries BGPU Cloud / AI Infrastructure Platform+1SEP 04$300M
Celero CommunicationsSourceSeries CSEP 08$275M
SwarmerSourceM&ADefense Command & Control / Battle Management System (C4ISR)SEP 10$224M
Ayar LabsSourceSeries E extensionSEP 10$150M
Forus (formerly Tandem)SourceSeries CAI Prior Authorization Automation+1SEP 08$150M
ClaySourceSeries DAI SDR / Autonomous Outbound Prospecting Agent+2SEP 09$115M
Maven RoboticsSourceSeries ARobotics Fleet Management / Multi-Robot OrchestrationSEP 10$100M
InspirenSourceSeries CAssisted Living / Senior Living Management PlatformSEP 10$70M
ArchySourceSeries CDental Practice Management Software for DSOsSEP 11$50M
Bluecore EnergySourceSeedSEP 08$50M
CymphonySourceSeries AAI Agent Identity & Authorization Platform+1SEP 09$25M
CloudNCSourceSeries B extensionCAD/CAM Software (CNC Programming)SEP 08$20M
AtiraSourceseedConfigure, Price, Quote (CPQ)+1SEP 03$15M
BesxarSourceSeedSEP 09$10.3M
HelmGuardSourceseedGRC Automation (Compliance Automation)+2SEP 09$7.3M
HelmGuardSourceSeedAI Governance & Compliance+2SEP 09$7.3M
HelmGuard AISourceSeedAI Governance & Compliance+2SEP 09$7.3M
IVEXSourceSeries ASEP 10$5.8M
FluencifySourcepre-seedInfluencer Marketing PlatformSEP 07$4.3M
VeridueSourcepre-seedAlternative Investment CRM / Deal Management+1SEP 08$4M
sci2sciSourcepre-seedAI Governance & Compliance+1SEP 08$1.4M

Capability signals

2
  • DemoedSelf-maintaining repo documentation as agent context

    Evidence SEP 10 · Source

  • Production ProvenMulti-agent scaffolds running long campaigns end to end, observed in the wild

    Evidence SEP 10 · Source

Papers

11
Large Benchmark

Across 787,562 paired functions, AI-generated code is roughly half the size and branching of human code and differs in defect KIND rather than amount: more, and more severe, security findings in Python and Java, but fewer high-severity memory-safety findings than humans in C.

What is the Difference Between Me and You? Benchmarking the Quality Gap Between Human-Written and AI-Generated Code

What it means Tune static-analysis and review gates by language rather than applying one AI-code policy everywhere: tighten security review on generated Python and Java, and expect AI defects to cluster in repetitive boilerplate instead of the legacy-complexity hotspots reviewers are trained to look for. Once size is controlled, complexity metrics carry little signal, so stop using them to screen generated code.

arXiv preprint · SEP 11

Large Benchmark

57.5% of agent conversations rated satisfied by a blind rater panel failed the customer's task, and the LLM-as-judge promotion gate disagreed with a grounded verifiable reward on 31% of close-reward pairs versus under 1% on wide ones.

GAUGE: When Not to Trust LLM-as-a-Judge in User-Simulated Evaluation of Task-Oriented Agents

What it means Never promote one agent version over a near-equal rival on an LLM-judge satisfaction score; that is exactly the regime where the gate loses resolution. Keep the judge for coarse capability screening, add a judge-free completion bit as a zero-cost tripwire, and settle close calls with a grounded verifiable reward.

arXiv preprint · SEP 10

Small Benchmark

Of 654 Python samples that produced no findings under a combined Bandit + Semgrep gate, runtime verification confirmed or partially confirmed exploitability in 95 files — an inclusive pipeline rate of 14.53%, roughly 1 in 7 statically clean samples — and frequently confirmed weakness classes including CWE-338 and CWE-916 were flagged by neither scanner.

Beyond Static Guarantees: Measuring the Static-Pass Dynamic-Fail Gap in Security-Sensitive and LLM-Generated Python Code

What it means Stop shipping AI-generated code on a clean SAST run alone. Add runtime exploit verification for the weakness classes static tools structurally miss — weak randomness and missing password hashing showed up repeatedly in code both Bandit and Semgrep passed. Static and dynamic are layers, not substitutes.

arXiv preprint · SEP 09

Controlled Study

Conversational 'vibe coding' cut task completion time 27% against traditional coding and 12% against AI-assisted coding, but the same tasks came back with lower maintainability indices and higher security vulnerability counts, and perceived loss of control was associated with the increased security risk.

The Vibe Shift in Software Engineering: Evaluating AI-Led Conversational Programming for Performance, Cognition, and Responsible Adoption

What it means Budget the speed gain against a maintenance and security bill rather than banking it. If a team adopts conversational coding, add the review capacity back on the quality side — the measured 27% time saving arrives with code that scores worse on maintainability, so net throughput depends on what happens after the merge.

International Journal on Advanced Science, Engineering and Information Technology · SEP 09

Large Benchmark

Blocking evaluation-time leakage channels on SWE-Bench Pro dropped GLM-5.2 from 78.80% to 57.32% (-21.48 points, recovering to 59.51% once 102 flawed instances were also repaired), while DeepSeek-V4-Pro moved only 49.98% to 49.11%, so published SWE-Bench Pro numbers overstate real coding ability for some models by roughly twenty points and not at all for others.

SWE-Bench Pro Verified: A Reliable Benchmark for Software Engineering Agents

What it means Stop treating a vendor's SWE-Bench Pro number as a capability reading. Ask which harness and which leakage controls produced the score, and rank coding agents by how they hold up under sealed conditions rather than by the headline figure — the measured drop was 21 points for one model and under one point for another, so the gap between two vendors can invert once the environment is closed.

arXiv preprint · SEP 08

Controlled Study

Agent self-reports of task progress are stage-dependent and often wrong: the two compliant no-reasoning-token deployments (gpt-4.1, gpt-5.5) were correct at 90.6–99.4% of pre-action checkpoints but only 5.8–11.5% mid-task; across the five deployments whose worst stage is mid-task, 82–90% of well-formed mid-task errors named a pre-action state the task had already left; and a bundled intervention that withdraws a needed tool while reassigning the unfinished step to another system raised false completion claims from 6.2% to 64.4% across nine deployments with task truth unchanged.

The Unreliable Progress Bar: Can LLM Agents Reliably Report Task Progress Throughout Execution?

What it means Never let an agent framework decide to continue or stop on the model's own status report. Derive completion from environment state — tests, database, filesystem — and treat the agent's 'done' as a hint. The failure concentrates mid-task, which is precisely where an unattended long run is when a supervisor would want the signal.

arXiv preprint · SEP 08

Large Benchmark

Model rankings on benchmarks assigned the same capability concept are often as strongly correlated as rankings on benchmarks assigned different concepts, correlations within the same assigned safety concept are often weak, benchmarks sharing a design element such as score format sometimes correlate more strongly than benchmarks sharing a concept, and some benchmarks correlate better with a different concept than their own — BBQ-accuracy tracks reasoning benchmarks more closely than bias ones.

What AI Benchmarks Actually Measure: Adapting Convergent and Discriminant Validity to Interrogate Fifty-Six AI Benchmarks

What it means Do not build a procurement or deployment decision on a named capability score. Benchmarks labeled 'reasoning' or 'safety' do not separate the concepts they claim, so evaluate candidate models on a task set drawn from your own work — a bespoke eval on real tickets is a better predictor than any published leaderboard column.

arXiv preprint · SEP 08

Controlled Study

Across 1,920 pre-registered trials, AI coding assistants opened any provenance signal before installing in 9 trials (0.5%) and ran a verification command in none, so the presence of an SBOM, signed release, or build attestation had no measurable effect on what they installed — and the most capable/most expensive model verified nothing while the cheapest verified most.

Do AI Coding Assistants Check Before They Install? A Pre-Registered Demand-Side Audit of Trust Signals in the Research Software Supply Chain

What it means Put dependency verification in the runtime that executes the assistant, not in the prompt and not in the model choice. Gate installs behind a policy layer that checks signatures and attestations itself — buying a more capable model demonstrably does not buy you a more careful installer.

arXiv preprint · SEP 07

Large Benchmark

16.0% of AI coding-agent setups carry a confirmed security defect: 9.8% install an MCP server with no version pinned, 3.1% pre-approve arbitrary execution behind a scoped-looking grant such as Bash(python:*), and 3.8% ship a skill that pre-approves the shell for whoever installs it.

Scanning the Harness: An Empirical Study of Supply-Chain Defects in AI Coding-Agent Configurations

What it means Treat instruction files, skills, hooks, subagents and MCP declarations as a dependency layer with no lockfile, because that is what it is. Pin MCP server versions, audit every pre-approval grant for scoped-looking wildcards, and scan skill collections before installing — the third defect class is visible to a marketplace scan and ships in 3.7% of collections anyway.

arXiv preprint · SEP 07

Controlled Study

Human reviewers approved AI-authored pull requests at 30.5% in their early periods and 36.6% in their late periods (Wilcoxon p = 8.6e-8, Cohen's d = 0.25), and Granger analysis shows the approval-rate change predicts later shifts in review language rather than following it.

Beyond Lexical Metrics: Sentence-Embedding Detection of Reviewer Habituation in AI Code Review

What it means Assume your reviewers get looser on agent PRs over time and instrument for it. Track approval rate per reviewer against exposure volume as a standing metric, rotate reviewers on agent-heavy repos, and stop relying on comment quality as the health signal — the linguistic features people would eyeball showed nothing while approval behavior had already moved.

arXiv preprint · SEP 05

Small Benchmark

Asked to build a working customer-service agent under real engagement conditions, the strongest configuration — Claude Opus 5 under Claude Code — passed 23.9% of evaluation simulations against an expert-authored reference ceiling of 82.2%.

ττ-Bench: An Environment for End-To-End, Realistic Agent Construction

What it means Treat 'have a coding agent build the agent' as a supervised exercise, not a delegation. The named failure modes are the supervisable ones: shallow queries instead of reading the records, almost no communication back to the requirements holder (client calls were 0.3% of all tool calls), and shipping the first architecture that runs. Force the agent to report its understanding of the data and to try more than one design before it is allowed to ship.

arXiv preprint · SEP 04

Every row here cleared the pipeline's verification before it published, and the research is free. The full database and the score updates behind it are in a B4 subscription.